DevOps 2025 — GitOps, Platform Engineering & AI (Practical Guide)

Practical, step-by-step guide to implementing GitOps + a minimal internal developer platform (IDP) and an AI-assisted observability demo. Includes a runnable GitHub repo, code snippets, and a cheat-sheet CTA.
TL;DR
GitOps (Argo CD / Flux), Platform Engineering (IDPs), and AIOps are the highest-impact DevOps themes in 2025. This walkthrough shows a compact, practical mini-project you can run in minutes: build a sample Node microservice, deploy it via GitOps, scaffold environments with GitHub Actions, and add a simple alert -> LLM suggestion pipeline (human-in-the-loop). Repo link and assets at the end.
Why this matters
Repeatability: GitOps makes deployments auditable and reproducible.
Developer velocity: A minimal IDP gives developers safe self-service.
Reduced toil: AIOps can triage and suggest next steps so humans focus on decision-making.
What you’ll ship (quick)
A public GitHub repo with:
A sample Node app (Dockerfile + k8s manifests)
ArgoCD Application manifest
GitHub Actions workflow to build image & update manifests
A PrometheusRule alert and a webhook that sends alert context to an LLM
A short GIF showing
git push → ArgoCD sync → new podA cheat sheet you can give subscribers
Repo layout (what to push)
/README.md
/apps/sample-node/Dockerfile
/apps/sample-node/index.js
/apps/sample-node/k8s/deployment.yaml
/apps/sample-node/k8s/service.yaml
/gitops/argocd/application.yaml
/.github/workflows/deploy.yml
/observability/prometheus-rules.yaml
/observability/otlp-demo/README.md
/assets/cover.png
/assets/demo.gif
Step-by-step (the core bits)
1) ArgoCD Application (paste into /gitops/argocd/application.yaml)
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: sample-node
labels:
app.kubernetes.io/name: sample-node
spec:
project: default
source:
repoURL: 'https://github.com/<YOUR-USER>/devops-2025-demo'
path: apps/sample-node/k8s
targetRevision: HEAD
destination:
server: 'https://kubernetes.default.svc'
namespace: sample
syncPolicy:
automated:
prune: true
selfHeal: true
2) Minimal K8s deployment (paste into /apps/sample-node/k8s/deployment.yaml)
apiVersion: apps/v1
kind: Deployment
metadata:
name: sample-node
labels:
app: sample-node
spec:
replicas: 1
selector:
matchLabels:
app: sample-node
template:
metadata:
labels:
app: sample-node
spec:
containers:
- name: sample-node
image: ghcr.io/<YOUR-USER>/sample-node:latest
ports:
- containerPort: 3000
readinessProbe:
httpGet:
path: /health
port: 3000
initialDelaySeconds: 5
periodSeconds: 5
3) GitHub Actions workflow (paste into /.github/workflows/deploy.yml)
name: GitOps Deploy
on:
push:
branches: [ main ]
jobs:
build-and-push:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Log in to GHCR
uses: docker/login-action@v2
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Build and push image
run: |
docker build -t ghcr.io/${{ github.repository_owner }}/sample-node:${{ github.sha }} ./apps/sample-node
docker push ghcr.io/${{ github.repository_owner }}/sample-node:${{ github.sha }}
echo "IMAGE=ghcr.io/${{ github.repository_owner }}/sample-node:${{ github.sha }}" >> $GITHUB_ENV
- name: Update k8s manifest with new image
run: |
sed -i "s|image: .*|image: $IMAGE|" apps/sample-node/k8s/deployment.yaml
- name: Commit updated manifest
uses: EndBug/add-and-commit@v9
with:
message: 'ci: update image'
add: 'apps/sample-node/k8s/deployment.yaml'
4) Observability alert (example /observability/prometheus-rules.yaml)
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: sample-node-rules
spec:
groups:
- name: sample-node.rules
rules:
- alert: HighErrorRate
expr: rate(http_server_errors_total[2m]) > 0.05
for: 2m
annotations:
summary: "High error rate for sample-node"
description: "Error rate > 5% for 2 minutes"
5) Simple AIOps webhook (concept)
Create a tiny serverless function (AWS Lambda / Cloud Run) that:
Receives Prometheus Alert manager webhook.
Gathers context (service name, last 5 error logs/traces, recent deploys).
Calls an LLM with prompt: “Given these metrics and context, list 3 likely causes and 3 prioritized next actions.”
Sends the suggestion to Slack/Jira but requires a human approval step before automation.
Important: Always keep humans in the loop for remediation — show AI suggestions as guidance.
CTA
Repo & cheat sheet: https://github.com/eknathdj/devops-2025-demo
Subscribe for a downloadable cheat sheet and weekly practical guides: https://eknathdj.substack.com
Author
Eknath D J — DevOps engineer. I publish project-first tutorials on GitOps, Platform Engineering, and cloud cost optimization. Subscribe for the repo + cheat sheet: https://eknathdj.substack.com
