Observability Stack: Grafana, Prometheus & Loki
Setting up a complete monitoring and logging stack for your Kubernetes cluster using the Grafana ecosystem.
Ahmed
DevOps & Cloud Engineer
Deploying to Kubernetes without observability is flying blind. When something breaks at 3 AM, the difference between a 5-minute and a 5-hour incident is having the right metrics, logs, and traces at your fingertips. The Grafana stack — Prometheus, Loki, and Grafana itself — is the most widely adopted open-source observability solution for Kubernetes.
The Three Pillars of Observability
- Metrics (Prometheus) — numeric time-series data: CPU, memory, request rate, error rate, latency
- Logs (Loki) — structured event streams from your applications and infrastructure
- Traces (Tempo) — distributed request tracing across microservices
Installing the Stack with Helm
The kube-prometheus-stack Helm chart installs everything — Prometheus, Alertmanager, Grafana, and a set of pre-built dashboards — in one command:
# Add the Helm repo
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
# Install kube-prometheus-stack
helm upgrade --install monitoring prometheus-community/kube-prometheus-stack \
--namespace monitoring \
--create-namespace \
--values monitoring-values.yaml \
--wait# monitoring-values.yaml
grafana:
adminPassword: "changeme-in-vault"
ingress:
enabled: true
hosts: ["grafana.internal.mycompany.com"]
persistence:
enabled: true
size: 10Gi
prometheus:
prometheusSpec:
retention: 30d
storageSpec:
volumeClaimTemplate:
spec:
resources:
requests:
storage: 100Gi
alertmanager:
config:
route:
receiver: slack-notifications
receivers:
- name: slack-notifications
slack_configs:
- api_url: "https://hooks.slack.com/services/..."
channel: "#alerts"Adding Loki for Log Aggregation
Loki is Prometheus for logs — it indexes only metadata (labels) rather than full text, making it incredibly storage-efficient:
helm repo add grafana https://grafana.github.io/helm-charts
# Install Loki stack (Loki + Promtail log collector)
helm upgrade --install loki grafana/loki-stack \
--namespace monitoring \
--set loki.persistence.enabled=true \
--set loki.persistence.size=50Gi \
--set promtail.enabled=trueWriting Effective PromQL Queries
PromQL is the query language for Prometheus. Here are the queries I use in every production dashboard:
# Request rate per second (5m rolling window)
sum(rate(http_requests_total{job="my-service"}[5m])) by (status_code)
# 99th percentile latency
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket{job="my-service"}[5m])) by (le)
)
# Error rate as a percentage
100 * sum(rate(http_requests_total{status_code=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
# Memory usage vs limit
container_memory_working_set_bytes{namespace="production"}
/
container_spec_memory_limit_bytes{namespace="production"}Tip
Use the RED method for service dashboards: Rate (requests/sec), Errors (error %), Duration (latency percentiles). These three metrics surface 90% of production issues.
Alerting That Doesn't Cry Wolf
Alert fatigue kills on-call effectiveness. Each alert should be actionable and map to a customer-impacting condition:
# alerting-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: service-alerts
namespace: monitoring
spec:
groups:
- name: service.rules
rules:
- alert: HighErrorRate
expr: |
(
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "Error rate above 5% for 5 minutes"
runbook: "https://runbooks.internal/high-error-rate"“If your on-call engineer is getting paged more than once a night on average, your alerting is broken — not your system.
— Google SRE Book