A

Ahmed

DevOps Engineer · Software Engineer

Egyptahmed@example.comOperational
Back to Blog
MonitoringGrafanaKubernetes

Observability Stack: Grafana, Prometheus & Loki

Setting up a complete monitoring and logging stack for your Kubernetes cluster using the Grafana ecosystem.

A

Ahmed

Nov 28, 20257 min read

Deploying to Kubernetes without observability is flying blind. When something breaks at 3 AM, the difference between a 5-minute and a 5-hour incident is having the right metrics, logs, and traces at your fingertips. The Grafana stack — Prometheus, Loki, and Grafana itself — is the most widely adopted open-source observability solution for Kubernetes.

The Three Pillars of Observability

  • Metrics (Prometheus) — numeric time-series data: CPU, memory, request rate, error rate, latency
  • Logs (Loki) — structured event streams from your applications and infrastructure
  • Traces (Tempo) — distributed request tracing across microservices

Installing the Stack with Helm

The kube-prometheus-stack Helm chart installs everything — Prometheus, Alertmanager, Grafana, and a set of pre-built dashboards — in one command:

bash
# Add the Helm repo
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update

# Install kube-prometheus-stack
helm upgrade --install monitoring prometheus-community/kube-prometheus-stack \
  --namespace monitoring \
  --create-namespace \
  --values monitoring-values.yaml \
  --wait
yaml
# monitoring-values.yaml
grafana:
  adminPassword: "changeme-in-vault"
  ingress:
    enabled: true
    hosts: ["grafana.internal.mycompany.com"]
  persistence:
    enabled: true
    size: 10Gi

prometheus:
  prometheusSpec:
    retention: 30d
    storageSpec:
      volumeClaimTemplate:
        spec:
          resources:
            requests:
              storage: 100Gi

alertmanager:
  config:
    route:
      receiver: slack-notifications
    receivers:
      - name: slack-notifications
        slack_configs:
          - api_url: "https://hooks.slack.com/services/..."
            channel: "#alerts"

Adding Loki for Log Aggregation

Loki is Prometheus for logs — it indexes only metadata (labels) rather than full text, making it incredibly storage-efficient:

bash
helm repo add grafana https://grafana.github.io/helm-charts

# Install Loki stack (Loki + Promtail log collector)
helm upgrade --install loki grafana/loki-stack \
  --namespace monitoring \
  --set loki.persistence.enabled=true \
  --set loki.persistence.size=50Gi \
  --set promtail.enabled=true

Writing Effective PromQL Queries

PromQL is the query language for Prometheus. Here are the queries I use in every production dashboard:

promql
# Request rate per second (5m rolling window)
sum(rate(http_requests_total{job="my-service"}[5m])) by (status_code)

# 99th percentile latency
histogram_quantile(0.99,
  sum(rate(http_request_duration_seconds_bucket{job="my-service"}[5m])) by (le)
)

# Error rate as a percentage
100 * sum(rate(http_requests_total{status_code=~"5.."}[5m]))
  /
sum(rate(http_requests_total[5m]))

# Memory usage vs limit
container_memory_working_set_bytes{namespace="production"}
  /
container_spec_memory_limit_bytes{namespace="production"}

Tip

Use the RED method for service dashboards: Rate (requests/sec), Errors (error %), Duration (latency percentiles). These three metrics surface 90% of production issues.

Alerting That Doesn't Cry Wolf

Alert fatigue kills on-call effectiveness. Each alert should be actionable and map to a customer-impacting condition:

yaml
# alerting-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: service-alerts
  namespace: monitoring
spec:
  groups:
    - name: service.rules
      rules:
        - alert: HighErrorRate
          expr: |
            (
              sum(rate(http_requests_total{status=~"5.."}[5m]))
              /
              sum(rate(http_requests_total[5m]))
            ) > 0.05
          for: 5m
          labels:
            severity: critical
          annotations:
            summary: "Error rate above 5% for 5 minutes"
            runbook: "https://runbooks.internal/high-error-rate"
“

If your on-call engineer is getting paged more than once a night on average, your alerting is broken — not your system.

— Google SRE Book