Back to Lessons
Intermediate 25 min
Monitoring & Observability
Metrics, logs, traces, dashboards, alerting - know what your systems are doing at 3 AM.
“Monitoring tells you something is broken. Observability tells you what specifically is broken and why. You want the second one.”
What you'll learn
Metrics (Prometheus)Logs (Loki/ELK)Traces (Jaeger)Dashboards (Grafana)Alerting
The three pillars: metrics (what's happening), logs (what happened), traces (why it happened). You need all three. Prometheus for metrics, Loki or ELK for logs, Jaeger or Tempo for traces.
Grafana dashboards: create at least one dashboard per service with: request rate, error rate, latency (p50/p95/p99), and resource utilization. Alert on error rate spikes and latency degradation.