Reliability & Observability Baseline
A shared SLO, alerting and instrumentation baseline that made service health legible across teams.
- Domains
- Reliability
- Cloud
- Platform
- Technologies
- Prometheus
- Grafana
- OpenTelemetry
- Loki
- Alertmanager
Overview
One instrumentation and alerting contract every service inherits, so on-call does not require service-specific tribal knowledge.
Context
Each service had its own dashboards, its own alert thresholds and its own definition of healthy.
Problem
Incidents were slow to diagnose because responders had to learn the service before they could read its signals.
Constraints
Cardinality and retention cost had to stay bounded, and instrumentation could not require rewriting existing services.
Architecture
OpenTelemetry instrumentation feeds a shared metrics and logs stack. SLOs are declared next to the service and rendered into dashboards and alerts automatically.
Decisions
Alerted on symptoms defined by SLO burn rate rather than on resource-level thresholds.
Implementation
Started with the three services that generated the most pages, then generalised the pattern.
Results
Placeholder — measured outcomes will be published with the real case study.
Lessons
Deleting alerts improved reliability more than adding them.