All case studiesRegulated industries · Multi-cloud
Took mean time to resolution from hours to minutes
When something broke across a hundred-plus clusters and two cloud providers, finding out where took longer than fixing it.
hours → minsMean time to resolution
100+Clusters under one view
The challenge
- Telemetry was siloed per operating company, each with its own stack and no shared view across the estate.
- Incidents spanned two cloud providers and 100+ clusters.
- Mean time to resolution ran to hours, most of it spent locating the fault.
- Regulatory separation between operating companies meant a single shared monitoring stack was not an option.
What we did
- Architected a federated observability platform, preserving tenant separation while allowing a cross-estate view.
- Standardised metrics, logs and traces on Prometheus, Grafana, Loki, Tempo and Mimir.
- Adopted OpenTelemetry so instrumentation was consistent regardless of language or cloud.
- Built dashboards and alerting around service ownership so alerts reached whoever could act on them.
Technology
PrometheusGrafanaLokiTempoMimirOpenTelemetryAWSGCP
Facing something similar?
Tell us what you are working on and we will come back with a straight answer on whether we can help, and what it would take.
Schedule Free Consultation


