Correlation & Root Cause
Alerts that fired together, collapsed into one incident with a probable cause and the evidence behind it
Raw alerts (24h)
1,904
Correlated incidents
46
Open incidents
7
Mean time to RCA
6m
INC-4821 · Checkout failures across three services
Correlated from 34 alerts spanning 11 minutes
Signal timeline
payments-db replication lag 14.8sdatabase · orders (MySQL 8.0) — first signal
checkout-api p99 latency 4.2sapi · 22 alerts collapsed
AWS API Gateway 5xx rate 6.1%gateway · api-prod-us-east-1
Replication recovered, latency normalauto-resolved
Probable root cause
The orders replica fell behind during a bulk backfill. Checkout reads served from the replica timed out, and the gateway surfaced them as 5xx. Latency and error rate both recovered within 90s of replication catching up.
Evidence
- Lag rose before any API alert fired
- Backfill job
nightly-customer-loadoverlapped the window - No deploy, config change, or scaling event in the window
- Recovery order matched the failure order in reverse
Alert clusters
Near-identical alerts collapsed by message template — open a brief to hand one to a coding agent
No matching results
Correlated incidents
Alerts grouped by shared service, time proximity, and dependency edges in the service graph
No matching results


