Your dashboard says MTTR is 12 minutes. Ask your customer how long the incident actually felt — the two numbers are rarely close.
The problem
MTTR — mean time to resolve — usually gets measured from “alert fired” to “alert cleared.” That's a clean number, and it's misleading. It ignores the time between when the problem actually started and when it got detected, and it ignores the time between “the fix deployed” and “the system is verifiably healthy again.” Teams optimize the middle of the incident and quietly ignore the two ends, because the two ends are harder to measure and less flattering.
What MTTR should actually capture
Real incident lifecycle has four phases: time to detect, time to acknowledge, time to remediate, and time to verify. A team can have a blazing-fast “time to remediate” and still leave customers in a degraded state for 40 minutes because detection was slow or verification was skipped. Measuring only the phase you're good at is how a dashboard lies to leadership with a straight face.
How I've applied it
Improving MTTR through incident response work only became meaningful once detection and verification were pulled into the same conversation as remediation — better monitoring and alerting (Splunk, AppDynamics) closed the detection gap, and standardized SLOs gave a real, customer-felt definition of “verified healthy,” rather than “the error rate graph looks fine to me.”
Takeaways
- If you're not measuring time-to-detect, your MTTR is fiction
- “Fixed” and “verified” are two different timestamps — track both
- A fast remediation with slow detection is still a bad customer experience
- Good observability shrinks detection time more than good engineers shrink remediation time
- Report all four phases to leadership, not just the flattering one
Ravi Rapolu is a Senior Site Reliability / DevOps Engineer, currently open to AWS / Azure platform engineering and AIOps / LLMOps roles.
Connect on LinkedIn ↗ → page Ravi