Your dashboard says MTTR is 12 minutes. Ask your customer how long the incident actually felt — the two numbers are rarely close.

The problem

MTTR — mean time to resolve — usually gets measured from “alert fired” to “alert cleared.” That's a clean number, and it's misleading. It ignores the time between when the problem actually started and when it got detected, and it ignores the time between “the fix deployed” and “the system is verifiably healthy again.” Teams optimize the middle of the incident and quietly ignore the two ends, because the two ends are harder to measure and less flattering.

What MTTR should actually capture

Real incident lifecycle has four phases: time to detect, time to acknowledge, time to remediate, and time to verify. A team can have a blazing-fast “time to remediate” and still leave customers in a degraded state for 40 minutes because detection was slow or verification was skipped. Measuring only the phase you're good at is how a dashboard lies to leadership with a straight face.

How I've applied it

Improving MTTR through incident response work only became meaningful once detection and verification were pulled into the same conversation as remediation — better monitoring and alerting (Splunk, AppDynamics) closed the detection gap, and standardized SLOs gave a real, customer-felt definition of “verified healthy,” rather than “the error rate graph looks fine to me.”

Takeaways

When your team reports MTTR, does the number include how long it took to notice the problem in the first place?

Ravi Rapolu is a Senior Site Reliability / DevOps Engineer, currently open to AWS / Azure platform engineering and AIOps / LLMOps roles.

Connect on LinkedIn ↗ → page Ravi