Most teams treat 100% uptime as the goal. That's not reliability — that's fear dressed up as ambition.
The problem
I've sat in enough incident reviews to know the pattern: a service goes down, leadership asks “how do we make sure this never happens again,” and the team responds by adding more approval gates, more manual checks, more caution. Six months later, releases take twice as long — and outages still happen, just less often and more painfully.
Chasing zero downtime doesn't make systems more reliable. It makes teams slower without making anything safer.
What an error budget actually is
An error budget flips the question. Instead of “how do we avoid all failure,” it asks: “how much failure can we afford, and what do we do with that room?”
If your SLO is 99.9% availability, your error budget is the 0.1% you're allowed to spend. That's roughly 43 minutes of downtime a month — and it's not a target to avoid. It's a resource to spend deliberately, on the releases and experiments that move the product forward.
Burn through the budget, and the rules change: freeze risky releases, prioritize reliability work. Have budget left over? Ship faster, take more risks, innovate harder.
How I've applied it
At NAB, introducing standardized SLOs, SLIs, and error budgets across business-critical services changed the conversation entirely. Instead of every outage triggering a blanket “slow everything down” reaction, teams had a shared, numeric answer to “are we reliable enough right now?” — and could make release decisions based on data instead of anxiety.
Takeaways
- 100% uptime is the wrong goal — it's usually not even what your customers need
- An error budget turns “reliability vs. velocity” from a debate into a formula
- When the budget's healthy, ship. When it's spent, stabilize. No politics required
- SLOs only work if they're tied to what customers actually feel — not vanity metrics
- The budget is a tool for trust between engineering and the business, not a punishment
Ravi Rapolu is a Senior Site Reliability / DevOps Engineer, currently open to AWS / Azure platform engineering and AIOps / LLMOps roles.
Connect on LinkedIn ↗ → page Ravi