The Ratio

Our weekly newsletter on reliability economics.

I run a benchmark that nobody asked for. 121 enterprise teams have taken it anyway. Every Tuesday I send you the one number that surprised me and the seven links that explain why it matters to me.

The newsletter is how I think out loud about what the data says.

No sponsors. No AI slop. Hit reply any time — I read everything.
Prefer RSS? reliabilityeconomics.com/blog/feed/the-ratio.xml

Issue №12August 4, 2026

The Ratio

A weekly newsletter on reliability economics


The Number

23 of 101

Nearly 1 in 4 organizations allocates less than 10% of their reliability budget to prevention, making their program structurally reactive by design.

23 of 101 organizations in the benchmark allocate under 10% of their reliability spend to prevention. Another 26 allocate between 10% and 24%. Combined, 49 organizations, close to half the benchmark, spend less than a quarter of their reliability budget on stopping failures before they happen.

Firefighting

A program where 90% of spend is reactive isn't a reliability program. It's a cleanup crew with a budget line. These organizations didn't choose to be reactive. They built a funding structure that can only produce reactive outcomes. This is the equivalent of a hospital that spends almost nothing on diagnostics and almost everything on the emergency room, then wonders why the ER is always full.

1 in 4 organizations doesn't have a reliability program. It has a faster way to clean up after things break.



The Crowd Favorite

  1. Sabotage — Beastie Boys — Most outages start at deploy time. Rollback velocity is your real MTTR metric.
  2. The Chain - 2004 Remaster — Fleetwood Mac — One broken upstream dependency collapses every downstream SLO. Circuit breakers are not optional.
  3. Danger Zone - From "Top Gun" Original Soundtrack — Kenny Loggins — Zero error-budget margin means the next routine change has nowhere to absorb failure.
  4. Immigrant Song - Remaster — Led Zeppelin — Sustained on-call without recovery windows compounds cognitive load and degrades MTTR across consecutive incidents.
  5. Free Bird — Lynyrd Skynyrd — A runbook too long to execute during an outage extends the outage.
Prevention

Five failure modes every deploy should survive


The Challenger — Comment of the Week

"Nobody ever got fired for adding another alert. So we have 847 of them and nobody knows which three actually matter."

847 alerts. That's not monitoring. That's noise with a pager attached.

Alert volume and signal stop moving together past a certain point. The fix isn't tuning thresholds. It's working backwards from SLOs. If an alert can't be mapped to something a user would actually notice breaking, it shouldn't wake anyone up.

Fewer alerts, anchored to SLOs, resolve faster. The audit is simple: name the SLO violation this alert represents. Can't? Silence it.

Firefighting

847 alerts. Zero signal.


The Ratio is a weekly newsletter by Florian Hoeppner.

Take the assessment → reliabilityeconomics.com/benchmark
Reply to this email with your take.

The Ratio

Our weekly newsletter on reliability economics.

I run a benchmark that nobody asked for. 121 enterprise teams have taken it anyway. Every Tuesday I send you the one number that surprised me and the seven links that explain why it matters to me.

No sponsors. No AI slop. Hit reply any time — I read everything.
Prefer RSS? reliabilityeconomics.com/blog/feed/the-ratio.xml