The Ratio
A weekly newsletter on reliability economics
The Number
23 of 101
Nearly 1 in 4 organizations allocates less than 10% of their reliability budget to prevention, making their program structurally reactive by design.
23 of 101 organizations in the benchmark allocate under 10% of their reliability spend to prevention. Another 26 allocate between 10% and 24%. Combined, 49 organizations, close to half the benchmark, spend less than a quarter of their reliability budget on stopping failures before they happen.
A program where 90% of spend is reactive isn't a reliability program. It's a cleanup crew with a budget line. These organizations didn't choose to be reactive. They built a funding structure that can only produce reactive outcomes. This is the equivalent of a hospital that spends almost nothing on diagnostics and almost everything on the emergency room, then wonders why the ER is always full.
1 in 4 organizations doesn't have a reliability program. It has a faster way to clean up after things break.
This Week in Reliability
Observability Enters the Agent Era
The observability stack is being rebuilt for AI agents, not humans. From causal reasoning platforms that automatically trace root cause chains to LLM-native monitoring that treats prompts as infrastructure, the shift moves observability spend from reactive dashboards toward preventive agent context — but only if you can trust the reasoning layer.
Deep Reads
Turbo Charge Application Instrumentation & Root Cause Analysis
Causely Blog · Primary evidence — causal reasoning for agents
Causely combines with Odigos distributed tracing to provide automated root cause analysis through causal reasoning. The platform builds a live causal model from telemetry, identifying cause-and-effect chains between problems and symptoms in real time without manual instrumentation.
Causal AI moves observability from 'here's what broke' to 'here's why it broke and what to fix' — the difference between a signal and a diagnosis. This is preventive spend disguised as reactive tooling: if your agent knows the failure chain before the engineer opens a dashboard, you've shifted left on MTTR.
Root cause automation is the new alert rule.
How to build a trust platform for your agent with Grafana Agent Observability
Grafana Labs Blog · Primary evidence — LLM-native observability
Grafana Agent Observability provides monitoring infrastructure purpose-built for LLM workloads and agentic systems. Grafana Assistant evolved from an internal chatbot to a production agent requiring dedicated observability tooling beyond traditional metrics designed for pre-LLM infrastructure.
If you're running agents in production, your existing observability stack is blind to the failure modes that matter — prompt latency, token budget exhaustion, context window drift. This isn't incremental tooling; it's the recognition that agent reliability is a different problem with different economics. The trust tax is real, and it compounds.
Agent observability is not dashboards with LLM labels.
Reflections on AI Week, and the future of solving problems with observability and AI
Grafana Labs Blog · Vendor response — observability roadmap shift
Grafana Labs hosted AI Week focused on the intersection of observability and AI agents, with community engagement around agent-driven problem solving.
When the observability vendor hosts 'AI Week,' the category has officially pivoted — your metrics strategy is now an agent strategy.
Observability vendors are betting their roadmaps on agents.
Cortex completes OSTIF security audit
CNCF Blog · Adjacent signal — observability security baseline
Cortex, the long-term multi-tenant storage backend for Prometheus and OpenTelemetry, completed an Open Source Technology Improvement Fund security audit conducted by Quarkslab.
When your observability backend becomes agent infrastructure, security audits shift from nice-to-have to table stakes — agent trust depends on telemetry integrity.
Observability infrastructure is now agent attack surface.
Traditional versus resilience engineering views
Lorin Hochstein · Counter-argument — measurement philosophy matters
Lorin Hochstein compares traditional reliability engineering focus areas with resilience engineering priorities, arguing for differences in where teams should allocate scarce engineering cycles to improve reliability.
If agents are optimizing for traditional metrics (MTTR, SLOs), they inherit traditional blindspots — resilience engineering asks what happens when the model of the system is wrong.
Agent optimization is only as good as the metrics.
My talk from the Software Should Work conference
Lorin Hochstein · Adjacent signal — saturation as observability gap
Lorin Hochstein presented on saturation as a reliability concept at the new Software Should Work conference, a talk focused on one of his recurring themes in system reliability.
Saturation — the point where adding capacity stops helping — is the failure mode your agent won't see in the telemetry until it's too late.
Agents inherit the same saturation blindspots humans have.
The Crowd Favorite
- Sabotage — Beastie Boys ↗ — Most outages start at deploy time. Rollback velocity is your real MTTR metric.
- The Chain - 2004 Remaster — Fleetwood Mac ↗ — One broken upstream dependency collapses every downstream SLO. Circuit breakers are not optional.
- Danger Zone - From "Top Gun" Original Soundtrack — Kenny Loggins ↗ — Zero error-budget margin means the next routine change has nowhere to absorb failure.
- Immigrant Song - Remaster — Led Zeppelin ↗ — Sustained on-call without recovery windows compounds cognitive load and degrades MTTR across consecutive incidents.
- Free Bird — Lynyrd Skynyrd ↗ — A runbook too long to execute during an outage extends the outage.
Five failure modes every deploy should survive
The Challenger — Comment of the Week
"Nobody ever got fired for adding another alert. So we have 847 of them and nobody knows which three actually matter."
847 alerts. That's not monitoring. That's noise with a pager attached.
Alert volume and signal stop moving together past a certain point. The fix isn't tuning thresholds. It's working backwards from SLOs. If an alert can't be mapped to something a user would actually notice breaking, it shouldn't wake anyone up.
Fewer alerts, anchored to SLOs, resolve faster. The audit is simple: name the SLO violation this alert represents. Can't? Silence it.
847 alerts. Zero signal.
The Ratio is a weekly newsletter by Florian Hoeppner.
Take the assessment → reliabilityeconomics.com/benchmark
Reply to this email with your take.