Blog

The 1-10-100 Rule: Why a Dollar of Prevention Beats a Hundred Dollars of Outage

1 10 100 image

Nick Caravella

July 15, 2026

A defect costs about a dollar to prevent, ten to catch, and a hundred to live with. In a data center, that math isn’t a metaphor.

By Nick Caravella, Sr. Director of Growth, Cumulus

Sit in on the root-cause review after a serious outage and you rarely hear a story about something exotic. You hear about something small. A connection that was never verified to spec. A checkpoint skipped under schedule pressure. A test that got signed off on paper but was never really proven. A scope gap everyone assumed someone else owned.

The failure was expensive. The thing that caused it was cheap.

That gap, between what the cause cost and what the consequence cost, has a name. It’s the 1-10-100 rule, and once you see it, you can’t unsee it on a jobsite.

What the rule says

The 1-10-100 rule describes how the cost of a defect escalates the longer it survives. Spend $1 to prevent a problem before it’s built in. Spend $10 to catch and correct it during verification. Or spend $100 or more to deal with it once it’s reached operations and your customer is the one who found it.

The numbers are directional, not literal. The point isn’t the exact multiple: it’s the shape of the curve. Cost doesn’t rise gently as a defect moves downstream. It compounds. Every stage a problem survives adds people, rework, urgency, and risk to the eventual bill.

For most industries, that’s a useful way to think about quality. In data center construction, it’s closer to a law of physics, because the last stage isn’t a product recall. It’s tens of megawatts of critical load going dark.

The three stages, on a real project

Prevention: the $1 stage. This is the work on paper and in planning: clear scope, agreed quality controls, the right tools in the right hands before anyone installs anything. It’s the cheapest place in the entire lifecycle to fix a problem, because here a “fix” is an edit, not a tear-out. It’s also the hardest to value, because prevention is invisible when it works. Nobody writes an incident report about the outage that didn’t happen.

Verification: the $10 stage. This is QA/QC and commissioning: proving that what got built actually works before it carries load. Torque verified to spec. Insulation resistance tested. Systems started, integrated, and run at full load. Done well, verification catches the designed-in and installed-wrong defects while a fix still costs hours and a change order. Done as a paperwork exercise, it produces a binder full of signatures and a false sense of readiness.

I’ve written before about why a checked box isn’t proof: a signature tells you someone had a pen, not that the connection will hold under load. The 1-10-100 rule is the economic reason that distinction matters so much. A checkbox that lets a defect pass isn’t a $10 miss anymore. It’s a $100 one waiting for go-live.

Operations: the $100+ stage. This is the stage nobody budgets for and everybody pays. Systems down. Customers impacted. Emergency crews at premium rates. Reputation and SLA costs that don’t hit a line item until they do. This is where a quality issue stops being a quality issue and becomes a business event.

Why commissioning is where the bill comes due

My colleague Matt Kleiman recently walked through the five levels of data center commissioning, and it’s the clearest illustration of the 1-10-100 rule I know. Commissioning doesn’t create quality. It reveals whether quality was built in. By the time you reach full-load testing, you aren’t building anything. You’re finding out what you built.

And the level that decides everything is L2, installation verification: torque, wiring, alignment checked against spec. It’s the most underinvested level in the process, and it’s exactly the $1-to-$10 stage where a defect is still cheap to fix. Skip the rigor there and the defect doesn’t disappear. It hides, then resurfaces as a start-up failure, a cascading fault, or an outage, at the most expensive possible moment to discover it. Fix it at L2, or pay for it at every level above.

That’s the 1-10-100 rule, wearing a commissioning schedule.

Why the curve is getting steeper

It’s tempting to think the industry has this handled. In one sense it does: Uptime Institute reports that outage frequency on a per-site basis has been declining for several years. Design and operations really are improving.

But the cost per event is moving the other way. In Uptime’s 2025 survey, published in its 2026 Annual Outage Analysis, 57% of operators said their most recent major outage cost more than $100,000, and for the second year running, one in five put it above $1 million. Power remains the leading cause, with failures in UPS systems, transfer switches, and generators dominant. Those are precisely the systems whose fate is decided at installation verification.

Meanwhile the environment is getting harder, not easier: AI workloads, rising rack densities, grid instability, deeper interdependencies. More complexity means more places for a $1 miss to hide until it becomes a $100 event. The curve isn’t flattening. The stakes of a surviving defect are climbing even as raw failure rates fall.

The questions worth sitting with

Most teams can’t actually say where on this curve their defects are being caught, and that uncertainty is the risk. A few questions I’d ask of any quality program:

  • When a defect is found, do you know which stage it originated in, and how far it traveled before anyone noticed?
  • Is your verification proving readiness, or documenting that a task occurred?
  • If an outage happened next week and you ran the root cause, would the evidence still exist, with real values and timestamps, or just a signature?
  • Can the people verifying the work see what the people who specified it decided?
  • What did your last “small” issue actually cost, fully loaded: the rework, the retest, the schedule slip, the risk it introduced downstream?

If those are hard to answer, it’s usually not a failure of effort. It’s a failure of information flow. A defect stays cheap when the person best positioned to catch it can see it in time. It gets expensive when quality data lives in disconnected tools and nobody has a current, trustworthy view of whether the facility was built the way it was designed and proven the way it was promised.

Moving work down the curve

You don’t move defects from the $100 column to the $1 column by trying harder at the end. You do it by making prevention and verification systematic, connected, and provable, capturing quality data at the point of work, so commissioning confirms what was built instead of exposing what was missed.

That’s the whole idea behind what we build at Cumulus: verify the work where it happens, flag issues before they become defects, and generate audit-ready proof automatically, so a $1 miss never gets the chance to grow up. Because rework isn’t just slow; it compounds. And the most expensive defect in any data center is the one you’re still paying for at load.

It started as the cheapest one you never caught.

 

Nick Caravella is Sr. Director of Growth at Cumulus, a quality execution platform built for data center construction. He writes and speaks on quality systems, workforce development, and what it takes to deliver critical infrastructure at scale. Want to see how Cumulus verifies quality at the point of work? Talk with us

Share:

Schedule a demo to see Cumulus in action

Build custom digital workflows to guide field workers

Automate work completion and quality reporting

Integrate with powerful tools and technologies to get real-time work visibility