The Outage That Didn’t Happen
There is a meeting for every outage. The postmortem, the incident review, the root-cause writeup circulated to people who weren’t even on the call. When something breaks badly enough, an organization convenes, examines, documents, and files the lesson away. The ritual is well worn because failure earns it.
There is no meeting for the outage that didn’t happen.
No one gathers to study the quarter that closed without an incident. There is no writeup for the release that shipped without a hitch, no review of the launch that went so smoothly it produced no story worth telling. The entire apparatus of operations is built to examine failure, and it has almost nothing that does the same for success. The best thing an operation can produce is a non-event—and a non-event is the one outcome the organization has no reliable way to see.
Uptime is not the default
The reflex is to call that reasonable. Reliability is the baseline. Attention belongs on failure, because failure is where the problem lives. Why hold a meeting about a day when nothing went wrong?
The answer is that nothing going wrong is not the natural state of anything. A system left alone does not stay up. It drifts—dependencies shift beneath it, load climbs past what it was sized for, a certificate expires, a disk fills, an assumption that held for three years quietly stops holding. Uptime is not the absence of activity. It is the product of continuous, deliberate work that happens to leave no visible trace when it succeeds.
The standard figure makes the point without meaning to. Five nines—99.999 percent availability—is about five minutes of downtime a year, and it is defined entirely as a subtraction from perfect. There is no matching metric for the outages that were designed out, no positive score for the failure prevented six months before it would have occurred. The target is stated as minimized loss, never as manufactured strength, because the strength doesn’t show up anywhere it could be counted. Treat uptime as the default and you have already stopped seeing the labor that produces it.
The labor behind a quiet day
That labor starts long before anything is running, in the design.
A large enterprise network came under my management some years ago in worse shape than it looked. The design itself wasn’t bad, but it had no true redundancy—single paths where there should have been two, places where one failure would have taken everything downstream with it. Building the redundancy in ran across every layer: the VPN concentrators, the switches, the routers, the load balancers. Some of it wasn’t even a matter of new hardware. In places the second path was already paid for and simply unused—two network connections provisioned, one of them carrying the entire load while the other sat dark. Balancing traffic properly across both did two things at once: it turned a wasted link into a live second path, and it relieved the one that had been quietly running at its limit. Redundancy and better performance, out of capacity the organization already owned and wasn’t using. When the work was done, a device could drop and the next one would carry the load immediately, and the people depending on the network would stay asleep. The one exception proved the rule—a user tunneled in over VPN would see the client blink and reconnect against the other concentrator, and even that faint trace would be gone by morning.
That is the shape of the work. The engineer who argues for the second path, for the expense of capacity whose value only shows the day something fails, produces nothing anyone will ever point to. The failover fires in the dark. On-call sees it; the organization never registers it as a moment at all.
The same holds for how a system gets built. Careful implementation—a real test plan, the discipline to exercise the edge cases before production does it for you—buys a launch nobody remembers. Clean releases are forgettable by nature. The deployment that goes exactly as planned generates no war room, and therefore no memory. Nobody recounts the migration that worked. They recount the one that didn’t. The better the preparation, the less there is to say afterward.
The save that never becomes a story
Then the system is live, and the work changes shape again.
Monitoring is usually described as knowing when something is down. That is the smallest part of it. Knowing something is down is easy—the users will tell you. The real value sits upstream of the outage, in the drift you can see before it becomes a failure.
The monitoring on that same network was telling a slower, more specific story. It was running five to seven releases behind on its operating system—an unacceptable gap in a managed-services environment, and a dangerous one—and the graphs that charted performance and utilization showed the reason it mattered: known defects, already fixed in later releases, were quietly consuming resources and dragging the equipment toward instability. Nothing had failed. Everything was heading that way. While the upgrade plan worked its way through the change board, I scheduled device power-downs inside the maintenance windows to clear the degraded state and buy time—an unglamorous holding action no one outside the team would ever have reason to know about. Then we brought the whole environment onto a single stable release, ahead of the crash it was sliding toward.
Catch a slope like that in time and you do something strange: you convert a disaster into a Tuesday. The fix becomes a routine change made inside a maintenance window by someone who read a graph correctly. No incident is declared, because no incident occurred. No save is recorded, because from the outside there was nothing to be saved from.
And here is the part with no answer. A prevented outage and plain good luck look identical on the availability dashboard—both read one hundred percent. The engineer who read the slope and quietly headed off the failure, and the one who simply wasn’t unlucky that week, are indistinguishable there. The organization may keep the record—the change ticket and the monitoring history sit in a system somewhere—but it has no ritual for recognizing what the work prevented, the way it has one for everything that breaks. Whatever separates skill from fortune lives inside a judgment almost no one outside the room ever sees exercised.
Fixing what hasn’t broken yet
The slope on a monitoring graph runs on the scale of days. There is a slower slope underneath it, measured in years, and it is the harder one to act on.
Systems that run well for a long time set their own trap. The longer something works, the more finished it looks, and the more finished it looks, the harder it becomes to justify touching it. But the ground underneath keeps moving. The load that was comfortable becomes marginal. The architecture that fit the business three years ago fits a business that no longer exists. Someone has to decide, while everything is still working, that it is time to rebuild the part that hasn’t broken yet—to spend real effort and accept real risk on a system whose only visible problem is that it will have one eventually.
That decision produces the most invisible result of all. When it goes right, the failure it was meant to prevent simply never arrives. There is no before and after, because the after is the outcome that never came.
The network I have been describing was not a weekend’s work. Bringing it to heel took the better part of six months, and it was not happening in isolation—it was one piece of separating two large organizations during a parent-company divestiture, the kind of program where a single bad night becomes a story people tell for years. It ran straight through the client’s busiest, least forgiving stretch of the year, the season the entire business was built around, when an outage would not have been an inconvenience but a catastrophe.
Six months of upgrades, redundancy, and rebuilding, on critical infrastructure, during the worst possible window to touch any of it.
It went off without so much as a blip in service delivery.
Which means that to everyone the network served, and nearly everyone outside the team doing the work, nothing happened. There was no event. There is no anniversary of the outage that would have ended someone’s quarter, because the outage never came. The larger and more dangerous the save, the more total the silence it leaves behind.
Success subtracts itself
Every one of these is a win, and not one of them leaves a mark. The redundant design proves itself by going unnoticed. The clean launch is forgotten because it was clean. The prevented outage is indistinguishable from luck. The system rebuilt in time averts a crisis that never existed to be averted. Operational success is self-erasing. Done well, the work removes the evidence that it was ever needed.
That is the strange arithmetic of keeping things running, and it falls hardest on the people who are best at it. The reward for doing the job at its highest level is a long, quiet stretch of days in which nothing happens—which is indistinguishable, to almost everyone, from a job that was never hard to begin with. The better you are, the less anyone can tell you were there.
None of this is anyone’s cruelty. Absence cannot compete with catastrophe for attention; the failure that didn’t happen will never hold a room the way the one that did. But an operation that only ever gathers to study what broke is reading half its own history. It can tell you, in exhaustive detail, everything that has ever gone wrong. It can tell you almost nothing about what kept going right, or who kept it that way.
The outage gets a meeting. The disaster that never arrived gets nothing—no review, no story, no name. And the people who spend their careers producing those non-events come to understand the terms of the work: the surest sign you did it right is that no one can point to what you did. Success subtracts itself from view. The only proof it happened is a quiet you have to know how to read.
Discover more from At Ground Level
Subscribe to get the latest posts sent to your email.
