
I spend a slightly ridiculous amount of my life training for running events. Marathons in particular require months of following a plan, looking at data and trying to predict what you might eventually be capable of on race day. Apps like Runna can take your recent performance, training volume and pace and tell you that, if everything goes according to plan, you should be capable of running a particular time.
The important bit there is if everything goes according to plan.
Anyone who has trained for a marathon knows it rarely does. You get ill. Your hip starts hurting. You miss a few sessions. Work gets in the way. A two-week injury costs you fitness. Or you arrive on race day after the perfect training block and discover it is unexpectedly hot, windy or you simply feel terrible.
The prediction was not necessarily wrong. It was based on the evidence available at the time. The problem is treating the prediction as though it were the result.
I think we make exactly the same mistake with resilience.
We write an RTO of four hours into a recovery plan and, over time, four hours stops being treated as an objective and starts being treated as a fact. The business plans around it. Risk gets calculated around it. Someone may even have tested it once.
Then the incident happens.
Reality does not care what was written in the plan.
I sit with a lot of CISOs. Peer dinners, hallway conversations at Black Hat, WhatsApp groups and those calls where nobody is presenting anything to anybody. Ask what threats people are worried about and we can happily talk for an hour. Ask a much simpler question, what does an hour of downtime actually cost the business, and how long until the revenue comes back, and the conversation gets considerably harder.
I include myself in that. For a long stretch of my career I could have walked you through our threat landscape, controls and vulnerabilities in enormous detail. I would have been much less comfortable sitting opposite a CFO and defending exactly what twelve hours without a critical system would cost us.
Once you notice that gap, it feels strange. We spend enormous amounts of time reporting threats to boards while often struggling to report downtime in the language the board already uses to run the company.
Start with what the market is measuring. Splunk and Oxford Economics surveyed 2,000 executives at the world's largest companies for their 2026 study and put the aggregate cost of unplanned downtime across the Global 2000 at $600 billion a year, a 50 percent increase in two years. The average comes out at around $15,000 per minute, while a single incident was associated with a 3.4 percent decline in stock price.
Those are enormous companies, so it is worth pressure-testing that against other research. ITIC's 2024 hourly cost of downtime survey found that 41 percent of enterprises put a single hour of downtime somewhere between $1 million and more than $5 million. Uptime Institute's 2024 outage analysis found outages becoming less frequent but more expensive, with 54 percent of significant outages exceeding $100,000 and 16 percent exceeding $1 million.
Then there are the incidents we can actually watch play out. When ransomware hit Marks & Spencer in 2025, the company initially warned that the incident could cost around £300 million in operating profit. Once the year closed, M&S disclosed £136 million in gross costs associated with the attack.
There are obviously differences between all of those numbers, companies and methodologies. The interesting part for me is the direction. Downtime has a measurable financial cost, and the longer recovery takes, the larger that number becomes.
Which means recovery time is not simply an engineering metric.
It is a financial one.
Most recovery plans already contain two numbers we tend to treat as technical targets living somewhere inside a runbook.
Recovery Time Objective is how long we are willing to be down. Recovery Point Objective is how much data we are willing to lose, measured in time.
Four hours RTO. One hour RPO. Nice, neat numbers.
But translate them into money and they start to look very different.
If a critical service generates or supports $100,000 of revenue every hour, a four-hour RTO means the business is accepting up to $400,000 of direct revenue exposure before we have considered missed transactions, recovery labour, contractual penalties, regulatory consequences or customers deciding to go elsewhere.
An RPO of one hour means accepting the potential loss or reconstruction of an hour of business data. Depending on the system, that could mean lost transactions, customer activity or operational records that now have to be reconstructed.
These are not really technology decisions. They are business decisions expressed through technology metrics.
That is why I think CISOs should be able to translate them.
Instead of telling the board, “Our RTO is four hours,” we should be capable of saying, “Every hour this service is unavailable costs approximately $X. We currently expect to restore it within four hours, meaning our direct exposure is approximately $Y.”
There is a very important word in that sentence though.
Expect.
Because your RTO is still only an objective.
You do not need another vendor tool or a six-month consulting engagement to produce a useful downtime number. You can build a defensible starting point from five inputs, and I would begin with the three systems most directly connected to your ability to operate or generate revenue.
Start with revenue per hour. Take the annual revenue that depends on the system and divide it by the number of hours it operates. For an always-on platform that is 8,760 hours per year. A company generating $200 million annually through that platform therefore generates roughly $22,800 an hour. It is deliberately simple, but it gives you somewhere sensible to start.
Then consider the dependency factor. An outage does not necessarily mean every dollar disappears. Some systems stop revenue completely, while others cause transactions to queue or customers to move to another channel. Estimate the percentage of revenue you genuinely lose while that system is unavailable. An ecommerce checkout system during peak trading might be close to 90 percent. An internal reporting platform might be closer to 10 percent. Multiply the revenue-per-hour number by that percentage and you have revenue genuinely at risk each hour.
Now apply your RTO and RPO. Multiply revenue at risk per hour by your stated RTO and you have the direct revenue exposure associated with one incident. Then consider the value of the data covered by your RPO. If your RPO is one hour, what does losing or reconstructing that hour actually cost the business?
Then add the things that never appear neatly on the revenue line. Engineers and incident responders working through the night cost money. There may be SLA penalties, contractual payments, regulatory consequences and external recovery costs. The same Splunk research puts regulatory fines arising from downtime at an average of $51 million per organisation for the largest firms. It also found 81 percent of technology leaders citing customer attrition as a consequence of downtime, while nearly one in five marketing professionals said brand recovery can take a full quarter.
You are unlikely to know every one of those numbers perfectly. That is fine. Put a conservative estimate against them and keep the assumptions visible. I would rather take an imperfect number into a boardroom and explain exactly how we calculated it than present a precise-looking RTO that nobody has ever proved.
Finally, multiply that number by realistic incident frequency. How many incidents of that severity can you honestly expect in your environment? One serious event a year is hardly an outrageous assumption for a large organisation. Multiply the per-incident exposure by realistic frequency and you now have an annualised cost of downtime you can defend line by line.
Published models approach the problem in similar ways. Atlassian, for example, uses minutes of downtime multiplied by a cost-per-minute estimate, with figures around $427 for small businesses and $9,000 for mid-to-large organisations. I would treat calculations like these as a floor because real incidents accumulate costs through labour, recovery, contractual impact and downstream consequences.
The important thing is not pretending your number is perfect.
It is understanding what downtime means financially before an incident forces somebody else to calculate it for you.
There is an obvious problem with everything I have just calculated.
It assumes the RTO is true.
If our plan says four hours, the financial model assumes that when something goes wrong the system will actually be back within four hours.
The evidence suggests that assumption deserves considerably more scrutiny.
Across surveyed organisations, the mean success rate for meeting a defined RTO on mission-critical applications is around 64 percent. In other words, teams miss their own recovery targets roughly a third of the time.
Ransomware shows how extreme the difference can become. Sophos found that 53 percent of victims were back within a week in 2025, while around 18 percent took longer than a month to recover.
Think about the difference between those numbers and the recovery objectives sitting in most business continuity plans.
If your financial model assumes four hours of downtime and recovery actually takes twelve, your direct downtime exposure is three times the number you presented.
If it takes four days, you have a completely different business problem.
And if ransomware recovery takes several weeks, an RTO measured in hours starts looking rather theoretical.
This is where the marathon analogy comes back for me.
Imagine my training data says I can run a three-hour marathon. Then I get injured, miss three weeks of training and lose a significant amount of fitness. It would be ridiculous for me to keep telling everyone I am a three-hour marathon runner simply because that is what the plan predicted before the injury.
Yet our infrastructure changes constantly and we often continue treating the recovery number exactly the same way.
Infrastructure as code redeploys environments. Backups rotate. Applications change. Dependencies move. Permissions change. Engineers join and leave. SaaS providers update their platforms. The environment that existed when the recovery plan was written or tested is not necessarily the environment we need to recover today.
This is not about teams failing.
It is about environments moving faster than traditional recovery testing.
Traditional disaster recovery testing can make this worse because we normally control the test.
We choose when it happens. We know the scenario. We can make sure the right engineers are available. We prepare beforehand. We test against a relatively clean environment and, unsurprisingly, very few organisations deliberately construct a DR exercise they expect themselves to fail.
That is useful for validating parts of a recovery process.
It is not the same thing as proving what will happen during a real incident.
Real incidents do not wait until the right engineer is online. They do not give you two weeks to prepare. They may affect the backup infrastructure you planned to rely on. They happen against whatever version of your environment happens to exist at that exact moment.
That is why I increasingly think resilience has to be treated as something continuous.
Can we recover today?
What changed since yesterday?
Did the latest deployment introduce a dependency we have not accounted for?
Are the backups we expect to use actually recoverable?
If a significant part of the environment disappeared, what is the Minimum Viable Business we would have to restore first?
And, most importantly, does the recovery time we can actually demonstrate match the recovery time the business has priced into its risk?
The difference between those two numbers is the interesting part.
Because that gap is effectively unbudgeted risk.
Downtime is coming. No CISO can promise otherwise.
What we can do is understand what that downtime costs, decide what absolutely has to keep running and continuously improve our confidence in how quickly we can recover it.
That changes the conversation with the board.
Instead of reporting threats and controls, we can talk about business outcomes. We know approximately what an hour costs. We know which services create the largest exposure. We know what recovery objectives the business has accepted. And, critically, we know whether the environment currently supports those objectives.
This is also a big part of why I joined Gambit.
Gambit was built around the gap between the recovery you plan for and the recovery you would actually get. It continuously maps the live environment across cloud, infrastructure as code, backups, security controls and on-prem infrastructure so organisations can see where recovery would hold and where it would break before an incident makes the discovery considerably more expensive.
For me, that comes back to running.
Before a marathon I can have all the predictions I want. Runna can predict a time. Garmin can predict one. My training data can say I am in the shape of my life.
Then the gun goes off.
At that point the prediction is irrelevant. Eventually I have to cross the finish line and look at the clock.
We need the same relationship with recovery.
An RTO is useful for planning. An RPO is useful for understanding tolerance. A recovery plan is absolutely worth having.
But none of them proves you can recover.
Run the downtime calculation against your most important systems. Work out what every additional hour really costs. Then ask the harder question: is the recovery time we have priced into the business the recovery time we can actually achieve today?
Until you have proved it, it is still just a prediction.
Recovery Time Objective is how long you are willing to be down. Recovery Point Objective is how much data you are willing to lose, measured in time. RTO covers the duration of the outage, while RPO covers the gap between your last good copy of the data and the moment of failure. Both are usually written as engineering targets. In reality, both are financial commitments the business has made, often without explicitly pricing them.
Splunk and Oxford Economics put the average cost of downtime across the Global 2000 at around $15,000 per minute, which works out at approximately $900,000 an hour. ITIC found that 41 percent of enterprises put the cost of a single hour somewhere between $1 million and more than $5 million. Those are useful reference points, but averages at that scale should be a starting point rather than your number. Your actual exposure depends on revenue, system dependency, time of day, contractual commitments and the wider consequences of the outage.
Start with the annual revenue that depends on the system and divide it by the number of hours the system operates. Multiply that by the percentage of revenue genuinely lost while the service is unavailable to calculate revenue at risk per hour. Then apply your RTO and RPO to estimate exposure per incident. Add recovery labour, SLA penalties, regulatory consequences, customer attrition and other downstream costs, then multiply the total by a realistic annual incident frequency. The goal is not perfect precision. It is a defensible number with assumptions the business can understand.
No. RTO and RPO are objectives. They describe the recovery performance the organisation wants or has agreed to tolerate. They do not prove the environment can achieve it. Across surveyed organisations, teams meet their defined RTO on mission-critical applications around 64 percent of the time. The difference between the objective on paper and the recovery time you actually achieve is where a significant amount of unbudgeted downtime risk sits.
One major reason is that environments move faster than runbooks. Infrastructure as code redeploys environments, backups rotate, dependencies change, applications are updated and people move roles. A recovery process successfully validated months ago may no longer represent the environment that exists today. Traditional DR exercises also normally happen under controlled conditions. Teams know a test is happening, the right people can be available and preparation can take place beforehand. A real incident provides none of those guarantees. A recovery plan rarely fails while somebody is reading it. It fails when it meets reality.
The latest from Gambit: research, insights, and live sessions

