Skip to content
DevOps
Blog
DevOps11 min

The Real Cost of Downtime: A Board-Level Case

Matthieu Robin14 August 2025

You are in a board meeting. There has been a 2-hour outage. The finance director asks the obvious question: "How much did that cost us?" You do not have a good answer. You know it was bad, you remember the chaos, but you cannot put a real number on it.

This is exactly the problem that blocks reliability investment in most organisations: when you later ask for a major budget to invest in redundancy, disaster recovery and reliability infrastructure, your board asks the same question in reverse ("What's the ROI?") and the conversation stalls because nobody on either side has the underlying numbers. If you cannot calculate the cost of downtime, you cannot make a financial case for preventing it.

This article walks you through the full calculation methodology so that, by the end, you have a defensible board-ready figure: "Every hour of downtime costs us X, therefore investing Y in reliability infrastructure has a payback period of Z months."

The Obvious Costs

Most companies only count the most obvious cost: Lost sales during the outage.

For an e-commerce company:

  • Annual revenue: in the tens of millions
  • Daily revenue: annual revenue divided by 365
  • Hourly revenue: a few thousand per hour
  • 2-hour outage: twice the hourly figure in lost revenue

This is real, but it's not the whole story.

The Hidden Costs (Usually Larger)

1. Lost Productivity of Affected Users

If your application is a business tool (Salesforce, ERP, accounting software), users can't work during downtime.

Example: SaaS CRM used by 100 sales reps

  • Average sales rep salary: a typical mid-level salary
  • Hourly cost: the corresponding hourly rate
  • 2-hour outage: 100 reps × hourly rate × 2 hours = several thousand
  • Plus they need to catch up (additional 4 hours of cleanup): roughly twice as much again
  • Total productivity cost: about three times the direct outage figure

2. Customer Churn and Reputation Damage

Downtime damages customer relationships.

Typical customer behavior after a major outage:

  • High-touch customers (50% of revenue): Don't churn immediately, but increase risk of future churn
  • Commodity customers (25% of revenue): Evaluate alternatives, 5-10% churn within 30 days
  • Enterprise customers (25% of revenue): File complaints, expect service credits, increase scrutiny

Calculating churn cost:

  • Assume 5% of customers churn within 30 days
  • Average customer lifetime value: a mid-range lifetime value
  • Churn cost: 50,000 customers × 5% × lifetime value = tens of millions

That's catastrophic. But if we assume a 2-hour outage causes only 0.5% churn:

  • Churn cost: 50,000 × 0.5% × lifetime value = a few million

3. Customer Support Escalation

Support volume spikes during outages.

During a 2-hour outage:

  • Normal support volume: 20 tickets/hour = 40 tickets over 2 hours
  • Outage support volume: 200 tickets/hour = 400 tickets over 2 hours
  • Plus 6 hours of overflow after outage as backlog clears

Cost calculation:

  • Support staff: 5 agents at a standard hourly rate
  • Normal 2-hour cost: a small baseline amount
  • Outage 2-hour cost: roughly ten times that baseline
  • Cleanup overflow (6 hours): a modest additional amount
  • Escalation (customers angry, longer calls): another modest amount
  • Total support cost: several thousand

4. Incident Response Costs

Your team spends time responding to the outage.

During the incident:

  • 5 engineers on-call: 2 hours
  • 2 manager/leads coordinating: 2 hours
  • CTO/VP involved: 0.5 hours
  • Total engineer hours: 11.5 hours

Cost:

  • 11.5 hours × (average cost per engineer hour): a standard hourly rate = under a thousand

5. Post-Incident Investigation

After the outage, you investigate root cause.

Typical post-incident work:

  • Root cause analysis: 4-6 hours
  • Documentation: 2-3 hours
  • Engineering improvements: 20-40 hours
  • Total: 30-50 hours

Cost:

  • 40 hours × the same hourly rate = a few thousand

6. Compliance and SLA Penalties

If you have SLA commitments, you owe credits.

Example SLA:

  • 99.5% uptime commitment
  • A fixed penalty for each 0.1% below SLA
  • 2-hour outage on system running 99.95% SLA baseline
  • Drops to 99.93% SLA
  • Penalty: twice that per-increment penalty

7. Lost Business Opportunities

During downtime, you can't sign new customers or close deals in progress.

Example:

  • Sales team is in middle of 5 deal closures
  • Total deal value: a few million across the pipeline
  • 2-hour delay could result in 1 deal lost (20% probability)
  • Expected loss: 20% of that pipeline value

This is harder to quantify but real.

Building Your Downtime Cost Model

Here's a framework to calculate your specific downtime cost:

For each of your critical systems, estimate:

Cost Factor Formula Your Number
Lost revenue per hour (Annual revenue / 8760 hours) _____
User productivity loss (# users × hourly cost × hours affected) _____
Support escalation (Support staff × hourly cost × surge hours) _____
Customer churn (Lost customers × LTV) _____
Incident response (Engineer hours × avg cost) _____
Post-incident cleanup (Investigation hours × avg cost) _____
SLA penalties (# customers × contract penalty) _____
Total per hour _____

Example company numbers:

  • Lost revenue: a few thousand/hour
  • Productivity loss: the largest single line, several thousand/hour
  • Support escalation: a modest amount/hour
  • Customer churn risk: several thousand/hour (conservative)
  • Incident response: a minor amount/hour
  • Post-incident (amortized): a modest amount/hour
  • SLA penalties: a modest amount/hour
  • Total: a substantial hourly figure

A 2-hour outage costs this company twice that hourly figure.

A 4-hour outage costs four times as much.

The Financial Case for Reliability Investment

Once you know your downtime cost, you can justify reliability investments.

Example: Disaster Recovery Infrastructure

Your team proposes: a major upfront infrastructure investment plus modest annual ops costs

Reliability improvement:

  • Current system: 99.5% uptime (3.7 hours downtime/year)
  • With DR infrastructure: 99.99% uptime (0.36 hours downtime/year)
  • Improvement: 3.34 fewer hours of downtime/year

Financial benefit:

  • 3.34 hours × the hourly downtime cost = a five-figure annual saving in downtime cost avoided

ROI calculation:

  • Year 1: annual benefit minus the large upfront cost = a large net loss (investment year)
  • Year 2: annual benefit minus modest ops cost = a modest net gain
  • Year 3: annual benefit minus modest ops cost = a modest net gain
  • Payback period: 6 years
  • 10-year NPV: modestly positive

This is a weak case financially (6-year payback), so you need another angle:

  • Reputation value (retention, future revenue)
  • Competitive advantage (market cap impact)
  • Compliance requirements

The Board Presentation

Here's how to present this to your board:

1. Lead with the number: "A single outage costs us a substantial amount per hour in lost revenue, support costs, and customer churn risk."

2. Show historical impact: "In the last year, we experienced 8 hours of downtime. That's a six-figure sum in direct cost, not including reputation damage and customer churn."

3. Show industry benchmarks: "Amazon reported that a 1-hour outage costs them several million. Our hourly cost puts us in line with businesses our size."

4. Present the investment case: "Investing in disaster recovery infrastructure would reduce downtime from 3.7 hours/year to 0.36 hours/year. That's a five-figure annual benefit, with payback in 6 years."

5. Add strategic value: "Beyond financials, reliable infrastructure is a competitive advantage. It improves customer retention, enables us to win deals against less-reliable competitors, and gives us confidence to scale."

Avoiding Common Mistakes in This Calculation

Mistake 1: Only Counting Lost Revenue

The most common mistake: Only counting the transactions you lose during downtime.

Problem: For businesses with thin margins (like SaaS), lost revenue understates actual cost.

Solution: Include all costs: productivity, support, reputation, churn risk.

Mistake 2: Being Too Conservative

Some teams assume "nobody will churn over a 2-hour outage."

Problem: This leads to undervaluing reliability.

Solution: Base churn assumptions on actual data (survey customers, analyze historical churn).

Mistake 3: Ignoring Indirect Costs

The most expensive cost is often reputation damage and customer churn.

Solution: Model this explicitly, even if you have to estimate.

Mistake 4: Underestimating Investigation Time

Post-incident investigation and engineering often takes 40+ hours.

Solution: Track actual post-incident time, include in calculations.

Using This Number Operationally

Once you know your downtime cost, use it:

1. For investment decisions:

  • "This reliability project costs a six-figure sum and prevents 1 hour of downtime/year"
  • "That's the value of one hour of downtime avoided, so ROI is 24%, payback is 4 years"
  • Reject projects with poor ROI

2. For on-call rotation decisions:

  • If on-call engineer burn-out causes missed incident response (adding 30 min to MTTR)
  • That's half an hour of downtime cost added
  • Worth investing a modest annual budget in better on-call tooling/processes

3. For customer communication:

  • If customer complains about reliability, show them the improvement plan
  • "We experienced 2 hours of downtime at a significant cost. We're investing in redundancy to prevent this."

4. For hiring and staffing:

  • "A 2-hour incident requires 3 engineers for 4 hours each (12 hours total)"
  • "The cost of that incident is a five-figure sum, cost of engineering is under a thousand"
  • "Investing in better monitoring and automation would be 20:1 ROI"

Building the Business Case for Reliability

Downtime has a real, calculable financial cost.

For most companies, that cost is much larger than they assume.

Once you quantify it, you can make rational financial decisions about reliability investment.

Your board won't approve a major infrastructure project based on "we need better reliability."

But they will approve a "major project prevents a five-figure annual downtime cost (6-year payback, plus strategic benefits)."

Calculate your number. Use it.

It transforms reliability from a cost center to an investment with measurable ROI.

Related reading:


Ready to build reliable infrastructure? Hidora helps enterprises design resilient systems: Infrastructure Consulting · Managed Services · Disaster Recovery Solutions

Matthieu Robin

CEO & Co-founder

Founder of Hidora, passionate about cloud-native and Swiss digital sovereignty. 15+ years in the cloud ecosystem.

Does this article resonate?

Hidora can support you on this topic.

Need support?

Let's talk about your project. 30 minutes, no strings attached.