Home / Articles / Disaster recovery plan
Business Continuity

A disaster recovery plan that actually works

Most companies have some kind of DR document - but it sits in a drawer, nobody has tested it, and during a real outage it turns out the contacts are a year out of date and nobody knows who's in charge. What a functional plan must include, and how it differs from business continuity.

A disaster recovery plan usually gets written once - typically for an audit or a customer requirement - and then nobody touches it again. The problem is infrastructure changes, people leave, and contacts go stale. A plan nobody has tested is more of an obstacle than a help during a real outage: the team spends the first hours figuring out whether the plan even still matches reality.

What a functional DR plan must include

01

RTO and RPO targets

Recovery Time Objective (how long a system can be down) and Recovery Point Objective (how much data loss is acceptable) - defined per critical system, not as a blanket number for the whole company.

02

Critical systems inventory

A list of systems with recovery priority - which must come back first, which can wait hours or days. Without priorities, the team improvises decisions during the outage itself.

03

Escalation chain

Real names, phone numbers and backups in case the key person isn't reachable. Not roles in an org chart, but actual contacts current as of today.

04

Failover infrastructure and procedure

Where the system switches to - a backup region, another provider, offline mode - and exact steps, not just a reference to "the infrastructure being redundant".

05

Communication plan

Who informs customers, who informs leadership, through which channel and at what interval - message templates prepared in advance instead of drafted under pressure mid-outage.

06

Data recovery procedure

Where backups are restored from, how their integrity is verified, and who has access to backups - including the case where the primary administrator is unavailable.

A plan without a clear owner is just a document. Someone specific has to be responsible for the plan existing, staying current, and having been tested - otherwise it goes stale like anything else without an owner.

Disaster recovery vs. business continuity

These two terms get conflated, but they answer different questions. Disaster recovery asks: how fast do we get systems and data back online? It's a technical answer - failover, restoring from backup, switching to backup infrastructure.

Business continuity asks: how does the company operate while that's happening, before everything is restored? That covers fallback manual processes, customer communication, deciding which activities get scaled back temporarily, and who runs operations while the IT team works on recovery. A functional DR plan is necessary but not sufficient on its own - a company needs both.

How often to test

A plan that isn't tested is just theory. I recommend at least once a year via a drill (tabletop or partial simulation), and every six months for critical systems. Testing usually surfaces more problems than writing the plan did in the first place - stale contacts, missing access, steps that don't actually take as long in reality as the plan assumes.

Frequently asked questions

What must a disaster recovery plan include?

RTO and RPO targets for each system, an inventory of critical systems with recovery priorities, an escalation chain with real names and contacts, a description of failover infrastructure and the failover procedure, and a communication plan for both the internal team and customers.

What is the difference between disaster recovery and business continuity?

Disaster recovery covers the technical restoration of systems and data after an outage - how quickly they come back online. Business continuity covers the broader question of how the company operates during the outage, before everything is restored - fallback processes, customer communication, running the business without key systems.

How often should a DR plan be tested?

At least once a year through a drill, ideally every six months for critical systems. A plan that isn't tested is just a document - in practice it's often outdated within a few months due to infrastructure changes.

Why do most DR plans fail in practice?

The most common reasons are a plan with no clear owner, an untested procedure, outdated contacts and systems, and missing decisions about priorities - who decides which system gets restored first when not everything can be recovered at once.

Have a DR plan you haven't tested in years?

On a free consultation we'll go through whether your current plan would hold up in reality - and what it needs to cover.

Book a free consultation →