Disaster recovery planning and testing — recovery objectives agreed, restore sequence documented, dependencies mapped and the whole thing rehearsed before it is needed.
The gap between backup and recovery
Ask most businesses whether they could recover from losing their systems and the answer is “yes, we have backups”. Ask how long it would take, in what order things would come back, who would do it, and where the software licence keys are, and the answer changes.
Backup is a prerequisite. Disaster recovery is the capability, and the difference is measured in days of downtime.
What a plan has to answer
What comes back first. Ranked by business impact, not by what is easiest. If invoicing stops the business and the intranet does not, invoicing is first.
In what dependency order. This is where recovery goes wrong most often. Restoring the application server before the authentication and database services it relies on means restoring it twice. The order needs to be mapped once, in advance, by someone who knows the environment.
How long each step actually takes. Restoring several terabytes from an offsite copy over a business internet connection takes as long as it takes. Measuring that once means the plan reflects reality instead of optimism.
Who does what. Named people, with a fallback for each. Who can authorise emergency spend. Who contacts the insurer — often required early under the policy. Who tells staff, and what they are told.
How anyone gets access. Credentials for the recovery systems, stored so they are available when the environment holding them is down. This sounds obvious and is missed constantly.
Objectives, set per system
Not everything deserves the same target, and treating it that way is how recovery plans become unaffordable and get abandoned.
Your finance system might justify a four-hour recovery time and near-zero data loss. A document archive might tolerate two days. Deciding this system by system is what keeps the design proportionate — the expensive engineering goes only where the business genuinely needs it.
Local conditions belong in the plan
South-East Queensland businesses are considerably more likely to lose a site to flooding or storm than to a targeted cyber attack. A recovery plan that only contemplates ransomware is half a plan.
The practical differences are worth thinking through: whether your offsite copy is far enough away to be unaffected by the same weather event, whether staff can work productively without the office, whether your phone system follows you, and whether anyone needs physical access to premises that may be inaccessible for a week.
Testing is where the value is
Every disaster recovery test we run finds something. A dependency nobody documented. A licence tied to a decommissioned server. A credential held only by someone who left. A restore that takes four times longer than assumed.
None of those are failures of the test — they are the reason the test exists. Each one found in a rehearsal is one not discovered during an actual outage, when it costs a day.