A disaster recovery plan explains how an organisation will restore technology services after a major interruption.
Disruptions may be caused by hardware failure, ransomware, fire, theft, power problems, human error, software corruption, connectivity failure or cloud-service outages.
A useful recovery plan identifies critical services, recovery priorities, responsible people, backup locations, communication procedures and testing requirements.
Identify critical business services
Begin by identifying the systems that the organisation requires to operate. These may include email, accounting systems, customer databases, file storage, telephony, websites, cloud platforms and line-of-business applications.
- Which services are essential?
- How long can each service remain unavailable?
- How much data can the business afford to lose?
- Which users and departments depend on each service?
- Which suppliers or platforms are required for recovery?
Define recovery objectives
The recovery time objective defines the target period for restoring a service. The recovery point objective defines the acceptable amount of data loss measured in time.
A system requiring restoration within two hours normally needs a different design and budget from a system that can remain offline for two days.
Build a reliable backup strategy
- Maintain more than one copy of important data
- Keep at least one copy in a separate location
- Protect backups from unauthorised deletion
- Encrypt sensitive backup data
- Monitor backup jobs and investigate failures
- Test restoration regularly
- Document retention requirements
Document recovery responsibilities
The plan should identify who declares a disaster, who contacts suppliers, who restores systems, who communicates with employees and customers, and who authorises emergency expenditure.
Contact information should be available outside the systems that may be affected by the outage.
Test the plan
A recovery plan should not be trusted until it has been tested. Testing may include individual file restoration, server recovery, cloud failover, communication exercises or full operational simulations.
Each test should record results, delays, missing information and corrective actions.
Disaster recovery checklist
- Critical systems have documented owners
- Recovery priorities are approved by management
- Recovery time and recovery point targets are defined
- Backups are monitored and protected
- Restoration tests are completed
- Supplier contacts are documented
- Administrative credentials are securely available
- Alternative communication methods are defined
- The plan is reviewed after major changes
- Employees understand their responsibilities