Backup & Continuity
Why IT recovery plans fail (and a test checklist that finds the gaps first)
Most IT recovery plans fail for the same reasons: untested backups, one person with the passwords, no priorities, no communication plan. Find the gaps first.
By Dig IT SolutionsUpdated 8 September 20266 min read
Short answer
IT recovery plans fail for predictable reasons: backups that were never test-restored, documentation that assumes the author is present, admin access held by one person, no agreed order of restoration, decisions that wait for an unreachable director, and plans written for a business that has since changed. A realistic test finds each of these before an incident does.
Most businesses that suffer a serious IT incident had a recovery plan of some kind. What they discover is that a document written on a quiet afternoon does not survive contact with a Friday-night ransomware attack when the IT manager is on holiday. The reasons plans fail are consistent, and every one can be found in advance with a realistic test. This article sets out the failure modes and the checklist that exposes them.
Failures in backups, documentation and access
Backups that were never restored
The commonest failure. Backup jobs run, the dashboard is green, and nobody has attempted a full restore since the system was installed. On the day it turns out that a folder was excluded, the job has been failing silently since March, the backup is on the same network with the same credentials and has been encrypted too, or a full restore takes 30 hours over the internet when everyone assumed three. A backup that has not been restored is a hope, not a plan. The distinction is set out in backup vs disaster recovery.
Documentation that assumes its author is present
Recovery documents tend to say "restore the server from backup" and stop. Where is the backup? Which credentials? What order? How do you know it worked? The author knows, so the gaps never mattered. When the author is unreachable, a competent engineer with the document in hand still cannot start. Good documentation lets a suitably skilled person who has never seen the environment follow it to a working result.
One person holds the keys
In most SMEs one individual has the domain admin password, the backup console login, the cloud tenant admin, the firewall and the supplier relationships. That is sensible for security and fatal for recovery if they are on a plane. The same pattern appears outside IT: one finance manager who knows the payroll run, one director who can authorise emergency spending. Recovery planning has to assume that some key people will be unavailable, because incidents do not schedule themselves around annual leave.
The fix is not to spread admin rights widely. It is a shared password manager with break-glass access, a named deputy for each role who has actually used the access, and delegation of spending and communication authority written down.
Failures in priorities, decisions and communication
No agreed order of restoration
Without priorities, teams restore whatever they understand first. The file server comes back before the domain controller it depends on, the application server before the database, the archive before the phones. Each wrong order costs hours. A plan should list systems in restoration order with their dependencies and their recovery time objective, so nobody has to decide under pressure. How to set those targets is in RTO vs RPO explained.
Decisions that wait
Do we invoke the plan? Do we pay for emergency hardware? What do we tell clients? Do we notify the ICO? If those decisions sit with one director who is unreachable, recovery stalls with engineers ready to proceed and nobody authorised to say go. Delegation, agreed in advance and written into the plan, removes the wait.
Communication that is improvised
Email and Teams are down, so the usual channels are gone. Staff hear rumours. Customers ring a phone nobody is answering. Clients learn about the incident from a competitor. The plan needs a way to reach staff that does not depend on the failed systems, a holding message for customers agreed in advance, and a named person who speaks for the business.
Failures of maintenance and testing
Plans written for a business that no longer exists
The plan was written when there were two servers and no cloud. Since then there has been a Microsoft 365 migration, a new job management system, a second site and three staff changes. The contact list is out of date, the procedures describe systems that were decommissioned, and the recovery targets were never revisited. A plan should be reviewed at every significant change, not on an annual anniversary.
Cloud assumed to be someone else's problem
"It is in the cloud" is often the whole plan. Microsoft 365 is resilient as a service, but a deleted mailbox, a ransomware-encrypted OneDrive synced to SharePoint or a compromised admin account are your problem, and Microsoft's retention is not designed to solve them. Cloud recovery needs an independent backup, account recovery procedures and a way for staff to work when the connection is down.
Tests that prove nothing
Some businesses do test, in a meeting room, by reading the plan aloud and agreeing it seems fine. That confirms the document exists. It does not confirm that the backup restores, that the deputy can log in, or that the recovery time is achievable. The tests that find problems are the uncomfortable ones: real restores, against the clock, with the usual person excluded.
A recovery test checklist
Run this at least annually and after every major change. Score each item pass or fail and fix the failures.
Before the test
- Choose a realistic scenario (ransomware on the main server, loss of the office, Microsoft 365 admin account compromised).
- Nominate the person who usually does recovery and exclude them for the exercise.
- Set a target time from the recovery time objectives.
During the test
- Locate the plan without using systems that would be down in the scenario.
- Find the backup and confirm its most recent successful run and its retention.
- Confirm an immutable or offline copy exists that the scenario could not have destroyed.
- Restore a full server image to real or temporary hardware or cloud, using only the documentation.
- Restore a Microsoft 365 mailbox and a SharePoint library from the independent backup.
- Log in to every critical system as the deputy, using the shared password manager.
- Restore systems in the documented order and confirm dependencies are correct.
- Switch the phones to the divert plan and ring the main number.
- Send the staff notification through the out-of-band channel.
- Read the customer holding message and confirm who would send it.
- Identify who authorises emergency spending and confirm they know.
- Check the supplier and insurer contact list is current.
- Time every step.
After the test
- Compare each timing with the recovery time objective.
- List every question nobody could answer and every step that needed the excluded person.
- Update the documentation, access arrangements and contact lists.
- Record the test date and results for insurers and auditors.
- Schedule the next test.
The National Cyber Security Centre's guidance on offline backups in an online world is a useful reference for the backup elements, particularly the requirement for a copy an attacker cannot reach.
Building a plan that survives
Plans that work are built around process rather than people, kept short enough to be read during an incident, stored somewhere that does not depend on the systems that might fail, tested with the usual people absent, and revised whenever the business changes. Our IT disaster recovery plan template provides the structure, and the wider business continuity plan for SMEs covers the people and communication elements around it. Dig IT's backup and disaster recovery service includes the scheduled test restores that keep the plan honest.
What to do next
If your recovery plan has never been tested with its author absent, that is the test to run first. Book an IT health check and Dig IT will review your backups, documentation and access arrangements, and tell you which of these failure modes apply before an incident does.

