Skip to main content
Dig IT Solutions logo

Backup & Continuity

Why IT recovery plans fail (and a test checklist that finds the gaps first)

Most IT recovery plans fail for the same reasons: untested backups, one person with the passwords, no priorities, no communication plan. Find the gaps first.

By Dig IT SolutionsUpdated 8 September 20266 min read

Short answer

IT recovery plans fail for predictable reasons: backups that were never test-restored, documentation that assumes the author is present, admin access held by one person, no agreed order of restoration, decisions that wait for an unreachable director, and plans written for a business that has since changed. A realistic test finds each of these before an incident does.

Most businesses that suffer a serious IT incident had a recovery plan of some kind. What they discover is that a document written on a quiet afternoon does not survive contact with a Friday-night ransomware attack when the IT manager is on holiday. The reasons plans fail are consistent, and every one can be found in advance with a realistic test. This article sets out the failure modes and the checklist that exposes them.

Failures in backups, documentation and access

Backups that were never restored

The commonest failure. Backup jobs run, the dashboard is green, and nobody has attempted a full restore since the system was installed. On the day it turns out that a folder was excluded, the job has been failing silently since March, the backup is on the same network with the same credentials and has been encrypted too, or a full restore takes 30 hours over the internet when everyone assumed three. A backup that has not been restored is a hope, not a plan. The distinction is set out in backup vs disaster recovery.

Documentation that assumes its author is present

Recovery documents tend to say "restore the server from backup" and stop. Where is the backup? Which credentials? What order? How do you know it worked? The author knows, so the gaps never mattered. When the author is unreachable, a competent engineer with the document in hand still cannot start. Good documentation lets a suitably skilled person who has never seen the environment follow it to a working result.

One person holds the keys

In most SMEs one individual has the domain admin password, the backup console login, the cloud tenant admin, the firewall and the supplier relationships. That is sensible for security and fatal for recovery if they are on a plane. The same pattern appears outside IT: one finance manager who knows the payroll run, one director who can authorise emergency spending. Recovery planning has to assume that some key people will be unavailable, because incidents do not schedule themselves around annual leave.

The fix is not to spread admin rights widely. It is a shared password manager with break-glass access, a named deputy for each role who has actually used the access, and delegation of spending and communication authority written down.

Failures in priorities, decisions and communication

No agreed order of restoration

Without priorities, teams restore whatever they understand first. The file server comes back before the domain controller it depends on, the application server before the database, the archive before the phones. Each wrong order costs hours. A plan should list systems in restoration order with their dependencies and their recovery time objective, so nobody has to decide under pressure. How to set those targets is in RTO vs RPO explained.

Decisions that wait

Do we invoke the plan? Do we pay for emergency hardware? What do we tell clients? Do we notify the ICO? If those decisions sit with one director who is unreachable, recovery stalls with engineers ready to proceed and nobody authorised to say go. Delegation, agreed in advance and written into the plan, removes the wait.

Communication that is improvised

Email and Teams are down, so the usual channels are gone. Staff hear rumours. Customers ring a phone nobody is answering. Clients learn about the incident from a competitor. The plan needs a way to reach staff that does not depend on the failed systems, a holding message for customers agreed in advance, and a named person who speaks for the business.

Failures of maintenance and testing

Plans written for a business that no longer exists

The plan was written when there were two servers and no cloud. Since then there has been a Microsoft 365 migration, a new job management system, a second site and three staff changes. The contact list is out of date, the procedures describe systems that were decommissioned, and the recovery targets were never revisited. A plan should be reviewed at every significant change, not on an annual anniversary.

Cloud assumed to be someone else's problem

"It is in the cloud" is often the whole plan. Microsoft 365 is resilient as a service, but a deleted mailbox, a ransomware-encrypted OneDrive synced to SharePoint or a compromised admin account are your problem, and Microsoft's retention is not designed to solve them. Cloud recovery needs an independent backup, account recovery procedures and a way for staff to work when the connection is down.

Tests that prove nothing

Some businesses do test, in a meeting room, by reading the plan aloud and agreeing it seems fine. That confirms the document exists. It does not confirm that the backup restores, that the deputy can log in, or that the recovery time is achievable. The tests that find problems are the uncomfortable ones: real restores, against the clock, with the usual person excluded.

A recovery test checklist

Run this at least annually and after every major change. Score each item pass or fail and fix the failures.

Before the test

  • Choose a realistic scenario (ransomware on the main server, loss of the office, Microsoft 365 admin account compromised).
  • Nominate the person who usually does recovery and exclude them for the exercise.
  • Set a target time from the recovery time objectives.

During the test

  • Locate the plan without using systems that would be down in the scenario.
  • Find the backup and confirm its most recent successful run and its retention.
  • Confirm an immutable or offline copy exists that the scenario could not have destroyed.
  • Restore a full server image to real or temporary hardware or cloud, using only the documentation.
  • Restore a Microsoft 365 mailbox and a SharePoint library from the independent backup.
  • Log in to every critical system as the deputy, using the shared password manager.
  • Restore systems in the documented order and confirm dependencies are correct.
  • Switch the phones to the divert plan and ring the main number.
  • Send the staff notification through the out-of-band channel.
  • Read the customer holding message and confirm who would send it.
  • Identify who authorises emergency spending and confirm they know.
  • Check the supplier and insurer contact list is current.
  • Time every step.

After the test

  • Compare each timing with the recovery time objective.
  • List every question nobody could answer and every step that needed the excluded person.
  • Update the documentation, access arrangements and contact lists.
  • Record the test date and results for insurers and auditors.
  • Schedule the next test.

The National Cyber Security Centre's guidance on offline backups in an online world is a useful reference for the backup elements, particularly the requirement for a copy an attacker cannot reach.

Building a plan that survives

Plans that work are built around process rather than people, kept short enough to be read during an incident, stored somewhere that does not depend on the systems that might fail, tested with the usual people absent, and revised whenever the business changes. Our IT disaster recovery plan template provides the structure, and the wider business continuity plan for SMEs covers the people and communication elements around it. Dig IT's backup and disaster recovery service includes the scheduled test restores that keep the plan honest.

What to do next

If your recovery plan has never been tested with its author absent, that is the test to run first. Book an IT health check and Dig IT will review your backups, documentation and access arrangements, and tell you which of these failure modes apply before an incident does.

Frequently asked questions

We have backups. Why is that not a recovery plan?
A backup is a copy of data. A recovery plan is how you get a working business back from it: which systems first, restored to what hardware or cloud, using which credentials, by whom, in how long, and how staff and customers are handled meanwhile. Backups that have never been restored, and procedures that exist only in one person's head, are the two commonest reasons a business with backups still loses days.
What is the single biggest reason recovery plans fail?
Dependence on one person. In most SMEs one individual holds the admin passwords, understands the server, knows the suppliers and is the only one who has ever run a restore. Plans are written assuming that person is available. Incidents happen at weekends, overnight and during holidays. Testing with that person deliberately excluded is the fastest way to discover how much of the plan lives in their head.
How often should a recovery plan be tested?
Test a file restore monthly, a full server restore quarterly and a whole scenario at least once a year. Also retest after any significant change: a new server, a cloud migration, a new line-of-business system, an office move or the departure of anyone with admin access. A plan that has not been tested since the environment changed should be assumed not to work.
What should a recovery test actually involve?
Pick a scenario, such as ransomware on the main server on a Friday evening, and work it against the clock. Restore from the actual backups to real or temporary hardware, using only the documentation, with a key person absent. Time each step, note every question nobody could answer, and compare the result with the recovery time objective. Then fix the gaps and record what was learned.
Does cloud software mean we no longer need a recovery plan?
No. Cloud removes the risk of a server dying and adds others: account compromise, deleted or encrypted files syncing everywhere, total dependence on internet access, and a Microsoft 365 tenant that is not backed up by default. The plan needs to cover how staff work when the connection or the service is down, how accounts are recovered, and where the independent backup of cloud data lives.
Why do cyber insurers ask about recovery testing?
Because tested backups and a rehearsed plan are the difference between a claim for a weekend's disruption and a claim for three weeks of lost trading. Many policies now ask whether backups are immutable, whether restores are tested and whether an incident response plan exists. Answering honestly, and being able to show evidence, affects both cover and premiums.

Next step

Not sure how exposed you are?

An IT health check reviews your security, backups, Microsoft 365 and network and gives you a prioritised list, whether or not you work with us afterwards.

WhatsApp us