Managed backup with immutable offsite copies, Microsoft 365 protection and periodic tested restores — monitored daily rather than assumed to be working.
The three things that go wrong
Backup is one of the few areas of IT where the failure is invisible until the exact moment it is catastrophic. Three patterns account for nearly everything we find.
The job has been failing for weeks. Something changed — a credential rotated, a volume was added, a software update altered a path — and the job started failing. Nobody was checking, because it had worked for two years. This is the single most common finding on a new client audit.
The backup completed but is not recoverable. Job status green, data unusable. Corrupted, incomplete, or missing the one database that mattered because it was added after the job was configured. Only a restore test reveals this.
The ransomware deleted the backups first. Backups on a network share reachable with domain credentials, or on a NAS joined to the domain. An attacker who has administrative access has access to those too, and destroying them is a deliberate step in the playbook, because it is what makes paying the only option.
What actually protects you
Separation. The backup must not be reachable using credentials from the environment it protects. If your domain admin account can delete the backups, so can whoever obtains that account.
Immutability. Copies written so they cannot be modified or deleted for a retention window, regardless of privilege. This is the specific control that defeats the delete-the-backups step.
Multiple copies, at least one offsite. Local for restore speed, offsite for survivability against fire, flood and theft as well as ransomware.
Microsoft 365 coverage. Treated as its own workload with its own retention, because it is not covered by anything else you have.
Recovery objectives are a business decision
Two numbers determine the design, and neither is technical.
How much data can you afford to lose? If backups run nightly and you are hit at 4pm, you lose a day. For some businesses that is an annoyance; for others it is unrecoverable. That tolerance sets the backup interval.
How long can you afford to be down? Restoring a large file server from an offsite copy over a business internet connection takes real time. If the answer is “hours, not days”, the design needs a local copy and a rehearsed sequence, and it costs more.
Deciding these deliberately is the difference between a backup product and a recovery capability. Discovering them during an incident is how businesses find out their backup was technically fine and commercially useless.
Testing is the whole point
A backup you have never restored is a hypothesis.
We run periodic test restores against real backup data and report the outcome in writing. When a test fails — and on inherited environments they do — that is the finding the service exists to produce. Far better to discover it on a Tuesday with time to fix it than at 6am during an incident with the business stopped.