Skip to content
Infrastructure

The backup you have never restored from

A backup job reporting success is not evidence of anything. Here is what to test, how to test it, and what we usually find.

SYSGOT Engineering· Infrastructure team11 February 20265 min read

Every engagement we run that touches backup begins the same way: before changing anything, we try to restore.

It is the most productive hour of the project, and the results are consistently uncomfortable.

What we find

Backups that complete but cannot restore. Database files captured without transaction logs, so the database cannot be brought to a consistent state. The job is green every night. The data is unusable.

Retention that silently shortened. The target filled up, the software started overwriting older sets to make room, and nobody changed the alerting. Thirty days of retention became four.

Scope that drifted. A new server, a new share, a new database added over the years and never added to the backup selection.

Only online copies. Replication to a second location, which is excellent for hardware failure and worthless against ransomware: the encryption replicates along with everything else.

Nobody knowing the recovery time. Not an estimate that is wrong; no estimate at all.

The two numbers you need

Write these down before discussing technology:

  • Recovery time objective. How long can this system be unavailable before it hurts the business?
  • Recovery point objective. How much data can you afford to lose — an hour, a day, a week?

An hour of tolerance and a day of tolerance describe completely different architectures, and only one of them is expensive. Most companies have never had the conversation, so they buy in the middle and get neither.

How to test properly

Restore to somewhere else. Not over the top of production. An isolated environment where you can verify without risk.

Restore the whole thing. A single file coming back proves very little. The question is whether the application starts, the database is consistent, and the users can log in.

Time it. Wall clock, from decision to working. Include finding the media, the credentials and the documentation, because in a real event you will be doing that too, possibly at 3am.

Write down what went wrong. The first rehearsal always finds something. A missing licence key, an undocumented dependency, a service account nobody has the password for. That list is the actual output of the exercise.

Do it again, on a schedule. Quarterly for anything critical. It takes an afternoon and it is the only thing that keeps the number honest.

3-2-1, and why the "1" matters most

Three copies, on two different media, with one off-site. The version worth insisting on now adds: one offline or immutable.

Ransomware operators specifically hunt for backup infrastructure. If your backup target is reachable with credentials that exist on a compromised machine, it is part of the attack surface, not protection from it.

Immutable storage, where a written object cannot be changed or deleted until a retention period runs out, is the most valuable backup feature of the last decade. Most platforms now support it. Turn it on.

The finding that justifies the exercise

On one engagement we found four years of nightly backups that could not be restored from at all. The client had done everything they thought was required: bought the software, scheduled the job, monitored the result.

Finding that in a test cost an afternoon. Finding it in an incident would have cost the business.

Related work

Recognise the problem in your own estate?

Happy to talk through how it maps to your situation.

  • A reply within one working dayFrom an engineer.
  • We look before we quoteA call, and a site visit if needed.
  • The recommendation is yoursYours to take elsewhere.
  • Or call +20 109 777 8090