Skip to content
Healthcare

A backup that had never been restored

Discovering that four years of nightly backups could not be recovered from, then building and rehearsing something that could.

Client
Diagnostic clinic groupNamed once permission is confirmed
Year
2024
Duration
5 weeks
Stack
AzureVeeamFortinetZKTeco
  • Unrecoverable → 90 min

    Measured full recovery time

  • 3-2-1

    Backup copies, with off-site immutability

  • 0

    Shared logins remaining

  • Drilled

    Downtime procedure, rehearsed with clinical staff

The situation

The group ran a clinical system across four sites with a nightly backup job that had reported success every night for four years. Nobody had ever attempted a restore. Staff shared logins at reception workstations because individual logins took too long, and there was no written procedure for what to do if the system went down.

What we did

  1. 1Attempted a full restore into an isolated environment as the first task of the engagement
  2. 2Found the backup contained the database files but not the transaction logs, making it unrecoverable to any consistent point
  3. 3Rebuilt backup on a 3-2-1 basis with off-site immutable copies, then rehearsed a full recovery and published the measured recovery time
  4. 4Replaced shared reception logins with badge authentication and session roaming, removing the friction that caused the workaround
  5. 5Wrote and drilled a downtime procedure, with printed forms located where staff actually work

Finding out in a test rather than in an emergency is the best money we have spent on IT.

— Group Operations Director (attribution pending approval)

The restore test came first

Before changing anything, we tried to restore from the backups.

The nightly job had reported success for four years. It had been copying the database files without the transaction logs, so not one of those backups could be restored to a working state. They existed, and they were useless.

This is not rare. It is the single most common serious finding in our assessments, and it is invisible until someone tries.

Fixing the shared logins by removing the reason

Reception staff shared a login. Not out of carelessness: logging in individually took too long, dozens of times a shift, with a patient waiting.

We did not run an awareness session. We put badge readers on the workstations with session roaming, so a clinician's desktop follows them and authentication takes under two seconds.

Shared logins disappeared within a week, because the workaround stopped being worth the effort. Policy without workflow change would have achieved nothing.

A rehearsed downtime procedure

The recovery time we quote, 90 minutes, is one we measured in a rehearsal, not one we worked out from a datasheet.

The rehearsal also produced the parts nobody thinks about: which forms staff need on paper, where they are kept, who calls whom, and in what order systems come back. Then we drilled it with clinical staff, which found two wrong assumptions in our own document.

What is still open

The group runs one clinical application that only supports authentication we would rather not rely on. We have documented the risk, compensated with network controls, and it is on the roadmap with the vendor. We are noting it here so nobody reads this as a finished estate.

Comparable project

Something like this, for your business?

We will say whether the same approach applies.

  • A reply within one working dayFrom an engineer.
  • We look before we quoteA call, and a site visit if needed.
  • The recommendation is yoursYours to take elsewhere.
  • Or call +20 109 777 8090