The old advice was short. Keep backups. Get hit, restore, move on.
It isn't wrong. It's just not enough anymore. These days an attack doesn't sit politely in one folder. It moves through servers, user accounts and cloud services, and often the attackers copy your data out before they encrypt anything. A company with a perfectly sensible backup routine can still spend weeks digging itself out.
I'm not knocking backups. A clean copy might mean you never have to deal with the criminals at all. But having a backup and being able to recover from it are two different things, and most people find that out the hard way.
The gap shows up in familiar ways. The backup got encrypted along with everything else. It's corrupted. It's a month old when you needed yesterday's. Nobody has ever run a restore, and the one person who knew how left last spring. Or the data comes back fine and the login system it depends on doesn't.
Then there's the nasty one. The attacker might still be in your network. Rebuild everything while their stolen admin account is alive and you've just handed them fresh systems to break.
Maersk is the example I keep coming back to. When NotPetya hit in 2017, the shipping giant lost most of its Windows environment. As widely reported, what saved them was a single domain controller in Ghana that happened to be offline because of a power outage. Luck, basically. You don't want your recovery plan to depend on a blackout.
A green checkmark proves less than you think
Plenty of teams can show you a dashboard full of successful backup jobs. Nice. All it proves is that the software ran. It says nothing about whether you could bring payroll back at 3 a.m. with half the network dead.
So try it. There's no other way to find out.
Attackers go for the backups first
Ransomware crews aren't stupid. They know the backups are what stand between them and their payout, so once they have privileged access, that's where they head. They delete copies, encrypt repositories, switch off jobs, purge snapshots, or quietly shorten retention so the older restore points just disappear. Nobody notices until they need them.
The classic mistake is running production and backups under the same powerful admin account. Steal one login and you own both. The backups still exist, on paper. Can you trust them? Different question.
That's why I think of backup security as partly an identity problem. Locked-down storage doesn't help if someone can sign in as an admin and delete everything on it.
Getting the data back isn't getting the service back
Say the database restores fine, but the application that uses it can't be rebuilt. You've got your data and no working business.
Monitoring that only asks "did the job finish?" won't catch this. It won't flag the missing encryption key, the software versions that no longer match, the server waiting on another server that's still down, or the documentation nobody wrote. Test your restores on a schedule. Not once, because an auditor asked.
RPO and RTO, plainly
Two acronyms come up in every recovery conversation.
RPO (Recovery Point Objective) is how much data you can stand to lose, measured in time. A four-hour RPO means you should be able to restore to within about four hours of the incident. Back up a database once a day, get hit just before the next run, and you've lost close to a full day of transactions. A small office with mostly static files might live with that. A bank or a factory can't. There's no right number in general, only the one that fits your business, your regulators and what lost data would cost you.
RTO (Recovery Time Objective) is how long you can be down. Four-hour RPO, eight-hour RTO: you accept losing four hours of data and want service back within eight. Here's the catch. An RTO nobody has tested is a guess. In a real incident, teams discover they don't know what to restore first, the recovery passwords live on a dead server, and DNS and authentication need rebuilding before anything else works. Eight hours turns into days. Get your number from an actual recovery run.
Getting the attacker out
Picture it. Servers rebuilt from clean backups, everything green, everyone relieved. Then the intruder logs in with the same stolen credentials. Back to square one.
Recovery has to include containment and cleanup, not just restoration. Somebody has to work out how they got in and what they left behind, which means combing through privileged and VPN accounts, live sessions, remote access tools, authentication logs, endpoint alerts, scheduled tasks, service accounts, cloud identities, API keys and firewall logs. It's slow, boring work. It's also what stops you getting hit a second time.
And don't aim to recreate the old environment exactly. That's the one that got breached. You want a known-clean state with the doors they used closed as tightly as you can manage.
Immutable and offline copies
An immutable backup can't be altered or deleted for a set period, even by an admin on the production side. You can get there with object storage that supports immutability, WORM storage, protected snapshots or cloud retention locks. But "it's in the cloud" doesn't mean "it's immutable." Find out who can delete data or change retention settings, ask what happens if one of those accounts gets stolen, and test that the protection really works.
Offline copies feel old-fashioned, and I'd still keep one. A backup that isn't connected to the network can't be reached by malware on that network. Removable media works, and so does storage that's only attached while the backup runs. You pay for it in manual effort, slower restores, and somewhere physical to store the media and check on it. Worth it, in my view, for the one copy that survives when everything else doesn't.
3-2-1
You've probably heard it: three copies, two kinds of storage, one offsite. Live data, a local backup, and a third copy elsewhere or with a cloud provider. Make one immutable or offline and it gets stronger.
What matters is that the copies don't all share the same infrastructure, credentials or network path. If one admin account reaches every one of them, you effectively have a single copy.
Who controls the backups
If one account can reach production, the cloud and the backup system, then stealing it lets someone wreck your recovery along with everything else. The usual measures apply: multi-factor authentication, separate admin accounts, least privilege, privileged access management, conditional access, segmentation, logging and regular access reviews. Keep backup administration apart from everyday accounts wherever you can. No single login should be able to destroy both the systems and the thing meant to save them.
Plan around services, not files
It's natural to think in files and databases. The applications on top get forgotten. A perfectly restored database is useless if the app server is encrypted, authentication is offline, or a certificate has gone missing.
For each critical service, write down what it needs: servers, databases, authentication, DNS, certificates, keys, service accounts, outside integrations. Having that map before an incident can save you hours when it counts.
Check before you reconnect
Restored doesn't mean safe. If you can, bring systems up in an isolated environment first. Verify the backups, scan, patch, reset credentials, review configurations, test the applications, confirm connectivity. Then reconnect.
Practice
Plans look great on paper and fall apart on contact. So run a drill. Pretend the file servers and identity systems are gone and have the team restore critical services with only what they'd really have.
Drills show you what documents hide. Missing recovery credentials. A runbook two years out of date. A backup job that's been failing for weeks while nobody looked. An app that depends on a system nobody remembered. Not enough storage, crawling restores, lapsed licenses, key people unreachable, an emergency contact list full of people who've left. Far better to learn all that on a quiet Tuesday than mid-attack.
Keep the plan where you can reach it
Write down where the backups are, who owns them, which systems matter most, what order to restore in and what depends on what. Add RPO and RTO targets, emergency contacts, vendor details, procedures, and the access needed to carry them out.
Then protect it. A recovery plan that lives only on an encrypted file server is no plan at all. Keep an emergency copy somewhere that survives a big outage.
Decide the order now
A big incident can knock out hundreds of systems, and you can't restore them all at once. A common order: identity and authentication first, then critical network services, databases, key business applications, file services, collaboration tools, and the low-priority stuff last. Yours will differ. A hospital isn't a manufacturer, and a bank isn't a software company.
Agree on it while things are calm. Arguing about what the business can't live without is miserable when everything is down.
One layer among many
Backups sit inside a wider defense; they don't replace one. Endpoint detection, segmentation, MFA, patching, email security, vulnerability management, monitoring, staff training, application control and incident response planning all reduce the odds and the damage. Prevention makes attacks harder, detection catches them, containment limits the spread, and if production goes down anyway, tested backups are your way back. People call this defense in depth.
Quick checklist
Protect the backups: several copies, one offsite, immutable or offline where it makes sense, backup admin separate from normal accounts, MFA, alerts on deletions and config changes.
Test them: regular restores of files and whole systems, real restore times measured, integrity checked, results written down.
Guard the identities: protected privileged accounts, MFA, regular admin reviews, backup credentials kept apart from production, and a plan to reset anything possibly compromised.
Be ready to respond: a ransomware plan, defined roles, reachable emergency contacts, a way to isolate systems, and procedures for preserving evidence.
Keep the business running: critical services identified, realistic RPO and RTO, documented dependencies, agreed priorities, recovery drills on the calendar.
The bottom line
Backups are still among your best tools against ransomware. On their own they aren't a recovery strategy. They can be deleted, corrupted, out of reach or too old, and even a perfect restore fails if the attacker is still inside.
What works is the combination: protected backups, tested restores, locked-down privileged accounts, incident response, mapped dependencies and clear priorities.
So don't stop at "Do we have backups?"
Ask this instead:
If our production environment were compromised today, could we restore our critical operations from clean, trusted recovery points?
If you've never proven the answer with a real test, then what you have is backups. Whether they'd get you through an attack is still unknown.