Introduction
Disaster recovery evidence for SOC 2 is the dated proof that your organization planned how to restore critical systems and data, set recovery targets, ran backups, tested restores and recovery procedures, and fixed what the tests exposed, across the audit period.
Many teams have a recovery plan and nightly backups. They still struggle when auditors ask for the plan version that was current in March, proof that a restore actually worked, the recovery time objective (RTO) and recovery point objective (RPO) for each critical system, and the follow-up tickets from the last tabletop exercise. A green backup dashboard is not the same as proof that you can recover.
In plain language: this guide shows how to operate and evidence disaster recovery so SOC 2 sampling is routine. It is an evidence playbook, not a second definition page. For the definitions, see What Is Disaster Recovery?, What Is a Business Continuity Plan?, and What Is Availability?. Related guides include Incident Response Evidence for SOC 2, What Is SOC 2 Evidence?, and SOC 2 Trust Services Criteria Explained.
What Disaster Recovery Evidence Means
Disaster recovery (DR) evidence is the system-of-record trail that shows recovery was planned, resourced, tested, and improved, with dates and owners attached.
It usually includes:
- A current, versioned DR plan with named owners and an approval date.
- A list of critical systems and dependencies, often from a business impact analysis.
- RTO and RPO targets for each critical system, with an approval record.
- Backup configuration: scope, frequency, retention, encryption, and storage location.
- Backup job logs or reports that cover the period, including failures and how they were handled.
- Restore test records: what was restored, when, by whom, how long it took, and whether data was complete.
- Tabletop or failover exercise records with participants, scenario, results, and follow-up actions.
- Remediation tickets for gaps that tests or real events exposed.
- Plan review records after major architecture, vendor, or staff changes.
- Period exports so auditors can sample across the full window.
DR evidence is not the same as uptime monitoring. Monitoring tells you something is wrong. DR evidence shows you can bring systems and data back within agreed targets.
Why DR Evidence Matters for SOC 2
Customers who depend on your service want to know you can recover from a major outage, ransomware event, or region failure. SOC 2 gives them a way to check.
It matters because:
- Backups that have never been restored are an assumption, not a control.
- RTO and RPO targets mean little without a test that measured them.
- A plan that does not match current architecture fails at the worst time.
- Type 2 reports test whether recovery controls operated over the period, not just on one good day.
- Gaps found in a test and never fixed look like control drift.
For Type 1, a current plan, defined targets, and backup configuration may show the design exists. For Type 2, auditors look for operation: backup logs across the period, restore tests on the stated schedule, and follow-up on findings. See SOC 2 Type 1 vs Type 2.
What SOC 2 Requires vs Common Practice
Keep the criteria and the habits apart.
What the criteria address. The AICPA 2017 Trust Services Criteria with revised points of focus (2022) include:
- A1.2 (Availability): environmental protections, software, data backup processes, and recovery infrastructure are authorized, designed, developed or acquired, implemented, operated, approved, maintained, and monitored to meet objectives.
- A1.3 (Availability): recovery plan procedures that support system recovery are tested to meet objectives.
- CC7.5 (Security, common criteria): the entity identifies, develops, and implements activities to recover from identified security incidents.
- CC9.1 (Security, common criteria): the entity identifies, selects, and develops risk mitigation activities for risks arising from potential business disruptions.
The A1 criteria apply only when Availability is in scope. CC7.5 and CC9.1 are common criteria, so recovery and disruption themes can come up even in a Security-only report.
What is common practice, not a fixed SOC 2 rule:
- Annual DR plan review and approval
- Quarterly or semiannual restore tests for critical data stores
- An annual tabletop or failover exercise
- Specific RTO and RPO values
- Specific backup tools, regions, or retention lengths
The criteria do not set test frequencies or target values. You set them in your plan, and your auditor tests whether you followed what you wrote. NIST guidance is a useful reference for structure: NIST SP 800-34 Rev. 1 covers contingency planning, business impact analysis, and plan testing, and the CP (Contingency Planning) family in NIST SP 800-53 Rev. 5 includes controls such as CP-2 (Contingency Plan), CP-4 (Contingency Plan Testing), CP-9 (System Backup), and CP-10 (System Recovery and Reconstitution). These are reference frameworks, not SOC 2 requirements.
DR Plan vs BCP vs Backup vs Restore Test vs Incident Recovery
Keep related artifacts distinct so owners and auditors share meaning.
| Artifact | Primary job | Typical evidence | Common confusion |
|---|---|---|---|
| DR plan | Define how technology and data are restored | Versioned plan, owners, approval, RTO and RPO | Treating a short wiki page as a tested plan |
| Business continuity plan | Keep business operations running during disruption | BCP document, roles, communication trees | Assuming a BCP covers technical restore steps |
| Backup job | Copy data on a schedule | Job config, success and failure logs, retention settings | Equating "job succeeded" with "data is recoverable" |
| Restore test | Prove data and systems can actually be recovered | Test record with scope, time taken, integrity check, result | Skipping tests because backups look healthy |
| Tabletop or failover exercise | Practice decisions and procedures under a scenario | Scenario, participants, timeline, findings, actions | Holding a meeting with no notes or follow-ups |
| Incident recovery | Real-world recovery from an actual event | Incident ticket, recovery steps, PIR | Counting a real outage as a test with no review |
Real recoveries can count as useful evidence when you document them well. See Incident Response Evidence for SOC 2 for how to record them.
Setting and Recording RTO and RPO
RTO is how long a system can stay down before the impact becomes unacceptable. RPO is the point in time to which data must be recovered, which sets how much data loss is tolerable. NIST defines both in its glossary: recovery time objective and recovery point objective.
For evidence, record:
- The target RTO and RPO for each critical system, not one number for the whole company
- Who approved the targets, and when
- The basis for each target, such as customer commitments or a business impact analysis
- How backup frequency supports the RPO (for example, hourly snapshots for a one-hour RPO)
- Measured results from the latest restore or failover test, next to the target
A labeled hypothetical: a team sets a four-hour RTO for its primary database. Its restore test takes six hours because nobody had access to the restore role. The honest record shows the miss, the access fix ticket, and a retest. That is stronger evidence than a test report that quietly drops the timing.
Fields and Artifacts That Survive Sampling
Prefer one place where DR records live, such as a ticket project, a GRC tool, or a controlled document repository.
| Field / artifact | Why auditors care |
|---|---|
| Plan version and approval date | Shows which plan applied at each point in the period |
| Critical system list with owners | Ties recovery scope to the system description |
| RTO and RPO per system | Gives a target to test against |
| Backup configuration export | Shows scope, frequency, retention, and encryption |
| Backup job history for the period | Shows backups ran, and how failures were handled |
| Restore test ID, date, scope, and duration | Proves recovery works and measures it |
| Data integrity check result | Shows restored data was complete and usable |
| Exercise record with participants | Shows people know their roles |
| Findings and remediation ticket IDs | Shows the learning loop closed |
| Plan review log after major changes | Shows the plan tracks the real system |
Screenshots of a backup console with no date, no scope, and no restore result are weak. Exports and tickets with stable IDs are stronger.
Operating Model: Run DR Evidence Across the Audit Period
1. Name owners and the system of record
Assign an owner for the DR plan and an owner for each critical system's recovery. Point the plan at where test records and tickets live.
2. Identify critical systems and dependencies
List the systems customers rely on, plus their dependencies: identity, DNS, secrets, third-party services, and deploy pipelines. Many failed tests trace back to a dependency nobody listed.
3. Set and approve RTO and RPO
Record targets per system with an approval date. Align them with what you tell customers.
4. Configure and monitor backups
Document scope, frequency, retention, encryption, and storage location. Keep at least one copy outside the failure domain it protects, and protect restore credentials. Alert on failed jobs and record how each failure was handled.
5. Test restores on a schedule
Restore real data to a test environment. Record start and end time, what was restored, integrity checks, and the result. NIST SP 800-53 control enhancement CP-9(1) describes testing backups for reliability and integrity, which is a useful reference for what a restore test should check.
6. Run exercises
Hold tabletop or failover exercises against realistic scenarios. NIST SP 800-84 describes test, training, and exercise programs, and NIST's glossary defines a tabletop exercise as a discussion-based exercise. Record participants, the scenario, decisions, and gaps.
7. Fix findings
Open remediation tickets for every gap. If you accept a gap for now, record it as a governed exception with an owner and expiry.
8. Review the plan after change
Update the plan after major architecture, vendor, region, or team changes. Log the review even when nothing changes.
Evidence Examples by Scenario
| Scenario | Stronger evidence | Weaker evidence |
|---|---|---|
| Quarterly database restore test | Test ticket with timing, row counts or checksums, owner, result | "Restore worked" message in chat |
| Annual tabletop exercise | Scenario doc, attendee list, decision log, action tickets | Calendar invite only |
| Region failover drill | Runbook version, start and end times, measured RTO, issues found | Architecture diagram claiming multi-region |
| Backup job failure | Alert, ticket, root cause, rerun confirmation | Silent failure discovered at audit time |
| Real outage recovery | Incident ticket, recovery steps, PIR, plan update | Status page post with no internal record |
| Cloud provider change mid-period | Plan review log, updated dependencies, new test | Old plan still naming the retired provider |
| Missed RTO in a test | Recorded miss, fix ticket, retest result | Test report that omits timing |
| Vendor-hosted critical system | Vendor SOC report review, contract recovery terms, your own export tests | Assumption that the vendor handles it all |
How Auditors May Sample Availability Controls
Sampling depends on control frequency and your auditor's approach. Common patterns include:
- Selecting backup job dates across the period and asking for logs
- Asking for every restore test the plan says should have happened
- Reviewing the DR plan version and approval for the period
- Checking that exercise findings were remediated or governed
- Comparing the critical system list with the system description
Write your plan so the cadence is clear. If the plan says quarterly restore tests, expect a request for each quarter in the period.
Common Mistakes and Why Teams Struggle
- Backups run, but nobody has restored from them in the period
- One company-wide RTO that no system has ever been tested against
- Restore credentials stored inside the environment that would be lost
- DR plan last updated before a major migration
- Tabletop exercises with no notes, owners, or follow-ups
- Test results that drop failures or missed targets
- Vendor-hosted systems left out of recovery planning
- Treating the BCP and the DR plan as the same document
- Findings that never become tickets
These patterns create diligence friction and real recovery risk.
Linking DR Evidence to Neighboring Controls
Disaster recovery is stronger when it connects to the rest of the program.
- Incident response. Real recoveries should have tickets and PIRs (incident response evidence).
- Change management. Failover changes and plan updates should leave change records (change management evidence for SOC 2).
- Vendor oversight. Critical vendors need recovery terms and reviews (Vendor Management for SOC 2).
- Risk assessment. Business disruption risks should appear in your risk work (risk assessment evidence for SOC 2).
- Access control. Restore roles and backup storage need restricted access (access control).
- Encryption. Backups that hold sensitive data should follow your encryption standard.
Internal Pre-Audit Sampling Checklist
Before fieldwork, run a short internal check:
- Confirm the DR plan is current, approved, and names owners.
- Confirm each critical system has an approved RTO and RPO.
- Pull backup job history for the period and check failures were handled.
- Gather every restore test the plan required, with timing and integrity results.
- Gather exercise records and confirm findings have tickets.
- Check that the plan was reviewed after major changes in the period.
- Compare the critical system list with the system description.
- Fix gaps, then re-export so the packet matches reality.
How AuditFlo Helps
AuditFlo (auditflo.co) helps teams retain continuous evidence collection and audit-period history for control and readiness artifacts when those records are stored or linked through connected systems and workflows your team already uses, such as AWS, GitHub, Jira, Okta, and Google Workspace.
The focus is organizing dated proof so backup, restore test, exercise, and remediation records can be shown with history instead of last-minute screenshots. AuditFlo is positioned here as evidence and readiness support for the recovery program your organization defines. It does not run your backups, does not perform restores or failovers, does not issue SOC 2 reports, does not certify compliance, and does not replace auditors.
To see the workflow for your stack, request a demo.
Final Thoughts
Disaster recovery evidence is an operating habit: a current plan, approved RTO and RPO targets, backups with handled failures, restore tests that measure real recovery, exercises with follow-ups, and plan reviews after change. Keep it distinct from uptime dashboards and from the business continuity plan, connect it to incident response and vendor oversight, and retain dated records across the period.
If you are building the broader readiness path, start from What Is Disaster Recovery?, What Is Availability?, What Is SOC 2 Evidence?, and how to prepare for a SOC 2 audit.
FAQ
Is disaster recovery required for SOC 2?
When Availability is in scope, the A1 criteria address backup, recovery infrastructure, and recovery plan testing. Even in a Security-only report, common criteria such as CC7.5 and CC9.1 address recovery from incidents and business disruption risk. Exact auditor focus depends on your system description, risks, and control design. This page describes common practice, not a guarantee of any report opinion.
How often should we test restores?
SOC 2 does not set a frequency. Many teams test critical data restores quarterly or semiannually and run a broader exercise once a year. Write your cadence into the plan and follow it, because auditors test against what you wrote.
What is the difference between DR evidence and a business continuity plan?
A business continuity plan covers how the business keeps operating during disruption. DR evidence shows how technology and data are restored, measured against RTO and RPO.
Does a successful backup job count as a recovery test?
No. A backup job shows data was copied. A restore test shows data can be recovered, intact, within your targets. Keep both.
Can a real outage count as a DR test?
It can support your evidence if you record the recovery steps, timing, and post-incident review. Many teams still run planned tests so coverage does not depend on luck.
What if a test misses our RTO?
Record the miss honestly, open a fix ticket, and retest. A documented miss with follow-through is stronger evidence than a report that hides the result.
Does AuditFlo replace our backup or DR tooling?
No. AuditFlo helps organize evidence and readiness artifacts around the recovery program you define. It does not replace backup tools, failover systems, or your auditor.