Module 5 · Recovery Testing — Tabletop, Functional, Full Failover

Manish Garg
Manish Garg Associate of (ISC)² · RingSafe
May 14, 2026
4 min read
Read as
100% Free

No signup. No paywall. No catch. One of our 10 most-requested practitioner modules — published in full so anyone can learn for free. We earn through consulting, not by gating knowledge.

See all 10 free modules →

Why this module exists. A BCDR plan that has never been tested is a hypothesis. Recovery testing — tabletop, functional, full failover — is what turns the plan into a verified capability. This module covers the four testing tiers, what each catches, and the cadence that maintains real recovery readiness.

Why this module exists. Indian enterprises that test their DR architecture annually consistently outperform those that test it on the day of incident. The difference is operational muscle memory. This module covers building it.

The four testing tiers

Tier Disruption What it tests
Tabletop None Plan logic, role clarity, decision-making
Walk-through None Detailed step verification without execution
Functional / Simulation Partial — isolated environment Technical recovery procedures, RTO verification
Full failover Significant — production switched to DR End-to-end capability, real-world dependencies

Tabletop — the cheapest highest-value test

2-3 hour facilitated session with all responsible parties. Scenario presented; participants walk through what they would do. Facilitator captures decisions, gaps, conflicts.

Cadence: quarterly for any organisation operating critical IT. Scenarios rotate — natural disaster, cyber-attack-driven recovery, vendor failure, regulatory-driven shutdown.

Output: list of identified gaps. Each gap becomes a ticket; closed before the next tabletop.

Tabletop scenarios that work

  • “Mumbai data centre flood at 14:00 IST during peak transaction processing.”
  • “Domain Admin compromise discovered Friday evening; ransomware encryption observed on Monday morning.”
  • “Primary cloud-region (AWS Mumbai) reports multi-AZ outage with 8-hour estimated recovery.”
  • “Critical SaaS vendor announces 48-hour outage for emergency security patching.”
  • “Key person (CISO or Head of Infra) unreachable for 72 hours during active incident.”

The scenarios should map to real risks from the BIA, with realistic timing constraints.

Functional / simulation testing

Restore specific workloads in an isolated environment without affecting production. Measure: did the restore succeed? How long did it take vs the RTO target? Did the data integrity verification pass? Cadence: monthly for critical workloads, quarterly for others.

Methodology:

  1. Pick a workload (one application + its dependencies).
  2. Stand up an isolated environment.
  3. Execute restore procedure from documentation.
  4. Verify data integrity (post-restore validation queries).
  5. Time-stamp every step; measure against RTO target.
  6. Document findings: what worked, what did not, gaps in the procedure.

Full failover testing

The real thing: shift production to the DR site, run there for a defined window, then shift back. Disruptive; high signal. Cadence: annually for regulated entities (RBI mandates this), biennial for less-regulated.

Pre-conditions for a successful full-failover test:

  • Tabletop and functional tests pass in the preceding months.
  • Maintenance window communicated to customers (regulators allow planned-DR-test windows).
  • Rollback plan documented if the failover encounters issues.
  • Senior leadership awareness — board notified if material customer-facing impact.

What full-failover catches

  • Hardcoded references to primary-site URLs, IP addresses, hostnames.
  • Stale DR copies — application config differs between primary and DR.
  • Network-routing assumptions that work only in primary topology.
  • Third-party integrations whose endpoints are not failover-aware.
  • Vendor SLAs that only commit to primary site.
  • Capacity gaps — DR site sized for 80% of primary load.

Chaos engineering — the continuous version

For cloud-native architectures, chaos engineering exercises failure modes continuously. Tools (Chaos Monkey, Litmus, Gremlin) inject failures into production. This is the modern equivalent of constant testing — not separate tests but a continuous validation of resilience.

Indian regulated entities have been slow to adopt chaos engineering due to disruption concerns. The pragmatic compromise: chaos in pre-production environments continuously, full-failover annually, tabletop quarterly.

Documentation of test results

Every test produces a report. Standard fields:

  • Test name, type, date, duration.
  • Scenario tested.
  • Participants and roles.
  • RTO target vs actual.
  • RPO target vs actual.
  • Functional verification results.
  • Gaps identified.
  • Remediation tickets created.
  • Sign-off — test passed / partial / failed.

Reports go to: Risk Committee (quarterly summary), regulator if applicable, internal audit.

Indian regulatory mandates

  • RBI Cyber Security Framework: annual DR drill mandatory.
  • SEBI CSCRF: business-continuity / DR drills mandatory; reporting to SEBI.
  • IRDAI: DR drill annually; coverage extends to outsourced functions.
  • NCIIPC for designated CII: DR drill annually with NCIIPC-approved scope.

Common failure modes

  • Tabletop without business participants — IT-only exercise misses operational reality.
  • Functional test in an environment that does not match production.
  • Full failover that excludes critical SaaS dependencies — not a realistic test.
  • Findings logged but not closed — same gaps in successive tests.
  • Annual drill at the same time of year — attackers and disasters do not honour calendars.

Key takeaways

  • Four tiers: tabletop, walk-through, functional, full failover.
  • Cadence: tabletop quarterly, functional monthly for critical, full failover annually.
  • Tabletops with business + IT + vendor reps; scenarios from BIA.
  • Full failover catches integration assumptions invisible to other tests.
  • RBI / SEBI / IRDAI mandate annual DR drills; documentation drives regulator response.
Worried about your exposure?

Get a free attack-surface review

We check what an attacker would see about your business — leaked credentials, exposed services, dark-web mentions. 30 minutes, no obligation.

Book exposure review Replies in 4 working hrs · India-only · Senior consultants