One Of India's Largest Public Sector Banks.
The Bank Runs Core Banking, Payments And Digital Channels For Millions Of Account Holders Across Thousands Of Branches. Every Hour Of Downtime Interrupts Transactions Nationwide, And Regulatory Guidance Requires Demonstrable Recovery Capability At A Separate Geographic Site.
Its Primary Data Centre Had Grown Through Successive Additions Rather Than Design. Power And Cooling Capacity Were Fully Consumed, Cabling Had Become Untraceable, And The Secondary Site Existed On Paper But Had Never Been Tested Under Real Failover Conditions.
PROBLEM STATEMENT
Despite Continuous Investment, The Bank Struggled With:
- 01
Exhausted Capacity
Power And Cooling Headroom Left No Room For New Workloads Or Refresh.
- 02
Untested Recovery
The Secondary Site Had Never Completed A Full Production Failover Exercise.
- 03
Reactive Operations
Faults Were Discovered By Users Rather Than By Any Monitoring Layer.
PROPOSED SOLUTION
The Engagement Rebuilt Data Centre Operations Around:
-
Redesigned Power And Cooling Distribution With Measured Headroom For Planned Refresh Cycles.
-
Rebuilt Structured Cabling And Rack Layout With Full Documentation And Port Level Traceability.
-
Established A Geographically Separate DR Site With Replicated Storage And Tested Failover Runbooks.
-
Deployed A 24x7 NOC Monitoring Power, Cooling, Network, Compute And Storage Continuously.
-
Introduced Scheduled Failover Drills Validating Recovery Objectives Against Documented Regulatory Commitments.
-
Implemented Capacity Forecasting Linking Workload Growth To Power, Cooling And Rack Planning.
Engineering For Failure Before It Happens
Every Layer Now Has A Documented Failure Mode And A Tested Response. Monitoring Detects Degradation Before Service Is Affected, And Failover To The Secondary Site Is Rehearsed On A Fixed Schedule Rather Than Assumed.
Redundant Power And Cooling Paths Remove Single Points Of Failure Across Every Critical Rack Row.
Documented Cabling And Rack Records Make Fault Isolation A Minutes-Long Task Rather Than An Investigation.
Scheduled Failover Drills Prove Recovery Objectives Instead Of Leaving Them As Untested Design Assumptions.
Continuous Monitoring Escalates Threshold Breaches To The NOC Before Users Notice Any Degradation.
RESULT
Proven Recovery, Measured Capacity
The Bank Now Demonstrates Tested Failover To Regulators Rather Than Describing Intent, And Capacity Planning Runs On Measured Data. Incidents Are Detected And Contained Before Branch And Digital Channels Are Affected.
Infrastructure Uptime
Minute Tested Failover
Faster Incident Detection
Reclaimed Rack Capacity
LESSONS LEARNED
The Engagement Highlighted Three Lasting Takeaways:
-
Untested Is Unproven
A DR Site Without A Completed Drill Is Documentation, Not Capability.
-
Document The Physical Layer
Most Incident Delay Came From Not Knowing What Connected Where.
-
Plan Capacity, Not Purchases
Power And Cooling Constrain Growth Long Before Rack Space Does.
TECHNOLOGIES - TOOLS USED
Operations Run On A Monitoring Stack Covering Environmental, Network, Compute And Storage Layers With Alerting Into A Single NOC Console. Configuration Records, Replication Tooling And Runbook Automation Sit Alongside It, So Detection, Diagnosis And Failover Follow One Documented Path.
- Infrastructure Monitoring
- DCIM
- Environmental Sensors
- Virtualisation Platform
- Storage Replication
- Backup & Recovery
- Network Management
- Runbook Automation
- Ticketing System
- Capacity Planning
- NOC Dashboards
CONCLUSION
Resilience Is Not A Purchase, It Is A Practice. Rebuilding The Physical Layer With Documented Headroom, Standing Up A Genuinely Separate Recovery Site And Rehearsing Failover On Schedule Gave The Bank Something It Had Never Held Before - Evidence That Its Recovery Commitments Actually Work.