Mainframe Failover Orchestration for Rapid Disaster Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current disaster recovery testing for mainframe systems is isolated and lacks comprehensive validation, failing to ensure consistent and rapid failover and recovery across alternate regions, which is critical for enterprise-class systems facing cyber threats.
Innovation Solution
A system and method that involves invoking a failover script, validating system shutdown, preventing workload movement, holding batch work, initiating a system close with memory de-stage, and passing control to a remote region for consistent hardware replication, ensuring rapid and automated recovery of mainframe systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If current disaster recovery testing process is used, then platform recovery can be demonstrated, but the testing is isolated by firewalls and lacks comprehensive validation across alternate regions
Solution Approach 1:
The system segments the disaster recovery validation into distinct phases: local platform recovery testing, cross-region failover testing, and comprehensive end-to-end validation. This segmentation allows each aspect to be tested independently while ensuring overall system reliability across distributed regions.
Solution Approach 2:
An intermediary failover orchestration system is introduced that coordinates between the primary mainframe system, alternate regional systems, and validation mechanisms. This intermediary enables comprehensive cross-region validation without requiring direct exposure between systems, resolving the contradiction between isolation and versatility.
2Loss of time
If rapid failover is implemented to ensure sustained resiliency, then recovery time is reduced, but system complexity increases due to automated processes and controls
Solution Approach 1:
The system performs preliminary actions by pre-configuring failover scripts, pre-validating system shutdown procedures, and pre-establishing alternate regional environments. This preliminary preparation enables rapid failover execution without complex real-time decision-making, reducing recovery time while managing system complexity.
Solution Approach 2:
The automated failover system implements self-service mechanisms where the system automatically validates its own shutdown state, self-coordinates the failover process, and autonomously manages the transition to alternate regions. This self-service approach reduces the need for external control complexity while maintaining rapid recovery capability.
3Manufacturing precision
If comprehensive system validation is performed across alternate regions, then data integrity is ensured, but the testing process becomes more complex and time-consuming
Solution Approach 1:
The system implements feedback mechanisms that automatically validate data integrity during failover testing by comparing system states, verifying data consistency, and confirming successful replication to alternate regions. This automated feedback ensures data integrity without requiring manual verification, maintaining testing efficiency while ensuring precision.
Solution Approach 2:
The validation process uses copying mechanisms to replicate system states and data to alternate regions, then verifies the copies against the original. This copying approach ensures data integrity through systematic comparison while automating the verification process, thus maintaining testing efficiency alongside precision.
Data Source
AI summary
An embodiment of the present invention is directed to enabling a mainframe system to be shutdown and restarted in an alternate region within minutes in a consistent and demonstrated manner ensuring data consistency for various components including disk, storage, coupling facility, etc. This enhances and packages together various software products from a mainframe platform in order to deliver a solution. An embodiment of the present invention is directed to an integrated automation that validates the integrity of the systems after restarting in remote regions.


