Distributed Computing Clusters for Rapid Catastrophic Failure Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for recovering computing capacity and critical software applications after a catastrophic failure are time-consuming and burdensome, requiring manual intervention or frequent updates of script files, which can lead to prolonged downtime and resource exhaustion.
Innovation Solution
A method and system that distribute computing capacity across multiple clusters with shared data and software applications, allowing for automatic detection and rapid recovery of workloads to backup capacity, enabling quick restoration of computing capacity and critical software applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual recovery procedures or script files are used to restore computing capacity after catastrophic failure, then recovery can be performed, but the process is time-consuming and requires frequent maintenance updates
Solution Approach 1:
The patent applies preliminary action by pre-configuring backup computing clusters with identical software applications and data before catastrophic failure occurs. The backup infrastructure is prepared in advance with all necessary components (software, data, configuration) so that recovery simply requires activating pre-prepared resources rather than rebuilding from scratch. This eliminates the time-consuming manual setup and script execution required by conventional methods.
Solution Approach 2:
The patent uses copying by maintaining duplicate copies of critical software applications and data on backup computing clusters. Instead of manually recreating or restoring applications after failure, the system has ready-made copies that can be immediately activated. The backup clusters contain replicated versions of all critical applications and their associated data, enabling instant failover without reconstruction.
2Reliability
If manual recovery procedures are used, then recovery can be performed, but technical support personnel are burdened with tedious and time-consuming maintenance
Solution Approach 1:
The patent implements self-service by enabling automatic failover detection and execution without requiring technical support personnel intervention. The system continuously monitors the operational status of primary computing clusters and automatically activates backup clusters when failures are detected. This eliminates the need for technical staff to manually detect failures, execute recovery commands, or maintain recovery scripts, thereby reducing their burden significantly.
Solution Approach 2:
The patent uses feedback mechanisms by continuously monitoring the operational status of computing applications and automatically triggering recovery procedures when failures are detected. The system has built-in detection capabilities that provide real-time feedback on application health, and this feedback automatically initiates the failover process to backup clusters, eliminating the need for manual monitoring and intervention by technical support personnel.
3Device complexity
If computing capacity is concentrated in a single location, then system management is simplified, but the system is vulnerable to catastrophic failures that wipe out all computing capacity
Solution Approach 1:
The patent applies segmentation by dividing the computing infrastructure into separate primary and backup clusters located at different physical sites. Instead of having all computing capacity in one location, the system segments functionality across geographically distributed clusters. This segmentation ensures that a catastrophic failure at one location does not affect the other, providing resilience while maintaining manageable complexity through clear separation of primary and backup roles.
Solution Approach 2:
The patent uses another dimension by adding a geographic dimension to the computing architecture. Instead of simply having more computing resources at the same location, the system distributes clusters across different physical locations and networks. This spatial dimensionality provides natural protection against localized catastrophic events while the automated failover mechanism manages the complexity of coordinating across multiple locations.
Data Source
AI summary
Method/system is disclosed for recovering computing capacity and critical applications after a catastrophic failure. The method/system involves distributing the computing capacity over multiple computing clusters, each computing cluster having concurrent access to shared data and software applications of other computing clusters. Sufficient backup computing capacity is reserved on each computing cluster to recover some or all active computing capacity on the other computing clusters. Message traffic throughout the computing clusters is monitored for indications of a catastrophic failure. Upon confirmation of a catastrophic failure at one computing cluster, the workloads of that computing cluster are transferred to the backup computing capacity of the other computing clusters. Software applications that have been designated for recovery are then brought up on the backup computing capacity of the other computing clusters. Such an arrangement allows computing capacity and critical software applications to be quickly recovered after a catastrophic failure.


