Data Center Recovery Agents for Catastrophic Failure
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for recovering computing capacity and critical software applications after a catastrophic failure are time-consuming and burdensome, requiring manual intervention or frequent updates of script files, which can lead to prolonged downtime and resource exhaustion.
Innovation Solution
A method and system that distribute computing capacity across multiple clusters, allowing for automatic detection and rapid recovery of workloads to backup capacity, with recovery agents monitoring message traffic and activating capacity backup units to bring critical software applications back online.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual recovery procedures are used with technical support personnel, then recovery can be performed with human judgment and adaptability, but recovery time is prolonged and operational efficiency decreases
Solution Approach 1:
The patent implements pre-configured recovery scripts that contain all necessary commands and parameters for application recovery. These scripts are prepared in advance and stored in the system, allowing automatic execution when failures occur without requiring manual intervention. This preliminary preparation eliminates the time-consuming manual command entry while maintaining accurate recovery procedures.
Solution Approach 2:
The system enables self-service recovery through automatic detection of catastrophic failures and autonomous execution of recovery scripts. The technical support system automatically monitors system status, detects failures, and triggers recovery procedures without requiring human operators to manually initiate or manage the recovery process, significantly reducing downtime.
2Productivity
If script files are used for automatic recovery, then recovery time is reduced and automation is improved, but script files require frequent maintenance and updates which increases operational burden
Solution Approach 1:
The patent implements a feedback mechanism where the system automatically monitors application performance and system status, comparing actual state against expected state. When discrepancies are detected, the system automatically updates recovery scripts to reflect current application configurations and parameters, eliminating the need for manual script maintenance while ensuring scripts remain accurate and up-to-date.
Solution Approach 2:
The system performs self-maintenance of recovery scripts by automatically detecting configuration changes in applications and updating the embedded recovery commands accordingly. This self-updating capability eliminates the operational burden of manual script maintenance while maintaining accurate recovery procedures that reflect current system state.
3Reliability
If excess computing capacity is reserved offsite at remote locations, then business continuity is ensured during catastrophic failures, but recovery time increases due to distance and manual setup requirements
Solution Approach 1:
The patent implements pre-configured recovery scripts that are distributed to and stored at remote backup locations in advance. These scripts contain all necessary commands, parameters, and configuration information needed for immediate application recovery. When failures occur, the scripts are automatically executed at the remote location without requiring manual setup or configuration, enabling rapid recovery despite the physical distance.
Solution Approach 2:
The patent uses automated recovery scripts as an intermediary between the failure detection system and the remote backup computing capacity. The scripts serve as a bridge that automatically translates failure detection into coordinated recovery actions at remote locations, eliminating the need for manual coordination and significantly reducing the time required to activate remote backup resources.
4Reliability
If technical support personnel continuously monitor and maintain recovery scripts, then recovery accuracy is maintained, but operational efficiency and resource utilization decrease
Solution Approach 1:
The patent implements automatic feedback mechanisms that continuously monitor application configurations, system parameters, and performance metrics. This feedback is automatically fed into the recovery script management system, which compares current state against stored script definitions and automatically updates scripts when changes are detected. This eliminates the need for continuous manual monitoring while maintaining high recovery accuracy through real-time configuration tracking.
Solution Approach 2:
The patent replaces the manual mechanical process of script maintenance with an automated electronic system. Instead of technical personnel manually reviewing and updating recovery scripts, the system uses automated software agents to detect configuration changes, generate updated scripts, and validate their correctness. This substitution of manual labor with automated electronic processes maintains recovery accuracy while dramatically improving personnel efficiency.
Data Source
AI summary
Method/system is disclosed for recovering computing capacity and critical applications after a catastrophic failure. The method/system involves distributing the computing capacity over multiple computing clusters, each computing cluster having concurrent access to shared data and software applications of other computing clusters. Sufficient backup computing capacity is reserved on each computing cluster to recover some or all active computing capacity on the other computing clusters. Message traffic throughout the computing clusters is monitored for indications of a catastrophic failure. Upon confirmation of a catastrophic failure at one computing cluster, the workloads of that computing cluster are transferred to the backup computing capacity of the other computing clusters. Software applications that have been designated for recovery are then brought up on the backup computing capacity of the other computing clusters. Such an arrangement allows computing capacity and critical software applications to be quickly recovered after a catastrophic failure.


