Automated Crash Recovery in Hyperconverged Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing tools for crash recovery in complex clustered server environments are inadequate as they fail to collect and analyze data from multiple machines and virtual environments, and cannot perform automatic recovery actions, relying on human intervention and knowledge bases for resolving crashes.
Innovation Solution
A system and method for detecting crashes in a cluster environment that generates a call trace, creates a crash ID, checks against a knowledge base for known issues, and applies automatic recovery procedures such as restarting services, updating software, or rebooting machines, while collecting logs from machines and virtual environments to identify and address crashes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If existing crash detection tools are used in cluster environments, then crash data from a single machine can be collected, but data from multiple machines and virtual environments cannot be comprehensively analyzed
Solution Approach 1:
The patent merges crash detection and data collection capabilities across multiple machines and virtual environments into a unified cluster-wide system. The crash detection agent on each node collects local crash information and sends it to a central server, which aggregates and analyzes data from the entire cluster, enabling comprehensive crash analysis that spans physical machines, virtual machines, and containers.
Solution Approach 2:
The patent implements a nested data collection structure where crash information is collected at multiple levels: individual application crashes within virtual machines, virtual machine crashes within host systems, and host system crashes within the cluster. Each level nests its data collection within the broader context of the next level, allowing hierarchical analysis from application to cluster level.
2Productivity
If manual intervention is used for crash recovery, then recovery actions can be customized, but recovery speed is reduced due to human response time
Solution Approach 1:
The patent implements self-service automated recovery where the crash detection system automatically analyzes crash information, identifies the cause, and executes appropriate recovery actions without human intervention. The system maintains a knowledge base of known crash scenarios and automatically applies appropriate fixes, restarting services, or isolating affected components, enabling rapid self-healing of the cluster.
Solution Approach 2:
The patent establishes a feedback loop where crash detection triggers automated analysis and recovery actions, and the results are fed back to update the knowledge base. This feedback mechanism allows the system to learn from each crash event and improve future recovery actions, creating a continuously improving automated recovery system.
3Measurement precision
If comprehensive log collection from all cluster components is performed, then crash analysis accuracy is improved, but data collection time and processing overhead increase
Solution Approach 1:
The patent implements preliminary action by having crash detection agents continuously monitor and buffer crash-relevant information before actual crashes occur. When a crash is detected, the pre-collected data is immediately available for analysis, eliminating the need for time-consuming post-crash data gathering. The system maintains ready-to-analyze buffers of logs, metrics, and state information for all cluster components.
4Productivity
If automated recovery procedures are implemented, then recovery speed is improved, but the ability to handle unique crash scenarios is reduced
Solution Approach 1:
The patent creates a universal recovery system that can handle multiple types of crash scenarios through a single automated framework. The knowledge base stores diverse recovery procedures for different crash types (application crashes, virtual machine crashes, host system crashes, network failures), and the automated system selects and executes the appropriate procedure based on the detected crash scenario, providing versatile coverage across the entire cluster ecosystem.
Data Source
AI summary
A system and method for detecting and fixing crashes in a cluster environment, including detecting a crash; generating a call trace of the crash; generating a crash ID based on the call trace; checking if the crash ID matches a known crash ID from a knowledge base; when the crash ID matches, applying an automatic recovery procedure, including any of (a) restarting a service that caused the crash; (b) removing and replacing a software package that caused the crash; (c) updating software that caused the crash; and (d) rebooting a machine where the crash occurred; when the crash ID does not match, (a) collecting logs on the machine where the crash occurred; (b) collecting logs from any virtual environments on the machine where the crash occurred; and (c) generating crash ID and sending the crash ID and the logs to the knowledge base.


