Automated Crash Recovery in Hyperconverged Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing tools for crash recovery in complex clustered server environments are inadequate as they fail to collect and analyze data from multiple machines and virtual environments, and cannot perform automatic recovery actions, relying on human intervention and knowledge bases for resolving crashes.

Innovation Solution

A system and method for detecting crashes in a cluster environment that generates a call trace, creates a crash ID, checks against a knowledge base for known issues, and applies automatic recovery procedures such as restarting services, updating software, or rebooting machines, while collecting logs from machines and virtual environments to identify and address crashes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If existing crash detection tools are used in cluster environments, then crash data from a single machine can be collected, but data from multiple machines and virtual environments cannot be comprehensively analyzed

Engineering Contradiction:
Improvecrash information completenessVSAvoidsystem architecture complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges crash detection and data collection capabilities across multiple machines and virtual environments into a unified cluster-wide system. The crash detection agent on each node collects local crash information and sends it to a central server, which aggregates and analyzes data from the entire cluster, enabling comprehensive crash analysis that spans physical machines, virtual machines, and containers.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements a nested data collection structure where crash information is collected at multiple levels: individual application crashes within virtual machines, virtual machine crashes within host systems, and host system crashes within the cluster. Each level nests its data collection within the broader context of the next level, allowing hierarchical analysis from application to cluster level.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Productivity

If manual intervention is used for crash recovery, then recovery actions can be customized, but recovery speed is reduced due to human response time

Engineering Contradiction:
Improverecovery speedVSAvoidrecovery automation level
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The patent implements self-service automated recovery where the crash detection system automatically analyzes crash information, identifies the cause, and executes appropriate recovery actions without human intervention. The system maintains a knowledge base of known crash scenarios and automatically applies appropriate fixes, restarting services, or isolating affected components, enabling rapid self-healing of the cluster.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent establishes a feedback loop where crash detection triggers automated analysis and recovery actions, and the results are fed back to update the knowledge base. This feedback mechanism allows the system to learn from each crash event and improve future recovery actions, creating a continuously improving automated recovery system.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If comprehensive log collection from all cluster components is performed, then crash analysis accuracy is improved, but data collection time and processing overhead increase

Engineering Contradiction:
Improvecrash analysis accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by having crash detection agents continuously monitor and buffer crash-relevant information before actual crashes occur. When a crash is detected, the pre-collected data is immediately available for analysis, eliminating the need for time-consuming post-crash data gathering. The system maintains ready-to-analyze buffers of logs, metrics, and state information for all cluster components.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If automated recovery procedures are implemented, then recovery speed is improved, but the ability to handle unique crash scenarios is reduced

Engineering Contradiction:
Improverecovery speedVSAvoidrecovery scenario coverage
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal recovery system that can handle multiple types of crash scenarios through a single automated framework. The knowledge base stores diverse recovery procedures for different crash types (application crashes, virtual machine crashes, host system crashes, network failures), and the automated system selects and executes the appropriate procedure based on the detected crash scenario, providing versatile coverage across the entire cluster ecosystem.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12019504B1Automated software crash recovery in hyperconverged systems using centralized knowledge database
Publication Date: 2024.06.25 VIRTUOZZO INT GMBH
  • US12019504B1 patent drawing
  • US12019504B1 patent drawing
  • US12019504B1 patent drawing

AI summary

A system and method for detecting and fixing crashes in a cluster environment, including detecting a crash; generating a call trace of the crash; generating a crash ID based on the call trace; checking if the crash ID matches a known crash ID from a knowledge base; when the crash ID matches, applying an automatic recovery procedure, including any of (a) restarting a service that caused the crash; (b) removing and replacing a software package that caused the crash; (c) updating software that caused the crash; and (d) rebooting a machine where the crash occurred; when the crash ID does not match, (a) collecting logs on the machine where the crash occurred; (b) collecting logs from any virtual environments on the machine where the crash occurred; and (c) generating crash ID and sending the crash ID and the logs to the knowledge base.