Auto-Recovery for Software Systems via Event Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technology remediation systems are inefficient in quickly identifying and addressing root causes of issues in complex, cloud-based computer systems, relying heavily on human intervention and manual processes, which increases Mean Time to Repair (MTTR) and can lead to false positives and unaddressed systemic issues.
Innovation Solution
A system that correlates break events in new processes with existing processes to automate the creation of runbooks, using monitoring systems to generate detailed alerts and associate them with automated fix events from similar processes, allowing for automated responses without human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual break-fix processes are used with human intervention, then flexibility and adaptability are maintained, but Mean Time to Repair (MTTR) increases and productivity decreases
Solution Approach 1:
The system enables automated self-service through correlation between break events and runbooks. When a break event occurs, the system automatically correlates it with existing runbooks and executes remediation actions without human intervention, allowing the system to fix itself while maintaining operational flexibility through intelligent correlation algorithms.
Solution Approach 2:
The system performs preliminary actions by pre-correlating break events with runbooks during system operation. Runbooks are created and stored in advance with associated break events, so when a break event occurs, the correlated runbook is already ready for immediate execution, eliminating manual lookup and decision time.
2Productivity
If automated runbook creation is implemented, then productivity and MTTR improvement are achieved, but system complexity increases
Solution Approach 1:
The correlation system serves multiple functions: it creates runbooks automatically, stores them in a library, correlates break events with appropriate runbooks, and executes remediation actions. This multi-functional approach consolidates what could be separate complex systems into a unified platform, managing complexity through integration rather than proliferation of components.
Solution Approach 2:
The system implements nesting by embedding the runbook creation and correlation logic within the existing monitoring and incident management infrastructure. The correlation system operates as a nested layer that leverages existing break event detection and runbook execution capabilities, avoiding the need for completely separate complex systems.
3Measurement precision
If detailed monitoring and correlation systems are deployed, then measurement precision and detection accuracy improve, but device complexity and resource requirements increase
Solution Approach 1:
The correlation system acts as an intermediary layer between break event detection and runbook execution. It receives break events from monitoring systems, performs correlation analysis with stored runbooks, and triggers appropriate remediation actions. This intermediary approach enhances detection precision without requiring the monitoring system itself to become more complex.
Data Source
AI summary
Disclosed are hardware and techniques for building runbooks for new computer-implemented processes by correlating break events from the new processes with break events extant in existing runbooks for existing computer-implemented processes. In addition, fix events associated with the correlated break events are evaluated to determine the likelihood that they will be able to fix the error condition which caused the break event from the new process. The fix events are presented to a human operator who may select and test each fix event to determine if the error condition is directed and, if so, the correlated break event associated with the fix event are merged together and added to a new runbook for the new computer-implement process.


