Auto-Recovery for Software Systems via Event Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technology remediation systems are inefficient in quickly identifying and addressing root causes of issues in complex, cloud-based computer systems, relying heavily on human intervention and manual processes, which increases Mean Time to Repair (MTTR) and can lead to false positives and unaddressed systemic issues.

Innovation Solution

A system that correlates break events in new processes with existing processes to automate the creation of runbooks, using monitoring systems to generate detailed alerts and associate them with automated fix events from similar processes, allowing for automated responses without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual break-fix processes are used with human intervention, then flexibility and adaptability are maintained, but Mean Time to Repair (MTTR) increases and productivity decreases

Engineering Contradiction:
Improveflexibility in handling break eventsVSAvoidMean Time to Repair (MTTR)
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system enables automated self-service through correlation between break events and runbooks. When a break event occurs, the system automatically correlates it with existing runbooks and executes remediation actions without human intervention, allowing the system to fix itself while maintaining operational flexibility through intelligent correlation algorithms.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by pre-correlating break events with runbooks during system operation. Runbooks are created and stored in advance with associated break events, so when a break event occurs, the correlated runbook is already ready for immediate execution, eliminating manual lookup and decision time.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If automated runbook creation is implemented, then productivity and MTTR improvement are achieved, but system complexity increases

Engineering Contradiction:
Improveautomation of remediation processesVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The correlation system serves multiple functions: it creates runbooks automatically, stores them in a library, correlates break events with appropriate runbooks, and executes remediation actions. This multi-functional approach consolidates what could be separate complex systems into a unified platform, managing complexity through integration rather than proliferation of components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system implements nesting by embedding the runbook creation and correlation logic within the existing monitoring and incident management infrastructure. The correlation system operates as a nested layer that leverages existing break event detection and runbook execution capabilities, avoiding the need for completely separate complex systems.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Measurement precision

If detailed monitoring and correlation systems are deployed, then measurement precision and detection accuracy improve, but device complexity and resource requirements increase

Engineering Contradiction:
Improvebreak event detection accuracyVSAvoidmonitoring system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The correlation system acts as an intermediary layer between break event detection and runbook execution. It receives break events from monitoring systems, performs correlation analysis with stored runbooks, and triggers appropriate remediation actions. This intermediary approach enhances detection precision without requiring the monitoring system itself to become more complex.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11314610B2Auto-recovery for software systems
Publication Date: 2022.04.26 CAPITAL ONE SERVICES LLC
  • US11314610B2 patent drawing
  • US11314610B2 patent drawing
  • US11314610B2 patent drawing

AI summary

Disclosed are hardware and techniques for building runbooks for new computer-implemented processes by correlating break events from the new processes with break events extant in existing runbooks for existing computer-implemented processes. In addition, fix events associated with the correlated break events are evaluated to determine the likelihood that they will be able to fix the error condition which caused the break event from the new process. The fix events are presented to a human operator who may select and test each fix event to determine if the error condition is directed and, if so, the correlated break event associated with the fix event are merged together and added to a new runbook for the new computer-implement process.