Auto-Recovery Engine for Cloud Process Fault Remediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technology remediation systems are inefficient in identifying and addressing the root cause of process breakdowns in complex computer systems, particularly in cloud-based environments, due to their reliance on human expertise and manual processes, which leads to prolonged Mean Time to Repair (MTTR) and increased risk of false positives.
Innovation Solution
An auto-recovery and optimality engine that utilizes a rules engine to evaluate process break events, assign risk assessment values, and generate a dynamic runbook to identify optimal corrective actions, accounting for interdependencies between systems and processes, thereby automating the remediation process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human expertise and manual processes are used to identify and address root causes of process breakdowns, then the accuracy of root cause analysis is improved, but the Mean Time to Repair (MTTR) increases
Solution Approach 1:
The system enables self-service by automatically analyzing process break events, evaluating correlations to possible causes, and generating corrective actions without requiring human intervention. The rules engine autonomously processes break events, assigns risk assessment values, and produces optimized response strategies, allowing the system to remediate itself rather than requiring human expertise for each step of the analysis and repair process.
Solution Approach 2:
The patent replaces the mechanical manual process of human analysis and decision-making with an automated computational system. The rules engine performs correlation analysis, risk assessment, and corrective action generation through computer processing rather than human cognitive functions, substituting the mechanical manual system with an automated digital system that operates continuously without human intervention.
2Loss of time
If automated systems are used to respond to process break events, then the Mean Time to Repair (MTTR) is reduced, but the risk of false positives increases
Solution Approach 1:
The system incorporates feedback mechanisms where break events are continuously monitored, evaluated against stored correlations, and used to refine future responses. The rules engine analyzes the relationship between break events and possible causes, using feedback from actual system behavior to improve the accuracy of automated responses and reduce false positives over time through iterative learning and refinement of correlation data.
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting risk assessment values and correlation thresholds based on system state. The rules engine modifies evaluation parameters such as risk assessment scores and correlation strength requirements to optimize the balance between rapid automated response and accuracy, adapting parameters to current system conditions to minimize false positives while maintaining fast MTTR.
3Quantity of substance
If manual runbooks are used to document fix procedures, then the completeness of corrective actions is improved, but the time required to locate and implement fixes increases
Solution Approach 1:
The patent replaces manual runbook searching and implementation with automated computer-based procedures. The rules engine electronically searches, evaluates, and selects appropriate corrective actions from stored correlations, eliminating the need for manual document review. The system automatically retrieves and executes fixes through computer processes rather than human operators manually searching and implementing procedures.
Solution Approach 2:
The rules engine serves as an intermediary between the stored corrective action data and the actual fix implementation. Rather than requiring direct human access to complete runbooks, the rules engine mediates by processing break events, filtering relevant corrective actions based on correlations and risk assessments, and presenting optimized responses, thereby reducing the time humans need to locate and implement fixes while maintaining completeness.
4Reliability
If system interdependencies are considered in remediation decisions, then the overall system resilience is improved, but the complexity of the remediation process increases
Solution Approach 1:
The patent replaces complex manual analysis of system interdependencies with automated computational evaluation. The rules engine uses computer processing to analyze relationships between processes, services, and systems, automatically determining interdependency impacts without requiring human experts to manually trace and evaluate complex system relationships, thereby reducing the perceived complexity while maintaining comprehensive resilience consideration.
Solution Approach 2:
The rules engine provides multi-functional capability by simultaneously performing correlation analysis, risk assessment, interdependency evaluation, and corrective action generation. Rather than requiring separate specialized processes for each function, the rules engine integrates multiple functions into a single unified system that handles all aspects of remediation decision-making, including interdependency analysis, thereby managing complexity through consolidation rather than multiplication of separate complex processes.
Data Source
AI summary
Disclosed are hardware and techniques for correcting computer process faults by identifying risk associated with correcting a computer process fault and computer processes that may depend on the corrected computer process. The interdependent computer processes in a network may be determined by evaluating a stream of process break flags from a monitoring component coupled to the network. Each computer process break flag in the stream of computer process break flags indicates a process fault detected by the monitoring component and is correlated to a corrective response. The break flag and the corrective response are assigned a risk. A risk matrix accounts for interdependencies between computer processes and identified corrective actions. A final response strategy that corrects the computer process faults is determined using the assigned risk and computer system interdependence. A runbook stores the final response strategy, which may be updated based on changing computer process interdependencies and assigned risk.


