Autonomous System Anomaly Remediation via ML Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing monitoring and logging tools in IT operations fail to capture information related to remediation actions, leading to inconsistent and time-consuming human-driven processes that negatively impact platform availability and increase support costs.
Innovation Solution
The implementation of cognitively assorted machine learning algorithms that identify system anomalies by generating inference models and using machine reinforcement learning to determine remediation actions, predicting adverse states, and automatically executing actions to address these issues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If human-driven remediation is used to identify and address system issues, then flexibility in handling diverse problems is improved, but consistency and speed of remediation deteriorate
Solution Approach 1:
The system uses machine learning models to autonomously identify system anomalies and select appropriate remediation actions without human intervention. The inference model automatically analyzes system log data, identifies adverse states and root causes, while the action policy automatically selects and executes remediation actions, enabling the system to serve itself and achieve consistent, scalable remediation across diverse problems
Solution Approach 2:
The system implements a feedback loop where the inference model continuously monitors system log data, predicts adverse states, identifies root causes, and the action policy executes remediation actions based on this analysis. The system learns from historical data and feedback to improve its anomaly detection and remediation capabilities over time, maintaining both consistency and adaptability
2Adaptability or versatility
If human-driven remediation is used to address system issues, then complex problem-solving capability is improved, but time consumption and support costs increase
Solution Approach 1:
The system pre-trains inference models using historical system log data before deployment. The models are prepared in advance to quickly identify anomalies and predict adverse states when deployed, eliminating the need for time-consuming human analysis during actual remediation events while maintaining complex problem-solving capability
Solution Approach 2:
The system replaces human engineers with machine learning-based automated remediation. The inference model and action policy constitute an automated mechanical system that processes system log data and executes remediation actions without human intervention, dramatically reducing time consumption and support costs while maintaining or improving remediation quality
3Reliability
If automated remediation systems are implemented, then speed and consistency of remediation are improved, but system complexity and development costs increase
Solution Approach 1:
The automated remediation system is segmented into distinct functional modules: the inference model for anomaly detection and root cause identification, and the action policy for remediation selection and execution. This modular architecture separates concerns, making the system easier to develop, maintain, and update while maintaining high consistency in remediation operations
Solution Approach 2:
The system introduces machine learning models as intermediaries between system log data and remediation actions. The inference model acts as an intermediary to translate raw log data into meaningful anomaly detections and root cause identifications, while the action policy serves as an intermediary to translate system states into appropriate remediation actions, simplifying the overall system architecture
Data Source
AI summary
Methods, apparatus, and processor-readable storage media for identifying and remediating anomalies through cognitively assorted machine learning algorithms are provided herein. A computer-implemented method includes: identifying, using system log data, a target variable based at least in part on correlations between a set of performance indicators of a system and the target variable, and threshold values for the performance indicators relative to the target variable; generating an inference model to predict when the system will enter an adverse state and identify one or more root causes of the system entering the adverse state; using machine reinforcement learning to determine an action policy including actions that remediate the adverse state; predicting that the system will enter the adverse state by applying the inference model to further system log data; and automatically executing one or more actions of the action policy in response to the prediction.


