Computing Service Mitigation Workflow Generation from Historical Incidents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing incident management techniques for computing services, such as human-driven and hybrid mitigation, are inefficient and labor-intensive, leading to prolonged downtime and increased risk of errors in addressing technical issues.
Innovation Solution
Automatically generating a mitigation workflow using historical mitigation workflows, leveraging machine learning to identify relevant operations based on confidence factors and relevance criteria, thereby reducing manual effort and improving the speed and effectiveness of issue resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human-driven mitigation is used, then the on-call engineer can perform all tasks required to mitigate an incident, but the speed and effectiveness are compromised due to limited tooling and lack of domain knowledge
Solution Approach 1:
The patent introduces an automated incident management system that acts as an intermediary between the on-call engineer and the computing service. This system uses machine learning models to analyze incident data, determine root causes, and execute mitigation workflows automatically, thereby reducing the time loss while maintaining or improving effectiveness through consistent, data-driven decisions.
Solution Approach 2:
The incident management system enables self-service by automatically performing mitigation tasks without requiring continuous human intervention. The system can autonomously execute pre-configured workflows, apply fixes, and restore service based on learned patterns from historical incident data, significantly reducing downtime while maintaining reliability.
2Productivity
If pre-configured workflows are used in hybrid mitigation, then some tasks are automated, but manual creation and maintenance of workflows requires substantial effort
Solution Approach 1:
The patent implements dynamic workflow generation where the incident management system learns from historical incident data and automatically adapts workflows based on new patterns. Rather than requiring manual updates for every service change, the system dynamically adjusts its mitigation strategies, reducing the complexity of workflow maintenance while maintaining high automation levels.
Solution Approach 2:
The system incorporates feedback loops where outcomes of automated mitigations are continuously analyzed to improve future workflows. This self-learning mechanism allows the system to automatically refine its processes based on real-world performance data, reducing manual maintenance effort while improving productivity over time.
3Extent of automation
If pre-configured workflows are maintained across evolving computing services, then automation is achieved, but consistency and coordination become challenging as services grow in complexity
Solution Approach 1:
The patent creates a universal incident management system that can handle multiple types of incidents across different computing services through a common framework. The machine learning models are trained on diverse historical data and can generalize to new service configurations, maintaining automation consistency while adapting to service evolution without requiring service-specific workflow maintenance.
Solution Approach 2:
The system performs preliminary actions by pre-configuring mitigation workflows based on historical incident patterns before actual incidents occur. These pre-learned workflows are automatically applied when similar incidents are detected, enabling consistent automated response across evolving services without requiring manual reconfiguration for each new service version.
Data Source
AI summary
Techniques are described herein that are capable of generating a mitigation workflow for a computing service using historical mitigation workflows. A determination is made that a historical technical issue that was encountered by a first computing service corresponds to a current technical issue that is encountered by a second computing service. A workflow, which is configured to mitigate the current technical issue, is generated to include historical mitigation operations that are included in historical mitigation workflows that were performed to mitigate the historical technical issue based at least in part on the historical technical issue corresponding to the current technical issue.


