Cross-Layer Incident Investigation With Automated Root Cause Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing incident investigation systems are domain-based and lack awareness of application and platform layers, often providing inaccurate responses and requiring manual processes.
Innovation Solution
An automated incident investigation system that performs end-to-end analysis across application, managed infrastructure, and platform layers, utilizing anomaly detection, troubleshooting, and enrichment to identify root causes with minimal manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual incident investigation processes are used, then human expertise can be applied to complex incidents, but investigation time and productivity are reduced
Solution Approach 1:
The system performs self-service by automatically collecting incident data from multiple sources, detecting anomalies, running diagnostics, and generating root cause analysis reports without requiring manual intervention. The automated incident investigation system collects data from alerts, metrics, logs, and change data, processes this information through anomaly detection and diagnostic tools, and produces comprehensive investigation summaries that would otherwise require manual analysis by human experts.
Solution Approach 2:
The patent replaces manual mechanical investigation processes with automated computer-based systems. Instead of human investigators manually analyzing incident data across multiple domains, the system uses automated data collection mechanisms, anomaly detection algorithms, diagnostic tools, and report generation systems to perform the same functions, thereby increasing productivity while maintaining reliability through comprehensive automated analysis.
2Reliability
If domain-based incident investigation systems are used, then specific technical expertise can be applied, but awareness of application and platform layers is lacking
Solution Approach 1:
The automated incident investigation system achieves universality by collecting and analyzing incident data from multiple domains and layers simultaneously. It integrates alert data, metrics data, log data, and change data from application, managed infrastructure, and platform layers into a unified investigation process. The system uses a single consolidated explainer that processes all this multi-layer data to produce comprehensive root cause analysis, enabling it to adapt to various incident types across different layers without requiring separate domain-specific systems.
3Productivity
If automated anomaly detection is implemented, then investigation speed increases, but system complexity increases
Solution Approach 1:
The system segments the incident investigation process into distinct functional modules: data collection from multiple sources, anomaly detection, diagnostic execution, and report generation. Each module handles a specific aspect of the investigation independently, making the overall complex system manageable through modular architecture. The investigation orchestrator coordinates these segmented functions while maintaining overall system coherence and productivity.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
Systems and methods are provided for automated incident investigation. Anomaly detection is used to identify anomalies in incident data (e.g., alerts, changes, metrics, logs, and/or system health), and the identified anomalies are converted into facts (or textual prompt inputs for a large language model ("LLM")). A troubleshooting or diagnostic system is run on the anomalies to provide additional facts to identify a root cause of an incident. The facts from the diagnostics, the facts from the anomaly detections are entered into a consolidated explainer that generates a summary of what happened, what is a likely cause, and what to do next to resolve the issue. In examples, anomaly enrichment data including a time correlation result, a weighted list of abnormal transaction patterns, a list of abnormal trace patterns, a list of exception patterns, a difference pattern, and/or region data are input as further facts to enhance the incident investigation process.