Root Cause Detection Using Service Graph Causal Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional root cause detection techniques in large and complex IT systems are time-consuming and inefficient, often leading to 'alarm storms' where a single anomaly triggers numerous alarms, obscuring the actual root cause due to cascading effects, making it difficult to identify and resolve issues effectively.

Innovation Solution

A root cause detection and diagnosis system that utilizes a service graph to group alarms based on dependencies and applies a causal analysis model to infer probable root causes, leveraging a knowledge database of past anomaly events to generate corrective actions, thereby reducing the complexity of identifying and resolving anomalies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional root cause detection techniques are used in large and complex IT systems, then comprehensive monitoring of all alarms is achieved, but the time required to resolve anomalies increases significantly

Engineering Contradiction:
Improvealarm detection completenessVSAvoidanomaly resolution time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system segments the complex alarm monitoring task by introducing intermediate components (alarm correlation module, causal analysis engine) that divide the problem into manageable parts: collecting alarms, correlating them with service dependencies, analyzing causal relationships, and identifying root causes. This segmentation enables efficient processing without compromising detection completeness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs intermediary elements such as the service dependency graph and alarm correlation rules that mediate between raw alarm data and root cause identification. These intermediaries transform complex alarm streams into structured information that can be quickly analyzed, reducing resolution time while maintaining detection accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If traditional alarm monitoring methods are used, then all alarm events are captured, but alarm storms obscure the actual root cause

Engineering Contradiction:
Improvealarm event coverageVSAvoidroot cause visibility
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The system merges multiple correlated alarms into unified root cause identifications by leveraging service dependency relationships. Alarms that are causally related are combined and analyzed together, preventing alarm storms from obscuring the actual root cause while maintaining complete coverage of individual alarm events.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The causal analysis engine uses feedback from service dependency graphs and alarm correlation results to continuously refine root cause identification. When alarms are correlated, the system feeds this information back into the analysis process to adjust probability assessments and improve root cause visibility amidst alarm storms.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If manual root cause analysis is performed, then detailed investigation of each alarm is possible, but productivity decreases due to time-consuming processes

Engineering Contradiction:
Improveroot cause analysis accuracyVSAvoidanomaly resolution efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically analyzing alarms, correlating them with service dependencies, and identifying root causes without requiring manual investigation. The causal analysis engine and knowledge base enable the system to autonomously determine root causes with high accuracy, significantly improving productivity while maintaining analysis precision.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical analysis processes with automated computational systems. The causal analysis engine, alarm correlation module, and knowledge base constitute a mechanical-substitution system that performs root cause analysis automatically, eliminating time-consuming manual processes while preserving or enhancing analysis accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Reliability

If comprehensive alarm monitoring is implemented, then system-wide issues are detected, but device complexity increases

Engineering Contradiction:
Improvesystem-wide detection capabilityVSAvoidmonitoring system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The monitoring system achieves universality by implementing multi-functional components that handle multiple tasks. The alarm correlation module simultaneously collects alarms, correlates them with service dependencies, and feeds information to the causal analysis engine. This multi-functionality enables system-wide detection without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces a new dimension of analysis by incorporating service dependency graphs and causal relationships into alarm monitoring. Instead of analyzing alarms in isolation, the system adds the dimension of service relationships, enabling system-wide detection while managing complexity through structured dimensional expansion.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11269718B1Root cause detection and corrective action diagnosis system
Publication Date: 2022.03.08 AMAZON TECH INC
  • US11269718B1 patent drawing
  • US11269718B1 patent drawing
  • US11269718B1 patent drawing

AI summary

Methods, systems, and computer-readable media for automatically detecting root causes of anomalies occurring in information technology (IT) systems are disclosed. In some embodiments, data of a service graph depicting dependencies between nodes or services of the IT infrastructure is traversed to determine propagation patterns of anomaly symptoms/alarms through the IT infrastructure. Also, a causal inference model is used to determine probabilities that an observed propagation pattern corresponds to a stored propagation pattern, wherein a close correspondence indicates that the current anomaly is likely caused by a similar root cause as a past anomaly that caused the stored propagation pattern.