Data Center Root Cause Detection Using Automated Alert Patterns

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data center systems lack an efficient method to automatically detect and address root cause failures, which can lead to critical outages due to interconnected device failures.

Innovation Solution

A root cause detector system that analyzes device alerts using machine learning classifiers and large language models to identify the initial failure causing a chain reaction, enabling proactive detection and resolution of potential system failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual troubleshooting methods are used to identify root cause failures, then human analysis can be applied to complex failure patterns, but troubleshooting time and effort increase significantly

Engineering Contradiction:
Improveroot cause detection accuracyVSAvoidtroubleshooting time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service by having the monitoring system automatically analyze its own alert data to identify root causes without requiring manual intervention. The machine learning model processes alerts autonomously, matching patterns against historical data to detect root cause failures automatically.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical troubleshooting with automated machine learning-based detection. The system uses ML algorithms to process alert patterns, substitute human analytical effort with computational analysis, and automatically identify root causes based on learned patterns from historical failure data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated root cause detection systems are implemented, then troubleshooting efficiency improves, but system complexity increases due to machine learning components

Engineering Contradiction:
Improvetroubleshooting efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system achieves universality by using a single machine learning model that handles multiple types of failures and alert patterns. The model is trained on diverse historical failure data and can generalize to detect various root cause failures across different device types and failure modes, eliminating the need for separate specialized detectors.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies parameter changes by transforming raw alert data into standardized feature representations that the machine learning model can process. The system adjusts and optimizes model parameters during training to efficiently capture failure patterns, enabling accurate detection without requiring overly complex system architecture.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If historical failure data is analyzed to train detection models, then detection accuracy improves, but data storage and processing requirements increase

Engineering Contradiction:
Improvefailure detection reliabilityVSAvoiddata volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system extracts only the critical information from historical failure data by identifying and isolating key alert patterns and failure signatures. Rather than storing and processing all raw historical data, the system extracts and stores only the essential feature representations and pattern templates that are needed for accurate root cause detection.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses copying by creating simplified representations of historical failure patterns as training data for the machine learning model. Instead of working with the full complexity of historical datasets, the system creates condensed copies of failure patterns that capture essential characteristics while reducing data volume and processing requirements.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250272175A1Systems and methods for automatically detecting root cause failures
Publication Date: 2025.08.28 SAP SE
  • US20250272175A1 patent drawing
  • US20250272175A1 patent drawing
  • US20250272175A1 patent drawing

AI summary

Described herein are techniques for automatically detecting root cause failures in a computing environment. A data center may experience a critical system failure that renders the data center inoperable. An administrator may utilize a root cause detector to analyze alerts generated from the data center to automatically detect the root cause of the critical system failure. Once detected, the administrator may investigate the device with the root cause failure to repair the data center. In some examples, the root cause detector may compare alerts received from the data center with root cause failures in a failure repository that it is familiar with to determine whether the present sequence of alerts is similar to alert patterns it has seen before. The root cause detector may be implemented with a classifier model, a large language model, a rule-based heuristics identifying spurious alert patterns, or combinations of these techniques.