Datacenter Outage Detection Using ML Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As datacenters incorporate more devices, processing resources, and applications, efficiently identifying the source of a datacenter outage becomes increasingly difficult, leading to prolonged downtime and decreased user experience.

Innovation Solution

The use of machine learning models that analyze near real-time and offline data to detect datacenter mass outages, processing data from various sources such as server temperatures, power usage, and ticketing data to identify anomalous parameters and project the sources of the outage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If more devices and applications are implemented in a datacenter, then the datacenter's processing capability and functionality are improved, but the difficulty of identifying the source of an outage increases

Engineering Contradiction:
Improvedatacenter functionalityVSAvoidoutage source identification
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the complex datacenter monitoring problem into multiple independent analysis components: collecting data from diverse sources (servers, network devices, applications), analyzing each data type separately using appropriate methods, and then integrating results to identify outage sources. This segmentation makes the overall system manageable despite the increasing complexity of datacenter environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary analysis system that sits between the raw data from multiple datacenter sources and the final outage identification. This intermediary layer collects, processes, and correlates data from various sources using multiple analysis techniques, thereby mediating the complexity between diverse data sources and the need for clear outage source identification.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If traditional monitoring methods are used to identify outage sources, then system simplicity is maintained, but the time to detect and resolve outages increases

Engineering Contradiction:
Improvemonitoring system complexityVSAvoidoutage resolution time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent merges multiple analysis techniques (statistical analysis, machine learning, rule-based analysis) and multiple data sources into a single integrated monitoring system. This combination enables comprehensive outage detection and source identification that is both fast and accurate, overcoming the limitations of traditional single-method approaches while managing complexity through unified architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary actions by continuously collecting and pre-processing data from all datacenter sources before outages occur. Historical data is stored and analyzed in advance, establishing baselines and patterns that enable rapid outage detection and source identification when anomalies occur, thereby reducing overall resolution time.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If comprehensive data collection from multiple sources is performed, then the accuracy of outage source identification is improved, but the processing complexity and computational resources required increase

Engineering Contradiction:
Improveoutage source identification accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the comprehensive data processing task into distinct processing pipelines for different data types (server metrics, network data, application logs). Each segment is processed using specialized techniques appropriate to its format and characteristics, improving accuracy while managing complexity through modular processing architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms raw data from multiple sources into standardized parameters and features that can be uniformly analyzed. By changing the parameter representation of diverse data types into common analytical formats, the system achieves high identification accuracy while simplifying the processing complexity through parameter standardization.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4348429B1Detecting datacenter mass outage with near real-time/offline data using ML models
Publication Date: 2025.05.07 ORACLE INT CORP
  • EP4348429B1 patent drawingFigure 1
  • EP4348429B1 patent drawingFigure 2
  • EP4348429B1 patent drawingFigure 3

AI summary

The present embodiments relate to data center outage detection and alert generation. An outage detection service as described herein can process near real-time data from various sources in a datacenter and process the data using a model to determine one or more projected sources of a detected outage. The model as described herein can include one or more machine learning models incorporating a series of rules to process near-real time data and offline data and determine one or more projected sources of an outage. An alert message can be generated to provide the projected sources of the outage and other data relevant to the outage.