Data Center Telemetry Grouping for DU/DL Failure Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data center management systems struggle to accurately predict and prevent data unavailability and data loss (DU/DL) events, which can cause significant disruptions, often leading to costly and time-consuming manual restoration processes.

Innovation Solution

Implementing a data-driven deep learning approach using a residual neural network to analyze telemetry metrics from data center assets, grouping them based on functionality, and using service level agreements (SLAs) to prioritize predictive actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual monitoring and restoration processes are used for data center failures, then operational simplicity is maintained, but response time and reliability deteriorate due to slower detection and restoration speeds

Engineering Contradiction:
Improvefailure prediction accuracyVSAvoidmonitoring system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces manual monitoring and restoration processes with an automated deep learning-based system. The residual neural network automatically analyzes telemetry metrics, predicts failures, and triggers restoration actions, substituting human operations with an intelligent automated system that operates continuously without intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs preliminary actions by predicting failures before they actually occur. The deep learning model analyzes historical and real-time telemetry data to identify patterns that indicate impending failures, allowing the system to take preventive restoration actions before data unavailability or loss occurs.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If real-time telemetry analysis is performed using deep learning, then failure prediction accuracy is improved, but computational resource consumption increases

Engineering Contradiction:
Improvetelemetry metric analysis precisionVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the telemetry metrics into different categories (storage metrics, network metrics, compute metrics) and processes them through separate analysis paths in the residual neural network. This segmentation allows the system to focus computational resources on the most critical metrics and reduce overall processing load while maintaining high prediction accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts analysis parameters based on the current state of the data center. The residual neural network modifies its processing depth and computational intensity based on the urgency and severity of detected anomalies, allocating more computational resources to critical failures and fewer resources to less severe issues.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If comprehensive telemetry metrics are collected from all data center assets, then prediction completeness is improved, but data processing complexity increases

Engineering Contradiction:
Improvetelemetry data volumeVSAvoiddata processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent extracts and focuses on the most critical telemetry metrics related to storage, network, and compute functions that are directly associated with data unavailability and loss. The residual neural network is designed to process only these essential metrics rather than all possible telemetry data, reducing processing complexity while maintaining comprehensive prediction capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments telemetry data by functional category (storage, network, compute) and processes each segment through dedicated analysis paths. This segmentation organizes the large volume of telemetry data into manageable groups, reducing processing complexity while maintaining complete coverage of all critical asset types.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250254098A1Data Center Monitoring and Management Operation for Data Unavailable/Data Loss (DU/DL) Predictions, Including Service Level Agreement (SLA) Failure Prediction
Publication Date: 2025.08.07 DELL PROD LP
  • US20250254098A1 patent drawing
  • US20250254098A1 patent drawing
  • US20250254098A1 patent drawing

AI summary

A system, method, and computer-readable medium for performing a data center management and monitoring operation to predict event failures, comprising taking time stamped telemetry metrics of data asset clusters; sampling the time stamped telemetry metrics; separating the sampled time stamped telemetry metrics into groups based on metric functionality; providing data of the groups into a residual neural network to learn patterns of the three groups; and concatenating output of the residual neural network of the groups, wherein concatenated output is feed back to the residual neural network to learn relationships of the three groups.