Data Center Telemetry Grouping for DU/DL Failure Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data center management systems struggle to accurately predict and prevent data unavailability and data loss (DU/DL) events, which can cause significant disruptions, often leading to costly and time-consuming manual restoration processes.
Innovation Solution
Implementing a data-driven deep learning approach using a residual neural network to analyze telemetry metrics from data center assets, grouping them based on functionality, and using service level agreements (SLAs) to prioritize predictive actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual monitoring and restoration processes are used for data center failures, then operational simplicity is maintained, but response time and reliability deteriorate due to slower detection and restoration speeds
Solution Approach 1:
The patent replaces manual monitoring and restoration processes with an automated deep learning-based system. The residual neural network automatically analyzes telemetry metrics, predicts failures, and triggers restoration actions, substituting human operations with an intelligent automated system that operates continuously without intervention.
Solution Approach 2:
The system performs preliminary actions by predicting failures before they actually occur. The deep learning model analyzes historical and real-time telemetry data to identify patterns that indicate impending failures, allowing the system to take preventive restoration actions before data unavailability or loss occurs.
2Measurement precision
If real-time telemetry analysis is performed using deep learning, then failure prediction accuracy is improved, but computational resource consumption increases
Solution Approach 1:
The patent segments the telemetry metrics into different categories (storage metrics, network metrics, compute metrics) and processes them through separate analysis paths in the residual neural network. This segmentation allows the system to focus computational resources on the most critical metrics and reduce overall processing load while maintaining high prediction accuracy.
Solution Approach 2:
The system dynamically adjusts analysis parameters based on the current state of the data center. The residual neural network modifies its processing depth and computational intensity based on the urgency and severity of detected anomalies, allocating more computational resources to critical failures and fewer resources to less severe issues.
3Quantity of substance
If comprehensive telemetry metrics are collected from all data center assets, then prediction completeness is improved, but data processing complexity increases
Solution Approach 1:
The patent extracts and focuses on the most critical telemetry metrics related to storage, network, and compute functions that are directly associated with data unavailability and loss. The residual neural network is designed to process only these essential metrics rather than all possible telemetry data, reducing processing complexity while maintaining comprehensive prediction capability.
Solution Approach 2:
The system segments telemetry data by functional category (storage, network, compute) and processes each segment through dedicated analysis paths. This segmentation organizes the large volume of telemetry data into manageable groups, reducing processing complexity while maintaining complete coverage of all critical asset types.
Data Source
AI summary
A system, method, and computer-readable medium for performing a data center management and monitoring operation to predict event failures, comprising taking time stamped telemetry metrics of data asset clusters; sampling the time stamped telemetry metrics; separating the sampled time stamped telemetry metrics into groups based on metric functionality; providing data of the groups into a residual neural network to learn patterns of the three groups; and concatenating output of the residual neural network of the groups, wherein concatenated output is feed back to the residual neural network to learn relationships of the three groups.


