Predictive Maintenance for Distributed Data Centers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large-scale and distributed data centers face challenges in predictive maintenance due to the complexity of monitoring and analyzing vast amounts of data, making it difficult for human actors to identify root causes of issues and leading to error-prone maintenance processes.
Innovation Solution
The implementation of a predictive maintenance system using data science and machine learning (ML) to perform condition-based monitoring, where 'normal' system behavior is learned, and anomalies are detected, allowing for proactive maintenance by developing ML models specific to host groups or cliques, reducing the number of models needed while increasing their specificity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human actors manually monitor and analyze data center system stacks and logs, then they can identify root causes of issues, but the process becomes error-prone and inefficient due to the vast amount of data
Solution Approach 1:
The patent replaces manual human analysis of system stacks and logs with automated machine learning models. The ML models automatically ingest, parse, and analyze telemetry data, replacing the mechanical process of human actors manually examining vast amounts of data, thereby eliminating human error while maintaining high productivity
Solution Approach 2:
The system enables self-service through automated anomaly detection and root cause analysis. The ML models autonomously monitor system behavior, detect deviations from normal patterns, and identify root causes without requiring human intervention, allowing the system to serve itself in terms of predictive maintenance
2Reliability
If predictive maintenance is implemented in large-scale distributed data centers, then system reliability can be improved, but the complexity of monitoring and analyzing vast amounts of data increases
Solution Approach 1:
The patent segments the monitoring system into modular components: data collection agents distributed across hosts, centralized telemetry processing, and separate ML model modules for anomaly detection and root cause analysis. This segmentation allows the system to handle large-scale data centers while managing complexity through divided responsibilities and independent, interchangeable components
Solution Approach 2:
The patent introduces intermediary components including telemetry processing layers that aggregate and normalize data before ML analysis, and ML models that act as intermediaries between raw data and actionable insights. These intermediaries simplify the overall system by breaking down the complex task of predictive maintenance into manageable stages
3Measurement precision
If machine learning models are developed for each individual host, then model specificity increases, but the number of models needed increases significantly
Solution Approach 1:
The patent creates universal ML models that can be applied across multiple hosts with similar characteristics. Instead of building individual models for each host, the system develops models that work universally across host groups, reducing the total number of models while maintaining effectiveness through group-based generalization
4Loss of time
If condition-based monitoring is performed to detect anomalies, then maintenance costs and downtime can be reduced, but the system requires learning normal behavior patterns which increases initial complexity
Solution Approach 1:
The patent implements preliminary action by training ML models on historical telemetry data to learn normal behavior patterns before deployment. This preliminary training phase allows the models to establish baseline expectations of system behavior, enabling them to detect anomalies more effectively once deployed in production environments
Data Source
AI summary
Predictive maintenance can be achieved by fetching log data from a monitoring application monitoring one or more hosts of a data center site. Based on the log data, a clique comprising a subset of the one or more hosts exhibiting sufficiently similar behavior is created. The log data may also be used to train a machine learning model configured to predict a need for predictive maintenance/existence of a predictive maintenance state for any of the subset of the one or more hosts of the clique. The machine learning model is trained with the log data, and the machine learning model is operationalized to predict the existence of anomalous data in further log data collected from the monitoring application. The existence of anomalous data reflects a need for predictive maintenance.


