Data Center Behavioral Modeling for Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modeling a data center is computationally intensive due to its complex and dynamic nature, making it time-consuming and laborious to detect root causes of failures and maintain normal operations.
Innovation Solution
A behavioral modeling method using human knowledge to enhance machine learning algorithms, where human modelers decompose the data center into connected nodes, detect anomalies, and automatically recommend actions to maintain normal operations by recursively applying the model and updating the system model based on dynamic changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a machine learning algorithm is applied to model the entire data center system, then the model comprehensively captures system behavior, but the computational complexity and time required for modeling and anomaly detection increase significantly
Solution Approach 1:
The patent divides the data center system into hierarchical levels (data center level, rack level, component level) and further segments nodes into sub-nodes and simple components. This segmentation allows the machine learning algorithm to process smaller, manageable subsets of the system independently, reducing overall computational complexity while maintaining comprehensive anomaly detection capability across the entire system.
Solution Approach 2:
The patent introduces a hierarchical dimension to the anomaly detection process by organizing nodes across multiple levels (data center, rack, component). This dimensional transformation allows the system to detect anomalies at different granularities simultaneously, achieving comprehensive coverage without requiring the algorithm to process all millions of features at a single level, thus reducing modeling time.
2Device complexity
If the data center is decomposed into numerous nodes and components by human modelers, then the system structure becomes more manageable, but the device complexity and manual effort required increase
Solution Approach 1:
The patent creates a universal hierarchical template that can be applied across different data center configurations. The same segmentation framework (data center level, rack level, component level) and node decomposition approach can be reused regardless of the specific data center architecture, reducing the need for custom manual modeling efforts while maintaining organized system structure.
Solution Approach 2:
The patent uses template-based node definitions and reusable decomposition patterns that can be copied and adapted across different data center scenarios. Once the hierarchical structure and node types are defined for one data center, they can be replicated and modified for other data centers, significantly reducing the manual modeling effort required while maintaining consistent system organization.
3Measurement precision
If the behavioral model recursively processes each node and simple component, then root cause identification becomes more accurate, but the computational resources and processing time required increase
Solution Approach 1:
The recursive processing is applied segment by segment through the hierarchical levels rather than processing all nodes simultaneously. The system processes data center level nodes, then drills down to rack level sub-nodes, and finally to component level simple components only when anomalies are detected at higher levels. This segmented recursive approach maintains high root cause identification accuracy while significantly reducing computational resource usage compared to processing all nodes at once.
Solution Approach 2:
The patent applies recursive processing selectively rather than uniformly across all nodes. Full recursive analysis is performed only on nodes where anomalies are detected at higher hierarchical levels, while normal nodes are processed more efficiently. This partial application of recursive processing maintains accurate root cause identification for problematic areas while conserving computational resources in normal operating conditions.
Data Source
AI summary
A method generates a behavioral model of a data center when a machine learning algorithm is applied. A team of human modelers that partition the data center into a plurality of connected nodes is analyzed by a behavioral model. The behavioral model of the data center detects an anomaly in a system behavior center by recursively applying the behavioral model to each node and simple component. A compressed metric vector for the node is generated by reducing a dimension of an input metric vector. A root cause of a failure caused is determined by the anomaly and an action is automatically recommended to an operator to resolve a problem caused by the failure. The proactively actions are taken to keep the data center in a normal state based on the behavioral model using the machine learning algorithm.


