Root Cause Analysis for Non-Deterministic Datacenter Anomalies
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional performance analysis methods in datacenters are not accurate and precise in identifying the root cause of non-deterministic performance anomalies, making it difficult to manage complex and dynamic heterogeneous infrastructures.
Innovation Solution
A method that collects resource consumption data from compute, network, and service components, generates digital signatures, and compares them with pre-tabulated signatures to identify the root cause of performance degradation, using a graph-based approach to filter and analyze data for root cause analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional performance analysis methods are used in datacenters, then the system can operate with simple analysis processes, but the accuracy and precision in identifying root cause of non-deterministic performance anomalies deteriorates
Solution Approach 1:
The patent segments the complex analysis task into distinct phases: data collection from multiple sources, digital signature generation from collected data, signature comparison against known anomaly patterns, and root cause identification. This segmentation transforms an overwhelming complex analysis into manageable discrete steps, improving measurement precision without requiring the entire system to be simultaneously complex.
Solution Approach 2:
The patent introduces digital signatures as an intermediary representation between raw performance data and root cause identification. Instead of directly analyzing complex heterogeneous data from multiple sources, the system transforms data into signature form that can be systematically compared against known anomaly patterns, thereby improving accuracy while managing complexity through this intermediate abstraction layer.
2Loss of information
If comprehensive resource consumption data is collected from all components, then the completeness of analysis information improves, but the complexity of data processing and noise increases
Solution Approach 1:
The patent extracts only the essential characteristics from comprehensive resource consumption data by generating digital signatures that capture the fundamental performance patterns. This extraction process removes redundant and noisy information while retaining the critical features needed for anomaly detection, thus maintaining information completeness while reducing processing complexity.
Solution Approach 2:
The patent transforms raw resource consumption data into a different parameter space through signature generation. By changing the parameters from detailed resource metrics to signature representations, the system maintains the essential information content while reducing the dimensionality and complexity of subsequent data processing operations.
3Measurement precision
If detailed granular data is collected from numerous components, then the precision of performance monitoring improves, but the difficulty of filtering relevant data increases
Solution Approach 1:
The patent creates signature copies that represent the essential characteristics of detailed granular data. Instead of directly processing the original voluminous granular data, the system works with these signature copies that preserve the critical performance patterns, thereby maintaining monitoring precision while significantly reducing the difficulty of data filtering and analysis.
Data Source
AI summary
Some embodiments of the invention provide methods for performing root cause analysis for non-deterministic anomalies in a datacenter. For instance, the method of some embodiments identifies a root cause for degradation in performance of one or more components in a network of the datacenter. This method collects and generates resource consumption data regarding resources consumed by a set of components in this network. The method performs a first analysis on the collected and/or generated data to identify an instance in time when one or more components, while still operational, are possibly suffering from performance degradation. The method then performs a second analysis on the collected and/or generated data associated with the identified time instance to identify a root cause of a performance degradation of at least one component in the network.


