Correlated-Extreme Behavior Detection in Data Center Resource Consumers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center management tools rely on historical metric data to set thresholds, which may not accurately identify current resource congestion, leading to inefficiencies in identifying consumers causing alerts and troubleshooting.
Innovation Solution
The method involves analyzing current metric data to identify consumers exhibiting correlated-extreme behavior with their resource providers, using correlation coefficients, data tails, and probability density functions to pinpoint consumers contributing to alerts and extreme behavior.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If historical metric data is used to set thresholds, then threshold calculation is simplified, but accuracy in identifying current resource congestion deteriorates
Solution Approach 1:
The system performs preliminary analysis of historical metric data to establish baseline patterns and relationships between providers and consumers before congestion occurs. This preliminary action enables the system to quickly identify correlated-extreme behavior when congestion happens, resolving the contradiction by preparing advance knowledge that improves accuracy without increasing real-time calculation complexity
Solution Approach 2:
The patent introduces an intermediary analysis layer that examines relationships between multiple data sets (provider metric data and consumer metric data) to identify correlated-extreme behavior. This intermediary approach bridges the gap between simple historical thresholds and complex real-time analysis, improving congestion identification accuracy while maintaining manageable system complexity
2Measurement precision
If all consumers are monitored to identify root cause, then comprehensive analysis is achieved, but troubleshooting time increases
Solution Approach 1:
The system extracts and identifies only the specific consumers exhibiting correlated-extreme behavior with the provider, separating these problematic consumers from the normal population. This extraction approach enables focused troubleshooting on only the relevant consumers rather than analyzing all consumers, achieving complete root cause identification while minimizing troubleshooting time
Solution Approach 2:
The patent metaphorically applies 'color changes' by identifying and highlighting consumers with correlated-extreme behavior through distinct metric patterns. This allows administrators to quickly distinguish problematic consumers from normal ones, achieving comprehensive root cause analysis while reducing the time needed to locate issues among numerous consumers
3Device complexity
If fixed thresholds are used for alert generation, then alert generation is simple, but reliability in identifying current congestion deteriorates
Solution Approach 1:
The system transitions from static fixed thresholds to dynamic threshold generation based on current metric data and identified correlated-extreme behavior. This dynamic approach allows thresholds to adapt to changing conditions, improving congestion detection reliability while maintaining relatively simple alert generation through automated pattern recognition rather than manual threshold management
Data Source
AI summary
Methods and systems that identify objects of a data center that exhibit correlated-extreme behavior are described. The objects may be, but are not limited to, virtual machines (“VMs”), containers, server computers, clusters of server computers, and the data center itself. Metric data is collected for the various objects and the methods identify the objects that exhibit correlated-extreme behavior. In particular, the methods and systems narrow a search for correlated-extreme behavior of consumers of computational resources of a data center when a provider of the computational resources exhibits unexpected or extreme behavior.


