Automated Root Cause Identification in Data Center Performance Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center management tools are unable to timely identify the root causes of performance problems, requiring manual filtering by teams of engineers, which is time-consuming and costly, and often results in prolonged downtime and revenue loss.
Innovation Solution
An automated system using an operations management server that employs machine learning to construct models for identifying performance issues by analyzing historical and real-time data, providing immediate alerts and recommendations for resolving problems through a graphical user interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual filtering of metrics and log messages by teams of engineers is used to identify root causes, then measurement precision is improved, but loss of time increases significantly
Solution Approach 1:
An automated analysis system acts as an intermediary between the vast amounts of metrics/log messages and the engineering teams. The system automatically filters, correlates, and analyzes data to identify potential root causes, presenting refined results to engineers for validation. This intermediary processing dramatically reduces the time engineers spend manually sifting through data while maintaining high accuracy through automated pattern recognition and human expertise verification.
Solution Approach 2:
The system performs preliminary analysis of metrics and log messages before presenting them to engineering teams. By pre-filtering, pre-correlating, and pre-analyzing the data to identify potential root causes and their likelihood, the system prepares refined information in advance. This preliminary action reduces the subsequent manual effort required and accelerates the overall troubleshooting process while maintaining measurement precision.
2Loss of time
If automated methods are implemented to identify root causes quickly, then loss of time is reduced, but device complexity increases
Solution Approach 1:
The automated analysis system is designed to handle multiple types of data (metrics, log messages, performance data) and perform multiple functions (filtering, correlation, analysis, root cause identification) through a unified platform. This multi-functional approach consolidates what would otherwise require multiple separate tools and processes, reducing overall system complexity while enabling rapid automated troubleshooting across diverse data sources and problem types.
Solution Approach 2:
The system incorporates feedback mechanisms where engineering team validations and corrections of automated root cause analyses are fed back into the system to improve future automated detections. This learning feedback loop allows the system to become more accurate over time without requiring proportional increases in complexity, as the system adapts to organizational knowledge and patterns.
3Measurement precision
If teams of engineers manually search for root causes, then measurement precision is improved, but productivity decreases
Solution Approach 1:
The troubleshooting process is segmented into distinct phases: automated data collection and initial analysis, automated pattern recognition and root cause hypothesis generation, and human validation and decision-making. This segmentation allows high-volume processing to be handled automatically while reserving human expertise for critical validation steps, thereby increasing overall productivity without sacrificing measurement precision in the final root cause identification.
Solution Approach 2:
The automated system serves as an intermediary that processes and prepares data before human review, increasing the productivity of engineering teams by eliminating repetitive manual filtering tasks. The system handles the time-consuming preliminary analysis at scale, allowing engineers to focus their productivity on higher-value validation and decision-making activities while maintaining accurate root cause identification through the combined automated-human workflow.
Data Source
AI summary
Automated methods and systems for identifying and resolving performance problems of objects of a data center are described. The automated methods and systems construct a model for identifying objects of the datacenter that are experiencing performance problems based on baseline distributions of events of the objects in a historical time period and event distributions of events of the objects in a time window located outside the historical time period. A root causes and recommendations database is constructed for resolving performance problems based on remedial measures previously performed for resolving performance problems. The model is used to monitor the objects of data center for runtime performance problems. When a performance problem with an object is detected, the root causes and recommendations database is used to identify a root cause of the performance problem and generate a recommendation for resolving the performance problem in near real time.


