Operations Management Server for Data Center Root Cause Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center management tools are ineffective in timely troubleshooting of performance issues, often requiring extensive manual effort from teams of engineers to identify root causes, leading to prolonged downtime and increased costs due to their inability to accurately and efficiently analyze vast amounts of metrics and log messages.
Innovation Solution
An automated system utilizing an operations management server that determines baseline and runtime distributions of events to identify performance problems, allowing for real-time monitoring and alerting of root causes through graphical user interfaces, significantly reducing the time and reliance on human teams.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If automated distribution-based detection is implemented, then problem identification speed improves, but system complexity increases
Solution Approach 1:
The system performs preliminary actions by collecting historical event data and establishing baseline distributions before problems occur. The operations management server continuously monitors and stores normal operational patterns, enabling rapid root cause identification when anomalies occur without requiring complex real-time analysis during incidents.
Solution Approach 2:
The patent introduces an intermediary operations management server that acts as a mediator between raw metrics/logs and root cause identification. This server implements distribution-based detection algorithms, transforming complex data analysis into manageable statistical comparisons between baseline and runtime distributions, thereby reducing overall system complexity.
2Productivity
If manual troubleshooting by engineering teams is used, then system complexity remains low, but productivity decreases
Solution Approach 1:
The system implements self-service by automatically performing root cause identification without requiring engineering team intervention. The operations management server autonomously compares runtime event distributions against baseline distributions, identifies anomalies, and determines root causes, freeing engineering teams from manual troubleshooting tasks and significantly improving productivity.
Solution Approach 2:
The patent replaces the mechanical system of manual troubleshooting with an automated computational system. Instead of engineers manually analyzing metrics and logs, the operations management server uses distribution-based detection algorithms to automatically identify root causes, substituting human analytical processes with automated statistical methods.
3Measurement precision
If traditional alerting mechanisms are used, then ease of operation is maintained, but measurement precision deteriorates
Solution Approach 1:
The system transitions from traditional single-threshold alerting to multi-dimensional distribution-based detection. Instead of comparing single metric values against fixed thresholds, the operations management server analyzes entire event distributions across multiple dimensions, comparing baseline distributions with runtime distributions to achieve more precise root cause identification while maintaining operational simplicity through automated GUI alerts.
Data Source
AI summary
Automated methods and systems for identifying problems associated with objects of a data center are described. Automated methods and systems are performed by an operations management server. For each object, the server determines a baseline distribution from historical events that are associated with a normal operational state of an object. The server determines a runtime distribution of runtime events that are associated with the object and detected in a runtime window of the object. The management server monitors runtime performance of the object while the object is running in the datacenter. When a performance problem is detected, the management server determines a root cause of a performance problem based on the baseline distribution and the runtime distribution and displays an alert in a graphical user interface of a display.


