Machine Learning System for Data Center Log Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center management tools cannot effectively identify the root causes of performance problems, leading to lengthy troubleshooting processes that are costly and error-prone, resulting in downtime and revenue loss.
Innovation Solution
An automated system using machine learning to analyze log messages and key performance indicators (KPIs) to determine the probable root cause of performance issues, enabling rapid diagnosis and remediation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual troubleshooting by software engineers is used to identify root causes, then diagnostic accuracy can be maintained through human analysis, but the time required increases to days or weeks
Solution Approach 1:
The patent replaces the mechanical system of manual human analysis with an automated machine learning system that processes log messages and metrics. The ML model automatically identifies root causes by analyzing patterns in operational data, substituting human engineers' manual troubleshooting with an automated computational system that operates faster and at scale.
Solution Approach 2:
The system enables self-service troubleshooting by automatically diagnosing root causes without requiring human intervention. The ML-powered operations manager independently analyzes performance data, identifies problems, and provides diagnostic information, allowing the system to serve itself in the troubleshooting process rather than relying on external human expertise.
2Reliability
If teams of software engineers are employed to search for root causes, then comprehensive analysis can be performed, but operational costs increase significantly
Solution Approach 1:
The patent replaces expensive human engineering teams with an automated machine learning system. The ML model performs comprehensive analysis of operational data at a fraction of the cost of human salaries, benefits, and training, while maintaining or improving diagnostic reliability through consistent application of learned patterns across all data points.
Solution Approach 2:
The system changes the parameter of resource allocation from human capital (expensive, limited availability) to computational resources (cheaper, scalable). By transforming the troubleshooting function into an automated computational process, the system achieves comprehensive analysis at lower operational costs while maintaining reliability through repeated validation against training data.
3Productivity
If automated machine learning systems are used to identify root causes, then troubleshooting time is reduced to minutes, but system complexity increases
Solution Approach 1:
The system performs preliminary action by pre-training machine learning models on historical operational data before deployment. The ML models are trained in advance to recognize patterns and root causes, so when production issues occur, the pre-trained models can immediately analyze new data and identify problems in minutes without requiring complex real-time computation or human intervention.
Solution Approach 2:
The patent introduces an intermediary layer between raw operational data and human operators. The ML system acts as a mediator that processes complex log messages and metrics, transforming them into simplified root cause diagnoses. This intermediary handles the complexity internally while presenting simplified results to users, masking the underlying system complexity from end users.
4Device complexity
If manual analysis of log messages and metrics is performed, then system complexity remains low, but diagnostic precision deteriorates due to human error and time constraints
Solution Approach 1:
The patent replaces manual human analysis with automated machine learning processing. The ML system eliminates human errors such as fatigue, bias, and inconsistency by applying the same analytical algorithms uniformly to all data. The automated system processes log messages and metrics with consistent precision, removing the variability inherent in human analysis while maintaining manageable system complexity through established ML techniques.
Data Source
AI summary
Automated computer-implemented methods and systems for resolving performance problems with objects executing in a data center are described. The automated methods use machine learning to obtain rules defining relationships between probabilities of event types of in log messages and performance problems identified by a key performance indictor (“KPI”) of the object. When a KPI violates a corresponding threshold, the rules are used to evaluate run time log messages that describe the probable root cause of the performance problem. An alert identifying the KPI threshold violation, and the log messages are displayed in a graphical user interface of an electronic display device.


