ML Log Analysis for Data Center Root Cause Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center management tools are ineffective in identifying the root cause of performance problems, leading to lengthy troubleshooting processes that are costly and error-prone, resulting in downtime and revenue loss.
Innovation Solution
An automated system using machine learning to analyze log messages and key performance indicators (KPIs) to determine the probable root cause of performance issues, reducing reliance on manual searches by teams of engineers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual searching through metrics and log messages is performed by teams of engineers, then root cause identification may be thorough, but the time required increases to days or weeks
Solution Approach 1:
The patent introduces an automated analysis system that acts as an intermediary between the complex data center environment and the engineers. This system automatically collects metrics, generates log messages, analyzes patterns, and presents findings, thereby eliminating the need for engineers to manually search through vast amounts of data while maintaining high accuracy in root cause identification
Solution Approach 2:
The patent replaces the mechanical manual process of engineers searching through metrics and logs with an automated computational system. The system uses algorithms to automatically collect, correlate, and analyze data from multiple sources, substituting human manual effort with automated processing that is both faster and equally thorough
2Productivity
If automated methods are implemented to identify root causes, then troubleshooting time is significantly reduced, but the complexity of the system increases
Solution Approach 1:
The automated analysis system is divided into distinct functional modules: a data collection module that gathers metrics and logs, an analysis module that processes the data using algorithms, and a reporting module that presents findings. This segmentation allows each component to be developed and maintained independently, managing overall system complexity while achieving high troubleshooting speed
Solution Approach 2:
The system is designed to autonomously collect data, analyze patterns, and generate reports without requiring external intervention. The automated nature of the system enables it to service itself in terms of data collection and analysis, reducing the operational complexity despite the sophisticated algorithms employed
3Loss of information
If comprehensive monitoring of all metrics and log messages is performed, then complete visibility into performance problems is achieved, but the amount of data to be analyzed increases significantly
Solution Approach 1:
The system extracts only the most relevant and actionable information from the vast volume of collected metrics and log messages. Rather than presenting all raw data, the analysis algorithms identify and extract key patterns, anomalies, and root cause indicators, maintaining complete visibility into performance problems while reducing the presented data volume to manageable levels
Data Source
AI summary
Automated, computer-implemented methods and systems for resolving performance problems with objects executing in a data center are described. The automated methods use machine learning to train a model that comprises rules defining relationships between probabilities of event types of in log messages and values of a key performance indictor (“KPI”) of the object over a historical time period. When a KPI violates a corresponding threshold, the rules are used to evaluate run time log messages that describe the probable root cause of the performance problem. An alert identifying the KPI threshold violation, and the log messages are displayed in a graphical user interface of an electronic display device.


