Machine Learning System for Data Center Log Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data center management tools cannot effectively identify the root causes of performance problems, leading to lengthy troubleshooting processes that are costly and error-prone, resulting in downtime and revenue loss.

Innovation Solution

An automated system using machine learning to analyze log messages and key performance indicators (KPIs) to determine the probable root cause of performance issues, enabling rapid diagnosis and remediation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual troubleshooting by software engineers is used to identify root causes, then diagnostic accuracy can be maintained through human analysis, but the time required increases to days or weeks

Engineering Contradiction:
Improveroot cause identification accuracyVSAvoidtroubleshooting time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical system of manual human analysis with an automated machine learning system that processes log messages and metrics. The ML model automatically identifies root causes by analyzing patterns in operational data, substituting human engineers' manual troubleshooting with an automated computational system that operates faster and at scale.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service troubleshooting by automatically diagnosing root causes without requiring human intervention. The ML-powered operations manager independently analyzes performance data, identifies problems, and provides diagnostic information, allowing the system to serve itself in the troubleshooting process rather than relying on external human expertise.

Inventive Principle:
Principle #25Self-service

2Reliability

If teams of software engineers are employed to search for root causes, then comprehensive analysis can be performed, but operational costs increase significantly

Engineering Contradiction:
Improveproblem identification reliabilityVSAvoidoperational cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent replaces expensive human engineering teams with an automated machine learning system. The ML model performs comprehensive analysis of operational data at a fraction of the cost of human salaries, benefits, and training, while maintaining or improving diagnostic reliability through consistent application of learned patterns across all data points.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the parameter of resource allocation from human capital (expensive, limited availability) to computational resources (cheaper, scalable). By transforming the troubleshooting function into an automated computational process, the system achieves comprehensive analysis at lower operational costs while maintaining reliability through repeated validation against training data.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated machine learning systems are used to identify root causes, then troubleshooting time is reduced to minutes, but system complexity increases

Engineering Contradiction:
Improvetroubleshooting speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-training machine learning models on historical operational data before deployment. The ML models are trained in advance to recognize patterns and root causes, so when production issues occur, the pre-trained models can immediately analyze new data and identify problems in minutes without requiring complex real-time computation or human intervention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary layer between raw operational data and human operators. The ML system acts as a mediator that processes complex log messages and metrics, transforming them into simplified root cause diagnoses. This intermediary handles the complexity internally while presenting simplified results to users, masking the underlying system complexity from end users.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Device complexity

If manual analysis of log messages and metrics is performed, then system complexity remains low, but diagnostic precision deteriorates due to human error and time constraints

Engineering Contradiction:
Improvesystem complexityVSAvoidroot cause detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent replaces manual human analysis with automated machine learning processing. The ML system eliminates human errors such as fatigue, bias, and inconsistency by applying the same analytical algorithms uniformly to all data. The automated system processes log messages and metrics with consistent precision, removing the variability inherent in human analysis while maintaining manageable system complexity through established ML techniques.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12056002B2Methods and systems for using machine learning to resolve performance problems with objects of a data center
Publication Date: 2024.08.06 VMWARE INC
  • US12056002B2 patent drawing
  • US12056002B2 patent drawing
  • US12056002B2 patent drawing

AI summary

Automated computer-implemented methods and systems for resolving performance problems with objects executing in a data center are described. The automated methods use machine learning to obtain rules defining relationships between probabilities of event types of in log messages and performance problems identified by a key performance indictor (“KPI”) of the object. When a KPI violates a corresponding threshold, the rules are used to evaluate run time log messages that describe the probable root cause of the performance problem. An alert identifying the KPI threshold violation, and the log messages are displayed in a graphical user interface of an electronic display device.