Operations Management Server for Data Center Root Cause Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data center management tools are ineffective in timely troubleshooting of performance issues, often requiring extensive manual effort from teams of engineers to identify root causes, leading to prolonged downtime and increased costs due to their inability to accurately and efficiently analyze vast amounts of metrics and log messages.

Innovation Solution

An automated system utilizing an operations management server that determines baseline and runtime distributions of events to identify performance problems, allowing for real-time monitoring and alerting of root causes through graphical user interfaces, significantly reducing the time and reliance on human teams.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If automated distribution-based detection is implemented, then problem identification speed improves, but system complexity increases

Engineering Contradiction:
Improvetime to identify root causeVSAvoidcomplexity of automated system
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by collecting historical event data and establishing baseline distributions before problems occur. The operations management server continuously monitors and stores normal operational patterns, enabling rapid root cause identification when anomalies occur without requiring complex real-time analysis during incidents.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary operations management server that acts as a mediator between raw metrics/logs and root cause identification. This server implements distribution-based detection algorithms, transforming complex data analysis into manageable statistical comparisons between baseline and runtime distributions, thereby reducing overall system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If manual troubleshooting by engineering teams is used, then system complexity remains low, but productivity decreases

Engineering Contradiction:
Improvetroubleshooting efficiencyVSAvoidcomplexity of management system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements self-service by automatically performing root cause identification without requiring engineering team intervention. The operations management server autonomously compares runtime event distributions against baseline distributions, identifies anomalies, and determines root causes, freeing engineering teams from manual troubleshooting tasks and significantly improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual troubleshooting with an automated computational system. Instead of engineers manually analyzing metrics and logs, the operations management server uses distribution-based detection algorithms to automatically identify root causes, substituting human analytical processes with automated statistical methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If traditional alerting mechanisms are used, then ease of operation is maintained, but measurement precision deteriorates

Engineering Contradiction:
Improveaccuracy of root cause identificationVSAvoidoperational simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system transitions from traditional single-threshold alerting to multi-dimensional distribution-based detection. Instead of comparing single metric values against fixed thresholds, the operations management server analyzes entire event distributions across multiple dimensions, comparing baseline distributions with runtime distributions to achieve more precise root cause identification while maintaining operational simplicity through automated GUI alerts.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11815989B2Automated methods and systems for identifying problems in data center objects
Publication Date: 2023.11.14 VMWARE INC
  • US11815989B2 patent drawing
  • US11815989B2 patent drawing
  • US11815989B2 patent drawing

AI summary

Automated methods and systems for identifying problems associated with objects of a data center are described. Automated methods and systems are performed by an operations management server. For each object, the server determines a baseline distribution from historical events that are associated with a normal operational state of an object. The server determines a runtime distribution of runtime events that are associated with the object and detected in a runtime window of the object. The management server monitors runtime performance of the object while the object is running in the datacenter. When a performance problem is detected, the management server determines a root cause of a performance problem based on the baseline distribution and the runtime distribution and displays an alert in a graphical user interface of a display.