Alert Clustering for Data Center Incident Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data center administrators face challenges in understanding and prioritizing high-volume 'event storms' that generate numerous alerts in a short period, making it difficult to determine the source of issues and take appropriate remedial actions in real time, as existing operations management tools lack features to aid in this process.

Innovation Solution

An automated computer-implemented method and system that detects clusters of alerts based on start times and topological proximity within a data center, identifying incidents and determining high-priority alerts, which are then displayed in a graphical user interface for immediate action.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If automated clustering of alerts is implemented, then understanding and prioritization of event storms is improved, but system complexity increases

Engineering Contradiction:
Improveunderstanding of event stormsVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the complex task of event storm analysis into distinct functional modules: alert retrieval module, clustering module, incident detection module, and prioritization module. Each module handles a specific aspect of the analysis, making the overall system more manageable and maintainable while effectively reducing information loss in event storm understanding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an operations manager as an intermediary system between the data center infrastructure and system administrators. This intermediary automatically performs clustering and analysis of alerts, mediating the complex processing requirements and presenting simplified incident information to administrators, thereby improving understanding without requiring administrators to directly manage the complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If real-time analysis of event storms is performed, then response time is reduced, but computational resources are consumed

Engineering Contradiction:
Improveresponse timeVSAvoidcomputational resources
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary clustering of alerts based on topological proximity and temporal patterns before full incident analysis. By pre-organizing alerts into clusters using efficient algorithms, the system reduces the computational burden of subsequent real-time analysis, enabling faster response times without excessive resource consumption.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically adjusts analysis parameters such as clustering thresholds, time window sizes, and priority weighting based on system conditions and incident characteristics. This allows the system to optimize computational resource usage while maintaining real-time analysis capabilities, scaling the level of detail analyzed based on the severity and nature of detected incidents.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If clusters of alerts are analyzed to detect incidents, then accuracy of incident identification is improved, but processing time increases

Engineering Contradiction:
Improveaccuracy of incident identificationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies different analysis depths and clustering algorithms to different regions of the data center topology based on local characteristics. High-criticality components receive more detailed analysis with higher accuracy algorithms, while less critical components use faster, lighter-weight analysis methods. This local differentiation improves overall incident identification accuracy for critical systems without uniformly increasing processing time across the entire system.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12009965B2Methods and systems for discovering incidents through clustering of alert occurring in a data center
Publication Date: 2024.06.11 VMWARE INC
  • US12009965B2 patent drawing
  • US12009965B2 patent drawing
  • US12009965B2 patent drawing

AI summary

Automated computer-implemented methods and systems for discovering clusters of alerts triggered by abnormal events occurring with objects in a data center are described. In one aspect, alerts with start times in a sliding run-time window are retrieved from an alerts database. Each alert corresponds to a run-time event occurring with an object of the data center. Clusters of alerts in the sliding run-time window are detected based on the start times of the alerts and topological proximity of the objects. High priority alerts in the clusters of alerts are determined based on alert types. The events associated with discovered clusters of alerts and high priority alerts are displayed in a graphical user interface (“GUI”). Time evolution clustering of alerts and coverage evolution of alerts are over time based on the start times of the alerts and topological proximity of objects exhibiting abnormal behavior in the data center.