Alert Clustering for Data Center Incident Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data center administrators face challenges in understanding and prioritizing high-volume 'event storms' that generate numerous alerts in a short period, making it difficult to determine the source of issues and take appropriate remedial actions in real time, as existing operations management tools lack features to aid in this process.
Innovation Solution
An automated computer-implemented method and system that detects clusters of alerts based on start times and topological proximity within a data center, identifying incidents and determining high-priority alerts, which are then displayed in a graphical user interface for immediate action.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If automated clustering of alerts is implemented, then understanding and prioritization of event storms is improved, but system complexity increases
Solution Approach 1:
The patent segments the complex task of event storm analysis into distinct functional modules: alert retrieval module, clustering module, incident detection module, and prioritization module. Each module handles a specific aspect of the analysis, making the overall system more manageable and maintainable while effectively reducing information loss in event storm understanding.
Solution Approach 2:
The patent introduces an operations manager as an intermediary system between the data center infrastructure and system administrators. This intermediary automatically performs clustering and analysis of alerts, mediating the complex processing requirements and presenting simplified incident information to administrators, thereby improving understanding without requiring administrators to directly manage the complexity.
2Loss of time
If real-time analysis of event storms is performed, then response time is reduced, but computational resources are consumed
Solution Approach 1:
The patent performs preliminary clustering of alerts based on topological proximity and temporal patterns before full incident analysis. By pre-organizing alerts into clusters using efficient algorithms, the system reduces the computational burden of subsequent real-time analysis, enabling faster response times without excessive resource consumption.
Solution Approach 2:
The patent dynamically adjusts analysis parameters such as clustering thresholds, time window sizes, and priority weighting based on system conditions and incident characteristics. This allows the system to optimize computational resource usage while maintaining real-time analysis capabilities, scaling the level of detail analyzed based on the severity and nature of detected incidents.
3Measurement precision
If clusters of alerts are analyzed to detect incidents, then accuracy of incident identification is improved, but processing time increases
Solution Approach 1:
The patent applies different analysis depths and clustering algorithms to different regions of the data center topology based on local characteristics. High-criticality components receive more detailed analysis with higher accuracy algorithms, while less critical components use faster, lighter-weight analysis methods. This local differentiation improves overall incident identification accuracy for critical systems without uniformly increasing processing time across the entire system.
Data Source
AI summary
Automated computer-implemented methods and systems for discovering clusters of alerts triggered by abnormal events occurring with objects in a data center are described. In one aspect, alerts with start times in a sliding run-time window are retrieved from an alerts database. Each alert corresponds to a run-time event occurring with an object of the data center. Clusters of alerts in the sliding run-time window are detected based on the start times of the alerts and topological proximity of the objects. High priority alerts in the clusters of alerts are determined based on alert types. The events associated with discovered clusters of alerts and high priority alerts are displayed in a graphical user interface (“GUI”). Time evolution clustering of alerts and coverage evolution of alerts are over time based on the start times of the alerts and topological proximity of objects exhibiting abnormal behavior in the data center.


