Event Clustering for Anomaly Detection Training Samples
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current monitoring systems in large computing environments face challenges in accurately detecting anomalies due to the vast number of metrics and the need for manual rule-setting, which is tedious and prone to errors, especially when dealing with influencing events that have a predictable impact on metrics.
Innovation Solution
A method and system for generating samples for an anomaly detection system by grouping events with similar influence patterns, determining impact scores, and clustering these groups to create training samples that account for the influence of events on target metrics, thereby improving the accuracy of anomaly detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual rule-setting is used for anomaly detection, then the system can detect anomalies based on administrator expertise, but the process becomes tedious and error-prone
Solution Approach 1:
The system automatically generates anomaly detection rules by analyzing historical metric data and identifying patterns, eliminating the need for manual rule configuration by administrators. The anomaly detection model self-learns from data and autonomously creates detection logic.
Solution Approach 2:
The patent replaces manual mechanical rule-setting processes with an automated machine learning system that uses algorithms to analyze data and generate detection rules, substituting human cognitive work with computational processes.
2Reliability
If administrators manually monitor multiple metrics, then they can use experience to identify problems, but the number of visually trackable metrics is limited
Solution Approach 1:
The system replaces human visual monitoring capability with automated computational analysis that can process and analyze thousands of metrics simultaneously, far exceeding human visual tracking limits while maintaining expert-level detection capability.
Solution Approach 2:
The anomaly detection model serves multiple functions: it monitors numerous metrics across different systems, identifies various types of anomalies, and adapts to different metric patterns, making the system highly versatile and scalable.
3Productivity
If automated processes evaluate every metric, then comprehensive coverage is achieved, but the system lacks administrator insight and experience
Solution Approach 1:
The patent replaces automated rule-based evaluation with a machine learning model that learns complex patterns and relationships from historical data, capturing nuanced insights that simple automated rules cannot detect.
Solution Approach 2:
The system transforms static automated evaluation rules into dynamic, adaptive detection parameters that evolve by learning from historical metric data, allowing the model to improve its detection capability over time while processing large volumes of metrics.
4Reliability
If thresholds are set for multiple metrics, then anomaly detection coverage increases, but determining correct thresholds becomes difficult and time-consuming
Solution Approach 1:
The system automatically determines optimal detection thresholds by analyzing historical metric data and learning normal behavior patterns, eliminating the time-consuming manual threshold configuration process while maintaining comprehensive detection coverage.
Solution Approach 2:
The model performs preliminary analysis of historical data to establish baseline behavior patterns and detect deviations, proactively identifying anomalies before they become critical issues and reducing the need for manual threshold adjustments.
Data Source
AI summary
A method for generating samples for an anomaly detection system includes receiving events that occurred during a time series of a target metric. Each respective event includes an event attribute characterizing the respective event. The method includes generating a set of event groups for the events. Each event shares a respective attribute with one or more other events of the respective event group. For each respective event group of the set of event groups, the method includes determining an influence pattern that identifies an influence of the respective event group on the target metric. The method includes clustering the set of event groups into event clusters based on a respective influence pattern of each respective event group. Each event cluster includes one or more event groups that share a similar influence pattern. The method includes generating training samples for an anomaly detection system based on a respective event cluster.


