Anomaly Detection System Using Temporal Correlation Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual detection of anomalous behavior in large and complex computing systems is not scalable, as system administrators face challenges in monitoring multiple resources and identifying interdependent anomalies, leading to inefficiencies and false positives with existing statistical anomaly detectors.
Innovation Solution
The system automatically detects, summarizes, and responds to anomalies by correlating anomalies across various entities, using score-based summarization and correlation-based methods to identify linked issues, allowing for timely and efficient corrective actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If system administrators manually monitor computing applications to detect anomalies, then detection accuracy may be maintained, but scalability deteriorates and time consumption increases
Solution Approach 1:
The system enables self-service anomaly detection through automated machine learning models that continuously monitor computing applications without human intervention. The unsupervised learning algorithms automatically identify anomalies by comparing current system state against learned baselines, eliminating the need for manual monitoring while maintaining detection accuracy across large-scale distributed systems.
Solution Approach 2:
The patent replaces manual administrative actions with automated computational systems. Machine learning models substitute human analysts by processing telemetry data, detecting patterns, and identifying anomalies algorithmically. This substitution enables scalable monitoring of numerous applications simultaneously while preserving detection capabilities through sophisticated statistical analysis.
2Reliability
If statistical anomaly detectors are used to monitor multiple resources, then detection coverage is improved, but false positives increase and operational complexity worsens
Solution Approach 1:
The system segments anomaly detection into hierarchical levels: individual metric detection, time-series pattern recognition, and contextual correlation analysis. By dividing the detection process into discrete analytical stages, the system reduces false positives through progressive filtering while maintaining comprehensive coverage across multiple computing resources.
Solution Approach 2:
The patent employs composite detection methodologies that combine multiple statistical techniques and machine learning algorithms. By integrating diverse analytical approaches (e.g., baseline comparison, pattern recognition, correlation analysis), the system achieves robust anomaly detection with reduced false positives through mutual validation of detection signals.
3Measurement precision
If comprehensive monitoring of interdependent entities is implemented, then root cause identification accuracy is improved, but system complexity and computational resources increase
Solution Approach 1:
The system introduces intermediary components including telemetry collectors, normalization layers, and correlation engines that mediate between raw data sources and analysis functions. These intermediaries simplify the monitoring architecture by standardizing data formats and pre-processing relationships, enabling accurate root cause identification without proportionally increasing overall system complexity.
Solution Approach 2:
The patent analyzes interdependent entities across multiple dimensions including temporal relationships, causal pathways, and contextual metadata. By examining anomalies through additional analytical dimensions rather than simple linear monitoring, the system achieves superior root cause identification accuracy while managing complexity through structured multi-dimensional analysis frameworks.
Data Source
AI summary
Techniques are disclosed for summarizing, diagnosing, and correcting the cause of anomalous behavior in computing systems. In some embodiments, a system identifies a plurality of time series that track different metrics over time for a set of one or more computing resources. The system detects a first set of anomalies in a first time series that tracks a first metric and assigns a different respective range of time to each anomaly. The system determines whether the respective range of time assigned to an anomaly overlaps with timestamps or ranges of time associated with anomalies from one or more other time series. The system generates at least one cluster that groups metrics based on how many anomalies have respective ranges of time and/or timestamps that overlap. The system may preform, based on the cluster, one or more automated actions for diagnosing or correcting a cause of anomalous behavior.


