Anomaly Detection Using Representative Metrics Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current anomaly detection algorithms for large datasets in networked computing environments require extensive processing resources, leading to inefficiencies and increased complexity, especially when analyzing metrics data across multiple dimensions.
Innovation Solution
The approach involves selecting representative subsets of metrics by grouping similar data into clusters, determining principal component datasets, and performing anomaly detection on these subsets, reducing the need to analyze all metrics data, thereby optimizing processing resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If anomaly detection algorithms analyze all metrics data generated by reporting tools, then detection accuracy is improved, but processing resources and complexity increase significantly
Solution Approach 1:
The patent segments the large metrics dataset into multiple smaller partitions, each processed by a separate anomaly detection algorithm instance. This division reduces the computational burden on each processing unit while collectively maintaining comprehensive coverage of all metrics data, thus resolving the contradiction between detection accuracy and processing complexity.
Solution Approach 2:
The patent applies anomaly detection algorithms to multiple partitions of the metrics data rather than requiring all algorithms to process the complete dataset simultaneously. This partial action approach ensures that anomalies are detected across the entire dataset through aggregated results from partial processing, reducing overall processing complexity while preserving detection accuracy.
2Measurement precision
If anomaly detection algorithms analyze all metrics data across multiple dimensions, then detection accuracy is improved, but processing time increases
Solution Approach 1:
The patent segments both the metrics data and the processing timeline by dividing data into partitions and executing anomaly detection algorithms in parallel across multiple processing units. This segmentation enables simultaneous processing of different data portions, significantly reducing total processing time while maintaining comprehensive anomaly detection across all dimensions.
Solution Approach 2:
The patent performs anomaly detection on multiple partitions of the metrics data in parallel, where each processing unit handles a subset of the total data. The aggregation of results from these partial processing actions provides comprehensive anomaly detection coverage across all dimensions without requiring sequential processing of the entire dataset, thus reducing processing time while preserving detection accuracy.
3Measurement precision
If anomaly detection is performed on large datasets with multiple dimensions, then detection accuracy is improved, but memory requirements increase
Solution Approach 1:
The patent segments the large metrics dataset into multiple smaller partitions that can be loaded into memory simultaneously by multiple processing units. Each partition requires only a fraction of the total memory resources, enabling parallel processing of dimensional data without requiring the entire dataset to reside in memory at once, thus reducing overall memory requirements while maintaining detection accuracy.
Solution Approach 2:
The patent processes multiple partitions of the metrics data in parallel, where each processing unit loads and analyzes only its assigned subset of data in memory. The collective results from these partial processing actions provide comprehensive anomaly detection across all dimensions without requiring the full dataset to be loaded into memory simultaneously, effectively reducing memory resource requirements while preserving detection accuracy.
Data Source
AI summary
Certain embodiments involve selecting metrics that are representative of large metrics datasets and that are usable for efficiently performing anomaly detection. For example, metrics datasets are grouped into clusters based on, for each of the clusters, a similarity of data values in a respective pair of datasets from the metrics datasets. Principal component datasets are determined for the clusters. A principal component dataset for a cluster includes a linear combination of a subset of metrics datasets included in the cluster. Each representative metric is selected based on the metrics dataset having a highest contribution to a principal component dataset in the cluster into which the metrics dataset is grouped. An anomaly detection is executed in a manner that is restricted to a subset of the metrics datasets corresponding to the representative metrics.


