ML Metric Selection for Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cloud computing environments, manually specifying metrics and statistics for anomaly detection is time-consuming and resource-intensive, requiring significant expertise and potentially leading to missed anomalies or false positives due to the complexity of managing large-scale computing resources with varying workloads.
Innovation Solution
A machine learning model is trained using metadata from existing applications to automatically select relevant metrics and statistics for anomaly detection, predicting anomaly analysis relevance scores without requiring manual input from application owners, thereby reducing the workload and improving detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual specification of metrics and statistics is used for anomaly detection, then detection accuracy can be maintained through expert knowledge, but time consumption and resource requirements increase significantly
Solution Approach 1:
The system enables self-service by automatically selecting metrics and statistics for anomaly detection without requiring manual expert input. The machine learning model autonomously analyzes application behavior patterns and identifies relevant metrics, allowing the system to serve itself rather than relying on external expert configuration.
Solution Approach 2:
The patent replaces the mechanical process of manual metric specification with an automated machine learning-based selection process. Instead of experts manually analyzing and selecting metrics, the system uses computational algorithms to automatically identify and select the most relevant metrics for anomaly detection.
2Reliability
If manual specification of metrics is performed, then relevant anomalies can be detected, but resource intensity and complexity increase
Solution Approach 1:
The complex manual process of metric specification is replaced with an automated machine learning system that handles the complexity internally. The machine learning model absorbs the complexity of metric selection, presenting a simplified interface to users while maintaining high detection reliability through algorithmic analysis.
Solution Approach 2:
The machine learning model acts as an intermediary between the raw application data and the anomaly detection process. It mediates the complex task of metric selection by automatically identifying relevant metrics and statistics, thereby reducing the complexity exposed to users while maintaining detection reliability.
3Productivity
If automated metric selection is implemented, then time and resources are reduced, but risk of false positives or missed anomalies increases
Solution Approach 1:
The machine learning model performs preliminary analysis of application behavior patterns before anomaly detection begins. By pre-learning normal behavior characteristics and automatically selecting relevant metrics during a preparation phase, the system establishes a foundation for accurate detection that reduces false positives while maintaining high productivity.
Solution Approach 2:
The system incorporates feedback mechanisms where the machine learning model continuously learns from detected anomalies and adjusts its metric selection accordingly. This feedback loop improves detection accuracy over time by refining the understanding of what constitutes a true anomaly versus a false positive, thereby maintaining reliability while preserving automation benefits.
Data Source
AI summary
A plurality of metrics records, including some records indicating metrics for which anomaly analysis has been performed, is obtained. Using a training data set which includes the metrics records, a machine learning model is trained to predict an anomaly analysis relevance score for an input record which indicates a metric name. Collection of a particular metric of an application is initiated based at least in part on an anomaly analysis relevance score obtained for the particular metric using a trained version of the model.


