Confidence-Controlled Sampling for Distributed System Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributed computing systems face challenges in efficiently analyzing and detecting anomalous behavior due to the high volume and frequency of metric data and event messages, leading to increased storage costs and processing delays, which hinder the detection of anomalies and characterization of behavior patterns.
Innovation Solution
Implementing confidence-controlled sampling to randomly select a small number of data points with a selected confidence level, allowing for the determination of data characteristics, identification of periodic patterns, and comparison of behavior between sources, thereby speeding up data characterization and anomaly detection without compromising accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-frequency monitoring data is collected at sub-second frequency, then detection accuracy of anomalous behavior is improved, but data storage cost and processing burden increase significantly
Solution Approach 1:
The patent segments the monitoring data into different categories (metric data, event messages, logs) and applies different sampling strategies to each type. Confidence-controlled sampling divides the data processing into manageable segments that can be analyzed independently, reducing the overall processing burden while maintaining detection accuracy.
Solution Approach 2:
The patent implements partial action by sampling only a confidence-controlled portion of the data rather than processing all high-frequency data. This allows the system to achieve sufficient anomaly detection accuracy with a subset of data, reducing storage and processing requirements while maintaining effective monitoring.
2Reliability
If all monitoring data is processed to ensure accurate anomaly detection, then detection reliability is improved, but processing time increases drastically
Solution Approach 1:
The patent applies preliminary action by pre-calculating confidence intervals and determining sampling rates before data analysis. This allows the system to quickly process data according to pre-determined parameters without performing complex calculations on the entire dataset, significantly reducing processing time while maintaining reliable detection.
Solution Approach 2:
The patent changes the processing parameter from analyzing all data points to analyzing a confidence-controlled sample size. By adjusting the confidence level parameter, the system can balance between detection reliability and processing speed, achieving reliable anomaly detection with reduced processing time.
3Measurement precision
If high-frequency data collection is implemented, then behavior pattern characterization accuracy is improved, but system resource consumption exceeds limits
Solution Approach 1:
The patent extracts only the essential features and patterns from the high-frequency data through confidence-controlled sampling. By taking out and analyzing only the representative data points needed for behavior pattern characterization, the system maintains accuracy while reducing CPU and memory consumption to acceptable levels.
Data Source
AI summary
Methods and systems of automatic confidence-controlled sampling to analyze, detect anomalies and problems in monitoring data and event messages generated by sources of a distributed computing system are described. A source can be virtual or physical object of the distributed computing system, a resource of the distributed computing system, or an event source running in the distributed computing. Monitoring data includes metric data generated by resources and data that represents meta-data properties of event sources. Confidence-controlled sampling is used to determine characteristics of the monitoring data, identify periodic patterns in the behavior of a source, detect changes in behavior of a source, and compare the behavior of two sources. Confidence-controlled sampling speeds up characterization the data sets, determination of behavior patterns, and detection and reporting of anomalies and problems of the resources and event sources of the distributed computing system.


