Adaptive Sampling for Application Throughput Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network assurance systems face challenges in accurately detecting application throughput anomalies due to irregular applicative patterns and the inefficiency of traditional sampling methods, which can lead to excessive telemetry data generation and increased computational overhead.
Innovation Solution
An adaptive sampling mechanism is introduced, where nodes in a network report histograms of application-specific throughput metrics to a supervisory service, which merges these metrics for anomaly detection and adjusts the reporting strategy to minimize overhead while maintaining accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional sampling methods are used to monitor application throughput, then network assurance systems can operate with simple rules, but measurement precision and anomaly detection accuracy deteriorate due to irregular applicative patterns
Solution Approach 1:
The patent implements adaptive sampling where the sampling rate is dynamically adjusted based on network conditions and applicative patterns. The system transitions from static predefined sampling intervals to dynamic sampling that responds to detected anomalies and network state, thereby improving measurement precision without requiring overly complex fixed-rule systems
Solution Approach 2:
The system employs feedback mechanisms where measured throughput data is used to refine future sampling decisions. Anomaly detection results feed back into the sampling strategy, allowing the system to focus sampling resources on periods and conditions most likely to yield informative data, thus improving accuracy while managing complexity through learned patterns
2Measurement precision
If sampling frequency is increased to improve throughput modeling accuracy, then measurement precision improves, but bandwidth consumption and computational overhead increase
Solution Approach 1:
The patent changes the sampling parameter from fixed frequency to adaptive frequency based on network conditions. During normal operation, sampling occurs at lower frequencies to conserve bandwidth, while sampling intensity increases automatically when anomalies are detected or network conditions indicate potential issues, thereby maintaining accuracy when needed while reducing overall bandwidth consumption
Solution Approach 2:
The system applies partial sampling rather than continuous full sampling. By sampling only at critical moments (anomaly detection points, threshold violations, or when network state changes), the system achieves sufficient accuracy for assurance purposes without the excessive bandwidth consumption that would result from continuous high-frequency sampling
3Measurement precision
If more comprehensive metrics are collected to improve anomaly detection, then measurement precision improves, but device complexity and computational overhead increase
Solution Approach 1:
The patent extracts and focuses on the most critical throughput metrics rather than collecting all possible network parameters. By identifying and measuring only the essential throughput characteristics (application-level throughput, packet loss, latency) rather than all底层 network metrics, the system achieves adequate anomaly detection capability with reduced processing complexity
Solution Approach 2:
The system segments throughput measurement into application-specific components, measuring throughput separately for different applications rather than treating network traffic as a monolithic stream. This segmentation allows targeted collection of relevant metrics for each application type, improving detection accuracy for specific applications while avoiding the computational overhead of processing all possible network metrics uniformly
Data Source
AI summary
In one embodiment, a node in a network reports, to a supervisory service, histograms of application-specific throughput metrics measured from the network. The node receives, from the supervisory service, a merged histogram of application-specific throughput metrics. The supervisory service generated the merged histogram based on a plurality of histograms reported to the supervisory service by a plurality of nodes. The node performs, using the merged histogram, application throughput anomaly detection on traffic in the network. The node causes performance of a mitigation action in the network when an application throughput anomaly is detected. The node adjusts, based on a control command sent by the supervisory service, a histogram reporting strategy used by the node to report the histograms of application-specific throughput metrics to the supervisory service.


