Adaptive Probing Frequency for Micro-Failure Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current software-defined wide area networks (SD-WANs) face challenges in detecting transient 'micro-failures' in network paths, which are undetected due to the tradeoff in probing frequency, leading to potential impacts on the quality of experience for online applications.
Innovation Solution
A device obtains telemetry data to identify network paths potentially experiencing micro-failures, determines a new higher probing frequency, and sends probes along these paths to detect transient events, using machine learning models to classify and prioritize paths for high-frequency probing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If probing frequency is increased to detect micro-failures, then detection capability is improved, but network traffic impingement worsens
Solution Approach 1:
The system applies different probing frequencies to different network paths based on their sensitivity requirements. Sensitive paths receive high-frequency probing to detect micro-failures, while non-sensitive paths use lower frequencies, thereby optimizing detection capability where needed while minimizing overall network traffic impingement.
Solution Approach 2:
Instead of uniformly increasing probing frequency across all paths, the system applies excessive (high-frequency) probing only partially to specifically identified sensitive paths where micro-failure detection is critical, rather than applying it system-wide.
2Productivity
If probing frequency is decreased to reduce network traffic impingement, then network performance is preserved, but micro-failure detection capability deteriorates
Solution Approach 1:
The system maintains high network performance by using low probing frequencies on non-sensitive paths, while locally increasing probing frequency only on sensitive paths where micro-failure detection is required, thus preserving overall network performance while maintaining necessary detection capability.
Solution Approach 2:
The system applies minimal (low-frequency) probing universally to preserve network performance, but applies excessive (high-frequency) probing partially only to sensitive paths where micro-failure detection is critical, balancing overall performance with specific detection needs.
3Measurement precision
If high frequency probing is applied to all paths, then micro-failure detection is improved, but system complexity increases
Solution Approach 1:
The system identifies and classifies specific network paths as sensitive or non-sensitive based on their characteristics and SLA requirements. This local differentiation allows high-frequency probing to be applied only where necessary, reducing the overall complexity of the probing system while maintaining micro-failure detection capability on sensitive paths.
Solution Approach 2:
Instead of applying high-frequency probing to all paths (which would maximize complexity), the system applies it partially only to sensitive paths, thereby reducing system complexity while still achieving micro-failure detection where it matters most.
Data Source
AI summary
In one embodiment, a device obtains telemetry data associated with application traffic for an online application. The device identifies, based on the telemetry data, a network path as potentially exhibiting transient events during which a performance metric of the network path is degraded that are undetected at a first probing frequency for the network path. The device determines a new probing frequency for the network path, to detect the transient events. The device causes probes to be sent along the network path according to the new probing frequency.


