Network Dynamics Forecasting with Triggered Models for SLA Violations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing network prediction systems in SD-WANs face challenges in accurately predicting service level agreement (SLA) violations due to the selection of inappropriate timescales for failure prediction models, leading to reactive decision-making and high false positive rates, which can negatively impact user experience.
Innovation Solution
Implementing a combination of short and long timescale prediction models in a network, where the long timescale model triggers the short timescale model to activate predictions when certain conditions are met, allowing for collaborative operation and reducing false positives by leveraging high-frequency telemetry data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a short timescale prediction model is used to predict network failures, then the predictive accuracy for certain failure types improves, but the false positive rate increases and system complexity increases
Solution Approach 1:
The prediction system is segmented into multiple independent prediction models, each specialized for predicting specific failure types. This allows each model to be optimized for its specific domain while maintaining overall system reliability through selective activation based on failure type identification.
Solution Approach 2:
Different prediction models are assigned different timescales and characteristics tailored to their specific failure type. Each model operates with locally optimized parameters (timescale, sensitivity) appropriate for its target failure type, rather than using a uniform approach across all predictions.
2Measurement precision
If high frequency telemetry collection is implemented to capture short-term network dynamics, then the predictive capability for certain failures improves, but the data processing complexity and resource consumption increase
Solution Approach 1:
The system collects telemetry at different frequencies tailored to the specific requirements of each prediction model. Critical failure types requiring short-term detection use high-frequency collection, while less critical types use lower frequency, optimizing the balance between predictive capability and processing overhead.
Solution Approach 2:
Instead of continuously collecting high-frequency telemetry for all failure types, the system implements partial monitoring - high frequency only where needed based on failure type analysis. This reduces overall data processing complexity while maintaining predictive accuracy for critical failures.
3Device complexity
If a single prediction model is used for all failure types, then the system complexity is reduced, but the predictive accuracy for specific failure types deteriorates
Solution Approach 1:
The prediction system is divided into multiple specialized models, each trained and optimized for specific failure types. This segmentation allows each model to achieve high accuracy for its target domain while the overall system architecture remains manageable through modular design and selective activation.
Data Source
AI summary
In one embodiment, a device deploys short timescale prediction model and a long timescale prediction model to one or more hosts in a network, whereby the short timescale prediction model predicts failure conditions for an online application that are attributable to the network on a timescale that is shorter than that of the long timescale prediction model. The device configures a trigger that causes the long timescale prediction model to activate predictions by the short timescale prediction model. The device evaluates performance of the short timescale prediction model. The device adjusts the trigger, when the performance of the short timescale prediction model is unacceptable.


