Self-Healing SD-WAN Flow Testing for Anomaly Remediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing SD-WAN networks lack effective mechanisms for autonomously detecting and remediating anomalies in real-time to maintain optimal performance and user experience, particularly in distributed workforces and multi-cloud environments.
Innovation Solution
A self-healing SD-WAN system utilizing machine-trained processes to analyze flow data from multiple forwarding elements, identify anomalies through Gaussian distribution analysis and topology-based detection, and implement remedial actions via an SD-WAN controller to optimize network performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If traditional network monitoring is used, then basic network status can be tracked, but real-time anomaly detection and automated remediation capability is insufficient
Solution Approach 1:
The system enables self-service through automated anomaly detection and remediation. The analytics system continuously monitors network metrics, detects anomalies using machine learning models, and automatically executes remediation actions without human intervention. This allows the network to self-diagnose and self-heal, maintaining reliability while increasing automation extent.
Solution Approach 2:
The system implements feedback loops where network performance data is continuously collected, analyzed, and used to trigger remediation actions. The closed-loop architecture ensures that anomalies are detected, remediation is executed, and the impact is measured, creating a self-correcting system that maintains network reliability through automated feedback mechanisms.
2Measurement precision
If comprehensive flow data collection from multiple FEs is implemented, then detailed network analysis is possible, but data processing complexity increases
Solution Approach 1:
The system segments the data processing workload by implementing distributed collection at FEs, centralized aggregation at the analytics system, and specialized analysis functions. Flow data is collected from multiple FEs, aggregated to a central analytics system, and then processed through specialized machine learning models, dividing the complex processing task into manageable segments that improve accuracy while controlling complexity.
Solution Approach 2:
The analytics system acts as an intermediary between data collection and anomaly detection. It aggregates and pre-processes flow data from multiple FEs before feeding it to machine learning models, simplifying the overall architecture by introducing a dedicated layer that handles data transformation and preparation, thereby reducing direct processing complexity.
3Reliability
If machine-trained processes are used for anomaly detection, then detection capability is enhanced, but system configuration and maintenance difficulty increases
Solution Approach 1:
The system performs preliminary action by pre-training machine learning models during system deployment before actual anomaly detection begins. Baseline behavior patterns are established in advance, and detection thresholds are configured beforehand, allowing the system to start operating with high reliability while reducing the complexity of ongoing configuration and maintenance.
Solution Approach 2:
The system manages complexity by dynamically adjusting detection parameters and model thresholds based on network conditions. Rather than requiring manual reconfiguration, the system automatically adapts its anomaly detection parameters to changing network environments, maintaining high reliability while simplifying operational maintenance through automated parameter optimization.
Data Source
AI summary
Some embodiments of the invention provide a method of remediating anomalies in an SD-WAN implemented by multiple forwarding elements (FEs) located at multiple sites connected by the SD-WAN. The method determines that a particular anomaly detected in the SD-WAN requires remediation to improve performance for a set of one or more flows traversing through the SD-WAN. The method identifies a set of two or more remedial actions for remediating the particular anomaly in the SD-WAN. For each identified remedial action in the set, the method selectively implements the identified remedial action for a subset of the set of flows for a duration of time in order to collect a set of performance metrics associated with SD-WAN performance during the duration of time for which the identified remedial action is implemented. Based on the collected sets of performance metrics, the method uses a machine-trained process to select one of the identified remedial actions as an optimal remedial action in the set to implement for all of the flows in the set of flows. The method implements the selected remedial action for all of the flows in the set of flows.


