Network Pacemaker Agents for SDN Fault Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In software-defined networks (SDNs) with virtualized network functions (VNFs) on general-purpose hardware, it is challenging to assess root causes of failures and ensure performance assurance due to the complexity of monitoring various network elements without interfering with normal network functions, especially in cloud infrastructure where traditional passive monitoring methods often miss failures at different layers.
Innovation Solution
A network assurance system that includes a pacemaker, a configuration analysis microservice, and agents to selectively apply an atomic workload to network paths, monitor the pulse, and adjust the workload based on path health, using active probing in conjunction with passive monitoring to improve fault detection and minimize resource costs by distributing probing activities across different times.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If passive monitoring methods are used to monitor network elements, then resource usage is minimized, but fault detection capability is insufficient and failures at different layers are missed
Solution Approach 1:
The monitoring system is segmented into multiple pacemaker agents distributed across different network elements, each independently generating atomic workloads and collecting metrics. This segmentation enables comprehensive fault detection across multiple layers without requiring a single complex centralized monitoring system.
Solution Approach 2:
Pacemaker agents act as intermediary components that are integrated into the existing VNF architecture. These agents generate atomic workloads and collect performance metrics without requiring external monitoring infrastructure, thereby improving fault detection capability while avoiding additional system complexity.
2Reliability
If active probing is used to monitor network paths, then fault detection is improved, but resource usage and interference with normal network functions increases
Solution Approach 1:
The system applies partial probing by selecting specific atomic workloads from a predefined set based on current network conditions and requirements. Not all possible workloads are executed simultaneously, but only those necessary for detecting specific fault conditions, thereby reducing resource consumption while maintaining effective fault detection.
Solution Approach 2:
The pacemaker agents dynamically adjust workload parameters such as probing frequency, intensity, and type based on network conditions. When faults are detected, the system increases probing intensity; when network is stable, it reduces probing activity, thereby optimizing the balance between fault detection accuracy and resource usage.
3Reliability
If comprehensive monitoring of all network elements is implemented, then performance assurance is improved, but the scope of monitoring creates interference with normal network functions
Solution Approach 1:
The pacemaker agents execute atomic workloads periodically rather than continuously, with configurable intervals between probing activities. This periodic execution provides comprehensive performance assurance data while allowing the network to operate normally between probing cycles, thereby minimizing interference with network functions.
Solution Approach 2:
The pacemaker agents are designed as multi-functional components that can perform multiple types of atomic workloads (network connectivity testing, service performance measurement, fault detection) within a single unified framework. This universality enables comprehensive monitoring without requiring separate specialized monitoring systems for each function, reducing overall interference.
Data Source
AI summary
A system for performing assurance of at least one network element in a network, the system comprising: at least one pacemaker, wherein the at least one pacemaker is a virtual machine in communication with a path to the at least one network element in the network, the at least one pacemaker configured to selectively provide an atomic workload to the path; a configuration analysis microservice in communication with the at least one pacemaker, the configuration analysis microservice is configured to provide a policy to the at least one pacemaker for application of the atomic workload and monitor a pulse from the atomic workload on the path.


