SDN Datapath Service Rings for Zero-Downtime Maintenance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing software defined networks (SDN) intermediate devices experience disruptions due to issues that cause downtime, affecting user traffic and workflows, necessitating a solution for maintaining device health and ensuring zero downtime.
Innovation Solution
Implementing a system where each processing unit in an SDN intermediate device is paired with another card to form a logical ring, with health monitoring and mapping to ensure at most one card can be down at a time, and maintaining a threshold service availability guarantee through health signal monitoring and proactive maintenance management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If SDN intermediate devices are maintained and serviced, then device health is improved, but downtime and service disruption occur
Solution Approach 1:
The system performs preliminary actions by maintaining a hot standby processing unit that is pre-configured and ready to immediately take over if the active unit fails or requires maintenance. This eliminates downtime by having the backup unit prepared in advance, allowing seamless failover without service interruption.
Solution Approach 2:
The patent applies local quality by implementing health monitoring and failover capabilities at the processing unit level rather than requiring full system downtime. Individual processing units can be serviced independently while others continue operating, allowing localized maintenance without global service disruption.
2Reliability
If health monitoring and ring management are implemented, then service availability is improved, but system complexity increases
Solution Approach 1:
The system implements feedback through health signal monitoring between processing units and the switch. Health signals are continuously exchanged to detect failures or maintenance needs, triggering automatic failover when thresholds are exceeded. This feedback mechanism improves service availability by enabling proactive response to issues while managing complexity through standardized signal protocols.
Solution Approach 2:
The patent uses an intermediary approach by introducing a switch as a central coordinating element that manages the logical ring of processing units. The switch mediates health signal exchange and coordinates failover operations, simplifying the complexity of direct peer-to-peer communication between multiple processing units while maintaining high service availability.
3Reliability
If network cards are paired in logical rings with backup pairing, then fault tolerance is improved, but device complexity increases
Solution Approach 1:
The patent applies merging by combining multiple processing units into a logical ring structure where they share common health monitoring and failover protocols. By merging their operations under unified management rules (at most one card down at a time), the system achieves enhanced fault tolerance while controlling complexity through standardized interaction patterns rather than independent complex configurations for each unit.
Data Source
AI summary
High availability network services are provided in a in a virtual computing environment comprising a plurality of network devices running in a software defined network (SDN) of the virtual computing environment, the network devices comprising a plurality of SDN appliances configured to disaggregate enforcement of policies of the SDN from hosts of the virtual computing environment, the hosts implemented on servers communicatively coupled to network interfaces of the SDN appliances, the servers hosting a plurality of virtual machines, the SDN appliance comprising a plurality of smart network interface cards (fNICs) configured to implement functionality of the SDN appliances.


