MLAG Packet Redirection During Hitless Reboot
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In a multi-chassis system undergoing a hitless reboot, the control plane of the restarting peer becomes unavailable, leading to issues with handling CPU-bound packets like ARP refresh replies, which can cause communication failures and downstream effects such as EVPN route withdrawals.
Innovation Solution
Implement modified shutdown and bootup processing logic to automatically redirect CPU-bound packets received during the reboot to another fully operational peer, by reprogramming hardware destination interface rules in the packet processor to forward these packets over an inter-chassis link.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a hitless reboot is performed on a multi-chassis system, then the network device can be upgraded or rebooted without interrupting network traffic flow, but the restarting peer cannot handle CPU-bound packets requiring control plane processing
Solution Approach 1:
The patent introduces a peer selection mechanism that acts as an intermediary to redirect CPU-bound packets from the restarting peer to a selected peer that is fully operational. The packet processor identifies CPU-bound packets and redirects them to an alternative peer, allowing the rebooting peer to maintain data plane functionality while another peer handles control plane processing temporarily
Solution Approach 2:
The system segments the handling of different packet types by routing data plane traffic through the restarting peer's data plane while redirecting control plane traffic to a selected peer. This segmentation allows the control plane and data plane to operate independently during reboot, maintaining network continuity
2Duration of action of moving object
If the control plane of the restarting peer is down during hitless reboot, then the peer can undergo software updates without interruption, but communication failures occur when receiving ARP refresh replies or other CPU-bound packets
Solution Approach 1:
The system performs preliminary actions by selecting a peer to handle control plane packets before the reboot completes. The peer selection occurs during the shutdown phase, ensuring that when the control plane is down, packets are already configured to be handled by the selected peer, preventing communication failures during the reboot window
3Productivity
If manual intervention is required to handle CPU-bound packets during reboot, then packet processing can be controlled, but system complexity and operational overhead increase
Solution Approach 1:
The system implements self-service by automatically selecting a peer to handle control plane packets and dynamically routing packets based on the reboot state. The packet processor automatically identifies CPU-bound packets and redirects them without manual intervention, and automatically updates routing rules when the reboot completes, reducing operational overhead while maintaining processing efficiency
Data Source
AI summary
In one set of embodiments, at the time the control plane of a peer in a multi-chassis system is shut down as part of a hitless reboot, the peer can identify one or more rules programmed in its data plane that (1) are configured on an ingress interface of a multi-chassis link aggregation group (MLAG) to which the peer is connected, and (2) specify a central processing unit (CPU) of the peer as a destination for matched network traffic. The peer can then change each identified rule to specify an inter-chassis link between the peer and another peer in the multi-chassis system, rather than the CPU of the peer, as the destination for matched network traffic.


