VLT Fabric Zero Traffic Loss Sync Mechanism
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions for managing communications links in hyper-converged infrastructure (HCI) systems, such as those using virtual link trunking (VLT), often result in traffic loss when an inter-chassis link goes down or a VLT node reboots, as they fail to ensure synchronization of routing information before enabling traffic flow.
Innovation Solution
An information handling system is configured to detect when a VLT node has malfunctioned and recovered, and it prevents traffic over specific links until all necessary information is synced between the VLT nodes, using a Smart Fabric Services controller to manage the network and ensure zero traffic loss by delaying the restoration of orphan ports until routing protocols converge and data is synchronized.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If VLT nodes are connected to downstream devices with redundant links for high availability, then network redundancy is improved, but traffic loss occurs when a VLT node recovers without proper synchronization
Solution Approach 1:
The system performs preliminary synchronization of routing and MAC information between VLT nodes before allowing traffic flow to be restored. When a VLT node recovers, the controller detects the recovery and initiates a sync process before enabling traffic over the first set of links, preventing traffic loss due to unsynchronized forwarding information.
Solution Approach 2:
The system implements a feedback mechanism where the controller continuously monitors the operational status of VLT nodes and inter-chassis links. Upon detecting node recovery, the controller initiates synchronization and only after confirming sync completion does it allow traffic restoration, creating a closed-loop control system that prevents traffic loss.
2Productivity
If traffic is immediately restored after VLT node recovery, then network availability is improved, but routing information may be outdated causing traffic loss
Solution Approach 1:
The system performs preliminary synchronization of routing and MAC information between VLT nodes before allowing traffic flow to be restored. When a VLT node recovers, the controller detects the recovery and initiates a sync process before enabling traffic over the first set of links, preventing traffic loss due to unsynchronized forwarding information.
Solution Approach 2:
The system maintains continuous monitoring of VLT node status and routing information synchronization. The controller ensures that the synchronization process completes without interruption before traffic restoration, maintaining the continuity of useful action by ensuring forwarding information is always current when traffic flows.
3Loss of information
If synchronization checking is performed before traffic restoration, then traffic loss is prevented, but network recovery time is increased
Solution Approach 1:
The system implements self-service mechanisms where VLT nodes automatically detect their own recovery status and initiate synchronization requests without manual intervention. The controller automatically manages the sync process and traffic restoration, reducing overall recovery time while ensuring synchronization is completed.
Solution Approach 2:
The system dynamically adjusts operational parameters such as synchronization timeout values and traffic restoration thresholds based on network conditions. This allows the system to optimize the balance between synchronization completeness and recovery speed, reducing unnecessary delays while maintaining traffic loss prevention.
Data Source
AI summary
An information handling system may include at least one processor; and a memory; wherein the information handling system is configured to manage a network that includes a first virtual link trunking (VLT) node, a second VLT node, and a plurality of devices that are communicatively coupled to the first VLT node via a first set of links and to the second VLT node via a second set of links, wherein the managing includes: detecting that the first VLT node has malfunctioned; detecting that the first VLT node has recovered; and after the first VLT node has recovered, preventing traffic over the first set of links until determining that all information needed to forward the traffic has been synced between the first VLT node and the second VLT node.

