Ethernet Backplane Failover via Heartbeat Link Integrity Checks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Ethernet backplane systems in chassis-based systems face challenges with long convergence times in link failure detection and failover switching, which are not suitable for high availability requirements, as conventional spanning tree protocols and link aggregation methods are inefficient in these environments.
Innovation Solution
A high availability backplane architecture with redundant node boards and switch fabric boards that perform link integrity checks and automatically switch to backup links upon failure, using heartbeat messages and frame error rates for rapid failover, minimizing CPU processing load and enabling fast recovery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If spanning tree protocol is used for link failure detection and failover, then network reliability is improved, but convergence time becomes too long (20-50 seconds)
Solution Approach 1:
The patent implements preliminary action by pre-configuring redundant paths and pre-establishing failure detection mechanisms. The system proactively monitors link status using heartbeat packets before failures occur, and pre-computes alternative paths so that when a failure is detected, the switch can immediately activate the backup path without waiting for protocol convergence. This eliminates the 20-50 second delay inherent in traditional spanning tree protocols.
Solution Approach 2:
The patent employs feedback mechanisms through continuous heartbeat packet exchanges between switches. When a link failure occurs, the absence of heartbeat packets provides immediate feedback about the failure condition. The system uses this feedback to trigger rapid failover to pre-configured redundant paths, achieving convergence in milliseconds rather than seconds or tens of seconds.
2Loss of time
If fast spanning tree protocol is used, then convergence time is reduced to 50 msec, but complexity of the system increases
Solution Approach 1:
The patent segments the failover process into independent, simple components: (1) continuous heartbeat packet transmission for failure detection, (2) pre-configured redundant paths in the network topology, and (3) simple path switching logic. This segmentation avoids the complex protocol state machines and calculations required by fast spanning tree, achieving rapid convergence through modular, independent functions that are easier to implement and maintain.
Solution Approach 2:
The patent introduces heartbeat packets as an intermediary mechanism for failure detection. Instead of using complex protocol negotiations and calculations, the system uses simple periodic heartbeat packets to probe link status. The presence or absence of these packets provides clear, unambiguous failure signals that trigger pre-planned failover actions, simplifying the overall system while achieving fast convergence.
3Reliability
If link aggregation is used for failover, then availability is improved, but it only provides failover among parallel connections shared with the same end nodes
Solution Approach 1:
The patent creates a universal failover mechanism that works across all connection types in the network. The failure detection method using heartbeat packets and the failover mechanism using pre-configured redundant paths are not limited to specific connection patterns. This approach can handle point-to-point links, hub-and-spoke topologies, and mesh networks equally well, providing adaptability to various network architectures while maintaining high availability.
Data Source
AI summary
A high availability backplane architecture. The backplane system includes redundant node boards operatively communicating with redundant switch fabric boards. Uplink ports of the node boards are logically grouped into trunk ports at one end of the communication link with the switch fabric boards. The node boards and the switch fabric boards routinely perform link integrity checks when operating in a normal mode such that each can independently initiate failover to working ports when a link failure is detected. Link failure is detected either by sending a link heartbeat message after the link has had no traffic for a predetermined interval, or after receiving a predetermined consecutive number of invalid packets. Once the link failure is resolved, operation resumes in normal mode.


