TCP High Availability via Multi-Dimensional Network Element Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing TCP High Availability systems fail when both control boards in a router fail simultaneously, leading to a loss of routing functionality, as they are not designed to handle such dual failures effectively.
Innovation Solution
A system with multiple network elements, including a primary and standby board configuration, where data and acknowledgments are transferred via TCP modules, allowing the system to reconfigure and maintain functionality even if one or two boards fail, by using sequence numbers for synchronization and data aggregation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If two control boards are used in existing TCP High Availability systems, then basic redundancy is achieved, but the system fails when both boards fail simultaneously
Solution Approach 1:
The system divides the TCP High Availability functionality into separate network elements (active network element and standby network elements) rather than relying on a single multi-functional control board. Each network element can independently handle TCP connections, allowing the system to survive multiple failures while maintaining manageable complexity through functional segmentation.
Solution Approach 2:
The patent transitions from a two-dimensional redundancy model (two control boards) to a multi-dimensional architecture where multiple standby network elements are distributed across different physical or logical dimensions. This allows failures in one dimension (board) to be compensated by redundancy in other dimensions (separate network elements), achieving higher reliability without proportionally increasing overall system complexity.
2Reliability
If multiple standby network elements are added to handle dual failures, then reliability improves to 6 nines, but system complexity increases
Solution Approach 1:
Standby network elements are pre-configured with synchronization capabilities before failures occur. The active network element continuously synchronizes state information to standby elements, so when failures happen, the standbys are already prepared to take over immediately. This preliminary preparation allows the system to achieve 6-nine reliability without requiring complex real-time decision-making during failure events.
Solution Approach 2:
The system creates simplified copies of the active network element's functionality in standby elements. Rather than each standby needing full complexity, they maintain synchronized copies of essential state information and can assume the active role through straightforward state transfer. This copying approach enables high reliability with manageable complexity by avoiding the need for complex coordination among multiple standbys.
Data Source
AI summary
According to one aspect of the present disclosure, a system includes an active network element having circuitry for executing a primary application and a transmission control protocol (TCP) module, multiple standby network elements having circuitry for executing a secondary copy of the primary application and a secondary TCP module, and a network connection coupled to one or more of the active and standby network elements, wherein the active network element and standby network elements are coupled to transfer data and acknowledgments via their respective TCP modules, and wherein the standby network elements are reconfigurable to communicate via the network connection to a peer regardless of the failure of one or two of the network elements.


