Network Recovery from Multiple Link Failures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Ethernet-based networks are slow to respond and recover from multiple link failures, leading to unreliable self-recovery and noticeable disruptions to subscribers, with existing protocols like STP/RSTP and EPSR failing to provide fast and reliable recovery from multiple link failures.
Innovation Solution
A system and method utilizing a master node and transit nodes in a ring configuration, where the master node initiates health-check messages and timers to determine restored links, allowing blocked ports to be reopened for data forwarding, preventing network loops and ensuring quick recovery from single and multiple link failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If STP/RSTP protocols are used for loop prevention and backup path assurance, then network availability is maintained, but response time to network failures increases to up to 30 seconds
Solution Approach 1:
The patent pre-configures alternative active paths and designates master and slave ports before failures occur. When a failure is detected, the pre-established alternative path is immediately activated without requiring complex real-time calculations, thus achieving fast recovery while maintaining network availability.
Solution Approach 2:
The system dynamically adjusts port states (blocked vs. active) based on real-time network conditions and failure detection. The master node can dynamically switch between primary and alternative paths, and ports can transition between blocked and active states, providing adaptive response to changing network topology.
2Loss of time
If EPSR protocol is used for fast single link failure recovery, then recovery speed improves, but recovery from multiple link failures becomes impossible until all failed links are restored
Solution Approach 1:
The patent segments the network recovery process by having each node independently detect and report its local link status to the master node. The master node then coordinates recovery by selectively activating alternative paths for each failed link independently, allowing multiple link failures to be recovered simultaneously rather than requiring sequential recovery.
Solution Approach 2:
The system implements a feedback mechanism where nodes continuously monitor link status and report to the master node. The master node receives feedback about which links are failed and which alternative paths are available, then makes informed decisions about which ports to activate, enabling coordinated recovery from multiple simultaneous failures.
3Measurement precision
If health-check messages are transmitted to determine restored links, then accurate detection of link restoration is achieved, but network traffic increases
Solution Approach 1:
The master node transmits health-check messages periodically at predetermined intervals rather than continuously. This periodic transmission provides sufficient information to detect link restoration events while limiting the total volume of control traffic on the network, achieving a balance between detection accuracy and traffic overhead.
Data Source
AI summary
Methods and systems for fast and reliable network recovery from multiple link failures. In accordance with one example of the method, a master node receives a request from a transit node having a blocked port to open the blocked port for data forwarding, wherein the blocked port is associated with a restored link. The master node starts a health-check timer and transmits a health-check message on its primary port to determine that all failed links of the network are restored. Upon determining that the health-check message is received at its secondary port, the master node transmits a message to each transit node indicating that all failed links are restored. Upon determining that the health-check message is not received at its secondary port before the health-check timer lapsed, the master node transmits a message to the transit node to open the blocked port associated with the restored link for data forwarding.


