Spray Network Fault Recovery Through Entropy-Based Port Rerouting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing network fault recovery methods in computer systems are time-consuming and resource-intensive, particularly in large-scale networks, due to the frequent occurrence of link failures and hardware errors, which disrupt data transmission and require manual remapping of faulty links.
Innovation Solution
A system and method for detecting and recovering from network faults by storing entropy values in packets that map to specific ports, monitoring acknowledgement packets for congestion or failure, and dynamically rerouting traffic around problematic ports using a spray network approach.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional fault recovery methods are used to detect and isolate path failures in software, then fault detection capability is improved, but the recovery process becomes time-consuming and resource-intensive
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing multiple alternative paths between source-destination pairs in a path cache before faults occur. When a fault is detected, the system can immediately switch to a pre-computed alternative path without performing time-consuming real-time path computation, thus resolving the contradiction between reliable fault detection and fast recovery.
2Adaptability or versatility
If multiple network switches and links are used to connect processors, then network connectivity and data transmission capability are improved, but the complexity of detecting and recovering from link failures increases
Solution Approach 1:
The patent implements feedback mechanisms where destination processors send path status information back to source processors, and sources monitor the health of active paths. This feedback loop enables automatic detection of link failures and triggers seamless switching to alternative paths, managing the complexity of fault detection in large-scale networks while maintaining high connectivity.
3Reliability
If manual remapping of faulty links to spare devices is performed, then network fault recovery is achieved, but network operating costs and resource consumption increase significantly
Solution Approach 1:
The patent enables self-service fault recovery by implementing automatic path failure detection and recovery mechanisms at the network endpoints. Source and destination processors work autonomously to detect faults, select alternative paths from cached options, and resume data transmission without requiring manual intervention or significant network resource consumption, thus maintaining reliability while improving operational efficiency.
Data Source
AI summary
Embodiments of the present disclosure include systems and methods for fault detection and recovery over a network. A value of a set of values is stored in packets transmitted during a data transaction between a source and destination. The value corresponds to ports used by one or more switches in the path between the source and destination. The destination includes the value in an acknowledgement packet. Logic circuits in the source device track packets and corresponding values. When a status indicates a particular packet has not received an acknowledgement, the value for the packet may be removed from the set of values. Particular ports that may be congested or down may be detected and the packets re-routed using the logic circuits in the source device.


