Spray Network Fault Recovery Through Entropy-Based Port Rerouting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing network fault recovery methods in computer systems are time-consuming and resource-intensive, particularly in large-scale networks, due to the frequent occurrence of link failures and hardware errors, which disrupt data transmission and require manual remapping of faulty links.

Innovation Solution

A system and method for detecting and recovering from network faults by storing entropy values in packets that map to specific ports, monitoring acknowledgement packets for congestion or failure, and dynamically rerouting traffic around problematic ports using a spray network approach.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional fault recovery methods are used to detect and isolate path failures in software, then fault detection capability is improved, but the recovery process becomes time-consuming and resource-intensive

Engineering Contradiction:
Improvefault detection capabilityVSAvoidrecovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing multiple alternative paths between source-destination pairs in a path cache before faults occur. When a fault is detected, the system can immediately switch to a pre-computed alternative path without performing time-consuming real-time path computation, thus resolving the contradiction between reliable fault detection and fast recovery.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If multiple network switches and links are used to connect processors, then network connectivity and data transmission capability are improved, but the complexity of detecting and recovering from link failures increases

Engineering Contradiction:
Improvenetwork connectivityVSAvoidfault detection complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements feedback mechanisms where destination processors send path status information back to source processors, and sources monitor the health of active paths. This feedback loop enables automatic detection of link failures and triggers seamless switching to alternative paths, managing the complexity of fault detection in large-scale networks while maintaining high connectivity.

Inventive Principle:
Principle #23Feedback

3Reliability

If manual remapping of faulty links to spare devices is performed, then network fault recovery is achieved, but network operating costs and resource consumption increase significantly

Engineering Contradiction:
Improvefault recovery capabilityVSAvoidnetwork operating efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent enables self-service fault recovery by implementing automatic path failure detection and recovery mechanisms at the network endpoints. Source and destination processors work autonomously to detect faults, select alternative paths from cached options, and resume data transmission without requiring manual intervention or significant network resource consumption, thus maintaining reliability while improving operational efficiency.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12413508B2System and method for fault recovery in spray based networks
Publication Date: 2025.09.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12413508B2 patent drawing
  • US12413508B2 patent drawing
  • US12413508B2 patent drawing

AI summary

Embodiments of the present disclosure include systems and methods for fault detection and recovery over a network. A value of a set of values is stored in packets transmitted during a data transaction between a source and destination. The value corresponds to ports used by one or more switches in the path between the source and destination. The destination includes the value in an acknowledgement packet. Logic circuits in the source device track packets and corresponding values. When a status indicates a particular packet has not received an acknowledgement, the value for the packet may be removed from the set of values. Particular ports that may be congested or down may be detected and the packets re-routed using the logic circuits in the source device.