Multi-path Network Fault Bypass via Distributed Switch Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional multi-path networks face significant delays and disruptions when re-evaluating routes due to faults, leading to performance issues and increased error rates, especially in large networks with many components.

Innovation Solution

A multi-path network with fault monitors and dynamically selectable output ports that communicate fault information directly between network elements to bypass faulty links, using an Encapsulation Layer and Fabric Protocol Data Units to manage faults without a supervisory manager and re-evaluation of the entire network.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional error detection and retry protocols (TCP, CRC) are used to ensure data reliability, then data transmission reliability is improved, but network performance and bandwidth utilization deteriorate due to significant overhead and repeated transmissions

Engineering Contradiction:
Improvedata transmission reliabilityVSAvoidnetwork performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements preliminary fault detection by monitoring error rates on network links and proactively identifying faulty connections before they cause widespread data corruption. By detecting faults early and preemptively rerouting traffic around problematic links, the system prevents repeated transmission attempts on failing connections, thereby maintaining high reliability without the performance penalty of conventional retry protocols

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary fault detection and management layer that sits between the data transmission layers and the physical network infrastructure. This intermediary monitors link health, identifies faulty connections, and orchestrates dynamic rerouting decisions, separating the reliability function from the transmission protocol itself and eliminating the need for overhead-heavy retry mechanisms

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If a supervisory manager re-evaluates the entire network to reroute traffic around faults, then fault tolerance is improved, but network disruption and reconfiguration time increase significantly

Engineering Contradiction:
Improvefault toleranceVSAvoidreconfiguration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the fault management function into distributed components at individual network switches rather than relying on a centralized supervisory manager. Each switch independently monitors its own uplink connections and makes local rerouting decisions, dividing the global reconfiguration problem into smaller, parallel local decisions that can be executed simultaneously without coordination overhead

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enables network switches to autonomously detect faults on their own connections and self-correct by dynamically rerouting traffic around faulty links without external intervention. This self-service capability eliminates the time-consuming centralized re-evaluation process while maintaining comprehensive fault tolerance across the entire network

Inventive Principle:
Principle #25Self-service

3Reliability

If dynamic rerouting is implemented to bypass faulty links, then error rates are reduced, but network complexity and control mechanisms increase

Engineering Contradiction:
Improveerror rateVSAvoidnetwork control complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by implementing fault detection and rerouting capabilities at each individual network switch rather than through a centralized system. Each switch independently monitors its own uplink error rates and makes localized rerouting decisions, which distributes the control complexity across multiple simple nodes rather than concentrating it in one complex controller

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent merges the fault detection, fault classification, and rerouting decision-making functions into a single integrated process at each network switch. By combining these functions locally, the system avoids the complexity of separate centralized monitoring and control mechanisms while achieving effective dynamic rerouting to maintain low error rates

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9954800B2Multi-path network with fault detection and dynamic adjustments
Publication Date: 2018.04.24 CRAY UK
  • US9954800B2 patent drawing
  • US9954800B2 patent drawing
  • US9954800B2 patent drawing

AI summary

A multi-path network for use in a bridge, switch, router, hub or the like, includes a plurality of network ports; a plurality of network elements; and a plurality of network links interconnecting the network elements and the network ports, for transporting data packets. Each network element has a fault monitor for detecting faults on the network links to which the network element is connected, a plurality of dynamically selectable output ports and output port selection mechanism. Each network element also being adapted to communicate the existence of a fault back to one or more other network elements so that network elements connected to the faulty network link can be bypassed, and each network element being adapted to update the output port selection mechanism when communication of the existence of a fault is received so that only those output ports which ensure the faulty network link is bypassed are dynamically selectable. Also, a method of managing faults in a multi-path network utilizes a similar methodology to insure that identified faults are avoided.