Selective Multipath Packet Spraying for Data Center Resilience

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data center networks face significant challenges in detecting path failures and reducing packet loss due to link or node failures, as they rely on slow control plane updates of forwarding tables to adapt to network topology changes.

Innovation Solution

Implementing techniques where source and destination network devices maintain information about path health and connectivity, allowing them to spray packets over healthy paths and avoid failed ones, enabling fast recovery and reduced packet loss without waiting for control plane updates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If control plane software updates forwarding tables to address detected failures, then network reliability is improved, but recovery time increases

Engineering Contradiction:
Improvenetwork reliabilityVSAvoidrecovery time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent pre-establishes multiple forwarding paths and maintains path health information before failures occur. When a failure is detected, the system can immediately switch to pre-planned alternative paths without waiting for control plane updates, thus achieving fast recovery while maintaining reliability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The data plane devices autonomously detect path failures and make routing decisions based on local path health information without requiring control plane intervention. This self-service mechanism enables immediate local response to failures, reducing recovery time while maintaining network reliability

Inventive Principle:
Principle #25Self-service

2Device complexity

If packets are forwarded along a single path through the switching fabric, then device complexity is reduced, but network reliability deteriorates due to significant failure rates

Engineering Contradiction:
Improveforwarding complexityVSAvoidnetwork reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments the network into multiple independent forwarding paths between source and destination. By dividing the single path into multiple parallel paths, the system achieves both simplicity (each path uses standard forwarding) and reliability (multiple paths provide redundancy against failures)

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If forwarding tables are updated to reflect network topology changes, then adaptability is improved, but productivity decreases due to long update times

Engineering Contradiction:
Improvetopology adaptabilityVSAvoiddata transmission efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system pre-computes and caches multiple alternative forwarding paths along with their health status before topology changes occur. When failures are detected, the system can immediately activate pre-planned alternative paths without waiting for control plane topology updates, thus maintaining both adaptability and productivity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Data plane devices autonomously track path health and make adaptive routing decisions without requiring control plane updates. This self-service approach enables immediate adaptation to topology changes while maintaining high data transmission efficiency by avoiding control plane update delays

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12166666B2Resilient network communication using selective multipath packet flow spraying
Publication Date: 2024.12.10 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12166666B2 patent drawing
  • US12166666B2 patent drawing
  • US12166666B2 patent drawing

AI summary

Techniques for detecting path failures and reducing packet loss as a result of such failures are described for use within a data center or other environment. For example, a source and/or destination access node may create and/or maintain information about health and/or connectivity for a plurality of ports or paths between the source and destination device and core switches. The source access node may spray packets over a number of paths between the source access node and the destination access node. The source access node may use the information about connectivity for the paths between the source or destination access nodes and the core switches to limit the paths over which packets are sprayed. The source access node may spray packets over paths between the source access node and the destination access node that are identified as healthy, while avoiding paths that have been identified as failed.