Dynamic Fabric Scheduling for Small-Flow Congestion Bypass

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing networking systems experience high latency and increased flow completion time due to larger-sized data flows during all-to-all communications, particularly in training deep learning models like DLRM, which consume a majority of training time and represent a bottleneck in data-parallel synchronized distributed deep learning.

Innovation Solution

Networking devices determine the size of data flows and bypass congestion controllers for small data flows, implementing time synchronization to divide the processing network into smaller networks when necessary, and adjust Quality of Service (QoS) adaptations to optimize data flow management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If congestion controllers manage all data flows in the network, then network traffic is controlled and prevented from overwhelming the system, but latency and flow completion time increase for small data flows

Engineering Contradiction:
Improvenetwork traffic controlVSAvoidlatency and flow completion time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies different quality of service treatments to different data flows based on their characteristics. Small data flows (below threshold) receive expedited handling by bypassing congestion controllers, while large data flows continue to receive full congestion control management. This local differentiation resolves the contradiction by optimizing latency for small flows without compromising the reliability of overall network traffic control.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments data flows into two categories based on size thresholds: small data flows that bypass congestion controllers and large data flows that are managed by them. This segmentation allows the system to apply appropriate handling strategies to each category, reducing latency for small flows while maintaining reliable traffic control for large flows.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If the processing network operates as a single large network, then all-to-all communications can occur across the entire network, but latency increases due to the size of the network

Engineering Contradiction:
Improveall-to-all communication capabilityVSAvoidcommunication latency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent divides the large processing network into multiple smaller subnetworks when the total network size exceeds a threshold. This segmentation reduces communication latency within each subnetwork while maintaining the overall all-to-all communication capability across the entire network through coordinated operations between subnetworks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a temporal dimension to network operations by implementing time synchronization across subnetworks. This allows the system to maintain large-scale all-to-all communication capabilities while achieving small-network latency characteristics through time-division multiplexing of communication operations across subnetworks.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of operation

If uniform quality of service is applied to all data flows, then network management is simplified, but performance is not optimized for different flow sizes

Engineering Contradiction:
Improvenetwork management simplicityVSAvoidtraining efficiency for deep learning models
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent implements differentiated quality of service policies based on data flow characteristics, specifically flow size. Small data flows receive prioritized handling with bypassed congestion control, while large data flows receive standard congestion control management. This local quality differentiation optimizes training efficiency for deep learning models without significantly complicating network management, as the differentiation is based on simple threshold criteria.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250240243A1Dynamic fabric reaction for optimized collective communication
Publication Date: 2025.07.24 MELLANOX TECHNOLOGIES LTD(IL)
  • US20250240243A1 patent drawing
  • US20250240243A1 patent drawing
  • US20250240243A1 patent drawing

AI summary

A networking device and system are described, among other things. An illustrative system is disclosed to include a congestion controller that manages traffic across a network fabric using receiver-based packet scheduling and a networking device that employs the congestion controller for data flows qualified as a large data flow but bypasses the congestion controller for data flows qualified as a small data flow. For example, the networking device may receive information describing a data flow directed toward a processing network; determine, based on the information describing the data flow, a size of the data flow; determine the size of the data flow is below a predetermined flow threshold; and in response to determining that the size of the data flow is below a predetermined threshold, bypass the congestion controller.