Multi-layered Ring Network for Asymmetric Bandwidth Utilization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently distributing delta weights across processing units in distributed neural network training, particularly in machine learning applications, where the Allreduce collective operation is crucial but often inefficient.
Innovation Solution
A computer system with interconnected processing nodes arranged in a multi-layered configuration, where each layer forms a ring connected by intralayer links, and interlayer links connect nodes across layers, enabling simultaneous data transmission around embedded one-dimensional logical rings with asymmetric bandwidth utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional Allreduce collective operation is used in distributed neural network training, then data exchange between processing nodes is enabled, but bandwidth utilization is inefficient and data exchange speed is limited
Solution Approach 1:
The patent segments the data exchange process by dividing the processing nodes into multiple rings (first ring, second ring, third ring) that operate simultaneously. Each ring handles a portion of the Allreduce operation independently, allowing parallel data exchange across different node groups without blocking other rings, thereby improving overall bandwidth utilization and data exchange efficiency
Solution Approach 2:
The patent introduces a multi-dimensional ring structure where processing nodes are arranged not just in a single sequence but across multiple concurrent rings. This dimensional expansion allows simultaneous data exchange operations across different rings, transforming the traditional one-dimensional sequential Allreduce into a multi-dimensional parallel operation that significantly boosts productivity and bandwidth efficiency
2Speed
If processing nodes are connected in a simple ring topology, then device complexity is low, but data exchange speed and bandwidth utilization are limited
Solution Approach 1:
The patent segments the single ring topology into multiple concurrent rings (first ring, second ring, third ring), each handling data exchange independently. This segmentation allows simultaneous operations across multiple rings, increasing data exchange speed while maintaining manageable complexity through modular ring structures
Solution Approach 2:
The patent transitions from a simple one-dimensional ring to a multi-dimensional concurrent ring structure. By adding the dimension of time concurrency and multiple parallel rings, the system achieves significantly higher data exchange speeds without proportionally increasing complexity, as each individual ring remains a simple topology
Data Source
Figure 1
Figure 1A
Figure 1B
AI summary
A computer comprising a plurality of interconnected processing nodes arranged in a configuration with multiple layers, arranged along an axis, comprising first and second endmost layers and at least one intermediate layer between the first and second endmost layers is provided. Each layer comprises a plurality of processing nodes connected in a ring by an intralayer respective set of links between each pair of neighbouring processing nodes, the links adapted to operate simultaneously. Nodes in each layer are connected to respective corresponding nodes in each adjacent layer by an interlayer link. Each processing node in the first endmost layer is connected to a corresponding node in the second endmost layer. Data is transmitted around a plurality of embedded one-dimensional logical rings with an asymmetric bandwidth utilisation, each logical ring using all processing nodes of the computer in such a manner that the plurality of embedded one-dimensional logical rings operate simultaneously.