Multi-layered Ring Network for Asymmetric Bandwidth Utilization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently distributing delta weights across processing units in distributed neural network training, particularly in machine learning applications, where the Allreduce collective operation is crucial but often inefficient.

Innovation Solution

A computer system with interconnected processing nodes arranged in a multi-layered configuration, where each layer forms a ring connected by intralayer links, and interlayer links connect nodes across layers, enabling simultaneous data transmission around embedded one-dimensional logical rings with asymmetric bandwidth utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional Allreduce collective operation is used in distributed neural network training, then data exchange between processing nodes is enabled, but bandwidth utilization is inefficient and data exchange speed is limited

Engineering Contradiction:
Improvedata exchange efficiencyVSAvoidbandwidth utilization efficiency
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent segments the data exchange process by dividing the processing nodes into multiple rings (first ring, second ring, third ring) that operate simultaneously. Each ring handles a portion of the Allreduce operation independently, allowing parallel data exchange across different node groups without blocking other rings, thereby improving overall bandwidth utilization and data exchange efficiency

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a multi-dimensional ring structure where processing nodes are arranged not just in a single sequence but across multiple concurrent rings. This dimensional expansion allows simultaneous data exchange operations across different rings, transforming the traditional one-dimensional sequential Allreduce into a multi-dimensional parallel operation that significantly boosts productivity and bandwidth efficiency

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Speed

If processing nodes are connected in a simple ring topology, then device complexity is low, but data exchange speed and bandwidth utilization are limited

Engineering Contradiction:
Improvedata exchange speedVSAvoidnetwork topology complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent segments the single ring topology into multiple concurrent rings (first ring, second ring, third ring), each handling data exchange independently. This segmentation allows simultaneous operations across multiple rings, increasing data exchange speed while maintaining manageable complexity through modular ring structures

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from a simple one-dimensional ring to a multi-dimensional concurrent ring structure. By adding the dimension of time concurrency and multiple parallel rings, the system achieves significantly higher data exchange speeds without proportionally increasing complexity, as each individual ring remains a simple topology

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP3948562B1Embedding rings on a toroid computer network
Publication Date: 2025.04.30 GRAPHCORE LTD
  • EP3948562B1 patent drawingFigure 1
  • EP3948562B1 patent drawingFigure 1A
  • EP3948562B1 patent drawingFigure 1B

AI summary

A computer comprising a plurality of interconnected processing nodes arranged in a configuration with multiple layers, arranged along an axis, comprising first and second endmost layers and at least one intermediate layer between the first and second endmost layers is provided. Each layer comprises a plurality of processing nodes connected in a ring by an intralayer respective set of links between each pair of neighbouring processing nodes, the links adapted to operate simultaneously. Nodes in each layer are connected to respective corresponding nodes in each adjacent layer by an interlayer link. Each processing node in the first endmost layer is connected to a corresponding node in the second endmost layer. Data is transmitted around a plurality of embedded one-dimensional logical rings with an asymmetric bandwidth utilisation, each logical ring using all processing nodes of the computer in such a manner that the plurality of embedded one-dimensional logical rings operate simultaneously.