AI Training Network With Pre-Established Optical GPU Channels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing scale of neural network and data sets in AI training leads to frequent data transmission between GPU chips, which significantly impacts the duration of the training process due to the time required for establishing optical channels.

Innovation Solution

Implementing an optical cross-connect (OXC) management system that initiates channel switching before data transmission is needed, utilizing MEMS or SiP technology to establish channels ahead of time, allowing immediate data transfer between GPUs without waiting for channel completion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If channel switching is performed only after calculation completion, then data transmission can occur, but the overall training duration increases due to channel establishment time

Engineering Contradiction:
Improvechannel establishment timeVSAvoiddata transmission efficiency
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The OXC performs channel switching in advance before the GPU completes its calculation. The channel establishment is triggered when the GPU finishes computing, allowing the optical channel to be ready before data transmission begins, thereby eliminating idle waiting time and reducing overall training duration

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By overlapping the channel establishment process with the GPU calculation process, the system ensures continuous useful action. The OXC channel setup occurs concurrently with GPU computation, so that when calculation completes, the transmission channel is already established and data can flow immediately without interruption

Inventive Principle:
Principle #20Continuity of useful action

2Power

If accelerator cluster scale is increased to improve computing power, then processing capability increases, but data transmission frequency between GPU chips increases, worsening the impact of channel establishment time

Engineering Contradiction:
Improvecomputing powerVSAvoidtraining process duration
Core Design Contradiction:
PowerVSLoss of time

Solution Approach 1:

The system proactively initiates channel switching operations before data transmission is required. When GPU scale increases and transmission frequency increases, this preliminary action ensures channels are pre-established, preventing transmission bottlenecks that would otherwise worsen with higher transmission frequencies

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The OXC acts as an intermediary between GPUs, managing optical channel establishment independently. This mediator handles the channel setup process separately from GPU computation, allowing multiple GPUs to communicate efficiently without each transmission cycle requiring sequential channel establishment, thus mitigating the time loss from increased transmission frequency

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Reduces the overall duration of AI training by minimizing the time spent on channel establishment, thereby enhancing the efficiency of data transfer and computation.

Implementation Method 1

utilizing MEMS or SiP technology to establish channels ahead of time

Methodology Applied
Scientific EffectMEMS (Microelectromechanical Systems): Microelectromechanical Systems

Implementation Method 2

utilizing MEMS or SiP technology to establish channels ahead of time

Methodology Applied
Scientific EffectSiP (Silicon Photonics):

Data Source

PatentUS12430553B2AI training network and method
Publication Date: 2025.09.30 HUAWEI TECH CO LTD
  • US12430553B2 patent drawing
  • US12430553B2 patent drawing
  • US12430553B2 patent drawing

AI summary

An artificial intelligence training technology, applied to an artificial intelligence training network. Before graphics processing units located on different servers need to communicate with each other, an optical channel used for communication is established in advance. Once a graphics processing unit of a previous server completes calculation of the graphics processing unit, a calculation result can be immediately sent to a graphics processing unit on a next server without waiting or only in a short time period, to reduce duration of artificial intelligence training.