Novel data center network topology architecture for large model training
By designing ZCube topology and ZCube-partial topology, the problem that the existing data center network topology architecture cannot effectively support full-to-full-collection communication in large model training is solved, and efficient model training and good fault tolerance are achieved, which is highly cost-effective.
Patent Information
- Application Number
- CN202510190628.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing data center network topology architecture, especially the rail-optimized fat tree architecture, cannot effectively support the full-to-full-collection communication generated in large-scale model training, resulting in inefficient training of hybrid expert models and poor fault tolerance. Single-point network failures will greatly reduce the performance of ensemble communication.
A new data center network topology architecture is designed, called ZCube topology and ZCube-partial topology, which connects thousands to hundreds of thousands of GPUs through recursive construction methods, with low network diameter characteristics and strong scalability.
The ZCube topology can significantly accelerate the full-standard and full-to-full traffic in large-model training, improve computing resource utilization, reduce communication delay between nodes, and show good fault tolerance in the event of single-point network failure, which is cost-effective.
Smart Images

Figure CN120034480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data center network technology, and specifically to a novel data center network topology architecture for large model training. Background Art
[0002] With the rapid development of generative artificial intelligence, the number of model parameters and the amount of computing required for model training and inference have increased exponentially, resulting in the need for increasingly larger and more expensive computing clusters; the increase in the number of model parameters means that the video memory of a single GPU is usually unable to store the complete large model parameters. Training large models requires the use of a large number of GPUs at the same time and a series of parallel strategies; these parallel strategies will generate collective communication operations to synchronize model data on different GPUs; the most commonly used collective communication operation is all-reduce, which reduces and synchronizes data on multiple GPUs and is mainly used for data parallelism and tensor parallelism; in recent years, with the rise of the mixture-of-expert model architecture, expert parallelism is usually used when training such models. This parallel strategy uses all-to-all collective communication operations.
[0003] The network interconnection between GPUs will affect the efficiency of collective communication, and thus affect the efficiency of large model training. Generally, GPUs on the same server are interconnected through NVLink or PCIe, which is called a scale-up network. Different servers are interconnected through Ethernet, Infinite Band (IB) or RDMA over Converged Ethernet (RoCE), which is called a scale-out network, that is, a data center network. Different data center network topologies will have a huge impact on the efficiency of large model training and inference. How to design an economical and efficient data center network topology for large model training scenarios has become a key issue of concern in the industry.
[0004] Among the existing mainstream topology design schemes, the most advanced topology architecture is Rail-Optimized Fat-Tree; it connects GPUs with the same ID on different servers through rail switches, allowing them to communicate within 1 hop, thereby improving the performance of full-protocol collective communication. However, the Rail-Optimized Fat-Tree architecture is only optimized for full-protocol collective communication, and does not consider the all-to-all collective communication generated by the new model architecture, resulting in low efficiency in hybrid expert model training; in addition, the Rail-Optimized Fat-Tree architecture has poor fault tolerance, and a single point network failure will affect a large number of GPU communications, significantly reducing the collective communication performance.
[0005] Based on this, the present invention aims to design a new data center network topology to connect thousands to hundreds of thousands of GPUs in an efficient and economical manner. Summary of the invention
[0006] The purpose of the present invention is to provide a new data center network topology architecture for large model training to solve the problems raised in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions: a new data center network topology architecture for large model training, including ZCube topology and ZCube-partial topology, wherein the ZCube topology includes ZCube(n,1) and ZCube(n,k+1), wherein the ZCube(n,1) is composed of a switch connecting n GPUs, and the ZCube(n,k+1) is composed of n ZCube(n,k) and n k It consists of switches;
[0008] Each GPU in the ZCube (n, k+1) is configured with k+1 network cards or ports, denoted as level-0 to level-k, which are connected to the switches of the corresponding layers respectively;
[0009] The ZCube-partial topology includes ZCube(n,3)-partial, which includes n ZCube(n,2) and n 2 / 2 level-2 switches.
[0010] Preferably, the connection method between the network card on the GPU and the switch is as follows: mark n ZCubes (n, k) as 0th to n-1th ZCubes (n, k), mark the GPUs in each ZCube (n, k) as 0th to nk-1th GPUs, and connect the level-k network card of the i-th GPU in the j-th ZCube (n, k) to the j-th port of the i-th level-k switch.
[0011] Preferably, the switches in the ZCube (n, k) are connected to the switches in the ZCube (n, k+1) as follows: for the level-(k-1) switch, it will be interconnected with n level-k switches, and for the i-th level-(k-1) switch in the j-th ZCube (n, k), it will be interconnected with the n switches from i×n to (i+1)×n-1 in level-k respectively.
[0012] Preferably, the ZCube topology has a low network diameter characteristic, the network diameter of ZCube(n,k) is k, the network diameter of ZCube(n,k)-partial is k+1, and k is 2 to 4.
[0013] Preferably, the ZCube topology has strong scalability: a ZCube (42,2) can be constructed using a 128-port switch, and 3,111,696 GPUs can be interconnected within a network diameter of 4.
[0014] Compared with the prior art, the new data center network topology architecture ZCube designed by the present invention has the following advantages:
[0015] First, ZCube can accelerate the full-reduction traffic and full-to-full traffic generated during large model training, thereby accelerating model training and improving computing resource utilization;
[0016] Second, ZCube has the characteristics of low network diameter, low communication latency between nodes, and can quickly respond to user requests;
[0017] Third, ZCube has good fault tolerance and graceful performance degradation in the scenario of single-point network failure.
[0018] The present invention improves the model training efficiency and computing resource utilization by improving the full protocol and all-to-all set communication capabilities, while showing good fault tolerance in the face of network failures and having high cost-effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Schematic diagram of the ZCube topology construction method of the present invention;
[0020] Figure 2 This is a schematic diagram of the ZCube (3,2) topology of the present invention;
[0021] Figure 3 This is a schematic diagram of the ZCube (84, 3)-partial topology of the present invention. DETAILED DESCRIPTION
[0022] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] See also Figure 1-3The present invention provides a technical solution: a new data center network topology architecture for large model training, including ZCube topology and ZCube-partial topology. The ZCube topology adopts a recursive construction method, including ZCube (n, 1) and ZCube (n, k+1). ZCube (n, 1) is the smallest unit. The ZCube (n, 1) is composed of a switch connecting n GPUs. The ZCube (n, k+1) is composed of n ZCube (n, k) and n k It consists of switches;
[0024] Each GPU in the ZCube (n, k+1) is configured with k+1 network cards or ports, denoted as level-0 to level-k, which are connected to the switches of the corresponding layers respectively;
[0025] The ZCube-partial topology includes ZCube(n,3)-partial, which includes n ZCube(n,2) and n 2 / 2 level-2 switches.
[0026] In the present invention, the connection method between the network card on the GPU and the switch is as follows: mark n ZCubes (n, k) as 0th to n-1th ZCubes (n, k), mark the GPUs in each ZCube (n, k) as 0th to nk-1th GPUs, and connect the level-k network card of the i-th GPU in the j-th ZCube (n, k) to the j-th port of the i-th level-k switch.
[0027] In the present invention, the switches in the ZCube (n, k) are connected to the switches in the ZCube (n, k+1) as follows: for the level-(k-1) switch, it will be interconnected with n level-k switches, and for the i-th level-(k-1) switch in the j-th ZCube (n, k), it will be interconnected with the n switches from i×n to (i+1)×n-1 in level-k respectively.
[0028] In the present invention, the ZCube topology has a low network diameter characteristic, the network diameter of ZCube(n,k) is k, the network diameter of ZCube(n,k)-partial is k+1, k is 2 to 4, and the network diameter of the three-layer guide rail optimized fat tree is 5.
[0029] In the present invention, the ZCube topology has strong scalability: using a 128-port switch, ZCube (42,2) can be constructed, and 3,111,696 GPUs can be interconnected within a network diameter of 4; while using a 128-port switch to construct a three-layer rail optimized fat tree topology, only 524,288 GPUs can be interconnected, and the network diameter is 5.
[0030] The present invention: uses 3 servers and 6 Mellanox QM9790 IB switches to build a ZCube topology, each server contains 3 NVIDIA H800 GPUs, and the GPUs on the same server are interconnected through 200GB / s NVLink. 9 NVIDIA ConnectX-7 400GbE network cards are used to build a rail-optimized fat tree topology, and 18 NVIDIAConnectX-6 200GbE network cards are used to build a ZCube topology. In the NCCL 2.21.5 collection communication library, the full protocol and full-to-full communication performance are tested; the experimental results show that compared with the rail-optimized fat tree topology of the same scale and cost, ZCube can provide 31% acceleration for full protocol communication and 68% acceleration for full-to-full communication. Further, we built a large-scale ZCube topology and rail-optimized fat tree topology based on the simulation platform, which contains 4096 GPUs. Experimental results show that the ZCube topology has a 31% acceleration for GPT-3 175B training and a 49% acceleration for MoE-GPT training while saving 12% of network costs. In the event of a single-point network failure, the average communication path length extension of the ZCube topology is only 14% of that of the rail-optimized fat tree.
[0031] The present invention designs ZCube, a new data center network topology architecture for large model training. The topology has the characteristics of low network diameter, can provide powerful full-protocol and all-to-all collection communication capabilities, and at the same time exhibits good fault tolerance in the face of network failures, and is cost-effective; considering the limitations of GPU network cards, the present invention also constructs a variant of ZCube, ZCube-partial topology, which reduces the interconnection between some GPUs and switches on the basis of ZCube; ZCube of the present invention can provide acceleration for the full-protocol traffic and all-to-all traffic generated during large model training, thereby accelerating model training and improving computing resource utilization; ZCube has the characteristics of low network diameter, low communication latency between nodes, and can quickly respond to user requests; ZCube has good fault tolerance and has elegant performance degradation in the scenario of single-point network failure.
[0032] The contents not described in detail in this specification belong to the prior art known to professional and technical personnel in the field. Although the embodiments of the present invention have been shown and described, it is understood by ordinary technicians in the field that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the attached claims and their equivalents.
Claims
1. A new data center network topology architecture for large model training, characterized by: It includes ZCube topology and ZCube-partial topology. The ZCube topology includes ZCube (n, 1) and ZCube (n, k+1). The ZCube (n, 1) is composed of a switch connecting n GPUs, and the ZCube (n, k+1) is composed of n ZCube (n, k) and n k It consists of switches; Each GPU in the ZCube (n, k+1) is configured with k+1 network cards or ports, denoted as level-0 to level-k, which are connected to the switches of the corresponding layers respectively; The ZCube-partial topology includes ZCube(n,3)-partial, which includes n ZCube(n,2) and n 2 / 2 level-2 switches.
2. The novel data center network topology architecture for large model training according to claim 1 is characterized in that: The connection method between the network card on the GPU and the switch is as follows: mark n ZCubes (n, k) as 0th to n-1th ZCubes (n, k), mark the GPUs in each ZCube (n, k) as 0th to nk-1th GPUs, and connect the level-k network card of the i-th GPU in the j-th ZCube (n, k) to the j-th port of the i-th level-k switch.
3. The novel data center network topology architecture for large model training according to claim 1 is characterized in that: The switches in the ZCube (n, k) are connected to the switches in the ZCube (n, k+1) as follows: for the level-(k-1) switch, it will be interconnected with n level-k switches, and for the i-th level-(k-1) switch in the j-th ZCube (n, k), it will be interconnected with the n switches from i×n to (i+1)×n-1 in level-k respectively.
4. The novel data center network topology architecture for large model training according to claim 1 is characterized in that: The ZCube topology has a low network diameter characteristic. The network diameter of ZCube(n,k) is k, the network diameter of ZCube(n,k)-partial is k+1, and k is 2 to 4.
5. The novel data center network topology architecture for large model training according to claim 1 is characterized in that: The ZCube topology has strong scalability: a ZCube (42,2) can be constructed using a 128-port switch, which can interconnect 3,111,696 GPUs within a network diameter of 4.
Citation Information
Cited By
Intelligent computing cluster networking topology method and device, electronic equipment and storage medium
CN121357190A
Artificial intelligence cluster networking topology method and device, electronic equipment and storage medium
CN121357190B