A multi-gpu interconnect system
By employing a pre-defined connection relationship between a PCIe switch and a GPU chip set in a multi-GPU system, the problems of low data transmission efficiency and high cost are solved, achieving low-cost and high-efficiency GPU chip interconnection, which is suitable for training and inference tasks of large language models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD
- Filing Date
- 2024-11-21
- Publication Date
- 2026-05-29
AI Technical Summary
Existing multi-GPU systems suffer from low data transmission efficiency and high cost due to the bandwidth limitations of the interface between the PCIe switch and the CPU. This is especially true in large language model training and inference tasks, where achieving low-cost and efficient data transmission remains a challenge.
The architecture employs M PCIe switches and M GPU chip sets. By pre-setting connection relationships, the GPU chip sets are interconnected internally, reducing the dependence on CPU and PCIe switches. Interconnection is only achieved through the PCIe switches within the GPU chip sets, reducing the number of interconnection ports and cables.
It achieves efficient GPU chip interconnection, reduces costs, and ensures data transmission efficiency, making it suitable for training and inference tasks of large language models.
Smart Images

Figure CN122111899A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of GPU cluster design technology, and in particular to a multi-GPU interconnect system. Background Technology
[0002] With the rapid development of fields such as artificial intelligence and machine learning, the application scenarios of GPU clusters are constantly expanding. In these fields, GPU clusters are widely used for tasks such as image recognition, speech recognition, and natural language processing. Taking natural language processing tasks as an example, large language models have hundreds of billions or even trillions of parameters. Therefore, for the training and inference of large language models, distributed processing using multiple GPUs is necessary to complete the task within a reasonable timeframe. Thus, designing efficient multi-GPU interconnects based on application requirements has become a problem that needs to be solved.
[0003] In existing technologies, one approach involves interconnecting multiple GPU cards within a single server via PCIe switches, with interconnections between these PCIe switches to increase the scale of GPU card interconnection within the server. For example, GPU cards can communicate via the UPI bus between CPUs connected to each PCIe switch or via the PCIe interfaces between the PCIe switches. Building upon this, some methods propose adding additional connection methods, such as bridge boards, to improve the transmission bandwidth of multi-GPU interconnection within a single server when using PCIe switches for multi-card interconnection.
[0004] However, when multiple GPUs perform distributed data processing, if data transmission is performed using the PCIe interface between PCIe switches, the data transmission efficiency will be limited by the bandwidth of the PCIe interface between PCIe switches. If data transmission is performed using the UPI bus between CPUs, the data transmission efficiency will be limited by the bandwidth of the PCIe interface between the PCIe switch and the CPU, as well as the bandwidth of the UPI bus.
[0005] To address the aforementioned issues, existing technologies propose adding a new interconnect topology between all GPU chips contained in each server. This allows data transmission between servers to be achieved through the interconnect topology without the need for PCIe interfaces between PCIe switches or UPI buses between CPUs. However, existing interconnect topologies require each GPU chip to provide a large number of interconnect interfaces to connect with other GPU chips. For example, in a Hybrid Cube Mesh (HCM) architecture that interconnects eight GPU chips, each GPU chip needs to provide six interconnect interfaces. A large number of interconnect interfaces also means higher design and manufacturing costs for GPU chips and higher interconnection costs. Therefore, how to achieve a low-cost GPU interconnect system has become an urgent problem to be solved. Summary of the Invention
[0006] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:
[0007] A multi-GPU interconnect system, the system comprising: M PCIe switches {a1, a2, ..., a...} m , ..., a M}, a set of M GPU chips {B1, B2, ..., B m B M}, where a m Let B be the m-th PCIe switch, where m is an integer in the range [1, M]. m Let a be the m-th GPU chip set. m and B m There is a corresponding relationship.
[0008] GPU chip assembly B m Including N GPU chips {b m 1, b m 2, ..., b m n , ..., b m N}, where b m n GPU chip assembly B m The nth GPU chip in the set B of GPU chips. m The included N GPU chips are all related to B m The corresponding PCIe switch a m connect.
[0009] The nth GPU chip contained in each of the M GPU chip sets 1 n b 2 n ... b m n ... b M n The connection is established through a preset connection relationship, in which b s n and b s+1 n Connection, b M n and b 1 n Connect s, where s is an integer in the range [1, M-1].
[0010] Compared with the prior art, the present invention has significant advantages. Through the above technical solution, the multi-GPU interconnect system provided by the present invention achieves considerable technological progress and practicality, and has broad industrial application value. It has at least the following advantages:
[0011] This invention provides a multi-GPU interconnect system, the system comprising: M PCIe switches {a1, a2, ..., a...} m , ..., a M}, a set of M GPU chips {B1, B2, ..., B m B M}, where a m Let B be the m-th PCIe switch, where m is an integer in the range [1, M]. m Let a be the m-th GPU chip set. m and B m There is a corresponding relationship; GPU chip set B m Including N GPU chips {b m 1, b m 2, ..., b m n , ..., b m N}, where b m n GPU chip assembly B m The nth GPU chip in the set B of GPU chips. m The included N GPU chips are all related to B m The corresponding PCIe switch a m Connect the nth GPU chip b contained in each of the M GPU chip sets. 1 n b 2 n ... b m n ... b M n The connection is established through a preset connection relationship, in which b s n and b s +1 n Connection, b M n and b 1 n Connect s, where s is an integer in the range [1, M-1].
[0012] As can be seen, the interconnection between GPU chip sets is achieved through the preset connection relationship between the GPU chips contained in each GPU chip set. The interconnection of each GPU chip within a GPU chip set is carried out through a single PCIe switch corresponding to the GPU chip set. Data transmission between GPU chip sets does not need to go through the PCIe interface between the CPU or the PCIe switch, thus ensuring the data transmission efficiency of the GPU chip interconnection architecture. Moreover, a single GPU chip needs to be connected to at most two other GPU chips, reducing the wiring cost of GPU chip interconnection and enabling low-cost GPU chips to form GPU clusters. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a schematic diagram of the architecture design of a GPU cluster in a multi-GPU interconnect system provided by an embodiment of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] This embodiment provides a multi-GPU interconnect system, characterized in that the system includes: M PCIe switches {a1, a2, ..., a...} m , ..., a M}, a set of M GPU chips {B1, B2, ..., B m B M}, where a m Let B be the m-th PCIe switch, where m is an integer in the range [1, M]. m Let a be the m-th GPU chip set. m and B m There is a corresponding relationship;
[0017] GPU chip assembly B m Including N GPU chips {b m 1, b m 2, ..., b mn , ..., b m N}, where b m n GPU chip assembly B m The nth GPU chip in the set B of GPU chips. m The included N GPU chips are all related to B m The corresponding PCIe switch a m connect;
[0018] The nth GPU chip contained in each of the M GPU chip sets 1 n b 2 n ... b m n ... b M n The connection is established through a preset connection relationship, in which b s n and b s+1 n Connection, b M n and b 1 n Connect s, where s is an integer in the range [1, M-1].
[0019] Among them, PCIe switch can refer to a customizable, multi-port embedded PCIe switch that can support the connection of uplink ports and several downlink ports. GPU chip set can refer to a server formed by N GPU chips. The preset connection relationship can be implemented through high-speed interconnect interface, optical communication port, PCIe port, etc., without any restrictions.
[0020] Specifically, in this embodiment, the interconnection of N GPU chips in a single GPU chip set is achieved only through the PCIe switch corresponding to the GPU chip set. There is no need to interconnect the GPU chips within the GPU chip set through the interconnection between GPU chips, thereby reducing the number of GPU chip interconnection ports required and the number of interconnection cables between GPU chips, thus reducing interconnection costs.
[0021] When M=2, the preset connection relationship only contains b. 1 n and b 2 nIn contrast, in the existing HCM8 architecture, a single GPU chip only needs to connect to a single GPU chip in another GPU chip set. In this case, a single GPU chip needs to connect to other GPU chips in its own GPU chip set, as well as to GPU chips in another GPU chip set. The number of connections for GPU chip connections across GPU chip sets may be one or two.
[0022] When M is greater than 2, in the preset connection relationship, a single GPU chip only needs to be connected to the single GPU chips contained in the other two GPU chip sets respectively. That is, a single GPU chip corresponds to only two connections. In the scenario of using low-cost GPU chips, low-cost GPU chips can refer to GPU chips that contain only a small number of interconnect ports, which can still form a GPU cluster.
[0023] See Figure 1 This is a schematic diagram of the architecture of a GPU cluster design in a multi-GPU interconnect system provided by an embodiment of the present invention. The schematic diagram takes the case of two GPU chip sets and their corresponding PCIe switches, with each GPU chip set containing four GPU chips and eight GPU chips interconnected as an example.
[0024] It should be noted that this embodiment assumes that each GPU chip set contains the same number of GPU chips, and that any two GPU chips in a single GPU chip set are equivalent to each other. Implementers can adjust the identifier of a GPU chip in its respective GPU chip set according to the actual situation.
[0025] In one specific implementation, the nth GPU chip b included in the M GPU chip sets 1 n b 2 n ... b m n ... b M n Connections are made through preset connection relationships, including:
[0026] The nth GPU chip b contained in each of the M GPU chip sets 1 n b 2 n ... b m n ... b M n A pre-defined connection relationship is established through a high-speed interconnect interface.
[0027] In one specific implementation, GPU chip b mn Corresponding to the sub-data c to be processed m n ;
[0028] GPU chip b m n For its corresponding sub-data c to be processed m n After processing, we get b m n The corresponding processed sub-data d m n .
[0029] In scenarios where GPU clusters are used to process large model tasks, the sub-data c to be processed m n This could refer to a subset of parameters in a large model inference or training task, or the processed sub-data d. m n It can refer to the sub-data c to be processed. m n The result obtained through forward computation in the inference task or through backward gradient computation in the training task.
[0030] In one specific implementation, the nth GPU chip b is included in each of the M GPU chip sets. 1 n b 2 n ... b m n ... b M n Through the preset connection relationship, d is performed. 1 n d 2 n 、…、d m n 、…、d M n The full aggregation operation yields the first intermediate data e. n b 1 n b 2 n ... b m n ... b M n All contain the intermediate data e n .
[0031] Where, d 1 n d 2 n 、…、dm n 、…、d M n The full aggregation operation can refer to the operation of d 1 n d 2 n 、…、d m n 、…、d M n In b 1 n b 2 n ... b m n ... b M n The average is aggregated into intermediate data e n .
[0032] In one specific implementation, with the m-th PCIe switch a m Connected N GPU chips b m 1, b m 2, ..., b m n , ..., b m N , through a m Perform the first intermediate data e1, e2, ..., e n , ..., e N The full aggregation operation yields the target data F, b. m 1, b m 2, ..., b m n , ..., b m N All of them contain the target data F.
[0033] Where e1, e2, ..., e n , ..., e N The full aggregation operation can refer to the aggregation of e1, e2, ..., e n , ..., e N In GPU chip set B m The target data F is aggregated in each of the included GPU chips.
[0034] In one specific implementation, with the m-th PCIe switch a m Connected N GPU chips b m 1, b m 2, ..., b m n , ..., b m N , through am Perform d m 1,d m 2, ..., d m n , ..., d m N The full aggregation operation yields the second intermediate data g. m b m 1, b m 2, ..., b m n , ..., b m N All contain the second intermediate data g m .
[0035] Where, d m 1,d m 2, ..., d m n , ..., d m N The full aggregation operation can refer to the operation of d m 1,d m 2, ..., d m n , ..., d m N In GPU chip set B m The data is aggregated into a second intermediate data g in each of the included GPU chips. m .
[0036] In one specific implementation, the nth GPU chip b is included in each of the M GPU chip sets. 1 n b 2 n ... b m n ... b M n Through the preset connection relationship, g1, g2, ..., g m , ..., g M The full aggregation operation is performed to obtain the target data F, b. 1 n b 2 n ... b m n ... b M n All of them contain the target data F.
[0037] Where g1, g2, ..., g m , ..., g M The full aggregation operation can refer to the aggregation of g1, g2, ..., g m, ..., g M In b 1 n b 2 n ... b m n ... b M n The data are all aggregated into the target data F.
[0038] Both of the above-mentioned full aggregation methods can realize data interconnection and transmission between all GPU chips in each GPU chip set, and implementers can choose to use them according to the actual situation.
[0039] In one embodiment, multiple communication loops can be formed by PCIe switches and GPU chips. Within each communication loop, a single GPU chip is connected to two units, which can be either PCIe switches or GPU chips themselves. GPU chips belonging to the same GPU chip set can be connected via PCIe switches, while GPU chips not belonging to the same set can be connected via pre-defined connections. When using communication loops for data interconnection and transmission, it is not necessary to use the two fully aggregated methods described above. A ring algorithm effect similar to a hybrid three-dimensional interconnect can be achieved by switching between PCIe switches and GPU chip interconnections. The implementer can determine the specific method of data interconnection and transmission based on the actual interconnection topology.
[0040] In one specific implementation, M is set to 2 and N is set to 4.
[0041] When M is set to 2 and N is set to 4, it corresponds to a structure where 8 GPU chips are interconnected.
[0042] In this embodiment, the interconnection between GPU chip sets is achieved through the preset connection relationship between the GPU chips contained in each GPU chip set. The interconnection of each GPU chip within the GPU chip set is carried out through a single PCIe switch corresponding to the GPU chip set. Data transmission between GPU chip sets does not need to go through the PCIe interface between the CPU or the PCIe switch, thereby ensuring the data transmission efficiency of the GPU chip interconnection architecture. Moreover, a single GPU chip needs to be connected to at most two other GPU chips, which reduces the wiring cost of GPU chip interconnection and enables low-cost GPU chips to form GPU clusters.
[0043] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.
Claims
1. A multi-GPU interconnect system, characterized in that, The system includes: M PCIe switches {a1, a2, ..., a...} m , ..., a M }, a set of M GPU chips {B1, B2, ..., B m B M }, where a m Let B be the m-th PCIe switch, where m is an integer in the range [1, M]. m Let a be the m-th GPU chip set. m and B m There is a corresponding relationship; GPU chip assembly B m Including N GPU chips {b m 1, b m 2, ..., b m n , ..., b m N }, where b m n GPU chip assembly B m The nth GPU chip in the set B of GPU chips. m The included N GPU chips are all related to B m The corresponding PCIe switch a m connect; The nth GPU chip contained in each of the M GPU chip sets 1 n b 2 n ... b m n ... b M n The connection is established through a preset connection relationship, in which b s n and b s+1 n Connection, b M n and b 1 n Connect s, where s is an integer in the range [1, M-1].
2. The multi-GPU interconnect system according to claim 1, characterized in that, The M GPU chip sets each contain the nth GPU chip b 1 n b 2 n ... b m n ... b M n Connections are made through preset connection relationships, including: The M GPU chip sets each contain the nth GPU chip b 1 n b 2 n ... b m n ... b M n A pre-defined connection relationship is established through a high-speed interconnect interface.
3. The multi-GPU interconnect system according to claim 1, characterized in that, GPU chip b m n Corresponding to the sub-data c to be processed m n ; GPU chip b m n For its corresponding sub-data c to be processed m n After processing, we get b m n The corresponding processed sub-data d m n .
4. The multi-GPU interconnect system according to claim 3, characterized in that, The nth GPU chip contained in each of the M GPU chip sets 1 n b 2 n ... b m n ... b M n Through the preset connection relationship, d is performed. 1 n d 2 n 、…、d m n 、…、d M n The full aggregation operation yields the first intermediate data e. n b 1 n b 2 n ... b m n ... b M n All contain the intermediate data e n .
5. The multi-GPU interconnect system according to claim 4, characterized in that, With the m-th PCIe switch a m Connected N GPU chips b m 1, b m 2, ..., b m n , ..., b m N , through a m Perform the first intermediate data e1, e2, ..., e n , ..., e N The full aggregation operation yields the target data F, b. m 1, b m 2, ..., b m n , ..., b m N All of them contain the target data F.
6. The multi-GPU interconnect system according to claim 3, characterized in that, With the m-th PCIe switch a m Connected N GPU chips b m 1, b m 2, ..., b m n , ..., b m N , through a m Perform d m 1,d m 2, ..., d m n , ..., d m N The full aggregation operation yields the second intermediate data g. m b m 1, b m 2, ..., b m n , ..., b m N All contain the second intermediate data g m .
7. The multi-GPU interconnect system according to claim 6, characterized in that, The nth GPU chip contained in each of the M GPU chip sets 1 n b 2 n ... b m n ... b M n Through the preset connection relationship, g1, g2, ..., g m , ..., g M The full aggregation operation is performed to obtain the target data F, b. 1 n b 2 n ... b m n ... b M n All of them contain the target data F.
8. The multi-GPU interconnect system according to claim 1, characterized in that, M is set to 2, and N is set to 4.