Collaborative training method for computing system, computing device, computer readable storage medium and computer program product

By designing an abstraction layer and a global-local communication group in the computing system, the problem of resource reuse in heterogeneous GPU systems is solved, and efficient heterogeneous GPU communication and performance optimization are achieved.

CN120950443APending Publication Date: 2025-11-14SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411661014.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In heterogeneous GPU systems, existing technologies struggle to efficiently reuse GPU resources from different manufacturers and models, resulting in complex communication implementations, high costs, and difficult maintenance.

Method used

Design an abstraction layer below the computational model layer and framework layer, and above the hardware layer, to enable collaborative training between GPUs in heterogeneous systems. Through the creation of global and local communication groups, integrate communication libraries of different architectures for lossless integration.

Benefits of technology

It enables seamless communication bridging between GPUs from different manufacturers and models, improving resource utilization efficiency and reducing maintenance difficulty and performance loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950443A_ABST
    Figure CN120950443A_ABST
Patent Text Reader

Abstract

The invention provides a cooperative training method for a computing system, computing equipment, a computer readable storage medium and a computer program product. The method comprises the steps of creating a global communication group for all GPUs of the computing system for a training task so as to be used for global communication of all GPUs of the computing system; determining a GPU corresponding to at least one subtask of the training task; for each sub-task, determining whether the GPU corresponding to the sub-task contains a heterogeneous GPU or not; and if the GPUs corresponding to the subtasks comprise heterogeneous GPUs, creating a local heterogeneous communication group for the heterogeneous GPUs for communication among the heterogeneous GPUs in the execution process of the subtasks, and respectively creating a first local isomorphic communication group for each type of GPU in the heterogeneous GPUs for communication between the types of GPUs in the subtask execution process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of processors, and more specifically, to a method for collaborative training of computing systems, a computing device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] With the rapid development of artificial intelligence technology, especially the rise of large-scale models, the demand for computing resources has exploded. The size of these models, the scale of training data, and the number of graphics processing units (GPUs) required are all growing exponentially. In some cases, thousands or even tens of thousands of GPUs are needed to meet training requirements. However, in the current GPU cloud service and resource usage environment, these tens of thousands of GPUs may come from different manufacturers or different product models from the same manufacturer, with different hardware architectures—that is, they are heterogeneous.

[0003] Against this backdrop, how to efficiently reuse existing resource pools and mix GPUs of different models and manufacturers to support the training of large-scale models has become a hot research topic in the industry.

[0004] For homogeneous systems consisting of multiple identical GPUs, the communication library designed by the GPU manufacturer is typically used to solve the communication problem between the multiple GPUs.

[0005] Therefore, one current solution for heterogeneous systems composed of multiple different types of GPUs is to utilize the idea of ​​a unified communication library, that is, to design a unified communication library for GPUs from different manufacturers and models so that they can all strictly follow the same interface definition and communication implementation.

[0006] However, this approach presents several significant challenges in its implementation. For instance, GPUs from different manufacturers have varying hardware architectures and corresponding incompatible software stacks, making the adaptation of a unified communication library complex and costly. Furthermore, the differences in communication implementations and hardware parameters across GPUs, along with variations in computing power and memory size, lead to distinct communication implementations and optimization strategies for each GPU. Consequently, using a unified communication library necessitates that at least some GPUs sacrifice some performance to meet the requirements of others.

[0007] Furthermore, the characteristics of various GPUs need to be considered when maintaining such a unified communications library. With the development of technology and the introduction of new GPU models, the difficulty of maintaining and updating the unified communications library continues to increase. Summary of the Invention

[0008] To address at least one of the aforementioned problems, the present disclosure provides a solution that enables collaborative training between GPUs in a heterogeneous system by designing an abstraction layer below the computational model and framework layers and above the hardware layer. This abstraction layer can reuse and integrate communication libraries from GPUs with different architectures, achieving lossless integration and functional expansion of these communication libraries.

[0009] According to one aspect of this disclosure, a collaborative training method for a computing system is provided. The method includes: creating a global communication group for all GPUs of the computing system for a training task to facilitate global communication among all GPUs of the computing system; determining GPUs corresponding to at least one subtask of the training task; for each subtask, determining whether the GPUs corresponding to the subtask include heterogeneous GPUs; and if the GPUs corresponding to the subtask include heterogeneous GPUs, creating a local heterogeneous communication group for the heterogeneous GPUs to facilitate communication between the heterogeneous GPUs during the execution of the subtask, and creating a first local homogeneous communication group for each type of GPU in the heterogeneous GPUs to facilitate communication between GPUs of that type during the execution of the subtask.

[0010] In some implementations, creating a global communication group includes: sensing the types of all GPUs in the computing system; determining whether the computing system is a heterogeneous system including multiple types of GPUs based on the sensed types of all GPUs; and if the computing system is determined to be a heterogeneous system, creating the global heterogeneous communication group.

[0011] In some implementations, perceiving the type of all GPUs in the computing system includes: during process initialization for the training task, acquiring scale information of the computing system, the scale information including the number of GPUs contained in the computing system and the global parallelism of the training task; and acquiring device information of each GPU, wherein the device information includes at least the type of the GPU.

[0012] In some implementations, the method further includes: if it is determined that the computing system is a homogeneous system, creating a global homogeneous communication group for all GPUs based on the type of GPUs in the computing system.

[0013] In some implementations, the method further includes: if the GPU corresponding to the subtask contains only homogeneous GPUs, then creating a second local homogeneous communication group for the homogeneous GPUs to be used for communication between the homogeneous GPUs during the execution of the subtask.

[0014] In some implementations, the heterogeneous system includes multiple GPUs of type 1 and multiple GPUs of type 2.

[0015] In some implementations, the type of GPU is determined by the GPU's manufacturer or product model.

[0016] In some implementations, the method further includes: performing a Reduce-Scatter reduction distribution operation on slices of input data for the subtask using a first local homogeneous communication group created for each type of GPU to obtain the reduction result of the slice data; one GPU of each type of heterogeneous GPU performing an AllReduce full reduction operation on both GPUs through the local heterogeneous communication group with one GPU of the other type, so that the two GPUs each have the reduction results of the slice data of the two GPUs; and each type of GPU performing an All-Gather full collection operation on its respective reduction results through the first local homogeneous communication group created for the type of GPU to obtain the full reduction result of the input data for the subtask.

[0017] According to another aspect of the present invention, a computing device is provided, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the computing device to perform the steps of the method described above.

[0018] According to another aspect of the present invention, a computing system is provided, comprising: a plurality of GPUs of different types; and a host, wherein the host runs a software stack, the top layer of the software stack being a computing model layer including various computing models, the computing model layer supporting a learning framework layer for the computing models during execution, and below the learning framework layer being a communication layer for providing a communication bridge between the computing model layer and the learning framework layer and the plurality of GPUs of different types, wherein the communication layer is configured to cause the host to perform the steps of the method described above.

[0019] According to another aspect of this disclosure, a computer-readable storage medium is provided having computer program code stored thereon, which, when run, performs the method described above.

[0020] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a machine, performs the method described above.

[0021] By using the scheme disclosed herein, a communication abstraction layer can be designed in the software stack to provide communication semantics compatible with existing frameworks, thereby enabling communication bridging between multiple heterogeneous GPUs from different manufacturers and models without the user's awareness. Attached Figure Description

[0022] This disclosure will be better understood by referring to the following description of specific embodiments given in the accompanying drawings, and other objects, details, features, and advantages of this disclosure will become more apparent.

[0023] Figure 1 A schematic diagram of an exemplary computing system is shown.

[0024] Figure 2 An exemplary flowchart of a collaborative training method for a computing system according to an embodiment of the present invention is shown.

[0025] Figure 3 An exemplary flowchart of a process for creating a global communication group according to an embodiment of the present invention is shown.

[0026] Figure 4 An exemplary flowchart illustrating the execution process of a subtask according to an embodiment of the present invention is shown.

[0027] Figure 5 An example is shown Figure 4 The diagram shows the GPU corresponding to the subtask and the data processing results.

[0028] Figure 6 A block diagram of a computing device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0029] Preferred embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0030] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one embodiment" and "some embodiments" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects.

[0031] Figure 1 A schematic diagram of an exemplary computing system 100 is shown. The computing system 100 may include multiple GPUs 110 of different types, i.e., heterogeneous GPUs. Figure 1As shown, the computing system 100 includes multiple GPUs 110-1 of a first type and multiple GPUs 110-2 of a second type. In this document, a GPU may be, for example, a GPU chip or a GPU board.

[0032] In addition, the computing system 100 may also include a host 120. The host 120 is connected to multiple GPUs 110 to control these GPUs 110 to perform various computing tasks.

[0033] The host 120 runs a software stack. At the top of the software stack is a computational model layer 122, which includes various computational models. Below the computational model layer 122 is a learning framework layer 124 that supports these computational models. The computational model layer 122 may include, for example, AIGC (Artificial Intelligence Generated Content) models, LLM (Large Language Model), and AI4S (AI for Science) models. The learning framework layer 124 may include, for example, open-source deep learning frameworks such as PyTorch for machine learning and deep learning.

[0034] Furthermore, below the learning framework layer 124 is the communication layer 126 proposed in this invention. The communication layer 126 provides a communication bridge between the computation model layer 122 and the learning framework layer 124 and the hardware layer composed of the various GPUs 110. In the case where the computing system 100 includes heterogeneous GPUs, the communication layer 126, in addition to handling homogeneous GPUs 110 (such as...), also handles communication between the heterogeneous GPUs 110 and the GPUs 110. Figure 1 The communication between the first type of GPU 110-1 shown in the diagram is also responsible for communication between heterogeneous GPUs 110 (such as...). Figure 1 The communication between the first type of GPU 110-1 and the second type of GPU 110-2 shown in the figure.

[0035] Furthermore, the software stack on host 120 may also include a heterogeneous collaborative training splitting layer (not shown) for splitting the training task based on the heterogeneous topology of GPU 110. This splitting layer may be part of communication layer 126, or it may be located between computation model layer 122 and learning framework layer 124.

[0036] Figure 2 An exemplary flowchart of a collaborative training method 200 for a computing system 100 according to an embodiment of the present invention is shown. Method 200 can run on a host 120, for example, it can be implemented by the communication layer 126 of the software stack of host 120. Method 200 can implement collaborative training of heterogeneous GPUs, also known as heterogeneous training.

[0037] like Figure 2 As shown in block 210, communication layer 126 can create a global communication group for all GPUs 110 of computing system 100 for a training task to facilitate global communication among all GPUs 110 of computing system 100.

[0038] This global communication group plays a central role throughout the entire lifecycle of the distributed training task, responsible for coordinating global communication and synchronization tasks.

[0039] Figure 3 An exemplary flowchart (block 210) of the process for creating a global communication group according to an embodiment of the present invention is shown.

[0040] like Figure 3 As shown in block 212, the communication layer 126 can be aware of the types of all GPUs 110 in the computing system 100. Here, the type of GPU 110 is determined by the manufacturer or product model of the GPU 110. For example, assuming... Figure 1 The first type of GPU 110-1 shown is a GPU manufactured by one company, and the second type of GPU 110-2 is a GPU manufactured by another company. Alternatively, suppose the first type of GPU 110-1 is a model of GPU manufactured by one company, and the second type of GPU 110-2 is a different model of GPU manufactured by the same company.

[0041] Furthermore, the software stack of the present invention can be used in any computing system 100. Therefore, it is assumed here that the topology of the computing system 100 is unknown in advance as to whether it is a homogeneous or heterogeneous system, and needs to be determined through topology awareness.

[0042] In one specific implementation, the communication layer 126 can obtain the scale information of the computing system 100 during the process initialization for a training task. This scale information may include, for example, the number of GPUs 110 included in the computing system 100 and the global parallelism of the training task.

[0043] Then, each GPU 110 can, for example, be driven by a training script to transmit its own device information to the communication layer 126, wherein the device information includes at least the type of the GPU 110.

[0044] In addition, in some embodiments, the device information may also include the connection relationship and connection type between the various GPUs 110, wherein the connection relationship indicates which one or more other GPUs 110 are connected to one GPU 110, and the connection type indicates the connection method used between the connected GPUs, such as PCIe interconnect, NVLink interconnect, NVSwitch interconnect, etc.

[0045] Next, in box 214, it can be determined whether computing system 100 is a heterogeneous system that includes multiple types of GPUs 110 based on the types of all GPUs 110 perceived in box 212.

[0046] As mentioned earlier, if the type of GPU 110 perceived in box 212 is different (e.g. Figure 1 As shown in the diagram, if the GPUs (including the first type and the second type) are all of the same type, then the computing system 100 can be determined to be a heterogeneous system. Conversely, if the GPUs 110 perceived in box 212 are of the same type, then the computing system 100 can be determined to be a homogeneous system.

[0047] If it is determined that computing system 100 is a heterogeneous system, then in box 216, a global heterogeneous communication group can be created for all GPUs 110.

[0048] Here, the global heterogeneous communication group can be established using a CPU-based communication backend, such as Gloo. Gloo is an open-source library focused on collective communication, providing algorithms for machine learning applications, including barriers, broadcasts, and all-reduce.

[0049] On the other hand, if it is determined that the computing system 100 is a homogeneous system, then in block 218, a global homogeneous communication group can be created for all GPUs 110 based on the type of GPU 110.

[0050] Here, the global homogeneous communication group can be determined based on the type of GPU 110. For example, if all GPU 110s are Nvidia GPUs, the global homogeneous communication group can be created based on NCCL.

[0051] continue Figure 2 After the global communication group is established in box 210, in box 220, the GPU required for at least one subtask of the training task is determined.

[0052] In some embodiments, the training task can be divided into at least one subtask based on the dimension of the training task and the GPU 110 included in the computing system 100. This division can be performed, for example, by a user of the host 120.

[0053] For example, when the computing system 100 contains 16 GPUs 110 and the training task is a three-dimensional parallel training task, the training task can be divided into three sub-tasks: Tensor Parallelism (TP) performed by 4 GPUs 110, Pipeline Parallelism (PP) performed by 2 GPUs 110, and Data Parallelism (DP) performed by 2 GPUs 110.

[0054] Meanwhile, the number of GPUs 110 required to execute each subtask can be determined by the rank of these GPUs 110.

[0055] Next, in box 230, for each subtask, it can be determined whether the GPU 110 corresponding to the subtask contains a heterogeneous GPU.

[0056] As mentioned earlier, during the process initialization for this training task, device information for each GPU 110 can be obtained, including the type of the GPU 110.

[0057] Therefore, in block 230, the type of GPU 110 corresponding to the subtask can be determined based on the acquired device information of each GPU 110, thereby determining whether these GPUs 110 contain heterogeneous GPUs.

[0058] For example, such as Figure 1 As shown, if it is determined that the GPU 110 corresponding to the subtask includes both GPU 110-1 of the first type and GPU 110-2 of the second type, then it can be determined that the GPU corresponding to the subtask contains heterogeneous GPUs.

[0059] In this case, in box 240, local heterogeneous communication groups can be created for these heterogeneous GPUs 110 for communication between these heterogeneous GPUs during the execution of this subtask.

[0060] In addition, in some embodiments, besides creating the local heterogeneous communication group for all heterogeneous GPUs 110 corresponding to the subtask, in block 240, a first local homogeneous communication group is also created for each type of GPU 110 among all heterogeneous GPUs 110 corresponding to the subtask for communication between GPUs of that type during the execution of the subtask.

[0061] On the other hand, if it is determined that the GPU 110 corresponding to the subtask only contains homogeneous GPUs, such as only GPU 110-1 of the first type or GPU 110-2 of the second type, then in block 250, a second local homogeneous communication group can be built for these homogeneous GPUs for communication between these homogeneous GPUs during the execution of the subtask.

[0062] Similarly, the second local homogeneous communication group can also be determined based on the type of these GPUs 110. For example, if all of these homogeneous GPUs 110 are Nvidia GPUs, the second local homogeneous communication group can be created based on NCCL.

[0063] The communication groups described above, including global homogeneous communication groups, global heterogeneous communication groups, first local homogeneous communication groups, second local homogeneous communication groups, and local heterogeneous communication groups, are used to handle communication between corresponding GPUs 110 during the execution of different subtasks. Each communication group defines various communication operations during the subtask execution process, such as Send, Receive, and AllReduce, also known as communication primitives. Furthermore, these communication operations may also include Barrier, Broadcast, Scatter, Gather, All-Gather, Reduce-Scatter, Reduce, and All-To-All.

[0064] Figure 4 An exemplary flowchart of the execution process 400 of a subtask according to an embodiment of the present invention is shown. The use of various communication groups is described below in conjunction with this execution process.

[0065] Here, we assume that the subtask requires a full reduction operation of 2n GPUs (n is a positive integer, and further, n = 2, 4, 8, 16...), and that the 2n GPUs include n GPUs of type 110-1 (hereinafter referred to as the first group of GPUs) and n GPUs of type 210-2 (hereinafter referred to as the second group of GPUs). Figure 5 An exemplary diagram illustrates the GPU corresponding to this subtask and its data processing results. Figure 5In this example, the subtask requires a full reduction operation of 16 GPUs (i.e., n=8), with the first group of GPUs comprising 8 GPUs of type 110-1 and the second group comprising 8 GPUs of type 210-2. The 8 GPUs of type 110-1 in the first group are labeled B1, B2, ..., B8, and the 8 GPUs of type 210-2 in the second group are labeled N1, N2, ..., N8. Furthermore, assuming that through the execution of method 200, a first local homogeneous communication group Group(B1, B2, ..., B8) (hereinafter referred to as G1) is created for the first group of GPUs, another first local homogeneous communication group Group(N1, N2, ..., N8) (hereinafter referred to as G2) is created for the second group of GPUs, and a local heterogeneous communication group Group(B1, B2, ..., B8, N1, N2, ..., N8) (hereinafter referred to as G3) is created for these 16 heterogeneous GPUs.

[0066] like Figure 4 As shown in block 410, a reduce-scatter operation is performed on the slice data of the input data of the subtask using a first local homogeneous communication group created for each type of GPU to obtain the reduction result of the slice data.

[0067] Specifically, firstly, the input data can be divided into multiple slices based on the number of first-type GPUs 110-1 in the first group of GPUs and the number of second-type GPUs 110-2 in the second group of GPUs.

[0068] For example, assuming the input data for this subtask is D, and each GPU group contains 8 GPUs 110, the input data D is divided into 8 slices S1, S2...S8, and each GPU 110-1 and 110-2 processes one slice of data. Figure 5 In order to distinguish which GPU the slice data belongs to, they are represented as B_S1, B_S2...B_S8 and N_S1, N_S2...N_S8 respectively.

[0069] In box 410, for each subtask, each first-type GPU 110-1 performs a reduce-scatter operation on the corresponding slice data S1, S2...S8 through the first local homogeneous communication group G1 to obtain the reduction results B_RS1, B_RS2...B_RS8 of the corresponding slice data, such as... Figure 5 As shown in the image.

[0070] Simultaneously, each of the second-type GPUs 110-2 performs a reduce-scatter operation on the corresponding slice data S1, S2...S8 through the first local homogeneous communication group G2 to obtain the reduction results N_RS1, N_RS2...N_RS8 of the corresponding slice data, such as... Figure 5 As shown in the image.

[0071] Next, in box 420, each type of GPU corresponding to a slice of data performs an AllReduce operation via the local heterogeneous communication group G3, so that each GPU has the reduction result of its own slice of data.

[0072] Specifically, each first-type GPU 110-1 performs a full reduction operation on a corresponding second-type GPU 110-2 via a local heterogeneous communication group G3, so that each GPU has the reduction results of its own slice data. For example, as... Figure 5 As shown, after a full reduction operation is performed on GPU 110-1 (e.g., B1) in the first group of GPUs and GPU 110-2 (e.g., N1) in the second group of GPUs, B1 and N1 will respectively have the reduction result B_RS1+N_RS1, and so on.

[0073] Finally, in box 430, each type of GPU performs an all-gather operation on its respective reduction result through a first local homogeneous communication group created for that type of GPU to obtain the fully reduced result of the input data D for that subtask.

[0074] Specifically, each first-type GPU 110-1 performs an all-gather operation on the reduction results of all first-type GPUs 110-1 through a first local homogeneous communication group G1 to obtain the reduction result of the input data D for the subtask, and each second-type GPU 110-2 performs an all-gather operation on the reduction results of all second-type GPUs 110-2 through another first local homogeneous communication group G2 to obtain the fully reduced result of the input data D for the subtask.

[0075] For example, after the first group of GPUs performs a full collection operation, each GPU 110-1 will have a full reduction result of (B_RS1+N_RS1)+(B_RS2+N_RS2)+……+(B_RS8+N_RS8). Figure 5 For convenience, the letter A will be used to represent the middle part.

[0076] Similarly, after the second group of GPUs performs the full collection operation, each GPU 110-2 will also have the full reduction result (B_RS1+N_RS1)+(B_RS2+N_RS2)+……+(B_RS8+N_RS8). Figure 5 The letter A is also used to represent it.

[0077] Figure 6 A block diagram of a computing device 600 suitable for implementing embodiments of the present disclosure is shown. The computing device 600 can be used to implement, for example... Figure 1 The host 120 is shown in the image. Figure 6 As shown, computing device 600 may include processor 610. Processor 610 controls the operation and functions of computing device 600. For example, in some embodiments, processor 610 may perform various operations by means of instructions 630 stored in memory 620 coupled thereto. Memory 620 may be any suitable type suitable for the local technical environment and may be implemented using any suitable data storage technology, including but not limited to semiconductor-based memory devices, magnetic storage devices and systems, optical storage devices and systems. Although Figure 6 Only one memory 620 is shown, but those skilled in the art will understand that the computing device 600 may include more physically different memories 620.

[0078] Processor 610 can be any suitable type for the local technical environment and can include, but is not limited to, one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and a processor-based multi-core processor architecture. Computing device 600 may also include multiple processors 610.

[0079] The present invention utilizes an abstraction layer designed below the computational model layer and above the hardware layer to achieve collaborative training between GPUs in a heterogeneous system. This abstraction layer can reuse and integrate communication libraries of GPUs with different architectures, enabling lossless integration and functional expansion of these communication libraries.

[0080] This invention can be implemented as a method, a computing device, a computer-readable storage medium, and / or a computer program product. The computer-readable storage medium stores computer program code, which, when run, performs the methods of this disclosure. The computer program product includes a computer program that, when run, performs the methods of this disclosure. The computing device may include at least one processor and at least one memory coupled to the at least one processor, the memory storing instructions for execution by the at least one processor. When the instructions are executed by the at least one processor, the computing device can perform the methods described above.

[0081] In one or more exemplary designs, the functions described herein may be implemented using hardware, software, firmware, or any combination thereof. For example, if implemented in software, the functions may be stored as one or more instructions or code on a computer-readable medium, or transmitted as one or more instructions or code on a computer-readable medium.

[0082] Those skilled in the art should also understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with embodiments of this disclosure can be implemented as electronic hardware, computer software, or a combination of both.

[0083] The foregoing description of this disclosure is intended to enable any person skilled in the art to implement or use this disclosure. Various modifications to this disclosure will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other variations without departing from the spirit and scope of this disclosure. Therefore, this disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope of the principles and novel features disclosed herein.

Claims

1. A method for collaborative training of a computing system, comprising: For a training task, a global communication group is created for all GPUs of the computing system for global communication among all GPUs of the computing system; Identify the GPU corresponding to at least one subtask of the training task; For each subtask, determine whether the GPU corresponding to the subtask contains heterogeneous GPUs; as well as If the GPU corresponding to the subtask includes heterogeneous GPUs, a local heterogeneous communication group is created for the heterogeneous GPUs to facilitate communication between the heterogeneous GPUs during the execution of the subtask, and a first local homogeneous communication group is created for each type of GPU in the heterogeneous GPUs to facilitate communication between GPUs of that type during the execution of the subtask.

2. The method of claim 1, wherein creating a global communication group comprises: Perceive the type of all GPUs in the computing system; Determine whether the computing system is a heterogeneous system that includes multiple types of GPUs based on the types of all GPUs perceived; as well as If the computing system is determined to be a heterogeneous system, the global heterogeneous communication group is created.

3. The method of claim 2, wherein sensing the type of all GPUs in the computing system includes: During process initialization for the training task, the scale information of the computing system is obtained, including the number of GPUs contained in the computing system and the global parallelism of the training task. as well as Obtain device information for each GPU, wherein the device information includes at least the type of the GPU.

4. The method of claim 2, further comprising: If the computing system is determined to be a homogeneous system, a global homogeneous communication group is created for all GPUs based on the type of GPUs in the computing system.

5. The method of claim 1, further comprising: If the GPU corresponding to the subtask only contains homogeneous GPUs, then a second local homogeneous communication group is created for the homogeneous GPUs to be used for communication between the homogeneous GPUs during the execution of the subtask.

6. The method of claim 1, wherein the heterogeneous system comprises a plurality of first-type GPUs and a plurality of second-type GPUs.

7. The method of claim 2, wherein the type of the GPU is determined by the manufacturer or product model of the GPU.

8. The method of claim 1, further comprising: In the case where the subtask is a full reduction operation requiring 2n GPUs, and the 2n GPUs include n first-type GPUs and n second-type GPUs, the Reduce-Scatter reduction distribution operation is performed on the slice data of the input data of the subtask using the first local homogeneous communication group created for each type of GPU to obtain the reduction result of the slice data; Each type of GPU corresponding to a slice of data performs an AllReduce full reduction operation through the local heterogeneous communication group, so that each GPU has the reduction result of its own slice of data; and Each type of GPU performs an All-Gather operation on its respective reduction results through a first local homogeneous communication group created for that type of GPU to obtain the full reduction results of the input data of the subtask.

9. A computing device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, causing the computing device to perform the steps of the method according to any one of claims 1 to 8.

10. A computing system, comprising: Multiple different types of GPUs; and A host computer, wherein a software stack runs on the host computer, the top layer of the software stack being a computation model layer including various computation models, the computation model layer supporting a learning framework layer when executed, and below the learning framework layer being a communication layer for providing communication bridges between the computation model layer and the learning framework layer and the plurality of different types of GPUs, wherein the communication layer is configured to cause the host computer to perform the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having computer program code stored thereon, the computer program code performing the method as described in any one of claims 1 to 8 when executed.

12. A computer program product comprising a computer program that, when executed by a machine, performs the method as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Collaborative training method for computing system, and computing device, computer-readable storage medium and computer program product

    WO2026108832A1