Heterogeneous hardware cluster distributed training method and device, electronic equipment and medium
Through the concept of process mesh and multiple heterogeneous parallel training strategies, the problems of single parallel training strategies and communication difficulties in the existing technology are solved, and efficient expansion of heterogeneous clusters and improvement of resource utilization are achieved.
Patent Information
- Application Number
- CN202510229319.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
In the prior art, parallel training strategies are single, and parallel training strategies cannot be flexibly set based on heterogeneous hardware resources. In addition, multi-dimensional hybrid heterogeneous parallel training strategies bring difficulties to communication libraries and communication algorithms.
By proposing the concept of process mesh (ProcessMesh), shielding the differences in underlying hardware, achieving any multiple hardware cluster mixing, providing a variety of optional heterogeneous parallel training strategies, and determining the preferred strategies based on the model to be trained and heterogeneous hardware to be trained, building a heterogeneous communication group between process mesh.
It realizes efficient expansion of heterogeneous clusters, provides more flexible parallel strategies, improves the resource utilization rate of heterogeneous clusters, and solves the problem of communication algorithms.
Smart Images

Figure CN120066731A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to a distributed training method, device, electronic device and medium for heterogeneous hardware clusters. Background Art
[0002] The underlying hardware topology involved in the parallel training of large-scale models usually corresponds one-to-one with the parallel training strategy. Due to the complex topology under multiple hardware, it brings difficulties to the expression of the top-level parallel training strategy and is not easy to expand the cluster.
[0003] In the prior art, large-scale heterogeneous parallel training strategies usually include data parallel heterogeneous mode and pipeline parallel heterogeneous mode. However, according to the underlying hardware topology, only one of these two parallel training strategies can be used, making the parallel training strategy single and unable to flexibly set the parallel training strategy based on heterogeneous hardware resources. Moreover, large-scale multi-dimensional hybrid heterogeneous parallel training strategies also bring difficulties to the communication library and communication algorithms, and the existing communication methods are not applicable to multi-dimensional hybrid heterogeneous parallel training strategies. Summary of the Invention
[0004] The purpose of this application is to provide a distributed training method, device, electronic device and medium for heterogeneous hardware clusters, aiming to solve the problem that the parallel training strategy in the prior art is single and unable to flexibly set the parallel training strategy based on the resources of heterogeneous hardware.
[0005] According to the first aspect of this application, a distributed training method for heterogeneous hardware clusters is provided, including:
[0006] Determine a heterogeneous parallel training strategy based on the model to be trained and heterogeneous hardware;
[0007] Based on the heterogeneous parallel training strategy and heterogeneous hardware, construct a process grid for executing the model training task of the model to be trained; wherein, different process grids have a mapping relationship with different hardware types in the heterogeneous hardware; after being called, the process grid executes on the heterogeneous hardware corresponding to the hardware type with the mapping relationship;
[0008] Based on the heterogeneous parallel training strategy and the mapping relationship, construct a heterogeneous communication group between process grids;
[0009] Call the process grid to execute the model training task of the model to be trained.
[0010] In an optional implementation, determining a heterogeneous parallel training strategy based on the model to be trained and heterogeneous hardware includes:
[0011] Determine the heterogeneous parallel training strategy based on the composition of the model to be trained and the resource information of the heterogeneous hardware.
[0012] In an alternative embodiment, based on the heterogeneous parallel training strategy and heterogeneous hardware, a process grid for executing the model training task of the to-be-trained model is constructed, including:
[0013] Establish a process grid corresponding one-to-one to the chip types included in the heterogeneous hardware;
[0014] In each group of process grids, converge the chip types and serial numbers corresponding to other process grids;
[0015] In each group of process grids, establish a mapping relationship between the chip physical identifiers corresponding to the chip type and the process logical identifiers corresponding to the process grid according to the chip type and serial number; wherein, the chip physical identifiers of all chips of the same chip type correspond to the process logical identifiers of the same group of process grids.
[0016] In an alternative embodiment, when the heterogeneous parallel training strategy includes a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy, based on the heterogeneous parallel training strategy and the mapping relationship, a heterogeneous communication group between process grids is constructed, including:
[0017] Based on the pipeline parallel strategy, data parallel strategy, and tensor parallel strategy, determine the data correspondence relationship and data rearrangement result between the process grids corresponding to different pipeline stages;
[0018] Based on the data correspondence relationship and the data rearrangement result, establish a heterogeneous communication group in which one process sends data to multiple processes and / or multiple processes send data to one process.
[0019] In an alternative embodiment, constructing a heterogeneous communication group between process grids based on the heterogeneous parallel training strategy and the mapping relationship further includes:
[0020] According to the data parallel dimension and tensor parallel dimension on the heterogeneous hardware corresponding to different process grids, determine the data correlation between different process grids;
[0021] Based on the data correlation, establish a heterogeneous communication group required for global gradient normalization during the training of the to-be-trained model.
[0022] In an alternative embodiment, the heterogeneous parallel training strategy includes one or a combination of more of a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy.
[0023] According to the second aspect of the present application, a heterogeneous hardware cluster distributed training device is provided, including:
[0024] A determination module, configured to determine a heterogeneous parallel training strategy based on a model to be trained and heterogeneous hardware;
[0025] A first construction module, configured to construct a process grid for executing a model training task of the model to be trained based on the heterogeneous parallel training strategy and the heterogeneous hardware; wherein, different process grids have a mapping relationship with different hardware types in the heterogeneous hardware; after being called, the process grid is executed on the heterogeneous hardware corresponding to the hardware type with the mapping relationship;
[0026] A second construction module, configured to construct a heterogeneous communication group between process grids based on the heterogeneous parallel training strategy and the mapping relationship;
[0027] A calling module, configured to call the process grid to execute the model training task of the model to be trained.
[0028] In an optional implementation manner, the determination module includes:
[0029] A first determination sub-module, configured to determine the heterogeneous parallel training strategy based on the composition of the model to be trained and the resource information of the heterogeneous hardware.
[0030] In an optional implementation manner, the first construction module includes:
[0031] A first establishment sub-module, configured to establish process grids in one-to-one correspondence with the chip types included in the heterogeneous hardware;
[0032] A convergence sub-module, configured to converge the chip types and serial numbers corresponding to other process grids in each group of process grids;
[0033] A second establishment sub-module, configured to establish a mapping relationship between the chip physical identifiers corresponding to the chip type and the process logical identifiers corresponding to the process grid in each group of process grids according to the chip type and serial number; wherein, the chip physical identifiers of all chips of the same chip type correspond to the process logical identifiers of the same group of process grids.
[0034] In an optional implementation manner, when the heterogeneous parallel training strategy includes a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy, the second construction module includes:
[0035] A second determination sub-module, configured to determine the data correspondence relationship and data rearrangement result between process grids corresponding to different pipeline stages based on the pipeline parallel strategy, the data parallel strategy, and the tensor parallel strategy;
[0036] A third establishment sub-module, configured to establish a heterogeneous communication group in which one process sends data to multiple processes and / or multiple processes send data to one process based on the data correspondence relationship and the data rearrangement result.
[0037] In an alternative embodiment, the second construction module further includes:
[0038] A third determination sub-module, configured to determine the data correlation between different process meshes according to the data parallelism dimension and the tensor parallelism dimension on heterogeneous hardware of different hardware types corresponding to different process meshes;
[0039] A fourth establishment sub-module, configured to establish a heterogeneous communication group required for global gradient normalization during the training of the to-be-trained model based on the data correlation.
[0040] In an alternative embodiment, the heterogeneous parallel training strategy includes one or a combination of a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy.
[0041] A third aspect of the present application provides an electronic device, including a processor and a memory, where the memory stores multiple instructions, and the processor is configured to read the instructions and execute the method of the foregoing first aspect.
[0042] A fourth aspect of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores multiple instructions, and the multiple instructions can be read and executed by a processor to execute the method of the foregoing first aspect.
[0043] Compared with the related art, the technical solution of the present application has at least the following advantages:
[0044] Based on the above solution, first, by proposing the concept of ProcessMesh, the present application shields the underlying hardware differences, can realize the mixing of any multiple hardware clusters, and realizes the efficient expansion of heterogeneous clusters. Second, the present application provides multiple optional heterogeneous parallel training strategies, and determines the preferred heterogeneous parallel training strategy based on the to-be-trained model and heterogeneous hardware, providing a more flexible parallel strategy, and further improving the resource utilization rate of heterogeneous clusters. Third, according to the heterogeneous parallel dimension strategy, the present application constructs a heterogeneous communication group among processes in the process mesh in the corresponding parallel dimension to solve the communication algorithm problem.
[0045] Other features and advantages of the present application will be described in the subsequent specification, and will be partially obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures and processes pointed out in the specification and the drawings. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 It is a schematic flowchart of a heterogeneous hardware cluster distributed training method according to an exemplary embodiment of the present application.
[0048] Figure 2 It is a schematic diagram of the effect of a process grid according to an exemplary embodiment of the present application.
[0049] Figure 3 It is a schematic diagram of the mapping relationship between a process grid and heterogeneous hardware according to an exemplary embodiment of the present application.
[0050] Figure 4 It is a schematic diagram of the construction principle of a heterogeneous communication group between different pipeline stages according to an exemplary embodiment of the present application.
[0051] Figure 5 It is a schematic diagram of the construction principle of aggregating heterogeneous communication groups during the global gradient normalization process according to an exemplary embodiment of the present application. Detailed implementation manners
[0052] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0053] The method provided by the present application can be implemented in the following terminal environment. The terminal may include one or more of the following components: a processor, a memory, and a display screen. Among them, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method described in the following embodiments.
[0054] The processor may include one or more processing cores. The processor connects various parts within the entire terminal through various interfaces and lines, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by calling the data stored in the memory, it executes various functions of the terminal and processes data.
[0055] The memory may include a Random Access Memory (RAM), and may also include a Read-Only Memory (ROM). The memory can be used to store instructions, programs, code, code sets, or instructions.
[0056] The display screen is used to display the user interfaces of various applications.
[0057] In addition, those skilled in the art can understand that the structure of the above terminal does not constitute a limitation on the terminal. The terminal may include more or fewer components, or combine certain components, or have different component arrangements. For example, the terminal also includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, and a power supply, which will not be elaborated here.
[0058] This application proposes a heterogeneous hardware cluster distributed training method and device. Among them, the method includes: determining a heterogeneous parallel training strategy based on the model to be trained and heterogeneous hardware; constructing a process grid for executing the model training task of the model to be trained based on the heterogeneous parallel training strategy and heterogeneous hardware, where different process grids have a mapping relationship with different hardware types in the heterogeneous hardware; the process grid is executed on the heterogeneous hardware corresponding to the hardware type with the mapping relationship after being called; constructing a heterogeneous communication group between process grids based on the heterogeneous parallel training strategy and the mapping relationship; calling the process grid to execute the model training task of the model to be trained. Based on the above solution, first, this application shields the underlying hardware differences by proposing the concept of ProcessMesh, can realize the mixing of any number of hardware clusters, and achieve the efficient expansion of heterogeneous clusters. Second, this application provides a variety of optional heterogeneous parallel training strategies, and determines the optimal heterogeneous parallel training strategy based on the model to be trained and heterogeneous hardware, providing a more flexible parallel strategy, and further improving the resource utilization rate of heterogeneous clusters. Third, this application constructs a heterogeneous communication group between processes in the process grid in the corresponding parallel dimension according to the heterogeneous parallel dimension strategy to solve the communication algorithm problem.
[0059] The technical solution of this application will be introduced below through specific embodiments.
[0060] See Figure 1 In the flowchart of, the heterogeneous hardware cluster distributed training method provided by this application includes:
[0061] Step 101: Determine a heterogeneous parallel training strategy based on the model to be trained and heterogeneous hardware;
[0062] Step 102: Based on the heterogeneous parallel training strategy and heterogeneous hardware, construct a process grid for executing the model training task of the to-be-trained model; wherein, different process grids have a mapping relationship with different hardware types in the heterogeneous hardware; after being called, the process grid is executed on the heterogeneous hardware corresponding to the hardware type with the mapping relationship.
[0063] Step 103: Based on the heterogeneous parallel training strategy and the mapping relationship, construct a heterogeneous communication group between process grids.
[0064] Step 104: Call the process grid to execute the model training task of the to-be-trained model.
[0065] In this embodiment, the training method can be applied to a heterogeneous hardware cluster, so a preferred heterogeneous parallel training strategy can be selected according to the to-be-trained model and the heterogeneous hardware for executing the model training task of the to-be-trained model. In some embodiments, optional heterogeneous parallel training strategies may include, but are not limited to, one or a combination of more of a pipelining parallel strategy, a data parallel strategy, and a tensor parallel strategy.
[0066] In some embodiments, a combination of one or more heterogeneous parallel training strategies can be selected to complete the training of the to-be-trained model.
[0067] After determining the heterogeneous parallel training strategy, multiple groups of process grids for executing the model training task of the to-be-trained model can also be constructed according to the heterogeneous parallel training strategy and heterogeneous hardware. A group of process grids can include multiple processes. Each group of process grids corresponds one-to-one with a hardware type in the heterogeneous hardware, that is, one hardware type corresponds to a group of process grids, and different hardware types correspond to different groups of process grids. During the model training process, the training task to be executed on the hardware of a certain hardware type is executed by calling the process with the mapping relationship with this hardware type. It should be noted that the hardware type can be the chip type in the heterogeneous hardware cluster.
[0068] See Figure 2 As shown, if there are three different types of chips in the heterogeneous hardware cluster for executing the model training task of the to-be-trained model, three groups of process grids A, B, and C can be constructed. Each group of process grids corresponds to a heterogeneous parallel training strategy and also corresponds to a chip sub-cluster of a chip type.
[0069] In addition, when considering training a model on hybrid heterogeneous hardware, a heterogeneous communication library that supports multiple hardware devices is required. Corresponding communication groups also need to be constructed for heterogeneous parallel training strategies, and the data needs to be re-partitioned to ensure the correctness of communication results. Communication across different process grids involves heterogeneous communication. To this end, the present application creates a heterogeneous communication group between process grids, which ensures normal communication between training tasks executed on different process grids.
[0070] After creating the heterogeneous communication group between process grids, the process grids can be called to execute the model training task of the model to be trained, so as to complete the training of the model to be trained by executing the corresponding process grids on heterogeneous hardware at the corresponding stage.
[0071] In some alternative implementation manners of this embodiment, step S101, that is, the step of determining a heterogeneous parallel training strategy based on the model to be trained and heterogeneous hardware, can be implemented as follows:
[0072] Determine the heterogeneous parallel training strategy based on the composition of the model to be trained and the resource information of the heterogeneous hardware.
[0073] In this alternative embodiment, the composition of the model to be trained may include the model structure and scale of the model to be trained, and the resource information of the heterogeneous hardware may include key information of the heterogeneous hardware cluster resources, such as computing power, communication bandwidth, topological relationship, etc. Based on the composition of the heterogeneous hardware and the resource information of the heterogeneous hardware, a preferred heterogeneous parallel training strategy for training the model to be trained can be determined. For example, when the model structure of the model to be trained is relatively complex and the computing power of the heterogeneous hardware is relatively strong, pipeline parallel training strategy, data parallel training strategy, and tensor parallel training strategy can be used simultaneously.
[0074] In some alternative implementation manners of this embodiment, step S102, that is, the step of constructing a process grid for executing the model training task of the model to be trained based on the heterogeneous parallel training strategy and heterogeneous hardware, can be implemented as follows:
[0075] Establish process grids that correspond one-to-one to the chip types included in the heterogeneous hardware;
[0076] Converge the chip types and serial numbers corresponding to other process grids in each group of process grids;
[0077] In each group of process grids, establish a mapping relationship between the chip physical identifiers corresponding to the chip type and the process logical identifiers corresponding to the process grids according to the chip type and serial number; among them, the chip physical identifiers of all chips of the same chip type correspond to the process logical identifiers of the same group of process grids.
[0078] In this alternative implementation, the process grid can be understood as a virtual representation of a physical chip. Before constructing a heterogeneous communication group and starting a training task, the process grid needs to be mapped to the corresponding physical chip.
[0079] As Figure 3 shown, the heterogeneous hardware includes three different types of chips. Therefore, three groups of process grids can be constructed. Process grid A corresponds to the chips of the first chip type, process grid B corresponds to the chips of the second chip type, and process grid C corresponds to the chips of the third chip type. In some embodiments, each process in a group of process grids can correspond to a chip of the corresponding chip type (as shown by the small squares distributed in an array in Figure 3 ), that is, the process grid corresponds one-to-one with the chip type, and the processes and chips in each group of process grids also correspond one-to-one.
[0080] It should be noted that since the original training platform does not guarantee that the chip physical identifier of the chip and the serial number of the process grid (i.e., the process logical identifier) correspond one-to-one during initialization, in this embodiment, the chip physical serial numbers of the chips can be reordered and bound to the process logical identifiers of the process grid. During this process, the chip physical identifiers of all chips can be aggregated first, and then the chip physical identifiers can be reordered according to the chip type. After that, they are bound in the way that a group of process grids corresponds to the chips of the same chip type.
[0081] Among them, when aggregating the physical chip identifiers of all chips, a collective communication (All_Gather) operation can be used in each process to gather the chip types and chip physical identifiers corresponding to other processes to the current process, so that each process has a complete correspondence between chips and processes.
[0082] In addition, when reordering the process grids according to the chip type, all chips can be reordered with the chip type as the primary order and the chip physical identifier as the secondary order, and a mapping relationship table between the process logical identifier and the chip physical identifier can be established. In this way, it is equivalent to binding the logical sequence number of the chip (the corresponding process logical identifier) and the physical sequence number (i.e., the chip physical identifier). The communication group is constructed according to the logical sequence number, and the corresponding chip can be found through the mapping relationship table.
[0083] In some alternative implementations of this embodiment, when the heterogeneous parallel training strategy includes a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy, step S103, that is, the step of constructing a heterogeneous communication group between process grids based on the heterogeneous parallel training strategy and the mapping relationship, can be implemented as follows:
[0084] Determine the data correspondence and data rearrangement results between process grids corresponding to different pipeline stages based on the pipeline parallel strategy, data parallel strategy, and tensor parallel strategy;
[0085] Based on the data correspondence and the data rearrangement results, establish a heterogeneous communication group in which one process sends data to multiple processes and / or multiple processes send data to one process.
[0086] In this alternative implementation, there is heterogeneous communication between pipeline stages that span different process grids. Different pipelines rely on P2P (send / recv) communication, so the heterogeneous communication library needs to support P2P communication. However, just having the support of the heterogeneous communication library does not guarantee the correctness of the data.
[0087] As Figure 4 shown, layers 1, 2, and 3 of the model and layers 4 and 5 of the model belong to two adjacent pipeline stages. These two pipeline stages adopt the data parallel strategy and the tensor parallel strategy, and the data parallel strategy and the tensor parallel strategy are different between different pipeline stages. For example, the data parallel strategy with two parallel degrees and the tensor parallel strategy with three parallel degrees are adopted for layers 1, 2, and 3 of the model, while the data parallel strategy with four parallel degrees and the tensor parallel strategy with two parallel degrees are adopted for layers 4 and 5 of the model.
[0088] To ensure the correctness of communication, according to the data correspondence and rearrangement results, a P2P heterogeneous communication group of "one chip sending to multiple downstream chips" and "multiple chips sending to one downstream chip" can be created to complete the forward propagation and backward propagation processes of training.
[0089] In some alternative implementation manners of this embodiment, step S103, that is, the step of constructing a heterogeneous communication group between process grids based on the heterogeneous parallel training strategy and the mapping relationship, can also be implemented in the following manner:
[0090] Determine the data correlation between different process grids according to the data parallel dimension and tensor parallel dimension on heterogeneous hardware of different hardware types corresponding to different process grids;
[0091] Based on the data correlation, establish the heterogeneous communication group required for global gradient normalization during the training of the model to be trained.
[0092] In this alternative implementation, when calculating global gradient normalization, Allreduce (aggregation) communication needs to span all pipeline stages.
[0093] The data parallelism on heterogeneous hardware of different hardware types may not be consistent. Heterogeneous communication groups required for global gradient normalization can be established based on the correlation between data in different pipeline stages. Heterogeneous hardware with low data parallelism can correspond to multiple heterogeneous communication groups. For example, Figure 5 As shown, two adjacent pipeline stages are respectively the first pipeline stage corresponding to model layers 1, 2, and 3 and the second pipeline stage corresponding to model layers 4 and 5. The data parallelism on the first process grid corresponding to the first pipeline stage is 2, and the data parallelism on the second process grid corresponding to the second pipeline stage is 4. Then each communication group of the first process grid corresponds to two communication groups of the second process grid. In some embodiments, to ensure the correctness of the results of multiple Allreduce operations, a reentrant communication method can be constructed to ensure that the effect of multiple Allreduce operations is equivalent to that of a single Allreduce operation, ensuring the correctness of the global gradient normalization calculation results.
[0094] Correspondingly, the present application provides a heterogeneous hardware cluster distributed training device in a second aspect, including:
[0095] A determination module, configured to determine a heterogeneous parallel training strategy based on a model to be trained and heterogeneous hardware;
[0096] A first construction module, configured to construct a process grid for executing the model training task of the model to be trained based on the heterogeneous parallel training strategy and the heterogeneous hardware; wherein, different process grids have a mapping relationship with different hardware types in the heterogeneous hardware; after being called, the process grid executes on the heterogeneous hardware corresponding to the hardware type with the mapping relationship;
[0097] A second construction module, configured to construct heterogeneous communication groups between process grids based on the heterogeneous parallel training strategy and the mapping relationship;
[0098] An invocation module, configured to invoke the process grid to execute the model training task of the model to be trained.
[0099] In an optional implementation manner, the determination module includes:
[0100] A first determination sub-module, configured to determine the heterogeneous parallel training strategy based on the composition of the model to be trained and the resource information of the heterogeneous hardware.
[0101] In an optional implementation manner, the first construction module includes:
[0102] A first establishment sub-module, configured to establish process grids corresponding one-to-one to the chip types included in the heterogeneous hardware;
[0103] An aggregation sub-module, configured to aggregate the chip types and serial numbers corresponding to other process grids in each group of process grids;
[0104] A second establishment sub-module, configured to establish a mapping relationship between the chip physical identifiers corresponding to the chip type and the process logical identifiers corresponding to the process grids in each group of process grids according to the chip type and the serial number; wherein, the chip physical identifiers of all chips of the same chip type correspond to the process logical identifiers of the same group of process grids.
[0105] In an optional implementation manner, when the heterogeneous parallel training strategy includes a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy, the second construction module includes:
[0106] A second determination sub-module, configured to determine the data correspondence relationship and data rearrangement result between the process grids corresponding to different pipeline stages based on the pipeline parallel strategy, the data parallel strategy, and the tensor parallel strategy;
[0107] A third establishment sub-module, configured to establish a heterogeneous communication group in which one process sends data to multiple processes and / or multiple processes send data to one process based on the data correspondence relationship and the data rearrangement result.
[0108] In an optional implementation manner, the second construction module further includes:
[0109] A third determination sub-module, configured to determine the data correlation between different process grids according to the data parallel dimension and the tensor parallel dimension of the heterogeneous hardware of different hardware types corresponding to different process grids;
[0110] A fourth establishment sub-module, configured to establish a heterogeneous communication group required for global gradient normalization during the training of the to-be-trained model based on the data correlation.
[0111] In an optional implementation manner, the heterogeneous parallel training strategy includes one or a combination of the pipeline parallel strategy, the data parallel strategy, and the tensor parallel strategy.
[0112] The above device corresponds to the heterogeneous hardware cluster distributed training method provided in the above embodiment, and specific details can be referred to the description of the heterogeneous hardware cluster distributed training method in the above embodiment, which will not be elaborated here.
[0113] It can be understood that the circuit structures, names, and parameters described in the above embodiments are only examples. Those skilled in the art can also make combinations and adjustments of the structural features of the above multiple embodiments that are easily conceivable according to the usage needs, and should not limit the concept of the present application to the specific details of the above examples.
[0114] The present application also provides an electronic device, including a processor and a memory. The memory stores multiple instructions, and the processor is configured to read the instructions and execute any of the methods in the foregoing first aspect. The processor and the memory may be connected through a bus or other means. Taking the connection through the bus as an example, the processor may be a Central Processing Unit (CPU). The processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips.
[0115] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor can execute various functional applications and data processing of the processor, that is, implement the methods in the above method embodiments.
[0116] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0117] As another aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium may be the computer-readable storage medium included in the device in the above embodiments; or it may exist separately and not be assembled into the device. The computer-readable storage medium stores one or more programs, and the one or more programs are used by one or more processors to execute the methods described in the present application.
[0118] Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A distributed training method for heterogeneous hardware clusters, characterized in that: include: Determine the heterogeneous parallel training strategy based on the model to be trained and the heterogeneous hardware; Based on the heterogeneous parallel training strategy and heterogeneous hardware, a process grid for executing the model training task of the model to be trained is constructed; wherein different process grids have a mapping relationship with different hardware types in the heterogeneous hardware; after being called, the process grid is executed on the heterogeneous hardware corresponding to the hardware type with the mapping relationship; Based on the heterogeneous parallel training strategy and the mapping relationship, a heterogeneous communication group between process grids is constructed: The process grid is called to execute the model training task of the model to be trained.
2. The heterogeneous hardware cluster distributed training method according to claim 1, characterized in that: Based on the model to be trained and the heterogeneous hardware, determine the heterogeneous parallel training strategy, including: The heterogeneous parallel training strategy is determined based on the composition of the model to be trained and the resource information of the heterogeneous hardware.
3. The heterogeneous hardware cluster distributed training method according to claim 1, characterized in that: Based on the heterogeneous parallel training strategy and heterogeneous hardware, a process grid for executing the model training task of the model to be trained is constructed, including: Establishing a process grid corresponding to chip types included in the heterogeneous hardware; In each group of process grids, the chip types and serial numbers corresponding to other process grids are gathered; In each group of process grids, a mapping relationship between the chip physical identifier corresponding to the chip type and the process logical identifier corresponding to the process grid is established according to the chip type and serial number; wherein the chip physical identifiers of all chips of the same chip type correspond to the process logical identifiers of the same group of process grids.
4. The heterogeneous hardware cluster distributed training method according to claim 1, characterized in that: When the heterogeneous parallel training strategy includes a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy, a heterogeneous communication group between process grids is constructed based on the heterogeneous parallel training strategy and the mapping relationship, including: Based on the pipeline parallel strategy, data parallel strategy and tensor parallel strategy, determine the data correspondence relationship and data rearrangement result between process grids corresponding to different pipeline stages; Based on the data correspondence and the data rearrangement result, a heterogeneous communication group is established in which one process sends data to multiple processes and / or multiple processes send data to one process.
5. The heterogeneous hardware cluster distributed training method according to claim 4, characterized in that: Based on the heterogeneous parallel training strategy and the mapping relationship, a heterogeneous communication group between process grids is constructed, which also includes: Determine the correlation of data between different process grids based on the data parallelism and tensor parallelism dimensions on heterogeneous hardware of different hardware types corresponding to different process grids; Based on the correlation of the data, a heterogeneous communication group required for global gradient normalization of the model to be trained during the training process is established.
6. The heterogeneous hardware cluster distributed training method according to any one of claims 1 to 3, characterized in that: The heterogeneous parallel training strategy includes a combination of one or more of a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy.
7. A heterogeneous hardware cluster distributed training device, characterized in that: include: A determination module, used to determine a heterogeneous parallel training strategy based on a model to be trained and heterogeneous hardware; A first construction module is used to construct a process grid for executing the model training task of the model to be trained based on the heterogeneous parallel training strategy and the heterogeneous hardware; wherein different process grids have a mapping relationship with different hardware types in the heterogeneous hardware; after being called, the process grid is executed on the heterogeneous hardware corresponding to the hardware type with the mapping relationship; A second construction module is used to construct a heterogeneous communication group between process grids based on the heterogeneous parallel training strategy and the mapping relationship; The calling module is used to call the process grid to execute the model training task of the model to be trained.
8. The heterogeneous hardware cluster distributed training device according to claim 7, characterized in that: The determining module comprises: The first determination submodule is used to determine the heterogeneous parallel training strategy based on the composition of the model to be trained and the resource information of the heterogeneous hardware.
9. The heterogeneous hardware cluster distributed training device according to claim 7, characterized in that: The first building block comprises: A first establishing submodule is used to establish a process grid corresponding to the chip types included in the heterogeneous hardware; The aggregation submodule is used to aggregate the chip types and serial numbers corresponding to other process grids in each group of process grids; The second establishment submodule is used to establish a mapping relationship between the chip physical identifier corresponding to the chip type and the process logical identifier corresponding to the process grid in each group of process grids according to the chip type and serial number; wherein the chip physical identifiers of all chips of the same chip type correspond to the process logical identifiers of the same group of process grids.
10. The heterogeneous hardware cluster distributed training device according to claim 7, characterized in that: When the heterogeneous parallel training strategy includes a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy, the second building module includes: A second determination submodule, used to determine the data correspondence and data rearrangement results between process grids corresponding to different pipeline stages based on the pipeline parallel strategy, data parallel strategy and tensor parallel strategy; The third establishing submodule is used to establish a heterogeneous communication group in which a process sends data to multiple processes and / or multiple processes send data to one process based on the data corresponding relationship and the data rearrangement result.
11. The heterogeneous hardware cluster distributed training device according to claim 10, characterized in that: The second building block further includes: A third determination submodule is used to determine the correlation of data between different process grids according to the data parallel dimension and tensor parallel dimension on heterogeneous hardware of different hardware types corresponding to different process grids; The fourth establishing submodule is used to establish a heterogeneous communication group required for global gradient normalization of the model to be trained during the training process based on the correlation of the data.
12. The heterogeneous hardware cluster distributed training device according to any one of claims 7 to 9, characterized in that: The heterogeneous parallel training strategy includes a combination of one or more of a pipeline parallel strategy, a data parallel strategy, and a tensor parallel strategy.
13. An electronic device, characterized in that: It includes a processor and a memory, the memory stores multiple instructions, and the processor is used to read the instructions and execute the heterogeneous hardware cluster distributed training method as described in any one of claims 1-6.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the plurality of instructions can be read by a processor and executed by the heterogeneous hardware cluster distributed training method as described in any one of claims 1-6.
Citation Information
Patent Citations
Computational grid parallel region division method and device based on reinforcement learning
CN111353260A
Assembly line parallel method for accelerating neural network training in heterogeneous GPU cluster
CN116883229A
Large model distributed training method and system for heterogeneous hardware cluster
CN117909742A
Distributed training method, device and equipment based on heterogeneous equipment and medium
CN118733282A
Communication method, system and device, equipment and storage medium
CN118802909A