Distributed training method and system, electronic equipment and storage medium

By triggering the optimization update task immediately when the reverse computing task is completed and migrating it to the host for processing, the problem of communication bottleneck in large-scale distributed deep learning training is solved, and training efficiency and parallelism are improved.

CN120218190AActive Publication Date: 2025-06-27SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510667964.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-27
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

In large-scale distributed deep learning training, communication overhead becomes a bottleneck that restricts overall training performance, especially in data parallel mode, communication accounts for a high proportion and it is difficult to make full use of network bandwidth, resulting in a low transmission rate.

Method used

By triggering the execution of the optimization update task immediately when the reverse computing task is completed, and migrating the optimization update task from the device to the host, the host completes the optimization update task, saving the device's computing resources, and achieving efficient overlap between communication and computing.

Benefits of technology

Reduce waiting time, improve pipeline efficiency and parallelism, and alleviate the communication bottleneck problem in large-scale distributed deep learning training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218190A_ABST
    Figure CN120218190A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a distributed training method and system, electronic equipment and a storage medium. The distributed training method is applied to a distributed training system, in a current round of iterative training, training of at least one stage of a target model is cooperatively executed by at least one corresponding device and at least one corresponding host, and the distributed training method comprises the following steps: for each stage of the target model, the device executes a forward calculation task, obtaining a forward calculation result; the equipment executes a reverse calculation task based on the forward calculation result to obtain gradient data; and in response to the completion of the reverse calculation task, the equipment triggers to execute the optimization update task and obtains the updated parameters for the use of the next round of iterative training. According to the distributed training method, waiting time can be reduced, computing resources of equipment can be saved, efficient overlapping of communication and computing can be realized, assembly line efficiency and parallelism can be effectively improved, and thus the communication bottleneck problem in large-scale distributed deep learning training can be relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a distributed training method and system, an electronic device, and a storage medium. Background Art

[0002] With the development of technology, Artificial Intelligence (AI) technology has been widely applied in multiple fields. Deep Learning is one of the important technologies of AI technology. Deep learning technology based on artificial neural networks has made great progress in fields such as object classification, text processing, image search, and human-computer dialogue.

[0003] With the increase in problem complexity, the depth and scale of deep learning models have also been continuously improved, and distributed training is widely used for large-scale model training. Summary of the Invention

[0004] At least one embodiment of the present disclosure provides a distributed training method applied to a distributed training system. The distributed training system is configured to perform distributed training on a target model. The distributed training system includes multiple computing nodes, and each computing node includes at least one host and at least one device. The target model is divided into multiple stages in a pipeline parallel mode, and each stage includes at least one layer of the target model. The multiple stages are respectively distributed in multiple devices of the multiple computing nodes. In the current round of iterative training, at least one stage of the target model is trained by at least one corresponding device and at least one corresponding host in cooperation.

[0005] The distributed training method includes: for each stage of the target model, the device performs a forward calculation task to obtain a forward calculation result; the device performs a backward calculation task based on the forward calculation result to obtain gradient data; in response to the completion of the backward calculation task, the device triggers the execution of an optimization update task and obtains updated parameters for use in the next round of iterative training; for at least one stage of the target model, in response to the completion of the backward calculation task, the device triggers the execution of an optimization update task and obtains updated parameters for use in the next round of iterative training, including: in response to the completion of the backward calculation task, the device transfers at least part of the gradient data to the host and triggers the host to execute the optimization update task; the host performs the optimization update task based on at least part of the gradient data to obtain at least part of the updated parameters, and transfers the at least part of the updated parameters to the device for use in the next round of iterative training.

[0006] In the distributed training method provided by at least one embodiment of the present disclosure, devices in the multiple computing nodes are divided into multiple data parallel groups in a data parallel mode, and gradient aggregation is performed among devices within the same data parallel group. In response to completion of the backward calculation task, the device transfers at least part of the gradient data to the host, including: for each data parallel group related to the stage, in response to completion of the backward calculation task, each device in the data parallel group transfers the gradient data to the corresponding host through device-to-host transmission.

[0007] In the distributed training method provided by at least one embodiment of the present disclosure, the host performs an optimization update task based on at least part of the gradient data, obtains at least part of the updated parameters, and transfers the at least part of the updated parameters to the device for use in the next round of iterative training, including: in each data parallel group related to the stage, for the host corresponding to each device in the data parallel group, the host divides the gradient data transferred by the device into N parts of data, and sends N-1 parts of the N parts of data to the hosts corresponding to other devices in the data parallel group respectively, where N is the number of devices in the data parallel group; the host performs an optimization update task according to all the received gradient data to obtain at least part of the updated parameters; the host obtains the updated parameters generated by the hosts corresponding to other devices in the data parallel group, and transfers all the updated parameters to the device through host-to-device transmission for use in the next round of iterative training.

[0008] In the distributed training method provided by at least one embodiment of the present disclosure, devices in the multiple computing nodes are divided into multiple data parallel groups in a data parallel mode, and gradient aggregation is performed among devices within the same data parallel group. In response to completion of the backward calculation task, the device transfers at least part of the gradient data to the host, including: for each data parallel group related to the stage, in response to completion of the backward calculation task, each device in the data parallel group divides the gradient data into N parts of data, transfers one part of the N parts of data to the corresponding host through device-to-host transmission, and sends the remaining N-1 parts of the N parts of data to the hosts corresponding to other devices in the data parallel group respectively, where N is the number of devices in the data parallel group.

[0009] In the distributed training method provided by at least one embodiment of the present disclosure, the host performs an optimization update task based on at least part of the gradient data, obtains at least part of the updated parameters, and transfers the at least part of the updated parameters to the device for use in the next round of iterative training, including: in each data parallel group related to the stage, for the host corresponding to each device in the data parallel group, the host performs an optimization update task based on all the received gradient data, obtains at least part of the updated parameters, and sends the at least part of the updated parameters to all the devices in the data parallel group for use in the next round of iterative training.

[0010] The distributed training method provided by at least one embodiment of the present disclosure further includes: allocating the i-th stage and the (M - i + 1)-th stage to the same computing node, where 0 < i ≤ M / 2, M is the total number of stages obtained by dividing the target model, and the i-th stage represents the i-th stage sequentially executed during the forward propagation of the target model.

[0011] In the distributed training method provided by at least one embodiment of the present disclosure, the at least one stage includes a first stage, and the first stage includes the layer at the very front of the target model in the forward propagation order.

[0012] In the distributed training method provided by at least one embodiment of the present disclosure, the configuration of the host corresponding to the first stage is higher than the configurations of the hosts corresponding to other stages.

[0013] In the distributed training method provided by at least one embodiment of the present disclosure, each computing node further includes at least one network card, and the at least one host is configured to drive at least one network card to implement network communication between the multiple computing nodes.

[0014] In the distributed training method provided by at least one embodiment of the present disclosure, the network communication is based on the remote direct memory access mechanism.

[0015] At least one embodiment of the present disclosure provides a distributed training system configured to perform distributed training on a target model. The distributed training system includes a plurality of computing nodes, each computing node including at least one host and at least one device. The target model is divided into multiple stages in a pipeline parallel mode, each stage including at least one layer of the target model, and the multiple stages are respectively distributed in multiple devices of the multiple computing nodes. During the current round of iterative training, the training of at least one stage of the target model is jointly executed by at least one corresponding device and at least one corresponding host. For each stage of the target model, the device is configured to: perform a forward calculation task to obtain a forward calculation result; perform a backward calculation task based on the forward calculation result to obtain gradient data; in response to the completion of the backward calculation task, trigger the execution of an optimization update task and obtain updated parameters for use in the next round of iterative training; for at least one stage of the target model, the device is further configured to: in response to the completion of the backward calculation task, transfer at least part of the gradient data to the host and trigger the host to execute the optimization update task; the host is configured to: perform the optimization update task based on at least part of the gradient data to obtain at least part of the updated parameters, and transfer the at least part of the updated parameters to the device for use in the next round of iterative training.

[0016] At least one embodiment of the present disclosure provides an electronic device, including: at least one processor; at least one memory including one or more computer program modules. The one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the distributed training method provided by at least one of the above embodiments of the present disclosure.

[0017] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions, when executed by at least one processor, execute the distributed training method provided by at least one of the above embodiments of the present disclosure.

[0018] The distributed training method, system, electronic device, and storage medium provided by at least one embodiment of the embodiments of the present disclosure can reduce the waiting time and effectively improve the pipeline efficiency and parallelism by immediately triggering the execution of the optimization update task when the backward calculation task is completed. Moreover, by migrating the optimization update task from the device to the host and having the host complete the optimization update task, the computing resources of the device can be saved, and the efficient overlap of communication and computing can be achieved, thereby alleviating the communication bottleneck problem in large-scale distributed deep learning training. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0020] Figure 1A It is a schematic diagram of a data parallel mode;

[0021] Figure 1B It is a schematic diagram of a tensor parallel mode;

[0022] Figure 1C It is a schematic diagram of a pipeline parallel mode;

[0023] Figure 1D It is a schematic diagram of a micro-batch pipeline parallel mode;

[0024] Figure 2A It is a schematic block diagram of a distributed training system provided by at least one embodiment of the present disclosure;

[0025] Figure 2B It is a schematic block diagram of a computing node provided by at least one embodiment of the present disclosure;

[0026] Figure 3A It is a flowchart of a distributed training method provided by at least one embodiment of the present disclosure;

[0027] Figure 3B It is a timing schematic diagram of a distributed training method provided by at least one embodiment of the present disclosure;

[0028] Figure 3C It is a flowchart of a distributed training method provided by at least one embodiment of the present disclosure;

[0029] Figure 4 It is an exemplary schematic diagram of a distributed training method provided by at least one embodiment of the present disclosure;

[0030] Figure 5 It is an exemplary schematic diagram of a distributed training method provided by at least one embodiment of the present disclosure;

[0031] Figure 6 It is a schematic diagram of a folded-back pipeline provided by at least one embodiment of the present disclosure;

[0032] Figure 7 It is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure;

[0033] Figure 8 It is a schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure;

[0034] Figure 9Schematic block diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners

[0035] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0036] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second" and similar terms used in the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. The terms "including" or "comprising" and the like mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0037] The present disclosure will be described below through several specific embodiments. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of the embodiments of the present disclosure appears in more than one drawing, the component is denoted by the same or similar reference numerals in each drawing.

[0038] There is a wide variety of processors for model training, such as Graphics Processing Unit (GPU), General-purpose Graphics Processing Unit (GPGPU), Tensor Processing Unit (TPU), Application-Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA), Artificial Intelligence (AI) accelerator or coprocessor, etc. Currently, GPUs are usually used as the main training devices, combined with Central Processing Unit (CPU) for auxiliary tasks and overall process management. For specific scenarios or requirements, ASICs, FPGAs or AI accelerators may be selected for customized acceleration.

[0039] Since a large-scale model may have up to hundreds of billions of parameters, a single machine with a single card is no longer capable of handling the training tasks of such large-scale models. Therefore, distributed training is widely used for training large-scale models in artificial intelligence. Distributed training refers to the use of multiple machines working together to accelerate the training process of large deep learning models. In distributed training, the original model, dataset, and training process are decomposed and distributed to multiple machines for simultaneous processing, thus effectively utilizing more computing resources and shortening the training time.

[0040] In the distributed training of large models, common parallel techniques include data parallelism and model parallelism. Model parallelism includes tensor parallelism and pipeline parallelism.

[0041] The data parallelism mode means that the training dataset of the model is divided into multiple sub-datasets. The divided multiple sub-datasets are distributed to multiple devices, and each device holds a complete copy of the model, so that it can independently train the sub-dataset assigned to it. For example, each device can independently complete the forward propagation and backward propagation calculations of the sub-dataset.

[0042] Figure 1A It is a schematic diagram of a data parallelism mode. As Figure 1A shown, a training dataset is divided into 3 parts, which are respectively distributed to device 0, device 1, and device 2. Each of device 0, device 1, and device 2 has a complete model. After the backward propagation is completed on each device, it is necessary to aggregate all the gradients calculated by each device and update the global model parameters.

[0043] Figure 1B It is a schematic diagram of a tensor parallelism mode. The tensor parallelism mode means splitting a layer (or operator) of a model and placing partial weights of the operator on different devices, thereby reducing the memory footprint on each device (e.g., GPU). A neural network model includes multiple layers ( Figure 1B 4 layers are shown as an example in the figure), and each layer can be understood as a function that can perform specific mathematical operations on the input tensor. For example, a fully connected layer can perform a linear transformation operation on the tensor, a convolutional layer can perform a convolution operation on the tensor, a pooling layer can perform a pooling operation on the tensor, etc., to extract or transform features. These operations can change the content and shape of the tensor, realizing the gradual transformation from the original data to the high-level abstract representation. As Figure 1B shown, in the tensor parallelism mode, the tensor of a layer is split into multiple parts, and the multiple split partial tensors are respectively assigned to multiple devices, such as device 0, device 1, and device 2. Each device is only responsible for a part of the model's calculation and exchanges necessary intermediate results through communication.

[0044] The pipeline parallelism mode means dividing different layers of the model into multiple stages, and the multiple stages are respectively assigned to multiple devices to form a pipeline. The activation values and gradients are passed in sequence between different devices. For example, during the forward propagation process, each device passes the activation value to the device where the next pipeline stage is located. During the backward propagation process, each device passes the gradient back to the device where the previous pipeline stage is located, so that multiple devices can be utilized for calculation at the same time, improving the calculation efficiency.

[0045] Figure 1C It is a schematic diagram of a pipeline parallelism mode. As Figure 1C shown, a model includes 4 layers, and these 4 layers are divided into 3 stages. Each stage is assigned to a device. For example, the first layer is assigned to device 0, the second layer is assigned to device 1, and the third and fourth layers are assigned to device 2.

[0046] Figure 1D It is a schematic diagram of a micro-batch pipeline parallelism mode. By dividing the global batch data into micro-batch data and artificially creating a pipeline to solve the problem of device idleness, different devices are allowed to participate in the calculation process simultaneously, which can significantly improve the utilization rate of pipeline parallel devices and reduce the time of the device idle state.

[0047] As Figure 1D shown, multiple layers of the target model are split into 4 stages, and each stage includes, for example, 4 micro-batch data. F i,j represents the forward propagation of the j-th micro-batch data in the i-th stage. B x,yRepresents the backpropagation of the y-th micro-batch data in the x-th stage. During the forward calculation process, the calculations in the subsequent stages depend on the calculations in the previous stage. For example, F 1,0 must be calculated before F 0,0 can start. Similarly, during the backpropagation process, the calculations in the previous stage depend on the calculations in the subsequent stage. For example, B 2,3 must be calculated before B 3,3 can start.

[0048] In the pipeline parallelism mode, data is transmitted between adjacent devices through communication links. For example, during the forward calculation process, the input data first obtains intermediate results through the calculation of the first layer on device 0, and the intermediate results are transmitted to device 1. Then, the output of the second layer is calculated on device 1, and the output result of the second layer of the model is transmitted to device 2. The final result of the forward calculation is obtained through the calculation of the last layer on device 2. The backpropagation process is similar. Finally, the network layers on each device will update the parameters using the gradients calculated during the backpropagation process.

[0049] In Figure 1D the pipeline parallelism mode shown, the backpropagation calculation cannot start until all devices have completed the forward calculation. This causes the intermediate results of the forward calculation in the first stage to be retained until all stages of the forward calculation are completed before being released, resulting in a large number of pipeline bubbles, that is, idle cycles or gaps that appear in the pipeline. Excessive bubbles will lead to a decrease in processor performance. Therefore, reducing pipeline bubbles has become a key goal for optimizing processor performance.

[0050] As the model scale gradually increases, simply using a single parallelism mode often cannot simultaneously meet the requirements of device memory limitations and high computational efficiency. Therefore, for large-scale models, it is usually necessary to combine multiple parallel techniques such as data parallelism, tensor parallelism, and pipeline parallelism for distributed training.

[0051] However, the inventors of the present disclosure have noticed that in large-scale distributed parallel training scenarios, communication overhead has become a bottleneck restricting the overall training performance. In the data parallel mode, as the scale of the deep learning model increases, communication overhead has become a key factor limiting performance improvement. For example, the proportion of data parallel communication is relatively high, making it difficult to fully utilize the network bandwidth, resulting in a low transmission rate. In addition, as the scale of data parallelism expands, the bandwidth of the collective communication algorithm significantly decreases. Although some methods based on Streaming Multiprocessors can improve the bandwidth, they cannot achieve the overlap of communication and computing. Other methods use a CPU optimizer to offload parameters to the CPU. Although this can relieve the burden on the GPU to a certain extent, the communication process still needs to be completed through the GPU, and the parallelism of backpropagation and communication cannot be achieved.

[0052] At least one embodiment of the present disclosure provides a distributed training method applied to a distributed training system. The distributed training system is configured to perform distributed training on a target model. The distributed training system includes multiple computing nodes, and each computing node includes at least one host and at least one device. The target model is divided into multiple stages in a pipeline parallel mode, each stage includes at least one layer of the target model, and the multiple stages are respectively distributed in multiple devices of multiple computing nodes. In the current round of iterative training, the training of at least one stage of the target model is jointly executed by at least one corresponding device and at least one corresponding host.

[0053] The distributed training method includes: for each stage of the target model, the device performs a forward calculation task to obtain a forward calculation result; the device performs a backward calculation task based on the forward calculation result to obtain gradient data; in response to the completion of the backward calculation task, the device triggers the execution of an optimization update task and obtains updated parameters for use in the next round of iterative training.

[0054] For at least one stage of the target model, in response to the completion of the backward calculation task, the device triggers the execution of an optimization update task and obtains updated parameters for use in the next round of iterative training, including: in response to the completion of the backward calculation task, the device transfers at least part of the gradient data to the host and triggers the host to execute the optimization update task; the host performs the optimization update task based on at least part of the gradient data to obtain at least part of the updated parameters, and transfers at least part of the updated parameters to the device for use in the next round of iterative training.

[0055] The distributed training method provided by at least one embodiment of the present disclosure can reduce the waiting time and effectively improve the pipeline efficiency and parallelism by immediately triggering the execution of the optimization update task when the reverse calculation task is completed. Moreover, by migrating the optimization update task from the device to the host and having the host complete the optimization update task, the computing resources of the device can be saved, and efficient overlap of communication and computing can be achieved, thereby alleviating the communication bottleneck problem in large-scale distributed deep learning training.

[0056] The distributed training method provided by at least one embodiment of the present disclosure can be applied to a distributed training system. In an embodiment of the present disclosure, the distributed training system is configured to perform distributed training on a target model. The distributed training system includes multiple computing nodes, and each computing node includes at least one host and at least one device.

[0057] Figure 2A It is a schematic block diagram of a distributed training system provided by at least one embodiment of the present disclosure. For example, Figure 2A The illustrated distributed training system includes multiple computing nodes (computing node 0 to computing node x), each computing node includes multiple hosts (host 0 to host y), and each host is physically connected to multiple devices (device 0 to device z) (for example, connected through physical lines such as Ethernet cables, optical fibers, and network hardware such as switches). Here, x, y, and z are all positive integers and can be set according to actual needs. In each computing node, data can be transmitted between the host and the device in different ways. The device (also referred to as a "slave device") can be directly connected to the host through a hardware interface (for example, connected through a Peripheral Component Interconnect Express (PCIe) bus) to achieve low-latency data transmission such as Device-to-Host (D2H) transmission and Host-to-Device (H2D) transmission; it can also work with the host in a software-defined manner through a network (such as virtualization technology). It should be noted that the hosts in different computing nodes are different, and the devices corresponding to different hosts are also different. Specifically, host 0 in computing node 0 and host 0 in computing node x are not the same host, and device 0 connected to host 0 and device 0 connected to host 1 are not the same device, and so on.

[0058] It should be noted that Figure 2A Only as an example, in the above distributed training system, each computing node can also include only one host, and each host can also be physically connected to only one device. Moreover, the number of hosts in each computing node can be unequal, and the number of devices physically connected to each host can also be unequal.

[0059] For example, in at least one example of the embodiments of the present disclosure, the host may include a central processing unit (CPU), and the device may include an accelerator such as a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a tensor processing unit (TPU), an AI accelerator, an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA). For example, the distributed training system may include 8, 16, or 32 computing nodes, and each computing node may include 4 or 8 GPUs.

[0060] In the embodiments of the present disclosure, each computing node may further include at least one network card, and at least one host is configured to drive one network card to implement network communication between multiple computing nodes.

[0061] Figure 2B A schematic block diagram of a computing node provided for at least one embodiment of the present disclosure.

[0062] For example, as Figure 2B shown, the computing node includes hosts 0 to 1, switches 0 to 3, devices 0 to 7, and network cards 0 to 7. For example, host 0, device 0, and device 1, and network cards 0 and 1 are connected through switch 0. Host 0 may drive network cards 0 to 3 to implement communication between this computing node and other computing nodes in the distributed training system. For the network cards and devices connected through the same switch, the network card may be referred to as the affinity network card of the device. In the Figure 2B example, network cards 0 and 1 are the affinity network cards of devices 0 and 1. The switch may be, for example, a PCIe switch. In the same computing node, the hosts may be connected through an Ultra Path Interconnect (UPI) bus, and the devices may be connected through a high-speed bus or other GPU interconnection technologies (not shown in the figure).

[0063] It should be noted that Figure 2B this is only an example. According to actual needs, the host may be connected to more or fewer switches, and each switch may be connected to more or fewer devices and network cards. The embodiments of the present disclosure do not limit this. Moreover, the number of devices and network cards connected to each switch may not be equal. For example, in some examples, a switch may be connected to two devices and one network card, that is, these two devices communicate with other computing nodes through the same network card, and this network card may be referred to as the affinity network card of these two devices.

[0064] In the distributed training system provided by at least one embodiment of the present disclosure, the host drives the network card to implement communication between computing nodes, which can make full use of the multi-threaded processing ability of the host, thereby efficiently utilizing the network bandwidth and improving the communication efficiency. In this architecture, the device can focus on executing computing tasks, and at the same time effectively solve the problem that it is difficult to fully utilize the network bandwidth when the device drives the network card, and improve the overall performance of the distributed training system under high load.

[0065] The distributed training method provided by at least one embodiment of the present disclosure can be applied to various types of distributed training systems, including but not limited to Figure 2A the distributed training system shown, nor limited to Figure 2B the distributed training system including the computing nodes shown.

[0066] In a distributed training method provided by at least one embodiment of the present disclosure, the target model is divided into multiple stages in a pipelined parallel mode, each stage includes at least one layer of the target model, and the multiple stages are respectively distributed in multiple devices of multiple computing nodes. For example, in the case of adopting the micro-batch pipelined parallel mode, each stage can include multiple micro-batches. For example, taking the target model including 80 layers as an example, assuming that the target model is divided into 8 stages according to the pipelined parallel mode, then each stage includes 10 layers of the target model.

[0067] It should be noted that the above-mentioned multiple stages and multiple computing nodes do not need to correspond one by one. One computing node can correspond to one or more stages, one stage can also correspond to one or more computing nodes, or one or more devices in one computing node.

[0068] For example, the target model described above can be a machine learning model to be trained, such as a deep learning model. The deep learning model can be a neural network structure including multiple layers (3 layers, 4 layers, 8 layers or more layers), such as a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory network (LSTM), etc. The structures of these neural networks generally include an input layer, a hidden layer, and an output layer. The hidden layer refers to those layers in the neural network located between the input layer and the output layer, and the hidden layer is also called the processing layer. The input layer is used to receive the data to be processed, such as the image to be processed, etc., the output layer is used to output the processing result, such as the processed image, etc., and the processing layer can include a convolutional layer, a pooling layer, a batch normalization layer, a fully connected layer, etc. According to the different structures of the neural network, the processing layer can include different contents and combination methods. In some examples, the target model can be, for example, a large language model (LLM). The large language model is a deep learning model trained based on a large amount of text data and can understand and generate natural language text. The large language model is usually based on the Transformer architecture and learns the statistical laws and semantic information of language from a large amount of text data through self-supervised learning, and is usually used for tasks such as text generation, translation, question answering, and summarization.

[0069] Figure 3A Flowchart of a distributed training method provided by at least one embodiment of the present disclosure.

[0070] For example, as Figure 3A shown, in the current round of iterative training, the training of at least one stage of the target model is jointly executed by at least one corresponding device and at least one corresponding host. For each stage of the target model, the distributed training method provided by at least one embodiment of the present disclosure includes the following steps S101 to step S103.

[0071] Step S101: The device executes the forward calculation task to obtain the forward calculation result.

[0072] Step S102: The device executes the backward calculation task based on the forward calculation result to obtain the gradient data.

[0073] Step S103: In response to the completion of the backward calculation task, the device triggers the execution of the optimization update task and obtains the updated parameters for use in the next round of iterative training.

[0074] For example, in step S101, forward calculation (also known as forward propagation) is a calculation method in a neural network used to calculate the relationships between the input layer, hidden layers, and output layer. It involves a series of processes of starting from the input data of the neural network's input layer, passing through a series of weight matrices and activation function calculations, and finally obtaining the output result. In forward calculation, the neural network takes the output of each layer as the input of the next layer until the final calculation result is obtained at the output layer.

[0075] Activation values are the intermediate results generated by each layer of the network during forward propagation, and these results will be used as the basis for gradient calculation during backpropagation. In the forward propagation stage, the activation value refers to the result after weighted input and activation function operations. For a neuron, it first receives signals from neurons in the previous layer (usually the weighted sum plus the bias term), and then calculates the activation value through a non-linear activation function (such as Sigmoid, ReLU, etc.). The activation value represents the "activation degree" of the neuron at a specific moment.

[0076] For example, in step S102, back calculation (also known as backpropagation) is an optimization algorithm used to adjust the weights and biases in a neural network, which can be carried out based on, for example, minimizing the loss function. The loss function is a mathematical function used to measure the error between the output of a neural network model and the actual label or target value. During the neural network training process, for example, backpropagation can calculate the gradient of the loss function and propagate the gradient backward from the output layer to the input layer, thereby updating the weight parameters of each layer. The weight is a parameter that connects neurons in a neural network. Each weight corresponds to the connection strength from a neuron in the previous layer to a neuron in the current layer. In forward propagation, the weights are used to adjust the contribution size of the input signal, affecting the change of the neuron's output activation value.

[0077] The gradient is the local rate of change of the objective function with respect to the model parameters, reflecting the change trend of the loss function value when the parameters change. The gradient is used to guide the direction and magnitude of weight update during the optimization process, enabling the entire network to gradually improve its performance in the direction of reducing the prediction error. In the backpropagation stage, the core of gradient calculation is the partial derivatives of the loss function with respect to each weight and bias. Once the forward propagation has completed the network output calculation and calculated the value of the loss function, the backpropagation process will start, calculating the gradients of the loss function with respect to the weights and biases of each layer layer by layer.

[0078] For example, in step S103, the parameters refer to the model parameters, including the trainable weights and biases in a neural network. During the training process, the model parameters are not only used multiple times in forward propagation and backpropagation, but also change with the update of the optimizer. The updated parameters refer to the model parameters obtained after the optimization update task and are used for the next round of iterative training.

[0079] For example, in step S103, the optimization and update task may include, for example, gradient synchronization, parameter optimization, and parameter update.

[0080] Gradient synchronization refers to the process of aggregating the gradient values calculated by multiple devices in a distributed training scenario to ensure the consistency of the global gradient when multiple devices perform parallel computing of local gradients. Gradient synchronization is usually achieved through communication operations such as All-Reduce, so that the gradient data on multiple devices is consistent.

[0081] Parameter optimization refers to the process of iteratively adjusting model parameters based on gradient data. Its essence is to search for the optimal solution that minimizes the loss function in the parameter space through optimization algorithms. Optimization algorithms can be, for example, the Stochastic Gradient Descent (SGD) algorithm, the Adaptive Moment Estimation (Adam) optimizer, etc.

[0082] Parameter update refers to the process of actually adjusting model parameters according to the update rules provided by the optimization algorithm. This means modifying the current model weights and biases in the manner determined by the optimization algorithm, so as to gradually approach the optimal solution. Parameter update is directly related to the improvement of model performance and is the specific operation to improve the model after each iteration.

[0083] For example, in step S103, for each stage of the target model, once the backpropagation task is completed, the device immediately triggers the execution of the optimization and update task. It should be noted that the optimization and update task can be executed by the device itself or triggered by the device to the host for execution. For example, in the case of adopting the micro-batch pipeline parallel mode, for each stage of the target model, when the backpropagation task of the last micro-batch is completed, the device immediately triggers the execution of the optimization and update task. Compared with Figure 1D the need to wait for all micro-batches of all stages to complete the backpropagation task (that is, after B 0,0 completes the backpropagation task) and then uniformly execute the optimization and update tasks of each stage, the distributed training method proposed in this embodiment of the present disclosure can reduce the waiting time and effectively improve the pipeline efficiency and parallelism.

[0084] Figure 3B It is a timing diagram of a distributed training method provided by at least one embodiment of the present disclosure.

[0085] For example, Figure 3BThe timing relationship between the backward calculation task and the optimization update task in four stages (Stage 0 to Stage 3) is shown. Through the above step S103, the optimization update task in Stage 3 overlaps with the backward calculation task in Stage 2 in time, the optimization update task in Stage 2 overlaps with the backward calculation task in Stage 1 in time, and the optimization update task in Stage 1 overlaps with the backward calculation task in Stage 0 in time. This way effectively hides the communication overhead brought by the optimization update tasks in Stage 1 to Stage 3, thereby significantly improving the resource utilization rate and the overall training efficiency.

[0086] Figure 3C The flowchart of a distributed training method provided by at least one embodiment of the present disclosure.

[0087] For example, as Figure 3C shown, for at least one stage of the target model, an example of step S103 may include the following steps S1031 to S1032.

[0088] Step S1031: In response to the completion of the backward calculation task, the device transfers at least part of the gradient data to the host and triggers the host to execute the optimization update task.

[0089] Step S1032: The host executes the optimization update task based on at least part of the gradient data, obtains at least part of the updated parameters, and transfers at least part of the updated parameters to the device for use in the next round of iterative training.

[0090] For example, in step S1031 and step S1032, for the current stage, once the backward calculation task is completed, the device immediately transfers at least part of the gradient data to the host, and at the same time triggers the host to execute the optimization update task. For example, the device can send a trigger signal to the host to trigger the host to execute the optimization update task. When receiving the trigger signal from the device, the host can execute the optimization update task based on at least part of the gradient data, obtain at least part of the updated parameters, and transfer at least part of the updated parameters to the device for use in the next round of iterative training.

[0091] For example, the device can transfer gradient data to the host through D2H transfer or network communication, and the host can transfer updated parameters to the device through H2D transfer or network communication. Specifically, within the same computing node, the device can transfer gradient data to the host through D2H transfer, and the host can transfer updated parameters to the device through H2D transfer. Between different computing nodes, the device can transfer gradient data to the host through network communication, and the host can transfer updated parameters to the device through network communication. For example, network communication can be implemented based on the network card and GPU Direct mechanism, and the GPU Direct mechanism allows high-speed data transfer directly from the GPU to the network card. According to actual needs, the device can transfer all gradient data obtained through the reverse calculation task to the host, or only transfer partial gradient data to the host.

[0092] For example, in the data parallel mode, there may be multiple data parallel groups related to the current stage, and gradient aggregation is required among devices within the same data parallel group. Gradient aggregation refers to the process of summarizing gradient data calculated on multiple devices, usually involving merging local gradients on multiple devices into a global gradient. These devices may be located in multiple different computing nodes, and each device has a host physically connected to it. For each device within the same data parallel group, the device can transfer all gradient data obtained through the reverse calculation task to each host (here the host refers to the host physically connected to the above-mentioned devices); the device can also divide all gradient data into multiple parts and then transfer each part to different hosts respectively.

[0093] For example, in the case of not adopting the data parallel mode and the tensor parallel mode, there is only one device corresponding to the current stage. This device can transfer all gradient data to the host physically connected to the device through D2H transfer and trigger the host to execute the optimization update task. For example, the host can execute the optimization update task based on the above-mentioned all gradient data, obtain the updated parameters, and transfer the updated parameters to the device through H2D transfer for use in the next round of iterative training.

[0094] It should be noted that the current stage includes multiple layers. When the reverse calculation task of each layer is completed, the device will immediately transfer at least partial gradient data to the host and trigger the host to execute the optimization update task, without waiting for the reverse calculation tasks of all layers in the current stage to be completed. Through the above method, the pipeline efficiency and parallelism can be effectively improved, and efficient overlap of communication and computing can be achieved.

[0095] For example, in the distributed training method provided by at least one embodiment of the present disclosure, the "at least one stage" in the above text may include a first stage, that is, the above steps S1031 to S1032 may be executed only for the first stage. The first stage here refers to the first stage executed in sequence during the forward propagation of the target model (for example Figure 3B Stage 0 in

[0096] , which includes several layers at the very front of the target model in the forward propagation order. Figure 3B By the above steps S1031 to S1032, the overlap of communication and computing can be achieved, effectively alleviating the communication bottleneck problem introduced by the lack of overlap in the pipeline timing of the optimization update task in the first stage (for example

[0097] Stage 0 in

[0098] ), and significantly improving the efficiency and performance of the entire pipeline.

[0099] It should be noted that in addition to the first stage, for other stages of the target model, the above steps S1031 to S1032 may also be executed according to actual needs, and the embodiments of the present disclosure do not limit this.

[0100] In the distributed training method provided by one of the above embodiments of the present disclosure, by migrating the optimization update task from the device to the host and having the host complete the optimization update task, the computing resources of the device can be saved, the pipeline efficiency and parallelism can be effectively improved, the efficient overlap of communication and computing can be achieved, thereby alleviating the communication bottleneck problem in large-scale distributed deep learning training. Moreover, the above method makes full use of the host and memory resources, realizes heterogeneous parallel computing with the device, reduces the interference between hardware, and avoids the influence brought by resource competition and potential hardware failures.

[0101] It should be noted that if, according to actual needs, the above steps S1031 to S1032 are also executed for some stages other than the first stage, the configuration of the corresponding host can also be improved accordingly. The embodiments of the present disclosure do not limit this.

[0102] In the distributed training method provided by at least one embodiment of the present disclosure, while adopting the pipeline parallel mode, the devices in multiple computing nodes can also be divided into multiple data parallel groups according to the data parallel mode.

[0103] The following gives an example of the tensor parallel - data parallel - pipeline parallel mode, which is divided in the order of tensor parallel, data parallel, and pipeline parallel. In this example, it is assumed that the target model is divided in the manner of TP4 - DP2 - PP2 (tensor parallelism is 4, data parallelism is 2, and pipeline parallelism is 2), then a total of 16 devices (denoted as G0 to G15) are required. Among them, the tensor parallelism of 4 means that a single model layer is divided into 4 parts and calculated on 4 devices; the data parallelism of 2 means that each model copy uses 2 different data for training, and finally gradient aggregation is required to update the model parameters; the pipeline parallelism of 2 means that the entire target model is divided into two stages and executed in a pipeline.

[0104] Through the above partitioning operation, 4 tensor parallel groups, 8 data parallel groups, and 8 pipeline parallel groups can be obtained.

[0105] The 4 tensor parallel groups include: [G0, G1, G2, G3], [G4, G5, G6, G7], [G8, G9, G10, G11], [G12, G13, G14, G15]. The devices within each tensor parallel group jointly complete the tensor operations in a model layer.

[0106] The 8 data parallel groups include: [G0, G4], [G1, G5], [G2, G6], [G3, G7], [G8, G12], [G9, G13], [G10, G14], [G11, G15]. Gradient aggregation is required between the devices within the same data parallel group. For example, gradient aggregation is required between G0 and G4.

[0107] The 8 pipeline parallel groups include: [G0, G8], [G1, G9], [G2, G10], [G3, G11], [G4, G12], [G5, G13], [G6, G14], [G7, G15]. For example, G0 to G7 correspond to the first stage, and G8 to G15 correspond to the second stage.

[0108] In the above example, the data parallel groups participating in the first-stage gradient aggregation include [G0, G4], [G1, G5], [G2, G6], and [G3, G7]. The above 4 data parallel groups can be referred to as the data parallel groups related to the first stage.

[0109] It should be noted that the above partitioning method is only an example, and the embodiments of the present disclosure do not limit the partitioning mode, parallelism degree, and order. For example, it can also be partitioned in the order of tensor parallelism - pipeline parallelism - data parallelism. The specific partitioning method is similar to the above, and will not be elaborated here.

[0110] In the distributed training method provided by at least one embodiment of the present disclosure, an example of the above step S1031 may include the following step S201. It should be noted that since step S1031 is executed for "at least one stage" of the target model, the "data parallel group" below should be understood as the data parallel group related to the current stage (i.e., the "at least one stage" targeted by step S1031), and does not involve other stages.

[0111] Step S201: For each data parallel group related to the current stage, in response to the completion of the backpropagation calculation task, each device in the data parallel group transfers the gradient data to the corresponding host through D2H transmission.

[0112] For example, in step S201, the "corresponding host" refers to the host physically connected to the device, such as the host connected through a switch. For example, in each data parallel group related to the current stage, the following operations are performed: Each device in the data parallel group first completes the forward calculation and backpropagation calculation locally to obtain the gradient data. When the backpropagation calculation task is completed, each device transfers all the gradient data it generates to the host physically connected to it through D2H transmission.

[0113] Corresponding to step S201, an example of the above step S1032 may include performing the following operations in each data parallel group related to the current stage: The host corresponding to each device in the data parallel group executes the following steps S202 to step S204.

[0114] Step S202: The host divides the gradient data transferred by the device into N pieces of data, and sends N - 1 pieces of the N pieces of data to the hosts corresponding to the other devices in the data parallel group, where N is the number of devices in the data parallel group.

[0115] For example, in step S202, for each host (i.e., the host corresponding to each device in the data parallel group), the received gradient data is divided into N parts, one part is retained by itself, and the remaining N - 1 parts are respectively sent to the hosts corresponding to the other devices in the data parallel group. For the convenience of description, the host that sends the data is referred to as the source host, and the host that receives the data is referred to as the destination host. If the source host and the destination host are not within the same computing node, the source host can send the gradient data through network communication. If they are within the same computing node, the source host can send the gradient data through intra-machine interconnection (such as the UPI bus introduced above).

[0116] Step S203: The host performs an optimization update task based on all the received gradient data to obtain at least partially updated parameters.

[0117] For example, in step S203, for each host, all the received gradient data includes one part of data retained by itself (i.e., the one part of data not sent in step S202), and the data sent by the hosts corresponding to the other devices in the data parallel group (a total of N - 1 parts).

[0118] Step S204: The host obtains the updated parameters generated by the hosts corresponding to the other devices in the data parallel group, and transfers all the updated parameters to the device through H2D transfer for use in the next round of iterative training.

[0119] For example, in step S204, if the source host and the destination host are not within the same computing node, the destination host can receive the updated parameters through network communication. If they are within the same computing node, the destination host can receive the updated parameters through intra-machine interconnection (such as the UPI bus introduced above). For each host, after integrating all the received updated parameters, the integrated parameters can be transferred back to the device physically connected to the host through H2D transfer for use in the next round of iterative training.

[0120] In the distributed training method provided in the above embodiments of the present disclosure, through data partitioning, the parameter optimization task can be dispersed to multiple hosts for collaborative completion, which can effectively improve the computing efficiency and reduce the redundancy of data transmission.

[0121] Figure 4 It is an exemplary schematic diagram of a distributed training method provided by at least one embodiment of the present disclosure. Figure 4 It is an example of the above steps S201 to S204.

[0122] For example, Figure 4 shows Figure 3BThe execution process of the reverse calculation task and the optimization update task in the first stage (Stage 0). In this example, Stage 0 includes multiple layers, Figure 4 and only four of them (Layer 0 to Layer 3) are shown. The reverse calculation task is performed layer by layer.

[0123] The following takes the execution process of the reverse calculation task and the optimization update task of Layer 3 as an example for introduction. The execution processes of Layer 0 to Layer 2 are the same and will not be elaborated here. The following operations are performed for each data parallel group related to Stage 0:

[0124] For example, as Figure 4 shown, when the reverse calculation task of Layer 3 is completed, each device in the data parallel group transfers the gradient data of Layer 3 to the corresponding host through D2H transfer.

[0125] It should be noted that since the D2H transfer time of the gradient data of each layer is longer than the execution time of the reverse calculation task of that layer, in the Figure 4 shown timing diagram, "Layer 3 gradient data D2H" does not follow "Stage 0 Layer 3" closely in time, but there is a waiting time. Specifically, the dashed box before "Layer 3 gradient data D2H" in the figure represents the D2H transfer processes of the gradient data of Layer 5 and Layer 4 that are not shown. When the D2H transfer of Layer 4 has not been completed, the reverse calculation of Layer 3 has been completed, resulting in the D2H transfer of the gradient data of Layer 3 needing to wait for the transfer of Layer 4 to complete. This phenomenon is manifested as a delay before the D2H transfer of the gradient data of Layer 3 in the timing diagram.

[0126] In the Figure 4 Reduce-Scatter stage, each host (referring to the host corresponding to each device in the data parallel group) divides the gradient data of Layer 3 transferred by the device into N pieces of data, and sends N - 1 pieces of the N pieces of data to the hosts corresponding to other devices in the data parallel group respectively, where N is the number of devices in the data parallel group.

[0127] In the Figure 4 sum stage and opt stage, each host performs the optimization update task based on all the received gradient data to obtain at least part of the updated parameters of Layer 3. For example, the sum stage corresponds to gradient synchronization, and the opt stage corresponds to parameter optimization.

[0128] In the Figure 4In the All-gather phase, each host obtains the updated parameters generated by the hosts corresponding to the other devices in the data parallel group, obtains all the updated parameters of Layer 3, and transfers all the updated parameters of Layer 3 to the corresponding device through H2D transfer for use in the next round of iterative training.

[0129] As Figure 4 shown, when the backward calculation task of each layer is completed, the device immediately transfers the gradient data to the host and triggers the host to execute the optimization update task, without waiting for all layers in the current phase to complete the backward calculation task. By the above method, the pipeline efficiency and parallelism can be effectively improved, and the efficient overlap of communication and calculation can be realized.

[0130] In the above method, each device in the data parallel group needs to transfer all the gradient data generated by itself to the corresponding host through D2H transfer, and then the host divides the data and distributes it. To further improve the data transfer efficiency, the data can be directly divided at the device and distributed to different hosts, reducing the amount of data transferred to the corresponding host and effectively reducing the time overhead in the data transfer process. Based on this, another example is provided below.

[0131] Another example of the above step S1031 may include the following step S301. It should also be noted that since step S1031 is executed for "at least one phase" of the target model, the "data parallel group" below should be understood as the data parallel group related to the current phase and does not involve other phases.

[0132] Step S301: For each data parallel group related to the current phase, in response to the completion of the backward calculation task, each device in the data parallel group divides the gradient data into N pieces of data, transfers one piece of data among the N pieces of data to the corresponding host through D2H transfer, and sends the remaining N - 1 pieces of data among the N pieces of data to the hosts corresponding to the other devices in the data parallel group, where N is the number of devices in the data parallel group.

[0133] For example, in step S301, the "corresponding host" refers to the host physically connected to the device, such as the host connected through a switch. For example, in the data parallel group related to the current stage, the following operations are all performed: Each device in the data parallel group first completes the forward calculation and the backward calculation locally to obtain gradient data. When the backward calculation task is completed, each device divides all the gradient data it generates into N parts of data, and transmits one part of the data to the host physically connected to it through D2H transmission, and sends the remaining N - 1 parts of data to the hosts corresponding to other devices in the data parallel group respectively. For the convenience of description, the device that sends the data is referred to as the source device here, and the host that receives the data is referred to as the destination host. If the source device and the destination host are not within the same computing node, the source device can send the gradient data through network communication. If they are within the same computing node, the source device can send the gradient data through intra-machine interconnection. For example, within the same computing node, multiple devices can be connected through a high-speed bus. An example of a source device within the same computing node sending gradient data through intra-machine interconnection is as follows: Taking Figure 2B device 0 in Figure 2B as the source device and host 1 as the destination host as an example, device 0 can send the gradient data to host 0 through D2H transmission, and then send the gradient data to host 1 through the UPI bus (for example, the transmission between host 0 and host 1 can be performed implicitly without manual control); or, device 0 can send the gradient data to device 4 through the high-speed bus, and device 4 then sends the gradient data to host 1 through D2H transmission.

[0134] Corresponding to step S301, an example of the above step S1032 can include performing the following operations in each data parallel group related to the current stage: The host corresponding to each device in the data parallel group executes the following step S302.

[0135] Step S302: The host performs an optimization update task based on all the received gradient data to obtain at least some updated parameters, and sends at least some of the updated parameters to all the devices in the data parallel group for use in the next round of iterative training.

[0136] For example, in step S302, for each host (referring to the host corresponding to each device in the data parallel group), all the received gradient data includes a copy of the data transmitted by the device physically connected to it, as well as the data sent by other devices in the data parallel group (a total of N - 1 copies). Each host can perform an optimization update task based on all the received gradient data, obtain partially updated parameters, and send them to all the devices in the data parallel group. For the convenience of description, the host that sends the parameters is referred to as the source host here, and the device that receives the parameters is referred to as the destination device. If the source host and the destination device are not within the same computing node, the source host can send the updated parameters through network communication. If they are within the same computing node, the source host can send the updated parameters through intra-machine interconnection. For example, within the same computing node, multiple devices can be connected through a high-speed bus. An example of the source host sending the updated parameters through intra-machine interconnection within the same computing node is as follows: Taking Figure 2B host 0 in the middle as the source host and device 4 as the destination device as an example, host 0 can send the updated parameters to device 0 through H2D transfer, and device 0 can send the updated parameters to device 4 through the high-speed bus; alternatively, host 0 can send the updated parameters to host 1 through the UPI bus, and host 1 can then send the updated parameters to device 4 through H2D transfer. Through the above method, each device in the data parallel group can obtain the complete updated parameters for use in the next round of iterative training.

[0137] Through the above method, directly partitioning the data at the device and distributing it to different hosts can reduce the amount of data transmitted to the corresponding host, effectively reducing the time overhead during the data transmission process and further improving the data transmission efficiency.

[0138] Figure 5 It is an exemplary schematic diagram of a distributed training method provided by at least one embodiment of the present disclosure. Figure 5 It is an example of the above steps S301 to S302.

[0139] For example, as Figure 5 shown, a data parallel group includes four devices (device 0 to device 3). Here, it is assumed that the four devices are located on different computing nodes respectively.

[0140] For example, as Figure 5 shown, in response to the completion of the reverse calculation task, device 0 divides all the gradient data it generates (stored in the device memory) into 4 copies of data ( Figure 5 data a0 to data a3 in Figure 5Host 0, which is connected to Device 0 through a switch, and is stored in the host memory. Device 0 sends Data a1 to Data a3 to the corresponding hosts of Device 1 to Device 3 respectively through network communication.

[0141] To avoid the information in the figure being too complex and affecting understanding, Figure 5 only the process of Device 0 sending gradient data is shown in the figure, while the processes of Device 1 to Device 3 sending gradient data are not shown. These processes will be described in detail below in words.

[0142] In response to the completion of the reverse calculation task, Device 1 divides all the gradient data it generates (stored in the device memory) into 4 portions of data ( Figure 5 Data b0 to Data b3 in it), and transfers one portion of data b1 to the corresponding host through D2H transmission (that is, Figure 5 Host 1 connected to Device 1 through a switch in it), and is stored in the host memory. Device 1 sends Data b0, Data b2, and Data b3 to the corresponding hosts of Device 0, Device 2, and Device 3 respectively through network communication.

[0143] In response to the completion of the reverse calculation task, Device 2 divides all the gradient data it generates (stored in the device memory) into 4 portions of data ( Figure 5 Data c0 to Data c3 in it), and transfers one portion of data c2 to the corresponding host through D2H transmission (that is, Figure 5 Host 2 connected to Device 2 through a switch in it), and is stored in the host memory. Device 2 sends Data c0, Data c1, and Data c3 to the corresponding hosts of Device 0, Device 1, and Device 3 respectively through network communication.

[0144] In response to the completion of the reverse calculation task, Device 3 divides all the gradient data it generates (stored in the device memory) into 4 portions of data ( Figure 5 Data d0 to Data d3 in it), and transfers one portion of data d3 to the corresponding host through D2H transmission (that is, Figure 5 Host 3 connected to Device 3 through a switch in it), and is stored in the host memory. Device 3 sends Data d0 to Data d2 to the corresponding hosts of Device 0 to Device 2 respectively through network communication.

[0145] Through the above operations, data a0, data b0, data c0, and data d0 are stored in the host memory of host 0. Host 0 performs an optimization update task based on the above data, obtains the updated parameter p0, and transfers the updated parameter p0 to device 0 through H2D transfer, and broadcasts and sends the updated parameter p0 to devices 1 to 3 through network communication. Data a1, data b1, data c1, and data d1 are stored in the host memory of host 1. Host 1 performs an optimization update task based on the above data, obtains the updated parameter p1, and transfers the updated parameter p1 to device 1 through H2D transfer, and broadcasts and sends the updated parameter p1 to devices 0, 2, and 3 through network communication. The situations of host 2 and host 3 are the same and will not be elaborated here. Through the above operations, each device obtains the updated parameters p0 to p3 for use in the next round of iterative training.

[0146] It should be noted that the above method is not only applicable to the architecture adopting the data parallel - pipeline parallel strategy, but can also be further applicable to the architecture adopting the tensor parallel - data parallel - pipeline parallel strategy.

[0147] In the distributed training method provided in at least one embodiment of the present disclosure, network communication can be based on the Remote Direct Memory Access (RDMA) mechanism. For example, the RDMA over Converged Ethernet v2 (RoCEv2) network protocol based on Ethernet can support the RDMA mechanism. For example, the above-mentioned broadcast sending can be implemented through the RDMA mechanism. By utilizing the broadcast function of RDMA, the number of transfers required during data transmission can be significantly reduced. In large-scale distributed training scenarios (such as N > 16), the RDMA broadcast mechanism can reduce the number of data transfers to 1 / (N - 1) of the original, thereby greatly improving the data transmission efficiency. Theoretically, the larger the scale, the lower the communication overhead.

[0148] In the distributed training method provided in at least one embodiment of the present disclosure, a U-shaped pipeline can also be constructed. For example, the distributed training method can further include the following step S104.

[0149] Step S104: Assign the i-th stage and the (M - i + 1)-th stage to the same computing node, where 0 < i ≤ M / 2, M is the total number of stages obtained by dividing the target model, and the i-th stage represents the i-th stage executed in sequence during the forward propagation of the target model.

[0150] For example, in step S104, each computing node can undertake the computing tasks of two stages, and the computing tasks of the two stages are staggered in time. In this way, the first stage of the target model (i.e., the first stage mentioned above) and the last stage (in the case of having M stages, the last stage is the Mth stage) can be allocated to the same computing node. For ease of description, assume that the devices in a computing node are divided into two groups, where the first group of devices undertakes the computing tasks of the first stage of the target model, and the second group of devices undertakes the computing tasks of the last stage. In the forward computing stage, when the second group of devices starts to execute the forward computing tasks, the forward computing tasks of the first group of devices have been completed and stopped running. At this time, the first group of devices will not compete with the second group of devices for computing resources and communication resources, so that the computing resources and communication resources available to the second group of devices (such as available network bandwidth) are increased to twice the original. The same is true for the backward computing stage. When the first group of devices starts to execute the backward computing tasks, the backward computing tasks of the second group of devices have been completed and stopped running. At this time, the second group of devices will not compete with the first group of devices for computing resources and communication resources, so that the computing resources and communication resources available to the first group of devices (such as available network bandwidth) are increased to twice the original. For example, step S104 can be applied to an architecture adopting the tensor parallel - pipeline parallel strategy, or can also be applied to an architecture adopting the tensor parallel - data parallel - pipeline parallel strategy.

[0151] Through the above - mentioned method, the available computing resources and communication resources of the computing node where the first stage is located can be effectively improved, further alleviating the communication bottleneck problem in large - scale distributed deep learning training.

[0152] Figure 6 It is a schematic diagram of a folded pipeline provided by at least one embodiment of the present disclosure.

[0153] For example, as Figure 6 shown, this example adopts the method of TP4 - PP8 (tensor parallelism degree is 4, pipeline parallelism degree is 8) for partitioning. Figure 6 It is a simplified schematic diagram, only showing the devices in each computing node, and components such as hosts and network cards are not shown in the figure. Figure 6 Four computing nodes (computing node 0 to computing node 3) are shown, and each computing node includes 8 devices (device 0 to device 7). Figure 6 Each elliptical dotted box in it represents a stage of the target model, and the order of the arrows represents the execution order of each stage during forward computing. Figure 6Eight stages are shown, where within each stage, four devices are used to achieve tensor parallelism. For example, devices 0 to 3 of computing node 0 implement tensor parallelism in the first stage, devices 0 to 3 of computing node 1 implement tensor parallelism in the second stage, and so on. Devices 4 to 7 of computing node 0 implement tensor parallelism in the eighth stage. Thus, it can be seen that the first stage (the first stage) and the last stage (the eighth stage) of the target model are both executed on computing node 0. In the forward calculation stage, when devices 4 to 7 of computing node 0 start to execute the forward calculation task, the forward calculation tasks of devices 0 to 3 of computing node 0 have been completed and stopped running. At this time, devices 0 to 3 do not compete with devices 4 to 7 for computing resources and communication resources, so that the computing resources and communication resources available to devices 4 to 7 are increased to twice the original. The same is true for the backward calculation stage. When devices 0 to 3 of computing node 0 start to execute the backward calculation task, the backward calculation tasks of devices 4 to 7 of computing node 0 have been completed and stopped running. At this time, devices 4 to 7 do not compete with devices 0 to 3 for computing resources and communication resources, so that the computing resources and communication resources available to devices 0 to 3 are increased to twice the original. For example, assuming that each computing node includes eight network cards, in the backward calculation stage, devices 0 to 3 can use all eight network cards for communication, and the available network bandwidth is increased to twice the original, effectively improving the network bandwidth utilization rate.

[0154] It should also be noted that in various embodiments of the present disclosure, the execution order of the steps of the distributed training method is not limited. Although the execution processes of the steps are described in a specific order above, this does not constitute a limitation on the embodiments of the present disclosure. The steps in the distributed training method can be executed serially or in parallel, which can be determined according to actual needs.

[0155] For example, compared with the above description, the distributed training method provided by at least one embodiment of the present disclosure may further include more or fewer steps, and the embodiments of the present disclosure do not limit this.

[0156] At least one embodiment of the present disclosure further provides a distributed training system. For example, referring to the Figure 2A and Figure 2B described above, the distributed training system provided by at least one embodiment of the present disclosure includes a plurality of computing nodes, and each computing node includes at least one host and at least one device. The distributed training system is configured to perform distributed training on a target model.

[0157] For example, the target model is divided into multiple stages in a pipeline parallel mode, each stage includes at least one layer of the target model, and the multiple stages are respectively distributed in multiple devices of multiple computing nodes. In the current round of iterative training, the training of at least one stage of the target model is jointly executed by at least one corresponding device and at least one corresponding host.

[0158] For each stage of the target model, the device is configured to: perform a forward computing task to obtain a forward computing result; perform a backward computing task based on the forward computing result to obtain gradient data; in response to the completion of the backward computing task, trigger the execution of an optimization update task and obtain updated parameters for use in the next round of iterative training.

[0159] For at least one stage of the target model, the device is further configured to: in response to the completion of the backward computing task, transfer at least part of the gradient data to the host and trigger the host to execute an optimization update task; the host is configured to: perform an optimization update task based on at least part of the gradient data to obtain at least part of the updated parameters, and transfer at least part of the updated parameters to the device for use in the next round of iterative training.

[0160] For example, in at least one embodiment of the present disclosure, the devices in multiple computing nodes are divided into multiple data parallel groups in a data parallel mode, and gradient aggregation is performed between the devices within the same data parallel group. For each data parallel group related to the above at least one stage, each device in the data parallel group is further configured to: in response to the completion of the backward computing task, transfer the gradient data to the corresponding host through D2H transmission. In each data parallel group related to the above at least one stage, the host corresponding to each device in the data parallel group is further configured to: divide the gradient data transferred by the device into N pieces of data, and send N-1 pieces of the N pieces of data to the hosts corresponding to other devices in the data parallel group respectively, where N is the number of devices in the data parallel group; perform an optimization update task based on all the received gradient data to obtain at least part of the updated parameters; obtain the updated parameters generated by the hosts corresponding to other devices in the data parallel group, and transfer all the updated parameters to the device through H2D transmission for use in the next round of iterative training.

[0161] For example, in at least one embodiment of the present disclosure, for each data parallel group associated with the above at least one stage, each device in the data parallel group is further configured to: in response to the completion of the reverse calculation task, divide the gradient data into N pieces of data, transfer one piece of data out of the N pieces of data to the corresponding host through D2H transmission, and send the remaining N - 1 pieces of data out of the N pieces of data to the hosts corresponding to other devices in the data parallel group, where N is the number of devices in the data parallel group. In each data parallel group associated with the above at least one stage, the host corresponding to each device in the data parallel group is further configured to: perform an optimization update task based on all the received gradient data to obtain at least partial updated parameters, and send the at least partial updated parameters to all the devices in the data parallel group for use in the next round of iterative training.

[0162] The distributed training system provided by at least one embodiment of the present disclosure further includes an allocation module, and the allocation module is configured to allocate the i-th stage and the (M - i + 1)-th stage to the same computing node, where 0 < i ≤ M / 2, M is the total number of stages obtained by dividing the target model, and the i-th stage represents the i-th stage sequentially executed during the forward propagation process of the target model.

[0163] For example, in at least one embodiment of the present disclosure, the at least one stage includes a first stage, and the first stage includes the layer at the very front of the target model in the forward propagation order.

[0164] For example, in at least one embodiment of the present disclosure, the configuration of the host corresponding to the first stage is higher than the configurations of the hosts corresponding to other stages.

[0165] For example, in at least one embodiment of the present disclosure, each computing node further includes at least one network card, and at least one host is configured to drive at least one network card to implement network communication between multiple computing nodes.

[0166] For example, in at least one embodiment of the present disclosure, the network communication is based on the remote direct memory access mechanism.

[0167] It should be noted that the above various modules can be implemented by software, hardware, firmware, or any combination thereof. For example, the allocation module can be implemented as an allocation circuit, and the embodiments of the present disclosure do not limit the specific implementation manners thereof.

[0168] It should be understood that the distributed training system provided by at least one embodiment of the present disclosure can be used to implement the foregoing distributed training method, and can also achieve technical effects similar to those of the foregoing distributed training method, which will not be elaborated herein.

[0169] It should be noted that, in the embodiments of the present disclosure, the distributed training system may include more or fewer modules or units, and the connection relationships between the various modules or units are not limited and may be determined according to actual needs. The specific composition manners of the various modules or units are not limited and may be composed of analog devices according to circuit principles, or may be composed of digital chips, or may be composed in other applicable manners.

[0170] Figure 7 Schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.

[0171] For example, as Figure 7 shown, the electronic device 700 includes at least one processor 701 and at least one memory 702. The at least one memory 702 includes one or more computer program modules. The one or more computer program modules are stored in the memory 702 and are configured to be executed by the at least one processor 701. The one or more computer program modules include instructions for executing the above-mentioned distributed training method. When executed by the at least one processor 701, one or more steps in the distributed training method provided by at least one embodiment of the present disclosure can be executed. The memory 702 and the processor 701 may be interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0172] For example, the processor 701 may be a central processing unit (CPU), a digital signal processor (DSP), an image processor (GPU), a general-purpose graphics processing unit (GPGPU), an artificial intelligence (AI) accelerator, or other forms of processing units having data processing capabilities and / or program execution capabilities, such as a field programmable gate array (FPGA), etc.; for example, the central processing unit (CPU) may be of the X86, ARM, RISC-V architectures, etc. The processor 701 may be a general-purpose processor or a dedicated processor and may control other components in the electronic device 700 to perform desired functions.

[0173] For example, the memory 702 may include any combination of one or more computer program products. The computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc.

[0174] Figure 8Schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure.

[0175] The electronic device in at least one embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The illustrated electronic device is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0176] The electronic device includes at least one processor and a memory. The processor here may be referred to as the processing device 801 described below, and the memory may include at least one of the read-only memory (ROM), random access memory (RAM), and storage device 808 described below. The memory is used to store programs for executing the methods described in the above various method embodiments; the processor is configured to execute the programs stored in the memory. The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0177] As Figure 8 shown, the electronic device 800 may include a processing device 801 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to programs stored in the read-only memory (ROM) or programs loaded from the storage device 808 into the random access memory (RAM). In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, ROM 802, and RAM 803 are connected to each other via a bus 804. The input / output (I / O) interface is also connected to the bus 804.

[0178] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a display, a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 the electronic device 800 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0179] In particular, according to at least one embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, at least one embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above functions defined in the method of at least one embodiment of the present disclosure are performed.

[0180] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In at least one embodiment of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in at least one embodiment of the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0181] The above computer-readable medium can be included in the above electronic device 800; or it can exist separately and not be assembled into the electronic device 800.

[0182] Figure 9 A schematic block diagram of a non-transitory computer-readable storage medium provided for at least one embodiment of the present disclosure.

[0183] For example, as Figure 9 shown, computer-readable instructions 901 are stored on a non-transitory computer-readable storage medium 900. When the computer-readable instructions 901 are executed by at least one processor, one or more steps of the above-described distributed training method are performed.

[0184] For example, the storage medium may include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a flash memory, or any combination of the above storage media, and may also be other applicable storage media. For example, the readable storage medium may also be the Figure 7 memory 702 in, and the related description may refer to the foregoing content and will not be elaborated herein.

[0185] Although the present disclosure has been described in detail above with general descriptions and specific embodiments, on the basis of the embodiments of the present disclosure, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present disclosure all fall within the scope of protection required by the present disclosure.

[0186] For the present disclosure, the following points need to be noted:

[0187] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0188] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present disclosure, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn according to the actual scale.

[0189] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0190] The above is only the specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A distributed training method, applied to a distributed training system, wherein, the distributed training system is configured to perform distributed training on a target model, the distributed training system includes a plurality of computing nodes, each computing node includes at least one host and at least one device, wherein, the target model is divided into multiple stages in a pipeline parallel mode, each stage includes at least one layer of the target model, and the multiple stages are respectively distributed in multiple devices of the multiple computing nodes, wherein, in the current round of iterative training, the training of at least one stage of the target model is jointly executed by at least one corresponding device and at least one corresponding host, the distributed training method includes: For each stage of the target model, the device performs a forward calculation task to obtain a forward calculation result; the device performs a backward calculation task based on the forward calculation result to obtain gradient data; in response to the completion of the backward calculation task, the device triggers the execution of an optimization update task and obtains updated parameters for use in the next round of iterative training; For at least one stage of the target model, the step of, in response to the completion of the backward calculation task, the device triggers the execution of an optimization update task and obtains updated parameters for use in the next round of iterative training, includes: in response to the completion of the backward calculation task, the device transfers at least part of the gradient data to the host and triggers the host to execute the optimization update task; the host performs an optimization update task based on at least part of the gradient data to obtain at least part of the updated parameters, and transfers the at least part of the updated parameters to the device for use in the next round of iterative training.

2. The method according to claim 1, wherein The devices in the multiple computing nodes are divided into multiple data parallel groups in a data parallel mode, and gradient aggregation is performed between the devices within the same data parallel group, the step of, in response to the completion of the backward calculation task, the device transfers at least part of the gradient data to the host, includes: For each data parallel group related to the stage, in response to the completion of the backward calculation task, each device in the data parallel group transfers the gradient data to the corresponding host through device-to-host transmission.

3. The method according to claim 2, wherein, The step of, the host performs an optimization update task based on at least part of the gradient data to obtain at least part of the updated parameters, and transfers the at least part of the updated parameters to the device for use in the next round of iterative training, includes: In each data parallel group related to the stage, for the host corresponding to each device in the data parallel group, the host divides the gradient data transferred by the device into N parts of data, and sends N - 1 parts of the N parts of data to the hosts corresponding to other devices in the data parallel group respectively, where N is the number of devices in the data parallel group; the host performs an optimization update task according to all the received gradient data to obtain at least part of the updated parameters; the host obtains the updated parameters generated by the hosts corresponding to other devices in the data parallel group, and transfers all the updated parameters to the device through host-to-device transmission for use in the next round of iterative training.

4. The method according to claim 1, wherein The devices in the multiple computing nodes are divided into multiple data parallel groups in a data parallel mode, and gradient aggregation is performed among the devices within the same data parallel group. In response to the completion of the backward calculation task, the device transfers at least part of the gradient data to the host, including: For each data parallel group related to the stage, In response to the completion of the backward calculation task, each device in the data parallel group divides the gradient data into N parts of data, transfers one part of the N parts of data to the corresponding host through device-to-host transmission, and sends the remaining N-1 parts of the N parts of data to the corresponding hosts of other devices in the data parallel group, where N is the number of devices in the data parallel group.

5. The method according to claim 4, wherein, The host performs an optimization update task based on at least part of the gradient data, obtains at least part of the updated parameters, and transfers the at least part of the updated parameters to the device for use in the next round of iterative training, including: In each data parallel group related to the stage, for the host corresponding to each device in the data parallel group, The host performs an optimization update task based on all the received gradient data, obtains at least part of the updated parameters, and sends the at least part of the updated parameters to all the devices in the data parallel group for use in the next round of iterative training.

6. The method according to claim 1, further comprising: Assigning the i-th stage and the (M - i + 1)-th stage to the same computing node, where 0 < i ≤ M / 2, M is the total number of stages obtained by dividing the target model, and the i-th stage represents the i-th stage executed in sequence during the forward propagation of the target model.

7. The method according to claim 1, wherein The at least one stage includes a first stage, and the first stage includes the layer at the very front in the forward propagation order of the target model.

8. The method according to claim 7, wherein The configuration of the host corresponding to the first stage is higher than the configurations of the hosts corresponding to other stages.

9. The method according to claim 1, wherein, Each computing node further includes at least one network card, and the at least one host is configured to drive at least one network card to implement network communication among the multiple computing nodes.

10. The method according to claim 9, wherein, The network communication is based on the Remote Direct Memory Access mechanism.

11. A distributed training system configured to perform distributed training on a target model, where, The distributed training system includes: Multiple computing nodes, each computing node including at least one host and at least one device, wherein the target model is divided into multiple stages in a pipeline parallel mode, each stage includes at least one layer of the target model, and the multiple stages are respectively distributed in the multiple devices of the multiple computing nodes; wherein, in the current round of iterative training, the training of at least one stage of the target model is jointly executed by at least one corresponding device and at least one corresponding host, For each stage of the target model, the device is configured to: Execute a forward calculation task to obtain a forward calculation result; Execute a backward calculation task based on the forward calculation result to obtain gradient data; In response to the completion of the backward calculation task, trigger the execution of an optimization update task and obtain updated parameters for use in the next round of iterative training; For at least one stage of the target model, the device is further configured to: in response to the completion of the reverse calculation task, transfer at least part of the gradient data to the host and trigger the host to execute an optimization update task; the host is configured to: execute an optimization update task based on at least part of the gradient data, obtain at least part of the updated parameters, and transfer the at least part of the updated parameters to the device for use in the next round of iterative training.

12. An electronic device, comprising: at least one processor; at least one memory, including one or more computer program modules; wherein, the one or more computer program modules are stored in the at least one memory and configured to be executed by the at least one processor, and the one or more computer program modules are used to implement the distributed training method according to any one of claims 1-10.

13. A non-transitory computer-readable storage medium having computer instructions stored thereon, wherein, When the computer instructions are executed by at least one processor, the distributed training method according to any one of claims 1-10 is executed.

Citation Information

Patent Citations

  • Dynamic minibatch sizes

    CN112655005A

  • Distributed training evaluation method and system for deep learning

    CN117093871A

  • Heterogeneous environment-oriented large model hybrid parallel training method and system

    CN117633527A

  • Data handling method, distributed training system, electronic equipment and storage medium

    CN118093203A

  • Model training method and device, electronic equipment and storage medium

    CN118863000A

Cited By

  • Model distributed training optimization method, electronic equipment and storage medium

    CN120803677A

  • Parameter optimization method and device, electronic equipment, storage medium and computer program product

    CN120875051A

  • Model updating method and device, electronic equipment, medium and product

    CN121480594A

  • Model distributed training optimization method and device, equipment and storage medium

    CN121683918A

  • Data communication method and device under distributed training task, equipment, storage medium and program product

    CN121792535A