Optimization method and device for model distributed training, equipment and storage medium

CN121683918BActive Publication Date: 2026-09-08BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511738417.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-09-08
Estimated Expiration
2045-11-24

AI Technical Summary

Technical Problem

在训练大模型时,这种跨设备同步参数梯度的通信开销巨大,其耗时(例如All-Reduce延迟)往往占到总训练时间的10%甚至更高

Benefits of technology

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121683918B_ABST
    Figure CN121683918B_ABST
Patent Text Reader

Abstract

The present disclosure provides an optimization method, device and equipment for model distributed training and a storage medium, and relates to the technical field of computers, in particular to the technical field of model training. The specific implementation scheme is as follows: a batch of data blocks for pipeline parallel processing is obtained; a reverse calculation process of at least one first data block in the batch of data blocks is decoupled into an activation value gradient calculation stage and a parameter gradient calculation stage; the parameter gradient calculation stage is executed after the completion of the activation value gradient calculation stage; and the execution process of the parameter gradient calculation stage is time overlapped with a data parallel communication process. The technical scheme of the present disclosure can reduce the total time of model training and improve the utilization rate of hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more particularly to the field of model training technology. Background Technology

[0002] Artificial intelligence models, exemplified by Large Language Models (LLMs), are growing exponentially in size. Due to the limitations of the memory and computing power of a single computing device (such as a GPU or TPU), it is impossible to support the training of a complete, ultra-large model. Therefore, distributed training has become the standard technique in this field.

[0003] In distributed training strategies, data parallelism (DP) and pipeline parallelism (PP) are two of the most core and widely used strategies. Data parallelism (DP) refers to creating replicas of the model (or a portion of the model) on multiple computing devices, with each replica independently processing a different set of training data. After the backpropagation of each training step, all replicas need to synchronize their computed parameter gradients (P-grad). This synchronization process is typically accomplished through ensemble communication operations such as all-reduce or reduce-scatter to ensure that the parameter updates of all replicas remain consistent. When training large models, the communication overhead of synchronizing parameter gradients across devices is enormous, and its time consumption (e.g., all-reduce latency) often accounts for 10% or even higher of the total training time. Summary of the Invention

[0004] This disclosure provides an optimization method, apparatus, device, and storage medium for distributed training of models.

[0005] According to one aspect of this disclosure, an optimization method for distributed training of a model is provided, applied in a computing device including at least one processor, comprising: Acquire a batch of data blocks for pipelined parallel processing; The reverse computation process of at least one first data block in a batch of data blocks is decoupled into an activation value gradient computation stage and a parameter gradient computation stage. After the activation value gradient calculation phase is completed, the parameter gradient calculation phase is executed. The execution process of the parameter gradient calculation stage is time-overlapped with the data parallel communication process.

[0006] According to another aspect of this disclosure, an optimization apparatus for distributed training of a model is provided, applied in a computing device including at least one processor, comprising: The acquisition module is used to acquire a batch of data blocks for pipelined parallel processing. The decoupling module is used to decouple the reverse calculation process of at least one first data block in a batch of data blocks into an activation value gradient calculation stage and a parameter gradient calculation stage. The execution module is used to perform the parameter gradient calculation phase after the activation value gradient calculation phase is completed. The overlap module is used to overlap the execution process of the parameter gradient calculation stage with the data parallel communication process in time.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.

[0010] The technical solution disclosed herein can reduce the total training time of the model and improve the utilization rate of hardware resources.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a flowchart illustrating an optimization method for distributed training of a model according to an embodiment of the present disclosure. Figure 2 This is a flowchart illustrating an optimization method for distributed training of a model according to another embodiment of the present disclosure; Figure 3 This is a flowchart illustrating the pipeline parallel scheduling strategy provided according to embodiments of this disclosure; Figure 4This is a schematic diagram of the structure of an optimization device for distributed training of a model according to an embodiment of the present disclosure; Figure 5 This is a block diagram of an electronic device used to implement the optimization method for distributed training of models according to embodiments of the present disclosure. Detailed Implementation

[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0014] Pipeline parallelism (PP) divides the different layers of the model into multiple stages and distributes these stages across different computing devices. Training data is divided into micro-batches (also called data chunks or chunks) that flow sequentially through all stages like a pipeline. After a stage completes its forward computation, it passes its output activation values ​​to the next stage. During backpropagation, a stage receives activation gradients (A-grad) from the next stage and locally computes its parameter gradients (P-grad) as well as the activation gradients to pass to the previous stage.

[0015] In actual large-scale model training, a single strategy is often insufficient to meet the requirements. Therefore, the industry usually adopts a hybrid parallel strategy that combines data parallelism and pipeline parallelism.

[0016] In the hybrid parallel backpropagation process, after a pipeline stage has completed the backpropagation of a microbatch, it typically needs to perform two key operations: The calculated activation gradient is passed to the previous stage of the pipeline in parallel.

[0017] The calculated parameter gradients are globally reduced within the data parallel group (DP Group).

[0018] In conventional implementations, backward computation (including A-grad and P-grad computations) is a complete computational step, while data-parallel All-Reduce communication only begins after backward computation is complete. Therefore, the high latency introduced by All-Reduce communication is typically exposed as a separate overhead in the training process.

[0019] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, the embodiments of this disclosure provide an optimization method for distributed training of models. By utilizing the technical solutions of the embodiments of this disclosure, communication overhead can be further hidden and resource competition between different communication types can be avoided.

[0020] Figure 1 This is a flowchart illustrating an optimization method for distributed training of a model according to an embodiment of the present disclosure. The method is applied to a computing device including at least one processor. The computing device may be one or more servers equipped with accelerators such as graphics processing units (GPUs) or tensor processing units (TPUs).

[0021] S110: Obtain a batch of data blocks for pipelined parallel processing.

[0022] In this embodiment of the disclosure, distributed training of models can be understood as a technique that distributes the training task of a large artificial intelligence model (such as a large language model) to multiple computing devices for execution.

[0023] Pipeline parallelism is a strategy for distributed training that places different layers of the model on different computing devices, with data flowing through these devices sequentially like a pipeline. A batch of chunks refers to the process in pipeline parallelism where a large training batch is divided into multiple smaller data chunks to improve device utilization. These chunks are then fed into the pipeline one by one for processing.

[0024] For example, a batch of data with a global size of 1024 is divided into 32 data blocks, each with a size of 32. The training system sequentially acquires these 32 data blocks and feeds them into the pipeline's parallel channel.

[0025] S120: Decouple the reverse calculation process of at least one first data block in a batch of data blocks into an activation value gradient calculation stage and a parameter gradient calculation stage.

[0026] In this step, the Backward Pass is the process used during model training to calculate the gradients of the loss function with respect to the model parameters and activation values. In a standard (undecoupled) Backward Pass, the activation gradients (A-grad) (i.e., the partial derivatives of the loss with respect to the activations) and parameter gradients (P-grad) (i.e., the partial derivatives of the loss with respect to the model weights) are usually calculated together in the same computational operation. Decoupling refers to breaking the computational dependency between these two by modifying the computation graph or computation kernel, so that they can be scheduled and executed independently at different times. The first data block refers to the data block selected to execute this decoupling strategy. The first data block can be at least a portion of a batch of data blocks, such as 1 / 10, 1 / 5, 1 / 4, 1 / 3, 1 / 2, 2 / 3, or all of them.

[0027] For example, for a linear layer, its standard backward computation outputs both A-grad and P-grad. After decoupling, it is split into two independent computation kernels: the first kernel only computes A-grad; the second kernel (which can be called later) uses A-grad and the activation value from the forward computation to compute P-grad.

[0028] S130: After the activation value gradient calculation phase is completed, the parameter gradient calculation phase is executed.

[0029] This step defines the execution timing of the two stages in S120. The system waits for the (decoupled) activation value gradient calculation stage (A-grad stage) of the first data block to complete on the computing device (e.g., GPU). It should be noted that "completion" here refers to the completion of the computation action itself, not the completion of its subsequent communication actions (such as gradient transfer).

[0030] In a preferred embodiment, the system waits until the activation value gradient calculation phase of all decoupled first data blocks has been completed before triggering the execution of the parameter gradient calculation phase.

[0031] For example, if the system detects that the A-grad computation kernel for the first data block has been completed, it sets an "A-grad complete" flag. A scheduler monitors this flag and only puts the corresponding P-grad computation kernel into the execution queue after the flag is set.

[0032] S140: Overlap the execution process of the parameter gradient calculation stage with the data parallel communication process in time.

[0033] In this embodiment of the disclosure, data parallel communication refers to network communication between multiple devices holding copies of the same model in order to synchronize parameter gradients (P-grad) in a data parallelism strategy. Temporal overlap refers to the scheduling system executing computational tasks (parameter gradient calculation) and network communication tasks (data parallel communication) concurrently in time, so that while the computing devices are performing computations, the network devices (such as network interface cards) are also processing communication data.

[0034] For example, when S130 triggers the start of the P-grad computation phase, the system simultaneously (or immediately) initiates parallel data communication (e.g., an All-Reduce operation). Since P-grad computation (such as matrix multiplication) mainly occupies the GPU's computing units, while All-Reduce mainly occupies network bandwidth and communication buses, the two are executed on different hardware resources, thus achieving time overlap.

[0035] According to the scheme of this disclosure, by decoupling the reverse computation process into an activation value gradient calculation stage and a parameter gradient calculation stage, and delaying the execution of the parameter gradient calculation stage, the purely local computation task of parameter gradient can overlap with the network task of data parallel communication in time. This overlap hides the high network latency overhead caused by data parallel communication, solves the problem of low training efficiency due to high communication overhead, significantly reduces the total model training time, and improves the utilization rate of hardware resources.

[0036] In one possible implementation, after the activation value gradient calculation phase is completed, step S130, which performs the parameter gradient calculation phase, may include: S131: Perform the activation value gradient calculation phase to generate activation value gradient data.

[0037] In this embodiment, after decoupling the reverse computation process of the first data block into an activation value gradient computation stage and a parameter gradient computation stage, the activation value gradient computation stage is executed first to generate activation value gradient data. The activation value gradient data is a tensor representing the gradient of the loss function with respect to the current layer's output activation value, which will be used for computation in the pipelined parallel previous computation stage (i.e., the layer above the model).

[0038] S132: Send the activation value gradient data to the previous computation stage in the pipeline.

[0039] This step is a pipelined parallel communication process. Unlike the data-parallel communication of S140 (broadcasting within the data-parallel group), this is a point-to-point communication, for example, sending the activation gradient data generated in S131 from the current computing device (e.g., holding layer N) to the previous computing device in the pipeline (e.g., holding layer N-1) via a Send / Recv operation.

[0040] S133: Perform the parameter gradient calculation phase to generate parameter gradient data.

[0041] After the activation gradient calculation phase is completed, the parameter gradient calculation phase is executed. The parameter gradient data (P-graddata) is the output of this phase; it is a tensor representing the gradient of the loss function with respect to the current layer's model parameters (such as weights and biases). This data will be used later in the data-parallel communication process.

[0042] According to the scheme of this disclosure, by clearly defining the generation and transmission of activation value gradient data and the generation of parameter gradient data, the flow of different data is clearly distinguished. More importantly, this scheme delays the execution of parameter gradient calculation, so that pipeline parallel communication and data parallel communication are staggered in time. This solves the hardware resource (network bandwidth) competition problem between pipeline activation value communication and data parallel communication, avoids computational pauses caused by pipeline being forced to wait due to network congestion, and significantly improves the effective working time of computing devices and overall training performance.

[0043] In one possible implementation, the step of performing the parameter gradient calculation phase in S133 to generate parameter gradient data further includes the following steps: S133_1: Retrieve the forward activation value data cached during the forward computation of the first data block.

[0044] In this embodiment, forward activation data refers to the input activation values ​​of the current layer during the forward pass of the model. In standard (undecoupled) backward computation, the forward activation data required to calculate the parameter gradients is typically used and released immediately after the forward pass. However, in this embodiment, since the calculation of the parameter gradients is delayed, this forward activation data needs to be cached in GPU memory during the forward pass phase so that it remains available during subsequent parameter gradient calculation phases. The retrieval operation involves reading this cached data from GPU memory.

[0045] S133_2: Perform parameter gradient calculation using cached forward activation value data to generate parameter gradient data.

[0046] This step is the standard procedure for calculating the parameter gradient. For example, for a linear layer, its parameter gradient (the gradient of the weights) is obtained by calculating the outer product of the cached forward activation data and the activation gradient data. The execution of this step depends on the cached data being retrieved.

[0047] According to the embodiments of this disclosure, by explicitly caching forward activation value data and using this cached data to perform parameter gradient calculation, the specific technical means for achieving decoupling and delayed computation are revealed in detail. It is pointed out that the cost of delaying the execution of parameter gradient calculation is the need for additional caching of forward activation value data, thus enabling decoupling.

[0048] In one possible implementation, S140, which involves time-overlapping the execution of the parameter gradient calculation phase with the data parallel communication phase, further includes the following steps: S141: After generating parameter gradient data during the parameter gradient calculation stage, the parameter gradient data is used to perform a global reduction operation or a reduction-distribution operation within the data parallel group.

[0049] In this step, a data parallel group refers to a collection of computing devices that hold copies of the same model (or the same slice of the model) during distributed training. All-Reduce is a collective communication operation where all devices in the group provide an input tensor (parameter gradient data), and this operation performs some operation (such as summation or averaging) on ​​all input tensors and returns the final result to all devices in the group.

[0050] Reduce-Scatter is another group communication operation that also reduces the input tensor, but the result is split and sent to different devices within the group. It is often used in scenarios where model parameters are sharded. The execution of this step (communication) overlaps with the execution of the parameter gradient calculation stage (computation) in time.

[0051] According to the scheme of the present disclosure, by overlapping the parameter gradient calculation with the global reduction or reduction-distribution operation, the network interface is also processing communication while the computing unit (such as GPU) is performing calculations, thereby directly hiding this part of the hardware communication time, reducing the training time ratio caused by the reduction operation, and improving the computational throughput of training.

[0052] In one possible implementation, at least one first data block is a data block that accounts for a predetermined proportion, counted from the end of the processing flow in a batch of data blocks.

[0053] In this embodiment of the disclosure, S120 does not perform decoupling operations on all data blocks in a batch of acquired data blocks. Instead, it applies the decoupling strategy only to a preset proportion of data blocks in the batch, counting backwards from the end of the processing (i.e., the last data block processed in the pipeline). This preset proportion can be configured according to hardware conditions (such as video memory capacity) and performance requirements. For example, the preset proportion can be one-third, one-half, or 1 (i.e., all data blocks) when video memory is sufficient.

[0054] According to the scheme of the present disclosure, by performing a high-memory-overhead decoupling operation only on data blocks that account for a preset proportion counting from the end of the processing, and by adopting different strategies for other data blocks (e.g., the remaining first part of the batch), a balance can be achieved between communication overlap benefits and hardware resource (memory) overhead.

[0055] In one possible implementation, such as Figure 2 As shown, the method also includes the following steps: During reverse computation, it is determined whether the current data block is the second data block in a batch of data blocks. If not (i.e., the first data block), steps S120~S140 are executed; if it is the second data block, the following steps are executed: S210: Perform an undecoupled reverse computation process on the second data block in a batch of data blocks.

[0056] The second data block is the data block that precedes the first data block.

[0057] In this embodiment, the second data block can be understood as a data block other than the first data block. The undecoupled reverse computation process refers to standard reverse computation, where the activation gradient (A-grad) and parameter gradient (P-grad) are calculated simultaneously in the same computation kernel without splitting. This is a traditional computation method with low memory usage.

[0058] This step clarifies the execution order. The system employs a hybrid strategy: for a batch of acquired data blocks, the second data blocks (e.g., the first 2 / 3 of the data blocks) are processed sequentially, performing undecoupled reverse computation on them; then, the first data blocks (e.g., the last 1 / 3 of the data blocks) are processed, performing decoupled reverse computation on them.

[0059] According to the scheme of this disclosure embodiment, by adopting this hybrid strategy, the system performs decoupling operations only on the first data block (to obtain communication overlap benefits), while performing standard operations on the preceding data blocks (to maintain low memory usage). This lays the foundation for subsequent memory optimization.

[0060] In one possible implementation, the method further includes the following steps: S160: After the undecoupled reverse computation process is performed on the second data block, the memory associated with the activation value it occupies is released.

[0061] In this embodiment, the activation value-related video memory mainly refers to the forward activation value data cached for the second data block during the forward computation. Since the second data block performs undecoupled backward computation, its parameter gradient is calculated immediately, therefore its corresponding forward activation value data is no longer needed after the backward computation is completed. This step describes the operation of releasing this portion of video memory back to the system video memory pool.

[0062] S170: Utilize activation-related memory to store intermediate computation data for the activation-gradient calculation stage and parameter-gradient calculation stage of at least one first data block.

[0063] In this step, the intermediate computation data can be the additional GPU memory usage caused by the decoupling strategy described in the previous embodiments, specifically including the forward activation value data and activation value gradient data that need to be cached (which must also be retained before the parameter gradient calculation is completed).

[0064] In one specific implementation, when the system starts processing the first data block, the additional video memory (intermediate computation data) it needs no longer needs to request new video memory space from the system, but directly reuses the video memory related to the activation value that was just released by the second data block.

[0065] According to the scheme of this disclosure embodiment, this memory reuse mechanism cleverly utilizes the time difference in the computation process. When the first data block begins to perform decoupling operations, causing an increase in memory demand, the second data block has just completed its standard operations, releasing memory. This released memory is used to meet the additional demands of the decoupling operations, ensuring that the peak memory usage throughout the training process remains unchanged compared to the peak memory usage of using only the standard reverse operation. This scheme ultimately achieves both communication overlap and avoidance of communication contention without increasing peak memory usage, significantly improving training efficiency and resource utilization.

[0066] In one possible implementation, the preset ratio of the first data block in the foregoing embodiments is further limited to: The number of the first data block is one-third of the number of data blocks in a batch.

[0067] In this embodiment of the disclosure, the number of first data blocks is determined to be 1 / 3 of the total number of data blocks. Correspondingly, the number of second data blocks is the first 2 / 3 of the total number of data blocks. Preferably, the first data block is a data block that is counted backward from the end of the batch of data blocks, accounting for 1 / 3 of the total number of data blocks in that batch, that is, the last 1 / 3 of the data blocks in the batch.

[0068] This ratio (1 / 3) is a preferred balance. The amount of intermediate computation data memory required by the decoupling strategy is approximately twice the memory required for forward activation values ​​(one part is used to cache forward activation values, and the other part is used to cache activation value gradients). When the first 2 / 3 of the data block performs standard reverse computation and releases memory, it can release exactly twice the memory required for forward activation values ​​(assuming that the activation value memory usage of each data block is similar). Therefore, when the first 2 / 3 of the second data block releases memory, it provides exactly the additional memory space required for the last 1 / 3 of the first data block.

[0069] According to the scheme of this disclosure embodiment, by adopting a specific ratio of one-third, the memory balance can be optimally achieved, ensuring that the peak memory usage remains strictly unchanged, thereby maximizing the training acceleration benefits while maintaining memory stability.

[0070] It should be noted that one-third is a preferred example for achieving a constant peak video memory usage. In other possible implementations, especially in scenarios where the computing device has ample video memory and a constant peak video memory usage is not strictly required, this ratio can be adjusted flexibly. For example, the ratio of the first data block could be one-half, two-thirds, or more.

[0071] In one scenario, the first data block can represent a 1:1 ratio of the batch of data blocks, meaning decoupling is performed on all data blocks within the batch. Choosing a larger ratio (such as 1 / 2, 2 / 3, or all) allows more parameter gradient computation to overlap with data-parallel communication, potentially resulting in a greater training speedup, at the cost of potentially higher peak GPU memory usage. This approach also covers other implementations that allow for trade-offs between GPU memory and speed to accommodate different hardware environments and performance requirements.

[0072] This embodiment provides a method such as Figure 3 The pipelined parallel scheduling strategy shown is used to illustrate the optimization method of this disclosure. In this embodiment, a global batch is divided into 27 data chunks, which are further divided into three groups (e.g., chunk0, chunk1, and chunk2), each containing 9 data chunks (numbered 1 to 9 in the figure). These 27 data chunks together constitute a "batch of data chunks". Pipeline parallelism includes 4 computation stages, as shown in the 4 rows (representing 4 computing devices) in the figure. The execution process of this method includes: S310: Obtain a batch of data blocks for pipelined parallel processing.

[0073] This involves acquiring these 27 data blocks. As shown by the white squares on the right side of the diagram, all 27 data blocks in chunk0, chunk1, and chunk2 are sequentially computed forward in four pipeline stages.

[0074] S320: Perform an undecoupled reverse computation process on the second data block in a batch of data blocks.

[0075] As shown by the light gray squares on the right side of the diagram, for the first two-thirds of the batch, i.e., 18 data blocks (corresponding to chunk1 and chunk2 here, as the second data block), the system performs a standard, undecoupled reverse computation process. During this process, the activation gradient and parameter gradient are calculated simultaneously.

[0076] S330: Decouple the reverse calculation process of at least one first data block in a batch of data blocks into an activation value gradient calculation stage and a parameter gradient calculation stage.

[0077] For the last third of the batch, as shown by the square corresponding to chunk0 on the right side of the figure, these 9 data blocks (as the first data block) undergo decoupling by the system. This step first performs the decoupled activation value gradient calculation stage (A-grad).

[0078] S340: After the activation value gradient calculation phase is completed, the parameter gradient calculation phase is executed.

[0079] The system waits until the activation gradient calculation phase of all 9 data blocks in the first data block (chunk0) is completed before uniformly and delayedly executing the parameter gradient calculation phase (P-grad) of these 9 data blocks.

[0080] S350: Overlap the execution process of the parameter gradient calculation stage with the data parallel communication process in time.

[0081] The parameter gradient calculation for the first data block (chunk0) is deferred and bundled for execution, allowing it to overlap with subsequent operations (such as forward computation for the next batch or data-parallel communication processes) in time.

[0082] According to the scheme of this disclosure embodiment, a batch of data blocks is divided into two parts by the scheduling strategy shown in the figure: the first 2 / 3 performs standard reverse computation; the last 1 / 3 performs decoupled reverse computation.

[0083] First and foremost, it solves the problems of "resource contention bubbles" and "pipeline-forced waiting." In conventional hybrid parallel implementations, the system attempts to overlap communication while computing backwards. This results in two different types of communication—"activation gradient propagation" (pipeline parallel communication, corresponding to the black squares in the diagram) and "parameter gradient reduction" (data parallel communication, such as All-Reduce)—occurring simultaneously. These two types of communication compete for hardware resources (such as network bandwidth), slowing down activation gradient propagation. Since the next stage in the pipeline must wait for this slowed-down activation gradient data, "pipeline-forced waiting" occurs. This waiting time forms "computation bubbles" that propagate through the pipeline, ultimately reducing overall performance.

[0084] This scheme completely separates the two conflicting communications in time by delaying the calculation of the parameter gradient of the first data block (chunk0) (S130). This strategy brings several key technical benefits: Phase 1: Activation value gradient calculation and propagation (pipeline communication). At this stage, there is no parameter gradient reduction (data parallel communication), therefore no resource contention.

[0085] The second stage: After the activation value gradient calculation is completed, the parameter gradient is calculated and overlapped with the parameter gradient reduction (data-parallel communication). At this stage, there is no activation value gradient propagation, and therefore no resource contention. In this way, this scheme fundamentally avoids communication resource contention, eliminates the bubble of forced pipeline waiting, and improves training performance.

[0086] Secondly, this scheme effectively hides the high time overhead of data parallel communication itself by bundling parameter gradient calculation with data parallel communication.

[0087] Finally, as Figure 3 The relationship between memory usage and time in the lower half is shown: after the second data block (chunk1 and chunk2) completes the standard reverse computation, the memory associated with its activation value is released; when the system starts processing the first data block (chunk0), the additional intermediate computation data memory required for its decoupling operation can be utilized (reused) from the previously released memory space. Ultimately, this achieves the acceleration benefits of communication overlap without increasing peak memory usage.

[0088] Figure 4 This is a schematic diagram of the structure of an optimization device for distributed training of a model according to an embodiment of this disclosure. The device is applied in a computing device including at least one processor, such as... Figure 4 As shown, the device 400 includes: The acquisition module 401 is used to acquire a batch of data blocks for pipelined parallel processing. The decoupling module 402 is used to decouple the reverse calculation process of at least one first data block in a batch of data blocks into an activation value gradient calculation stage and a parameter gradient calculation stage. Execution module 403 is used to execute the parameter gradient calculation stage after the activation value gradient calculation stage is completed; The overlap module 404 is used to overlap the execution process of the parameter gradient calculation stage with the data parallel communication process in time.

[0089] In one possible implementation, execution module 403 is used for: Perform the activation value gradient calculation phase to generate activation value gradient data; Send the activation value gradient data to the previous computation stage in the pipeline; Perform the parameter gradient calculation phase to generate parameter gradient data.

[0090] In one possible implementation, execution module 403 is used for: Retrieve the forward activation value data cached during the forward computation of the first data block; Parametric gradient calculations are performed using cached forward activation value data to generate parametric gradient data.

[0091] In one possible implementation, the overlapping module 404 is used for: After generating parameter gradient data during the parameter gradient calculation stage, the parameter gradient data is used to perform a global reduction operation or a reduction-distribution operation within the data parallel group.

[0092] In one possible implementation, at least one first data block is a data block that accounts for a predetermined proportion, counted from the end of the processing flow in a batch of data blocks.

[0093] In one possible implementation, the device further includes a computing module for: Perform an undecoupled reverse computation process on the second data block in a batch of data blocks; wherein the second data block is the data block that precedes the first data block.

[0094] In one possible implementation, the device further includes a video memory management module for: After the undecoupled reverse computation process is performed on the second data block, the memory associated with the activation value it occupies is released; By utilizing activation-related memory, intermediate computation data is stored for the activation gradient calculation stage and parameter gradient calculation stage of at least one first data block.

[0095] In one possible implementation, the number of the first data block is one-third of the number of data blocks in a batch.

[0096] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0097] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0098] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0099] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0100] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0101] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0102] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as optimization methods for distributed model training. For example, in some embodiments, the optimization methods for distributed model training can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the optimization methods for distributed model training described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform optimization methods for distributed model training by any other suitable means (e.g., by means of firmware).

[0103] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0104] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0105] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0107] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0108] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0109] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0110] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An optimization method for distributed model training, applied in a computing device including at least one processor, comprising: Acquire a batch of data blocks for pipelined parallel processing; The reverse calculation process of at least one first data block in the batch of data blocks is decoupled into an activation value gradient calculation stage and a parameter gradient calculation stage; wherein, the at least one first data block is a data block in the batch of data blocks that accounts for a preset proportion and is counted from the end of the processing. An undecoupled reverse computation process is performed on the second data block in the batch of data blocks; wherein the second data block is a data block located before the first data block; After the undecoupled reverse calculation process is executed in the second data block, the activation value-related video memory it occupies is released; The activation value-related memory is used to store intermediate calculation data for the activation value gradient calculation stage and parameter gradient calculation stage of the at least one first data block. After the activation value gradient calculation phase is completed, the parameter gradient calculation phase is executed. The execution process of the parameter gradient calculation stage is time-overlapped with the data parallel communication process.

2. The method according to claim 1, wherein, After the activation value gradient calculation phase is completed, the parameter gradient calculation phase is executed, including: The activation value gradient calculation phase is performed to generate activation value gradient data; The activation value gradient data is sent to the previous computation stage in the pipeline; The parameter gradient calculation phase is performed to generate parameter gradient data.

3. The method according to claim 2, wherein performing the parameter gradient calculation stage to generate parameter gradient data includes: Obtain the forward activation value data cached during the forward computation of the first data block; The parameter gradient is calculated using the cached forward activation value data to generate parameter gradient data.

4. The method according to claim 1 or 2, wherein, The process of overlapping the execution of the parameter gradient calculation stage with the data parallel communication process in time includes: After generating parameter gradient data in the parameter gradient calculation stage, the parameter gradient data is used to perform a global reduction operation or a reduction-distribution operation within the data parallel group.

5. The method according to claim 1, wherein, The number of the first data block is one-third of the number of data blocks in the batch.

6. An optimization apparatus for distributed training of a model, applied in a computing device including at least one processor, comprising: The acquisition module is used to acquire a batch of data blocks for pipelined parallel processing. The decoupling module is used to decouple the reverse calculation process of at least one first data block in the batch of data blocks into an activation value gradient calculation stage and a parameter gradient calculation stage; wherein, the at least one first data block is a data block in the batch of data blocks that accounts for a preset proportion and is counted from the end of the processing. A calculation module is used to perform an undecoupled reverse calculation process on a second data block in the batch of data blocks; wherein the second data block is a data block located before the first data block; The video memory management module is used for: After the undecoupled reverse computation process is performed on the second data block, the memory associated with the activation value it occupies is released; By utilizing activation-related memory, intermediate computation data is stored for the activation-value gradient calculation stage and parameter gradient calculation stage of at least one first data block; An execution module is used to execute the parameter gradient calculation phase after the activation value gradient calculation phase is completed; The overlap module is used to overlap the execution process of the parameter gradient calculation stage with the data parallel communication process in time.

7. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.

8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.

9. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5.