Model training method and device, electronic device, and computer-readable storage medium

CN119180317BActive Publication Date: 2026-09-25LYNXI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411230406.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2026-09-25
Estimated Expiration
2044-09-03

AI Technical Summary

Technical Problem

但是,在基于反向传播的实现方式中,需要保持大量的模型中间状态,因而需要耗费大量的存储资源

Benefits of technology

[0013]在本公开所提供的实施例中,模型中的多个模型层级按照流水并行方式分别运行在多个不同设备节点中,由于多个设备节点之间存在数据同步需求,因此,势必会在每个设备节点中产生空闲时间分片。为了提升运算效率和准确性,在多个设备节点根据前向自动微分方式确定多个批次数据对应的第一梯度结果的过程中,进一步在多个设备节点的空闲时间分片中根据反向自动微分方式确定多个批次数据中的指定批次数据对应的第二梯度结果,从而根据第二梯度结果对第一梯度结果进行修正,以得到更加准确的目标梯度结果,进而提升模型训练的准确性。由此可见,该方式将前向自动微分运算与反向自动微分运算相结合,主要根据前向自动微分运算计算梯度结果,从而无需大量缓存中间结果,降低了对存储资源的占用,提升了计算效率。并且,通过在前向自动微分运算的空闲时间分片中插入反向自动微分运算,从而能够充分利用空闲时间分片,避免时间分片的浪费,提升计算结果的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119180317B_ABST
    Figure CN119180317B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model training method and device, an electronic device and a computer readable storage medium, wherein the model comprises: a plurality of model layers running in a plurality of device nodes in a pipeline parallel manner; the method comprises: in the process of determining the first gradient result corresponding to the plurality of batches of data in the plurality of device nodes according to the forward automatic differentiation manner, determining the second gradient result corresponding to the specified batch of data in the plurality of batches of data in the idle time slice of the plurality of device nodes according to the backward automatic differentiation manner; correcting the first gradient result according to the second gradient result to obtain the target gradient result, and updating the model parameter of the model according to the target gradient result. This way inserts the backward automatic differentiation operation in the idle time slice of the forward automatic differentiation operation, so as to make full use of the idle time slice, avoid the waste of the time slice, and improve the accuracy of the calculation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a model training method and apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] Model training is a crucial step in machine learning and deep learning, involving the process of optimizing model parameters using training data. This process typically includes preparing data, selecting or designing a model architecture, defining a loss function and optimization algorithm, and then iteratively tuning the model parameters to minimize the loss function.

[0003] In related technologies, model training can be implemented using the gradient descent algorithm, with the gradients of the model parameters solved using backpropagation (BP). BP is essentially a backpropagation mode with automatic differentiation. However, the backpropagation-based implementation requires maintaining a large number of intermediate model states, thus consuming significant storage resources. For example, large models (LMs), especially large language models such as GPT-3 (Generative Pre-trained Transformer-3) and GLM (General Language Model), have a massive number of parameters exceeding hundreds of billions. Therefore, training large models via backpropagation inevitably consumes enormous amounts of storage resources. Therefore, reducing the storage resources required during model training has become a pressing technical challenge. Summary of the Invention

[0004] This disclosure provides a model training method and apparatus, an electronic device, and a computer-readable storage medium.

[0005] In a first aspect, this disclosure provides a model training method, wherein the model comprises: multiple model layers running in a pipelined parallel manner across multiple device nodes; the method comprises:

[0006] During the process of determining the first gradient result corresponding to multiple batches of data by the multiple device nodes according to the forward automatic differentiation method, the second gradient result corresponding to a specified batch of data in the multiple batches of data is determined by the backward automatic differentiation method in the idle time slice of the multiple device nodes.

[0007] Based on the second gradient result, the first gradient result is corrected to obtain the target gradient result, and the model parameters of the model are updated based on the target gradient result.

[0008] Secondly, this disclosure provides a model training apparatus, wherein the model comprises: multiple model layers running in a pipelined parallel manner across multiple device nodes; the apparatus comprises:

[0009] The gradient determination module is adapted to determine the second gradient result corresponding to a specified batch of data in the multiple batches of data in the idle time slices of the multiple device nodes in the process of determining the first gradient result corresponding to multiple batches of data ...

[0010] The update module is adapted to correct the first gradient result based on the second gradient result to obtain a target gradient result, and update the model parameters of the model based on the target gradient result.

[0011] Thirdly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the model training method described above.

[0012] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-described model training method when executed by a processor / processor core.

[0013] In the embodiments provided in this disclosure, multiple model layers in the model run in a pipelined parallel manner on multiple different device nodes. Since there are data synchronization requirements between the multiple device nodes, idle time slices are inevitably generated on each device node. To improve computational efficiency and accuracy, during the process of determining the first gradient results corresponding to multiple batches of data on multiple device nodes using forward automatic differentiation, the second gradient results corresponding to a specified batch of data are further determined in the idle time slices of multiple device nodes using backward automatic differentiation. The first gradient results are then corrected based on the second gradient results to obtain a more accurate target gradient result, thereby improving the accuracy of model training. Thus, this method combines forward automatic differentiation with backward automatic differentiation, primarily calculating gradient results based on forward automatic differentiation, thereby eliminating the need for extensive caching of intermediate results, reducing storage resource consumption, and improving computational efficiency. Furthermore, by inserting backward automatic differentiation into the idle time slices of forward automatic differentiation, the idle time slices can be fully utilized, avoiding wasted time slices and improving the accuracy of the calculation results.

[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:

[0016] Figure 1 A schematic flowchart of a model training method provided in an embodiment of this disclosure is shown;

[0017] Figure 2 A schematic diagram of a pipelined parallel scheme is shown;

[0018] Figure 3 A schematic diagram is shown in one example of this disclosure, illustrating the determination of gradient results using a forward automatic differentiation method with multiple device nodes;

[0019] Figure 4 A schematic diagram illustrates one possible scheme for performing reverse automatic differentiation and forward automatic differentiation in a time-division manner;

[0020] Figure 5 A schematic diagram is shown of another alternative scheme in which reverse automatic differentiation and forward automatic differentiation are performed in a time-division manner;

[0021] Figure 6 A block diagram of a model training apparatus provided in an embodiment of this disclosure;

[0022] Figure 7 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0023] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0024] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0025] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0027] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0028] In related technologies, model training can be implemented using the gradient descent algorithm, and the gradient of the model parameters can be solved based on backpropagation (BP). The BP algorithm is the backpropagation mode of automatic differentiation. However, the backpropagation-based implementation requires maintaining a large number of intermediate model states, thus consuming significant storage resources. To address this issue, this disclosure combines forward automatic differentiation with backward automatic differentiation. Gradient results are primarily calculated based on forward automatic differentiation, eliminating the need for extensive caching of intermediate results, reducing storage resource consumption, and improving computational efficiency. Furthermore, by inserting backward automatic differentiation into the idle time slices of forward automatic differentiation, idle time slices can be fully utilized, avoiding wasted time slices and improving the accuracy of the calculation results.

[0029] Figure 1 The diagram illustrates a flowchart of a model training method provided in an embodiment of this disclosure, wherein the model includes multiple model layers running in a pipelined parallel manner across multiple device nodes. Figure 1 As shown, the training method includes:

[0030] Step S110: During the process of determining the first gradient result corresponding to multiple batches of data in multiple device nodes according to the forward automatic differentiation method, the second gradient result corresponding to the specified batch of data in multiple batches of data is determined in the idle time slice of multiple device nodes according to the backward automatic differentiation method.

[0031] In this process, multiple device nodes train the corresponding model layers in a pipelined parallel manner. To avoid excessive storage overhead due to caching intermediate results, in this disclosure, multiple device nodes determine the first gradient results for multiple batches of data using forward automatic differentiation. The first gradient result refers to the gradient result of the model parameters determined by forward automatic differentiation. Since multiple device nodes need to synchronize data with each other, and following the order from input layer to output layer, device nodes at later model layers (i.e., closer to the output layer and farther from the input layer) depend on the computation results of device nodes at earlier model layers (i.e., closer to the input layer and farther from the output layer). For example, for any batch of data, device node 2 corresponding to the second model layer needs to depend on the computation results generated by device node 1 corresponding to the first model layer for the same batch of data. Therefore, each device node will have some idle time slices, the number of which depends on the number of device nodes. An idle time slice refers to the time interval during which the forward automatic differentiation operation for any batch of data is not being performed, and the node is waiting for computation results from other device nodes.

[0032] To improve efficiency and accuracy, this disclosure further determines the second gradient result corresponding to a specified batch of data from multiple batches of data using a reverse automatic differentiation method within the idle time slices of multiple device nodes. Considering the limited number of idle time slices, the specified batch of data is typically a subset of all batches. By performing reverse automatic differentiation operations within the idle time slices, idle time can be fully utilized, improving the accuracy of the final result. The second gradient result refers to the gradient result of the model parameters determined by the reverse automatic differentiation method.

[0033] Step S120: Based on the second gradient result, correct the first gradient result to obtain the target gradient result, and update the model parameters based on the target gradient result.

[0034] Since the second gradient result is obtained through inverse automatic differentiation, combining the second gradient result yields a more accurate target gradient result, thereby improving the accuracy of model training. This disclosure does not limit the specific implementation method of correcting the first gradient result based on the second gradient result. Those skilled in the art can flexibly choose various methods. For example, a preset operation can be performed on the first gradient result based on the second gradient result, or the first gradient result can be directly replaced with the second gradient result. In short, any method that achieves the goal of improving the accuracy of gradient calculation is acceptable.

[0035] In the embodiments provided in this disclosure, multiple model layers in the model run in a pipelined parallel manner on multiple different device nodes. Since there are data synchronization requirements between the multiple device nodes, idle time slices are inevitably generated on each device node. To improve computational efficiency and accuracy, during the process of determining the first gradient results corresponding to multiple batches of data on multiple device nodes using forward automatic differentiation, the second gradient results corresponding to a specified batch of data are further determined in the idle time slices of multiple device nodes using backward automatic differentiation. The first gradient results are then corrected based on the second gradient results to obtain a more accurate target gradient result, thereby improving the accuracy of model training. Thus, this method combines forward automatic differentiation with backward automatic differentiation, primarily calculating gradient results based on forward automatic differentiation, thereby eliminating the need for extensive caching of intermediate results, reducing storage resource consumption, and improving computational efficiency. Furthermore, by inserting backward automatic differentiation into the idle time slices of forward automatic differentiation, the idle time slices can be fully utilized, avoiding wasted time slices and improving the accuracy of the calculation results.

[0036] Furthermore, those skilled in the art can make various modifications and variations to the embodiments disclosed herein:

[0037] In one optional implementation, the model is trained through an iterative process across multiple training epochs. Correspondingly, when multiple device nodes determine the first gradient results corresponding to multiple batches of data using forward automatic differentiation, this is achieved as follows: multiple device nodes, based on the first gradient results of multiple batches of data in the previous training epoch, use forward automatic differentiation to determine the first gradient results corresponding to multiple batches of data in the current training epoch. Therefore, the forward automatic differentiation operation is executed periodically and cyclically.

[0038] In any training cycle, the start time of the first non-idle time slice of the m-th device node corresponding to the m-th model level is earlier than the start time of the first non-idle time slice of the (m+1)-th device node corresponding to the m-th model level; and the end time of the last non-idle time slice of the m-th device node corresponding to the m-th model level is earlier than the end time of the last non-idle time slice of the (m+1)-th device node corresponding to the (m+1)-th model level; where m is a natural number. As mentioned above, the device node corresponding to the later model level depends on the computation results of the device node corresponding to the earlier model level. Therefore, the times of the first non-idle time slice and the last idle time slice of the m-th device node corresponding to the m-th model level are both earlier than the times of the (m+1)-th device node corresponding to the (m+1)-th model level. For example, at the beginning of any training cycle, the start time of the first non-idle time slice of the first device node corresponding to the first model level is time 1, the start time of the first non-idle time slice of the second device node corresponding to the second model level is time 2, and the start time of the first non-idle time slice of the third device node corresponding to the third model level is time 3. Then time 1 is earlier than time 2, and time 2 is earlier than time 3.

[0039] In one optional implementation, multiple batches of data are arranged sequentially according to a specified batch order; and each device node processes multiple batches of data sequentially within multiple time slices according to the specified batch order. For example, assuming there are four batches of data, the specified batch order is: batch 1, batch 2, batch 3, batch 4, so that each batch of data is processed sequentially. In any training cycle, each device node has the same number of non-idle time slices, and the number of non-idle time slices is determined based on the number of batches of data. Typically, time slices correspond one-to-one with batches of data, that is, one time slice is used to process one batch of data. In this case, the number of non-idle time slices equals the number of batches of data. Furthermore, in any training cycle, each device node has the same number of idle time slices, and the number of idle time slices is determined based on the number of nodes in the device nodes. Typically, the number of idle time slices is less than the number of nodes in the device nodes. For example, the number of idle time slices can be one less than the number of nodes in the device nodes.

[0040] In one optional implementation, when determining the second gradient result corresponding to a specified batch of data from multiple batches of data using the inverse automatic differentiation method in the idle time slices of multiple device nodes, it can be achieved as follows: In any training cycle, for the i-th device node corresponding to the i-th model level, select the i-th set of specified batch data based on the second gradient result corresponding to the i+1 set of specified batch data determined by the inverse automatic differentiation method in the idle time slice of the i+1 device node corresponding to the i+1 model level; determine the second gradient result corresponding to the i-th set of specified batch data in the idle time slice corresponding to the i-th device node using the inverse automatic differentiation method; where i is a natural number.

[0041] Therefore, following the order from the input layer to the output layer (i.e., from front to back), the specified batch of data corresponding to the device node of the earlier model layer is determined based on the specified batch of data corresponding to the device node of the later model layer. In other words, the specified batch of data used to perform inverse automatic differentiation is determined sequentially from back to front of the model layers. That is: starting from the device node corresponding to the last model layer, the specified batch of data used to perform inverse automatic differentiation in that device node is determined; then, the specified batch of data used to perform inverse automatic differentiation in the device node corresponding to the second-to-last model layer is determined, and so on, until the specified batch of data used to perform inverse automatic differentiation in the device node corresponding to the first model layer is determined.

[0042] In one optional implementation, the i-th specified batch of data is a subset of the (i+1)-th specified batch of data. In other words, the specified batch of data used for performing inverse automatic differentiation in the device node corresponding to the preceding model level is typically less than the specified batch of data used for performing inverse automatic differentiation in the device node corresponding to the following model level. Alternatively, the specified batch of data used for performing inverse automatic differentiation in the device node corresponding to the preceding model level is a subset of the specified batch of data used for performing inverse automatic differentiation in the device node corresponding to the following model level. Furthermore, the sharding time of any specified batch of data in the i-th specified batch of data in the (i+1)-th device node is earlier than the sharding time of any specified batch of data in the i-th device node. Since in the inverse automatic differentiation method, for the same batch of data, the data in the preceding model level needs to use the computational results of the same batch of data in the following model level, there are certain limitations on the time for performing inverse automatic differentiation in different model levels for the same batch of data.

[0043] In one optional implementation, the number of nodes in the multiple device nodes is M, and the number of batches in the multiple batches of data is N; i is a natural number greater than or equal to 1 and less than M; then, when determining the second gradient result corresponding to a specified batch of data in the multiple batches of data according to the inverse automatic differentiation method in the idle time slice of the multiple device nodes, it can also be implemented in the following way:

[0044] In any training cycle, for the Mth device node corresponding to the Mth model level, the number of idle time slices for the Mth device node is determined. Based on the number of idle time slices, starting from the Nth batch of data, the Mth set of specified batch data matching the number of slices is selected from multiple batches of data in reverse order of the specified batch order. The second gradient result corresponding to the Mth set of specified batch data is determined in the idle time slices corresponding to the Mth device node using the reverse automatic differentiation method. Where one time slice corresponds to one batch of data, the number of idle time slices for the Mth device node can be M-1, and correspondingly, the number of batches included in the Mth set of specified batch data can be M-1, so as to fully utilize each idle time slice in the Mth device node. Furthermore, considering that the reverse automatic differentiation method needs to utilize intermediate results generated during the forward computation, in order to minimize the storage resource consumption caused by caching intermediate results, preferably, starting from the Nth batch of data, M-1 batches of data are selected as the Mth set of specified batch data in reverse order of the specified batch order. Since forward computation is performed in the order of specified batches, using an order that is the reverse of the specified batch order can ensure that the caching time of intermediate results required for the selected Mth specified batch of data is shorter, thus resulting in lower resource consumption.

[0045] In one optional implementation, the above method further includes the following process: During the process of multiple device nodes determining the first gradient results corresponding to multiple batches of data using a forward automatic differentiation method, for the Mth specified batch of data, intermediate result data of the Mth specified batch of data during the forward computation process is cached. Correspondingly, when determining the second gradient results corresponding to the Mth specified batch of data in the idle time slice corresponding to the Mth device node using a backward automatic differentiation method, the intermediate result data of the Mth specified batch of data during the forward computation process can be used to determine the second gradient results corresponding to the Mth specified batch of data in the idle time slice corresponding to the Mth device node using a backward automatic differentiation method. Therefore, during the process of multiple device nodes determining the first gradient results corresponding to multiple batches of data using a forward automatic differentiation method, forward computation (i.e., forward propagation, also called forward inference) is further performed for the multiple batches of data. The forward automatic differentiation process and the forward computation process can be performed simultaneously. Furthermore, since this application only performs inverse automatic differentiation on a portion of the batch data, there is no need to cache all intermediate results of the forward calculation process; only a small portion of the intermediate results associated with the inverse automatic differentiation need to be cached.

[0046] In one optional implementation, when correcting the first gradient result based on the second gradient result to obtain the target gradient result, this can be achieved as follows: For any device node, perform a preset operation on the second gradient result and the first gradient result corresponding to the same batch of data, and obtain the target gradient result corresponding to the same batch of data based on the operation result. The preset operation can be various operations such as averaging. Since the second gradient result is determined through inverse automatic differentiation, it has better accuracy. Using the second gradient result, the final target gradient result can be more accurate.

[0047] In one alternative implementation, the device node may include: a processing core in a many-core system, a chip device containing multiple processing cores, and / or a server device containing multiple chips. In summary, this application does not limit the specific implementation form of the device node.

[0048] To facilitate understanding, an example is given below to illustrate the implementation of the model training method provided in this embodiment.

[0049] First, the technical implementation scheme related to this example will be introduced.

[0050] In one related technique, the gradient descent algorithm is used, and the gradient of the model parameters is solved by backpropagation (BP) to train large models. The BP algorithm is the backpropagation mode of automatic differentiation. However, the above method has at least the following problems: (1) It is necessary to save the intermediate state of the model, which causes a lot of storage overhead. Especially when the batch size is very large, it causes a huge storage overhead. (2) The dynamic range of the gradient is large, which leads to the need for high-precision data type (such as fp32) to save the gradient. In addition, considering the first moment, second moment and other gradients required by the optimizer, the storage requirements are further increased. Therefore, the memory overhead of training large models is about 10 times more than that of the model parameters, which greatly restricts the growth of parameters of large models.

[0051] On the other hand, large models, due to their massive number of parameters, often require multi-level (i.e., multiple model layers) and multi-GPU (i.e., multiple computing devices) parallelism. One parallelism approach is pipelined parallelism, which involves dividing the entire model network into multiple segments and running them on different nodes in a phased, pipelined manner. For example, in the case of a large model with more than a specified threshold of parameters, it is necessary to pre-divide the large model network into multiple model layers (i.e., multiple segments) and distribute these multiple model layers to multiple device nodes, so that the multiple device nodes can run the multiple model layers in a pipelined parallel manner in stages. Each model layer corresponds to one training phase (which can be represented by corresponding time slices), and multiple model layers are run sequentially through multiple training phases. Figure 2 A schematic diagram of a pipelined parallel scheme is shown. (For example...) Figure 2 As shown, the large model is divided into four segments, also called four model layers, running on four devices: device 0, device 1, device 2, and device 3. Furthermore, the model's training data is divided into four batches: batch 0, batch 1, batch 2, and batch 3. Correspondingly, F(i,j) represents performing forward computation on device i for batch j; for example, F(0,0) represents performing forward computation on device 0 for batch 0. B(i,j) represents performing backward computation on device i for batch j; for example, B(0,0) represents performing backward computation on device 0 for batch 0.

[0052] from Figure 2As can be seen, due to data synchronization and dependencies among multiple devices, for device 0, time slices 1 to 4 are used for forward computation, and time slices 11 to 14 are used for backward computation. Correspondingly, time slices 5 to 10 are idle time slices (also called "bubbles"). For device 1, time slices 2 to 5 are used for forward computation, and time slices 10 to 13 are used for backward computation. Correspondingly, time slices 6 to 9, as well as time slices 1 and 14, are idle time slices… and so on. It can be seen that due to the synchronization characteristics of backward propagation, device nodes have a lot of idle time, forming many "bubbles." Therefore, reducing the bubble ratio is beneficial to improving the utilization of computing resources and accelerating model convergence.

[0053] In some related techniques, mixed-precision training can be used to store intermediate states and other data that are relatively insensitive to data precision using fp8 or even lower precision data types, thereby reducing memory usage. However, the convergence performance is far inferior to models represented by high-precision data, and the training stability is poor. In other related techniques, recomputation can be used, which trades time for space by not storing some intermediate results. During backpropagation, the unsaved intermediate states are calculated from other state data. This method can save storage space but significantly increases the computational load. Therefore, mixed ultra-low precision training generally has poor convergence performance, while recomputation leads to significant additional computational overhead.

[0054] To address the aforementioned issues, this example aims to adjust the gradient calculation method to avoid the excessive storage requirements caused by backpropagation (BP) gradient calculation, while also reducing the idle rate of computing resources in pipelining parallelism.

[0055] The core of this example lies in employing forward automatic differentiation (FAD) to solve for the gradients of large model parameters, thereby reducing the storage overhead caused by caching numerous intermediate results. Forward Automatic Differentiation (FAD) is a computer program algorithm for calculating the derivative of a function. By applying the chain rule and without using numerical differencing methods, it can calculate the derivative of a function with high accuracy. FAD is commonly used in scientific computing, machine learning, and other fields, and is particularly effective in calculating gradients, making it crucial for optimization problems.

[0056] Forward automatic differentiation can calculate the directional derivatives of model parameters with respect to a given initial direction, thus obtaining an approximate gradient. Furthermore, forward automatic differentiation eliminates the need for backpropagation, avoiding the storage of intermediate states, and allows control over the numerical range of the initial direction, thus avoiding the problem of high-precision storage of gradient and other information. This avoids the excessive storage requirements caused by backpropagation (BP) gradient calculation, significantly reducing storage needs. This example demonstrates how forward automatic differentiation can be adapted to the training of large models.

[0057] In each iteration, the forward automatic differentiation operation can be implemented in the following way:

[0058] First, determine the initial gradient that matches the preset numerical precision. That is, the initial gradient should be accurately represented by the preset numerical precision and should not overflow in gradient update and other operations. For example, if the gradient is stored with fp16 precision, the magnitude of the initial gradient can be on the order of 0.01.

[0059] Then, according to certain rules, the initial gradient is obtained. Specifically, the initial gradient can be obtained in the following ways: (1) Set a random initial gradient: for example, generate a binary perturbation that conforms to the Bernoulli distribution for each parameter. (2) Generate based on the input and output of the operation involved in the parameter, such as the parameter update amount in the form of hebb. (3) Obtain by combining the above two methods.

[0060] Next, the model performs an inference operation on a given dataset, which can be done simultaneously with forward automatic differentiation. Alternatively, model inference can be performed first, followed by forward automatic differentiation. Then, the directional derivative is derived from the forward automatic differentiation, truncated, and multiplied by the initial gradient to obtain the final approximate gradient. After obtaining the approximate gradient, the subsequent optimizer is adapted to complete the gradient update. The directional derivative should be truncated to an appropriate numerical range to prevent numerical overflow during the gradient update process.

[0061] Finally, perform the next iteration until the preset convergence condition is met.

[0062] Furthermore, various other implementation methods can be flexibly adopted to implement the model training process. For example, a hybrid forward and backward automatic differentiation training method can be introduced. One hybrid method involves obtaining the gradients of some model parameters using forward automatic differentiation and the gradients of the other part using backward automatic differentiation (i.e., BP). For instance, in one implementation, after model inference is completed, the gradients of the parameters of the second half of the network are obtained by BP from back to front, while the gradients of the parameters of the first half of the network are obtained by forward automatic differentiation from front to back. These two methods can be used in parallel, and the specific ratio can be set according to requirements and hardware resources.

[0063] In another implementation, the gradients of the parameters in the first half of the network can be calculated from back to front using backpropagation (BP), while the gradients of the parameters in the second half of the network can be calculated from front to back using forward automatic differentiation (AE). In this case, it is necessary to first determine the boundary between the two methods (i.e., the boundary between AE and BP methods for solving the gradients), and first use the initial gradient at the boundary to calculate the error signal of the network state (neuron activation) at the boundary before performing BP.

[0064] Alternatively, the two methods described above can be used alternately. By combining forward and backward automatic differentiation, a balance can be struck between storage requirements and convergence accuracy and speed. Furthermore, the ratio of forward to backward gradient solving can be dynamically adjusted. For example, when hardware resources are limited, the proportion of forward automatic differentiation can be increased, and vice versa.

[0065] Another hybrid approach is to utilize both forward automatic differentiation and backward automatic differentiation in the entire model, performing them in a time-sharing manner and calculating them separately, with the aim of increasing the utilization of computational resources. This example mainly adopts a time-sharing scheme for forward and backward automatic differentiation.

[0066] Figure 3 This diagram illustrates how multiple device nodes are used to determine the gradient results using a forward automatic differentiation method, as shown in this example. Figure 3As shown, forward automatic differentiation can be performed on four batches of data using a pipelined parallel approach across four devices. The large model consists of four model layers, running on four devices: device 0, device 1, device 2, and device 3. The training data is divided into four batches: batch a, batch b, batch c, and batch d. Correspondingly, F / J indicates that forward computation and forward automatic differentiation are performed separately, which can be performed simultaneously or sequentially. For example, F / J 0a indicates that forward computation and forward automatic differentiation are performed on device 0 for batch a (i.e., sub-batch a). Similarly, F / J 1a indicates that forward computation and forward automatic differentiation are performed on device 1 for sub-batch a. The same logic applies to other devices and other batches, and will not be elaborated further here. Therefore, F represents the forward computation of the model, J represents the forward automatic differentiation, and F / J indicates that the forward computation and forward automatic differentiation are performed simultaneously. When both are performed simultaneously, there is no need to cache intermediate states, thus significantly reducing storage requirements. Specifically, for device 0, time slices 1 to 4 are used to perform forward computation and forward automatic differentiation, while time slices 5 to 7 are used to wait for the results of other devices' calculations. Correspondingly, time slices 5 to 7 are the idle time slices (also called "bubbles") of device 0. For device 1, time slices 2 to 5 are used to perform forward computation and forward automatic differentiation, while time slices 6 to 7 and time slice 1 are used to wait for the results of other devices' calculations. Correspondingly, time slices 6 to 7 and time slice 1 are the idle time slices of device 1.

[0067] In addition, during time slice 8, all devices completed the first round of computation results, and therefore, all devices performed a unified update. This update refers to updating the model's weights and parameters based on the gradient data obtained in this round of computation.

[0068] Therefore, in this example, the entire batch is divided into four sub-batches (a, b, c, d) for time-sharing computation. Since only forward computation is needed and backward computation is not required, compared to... Figure 2 compared to, Figure 3 The number of bubbles caused by data synchronization decreased by 50%. Among them, Figure 3 The bubbles in the text represent the idle time fragments of the multiple device nodes mentioned above.

[0069] In addition, such as Figure 3As shown, time slices 1 to 8 constitute one complete training cycle. Similarly, time slices 9 to 16 constitute the next complete training cycle. Furthermore, within any training cycle, the start time of the first non-idle time slice (time slice 1 when m=1) of the m-th model level corresponding to the m-th device node is earlier than the start time of the first non-idle time slice (time slice 2) of the (m+1)-th model level corresponding to the (m+1)-th device node. Also, the end time of the last non-idle time slice (time slice 4) of the m-th model level corresponding to the m-th device node is earlier than the end time of the last non-idle time slice (time slice 5) of the (m+1)-th model level corresponding to the (m+1)-th device node.

[0070] In addition, to further increase the utilization of computing resources and reduce bubbles, Figure 3 Further inverse calculations (also called inverse automatic differentiation) are inserted into the bubble, such as... Figure 4 As shown. Figure 4 This illustrates an alternative approach where backward automatic differentiation and forward automatic differentiation are performed in a time-sharing manner.

[0071] like Figure 4 As shown, in any training cycle, with M=4 and N=4 as mentioned above, for the Mth device node corresponding to the Mth model level, the number of idle time slices for the Mth device node is determined to be M-1 (i.e., 3). Then, starting from the Nth batch of data, M-1 batches of data are selected from multiple batches of data in the reverse order of the specified batches as the Mth specified batch of data. For example, the Mth specified batch of data is batch d, batch c, and batch b in sequence. Accordingly, the second gradient result corresponding to the Mth specified batch of data is determined in the idle time slice corresponding to the Mth device node using the inverse automatic differentiation method.

[0072] Additionally, in the case of i=3 mentioned above, for the i-th device node corresponding to the i-th model level, the i-th specified batch of data is selected based on the second gradient result corresponding to the i+1th specified batch of data (i.e., the M-th specified batch of data) determined by the inverse automatic differentiation method in the idle time slice of the i+1th device node corresponding to the i+1th model level. Since the i-th specified batch of data is a portion of the batches in the i+1th specified batch of data, and the sharding time of any specified batch of data in the i-th specified batch of data in the i+1th device node is earlier than the sharding time of any specified batch of data in the i-th device node, therefore, as... Figure 4 As shown, when i=3, the i-th specified batch of data can be batch d (e.g., B). 2d ), batch c (e.g., B)2c ).

[0073] Similarly, when i=2, the i-th specified batch data can be batch d (e.g., B1d). Here, B1d indicates that device 1 performs reverse calculation on sub-batch d. When i=1, the i-th specified batch data can be batch d (e.g., B... 0d B 0d This indicates that device 0 performs reverse computation on sub-batch d. Therefore, the i-th specified batch data is a portion of the i+1-th specified batch data, and the sharding time of any specified batch data in the i-th specified batch data on the i+1-th device node is earlier than the sharding time of any specified batch data on the i-th device node (e.g., B). 3C The fragmentation time is earlier than B 2C And, B 3d The fragmentation time is earlier than B 2d B 2d The fragmentation time is earlier than B 1d B 1d The fragmentation time is earlier than B 0d ), among which, B 3c This indicates that device 3 performs reverse computation on sub-batch c, B 3d This indicates that device 3 performs reverse computation on sub-batch d; B 2c This indicates that device 2 performs reverse computation on sub-batch c, B 2d This indicates that device 2 performs the reverse computation on sub-batch d... and so on.

[0074] It should be noted that the method of inserting inverse calculation is not unique, as long as the constraints mentioned above are met. For example, Figure 5 This illustrates another alternative scheme where backward automatic differentiation and forward automatic differentiation are performed in a time-division manner. For example... Figure 5 As shown, in any training cycle, the fourth batch of data is B. 3C B 3d The third group of designated batch data is B. 2C B 2d The second group of designated batch data is B. 1C The first batch of data is designated as B. 0C .

[0075] Therefore, in the above example, reverse computation can be inserted during the idle time of the computation node (i.e., bubble time). When calculating the gradient, the second gradient result obtained from the reverse computation is used as the correction amount for the first gradient result obtained from the forward automatic differentiation. For example, the average of the two can be taken as the final gradient when updating the weights next time. It should be noted that the inserted reverse computation has the following characteristics: 1. Due to the limitation of the number of bubbles, reverse computation is only performed on a portion of the sub-batch (because there are not enough bubbles to insert reverse computation for all sub-batch). However, this portion of the sub-batch also has a positive effect on gradient correction. 2. There are some sub-batches that only proceed to the intermediate layers of the network during reverse computation, but not to the network input layer, such as sub-batch d. This is because there are not enough bubbles. However, this portion of the sub-batch will also have a positive effect on gradient correction. In addition, the intermediate states of the sub-batches that need to be reverse computed need to be cached during the forward computation process. However, in this case, since it is not necessary to reverse all network stages of all sub-batches, the number of caches is still significantly reduced compared to the complete reverse computation method.

[0076] In correcting the gradient, the first and second gradient results obtained from the same batch can be averaged using the same device to obtain the final gradient result for that batch. For example, in Figure 4 In the context of device 2, the forward differential result (i.e., the first gradient result) and the backward differential result (i.e., the second gradient result) of batch d can be averaged to obtain the final gradient result of batch d.

[0077] Furthermore, to increase the proportion of sub-batches performing backpropagation and reduce the proportion of bubble-like computations, the number of model stages (i.e., the number of model layers) can be appropriately increased, thereby inserting more backpropagation computations. This parallel scheme is applicable to many-core processors, as well as computing platforms such as GPUs and AI accelerators. This approach can avoid or alleviate the problem of excessive storage requirements caused by using backpropagation to solve gradients during large model training, while improving the utilization of computing resources in pipelined parallelism.

[0078] Therefore, this approach not only enables large-scale model training using forward automatic differentiation, but also allows for large-scale model training using a hybrid forward and backward automatic differentiation method, as well as pipelined parallel schemes for forward and hybrid automatic differentiation. This approach is adaptable to many-core chips, enabling high parallelism by performing physical mapping on many cores for computationally independent paths.

[0079] Additionally, the processing device in this example can be one or more processing cores in a many-core system. Alternatively, a processing device can be a chip in a many-core system that contains multiple processing cores. Or, a processing device can be a server device that contains multiple chips.

[0080] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0081] In addition, this disclosure also provides a model training device, an electronic device, and a computer-readable storage medium, all of which can be used to implement any of the model training methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the relevant section on methods and will not be repeated here.

[0082] Figure 6 This is a block diagram of a model training apparatus provided in an embodiment of the present disclosure.

[0083] Reference Figure 6 This disclosure provides a model training apparatus, wherein the model includes: multiple model layers running in a pipelined parallel manner across multiple device nodes; the apparatus includes:

[0084] The gradient determination module 61 is adapted to determine the second gradient result corresponding to a specified batch of data in the multiple batches of data in the idle time slices of the multiple device nodes in the process of determining the first gradient result corresponding to multiple batches of data ...

[0085] The update module 62 is adapted to correct the first gradient result based on the second gradient result to obtain the target gradient result, and update the model parameters of the model based on the target gradient result.

[0086] In one alternative implementation, the first gradient result is determined in the following way:

[0087] The multiple device nodes determine the first gradient result of the multiple batches of data in the current training cycle based on the first gradient result of the multiple batches of data in the previous training cycle using a forward automatic differentiation method.

[0088] In any training cycle, the start time of the first non-idle time segment of the mth device node corresponding to the mth model level is earlier than the start time of the first non-idle time segment of the (m+1)th device node corresponding to the (m+1)th model level; and the end time of the last non-idle time segment of the mth device node corresponding to the mth model level is earlier than the end time of the last non-idle time segment of the (m+1)th device node corresponding to the (m+1)th model level; where m is a natural number.

[0089] In one optional implementation, the multiple batches of data are arranged sequentially according to a specified batch order; and each device node is used to process the multiple batches of data sequentially in multiple time slices according to the specified batch order.

[0090] In any training cycle, the number of non-idle time shards for each device node is the same, and the number of non-idle time shards is determined according to the batch number of the multiple batches of data; the number of idle time shards for each device node is the same, and the number of idle time shards is determined according to the number of the multiple device nodes.

[0091] In one alternative implementation, the second gradient result is determined in the following way:

[0092] In any training cycle, for the i-th device node corresponding to the i-th model level, the i-th specified batch of data is selected based on the second gradient result of the i+1th specified batch of data determined by the inverse automatic differentiation method in the idle time slice of the i+1th device node corresponding to the i+1th model level.

[0093] In the idle time slice corresponding to the i-th device node, the second gradient result corresponding to the i-th group of specified batch data is determined according to the inverse automatic differentiation method; where i is a natural number.

[0094] In one optional implementation, the i-th specified batch data is a portion of the batch data in the (i+1)-th specified batch data;

[0095] Furthermore, the sharding time of any specified batch of data in the i-th group of specified batches in the i+1-th device node is earlier than the sharding time of any specified batch of data in the i-th device node.

[0096] In one optional implementation, the number of nodes of the plurality of device nodes is M, and the number of batches of the plurality of batch data is N; i is a natural number greater than or equal to 1 and less than M;

[0097] The second gradient result is determined in the following way:

[0098] In any training cycle, for the Mth device node corresponding to the Mth model level, determine the number of idle time slices for the Mth device node;

[0099] Based on the number of shards in the idle time sharding, starting from the Nth batch of data, in the reverse order of the specified batch order, select the Mth set of specified batch data that matches the number of shards from the plurality of batch data;

[0100] In the idle time slice corresponding to the Mth device node, the second gradient result corresponding to the Mth group of specified batch data is determined according to the inverse automatic differentiation method.

[0101] In an alternative implementation, the gradient determination module is further configured to:

[0102] During the process of determining the first gradient results corresponding to multiple batches of data by the multiple device nodes according to the forward automatic differentiation method, for the specified batch of data in the Mth group, the intermediate result data of the specified batch of data in the forward calculation process is cached;

[0103] The second gradient result is determined in the following way:

[0104] Based on the intermediate results of the specified batch of data in the forward computation process, the second gradient result corresponding to the specified batch of data in the idle time slice corresponding to the Mth device node is determined according to the reverse automatic differentiation method.

[0105] In one alternative implementation, the update module is specifically used for:

[0106] For any device node, perform a preset operation on the second gradient result and the first gradient result corresponding to the same batch of data, and obtain the target gradient result corresponding to the same batch of data based on the operation result.

[0107] In one alternative implementation, the device node includes: a processing core in a many-core system, a chip device containing multiple processing cores, and / or a server device containing multiple chips.

[0108] Figure 7 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.

[0109] Reference Figure 7This disclosure provides an electronic device, which includes: at least one processor 901; at least one memory 902; and one or more I / O interfaces 903 connected between the processor 901 and the memory 902; wherein the memory 902 stores one or more computer programs that can be executed by the at least one processor 901, and the one or more computer programs are executed by the at least one processor 901 to enable the at least one processor 901 to perform the above-described model training method.

[0110] In some embodiments, the processing device can be a neuromorphic chip. Since neuromorphic chips can employ vectorized computation and require external memory, such as Double Data Rate (DDR) synchronous dynamic random access memory, to load parameters like weights of the neural network model, the batch processing method used in this disclosure provides higher computational efficiency.

[0111] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor / processor core, implements the model training method described above. The computer-readable storage medium may be volatile or non-volatile.

[0112] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the above-described model training method.

[0113] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).

[0114] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0115] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0116] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0117] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0118] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0119] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0120] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0122] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A model training method, wherein the model comprises: Multiple model levels run in a pipelined parallel manner across multiple device nodes; The method includes: During the process of determining the first gradient result corresponding to multiple batches of data by the multiple device nodes according to the forward automatic differentiation method, the second gradient result corresponding to a specified batch of data in the multiple batches of data is determined by the backward automatic differentiation method in the idle time slice of the multiple device nodes. Based on the second gradient result, the first gradient result is corrected to obtain the target gradient result, and the model parameters of the model are updated based on the target gradient result.

2. The method according to claim 1, wherein, The multiple device nodes determine the first gradient results corresponding to multiple batches of data according to the forward automatic differentiation method, including: The multiple device nodes determine the first gradient result of the multiple batches of data in the current training cycle based on the first gradient result of the multiple batches of data in the previous training cycle using a forward automatic differentiation method. In any training cycle, the start time of the first non-idle time segment of the mth device node corresponding to the mth model level is earlier than the start time of the first non-idle time segment of the (m+1)th device node corresponding to the (m+1)th model level; and the end time of the last non-idle time segment of the mth device node corresponding to the mth model level is earlier than the end time of the last non-idle time segment of the (m+1)th device node corresponding to the (m+1)th model level; where m is a natural number.

3. The method according to claim 2, wherein, The multiple batches of data are arranged sequentially according to a specified batch order; and each device node is used to process the multiple batches of data sequentially in multiple time slices according to the specified batch order. In any training cycle, the number of non-idle time shards for each device node is the same, and the number of non-idle time shards is determined according to the batch number of the multiple batches of data; the number of idle time shards for each device node is the same, and the number of idle time shards is determined according to the number of the multiple device nodes.

4. The method according to claim 2 or 3, wherein, The determination of the second gradient result corresponding to a specific batch of data in the multiple batches of data according to the inverse automatic differentiation method in the idle time shards of the multiple device nodes includes: In any training cycle, for the i-th device node corresponding to the i-th model level, the i-th specified batch of data is selected based on the second gradient result of the i+1th specified batch of data determined by the inverse automatic differentiation method in the idle time slice of the i+1th device node corresponding to the i+1th model level. In the idle time slice corresponding to the i-th device node, the second gradient result corresponding to the i-th group of specified batch data is determined according to the inverse automatic differentiation method; where i is a natural number.

5. The method according to claim 4, wherein, The i-th specified batch data is a portion of the batch data in the (i+1)-th specified batch data; Furthermore, the sharding time of any specified batch of data in the i-th group of specified batches in the (i+1)-th device node is earlier than the sharding time of any specified batch of data in the i-th device node.

6. The method according to claim 4, wherein, The number of nodes in the plurality of device nodes is M, and the number of batches in the plurality of batch data is N; i is a natural number greater than or equal to 1 and less than M; The determination of the second gradient result corresponding to a specified batch of data in the multiple batches of data using the inverse automatic differentiation method in the idle time shards of the multiple device nodes includes: In any training cycle, for the Mth device node corresponding to the Mth model level, determine the number of idle time slices for the Mth device node; Based on the number of shards in the idle time sharding, starting from the Nth batch of data, in the reverse order of the specified batch order, select the Mth set of specified batch data that matches the number of shards from the plurality of batch data; In the idle time slice corresponding to the Mth device node, the second gradient result corresponding to the Mth group of specified batch data is determined according to the inverse automatic differentiation method.

7. The method according to claim 6, wherein, The method further includes: During the process of determining the first gradient results corresponding to multiple batches of data by the multiple device nodes according to the forward automatic differentiation method, for the specified batch of data in the Mth group, the intermediate result data of the specified batch of data in the forward calculation process is cached; The determination of the second gradient result corresponding to the specified batch of data in the idle time slice corresponding to the Mth device node according to the inverse automatic differentiation method includes: Based on the intermediate results of the specified batch of data in the forward computation process, the second gradient result corresponding to the specified batch of data in the idle time slice corresponding to the Mth device node is determined according to the reverse automatic differentiation method.

8. The method according to any one of claims 1-3, wherein, The step of correcting the first gradient result based on the second gradient result to obtain the target gradient result includes: For any device node, perform a preset operation on the second gradient result and the first gradient result corresponding to the same batch of data, and obtain the target gradient result corresponding to the same batch of data based on the operation result.

9. The method according to any one of claims 1-3, wherein, The model is a large model, and the method further includes: The model network of the large model is pre-divided into multiple model layers, and the multiple model layers are allocated to multiple device nodes so that the multiple device nodes run the multiple model layers in a pipelined parallel manner in stages.

10. The method according to any one of claims 1-3, wherein, The device nodes include: processing cores in a many-core system, chip devices containing multiple processing cores, and / or server devices containing multiple chips.

11. A model training apparatus, the model comprising: Multiple model levels run in a pipelined parallel manner across multiple device nodes; The device includes: The gradient determination module is adapted to determine the second gradient result corresponding to a specified batch of data in the multiple batches of data in the idle time slices of the multiple device nodes in the process of determining the first gradient result corresponding to multiple batches of data ... The update module is adapted to correct the first gradient result based on the second gradient result to obtain a target gradient result, and update the model parameters of the model based on the target gradient result.

12. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the model training method as described in any one of claims 1-10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the model training method as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Model training system and method, storage medium and electronic equipment

    CN118378726A

  • Training giant neural networks using pipeline parallelism

    US20210042620A1