Model training method and related equipment
By using the aggregation average gradient method to update parameters in distributed data parallel or mixed parallel training of large models, the problem of high training costs and low resource utilization is solved, and a more efficient training process is achieved.
Patent Information
- Application Number
- CN202411998647.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Training large models on a single computing device and a single computing node is expensive, and existing parallel and distributed algorithms are difficult to effectively improve resource utilization.
A model training method is provided, by reducing data transmission between multiple training instances in distributed data parallel or distributed hybrid parallel training, and the parameter update is used to optimize resource utilization.
It shortens the training time of large models, improves resource utilization during large models training, and reduces system complexity and data transmission.
Smart Images

Figure CN119939246A_ABST
Abstract
Description
[0001] This application claims the priority of the Chinese patent application filed with the China Patent Office on May 23, 2024, with application number 202410650013.7 and invention name “Model training method and related equipment based on distributed data parallelism”, the entire contents of which are incorporated by reference in this application. Technical Field
[0002] The embodiments of the present application relate to the field of neural networks, and in particular to model training methods and related equipment. Background Art
[0003] Deep learning models are profoundly and dramatically changing all aspects of today's society, work, and life. However, as the model size (parameters) grows, it is no longer possible to train the model on a single computing device (such as a graphics card) or a single computing node. To solve this problem, the industry and academia have proposed a variety of parallel and distributed algorithms, among which the most representative algorithms are distributed data parallelism, distributed model parallelism (pipeline parallelism, tensor parallelism), etc.
[0004] In fact, even with existing technical solutions, the cost of training a large model is extremely high. For a large model with billions of parameters, the training cost, including hardware investment and daily operation and maintenance, is usually as high as tens of millions or even hundreds of millions of dollars. Therefore, how to improve resource utilization during large model training has become a major issue that needs to be solved urgently. Summary of the invention
[0005] The embodiments of the present application provide a model training method and related equipment for reducing data transmission between multiple training instances when training a large model in distributed data parallel or distributed hybrid parallel mode, thereby shortening the training time of the large model and improving resource utilization during the training of the large model.
[0006] A first aspect of an embodiment of the present application provides a model training method, which is applied to a large model including M layers of networks, each layer of the network runs in one or L instances, M is greater than 1, and L is greater than 1, and the method further includes:
[0007] Get the training loss of the large model in the Nth round of training;
[0008] If the number of j-th layer instances is 1 and the number of j+1-th layer instances is L, then when back propagation is performed to update the parameters, the average of the returned reverse gradients of the j+1-th layer network of the multiple j+1-th layer instances in the N-th round of training is used as the aggregated average gradient of the j+1-th layer network in the N-th round of training, and the product of the aggregated average gradient of the j+1-th layer network in the N-th round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance is used to calculate the parameter gradient of the j-th layer network of the j-th layer instance, and the returned reverse gradient of the j+1-th layer network of the multiple j+1-th layer instances in the N-th round of training is: the gradient of the training loss with respect to the input of the j+1-th layer network of the j-th layer instance during the forward propagation of the N-th round of training;
[0009] Based on the parameter gradient of the j-th layer network of the j-th layer instance, the j-th layer network of the j-th layer instance is updated.
[0010] In a specific implementation, the method further includes:
[0011] If the number of j-th layer instances is L and the number of j+1-th layer instances is 1, then for each j-th layer network of the j-th layer instance, the product of the backpropagated reverse gradient of the j+1-th layer network of the j+1-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance is used as the parameter gradient of the j-th layer network of the j-th layer instance.
[0012] In a specific implementation, the method further includes:
[0013] If the number of j-th layer instances is equal to the number of j+1-th layer instances, then for each j-th layer network of the j-th layer instance, the product of the backpropagated backward gradient of the j+1-th layer network of the j+1-th layer instance corresponding one-to-one to the j-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance is used as the parameter gradient of the j-th layer network of the j-th layer instance.
[0014] In a specific implementation, the method further includes:
[0015] If the number of j-1-th layer instances is L and the number of j-th layer instances is 1, then for the j-1-th layer network of each j-1-th layer instance, the product of the backpropagated reverse gradient of the j-th layer network of the j-th layer instance in the Nth round of training and the Jacobian matrix of the j-1-th layer network of the j-1-th layer instance is used as the parameter gradient of the j-1-th layer network of the j-1-th layer instance.
[0016] In a specific implementation, the method further includes:
[0017] If the large model meets the preset parameter synchronization conditions, each network layer running in the L instances is determined as a layer to be synchronized;
[0018] For each layer to be synchronized, L parameters to be synchronized corresponding to the layer to be synchronized are obtained from the L instances running the layer to be synchronized, and based on the L parameters to be synchronized corresponding to the layer to be synchronized, the aggregation parameters corresponding to the layer to be synchronized are determined, and the parameters to be synchronized in the L instances running the layer to be synchronized are updated to the aggregation parameters corresponding to the layer to be synchronized.
[0019] In a specific implementation, applied to a first training end, the large model is trained in parallel by multiple training ends, the multiple training ends include the first training end, the multiple training ends share each layer of the network running in one instance in the large model, and obtaining the training loss of the large model in the Nth round of training includes:
[0020] The N-th round weighted training loss is obtained as the training loss of the large model in the N-th round of training, wherein the weighted training loss is a weighted average of the N-th round training losses of multiple training ends, wherein the multiple training ends include the first training end, and the N-th round training loss of each of the training ends is determined based on a preset loss function and an output of the M-th layer network in the N-th round of training of the training end.
[0021] A second aspect of an embodiment of the present application provides a training end, which is used to train a large model including M layers of networks, each layer of the network runs in one or L instances, L is greater than 1, including:
[0022] An acquisition unit, used to acquire the training loss of the large model in the Nth round of training;
[0023] A calculation unit, for, if the number of j-th layer instances is 1 and the number of j+1-th layer instances is L, taking the average of the returned reverse gradients of the j+1-th layer network of the multiple j+1-th layer instances in the N-th round of training as the aggregated average gradient of the j+1-th layer network in the N-th round of training when back propagation is performed to update the parameters, and using the product of the aggregated average gradient of the j+1-th layer network in the N-th round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance to calculate the parameter gradient of the j-th layer network of the j-th layer instance, the returned reverse gradient of the j+1-th layer network of the multiple j+1-th layer instances in the N-th round of training is: the gradient of the training loss with respect to the input of the j+1-th layer network of the j-th layer instance during the forward propagation of the N-th round of training;
[0024] An updating unit is used to update the j-th layer network of the j-th layer instance based on the parameter gradient of the j-th layer network of the j-th layer instance.
[0025] In a specific implementation, the computing unit is also used to, if the number of j-th layer instances is L and the number of j+1-th layer instances is 1, then, for each j-th layer network of the j-th layer instance, use the product of the backpropagated reverse gradient of the j+1-th layer network of the j+1-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance as the parameter gradient of the j-th layer network of the j-th layer instance.
[0026] In a specific implementation, the computing unit is further used to, if the number of j-th layer instances is equal to the number of j+1-th layer instances, then, for each j-th layer network of the j-th layer instance, use the product of the backpropagated backward gradient of the j+1-th layer network of the j+1-th layer instance corresponding one-to-one to the j-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance as the parameter gradient of the j-th layer network of the j-th layer instance.
[0027] In a specific implementation, the computing unit is further used to, if the number of j-1 layer instances is L and the number of j-layer instances is 1, then, for each j-1 layer network of the j-1 layer instance, use the product of the backpropagated reverse gradient of the j-layer network of the j-layer instance in the Nth round of training and the Jacobian matrix of the j-1 layer network of the j-1 layer instance as the parameter gradient of the j-1 layer network of the j-1 layer instance.
[0028] In a specific implementation, the training end further includes: a determination unit;
[0029] The determining unit is configured to determine each network layer running in the L instances as a layer to be synchronized if the large model meets a preset parameter synchronization condition;
[0030] The determination unit is further used to obtain, for each layer to be synchronized, L parameters to be synchronized corresponding to the layer to be synchronized from L instances running the layer to be synchronized, and determine the aggregation parameters corresponding to the layer to be synchronized based on the L parameters to be synchronized corresponding to the layer to be synchronized, and update the parameters to be synchronized in the L instances running the layer to be synchronized to the aggregation parameters corresponding to the layer to be synchronized.
[0031] In a specific implementation, applied to a first training end, the large model is trained in parallel by multiple training ends, the multiple training ends include the first training end, the multiple training ends share each layer of the network running in an instance in the large model, and the acquisition unit is specifically used to obtain the Nth round weighted training loss as the training loss of the large model in the Nth round of training, the weighted training loss is the weighted average of the Nth round training losses of multiple training ends, the multiple training ends include the first training end, and the Nth round training loss of each of the training ends is determined based on a preset loss function and the output of the Mth layer network in the Nth round of training of the training end.
[0032] A third aspect of an embodiment of the present application provides a training terminal, including:
[0033] CPU, memory and input / output interface;
[0034] The memory is a short-term storage memory or a persistent storage memory;
[0035] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method described in the first aspect.
[0036] A fourth aspect of the embodiments of the present application provides a computer program product comprising instructions, and when the computer program product is run on a computer, the computer is caused to execute the method described in the first aspect.
[0037] A fifth aspect of an embodiment of the present application provides a computer storage medium, wherein the computer storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the method described in the first aspect.
[0038] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages: In the existing hybrid parallel scheme, multiple training terminals can only be simply expressed as data parallelism or model parallelism internally or externally. If it is expressed as data parallelism externally and model parallelism internally, a complete large model needs to be run in each training terminal, and the multiple instances in the training terminal respectively run at least one network layer of the large model. The present application provides a new model training framework. In a scenario where multiple training terminals are expressed as data parallelism externally, if a certain network layer belongs to a network that occupies a large amount of storage resources but a small amount of computing resources, this network layer can be deleted from all training terminals, and this network layer can be run in a new instance and provided to all training terminals. It reduces the excessive occupation of storage resources when multiple network layers are redundantly deployed, and can effectively improve resource utilization. In addition, to perform model training under the training framework of the embodiment of the present application, only two places need to be used for gradient aggregation averaging and backpropagation, and synchronization is only required at these two places: 1. Aggregate and average the training losses of multiple training end forward propagations when the forward propagation of each training end is completed, and use the aggregated averaged loss value as the loss value of each j-th instance for independent backpropagation and parameter update; 2. Aggregate and average the backpropagated gradients on a single instance shared by multiple training ends when backpropagation is performed for parameter update. Therefore, compared with the existing training framework, the embodiment of the present application greatly reduces the amount of data transmission between multiple instances and the waiting time for synchronization between multiple instances, while greatly reducing the actual complexity of the system, which can effectively improve the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of an existing training framework disclosed in an embodiment of the present application;
[0040] Figure 2 A schematic diagram of the training framework of the present application disclosed in the embodiment of the present application;
[0041] Figure 3 A flow chart of the model training method disclosed in the embodiment of this application
[0042] Figure 4 A schematic diagram of the structure of the training end disclosed in the embodiment of the present application;
[0043] Figure 5 This is another structural schematic diagram of the training end disclosed in the embodiment of the present application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0045] The embodiments of the present application provide a model training method for efficiently performing model training in complex scenarios.
[0046] See also Figure 1 In existing distributed data parallel or hybrid parallel schemes, when backpropagation is used to update parameters, it is necessary to aggregate and average the gradients of each network layer of multiple instances of distributed training and transmit them back to each instance. This implementation method not only has a huge amount of data transmission between multiple instances, but also requires parameter synchronization between multiple instances in each network layer of the model during each training iteration, which increases the waiting time caused by synchronization and greatly increases the complexity of system design.
[0047] In order to solve the limitations of the existing training framework, based on the existing training framework, this application provides Figure 2 When the training framework of the embodiment of the present application is used for model training, multiple training ends for parallel training will share some network layers in the large model (such as Figure 2 The network layers 1, 2, and M-1 are deployed on a single instance. Figure 2 As shown in the M-3 layer network, the M-2 layer network and the M-th layer network), corresponding instances are deployed in each training end.
[0048] Specifically, other network layers in each training end can be deployed in one or more instances, which is not specifically limited in the embodiments of the present application.
[0049] See also Figure 3 The present application provides a model training method, which is applied to a large model including M layers of networks, each layer of the network runs in one or L instances, L is greater than 1, and includes the following steps:
[0050] 301. Obtain the training loss of the large model in the Nth round of training.
[0051] In practical applications, the large model will undergo multiple rounds of training during the training process, and after each round of training, the parameter gradient needs to be calculated through back propagation, and the network parameters are updated based on the calculated parameter gradient. In order to better illustrate the technical solution of the embodiment of the present application, the embodiment of the present application takes the Nth round as an example to describe the specific implementation of the model training method of the embodiment of the present application during the Nth round of training. In the following description, the traditional convention of the existing literature is followed, and the model layer count is based on the count during forward propagation, that is, the jth layer is calculated first and then the j+1th layer is calculated during forward propagation. On the contrary, during back propagation, the gradient of the j+1th layer is calculated first and the parameter update of the j+1th layer is completed, and then the gradient of the jth layer is calculated and the parameter update of the jth layer is completed. In addition, when calculating the gradient of each layer of network parameters, the forward gradient algorithm JVP and the backward gradient algorithm VJP can be used for implementation. In the following description, the backward gradient algorithm VJP, i.e., VJ=P, is used to calculate the gradient of the j-th layer, and V is called the reverse gradient of the j+1-th layer to the j-th layer in the directional propagation, J is called the parameter Jacobian matrix of the j-th layer, and P is called the parameter gradient of the j-th layer, which is referred to as the gradient of the j-th layer in the existing literature. The P value is also the reverse feedback gradient of the j-th layer to the j-1-th layer.
[0052] Similarly, the large model contains multiple layers of networks. Therefore, the jth layer in the embodiment of the present application can be any layer in the M-layer network of the large model, and no specific limitation is made here.
[0053] 302. If the number of j-th layer instances is 1 and the number of j+1-th layer instances is L, when backpropagation is performed to update the parameters, the average of the backpropagated reverse gradients of the j+1-th layer network of multiple j+1-th layer instances in the N-th round of training is used as the aggregated average gradient of the j+1-th layer network in the N-th round of training, and the product of the aggregated average gradient of the j+1-th layer network in the N-th round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance is used to calculate the parameter gradient of the j-th layer network of the j-th layer instance.
[0054] Specifically, in the embodiment of the present application, the only j-th layer instance is a single instance of the j-th layer network shared by multiple training terminals, and multiple j+1-th layer instances are specific instances of multiple j+1-th layer networks owned by each training terminal. This means that when back propagation is performed to update parameters, it is necessary to perform an aggregated average operation on the backpropagated reverse gradients of the j+1-th layer network of multiple j+1-th layer instances to obtain the aggregated average gradient of the j+1-th layer network. Then, the product of the aggregated average gradient of the j+1-th layer network in the Nth round of training and the Jacobian matrix of the j-th layer network of the single j-th layer instance is used as the parameter gradient of the j-th layer network of the single j-th layer instance.
[0055] If the embodiment j of the present application is Figure 2The corresponding relationship between the j-th layer instance and the j+1-th layer instance in the embodiment of the present application is M-1. Figure 2 The correspondence between the M-1th layer instance and the Mth layer instance is consistent. Figure 2 From the arrows between the Mth layer network of each Mth layer instance and the unique M-1th layer network, it can be seen that the gradient of the Mth layer network obtained by the M-1th layer network is the aggregated average gradient, which is the average of the sum of the returned reverse gradients of the Mth layer networks of multiple Mth layer instances.
[0056] It should be noted that the backpropagation gradient of any layer of the network is the gradient that needs to be transmitted back to the corresponding previous layer of the network during the gradient update process of backpropagation; and the parameter gradient of any layer of the network is the gradient used by this layer of the network when adjusting its own network parameters.
[0057] 303. Update the j-th layer network of the j-th layer instance based on the parameter gradient of the j-th layer network of the j-th layer instance.
[0058] Based on the above embodiments, it can be known that the parameter gradient of any layer of network is the gradient used by this layer of network when adjusting its own network parameters. Therefore, based on the parameter gradient of the jth layer of network of the jth layer instance, the new parameters of the jth layer of network of the jth layer instance can be calculated. Finally, the parameters of the jth layer of network of the jth layer instance are updated to the new parameters calculated above to complete the update of the jth layer of network of the jth layer instance.
[0059] In an embodiment of the present application, in a scenario where multiple training ends appear to be data parallel, if a certain network layer needs to occupy a large amount of memory storage resources but the required computing resources are relatively small, a solution can be adopted in which a single instance of the network layer is shared by all training ends, thereby reducing the excessive occupation of storage resources caused by each training end instance using a separate instance of the network layer, effectively improving the utilization rate of memory storage resources, and reducing the pressure on GPU memory resource usage.
[0060] In some specific implementations, if the number of j-th layer instances is L and the number of j+1-th layer instances is 1, then for the j-th layer network of each j-th layer instance, the product of the backpropagated backward gradient of the j+1-th layer network of the j+1-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance is used as the parameter gradient of the j-th layer network of the j-th layer instance.
[0061] Specifically, in the embodiment of the present application, multiple j-th layer instances are specific instances of the j-th layer network independently owned by multiple training terminals, and the unique j+1-th layer instance is a single instance of the j+1-th layer network shared by multiple training terminals. This means that when back propagation is performed to update parameters, the j-th layer network of each j-th layer instance calculates its own parameter gradient based on the product of the return reverse gradient of the unique j+1-th layer instance and its own Jacobian matrix. This is equivalent to sending a "copy" of the return reverse gradient of the unique j+1-th layer instance to the j-th layer network of each j-th layer instance, so that the j-th layer network of each j-th layer instance can calculate the parameter gradient.
[0062] If the embodiment j of the present application is Figure 2 The corresponding relationship between the j-th layer instance and the j+1-th layer instance in the embodiment of the present application is M-2. Figure 2 The correspondence between the M-2 layer instances and the M-1 layer instances is consistent. Figure 2 From the arrows between the M-2th layer network of each M-2th layer instance and the unique M-1th layer network, it can be seen that the gradient of the M-1th layer network obtained by the M-2th layer network of each M-2th layer instance is the returned reverse gradient of the unique M-1th layer instance.
[0063] In some other specific implementations, if the number of j-th layer instances is equal to the number of j+1-th layer instances, then for the j-th layer network of each j-th layer instance, the product of the backpropagated backward gradient of the j+1-th layer network of the j+1-th layer instance and the Jacobian matrix of the j-th layer network of the j-th layer instance corresponding one by one in the Nth round of training is used as the parameter gradient of the j-th layer network of the j-th layer instance.
[0064] Specifically, if the number of j-th layer instances is equal to the number of j+1-th layer instances, it means that each j-th layer instance has a corresponding j-th layer instance. This may be in the following two cases: 1. Multiple j-th layer instances are specific instances of multiple j-th layer networks owned by each training end alone, and multiple j+1-th layer instances are specific instances of multiple j+1-th layer networks owned by each training end alone, and the j-th layer network of each instance is uniquely connected to the j+1-th layer network of the same instance; 2. The only j-th layer instance is a single instance of the j-th layer network shared by multiple training ends, and the only j+1-th layer instance is a single instance of the j+1-th layer network shared by multiple training ends.
[0065] In the first case, this means that when backpropagation is performed to update parameters, the returned reverse gradient of the j+1th layer network of the j+1th layer instance corresponding to each training end only needs to be returned to the jth layer network of the j+1th layer instance corresponding to the training end. Then the jth layer network of the j+1th layer instance of each training end calculates the parameter gradient of the jth layer network based on the received returned reverse gradient of the j+1th layer network of the corresponding j+1th layer instance. It should be noted that in the first case, the embodiment of the present application does not need to aggregate and average the returned reverse gradients of the j+1th layer networks of multiple j+1th layer instances, nor does it need to "copy" the returned reverse gradient of the j+1th layer network of the j+1th layer instance and send it to the jth layer networks of multiple jth layer instances.
[0066] If the embodiment j of the present application is Figure 2 The corresponding relationship between the j-th layer instance and the j+1-th layer instance in the embodiment of the present application is as follows: Figure 2 The corresponding relationship between the M-3 layer instance and the M-2 layer instance is consistent. Figure 2 It can be seen from the arrows between the M-3th layer networks of multiple M-3th layer instances and the multiple M-2th layer networks that the gradient of the M-2th layer network obtained by the M-3th layer network of each M-3th layer instance corresponds one-to-one to the return reverse gradient of the M-2th layer network of the M-2th layer instance.
[0067] In the second case, this means that during backpropagation for parameter updates, the j-th layer of the unique j-th instance of the network calculates the parameter gradients based on the backpropagated gradients of the j+1-th layer of the unique j+1-th instance of the network.
[0068] If the embodiment j of the present application is Figure 2 The first layer network shown in , then the quantitative relationship between the jth layer instance and the j+1th layer instance of the embodiment of the present application and Figure 2 The corresponding relationship between the first layer network instance and the second layer network instance is consistent. Figure 2 From the arrow between the first-layer network of the only first-layer instance and the only second-layer network, it can be seen that the gradient of the second-layer network obtained by the first-layer network of the only first-layer instance is the returned reverse gradient of the second-layer network of the only second-layer instance.
[0069] In some specific implementations, if the large model meets the preset parameter synchronization conditions, each network layer running in the L instances is determined as a layer to be synchronized; for each layer to be synchronized, L parameters to be synchronized corresponding to the layer to be synchronized are obtained from the L instances running the layer to be synchronized, and based on the L parameters to be synchronized corresponding to the layer to be synchronized, the aggregation parameters corresponding to the layer to be synchronized are determined, and the parameters to be synchronized in the L instances running the layer to be synchronized are updated to the aggregation parameters corresponding to the layer to be synchronized.
[0070] Based on the foregoing embodiments, it can be known that when multiple j-th layer instances are specific instances of multiple j-th layer networks independently owned by each training end, or multiple j+1-th layer instances are specific instances of multiple j+1-th layer networks independently owned by each training end, the embodiments of the present application do not synchronize network parameters between multiple j+1-th layer networks independently owned by different training ends.
[0071] Therefore, in order to further improve the convergence speed of the large model, based on the above-mentioned embodiment, the embodiment of the present application also provides a technical solution to synchronize the network parameters between multiple layers to be synchronized that are independently owned by different training terminals. Figure 2 Taking the large model shown as an example, the M-3 layer network, the M-2 layer network, and the M layer network shown in the figure all belong to the layers to be synchronized.
[0072] Specifically, for the multi-layer network of the large model, in addition to the network layer shared by multiple training terminals, the embodiment of the present application will obtain L parameters to be synchronized corresponding to the layer to be synchronized from the L instances running the same layer to be synchronized. Then, weighted aggregation processing is performed based on the L parameters to be synchronized corresponding to the layer to be synchronized to determine the aggregation parameters corresponding to the layer to be synchronized. Specifically, the weighted weights corresponding to the L instances running the same layer to be synchronized can be configured as needed. For example, the weighted weights corresponding to multiple instances can be the same or determined according to the number of training samples used in each instance in the Nth round of training, which is not specifically limited here. Finally, the parameters to be synchronized in the L instances running the layer to be synchronized are updated to the aggregation parameters corresponding to the layer to be synchronized.
[0073] In practical applications, for a solution in which the Mth layer network runs in L Mth layer instances, an embodiment of the present application is applied to a first training end, and a large model is trained in parallel by multiple training ends, the multiple training ends include a first training end, and the multiple training ends share each layer of the network running in an instance in the large model. The aforementioned step 301 can be specifically implemented in the following manner: obtaining the Nth round weighted training loss as the training loss of the large model in the Nth round of training, the weighted training loss is the weighted average of the Nth round training losses of multiple training ends, the multiple training ends include the first training end, and the Nth round training loss of each training end is determined based on a preset loss function and the output of the Mth layer network in the Nth round of training of the training end.
[0074] Specifically, the training loss of each training end in the Nth round of training is determined based on a preset loss function and the output of the Mth layer network of the training end in the Nth round of training. The weighted training loss is the weighted average of the Nth round of training losses of multiple training ends, wherein the weighted weight of each training end can be determined based on the amount of data of its own Nth round of training data, and is positively correlated with the amount of data used in its own Nth round of training data, which is not specifically limited here. It should be noted that if the weighted weights of each training end are the same, the weighted average processing is equivalent to the sum average processing.
[0075] To train the model under the training framework of the embodiment of the present application, only two places are needed to aggregate and average the gradients and return them, and synchronization is only required at these two places: 1. Aggregate and average the training losses of multiple training ends when the forward propagation is completed at each training end, and use the aggregated averaged loss value as the loss value of each j-th instance for independent back propagation and parameter update; 2. Aggregate and average the back propagated gradients on a single instance shared by multiple training ends when back propagation is performed for parameter update. Therefore, compared with the existing training framework, the embodiment of the present application greatly reduces the amount of data transmission between multiple instances and the waiting time for synchronization between multiple instances, while greatly reducing the actual complexity of the system, which can effectively improve the training efficiency.
[0076] See also Figure 4 The embodiment of the present application provides a training end, which is used to train a large model including M layers of networks, each layer of the network runs in one or L instances, L is greater than 1, including:
[0077] An acquisition unit 401 is used to acquire the training loss of the large model in the Nth round of training;
[0078] The calculation unit 402 is used for, if the number of j-th layer instances is 1 and the number of j+1-th layer instances is L, then, when back propagation is performed to update the parameters, using the average of the returned reverse gradients of the j+1-th layer network of multiple j+1-th layer instances in the N-th round of training as the aggregated average gradient of the j+1-th layer network in the N-th round of training, and using the product of the aggregated average gradient of the j+1-th layer network in the N-th round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance to calculate the parameter gradient of the j-th layer network of the j-th layer instance, and the returned reverse gradient of the j+1-th layer network of multiple j+1-th layer instances in the N-th round of training is: the gradient of the training loss with respect to the input of the j+1-th layer network of the j-th layer instance during the forward propagation of the N-th round of training;
[0079] The updating unit 403 is used to update the j-th layer network of the j-th layer instance based on the parameter gradient of the j-th layer network of the j-th layer instance.
[0080] In a specific implementation, the computing unit 402 is also used to, if the number of j-th layer instances is L and the number of j+1-th layer instances is 1, then, for each j-th layer network of the j-th layer instance, use the product of the backpropagated reverse gradient of the j+1-th layer network of the j+1-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance as the parameter gradient of the j-th layer network of the j-th layer instance.
[0081] In a specific implementation, the computing unit 402 is further used to, if the number of j-th layer instances is equal to the number of j+1-th layer instances, then, for each j-th layer network of the j-th layer instance, use the product of the backpropagated backward gradient of the j+1-th layer network of the j+1-th layer instance corresponding one-to-one to the j-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance as the parameter gradient of the j-th layer network of the j-th layer instance.
[0082] In a specific implementation, the computing unit 402 is also used to, if the number of j-1 layer instances is L and the number of j-layer instances is 1, then, for each j-1 layer network of the j-1 layer instance, use the product of the backpropagated reverse gradient of the j-layer network of the j-layer instance in the Nth round of training and the Jacobian matrix of the j-1 layer network of the j-1 layer instance as the parameter gradient of the j-1 layer network of the j-1 layer instance.
[0083] In a specific implementation, the training end further includes: a determination unit;
[0084] A determination unit, configured to determine each network layer running in the L instances as a layer to be synchronized if the large model meets a preset parameter synchronization condition;
[0085] The determination unit is also used to obtain, for each layer to be synchronized, L parameters to be synchronized corresponding to the layer to be synchronized from the L instances running the layer to be synchronized, and determine the aggregation parameters corresponding to the layer to be synchronized based on the L parameters to be synchronized corresponding to the layer to be synchronized, and update the parameters to be synchronized in the L instances running the layer to be synchronized to the aggregation parameters corresponding to the layer to be synchronized.
[0086] In a specific implementation, applied to a first training end, a large model is trained in parallel by multiple training ends, the multiple training ends include a first training end, and the multiple training ends share each layer of the network running in an instance in the large model. An acquisition unit 401 is specifically used to obtain an N-th round weighted training loss as the training loss of the large model in the N-th round of training. The weighted training loss is a weighted average of the N-th round training losses of multiple training ends. The multiple training ends include the first training end. The N-th round training loss of each training end is determined based on a preset loss function and an output of the M-th layer network in the N-th round of training of the training end.
[0087] Figure 51 is a schematic diagram of a training terminal structure provided in an embodiment of the present application. The training terminal 500 may include one or more central processing units (CPU) 501 and a memory 505. The memory 505 stores one or more application programs or data.
[0088] The memory 505 may be a volatile storage or a persistent storage. The program stored in the memory 505 may include one or more modules, each of which may include a series of instruction operations on the training end. Furthermore, the central processor 501 may be configured to communicate with the memory 505 and execute a series of instruction operations in the memory 505 on the training end 500.
[0089] The training end 500 may also include one or more power supplies 502, one or more wired or wireless network interfaces 503, one or more input and output interfaces 504, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, NinuxTM, FreeBSDTM, etc.
[0090] The CPU 501 can execute the aforementioned Figures 1 to 4 The operations performed by the training end in the illustrated embodiment are not described in detail here. In the embodiment of the present application, the central processing unit 501 can be replaced by a graphics processing unit (GPU).
[0091] It should be noted that, although the steps in the flowcharts involved in the embodiments are drawn in sequence according to the instructions of the arrows, unless otherwise clearly stated in this document, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0092] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0093] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0094] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0095] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a training terminal (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
[0097] An embodiment of the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the model training method as described above.
Claims
1. A model training method, characterized in that: Applied to a large model including M layers of networks, each layer of the network runs in one or L instances, M is greater than 1, and L is greater than 1, the method further includes: Get the training loss of the large model in the Nth round of training; If the number of j-th layer instances is 1 and the number of j+1-th layer instances is L, then when back propagation is performed to update the parameters, the average of the returned reverse gradients of the j+1-th layer network of the multiple j+1-th layer instances in the N-th round of training is used as the aggregated average gradient of the j+1-th layer network in the N-th round of training, and the product of the aggregated average gradient of the j+1-th layer network in the N-th round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance is used to calculate the parameter gradient of the j-th layer network of the j-th layer instance, and the returned reverse gradient of the j+1-th layer network of the multiple j+1-th layer instances in the N-th round of training is: the gradient of the training loss with respect to the input of the j+1-th layer network of the j-th layer instance during the forward propagation of the N-th round of training; Based on the parameter gradient of the j-th layer network of the j-th layer instance, the j-th layer network of the j-th layer instance is updated.
2. The model training method according to claim 1, characterized in that: The method further comprises: If the number of j-th layer instances is L and the number of j+1-th layer instances is 1, then for each j-th layer network of the j-th layer instance, the product of the backpropagated reverse gradient of the j+1-th layer network of the j+1-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance is used as the parameter gradient of the j-th layer network of the j-th layer instance.
3. The model training method according to claim 1, characterized in that: The method further comprises: If the number of j-th layer instances is equal to the number of j+1-th layer instances, then for each j-th layer network of the j-th layer instance, the product of the backpropagated backward gradient of the j+1-th layer network of the j+1-th layer instance corresponding one-to-one to the j-th layer instance in the Nth round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance is used as the parameter gradient of the j-th layer network of the j-th layer instance.
4. The model training method according to claim 1, characterized in that: The method further comprises: If the number of j-1-th layer instances is L and the number of j-th layer instances is 1, then for the j-1-th layer network of each j-1-th layer instance, the product of the backpropagated reverse gradient of the j-th layer network of the j-th layer instance in the Nth round of training and the Jacobian matrix of the j-1-th layer network of the j-1-th layer instance is used as the parameter gradient of the j-1-th layer network of the j-1-th layer instance.
5. The model training method according to claim 1, characterized in that: The method further comprises: If the large model meets the preset parameter synchronization conditions, each network layer running in the L instances is determined as a layer to be synchronized; For each layer to be synchronized, L parameters to be synchronized corresponding to the layer to be synchronized are obtained from the L instances running the layer to be synchronized, and based on the L parameters to be synchronized corresponding to the layer to be synchronized, the aggregation parameters corresponding to the layer to be synchronized are determined, and the parameters to be synchronized in the L instances running the layer to be synchronized are updated to the aggregation parameters corresponding to the layer to be synchronized.
6. The model training method according to any one of claims 1 to 5, characterized in that: Applied to a first training end, the large model is trained in parallel by multiple training ends, the multiple training ends include the first training end, the multiple training ends share each layer of the network running in one instance in the large model, and obtaining the training loss of the large model in the Nth round of training includes: The N-th round weighted training loss is obtained as the training loss of the large model in the N-th round of training, wherein the weighted training loss is a weighted average of the N-th round training losses of multiple training ends, wherein the multiple training ends include the first training end, and the N-th round training loss of each of the training ends is determined based on a preset loss function and an output of the M-th layer network in the N-th round of training of the training end.
7. A training terminal, characterized in that: The training end is used to train a large model including M layers of networks, each layer of the network runs in one or L instances, L is greater than 1, including: An acquisition unit, used to acquire the training loss of the large model in the Nth round of training; A calculation unit, for, if the number of j-th layer instances is 1 and the number of j+1-th layer instances is L, taking the average of the returned reverse gradients of the j+1-th layer network of the multiple j+1-th layer instances in the N-th round of training as the aggregated average gradient of the j+1-th layer network in the N-th round of training when back propagation is performed to update the parameters, and using the product of the aggregated average gradient of the j+1-th layer network in the N-th round of training and the Jacobian matrix of the j-th layer network of the j-th layer instance to calculate the parameter gradient of the j-th layer network of the j-th layer instance, the returned reverse gradient of the j+1-th layer network of the multiple j+1-th layer instances in the N-th round of training is: the gradient of the training loss with respect to the input of the j+1-th layer network of the j-th layer instance during the forward propagation of the N-th round of training; An updating unit is used to update the j-th layer network of the j-th layer instance based on the parameter gradient of the j-th layer network of the j-th layer instance.
8. A training terminal, characterized in that: include: CPU, memory and input / output interface; The memory is a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the model training method described in any one of claims 1 to 6.
9. A computer program product comprising instructions, characterized in that When the computer program product runs on a computer, the computer executes the model training method as described in any one of claims 1 to 6.
10. A computer storage medium, characterized in that: The computer storage medium stores instructions, and when the instructions are executed on a computer, the computer executes the model training method as described in any one of claims 1 to 6.