Method, apparatus, device, and storage medium for training a neural network

By dividing the network layer of the neural network into two groups, the local gradients of the current and previous training steps are obtained, and the gradient prediction compensation technology is used to solve the problem of communication delay in data parallel neural network training, and efficient model training is achieved.

CN115688867BActive Publication Date: 2025-07-18DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211431463.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-07-18
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

The existing data parallel neural network training method is difficult to achieve linear acceleration due to frequent communication and interactions between multiple computing devices, and the introduction of gradient delays adversely affects the convergence and accuracy of the model.

Method used

The network layer of the neural network is divided into two groups, and the local gradients of the current training step and the previous training step are obtained respectively. By introducing gradient delays for only some network layers, the communication overhead is hidden, and the gradient prediction compensation technology is used to reduce the latency impact.

Benefits of technology

While hiding communication overhead, it ensures the convergence and accuracy of the model, improves training speed, and reduces system resources and time overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115688867B_ABST
    Figure CN115688867B_ABST
Patent Text Reader

Abstract

According to various embodiments of the present disclosure, there are provided a method, an apparatus, a device, and a storage medium for training a neural network. In this method, at a first worker node among a plurality of worker nodes, a first set of global gradients for a first set of network layers in the neural network is obtained. The first set of global gradients is aggregated from local gradients determined by the plurality of worker nodes for the first set of network layers in the current training step. The plurality of worker nodes are configured to jointly train the neural network. Further, a second set of global gradients for a second set of network layers in the neural network is obtained. The second set of network layers is different from the first set of network layers. The second set of global gradients is aggregated from local gradients determined by the plurality of worker nodes for the second set of network layers in a previous training step before the current training step. In addition, based on the first set of global gradients and the second set of global gradients, the parameters of the neural network are updated. In this way, the convergence and accuracy of the trained neural network can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for training neural networks. Background Art

[0002] With the rapid development of computer technology, neural networks (NNs) are being increasingly widely used in fields such as computer vision and natural language processing. At the same time, the exponential growth of model size and data volume makes the training of neural networks time-consuming and resource-intensive.

[0003] Currently, the common method for accelerating neural network training is data parallelism, which uses multiple computing devices to train neural networks. Although the data parallelism method greatly speeds up the training speed, due to the frequent communication interactions between multiple computing devices, it is difficult for this method to achieve a linear acceleration effect. Summary of the Invention

[0004] According to a first aspect of the present disclosure, there is provided a method for training a neural network. In this method, at a first worker node among a plurality of worker nodes, a first set of global gradients for a first set of network layers in the neural network is obtained. The first set of global gradients is aggregated from local gradients determined by the plurality of worker nodes for the first set of network layers in the current training step. The plurality of worker nodes are configured to jointly train the neural network. Further, a second set of global gradients for a second set of network layers in the neural network is obtained. The second set of network layers is different from the first set of network layers. The second set of global gradients is aggregated from local gradients determined by the plurality of worker nodes for the second set of network layers in a previous training step before the current training step. In addition, based on the first set of global gradients and the second set of global gradients, the parameters of the neural network are updated.

[0005] According to a second aspect of the present disclosure, there is provided an apparatus for training a neural network. The apparatus includes a first acquisition module, a second acquisition module, and an update module. The first acquisition module is configured to: at a first worker node among a plurality of worker nodes, obtain a first set of global gradients for a first set of network layers in the neural network, the first set of global gradients being aggregated from local gradients determined by the plurality of worker nodes for the first set of network layers in the current training step, and the plurality of worker nodes being configured to jointly train the neural network. The second acquisition module is configured to: obtain a second set of global gradients for a second set of network layers in the neural network, the second set of network layers being different from the first set of network layers, and the second set of global gradients being aggregated from local gradients determined by the plurality of worker nodes for the second set of network layers in a previous training step before the current training step. The update module is configured to: based on the first set of global gradients and the second set of global gradients, update the parameters of the neural network.

[0006] According to a third aspect of the present disclosure, there is provided an electronic device. The electronic device includes at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to the first aspect of the present disclosure.

[0007] According to a fourth aspect of the present disclosure, there is provided a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to perform the method according to the first aspect of the present disclosure.

[0008] It should be understood that the content described in the summary of the invention section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In the following, in combination with the drawings and with reference to the following detailed description, the above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A flowchart showing a method for training a neural network according to some embodiments of the present disclosure;

[0012] Figure 3 A neural network training pipeline showing according to some embodiments of the present disclosure;

[0013] Figure 4 A block diagram showing an example device for training a neural network according to some embodiments of the present disclosure; and

[0014] Figure 5 A block diagram showing a computing device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0017] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.

[0018] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner according to the relevant laws and regulations.

[0019] For example, when receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0020] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0022] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of the subsequent action to be performed in response to the event or condition may not necessarily be strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, the subsequent action may be performed immediately when the event occurs or the condition is met; while in other cases, the subsequent action may be performed after a period of time after the event occurs or the condition is met.

[0023] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this document, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably herein.

[0024] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and it generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications usually include many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer. In the context of the present disclosure, the input layer, the output layer, and the hidden layer can also be referred to as network layers individually or collectively.

[0025] Generally, machine learning can roughly include three stages, namely, a training stage, a testing stage, and an application stage (also called an inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously iteratively updated until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also called the input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to test whether the model can provide correct outputs, so as to determine the performance of the model. In the application stage, the model can be used to process actual inputs based on the parameter values obtained from training and determine the corresponding outputs.

[0026] Figure 1FIG. 0 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Environment 100 relates to a neural network training environment based on data parallelism, which includes N worker nodes 110-1, 110-2, …… 110-N (where N is an integer greater than 1) and a service node 120. The worker nodes 110-1, 110-2 …… 110-N can respectively maintain their own local training data sets 112-1, 112-2, …… 112-N, and are configured to jointly train a neural network. For ease of discussion, the worker nodes 110-1, 110-2, …… 110-N can be individually or collectively referred to as worker node 110, and the local training data sets 112-1, 112-2, …… 112-N can be individually or collectively referred to as local training data set 112.

[0027] Data Parallelism is an acceleration method for neural network training. In data parallel neural network training, the training task is split among multiple worker nodes 110, and each worker node 110 maintains the same model parameters and the same computational task, but processes different training data. The service node 120 can aggregate the gradient data at each worker node 110 and synchronize the aggregated gradient data to each worker node 110. In this way, the training and computation under the same global training data are split among different worker nodes 110, thus reducing the computational and storage pressure at a single worker node 110. Worker nodes are sometimes also referred to as clients, worker node devices, clients, terminal nodes, terminal devices, edge devices, etc.

[0028] In data parallel neural network training, each worker node 110 stores its own local neural network 132, such as local neural networks 132-1, 132-2, …… 132-N (for ease of discussion, hereinafter referred to as local neural network 132 individually or collectively). The worker node 110 performs local training using its corresponding local training data set 112. The worker node 110 sends the gradients for each network layer determined during the local training process to the service node 120. The service node 120 obtains the global gradient by aggregating the gradients from each worker node 110, and synchronizes the aggregated global gradient to each worker node 110. The worker node 110 can update the local neural network 132 based on the synchronized global gradient to perform the next training step. The above process can be repeatedly executed until the neural network training converges. Since the corresponding neural networks finally trained at each worker node 110 are the same, for ease of description, the local neural network can also be simply referred to as the neural network hereinafter.

[0029] In some embodiments, the worker node 110 and / or the service node 120 may be implemented at a terminal device or a server. The terminal device may be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device is also capable of supporting any type of user interface (such as a "wearable" circuit, etc.). The server is various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and the like.

[0030] It should be understood that Figure 1 only an exemplary data-parallel based neural network training environment is shown. According to the needs of neural network training and actual applications, the environment may also be different. For example, although the service node 120 is Figure 1 shown as a separate node in, in some applications, in addition to acting as a central node, the service node 120 may also act as a worker node 110 for local training and the like. The scope of the present disclosure is not limited in this regard.

[0031] As described above, although the data-parallel based neural network training method can greatly accelerate the training speed, due to the frequent communication interactions between multiple worker nodes 110, it is difficult for this method to achieve a linear acceleration effect. Currently, existing methods update all network layers of the neural network by using the gradients determined in the previous training step in the current training step to eliminate the dependency between the communication process of the gradients and the calculation process of the current training step, and hide the communication overhead by constructing a communication and calculation parallel pipeline. However, introducing gradient delay for all network layers of the neural network will have an adverse impact on the convergence and accuracy of the model.

[0032] The inventors have found through research that due to the hierarchical structure of the neural network and the opposite order of forward propagation and backward propagation, it is only necessary to introduce gradients for some network layers in the neural network rather than all network layers to hide the communication overhead. In this regard, embodiments of the present disclosure propose an improved neural network training scheme based on gradient delay. Specifically, according to embodiments of the present disclosure, the network layers in the neural network are divided into a first group of network layers and a second group of network layers, and a first group of gradients for the first group of network layers and a second group of gradients for the second group of network layers are respectively obtained, where the first group of gradients is aggregated from the local gradients determined by multiple worker nodes 110 for the first group of network layers in the current training step, and the second group of gradients is aggregated from the local gradients determined by multiple worker nodes 110 for the second group of network layers in the previous training step. Further, the parameters of the neural network are updated based on the first group of gradients and the second group of gradients.

[0033] As will be more clearly understood from the following description, according to embodiments of the present disclosure, the communication overhead is hidden by introducing gradient delay only for some network layers in the neural network. In this way, the impact of the introduced gradient delay on model convergence and accuracy can be minimized as much as possible, so as to ensure model convergence and accuracy while hiding the communication overhead.

[0034] Some exemplary embodiments of the present disclosure will be described below with continued reference to the drawings.

[0035] Partial delay

[0036] Figure 2 The flowchart of a method 200 for training a neural network according to some embodiments of the present disclosure is shown. In some embodiments, the method 200 may be executed by a worker node 110 as shown in Figure 1 It should be understood that the method 200 may further include additional blocks not shown and / or some (or certain) of the shown blocks may be omitted, and the scope of the present disclosure is not limited in this regard.

[0037] At block 202, worker node 110 obtains a first set of global gradients for a first set of network layers in a neural network. The first set of global gradients is aggregated from local gradients determined by multiple worker nodes 110 for the first set of network layers in the current training step, that is, the global gradients for the first set of network layers are not delayed. At block 204, worker node 110 obtains a second set of global gradients for a second set of network layers in the neural network. The second set of network layers is different from the first set of network layers. The second set of global gradients is aggregated from local gradients determined by multiple worker nodes 110 for the second set of network layers in a previous training step before the current training step, that is, the gradients for the second set of network layers are delayed. In some embodiments, the previous training step may be the previous training step before the current training step. It should be understood that the previous training step may also include any other suitable training steps before the current training step, and the scope of the present disclosure is not limited in this regard.

[0038] For ease of explanation, hereinafter reference will be made to Figure 3 for description. Figure 3 FIG. 300 shows a neural network training pipeline 300 according to some embodiments of the present disclosure. The pipeline 300 is illustrated by taking a neural network with 3 network layers as an example. It should be understood that the solution according to the embodiments of the present disclosure can be applied to neural networks with any other suitable number of network layers, and the scope of the present disclosure is not limited in this regard.

[0039] In Figure 3 's example, the first set of network layers includes the output layer in the neural network, that is, the third network layer. The second set of network layers includes the first two network layers starting from the input layer in the neural network, that is, the first network layer and the second network layer. As Figure 3 shown, blocks 302-1, 302-2, and 302-3 schematically show the backpropagation calculation processes of the first network layer, the second network layer, and the third network layer in the neural network in the current training step. Note that since the backpropagation calculation process propagates backward from the output layer to the input layer, that is, first calculates the gradients for the third network layer, then calculates the gradients for the second network layer, and finally calculates the gradients for the first network layer, so in Figure 3The middle frame 302-3 is shown before the frame 302-2, and the frame 302-2 is shown before the frame 302-1. The frames 304-1, 304-2, and 304-3 schematically show the forward propagation calculation processes of the first network layer, the second network layer, and the third network layer in the neural network in the next training step after the current training step. The frames 306-1, 306-2, and 306-3 schematically show the round-trip communication processes of the first network layer, the second network layer, and the third network layer in the neural network in the current training step. This round-trip communication process corresponds to the worker node 110 sending locally determined local gradients to the server node 120 and receiving global gradients aggregated from the local gradients determined by multiple worker nodes 110 for the corresponding network layers from the server node 120. For the purpose of simplification, the "round-trip communication process" may also be referred to as the "communication process" hereinafter.

[0040] As Figure 3 shown, since the gradients for the first group of network layers (i.e., the third network layer) are not delayed, the worker node 110 can obtain the gradients for the first group of network layers from the server node 120 via the communication process 306-3. Further, since the gradients for the second group of network layers (i.e., the first network layer and the second network layer) are delayed, the worker node 110 can retrieve from the buffer 310 the second group of global gradients for the second group of network layers that were stored in the buffer 310 in the previous training step before the current training step.

[0041] Returning to reference Figure 2 , at block 206, the worker node 110 updates the parameters of the neural network based on the first group of global gradients and the second group of global gradients. In some embodiments, the worker node 110 can directly use the obtained global gradients to update the corresponding layers in the neural network. Referring Figure 3 , the worker node 110 can use the global gradients for the first network layer and the second network layer retrieved from the buffer 310 to update the parameters of the first network layer and the second network layer, and this update process corresponds to Figure 3 block 308-1 in Figure 3 . The worker node 110 can then perform the forward propagation calculation for the first network layer and the second network layer in the next training step based on the updated parameters. Further, the worker node 110 can use the global gradients for the third network layer obtained via the communication process 306-3 to update the parameters of the third network layer, and this update process corresponds to

[0042] By introducing gradient delay only for some network layers in the neural network rather than all network layers to hide the communication overhead, the impact of the introduced gradient delay on the model convergence and accuracy can be minimized as much as possible, so as to ensure the model convergence and accuracy while hiding the communication overhead.

[0043] In some embodiments, the minimum number of network layers for gradient delay can also be determined so that the communication overhead is completely hidden, thus ensuring that the training pipeline does not stall, that is, there is no waiting time at the worker node 110. The inventors have found through research that for a hierarchical neural network, in a non-stalled training pipeline, if a forward network layer is delayed, all forward network layers before that forward network layer are delayed. Therefore, the minimum number of network layers for introducing gradient delay required to completely hide the communication overhead can be determined by solving the following optimization problem:

[0044] Minimize k

[0045] Constraints

[0046] where k represents the number of network layers for which gradient delay needs to be introduced, v i represents the round-trip communication time between the worker node 110 and the server node 120 for the i-th network layer, b i represents the backpropagation calculation time for the i-th network layer, u i represents the forward propagation calculation time for the i-th network layer, and m represents the total number of network layers in the neural network. The first constraint in this optimization problem is used to ensure that communication and computation are fully overlapped, that is, the communication overhead is completely hidden; and the second constraint is used to ensure that the undelayed gradients are synchronized before the start of the forward propagation calculation of the network layer to which they correspond.

[0047] In some embodiments, the forward propagation calculation time, the backward propagation calculation time, and the round-trip communication time for each network layer in the neural network can be determined by monitoring the durations of the forward propagation calculation process, the backward propagation calculation process, and the communication process of the corresponding network layer during the warm-up training phase of the neural network. The worker node 110 can then determine the number of network layers for gradient delay based on the determined forward propagation calculation time, backward propagation calculation time, and round-trip communication time. In one example, the worker node 110 can obtain the minimum number of network layers that need to introduce gradient delay by solving the above optimization problem, and determine the number of network layers for gradient delay based on this minimum number. For example, the number of network layers for gradient delay can be determined to be any suitable value equal to or greater than this minimum number and less than the total number of network layers in the neural network.

[0048] Further, the worker node 110 can determine a first set of network layers and a second set of network layers based on the determined number of network layers. In some embodiments, the worker node 110 can select the determined number of network layers starting from the input layer of the neural network as the second set of network layers, and determine the remaining network layers in the neural network as the first set of network layers. Referring to Figure 3 , when the determined number of network layers for gradient delay is 2, the worker node 110 can determine the first 2 network layers (i.e., the first network layer and the second network layer) of the neural network starting from the input layer as the second set of network layers, and determine the remaining network layers of the neural network (i.e., the third network layer) as the first set of network layers.

[0049] Delay compensation

[0050] In some embodiments, the worker node 110 can predict a third set of global gradients determined for the second set of network layers in the current training step based on the second set of global gradients for the second set of network layers. In one example, the inventors have found through research that the third set of global gradients can be predicted based on the Taylor expansion of the gradient function, as shown in the following equation (1):

[0051]

[0052] where \(g_t\) represents the predicted third set of global gradients, \(g_t'\) represents the predicted third set of global gradients, \(\lambda\) is a pre-determined hyperparameter, represents the transpose of the matrix \(g_t'\), \(x\) t-1 represents the parameters of the second set of network layers in the current training step, and \(x\) t′-1Represent the parameters of the second set of network layers in the previous training step. Further, the worker node 110 can use the determined third set of global gradients to update the parameters of the second set of network layers. In this way, the introduced gradient delay can be compensated by means of the predicted global gradients, and the impact of the introduced gradient delay on model convergence and accuracy can be further reduced.

[0053] In some embodiments, to compensate for the lag introduced by gradient delay, the gradient one step backward in prediction can be obtained by predicting the model parameters in the next training step. For this purpose, the worker node 110 can locally store two identical neural networks, where one neural network is for training purposes (hereinafter also referred to as the training neural network), and the other neural network is for predicting model parameters (hereinafter also referred to as the prediction neural network).

[0054] In one embodiment, the worker node 110 can obtain a first set of local gradients for the second set of network layers. The first set of local gradients was determined for the second set of network layers in a previous training step. The worker node 110 can predict the parameters of the neural network in the next training step after the current training step based on the first set of local gradients, and determine a second set of local gradients for the second set of network layers in the current training step based on the predicted parameters, for aggregating with the local gradients determined by the remaining worker nodes 110 among the multiple worker nodes 110 to obtain the global gradient for the second set of network layers.

[0055] Exemplarily, the worker node 110 can retrieve from the buffer 310 the first set of local gradients determined for the second set of network layers in the previous training step, and use the first set of local gradients to update the parameters of the prediction neural network by considering the parameters and learning rate of the neural network in the current training step to obtain the parameters of the prediction neural network in the next training step as the prediction of the parameters of the neural network in the next training step. Further, the worker node 110 can determine a second set of local gradients for the second set of network layers in the current training step based on the updated parameters of the prediction neural network. Note that the second set of local gradients is the gradient one step backward in prediction. The worker node 110 can then use the second set of local gradients to aggregate with the local gradients determined by the remaining worker nodes 110 among the multiple worker nodes 110 to obtain the global gradient for the second set of network layers for updating the parameters of the training neural network in the next training step. In this way, the parameters of the network layer introduced with gradient delay can be updated using the gradient one step backward in prediction, so that the introduced gradient delay can be offset, and further the impact of the introduced gradient delay on model convergence and accuracy can be reduced.

[0056] In another embodiment, the worker node 110 may predict the parameters of the second set of neural network layers in the next training step after the current training step based on the second set of global gradients, and determine a set of local gradients for the second set of network layers in the current training step based on the predicted parameters, so as to aggregate with the local gradients determined by the remaining worker nodes 110 among the multiple worker nodes 110 to obtain the global gradient for the second set of network layers.

[0057] Exemplarily, the worker node 110 may update the parameters of the prediction neural network by considering the parameters and learning rate of the neural network in the current training step using the second set of global gradients obtained from the buffer 310, so as to obtain the parameters of the prediction neural network in the next training step as the prediction of the parameters of the neural network in the next training step. Further, the worker node 110 may determine a set of local gradients for the second set of network layers in the current training step based on the updated parameters of the prediction neural network. It should be noted that this set of local gradients is the gradient predicted one step backward. The worker node 110 may then aggregate this set of local gradients with the local gradients determined by the remaining worker nodes 110 among the multiple worker nodes 110 to obtain the global gradient for the second set of network layers, for updating the parameters of the training neural network in the next training step. In this way, the gradient predicted one step backward can be used to update the parameters of the network layer with introduced gradient delay, so as to offset the introduced gradient delay, and further reduce the impact of the introduced gradient delay on the model convergence and accuracy.

[0058] In yet another embodiment, the worker node 110 may obtain a first set of local gradients for the second set of network layers. This first set of local gradients was determined for the second set of network layers in a previous training step. The worker node 110 may predict a third set of global gradients determined for the second set of network layers in the current training step based on the first set of local gradients and the second set of global gradients. Further, the worker node 110 may predict the parameters of the neural network in the next training step after the current training step based on the third set of global gradients, and determine a second set of local gradients for the second set of network layers in the current training step based on the predicted parameters, so as to aggregate with the local gradients determined by the remaining worker nodes 110 among the multiple worker nodes 110 to obtain the global gradient for the second set of network layers.

[0059] Exemplarily, the worker node 110 may retrieve the first set of local gradients determined for the second set of network layers in a previous training step from the buffer 310. The worker node 110 may use the first set of local gradients and the second set of global gradients to predict the third set of global gradients determined for the second set of network layers in the current training step based on the following formula (2):

[0060]

[0061] where represents the predicted third set of global gradients, g t-1 represents the second set of global gradients, g i,t-1 represents the local gradients used to aggregate the second set of global gradients, g i,t represents the first set of local gradients, n represents the total number of multiple worker nodes 110, x t-1 represents the parameters of the neural network in the current training step, x t-2 represents the parameters of the neural network in the previous training step, and the function DC λ () is defined in equation (1).

[0062] Further, the worker node 110 can update the parameters of the prediction neural network by considering the parameters and learning rate of the neural network in the current training step using the predicted third set of global gradients, so as to obtain the parameters of the prediction neural network in the next training step as a prediction of the parameters of the neural network in the next training step. Further, the worker node 110 can determine the second set of local gradients for the second set of network layers in the current training step based on the updated parameters of the prediction neural network. Note that the second set of local gradients is the gradient predicted one step backward. The worker node 110 can then use the second set of local gradients to aggregate with the local gradients determined by the remaining worker nodes 110 among the multiple worker nodes 110 to obtain the global gradients for the second set of network layers for updating the parameters of the training neural network in the next training step. In this way, the gradients predicted one step backward can be used to update the parameters of the network layers with introduced gradient delay, so that the introduced gradient delay can be offset, and further the influence of the introduced gradient delay on the model convergence and accuracy can be reduced.

[0063] Multiple possible ways for compensating gradient delay are described above. It should be understood that the introduced gradient delay can also be compensated in any other suitable way, and the scope of the present disclosure is not limited in this regard.

[0064] System design

[0065] By introducing gradient delay for the second set of network layers, the second set of global gradients for the second set of network layers can be stored in the buffer 310. This enables the second set of global gradients to be retrieved simultaneously. Therefore, in the case where the second set of network layers includes multiple network layers, the parameter updates for the multiple network layers can be performed simultaneously. In some embodiments, the worker node 110 can use the same kernel to update the parameters of the second set of network layers. In this way, compared with the existing solution that needs to call different kernels to update parameters for different network layers, system resources and time overhead can be saved, and further the training speed of the neural network can be improved.

[0066] In some embodiments, buffer 310 may include two different buffer areas, namely a first buffer area and a second buffer area. Worker node 110 may store a second set of global gradients in the first buffer area of buffer 310 in a previous training step, and directly read the second set of global gradients from the first buffer area in the current training step. Further, in the current training step, worker node 110 obtains a third set of global gradients for the second set of network layers through a communication process. The third set of global gradients is aggregated from the local gradients determined by multiple worker nodes 110 for the second set of network layers in the current training step. Worker node 110 may then store the third set of global gradients in the second buffer area of buffer 310, and directly read the third set of global gradients from the second buffer area in the next training step. In this way, compared with existing solutions, the system overhead caused by copying gradient data from one buffer area in the buffer to another can be avoided, thereby improving the training speed of the neural network.

[0067] In some embodiments, in a previous training step, worker node 110 may store a second set of global gradients in buffer 310, and lock buffer 310 after the second set of global gradients is stored in buffer 310. Further, when the second set of global gradients needs to be obtained in the current training step, worker node 110 may unlock buffer 310 and read the second set of global gradients from buffer 310. Subsequently, worker node 110 may store a third set of global gradients for the second set of network layers obtained in the current training step in buffer 310. In this way, sharing of buffer 310 can be achieved, thereby reducing the overhead of buffer 310 on memory resources.

[0068] In some embodiments, worker node 110 may include two different processing units, namely a first processing unit and a second processing unit, where the first processing unit is configured to implement the training of the neural network, and the second processing unit is configured to control the communication at worker node 110. Worker node 110 may determine the total computation time and total communication time for the neural network in a training step based on warm-up training of the neural network, and select a buffer from a first buffer associated with the first processing unit and a second buffer associated with the second processing unit based on a comparison of the total computation time and the total communication time. Further, worker node 110 may store the second set of global gradients in the selected buffer.

[0069] Exemplarily, the working node 110 may include a Central Processing Unit (CPU) and a Graphics Processing Unit (GPU). The CPU may be used to control the communication at the working node 110, and the GPU may be used to train the neural network. If it is determined that the total computation time in a training step of the neural network is greater than the total communication latency, a buffer associated with the GPU may be selected to store the second set of global gradients. If it is determined that the total computation time in a training step of the neural network is less than the total communication latency, a buffer associated with the CPU may be selected to store the second set of global gradients. In this way, the transmission overhead can be shifted from the relatively slower pipeline stage to the relatively faster pipeline stage, thereby balancing the neural network training pipeline and further improving the training speed of the neural network.

[0070] As described above in combination with Figures 1 to 3 it can be seen that, for the method for training a neural network according to an embodiment of the present disclosure, the communication overhead is hidden by introducing gradient latency only for some of the network layers in the neural network. In this way, the impact of the introduced gradient latency on the model convergence and accuracy can be minimized as much as possible, so as to ensure the model convergence and accuracy while hiding the communication overhead.

[0071] In the foregoing, examples of implementations of the method according to the present disclosure have been described in detail. In the following, implementations of the corresponding apparatuses and devices will be described with reference to Figures 1 to 3 and Figure 4 and Figure 5 describe the implementations of the corresponding apparatuses and devices.

[0072] Example devices and equipment

[0073] Figure 4 FIG. shows a block diagram of an exemplary apparatus 400 for training a neural network according to some embodiments of the present disclosure. The apparatus 400 may be used, for example, to implement as Figure 1The working node 110 shown. The apparatus 400 may include a first acquisition module 402 configured to: at a first working node among a plurality of working nodes, acquire a first set of global gradients for a first set of network layers in a neural network, the first set of global gradients being aggregated from local gradients determined by the plurality of working nodes for the first set of network layers in a current training step, the plurality of working nodes being configured to jointly train the neural network. Further, the apparatus 400 may include a second acquisition module 404 configured to: acquire a second set of global gradients for a second set of network layers in the neural network, the second set of network layers being different from the first set of network layers, the second set of global gradients being aggregated from local gradients determined by the plurality of working nodes for the second set of network layers in a previous training step before the current training step. Additionally, the apparatus 400 may further include an update module 406 configured to: update parameters of the neural network based on the first set of global gradients and the second set of global gradients.

[0074] In some embodiments, the first set of global gradients and the second set of global gradients are aggregated at a service node, and the apparatus 400 further includes: a time determination module configured to: based on warm-up training of the neural network, determine a forward propagation calculation time, a backward propagation calculation time, and a round-trip communication time between the first working node and the service node for each network layer in the neural network; a number determination module configured to: based on the determined forward propagation calculation time, backward propagation calculation time, and round-trip communication time, determine a number of network layers for gradient delay; and a network layer determination module configured to: based on the determined number of network layers, determine the first set of network layers and the second set of network layers.

[0075] In some embodiments, the update module 406 includes: a first parameter update module configured to: update parameters of the first set of network layers using the first set of global gradients; a gradient prediction module configured to: predict a third set of global gradients determined for the second set of network layers in the current training step based on the second set of global gradients; and a second parameter update module configured to: update parameters of the second set of network layers using the third set of global gradients.

[0076] In some embodiments, the apparatus 400 further includes: a first gradient acquisition module configured to: acquire a first set of local gradients for the second set of network layers, the first set of local gradients being determined for the second set of network layers in a previous training step; a first parameter prediction module configured to: predict parameters of the neural network in a next training step after the current training step based on the first set of local gradients; and a first gradient determination module configured to: based on the predicted parameters, determine a second set of local gradients for the second set of network layers in the current training step for aggregation with local gradients determined by the remaining working nodes among the plurality of working nodes to obtain global gradients for the second set of network layers.

[0077] In some embodiments, the apparatus 400 further includes: a second parameter prediction module configured to predict, based on a second set of global gradients, the parameters of the neural network in the next training step after the current training step; and a second gradient determination module configured to determine, based on the predicted parameters, a set of local gradients for a second set of network layers in the current training step for aggregating with local gradients determined by the remaining worker nodes among the plurality of worker nodes to obtain a global gradient for the second set of network layers.

[0078] In some embodiments, the apparatus 400 further includes: a second gradient acquisition module configured to acquire a first set of local gradients for a second set of network layers, the first set of local gradients being determined for the second set of network layers in a previous training step; a global gradient prediction module configured to predict, based on the first set of local gradients and the second set of global gradients, a third set of global gradients determined for the second set of network layers in the current training step; a third parameter prediction module configured to predict, based on the third set of global gradients, the parameters of the neural network in the next training step after the current training step; and a third gradient determination module configured to determine, based on the predicted parameters, a second set of local gradients for the second set of network layers in the current training step for aggregating with local gradients determined by the remaining worker nodes among the plurality of worker nodes to obtain a global gradient for the second set of network layers.

[0079] The modules and / or units included in the apparatus 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to the machine-executable instructions, some or all of the units in the apparatus 400 can be at least partially implemented by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0080] Figure 4 These modules and / or units shown in can be partially or fully implemented as hardware modules, software modules, firmware modules, or any combination thereof. In particular, in certain embodiments, the processes, methods, or procedures described above can be implemented by hardware in a storage system or a host corresponding to the storage system or other computing devices independent of the storage system.

[0081] Figure 5 A block diagram of a computing device 500 is shown in which one or more embodiments of the present disclosure can be implemented. It should be understood that Figure 5The computing device 500 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 5 The computing device 500 shown can be used to implement Figure 1 the worker node 110.

[0082] As Figure 5 shown, the computing device 500 is in the form of a general-purpose computing device. The components of the computing device 500 can include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 500.

[0083] The computing device 500 generally includes multiple computer storage media. Such media can be any available media accessible to the computing device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, magnetic disks, or any other medium that can be used to store information and / or data (such as training data for training) and can be accessed within the computing device 500.

[0084] The computing device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules that are configured to execute the various methods or actions of the various embodiments of the present disclosure.

[0085] The communication unit 540 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 500 may be implemented in a single computing cluster or multiple computing machines that are capable of communicating via a communication link. Thus, the computing device 500 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.

[0086] The input device 550 may be one or more input devices, such as a mouse, keyboard, trackball, etc. The output device 560 may be one or more output devices, such as a display, speakers, printer, etc. The computing device 500 may also communicate, as needed, with one or more external devices (not shown) via the communication unit 540, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 500, or communicate with any device that enables the computing device 500 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0087] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, where the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the methods described above.

[0088] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0089] These computer-readable program instructions may be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when the instructions are executed by the processing unit of the computer or other programmable data processing apparatus, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which instructions cause a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0090] Computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0091] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0092] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled artisans in the art to understand the various implementations disclosed herein.

Claims

1. A method for training a neural network, comprising: at a first worker node among a plurality of worker nodes, obtaining a first set of global gradients for a first set of network layers in the neural network, the first set of global gradients being aggregated from local gradients determined for the first set of network layers by the plurality of worker nodes in a current training step, the plurality of worker nodes being configured to jointly train the neural network; obtaining a second set of global gradients for a second set of network layers in the neural network, the second set of network layers being different from the first set of network layers, the second set of global gradients being aggregated from local gradients determined for the second set of network layers by the plurality of worker nodes in a previous training step before the current training step; and updating parameters of the neural network based on the first set of global gradients and the second set of global gradients; wherein the first set of global gradients and the second set of global gradients are aggregated at a service node, and the method further comprises: based on warm-up training of the neural network, determining a forward propagation calculation time, a backward propagation calculation time, and a round-trip communication time between the first worker node and the service node for each network layer in the neural network; determining the number of network layers for gradient delay based on the determined forward propagation calculation time, the backward propagation calculation time, and the round-trip communication time; and determining the first set of network layers and the second set of network layers based on the determined number of network layers.

2. The method according to claim 1, wherein determining the first set of network layers and the second set of network layers comprises: starting from the input layer of the neural network, selecting the determined number of network layers as the second set of network layers; and determining the remaining network layers in the neural network as the first set of network layers.

3. The method according to claim 1, wherein updating parameters of the neural network comprises: updating parameters of the first set of network layers using the first set of global gradients; predicting a third set of global gradients determined for the second set of network layers in the current training step based on the second set of global gradients; and updating parameters of the second set of network layers using the third set of global gradients.

4. The method according to claim 1, further comprising: obtaining a first set of local gradients for the second set of network layers, the first set of local gradients being determined for the second set of network layers in the previous training step; predicting parameters of the neural network in a next training step after the current training step based on the first set of local gradients; and determining a second set of local gradients for the second set of network layers in the current training step based on the predicted parameters, for aggregating with local gradients determined by the remaining worker nodes among the plurality of worker nodes to obtain global gradients for the second set of network layers.

5. The method according to claim 1, further comprising: predicting parameters of the neural network in a next training step after the current training step based on the second set of global gradients; and Based on the predicted parameters, determine a set of local gradients for the second set of network layers in the current training step for aggregating with local gradients determined by the remaining worker nodes among the multiple worker nodes to obtain a global gradient for the second set of network layers.

6. The method according to claim 1, further comprising: Obtain a first set of local gradients for the second set of network layers, the first set of local gradients being determined for the second set of network layers in the previous training step; Based on the first set of local gradients and the second set of global gradients, predict a third set of global gradients to be determined for the second set of network layers in the current training step; Based on the third set of global gradients, predict the parameters of the neural network in the next training step after the current training step; and Based on the predicted parameters, determine a second set of local gradients for the second set of network layers in the current training step for aggregating with local gradients determined by the remaining worker nodes among the multiple worker nodes to obtain a global gradient for the second set of network layers.

7. The method according to any one of claims 1 to 6, wherein the parameters of the second set of network layers are updated by using the same kernel.

8. The method according to any one of claims 1 to 6, wherein the second set of global gradients is stored in a first buffer area of a buffer in the previous training step, and wherein obtaining the second set of global gradients comprises: Reading the second set of global gradients from the first buffer area.

9. The method according to claim 8, further comprising: Obtain a third set of global gradients for the second set of network layers in the neural network, the third set of global gradients being aggregated from local gradients determined by the multiple worker nodes for the second set of network layers in the current training step; and Store the third set of global gradients in a second buffer area of the buffer, the second buffer area being different from the first buffer area.

10. The method according to any one of claims 1 to 6, wherein the second set of global gradients is stored in a buffer in the previous training step, the method further comprising: In response to the second set of global gradients being stored in the buffer, lock the buffer, and wherein obtaining the second set of global gradients comprises: Unlock the buffer in the current training step; and Read the second set of global gradients from the buffer.

11. The method according to any one of claims 1 to 6, wherein the training of the neural network is implemented on a first processing unit, the method further comprising: Based on the warm-up training of the neural network, determine the total computation time and the total communication time in the training step of the neural network; Based on the comparison between the total computation time and the total communication time, select a buffer from a first buffer associated with the first processing unit and a second buffer associated with a second processing unit, the second processing unit being configured to control the communication at the first worker node; and Store the second set of global gradients in the selected buffer.

12. An apparatus for training a neural network, comprising: A first acquisition module, configured to: at a first worker node among a plurality of worker nodes, acquire a first set of global gradients for a first set of network layers in the neural network, where the first set of global gradients is obtained by aggregating local gradients determined by the plurality of worker nodes for the first set of network layers in the current training step, and the plurality of worker nodes are configured to jointly train the neural network; A second acquisition module, configured to: acquire a second set of global gradients for a second set of network layers in the neural network, where the second set of network layers is different from the first set of network layers, and the second set of global gradients is obtained by aggregating local gradients determined by the plurality of worker nodes for the second set of network layers in a previous training step before the current training step; And An update module, configured to: update parameters of the neural network based on the first set of global gradients and the second set of global gradients; Where the first set of global gradients and the second set of global gradients are aggregated at a service node, and the apparatus further comprises: A time determination module, configured to: based on warm-up training of the neural network, determine the forward propagation calculation time, the backward propagation calculation time, and the round-trip communication time between the first worker node and the service node for each network layer in the neural network; A number determination module, configured to: based on the determined forward propagation calculation time, the backward propagation calculation time, and the round-trip communication time, determine the number of network layers for gradient delay; and A network layer determination module, configured to: based on the determined number of network layers, determine the first set of network layers and the second set of network layers.

13. The apparatus according to claim 12, wherein the update module comprises: A first parameter update module, configured to: update parameters of the first set of network layers using the first set of global gradients; A gradient prediction module, configured to: based on the second set of global gradients, predict a third set of global gradients determined for the second set of network layers in the current training step; and A second parameter update module, configured to: update parameters of the second set of network layers using the third set of global gradients.

14. The apparatus according to claim 12, further comprising: A first gradient acquisition module, configured to: acquire a first set of local gradients for the second set of network layers, where the first set of local gradients is determined for the second set of network layers in the previous training step; A first parameter prediction module, configured to: based on the first set of local gradients, predict parameters of the neural network in the next training step after the current training step; And A first gradient determination module, configured to: based on the predicted parameters, determine a second set of local gradients for the second set of network layers in the current training step for aggregating with local gradients determined by the remaining worker nodes among the plurality of worker nodes to obtain global gradients for the second set of network layers.

15. The apparatus according to claim 12, further comprising: A second parameter prediction module, configured to: predict parameters of the neural network in a next training step after the current training step based on the second set of global gradients; and A second gradient determination module, configured to: determine a set of local gradients for the second set of network layers in the current training step based on the predicted parameters, for aggregating with local gradients determined by the remaining worker nodes among the multiple worker nodes to obtain a global gradient for the second set of network layers.

16. The apparatus according to claim 12, further comprising: A second gradient acquisition module, configured to: acquire a first set of local gradients for the second set of network layers, where the first set of local gradients are determined for the second set of network layers in the previous training step; A global gradient prediction module, configured to: predict a third set of global gradients determined for the second set of network layers in the current training step based on the first set of local gradients and the second set of global gradients; A third parameter prediction module, configured to: predict parameters of the neural network in a next training step after the current training step based on the third set of global gradients; and A third gradient determination module, configured to: determine a second set of local gradients for the second set of network layers in the current training step based on the predicted parameters, for aggregating with local gradients determined by the remaining worker nodes among the multiple worker nodes to obtain a global gradient for the second set of network layers.

17. An electronic device, comprising: At least one processing unit; and At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the device to execute the method according to any one of claims 1 to 11.

18. A computer-readable storage medium, having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Parameter updating method based on neural network, related platform and computer storage medium

    CN108960410A

  • Neural network model training method and related product

    CN111723933A