A Distributed Deep Learning Model Update Method and Device
By obtaining and calculating the true gradient value of the distributed deep learning model to predict the gradient value update parameters, the problem of low training efficiency in the existing technology is solved, and more efficient model training and accuracy improvement is achieved, which is suitable for traffic target recognition.
Patent Information
- Application Number
- CN202310129775.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-04
- Filing Date
- 2023-02-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-02-17
AI Technical Summary
The existing distributed deep learning model update method is difficult to effectively improve training efficiency while ensuring training accuracy. Especially in areas with high real-time demand such as traffic target recognition, asynchronous update causes parameter inconsistency to affect convergence speed and accuracy, and synchronous updates increase communication overhead.
By obtaining the true gradient value of the distributed deep learning model, using formulas to calculate the predicted gradient value to update the model parameters, reduce the iterative training time, and update the parameters without waiting for all nodes to complete the calculation.
It improves the efficiency and weak scalability of model training, ensures the accuracy of traffic target recognition, and reduces additional communication requirements, achieving more efficient model training.
Smart Images

Figure CN116306913B_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202210929897.0 and the invention title "A Method and Device for Updating a Distributed Deep Learning Model" submitted to the Chinese Patent Office on August 4, 2022, the entire content of which is incorporated herein by reference. Technical Field
[0002] The present invention relates to the technical field of distributed neural networks, and particularly to a method and device for updating a distributed deep learning model. Background Art
[0003] With the continuous development of deep learning networks, while the accuracy is continuously improved, the model and dataset become larger and larger, and the required training time is also increasing. Therefore, how to accelerate model training while ensuring training accuracy has become a research direction worthy of attention, especially in some fields with high real-time requirements and accuracy requirements, such as traffic target recognition. In traffic target recognition, it is necessary to identify vehicles, pedestrians, traffic lights, obstacles, etc. in real time. The distributed method can enable the deep learning network to speed up this recognition process. And in some scenarios, the deep learning model needs to be updated in real time, which poses a great real-time challenge to the distributed update of the model. For static models, excellent distributed update methods can improve the training efficiency of the model and reduce costs.
[0004] Currently, the update method is an important research direction in distributed deep learning. There are mainly two data parallel update methods: synchronous update and asynchronous update. In synchronous update, after each node calculates the gradient, it must wait for other nodes to complete the calculation. Then, the parameter server collects the gradients from each node and downloads the new gradients to them. After each node uniformly updates the parameters, it then performs the next iteration. Different from synchronous update, asynchronous update does not require all nodes to be consistent. Each node calculates the gradient and directly uploads it to the parameter server. Then, the node downloads the updated parameters from the parameter server, without waiting for other nodes and starts the next iteration. Asynchronous update can potentially achieve higher throughput in a distributed system: the nodes can spend more time performing useful calculations instead of waiting for gradient collection. However, another problem is caused by the asynchronous update of the parameter vector: when the node completes these calculations and applies the results to the parameter server, the parameters may have been updated several times by the gradients of other nodes. Therefore, the gradients submitted by the node are inconsistent with the model parameters, which will have a negative impact on the convergence speed and final accuracy.
[0005] Due to these shortcomings, there are some alternative methods that trade off convergence and utilization. Some methods use a mixture of synchronous and asynchronous methods to balance performance and efficiency. Some work uses fine-grained synchronous updates through multi-layer communication, which is more efficient than traditional synchronous updates while ensuring convergence, but this adds multiple additional communications. These additional communication overheads are very unfavorable for large-scale distributed deep learning, because multiple communications are more likely to cause communication congestion and anomalies.
[0006] Optimization algorithms are another research focus of distributed deep learning. In order to handle large data sets, synchronous SGD performs parallel processing between multiple workers with parameter servers. The convergence of synchronous SGD is the same as that of mini-batch SGD. In addition, while synchronous SGD ensures convergence, it will cause idle computing nodes. To solve this problem, some studies have also proposed an asynchronous SGD algorithm. The recently proposed asynchronous SGD method (Asynchronous stochastic gradient descent with delay compensation) has good performance, but performs poorly in large-scale distributed deep learning.
[0007] The prior art discloses an adaptive neural network training method, electronic device, medium and program product, wherein the method includes: training a target neural network based on the adaptive parameters of the current training round, wherein the adaptive parameters are used to determine the training task amount of each training node for training the target neural network; adjusting the adaptive model parameters of the current training round using the historical gradient value or the current training round gradient value based on the training time of each training node in the current training round, and determining the adjusted adaptive model parameters as the adaptive model parameters of the next training round. The prior art discloses adjusting the adaptive model parameters of the current training round and determining the adjusted adaptive model parameters as the adaptive model parameters of the next training round, but does not disclose how to specifically adjust the model parameters by combining the historical round gradient values to predict the current round gradient values, so as to ensure the model accuracy while improving the model training efficiency, which makes researchers urgently need an update method and device for accelerating the training of distributed deep learning models. Summary of the invention
[0008] Based on this, the purpose of the present invention is to provide a distributed deep learning model updating method and device that can be applied to a traffic target recognition model, which can ensure the accuracy of traffic target recognition while improving the training efficiency of the model.
[0009] To achieve the above object, the present invention provides a distributed deep learning model updating method, comprising:
[0010] Obtain the true gradient value g of the distributed deep learning model t ;
[0011] According to the true gradient value g t and the following formula, calculate the predicted gradient value
[0012] D t =(1 - β t )D t-1 + β t g t-1
[0013]
[0014]
[0015] where D t represents the smoothed gradient exponent of the t-th iteration, β t represents the first penalty coefficient, D t-1 represents the smoothed gradient exponent of the (t - 1)-th iteration, g t-1 represents the true gradient value of the (t - 1)-th iteration, v t represents the smoothed gradient exponent D of the t-th iteration t and the predicted error between the true gradient value g t , α represents the second penalty coefficient, represents the error between the true gradient value and the predicted gradient value in the (t - 1)-th iteration, represents the predicted gradient value of the t-th iteration;
[0016] Use the predicted gradient value to update the model parameters of the distributed deep learning model.
[0017] Optionally, after using the predicted gradient value to update the parameters of the distributed deep learning model, it further includes:
[0018] Judge whether the number of iterations reaches the preset number;
[0019] If so, complete the training of the distributed deep learning model;
[0020] If not, start the next iteration process.
[0021] Optionally, the obtaining of the true gradient value of the distributed deep learning model includes:
[0022] Obtain the model parameters of the distributed deep learning model;
[0023] According to the following formula and the model parameters, calculate the true gradient value:
[0024]
[0025] Among them, g t represents the true gradient value at the t-th iteration, and θ t represents the model parameter at the t-th iteration, and F θ (o(t,i)) represents the loss function with respect to the sample o(t,i).
[0026] Optionally, obtaining the true gradient value of the distributed deep learning model includes:
[0027] Obtaining the true gradient value g t-1 .
[0028] Optionally, before obtaining the predicted gradient value, it further includes:
[0029] Calculating the first penalty coefficient β using the following formula t :
[0030] e t-1 = g t-1 - D t-1
[0031]
[0032] Among them, e t-1 represents the true error between the smoothed gradient exponent D t-1 at the (t - 1)-th iteration and the true gradient value g t-1 , β t represents the first penalty coefficient, and k represents a hyperparameter.
[0033] Optionally, updating the model parameter of the distributed deep learning model using the predicted gradient value includes:
[0034] Calculating the model parameter of the distributed deep learning model according to the following formula and the predicted gradient value:
[0035]
[0036] Among them, θ t+1 represents the updated model parameter, θ t represents the model parameter at the t-th iteration, ∈ t represents the learning rate, represents the predicted gradient value at the t-th iteration.
[0037] The present invention also provides a distributed deep learning model updating device, including:
[0038] Gradient acquisition module, configured to acquire the true gradient value of a distributed deep learning model;
[0039] Gradient prediction module, configured to calculate a predicted gradient value according to the true gradient value and the following formula:
[0040] D t = (1 - β t )D t-1 + β t g t-1
[0041]
[0042]
[0043] where D t represents the smoothed gradient exponent of the t-th iteration, β t represents the first penalty coefficient, D t-1 represents the smoothed gradient exponent of the (t - 1)-th iteration, g t-1 represents the true gradient value of the (t - 1)-th iteration, v t represents the predicted error between the smoothed gradient exponent D t of the t-th iteration and the true gradient value g t , α represents the second penalty coefficient, represents the error between the true gradient value and the predicted gradient value in the (t - 1)-th iteration, represents the predicted gradient value of the t-th iteration;
[0044] Parameter update module, configured to update the model parameters of the distributed deep learning model by using the predicted gradient value.
[0045] Optionally, the apparatus further includes:
[0046] Judgment cut-off unit, configured to judge whether the number of iterations meets a preset number of iterations. If so, the training of the distributed deep learning model is completed; if not, the next iteration process is started, and the gradient acquisition module is returned to execute.
[0047] Optionally, the gradient acquisition module includes:
[0048] First gradient acquisition sub-module, configured to acquire the model parameters of the distributed deep learning model;
[0049] Second gradient acquisition sub-module, configured to calculate the true gradient value according to the following formula and the model parameters:
[0050]
[0051] where g tdenotes the true gradient value at the t-th iteration, θ t denotes the model parameters at the t-th iteration, F θ (o(t,i)) represents the loss function with respect to the sample o(t,i).
[0052] The present invention also provides an application method of a distributed deep learning model updating method, including:
[0053] Obtain a training sample set of traffic targets; the traffic targets include pedestrians, vehicles, and traffic lights;
[0054] Input the training sample set into a pre-selected and constructed distributed deep learning model for model training, and adopt the above-mentioned distributed deep learning model updating method during the training process;
[0055] Identify traffic targets through the trained distributed deep learning model.
[0056] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0057] The present invention obtains the true gradient value of the distributed deep learning model, calculates the predicted gradient value from the true gradient value, updates the model parameters of the distributed deep learning model through the predicted gradient value, reduces the training time of the model by predicting the gradient value, and the predicted gradient value is equivalent to the standard synchronous stochastic gradient descent (SGD), without waiting for all nodes to complete the calculation, that is, without additional communication, further improving the efficiency and weak scalability of model training, thereby being able to improve the accuracy of traffic target recognition.
[0058] Drawings of the specification
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0060] Figure 1 is the flowchart of the distributed deep learning model updating method provided in Embodiment 1 of the present invention;
[0061] Figure 2 is the flowchart of the distributed deep learning model updating method provided in Embodiment 2 of the present invention;
[0062] Figure 3 is the schematic diagram of the distributed deep learning model updating device provided in Embodiment 5 of the present invention;
[0063] Figure 4 It is a schematic diagram of the code for updating a distributed deep learning model provided by the present invention;
[0064] Figure 5 It is an application flowchart of the distributed deep learning model update method provided in the sixth embodiment of the present invention. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0066] The purpose of the present invention is to provide a distributed deep learning model update method and device that can be applied to a traffic target recognition model, which can improve the training efficiency of the model and thus improve the accuracy of traffic target recognition.
[0067] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0068] Embodiment 1
[0069] Figure 1 It is a flowchart of the distributed deep learning model update method provided in Embodiment 1 of the present invention, and the method may include the following steps:
[0070] S100: Obtain the true gradient value g of the distributed deep learning model t .
[0071] Specifically, the distributed deep learning model may be a deep learning model trained distributively. Deep learning may be a data-driven learning algorithm. The distributed deep learning model may include any one or more of three types: model parallelism, pipeline parallelism, and data parallelism. The true gradient value may be the actual gradient between the current parameters of the distributed deep learning model and the parameters of the distributed deep learning model in the previous iteration, and the true gradient value of the distributed deep learning model may be obtained.
[0072] S110: Calculate the predicted gradient value according to the true gradient value g t and the formula
[0073] Specifically, the formula may be as follows:
[0074] D t =(1 - β t )D t-1 +βt g t-1
[0075]
[0076]
[0077] wherein, D t represents the smoothed gradient exponent of the t-th iteration, and β t represents the first penalty coefficient, and D t-1 represents the smoothed gradient exponent of the (t - 1)-th iteration, and g t-1 represents the true gradient value of the (t - 1)-th iteration, and v t represents the smoothed gradient exponent D of the t-th iteration t and the prediction error between the true gradient value g t where α represents the second penalty coefficient, represents the error between the true gradient value and the predicted gradient value in the (t - 1)-th iteration, represents the predicted gradient value of the t-th iteration;
[0078] S120: Update the model parameters of the distributed deep learning model by using the predicted gradient value Specifically, the predicted gradient value can be the gradient obtained through the above calculation, and the model parameters of the distributed deep learning model can be updated by the predicted gradient value.
[0079] As can be seen from the above technical solutions, in the method for updating a distributed deep learning model provided in the first embodiment of the present invention, by obtaining the true gradient value of the distributed deep learning model, calculating the predicted gradient value according to the true gradient value and the formula, and updating the model parameters of the distributed deep learning model by using the predicted gradient value, so the present invention reduces the training time of the continuously iterated model by the predicted gradient value, and the training result by the predicted gradient value is equivalent to the standard synchronous stochastic gradient descent SGD, further improving the efficiency and weak scalability of model training.
[0080] Embodiment 2
[0081] The difference between this embodiment and Embodiment 1 is that, on the basis of Embodiment 1, this embodiment further describes the method for updating a distributed deep learning model.
[0082] In some embodiments of the present application, as
[0083] shown, after S102, the following steps may further be included: Figure 2 S130: Determine whether the number of iterations meets a preset number of iterations.
[0084]
[0085] Specifically, the judgment condition can be whether the number of training iterations of the current predicted gradient value reaches a preset value. If so, the termination condition is met. For example, if the number of training iterations of the current predicted gradient value is 10, the termination condition of 10 iterations is met.
[0086] S140: If so, complete the training of the distributed deep learning model.
[0087] Specifically, if the number of iterations meets the preset number of times, the training can be terminated, that is, the training of the distributed deep learning model is completed.
[0088] If not, start the next iteration.
[0089] Specifically, if the number of iterations does not meet the preset number of times, it is necessary to continue the iteration and return to execute the step of obtaining the true gradient value of the distributed deep learning model.
[0090] The other structures of this embodiment are the same as those of Embodiment 1 and will not be elaborated here.
[0091] Embodiment 3
[0092] The difference between this embodiment and Embodiment 2 is that, on the basis of Embodiment 2, this embodiment further explains the method for updating the distributed deep learning model.
[0093] In some embodiments of the present invention, S100 is further introduced below. Specifically, it may include the following steps:
[0094] S101: Obtain the model parameters of the distributed deep learning model.
[0095] Specifically, the model parameters can be the model parameters of the deep learning model in the distributed computing nodes, and the model parameters of the distributed deep learning model can be obtained through the parameter server.
[0096] S102: Calculate the true gradient value according to the formula and the model parameters.
[0097] Specifically, the formula can be as follows:
[0098]
[0099] where g t represents the true gradient value of the t-th iteration, θ t represents the model parameters of the t-th iteration, and F θ (o(t,i)) represents the loss function with respect to the sample o(t,i).
[0100] Furthermore, when first using the method for updating the distributed deep learning model of the present invention, the following steps can be executed:
[0101] Obtain the true gradient value in the previous iteration of the distributed deep learning model.
[0102] Specifically, the model parameters of the first iteration and the second iteration of the distributed deep learning model can be obtained, and the first true gradient value can be obtained through the model parameters of the first iteration and the second iteration to improve the accuracy of model update prediction.
[0103] The other structures of this embodiment are the same as those of Embodiment 2 and will not be elaborated here.
[0104] Embodiment 4
[0105] The difference between this embodiment and Embodiment 3 is that, on the basis of Embodiment 2, this embodiment further explains the distributed deep learning model update method.
[0106] In some embodiments of the present application, in order to achieve accurate prediction in the later stage of the training process, the present invention also proposes a more fine-grained coefficient adjustment method. Before obtaining the predicted gradient value, the following steps may further be included:
[0107] Calculate the first penalty coefficient β using the following formula t :
[0108] e t-1 = g t-1 - D t-1
[0109]
[0110] where e t-1 represents the smoothed gradient exponent D of the (t - 1)-th iteration t-1 and the true error between the true gradient value g t-1 , β t represents the first penalty coefficient, and k represents a hyperparameter.
[0111] In some embodiments of the present application, S120 is further introduced below, which may specifically include the following steps:
[0112] S121: Calculate the model parameters of the distributed deep learning model according to the formula and the predicted gradient value.
[0113] Specifically, the formula is as follows:
[0114]
[0115] where θ t+1 represents the updated model parameters, θ t represents the model parameters of the t-th iteration, ∈t represents the learning rate, represents the predicted gradient value at the t-th iteration.
[0116] The other structures of this embodiment are the same as those of Embodiment III, and will not be elaborated here.
[0117] Embodiment V
[0118] Figure 3 FIG. 5 is a schematic diagram of a distributed deep learning model update device proposed in Embodiment V of the present invention. The device may include:
[0119] A gradient acquisition module 10, configured to acquire the true gradient value of the distributed deep learning model;
[0120] A gradient prediction module 20, configured to calculate a predicted gradient value according to the true gradient value and the following formula:
[0121] D t =(1 - β t )D t-1 +β t g t-1
[0122]
[0123]
[0124] where D t represents the smoothed gradient exponent at the t-th iteration, β t represents the first penalty coefficient, D t-1 represents the smoothed gradient exponent at the (t - 1)-th iteration, g t-1 represents the true gradient value at the (t - 1)-th iteration, v t represents the predicted error between the smoothed gradient exponent D t at the t-th iteration and the true gradient value g t , α represents the second penalty coefficient, represents the error between the true gradient value and the predicted gradient value at the (t - 1)-th iteration, represents the predicted gradient value at the t-th iteration;
[0125] A parameter update module 30, configured to update the model parameters of the distributed deep learning model by using the predicted gradient value.
[0126] The device may further include:
[0127] A judgment cut-off unit, configured to judge whether the number of iterations meets a preset number of iterations. If so, the training of the distributed deep learning model is completed; if not, the next iteration is started, and the gradient acquisition module 10 is called back to execute.
[0128] The gradient acquisition module may include:
[0129] A first gradient acquisition sub-module, configured to acquire model parameters of a distributed deep learning model;
[0130] A second gradient acquisition sub-module, configured to calculate a true gradient value according to the following formula and the model parameters:
[0131]
[0132] where g t represents the true gradient value at the t-th iteration, θ t represents the model parameters at the t-th iteration, and F θ (o(t,i)) represents the loss function with respect to the sample o(t,i).
[0133] The apparatus may further include:
[0134] A first penalty coefficient calculation module, configured to calculate the first penalty coefficient β using the following formula before executing the gradient prediction module t :
[0135] e t-1 = g t-1 - D t-1
[0136]
[0137] where e t-1 represents the true error between the smoothed gradient exponent D t-1 at the (t - 1)-th iteration and the true gradient value g t-1 , β t represents the first penalty coefficient, and k represents a hyperparameter.
[0138] The parameter update module may include:
[0139] A parameter update sub-module, configured to calculate model parameters of the distributed deep learning model according to the following formula and the predicted gradient value:
[0140]
[0141] where θ t+1 represents the updated model parameters, θ t represents the model parameters at the t-th iteration, ∈ t represents the learning rate, represents the predicted gradient value at the t-th iteration.
[0142] In some embodiments of the present application, as Figure 4 shown, an optional operating program of the present invention is provided, where β0 ≤ βt ≤1. The present invention can adaptively adjust the confidence coefficient of historical gradients for accelerating large-scale distributed training of deep learning. It can be applied to video description and image classification tasks, greatly reducing the training time while ensuring convergence. Meanwhile, the present invention also has more effective weak scalability. When properly configured, the present invention can achieve nearly linear scalability when using 128 nodes.
[0143] Embodiment Six
[0144] Embodiment Six of the present invention provides an application method of a distributed deep learning model update method, applying the distributed deep learning model update method provided in Embodiment One to traffic target recognition. As Figure 5 shown, the application method includes:
[0145] S501: Obtain a training sample set of traffic targets; the traffic targets include pedestrians, vehicles, and traffic lights. The training sample set contains all required traffic targets, which can be a publicly available data set or a self-collected data set.
[0146] S502: Input the training sample set into a pre-selected and constructed distributed deep learning model for model training, and adopt the above-mentioned distributed deep learning model update method during the training process.
[0147] S503: Recognize traffic targets through the trained distributed deep learning model.
[0148] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0149] Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A distributed deep learning model update method, characterized in that Including: Obtain the true gradient value g of the distributed deep learning model t ; According to the true gradient value g t and the following formula, the predicted gradient value is calculated D t = (1 - β t )D t-1 + β t g t-1 Among them, D t represents the exponential of the smoothed gradient at the t-th iteration, β t represents the first penalty coefficient, D t-1 represents the exponential of the smoothed gradient at the (t - 1)-th iteration, g t-1 represents the true gradient value at the (t - 1)-th iteration, v t represents the exponential of the smoothed gradient at the t-th iteration D t and the prediction error between the true gradient value g t . α represents the second penalty coefficient, represents the error between the true gradient value and the predicted gradient value in the (t - 1)-th iteration, represents the predicted gradient value at the t-th iteration; Using the predicted gradient value Update the model parameters of the distributed deep learning model; Before obtaining the predicted gradient value, it further includes: The first penalty coefficient β is calculated using the following formula t :[[]]END]] e t-1 = g t-1 - D t-1 Among them, e t-1 represents the smoothed gradient exponent D of the (t - 1)-th iteration t-1 and the true error between the true gradient value g t-1 ; e t-2 represents the smoothed gradient exponent D of the (t - 2)-th iteration t-2 and the true error between the true gradient value g t-2 ; β t represents the first penalty coefficient of the t-th iteration, β t-1 represents the first penalty coefficient of the (t - 1)-th iteration, and k represents a hyperparameter; An application method of the distributed deep learning model update method, including: Obtain a training sample set of traffic targets; the traffic targets include pedestrians, vehicles, and traffic lights; Input the training sample set into a pre-selected and constructed distributed deep learning model for model training, and adopt the distributed deep learning model update method during the training process; Identify traffic targets through the trained distributed deep learning model.
2. The distributed deep learning model updating method according to claim 1, wherein After updating the parameters of the distributed deep learning model using the predicted gradient value, it further includes: Judge whether the number of iterations reaches a preset number; If so, complete the training of the distributed deep learning model; If not, start the next iteration process.
3. The distributed deep learning model update method according to claim 2, wherein The obtaining of the true gradient value of the distributed deep learning model includes: Obtain the model parameters of the distributed deep learning model; Calculate the true gradient value according to the following formula and the model parameters: Among them, g t represents the true gradient value at the t-th iteration, and θ t represents the model parameters at the t-th iteration, and F θ (o(t, i)) represents the loss function with respect to the sample o(t, i).
4. The distributed deep learning model updating method according to claim 3, wherein The obtaining of the true gradient value of the distributed deep learning model includes: Obtain the true gradient value g in the previous iteration of the distributed deep learning model t-1 .
5. The method for updating a distributed deep learning model according to claim 1, wherein Updating the model parameters of the distributed deep learning model using the predicted gradient value includes: Calculate the model parameters of the distributed deep learning model according to the following formula and the predicted gradient value: Among them, θ t+1 represents the updated model parameter, θ t represents the model parameter at the t-th iteration, ∈ t represents the learning rate, represents the predicted gradient value at the t-th iteration.
6. A distributed deep learning model update device, characterized in that Including: A gradient acquisition module for obtaining the true gradient value of the distributed deep learning model; A gradient prediction module for calculating a predicted gradient value according to the true gradient value and the following formula: D t = (1 - β t )D t-1 + β t g t-1 Among them, D t represents the smoothed gradient exponent of the t-th iteration, β t represents the first penalty coefficient, D t-1 represents the smoothed gradient exponent of the (t - 1)-th iteration, g t-1 represents the true gradient value of the (t - 1)-th iteration, v t represents the smoothed gradient exponent D of the t-th iteration t and the prediction error between the true gradient value g t . α represents the second penalty coefficient, represents the error between the true gradient value and the predicted gradient value in the (t - 1)-th iteration, represents the predicted gradient value of the t-th iteration; A parameter update module for updating the model parameters of the distributed deep learning model using the predicted gradient value; Before obtaining the predicted gradient value, it further includes: The first penalty coefficient β is calculated using the following formula t :[[]]END]] e t-1 = g t-1 - D t-1 Among them, e t-1 represents the smoothed gradient exponent D of the (t - 1)-th iteration t-1 and the true error between it and the true gradient value g t-1 e t-2 represents the smoothed gradient exponent D of the (t - 2)-th iteration t-2 and the true error between it and the true gradient value g t-2 β t represents the first penalty coefficient of the t-th iteration, β t-1 represents the first penalty coefficient of the (t - 1)-th iteration, and k represents a hyperparameter.
7. The distributed deep learning model updating device according to claim 6, wherein The device further includes: A judgment cut-off unit for judging whether the number of iterations meets the preset number of iterations. If so, complete the training of the distributed deep learning model. If not, start the next iteration process and return to execute the gradient acquisition module.
8. The distributed deep learning model updating device according to claim 7, wherein The gradient acquisition module includes: A first gradient acquisition sub-module for obtaining the model parameters of the distributed deep learning model; A second gradient acquisition sub-module for calculating the true gradient value according to the following formula and the model parameters: Among them, g t represents the true gradient value at the t-th iteration, and θ t represents the model parameters at the t-th iteration, and F θ (o(t, i)) represents the loss function with respect to the sample o(t, i).
Citation Information
Patent Citations
Model parameter updating system, method and device
CN112085074A