Model training method and device, storage medium and electronic equipment
By using the approximate calculation method of first-order gradients in deep model training to approximate the second-order gradient, the problem of low model training efficiency in the prior art is solved, and faster and more efficient model training is achieved.
Patent Information
- Application Number
- CN202411967383.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art only uses first-order gradients to optimize the parameter in deep model training, resulting in low model training efficiency and lack of solutions to effectively improve model training efficiency.
The second-order gradient is calculated by approximating the first-order gradient under two consecutive iterations, and the second-order gradient is used for parameter optimization, thereby accelerating convergence and improving the efficiency of parameter optimization.
It effectively improves the efficiency of model training, accelerates the convergence process, and improves the effect of the optimizer without consuming a lot of space resources.
Smart Images

Figure CN120069121A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a model training method, apparatus, storage medium, and electronic device. Background Art
[0002] Currently, in various technical fields of artificial intelligence (such as NLP (Natural Language Processing), images, speech, etc.), the application of deep models has been very extensive; and deep learning models often have a large number of parameters and are very deep, coupled with the complexity of the training sample distribution, which brings great challenges to model learning. The advent of the LLM (Large Language Models) era has made this challenge even more prominent. In the training of deep models, the effect of the optimizer is very important; however, related technologies usually only optimize parameters through first-order gradients, resulting in low model training efficiency. Based on this, there is currently no good solution to how to improve the efficiency of model training. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a model training method, apparatus, storage medium, and electronic device to solve problems such as low efficiency of model training caused by related technologies; that is, embodiments of the present invention can approximately calculate the second-order gradient through the first-order gradients in two consecutive iterations to optimize parameters using the second-order gradient, thereby effectively accelerating convergence, that is, effectively improving the efficiency of parameter optimization to improve the efficiency of model training; and, embodiments of the present invention do not consume a large amount of space resources when approximately calculating the second-order gradient, so on the basis of not consuming a large amount of space resources, the effect of the optimizer is effectively improved, and then the model training effect is effectively improved, and the model can be trained faster and better.
[0004] According to one aspect of the embodiments of the present invention, there is provided a model training method, the method including:
[0005] Obtain a first model and a training data set, where the first model is obtained based on the model parameter values of each model parameter in at least one model parameter in the previous iteration;
[0006] Call the first model to respectively determine the model operation results of each training data in the training data set; and based on the model operation results of each training data, calculate the first-order gradients of each model parameter in the current iteration;
[0007] Determine the first-order gradients of the respective model parameters at the previous iteration, and calculate the second-order gradients of the respective model parameters at the current iteration respectively based on the first-order gradients of the respective model parameters at the previous iteration and the first-order gradients of the respective model parameters at the current iteration;
[0008] Calculate the model parameter values of the respective model parameters at the current iteration respectively based on the first-order gradients and second-order gradients of the respective model parameters at the current iteration to obtain a second model, and thus determine a target model based on the second model; wherein, the second model is determined based on the model parameter values of the respective model parameters at the current iteration.
[0009] According to another aspect of the embodiments of the present invention, there is provided a model training apparatus, the apparatus includes:
[0010] An acquisition unit, configured to acquire a first model and acquire a training data set, where the first model is acquired based on the model parameter values of the respective model parameters in at least one model parameter at the previous iteration;
[0011] A processing unit, configured to call the first model to respectively determine the model running results of each training data in the training data set; and calculate the first-order gradients of the respective model parameters at the current iteration based on the model running results of each training data;
[0012] The processing unit is further configured to determine the first-order gradients of the respective model parameters at the previous iteration, and calculate the second-order gradients of the respective model parameters at the current iteration respectively based on the first-order gradients of the respective model parameters at the previous iteration and the first-order gradients of the respective model parameters at the current iteration;
[0013] The processing unit is further configured to calculate the model parameter values of the respective model parameters at the current iteration respectively based on the first-order gradients and second-order gradients of the respective model parameters at the current iteration to obtain a second model, and thus determine a target model based on the second model; wherein, the second model is determined based on the model parameter values of the respective model parameters at the current iteration.
[0014] According to another aspect of the embodiments of the present invention, there is provided an electronic device, the electronic device includes a processor and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the method mentioned above.
[0015] According to another aspect of the embodiments of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to cause a computer to execute the method mentioned above.
[0016] In the embodiments of the present invention, after obtaining the first model and the training data set, the first model can be called to respectively determine the model running results of each training data in the training data set. The first model is obtained based on the model parameter values of each model parameter in at least one model parameter in the previous iteration; and based on the model running results of each training data, the first-order gradients of each model parameter in the current iteration are calculated. Based on this, the first-order gradients of each model parameter in the previous iteration can be determined, and based on the first-order gradients of each model parameter in the previous iteration and the first-order gradients of each model parameter in the current iteration, the second-order gradients of each model parameter in the current iteration are respectively calculated. Further, based on the first-order gradients and second-order gradients of each model parameter in the current iteration, the model parameter values of each model parameter in the current iteration are respectively calculated to obtain a second model, so as to determine the target model based on the second model; wherein, the second model is determined based on the model parameter values of each model parameter in the current iteration. It can be seen that the embodiments of the present invention can approximately calculate the second-order gradient through the first-order gradients in two consecutive iterations, so as to use the second-order gradient for parameter optimization, thereby effectively accelerating convergence, that is, effectively improving the efficiency of parameter optimization to improve the efficiency of model training; and, the embodiments of the present invention do not consume a large amount of space resources when approximately calculating the second-order gradient, so on the basis of not consuming a large amount of space resources, the effect of the optimizer is effectively improved, and then the model training effect is effectively improved, and the model can be trained faster and better. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features and advantages of the present invention are disclosed. In the drawings:
[0018] Figure 1 A flowchart showing a model training method according to an exemplary embodiment of the present invention is shown;
[0019] Figure 2 A flowchart showing another model training method according to an exemplary embodiment of the present invention is shown;
[0020] Figure 3 A schematic block diagram showing a model training device according to an exemplary embodiment of the present invention is shown;
[0021] Figure 4 A structural block diagram of an exemplary electronic device capable of implementing the embodiments of the present invention is shown. DETAILED DESCRIPTION
[0022] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0023] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0024] As used herein, the term "comprising" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0025] It should be noted that the modifications of "one" and "a plurality" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless clearly indicated otherwise in the context, it should be understood as "one or more".
[0026] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0027] It should be noted that the execution subject of the model training method provided by the embodiments of the present invention can be one or more electronic devices, and the present invention does not limit this; among them, the electronic device can be a terminal (i.e., a client) or a server. Then, when the execution subject includes multiple electronic devices, and at least one terminal and at least one server are included in the multiple electronic devices, the model training method provided by the embodiments of the present invention can be jointly executed by the terminal and the server. Correspondingly, the terminal mentioned here can include, but is not limited to: smart phones, tablet computers, laptop computers, desktop computers, smart watches, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, and so on. The server mentioned here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, and so on.
[0028] Based on the above description, the embodiments of the present invention propose a model training method, which can be executed by the above-mentioned electronic device (terminal or server); or, this model training method can be jointly executed by the terminal and the server. For the convenience of description, in the following, it will be described by taking the electronic device executing this model training method as an example; as Figure 1 shown, this model training method may include the following steps S101 - S104:
[0029] S101, obtain a first model and obtain a training data set, where the first model is obtained based on the model parameter values of each model parameter in at least one model parameter in the previous iteration.
[0030] Optionally, a model (such as the first model described above, etc.) can be an Artificial Neural Networks (ANNs) model, a Convolutional Neural Networks (CNN) model, a BP (Back Propagation) neural network model, etc.; the embodiments of the present invention do not limit this. In the embodiments of the present invention, at least one model parameter can include all model parameters in a specified model structure; correspondingly, the model structures of the models (such as the first model or the second model described below, etc.) in an iterative process can all be the specified model structure, and the models in an iterative process can be obtained based on the model parameter values of each model parameter at the corresponding iteration. For example, the model parameter values of each model parameter in the previous iteration can be used as the model parameter values of the corresponding model parameter in the first model, and the first model can be the model in the previous iteration. Optionally, the specified model structure can be set according to experience or according to actual requirements; the embodiments of the present invention do not limit this; for example, the number of neural network layers of the specified model structure or the number of neurons in any layer of the neural network can be set, etc.
[0031] Optionally, the training data set can include at least one training data in a target application scenario. Optionally, the target application scenario can be any application scenario; the embodiments of the present invention do not limit this; for example, it can be any classification scenario or any target recognition scenario, etc.
[0032] In the embodiments of the present invention, the acquisition method of the training data set can include but is not limited to at least one of the following:
[0033] The first acquisition method: At least one training data in multiple application scenarios can be stored in the storage space of the electronic device itself. In this case, the electronic device can select at least one training data in the target application scenario from at least one training data in multiple application scenarios, and add at least one training data in the target application scenario to the training data set to obtain the training data set.
[0034] The second acquisition method: The electronic device can obtain a target training data download link, and add the training data downloaded based on the target training data download link to the training data set to obtain the training data set.
[0035] The third acquisition method: The electronic device can also acquire at least one initial training data in a target application scenario, perform data preprocessing on the at least one initial training data to obtain a data preprocessing result, and add the data preprocessing result to the training data set to implement the acquisition of the training data set, and so on. Exemplarily, the data preprocessing process may include but is not limited to at least one of the following: missing value processing, outlier detection and processing, data transformation, feature selection, data encoding, and data standardization, and so on; the embodiments of the present invention do not limit this.
[0036] Optionally, the acquisition method of the first model may include but is not limited to at least one of the following:
[0037] The first acquisition method: The previous iteration may be the (t - 1)-th iteration. Then when t takes the value of 1, the first model may be an initial model. In this case, the electronic device can initialize each model parameter in the specified model structure (i.e., each model parameter in the at least one model parameter) to obtain an initial model, and use the initial model as the first model to implement the acquisition of the first model. In this case, the model parameter values of each model parameter in the first model may be the model parameter values of the corresponding model parameter in the 0-th iteration, that is, the initial values of the corresponding model parameters.
[0038] The second acquisition method: The electronic device stores initial models in multiple application scenarios in its own storage space. Then when t takes the value of 1, the electronic device can select the initial model in the target application scenario from the initial models in multiple application scenarios, and use the initial model in the target application scenario as the first model.
[0039] The third acquisition method: When t is greater than 1, the first model may be a model obtained by performing (t - 1) iterations of training on the initial model. In this case, the electronic device can determine the model parameter values of each model parameter in the previous iteration, and use the model parameter values of each model parameter in the previous iteration as the model parameter values of the corresponding model parameter in the first model to implement the acquisition of the first model, and so on.
[0040] S102, call the first model to respectively determine the model operation results of each training data in the training data set; and calculate the first-order gradients of each model parameter in the current iteration based on the model operation results of each training data.
[0041] In an embodiment of the present invention, for any training data in a training data set, an electronic device may input any training data into a first model to call the first model to determine the model operation result of any training data, that is, the model operation result of any training data can be output through the first model. Exemplarily, taking a classification scenario as an example, the first model may be called to perform classification prediction on any training data to obtain the model operation result of any training data (which can be the classification prediction probability of any training data at this time), and so on.
[0042] S103. Determine the first-order gradient of each model parameter in the previous iteration, and calculate the second-order gradient of each model parameter in the current iteration based on the first-order gradient of each model parameter in the previous iteration and the first-order gradient of each model parameter in the current iteration.
[0043] In an embodiment of the present invention, for any one of at least one model parameter, an electronic device may calculate the second-order gradient of any model parameter in the current iteration based on the first-order gradient of any model parameter in the previous iteration and the first-order gradient of any model parameter in the current iteration.
[0044] S104. Calculate the model parameter value of each model parameter in the current iteration based on the first-order gradient and the second-order gradient of each model parameter in the current iteration to obtain a second model, and thus determine a target model based on the second model; wherein, the second model is determined based on the model parameter values of each model parameter in the current iteration.
[0045] In an embodiment of the present invention, for any one of at least one model parameter, an electronic device may calculate the model parameter value (i.e., the model parameter value) of any model parameter in the current iteration based on the first-order gradient and the second-order gradient of any model parameter in the current iteration.
[0046] Based on this, an electronic device may use the model parameter values of each model parameter in the current iteration as the model parameter values of the corresponding model parameters in the second model respectively to update the model parameter values of each model parameter; that is to say, the model parameter values of the corresponding model parameters in the first model may be updated based on the model parameter values of each model parameter in the current iteration to obtain a second model.
[0047] In an embodiment of the present invention, after obtaining the first model and the training data set, the first model may be called to respectively determine the model running results of each training data in the training data set. The first model is obtained based on the model parameter values of each model parameter in at least one model parameter in the previous iteration; and based on the model running results of each training data, the first-order gradients of each model parameter in the current iteration are calculated. Based on this, the first-order gradients of each model parameter in the previous iteration can be determined, and based on the first-order gradients of each model parameter in the previous iteration and the first-order gradients of each model parameter in the current iteration, the second-order gradients of each model parameter in the current iteration are respectively calculated. Further, based on the first-order gradients and the second-order gradients of each model parameter in the current iteration, the model parameter values of each model parameter in the current iteration are respectively calculated to obtain the second model, so as to determine the target model based on the second model; wherein, the second model is determined based on the model parameter values of each model parameter in the current iteration. It can be seen that the embodiment of the present invention can approximately calculate the second-order gradient through the first-order gradients in two consecutive iterations to use the second-order gradient for parameter optimization, thereby effectively accelerating convergence, that is, effectively improving the efficiency of parameter optimization to improve the efficiency of model training; and, when approximately calculating the second-order gradient, the embodiment of the present invention does not consume a large amount of space resources, so on the basis of not consuming a large amount of space resources, the effect of the optimizer is effectively improved, and then the model training effect is effectively improved, and the model can be trained faster and better.
[0048] Based on the above description, an embodiment of the present invention further proposes a more specific model training method. Correspondingly, this model training method can be executed by the above-mentioned electronic device (terminal or server); or, this model training method can be jointly executed by the terminal and the server. For the convenience of description, in the following, the example of the electronic device executing this model training method will be used for illustration; please refer to Figure 2 , this model training method may include the following steps S201-S206:
[0049] S201, obtain the first model and the training data set. The first model is obtained based on the model parameter values of each model parameter in at least one model parameter in the previous iteration.
[0050] S202, call the first model to respectively determine the model running results of each training data in the training data set; and based on the model running results of each training data, calculate the first-order gradients of each model parameter in the current iteration.
[0051] Optionally, the electronic device may determine a model loss function; and may calculate a first-order gradient of each model parameter at the current iteration based on the model running result of each training data, the labeled data of each training data, and the model loss function. It should be understood that in a deep model, the electronic device can usually obtain the first-order gradients of all model parameters in all layers according to the model loss function, by specifying the model structure, using the chain rule, and backpropagation, that is, the first-order gradients of each model parameter at the current iteration can be obtained.
[0052] Based on this, the electronic device may calculate a first-order gradient of each model parameter at the current iteration by using the model running result of each training data, the labeled data of each training data, the model loss function, and the model parameter value of each model parameter at the previous iteration. Exemplarily, the first-order gradient of any model parameter at the current iteration may be where J may be the model loss function, θ may be any model parameter, and θ t-1 may be the model parameter value of any model parameter at the (t - 1)-th iteration (i.e., the previous iteration); that is, the first-order partial derivative of the model loss function with respect to any model parameter can be calculated, and then substituted with the model running result of each training data, the labeled data of each training data, and the model parameter value of each model parameter at the previous iteration, etc., to obtain the first-order gradient of any model parameter at the current iteration. In other embodiments, the first-order partial derivative of the model loss function with respect to a parameter vector composed of at least one model parameter may also be calculated, so as to obtain the first-order gradients of each model parameter at the current iteration at one time, to obtain the first-order gradient of any model parameter at the current iteration, etc.; the present invention does not limit this.
[0053] S203. Determine the first-order gradients of each model parameter at the previous iteration.
[0054] In the embodiment of the present invention, the current iteration may be the t-th iteration, and the previous iteration may be the (t - 1)-th iteration, where t is a positive integer. Optionally, for any model parameter among at least one model parameter, after obtaining the first-order gradient of any model parameter at any iteration, the electronic device may store the first-order gradient of any model parameter at any iteration; that is, the electronic device may store the first-order gradients of each model parameter at the previous iteration, so that the first-order gradients of each model parameter at the previous iteration can be determined from the local storage space.
[0055] Optionally, when the value of t is 1, the first-order gradient of any model parameter in the previous iteration can be initialized to a preset first-order gradient, that is, the first-order gradient of any model parameter in the 0th iteration can be initialized to a preset first-order gradient. Optionally, the preset first-order gradient can be set according to experience or according to actual requirements, and the embodiments of the present invention do not limit this; exemplarily, the preset first-order gradient can be 0, that is, the first-order gradient of any model parameter in the 0th iteration can be 0.
[0056] S204. For any one of at least one model parameter, perform a difference operation on the first-order gradient of the any model parameter in the current iteration and the first-order gradient of the any model parameter in the previous iteration to obtain the iterative difference operation result of the any model parameter.
[0057] S205. Take the iterative difference operation result of any model parameter as the second-order gradient of the any model parameter in the current iteration.
[0058] Based on this, the electronic device can obtain the second-order gradients of each model parameter in the current iteration.
[0059] S206. Based on the first-order gradients and second-order gradients of each model parameter in the current iteration, calculate the model parameter values of each model parameter in the current iteration respectively to obtain a second model, and thus determine a target model based on the second model; wherein, the second model is determined based on the model parameter values of each model parameter in the current iteration.
[0060] In the embodiments of the present invention, for any one of at least one model parameter, the electronic device can calculate the second-order momentum of any model parameter in the current iteration (i.e., the tth iteration) according to Formula 1.1:
[0061] v t =β 2 v t-1 +(1 - β 2 )(g t -g t-1 ) 2 Formula 1.1
[0062] wherein, v t can be the second-order momentum of any model parameter in the current iteration, v t-1 can be the second-order momentum of any model parameter in the previous iteration (i.e., the (t - 1)th iteration), β 2 can be the exponential decay rate of the second-moment estimation, g t can be the first-order gradient of any model parameter in the current iteration, g t-1 can be the first-order gradient of any model parameter in the previous iteration, g t -g t-1can be the second-order gradient of any model parameter at the current iteration. Optionally, β 2 , v 0 , etc. can be set according to experience or according to actual requirements, and so on. The embodiments of the present invention do not limit this; for example, β 2 can be 0.999, and v 0 can be set to 0. It can be seen that the embodiments of the present invention can perform exponential smoothing on the square of the second-order gradient through Formula 1.1. Since the difference between two consecutive first-order gradients is closer to the differential calculation of the second-order gradient (the second-order gradient calculated by differentiation is equal to the difference between the first-order gradients divided by the difference of the independent variables. Due to the effect of exponential smoothing, the difference of the independent variables in each dimension will be greatly reduced, that is, the second-order gradient at the current iteration will not have a great impact on the second-order momentum. Therefore, the difference between the first-order gradients can be directly used to approximate the second-order gradient), the optimizer is closer to the ideal "Newton's method", and thus the effect will be better.
[0063] Furthermore, based on the second-order momentum of any model parameter at the current iteration and the first-order gradient of any model parameter at the current iteration, the model parameter value of any model parameter at the current iteration can be calculated to obtain a second model. Correspondingly, the electronic device can calculate the first-order momentum of any model parameter at the current iteration based on the first-order gradient of any model parameter at the current iteration, and calculate the first-order momentum bias correction result of any model parameter at the current iteration based on the first-order momentum of any model parameter at the current iteration. Optionally, the electronic device can use Formula 1.2 to calculate the first-order momentum of any model parameter at the current iteration:
[0064] m t = β 1 m t-1 +(1 - β 1 )g t Formula 1.2
[0065] where m t can be the first-order momentum of any model parameter at the current iteration, β 1 can be the exponential decay rate of the first-moment estimation, and m t-1 can be the first-order momentum of any model parameter at the previous iteration. Optionally, β 1 , m 0 , etc. can be set according to experience or according to actual requirements. The embodiments of the present invention do not limit this; for example, β 1 can be 0.9, and m 0It can be set to 0. Based on this, the embodiment of the present invention can implement exponential smoothing of the first-order gradient, thereby obtaining the first-order momentum; that is, the embodiment of the present invention can use the first-order gradient of any model parameter in the current iteration and the first-order momentum of any model parameter in the previous iteration through formula 1.2 to update the first-order momentum of any model parameter to obtain the first-order momentum of any model parameter in the current iteration.
[0066] Accordingly, the electronic device can use formula 1.3 to calculate the first-order momentum bias correction result of any model parameter in the current iteration:
[0067]
[0068] in, It can represent the first-order momentum bias correction result of any model parameter in the current iteration, and can also be called the normalized first-order momentum; that is, the embodiment of the present invention can normalize the first-order momentum to obtain the first-order momentum bias correction result of any model parameter in the current iteration.
[0069] Further, the electronic device may calculate the second-order momentum bias correction result of any model parameter in the current iteration based on the second-order momentum of any model parameter in the current iteration; illustratively, the electronic device may use formula 1.4 to calculate the second-order momentum bias correction result of any model parameter in the current iteration:
[0070]
[0071] in, It can represent the second-order momentum bias correction result of any model parameter in the current iteration, and can also be called the normalized second-order momentum; that is, the embodiment of the present invention can normalize the second-order momentum to obtain the second-order momentum bias correction result of any model parameter in the current iteration.
[0072] Based on this, the electronic device can use the first-order momentum bias correction result of any model parameter in the current iteration and the second-order momentum bias correction result of any model parameter in the current iteration to calculate the model parameter value of any model parameter in the current iteration. Optionally, the electronic device can use formula 1.5 to calculate the model parameter value of any model parameter in the current iteration:
[0073]
[0074] Among them, η can be the learning rate, ∈ can be a smaller parameter (used to ensure that the denominator is not 0), θ tcan be the model parameter value of any model parameter at the current iteration. Optionally, η and ∈ can be set according to experience, can be set according to actual requirements, can also be randomly generated within any range, and so on; the embodiments of the present invention do not make any limitations in this regard. Based on this, the embodiments of the present invention can use Formula 1.5 to update the model parameter value of any model parameter through the quasi-Newton method to obtain the model parameter value of any model parameter at the current iteration.
[0075] Further, the electronic device can use the second model (i.e., the model at the t-th iteration) as the first model (i.e., the model at the (t - 1)-th iteration), and iteratively execute the call to the first model to respectively determine the model operation results of each training data in the training dataset, so as to update the model parameter values of each model parameter at the current iteration until the convergence condition is reached; that is to say, t + 1 can be used as t (i.e., make t = t + 1), and iteratively execute the call to the model at the (t - 1)-th iteration to respectively determine the model operation results of each training data (i.e., can update the model operation results of each training data), so as to update the model parameter values of each model parameter at the current iteration. Optionally, it can be determined that the convergence condition is reached when the model loss value at the current iteration is less than the preset model loss threshold; or, it can be determined that the convergence condition is reached when the number of iterations is greater than the preset iteration number threshold, and so on; the embodiments of the present invention do not make any limitations in this regard. Optionally, both the preset model loss threshold and the preset iteration number threshold can be set according to experience or can be set according to actual requirements, and the embodiments of the present invention do not make any limitations in this regard.
[0076] Based on this, the electronic device can use the second model when the convergence condition is reached as the target model to realize determining the target model based on the second model; that is to say, the second model at the time of stopping iteration (i.e., the model at the t-th iteration at the time of stopping iteration) can be used as the target model.
[0077] In summary, the embodiments of the present invention can store the first-order gradient of any model parameter at any number of iterations, approximately obtain the second-order gradient of any model parameter at the current iteration by taking the difference, and then use the exponential smoothing algorithm for the first-order gradient and the second-order gradient, which can effectively eliminate the oscillations of the first-order gradient and the second-order gradient caused by the position (i.e., iteration), and then obtain the update amount of any model parameter (i.e., the model parameter value of any model parameter at the current iteration) through the quasi-Newton method. In other words, the embodiments of the present invention propose an approximately second-order Newton method optimizer, which can significantly improve the parameter learning effect of the optimizer without significantly increasing the computing and storage costs, and can be better used in many scenarios, that is, can be better applied to various scenarios, such as classification scenarios, etc.
[0078] In an embodiment of the present invention, after obtaining the first model and the training data set, the first model is called to determine the model operation results of each training data in the training data set respectively; and based on the model operation results of each training data, the first-order gradients of each model parameter in the current iteration are calculated. Based on this, the first-order gradients of each model parameter in the previous iteration can be determined; for any one of at least one model parameter, a difference operation can be performed on the first-order gradient of any model parameter in the current iteration and the first-order gradient of any model parameter in the previous iteration to obtain the iterative difference operation result of any model parameter; and the iterative difference operation result of any model parameter can be used as the second-order gradient of any model parameter in the current iteration. Further, based on the first-order gradients and second-order gradients of each model parameter in the current iteration, the model parameter values of each model parameter in the current iteration can be calculated respectively to obtain the second model, and thus based on the second model, the target model can be determined; wherein, the second model is determined based on the model parameter values of each model parameter in the current iteration. It can be seen that in the embodiment of the present invention, the second-order gradient of any model parameter can be approximately calculated through the difference operation result between the first-order gradient of any model parameter in the current iteration and the first-order gradient of any model parameter in the previous iteration, that is, the second-order gradient can be approximated by the first-order gradient, which can effectively reduce the deviation of the second-order gradient. Furthermore, through exponential smoothing and quasi-Newton method, the model parameters are iterated to continuously update the model parameter values of the model parameters during training, so that the model quickly reaches the optimum to obtain the target model; and, the embodiment of the present invention can accurately approximate the second-order gradient without bringing excessive computational and storage consumption. That is to say, compared with the complete second-order gradient, the embodiment of the present invention does not need to store a large second-order gradient matrix and does not require more additional calculations, which can effectively save storage and effectively save computing resources, thus making the embodiment of the present invention have great advantages, especially for the case of tight storage space. Based on this, the embodiment of the present invention can approximate the second-order gradient through the difference of the first-order gradient to calculate the second-order dynamics, which can reduce noise to a certain extent, accelerate convergence, and obtain the effect of an approximate second-order gradient optimizer, so as to quickly complete model training, obtain the target model, and effectively improve the model performance of the target model.
[0079] Based on the description of the related embodiments of the above model training method, an embodiment of the present invention also proposes a model training device, which can be a computer program (including program code) running in an electronic device; as Figure 3 shown, the model training device may include an acquisition unit 301 and a processing unit 302. The model training device can execute Figure 1 or Figure 2 the model training method shown, that is, the model training device can run the above units:
[0080] An acquisition unit 301, configured to acquire a first model and a training data set, where the first model is acquired based on the model parameter values of each model parameter in at least one model parameter in the previous iteration;
[0081] A processing unit 302, configured to call the first model to respectively determine the model operation results of each training data in the training data set; and calculate the first-order gradients of the respective model parameters in the current iteration based on the model operation results of each training data;
[0082] The processing unit 302 is further configured to determine the first-order gradients of the respective model parameters in the previous iteration, and calculate the second-order gradients of the respective model parameters in the current iteration respectively based on the first-order gradients of the respective model parameters in the previous iteration and the first-order gradients of the respective model parameters in the current iteration;
[0083] The processing unit 302 is further configured to calculate the model parameter values of the respective model parameters in the current iteration respectively based on the first-order gradients and second-order gradients of the respective model parameters in the current iteration to obtain a second model, and thus determine a target model based on the second model; where the second model is determined based on the model parameter values of the respective model parameters in the current iteration.
[0084] In one implementation, when calculating the second-order gradients of the respective model parameters in the current iteration based on the first-order gradients of the respective model parameters in the previous iteration and the first-order gradients of the respective model parameters in the current iteration, the processing unit 302 may specifically be configured to:
[0085] For any one of the at least one model parameter, perform a difference operation on the first-order gradient of the any one model parameter in the current iteration and the first-order gradient of the any one model parameter in the previous iteration to obtain an iterative difference operation result of the any one model parameter;
[0086] Use the iterative difference operation result of the any one model parameter as the second-order gradient of the any one model parameter in the current iteration.
[0087] In another implementation, the current iteration is the t-th iteration, and the previous iteration is the (t - 1)-th iteration, where t is a positive integer; when calculating the model parameter values of the respective model parameters in the current iteration based on the first-order gradients and second-order gradients of the respective model parameters in the current iteration to obtain a second model, the processing unit 302 may specifically be configured to:
[0088] Calculate the second-order momentum of any of the model parameters at the current iteration according to the formula; where, is the second-order momentum of any of the model parameters at the current iteration, is the second-order momentum of any of the model parameters at the previous iteration, is the exponential decay rate of the second-moment estimate, is the first-order gradient of any of the model parameters at the current iteration, is the first-order gradient of any of the model parameters at the previous iteration, and is the second-order gradient of any of the model parameters at the current iteration;
[0089] Based on the second-order momentum of any of the model parameters at the current iteration and the first-order gradient of any of the model parameters at the current iteration, calculate the model parameter value of any of the model parameters at the current iteration to obtain a second model.
[0090] In another implementation, when the processing unit 302 calculates the model parameter value of any of the model parameters at the current iteration based on the second-order momentum of any of the model parameters at the current iteration and the first-order gradient of any of the model parameters at the current iteration, it can be specifically used for:
[0091] Based on the first-order gradient of any of the model parameters at the current iteration, calculate the first-order momentum of any of the model parameters at the current iteration, and based on the first-order momentum of any of the model parameters at the current iteration, calculate the first-order momentum bias correction result of any of the model parameters at the current iteration;
[0092] Based on the second-order momentum of any of the model parameters at the current iteration, calculate the second-order momentum bias correction result of any of the model parameters at the current iteration;
[0093] Use the first-order momentum bias correction result of any of the model parameters at the current iteration and the second-order momentum bias correction result of any of the model parameters at the current iteration to calculate the model parameter value of any of the model parameters at the current iteration.
[0094] In another implementation, when the processing unit 302 calculates the first-order gradient of each of the model parameters at the current iteration based on the model running result of each training data, it can be specifically used for:
[0095] Determine the model loss function;
[0096] Based on the model running result of each training data, the label data of each training data, and the model loss function, calculate the first-order gradient of each of the model parameters at the current iteration.
[0097] In another implementation, when the processing unit 302 determines the target model based on the second model, it can be specifically used for:
[0098] Use the second model as the first model, and iteratively execute the step of calling the first model to respectively determine the model running results of each training data in the training dataset, so as to update the model parameter values of each model parameter under the current iteration until the convergence condition is reached;
[0099] Use the second model when the convergence condition is reached as the target model, so as to implement determining the target model based on the second model.
[0100] According to an embodiment of the present invention, Figure 1 or Figure 2 each step involved in the method shown can be executed by Figure 3 each unit in the model training device shown. For example, Figure 1 the step S101 shown in can be executed by Figure 3 the obtaining unit 301 shown in, and the steps S102 - S104 can all be executed by Figure 3 the processing unit 302 shown in. Another example, Figure 2 the step S201 shown in can be executed by Figure 3 the obtaining unit 301 shown in, and the steps S202 - S206 can all be executed by Figure 3 the processing unit 302 shown in, and so on.
[0101] According to another embodiment of the present invention, Figure 3 each unit in the model training device shown can be respectively or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units with more functions to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, any model training device may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.
[0102] According to another embodiment of the present invention, it can be achieved by running a computer program (including program code) that can execute the steps involved in the corresponding method shown in Figure 1 or Figure 2 on a general - purpose electronic device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random - access storage medium (RAM), and a read - only storage medium (ROM), to construct a model training device as shown in Figure 3The model training device shown in , and to implement the model training method of the embodiments of the present invention. The computer program can be recorded on, for example, a computer storage medium, loaded into the above-mentioned electronic device through the computer storage medium, and run therein.
[0103] In the embodiments of the present invention, after obtaining the first model and the training data set, the first model can be called to respectively determine the model operation results of each training data in the training data set. The first model is obtained based on the model parameter values of each model parameter in at least one model parameter in the previous iteration; and based on the model operation results of each training data, the first-order gradients of each model parameter in the current iteration are calculated. Based on this, the first-order gradients of each model parameter in the previous iteration can be determined, and based on the first-order gradients of each model parameter in the previous iteration and the first-order gradients of each model parameter in the current iteration, the second-order gradients of each model parameter in the current iteration are respectively calculated. Further, based on the first-order gradients and second-order gradients of each model parameter in the current iteration, the model parameter values of each model parameter in the current iteration are respectively calculated to obtain the second model, so as to determine the target model based on the second model; wherein, the second model is determined based on the model parameter values of each model parameter in the current iteration. It can be seen that the embodiments of the present invention can approximately calculate the second-order gradient by the first-order gradients in two consecutive iterations, so as to use the second-order gradient for parameter optimization, thereby effectively accelerating convergence, that is, effectively improving the efficiency of parameter optimization to improve the efficiency of model training; and, the embodiments of the present invention do not consume a large amount of space resources when approximately calculating the second-order gradient, so on the basis of not consuming a large amount of space resources, the effect of the optimizer is effectively improved, and then the model training effect is effectively improved, and the model can be trained faster and better.
[0104] Based on the descriptions of the above method embodiments and device embodiments, an exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to cause the electronic device to execute the method according to the embodiments of the present invention when executed by the at least one processor.
[0105] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is used to cause the computer to execute the method according to the embodiments of the present invention when executed by a processor of the computer.
[0106] An exemplary embodiment of the present invention further provides a computer program product, including a computer program, wherein the computer program is used to cause the computer to execute the method according to the embodiments of the present invention when executed by a processor of the computer.
[0107] Reference Figure 4 Now, the structural block diagram of the electronic device 400 that can be used as the server or client of the present invention will be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0108] As Figure 4 shown, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 402 or the computer program loaded from the storage unit 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0109] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, an output unit 407, a storage unit 408, and a communication unit 409. The input unit 406 can be any type of device that can input information into the electronic device 400. The input unit 406 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 407 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include but is not limited to a magnetic disk, an optical disk. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0110] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above. For example, in some embodiments, the model training method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. In some embodiments, the computing unit 401 can be configured to execute the model training method by any other suitable means (e.g., by means of firmware).
[0111] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable model training devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0112] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0113] As used in this invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0114] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0115] The systems and techniques described here can be implemented in a computing system including a back-end component (e.g., as a data server), or a computing system including a middleware component (e.g., an application server), or a computing system including a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described here), or in a computing system including any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0116] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0117] Also, it should be understood that the foregoing disclosure is only a preferred embodiment of the present invention and, of course, cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made in accordance with the claims of the present invention are still within the scope covered by the present invention.
Claims
1. A model training method, characterized in that: include: Acquire a first model and a training data set, wherein the first model is acquired based on a model parameter value of each model parameter in at least one model parameter in a previous iteration; Calling the first model to respectively determine the model running results of each training data in the training data set; and calculating the first-order gradient of each model parameter in the current iteration based on the model running results of each training data; Determine the first-order gradient of each model parameter in the previous iteration, and calculate the second-order gradient of each model parameter in the current iteration based on the first-order gradient of each model parameter in the previous iteration and the first-order gradient of each model parameter in the current iteration; Based on the first-order gradient and the second-order gradient of each model parameter in the current iteration, the model parameter values of each model parameter in the current iteration are calculated respectively to obtain a second model, and then based on the second model, the target model is determined; wherein the second model is determined based on the model parameter values of each model parameter in the current iteration.
2. The method according to claim 1, characterized in that The calculating the second-order gradients of the model parameters at the current iteration based on the first-order gradients of the model parameters at the previous iteration and the first-order gradients of the model parameters at the current iteration respectively includes: For any model parameter among the at least one model parameter, performing a difference operation on a first-order gradient of the any model parameter in the current iteration and a first-order gradient of the any model parameter in the previous iteration to obtain an iterative difference operation result of the any model parameter; The iterative difference calculation result of any model parameter is used as the second-order gradient of any model parameter in the current iteration.
3. The method according to claim 2, characterized in that The current iteration is the t-th iteration, and the previous iteration is the t-1-th iteration, where t is a positive integer; and based on the first-order gradient and the second-order gradient of each model parameter at the current iteration, respectively calculating the model parameter values of each model parameter at the current iteration to obtain the second model, comprises: According to the formula v t =β2v t-1 +(1-β2)(g t -g t-1 ) 2 Calculate the second-order momentum of any model parameter at the current iteration; wherein the v t is the second-order momentum of any model parameter at the current iteration, and v t-1 is the second-order momentum of any model parameter at the last iteration, β2 is the exponential decay rate of the second-order moment estimate, and g t is the first-order gradient of any model parameter at the current iteration, and g t-1 is the first-order gradient of any model parameter at the last iteration, and g t -g t-1 is the second-order gradient of any model parameter at the current iteration; Based on the second-order momentum of any model parameter at the current iteration and the first-order gradient of any model parameter at the current iteration, the model parameter value of any model parameter at the current iteration is calculated to obtain a second model.
4. The method according to claim 3, characterized in that: The calculating the model parameter value of any model parameter at the current iteration based on the second-order momentum of any model parameter at the current iteration and the first-order gradient of any model parameter at the current iteration comprises: Calculating a first-order momentum of any model parameter at the current iteration based on a first-order gradient of any model parameter at the current iteration, and calculating a first-order momentum bias correction result of any model parameter at the current iteration based on the first-order momentum of any model parameter at the current iteration; Based on the second-order momentum of any one of the model parameters at the current iteration, calculating a second-order momentum bias correction result of any one of the model parameters at the current iteration; The model parameter value of any model parameter in the current iteration is calculated by using the first-order momentum bias correction result of any model parameter in the current iteration and the second-order momentum bias correction result of any model parameter in the current iteration.
5. The method according to any one of claims 1 to 4, characterized in that: The step of calculating the first-order gradient of each model parameter in the current iteration based on the model running result of each training data includes: Determine the model loss function; Based on the model running results of each training data, the label data of each training data and the model loss function, the first-order gradient of each model parameter in the current iteration is calculated.
6. The method according to any one of claims 1 to 4, characterized in that: The step of determining a target model based on the second model includes: Using the second model as the first model, and iteratively executing the calling of the first model, respectively determining the model running results of each training data in the training data set, so as to update the model parameter values of each model parameter in the current iteration, until a convergence condition is reached; The second model that reaches the convergence condition is used as the target model to achieve the determination of the target model based on the second model.
7. A model training device, characterized in that: The device comprises: An acquisition unit, configured to acquire a first model and a training data set, wherein the first model is acquired based on a model parameter value of each model parameter in at least one model parameter in a previous iteration; A processing unit, configured to call the first model, respectively determine the model operation result of each training data in the training data set; and calculate the first-order gradient of each model parameter in the current iteration based on the model operation result of each training data; The processing unit is further used to determine the first-order gradient of each model parameter in the previous iteration, and calculate the second-order gradient of each model parameter in the current iteration based on the first-order gradient of each model parameter in the previous iteration and the first-order gradient of each model parameter in the current iteration; The processing unit is also used to calculate the model parameter values of each model parameter in the current iteration based on the first-order gradient and second-order gradient of each model parameter in the current iteration, so as to obtain a second model, and thus determine the target model based on the second model; wherein the second model is determined based on the model parameter values of each model parameter in the current iteration.
8. An electronic device, characterized in that: include: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-6.