Construction method of deep learning model parameter optimizer based on two-stage optimization architecture

By building a deep learning model parameter optimizer with a two-stage optimization architecture, combining traditional optimizers with periodic learning rates, determining the boundaries of better parameters of the model and optimizing parameters, the problem of existing optimizers relying on gradient direction updates is solved, and the prediction accuracy and stability of the model are improved.

CN120337991AActive Publication Date: 2025-07-18HUAZHONG UNIV OF SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510826609.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The existing deep learning model optimizer relies on gradient direction updates, resulting in unstable model prediction performance and large differences in optimization effects on different models, making it difficult to effectively improve prediction accuracy.

Method used

A deep learning model parameter optimizer based on a two-stage optimization architecture is built. Through the combination of traditional optimizers and periodic learning rates, the boundaries of the model's better parameters are determined, and the parameters are optimized by fitting the curve relationship between the model training parameters and the loss function value.

Benefits of technology

It breaks through the limitations of traditional optimizers and improves the prediction accuracy and stability of the model on time series prediction tasks. The optimization effect is better than traditional stochastic gradient descent and Adam and other optimizers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337991A_ABST
    Figure CN120337991A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of deep learning time sequence prediction, and discloses a construction method of a deep learning model parameter optimizer based on a two-stage optimization architecture. A model optimal parameter boundary determination method in which a traditional optimizer and a periodic learning rate are combined in the first step and a parameter optimization method in which a curvilinear relationship between a model training parameter and a loss function value is fitted in the second step are adopted. The construction method of the optimizer not only breaks through the limitation that a traditional optimizer only depends on gradient direction updating, but also can stably and effectively improve the prediction precision of the deep learning model on the time sequence, and provides a new research direction for improvement of the prediction performance of the subsequent deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning time series prediction, and particularly to a construction method of a deep learning model parameter optimizer based on a two-stage optimization architecture. Background Art

[0002] Currently, the research on deep learning model parameter optimization is mainly divided into two types. One is to use intelligent optimization algorithms to optimize model hyperparameters, such as particle swarm optimization algorithm, genetic algorithm, and sparrow search algorithm, etc.; the other is to use optimizers to optimize the internal parameters of the model, such as stochastic gradient descent (SGD), mini-batch gradient descent, RMSProp, and Adam, etc. Although both types of research are committed to exploring the prediction potential of the model, the latter involves the internal mechanism of the deep learning model and has relatively low innovation, so it is often ignored in the research process.

[0003] A deep learning model optimizer is an algorithm used to update the internal parameters of a neural network to reduce the difference between the model prediction result and the actual value. Different from intelligent optimization algorithms such as particle swarm and genetic algorithms that optimize model hyperparameters, the optimizer is mainly used to optimize the internal parameters during model training. It updates the model parameters during the training process through different gradient calculation methods to find the optimal parameters of the model in practical problems.

[0004] Currently, commonly used deep learning model optimizers such as Adam and SGD that only rely on gradient direction updates are convenient and fast to use, but the optimized model has poor stability. It can neither fully and effectively improve the prediction performance of the model, nor has a large difference in optimization performance when facing different models. Summary of the Invention

[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a construction method of a deep learning model parameter optimizer based on a two-stage optimization architecture, aiming to break the limitation of traditional optimizers that only rely on gradient direction updates by constructing a new type of deep learning model parameter optimizer with a two-stage optimization architecture, thereby improving the ability of the optimizer to explore the prediction performance of the model and stably and effectively improving the prediction accuracy of the deep learning model.

[0006] To achieve the above object, according to one aspect of the present invention, a construction method of a deep learning model parameter optimizer based on a two-stage optimization architecture is provided, including the following steps: Step 1: Determine the optimal parameter boundary of the model Select a traditional optimizer as the preliminary optimizer for determining the optimal parameter boundary of the deep learning model, and use a fixed learning rate Train the deep learning model. After reaching the preset number of iterations \(t\), adjust the learning rate to a periodic learning rate, and extract the model parameters during the periodic change of the learning rate to determine the optimal parameter boundary of the model. Step 2: Optimize the model parameters Fit the optimal parameters obtained in Step 1 with their corresponding model training loss values, and gradually find and obtain the optimal parameters of the deep learning model through the fitted curve.

[0007] Preferably, the traditional optimizer in Step 1 includes Adam optimizer, Stochastic Gradient Descent optimizer, Mini-batch Gradient Descent optimizer, RMSProp optimizer; preferably Adam optimizer.

[0008] Preferably, the periodic learning rate adjustment method in Step 1 includes periodic learning rate of semi-cosine annealing curve, linear decay learning rate, and stepped decay learning rate.

[0009] Preferably, select the periodic learning rate of the semi-cosine annealing curve in Step 1, and its calculation formula is: ; where is the period length, is the learning rate at the \(t\)-th moment, is the maximum learning rate, is the minimum learning rate, represents the cosine function, represents divided by the remainder of

[0010] Preferably, the specific steps to obtain the optimal parameter boundary of the model through the periodic change of the learning rate are as follows: Based on the periodic learning rate formula of the semi-cosine annealing curve, set the period length , the total number of periods \(k\), the number of iterations \(t\) before the start of the loop, the maximum learning rate and the minimum learning rate ; Select the traditional Adam optimizer as the early optimizer, and train the parameters of the deep learning model with the maximum learning rate ; When the training iteration reaches \(t\) times, start to adjust the learning rate to the periodic learning rate of the semi-cosine annealing curve, so that the learning rate periodically fluctuates between the maximum learning rate and the minimum learning rate during the model training process, and at the same time extract the model parameters at the last time point of each period during the periodic change of the learning rate ; After the model training is completed, k sets of model parameters and their loss values are obtained. The above model parameters surround and form a boundary of relatively optimal model parameters around the optimal parameters with smaller loss values on the loss surface of the deep learning model.

[0011] Preferably, the fitting formula in step two is: ; Where, is the model training loss value, is the model parameter, and are constants.

[0012] Preferably, the fitting formula in step two is: ; Where, is the model training loss value, is the model parameter, , and c are constants.

[0013] Preferably, the specific steps of the method for optimizing model parameters by fitting the curve relationship between model training parameters and loss function values in step two are as follows: Since the model parameters are gradually optimized during the training process of the deep learning model, in order to reduce the influence of poor model parameters on curve fitting, the last 2 sets of model parameters and obtained through the periodic learning rate and their loss values and are selected for curve fitting; then on the fitted curve, the early stopping mechanism is introduced to control the optimization process, that is, by gradually reducing the model loss value , to obtain the optimal parameter that can make the prediction effect evaluation index of the model on the training set or validation set reach the best. This optimal parameter is used as the model parameter obtained by optimizing the TSO method finally and is used as the prediction model parameter for the remaining test set.

[0014] Preferably, the prediction effect evaluation index includes the root mean square error of the prediction results on the training set or validation set or the sum of the root mean square errors of both.

[0015] Preferably, the prediction effect evaluation index also includes the mean absolute error, mean absolute percentage error or coefficient of determination.

[0016] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, due to the adoption of the method for determining the optimal model parameters boundary by combining the traditional gradient descent optimizer with the periodic learning rate in Step 1 and the parameter optimization method for fitting the curve relationship between the training parameters of the model and the loss function value in Step 2, the following beneficial effects can be achieved: 1. Creatively combine the curve function with progressive convergence in nature with the loss surface theory of the deep learning model to establish or mathematical models, breaking through the limitation that traditional optimizers only rely on gradient direction updates.

[0017] 2. Adopt a two-stage optimization architecture. In the first stage, use a traditional optimizer plus a periodic learning rate to explore the optimal model parameter boundary, and in the second stage, achieve fine-tuning of parameters through curve fitting, breaking through the rigidity of the single learning rate strategy of traditional optimizers.

[0018] 3. This optimizer performs better in optimizing various deep learning models for time series prediction tasks compared with traditional optimizers such as Stochastic Gradient Descent (SGD) and Adam, and the prediction accuracy of the optimized model is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a schematic diagram of the semi-cosine annealing learning rate.

[0020] Figure 2 is a schematic diagram of the deep learning model optimizer based on the two-stage optimization architecture on the loss surface.

[0021] Figure 3 is a comparison of the prediction performance of the optimizer proposed by the present invention with other traditional optimizers on multiple deep learning models. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0023] Example 1: Taking the hourly-scale load data prediction of a certain province as an example, the long short-term memory network (LSTM) and the gated recurrent unit (GRU) are selected as the deep learning models to be optimized to compare and verify the effect of the present invention. The data used includes the short-term load data of a certain province on an hourly time scale in 2021, the average temperature data per hour, and the working day mapping features. In addition, the data set is divided into a training set, a validation set, and a test set with a ratio of 8:1:1. When training and predicting the model, the historical data of 1 week (168h) is used as the model input, and the 1h future load is used as the output step length.

[0024] Define the hyperparameters of the basic model: The prediction models optimized by the two-stage architecture optimizer for the long short-term memory network (LSTM) and the gated recurrent unit (GRU) are TSO-LSTM and TSO-GRU respectively, and the hyperparameters of the models are as follows: TSO-LSTM model: Select Adam as the initial optimizer; select the semi-cosine annealing learning rate as the way of periodic fluctuation of the learning rate; the initial learning rate before fluctuation, that is, the maximum learning rate is set to 0.008, and the minimum learning rate during fluctuation is set to 0.0006; select the iteration number t to start the periodic change of the learning rate after 380 times, the cycle length is 10 generations, and a total of 2 cycles change.

[0025] TSO-GRU model: Select Adam as the initial optimizer; select the semi-cosine annealing learning rate as the way of periodic fluctuation of the learning rate; the initial learning rate before fluctuation, that is, the maximum learning rate is set to 0.006, and the minimum learning rate during fluctuation is set to 0.0008; select the iteration number t to start the periodic change of the learning rate after 370 times, the cycle length is 15 generations, and a total of 2 cycles change.

[0026] Note: The two models have no association during the optimization process and are independent individuals. Selecting the two models is only for comparing and demonstrating the applicability of the optimizer proposed by the present invention.

[0027] Step 1. Determine the boundary of the optimal parameters of the model Select the traditional Adam optimizer as the preliminary optimizer for determining the boundary of the optimal parameters of the deep learning model, and use a fixed learning rate to train the deep learning model. After reaching the preset iteration number t, adjust the learning rate to the periodic learning rate of the semi-cosine annealing curve, and extract the model parameters during the periodic change of the learning rate t to determine the boundary of the optimal parameters of the model; Among them, the formula for the periodic learning rate of the semi-cosine annealing curve is as follows. According to the set cycle length when the learning rate changes periodically , the total number of cycles k, the number of iterations t before the start of the loop, the maximum learning rate and the minimum learning rate and other hyperparameters are filled and the formula is used: ; where, is the cycle length, is the learning rate at time t, is the maximum learning rate, is the minimum learning rate, represents the cosine function, represents divided by the remainder.

[0028] The specific steps to obtain the better parameter boundary of the model through the periodic change of the learning rate are as follows: After setting the periodic learning rate formula of the semi-cosine annealing curve, select the traditional Adam optimizer as the pre-optimizer, and use the maximum learning rate to train the parameters of the deep learning model; when the number of training iterations reaches t times, start to adjust the learning rate to the periodic learning rate of the semi-cosine annealing curve, so that the learning rate periodically fluctuates between the maximum learning rate and the minimum learning rate during the model training process, and at the same time extract the model parameters at the last time point of each cycle during the periodic change of the learning rate and the corresponding loss value at its corresponding time point; after the model training is completed, k groups of model parameters and their loss values are obtained, and the above model parameters surround and form a better parameter boundary around the optimal parameters with smaller loss values on the loss surface of the deep learning model.

[0029] Since the model parameters of the deep learning model are gradually optimized during the training process, in order to reduce the influence of poor model parameters on curve fitting, let the learning rate change for 2 cycles. Through this process, two better parameters around the optimal parameters corresponding to smaller loss values in the deep neural network model can be obtained.

[0030] As Figure 1 shown, adjust the learning rate to the periodic learning rate of the semi-cosine annealing curve, so that the learning rate periodically fluctuates between the maximum learning rate and the minimum learning rate during the periodic change process, and at the same time extract the model parameters and at the last time point of each cycle, so as to obtain two better parameters around the optimal parameter and and its corresponding loss value and 。

[0031] Step 2: Optimize model parameters Take the two relatively optimal parameters obtained in Step 1 and and their corresponding model training loss values and Perform fitting according to the following formula, and gradually search for and obtain the optimal parameters of the deep learning model through the fitted curve. The fitting formula is: or

[0032] where is the model training loss value, is the model parameter, 、 and c are constants.

[0033] The method of optimizing model parameters by fitting the curve relationship between model training parameters and loss function values. The specific steps are as follows: Since the model parameters of the deep learning model are gradually optimized during the training process, in order to reduce the impact of poor model parameters on curve fitting, select the last 2 groups of model parameters obtained through the periodic learning rate and and their loss values and to perform curve fitting; then on the fitted curve, control the optimization process by introducing an early stopping mechanism, that is, by gradually reducing the model loss value to obtain the optimal parameter that can make the prediction effect evaluation index of the model on the training set or validation set reach the best. This optimal parameter is used as the model parameter optimized by the final TSO method and is used as the prediction model parameter for the remaining test set.

[0034] In this embodiment, the root mean square error (RMSE) of the validation set prediction result is used as the evaluation index. When the RMSE no longer improves, the parameter corresponding to this loss value is taken as the internal parameter of the model when finally predicting the test set.

[0035] In this example, the root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination ( ) are selected as the evaluation indexes for the prediction effect of this model. The Nesterov accelerated gradient (NAG) algorithm and Adam algorithm, which also belong to the field of model parameter optimization, are used as comparison algorithms. Figure 3It shows the comparison of the prediction performance of the optimizer proposed by the present invention with other traditional optimizers on LSTM and GRU deep learning models.

[0036] Except that the coefficient of determination ( ) index is better when the value is closer to 1, the smaller the values of the other three indicators, namely root mean square error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE), the better. Therefore, it can be seen from Figure 3 that whether LSTM or GRU is selected as the basic model to be optimized, the performance of the optimizer is that TSO is better than Adam, and Adam is better than NAG. It can be seen that the deep learning model parameter optimizer (TSO) based on the two-stage optimization architecture proposed by the present invention has certain advantages over traditional optimizers in terms of improving the model prediction performance.

[0037] In this embodiment, the traditional Adam optimizer is selected as the preliminary optimizer for determining the optimal parameter boundary of the deep learning model. In addition to selecting the Adam optimizer as the preliminary optimizer for optimizing the deep learning model parameters, other types of optimizers can also be selected, such as stochastic gradient descent (SGD), mini-batch gradient descent, RMSProp, etc. At the same time, in addition to selecting the periodic learning rate of the semi-cosine annealing curve, other periodic learning rate methods can also be selected to adjust the learning rate, such as linear decay learning rate and step decay learning rate.

[0038] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A construction method of a deep learning model parameter optimizer based on a two-stage optimization architecture, characterized in that It includes the following steps: Step 1: Determine the boundary of the better parameters of the model Select a traditional optimizer as the preliminary optimizer for determining the optimal parameter boundary of the deep learning model, and use a fixed learning rate Train the deep learning model. After reaching the preset number of iterations t, adjust the learning rate to a periodic learning rate, and extract the model parameters during the periodic change of the learning rate to determine the optimal parameter boundary of the model; Step 2: Optimize the model parameters Fit the better parameters obtained in Step 1 with their corresponding model training loss values, and gradually search and obtain the optimal parameters of the deep learning model through the fitted curve.

2. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 1, wherein, In Step 1, the traditional optimizers include Adam optimizer, stochastic gradient descent optimizer, mini-batch gradient descent optimizer, and RMSProp optimizer; preferably, the Adam optimizer is used.

3. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 1, characterized in that, In Step 1, the periodic learning rate adjustment methods include the periodic learning rate of the semi-cosine annealing curve, the linear decay learning rate, and the stepwise decay learning rate.

4. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 3, characterized in that, In Step 1, the periodic learning rate of the semi-cosine annealing curve is selected, and its calculation formula is: ; wherein, is the cycle length, is the learning rate at the t-th moment, is the maximum learning rate, is the minimum learning rate, represents the cosine function, represents the remainder of dividing by 5. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 4, characterized in that, The specific steps to obtain the optimal parameter boundary of the model through the periodic change of the learning rate are as follows: Based on the periodic learning rate formula of the semi-cosine annealing curve, set the period length when the learning rate changes periodically , the total number of periods k, the number of iterations t before the start of the loop, the maximum learning rate and the minimum learning rate ; Select the traditional Adam optimizer as the initial optimizer and train the parameters of the deep learning model with the maximum learning rate ; After the training iteration reaches t times, start to adjust the learning rate to the periodic learning rate of the semi-cosine annealing curve, so that the learning rate fluctuates periodically between the maximum learning rate and the minimum learning rate during the model training process. At the same time, extract the model parameters at the last time point of each period during the periodic change of the learning rate and the corresponding loss value at the time point where they are located; After the model training is completed, k groups of model parameters and their loss values are obtained. The above model parameters surround on the loss surface of the deep learning model and form a model better parameter boundary around the optimal parameters with smaller loss values.

6. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 1, characterized in that, In Step 2, the fitting formula is: ; Among them, is the model training loss value, are the model parameters, and are constants.

7. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 1, characterized in that, In Step 2, the fitting formula is: ; Among them, is the model training loss value, is the model parameter, , and c are constants.

8. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 6 or 7, characterized in that In step two, the method of optimizing the model parameters by fitting the curve relationship between the training parameters and the loss function values is as follows: select the last two sets of model parameters obtained through the periodic learning rate and and their loss values and to perform curve fitting; then, on the fitted curve, control the optimization process by introducing an early stopping mechanism, that is, by gradually reducing the model loss value , to obtain the optimal parameters that can make the prediction effect evaluation index of the model on the training set or the validation set reach the best , and this optimal parameter is used as the model parameter optimized by the final TSO method and used as the prediction model parameter for the remaining test set.

9. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 8, characterized in that, The prediction effect evaluation index includes the root mean square error of the prediction results of the training set or the validation set or the sum of the root mean square errors of both.

10. The construction method of the deep learning model parameter optimizer based on the two-stage optimization architecture according to claim 8, characterized in that, The prediction effect evaluation index also includes the mean absolute error, the mean absolute percentage error, or the coefficient of determination.

Citation Information

Patent Citations

  • Adaptive deep learning model optimization method based on Keras platform

    CN110245742A

  • Visual navigation method based on autonomous learning

    CN117288205A

  • Urban development boundary identification method based on deep learning

    CN119360147A

  • Deep learning model optimization method and apparatus for medical image segmentation

    KR102680328B1

  • System and method for efficient generation of machine-learning models

    US20200302234A1