Efficient optimization method based on local gradient information
By introducing the exponential moving average of gradient square and nonlinear transformation of local difference in the optimization algorithm, the DFGD algorithm is proposed, which solves the problems of slow convergence speed and large parameter fluctuations in the training process of the existing optimization algorithm, and achieves faster convergence and higher accuracy.
Patent Information
- Application Number
- CN202510297360.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-10
AI Technical Summary
The existing optimization algorithms are susceptible to noise and instability during training, resulting in slow convergence speed, frequent fluctuations in parameters, and it is difficult to gradually approach the short update step when approaching the local minimum value, resulting in oscillation.
An efficient optimization method based on local gradient information is proposed, called the difference gradient descent (DFGD) algorithm. By using the exponential moving average of the gradient square and the local difference of the parameter gradient, the coefficients after nonlinear transformation are used to adjust the step-by-step length of the parameter, reducing the use of Momentum momentum.
It effectively reduces the vertical fluctuations near the local optimal solution, speeds up the convergence speed, improves the algorithm accuracy, and shows higher accuracy and shorter training time in the experiment.
Smart Images

Figure CN120124702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gradient optimization, and more specifically, to an efficient optimization method based on local gradient information. Background Art
[0002] Stochastic Gradient Descent (SGD) is one of the most commonly used optimization methods and also one of the mainstream optimizers in the industry. Although SGD has been around for a long time, its classic theory and excellent experimental performance have withstood the test of time. It has extensive application potential in fields such as computer vision, natural language processing, speech recognition, reinforcement learning, and multi-modal tasks. SGD and its improved algorithms (such as Momentum, Adagrad, Adadelta, and Adam) can provide efficient parameter optimization strategies for deep learning task models such as trajectory prediction, image target recognition, and sensor data anomaly detection, thereby improving algorithm efficiency and stability.
[0003] The core principle of SGD is to iteratively adjust the model parameters to minimize the loss function. In each iteration, SGD randomly selects a small batch of training samples and calculates the gradients of these samples. Then, it updates the model parameters in the opposite direction of the gradient to reduce the value of the loss function. Although the concept of SGD is relatively simple, it faces some challenges in practical applications. First, the convergence speed of SGD may be slow, especially in complex deep neural networks. Second, SGD is vulnerable to noise and instability in the training data, resulting in oscillations during the training process. In addition, choosing an appropriate learning rate is also a key issue, because an inappropriate learning rate may lead to training failure or long convergence times.
[0004] The characteristic that SGD only uses the gradient of the current small batch for parameter update each time leads to frequent fluctuations of the parameters during training, especially in regions where the loss function surface has large variations, and also results in a slow convergence speed. To solve this problem, Boris T. Polyak proposed the Momentum algorithm, which introduced the physical concept of momentum into the optimization of deep learning. The core idea of the Momentum algorithm is to improve the performance of SGD by introducing historical gradient information. This historical gradient information is called momentum, which is similar to the physical concept of an object accumulating speed during motion. The momentum term is combined with the current gradient in each iteration to determine the direction and speed of parameter update.
[0005] Adaptive learning rate is another major direction for the improvement of SGD and is also the current main research direction. John C. Duchi et al. proposed the Adaptive Gradient (Adagrad) algorithm, which is an adaptive learning rate algorithm aiming to optimize model parameters in deep learning. Among the above-mentioned methods, whether it is traditional SGD or momentum methods such as Momentum and NAG, they all use a fixed learning rate. However, this approach is not suitable for the characteristics of different parameters and objective functions. A too large learning rate may lead to instability, while a too small learning rate will result in an overly slow convergence speed. The goal of Adagrad is to overcome these problems and improve the training efficiency through an adaptive learning rate. The core of the Adagrad algorithm lies in the adaptive learning rate mechanism, which automatically adjusts the learning rate for each parameter. The working principle of this mechanism is as follows: First, the algorithm maintains a growing "learning rate accumulation" for each parameter by accumulating the squared values of past gradients. As the training progresses, the learning rate of each parameter continuously decreases, thereby reducing the amplitude of the gradient descent step. This means that for parameters with frequently large gradients, the learning rate will decrease to reduce the fluctuations during the training process.
[0006] The adaptive learning rate mechanism of Adagrad can effectively handle the characteristics of different parameters, improve the convergence speed, and at the same time, it does not require manual setting of the global learning rate, reducing the burden of hyperparameter tuning. However, Adagrad also has some problems. The accumulation of the learning rate will cause the learning rate to be too small during long-term training and even tend to zero, resulting in the inability to update the parameters continuously. Later, Matthew D. Zeiler proposed the Adadelta algorithm. The Adadelta algorithm proposed a feasible solution to the problem of the monotonically decreasing learning rate in Adagrad. Adadelta limits the window size for calculating the historical gradients to a fixed value. Instead of storing the previous squared gradients, it recursively represents the squared gradient as the mean of all historical squared gradients. This adaptive learning rate mechanism enables Adadelta to better adapt to the characteristics of different parameters and improves the training stability and convergence speed.
[0007] Another mainstream optimizer in the industry currently is Adam, which effectively unifies the ideas of the Adadelta algorithm and the momentum method of Momentum. It not only has an excellent self-learning rate adjustment mechanism of the adaptive learning rate algorithm but also combines the consideration of historical cumulative gradients to form momentum, effectively accelerating the convergence speed. In addition, for the historical cumulative gradient, that is, the momentum term mt, and the weighted average term vt of the historical cumulative squared gradient, Adam performs bias correction by taking the unbiased estimates of the first moment and the second moment of the gradient to achieve the purpose of effectively correcting the large update error in the first few steps caused by the need to select a smaller hyperparameter to average the gradient when the gradient is sparse.
[0008] Optimizers need to possess the characteristics of high efficiency, stability, and precision to cope with the following complexities: complex task scenarios (such as trajectory optimization and image recognition), limited data resources (such as sparse or small-sample data), high computational costs, etc. There are certain problems with existing adaptive learning rate algorithms. For example, Adam uses the exponential moving average of past gradients to calculate the first- and second-moment estimates of gradients in the gradient update method, and all historical gradient information is used in each update step. However, after multiple iterations, historical gradients that are far from the current position may interfere with the parameter update at the current position. For example, if the process of descent is likened to a small ball rolling down a hillside, when the terrain is near a gentle valley, as the ball approaches the bottom of the valley, the terrain change near the bottom is very small, but the retention of all historical gradients makes the friction on the ball not obvious, and it cannot approach the bottom of the valley at a slower speed; that is, when the parameter is close to the local minimum, the current gradient and the gradients of the past few steps are relatively small. However, due to using all historical gradient information in the original technology, it is impossible to approach the local minimum with a shorter update step length when approaching the local minimum, and it will turn back to approach the local minimum after crossing the local minimum, and so on in a cycle. This will cause the parameter to oscillate and fluctuate near the local minimum.
[0009] Meanwhile, the existence of momentum makes it possible that in some cases, such as when the loss curve of the training data is very smooth and has not much fluctuation, the model may oscillate around the minimum value, making the convergence more unstable. When the loss curve of the training function has become smooth, this usually means that the model parameters are close to the local minimum or the optimal solution. At this time, the gradients of the model will be very small because the gradients are the derivatives of the loss function with respect to the parameters and they tend to zero. The role of momentum is to consider the previous gradient directions when updating the parameters to help the model continue to move forward when the gradients become small. However, in the case where the loss curve has become smooth, the gradients are already very small, and at this time the role of momentum is relatively small. If the momentum is set too large, it will introduce more historical gradient information when updating the parameters, resulting in the model oscillating around the minimum value because the historical gradient directions do not necessarily point to the minimum value. In addition, when the training data is too small, adding the Momentum term may lead to overfitting. If the training data is too small to cover various situations, boundary conditions, and feature variations of the problem, this will cause the model to only learn the specific patterns in the dataset and not be able to generalize to new, unseen data. The role of the Momentum term is to accumulate historical gradient information during the training process. When the training data is too small, the model may overly focus on the past gradient directions and thus be unable to flexibly respond to new and richer gradient information, increasing the risk of overfitting. And small-sample data may have a small number of samples or an uneven distribution of data among different classes, which results in large differences in gradients between different batches. The characteristic of the Momentum term to accumulate historical gradients makes it possible that after adding Momentum, the accumulation of historical gradient directions may have a negative impact. For example, during the descent process, the objective function passes through a region A with a larger curvature and a region B with a smaller curvature successively. Due to the accumulation of historical gradients in region A, when descending in region B, the parameters cannot move forward step by step with a shorter update step, which may lead to crossing the local minimum and thus oscillation. Therefore, Momentum may amplify the gradient differences between different batches of samples in small-sample data, resulting in the parameters crossing the local minimum and oscillating. This oscillation may cause the model to overfit the training data because the oscillation makes the parameter values of the model unstable, and the model may remember different training samples or patterns in different iterations. This may cause the model to overfit the training data because it iterates repeatedly on the training data and may adopt different parameter values in different iterations, without sufficient stability to capture the true data distribution and losing the stable estimation of the optimal model parameters. Summary of the Invention
[0010] Aiming at the deficiencies of the prior art, the object of the present invention is to propose an efficient optimization method based on local gradient information, including:
[0011] Step 1: Initialize the parameter vector of the Convolutional Neural Network (CNN) model as the parameter vector θ 0 , obtain the image dataset, and input the image data in the image dataset into the CNN model with the parameter vector θ 0 . Output the predicted classification result of the image data. The predicted classification result includes the predicted probabilities of the image data belonging to each category. Based on the predicted classification result and the true classification result of the image data, calculate the loss function L t (θ 0 ), calculate the gradient of the loss function L t (θ 0 ) Similarly, initialize the parameter vector of the CNN model as the parameter vector θ 1 , calculate the loss function L t (θ 1 ), calculate the gradient of the loss function L t )θ 1 (
[0012] Step 2: Initialize the step size η and the time step. The initial value of the time step is 0. Add 1 to the initial value of the time step as the current time step t. Initialize the exponentially weighted moving average v 0 of the squared gradient, and the initial value of v 0 is 0. Use the initialized exponentially weighted moving average v 0 as the exponentially weighted moving average v t-1 of the squared gradient at time step t - 1; Use the parameter vector θ 1 as the parameter vector at time step t; Use the gradient of the loss function L t (θ 0 0) as the gradient of the loss function at time step t - 1 Use the loss function L t (θ 1 ) as the gradient of the loss function at time step t
[0013] Step 3: According to the exponentially weighted moving average v t-1 of the squared gradient at time step t - 1, and the gradient of the loss function at the current time step t , calculate the exponentially weighted moving average v t of the squared gradient at time step t;
[0014] Step 4: Through the weighting coefficient β 2 , perform bias correction on the exponentially weighted moving average v t of the squared gradient at time step t to obtain the corrected weighted average
[0015] Step 5: Calculate the gradient of the loss function at the current time step t and the gradient of the loss function at time step t-1 to obtain the gradient difference ΔL t , and then calculate the transformed coefficient ξ based on the gradient difference ΔL t ; t
[0016] Step 6: According to the initialized step size η, the transformed coefficient ξ t , the parameter vector θ at time step t t , the gradient of the loss function at time step t and the exponentially weighted moving average of the squared gradient at the corrected time step t calculate the parameter vector θ at time step t+1 t+1 ;
[0017] Step 7: Input the image data in the image dataset into the convolutional neural network CNN model with the parameter vector θ t+1 , output the predicted classification result of the image data, and calculate the loss function L based on the predicted classification result and the true classification result of the image data t (θ t+1 (, and determine whether the value of the loss function L t (θ t+1 ) is less than or equal to a preset threshold. If the value of the loss function L t (θ t+1 ) is less than or equal to the preset threshold, take the parameter vector θ t+1 as the final parameter vector of the convolutional neural network CNN model. If the value of the loss function L t (θ t+1 ) is greater than the preset threshold, increment time step t by one and return to execute Step 3.
[0018] Optionally, Step 3 is specifically implemented by the following formula:
[0019]
[0020] where β 2 is the weighting coefficient.
[0021] Optionally, Step 4 is specifically implemented by the following formula:
[0022]
[0023] where β 2 is the weighting coefficient.
[0024] Optionally, Step 5 specifically includes:
[0025] Step 5.1: Subtract the gradient of the loss function at time step t - 1 from the gradient of the loss function at the current time step t to obtain the gradient difference ΔL of and the gradient difference; t ;
[0026] Step 5.2: Perform a non - linear transformation on the gradient difference ΔL of t t to obtain the transformed coefficient ξ of t t .
[0027] Optionally, Step 5.1 is specifically represented by the following formula:
[0028]
[0029] Optionally, Step 5.2 is specifically implemented by the following formula:
[0030]
[0031] Optionally, Step 6 is specifically implemented by the following formula:
[0032]
[0033] where ε is a smoothing term.
[0034] The beneficial effects of adopting the above - mentioned technical solution are as follows:
[0035] The present invention utilizes the exponential moving average of the squared gradient, adds the gradient difference of the parameter gradient, and uses the coefficient after non - linear transformation of the gradient difference during the parameter iteration process to adjust the step size of each step of the parameter, making the parameter with a faster gradient change have a larger update step size, and the parameter with a slower gradient change have a smaller update step size. At the same time, the Momentum momentum is omitted, effectively reducing the vertical fluctuations near the local optimal solution, accelerating the convergence speed, and improving the algorithm accuracy. At the same time, for the proposed algorithm, the present invention proves the upper bound of the parameter update step size and the convergence of the DFGD algorithm to ensure the theoretical rigor and completeness of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic flowchart of an efficient optimization method based on local gradient information in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] The following combines the drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention.
[0038] Aiming at the problems existing in the prior art, the object of the present invention is to explore and improve based on the two characteristics of using all historical gradient information and Momentum momentum in the existing adaptive optimization algorithm. The involved optimization algorithm can effectively improve the interference and oscillation problems existing in parameter update, can accelerate the training speed, reduce the number of iterations, enable the model to converge to the optimal solution faster, improve the accuracy of the model. In addition, the improved algorithm can be selected according to specific application scenarios to adapt to a wider range of application requirements.
[0039] The present invention designs a difference gradient descent (DFGD) algorithm. DFGD introduces a friction coefficient DFC, uses "short-term gradient behavior information" to control the learning rate, makes a non-linear function transformation on the difference between the current gradient and the previous gradient of the parameter, and the obtained result is used as the coefficient to control the change of the parameter step size. The addition of this coefficient makes the curvature change of the objective function near the local minimum smaller. When the gradient difference at this place is also small, the parameter update step size of DFGD will be significantly reduced. When approaching the local minimum, the prior art cannot approach the local minimum step by step with a short update step size, and will turn back to approach the local minimum after crossing the local minimum, resulting in oscillation in a cycle. However, DFGD can locally approach the local minimum with a short update step size, which significantly improves the oscillation suppression effect of DFGD; when the curvature change of the objective function near the local minimum is large, the coefficient after non-linear transformation will tend to 1, so the parameter update step size of DFGD will tend to be the same as that of the prior art and the performance will not be inferior; in addition, aiming at the problem that the Momentum momentum may cause the parameter to oscillate across the local minimum, we do not add Momentum momentum in DFGD.
[0040] Based on the design idea of the above difference gradient descent (DFGD) algorithm, the present invention provides an efficient optimization method based on local gradient information, combined with Figure 1 , which may include the following steps:
[0041] Step 1: Initialize the parameter vector of the convolutional neural network (CNN) model as the parameter vector θ 0 , obtain an image data set, input the image data in the image data set into the convolutional neural network (CNN) model with the parameter vector θ 0 , output the predicted classification result of the image data, where the predicted classification result includes the predicted probabilities of the image data belonging to each category. Based on the predicted classification result and the true classification result of the image data, calculate the loss function L t (θ 0 ), calculate the gradient of the loss function L t (θ 0 ) Similarly, initialize the parameter vector of the convolutional neural network (CNN) model as the parameter vector θ. 1 , and calculate the loss function L t (θ 1 ), and calculate the loss function L t )θ 1 (gradient of
[0042] Among them, the CNN model can be Vgg16 or GoogleNet, and the image dataset can be the CIFAR-10 and CIFAR-100 image datasets, the Fashion mnist clothing classification image dataset, and the mnist handwritten digit image dataset.
[0043] It should be noted that in the present invention, parameter updates are performed for the CNN model based on the image dataset. In actual use, the present invention can also be applied to other models and other datasets.
[0044] Step 2: Initialize the step size η and the time step. The initial value of the time step is 0. Increment the initial value of the time step by one to obtain the current time step t. Initialize the exponentially weighted moving average v of the squared gradient 0 , v 0 with an initial value of 0, and use the initialized exponentially weighted moving average v of the squared gradient 0 as the exponentially weighted moving average v of the squared gradient at time step t - 1 t-1 ; use the parameter vector θ 1 as the parameter vector at time step t; use the gradient of the loss function L t )θ 0 ) as the gradient of the loss function at time step t - 1 Use the loss function L as the gradient of the loss function at time step t t (θ 1 )
[0045] Among them, the step size η can also be understood as the learning rate, and the initial value needs to be set manually during initialization. The value of the step size η can be 0.001.
[0046] Step 3: Calculate the exponentially weighted moving average v of the squared gradient at time step t according to the exponentially weighted moving average v of the squared gradient at time step t - 1 t-1 , and the gradient of the loss function at the current time step t , specifically implemented through the following formula: t The specific formula is as follows:
[0047]
[0048] Among them, β 2 is a weighting coefficient, and the value of β 2 can be 0.999.
[0049] Among them, the exponential moving average of the squared gradient can be understood as the weighted average of the cumulative historical squared gradient and the current squared gradient. Using the exponential moving average of the squared gradient in the DFGD algorithm enables different weight values to be assigned to the current parameter gradients at different time steps t. Moreover, the larger the time step t, the larger the assigned weight value, while the weight value of the parameter gradient at the historical time step will gradually decrease as the iteration progresses. The specific weight assignment scheme is determined by the weighting coefficient β 2 .
[0050] In step 3, the second-order moment estimation of the gradient is used to perform an exponential moving average (EMA) on the squared gradient, thereby obtaining the second-order moment estimation of the gradient. At the same time, the learning rate can be automatically adjusted according to the gradient history to adapt to changes in different parameters. For example, in the initial stage of training, a larger learning rate is used for smaller gradient changes to accelerate convergence, while in the later stage of training, a smaller learning rate is used for possible larger gradient changes to avoid oscillations. Specifically, the exponential moving average causes the parameter gradient at time step t to continuously reduce the weight value at the decay rate of β 2 as the iteration progresses. This processing ensures that the parameter gradient at the current time step t has the maximum weight value to calculate the size of the next parameter update, while also enabling the historical gradient to possess the weight value at the rate of quadratic decay according to time step t, thereby enabling the learning rate to be continuously adjusted according to the change trend of the gradient.
[0051] Since the initial value of the moment estimation is usually 0, this makes the exponential moving average tend to 0, thereby resulting in a smaller estimated value. This deviation will affect the accuracy of parameter updates, especially in the initial stage of training. Based on this problem, bias correction is introduced to obtain a more accurate parameter estimated value in the initial stage of training. The initialization bias formula is obtained by calculating the difference between the expectation of the exponential moving average v t and the expectation of the true second-order moment as follows:
[0052]
[0053] Among them, E[v t represents the expectation of v t , represents 's expectation, g t represents the gradient ζ is the expected deviation generated when the gradient is non-stationary (ζ = 0 when the gradient is stationary), and it can be adjusted by selecting an appropriate decay rate β2 As a result, the exponential moving average assigns smaller weights to gradients in the more distant past, keeping ζ within a smaller range. Therefore, the bias correction formula can be derived as follows:
[0054]
[0055] Thus, the exponential moving average of the squared gradient is corrected based on the derived bias correction formula, specifically implemented through Step 4.
[0056] Step 4: By applying the weighting coefficient β 2 , the exponential moving average v t of the squared gradient at time step t is corrected for bias to obtain the corrected weighted average Specifically, it is achieved through the following formula:
[0057]
[0058] where β 2 is the weighting coefficient.
[0059] Step 5: Calculate the gradient of the loss function at the current time step t and the gradient difference ΔL between the gradient of the loss function at time step t - 1 t . Then, based on the gradient difference ΔL t , calculate the transformed coefficient ξ t ;
[0060] Step 5.1: Subtract the gradient difference between the gradient of the loss function at the current time step t and the gradient of the loss function at time step t - 1 to obtain the gradient difference ΔL t , specifically expressed by the following formula:
[0061]
[0062] Step 5.2: Apply a non - linear transformation to the gradient difference ΔL t to obtain the transformed coefficient ξ t , specifically achieved through the following formula:
[0063]
[0064] Among them, in Step 5.2, the non - linear function with a value range of (0.5, 1) is used to perform a non - linear transformation on the gradient difference ΔL t , mapping the gradient difference, which has large fluctuations and an uncertain value range, into the interval (0, 1), thus establishing a connection between the gradient difference and the adjustment of the learning rate. When the gradient difference is large, the obtained transformed coefficient ξ tis also relatively large and close to 1, which can make the learning rate almost unchanged so that the parameters converge at a relatively fast rate; when the gradient difference is small, the transformed coefficient ξ obtained t is relatively small and closer to 0.5, significantly suppressing the parameter descent speed, thereby reducing the oscillations generated near the extreme value.
[0065] Step 6: According to the initialized step size η, the transformed coefficient ξ t , the parameter vector θ of the time step t t , the gradient of the loss function at the time step t and the exponentially weighted moving average of the squared gradient of the corrected time step t calculate the parameter vector θ of the time step t + 1 t+1 , which is specifically implemented through the following formula:
[0066]
[0067] where ε is a smoothing term used to prevent the denominator from being zero, and the value of ε can be selected as 10 -8 .
[0068] Among them, this step realizes the refined adjustment of the learning rate by multiplying the coefficient ξ after the non - linear transformation t by the original learning rate η. Here, ε is a smoothing term to prevent the denominator from being zero, and generally takes 10 -8 , functions to normalize the gradient in the final parameter update formula, scale or adjust the gradient vector so that it has a unified scale or range. This can make the optimization path smoother, thereby accelerating the convergence speed, and at the same time can effectively avoid the problems of gradient explosion or gradient disappearance.
[0069] Step 7: Input the image data in the image dataset into the convolutional neural network CNN model with the parameter vector θ t+1 , output the predicted classification result of the image data, and based on the predicted classification result and the true classification result of the image data, calculate the loss function L t (θ t+1 ). Determine whether the value of the loss function L t (θ t+1 ) is less than or equal to a preset threshold. When the value of the loss function L t (θ t+1 ) is less than or equal to the preset threshold, take the parameter vector θ t+1 as the final parameter vector of the convolutional neural network CNN model. When the value of the loss function L t (θ t+1 ) is greater than the preset threshold, increment the time step t by one, and return to execute Step 3.
[0070] DFGD utilizes the exponential moving average of the historical squared gradients, adds the local difference of the parameter gradients, and uses the coefficient after non-linearly transforming the historical gradient differences during the parameter iteration process to adjust each step size of the parameters, making the update step size of the parameters with fast gradient changes larger and the update step size of the parameters with slow gradient changes smaller. At the same time, the Momentum momentum is omitted, effectively reducing the vertical fluctuations near the local optimal solution and accelerating the convergence speed.
[0071] The DFGD algorithm can also be understood as the DFGD optimizer. The present invention gives the upper bound of the parameter update step size and the convergence proof of the DFGD optimizer. Specifically, given a sequence of convex functions f 1 (θ), f 2 (θ), …, f T (θ), we use the statistic to determine convergence, where is the optimal parameter and X is the parameter space.
[0072] Theorem 1. Assume that the function f t (θ) has a bounded gradient, that is At the same time, the distance between any parameter θ t generated by the DFGD optimizer is bounded, that is And assume that when the gradient is severely sparse, we consider that the gradient is non-zero only at the current time step and zero at all other time steps, that is, it satisfies v t =(1 - β 2 )[g t 2 , then we have the following equation holds:
[0073] (1) When the gradient is severely sparse, the parameter update step size
[0074] (2) In other cases, the parameter update step size |Δt| ≤ α.[2]
[0075] Proof:
[0076] (1) From we know So
[0077] (2) Since and we have So
[0078] Theorem 2. Assume that the function f t (θ) has a bounded gradient, that is At the same time, the distance between any parameter θ t The distance between is bounded, i.e., Let Then for all T ≥ 1, we have the following equation holds:
[0079]
[0080] Proof:
[0081] First, we have
[0082] Since f t (θ) is a convex function,
[0083] f t (θ * ) ≥ f t (θ t ) + g t (θ * - θ t )
[0084] f t (θ t ) - f t (θ * ) ≤ g t (θ t - θ * )
[0085]
[0086] Therefore, next we derive the expression of . According to the parameter update formula of DFGD, we have: After transformation, we get After expansion, we can obtain
[0087] At this time, the left side of the equation contains the term g t (θ t - θ * ). Factoring it out, we can get
[0088]
[0089] Then we scale and respectively.
[0090] We first organize v t .
[0091]
[0092] Therefore, we have:
[0093]
[0094] Therefore, the scaling of (1) is as follows:
[0095]
[0096] Since terms appear in the denominator, the difference of squares formula is used for scaling to eliminate the denominator and then staggered subtraction is used to eliminate, as shown below:
[0097]
[0098] In summary, we have
[0099] Theorem 3. Assume that the function f t (θ) has a bounded gradient, that is At the same time, the distance between any parameter θ generated by the DFGD optimizer t is bounded, that is Then for all T≥1, we have:
[0100]
[0101] Proof: According to Theorem 1, we have
[0102]
[0103] Therefore, and
[0104] In summary, the parameter update step size of the DFGD optimizer has an upper bound and converges theoretically.
[0105] For the above technical solutions, the present invention has carried out the following experiments:
[0106] Set the batch size to 128, and the iteration termination condition is that the accuracy does not decrease for two consecutive epochs. The models selected are the most commonly used Vgg16 and GoogleNet in CNN. We compare the experimental results with multiple classic optimizers such as Adam, SGD, AdaMod, and AdaGC. The most classic Adam and SGD have been introduced in detail above. The datasets used are the CIFAR-10 and CIFAR-100 image datasets, the Fashion mnist clothing classification image dataset, and the mnist handwritten digit image dataset. The following is a brief introduction to the models and datasets.
[0107] VGG16 is a classic deep convolutional neural network model proposed by the Visual Geometry Group at the University of Oxford in 2014. The "16" indicates that this model has a total of 16 layers, including 13 convolutional layers and 3 fully connected layers. VGG16 mainly uses small 3x3 convolutional kernels to build the network and extracts image features by stacking multiple convolutional layers. The small convolutional kernels can capture more detailed local features, while the multiple stacks allow the network to learn more complex global features. VGG16 takes an image of a fixed size (usually 224x224 pixels) as input. The image passes through a series of convolutional layers and pooling layers, gradually extracting high-level features, and finally outputs the classification result through the fully connected layers. The core of VGG16 is the convolutional layer, which extracts features such as edges and textures by sliding a small window (convolutional kernel) over the image. After each layer of convolution, an activation function (such as ReLU) is usually connected to introduce non-linearity and enhance the expressive power of the model. The role of the pooling layer is to reduce the size of the image while retaining important feature information. For example, max pooling selects the maximum value from a region and ignores other values. The last few layers of VGG16 are fully connected layers, which combine the features extracted earlier and finally output a probability distribution indicating the likelihood that the input image belongs to each category. VGG16 achieved excellent results in the ImageNet image classification challenge in 2014, demonstrating its powerful feature extraction ability. Although its structure is relatively simple, due to the large number of layers, it requires a large amount of computing resources during training. An important feature of VGG16 is its strong versatility. It can not only be used for image classification tasks but also be applied to other computer vision tasks such as object detection and image segmentation through transfer learning. Generally speaking, VGG16 is a deep learning model with a clear structure and excellent performance, and is very suitable as an experimental model for studying optimizers.
[0108] GoogleNet is a deep convolutional neural network model proposed by the Google team in 2014, which introduced the "Inception module". Its design goal is to reduce the number of model parameters and computational complexity while maintaining high performance, enabling the network to be deeper but more efficient. Traditional convolutional neural networks usually only use a single-size convolutional kernel to extract features, while the Inception module simultaneously uses multiple sizes of convolutional kernels (such as 1x1, 3x3, 5x5) and max pooling operations to extract different-scale features of the image in parallel, and then stitches these features together. This design allows the model to capture both local details and global information of the image in the same layer, thus understanding the image content more comprehensively. The role of the 1x1 convolution is to reduce the dimension of the feature map and the number of channels, thereby reducing the computational complexity of subsequent convolutional operations. The overall structure of GoogleNet is composed of multiple stacked Inception modules, with some auxiliary classifiers interspersed to provide additional gradient signals during training and alleviate the vanishing gradient problem. The input of GoogleNet is an image of a fixed size (usually 224x224 pixels). The image goes through a series of convolutional layers, Inception modules, and pooling layers to gradually extract high-level features, and finally outputs the classification result through a fully connected layer. GoogleNet won the championship in the 2014 ImageNet image classification challenge, with a top-5 error rate of only 6.67%, far lower than other models at that time.
[0109] CIFAR-10 and CIFAR-100 are two classic image classification datasets, collected and released by Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. It provides a standardized dataset for training and testing image classification models. The CIFAR-10 dataset contains 60,000 color images of 32x32 pixels, which are evenly divided into 10 categories, with 6,000 images in each category. The categories include common objects in daily life, such as airplanes, cars, birds, cats, dogs, deer, frogs, horses, ships, and trucks. The CIFAR-10 dataset is divided into 50,000 training images and 10,000 test images. The training set is used to train the model, and the test set is used to evaluate the performance of the model. The CIFAR-100 dataset is similar to CIFAR-10, but it has more categories, a total of 100 categories, with 600 images in each category. These categories cover more fine-grained object classifications, such as apples, mushrooms, oranges, pears, sweet peppers, etc. The CIFAR-100 dataset is also divided into 50,000 training images and 10,000 test images. The characteristics of CIFAR-10 and CIFAR-100 are that the image size is small (32x32 pixels), which makes them very suitable for quick experiments and model verification. CIFAR-10 has fewer categories and relatively simple image content; CIFAR-100 has more categories and a more complex classification task. These two datasets are widely used in deep learning research, especially in the field of image classification, and have become one of the benchmark datasets for evaluating model performance.
[0110] Fashion-MNIST and MNIST are two classic image classification datasets, mainly used to test and verify the performance of image classification models. The MNIST dataset contains 70,000 grayscale images of 28x28 pixels, which are pictures of handwritten digits from 0 to 9, with 7,000 images in each category. The MNIST dataset is divided into 60,000 training images and 10,000 test images. Since the image content of MNIST is very simple (only handwritten digits), it is often used for beginners to get started with deep learning and image classification tasks, and is known as the "Hello World" in the field of deep learning. The Fashion-MNIST dataset is an alternative version of MNIST, which provides a more challenging dataset for researchers. Fashion-MNIST also contains 70,000 grayscale images of 28x28 pixels, but these images are pictures of various clothing items, divided into 10 categories, including T-shirts, trousers, pullovers, dresses, coats, sandals, shirts, sneakers, bags, and ankle boots. The Fashion-MNIST dataset is also divided into 60,000 training images and 10,000 test images. Compared with MNIST, the image content of Fashion-MNIST is more complex, and the classification task is more challenging because it requires the model to learn the detailed features of clothing. The common feature of Fashion-MNIST and MNIST is that the image size is small (28x28 pixels), which makes them very suitable for quick experiments and model verification. These two datasets are widely used in deep learning research, especially in the field of image classification, and have become one of the classic datasets for evaluating model performance.
[0111] Therefore, based on the above datasets and models, experiments are carried out according to the steps of the present invention, and the following experimental results are obtained:
[0112] Under the condition of batch = 128, the present invention updates the parameters of the Vgg16 model based on the CIFAR-10 and CIFAR-100 image datasets through the Adam optimizer, DFGD optimizer, SGD optimizer, AdaMod optimizer, and AdaGC optimizer. The experimental results are shown in Table 1. Further, the present invention also verifies the current accuracy comparison of each optimizer when the number of iteration steps is such that the total training time is similar. The experimental results are shown in Table 2.
[0113] Table 1 Training time and classification accuracy results of different optimizers on the CIFAR-10 and CIFAR-100 datasets
[0114]
[0115] Table 2 Classification accuracy results of different optimizers at the number of iterations with similar overall training time
[0116]
[0117] As can be seen from Table 1 and Table 2, DFGD maintains the highest accuracy while ensuring a relatively short training time, and its performance after balancing is also not inferior to other optimizers. If one wants to train an optimizer with high performance in a relatively short time, the advantages of DFGD are very prominent.
[0118] Under the condition of batch = 128, based on the Fashion mnist clothing classification image dataset and the mnist handwritten digit image dataset, the present invention updates the parameters of the GoogleNet model through the Adam optimizer, DFGD optimizer, SGD optimizer, AdaMod optimizer, and AdaGC optimizer. The experimental results are shown in Table 3. Further, the present invention also verifies the classification accuracy results of different optimizers at the number of iterations with similar overall training time, and the experimental results are shown in Table 4.
[0119] Table 3 Training time and classification accuracy results of different optimizers on the Fashion mnist and mnist datasets
[0120]
[0121] Table 4 Classification accuracy results of different optimizers at the number of iterations with similar overall training time
[0122]
[0123] As can be seen from Table 3 and Table 4, the performance of DFGD on the Fashion mnist and mnist datasets is still excellent. On the Fashion mnist dataset, DFGD achieves a significant improvement in accuracy with a relatively short additional time cost, and also has the highest accuracy after balancing the time; on the mnist dataset, it also achieves the best performance with the shortest training time.
[0124] Therefore, based on the existing optimization algorithms, the present invention proposes the DFGD algorithm, which utilizes the historical gradients without amplifying the differences in batch gradients, effectively reduces the vertical fluctuations near the optimal solution, and realizes local approximation of the local minimum with a shorter update step size, achieving a better effect of suppressing oscillations. Compared with the prior art in classic image recognition tasks and image segmentation tasks, the algorithm accuracy is optimally improved by nearly 1%, which indicates that the use of the DFGD optimizer has a relatively obvious positive effect on the improvement of the model's performance.
[0125] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.
Claims
1. An efficient optimization method based on local gradient information, characterized in that: include: Step 1: Initialize the parameter vector of the convolutional neural network CNN model as the parameter vector θ0, obtain an image data set, input the image data in the image data set into the convolutional neural network CNN model with the parameter vector θ0, and output the predicted classification results of the image data, wherein the predicted classification results include the predicted probabilities that the image data belongs to each category. Based on the predicted classification results and the actual classification results of the image data, calculate the loss function L t (θ0), calculate the loss function L t The gradient of (θ0) Similarly, the parameter vector of the convolutional neural network CNN model is initialized as parameter vector θ1, and the loss function L is calculated t (θ10, calculate the loss function L t (The gradient of θ10 Step 2: Initialize the step size η and time step. The initial value of the time step is 0. Add 1 to the initial value of the time step as the current time step t. Initialize the exponential moving average of the squared gradient v0. The initial value of v0 is 0. Use the exponential moving average of the squared gradient v0 after initialization as the exponential moving average of the squared gradient v at time step t-1. t-1 ; Set the parameter vector θ1 as the parameter vector of time step t; Set the loss function L t (The gradient of θ00 as the gradient of the loss function at time step t-1 The loss function L t (θ1) is the gradient of the loss function at time step t Step 3: Take the exponential moving average v of the squared gradient at time step t-1 t-1 , and the gradient of the loss function at the current time step t Compute the exponential moving average v of the squared gradient at time step t t ; Step 4: Take the exponential moving average v of the squared gradient at time step t by weighting it with coefficient β2 t Perform deviation correction to obtain the corrected weighted average Step 5: Calculate the gradient of the loss function at the current time step t and the gradient of the loss function at time step t-1 The gradient difference ΔL t , and then based on the gradient difference ΔL t Calculate the transformed coefficient ξ t ; Step 6: According to the initialized step size η and the transformed coefficient ξ t , the parameter vector θ at time step t t , the gradient of the loss function at time step t The exponential moving average of the squared gradient of the corrected time step t Calculate the parameter vector θ at time step t+1 t+1 ; Step 7: Input the image data in the image dataset into the parameter vector θ t+1 In the convolutional neural network CNN model, the predicted classification results of the image data are output. Based on the predicted classification results and the actual classification results of the image data, the loss function L is calculated. t (θ t+1 (, judge the loss function L t (θ t+1 ) is less than or equal to the preset threshold, in the loss function L t (θ t+1 ) is less than or equal to the preset threshold, the parameter vector θ t+1 As the final parameter vector of the convolutional neural network CNN model, in the loss function L t (θ t+1 ) is greater than the preset threshold, the time step t is increased by one, and the process returns to step 3.
2. The efficient optimization method based on local gradient information according to claim 1, characterized in that: Step 3 is specifically implemented by the following formula: Among them, β2 is the weighting coefficient.
3. The efficient optimization method based on local gradient information according to claim 1, characterized in that: Step 4 is specifically implemented by the following formula: Among them, β2 is the weighting coefficient.
4. The efficient optimization method based on local gradient information according to claim 1, characterized in that: Step 5 specifically includes: Step 5.1: Take the gradient of the loss function at the current time step t Subtract the gradient of the loss function at time step t-1 The gradient difference is ΔL. t ; Step 5.2: Gradient difference ΔL t Perform nonlinear changes to obtain the transformed coefficient ξ t .
5. The efficient optimization method based on local gradient information according to claim 4, characterized in that: Step 5.1 is specifically expressed by the following formula:
6. The efficient optimization method based on local gradient information according to claim 4, characterized in that: Step 5.2 is specifically implemented by the following formula:
7. The efficient optimization method based on local gradient information according to claim 1, characterized in that: Step 6 is specifically implemented by the following formula: Among them, ε is a smoothing term.