Deep learning adaptive learning rate optimization method based on gradient variance and time sequence attenuation
By introducing gradient dispersion regularity and confidence calibration variance, a dynamic learning rate scheduling mechanism is constructed, which solves the problem of inflexible learning rate adjustment in deep learning models, and improves the convergence stability and generalization ability of the model.
Patent Information
- Application Number
- CN202510376099.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
During the training process, the learning rate adjustment of existing deep learning models is inflexible and is easily disturbed by noise gradients, resulting in gradient variance distortion and rigid timing attenuation, affecting the convergence effect and generalization ability of the model.
By introducing a gradient dispersion regular mechanism to suppress the influence of noise gradients, combining confidence calibration variance and composite dynamic characteristics, a dynamic learning rate scheduling mechanism is built to realize dynamic calibration and timing scheduling of gradient variance volatility.
It improves the adaptive ability of learning rate scheduling, reduces the risk of learning rate oscillation and overfitting during the training process, and is suitable for deep learning training in large-scale models and dynamic task environments.
Smart Images

Figure CN120297350A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning adaptive learning rate optimization, and particularly relates to a deep learning adaptive learning rate optimization method based on gradient variance and temporal decay. Background Art
[0002] During the training process of deep learning models, the setting and dynamic adjustment of the learning rate have always been the key factors affecting the model convergence speed and performance. Existing mainstream adaptive optimizers, such as Adam, RMSProp, Adagrad, etc., generally rely on the first moment (momentum) and second moment (gradient variance) of the gradient to dynamically adjust the learning rate. Among them, the statistics of the gradient variance usually use exponential weighting to smooth the fluctuations, so as to avoid the learning rate oscillation problem caused by the drastic change of the gradient during the training process. However, these optimizers generally have two significant defects: First, the gradient variance is extremely vulnerable to the interference of noisy gradients during the training process. Especially in the initial stage of deep model training, when the data distribution is unstable or in mini-batch training, the high variance of the gradient does not necessarily truly reflect the learning difficulty of the model, but may amplify the variance signal due to noise, resulting in distorted learning rate adjustment; Second, most existing optimizers use a fixed temporal decay mechanism (such as a fixed exponential decay factor) to weight the gradient variance in history. This "rigid" strategy is difficult to achieve dynamic adaptation in different stages of model training (such as rapid exploration in the initial stage of training and stable convergence in the later stage of training), and is prone to insufficient or excessive learning rate decay, thus affecting the overall convergence effect and generalization ability of the model.
[0003] In addition, in recent years, some studies have tried to introduce methods such as cosine annealing, custom learning rate curves or staged scheduling strategies. Although the convergence effect in some tasks has been improved to a certain extent, most of these methods still rely on manually set hyperparameters and lack the adaptive perception ability of the actual gradient fluctuation characteristics and task dynamic changes during the training process, and it is difficult to fundamentally solve the two deep problems of "noisy variance distortion" and "temporal decay rigidity". Therefore, how to design a method that can dynamically and credibly calibrate the volatility of the gradient variance and implement a flexible and task-sensitive temporal scheduling mechanism has become a key technical bottleneck that urgently needs to be broken through in current deep learning optimization methods. The present invention precisely aims at this technical gap and proposes a targeted and innovative optimization strategy, aiming to systematically solve the above problems from two aspects: signal calibration and strategy scheduling. Summary of the Invention
[0004] The purpose of the present invention is to propose a deep learning adaptive learning rate optimization method based on gradient variance and temporal decay, which systematically solves the problems of inflexible learning rate adjustment and unstable model convergence in different training stages of existing methods.
[0005] To achieve the above object, the present invention provides a deep learning adaptive learning rate optimization method based on gradient variance and temporal decay. The method comprises the following steps:
[0006] S1. Obtain the deep learning model, the current batch of training data, and the model parameters at the current training round t. For the current batch of training data, calculate the single-sample gradient of each sample, collect the single-sample gradients of all samples to form a gradient set within the batch, and calculate the global variance signal of the batch gradient within the batch based on the single-sample gradient. Among them, for each parameter dimension, a gradient divergence regularization mechanism is introduced when extracting the global variance signal of the batch gradient to suppress the influence of noisy gradients.
[0007] S2. Perform random perturbation inference on the deep learning model to obtain the predicted output. Use the volatility of the multiple forward inference results of the model in the current batch combined with the predicted output to construct a confidence calibration variance, and calibrate the global variance signal of the batch gradient based on the confidence calibration variance to obtain a confidence trajectory unit. Among them, the confidence trajectory unit includes the global variance signal of the batch gradient, the dynamic confidence, and the confidence calibration variance.
[0008] S3. Based on the confidence trajectory unit and its historical trajectory, construct a composite dynamic feature including historical time step information.
[0009] S4. Jointly input the composite dynamic feature and the confidence calibration variance into a scheduler to generate a dynamic learning rate.
[0010] S5. Update the deep learning model according to the dynamic learning rate to generate new model parameters, combine the gradients calculated on the current batch of training data, complete the update of the model parameters, and transmit the training dynamic information to the next round of signal extraction process through a state feedback mechanism for closed-loop optimization.
[0011] Further, in S1, before calculating the global variance signal of the batch gradient within the batch based on the single-sample gradient, design the gradient divergence as a regularization term to suppress the abnormal gradient fluctuations within the batch. Among them, the gradient divergence is obtained by calculating and analyzing the mean of the batch gradient mean.
[0012] Further, for each parameter dimension, when extracting the global variance signal of the batch gradient, a gradient divergence regularization mechanism is introduced to suppress the influence of noisy gradients. Specifically:
[0013]
[0014] Among them, is the global variance signal of the batch gradient output in this step, Var(·) is the traditional gradient variance between samples under parameter dimension j, and respectively represent g t,i and μ t components in the j-th parameter dimension, λ is the weight hyperparameter of the regularization term, used to control the correction strength of the dispersion on the signal; d is the total number of parameter dimensions; i is the i-th sample; μ t is the mean vector of the gradients of all samples within the current batch.
[0015] Furthermore, the S2 includes:
[0016] Perform S times of random perturbation inference on the deep learning model on the current batch of training data, obtain the variances of the S times of predictions, and obtain a set of prediction outputs;
[0017] Perform uncertainty quantification calculation and analysis on the prediction outputs, and convert the uncertainty of the prediction results into a measurable dynamic confidence level, that is, the dynamic confidence level of this round of signal;
[0018] Perform confidence weighting on the batch gradient global variance signal based on the dynamic confidence level to calibrate the batch gradient global variance signal, and obtain the confidence-calibrated variance output in the current round;
[0019] Use the batch gradient global variance signal, the dynamic confidence level of this round of signal, and the confidence-calibrated variance output in the current round as the confidence trajectory unit.
[0020] Furthermore, the historical trajectory is a signal trajectory unit containing T + 1 historical time steps.
[0021] Furthermore, the S3 includes:
[0022] Construct a historical window based on the historical trajectory and generate a time series model;
[0023] Introduce the dynamic confidence level as the weight of the historical window in the time series model, and fuse the dynamic confidence level and the historical window through a composite trajectory encoder to generate a composite dynamic feature; among them, the composite trajectory encoder is used to jointly extract the time series features of the batch gradient global variance signal, the dynamic confidence level of this round of signal, and the confidence-calibrated variance output in the current round.
[0024] Furthermore, the composite trajectory encoder consists of a standard GRU network or a custom time series convolutional layer and a non-linear activation.
[0025] Furthermore, the scheduler is a deep learning rate scheduling module, which adopts a lightweight non-linear structure internally and is used to realize the fusion mapping of the two-way information of the confidence-calibrated variance output in the current round and the composite dynamic feature.
[0026] Furthermore, the scheduler also includes internally: a non-linear transformation mechanism for variance fluctuation suppression:
[0027] When generating a dynamic learning rate, an additional logarithmic buffer term is designed within the signal input dimension to enhance the robustness against high-variance signals.
[0028] Further, the S5 includes:
[0029] Perform a standard parameter adaptive update operation on the model parameters to generate new model parameters;
[0030] Transmit the training dynamics in the form of signal fluctuation amount to the signal link; wherein, the signal fluctuation amount is used to reflect the local signal sensitivity within the training rounds; the signal fluctuation amount ψ t is expressed as:
[0031]
[0032] wherein, and are the gradient change amounts before and after parameter update, respectively, on the same sample x t,i and N is the batch size;
[0033] wherein, the signal fluctuation amount does not directly participate in the scheduling or calibration mechanism.
[0034] The beneficial technical effects of the present invention are at least as follows:
[0035] Aiming at the deficiencies of the existing adaptive optimizers in terms of gradient variance distortion and temporal scheduling rigidity, the present invention proposes an adaptive learning rate optimization method based on gradient variance and temporal decay, systematically solving the problems such as inflexible learning rate adjustment and unstable model convergence in different training stages of the existing methods. The present invention introduces a new signal processing and dynamic scheduling mechanism in the optimizer design: firstly, by calibrating the original gradient variance signal, it can effectively identify and suppress the variance distortion problem caused by noise or non-stationary training data, improving the authenticity and effectiveness of the variance signal during the training process; secondly, based on the calibrated variance signal, an adaptive scheduling mechanism with temporal awareness ability is constructed, enabling the learning rate to balance between fast convergence and robust training according to the actual needs of the current training stage of the model. Through this innovative design, the present invention not only significantly improves the adaptive ability of the learning rate scheduling, but also effectively reduces the learning rate oscillation and overfitting risk during the training process, and is particularly suitable for deep learning training under large-scale models, dynamic task environments and high-noise data conditions, having wide practical application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The present invention will be further described with reference to the accompanying drawings. However, the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the following drawings without creative efforts.
[0037] Figure 1 It is a flowchart of a deep learning adaptive learning rate optimization method based on gradient variance and temporal decay disclosed in an embodiment of the present invention. Detailed implementation manners
[0038] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0039] As Figure 1 shown, an embodiment of the present invention provides a deep learning adaptive learning rate optimization method based on gradient variance and temporal decay. The method includes the following steps:
[0040] S1. Obtain the deep learning model, the current batch of training data, and the model parameters at the current training round t. For the current batch of training data, calculate the single-sample gradient of each sample, collect the single-sample gradients of all samples to form a gradient set within the batch, and calculate the global variance signal of the batch gradient within the batch based on the single-sample gradient. Among them, for each parameter dimension, a gradient dispersion regularization mechanism is introduced when extracting the global variance signal of the batch gradient to suppress the influence of noisy gradients.
[0041] Specifically, in this step, the input is the deep learning model with its parameters θ t , and the current batch of training data X t ={x t,1 , x t,2 ,..., x t,N}, where N is the batch size. This step is the signal entry of the patent solution. Its main task is to extract a more stable and reliable gradient variance signal for the core pain point of "gradient variance distortion" to serve the subsequent signal credibility and dynamic scheduling mechanisms, and directly lay a high-quality signal foundation for the overall patent process.
[0042] First, for each sample x t in the current batch X t,i , based on the current parameters θ of the model t , calculate its single-sample gradient to obtain:
[0043]
[0044] Among them, g t,i is the gradient vector of the i-th sample, represents the loss function with respect to the parameter θ t , where i = 1, 2, ..., N, d is the dimension of θ t .
[0045] Collect the gradients of all samples to form the gradient set G within the batch t ={g t,1 , g t,2 , ..., g t,N}, which is the input basis for subsequent variance extraction. To overcome the influence of noisy gradients on variance during the training process, this step proposes a "noise-aware regularization variance extraction mechanism". In traditional variance statistics, a "gradient dispersion" is superimposed as a regularization term to effectively suppress abnormal gradient fluctuations within the batch. We first define the batch gradient mean μ t as:
[0046]
[0047] Among them, μ t is the mean vector of the gradients of all samples within the current batch.
[0048] Subsequently, for each parameter dimension j, a novel gradient variance extraction formula is proposed by jointly designing variance and dispersion regularization:
[0049]
[0050] Among them, is the global variance signal of the batch gradient output in this step, Var(·) is the traditional gradient variance between samples under the parameter dimension j, and respectively represent the components of g t,i and μ t in the j-th parameter dimension, and λ is the weight hyperparameter of the regularization term, used to control the correction strength of dispersion on the signal.
[0051] It can be understood that this mechanism has two innovative features: on the one hand, it retains the conventional variance calculation framework to ensure that the output signal is compatible with the subsequent learning rate scheduling logic; on the other hand, it introduces a compensation design for gradient dispersion, enabling to dynamically suppress distorted fluctuations in the presence of noisy gradients and significantly improve signal stability.
[0052] For example, when there is g t,4When the abnormal gradient is [0.5, 0.8, -0.9] etc., the traditional Var will greatly inflate the variance value, while this regularization term captures the abnormal amplitude of |g t,4 -μ t |, automatically corrects the signal, avoids the noise from dominating the variance signal, and ensures its physical interpretability and practical usability.
[0053] It can be understood that this step, as the "signal starting point" in the patent, directly outputs As the input to the "signal credibility calibration" module in the subsequent step 2, it is the basic link to implement the "dynamic learning rate optimization" scheme. Its innovation lies in that, different from the existing optimizers that solely rely on Var, for the first time in the deep learning training process, a "gradient variance regularization mechanism based on dispersion" is introduced, effectively improving the robustness of variance extraction, and having clear technical breakthrough points and application values.
[0054] S2. Conduct random perturbation inference on the deep learning model to obtain the prediction output. Use the volatility of the multiple forward inference results of the model in the current batch combined with the prediction output to construct the confidence calibration variance, and calibrate the batch gradient global variance signal based on the confidence calibration variance to obtain the confidence trajectory unit; where the confidence trajectory unit includes the batch gradient global variance signal, the dynamic confidence, and the confidence calibration variance.
[0055] Specifically, this step receives the batch variance signal output from the previous step As the input basis for the subsequent signal link, aiming at the possible distortion and noise sensitivity problems, this step proposes a "dynamic confidence calibration mechanism based on the instantaneous uncertainty perception of the model" to improve the credibility and stability of the signal and provide more reliable input features for the next "time series modeling".
[0056] This step first designs a patent-specific instantaneous confidence factor u t . This factor does not depend on historical information and is constructed only based on the volatility of the multiple forward inference results of the model in the current batch X t . The specific process is as follows. On X t , perform S random perturbation inferences (such as dropout) on the model to obtain a set of prediction outputs From this, generate u t :
[0057]
[0058] Among them, u t is the dynamic confidence of this round of signal, and Var(·) represents the variance of the S predictions, reflecting the current state of the model in batch Xt For the prediction stability below, β is the regularization scaling factor, usually taking values in [0.5, 2.0]. If Var is low, it indicates that the model has good prediction consistency for the current data X t and u t tends to 1; conversely, if the model prediction is unstable, u t will tend to 0, which is used to limit the direct use of the variance signal.
[0059] Based on u t , we propose a calibration mechanism for signal dynamic compression to perform confidence weighting on the global variance signal σ of the batch gradient t 2 :
[0060]
[0061] Among them, is the confidence-calibrated variance output in the current round, is the moving average within the batch (the mean variance of multiple gradient sub-segments or subgroups within a single-round batch, not the cross-round historical signal), which is used to mitigate the t volatility distortion of when u is low and provides "signal smoothing within the batch".
[0062] For example, when if Var = 0.02, then u t ≈ 0.98, basically equal to if Var = 0.8, then u t ≈ 0.44, will be more inclined to thus reducing the pollution of the noise variance to the subsequent time series link.
[0063] It can be understood that the output of this step is the confidence trajectory unit in the structure of "signal original value + confidence factor + calibrated signal". The innovation of this step is to propose a dynamic confidence generation mechanism "based on the instantaneous prediction stability of the model", which is different from the passive smoothing or historical filtering of variance in existing optimizers. It first introduces a signal credibility control strategy "based on the actual data performance", with stronger noise adaptability and dynamics, ensuring that subsequent dynamic feature modeling is carried out based on high-confidence signals, solving the problems of early variance distortion and abnormal signal inflation during the training unstable period in training, and laying a reliable signal foundation for the adaptive learning rate optimization of the patent solution.
[0064] S3. Based on the confidence trajectory unit and its historical trajectory, construct a composite dynamic feature containing historical time step information.
[0065] Specifically, this step receives the confidence trajectory unit from the previous step The goal is to complete the modeling of "composite temporal features" based on the trajectory within the current time step t and its previous T historical time steps, and output the training dynamic feature h t for subsequent learning rate scheduling use.
[0066] Specifically, we first construct a historical window containing the signal trajectory units of T + 1 historical time steps:
[0067]
[0068] This historical window records u t and the time evolution process of the three core signals, which is a complete expression of the training dynamics.
[0069] For the "non-stationary + multi-signal trajectory" problem in the patent scenario, this step designs a "confidence-aware composite trajectory temporal modeling mechanism", and its innovation lies in introducing the confidence u t as the "signal weight guidance", enabling the system to dynamically focus on the confidence information of different historical time steps.
[0070] The following signal processing formula is used to generate the composite dynamic feature h t :
[0071]
[0072] where Encoder(·) is the "confidence-aware composite trajectory encoder", which internally adopts the mechanism of "signal-confidence joint embedding" to perform and u t joint temporal feature extraction. For example:
[0073] In the specific implementation, Encoder can be composed of a standard GRU network or a custom temporal convolutional layer + non-linear activation;
[0074] In the encoding stage, u t is used as an auxiliary attention factor to guide the model to dynamically focus on the signals of historical time steps with higher confidence, improving the feature anti-noise ability in the low-confidence stage.
[0075] Taking the actual scenario as an example, when the model training enters the unstable stage (such as data drift or new task switching), the confidence sequence u [t-T,t will show local troughs, and the encoder will adjust the attention weights to the signal trajectory accordingly at this stage, giving priority to focusing on the historical time steps with higher confidence and reducing the influence of noise signals on h tNegative impacts.
[0076] The output of this step is h t , as the "training dynamic composite feature" dedicated in the patent solution, is input into the subsequent learning rate scheduling module.
[0077] S4. Jointly input the composite dynamic feature and the confidence calibration variance, and input them into the scheduler to generate a dynamic learning rate.
[0078] Specifically, the goal of this step is to complete the generation of the adaptive learning rate η t at training round t + 1 based on the dynamic timing feature h output in step 3 and the confidence calibration variance t+1 in step 2. This scheduler not only needs to consider the responsiveness to the current signal fluctuations but also integrate the global trend perception of the historical dynamic features to ensure that the learning rate has the dual capabilities of "locally sensitive and globally stable" during the training process.
[0079] Therefore, this step proposes a "signal - history jointly driven deep learning rate scheduling mechanism". The specific approach is to directly use and h t as the joint input and input them into the scheduler f φ to generate the dynamic learning rate. The core of the solution is as follows:
[0080]
[0081] Among them, f φ (·) is a customized deep learning rate scheduling module, which can adopt a lightweight non - linear structure inside, such as one or more layers of non - linear activation networks, to realize the fusion mapping of and h t two - way information, and η t+1 is used as the output learning rate at the current moment t + 1.
[0082] Considering the common problem of "abnormal amplification of signal fluctuations" in the deep learning training process, we design a "non - linear transformation mechanism for variance fluctuation suppression" inside f φ (·), that is, when generating η t+1 , an additional logarithmic buffer term is designed within the signal input dimension to enhance the robustness to high - variance signals. The improved formula is as follows:
[0083]
[0084] Among them, W1 and W2 are learnable weight coefficients, b is the bias term, as a "dynamic smoothing mechanism" of the signal input, is used to relieve the direct impact when there are large fluctuations, and ReLU(·) ensures that ηt+1 Non-negativity
[0085] It can be understood that the innovation of this mechanism lies in that
[0086] the original dynamics of the retained signal makes the learning rate still sensitive to local signals, but through the log(1+·) mechanism, it avoids the problem of excessive learning rate caused by abnormal inflation during the training process ;
[0087] Through the long-term historical dynamic characteristics of h t provides "trend compensation" for the training process, avoiding policy oscillations caused by local signal noise when the learning rate switches or at the early stage of the training phase
[0088] Specific example: If h t = 0.05, and W1 = 2, W2 = 1, b = 0, then η t+1 is approximately equal to ReLU(2·log(1.1)+0.05)≈0.24; if is abnormally amplified to 0.9, the growth of η t+1 is also smoothed, avoiding the direct increase of the learning rate due to an extremely large variance and maintaining the smoothness of the scheduling
[0089] Finally, η t+1 as the output of step 4, is passed to the training process to perform parameter updates, forming a closed loop of "signal - history - scheduling - update".
[0090] S5. Update the deep learning model according to the dynamic learning rate, generate new model parameters, combine the gradients calculated on the current batch of training data, complete the model parameter update, and transmit the training dynamic information to the next round of signal extraction process through the state feedback mechanism for closed-loop optimization
[0091] Specifically, this step receives the learning rate η t+1 generated from the previous step, and based on the parameters θ of the model t at the current round t and the gradients t calculated on the training batch X complete the model parameter update. Subsequently, the system designs a lightweight but efficient state feedback mechanism to provide important training dynamic information for the next round of signal extraction process and construct the overall closed-loop structure of the patent
[0092] First, perform the standard parameter adaptive update operation, and the specific update formula is as follows
[0093]
[0094] where ηt+1 is the dynamic learning rate from the previous step, and θ t are the model parameters for the current round. is the current batch of data X t and the gradient under it.
[0095] After the model completes parameter update, this step proposes a "state feedback mechanism" to transmit the core training dynamics during this round of training to the signal link in the form of a feedback vector. Specifically, we adopt the following innovative feedback metric - the "update-aware signal fluctuation amount" ψ t , which reflects the local signal sensitivity within the training round:
[0096]
[0097] where and are the gradient change amounts before and after parameter update respectively, on the same sample x t,i , and N is the batch size.
[0098] It should be understood that ψ t , as the "overall perturbation metric of the signal field before and after model update", purely reflects the change in signal sensitivity to the current training data after update, and has the following characteristics:
[0099] It does not directly participate in the scheduling or calibration mechanism;
[0100] It only serves as a feedback input in the signal trajectory link to support the perception of the training dynamic fluctuation characteristics during the next round of signal extraction (step 1) or calibration (step 2).
[0101] For example, when ψ t continues to increase during training, it represents that the current update behavior leads to a decrease in signal stability, and the extraction strategy can be adaptively adjusted through the existing mechanism (such as the signal regularization term in step 1) during the next round of signal extraction; conversely, if ψ t remains stable or decreases, it indicates that the system enters a stable update stage and the closed-loop control process forms an effective dynamic regulation.
[0102] The above describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0103] The systems, devices, modules, or units described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0104] For convenience of description, when describing the above devices, they are described as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0105] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0106] This specification is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.
[0109] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0110] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0111] Computer-readable media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0112] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the said element.
[0113] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0114] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For related parts, reference can be made to the description of the method embodiments.
[0115] Finally, it should be noted that what is disclosed in an embodiment of a lithium battery pack chip equalization control platform of the present invention is only a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive learning rate optimization method for deep learning based on gradient variance and temporal decay, characterized in that The method includes the following steps: S1. Obtain the deep learning model, the current batch of training data, and the model parameters at the current training round t. For the current batch of training data, calculate the single-sample gradient of each sample, collect the single-sample gradients of all samples to form the gradient set within the batch, and calculate the global variance signal of the batch gradient within the batch based on the single-sample gradient. Among them, for each parameter dimension, a gradient dispersion regularization mechanism is introduced when extracting the global variance signal of the batch gradient to suppress the influence of noisy gradients; S2. Perform random perturbation inference on the deep learning model to obtain the prediction output. Use the volatility of the multiple forward inference results of the model in the current batch combined with the prediction output to construct the confidence calibration variance, and calibrate the global variance signal of the batch gradient based on the confidence calibration variance to obtain the confidence trajectory unit. Among them, the confidence trajectory unit includes the global variance signal of the batch gradient, the dynamic confidence, and the confidence calibration variance; S3. Based on the confidence trajectory unit and its historical trajectory, construct a composite dynamic feature including historical time step information; S4. Jointly input the composite dynamic feature and the confidence calibration variance into the scheduler to generate the dynamic learning rate; S5. Update the deep learning model according to the dynamic learning rate to generate new model parameters, combine the gradients calculated on the current batch of training data to complete the update of the model parameters, and transmit the training dynamic information to the next round of signal extraction process through the state feedback mechanism for closed-loop optimization.
2. The adaptive learning rate optimization method for deep learning based on gradient variance and temporal decay according to claim 1, wherein In S1, before calculating the global variance signal of the batch gradient within the batch based on the single-sample gradient, design the gradient dispersion as a regularization term to suppress the abnormal gradient fluctuations within the batch. Among them, the gradient dispersion is obtained by calculating and analyzing the mean of the batch gradient mean.
3. The adaptive learning rate optimization method for deep learning based on gradient variance and temporal decay according to claim 1, wherein For each parameter dimension, when extracting the global variance signal of the batch gradient, introduce a gradient dispersion regularization mechanism to suppress the influence of noisy gradients, specifically: Among them, is the batch gradient global variance signal output in this step, Var(·) is the traditional gradient variance between samples under parameter dimension j, and respectively represent the components of g t,i and μ t in the j-th parameter dimension. λ is the weight hyperparameter of the regularization term, used to control the correction strength of the dispersion on the signal; d is the total number of parameter dimensions; i is the i-th sample; μ t is the mean vector of the gradients of all samples within the current batch.
4. A deep learning adaptive learning rate optimization method based on gradient variance and temporal decay according to claim 1, characterized in that, S2 includes: Perform S random perturbation inferences on the deep learning model on the current batch of training data to obtain the variances of S predictions, which are a set of prediction outputs; Perform uncertainty quantification calculation and analysis on the prediction output, and convert the uncertainty of the prediction result into a measurable dynamic confidence, that is, the dynamic confidence of this round of signal; Perform confidence weighting on the global variance signal of the batch gradient based on the dynamic confidence to calibrate the global variance signal of the batch gradient to obtain the confidence calibration variance output in the current round; Use the global variance signal of the batch gradient, the dynamic confidence of this round of signal, and the confidence calibration variance output in the current round as the confidence trajectory unit.
5. A deep learning adaptive learning rate optimization method based on gradient variance and temporal decay according to claim 1, characterized in that, The historical trajectory is a signal trajectory unit including T + 1 historical time steps.
6. The adaptive learning rate optimization method for deep learning based on gradient variance and temporal decay according to claim 5, characterized in that S3 includes: Construct a historical window based on the historical trajectory to generate a time series model; Introduce dynamic confidence as the weight of the historical window in the timing model, and fuse the dynamic confidence and the historical window through a composite trajectory encoder to generate composite dynamic features; wherein, the composite trajectory encoder is used to jointly extract timing features from the batch gradient global variance signal, the dynamic confidence of the current round signal, and the confidence calibration variance output in the current round.
7. An adaptive learning rate optimization method for deep learning based on gradient variance and temporal decay according to claim 6, characterized in that, The composite trajectory encoder consists of a standard GRU network or a custom timing convolutional layer and a non-linear activation.
8. An adaptive learning rate optimization method for deep learning based on gradient variance and temporal decay according to claim 1, characterized in that The scheduler is a deep learning rate scheduling module, which adopts a lightweight non-linear structure internally and is used to realize the fusion mapping of the two-way information of the confidence calibration variance output in the current round and the composite dynamic features.
9. The adaptive learning rate optimization method for deep learning based on gradient variance and temporal decay according to claim 8, characterized in that The scheduler also includes internally: a non-linear transformation mechanism for variance fluctuation suppression: When generating the dynamic learning rate, an additional logarithmic buffer term is designed within the signal input dimension to enhance the robustness to high-variance signals.
10. A deep learning adaptive learning rate optimization method based on gradient variance and temporal decay according to claim 1, characterized in that The S5 includes: Perform a standard parameter adaptive update operation on the model parameters to generate new model parameters; Transmit the training dynamics to the signal link in the form of the signal fluctuation amount; wherein, the signal fluctuation amount is used to reflect the local signal sensitivity within the training round; the signal fluctuation amount ψ t is expressed as: wherein, and are the gradient change amounts before and after parameter update, respectively, on the same sample x t,i , and N is the batch size. Among them, the signal fluctuation amount does not directly participate in the scheduling or calibration mechanism.
Citation Information
Cited By
Aluminum electrolysis whole-process energy-saving scheduling method fused with deep reinforcement learning algorithm
CN122303972A