Transformer fault identification method
By introducing multi-scale feature extraction module and attention mechanism into the convolutional neural network, an improved convolutional neural network is solved, and the problem of inaccurate transformer fault recognition in actual scenarios in the existing technology is solved, achieving higher recognition accuracy and model generalization.
Patent Information
- Application Number
- CN202510030209.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to adapt to the transformer state threshold range in actual scenarios, resulting in insufficient identification of transformer faults.
By adding Fourier transform module, scale recognition module, multi-scale adaptive convolution module, multi-contact attention mechanism module and scale aggregation module to the convolution neural network, an improved convolution neural network is built and the recognition accuracy of different states of the transformer is improved using a combination of multi-scale features and deep learning.
It significantly improves the accuracy of transformer fault identification and generalization capabilities of the model, and can more effectively capture the correlations at different scales and overcome the shortcomings of traditional machine learning.
Smart Images

Figure CN119939470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and in particular to a transformer fault identification method. Background Art
[0002] As an indispensable core equipment in the power network, the transformer plays a vital role in ensuring the safety and stable operation of the power system. Once a fault occurs and the power supply is interrupted, it often brings immeasurable huge economic losses. The sound signal emitted by the power equipment during operation can effectively reflect the operating status of the equipment. Using non-contact sound sensors to collect the operating sound signals of power equipment, and using sound signals to analyze and diagnose the operating status has the advantages of flexible installation, rich information and high reliability. It is an important development direction for online monitoring of power equipment status.
[0003] With the development of technology, the diagnosis method of transformer faults has been deeply studied. The existing transformer voiceprint fault detection technology mainly uses short-time Fourier transform to extract feature information from transformer voiceprints, and models the feature information based on traditional machine learning (such as support vector machine, Gaussian mixture model, etc.), and uses the transformer state threshold range learned by the machine learning model to classify and identify the transformer voiceprint.
[0004] The defects of the above-mentioned existing technologies are: due to the different influences of environmental and other equipment noise in different actual scenarios, traditional machine learning is difficult to adapt to the transformer state threshold range, resulting in inaccurate transformer fault identification. Summary of the invention
[0005] Based on this, it is necessary to provide a transformer fault identification method to address the above technical problems. By combining multi-scale features and deep learning, the recognition accuracy of different transformer states (i.e. normal or different faults) in actual scenarios can be greatly improved, which has the technical advantages of high diagnostic accuracy and good model generalization.
[0006] An embodiment of the present invention provides a transformer fault identification method, comprising:
[0007] Obtain historical voiceprint signals of the transformer and the fault types corresponding to the historical voiceprint signals;
[0008] A Fourier transform module is added to the input layer of the convolutional neural network, and a scale recognition module, a multi-scale adaptive convolution module, a multi-touch attention mechanism module and a scale aggregation module are added to the convolutional layer of the convolutional neural network. The 1×1 convolution kernel of the convolutional layer in the convolutional neural network is replaced with a 4×4 convolution kernel, and the output ends of the Fourier transform module and the scale aggregation module are connected through the function module Softmax to construct an improved convolutional neural network.
[0009] Inputting the historical voiceprint signal into the improved convolutional neural network to obtain the fault recognition result of the historical voiceprint signal; training the improved convolutional neural network based on the fault recognition result of the historical voiceprint signal and the fault type corresponding to the historical voiceprint signal;
[0010] The real-time voiceprint signal of the transformer to be tested is input into the trained improved convolutional neural network, and the real-time voiceprint signal is fast Fourier transformed through the Fourier transform module to obtain the periodic characteristics of the real-time voiceprint signal; the correlation of the periodic characteristics is captured through the scale recognition module to obtain the correlation of the real-time voiceprint signal at multiple scales; the adjacency matrix of the correlation is learned through the multi-scale adaptive convolution module to obtain the sequence correlation of the real-time voiceprint signal; the sequence correlation is captured through the multi-contact attention mechanism module to obtain the multi-scale characteristics of the real-time voiceprint signal, and the multi-scale features are aggregated through the scale aggregation module to obtain the fusion characteristics of the real-time voiceprint signal; the fusion characteristics are predicted through the function module Softmax to obtain the fault recognition result of the real-time voiceprint signal, and the fault type of the transformer is judged according to the fault recognition result.
[0011] Optionally, the real-time voiceprint signal of the transformer to be tested is preprocessed by normalization, framing and windowing before being input into the trained convolutional neural network. The specific process includes:
[0012] Normalize the real-time voiceprint signal of the transformer to be tested to obtain voiceprint signal data in the same dimension;
[0013] The voiceprint signal data in the same dimension is framed and processed, and the formula is:
[0014]
[0015] Among them, num is the number of sub-frames, m is the frame length, s is the frame shift, and l is the total signal length;
[0016] The voiceprint signal data after frame processing is windowed, and the formula is:
[0017]
[0018] Among them, N is the window length and n is the index.
[0019] Optionally, a fast Fourier transform is performed on the real-time voiceprint signal through a Fourier transform module, which specifically includes:
[0020] Taking periodicity as the source of scale, the Fast Fourier Transform (FFT) is used to detect the prominent periodicity as the time scale:
[0021] F = Avg(Amp(FFT(Xemb)))
[0022] Where, FFT() is the fast Fourier transform of the input data, Amp() is the amplitude value of the frequency point after calculating the fast Fourier transform FFT, and Avg() is the averaging function.
[0023] Optionally, adjacency matrix learning is performed on the correlation through a multi-scale adaptive convolution module, which specifically includes:
[0024] The tensor corresponding to the i-th scale is projected back to a tensor containing N variables through a linear transformation, where N represents the number of time series; the projection is expressed as follows:
[0025] H i =W i X i
[0026] in, is a learnable weight matrix;
[0027] Generate two training parameters and Based on training parameters and Construct the adjacency matrix A i , the formula is:
[0028]
[0029] Among them, i is the scale, SoftMax() is the weight function;
[0030] Get the adjacency matrix A of the i-th scale i Finally, the Mixhop convolution method is used to capture the correlation between sequences. The formula is:
[0031]
[0032] in, is the output after fusion at scale i, σ() is the activation function, P is a set of hyperparameters consisting of integer adjacent powers, (A i ) j is the adjacency matrix A i Raise itself j times, || is a column-level join.
[0033] Optionally, the correlation between sequences is captured through a multi-touch attention mechanism module, which specifically includes:
[0034] At each time scale, a multi-touch attention mechanism is used to capture the correlation within the sequence. For each time scale tensor The expression is as follows:
[0035]
[0036] Among them, MHA s () is the multi-touch attention function in the scale dimension.
[0037] Optionally, the multi-scale features are aggregated through a scale aggregation module, which specifically includes:
[0038] Get k tensors of different scales Reshape each scale tensor back into a 2D matrix
[0039] Aggregating different scales according to amplitude, the formula is:
[0040]
[0041] Among them, F f1 ,…,F fk is the amplitude corresponding to each scale calculated using FFT, is the amplitude.
[0042] Optionally, the fusion feature is predicted by the function module Softmax, which specifically includes:
[0043] Linear projection is used in both the time dimension and the variable dimension. Convert to The conversion expression is:
[0044]
[0045] in, is the prediction result, T is the prediction range, W s , W t and b are learning parameters, and t is the future time step.
[0046] Optionally, the method further includes: adjusting the learning rate of each parameter by using the Adam optimization algorithm before training the model, and the specific process includes:
[0047] Calculate the first-order moment estimate, the formula is:
[0048] m t =β1·m t-1 +(1-β1)·g t
[0049] Among them, m t is the first-order moment estimate at time step t, m t-1 is the first-order moment estimate of the previous time step, g t is the gradient of the current time step, β1 is the first-order hyperparameter;
[0050] Calculate the unbiased estimate of the first-order moment estimate, the formula is:
[0051]
[0052] Calculate the second-order moment estimate, the formula is:
[0053] v t =β2·v t-1 +(1-β2)·(g t ) 2
[0054] Among them, v t is the second-order moment estimate of the time step, v t-1 is the second-order moment estimate of the previous time step, and β2 is the second-order hyperparameter;
[0055] Calculate the unbiased estimate of the second-order moment estimate, the formula is:
[0056]
[0057] The parameters are updated based on the unbiased estimate of the second-order moment estimate, and the formula is:
[0058]
[0059] Among them, θ t is the current parameter, α is the learning rate, and ε is a constant.
[0060] Compared with the prior art, the transformer fault identification method provided by the embodiment of the present invention has the following beneficial effects:
[0061] The present invention adds a Fourier transform module to the input layer of a convolutional neural network, adds a scale recognition module, a multi-scale adaptive convolution module, a multi-touch attention mechanism module and a scale aggregation module to the convolutional layer of the convolutional neural network, replaces the 1×1 convolution kernel of the convolutional layer in the convolutional neural network with a 4×4 convolution kernel, and connects the output ends of the Fourier transform module and the scale aggregation module through a function module Softmax to construct an improved convolutional neural network; utilizes the fast Fourier transform FFT technology to extract the periodic features in the time series, and maps them to a space closely related to the time scale, which can effectively capture the correlation at different scales, thereby overcoming the problem that traditional machine learning is difficult to adapt to the transformer state threshold range, and improving the accuracy of transformer fault identification.
[0062] In addition, the model's multi-scale adaptive convolution module can dynamically learn an exclusive adjacency matrix for each time scale, enabling the model to capture the correlation between sequences related to a specific scale; the multi-contact attention mechanism can synchronously capture the internal correlation of the sequence, further enhancing the accuracy of transformer fault identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1A schematic flow chart of a transformer fault identification method provided in an embodiment;
[0064] Figure 2 A schematic diagram of the structure of a transformer fault identification method provided in an embodiment. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0066] In one embodiment, a transformer fault identification method is provided, such as Figure 1 As shown, the method includes:
[0067] Step 1
[0068] The real-time voiceprint signal of the transformer to be tested is preprocessed by normalization, framing and windowing before being input into the trained convolutional neural network.
[0069] The specific process of step 1 includes:
[0070] Step 1.1: Normalize the real-time voiceprint signal of the transformer to be tested to obtain voiceprint signal data in the same dimension;
[0071] Step 1.2: Perform the following processing on the normalized data:
[0072] (1) Framing: Generally, considering that the collected voiceprint signal satisfies short-term stability, the voiceprint signal data in the same dimension is framed. The expressions of the number of frames num, frame length m, frame shift s, and total signal length l are as follows:
[0073]
[0074] (2) Windowing: In order to make the signal more continuous globally and avoid the Gibbs effect, the voiceprint signal data after frame processing is windowed. The window function usually sampled is the Hamming window, which is expressed as follows:
[0075]
[0076] Wherein, L represents the window length and n represents the index.
[0077] Step 2
[0078] Obtain historical voiceprint signals of the transformer and the fault types corresponding to the historical voiceprint signals;
[0079] A Fourier transform module is added to the input layer of the convolutional neural network, and a scale recognition module, a multi-scale adaptive convolution module, a multi-touch attention mechanism module and a scale aggregation module are added to the convolutional layer of the convolutional neural network. The 1×1 convolution kernel of the convolutional layer in the convolutional neural network is replaced by a 4×4 convolution kernel, and the output ends of the Fourier transform module and the scale aggregation module are connected through the function module Softmax to construct an improved convolutional neural network.
[0080] The historical voiceprint signal is input into the improved convolutional neural network to obtain the fault recognition result of the historical voiceprint signal. Based on the fault recognition result of the historical voiceprint signal and the fault type corresponding to the historical voiceprint signal, the improved convolutional neural network is trained.
[0081] The specific process of step 2 includes:
[0082] Step 2.1: Input Embedding and Concatenation (Input)
[0083] In the context of multivariate time series forecasting, assuming the number of variables is N, given the input data It represents the observations of a look-back window, which contains the observations of each variable i at time τ from tL to t-1. L represents the size of the look-back window, and t represents the starting position of the forecast window. The goal of the time series forecasting task is to predict the future values of N variables in the next T time steps. The predicted value is given by It means that it includes the effects of each time point τ on all variables from t to t+T-1. value.
[0084] Embed N variables at the same time step into a matrix of size d m In the vector: X t-L:t →X emb ,in The unified input representation is used to generate embedding. The expression is as follows:
[0085]
[0086] For input Normalize it and get Then use a one-dimensional convolution filter (kernel width is 3, step size is 1) to Projected to a d m The parameter α acts as a balancing factor to adjust the magnitude between the scalar projection and the local or global embedding. represents the position embedding of the input, Represents a learnable global temporal embedding with bounded capacity.
[0087] Set X0 = X emb, using the residual method to achieve multi-scale fusion. emb Represents the projection of the original input into the deep features through the embedding layer. In the lth layer of the multi-scale fusion model, the input is The expression is as follows:
[0088] X l =ScaleGraphBlock(X l-1 )+X l-1
[0089] Among them, ScaleGraphBlock represents the operations and calculations that constitute the core functions of the multi-scale fusion model.
[0090] Step 2.2: Scale Identification (Fourier Transform Module FFT, Scale Identification Module)
[0091] Taking periodicity as the source of scale, Fast Fourier Transformation (FFT) is used to detect prominent periodicity as the time scale:
[0092] F = Avg(Amp(FFT(Xemb)))
[0093] Where FFT() is the fast Fourier transform of the input data, which converts the time series from the time domain to the frequency domain. In the frequency domain, the periodic pattern of the data can be expressed as the amplitude of different frequency domains. Amp() is the amplitude value of the frequency point after calculating the fast Fourier transform FFT. The larger the amplitude, the more significant the periodic component of the frequency is in the time series. The vector F∈R L Contains the average amplitude of all frequencies, which is the amplitude at d m The dimension is averaged using the Avg() function.
[0094] Step 2.3: Multi-scale adaptive convolution (multi-scale adaptive convolution module)
[0095] The tensor corresponding to the i-th scale is projected back to a tensor containing N variables through a linear transformation, where N represents the number of time series. The projection is expressed as follows:
[0096] H i =W i X i
[0097] in, is the learnable weight matrix.
[0098] Two training parameters are generated during the learning process and and After multiplying these two parameter matrices, the adjacency matrix A is obtained according to the following formula i (Use SoftMax function to normalize the weights between different nodes):
[0099]
[0100] Get the adjacency matrix A of the i-th scale i Finally, the Mixhop convolution method is used to capture the correlation between sequences. The convolution is defined as follows:
[0101]
[0102] in, is the output after fusion at scale i, σ() is the activation function, P is a set of hyperparameters consisting of integer adjacent powers, (A i ) j is the adjacency matrix A i Raised j times, || is a column-level connection, connecting the intermediate variables generated during each iteration.
[0103] Then continue to use the multi-layer perceptron Projecting back to a 3D tensor
[0104] Step 2.4: Multi-touch attention (multi-touch attention mechanism module) and scale fusion (scale aggregation module)
[0105] At each time scale, a multi-touch attention mechanism is used to capture the correlation within the sequence. For each time scale tensor The expression is as follows:
[0106]
[0107] Among them, MHA s () is the multi-touch attention function in the scale dimension.
[0108] Before entering the next layer, k tensors of different scales need to be integrated First, reshape the tensor of each scale back into a 2D matrix Then, different scales are aggregated according to amplitude:
[0109]
[0110] In the formula, F f1 ,…,F fk is the amplitude corresponding to each scale calculated using FFT, is the amplitude.
[0111] Step 2.5: Predict output (Softmax function module)
[0112] To make predictions, the model uses linear projections in both the time dimension and the variable dimension. Convert to The conversion expression is as follows:
[0113]
[0114] in, is the prediction result, T is the prediction range, W s , W t and b are learning parameters, and t is the future time step. The matrix W s Perform a linear projection along the variable dimension, while W t The same operation is performed along the time dimension. It is the predicted data.
[0115] In the process, firstly, the matrix W s The features output by the model are mapped to the same dimension as the number of original variables, and then passed through the matrix W t These features are mapped to the prediction time range T. In this way, the model can generate predicted values for each variable between future time steps t and t+T from the comprehensive information extracted from multi-scale features.
[0116] This approach allows the model to exploit the deep temporal and inter-variable relationships learned throughout the training process, allowing for effective time series forecasting.
[0117] Furthermore, before training the model, the Adam optimization algorithm is used to adaptively adjust the learning rate of each parameter, thereby accelerating the convergence speed and improving the performance of the model.
[0118] The Adam algorithm is used to adjust the learning rate of each parameter, enhance the stability of the model, and then train the model. Based on multi-scale fusion feature reasoning, the inference probability of each category is output, and the voiceprint category corresponding to the largest inference probability is determined, and the voiceprint noise recognition result is output.
[0119] The specific steps are as follows:
[0120] (1) Calculate the first-order moment estimate:
[0121] m t =β1·m t-1 +(1-β1)·g t
[0122] Where m t is the first-order moment estimate at time step t, m t-1 is the estimate of the previous time step, g t is the gradient of the current time step, and β1 is a hyperparameter, usually set to 0.9.
[0123] (2) Calculate the unbiased estimate of the first-order moment estimate:
[0124]
[0125] (3) Calculate the second-order moment estimate:
[0126] v t =β2·v t-1 +(1-β2)·(g t ) 2
[0127] Among them, v t is the second-order moment estimate of the time step, v t-1 is the estimate of the previous time step and β2 is a hyperparameter, usually set to 0.999.
[0128] (4) Calculate the unbiased estimate of the second-order moment estimate:
[0129]
[0130] (5) Update parameters
[0131]
[0132] Among them, θ t is the current parameter, α is the learning rate, and ε is a constant, usually set to 1e -8 , in order to ensure numerical stability.
[0133] The steps are repeated in each iteration until the model's performance on the training set reaches a satisfactory level or the preset number of iterations is reached. In this way, the Adam algorithm can adaptively adjust the learning rate of each parameter, thereby accelerating convergence and improving the performance of the model.
[0134] Step 3
[0135] The real-time voiceprint signal of the transformer to be tested is input into the trained improved convolutional neural network, and the real-time voiceprint signal is subjected to fast Fourier transform through the Fourier transform module to obtain the periodic characteristics of the real-time voiceprint signal.
[0136] The scale recognition module is used to capture the correlation of periodic features and obtain the correlation of real-time voiceprint signals at multiple scales.
[0137] The adjacency matrix of the correlation is learned through a multi-scale adaptive convolution module to obtain the sequence correlation of the real-time voiceprint signal.
[0138] The correlation between sequences is captured through the multi-touch attention mechanism module to obtain the multi-scale features of the real-time voiceprint signal. The multi-scale features are aggregated through the scale aggregation module to obtain the fusion features of the real-time voiceprint signal.
[0139] The fusion features are predicted through the function module Softmax to obtain the fault recognition result of the real-time voiceprint signal, and the fault type of the transformer is determined based on the fault recognition result.
[0140] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A transformer fault identification method, characterized in that: include: Obtain historical voiceprint signals of the transformer and the fault types corresponding to the historical voiceprint signals; A Fourier transform module is added to the input layer of the convolutional neural network, and a scale recognition module, a multi-scale adaptive convolution module, a multi-touch attention mechanism module and a scale aggregation module are added to the convolutional layer of the convolutional neural network. The 1×1 convolution kernel of the convolutional layer in the convolutional neural network is replaced with a 4×4 convolution kernel, and the output ends of the Fourier transform module and the scale aggregation module are connected through the function module Softmax to construct an improved convolutional neural network. Input the historical voiceprint signal into the improved convolutional neural network to obtain the fault recognition result of the historical voiceprint signal; Based on the fault recognition results of historical voiceprint signals and the fault types corresponding to the historical voiceprint signals, the improved convolutional neural network is trained; The real-time voiceprint signal of the transformer to be tested is input into the trained improved convolutional neural network, and the real-time voiceprint signal is fast Fourier transformed through the Fourier transform module to obtain the periodic characteristics of the real-time voiceprint signal; The scale recognition module is used to capture the correlation of periodic features and obtain the correlation of real-time voiceprint signals at multiple scales. The multi-scale adaptive convolution module is used to learn the adjacency matrix of the correlation and obtain the sequence correlation of real-time voiceprint signals. The multi-touch attention mechanism module is used to capture the correlation between sequences and obtain the multi-scale features of the real-time voiceprint signal. The multi-scale features are aggregated through the scale aggregation module to obtain the fusion features of the real-time voiceprint signal. The fusion features are predicted through the function module Softmax to obtain the fault recognition result of the real-time voiceprint signal, and the fault type of the transformer is determined based on the fault recognition result.
2. A transformer fault identification method according to claim 1, characterized in that: The real-time voiceprint signal of the transformer to be tested is input into the trained convolutional neural network for normalization, framing and windowing preprocessing. The specific process includes: Normalize the real-time voiceprint signal of the transformer to be tested to obtain voiceprint signal data in the same dimension; The voiceprint signal data in the same dimension is framed and processed, and the formula is: Among them, num is the number of sub-frames, m is the frame length, s is the frame shift, and l is the total signal length; The voiceprint signal data after frame processing is windowed, and the formula is: Among them, N is the window length and n is the index.
3. A transformer fault identification method according to claim 1, characterized in that: The fast Fourier transform of the real-time voiceprint signal by the Fourier transform module specifically includes: Taking periodicity as the source of scale, the Fast Fourier Transform (FFT) is used to detect the prominent periodicity as the time scale: F = Avg(Amp(FFT(Xemb))) Where, FFT() is the fast Fourier transform of the input data, Amp() is the amplitude value of the frequency point after calculating the fast Fourier transform FFT, and Avg() is the averaging function.
4. A transformer fault identification method according to claim 1, characterized in that: The adjacency matrix learning of the correlation by the multi-scale adaptive convolution module specifically includes: The tensor corresponding to the i-th scale is projected back to a tensor containing N variables through a linear transformation, where N represents the number of time series; the projection is expressed as follows: H i =W i X i Among them, H i ∈R N×Si×fi , W i ∈R N×de is a learnable weight matrix; Generate two training parameters and Based on training parameters and Construct the adjacency matrix A i , the formula is: Among them, i is the scale, SoftMax() is the weight function; Get the adjacency matrix A of the i-th scale i Finally, the Mixhop convolution method is used to capture the correlation between sequences. The formula is: in, is the output after fusion at scale i, σ() is the activation function, P is a set of hyperparameters consisting of integer adjacent powers, (A i ) j is the adjacency matrix A i Raise itself j times, || is a column-level join.
5. A transformer fault identification method according to claim 1, characterized in that: The multi-touch attention mechanism module is used to capture the correlation between sequences, which specifically includes: At each time scale, a multi-touch attention mechanism is used to capture the correlation within the sequence. For each time scale tensor The expression is as follows: Among them, MHA s () is the multi-touch attention function in the scale dimension.
6. A transformer fault identification method according to claim 1, characterized in that: The multi-scale feature aggregation by the scale aggregation module specifically includes: Get k tensors of different scales Reshape each scale tensor back into a 2D matrix Aggregating different scales according to amplitude, the formula is: Among them, F f1 ,…,F fk is the amplitude corresponding to each scale calculated using FFT, is the amplitude.
7. A transformer fault identification method according to claim 1, characterized in that: The fusion feature is predicted by the function module Softmax, which specifically includes: Linear projection is used in both the time dimension and the variable dimension. Convert to The conversion expression is: in, is the prediction result, T is the prediction range, W s , W t and b are learning parameters, and t is the future time step.
8. A transformer fault identification method according to claim 1, characterized in that: Also includes: Before training the model, the Adam optimization algorithm is used to adjust the learning rate of each parameter. The specific process includes: Calculate the first-order moment estimate, the formula is: m t =β1·m t-1 +(1-β1)·g t Among them, m t is the first-order moment estimate at time step t, m t -1 is the first-order moment estimate of the previous time step, g t is the gradient of the current time step, β1 is the first-order hyperparameter; Calculate the unbiased estimate of the first-order moment estimate, the formula is: Calculate the second-order moment estimate, the formula is: v t =β2·v t-1 +(1-β2)·(g t ) 2 Among them, v t is the second-order moment estimate of the time step, v t-1 is the second-order moment estimate of the previous time step, and β2 is the second-order hyperparameter; Calculate the unbiased estimate of the second-order moment estimate, the formula is: The parameters are updated based on the unbiased estimate of the second-order moment estimate, and the formula is: Among them, θ t is the current parameter, α is the learning rate, and ε is a constant.