Multi-scale width ELM lithium battery RUL prediction model and method based on elastic network

By adopting the multi-scale width ELM model of elastic network in the lithium battery RUL prediction model, combining CEEMDAN, hollow convolution and multi-head attention mechanism, the accuracy and calculation complexity problems of traditional battery life prediction methods when processing complex data are solved, and high-precision, low-complexity and real-time prediction effects are achieved.

CN119986389APending Publication Date: 2025-05-13XUZHOU NORMAL UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510171310.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional machine learning battery life prediction methods face nonlinearity, single scale and noise interference problems when processing complex battery state data, resulting in limited prediction accuracy, and deep learning models have problems such as high computational complexity and insufficient real-time performance.

Method used

A multi-scale width ELM lithium battery RUL prediction model based on elastic network is adopted. This model includes a CEEMDAN multi-modal decomposition module, a multi-scale feature extraction module, a multi-limit learning machine (ELM) module and an elastic network regularized width neural network. Through the combination of these modules, multi-scale features are extracted and non-linear transformation is performed to generate the final prediction result.

Benefits of technology

It effectively reduces the amount of training data, improves prediction accuracy, significantly reduces the computational complexity, and ensures real-time application of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119986389A_ABST
    Figure CN119986389A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale width ELM lithium battery RUL prediction model and method based on an elastic network. The prediction model comprises a CEEMDAN multi-mode decomposition module, a multi-extreme learning machine module and an elastic network regularization width neural network module; a CEEMDAN multi-mode decomposition module decomposes battery recession data into a plurality of intrinsic mode IMFs, extracts multi-scale feature information through a multi-scale feature extraction module, inputs the extracted multi-scale feature information into an ELM module for training, selects a plurality of better ELMs based on a plurality of performance indexes, and finally, takes prediction results of the plurality of ELMs as new features, so that the performance of the battery is improved. And inputting the two features into an elastic network regularization width neural network module through enhancement features of width expansion to obtain a final prediction result. According to the method, the training data volume can be effectively reduced, the calculation complexity can be remarkably reduced while the prediction precision is improved, and the real-time application of the system is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remaining useful life (RUL) prediction of lithium batteries, and specifically refers to a multi-scale width ELM lithium battery RUL prediction model and method based on elastic network. Background Art

[0002] With the widespread application of agricultural drones in precision agriculture, their efficient data collection and operation capabilities have received great attention. Drones use a variety of sensors such as multispectral, infrared, and RGB cameras to collect farmland environmental data, providing strong support for tasks such as crop monitoring, pest and disease warning, and precision fertilization. However, the efficient operation of agricultural drones depends on the performance and remaining life (RUL) of their batteries, especially in scenarios that require long-term autonomous flight and the execution of complex tasks. The remaining life prediction of the battery can not only help to reasonably arrange flight missions, but also effectively avoid flight interruptions or data loss caused by battery failure. Therefore, accurately predicting the remaining life of the battery is an important task to improve the operating efficiency of drones and ensure their safety.

[0003] However, traditional machine learning battery life prediction methods often face problems of nonlinearity, single scale and noise interference when processing complex battery status data, resulting in limited prediction accuracy. Deep learning models have problems such as large model training data requirements, high computational complexity and lack of real-time performance. Summary of the invention

[0004] The purpose of the present invention is to provide a multi-scale width ELM lithium battery RUL prediction model and method based on an elastic network, which can effectively reduce the amount of training data. The method can significantly reduce the computational complexity while improving the prediction accuracy and ensure the real-time application of the system.

[0005] To achieve the above-mentioned purpose, the multi-scale width ELM lithium battery RUL prediction model based on elastic network of the present invention comprises a CEEMDAN multi-modal decomposition module, a multi-scale feature extraction module, a multi-extreme learning machine (ELM) module, and an elastic network regularized width neural network. The elastic network regularized width neural network comprises a feature width expansion module and an elastic network regression meta-learner. The CEEMDAN multi-modal decomposition module decomposes battery degradation data into multiple intrinsic mode IMFs, and extracts multi-scale feature information through a multi-scale feature extraction module. The extracted multi-scale feature information is input into the ELM module for training, and multiple ELMs with better performance indicators are selected. Finally, the prediction results of multiple ELMs are used as new features and enhanced features formed by mapping to high-dimensional space through nonlinear transformation of width expansion, and the two are input into the elastic network regularized width neural network module to obtain the final prediction result.

[0006] As a further solution of the present invention: the CEEMDAN multimodal decomposition module decomposes the intrinsic mode function IMF from the complex non-stationary signal through the CEEMDAN algorithm, thereby extracting feature information of different time scales. The final decomposition result of the original signal x(t) is expressed as

[0007]

[0008] Where N represents the number of modal functions obtained, IMF i (t) is the i-th intrinsic mode function of the decomposition, r(t) represents the residual sequence that cannot be decomposed; CEEMDAN is used to decompose the battery degradation data into eigenmodes of different frequencies. The implementation steps are as follows:

[0009] Add zero-mean Gaussian white noise n(t) to the original signal x(t) to get the noisy signal x n (t), expressed as

[0010] x n (t) = x(t) + n(t) (2)

[0011] Where n(t) is zero-mean Gaussian noise with variance Used to perturb the signal and improve the stability of the decomposition;

[0012] Add noise signal x n (t) Perform EMD decomposition to obtain the modal function of the kth decomposition and the residual term r (k) (t), expressed as

[0013]

[0014] Among them, M k is the number of mode functions obtained by decomposition, is the ith eigenmode function, r (k) (t) is the residual sequence that cannot be further decomposed;

[0015] Repeat the above steps to add multiple different noises n k (t), expressed as

[0016]

[0017] And for each noise signal Perform EMD decomposition to obtain the corresponding IMF components and the residual sequence Expressed as

[0018]

[0019] Where j=1,2,...,J represents the number of times noise is added. is the number of mode functions obtained by the j-th decomposition, is the i-th modal function;

[0020] For all modal functions The final IMF is obtained by averaging, which is expressed as

[0021]

[0022] Where J is the number of repetitions of noise addition;

[0023] When the residual signal r(t) becomes a monotonic signal, that is, its derivative is always positive or negative, it means that the signal has no fluctuations and no effective IMF components can be further extracted. At this time, the decomposition process is stopped, which is expressed as

[0024] or

[0025] If the standard deviation of the IMF components of two adjacent iterations meets a certain threshold criterion, it can also indicate that the IMF has converged, which can be expressed as

[0026]

[0027] in, is the n-1th iteration result of the k-th layer IMF component, is the nth iteration result of the kth layer IMF component, T is the signal length of the time series data, is the standard deviation threshold.

[0028] As a further solution of the present invention: the multi-scale feature extraction module performs multi-scale feature extraction through a dilated convolution layer with different dilation rates. The dilated convolution introduces a dilated factor, inserts intervals between convolution kernel elements, expands the receptive field of the convolution operation, and stacks multiple modal signals IMFs and residual sequences decomposed by CEEMDAN into a modal decomposition matrix M. The long-term and short-term time dependencies in the modal signals and residual sequences are captured through dilated convolutions with different dilation rates to obtain a feature matrix D. The dilated convolution operation is in the form of

[0029]

[0030] Where M(t) is the modal information of the signal decomposition at time t, w(i) is the i-th element of the convolution kernel, r is the dilation factor, which indicates the interval between the convolution kernel elements, K is the size of the convolution kernel, and D(t) is the result of the dilation convolution operation.

[0031] In the dilated convolution, the dilated factor r inserts gaps between the convolution kernel elements, allowing the convolution operation to cover a wider area; therefore, the receptive field size of the dilated convolution is proportional to the dilated factor r, expressed as

[0032] R=(k-1)·r+1(10)

[0033] Among them, R is the receptive field size, k is the convolution kernel size;

[0034] Next, a multi-head attention mechanism is introduced for the features extracted with different expansion rates. The multi-head attention mechanism maps the input features to multiple subspaces, generates queries, keys and values ​​in each subspace, and calculates the attention weight matrix. Based on this, the values ​​are weighted and summed to obtain the weighted feature matrix of each subspace. The execution process is as follows:

[0035] (1) Use Xavier to query the weight matrix W Q , key weight matrix W K Sum value weight matrix W V Perform initialization settings, which can be expressed as

[0036]

[0037] Among them, d in is the feature dimension of the input, d k is the dimension of query and key vectors, d v is the dimension of the value vector;

[0038] Output weight matrix W O Map the concatenated outputs of multiple attention heads back to the original dimension d k , expressed as

[0039]

[0040] Where h is the number of attention heads;

[0041] (2) The feature D extracted by the dilated convolutional layer is mapped to multiple subspaces through the learned weight matrix input, and different query, key, and value matrices are generated for each attention head. The query, key, and value corresponding to the i-th attention head can be expressed as

[0042] Q=DW Q , K=DW K , V=DW V (13)

[0043] in, represents the linear transformation weight of the query matrix, represents the linear transformation weights of the key matrix, represents the linear transformation weight of the value matrix, d k is the dimension of the key vector, d v is the dimension of the value vector, d k =d v ; Based on the relationship between the query and the key, the attention score of each head is calculated, normalized using the scaled dot product and the softmax activation function, expressed as

[0044]

[0045] in, is the dot product between the query and the key, which is used to measure their similarity. is a scaling factor used to prevent the dot product from being too large and causing the gradient to vanish or explode. The softmax operation is used to ensure that the sum of the attention weights is 1 and to normalize the attention scores of all heads.

[0046] By applying the attention weights to the corresponding value matrix V i , and the weighted sum of each head is obtained. Then the output of each head can be expressed as

[0047] O i =Α i V i (15)

[0048] in, is the output matrix of the i-th head, representing the value vector weighted by attention;

[0049] (3) Concatenate the output matrices of all attention heads by column to obtain a new matrix, expressed as

[0050] O MH =Concat(O1,O2,...,O h ) (16)

[0051] in, is the concatenation result of all attention heads’ outputs, and h is the number of attention heads;

[0052] Finally, through a linear transformation matrix Map the concatenated results to get the final output O, expressed as

[0053] O=O MH W O (17)

[0054] in, is the weight matrix of the linear transformation, It is the final output result of the multi-head attention mechanism;

[0055] (4) Select the root mean square error as the loss function L, expressed as

[0056]

[0057] Among them, y i is the true value, is the predicted value, n is the number of samples;

[0058] (5) In the back propagation and gradient update phase, the gradient of the loss function L with respect to the query, key, value matrix, and output weight matrix is ​​calculated by the chain rule. The gradient calculation of the query matrix is ​​expressed as

[0059]

[0060] The gradient calculation for the bond matrix is ​​expressed as

[0061]

[0062] The gradient calculation of the value matrix is ​​expressed as

[0063]

[0064] The gradient calculation for the output weight matrix is ​​expressed as

[0065]

[0066] (6) Using the calculated gradients, the Adam optimizer is used to update the query, key, value matrices, and output weight matrix; first, the first-order moment estimate m is calculated for each parameter θ t and the second-order moment estimate v t Expressed as

[0067]

[0068] Among them, g t is the gradient value calculated by the loss function for the current parameter at time t, β1 and β2 are the decay rates of the first-order moment and the second-order moment, β1 is set to 0.9, and β2 is set to 0.999;

[0069] (7) During the training process, the deviation correction term is used to correct m t and v t Corrected, expressed as

[0070]

[0071] in, and are bias-corrected estimates of momentum and RMS;

[0072] (8) Using the corrected moment and The updated parameters are calculated by the Adam optimizer and are expressed as

[0073]

[0074] Among them, θ t and θ t-1 are the parameter values ​​at time t and time t-1 respectively, and η is the learning rate, which is set to 10 -3 , ε is a very small constant to avoid division by zero error, set to 10 -8 .

[0075] As a further solution of the present invention: the multiple ELMs in the ELM module are compared by root mean square error (RMSE), mean absolute percentage error (MAPE) and coefficient of determination (R 2 ) Multiple evaluation indicators are used to select multiple ELMs with better performance. In order to give full play to the advantages of each model and further improve the prediction accuracy, the prediction results of the selected multiple ELMs are recorded as Y k ,k=1,…,K, K is the number of ELMs; the implementation steps are as follows:

[0076] The feature sample Input to the hidden layer, output through H hidden layer neurons, expressed as

[0077] z=WO+b (26)

[0078] in, is the weight matrix input to the hidden layer, is the bias corresponding to each neuron in the hidden layer;

[0079] Then, after the activation function g(·) is applied, the final hidden layer output can be expressed as

[0080] h=g(z) (27)

[0081] The output matrix of the hidden layer is recorded as Expressed as

[0082] H ij =g(Wx j +b) i (28)

[0083] Among them, x j is the jth input sample, H ij is the output value of the i-th hidden layer neuron for the j-th input sample;

[0084] The prediction result of the output layer is expressed as

[0085] Y=Hβ (29)

[0086] in, is the weight matrix of the output layer, is the predicted value of all samples;

[0087] The output layer weight matrix β is expressed as follows through the analytical solution of the least squares method:

[0088]

[0089] in, is the target value matrix, is the Moore-Penrose pseudo-inverse matrix of the hidden layer output matrix H, expressed as

[0090]

[0091] Where H is a full-rank square matrix. If H is not a square matrix or has insufficient rank, it can be solved by singular value decomposition (SVD) or regularization.

[0092] The training goal of ELM is to minimize the error between the predicted value Y and the true value T, which can be expressed as

[0093]

[0094] After training, for a new input x′, the output of the hidden layer is expressed as

[0095] h=g(Wx′+b) (33)

[0096] Then the predicted value is calculated by the trained output layer weight β, which is expressed as

[0097] y=hβ (34)

[0098] Next, we use RMSE, MAPE, and R 2 The evaluation index selects multiple models with better performance and gives full play to the advantages of each model to improve the prediction accuracy. The selected multiple ELM prediction results are stacked as Y = {y i ,i=1,2,...,n};

[0099] RMSE represents the square root of the mean square of the residual between the predicted value and the true value. It is used to measure the size of the prediction error and the degree of deviation of the distribution. It is expressed as:

[0100]

[0101] Among them, y i is the i-th true value, is the i-th predicted value, n is the number of samples;

[0102] MAPE is the average value of the ratio of the absolute value of the prediction error to the true value. It is used to measure the relative error of the prediction value and is expressed as:

[0103]

[0104] R 2 It is used to measure the explanatory power of the regression model on the target variable and is expressed as:

[0105]

[0106] Among them, SSR and SST are the residual ordinary sum and total sum of squares of the regression model, which are used to represent the variation that the model cannot explain and the total variation of the data, respectively. The form is:

[0107]

[0108] Among them, y i is the i-th true value, is the i-th predicted value, is the mean of the target variable y;

[0109] Choose a weighted approach to calculate the model score, expressed as

[0110] score=κ1·RMSE+κ2·MAPE-(1-κ1-κ2)·R 2 (39)

[0111] Among them, κ1 and κ2 are the weights of RMSE and MAPE, which are set to 0.4 and 0.3.

[0112] As a further solution of the present invention: the elastic network regularized width neural network maps the ELM module prediction results to a high-dimensional space through nonlinear transformation based on the width learning idea to capture more potential feature information and generate a feature enhancement matrix H. Then, the elastic network is used to remove redundant features in the feature enhancement matrix H and the feature matrix Y and integrate them; the selected multiple ELM prediction results Y={y i , i=1,2,...,n} for nonlinear transformation, and use different nonlinear activation functions to further map the output matrix to a high-dimensional space. The high-dimensional space feature enhancement matrix H={h i |i=1,2,...,n} is expressed as

[0113]

[0114] in, Non-linear activation functions include ReLU, ELU, Sigmoid, Swish, Tanh, and Mish;

[0115] The ReLU activation function is expressed as:

[0116] σ(x)=max(0,x) (41)

[0117] The ELU activation function is expressed as:

[0118]

[0119] The Sigmoid activation function is expressed as:

[0120]

[0121] The Swish activation function is expressed as:

[0122]

[0123] The Tanh activation function is expressed as:

[0124]

[0125] The Mish activation function is expressed as:

[0126] σ(x)=x·tanh(ln(1+e x )) (46)

[0127] The elastic net regularization width neural network training process is achieved by minimizing the data fitting term L1 regularization term and L2 regularization term The loss function consists of three parts; the data fitting term is used to measure the degree of fit of the model to the observed data, the L1 term is used to achieve feature selection and increase the sparsity of the model, and the L2 term is used to prevent the coefficients in the model from being too large and reduce overfitting linearity; the elastic net regularization width neural network loss function form is

[0128]

[0129] Among them, X is the feature input matrix of the sample, w is the coefficient vector of the model, y is the target value matrix of the sample, n is the number of samples, α is the regularization strength, and ρ is the L1 and L2 regularization ratio;

[0130] Use coordinate descent to iterate each weight w j Update; Since the elastic network contains L1 regularization, the coordinate descent method has lower computational complexity than the gradient descent method for weight update. For each w j The optimization problem is expressed as

[0131]

[0132] Expand it and calculate the partial derivatives to get

[0133]

[0134] Among them, sign(w j ) is w j The sign function of w j >0 when sign(w j )=1, when w j When <0, sign(w j )=-1;

[0135] Since the L1 regularization term makes the loss function non-smooth, a soft threshold function is used to update the weights. For the update, the new value calculation is expressed as

[0136]

[0137] Where η is the learning rate, which is used to control the step size of each update and is usually set to 0.01 or 0.05. soft(·) is the soft threshold function, expressed as

[0138] soft(x,λ)=sign(x)·max(|x|-λ,0) (51)

[0139] If |x|≤λ, then soft(x,λ)=0, that is, the coordinate is set to zero. Otherwise, if |x|>λ, then soft(x,λ)=x-λ·sign(x), that is, the coordinate value is shrunk. Repeat equations (48) to (50), and in each iteration process, all weights w j Updates are made until the loss function converges or the preset number of iterations is reached.

[0140] The multi-scale width ELM lithium battery RUL prediction model based on elastic network of the present invention adopts multi-modal integrated void convolution elastic network regularized width neural network regression (CDWE) method, and the method comprises the following steps:

[0141] Step 1: Decompose the original battery degradation signal x(t) into intrinsic mode IMFs components of different frequency-time scales through CEEMDAN, and calculate the residual sequence r(t);

[0142] Step 2: Stack all decomposed modes and residual sequences in time series to form a modal decomposition matrix M, and use a sliding window with a fixed length of L to divide the entire modal data and residual sequence into time series segments;

[0143] Step 3: The time series data fragments are passed through multiple layers of dilated convolutions with different expansion rates, and the features of each layer of modal data and residual sequence data are extracted in parallel, capturing the long-term and short-term time dependencies in the modal signal and the residual sequence to obtain the feature matrix D. Each layer of the dilated convolution network will extract features of different scales of the input fragments to help the model obtain diverse features; the extracted features are weighted through the multi-head attention mechanism of equations (11) to (17), so that the model can adaptively focus on key features at different time scales, that is, through the calculation of query, key and value matrices, the model can weight each feature and generate a weighted feature matrix O;

[0144] Step 4: Input the data features obtained by the feature extraction module into N extreme learning machine models for training, and use the RMSE of formula (35), the MAPE of formula (36) and the R 2 Multiple performance evaluation indicators are weighted by formula (39) to select n extreme learning machines with better performance;

[0145] Step 5: The prediction data of n extreme learning machines are used to form a feature matrix Y. Then, based on the idea of ​​width learning, the feature matrix Y is width-expanded using the nonlinear transformation method of equations (41) to (46), and expanded into m feature enhancement nodes to form an enhanced feature matrix H.

[0146] Step 6: Use the feature matrix Y and feature enhancement matrix H to train the elastic network meta-learner, integrate the prediction results, and finally output the lithium battery RUL prediction results.

[0147] Compared with the prior art, the present invention uses CEEMDAN to decompose the original signal data of the battery into behavioral modes of different influencing factors, removes the correlation between the modes and the interference of data noise; in order to capture the short-term and long-term dependencies in the battery capacity data, different expansion rate hole convolution modules are used to extract the features of multi-scale signals, and a multi-head attention mechanism is used to extract important data features. The RUL prediction accuracy of lithium batteries in complex nonlinear systems is improved by the width extreme learning machine module. The width extreme learning machine introduces the idea of ​​width learning into multiple ELMs, and introduces elastic networks. The regularization technology simplifies the output weights of the width extreme learning machine, and improves the prediction accuracy, robustness and generalization of the model; the model can effectively reduce the amount of training data. This method can significantly reduce the computational complexity while improving the prediction accuracy and ensure the real-time application of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0148] Figure 1It is a structural schematic diagram of the multi-scale width ELM lithium battery RUL prediction model based on elastic network of the present invention, wherein a is a schematic diagram of battery degradation data changes, b is a schematic diagram of the multi-scale feature advance module, c is a schematic diagram of the CEEMDAN multimodal decomposition module, d is the ELM module, e is the elastic network regularized width neural network, and e1 is the feature width expansion module.

[0149] Figure 2 It is a schematic diagram comparing the RMSE of different numbers of hidden units in the present invention.

[0150] Figure 3 It is a schematic diagram of the relationship between the number of ELM hidden units and the maximum RMSE / minimum RMSE in the present invention.

[0151] Figure 4 It is a schematic diagram of RMSE comparison of different numbers of hidden units ELM in the present invention.

[0152] Figure 5 2 is a schematic diagram comparing the running time of different numbers of hidden units in the present invention.

[0153] Figure 6 Schematic diagram of the comparison of average running time of different numbers of hidden units in the present invention.

[0154] Figure 7 It is a comparative schematic diagram of the remaining battery life prediction in the present invention.

[0155] Figure 8 It is a flow chart of the CDWE algorithm of the present invention. DETAILED DESCRIPTION

[0156] The present invention will be further described below in conjunction with the accompanying drawings.

[0157] In order to reduce the complexity of the lithium battery remaining life prediction model and improve the model prediction accuracy to adapt to the occasions with limited computing power and resources in practical applications, this paper proposes a multi-scale width ELM lithium battery RUL prediction model based on elastic network. This model combines frequency domain analysis, neural network and traditional machine learning algorithms, and can extract effective data features while fully considering the equipment resource requirements and accurately predict the RUL value. Its structure is as follows Figure 1As shown in the figure, the model includes battery degradation data, CEEMDAN multimodal decomposition module, multi-scale feature extraction module, multi-extreme learning machine (ELM) module, elastic network regularized width neural network module, wherein the elastic network regularized width neural network includes a feature width expansion module and an elastic network regression meta-learner. The model decomposes the complex battery degradation data into multiple intrinsic mode IMFs through CEEMDAN multi-scale decomposition, and extracts multi-scale feature information through a multi-scale feature extraction module composed of a hole convolution layer with different expansion rates and a multi-head attention mechanism; then the extracted feature information is input into multiple ELM network structures for training and based on the root mean square error (RMSE), mean absolute percentage error (MAPE) and determination coefficient (R 2 ) Multiple performance indicators are used to select multiple ELMs with better performance; finally, the prediction results of multiple ELMs are used as new features and the enhanced features formed by mapping the nonlinear transformation of width expansion to high-dimensional space, and the two are input into the elastic network regularized width neural network module to obtain the final prediction results. This prediction model uses elastic network regularized width learning and random ideas to expand feature nodes, and then integrates these features through the elastic network. Next, the CEEMDAN multimodal decomposition module, multi-scale feature extraction module, ELM module, and elastic network regularized width neural network will be introduced in detail.

[0158] CEEMDAN multimodal decomposition module: The Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN) algorithm is used to decompose the intrinsic mode function (IMF) from the complex non-stationary signal, thereby extracting feature information of different time scales, deeply exploring the dynamic changes in the battery capacity decay process, and helping to identify the key features of different time stages. The final decomposition result of the original signal x(t) is expressed as:

[0159]

[0160] Where N represents the number of modal functions obtained, IMF i (t) is the decomposed eigenmode function of the i-th frequency, and r(t) represents the residual sequence that cannot be decomposed. The present invention uses CEEMDAN to decompose the battery degradation data into eigenmodes of different frequencies, and the implementation steps are as follows:

[0161] Add zero-mean Gaussian white noise n(t) to the original signal x(t) to get the noisy signal x n (t), expressed as

[0162] x n(t) = x(t) + n(t) (2)

[0163] Where n(t) is zero-mean Gaussian noise with variance Used to perturb the signal and improve the stability of decomposition. By adding noise, the randomness that may appear in word decomposition can be reduced, thereby enhancing the reliability of modal decomposition.

[0164] Add noise signal x n (t) Perform EMD decomposition to obtain the modal function of the kth decomposition and the residual term r (k) (t), expressed as

[0165]

[0166] Among them, M k is the number of mode functions obtained by decomposition, is the ith eigenmode function, r (k) (t) is the residual sequence that cannot be further decomposed.

[0167] Repeat the above steps to add multiple different noises n k (t), expressed as

[0168]

[0169] And for each noise signal Perform EMD decomposition to obtain the corresponding IMF components and the residual sequence Expressed as

[0170]

[0171] Where j=1,2,...,J represents the number of times noise is added. is the number of mode functions obtained by the j-th decomposition, is the i-th modal function;

[0172] For all modal functions The final IMF is obtained by averaging, which is expressed as

[0173]

[0174] Among them, J is the number of repetitions of noise addition. Through this process, CEEMDAN effectively reduces the impact of random noise in a single decomposition process and enhances the stability of the decomposition results.

[0175] When the residual signal r(t) becomes a monotonic signal, that is, its derivative is always positive or negative, it means that the signal has no fluctuations and no effective IMF components can be further extracted. At this time, the decomposition process is stopped, which is expressed as

[0176] or

[0177] If the standard deviation of the IMF components of two adjacent iterations meets a certain threshold criterion, it can also indicate that the IMF has converged, which can be expressed as

[0178]

[0179] in, is the n-1th iteration result of the k-th layer IMF component, is the nth iteration result of the kth layer IMF component, T is the signal length of the time series data, is the standard deviation threshold.

[0180] Multi-scale feature extraction module: This module first performs multi-scale feature extraction through dilated convolution layers with different dilation rates to capture long-term and short-term temporal dependencies. Dilated convolution introduces a dilation factor to insert intervals between convolution kernel elements, expands the receptive field of the convolution operation, and captures dependencies over longer time spans without increasing the computational burden, while avoiding the loss of spatial information that may be caused by pooling operations. The multiple modal signal IMFs and residual sequences decomposed by CEEMDAN are stacked into a modal decomposition matrix M, and the long-term and short-term temporal dependencies in the modal signals and residual sequences are captured through dilated convolutions with different dilation rates to obtain the feature matrix D. The form of the dilated convolution operation is

[0181]

[0182] Among them, M(t) is the modal information of the signal decomposition at time t, w(i) is the i-th element of the convolution kernel, r is the dilation factor, which indicates the interval between the convolution kernel elements, K is the size of the convolution kernel, and D(t) is the result of the dilation convolution operation.

[0183] In the dilated convolution, the dilated factor r inserts gaps between the convolution kernel elements, allowing the convolution operation to cover a wider area. Therefore, the receptive field of the dilated convolution is different from that of the traditional convolution. Its receptive field size is proportional to the dilated factor r, expressed as

[0184] R=(k-1)·r+1 (10)

[0185] Where R is the receptive field size and k is the convolution kernel size. Dilated convolution is particularly effective in processing time series data, as it can extract features over a wider time range, thereby effectively modeling long-term and short-term temporal dependencies.

[0186] The comparison of parameters and calculation amount of dilated convolution and traditional convolution with the same receptive field is shown in Table 1:

[0187] Table 1 Comparison of parameters and computational complexity between dilated convolution and traditional convolution

[0188]

[0189] It can be seen that the total number of parameters of the method proposed in this invention is 4×154=616, and the total amount of calculation is 4×1470=5880 FLOPs, while the total number of parameters of the traditional convolution with the same receptive field is 154+252+350+448=1240, and the total amount of calculation is 1470+2540+3430+4410=11850 FLOPs. The total number of parameters and the amount of calculation of the dilated convolution are about twice that of the traditional convolution. Therefore, the method proposed in the present invention can extract the required feature information more quickly by using fewer computing resources under resource-constrained conditions.

[0190] Next, a multi-head attention mechanism is introduced for the features extracted with different expansion rates, so that the model can dynamically adjust the degree of attention to these features, effectively improving the ability to capture key time periods. The multi-head attention mechanism maps the input features to multiple subspaces, generates queries, keys and values ​​in each subspace, and calculates the attention weight matrix. Based on this, the values ​​are weighted and summed to obtain the weighted feature matrix of each subspace. Then, multiple heads calculate their respective weighted feature matrices in parallel and concatenate them, and the final global weighted feature matrix is ​​generated through a linear mapping layer. The error between the predicted value and the true value is calculated by the loss function L, and the gradient of the loss function for all learnable parameters in the multi-head attention mechanism is calculated using the back-propagation algorithm. These parameters are updated by the gradient descent method, and the query, key and value matrices and the output weight matrix are updated by the calculated gradient using the Adam optimizer, so that the model can adaptively focus on the key features at different time scales in the input data.

[0191] The execution process is as follows:

[0192] (1) Use Xavier to query the weight matrix W Q , key weight matrix W K Sum value weight matrix W V Perform initialization settings, which can be expressed as

[0193]

[0194] Among them, d in is the feature dimension of the input, d k is the dimension of query and key vectors, d v is the dimension of the value vector;

[0195] Output weight matrix W O Map the concatenated outputs of multiple attention heads back to the original dimension d k , expressed as

[0196]

[0197] Where h is the number of attention heads;

[0198] (2) The feature D extracted by the dilated convolutional layer is mapped to multiple subspaces through the learned weight matrix input, and different query, key, and value matrices are generated for each attention head. The query, key, and value corresponding to the i-th attention head can be expressed as

[0199] Q=DW Q , K = DW K , V=DW V (13)

[0200] in, represents the linear transformation weight of the query matrix, represents the linear transformation weights of the key matrix, represents the linear transformation weight of the value matrix, d k is the dimension of the key vector, d v is the dimension of the value vector, d k =d v ; Based on the relationship between the query and the key, the attention score of each head is calculated, normalized using the scaled dot product and the softmax activation function, expressed as

[0201]

[0202] in, is the dot product between the query and the key, which is used to measure their similarity. is a scaling factor used to prevent the dot product from being too large and causing the gradient to vanish or explode. The softmax operation is used to ensure that the sum of the attention weights is 1 and to normalize the attention scores of all heads.

[0203] By applying the attention weights to the corresponding value matrix V i , and the weighted sum of each head is obtained. Then the output of each head can be expressed as

[0204] O i =Αi V i (15)

[0205] in, is the output matrix of the i-th head, representing the value vector weighted by attention;

[0206] (3) Concatenate the output matrices of all attention heads by column to obtain a new matrix, expressed as

[0207] O MH =Concat(O1,O2,...,O h ) (16)

[0208] in, is the concatenation result of all attention heads’ outputs, and h is the number of attention heads;

[0209] Finally, through a linear transformation matrix Map the concatenated results to get the final output O, expressed as

[0210] O=O MH W O (17)

[0211] in, is the weight matrix of the linear transformation, This is the final output of the multi-head attention mechanism.

[0212] (4) Select the root mean square error as the loss function L, expressed as

[0213]

[0214] Among them, y i is the true value, is the predicted value, n is the number of samples;

[0215] (5) In the back propagation and gradient update phase, the gradient of the loss function L with respect to the query, key, value matrix, and output weight matrix is ​​calculated by the chain rule. The gradient calculation of the query matrix is ​​expressed as

[0216]

[0217] The gradient calculation for the bond matrix is ​​expressed as

[0218]

[0219] The gradient calculation of the value matrix is ​​expressed as

[0220]

[0221] The gradient calculation for the output weight matrix is ​​expressed as

[0222]

[0223] (6) Using the calculated gradients, the Adam optimizer is used to update the query, key, value matrices, and output weight matrix. First, the first-order moment estimate m is calculated for each parameter θ t and the second moment estimate v t It can be expressed as

[0224]

[0225] Among them, g t is the gradient value calculated by the loss function for the current parameter at time t, β1 and β2 are the decay rates of the first-order moment and the second-order moment, β1 is set to 0.9, and β2 is set to 0.999;

[0226] (7) During the training process, the deviation correction term is used to correct m t and v t Corrected, expressed as

[0227]

[0228] in, and are bias-corrected estimates of momentum and RMS;

[0229] (8) Using the corrected moment and The updated parameters are calculated by the Adam optimizer and are expressed as

[0230]

[0231] Among them, θ t and θ t-1 are the parameter values ​​at time t and time t-1, respectively, and η is the learning rate, usually 10 -3 or 10 -4 , the present invention is set to 10 -3 , ε is a very small constant to avoid division by zero error, and is set to 10 in the present invention. -8 .

[0232] Multiple Extreme Learning Machine (ELM) modules: The features D extracted by dilated convolution and the features O enhanced by the multi-head attention mechanism are input into ELM. ELM is an efficient single hidden layer feedforward neural network. It randomly initializes the weights and biases of hidden nodes and uses the least squares method to calculate the output layer weights, thereby avoiding the process of traditional neural networks relying on back propagation and gradient descent. Compared with traditional machine learning models, the training process of ELM does not require complex iterative optimization, and the training time is significantly reduced, especially when processing high-dimensional data. Compared with deep learning models, ELM is faster to train, and due to its simple structure and training process, it can effectively avoid overfitting, especially in small sample learning tasks. In this way, ELM can not only complete model training efficiently, but also show good prediction performance in a variety of application scenarios.

[0233] However, randomly initializing weights can cause ELM performance to exhibit certain instability, such as Figure 2 , Figure 3 and Figure 4 As shown in the figure, this instability will gradually decrease as the number of hidden units increases. In other words, fewer hidden units are often accompanied by greater randomness. On the contrary, increasing the number of hidden units makes the model performance more stable. However, for the small sample learning task involved in the present invention, it is more appropriate to select a smaller number of hidden units. Excessive increase in hidden units will inevitably increase the risk of overfitting.

[0234] In addition, the increase in the number of hidden units will lead to a sharp increase in the amount of system calculation, such as Figure 5 and Figure 6 As shown in the figure, as the computational burden increases, the system's operating time will also be extended. For agricultural drones, this performance loss makes it difficult to meet their mission requirements for real-time and rapid response capabilities. Therefore, when designing the ELM module, the present invention needs to balance the relationship between hidden units and computational complexity to ensure that the system can meet the real-time requirements in practical applications.

[0235] In the ELM module, through the ELM's randomly initialized weights and bias characteristics, different ELMs map input features differently. In other words, different ELMs pay different attention to the same features. This differentiated attention method enables each ELM to capture different information in the feature space, especially when dealing with complex nonlinear relationships.

[0236] To this end, the present invention selects fewer hidden units, such as 50, and uses the feature information in the feature matrix O to train multiple ELMs, effectively eliminating the instability caused by random initialization weights, thereby reducing the computational burden and overfitting risk caused by a large number of hidden units. 2 ) Multiple evaluation indicators are used to select multiple models with better performance, and the advantages of each model are fully utilized to improve the prediction accuracy. The selected multiple ELM prediction results are recorded as Y k ,k=1,…,K, K is the number of ELMs. The implementation steps are as follows:

[0237] The feature sample Input to the hidden layer, output through H hidden layer neurons, expressed as

[0238] z=WO+b (26)

[0239] in, is the weight matrix input to the hidden layer, is the bias corresponding to each neuron in the hidden layer. Then, after the activation function g(·) is applied, the final hidden layer output can be expressed as

[0240] h=g(z) (27)

[0241] The output matrix of the hidden layer is recorded as Expressed as

[0242] H ij =g(Wx j +b) i (28)

[0243] Among them, x j is the jth input sample, H ij is the output value of the i-th hidden layer neuron for the j-th input sample;

[0244] The prediction result of the output layer is expressed as

[0245] Y=Hβ (29)

[0246] in, is the weight matrix of the output layer, is the predicted value of all samples;

[0247] The output layer weight matrix β is expressed as follows through the analytical solution of the least squares method:

[0248]

[0249] in, is the target value matrix, is the Moore-Penrose pseudo-inverse matrix of the hidden layer output matrix H, expressed as

[0250]

[0251] Where H is a full-rank square matrix. If H is not a square matrix or has insufficient rank, it can be solved by singular value decomposition (SVD) or regularization.

[0252] The training goal of ELM is to minimize the error between the predicted value Y and the true value T, which can be expressed as

[0253]

[0254] After training, for a new input x′, the output of the hidden layer is expressed as

[0255] h=g(Wx′+b) (33)

[0256] Then the predicted value is calculated by the trained output layer weight β, which is expressed as

[0257] y=hβ (34)

[0258] Then, the root mean square error (RMSE), mean absolute percentage error (MAPE) and coefficient of determination (R 2 ) Multiple evaluation indicators are used to select multiple models with better performance, and the advantages of each model are fully utilized to improve the prediction accuracy. The selected multiple ELM prediction results are stacked as Y = {y i ,i=1,2,...,n}.

[0259] RMSE represents the square root of the mean square of the residual between the predicted value and the true value. It is used to measure the size of the prediction error and the degree of deviation of the distribution. It is expressed as:

[0260]

[0261] Among them, y i is the i-th true value, is the i-th predicted value, n is the number of samples;

[0262] MAPE is the average value of the ratio of the absolute value of the prediction error to the true value. It is used to measure the relative error of the prediction value and is expressed as:

[0263]

[0264] R 2It is used to measure the explanatory power of the regression model on the target variable, that is, the proportion of the total variation of the data that the model can explain, which can be expressed as:

[0265]

[0266] Among them, SSR and SST are the residual ordinary sum and total sum of squares of the regression model, which are used to represent the variation that the model cannot explain and the total variation of the data, respectively. The form is:

[0267]

[0268] Among them, y i is the i-th true value, is the i-th predicted value, is the mean of the target variable y;

[0269] In order to combine multiple performance evaluation indicators and select multiple ELMs with better performance, a weighted method is selected to calculate the model score, which can be expressed as

[0270] score=κ1·RMSE+κ2·MAPE-(1-κ1-κ2)·R 2 (39)

[0271] Among them, κ1 and κ2 are the weights of RMSE and MAPE, which are set to 0.4 and 0.3 in this patent.

[0272] Elastic network regularized wide neural network: When faced with high-dimensional features and complex nonlinear relationships, traditional machine learning has limited ability to process complex and changing data and cannot effectively capture the deep nonlinear structure in the data. The wide neural network can effectively improve the model's perception of data patterns by expanding the width of the input features through nonlinear transformation. In addition, the wide neural network can effectively avoid the computational burden and gradient vanishing problem of deep neural networks during training by increasing the network width rather than the network depth. The elastic network, by combining L1 and L2 regularization, can significantly improve the sparsity and feature selection capabilities of the model, simplify the model structure, and enhance the generalization ability of the model.

[0273] The elastic network regularized wide neural network combines width learning, deep learning and elastic network regularization technology, avoiding the complexity of iterative optimization, model generalization and robustness problems in deep neural networks. Based on the idea of ​​width learning, the ELM prediction results are mapped to a high-dimensional space through nonlinear transformation to capture more potential feature information and generate a feature enhancement matrix H. Subsequently, the elastic network is used to remove redundant features in the feature enhancement matrix H and the feature matrix Y and integrate them. The elastic network regularization method combines the advantages of LASSO regression and ridge regression, which can effectively improve the prediction accuracy, generalization and robustness of the model. The elastic network regularized wide neural network constructed in this patent can make full use of the diversity of ELM and improve the prediction accuracy, generalization and robustness of complex systems.

[0274] The elastic network regularized wide neural network includes feature width expansion and elastic network regression meta-learner. The wide neural network can improve the model's ability to learn complex patterns through feature width expansion, especially when facing nonlinear relationships. This process maps the input features to a high-dimensional space through nonlinear transformation, thereby generating enhanced feature nodes and expanding the model's ability to capture complex patterns. Feature width expansion includes feature mapping nodes and feature enhancement nodes. In this patent, multiple ELM prediction results Y={y i ,i=1,2,...,n} for nonlinear transformation, and use different nonlinear activation functions to further map the output matrix to high-dimensional space, improve the nonlinear fitting ability of the model, and better handle more complex data patterns. The high-dimensional space feature enhancement matrix H={h i |i=1,2,...,n} is expressed as

[0275]

[0276] in, Non-linear activation functions include ReLU, ELU, Sigmoid, Swish, Tanh, and Mish;

[0277] The ReLU activation function is expressed as:

[0278] σ(x)=max(0,x) (41)

[0279] The ELU activation function is expressed as:

[0280]

[0281] The Sigmoid activation function is expressed as:

[0282]

[0283] The Swish activation function is expressed as:

[0284]

[0285] The Tanh activation function is expressed as:

[0286]

[0287] The Mish activation function is expressed as:

[0288] σ(x)=x·tanh(ln(1+e x )) (46)

[0289] By using nonlinear activation functions other than those mentioned above, the model can capture more complex and delicate data patterns in high-dimensional space, further enhancing the model’s ability to learn complex relationships, thereby improving the overall accuracy and generalization of the model.

[0290] Elastic Net is a regression method that combines the advantages of L1 regularization (Lasso regression) and L2 regularization (Ridge regression), aiming to effectively handle highly correlated features and solve multicollinearity problems. L1 regularization achieves feature selection and improves the interpretability of the model by pushing the coefficients of certain features to zero, but when faced with highly correlated features, L1 regularization may randomly select certain features, resulting in model instability. In contrast, L2 regularization effectively reduces the overfitting problem by reducing the amplitude of all coefficients, but it cannot perform feature selection. Therefore, by combining the advantages of both and finding a balance between L1 and L2 regularization, Elastic Net can still handle highly correlated features while avoiding the random feature selection of L1 regularization, thereby improving the stability of the model and prediction accuracy.

[0291] The elastic net regularization width neural network training process is achieved by minimizing the data fitting term L1 regularization term and L2 regularization term The loss function consists of three parts. Among them, the data fitting term is used to measure the degree of fit of the model to the observed data, the L1 term is used to achieve feature selection and increase the sparsity of the model, and the L2 term is used to prevent the coefficients in the model from being too large and reduce overfitting linearity. The elastic net regularization width neural network loss function is

[0292]

[0293] Among them, X is the feature input matrix of the sample, w is the coefficient vector of the model, y is the target value matrix of the sample, n is the number of samples, α is the regularization strength, and ρ is the L1 and L2 regularization ratio.

[0294] For this non-smooth optimization problem, this patent uses the coordinate descent method to iterate each weight w j Because the elastic network contains L1 regularization, the coordinate descent method has lower computational complexity than the gradient descent method for weight updating. For each w j The optimization problem is expressed as

[0295]

[0296] Expand it and calculate the partial derivatives to get

[0297]

[0298] Among them, sign(w j ) is w j The sign function of w j >0 when sign(w j )=1, when w j When <0, sign(w j )=-1;

[0299] Since the L1 regularization term makes the loss function non-smooth, a soft threshold function is used to update the weights. For the update, the new value calculation can be expressed as

[0300]

[0301] Wherein, η is the learning rate, which is used to control the step size of each update. For the data set involved in this patent, a larger learning rate can be set, usually set to 0.01 or 0.05, and soft(·) is the soft threshold function, which is expressed as

[0302] soft(x,λ)=sign(x)·max(|x|-λ,0) (51)

[0303] If |x|≤λ, then soft(x,λ)=0, that is, the coordinate is set to zero. Otherwise, if |x|>λ, then soft(x,λ)=x-λ·sign(x), that is, the coordinate value is shrunk. Repeat equations (48) to (50), and in each iteration process, all weights w j Updates are made until the loss function converges or the preset number of iterations is reached.

[0304] like Figure 7The test results shown are based on the idea of ​​Stacking ensemble learning, and integrate the prediction results of multiple extreme learning machine models through the elastic network regularization width ELM network. This method can conduct a more comprehensive exploration in the feature space from multiple angles, make full use of the differences in the attention paid by different ELMs to the same features, and achieve complementary advantages between models. By integrating the outputs of different ELMs, it can not only enhance the robustness of the model to the data and improve its ability to learn complex patterns, but also effectively reduce the overfitting and underfitting problems that may occur in a single model, thereby significantly improving the overall prediction accuracy and generalization ability of the model.

[0305] The elastic network regularized width neural network proposed in this patent expands the feature width through nonlinear transformation of width learning, and uses elastic network to eliminate redundant features and simplify the model structure. In this patent, the d-dimensional input feature is first expanded to d through width learning. w Then, the elastic network is used to select features to eliminate redundant features and reduce the feature dimension to d f , which can not only use more important features, but also maintain the computational efficiency of the network; then through a single-layer network structure with a width of S, the total computational amount of the elastic network regularized width neural network mentioned in this patent is d f ×S. Although the dimension of the feature is expanded after width learning, the feature dimension actually used for model training and inference is small due to the feature selection of the elastic network. f When d is equal to d, the computational complexity of the regression method proposed in this patent is only 1 / L of that of the deep neural network. Therefore, the total computational complexity of the method proposed in this patent is significantly lower than that of the traditional deep neural network, especially in an environment with limited computing resources. It can significantly reduce the computational overhead while ensuring the prediction accuracy, thereby improving the real-time performance of the prediction.

[0306] The prediction model of this patent adopts the CEEMDAN Dilated-Convolution ELM of Width Expansion Elastic Net Regression Ensemble Learning (CDWE) method, which includes the following steps:

[0307] Step 1: Decompose the original battery degradation signal x(t) into intrinsic mode IMFs components of different frequency-time scales through CEEMDAN, and calculate the residual sequence r(t);

[0308] Step 2: Stack all decomposed modes and residual sequences in time series to form a modal decomposition matrix M, and use a sliding window with a fixed length of L to divide the entire modal data and residual sequence into time series segments;

[0309] Step 3: The time series data fragments are passed through multiple layers of dilated convolutions with different expansion rates, and the features of each layer of modal data and residual sequence data are extracted in parallel, capturing the long-term and short-term time dependencies in the modal signal and the residual sequence to obtain the feature matrix D. Each layer of dilated convolution network will extract features of different scales of the input fragments to help the model obtain diverse features. The extracted features are weighted through the multi-head attention mechanism of equations (11) to (17), so that the model can adaptively focus on key features at different time scales. That is, through the calculation of query, key and value matrices, the model can weight each feature and generate a weighted feature matrix O. Through the attention mechanism, the model can highlight the features of key moments and improve the model's sensitivity to important moments.

[0310] Step 4: Input the data features obtained by the feature extraction module into N extreme learning machine models for training, and use the root mean square error (RMSE) of formula (35), the mean absolute percentage error (MAPE) of formula (36) and the coefficient of determination (R 2 ) Multiple performance evaluation indicators are weighted by formula (39) to select n extreme learning machines with better performance;

[0311] Step 5: The prediction data of n extreme learning machines are used to form a feature matrix Y. Then, based on the idea of ​​width learning, the feature matrix Y is expanded into m feature enhancement nodes using the nonlinear transformation method of equations (41) to (46). This enhances the feature expression ability of the model, contains richer feature information, and can capture more complex models, thereby improving the prediction accuracy of the model.

[0312] Step 6: Use the feature matrix Y and the feature enhancement matrix H to train the elastic network meta-learner. By combining L1 and L2 regularization, it can effectively eliminate redundant features to simplify the model structure and improve the model robustness. By combining the prediction results of multiple basic models, the overall prediction accuracy can be improved. This stacking ensemble learning integration method can effectively reduce the deviation of a single model and improve the final prediction accuracy and stability.

[0313] The battery remaining life prediction method used in this patent performs excellently and has broad application prospects. Its low computational complexity and high prediction accuracy enable it to be effectively deployed in resource-constrained environments. It is particularly suitable for scenarios such as agricultural drones, industrial inspection robots, and portable medical equipment.

Claims

1. A multi-scale width ELM lithium battery RUL prediction model based on elastic network, characterized by: The invention comprises a CEEMDAN multimodal decomposition module, a multiscale feature extraction module, a multi-extreme learning machine (ELM) module, and an elastic network regularized width neural network module. The elastic network regularized width neural network comprises a feature width expansion module and an elastic network regression meta-learner. The CEEMDAN multimodal decomposition module decomposes battery degradation data into multiple intrinsic mode IMFs, extracts multiscale feature information through a multiscale feature extraction module, inputs the extracted multiscale feature information into an ELM module for training, and selects multiple ELMs with better performance based on multiple performance indicators. Finally, the prediction results of multiple ELMs are used as new features and enhanced features formed by mapping to a high-dimensional space through a nonlinear transformation of width expansion, and the two are input into the elastic network regularized width neural network module to obtain the final prediction result.

2. The multi-scale width ELM lithium battery RUL prediction model based on elastic network according to claim 1 is characterized in that: The CEEMDAN multimodal decomposition module decomposes the intrinsic mode function IMF from the complex non-stationary signal through the CEEMDAN algorithm, thereby extracting the characteristic information of different time scales. The final decomposition result of the original signal x(t) is expressed as Where N represents the number of modal functions obtained, IMF i (t) is the i-th intrinsic mode function of the decomposition, r(t) represents the residual sequence that cannot be decomposed; CEEMDAN is used to decompose the battery degradation data into eigenmodes of different frequencies. The implementation steps are as follows: Add zero-mean Gaussian white noise n(t) to the original signal x(t) to get the noisy signal x n (t), expressed as x n (t)=x(t)+n(t) (2) Where n(t) is zero-mean Gaussian noise with variance Used to perturb the signal and improve the stability of the decomposition; Add noise signal x n (t) Perform EMD decomposition to obtain the modal function of the kth decomposition and the residual term r (k) (t), expressed as Among them, M k is the number of mode functions obtained by decomposition, is the ith eigenmode function, r (k) (t) is the residual sequence that cannot be further decomposed; Repeat the above steps to add multiple different noises n k (t), expressed as And for each noise signal Perform EMD decomposition to obtain the corresponding IMF component IMF i (kj) (t) and the residual sequence Expressed as Where j=1,2,...,J represents the number of times noise is added. is the number of mode functions obtained by the j-th decomposition, is the i-th modal function; For all modal functions The final IMF is obtained by averaging, which is expressed as Where J is the number of repetitions of noise addition; When the residual signal r(t) becomes a monotonic signal, that is, its derivative is always positive or negative, it means that the signal has no fluctuations and no effective IMF components can be further extracted. At this time, the decomposition process is stopped, which is expressed as If the standard deviation of the IMF components of two adjacent iterations meets a certain threshold criterion, it can also indicate that the IMF has converged, which can be expressed as in, is the n-1th iteration result of the k-th layer IMF component, is the nth iteration result of the kth layer IMF component, T is the signal length of the time series data, is the standard deviation threshold.

3. The multi-scale width ELM lithium battery RUL prediction model based on elastic network according to claim 1 is characterized in that: The multi-scale feature extraction module performs multi-scale feature extraction through dilated convolution layers with different dilation rates. The dilated convolution introduces a dilated factor, inserts intervals between convolution kernel elements, expands the receptive field of the convolution operation, and stacks multiple modal signal IMFs and residual sequences decomposed by CEEMDAN into a modal decomposition matrix M. The long-term and short-term time dependencies in the modal signal and residual sequence are captured by dilated convolutions with different dilation rates to obtain the feature matrix D. The dilated convolution operation is in the form of Where M(t) is the modal information of the signal decomposition at time t, w(i) is the i-th element of the convolution kernel, r is the dilation factor, which indicates the interval between the convolution kernel elements, K is the size of the convolution kernel, and D(t) is the result of the dilation convolution operation. In the dilated convolution, the dilated factor r inserts gaps between the convolution kernel elements, allowing the convolution operation to cover a wider area; therefore, the receptive field size of the dilated convolution is proportional to the dilated factor r, expressed as R=(k-1)·r+1 (10) Among them, R is the receptive field size, k is the convolution kernel size; Next, a multi-head attention mechanism is introduced for the features extracted with different expansion rates. The multi-head attention mechanism maps the input features to multiple subspaces, generates queries, keys and values ​​in each subspace, and calculates the attention weight matrix. Based on this, the values ​​are weighted and summed to obtain the weighted feature matrix of each subspace. The execution process is as follows: (1) Use Xavier to query the weight matrix W Q , key weight matrix W K Sum value weight matrix W V Perform initialization settings, which can be expressed as Among them, d in is the feature dimension of the input, d k is the dimension of query and key vectors, d v is the dimension of the value vector; Output weight matrix W O Map the concatenated outputs of multiple attention heads back to the original dimension d k , expressed as Where h is the number of attention heads; (2) The feature D extracted by the dilated convolutional layer is mapped to multiple subspaces through the learned weight matrix input, and different query, key, and value matrices are generated for each attention head. The query, key, and value corresponding to the i-th attention head can be expressed as Q=DW Q ,K=DW K ,V=DW V (13) in, represents the linear transformation weight of the query matrix, represents the linear transformation weights of the key matrix, represents the linear transformation weight of the value matrix, d k is the dimension of the key vector, d v is the dimension of the value vector, d k =d v ; Based on the relationship between the query and the key, the attention score of each head is calculated, normalized using the scaled dot product and the softmax activation function, expressed as in, is the dot product between the query and the key, which is used to measure their similarity. is a scaling factor used to prevent the dot product from being too large and causing the gradient to vanish or explode. The softmax operation is used to ensure that the sum of the attention weights is 1 and to normalize the attention scores of all heads. By applying the attention weights to the corresponding value matrix V i , and the weighted sum of each head is obtained. Then the output of each head can be expressed as The i =A i V i (15) in, is the output matrix of the i-th head, representing the value vector weighted by attention; (3) Concatenate the output matrices of all attention heads by column to obtain a new matrix, expressed as O MH =Concat(O1,O2,...,O h ) (16) in, is the concatenation result of all attention heads’ outputs, and h is the number of attention heads; Finally, through a linear transformation matrix Map the concatenated results to get the final output O, expressed as O=O MH W O (17) in, is the weight matrix of the linear transformation, It is the final output result of the multi-head attention mechanism; (4) Select the root mean square error as the loss function L, expressed as Among them, y i is the true value, is the predicted value, n is the number of samples; (5) In the back propagation and gradient update phase, the gradient of the loss function L with respect to the query, key, value matrix, and output weight matrix is ​​calculated by the chain rule. The gradient calculation of the query matrix is ​​expressed as The gradient calculation for the bond matrix is ​​expressed as The gradient calculation of the value matrix is ​​expressed as The gradient calculation for the output weight matrix is ​​expressed as (6) Using the calculated gradients, the Adam optimizer is used to update the query, key, value matrices, and output weight matrix; first, the first-order moment estimate m is calculated for each parameter θ t and the second-order moment estimate v t Expressed as Among them, g t is the gradient value calculated by the loss function for the current parameter at time t, β1 and β2 are the decay rates of the first-order moment and the second-order moment, β1 is set to 0.9, and β2 is set to 0.999; (7) During the training process, the deviation correction term is used to correct m t and v t Corrected, expressed as in, and are bias-corrected estimates of momentum and RMS; (8) Using the corrected moment and The updated parameters are calculated by the Adam optimizer and are expressed as Among them, θ t and θ t-1 are the parameter values ​​at time t and time t-1 respectively, and η is the learning rate, which is set to 10 -3 , ε is a very small constant to avoid division by zero error, set to 10 -8 .

4. The multi-scale width ELM lithium battery RUL prediction model based on elastic network according to claim 1 is characterized in that: The multiple ELMs in the ELM module are analyzed by root mean square error (RMSE), mean absolute percentage error (MAPE) and coefficient of determination (R 2 ) Multiple evaluation indicators are used to select multiple ELMs with better performance, so as to give full play to the advantages of each model and further improve the prediction accuracy. The prediction results of the selected multiple ELMs are recorded as Y k ,k=1,…,K, K is the number of ELMs; the implementation steps are as follows: The feature sample Input to the hidden layer, output through H hidden layer neurons, expressed as z=WO+b (26) in, is the weight matrix input to the hidden layer, is the bias corresponding to each neuron in the hidden layer; Then, after the activation function g(·) is applied, the final hidden layer output can be expressed as h=g(z) (27) The output matrix of the hidden layer is recorded as Expressed as H ij =g(Wx j +b) i (28) Among them, x j is the jth input sample, H ij is the output value of the i-th hidden layer neuron for the j-th input sample; The prediction result of the output layer is expressed as Y=Hβ (29) in, is the weight matrix of the output layer, is the predicted value of all samples; The output layer weight matrix β is expressed as follows through the analytical solution of the least squares method: in, is the target value matrix, is the Moore-Penrose pseudo-inverse matrix of the hidden layer output matrix H, expressed as Where H is a full-rank square matrix. If H is not a square matrix or has insufficient rank, it can be solved by singular value decomposition (SVD) or regularization. The training goal of ELM is to minimize the error between the predicted value Y and the true value T, which can be expressed as After training, for a new input x′, the output of the hidden layer is expressed as h=g(Wx′+b) (33) Then the predicted value is calculated by the trained output layer weight β, which is expressed as y=hβ (34) Next, we use RMSE, MAPE, and R 2 The evaluation index selects multiple models with better performance and gives full play to the advantages of each model to improve the prediction accuracy. The selected multiple ELM prediction results are stacked as Y = {y i ,i=1,2,...,n}; RMSE represents the square root of the mean square of the residual between the predicted value and the true value. It is used to measure the size of the prediction error and the degree of deviation of the distribution. It is expressed as: Among them, y i is the i-th true value, is the i-th predicted value, n is the number of samples; MAPE is the average value of the ratio of the absolute value of the prediction error to the true value. It is used to measure the relative error of the prediction value and is expressed as: R 2 It is used to measure the explanatory power of the regression model on the target variable and is expressed as: Among them, SSR and SST are the residual ordinary sum and total sum of squares of the regression model, which are used to represent the variation that the model cannot explain and the total variation of the data, respectively. The form is: Among them, y i is the i-th true value, is the i-th predicted value, and y is the mean of the target variable y; Choose a weighted approach to calculate the model score, expressed as score=κ1·RMSE+κ2·MAPE-(1-κ1-κ2)·R 2 (39) Among them, κ1 and κ2 are the weights of RMSE and MAPE, which are set to 0.4 and 0.

3.

5. The multi-scale width ELM lithium battery RUL prediction model based on elastic network according to claim 1, characterized in that: The elastic network regularized width neural network maps the ELM module prediction results to a high-dimensional space through nonlinear transformation based on the width learning idea to capture more potential feature information and generate a feature enhancement matrix H. Then, the elastic network is used to remove redundant features in the feature enhancement matrix H and the feature matrix Y and integrate them. The selected multiple ELM prediction results Y = {y i , i=1,2,...,n} for nonlinear transformation, and use different nonlinear activation functions to further map the output matrix to a high-dimensional space. The high-dimensional space feature enhancement matrix H={h i |i=1,2,...,n} is expressed as in, Non-linear activation functions include ReLU, ELU, Sigmoid, Swish, Tanh, and Mish; The ReLU activation function is expressed as: σ(x)=max(0,x) (41) The ELU activation function is expressed as: The Sigmoid activation function is expressed as: The Swish activation function is expressed as: The Tanh activation function is expressed as: The Mish activation function is expressed as: σ(x)=x·tanh(ln(1+e x )) (46) The elastic net regularization width neural network training process is achieved by minimizing the data fitting term L1 regularization term and L2 regularization term The loss function consists of three parts; the data fitting term is used to measure the degree of fit of the model to the observed data, the L1 term is used to achieve feature selection and increase the sparsity of the model, and the L2 term is used to prevent the coefficients in the model from being too large and reduce overfitting linearity; the elastic net regularization width neural network loss function form is Among them, X is the feature input matrix of the sample, w is the coefficient vector of the model, y is the target value matrix of the sample, n is the number of samples, α is the regularization strength, and ρ is the L1 and L2 regularization ratio; Use coordinate descent to iterate each weight w j Update; Since the elastic network contains L1 regularization, the coordinate descent method has lower computational complexity than the gradient descent method for weight update. For each w j The optimization problem is expressed as Expand it and calculate the partial derivatives to get Among them, sign(w j ) is w j The sign function of w j >0 when sign(w j )=1, when w j When <0, sign(w j )=-1; Since the L1 regularization term makes the loss function non-smooth, a soft threshold function is used to update the weights. For the update, the new value calculation is expressed as Where η is the learning rate, which is used to control the step size of each update and is usually set to 0.01 or 0.

05. soft(·) is the soft threshold function, expressed as soft(x,λ)=sign(x)·max(x|-λ,0) (51) If |x|≤λ, then soft(x,λ)=0, that is, the coordinate is set to zero. Otherwise, if |x|>λ, then soft(x,λ)=x-λ·sign(x), that is, the coordinate value is shrunk. Repeat equations (48) to (50), and in each iteration process, all weights w j Updates are made until the loss function converges or the preset number of iterations is reached.

6. The multi-scale width ELM lithium battery RUL prediction model based on elastic network according to any one of claims 1 to 6 adopts a multi-modal integrated void convolutional elastic network regularized width neural network regression (CDWE) method, characterized in that: The method comprises the following steps: Step 1: Decompose the original battery degradation signal x(t) into intrinsic mode IMFs components of different frequency-time scales through CEEMDAN, and calculate the residual sequence r(t); Step 2: Stack all decomposed modes and residual sequences in time series to form a modal decomposition matrix M, and use a sliding window with a fixed length of L to divide the entire modal data and residual sequence into time series segments; Step 3: The time series data fragments are passed through multiple layers of dilated convolutions with different expansion rates, and the features of each layer of modal data and residual sequence data are extracted in parallel, capturing the long-term and short-term time dependencies in the modal signal and the residual sequence to obtain the feature matrix D. Each layer of the dilated convolution network will extract features of different scales of the input fragments to help the model obtain diverse features; the extracted features are weighted through the multi-head attention mechanism of equations (11) to (17), so that the model can adaptively focus on key features at different time scales, that is, through the calculation of query, key and value matrices, the model can weight each feature and generate a weighted feature matrix O; Step 4: Input the data features obtained by the feature extraction module into N extreme learning machine models for training, and use the RMSE of formula (35), the MAPE of formula (36) and the R 2 Multiple performance evaluation indicators are weighted by formula (39) to select n extreme learning machines with better performance; Step 5: The prediction data of n extreme learning machines are used to form a feature matrix Y. Then, based on the idea of ​​width learning, the feature matrix Y is expanded in width using the nonlinear transformation method of equations (41) to (46), and expanded into m feature enhancement nodes to form an enhanced feature matrix H. Step 6: Use the feature matrix Y and feature enhancement matrix H to train the elastic network meta-learner, integrate the prediction results, and finally output the lithium battery RUL prediction results.

Citation Information

Cited By

  • Wind power bearing fault diagnosis method based on dual-channel feature fusion

    CN120541716A

  • Battery life prediction method and system based on adaptive generative adversarial network

    CN120949092A

  • Method and device for predicting remaining life of battery and new energy automobile

    CN121232036A