Lithium battery life prediction method based on modal decomposition and sparse attention
By decomposing signals using 3σ-linear regression and ICEEMDAN, combined with GSA-optimized Stacking ensemble learning and sparse attention BiLSTM network, the problems of outliers and model mismatch in lithium battery life prediction are solved, achieving more accurate and stable predictions.
Patent Information
- Application Number
- CN202510679849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-09
AI Technical Summary
Existing lithium battery life prediction methods have problems such as outlier influence, model mismatch and limited feature extraction capabilities in capacity series, resulting in low prediction accuracy and poor stability.
The 3σ-linear regression detection and correction method is used to remove outliers. Combined with the ICEEMDAN decomposition signal, the GSA-optimized Stacking ensemble learning model and the sparse attention BiLSTM network are used for prediction. A lithium battery life prediction method based on modal decomposition and sparse attention is constructed.
It significantly improves the accuracy and stability of lithium battery life prediction, reduces noise interference, improves the generalization performance and prediction accuracy of the model, and overcomes the overfitting risk in traditional methods.
Smart Images

Figure CN120610164A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of lithium battery healthy life prediction, and specifically relates to a lithium battery life prediction method based on modal decomposition and sparse attention. Background Art
[0002] Lithium batteries are widely used in everyday life. Their high energy density, long lifespan, and low self-discharge rates have led to their widespread adoption in electric vehicles, mobile phones, and energy storage systems. However, constant charging and discharging of lithium batteries inevitably damages their health, leading to a continuous decrease in usable capacity. When a battery reaches a threshold of 70% to 80% of its rated capacity, it reaches the end of its service life. Battery degradation can pose safety risks and public safety incidents, making it crucial to predict battery health.
[0003] Early research on lithium-ion battery health life prediction methods relied primarily on empirical formulas and traditional mathematical models, such as those based on capacity decay curve fitting. However, these traditional methods often struggle to accurately describe the complex degradation mechanisms of lithium-ion batteries, have limited prediction accuracy, and are poorly adaptable to different types of lithium-ion batteries or varying operating conditions. In recent years, with the rapid development of machine learning and data mining technologies, data-driven approaches have gradually become a research hotspot.
[0004] In the prior art, Chinese patent CN117665627A discloses a method and system for predicting the remaining service life of a lithium battery based on an optimized neural network. Each battery capacity sub-time series is decomposed to obtain a residual component and a modal component; the residual component is input into a trained BiLSTM model to output a first predicted value of the dischargeable capacity of the lithium battery; the modal component is input into a trained BiGRU model to output a second predicted value of the dischargeable capacity of the lithium battery; the first predicted value of the dischargeable capacity of the lithium battery and the second predicted value of the dischargeable capacity of the lithium battery are integrated and added to obtain a predicted discharge capacity sequence of the future number of charge and discharge cycles of the lithium battery, and the predicted discharge capacity decays to a set ratio of the rated capacity as the end point of the remaining service life; a starting point of the remaining service life is set, and the number of charge and discharge cycles experienced between the starting point of the remaining service life and the end point of the remaining service life is calculated, and the obtained number of charge and discharge cycles is used as the remaining service life of the lithium battery.
[0005] However, while this method performs modal decomposition on the capacity series, outliers caused by measurement errors or unexpected events often appear in the capacity curve. Unfiltered, these outliers affect model training and prediction accuracy. Secondly, this method uses a fixed model structure (BiLSTM and BiGRU) to model the components, without conducting targeted analysis of the characteristics of different modal components and their appropriate adaptation models. This can easily lead to "model mismatch" problems, resulting in poor or overfitting.
[0006] In summary, existing technologies for lithium battery life prediction still suffer from issues such as low residual signal modeling accuracy, limited feature extraction capabilities, severe outlier interference, and weak model integration capabilities. Therefore, a new life prediction method combining efficient data decomposition methods with improved neural network architecture is urgently needed to improve prediction accuracy and model stability. Summary of the Invention
[0007] The purpose of this invention is to overcome the above-mentioned shortcomings of the existing technology and provide a lithium battery life prediction method based on modal decomposition and sparse attention. It is used to solve the technical problem that the capacity growth and nonlinear capacity decay caused by internal chemical reactions in lithium batteries lead to limited prediction accuracy.
[0008] The purpose of the present invention can be achieved by the following technical solutions:
[0009] The present invention provides a lithium battery life prediction method based on modal decomposition and sparse attention, comprising the following steps:
[0010] Obtaining a historical battery capacity time series of a battery to be predicted, wherein the historical battery capacity time series includes maximum battery capacities at multiple time points;
[0011] Constructing a battery capacity attenuation curve of the battery to be predicted according to the battery capacity time series;
[0012] Performing 3σ-linear regression detection and correction on the battery capacity decay curve to eliminate outliers in the capacity change process and retain valid values;
[0013] The improved fully adaptive noise ensemble empirical mode decomposition method ICEEMDAN is used to decompose the corrected battery capacity decay curve to obtain N intrinsic mode function (IMF) components and one residual component.
[0014] Inputting the N IMF components into the Stacking ensemble learning model optimized and trained by the gravitational search algorithm (GSA) for prediction, thereby obtaining a predicted IMF component sequence;
[0015] The residual component is input into a sparse attention bidirectional long short-term memory network model SBiLSTM for processing to obtain a predicted residual component sequence;
[0016] According to the predicted IMF component sequence and the predicted residual component sequence, a predicted value of the maximum battery capacity of the battery to be predicted at the next moment is obtained.
[0017] Furthermore, constructing a battery capacity attenuation curve of the battery to be predicted based on the battery capacity time series specifically includes:
[0018] Arrange the capacity values in the battery capacity time series in chronological order to generate a capacity-time data point sequence;
[0019] Applying a sliding average algorithm to perform numerical processing on the capacity values in the capacity-time data point sequence;
[0020] An interpolation method or a polynomial fitting method is used for the processed capacity value to generate a continuous curve showing the change of capacity over time as the battery capacity decay curve.
[0021] Furthermore, the IMF component reflects the high-frequency and low-frequency fluctuation characteristics in the battery capacity change, and the residual component reflects the overall decay trend of the battery capacity.
[0022] Furthermore, the 3σ-linear regression detection and correction of the battery capacity decay curve specifically includes:
[0023] Perform linear regression on each capacity data point in the battery capacity decay curve to obtain the linear regression equation:
[0024]
[0025] in, is the fitted capacity value at the i-th time point, x i is the i-th time point, a and b are the slope and intercept of the fitting curve respectively;
[0026] Calculate the fitting residual based on the fitted capacity value and the original capacity value:
[0027]
[0028] Among them, e i is the fitting residual at the i-th time point, y i is the actual battery capacity value at the i-th time point;
[0029] The standard deviation σ of the residual value is calculated based on the fitting residual, and 3σ is used as the threshold to eliminate data points that meet the following conditions:
[0030]
[0031] Where n is the number of residual data points, that is, the total number of time points of the fitting curve, and μ is the fitting residual e i The mean of , σ is the standard deviation of the fitting residuals;
[0032] Retention Satisfaction|e i The capacity data points with |≤3σ constitute the corrected battery capacity decay curve.
[0033] Furthermore, the improved fully adaptive noise ensemble empirical mode decomposition method ICEEMDAN is used to decompose the corrected battery capacity decay curve, specifically including:
[0034] Step A1: Assume that the corrected battery capacity decay curve is signal x(t), initialize the residual r0(t) = x(t), and the order k = 1;
[0035] Step A2: Generate M groups of different Gaussian white noise sequences w j (t), the statistical characteristics of each set of noise are the same and independent of the signal;
[0036] Step A3: For the extraction of the k-th order IMF component, the residual signal r k-1 (t) adds the adaptive white noise set w j (t), and obtain multiple disturbance signals:
[0037]
[0038] in, is the ith disturbance signal, r k-1 (t) is the k-1th order residual signal, w i (t) is the i-th white noise signal, β i is the corresponding white noise amplitude coefficient;
[0039] Step A4: For each disturbance signal Perform local extreme point detection and connect the local maximum points to obtain the upper envelope curve Connect the local minimum points to get the lower envelope curve Compute the local mean curve:
[0040]
[0041] in, represents the upper envelope curve of the i-th disturbance signal, represents the lower envelope curve of the i-th disturbance signal, is the local mean curve of the i-th disturbance signal;
[0042] Step A5: Calculate the mean of the local mean curves of all disturbance signals to obtain the local mean of the k-th order IMF:
[0043]
[0044] Among them, m k (t) represents the local mean function of the k-th order IMF;
[0045] Step A6: Calculate the k-th order IMF component based on the local mean function:
[0046] c k (t) = r k-1 (t)-m k (t)
[0047] Among them, c k (t) is the k-th order intrinsic mode function IMF component;
[0048] Step A7: Update the residual signal:
[0049] r k (t) = r k-1 (t)-c k (t)
[0050] Among them, r k (t) is the k-th order residual signal;
[0051] Step A8: Repeat steps A3 to A7 until the residual r k (t) Satisfy the monotonic function condition or reach the preset decomposition order, and finally obtain the decomposition form of the signal:
[0052]
[0053] Where N is the total number of IMF components extracted, r N (t) is the final residual component.
[0054] Furthermore, the N IMF components are input into a Stacking ensemble learning model optimized and trained by a gravitational search algorithm (GSA) for prediction, specifically including:
[0055] Input N IMF components into multiple trained base models in the Stacking ensemble learning model for feature extraction and preliminary prediction processing;
[0056] The prediction results output by each base model are used as new input features to construct a prediction sequence. The prediction sequence is input into the meta-model in the Stacking ensemble learning model, which is used to integrate the prediction results of each base model and output the final prediction sequence for each IMF component.
[0057] Furthermore, the Stacking ensemble learning model construction and training process includes:
[0058] Use multiple machine learning models with different learning mechanisms as candidate base learners;
[0059] The gravitational search algorithm (GSA) is used to optimize and screen the combination of candidate base learners. The optimal base model combination screened by GSA is used to construct the base model of the stacking ensemble learning model. Each base model is trained using cross-validation, and its predicted output is constructed as a meta-training set.
[0060] The meta-training set is used to train the meta-model based on the BiLSTM model to obtain the Stacking ensemble learning model.
[0061] Furthermore, the candidate base learners include support vector regression SVR, random forest RF, gradient boosting regression tree GBRT, extreme gradient boosting XGBoost and other models.
[0062] Furthermore, the gravitational search algorithm (GSA) is used to optimize and screen the combination of candidate base learners, specifically including:
[0063] Each candidate base learner is regarded as a particle, and its loss value on multiple loss functions is used as a feature representation to form a feature vector:
[0064]
[0065] Among them, X i is the feature vector of the i-th candidate base learner; are the root mean square error, absolute error, and mean absolute percentage error loss functions of the i-th candidate base learner on the training samples, respectively. N is the number of candidate base learners.
[0066] The fitness based on the comprehensive loss evaluation index is used as the basis of gravitational mass to calculate the fitness function of the i-th particle. The formula is:
[0067]
[0068] Among them, fit i (t) is the fitness function value of the i-th candidate learner in the t-th iteration, α, β, and γ are the loss weighting coefficients respectively;
[0069] Calculate the normalized mass value for the GSA mechanical model:
[0070]
[0071] Among them, best(t) and worst(t) are the best fitness and worst fitness of all particles in the tth iteration respectively; m i (t) is the temporary mass of the i-th particle, M i (t) is the normalized gravitational mass of the i-th particle;
[0072] Simulate the mutual attraction between candidate base learners and update their positions in the feature space:
[0073]
[0074] R ij (t)=||X i (t)-X j (t)||2
[0075] in, is the gravitational component exerted by the j-th candidate learner on the i-th learner in the loss dimension d; G(t) is the gravitational constant in the t-th iteration; are the loss values of the jth and ith learners in the dth loss dimension respectively; ε is a minimum constant; R ij (t) is the distance between the i-th and j-th learners;
[0076] The position and velocity of the particles are updated through gravity, and after multiple rounds of iterations, several learners with the best fitness are selected as the final base learners of the Stacking base model layer.
[0077] Furthermore, the sparse attention bidirectional long short-term memory network model SBiLSTM includes a BiLSTM model, a sparse attention module and a fully connected layer connected in sequence.
[0078] Compared with the prior art, the present invention has the following advantages:
[0079] (1) The present invention uses the 3σ-linear regression method to detect and correct outliers. The capacity attenuation data of the existing technology is affected by measurement errors, environmental interference, occasional abnormal events, etc., which appear as outliers or mutations. Simple smoothing filtering is difficult to distinguish between systematic trends and abnormal mutations. By assuming that the residuals approximately obey the normal distribution, the 3σ-linear regression dynamically defines the outlier threshold in combination with the trend model, and can effectively identify non-random, abnormally deviating sampling points from the trend. By eliminating these outliers, the corrected capacity curve more realistically reflects the degradation law of the lithium battery and reduces noise interference. It overcomes the problem that the traditional method handles outliers roughly, resulting in abnormal noise affecting the weight update during model training, thereby increasing the prediction error and instability. It significantly improves data quality and reduces the risk of overfitting of the model due to outliers; the subsequent decomposition and prediction model are based on a more representative and physically meaningful capacity curve, and the prediction results are more accurate and reliable.
[0080] (2) The present invention adopts the ICEEMDAN (Completely Integrated Empirical Mode Decomposition) method to introduce adaptive noise multiple decompositions, peel off the different intrinsic mode functions (IMFs) in the capacity decay signal layer by layer, and obtain multiple natural frequency components and residuals. For traditional EMD, mode aliasing is prone to occur, and the physical significance components in the signal cannot be completely distinguished, and the noise affects the decomposition stability. ICEEMDAN effectively reduces mode aliasing and suppresses noise interference in the decomposition process by iteratively adding adaptive white noise. The separated IMF components are more stable and each has independent physical information, such as short-term fluctuations, periodic decay and long-term trends. It improves the failure phenomenon of existing empirical mode decomposition methods in complex signals, reduces inter-modal interference, and ensures multi-scale and accurate decomposition of capacity decay signals. It can more finely distinguish the noise, abnormal fluctuations and decay trends in the battery capacity signal, and improve the feature expression ability and generalization performance of the subsequent prediction model based on modal components.
[0081] (3) The present invention constructs a bidirectional LSTM network combined with a sparse attention mechanism, and inputs the residual component sequence into it. Sparse attention strengthens the weight distribution of key information moments while weakening the response to redundant or noisy moments. The residual component of lithium battery capacity contains complex nonlinear weak signals and long-term dependencies. Although traditional LSTM can capture time series information, it is easily interfered by a large amount of redundant data and consumes a lot of computing resources. The sparse attention mechanism optimizes the attention weight distribution through sparse constraints, so that the model focuses on a very small number of key time points and feature dimensions, strengthens the focus on key patterns in the residual, suppresses the spread of invalid attention, improves prediction accuracy and reduces the risk of overfitting, thereby more effectively extracting potential signal structures and changing trends. The bidirectional structure uses the complementary information of the previous and next contexts to further enhance the ability to understand the sequence. The present invention overcomes the problems of large noise in the residual signal, sparse information and mixed long-term and short-term dependencies, and solves the bottleneck of traditional LSTM "blindly learning" irrelevant time series features, resulting in overfitting and low efficiency. It greatly improves the model's sensitivity and discrimination ability to important time series patterns, reduces computational complexity, and improves prediction accuracy and model generalization performance.
[0082] (4) The performance of multiple candidate base learners is evaluated and screened through the gravitational search algorithm, the optimal combination is dynamically selected, and an efficient Stacking ensemble learning model is constructed. The optimal combination is selected from multiple sub-models to improve the ensemble prediction performance. The model has better generalization and adaptability, and is used to predict the various modal components after the capacity signal is decomposed. In multi-model ensemble prediction, directly combining all candidate learners is likely to lead to the addition of redundant and noisy models, affecting the overall performance. Traditional screening methods mostly rely on empirical rules or single evaluation indicators, lacking a global perspective and dynamic adjustment capabilities. The gravitational search algorithm simulates the gravitational effect between objects, giving each learner "quality" (performance index), allowing excellent models to have a strong "attraction", gradually gathering a group of high-performance models, and achieving global optimal screening. This not only avoids the influence of models with poor performance, but also ensures model diversity and representativeness. It solves the problems in the existing technology that learner screening relies on fixed standards, cannot dynamically adapt, has low screening efficiency, and is prone to falling into local optimality. It realizes the automated intelligent screening of candidate learners, improves the overall prediction ability and generalization performance of the Stacking model; at the same time, it reduces the noise and computational overhead caused by redundant models, and enhances the stability and training efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 is a flow chart of the prediction method of the present invention;
[0084] Figure 2 This is a model structure diagram of the Stacking ensemble learning model of the present invention;
[0085] Figure 3 Schematic diagram of the improved BiLSTM neural network SBiLSTM model of the present invention. DETAILED DESCRIPTION
[0086] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0087] This embodiment provides a lithium battery life prediction method based on modal decomposition and sparse attention, such as Figure 1 As shown, the following steps are included:
[0088] Obtaining a battery capacity time series of a battery to be predicted, wherein the battery capacity time series includes battery capacity values at multiple time points;
[0089] Based on the battery capacity time series, the battery capacity attenuation curve of the battery to be predicted is constructed, including:
[0090] Arrange the capacity values in the battery capacity time series in chronological order to generate a capacity-time data point sequence;
[0091] The capacity values in the capacity-time data point series are numerically processed using a sliding average algorithm;
[0092] An interpolation method or a polynomial fitting method is used for the processed capacity value to generate a continuous curve showing the change of capacity over time as a battery capacity decay curve.
[0093] Perform 3σ-linear regression detection and correction on the battery capacity attenuation curve to eliminate outliers in the capacity change process and retain valid values, including:
[0094] Perform linear regression on each capacity data point in the battery capacity decay curve to obtain the linear regression equation:
[0095]
[0096] in, is the fitted capacity value at the i-th time point, x i is the i-th time point, a and b are the slope and intercept of the fitting curve respectively;
[0097] Calculate the fitting residual based on the fitted capacity value and the original capacity value:
[0098]
[0099] Among them, e i is the fitting residual at the i-th time point, y i is the actual battery capacity value at the i-th time point;
[0100] The standard deviation σ of the residual value is calculated based on the fitting residual, and 3σ is used as the threshold to eliminate data points that meet the following conditions:
[0101] |e i |>3σ
[0102]
[0103] Where n is the number of residual data points, that is, the total number of time points of the fitting curve, and μ is the fitting residual e i The mean of , σ is the standard deviation of the fitting residuals;
[0104] Retention Satisfaction|e i The capacity data points with |≤3σ constitute the corrected battery capacity decay curve.
[0105] The improved fully adaptive noise ensemble empirical mode decomposition method ICEEMDAN is used to decompose the corrected battery capacity decay curve to obtain N intrinsic mode function (IMF) components and a residual component. The IMF component reflects the high-frequency and low-frequency fluctuation characteristics of the battery capacity change, and the residual component reflects the overall decay trend of the battery capacity, including:
[0106] Step A1: Assume that the corrected battery capacity decay curve is signal x(t), initialize the residual r0(t) = x(t), and the order k = 1;
[0107] Step A2: Generate M groups of different Gaussian white noise sequences w j (t), the statistical characteristics of each set of noise are the same and independent of the signal;
[0108] Step A3: For the extraction of the k-th order IMF component, the residual signal r k-1 (t) adds the adaptive white noise set w j (t), and obtain multiple disturbance signals:
[0109]
[0110] in, is the ith disturbance signal, r k-1 (t) is the k-1th order residual signal, w i (t) is the i-th white noise signal, β i is the corresponding white noise amplitude coefficient;
[0111] Step A4: For each disturbance signal Perform local extreme point detection and connect the local maximum points to obtain the upper envelope curve Connect the local minimum points to get the lower envelope curve Compute the local mean curve:
[0112]
[0113] in, represents the upper envelope curve of the i-th disturbance signal, represents the lower envelope curve of the i-th disturbance signal, is the local mean curve of the i-th disturbance signal;
[0114] Step A5: Calculate the mean of the local mean curves of all disturbance signals to obtain the local mean of the k-th order IMF:
[0115]
[0116] Among them, m k (t) represents the local mean function of the k-th order IMF;
[0117] Step A6: Calculate the k-th order IMF component based on the local mean function:
[0118] c k (t) = r k-1 (t)-m k (t)
[0119] Among them, c k (t) is the k-th order intrinsic mode function IMF component;
[0120] Step A7: Update the residual signal:
[0121] r k (t) = r k-1 (t)-c k (t)
[0122] Among them, r k (t) is the k-th order residual signal;
[0123] Step A8: Repeat steps A3 to A7 until the residual r k (t) Satisfy the monotonic function condition or reach the preset decomposition order, and finally obtain the decomposition form of the signal:
[0124]
[0125] Where N is the total number of IMF components extracted, r N (t) is the final residual component.
[0126] Input N IMF components into the Stacking ensemble learning model optimized and trained by the gravitational search algorithm GSA to predict and obtain the predicted IMF component sequence. The Stacking ensemble learning model is as follows: Figure 2 As shown;
[0127] The N IMF components are input into the Stacking ensemble learning model optimized and trained by the gravitational search algorithm (GSA) for prediction, including:
[0128] Input N IMF components into multiple trained base models in the Stacking ensemble learning model for feature extraction and preliminary prediction processing;
[0129] The prediction results output by each base model are used as new input features to construct a prediction sequence. The prediction sequence is input into the meta-model in the Stacking ensemble learning model. The meta-model is used to integrate the prediction results of each base model and output the final prediction sequence for each IMF component.
[0130] The stacking ensemble learning model construction and training process includes:
[0131] Multiple machine learning models with different learning mechanisms are used as candidate base learners, including support vector regression SVR, random forest RF, gradient boosting regression tree GBRT, extreme gradient boosting XGBoost and other models.
[0132] The gravitational search algorithm (GSA) is used to optimize and screen the combination of candidate base learners. The optimal base model combination screened by GSA is used to construct the base model of the stacking ensemble learning model. Each base model is trained using cross-validation, and its predicted output is constructed as a meta-training set.
[0133] The gravitational search algorithm (GSA) is used to optimize and screen the combination of candidate base learners, including:
[0134] Each candidate base learner is regarded as a particle, and its loss value on multiple loss functions is used as a feature representation to form a feature vector:
[0135]
[0136] Among them, X i is the feature vector of the i-th candidate base learner; are the root mean square error, absolute error, and mean absolute percentage error loss functions of the i-th candidate base learner on the training samples, respectively. N is the number of candidate base learners.
[0137] The fitness based on the comprehensive loss evaluation index is used as the basis of gravitational mass to calculate the fitness function of the i-th particle. The formula is:
[0138]
[0139] Among them, fit i (t) is the fitness function value of the i-th candidate learner in the t-th iteration, α, β, and γ are the loss weighting coefficients respectively;
[0140] Calculate the normalized mass value for the GSA mechanical model:
[0141]
[0142] Among them, best(t) and worst(t) are the best fitness and worst fitness of all particles in the tth iteration respectively; m i (t) is the temporary mass of the i-th particle, M i (t) is the normalized gravitational mass of the i-th particle;
[0143] Simulate the mutual attraction between candidate base learners and update their positions in the feature space:
[0144]
[0145] R ij (t)=||X i (t)-X j (t)||2
[0146] in, is the gravitational component exerted by the j-th candidate learner on the i-th learner in the loss dimension d; G(t) is the gravitational constant in the t-th iteration; are the loss values of the jth and ith learners in the dth loss dimension respectively; ε is a minimum constant; R ij (t) is the distance between the i-th and j-th learners;
[0147] The position and velocity of the particles are updated through gravity, and after multiple rounds of iterations, several learners with the best fitness are selected as the final base learners of the Stacking base model layer.
[0148] The meta-training set is used to train the meta-model based on the BiLSTM model to obtain the Stacking ensemble learning model.
[0149] To avoid getting stuck in a local optimum, the algorithm must initially employ exploration. As iterations progress, exploration decreases while exploitation increases. To improve GSA performance by controlling exploration and exploitation, only the Kbest best masses can attract the others. Kbest is a function of time, starting with K0 and decreasing over time. Initially, all masses exert a force, and Kbest decreases linearly over time, until only one mass exerts a force on the others. Thus, as particles with smaller masses approach particles with larger masses, the optimal solution to the optimization problem is gradually approached.
[0150] Stacking is an ensemble learning method that improves prediction performance by combining multiple different base models. Unlike bagging and boosting, the main idea of stacking is to use the prediction results of several base models as new features and input them into a meta-model, which then makes the final prediction. This method can leverage the strengths of different models to improve the accuracy and generalization of the overall prediction. In stacking, the first layer contains multiple base models, each of which is trained independently and predicts the training data. To make the prediction results of the base models more robust, cross-validation is often used to reduce the risk of overfitting. Through cross-validation, the first layer model will make predictions on the training data set, and these predictions are called meta-features. Each base model generates a set of prediction results, and the prediction results of multiple models are combined together as input features for the second layer model. The second layer model, also called the meta-model, takes the prediction results of the first layer model as input and outputs the final prediction results.
[0151] In stacking, the first layer contains multiple independently trained base models that make predictions on the training data through cross-validation to reduce the risk of overfitting. Each base model generates a set of predictions called meta-features. The predictions from multiple base models are combined together as input features for the second layer model. The second layer model (meta-model) takes the meta-features as input and outputs the final predictions. The specific training process is as follows:
[0152] Step 1. Divide the original dataset into P j The non-overlapping training set (j=1,2,3,…,i) is denoted as T n (n=1,2,3,…,m) test set.
[0153] Step 2. Perform K-fold cross-validation on the training set. In each cross-validation, K-1 folds of data are used as the training set, and the remaining 1 fold is used as the validation set. The three selected base models are trained separately and predictions are generated on each validation set.
[0154] Step 3. Combine the predictions generated by each base model in the cross-validation to form a new feature set Q = [Q1, Q2, Q3]. Q contains the predicted output of each base model in the training set. This new feature set is called meta-features.
[0155] Step 4. The generated meta-feature set (Q) is input into the meta-model SBiLSTM to obtain the final prediction.
[0156] In the proposed framework, the GSA algorithm is used as a selection tool to determine the most effective combination of base models to ensure optimal performance. To evaluate the model's predictive performance, the stacking method employs K-fold cross-validation. To avoid overfitting, a 10-fold cross-validation approach was chosen in this study. Compared to previous single prediction models and other ensemble prediction methods, the optimized stacking model, applied to a public dataset, exhibits two significant advantages. First, during the model selection process, the GSA algorithm optimizes the combination of base models and meta-models under different machine learning algorithm configurations, thereby extensively exploring algorithm combinations. This globally adaptive cascade optimization provides a robust strategy for selecting hybrid models, resulting in the model combination with the best predictive accuracy. During the selection and training phase of the ten base models, all base models are first trained as input, and the prediction results of each model are compared. Subsequently, the optimal combination of three models—XGBoost, RF, and CatBoost—is selected to achieve higher battery SOH estimation accuracy. Furthermore, to prevent overfitting, cross-validation is employed to avoid reusing the same data, thereby improving generalization. These algorithms demonstrate excellent predictive performance and are commonly used in the field of SOH prediction.
[0157]
[0158]
[0159] The 10 machine learning algorithms mentioned above were selected to form a regressor pool. To optimize the prediction performance of these machine learning algorithms on the dataset, a random search method was used to search for the optimal hyperparameters of the machine learning algorithms, as shown in Table 1. Subsequently, the GSA algorithm was used to find the optimal model combination for the regressor pool, thereby fully leveraging the advantages of the machine learning algorithms.
[0160] The residual component is input into the sparse attention bidirectional long short-term memory network model SBiLSTM for processing to obtain the predicted residual component sequence, where, Figure 3 As shown in Figure 1, the sparse attention bidirectional long short-term memory network model SBiLSTM includes a BiLSTM model, a sparse attention module and a fully connected layer connected in sequence.
[0161] The BiLSTM neural network consists of two different LSTM layers, each of which includes multiple LSTM units. Each LSTM layer operates independently. The forward layer processes the input in chronological order, while the reverse layer processes the input in reverse order and outputs the output. These units can learn long-term dependencies from the input data and use memory units, input gates, and other methods to store and process the data. t 、Forget Gate t and output gate O tMechanisms such as LSTM control the flow of information. Traditional RNNs are prone to vanishing or exploding gradients when processing long sequences, making the model difficult to train. LSTM, through its designed gating mechanism, can effectively alleviate these problems. However, LSTM networks only have the ability to capture information in one direction and cannot simultaneously utilize both past and future information in the sequence. Input data flows not only during forward propagation but also during backward propagation, thereby fully utilizing historical and future information and significantly alleviating the vanishing and exploding gradient problems. The calculation formula is as follows:
[0162]
[0163] Where * represents convolution, W f 、W O 、W i and W C Represent the weight matrices of the forget gate, output gate, input gate and unit state respectively, and b f 、b O 、b i and b C Represent the bias vectors of these gates respectively. In addition, C t is the current state value of the cell unit, h t is the current output value of the cell.
[0164] Sparse attention is an optimized attention mechanism that maps a query vector and a set of key-value pairs to an output vector. However, unlike single-head attention and multi-head attention, it does not calculate the similarity between the query vector and all key vectors. Instead, it only calculates the similarity between the query vector and some key vectors, thereby reducing computation and memory consumption. The specific formula is as follows:
[0165]
[0166] In which, by i ) is applied to the softmax function, which can achieve idempotence even if the weights are non-negative. The attention mechanism is described using Query, Key, and Values vectors, where Q represents the query vector, K represents the key vector, and V is the value vector.
[0167] SBiLSTM is used to selectively focus on important parts of the BiLSTM output sequence, further enhancing the model's expressiveness and computational efficiency. Models combining BiLSTM and sparse attention mechanisms leverage the strengths of both, simultaneously capturing both global and local information in sequence data while improving computational efficiency and model performance. BiLSTM is used to initially extract features to capture global contextual information from the input sequence. Subsequently, sparse attention is used to extract and emphasize important features. Finally, a fully connected layer maps multiple important features to a one-dimensional space to produce the final output.
[0168] The predicted value of the battery capacity of the battery to be predicted at the next moment is obtained based on the predicted IMF component sequence and the predicted residual component sequence.
[0169] In order to analyze the prediction performance, the root mean square error (RMSE), mean absolute percentage error (MAPE), and mean absolute percentage error (MAPE) metrics were selected to evaluate the performance of the proposed prediction framework. The performance of the model was further verified by comparing it with existing research models using the same battery aging dataset. The parameter settings of other models were determined using a random search method to find the optimal values within a certain range to ensure that each model can perform predictions well. Based on the NASA dataset and the University of Maryland dataset, the method in this paper was compared with CEEMDAN-LSTM, BiLSTM-Attention, BiGRU-Transformer, CNN-LSTM, and others. By observing Table 2, the RMSE, MAE, and MAPE of the model in this paper are the lowest. The present invention outperforms other state-of-the-art methods and can be adapted to different types of batteries, showing good generalization performance in both short-term and long-term cycles.
[0170] Table 2 Prediction error results for five different battery models
[0171]
[0172]
[0173] The photovoltaic power generation prediction method of the present invention, within this framework, first uses 3σ linear regression to detect and correct outliers by analyzing values and trends. The data is then decomposed into multiple IMFs and a residual sequence using the ICEEMDAN decomposition method. In the first stacking layer, the GSA optimization algorithm is used to select the three best combinations of base models. Each base model combination models and trains the input data to generate a new dataset. This new dataset is then used as input features for the meta-model. In the second layer of the stacking model, BiLSTM is applied as the meta-model to predict the data, resulting in the final battery capacity prediction results. The prediction results were good for three battery types using NASA data. However, for two battery types using the University of Maryland data, the capacity growth showed significant fluctuations due to the longer measurement period. Although the prediction accuracy was not as good as on the NASA data, the overall trend of the capacity growth fluctuations was still predicted.
[0174] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0175] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A lithium battery life prediction method based on modal decomposition and sparse attention, characterized in that: The following steps are involved: Obtaining a historical battery capacity time series of a battery to be predicted, wherein the historical battery capacity time series includes maximum battery capacities at multiple time points; Constructing a battery capacity attenuation curve of the battery to be predicted according to the battery capacity time series; Performing 3σ-linear regression detection and correction on the battery capacity decay curve to eliminate outliers in the capacity change process and retain valid values; The improved fully adaptive noise ensemble empirical mode decomposition method ICEEMDAN is used to decompose the corrected battery capacity decay curve to obtain N intrinsic mode function (IMF) components and one residual component. Inputting the N IMF components into the Stacking ensemble learning model optimized and trained by the gravitational search algorithm (GSA) for prediction, thereby obtaining a predicted IMF component sequence; The residual component is input into a sparse attention bidirectional long short-term memory network model SBiLSTM for processing to obtain a predicted residual component sequence; According to the predicted IMF component sequence and the predicted residual component sequence, a predicted value of the maximum battery capacity of the battery to be predicted at the next moment is obtained.
2. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 1 is characterized in that: The step of constructing a battery capacity attenuation curve of the battery to be predicted based on the battery capacity time series specifically includes: Arrange the capacity values in the battery capacity time series in chronological order to generate a capacity-time data point sequence; Applying a sliding average algorithm to perform numerical processing on the capacity values in the capacity-time data point sequence; An interpolation method or a polynomial fitting method is used for the processed capacity value to generate a continuous curve showing the change of capacity over time as the battery capacity decay curve.
3. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 1 is characterized in that: The IMF component reflects the high-frequency and low-frequency fluctuation characteristics of the battery capacity change, and the residual component reflects the overall decay trend of the battery capacity.
4. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 1 is characterized in that: The performing 3σ-linear regression detection and correction on the battery capacity attenuation curve specifically includes: Perform linear regression on each capacity data point in the battery capacity decay curve to obtain the linear regression equation: in, is the fitted capacity value at the i-th time point, x i is the i-th time point, a and b are the slope and intercept of the fitting curve respectively; Calculate the fitting residual based on the fitted capacity value and the original capacity value: Among them, e i is the fitting residual at the i-th time point, y i is the actual battery capacity value at the i-th time point; The standard deviation σ of the residual value is calculated based on the fitting residual, and 3σ is used as the threshold to eliminate data points that meet the following conditions: Where n is the number of residual data points, that is, the total number of time points of the fitting curve, and μ is the fitting residual e i The mean of , σ is the standard deviation of the fitting residuals; Retention Satisfaction|e i The capacity data points with |≤3σ constitute the corrected battery capacity decay curve.
5. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 1 is characterized in that: The improved fully adaptive noise ensemble empirical mode decomposition method ICEEMDAN is used to decompose the corrected battery capacity decay curve, specifically including: Step A1: Assume that the corrected battery capacity decay curve is signal x(t), initialize the residual r0(t) = x(t), and the order k = 1; Step A2: Generate M groups of different Gaussian white noise sequences w j (t), the statistical characteristics of each set of noise are the same and independent of the signal; Step A3: For the extraction of the k-th order IMF component, the residual signal r k-1 (t) adds the adaptive white noise set w j (t), and obtain multiple disturbance signals: in, is the ith disturbance signal, r k-1 (t) is the k-1th order residual signal, w i (t) is the i-th white noise signal, β i is the corresponding white noise amplitude coefficient; Step A4: For each disturbance signal Perform local extreme point detection and connect the local maximum points to obtain the upper envelope curve Connect the local minimum points to get the lower envelope curve Compute the local mean curve: in, represents the upper envelope curve of the i-th disturbance signal, represents the lower envelope curve of the i-th disturbance signal, is the local mean curve of the i-th disturbance signal; Step A5: Calculate the mean of the local mean curves of all disturbance signals to obtain the local mean of the k-th order IMF: Among them, m k (t) represents the local mean function of the k-th order IMF; Step A6: Calculate the k-th order IMF component based on the local mean function: c k (t)=r k-1 (t)-m k (t) Among them, c k (t) is the k-th order intrinsic mode function IMF component; Step A7: Update the residual signal: r k (t)=r k-1 (t)-c k (t) Among them, r k (t) is the k-th order residual signal; Step A8: Repeat steps A3 to A7 until the residual r k (t) Satisfy the monotonic function condition or reach the preset decomposition order, and finally obtain the decomposition form of the signal: Where N is the total number of IMF components extracted, r N (t) is the final residual component.
6. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 1, characterized in that: Inputting the N IMF components into the Stacking ensemble learning model optimized and trained by the gravitational search algorithm (GSA) for prediction specifically includes: Input N IMF components into multiple trained base models in the Stacking ensemble learning model for feature extraction and preliminary prediction processing; The prediction results output by each base model are used as new input features to construct a prediction sequence. The prediction sequence is input into the meta-model in the Stacking ensemble learning model, which is used to integrate the prediction results of each base model and output the final prediction sequence for each IMF component.
7. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 1 is characterized in that: The Stacking ensemble learning model construction and training process includes: Use multiple machine learning models with different learning mechanisms as candidate base learners; The gravitational search algorithm (GSA) is used to optimize and screen the combination of candidate base learners. The optimal base model combination screened by GSA is used to construct the base model of the stacking ensemble learning model. Each base model is trained using cross-validation, and its predicted output is constructed as a meta-training set. The meta-training set is used to train the meta-model based on the BiLSTM model to obtain the Stacking ensemble learning model.
8. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 7, characterized in that: The candidate base learners include support vector regression SVR, random forest RF, gradient boosting regression tree GBRT, extreme gradient boosting XGBoost and other models.
9. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 7, characterized in that: The gravitational search algorithm GSA is used to optimize and screen the combination of candidate base learners, specifically including: Each candidate base learner is regarded as a particle, and its loss value on multiple loss functions is used as a feature representation to form a feature vector: Among them, X i is the feature vector of the i-th candidate base learner; are the root mean square error, absolute error, and mean absolute percentage error loss functions of the i-th candidate base learner on the training samples, respectively. N is the number of candidate base learners. The fitness based on the comprehensive loss evaluation index is used as the basis of gravitational mass to calculate the fitness function of the i-th particle. The formula is: Among them, fit i (t) is the fitness function value of the i-th candidate learner in the t-th iteration, α, β, and γ are the loss weighting coefficients respectively; Calculate the normalized mass value for the GSA mechanical model: Among them, best(t) and worst(t) are the best fitness and worst fitness of all particles in the tth iteration respectively; m i (t) is the temporary mass of the i-th particle, M i (t) is the normalized gravitational mass of the i-th particle; Simulate the mutual attraction between candidate base learners and update their positions in the feature space: R ij (T)=||X i (t)-X j (t)||2 in, is the gravitational component exerted by the j-th candidate learner on the i-th learner in the loss dimension d; G(t) is the gravitational constant in the t-th iteration; are the loss values of the jth and ith learners in the dth loss dimension respectively; ε is a minimum constant; R ij (t) is the distance between the i-th and j-th learners; The position and velocity of the particles are updated through gravity, and after multiple rounds of iterations, several learners with the best fitness are selected as the final base learners of the Stacking base model layer.
10. The lithium battery life prediction method based on modal decomposition and sparse attention according to claim 1, characterized in that: The sparse attention bidirectional long short-term memory network model SBiLSTM includes a BiLSTM model, a sparse attention module and a fully connected layer connected in sequence.
Citation Information
Patent Citations
Lithium battery remaining service life prediction method and system based on optimized neural network
CN117665627A