Lithium ion battery health state estimation method and system based on dual-drive interpretable integrated model

By constructing a dual-driven interpretable ensemble model, combining multi-source health features and ensemble learning methods, the problems of single feature sources and insufficient model generalization ability in lithium-ion battery health status monitoring are solved, achieving high-precision and interpretable battery aging status estimation.

CN120870926APending Publication Date: 2025-10-31HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510995654.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing methods for monitoring the health status of lithium-ion batteries suffer from problems such as limited feature sources, insufficient model generalization ability, and poor interpretability.

Method used

We adopt a dual-drive interpretable ensemble model, combined with a multi-source health feature space and an ensemble learning model. By extracting the peak features of the incremental capacity curve, the ohmic internal resistance parameters of the equivalent circuit model, and the time-domain statistical features of the voltage curve, we construct a multi-dimensional feature representation system. We then use a heterogeneous ensemble learning model with a stacking architecture to estimate the health status of lithium-ion batteries.

Benefits of technology

It improves the accuracy and generalization ability of lithium-ion battery health state estimation, reduces estimation error, enhances model interpretability and anti-interference ability, and achieves comprehensive and accurate assessment of battery aging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120870926A_ABST
    Figure CN120870926A_ABST
Patent Text Reader

Abstract

The invention provides a lithium ion battery health state estimation method and system based on a dual-drive interpretable integrated model, and belongs to the field of lithium ion battery health state estimation. The problems that an existing lithium battery health state monitoring method is single in feature source, insufficient in model generalization ability and poor in interpretability are solved. According to the method, based on an incremental capacity curve and a first-order RC equivalent circuit model, IC peak value features and ohmic internal resistance features are extracted, and a multi-source health feature space is constructed in combination with voltage statistical features; an integrated learning model based on Stacking is established, three basic models of a random forest, kernel ridge regression and an interpretable enhancement machine are integrated, and collaborative optimization of model hyper-parameters is realized by adopting a tree structure-based Bayesian optimization algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of lithium-ion battery health state estimation technology, and more specifically, to a lithium-ion battery health state estimation method and system based on a dual-drive interpretable ensemble model. Background Technology

[0002] With the acceleration of global industrialization, the excessive consumption of traditional fossil fuels has triggered a profound energy crisis and environmental problems. The fossil fuel-dominated energy structure not only exacerbates the depletion of non-renewable resources but also further worsens global warming due to the continuous emission of greenhouse gases such as carbon dioxide. To address this severe challenge, new energy vehicles have emerged. Powered by onboard batteries, they represent an innovative combination of electrochemical energy storage technology and novel power systems. By replacing traditional fuel-powered vehicles, new energy vehicles offer a viable path to energy conservation, emission reduction, and lower carbon intensity in the transportation sector.

[0003] In the electric vehicle sector, the performance of lithium-ion batteries directly affects a vehicle's range, power performance, and safety. In energy storage systems, lithium-ion batteries assist the power grid in load regulation and enhance its ability to absorb renewable energy. However, during use, lithium-ion batteries are affected by factors such as electrochemical side reactions and changes in electrode material structure. Increased charge-discharge cycles lead to a gradual decrease in capacity and an increase in internal resistance. Under complex environmental conditions, high temperatures accelerate the rate of internal chemical reactions, while high-rate charge-discharge cycles easily cause stress concentration in the electrode materials, thus accelerating battery aging. These problems not only reduce equipment performance but may also trigger serious safety accidents such as electric vehicle breakdowns and energy storage station fires. Therefore, accurately monitoring battery health is crucial. On the one hand, lithium-ion batteries inevitably age and experience capacity degradation after prolonged use, which severely affects their performance and safety, and may even lead to equipment failure or catastrophic accidents. On the other hand, premature battery replacement results in a significant waste of battery resources.

[0004] The Battery Management System (BMS), as the control core of the battery system, undertakes several key tasks, including signal acquisition and communication, state assessment, charge and discharge management, fault warning, and thermal regulation. Among these functions, the BMS needs to focus on monitoring the state parameters such as State of Charge (SOC), State of Health (SOH), and State of Energy (SOE). It is worth noting that accurate SOH estimation not only provides a basis for correcting SOC estimation but also directly affects the accuracy of remaining range prediction for electric vehicles. SOH estimation plays a central role in the energy management of battery packs and electric vehicles, effectively ensuring the safe operation of the battery system. Furthermore, SOH estimation lays an important foundation for formulating reasonable battery usage strategies and extending battery life. Therefore, in-depth exploration of the intrinsic mechanisms of lithium-ion battery degradation and the achievement of non-destructive monitoring and estimation of battery SOH are crucial not only for accurately assessing the remaining battery capacity and building an efficient fault warning system but also for ensuring stable battery operation.

[0005] State of Health (SOH) is an effective indicator of the aging state of lithium-ion batteries and a key metric in battery management systems. Accurate and efficient SOH estimation has long attracted in-depth research from researchers both domestically and internationally, resulting in numerous achievements and the development of various methods for SOH measurement and estimation. These methods can be categorized into experimental measurement methods, model-driven methods, and data-driven methods. However, while experimental measurement methods are simple and direct, they often require specialized laboratory equipment operated by professionals, making them highly susceptible to environmental limitations. Model-based methods offer high accuracy, but the computational burden of establishing accurate models is significant, accurate identification of model parameters is challenging, and generalization performance across different battery models is poor. Data-driven methods exhibit good adaptability and high fitting accuracy, but prediction accuracy is heavily influenced by the size of the training sample data and the selection of health features, and interpretability is poor. Summary of the Invention

[0006] The technical problem to be solved by this invention is:

[0007] To address the problems of existing lithium battery health status monitoring methods, such as limited feature sources, insufficient model generalization ability, and poor interpretability.

[0008] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:

[0009] This invention provides a method for estimating the state of health of lithium-ion batteries based on a dual-drive interpretable ensemble model, comprising the following steps:

[0010] S100. The average value, standard deviation, variance, root mean square, skewness, and kurtosis are obtained from the charging voltage curve. The peak characteristics of the IC are obtained based on the voltage curve, and the ohmic internal resistance characteristics are obtained based on the incremental capacity curve. The above characteristics are used to construct a multi-source health feature space.

[0011] S200. Build a dual-drive ensemble learning model, integrating three basic models: random forest, kernel ridge regression, and interpretable augmentation machine. Select logistic regression as the second layer model of the dual-drive ensemble learning model. Use the multi-source health feature space obtained in step S100 as the feature input. Use a tree-based Bayesian optimization algorithm to achieve collaborative optimization of model hyperparameters. Use the optimized dual-drive ensemble learning model to estimate the health status of lithium-ion batteries.

[0012] Further, in step S100, the statistical characteristics of the mean, standard deviation, variance, root mean square, skewness, and kurtosis are standardized using the following formula:

[0013]

[0014] In the formula: n is the number of sample points, V i V represents the specific voltage value corresponding to sample point i. mean V represents the sample mean. std V represents the sample standard deviation. var V represents the sample variance. rms V represents the root mean square of the sample. skew V represents sample skewness. kurt Represents the kurtosis of the sample.

[0015] Furthermore, in step S100, when obtaining the IC peak characteristics, according to the incremental capacity analysis method, the IC curve is obtained by calculating the first derivative of the VQ curve of the lithium-ion battery during the charging process:

[0016]

[0017] Where I is the current, t is the charging time, Q is the capacity, and V is the voltage;

[0018] Under constant current charging conditions, the battery terminal voltage is divided into several equally spaced voltage segments, including: recording the charging time Δt of each segment when the battery terminal voltage rises by ΔV; then, calculating the charged capacity ΔQ within the voltage segment using the product of Δt and the charging current; finally, approximating the differential capacity value dQ / dV at the corresponding voltage point by the ratio of ΔQ to ΔV; continuing this process until the battery reaches the set cutoff voltage, thus completing the entire charging process; when ΔV approaches zero, a series of continuous differential capacity data points can be obtained according to formula (8); after interpolating and fitting these discrete data points, a smooth incremental capacity curve is generated:

[0019]

[0020] Kalman filtering was used to smooth the IC curve obtained by the aforementioned formula, including prediction and updating. In the prediction step, the state and its covariance at the next time step were predicted based on the current state and the state transition model.

[0021]

[0022]

[0023] In the formula: It uses the result predicted from the previous state; It is the optimal result of the previous state; F is the state transition matrix; and P k-1 They are and The corresponding covariance; Q′ is the covariance of the system process;

[0024] In the update step, the prediction results are corrected using the newly acquired measurement data to obtain the state estimate. The update stage includes Kalman gain calculation, state update, and covariance update, with the corresponding formulas as follows:

[0025]

[0026] In the formula: K k H is the Kalman gain at time k; H is the observation matrix; R is the measurement noise covariance. It is the optimal estimate at time k; z k is the measurement value at time k; E is the identity matrix;

[0027] When the battery voltage curve shows a stable voltage plateau, the amount of charge entering the battery increases within the time range during which the plateau is maintained, while the change in voltage is close to zero. This electrochemical process is represented by a significant peak on the IC curve, and the peak in the IC curve is used as a health characteristic to characterize the battery state.

[0028] Furthermore, in step S100, when obtaining the ohmic internal resistance characteristics, the first-order RC equivalent circuit model is:

[0029]

[0030] In the formula: U t U is the terminal voltage, U1 is the polarization voltage, U oc R0 is the open-circuit voltage, R1 is the internal resistance in ohms, R1 is the polarization internal resistance, and C1 is the polarization capacitor; I t Represents the load current;

[0031] By using discrete polarization voltage, we can derive:

[0032]

[0033] Among them, U 1,k+1 U is the polarization voltage at time k+1. 1,k Let I be the polarization voltage at time k. t,k This represents the load current at time k;

[0034] Within a sampling time, U is approximately considered to be oc,k+1 ≈U oc,k According to equations (14) and (15), we get:

[0035]

[0036] Among them, U oc,k U is the open-circuit voltage at time k. t,k+1 Let U be the terminal voltage at time k+1. t,k Let I be the terminal voltage at time k. t,k+1 Represents the load current at time k+1; R p For load resistance;

[0037] Simplifying the above equation, we get:

[0038] U t,k+1 =a1U t,k +a2+a3I t,k+1 +a4I t,k (17)

[0039] in:

[0040]

[0041] The parameters a1 to a4 are identified to obtain the value of R0; the sum of squared prediction errors is minimized using the recursive least squares method, the mathematical expression of which is:

[0042]

[0043] Where: λ is the forgetting factor, 0 < λ ≤ 1, used to adjust the weight of historical data; θ is the parameter vector; J(θ) represents the objective function; U(i) is the true voltage value at time i. This represents the voltage estimate at time i when the selected parameter is θ;

[0044] The recursive formula consists of three parts: gain matrix update, covariance matrix update, and parameter estimation update. For the k-th time step, the specific formula is derived as follows:

[0045] (1) Gain matrix update:

[0046]

[0047] (2) Covariance matrix update:

[0048]

[0049] (3) Parameter estimation update:

[0050] θ(k)=θ(k-1)+K(k)(U(k)-φ T (k)θ(k-1)) (22)

[0051] Where: K(k) is the Kalman gain at the k-th time after the update, P(k) is the covariance matrix at the k-th time after the update, θ(k) is the parameter vector at the k-th time, and φ(k) is the regression vector at the k-th time;

[0052] After identifying the parameters of the equivalent circuit model using the recursive least squares method, the ohmic internal resistance R0 was obtained.

[0053] Further, in step S200, the steps of the Stacking integration architecture model are as follows:

[0054] (1) Dataset partitioning: First, the training dataset is divided into multiple subsets using a cross-validation strategy, and K-fold cross-validation is selected; the multi-source health features obtained in step S100 are divided into K mutually exclusive subsets, and each time a subset is selected as the validation set, and the remaining K-1 subsets are used as the training set;

[0055] (2) Basic model training and prediction: Under the cross-validation framework, multiple basic models are trained in parallel; for each basic model, in each round of cross-validation, it is trained using K - 1 fold training data; the trained basic model makes predictions on the validation set, and the prediction results of all rounds are combined to obtain the prediction result of this basic model on the entire training set; at the same time, the basic model is used to make predictions on the test set to obtain the corresponding prediction results;

[0056] (3) Feature combination: The prediction result of each basic model is used as a new feature to construct a meta-model training data set; the same operation is also performed on the test set, and the prediction result of the basic model on the test set is used as a new feature, which is combined with the original test set to form a new test data set;

[0057] (4) Meta-model training: Use the new training data set to train the meta-model, enabling it to learn how to combine the prediction results of the basic models to achieve complementary fusion of model advantages and maximize the accuracy of the overall model.

[0058] (5) Prediction: Use the trained meta-model to make predictions on the test data. The meta-model combines and weights according to the prediction results of the basic models to generate the final prediction result.

[0059] Furthermore, in step S200, the random forest constructs multiple decision trees by combining the bootstrap method and the random subspace method. From a set with a total of M features, m features are randomly selected each time for splitting, where m < M, and a single decision tree is generated by optimizing the information gain; finally, the average of multiple regression results is taken as the output prediction.

[0060] Furthermore, in step S200, kernel ridge regression借助核函数的力量,将原始特征空间映射至一个高维空间,随后在这个新构建的高维空间内执行岭回归分析;所采用的损失函数KernelRidge Loss表述为:

[0061]

[0062] In the formula: y is the target variable; X is the feature matrix; w is the regression coefficient; α is the regularization parameter; K(X, X) is the kernel matrix, and each of its elements K i,j is X i and X j The inner product in the high-dimensional space.

[0063] Furthermore, in step S200, the interpretable boosting machine uses modern machine learning techniques to learn each feature function f i , considering the interaction between two features ∑f ij (x i , x j It should be noted that there seems to be an incorrect expression "借助核函数的力量" in item (15). It might be a misspelling or an incomplete description. If it's a specific Chinese term that needs to be accurately translated, more context or clarification is required. The above translation is based on the existing text as accurately as possible.This is used to improve the accuracy of the model.

[0064]

[0065] In the formula: g(·) is the connection function; m is the number of features; x i It is the observed feature of feature x at time i; f i E(y) is the smoothing function of the feature at time i; E(y) is the expected value; β is the intercept; and ε is the residual.

[0066] Further, in step S200, a tree-structured Parzen estimator is used for hyperparameter optimization, including initial random sampling, classification of observations, construction of a probabilistic model, calculation of expected improvement, selection of new parameter combinations, performance evaluation, and updating of the dataset. Specifically, this includes...

[0067] (1) Initialize random sampling: Randomly sample the hyperparameter space, generate multiple sets of hyperparameter combinations and evaluate their performance, and accumulate initial data;

[0068] (2) Determine the initial sampling status: Check whether the preset number of initial samplings has been completed. If not, continue random sampling; if completed, proceed to the data partitioning stage.

[0069] (3) Divide into high-quality and ordinary parameter groups: Based on the evaluation results, the top γ% of hyperparameter combinations are divided into the Top group, and the rest are the remaining group, to distinguish the distribution characteristics of high-quality and ordinary parameters.

[0070] (4) Constructing a probability model: For the Top group, construct a probability model l(θ) to characterize the distribution of high-quality hyperparameters; for the remaining groups, construct a probability model g(θ) to describe the distribution of ordinary hyperparameters.

[0071] (5) Calculate the EI value: Calculate the EI value γ(θ) of each candidate hyperparameter using formula (25) to measure its probability advantage of belonging to the high-quality group and guide the search direction:

[0072]

[0073] (6) Select the optimal candidate parameters: Select the hyperparameter combination with the largest EI value and prioritize exploring the parameters most likely to optimize model performance;

[0074] (7) Evaluate the performance of the new parameters: Apply the selected hyperparameters to the target model and evaluate the performance through methods such as cross-validation to obtain new evaluation results;

[0075] (8) Determine the termination condition: Check whether the termination condition is met. If not, integrate the new data into the historical data, re-divide the groups, build the model, and iterate in a loop. If the condition is met, output the current optimal hyperparameter combination.

[0076] The lithium-ion battery health state estimation system based on the dual-drive interpretable ensemble model has a program module corresponding to the above steps, and executes the steps in the lithium-ion battery health state estimation method based on the dual-drive interpretable ensemble model during runtime.

[0077] A computer-readable storage medium storing a computer program configured to, when invoked by a processor, implement steps of a lithium-ion battery health state estimation method based on a dual-drive interpretable ensemble model.

[0078] Compared with the prior art, the beneficial effects of the present invention are:

[0079] (1) This invention integrates the advantages of model-driven and data-driven methods to construct a multi-source health feature space that includes physical mechanism features and statistical features. By extracting the peak features of the incremental capacity curve, the ohmic internal resistance parameters of the equivalent circuit model, and combining them with the time-domain statistical features of the voltage curve (mean, variance, skewness, etc.), a multi-dimensional feature representation system is formed. In the feature correlation analysis, based on the results of Kendall correlation coefficient and grey relational analysis, it is shown that the extracted features have a high correlation with SOH.

[0080] (2) This invention designs a heterogeneous ensemble learning model based on a Stacking architecture. This model combines three basic models—RF, KRR, and EBM—and introduces LR as a meta-model, achieving complementary advantages among different models. Furthermore, the model's hyperparameters are globally optimized using the TPE algorithm. Experimental results on the Oxford University dataset show that the ensemble learning model has an average RMSE of 0.331% and an average MAE of 0.234%. Compared to the basic models and several common neural network models, it has the lowest estimation error, and its training efficiency is more than 60% higher than that of neural network models. This verifies the model's dual advantages in accuracy and efficiency.

[0081] (3) The results of the feature ablation experiment show that the model estimation error of the fusion of three types of features (S+IC+R) is the lowest compared with other feature combinations. This indicates that the synergistic effect of multi-source features improves the limitations of single feature sources and enhances the comprehensiveness of battery aging characterization to a certain extent.

[0082] (4) This invention utilizes SHAP analysis to quantify the global contribution of each feature, revealing the physical mechanism of model decision-making. The analysis results show that the peak IC curve feature and the ohmic internal resistance feature play a core role, providing physical mechanism support for the model's decision-making. The peak IC curve feature and the ohmic internal resistance feature have the highest contribution to SOH prediction (their SHAP values ​​account for more than 60%), which further verifies their strong correlation with battery aging mechanisms (such as active lithium loss and electrode structure degradation).

[0083] (5) This invention verifies the adaptability of the proposed model under different battery types and operating conditions. Noise interference experiments show that, under 20% and 40% random noise, the relative increase in estimation error of the ensemble learning model is significantly lower than that of the base model. Cross-dataset validation results based on the Huazhong University of Science and Technology dataset show that the estimation error of the ensemble learning model on each battery sample is less than 2%, and lower than that of the neural network model. These experimental results demonstrate that the proposed model possesses excellent anti-interference ability and good generalization performance. Attached Figure Description

[0084] Figure 1 This is a diagram showing the change in battery terminal voltage during the charging phase in an embodiment of the present invention;

[0085] Figure 2 This is a trend graph showing the changes of statistical features F1-F6 with the number of iterations in an embodiment of the present invention;

[0086] Figure 3 This is a voltage-capacity relationship diagram during the charging process in an embodiment of the present invention;

[0087] Figure 4 This is a single-cycle IC curve diagram before filtering in an embodiment of the present invention;

[0088] Figure 5 This is the IC curve of Cell1 in the Oxford dataset after filtering and smoothing for all iterations in this embodiment of the invention.

[0089] Figure 6 This is a trend graph showing the change of IC peak value with the number of cycles in an embodiment of the present invention;

[0090] Figure 7 This is a trend graph showing the change of internal resistance value with the number of cycles in an embodiment of the present invention;

[0091] Figure 8 This is a heatmap of the Kendall correlation coefficients between battery features F1-F8 and SOH in an embodiment of the present invention.

[0092] Figure 9 This is a gray relational coefficient heatmap of battery features F1-F8 and SOH in an embodiment of the present invention.

[0093] Figure 10 This is a diagram of the Stacking algorithm structure for data training and testing based on cross-validation in an embodiment of the present invention.

[0094] Figure 11 This is a schematic diagram of the random forest principle in an embodiment of the present invention;

[0095] Figure 12 This is a flowchart illustrating the principle of the TPE optimization algorithm in this embodiment of the invention;

[0096] Figure 13 This is a diagram illustrating the structure of the dual-drive SOH estimation model constructed using the Stacking ensemble learning framework based on cross-validation in this embodiment of the invention.

[0097] Figure 14 This is a curve comparing the SOH prediction results of the integrated learning model and each neural network model in the embodiments of the present invention;

[0098] Figure 15 This is a SHAP result diagram using Cell1 battery as an example in an embodiment of the present invention; Detailed Implementation

[0099] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0100] Lithium-ion battery health status definition

[0101] The State of Health (SOH) of a lithium-ion battery is a key indicator for measuring the degree of battery performance degradation, usually expressed as a percentage. It reflects the changes in various performance parameters of the battery in its current state compared to its pristine state, encompassing multiple dimensions such as battery capacity, internal resistance, and charge / discharge efficiency. From the perspective of battery capacity, SOH equals the battery's current usable capacity C. now With the battery's initial rated capacity C rat The ratio of the two values ​​directly reflects the degree of battery capacity degradation, and its formula is shown in equation (1):

[0102]

[0103] Changes in internal resistance are also significant in SOH estimation. As the battery ages, its internal resistance gradually increases, leading to increased energy loss during charging and discharging, and consequently, a decline in performance. The SOH expression in terms of internal resistance is shown in equation (2):

[0104]

[0105] In the formula, Reol R represents the internal resistance of a battery when it reaches a failure state. new R represents the initial internal resistance of the battery, and R represents the internal resistance of the battery in its current state.

[0106] Quantitative evaluation of battery internal resistance relies on specific experimental methods, such as electrochemical impedance spectroscopy based on frequency domain response and hybrid pulse power detection based on time domain characteristics. It is worth noting that these methods are significantly dependent on environmental parameters (such as state of charge range and electrical contact impedance), and their measurement accuracy is easily affected by fluctuations in operating conditions. In contrast, capacity parameters are obtained through coulomb counting under constant current charge-discharge cycles, and their values ​​are strongly correlated with the core control variables of battery operation (current density, temperature field distribution, etc.). Therefore, this study selects SOH, defined based on capacity decay, as the core evaluation parameter for battery performance, as shown in equation (1).

[0107] Introduction to Battery Aging Dataset

[0108] Data, as a driving force for innovative technologies and interdisciplinary collaboration, is building bridges for collaborative innovation among experts from different fields. In the field of battery research, experimental datasets serve as a crucial foundation for developing predictive methods, and their quality and diversity directly impact the depth of technological breakthroughs. Lithium-ion battery research is currently a hot topic internationally, and numerous research institutions have constructed various types of battery aging datasets, providing important support for revealing the laws governing battery degradation. This invention will select two typical open-source datasets for analysis: one is the Oxford dataset provided by the University of Oxford, which is widely cited in the international academic community; the other is a battery aging dataset developed in recent years by a team from Huazhong University of Science and Technology. The specific information and corresponding aging experimental schemes for each dataset will be introduced below.

[0109] Model performance evaluation metrics

[0110] Model performance evaluation metrics are used to quantitatively analyze the deviation between model estimates and actual results, and to compare the performance of different models. Battery health state estimation is essentially a regression problem. In regression analysis, root mean square error (RMSE) and mean absolute error (MAE) are two commonly used evaluation metrics that effectively measure the deviation between model estimates and actual values.

[0111] RMSE is more sensitive to larger errors and can highlight the impact of outliers, while MAE is a better choice for assessing the overall error level while avoiding the influence of outliers. In practical applications, these two metrics can be used together to comprehensively evaluate the model's performance.

[0112] Furthermore, absolute error refers to the difference between the estimated value and the actual value. In this invention, its absolute value is calculated to reflect the volatility of the model estimation results.

[0113] Health Feature Extraction Based on Voltage Curve

[0114] In practical applications, due to varying usage scenarios, the battery discharge process is characterized by complex operating conditions and diverse circumstances, making it difficult to standardize and unify the discharge process, and also posing significant challenges to data collection. In contrast, the charging phase typically follows a relatively fixed pattern. Therefore, this invention conducts data analysis and extracts health characteristics based on the common constant current (CC) charging process.

[0115] During the charging and discharging process of a lithium-ion battery, external sensors can be used to collect voltage data from the battery. For example... Figure 1 As shown in the figure, taking Cell1 from the Oxford University dataset as an example, the graph illustrates the trend of battery terminal voltage over time at different cycle stages. It can be seen that during the constant current charging stage, the battery terminal voltage gradually increases from 2.7V to 4.2V, and the time consumed in this process gradually decreases with the increase of the number of cycles.

[0116] Statistical features can accurately characterize the shape and position of the battery voltage curve in a numerical way, revealing the evolution law of battery aging. Based on the voltage curve, this invention extracts a series of commonly used voltage statistical features, specifically including mean, standard deviation, variance, root mean square, skewness, and kurtosis, which are labeled as health features F1-F6 respectively. The mean (F1) is obtained by calculating the arithmetic mean of the voltage sequence, which reflects the overall voltage level of the battery during the charging and discharging process; the standard deviation (F2) and variance (F3) are used to quantify the dispersion of voltage data; the root mean square (F4) combines the amplitude and fluctuation characteristics of voltage; skewness (F5) and kurtosis (F6) are used to describe the morphological characteristics of voltage distribution, where skewness measures the asymmetry of data distribution, and kurtosis reflects the steepness of data distribution. These statistical features, through standardized calculations of equations (3) to (8), can effectively capture the voltage evolution law during the battery aging process.

[0117]

[0118] In the formula: n is the number of sample points, V i V represents the specific voltage value corresponding to the sample point. mean V represents the sample mean. std V represents the sample standard deviation. var V represents the sample variance. rms V represents the root mean square of the sample. skew V represents sample skewness.kurt Represents the kurtosis of the sample.

[0119] After calculating the voltage statistical characteristics, the above characteristics were normalized, and their relationship with the number of battery cycles was analyzed. Figure 2 This is a trend graph showing the changes of these six features with the number of iterations. As can be seen from the graph, each feature generally exhibits a monotonically increasing or decreasing trend. Specifically, the average value V... mean (F1), Root Mean Square V rms (F4) and kurtosis V kurt (F6) increases with the number of battery cycles, standard deviation V std (F2), Variance V var (F3) and skewness V skew (F5) decreases as the number of battery cycles increases. Therefore, it can be preliminarily determined that the extracted F1-F6 features can effectively reflect the aging status of the battery.

[0120] Health Feature Extraction Based on Incremental Capacity Curve

[0121] Incremental Capacity Analysis (ICA) is a non-invasive battery health assessment technique based on the differential processing of charge-discharge curves. This method obtains voltage-capacity (VQ) data through constant current charge-discharge experiments and generates an incremental capacity (IC) curve. The peak position, peak height, and peak shape evolution of the characteristic peaks in this curve are directly related to aging mechanisms such as active lithium loss and electrode structure degradation. By analyzing the morphological characteristics of the IC curve, the battery's health status can be monitored in real time without disassembling the battery, providing crucial data support for the Battery Management System (BMS). Its core advantage lies in transforming complex electrochemical degradation mechanisms into intuitive and measurable signals, making it an important analytical tool in the field of online health monitoring.

[0122] Collect relatively stable constant current charging data to perform capacity increment analysis on lithium-ion batteries. Figure 3 The voltage-capacity curves for a single charge process from the Oxford dataset are shown. Analysis of these curves reveals that they can be divided into three stages:

[0123] (1) Initial stage: In the initial stage where the capacity starts to increase from 0, the voltage rises rapidly. When the capacity is at a low level, the battery voltage rises rapidly from about 2.7V. No obvious voltage plateau appears in this stage, indicating that the electrochemical reaction inside the battery is relatively active and the electrode polarization is strong.

[0124] (2) Voltage Plateau Stage: As the capacity increases to a certain level, approximately between 0.1Ah and 0.4Ah, the voltage rise slows down, reaching a relatively stable stage, i.e., the voltage plateau. The voltage stabilizes at around 3.7V, which is the typical voltage plateau region of this battery. Within this plateau range, the battery stores energy at a relatively stable voltage, indicating that the electrochemical reactions inside the battery are relatively stable, and the lithium-ion insertion process between the positive and negative electrodes is relatively uniform.

[0125] (3) Post-plateau stage: After passing the voltage plateau, as the capacity further increases, the voltage begins to rise significantly again to 4.2V. This means that the battery is close to being fully charged, and the internal electrochemical reactions gradually change. Factors such as electrode polarization cause the voltage to rise again.

[0126] Based on the basic principle of incremental capacity analysis, the IC curve is obtained by calculating the first derivative of the VQ curve of the lithium-ion battery during the charging process. Its specific mathematical expression is shown in equation (9).

[0127]

[0128] Where I is the current, t is the charging time, Q is the capacity, and V is the voltage;

[0129] Under constant current charging conditions, the battery terminal voltage can be divided into several equally spaced voltage segments. The specific operation is as follows: First, when the battery terminal voltage rises by ΔV, the charging time Δt of that segment is recorded; then, the charged capacity ΔQ within that voltage segment is calculated using the product of Δt and the charging current; finally, the differential capacity value dQ / dV at the corresponding voltage point is approximately obtained by using the ratio of ΔQ to ΔV. This process continues until the battery reaches the set cutoff voltage, thus completing the entire charging process. Theoretically, when ΔV approaches zero, a series of continuous differential capacity data points can be obtained according to formula (10). After interpolating and fitting these discrete data points, a smooth incremental capacity curve can be generated. This method achieves accurate quantification of the charging process by dividing the voltage range, providing necessary data support for subsequent health feature extraction.

[0130]

[0131] In actual voltage sampling and calculation, due to limitations in the sampling accuracy of the hardware system, it is impossible to achieve a voltage interval ΔV that theoretically approaches zero. After experimental testing, ΔV = 20mV was selected as the calculation step size. Figure 4The calculated single-cycle IC curve is presented, and the results show that the curve clearly reflects the peaks during the charging process. However, due to sampling noise and calculation step size, the curve exhibits significant high-frequency noise interference, making specific feature identification difficult. To accurately extract the key feature parameters of the IC curve, further processing of the results is required, namely, using digital filtering techniques to smooth the original curve to suppress noise interference and retain effective features.

[0132] This invention employs Kalman filtering technology to smooth the IC curve obtained from the aforementioned formula. It utilizes a recursive method to solve the state estimation problem of dynamic systems. The core of this method lies in combining the system's state equations and observational information, continuously updating the state estimate based on previous estimates and current observational data, thereby achieving optimal filtering. Specifically, the IC curve is treated as a dynamic process affected by noise, and the state estimate is recursively updated to gradually approach the true situation. The Kalman filtering process mainly includes two steps: prediction and update. In the prediction step, the state and its covariance at the next moment are predicted based on the current state and the state transition model. In the update step, the prediction result is corrected using newly acquired measurement data to obtain a more accurate state estimate.

[0133] The prediction phase includes state prediction and covariance prediction, with the corresponding formulas as follows:

[0134]

[0135]

[0136] In the formula: It uses the result predicted from the previous state; It is the optimal result of the previous state; F is the state transition matrix; and P k-1 They are and The corresponding covariance; Q′ is the covariance of the system process; T is the transpose.

[0137] The update phase includes Kalman gain calculation, state update, and covariance update, with the corresponding formulas as follows:

[0138]

[0139] In the formula: K k H is the Kalman gain; H is the observation matrix; R is the measurement noise covariance. It is the optimal estimate at time k; z k is the measurement value at time k; E is the identity matrix.

[0140] Figure 5These are the IC curves of Cell1 from the Oxford dataset after filtering and smoothing across all cycles. Taking the 1000th and 7900th cycles as examples, it can be observed that the peak value of the curve at the 1000th cycle is significantly higher than that at the 7900th cycle. This indicates that the intensity of the electrochemical reaction inside the battery gradually weakens with increasing cycle count. Furthermore, the peak voltage at the 1000th cycle is higher than that at the 7900th cycle, indicating that the most active electrochemical reaction point shifts as the battery ages. In terms of IC curve shape, the curve at the 1000th cycle is sharper and steeper, indicating that the electrochemical reaction within this voltage range is more concentrated and rapid in a newer battery state; while the curve at the 7900th cycle is relatively wide and flat, reflecting a decrease in the uniformity of the electrochemical reaction and a slowdown in the reaction rate after aging.

[0141] When the battery voltage curve exhibits a stable voltage plateau, within the time range during which this plateau is maintained, the amount of charge deposited by the battery increases, while the voltage change is close to zero. This electrochemical process is represented by a distinct peak on the IC curve. This invention extracts this peak from the IC curve and uses it as a health characteristic characterizing the battery state, labeled F7. Figure 6 As shown, the peak value of the IC curve decreases significantly with the increase of the number of battery charge-discharge cycles. This phenomenon initially indicates that the extracted features can effectively reflect the aging condition of the battery, providing an important quantitative indicator for assessing the battery's health status.

[0142] Health Feature Extraction Based on Equivalent Circuit Model

[0143] In the study of nonlinear battery systems, in addition to mining the battery aging information contained in the voltage curve, it is also necessary to extract health features closely related to the battery aging mechanism. This process is of great significance for improving the accuracy of battery health state estimation. By identifying parameters through equivalent circuit models to extract health features, the model-driven method and the data-driven method are organically combined. Among many equivalent circuit models, the first-order RC model stands out due to its significant advantages in generalization ability, real-time performance and accuracy, making it the ideal choice for this invention to extract parameters highly related to the aging mechanism. By extracting these key parameters through the first-order RC equivalent circuit model, the intrinsic mechanism of battery SOH decline can be effectively revealed. The electrical behavior characteristics of this model can be represented by Equation (16).

[0144]

[0145] In the formula: U t U is the terminal voltage, U1 is the polarization voltage, U oc R0 is the open-circuit voltage, R1 is the internal resistance in ohms, C1 is the polarization internal resistance, and I is the polarization capacitance. tThis represents the load current.

[0146] By using discrete polarization voltage, equation (17) is derived:

[0147]

[0148] Among them, U 1,k+1 U is the polarization voltage at time k+1. 1,k Let I be the polarization voltage at time k. t,k The current at time k;

[0149] Assuming the sampling time interval Δt is small, then U can be approximated within one sampling time. oc,k+1 ≈U oc,k According to equations (16) and (17), we can obtain:

[0150]

[0151] Among them, U oc,k U is the open-circuit voltage at time k. t,k+1 Let U be the terminal voltage at time k+1. t,k Let I be the terminal voltage at time k. t,k+1 Represents the current at time k+1; R p For load resistance;

[0152] Simplify the above equation to equation (19):

[0153] U t,k+1 =a1U t,k +a2+a3I t,k+1 +a4I t,k (19)

[0154] in:

[0155]

[0156] Next, parameters a1 to a4 are identified to obtain the value of R0. Recursive Least Squares (RLS) is an efficient online parameter estimation method that dynamically updates model parameters through recursion, thus adapting to the time-varying characteristics of the system in real time. RLS does not require storing all historical data; instead, it iterative optimization is performed based on the current observation data and the parameter estimates from the previous time step, approximating the true values ​​by gradually correcting parameter deviations. This mechanism makes it particularly suitable for parameter identification in nonlinear time-varying systems such as lithium-ion batteries, continuously tracking the dynamic evolution of aging-sensitive parameters such as ohmic resistance. Compared to traditional least squares methods, RLS reduces computational complexity and storage requirements by not needing to store all historical data, significantly improving real-time performance and efficiency. The objective function of RLS is to minimize the sum of squared prediction errors, and its mathematical expression is:

[0157]

[0158] Where: λ is the forgetting factor (0 < λ ≤ 1), used to adjust the weight of historical data, here λ = 0.99; θ is the parameter vector; J(θ) represents the objective function; U(i) is the true voltage value at time i. This represents the voltage estimate at time i when the selected parameter is θ.

[0159] The recursive formula consists of three parts: gain matrix update, covariance matrix update, and parameter estimation update. For the k-th time step, the specific formula is derived as follows:

[0160] (1) Gain matrix update:

[0161]

[0162] (2) Covariance matrix update:

[0163]

[0164] (3) Parameter estimation update:

[0165] θ(k)=θ(k-1)+K(k)(U(k)-φ T (k)θ(k-1)) (24)

[0166] Where: K(k) is the Kalman gain at time k, P(k) is the covariance matrix at time k, θ(k) is the parameter vector at time k, and φ(k) is the regression vector at time k.

[0167] After parameter identification of the equivalent circuit model using RLS, the ohmic internal resistance R0 was successfully obtained. To construct a comprehensive battery state health assessment system, R0 was considered a key health characteristic and labeled F8. Figure 7 The correlation between resistance and cycle number was shown. Analysis reveals that R0 continuously increases with the number of charge-discharge cycles. This phenomenon indicates that the internal ohmic resistance characteristics of the battery change significantly with increasing cycle number during operation. This finding not only enriches our understanding of the internal physical processes of batteries but also provides important reference for a deeper understanding of battery aging processes and performance degradation mechanisms.

[0168] Feature correlation analysis

[0169] In many fields, especially in complex systems like battery research, feature correlation analysis plays a crucial role. Feature correlation analysis assesses the strength of the association between features and target variables using quantitative methods. After constructing a multi-source health feature set, it is necessary to further quantify the correlation between each feature and the battery's health state to provide a basis for model input selection. This section will use Kendall's correlation coefficient and grey relational analysis for a two-dimensional quantitative evaluation. Kendall's correlation coefficient captures the global monotonic relationship between features and target variables, while grey relational analysis reveals the similarity of local trends. The two complement each other, providing a more reliable and comprehensive basis for feature selection.

[0170] The Kendall Rank Correlation Coefficient (KRC) is a nonparametric statistical method primarily used to measure the degree of ordinal association between two variables. The core idea of ​​this method is to quantify the correlation by comparing the number of consistent and inconsistent pairs in a data pair. Specifically, for two sets of observations, if a pair of data points are ordained in the same direction in both variables (i.e., a larger value in one variable corresponds to a larger value in the other, or a smaller value corresponds to a smaller value), then the pair is called a consistent pair; conversely, if the ordinal directions are opposite, it is called an inconsistent pair. This invention uses the Kendall Rank Correlation Coefficient to quantitatively assess the association between multi-source health characteristics and SOH (Social Health Obstacles). The specific process is as follows:

[0171] (1) The extracted voltage statistical features (F1-F6), IC curve peak features (F7), and ohmic internal resistance features (F8) are aligned with the SOH sequence according to the number of battery cycles. The feature data and SOH values ​​are then rank-transformed to assign each data point its sorting position in the sequence. For parallel data, the average rank is used to eliminate bias.

[0172] (2) For any two pairs of observations (x) i ,yi ) and (x j ,y j If x i >x j time y i >y j , or x i <x j time y i <y j If the features are in the same direction as the SoH rank, they are called a consistent pair; otherwise, they are called an inconsistent pair. We iterate through all sample pairs and count the difference C between consistent pairs (features whose rank direction is the same as the SoH rank) and inconsistent pairs (features whose rank direction is opposite). The formula is as follows:

[0173]

[0174] Where *sign* is the sign function, contributing +1 to a consistent pair and -1 to an inconsistent pair. *m* represents the total number of observations of feature *x* at different times; *x*... i and x j Corresponding to two different observed features of feature x at time i and time j; y i and y j These correspond to two observed features of the actual SOH at time i and time j.

[0175] (3) Calculate the correlation coefficient τ using the following formula:

[0176]

[0177] The value of τ ranges from [-1, 1], where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no correlation. Generally, the larger |τ| is, the stronger the correlation. When |τ| > 0.7, the variables can be considered to have a strong correlation. Figure 8 This is a heatmap showing the Kendall correlation coefficients between the characteristics of each battery in the Oxford dataset and the State of Health (SOH). The graph shows that the correlation coefficients for the IC curve peak characteristics (F7) are mostly between 0.8 and 0.9, indicating a strong positive correlation with SOH; the correlation coefficients for the internal resistance R characteristics (F8) are mostly around -0.9, indicating a strong negative correlation with SOH; and the absolute values ​​of the correlation coefficients for each voltage statistical characteristic (F1-F6) also exceed 0.7, showing a strong correlation with the SOH of each battery.

[0178] Grey relational analysis

[0179] Grey Relational Analysis (GRA) is a quantitative analysis method for the degree of correlation between factors in an incomplete information system (grey system). Its key principle is to determine the degree of correlation between factors by calculating the geometric similarity of the curves of a reference sequence (e.g., State of Health, SOH) and comparison sequences (e.g., various feature sequences). Given that the reference sequence is the battery health state (SOH) and the comparison sequences are the various feature data, the following specific steps are:

[0180] (1) The original data is normalized to eliminate differences in dimensions and orders of magnitude between different variables, unify the data scale, and make the data comparable. This invention uses the mean method to perform dimensionless data processing, and the calculation formula is as follows:

[0181]

[0182] Where, x i (k) represents the i-th comparison sequence of the k-th sample, x′ i (k) represents the i-th comparison sequence of the k-th sample after normalization, where k is the index of the data point; here k and i can be understood as time k and time i respectively;

[0183] (2) Calculate the grey relational coefficient Y i (k), the correlation coefficient, reflects the local similarity between two sequences at a certain time, and is calculated using the following formula:

[0184]

[0185] In the formula: Δ i (k)=|x0(k)-x i (k)|(k=1,2,...,n) represents the absolute difference between the two sequences at time k, where x0(k) is the reference sequence (target sequence) being compared; Δ min Δ max These are the minimum and maximum absolute differences at all times, respectively; ρ is the resolution coefficient, which ranges from 0 to 1, and is usually taken as 0.5.

[0186] (3) Calculate the correlation degree by taking the mean of the above correlation coefficients, which is used to quantify the overall correlation strength Y. i The calculation formula is as follows:

[0187]

[0188] Overall correlation strength Y iThe closer the correlation is to 1, the stronger the association between the comparison sequence and the reference sequence. In most studies, a correlation greater than 0.8 is considered a strong correlation, indicating a close association between the sequences; 0.6-0.8 is a moderate correlation; and below 0.6 indicates a weak correlation. By calculating the grey relational degree, the correlation between the eight extracted features and the SOH of lithium-ion batteries was quantitatively analyzed. Combined with... Figure 9 As shown, the grey relational degree of all features is above 0.6, which preliminarily proves that they are all related to SOH.

[0189] Building an ensemble learning model

[0190] Choosing the right ensemble learning architecture is crucial for building a high-performance Somnolé (SOH) prediction model. Based on different ensemble strategies, current mainstream methods can be categorized into three types: Bagging, Boosting, and Stacking. Analysis shows that Stacking combines the advantages of both Bagging and Boosting; therefore, this invention will use the Stacking ensemble strategy to build the SOH prediction framework. The specific steps based on the Stacking ensemble architecture model are as follows:

[0191] (1) Dataset Partitioning: First, the training dataset is divided into multiple subsets using different methods. Here, a cross-validation strategy is used to partition the training set, typically K-fold cross-validation. The original data is divided into K mutually exclusive subsets, and each time one subset is selected as the validation set, while the remaining K-1 subsets are used as the training set. This mechanism ensures that each sample participates in only K-1 training iterations during the base model training phase, ultimately generating unbiased out-of-fold predictions.

[0192] (2) Base Model Training and Prediction: Multiple base models are trained in parallel within the cross-validation framework. For each base model, K-1 training data is used in each round of cross-validation. The trained base model makes predictions on the validation set, and the prediction results from all rounds are merged to obtain the prediction result of the base model on the entire training set. At the same time, the base model is used to make predictions on the test set to obtain the corresponding prediction results.

[0193] (3) Feature Combination: The prediction results of each base model are used as new features to construct a meta-model training dataset. This feature space integrates the differential capture capabilities of multiple base models for data patterns, forming a non-linear enhanced expression of the original feature space. The same operation is performed on the test set, using the prediction results of the base models on the test set as new features, and combining them with the original test set to form a new test dataset.

[0194] (4) Meta-model training: Use new training datasets to train meta-models (such as logistic regression, gradient boosting trees, etc.) so that they learn how to combine the prediction results of the base model to achieve complementary integration of model advantages and maximize the accuracy of the overall model.

[0195] (5) Prediction: The trained meta-model is used to predict the test data. The meta-model combines and weights the prediction results of the base model to generate the final prediction result.

[0196] Figure 10 This is a diagram of the Stacking algorithm structure, which uses cross-validation for data training and testing. Through Stacking, the advantages and characteristics of different base models can be fully utilized, thereby improving the overall model's performance and generalization ability.

[0197] Ensemble learning internal model selection

[0198] One of the key aspects of ensemble learning lies in the selection of the base model. Based on the type of base model, ensemble methods can be divided into two categories: one is homogeneous ensembles where all base models belong to the same type, such as using decision trees or neural networks as the base model; the other is heterogeneous ensembles where the base models include multiple types. Heterogeneous ensembles have the following advantages: First, heterogeneous models process data differently, and this diversity allows them to mine data patterns from multiple perspectives, thus covering a more comprehensive feature space. Second, homogeneous models (such as multiple decision trees) may produce highly correlated predictions due to their similar algorithmic principles, thus limiting the improvement of ensemble performance. Heterogeneous models, through the differences between their algorithms, naturally reduce covariance, thereby more effectively dispersing errors. Furthermore, a single model may be more sensitive to specific data distributions (such as outliers or noise), but heterogeneous models, with the adaptability of different algorithms, can effectively reduce dependence on a single model. Based on the aim of combining the advantages of different models and balancing their disadvantages, this study selects three different types of machine learning models as the base models for ensemble learning, as follows:

[0199] (1) Random Forest (RF) constructs multiple decision trees by combining bootstrapping and random subspace methods. In regression problems, its specific process is as follows: Figure 11As shown, RF randomly selects m features (m < M) each time it splits from a set of M total features, and generates a single decision tree by optimizing the information gain. Finally, the algorithm takes the average of multiple regression results as the output prediction. RF effectively reduces the variance by leveraging the double randomness of features and samples, thereby enhancing the robustness of the model in a noisy environment and alleviating the overfitting problem. However, due to the inherent characteristics of decision trees, their decision boundaries may exhibit stepwise discontinuities. In addition, when dealing with ultra-high dimensional data, random forests may be affected by computational resource limitations.

[0200] (2) Kernel Ridge Regression (KRR) is an advanced regression technique that combines ridge regression and the kernel trick. Its core purpose is to address non-linear problems and effectively avoid the risk of overfitting by implementing a data regularization strategy. KRR combines the advantages of ridge regression and the characteristics of kernel methods. When actually applying KRR, first, with the help of the kernel function, the original feature space is mapped to a high-dimensional space, and then ridge regression analysis is performed in this high-dimensional space. The loss function Kernel Ridge Loss used by KRR is expressed as:

[0201]

[0202] where: y is the target variable; X is the feature matrix; w is the regression coefficient; α is the regularization parameter; K(X, X) is the kernel matrix, and each of its elements K i,j is the inner product of X i and X j in the high-dimensional space. By incorporating the kernel function, KRR endows the model with the ability to capture the non-linear characteristics of the data. It also uses regularization to constrain the complexity of the model, thereby constructing a model with good generalization performance.

[0203] (3) Explainable Boosting Machine (EBM) is a new type of explainable algorithm proposed by Nori et al. in 2019. EBM is an improved version of the generalized additive model. This algorithm uses modern machine learning techniques, such as ensemble learning and gradient boosting, to learn each feature function f i . In addition, EBM also considers the interaction between two features

[0204] ∑f ij (x i )]], x j ), thereby further improving the accuracy of the model. Its mathematical expression is:

[0205]

[0206] In the formula: g(·) is the connection function; m is the number of features; x i It is a characteristic; f i y is the smoothing function of the feature; E(y) is the expected value; β is the intercept; ε is the residual.

[0207] The aforementioned foundational models address the problem based on different theoretical principles. Given their respective significant advantages, choosing these models as the foundation for ensemble learning is both reasonable and meaningful. By leveraging their characteristics, ensemble models can improve prediction accuracy and generalization ability, thereby enhancing the effectiveness of the ensemble learning framework.

[0208] After determining the base model, the next step is to select the Logistic Regression (LR) model as the second-layer model in the ensemble learning prediction framework proposed in this study. LR is a widely used linear model in machine learning. Compared to complex nonlinear models, the linearity of LR reduces the risk of overfitting in the meta-model. If the base model in the first layer includes a nonlinear model (such as decision trees or KRR), then the LR in the second layer can form a "nonlinear + linear" combination. This combination can cover a wider function space while preventing unstable prediction results due to excessive model complexity.

[0209] Model hyperparameter optimization

[0210] The selection of hyperparameters plays a crucial role in model performance. Different combinations of hyperparameters can lead to significant differences in the model's fitting ability and generalization ability. Inappropriate hyperparameters may cause overfitting or underfitting, thus affecting the accuracy of predicting targets such as battery state of harmonics (SOH). Therefore, employing efficient hyperparameter optimization algorithms is a necessary step to improve the performance of ensemble models.

[0211] In constructing a heterogeneous ensemble model composed of KRR, RF, and EBM, this study employs a Tree-structured Parzen Estimator (TPE) for hyperparameter optimization. This method achieves an optimal balance between the diversity and synergy of the base models by dynamically adjusting the sampling probability distribution. The TPE algorithm aims to solve the global optimization problem of black-box functions. This algorithm can adaptively adjust the parameter search space range according to the actual situation and locate the global optimum with fewer iterations. Compared with traditional grid search and random search methods, TPE effectively reduces the computational resource requirements through an intelligent sampling mechanism. Its core concept is to construct a probabilistic model based on historical evaluation data and dynamically adjust the direction of parameter search accordingly. Specifically, the main process of TPE includes the following aspects: initial random sampling, classification of observation results, construction of a probabilistic model, calculation of the Expected Improvement (EI) value, selection of new parameter combinations, performance evaluation, and updating the dataset. This method not only retains the ability of global search but also significantly improves optimization efficiency through local refined modeling, making it particularly suitable for discrete-continuous hybrid parameter spaces. Figure 12 The diagram shown illustrates the principle of the TPE optimization algorithm. The following are the specific parameter optimization steps:

[0212] (1) Initial random sampling: After the process starts, random sampling is performed on the hyperparameter space to generate multiple sets of hyperparameter combinations and evaluate their performance, accumulating initial data to lay the foundation for subsequent analysis.

[0213] (2) Determine the initial sampling status: Check whether the preset number of initial samplings has been completed. If not, continue random sampling; if completed, proceed to the data partitioning stage.

[0214] (3) Divide into high-quality and ordinary parameter groups: Based on the evaluation results, divide the top γ% (e.g., 20%) of hyperparameter combinations into the "Top Group" and the rest into the "Remaining Group" to distinguish the distribution characteristics of high-quality and ordinary parameters.

[0215] (4) Constructing a probability model: For the “Top group”, construct a probability model l(θ) to characterize the distribution of high-quality hyperparameters; for the “remaining group”, construct a probability model g(θ) to describe the distribution of ordinary hyperparameters.

[0216] (5) Calculate the EI value: Use formula (32) to calculate the EI value γ(θ) of each candidate hyperparameter, measure its probability advantage of belonging to the high-quality group, and guide the search direction.

[0217]

[0218] (6) Select the optimal candidate parameters: Select the hyperparameter combination with the largest EI value and explore the parameters most likely to optimize the model performance.

[0219] (7) Evaluate the performance of new parameters: Apply the selected hyperparameters to the target model and evaluate the performance (such as MAE, RMSE) through cross-validation and other methods to obtain new evaluation results.

[0220] (8) Determine the termination condition: Check whether the termination condition is met (such as reaching the maximum number of iterations or the performance improvement slowing down). If not, integrate the new data into the historical data, re-divide the groups, build the model, and iterate in a loop; if it is met, output the current optimal hyperparameter combination and the process ends.

[0221] Based on Bayesian optimization theory, the number of iterations was set to 25, and differentiated parameter search spaces were designed to account for the differences in parameter characteristics of each model, as shown in Table 1. Specifically:

[0222] (1) For KRR, the kernel function type is selected from the discrete space composed of linear kernel ('linear'), radial basis kernel ('rbf') and polynomial kernel ('poly'); the RBF kernel width parameter γ is sampled in the interval [1e-4, 1e1] using a log-uniform distribution to cover a wide range of nonlinear modes; the L2 regularization coefficient α is optimized by log_uniform(1e-5, 1e2) to balance the model complexity and the risk of overfitting.

[0223] (2) The optimization of RF focuses on tree complexity-related parameters. The number of trees n_estimators is uniformly sampled in integers with a step size of 10 in the range of 50-500 to ensure a balance between model capacity and computational efficiency; the tree depth max_depth is constrained by quniform(5,50,1) to avoid overfitting caused by excessively deep tree structures; the minimum number of samples for node splitting min_samples_split is searched in the range of 2-20 to enhance the generalization ability of the decision boundary.

[0224] (3) For the parameter search of EBM, the learning rate is optimized using the log_uniform(0.01,0.3) distribution, and the gradient boosting step size is finely adjusted; the interaction order of the feature interaction terms is controlled by quniform(0,5,1) to control the model complexity; the number of feature bins max_bins is searched in the range of 128-512 with a step size of 32 to balance the feature discretization accuracy and computational consumption.

[0225] Table 1. Search range of model hyperparameters

[0226]

[0227] Analysis of estimation results of ensemble learning models

[0228] This invention employs a stacking ensemble learning framework based on cross-validation to construct a dual-driven SOH estimation model. The final model structure diagram is shown below. Figure 13 As shown, three algorithms with different modeling characteristics—RF, KRR, and EBM—were selected as the base models. LR was introduced as the second-layer model. A stacking strategy with 10-fold cross-validation was used to integrate the prediction results of each base model on the validation set as second-layer features, effectively combining the advantages of different models in feature representation and generalization ability. Furthermore, the TPE algorithm was used to optimize the hyperparameter space of the ensemble learning model, improving the overall model performance through adaptive parameter tuning. This framework, by combining multi-source health features extracted based on equivalent circuit models with machine learning models, achieves an organic combination of physical models and data-driven learning, ultimately forming a SOH estimation model with dual driving characteristics.

[0229] Comparison with estimation results of the basic model

[0230] To comprehensively verify the predictive performance of the proposed Stacking ensemble learning model, experimental research was conducted based on the Oxford Battery Dataset. The data partitioning strategy was as follows: battery data from Cell1 to Cell4 were selected as the training set to support model training and parameter learning; battery data from Cell5 to Cell8 were used as the test set to evaluate the model's predictive ability for unknown samples. To verify the advantages of the ensemble learning model compared to individual basic models, the TPE optimization algorithm was used to optimize both the constructed ensemble model and the three basic models, and then applied to SOH prediction to compare the differences in prediction results between the ensemble learning model and each basic model.

[0231] Table 2 shows a comparison of the prediction errors of the ensemble learning model and the three base models on four battery samples (Cell5-Cell8). The analysis reveals that the ensemble learning model has lower RMSE and MAE on Cell5, Cell6, and Cell7 compared to other base models. Only Cell8 shows a low prediction error for the KRR model, which is attributed to the randomness of the battery. Furthermore, from an overall average perspective, the ensemble learning model exhibits the best comprehensive performance across all tested batteries, with an average RMSE of 0.331% and an average MAE of 0.234%, representing reductions of 28.4% and 40.6% respectively compared to the second-best base model, RF; 34.1% and 39.1% compared to the base model KRR; and 59.0% and 67.9% compared to the base model EBM. The highest reduction was observed in EBM, indicating that the ensemble learning model has a good ability to correct high-error base models. The above analysis validates the effectiveness of the ensemble method in integrating the advantages of heterogeneous models.

[0232] Table 2 Prediction errors of ensemble learning model and base model

[0233]

[0234] Comparison with estimation results from neural network models

[0235] The ensemble learning framework proposed in this invention uses RF, KRR, and EBM as its basic models, mainly considering the advantages of classic machine learning models in terms of computational efficiency, interpretability, and stability in small sample scenarios. To further verify the superiority of this framework, this section introduces neural network models to construct a comparative experimental system. Targeting the temporal characteristics of battery aging data, the following three types of neural network models are designed: Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and Long Short-Term Memory (LSTM).

[0236] To ensure the scientific rigor and reproducibility of the experimental comparisons, this study adopted the same TPE hyperparameter optimization process mentioned in Section 4.2 for the three neural network models and the ensemble model, and completed model training and validation on the Oxford dataset (Cell1-Cell4 as the training set and Cell5-Cell8 as the test set).

[0237] Figure 14The graphs show a comparison of the SOH prediction results of the ensemble learning model and various neural network models. As can be seen from the four graphs, the prediction curve of the ensemble learning model consistently and closely follows the trend of the true SOH, with its predicted values ​​deviating less from the true SOH than those of the neural network models. For example, in the 2000-4000 iteration range, LSTM exhibits significant fluctuations, while CNN and RNN also show considerable deviations. In the initial iterations (0-1000 iterations), the ensemble learning model quickly follows the changes in the true SOH, while LSTM lags behind early on. As the number of iterations increases, the prediction bias of the neural network models gradually accumulates. For instance, in Cell7 and Cell8, the deviation of CNN and LSTM intensifies after 3000 iterations, but the ensemble learning model consistently maintains a lower error.

[0238] Table 3 presents the prediction error and training time data. Experimental results show that the proposed ensemble learning model achieves higher prediction accuracy than neural network models, with lower RMSE and MAE compared to individual neural network models. It also demonstrates a significant advantage in training efficiency: the ensemble learning model reduces training time by approximately 73% compared to CNN, 67% compared to LSTM, and 72% compared to RNN. Further analysis indicates that the superiority of the proposed ensemble model lies in the fact that its basic models (such as KRR and RF) are simpler and have significantly fewer parameters than neural network models. For example, LSTM requires iterative optimization of gate unit parameters, while RF supports multi-threaded parallel construction of decision trees, thereby improving training efficiency. It is worth emphasizing that this ensemble model has a relatively lightweight architecture, making its deployment in embedded devices or edge computing scenarios more convenient and suitable for BMS with high real-time requirements. In contrast, neural network models such as LSTM rely on dedicated hardware and frameworks, which not only increases deployment costs in industrial environments but may also put pressure on the computational needs of the BMS.

[0239] Table 3. Errors and time between ensemble learning models and neural network models

[0240]

[0241] Health characteristic effectiveness assessment

[0242] Based on their different types, the eight extracted features were divided into four groups for comparative experiments. The specific groupings are as follows: Group 1 used all features (denoted as S+IC+R); Group 2 used voltage statistical features and IC curve peak features (denoted as S+IC); Group 3 used voltage statistical features and ohmic internal resistance features (denoted as S+R); and Group 4 used only voltage statistical features (denoted as S). Table 4 shows the specific estimation error data. For the four battery samples, when only voltage statistical features were used, the estimation error was the largest for all batteries, with average RMSE and average MAE of 0.986% and 0.909%, respectively. After combining voltage statistical features with IC curve peak features and ohmic internal resistance features, the RMSE and MAE were lower than other feature combinations, with average RMSE and average MAE of 0.331% and 0.234%, respectively. This fully demonstrates the advantage of multi-source feature fusion in improving the accuracy of SOH estimation.

[0243] Table 4. Estimation error under different feature combinations

[0244]

[0245] Model interpretability analysis

[0246] The interpretability of the proposed lithium-ion battery health state estimation model was analyzed using the SHAP method. The contribution of each feature parameter was systematically evaluated, and how changes in each feature affect the model output was investigated. The reasons were also explored from the perspective of battery aging mechanism. Figure 15 This is a SHAP result diagram using the Cell1 battery as an example. Figure 15 Figure (a) shows the importance ranking of the features, which is determined by calculating the average absolute value of the SHAP value of each feature. As can be seen from the figure, among the eight battery health features used in this invention, the IC curve peak feature (F7) and the ohmic internal resistance feature (F8) rank first and second in importance, indicating that they are key factors affecting the model's prediction results. Figure 15(b) in the diagram is the SHAP beehive plot, which further illustrates the contribution distribution of all features. In this plot, the X-axis represents the SHAP value, i.e., the degree of contribution of the feature to the prediction result (positive values ​​indicate improved prediction results, negative values ​​indicate decreased prediction results); the Y-axis is sorted according to the importance of the features. Each point represents a sample, with the number of samples stacked vertically, and the color of the point reflects the actual value of the feature in that sample (red corresponds to high values, blue corresponds to low values). The density of the points reflects the distribution of SHAP values; the larger the horizontal expansion range, the more significant and diverse the influence of the feature on the prediction. For the IC curve peak feature (F7), a high IC peak (red) has a positive impact on the prediction, i.e., the higher the IC peak, the larger the SOH value. Similarly, a low ohmic internal resistance value (blue) has a negative impact on the prediction, i.e., the smaller the internal resistance value, the larger the SOH value.

[0247] Further analysis revealed that among the eight extracted health characteristics, the IC curve peak characteristic (F7) and the ohmic internal resistance characteristic (F8) exhibited a significant dominant role. This phenomenon can be reasonably explained by the aging mechanism of lithium-ion batteries.

[0248] (1) The peak value in the IC curve, due to its unique shape, height, and position, can effectively reflect the electrochemical reaction characteristics during the charging and discharging process of lithium-ion batteries. Specifically, there is a close correlation between the peak value of the IC curve and the content of active material in the electrode material and the integrity of its crystal structure. During battery aging, the loss of active material, the thickening of the solid electrolyte interface film, and the phase transition of the electrode material all have a significant impact on the position and amplitude of the IC peak value.

[0249] (2) Ohmic internal resistance reflects the resistance to conduction between ions and electrons inside the battery, and its increase is a significant characteristic of battery aging. As the number of charge-discharge cycles increases, electrode material particles gradually break down, interfacial contact deteriorates, and the thickness of the SEI film on the negative electrode surface also increases. These factors collectively contribute to the continuous increase in ohmic internal resistance. From the perspective of the chemical reaction mechanism of lithium-ion battery charging and discharging, the composition of ohmic internal resistance mainly includes the impedance of the positive and negative electrode active materials, the resistance of the electrolyte, the resistance of the solid interface film, and the resistance of other battery components. During charging and discharging, particles may detach from the SEI film on the surface of the negative electrode active material. Once these detached particles enter the electrolyte, they will trigger electrophoresis, thus becoming an important factor in the increase of internal resistance. During continuous charge-discharge cycles, the internal resistance of lithium-ion batteries gradually increases, leading to increased battery heat generation, which in turn gradually reduces the battery's health and ultimately accelerates the aging process.

[0250] The remaining features (such as F1-F6) have a relatively balanced or weak impact on the model output. In the battery aging mechanism, the physical meaning reflected by these features is less sensitive and specific to aging characterization than the IC curve peak feature and ohmic resistance feature, thus their weight in model decision-making is relatively low. Furthermore, SHAP results show that features F7 and F8, through high internal weight allocation and nonlinear interaction, significantly improve the model's ability to capture the relationship between features and state of health (SOH), ultimately optimizing the estimation accuracy. In summary, the IC curve peak feature (F7) and ohmic resistance feature (F8) are core features characterizing the battery aging mechanism. They dominate the model's prediction logic for SOH from the perspective of the battery's internal reaction mechanism, decisively improving the performance of the ensemble learning model. This fully validates their effectiveness and importance in battery health state estimation. Voltage statistical characteristics play a fundamental role in the overall model. By calculating the mean, standard deviation, variance, root mean square, skewness, and slope, they comprehensively characterize the distribution and fluctuation of voltage, reflecting the basic state of the battery during operation and providing basic data support for health status estimation. Models built upon these characteristics can initially assess battery health. The newly added IC curve peak characteristics and ohmic internal resistance characteristics further supplement the battery's internal characteristics from a physical mechanism perspective. The combination of these two features enables a more accurate assessment of the battery's state of health (SOH).

[0251] Robustness analysis under noise interference

[0252] The original data was standardized to the [0,1] interval to eliminate dimensional differences. Then, for different levels of noise environments, uniformly distributed random perturbations were superimposed on the feature dimension. In the case of normal interference scenarios, a 20% amplitude control was used, that is, generating independent noise components that follow a uniform distribution in the [-0.2,0.2] interval. For extreme interference scenarios, the noise amplitude was increased to 40%, and the corresponding interval was expanded to [-0.4,0.4]. The perturbation intensity was controlled by adjusting the boundary parameters of the random number generator.

[0253] Experimental results show that, compared with the prediction results in Section 4.3 without random noise interference, the SOH prediction accuracy of all models decreases to varying degrees under noise interference. However, regardless of whether the noise level is 20% or 40% perturbation, the proposed ensemble learning model consistently maintains the best performance compared to the base models, with the lowest RMSE and MAE.

[0254] Cross-dataset generalization analysis

[0255] Using the Huazhong University of Science and Technology dataset (HZ dataset) as the validation object, the experimental design followed a standardized process: Of the eight batteries in the HZ dataset, four batteries (A1-A4) were selected as the training set, and the other four batteries (A5-A8) as the test set. This dataset will use the same feature extraction method to obtain the same health features as the Oxford dataset, specifically obtaining health state parameters through three approaches: statistical features based on voltage curves, peak features based on incremental capacity curves, and ohmic internal resistance features obtained through an equivalent circuit model. The estimation results of this invention will be compared with those of neural network models CNN, LSTM, and RNN.

[0256] The results show that the estimation results of the ensemble learning model consistently closely match the actual curve, especially in the later stages of the cycle, such as after more than 1000 cycles in A6 and A7, and after more than 1250 cycles in A8. It still accurately captures the decreasing trend of SOH with minimal fluctuations, demonstrating a strong ability to fit the entire battery lifecycle aging process. In contrast, the neural network model exhibits varying degrees of deviation across all battery samples. For example, in A5, the estimates of CNN and RNN deviate significantly from the actual values ​​in the middle of the cycle; in A7 and A8, the predictions of the neural network model deviate from the actual values ​​to varying degrees, with CNN and LSTM showing increased deviations in the later stages of the cycle, and RNN showing larger deviations in the early and middle stages. These phenomena indicate that the neural network model lacks the ability to track the complex trends of battery aging, exhibiting weak stability and generalization.

[0257] Table 5 details the specific error values ​​and time taken. The proposed estimation model maintains a small estimation error on the new dataset, with both RMSE and MAE not exceeding 2%. Although CNN exhibits the smallest error in SOH estimation for the A6 battery sample, the ensemble model's error is second only to it and the difference is small; in the estimation of the other three battery groups, the ensemble model's error is the lowest, while CNN's estimation error is the highest. Overall, the ensemble model demonstrates greater stability across different battery samples. It is worth emphasizing that compared to CNN, LSTM, and RNN, the proposed ensemble model not only has higher overall estimation accuracy but also has an advantage in processing time, reducing the total time by approximately 69%, 68%, and 60%, respectively. This further validates the stable tracking performance of the designed SOH estimation framework on different battery datasets, highlighting its efficiency in SOH estimation.

[0258] Table 5-4 Estimation error and time for the Hz dataset

[0259]

[0260] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.

Claims

1. A method for estimating the state of health of lithium-ion batteries based on a dual-drive interpretable ensemble model, characterized in that, Includes the following steps: S100. The average value, standard deviation, variance, root mean square, skewness, and kurtosis are obtained from the charging voltage curve. The peak characteristics of the IC are obtained based on the voltage curve, and the ohmic internal resistance characteristics are obtained based on the incremental capacity curve. The above characteristics are used to construct a multi-source health feature space. S200. Build a dual-drive ensemble learning model, integrating three basic models: random forest, kernel ridge regression, and interpretable augmentation machine. Select logistic regression as the second layer model of the dual-drive ensemble learning model. Use the multi-source health feature space obtained in step S100 as the feature input. Use a tree-based Bayesian optimization algorithm to achieve collaborative optimization of model hyperparameters. Use the optimized dual-drive ensemble learning model to estimate the health status of lithium-ion batteries.

2. The lithium-ion battery health state estimation method based on a dual-drive interpretable ensemble model according to claim 1, characterized in that: In step S100, the statistical characteristics of the mean, standard deviation, variance, root mean square, skewness, and kurtosis are standardized using the following formulas: In the formula: n is the number of sample points, V i V represents the specific voltage value corresponding to sample point i. mean V represents the sample mean. std V represents the sample standard deviation. var V represents the sample variance. rms V represents the root mean square of the sample. skew V represents sample skewness. kurt Represents the kurtosis of the sample.

3. The lithium-ion battery health state estimation method based on a dual-drive interpretable ensemble model according to claim 2, characterized in that: In step S100, when obtaining the IC peak characteristics, according to the incremental capacity analysis method, the IC curve is obtained by calculating the first derivative of the VQ curve of the lithium-ion battery during the charging process: Where I is the current, t is the charging time, Q is the capacity, and V is the voltage; Under constant current charging conditions, the battery terminal voltage is divided into several equally spaced voltage segments, including: recording the charging time Δt of each segment when the battery terminal voltage rises by ΔV; then, calculating the charged capacity ΔQ within the voltage segment using the product of Δt and the charging current; finally, approximating the differential capacity value dQ / dV at the corresponding voltage point by the ratio of ΔQ to ΔV; continuing this process until the battery reaches the set cutoff voltage, thus completing the entire charging process; when ΔV approaches zero, a series of continuous differential capacity data points can be obtained according to formula (8); after interpolating and fitting these discrete data points, a smooth incremental capacity curve is generated: Kalman filtering was used to smooth the IC curve obtained by the aforementioned formula, including prediction and updating. In the prediction step, the state and its covariance at the next time step were predicted based on the current state and the state transition model. In the formula: It uses the result predicted from the previous state; It is the optimal result of the previous state; F is the state transition matrix; and P k-1 They are and The corresponding covariance; Q′ is the covariance of the system process; In the update step, the prediction results are corrected using the newly acquired measurement data to obtain the state estimate. The update stage includes Kalman gain calculation, state update, and covariance update, with the corresponding formulas as follows: In the formula: K k H is the Kalman gain at time k; H is the observation matrix; R is the measurement noise covariance. It is the optimal estimate at time k; z k is the measurement value at time k; E is the identity matrix; When the battery voltage curve shows a stable voltage plateau, the amount of charge entering the battery increases within the time range during which the plateau is maintained, while the change in voltage is close to zero. This electrochemical process is represented by a significant peak on the IC curve, and the peak in the IC curve is used as a health characteristic to characterize the battery state.

4. The lithium-ion battery health state estimation method based on a dual-drive interpretable ensemble model according to claim 3, characterized in that: In step S100, when obtaining the ohmic internal resistance characteristics, the first-order RC equivalent circuit model is: In the formula: U t U is the terminal voltage, U1 is the polarization voltage, U oc R0 is the open-circuit voltage, R1 is the internal resistance in ohms, R1 is the polarization internal resistance, and C1 is the polarization capacitor; I t Represents the load current; By using discrete polarization voltage, we can derive: Among them, U 1,k+1 U is the polarization voltage at time k+1. 1,k Let I be the polarization voltage at time k. t,k This represents the load current at time k; Within a sampling time, U is approximately considered to be oc,k+1 ≈U oc,k According to equations (14) and (15), we get: Among them, U oc,k U is the open-circuit voltage at time k. t,k+1 Let U be the terminal voltage at time k+1. t,k Let be the terminal voltage at time k. I t,k+1 Represents the load current at time k+1; R p For load resistance; Simplifying the above equation, we get: U t,k+1 =a1U t,k +a2+a3I t,k+1 +a4I t,k (17) in: The parameters a1 to a4 are identified to obtain the value of R0; the sum of squared prediction errors is minimized using the recursive least squares method, the mathematical expression of which is: Where: λ is the forgetting factor, 0 < λ ≤ 1, used to adjust the weight of historical data; θ is the parameter vector; J(θ) represents the objective function; U(i) is the true voltage value at time i. This represents the voltage estimate at time i when the selected parameter is θ; The recurrence formula consists of three parts: gain matrix update, covariance matrix update, and parameter estimation update. For the k-th moment, the specific derived formula is as follows: (1) Gain matrix update: (2) Covariance matrix update: (3) Parameter estimation update: θ(k)=θ(k-1)+K(k)(U(k)-φ T (k)θ(k-1)) (22) where: K(k) is the Kalman gain at the updated k-th moment, P(k) is the covariance matrix at the updated k-th moment, θ(k) is the parameter vector at the k-th moment, and φ(k) is the regression vector at the k-th moment; After parameter identification of the equivalent circuit model using the recursive least squares method, the ohmic internal resistance R0 is obtained.

5. The lithium-ion battery health state estimation method based on a dual-drive interpretable ensemble model according to claim 4, characterized in that: In step S200, the steps of the Stacking integration architecture model are as follows. (1) Dataset division: First, the training dataset is divided into multiple subsets using the cross-validation strategy. K-fold cross-validation is selected; the multi-source health features obtained in step S100 are divided into K mutually exclusive subsets. Each time, one subset is selected as the validation set, and the remaining K - 1 subsets are used as the training set; (2) Basic model training and prediction: Under the cross-validation framework, multiple basic models are trained in parallel; for each basic model, in each round of cross-validation, it is trained using the K - 1-fold training data; the trained basic model predicts the validation set, and the prediction results of all rounds are combined to obtain the prediction result of the basic model on the entire training set; at the same time, the basic model is used to predict the test set to obtain the corresponding prediction result; (3) Feature combination: The prediction result of each basic model is used as a new feature to construct the meta-model training dataset; the same operation is also performed on the test set, and the prediction result of the basic model on the test set is used as a new feature and combined with the original test set to form a new test dataset; (4) Meta-model training: The new training dataset is used to train the meta-model so that it learns how to combine the prediction results of the basic models to achieve complementary fusion of model advantages and maximize the accuracy of the overall model. (5) Prediction: The trained meta-model is used to predict the test data. The meta-model combines and weights according to the prediction results of the basic models to generate the final prediction result.

6. The lithium-ion battery health state estimation method based on a dual-drive interpretable ensemble model according to claim 5, characterized in that: In step S200, the random forest constructs multiple decision trees by combining the bootstrap method and the random subspace method. From a set with a total of M features, m features are randomly selected each time for splitting, where m < M, and a single decision tree is generated by optimizing the information gain; finally, the average value of multiple regression results is taken as the output prediction; Kernel ridge regression借助核函数的力量,将原始特征空间映射至一个高维空间,随后在这个新构建的高维空间内执行岭回归分析;所采用的损失函数Kernel Ridge Loss表述为: In the formula: y is the target variable; X is the feature matrix; w is the regression coefficient; α is the regularization parameter; K(X,X) is the kernel matrix, and each element K i,j It is X i and X j Inner product in higher-dimensional space.

7. The lithium-ion battery health state estimation method based on a dual-drive interpretable ensemble model according to claim 6, characterized in that: In step S200, the interpretable augmentation machine learns each feature function f using modern machine learning techniques. i Consider the interaction between two features ∑f ij (x i ,x j This is used to improve the accuracy of the model. In the formula: g(·) is the connection function; m is the number of features; x i It is the observed feature of feature x at time i; f i E(y) is the smoothing function of the feature at time i; E(y) is the expected value; β is the intercept; and ε is the residual.

8. The lithium-ion battery health state estimation method based on a dual-drive interpretable ensemble model according to claim 7, characterized in that: In step S200, the tree-structured Parzen estimator is used for hyperparameter optimization, including random sampling in the initial stage, classifying the observed results, constructing a probability model, calculating the expected improvement value, selecting a new parameter combination, performing evaluation, and updating the dataset. Specifically, it includes It should be noted that there is an unclear part in the translation of item . It seems that some words are missing in the original Chinese description. You may need to check and correct it for a more accurate translation. (1) Initialize random sampling: Randomly sample the hyperparameter space, generate multiple sets of hyperparameter combinations and evaluate their performance, and accumulate initial data; (2) Determine the initial sampling status: Check whether the preset number of initial samplings has been completed. If not, continue random sampling. If completed, proceed to the data partitioning phase; (3) Divide into high-quality and ordinary parameter groups: Based on the evaluation results, the top γ% of hyperparameter combinations are divided into the Top group, and the rest are the remaining group, to distinguish the distribution characteristics of high-quality and ordinary parameters. (4) Constructing a probability model: For the Top group, construct a probability model l(θ) to characterize the distribution of high-quality hyperparameters; for the remaining groups, construct a probability model g(θ) to describe the distribution of ordinary hyperparameters. (5) Calculate the EI value: Calculate the EI value γ(θ) of each candidate hyperparameter using formula (25) to measure its probability advantage of belonging to the high-quality group and guide the search direction: (6) Select the optimal candidate parameters: Select the hyperparameter combination with the largest EI value and prioritize exploring the parameters most likely to optimize model performance; (7) Evaluate the performance of the new parameters: Apply the selected hyperparameters to the target model and evaluate the performance through methods such as cross-validation to obtain new evaluation results; (8) Determine the termination condition: Check whether the termination condition is met. If not, integrate the new data into the historical data, re-divide the groups, build the model, and iterate in a loop. If the condition is met, output the current optimal hyperparameter combination.

9. A lithium-ion battery health state estimation system based on a dual-drive interpretable ensemble model, characterized in that: The system has a program module corresponding to the steps described in any one of claims 1-8, and executes the steps in the above-described method for estimating the state of health of lithium-ion batteries based on a dual-drive interpretable ensemble model when it is run.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program configured to, when invoked by a processor, implement the steps of the lithium-ion battery health state estimation method based on the dual-drive interpretable ensemble model as described in any one of claims 1-8.

Citation Information

Cited By

  • Lithium ion battery health state evaluation method, system and equipment based on HPO-TabPFN interpretable model and storage medium

    CN121276379A

  • Zinc-bromine flow battery internal resistance calculation and state analysis method and system

    CN122063479A

  • A method and system for calculating internal resistance and analyzing state of zinc-bromine flow battery

    CN122063479B