Battery pack attenuation trajectory prediction method and system based on actual operation data

By constructing a multi-level feature set based on histogram data and a Transformer model, the nonlinear data processing problem in battery pack degradation trajectory prediction was solved, achieving accurate prediction of battery pack capacity degradation trajectory and improving prediction accuracy and model adaptability.

CN120993254APending Publication Date: 2025-11-21BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510836006.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing battery pack degradation trajectory prediction methods cannot effectively handle the nonlinear changes in random charging and discharging data of electric vehicles, and it is difficult to extract features that characterize the internal aging changes of the battery, resulting in low prediction accuracy and poor model robustness.

Method used

By converting actual operating data under aging factors such as current, voltage, SOC, and temperature into histogram data, a multi-level feature set is constructed. Combined with local weighted linear regression and the Transformer model, accurate prediction of capacity trajectory is achieved.

Benefits of technology

It improves the prediction accuracy and model robustness under complex discharge conditions, and can accurately predict the battery pack capacity decay trajectory, adapting to different vehicles and operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120993254A_ABST
    Figure CN120993254A_ABST
Patent Text Reader

Abstract

The invention discloses a battery pack attenuation trajectory prediction method and system based on actual operation data. The method comprises the following steps of: 1, preprocessing mileage, cycle index, SOC (State of Charge), voltage and current of a battery pack and voltage and temperature data of a single battery recorded by a battery management system in an actual operation condition; 2, current, voltage, SOC and temperature are selected as key factors influencing battery aging, and duration, ampere-hour throughput and watt-hour throughput under the condition of various types of factor partition intervals are calculated; step 3, constructing four levels of aging feature sets, namely a vehicle level, a battery pack level, a monomer level and an inconsistency level, by adopting the normalized charge and discharge use intensity histogram sequence data; step 4, carrying out smoothing processing on the capacity track of the actual vehicle in the service period by adopting local weighted linear regression; 5, performing feature screening on the aging feature libraries obtained in the step 3 and the step 4 by adopting a three-step feature screening method; and step 6, taking the feature space obtained by screening in the step 5 as the input of a Transform model, and predicting the capacity attenuation trajectory of the battery pack.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of lithium-ion battery health status prediction, specifically a method and system for predicting battery pack degradation trajectory based on actual operating data. Background Technology

[0002] With the rapid development of my country's electric vehicle market, lithium-ion batteries are widely used due to their excellent energy density and long service life. As the core power source of electric vehicles, the accurate assessment of the health status of power lithium batteries is a key technological bottleneck for reliable vehicle operation, efficient control, and value evaluation. However, the aging mechanism of electric vehicle power batteries is complex, and performance degradation is non-linear. Coupled with the randomness of actual operating conditions, this results in low accuracy and poor predictive ability in battery pack health status estimation. Compared to health status estimation, degradation trajectory prediction can obtain detailed aging paths, thereby providing comprehensive information support for battery management systems and enabling the formulation of precise maintenance strategies.

[0003] Existing battery pack degradation trajectory prediction methods cannot fully handle the nonlinear changes in random charging and discharging data of electric vehicles, and it is difficult to effectively extract features that characterize the aging changes inside the battery. This limits the predictive ability of the prediction model. There is an urgent need to develop feature construction methods that can regularize irregular data and accurately extract aging information, thereby improving the prediction accuracy and robustness of the model under complex discharge conditions. Summary of the Invention

[0004] To address the problems and deficiencies described in the prior art, this invention provides a battery pack degradation trajectory prediction method and system based on actual operating data. It utilizes a regularization method to transform complex time-series data into histogram data by considering aging factors such as current, voltage, SOC, and temperature, thereby reducing data fluctuations and retaining key information. A multi-level feature construction method based on charge-discharge histogram data is established to solve the problem of extracting effective aging features from irregular discharge data. Combined with a degradation trajectory prediction model, accurate prediction of capacity trajectories under different vehicle types, regularization cycles, and charging or discharging conditions is achieved.

[0005] Therefore, this invention provides a method for predicting battery pack degradation trajectory based on actual operating data, comprising: Step 1: Preprocess the mileage, cycle count, SOC, voltage, current of the battery pack, and voltage and temperature data of individual cells recorded by the battery management system under actual operating conditions; Step 2: Select current, voltage, SOC, and temperature as key factors affecting battery aging, and calculate the duration, ampere-hour throughput, and watt-hour throughput under the conditions of each type of factor interval. Step 3: Using the normalized charge and discharge usage intensity histogram sequence data, construct four levels of aging feature sets; the aging feature sets include vehicle level, battery pack level, single cell level, and inconsistency level; Step 4: Smooth the capacity trajectory of actual vehicles during their service life using locally weighted linear regression; Step 5: Use a three-step feature screening method to screen the aging feature library obtained from Step 3 and Step 4; Step 6: The feature space obtained from Step 5 will be used as the input to the Transformer model to predict the battery pack capacity decay trajectory.

[0006] Furthermore, the preprocessing of the mileage, cycle count, battery pack SOC, voltage, current, and individual cell voltage and temperature data recorded by the battery management system under actual operating conditions specifically includes: Data cleaning: Remove missing data from the collected battery voltage signals; Charge / discharge segment division: Based on the battery system's charging status flag, measurement data for the charging and discharging segments are obtained.

[0007] Furthermore, the specific calculation methods for the duration, ampere-hour throughput, and watt-hour throughput under the various factor interval conditions are as follows: The variable X includes battery pack level data and cell level data; the battery pack level data covers battery pack current, battery pack voltage, SOC, voltage range within the pack, maximum voltage within the pack, minimum voltage within the pack, maximum temperature difference within the pack, and maximum temperature within the pack; the cell level data covers the voltage and temperature of the cell at the temperature measurement point; the variable X is divided into several intervals, each interval representing a specific level, and the binning points for dividing each sub-interval are shown in equation (1): ; ; In the formula, Indicates the minimum current; This is the maximum current; This refers to the width of the container. The total voltage of the battery pack is divided into several compartments; After binning the data, the statistics for each interval are calculated. The cumulative working time, ampere-hour throughput, and watt-hour throughput; within the variable X interval Cumulative working hours ampere-hour throughput and watt-hour throughput The calculation formulas are shown in equations (3), (4), and (5): ; ; ; In the formula, Indicates the interval The time increment at each moment within the time frame; For in the interval Inner Time The measured current value; For in the interval Inner Time The measured voltage value; Finally, the time series data of variable X is converted into a structured cumulative time series. ampere-hour throughput sequence and watt-hour throughput sequence As shown in equations (6) to (8): ; ; .

[0008] Furthermore, step 3 specifically includes, The vehicle-level features include mileage. Number of charge-discharge cycles The description uses key intensity features, specifically: trip mileage within each regularization period. ,total Total experience time ampere-hour throughput watt-hour throughput The calculation formulas are shown in equations (9) to (13): ; ; ; ; ; In the formula, Indicates the regularization period; subscript For the first time in the regularization period One charge-discharge cycle; Indicates the first [number]th ... The driving distance of each driving act; as well as These represent the th [number] ... Subsequent charging or discharging behavior, depth of charge / discharge, ampere-hour throughput, and watt-hour throughput, superscript. and Used to distinguish between charging and discharging behavior; The battery pack-level features include multi-type battery pack data recorded by the battery management system. The battery pack-level features are converted into cumulative time series, ampere-hour throughput series and watt-hour throughput series respectively according to step 2. Then, eight statistical features of the three types of data series are calculated respectively, as shown in equation (14): ; In the formula, max, min, mean, var, kurt, skew, range, and slop represent the maximum, minimum, mean, variance, kurtosis, skewness, range, and slope of the eight statistical methods, respectively. express The first of the sequence One value; express The number of sampling points in the sequence; yes Standard deviation of the sequence; This represents the mean of the sequence; The single-cell level features include single-cell voltage and temperature point measurement data. According to step 2, the generated cumulative time, ampere-hour throughput and watt-hour throughput are converted into three types of data sequences respectively. Then, the eight statistical features of each data sequence are calculated by applying formula (14) to generate features that comprehensively characterize the usage intensity of each single cell. The inconsistency level features include eight statistical features of each feature category calculated by formula (14) based on the individual feature category, thereby analyzing the differences in usage between individual entities.

[0009] Furthermore, the capacity trajectory of actual vehicles during their service life is smoothed using locally weighted linear regression, specifically including: The expression for the locally weighted linear regression model is shown in equation (15): ; In the formula, For the regression coefficients, a weight matrix is ​​introduced when calculating the loss. Each element in the matrix Representing data points The weights are determined, therefore the weighted loss function is expressed as: ; Weights are determined using a Gaussian kernel function. The specific form is shown in (17): ; In the formula, The bandwidth parameter determines the model's focus on neighboring data points; ultimately, the parameters are estimated by minimizing the weighted loss function. The solution is in the form of (18): ; In the formula, For the input matrix, It is the target vector.

[0010] Furthermore, the three-step feature selection method specifically includes: The first step is to calculate the Pearson correlation coefficient. If the correlation coefficient is less than S1, it is considered an irrelevant feature and is deleted. After removing irrelevant features, the second step is to select the feature with the highest correlation score from the remaining feature library and designate it as the reference feature. Any feature whose Pearson correlation coefficient with the reference feature reaches or exceeds S2 is considered a redundant feature and is removed. The third step is to determine the best k features based on the maximum correlation-minimum redundancy algorithm. The method for determining the optimal k features based on the maximum correlation-minimum redundancy algorithm specifically includes, firstly, the mutual information between variables x and y, calculated as shown in equation (19): ; In the formula, Represents random variables and The joint probability density function; and Representing variables respectively and The marginal probability density function; Then, find the feature subset containing k features. As the final feature space, in , where n is the total number of input features, and the feature set is... With target variable Mutual information between Calculated using equation (20): ; Secondly, it is necessary to ensure that the redundancy between features is minimized, and that the mutual information between features within the same set is sufficient. Calculated using equation (21): ; In the formula, and They represent the first in the set. The and the first One characteristic, Features and characteristics Mutual information between them; Furthermore, given mutual information, the correlation and redundancy between features are evaluated using two methods: additive integral and multiplicative integral, denoted as follows: and ; Finally, an incremental search method is used to determine the optimal feature set, that is, from the remaining features in the existing feature set. Search the The features are shown in equation (22): ; In the formula, It is the input feature set In but not existing Features in It is a feature set Features of [the text].

[0011] Furthermore, step 6 specifically includes: The Transformer model consists of an input module, an encoder module, a decoder module, and an output module. In the input module, the normalized historical feature matrix Each eigenvector in the matrix is ​​transformed into a continuous vector of fixed dimensions, represented as a matrix. Furthermore, positional encoding is introduced to enable the model to distinguish different positions in the input sequence and generate the final input, as shown in equations (23) and (24): ; ; In the formula, It is the sequence length. Indicates the dimension of the vector; The input sequence is a high-dimensional vector matrix; It is the first A high-dimensional vector of input elements; This is the position encoding matrix; It is the first in the sequence Encoded vectors at each position; Indicates positional encoding The The position of the first Values ​​in each dimension; The encoder module consists of multiple identical stacked layers, each layer comprising two sub-layers. Each sub-layer includes a multi-head self-attention mechanism and a feedforward neural network. Residual connections are used in each sub-layer to add the sub-layer input and output, followed by layer normalization. For the input matrix... The output generated by the multi-head self-attention mechanism layer is represented as: ; In the formula, and It is the weight matrix for learning; and These are the query and the key-value matrix, respectively; the feedforward neural network consists of fully connected layers and nonlinear activation functions, and its mathematical expression is: ; In the formula, and As the weight for learning, and For the bias term; after using residual connections in each of the sub-layers, the output of each sub-layer is represented as: ; The decoder module consists of multiple identical layers stacked together, each layer comprising three sub-layers: two multi-head self-attention mechanisms and a feedforward neural network; the first sub-layer is a masked multi-head self-attention mechanism used to process the decoder's own input sequence, specifically, given the query matrix of the decoder input. Key matrix Sum matrix The formula for calculating attention weights has been revised as follows: ; In the formula, It's a mask matrix used to mask attention weights for future positions; the second sub-layer of the decoder is a multi-head encoder-decoder attention mechanism, whose query matrix... The key matrix is ​​the output from the previous layer of the decoder. Sum matrix This process originates from the encoder output and is represented as follows: ; The feedforward neural network of the decoder has the same function as the encoder structure; each decoder layer produces an output after being processed by three sub-layers. After being processed layer by layer by multiple decoder layers, the final output is transformed by a linear layer and a softmax layer to generate the predicted target sequence. The model performance evaluation metrics are root mean square error, maximum absolute error, root mean square percentage error, and maximum absolute percentage error, which are used to evaluate the predictive performance of the model. The calculation formulas are shown in equations (28) to (31), respectively: ; ; .

[0012] On the other hand, the present invention also provides a battery pack degradation trajectory prediction system based on actual operating data, comprising: Module 1: Used to preprocess the mileage, cycle count, SOC, voltage, current of the battery pack, and voltage and temperature data of individual cells recorded by the battery management system under actual operating conditions; Module 2: Used to select current, voltage, SOC, and temperature as key factors affecting battery aging, and calculate the duration, ampere-hour throughput, and watt-hour throughput under the conditions of each type of factor interval; Module 3: Used to construct four levels of aging feature sets using the normalized charge-discharge usage intensity histogram sequence data; the aging feature sets include vehicle level, battery pack level, single cell level and inconsistency level; Module 4: Used to smooth the capacity trajectory of actual vehicles during their service life using locally weighted linear regression; Module 5: Used to perform feature filtering on the aging feature library obtained from Modules 3 and 4 using a three-step feature filtering method; Module 6: Used to predict the battery pack capacity decay trajectory by using the feature space obtained from Module 5 as input to the Transformer model.

[0013] Furthermore, the preprocessing specifically includes, Data cleaning module: used to remove missing data from the collected battery voltage signals; Charge / discharge segment division module: used to obtain measurement data of charging and discharging segments based on the battery system charging status flag bits.

[0014] Furthermore, the three-step feature selection method specifically includes: The first step is to calculate the Pearson correlation coefficient. If the correlation coefficient is less than S1, it is considered an irrelevant feature and is deleted. After removing irrelevant features, the second step is to select the feature with the highest correlation score from the remaining feature library and designate it as the reference feature. Any feature whose Pearson correlation coefficient with the reference feature reaches or exceeds S2 is considered a redundant feature and is removed. The third step is to determine the best k features based on the maximum correlation-minimum redundancy algorithm. The method for determining the optimal k features based on the maximum correlation-minimum redundancy algorithm specifically includes, firstly, the mutual information between variables x and y, calculated as shown in equation (19): ; In the formula, Represents random variables and The joint probability density function; and Representing variables respectively and The marginal probability density function; Then, find the feature subset containing k features. As the final feature space, in (n is the total number of input features), feature set With target variable Mutual information between Calculated using equation (20): ; Secondly, it is necessary to ensure that the redundancy between features is minimized, and that the mutual information between features within the same set is sufficient. Calculated using equation (21): ; In the formula, and They represent the first in the set. The and the first One characteristic, Features and characteristics Mutual information between them; Furthermore, given mutual information, the correlation and redundancy between features are evaluated using two methods: additive integral and multiplicative integral, denoted as follows: and ; Finally, an incremental search method is typically used to determine the optimal feature set, that is, from the remaining features in the existing feature set. Search the The features are shown in equation (22): ; In the formula, It is the input feature set In but not existing Features in It is a feature set Features of [the text].

[0015] The present invention has the following beneficial technical effects: This invention transforms complex and irregular time-series data into histogram data describing battery usage intensity by statistically analyzing the duration, ampere-hour throughput, and watt-hour throughput of key aging parameters such as current, voltage, SOC, and temperature in different ranges. This effectively reduces data fluctuations while retaining key information. A multi-level feature construction method based on charge-discharge histogram data is proposed, which solves the problem of extracting effective aging features from irregular discharge data. The Transformer algorithm is used to accurately predict the capacity decay trajectory of real vehicle battery packs. Attached Figure Description

[0016] Figure 1 This is a schematic flowchart of the method of the present invention; Figure 2 This is a schematic diagram of the normalization processing results of current timing data under 5 charge-discharge cycles according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the normalization process of the timing data of the battery pack voltage under 5 charge-discharge cycles according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the normalization processing results of the battery pack SOC timing data under 5 charge-discharge cycles according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the timing data of the internal voltage difference of the battery pack and the normalization processing results under 5 charge-discharge cycles according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the time series data and normalization results of the temperature difference within the battery pack under 5 charge-discharge cycles according to an embodiment of the present invention. Figure 7 This is a schematic diagram of the time-series data and normalization results of the single cell temperature under 5 charge-discharge cycles in an embodiment of the present invention; Figure 8 This is a schematic diagram of the time-series data and normalization results of the single cell temperature under 5 charge-discharge cycles in an embodiment of the present invention; Figure 9 This is a schematic diagram illustrating the capacity change of a vehicle battery pack under different charge-discharge cycles according to an embodiment of the present invention. Figure 10 This is a schematic diagram illustrating the number of charge / discharge cycles for each vehicle during data acquisition, according to an embodiment of the present invention. Figure 11 This is a schematic diagram of the Pearson correlation matrix of the optimal feature subset in an embodiment of the present invention; Figure 12 This is a schematic diagram illustrating the error distribution of the vehicle capacity sequence prediction results for the test set in an embodiment of the present invention. Figure 13 This is a comparative diagram of RMSE, MAXE, RMSPE, and MAXPE in an embodiment of the present invention. Detailed Implementation

[0017] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0018] like Figure 1 The diagram shows a method and system flow chart for predicting battery pack degradation trajectory based on actual operating data, which proceeds according to the following steps: Step 1: Preprocess the mileage, cycle count, SOC, voltage, current of the battery pack, and voltage and temperature data of individual cells recorded by the battery management system under actual operating conditions.

[0019] Missing data in the collected battery voltage signal is deleted, and the measurement data of the charging and discharging stages are obtained based on the battery system charging status flag.

[0020] Step 2: Select current, voltage, SOC, and temperature as key factors affecting battery aging, and calculate the duration, ampere-hour throughput, and watt-hour throughput under the conditions of each type of factor interval.

[0021] The variable X includes battery pack level data and cell level data. Battery pack level data includes battery pack current, battery pack voltage, SOC, voltage range within the pack, maximum voltage within the pack, minimum voltage within the pack, maximum temperature difference within the pack, and maximum temperature within the pack. Cell level data includes cell voltage and temperature at the temperature measurement point. The variable X is divided into several intervals, each interval representing a specific level. The binning points for dividing each sub-interval are shown in equation (1): ; ; In the formula, Indicates the minimum current; This is the maximum current; This refers to the width of the container. This refers to the number of compartments for the total voltage of the battery pack.

[0022] After binning the data, the next step is to calculate the data for each interval. The cumulative working time, ampere-hour throughput, and watt-hour throughput. Within the variable X interval... Cumulative working hours ampere-hour throughput and watt-hour throughput The calculation formulas are shown in equations (3) to (5): ; ; ; In the formula, Indicates the interval The time increment at each moment within the time frame; For in the interval Inner Time The measured current value; For in the interval Inner Time The measured voltage value. Finally, the time series data of variable X is converted into a structured cumulative time series. ampere-hour throughput sequence and watt-hour throughput sequence As shown in equations (6) to (8): ; ; ; In this embodiment, the battery pack-level data includes battery pack current I, battery pack voltage volt, SOC, intra-pack voltage range diffvol, intra-pack maximum voltage maxvol, intra-pack minimum voltage minvol, intra-pack maximum temperature difference difftemp, and intra-pack maximum temperature maxtemp. The individual cell-level data includes the voltage vol_i of cell i and the temperature temp_j of temperature measurement point j. Therefore, X∈{I, volt, SOC, diffvol, maxvol, minvol, difftemp, maxtemp}∪{vol_i |i=1,2,3…,n}∪{temp_j | i=1,2,3…,m}, X name ∈{ I, volt, SOC, diffvol, maxvol,minvol, difftemp, maxtemp}∪{ vol_i | i=1,2,3…,n}∪{temp_j | i=1,2,3…,m}, where n represents the number of individual cells in the battery pack (96) and represents the number of temperature measurement points in the battery pack (32). When calculating the duration, ampere-hour throughput, and watt-hour throughput under various factor interval conditions, the minimum, maximum, and compartment widths of each factor variable are set as follows: The minimum battery pack current is -120A, the maximum is 160A, and the compartment width is set to 20A; the minimum battery pack voltage is 330V, the maximum is 420V, and the compartment width is set to 2V; the minimum battery pack SOC is 5%, the maximum is 95%, and the compartment width is set to 5%; the minimum internal voltage difference is 0mV, the maximum is 150mV, and the compartment width is set to 10mV; the minimum internal temperature difference is set to 0℃, the maximum temperature difference is set to 10℃, and the temperature difference range width is 1℃. The minimum maximum voltage of the battery pack is 3.3V, and the maximum is 4.3V, with a compartment width set to 0.01V. The minimum minimum voltage of the battery pack is 3.3V, and the maximum is 4.3V, with a compartment width set to 0.01V. The minimum maximum temperature of the battery pack is 0℃, and the maximum is 40℃, with a compartment width set to 1℃. The minimum voltage of a single cell is 3.3V, and the maximum is 4.3V, with a compartment width set to 0.01V. The minimum temperature of a single cell is 0℃, and the maximum is 40℃, with a compartment width set to 1℃. The normalization period for all variable factors is set to 5. The normalization processing of the current time series data under 5 charge-discharge cycles is shown in Figure 2. The normalization processing of the battery pack voltage time series data under 5 charge-discharge cycles is shown in Figure 2. Figure 3 Figure 4 shows the normalized processing of the battery pack's SOC timing data after 5 charge-discharge cycles. Figure 5 shows the timing data and normalized results of the internal voltage difference of the battery pack after 5 charge-discharge cycles. Figure 6 shows the timing data and normalized results of the internal temperature difference of the battery pack after 5 charge-discharge cycles. Figure 7 shows the timing data and normalized results of the individual cell temperature after 5 charge-discharge cycles. Figure 8 As shown.

[0023] Step 3: Using the normalized charge and discharge usage intensity histogram sequence data, construct four levels of aging feature sets: vehicle level, battery pack level, single cell level, and inconsistency level.

[0024] Using regularized charge-discharge intensity histogram data can address the issue of drastic performance changes in battery packs under dynamic operating conditions. It provides more stable and consistent data input under complex conditions, avoiding the volatility and uncertainty present in time-series data. The constructed aging feature set is divided into four levels: vehicle level, battery pack level, individual cell level, and inconsistency level.

[0025] Vehicle-level characteristics include mileage Number of charge-discharge cycles And five key characteristics describing usage intensity, specifically: trip mileage within each regularization cycle. ,total Total experience time ampere-hour throughput watt-hour throughput The calculation formulas are shown in equations (9) to (13): ; ; ; ; ; In the formula, Indicates the regularization period; subscript For the first time in the regularization period One charge-discharge cycle; Indicates the first [number]th ... The driving distance of each driving act; as well as These represent the th [number] ... Subsequent charging or discharging behavior, depth of charge / discharge, ampere-hour throughput, and watt-hour throughput, superscript. and Used to distinguish between charging and discharging behavior.

[0026] Battery pack-level features include multi-type battery pack data recorded by the battery management system. These data are converted into cumulative time series, ampere-hour throughput series, and watt-hour throughput series respectively according to step 2. Then, eight statistical features of the three types of data series are calculated respectively, as shown in the following calculation method. The calculation method is shown in equation (14).

[0027] ; In the formula, max, min, mean, var, kurt, skew, range, and slop represent the maximum, minimum, mean, variance, kurtosis, skewness, range, and slope of the eight statistical methods, respectively. express The first of the sequence One value; express The number of sampling points in the sequence; yes Standard deviation of the sequence; This represents the mean of the sequence.

[0028] The individual cell-level characteristics include individual cell voltage and temperature measurement data. Following step 2, the generated data sequences are converted into three categories: cumulative time, ampere-hour throughput, and watt-hour throughput. Then, formula (14) is applied to calculate eight statistical characteristics for each data sequence, generating features that comprehensively characterize the usage intensity of each individual cell. The inconsistency-level characteristics include individual cell-level characteristic categories. Formula (14) is applied to calculate eight statistical characteristics for each characteristic category, thereby analyzing the usage differences between individual cells.

[0029] In this example, the first layer of the actual vehicle battery pack feature set is vehicle-level features, totaling 7. The second layer is battery pack-level features, which uses 8 types of battery pack data recorded by the battery management system to generate 24 data sequences in 3 categories: cumulative time, ampere-hour throughput, and watt-hour throughput. The statistical method of formula (14) is applied to these 24 sequences, and a total of 192 features are extracted. The third layer is cell-level features, which includes 384 different types of data in 3 categories: usage time, ampere-hour throughput, and watt-hour throughput of 96 cells and 32 temperature measurement points under different voltage or temperature conditions. The statistical method of formula (14) is applied to extract a total of 3072 features. The fourth layer is inconsistency features, which includes 6 categories of usage time, ampere-hour throughput, and watt-hour throughput of each cell under different voltage and temperature levels. The statistical method of formula (14) is applied to obtain 48 categories. The feature values ​​under each category are then analyzed using the statistical method of formula (14) to analyze the usage differences between cells, resulting in a total of 384 features. Through the construction of the above four-layer feature sets, a total of 3,655 features were generated, covering multiple levels from macro to micro, and capturing the battery pack's performance under different conditions.

[0030] Step 4: Smooth the capacity trajectory of the actual vehicle during its service life using locally weighted linear regression.

[0031] The expression for the locally weighted linear regression model is shown in equation (15): ; In the formula, These are the regression coefficients. A weight matrix is ​​introduced when calculating the loss. Each element in the matrix Representing data points The weights are determined, therefore the weighted loss function can be expressed as: ; Weights are determined using a Gaussian kernel function. Its specific form is shown in (5-24): ; In the formula, The bandwidth parameter determines the degree to which the model pays attention to neighboring data points. The parameters are ultimately estimated by minimizing the weighted loss function. The solution takes the form shown in (18): ; In the formula, For the input matrix, It is the target vector.

[0032] In this example, take It is 65. Figure 9 This demonstrates the capacity change of a real vehicle's battery pack across different charge-discharge cycles. The fitted curve effectively smooths out the fluctuations in the original data, more intuitively reflecting the overall capacity decay trend with increasing charge-discharge cycles. After processing all vehicles in the dataset, based on the number of charge-discharge cycles for each vehicle during the data acquisition period, such as... Figure 10 As shown, the dataset is divided into three categories: less than 600 charge / discharge cycles, 600-800 cycles, and more than 800 cycles.

[0033] Step 5: Use a three-step feature screening method to screen the aging feature library obtained from Step 3 and Step 4.

[0034] The first step is to calculate the Pearson correlation coefficient. If the correlation coefficient is less than S1, it is considered an irrelevant feature and is deleted. After removing irrelevant features, the second step is to select the feature with the highest correlation score from the remaining feature library and designate it as the reference feature. Any feature whose Pearson correlation coefficient with the reference feature reaches or exceeds S2 is considered a redundant feature and is removed. The third step is to determine the best k features based on the maximum correlation-minimum redundancy algorithm. To determine the best k features based on the maximum correlation-minimum redundancy algorithm, the mutual information between variables x and y is required, and the calculation method is shown in equation (19).

[0035] ; In the formula, Represents random variables and The joint probability density function; and Representing variables respectively and The marginal probability density function. Then, find the feature subset containing k features. As the final feature space, in (n is the total number of input features), feature set With target variable Mutual information between It can be calculated using equation (20): ; Secondly, it is necessary to ensure that the redundancy between features is minimized. For features within the same set, mutual information between them is crucial. It can be calculated using equation (21): ; In the formula, and They represent the first in the set. The and the first One characteristic, Features and characteristics Mutual information between features. Furthermore, given mutual information, the correlation and redundancy between features can be evaluated using two methods: additive integral and multiplicative integral, denoted as... and Finally, an incremental search method is typically used to determine the optimal feature set, that is, from the remaining features in the existing feature set. Search the The process is shown in equation (22): ; In the formula, It is the input feature set In but not existing Features in It is a feature set Features of [the text].

[0036] In this embodiment, the correlation coefficient threshold S1 for irrelevant feature removal is set to 0.2; the correlation coefficient threshold S2 for redundant feature removal is set to 0.9; when the maximum correlation-minimum redundancy algorithm selects key features, k is initially set to... init The setting is 15. The first step of the screening process removed 3016 features with a relevance below 0.2, a removal rate of 82.5%, ultimately retaining 639 features. The second step removed 617 redundant features, ultimately retaining 22 features. The Pearson relevance matrix of the optimal feature subset selected through the maximum relevance-minimum redundancy algorithm is shown below. Figure 11 As shown. The final 22 retained features cover battery information at various levels: 3 vehicle-level features, 5 battery pack-level features, 9 cell-level features, and 5 inconsistency features.

[0037] Step 6: The feature space obtained in Step 5 will be used as input to the Transformer model to predict the battery pack capacity degradation trajectory. The Transformer model mainly consists of four parts: input module, encoder module, decoder module, and output module.

[0038] In the input module, the normalized historical feature matrix Each eigenvector in the matrix is ​​transformed into a continuous vector of fixed dimensions, represented as a matrix. Furthermore, positional encoding is introduced to enable the model to distinguish different positions in the input sequence and generate the final input, as shown in equations (23) and (24).

[0039] ; ; In the formula, It is the sequence length. Indicates the dimension of the vector; The input sequence is a high-dimensional vector matrix; It is the first A high-dimensional vector of input elements; This is the position encoding matrix; It is the first in the sequence A coding vector at each position. Indicates positional encoding The The position of the first The values ​​of each dimension.

[0040] The encoder module consists of multiple identical stacked layers, each containing two main sub-layers: a multi-head self-attention mechanism and a feedforward neural network. Within each sub-layer, residual connections are used to add the sub-layer's input and output, which are then normalized via layer normalization. For the input matrix... The output generated by the multi-head self-attention mechanism layer is represented as: ; In the formula, and It is the weight matrix for learning; and These are the query and the key-value matrix, respectively. A feedforward neural network consists of fully connected layers and a non-linear activation function (usually ReLU), and its mathematical expression is: ; In the formula, and As the weight for learning, and This is the bias term. After using residual connections in each sub-layer, the output of each sub-layer can be expressed as: ; The decoder module also consists of multiple identical layers stacked together, each layer comprising three sub-layers: two multi-head self-attention mechanisms and a feedforward neural network. The first sub-layer is a masked multi-head self-attention mechanism used to process the decoder's own input sequence. Specifically, given the query matrix of the decoder input... Key matrix Sum matrix The formula for calculating attention weights has been revised as follows: ; In the formula, This is a mask matrix used to mask the attention weights for future positions. The second sub-layer of the decoder is a multi-head encoder-decoder attention mechanism, whose query matrix... The key matrix is ​​the output from the previous layer of the decoder. Sum matrix The process, originating from the encoder output, can be represented as: ; The feedforward neural network of the decoder has the same structure as the encoder. Each decoder layer produces an output after being processed by three sub-layers. After being processed layer by layer by multiple decoder layers, the final output is transformed by a linear layer and a softmax layer to generate the predicted target sequence.

[0041] The model performance evaluation metrics are root mean square error (RMSE), maximum absolute error (MAXE), root mean square percentage error (RMSE), and maximum absolute percentage error (MAXPE) to evaluate the predictive performance of the model. The calculation formulas are shown in equations (28) to (31), respectively. ; ; ; ; In this embodiment, the rated capacity of the battery used for predicting the capacity degradation trajectory of the actual vehicle battery pack is 155Ah. The Transformer model parameters are set as follows: 2 layers, 256 high-dimensional vector dimensions, 8 heads for the multi-head self-attention mechanism, 1024 dimensions for the feedforward neural network, and a batch size of 64. The ratio of training set to validation set is 9:1, and the prediction step size and warping period are both set to 5 iterations. The error distribution of the vehicle capacity sequence prediction results on the test set is shown below. Figure 12 As shown. Figure 13As shown in the RMSE distribution, the prediction error of the model for most samples is concentrated in the range of 0 to 1 Ah, with an average RMSE of 0.87 Ah, indicating that the model can accurately capture the battery capacity degradation trajectory in most cases. Even in extreme cases, the maximum RMSE is only 2.78 Ah, indicating that the model can effectively control the error and will not produce excessive deviation. The average maximum error is 1.40 Ah, further verifying the stable performance of the model on most test samples. Even in the worst case, the maximum MAXE is only 4.34 Ah. Similarly, RMSPE and MAXPE, as relative error indicators, can reflect the model's predictive ability under different battery capacities and operating conditions. In the prediction of real vehicle battery packs, the average RMSPE is 0.61%, while the average MAXPE is 0.99%, which also indicates that the model can accurately predict capacity changes under different usage conditions. It is worth noting that the maximum RMSPE is 1.95%, while the maximum MAXPE is only 3.14%, both lower than the maximum RMSPE of 2.42% and the maximum MAXPE of 5.17% under laboratory conditions, respectively. This performance indicates that the model exhibits good adaptability and stronger error control capabilities when handling complex real-world vehicle conditions. Furthermore, this also verifies the effectiveness of the selected feature subset under complex driving conditions, accurately capturing the capacity changes of the battery during actual use, further demonstrating the model's feasibility and robustness in real-world environments.

[0042] The contents not described in detail in this specification are existing technologies known to those skilled in the art.

Claims

1. A method for predicting battery pack degradation trajectory based on actual operating data, characterized in that, include: Step 1: Preprocess the mileage, cycle count, SOC, voltage, current of the battery pack, and voltage and temperature data of individual cells recorded by the battery management system under actual operating conditions; Step 2: Select current, voltage, SOC, and temperature as key factors affecting battery aging, and calculate the duration, ampere-hour throughput, and watt-hour throughput under the conditions of each type of factor interval. Step 3: Using the normalized charge and discharge usage intensity histogram sequence data, construct four levels of aging feature sets; the aging feature sets include vehicle level, battery pack level, single cell level, and inconsistency level; Step 4: Smooth the capacity trajectory of actual vehicles during their service life using locally weighted linear regression; Step 5: Use a three-step feature screening method to screen the aging feature library obtained from Step 3 and Step 4; Step 6: The feature space obtained from Step 5 will be used as the input to the Transformer model to predict the battery pack capacity decay trajectory.

2. The battery pack degradation trajectory prediction method based on actual operating data as described in claim 1, characterized in that, The preprocessing of data on mileage, cycle count, battery pack SOC, voltage, current, and individual cell voltage and temperature recorded by the battery management system under actual operating conditions specifically includes: Data cleaning: Remove missing data from the collected battery voltage signals; Charge / discharge segment division: Based on the battery system's charging status flag, measurement data for the charging and discharging segments are obtained.

3. The battery pack degradation trajectory prediction method based on actual operating data as described in claim 1, characterized in that, The calculation methods for duration, ampere-hour throughput, and watt-hour throughput under various factor interval conditions are as follows: The variable X for each factor includes battery pack level data and cell level data; the battery pack level data covers battery pack current, battery pack voltage, SOC, voltage range within the pack, maximum voltage within the pack, minimum voltage within the pack, maximum temperature difference within the pack, and maximum temperature within the pack. The individual unit level data covers the voltage and temperature of the individual unit's temperature measurement points; the variable X is divided into several intervals, each interval representing a specific level, and the binning points for dividing each sub-interval are shown in equation (1): ; ; In the formula, Indicates the minimum current; This is the maximum current; This refers to the width of the container. The total voltage of the battery pack is divided into several compartments; After binning the data, the statistics for each interval are calculated. Cumulative working time, ampere-hour throughput, and watt-hour throughput; In variables interval Cumulative working hours ampere-hour throughput and watt-hour throughput The calculation formulas are shown in equations (3), (4), and (5): ; ; ; In the formula, Indicates the interval The time increment at each moment within the time frame; In the interval Inner Time The measured current value; In the interval Inner Time The measured voltage value; Finally, the variables Time series data are converted into structured cumulative time series. ampere-hour throughput sequence and watt-hour throughput sequence As shown in equations (6) to (8): ; ; 。 4. The battery pack degradation trajectory prediction method based on actual operating data as described in claim 1, characterized in that, Step 3 specifically includes, The vehicle-level features include mileage. Number of charge-discharge cycles The description uses key intensity features, specifically: trip mileage within each regularization period. ,total Total experience time ampere-hour throughput watt-hour throughput The calculation formulas are shown in equations (9) to (13): ; ; ; ; ; In the formula, Indicates the regularization period; subscript For the first time in the regularization period One charge-discharge cycle; Indicates the first [number]th ... The driving distance of each driving act; as well as These represent the th [number] ... Subsequent charging or discharging behavior, depth of charge / discharge, ampere-hour throughput, and watt-hour throughput, superscript. and Used to distinguish between charging and discharging behavior; The battery pack-level features include multi-type battery pack data recorded by the battery management system. The battery pack-level features are converted into cumulative time series, ampere-hour throughput series and watt-hour throughput series respectively according to step 2. Then, eight statistical features of the three types of data series are calculated respectively, as shown in equation (14): ; In the formula, max, min, mean, var, kurt, skew, range, and slop represent the maximum, minimum, mean, variance, kurtosis, skewness, range, and slope of the eight statistical methods, respectively. express The first of the sequence One value; express The number of sampling points in the sequence; yes Standard deviation of the sequence; This represents the mean of the sequence; The single-cell level features include single-cell voltage and temperature point measurement data. According to step 2, the generated cumulative time, ampere-hour throughput and watt-hour throughput are converted into three types of data sequences respectively. Then, the eight statistical features of each data sequence are calculated by applying formula (14) to generate features that comprehensively characterize the usage intensity of each single cell. The inconsistency level features include eight statistical features of each feature category calculated by formula (14) based on the individual feature category, thereby analyzing the differences in usage between individual entities.

5. The battery pack degradation trajectory prediction method based on actual operating data as described in claim 1, characterized in that, The capacity trajectory of actual vehicles during their service life is smoothed using locally weighted linear regression, specifically including: The expression for the locally weighted linear regression model is shown in equation (15): ; In the formula, For the regression coefficients, a weight matrix is ​​introduced when calculating the loss. Each element in the matrix Representing data points The weights are determined, therefore the weighted loss function is expressed as: ; Weights are determined using a Gaussian kernel function. The specific form is shown in (17): ; In the formula, The bandwidth parameter determines the model's focus on neighboring data points; ultimately, the parameters are estimated by minimizing the weighted loss function. The solution is in the form of (18): ; In the formula, For the input matrix, It is the target vector.

6. The battery pack degradation trajectory prediction method based on actual operating data as described in claim 1, characterized in that, The three-step feature selection method specifically includes: The first step is to calculate the Pearson correlation coefficient. If the correlation coefficient is less than S1, it is considered an irrelevant feature and is deleted. After removing irrelevant features, the second step is to select the feature with the highest correlation score from the remaining feature library and designate it as the reference feature. Any feature whose Pearson correlation coefficient with the reference feature reaches or exceeds S2 is considered a redundant feature and is removed. The third step is to determine the best k features based on the maximum correlation-minimum redundancy algorithm. The method for determining the optimal k features based on the maximum correlation-minimum redundancy algorithm specifically includes, firstly, the mutual information between variables x and y, calculated as shown in equation (19): ; In the formula, Represents random variables and The joint probability density function; and Representing variables respectively and The marginal probability density function; Then, find the feature subset containing k features. As the final feature space, in For the total number of input features, the feature set With target variable Mutual information between Calculated using equation (20): ; Secondly, it is necessary to ensure that the redundancy between features is minimized, and that the mutual information between features within the same set is sufficient. Calculated using equation (21): ; In the formula, and They represent the first in the set. The and the first One characteristic, Features and characteristics Mutual information between them; Furthermore, given mutual information, the correlation and redundancy between features are evaluated using two methods: additive integral and multiplicative integral, denoted as follows: and ; Finally, an incremental search method is used to determine the optimal feature set, that is, from the remaining features in the existing feature set. Search the The features are shown in equation (22): ; In the formula, It is the input feature set In but not existing Features in It is a feature set Features of [the text].

7. The battery pack degradation trajectory prediction method based on actual operating data as described in claim 6, characterized in that, Step 6 specifically includes: The Transformer model comprises an input module, an encoder module, a decoder module, and an output module; in the input module, the normalized historical feature matrix... Each eigenvector in the matrix is ​​transformed into a continuous vector of fixed dimensions, represented as a matrix. Furthermore, positional encoding is introduced to enable the model to distinguish different positions in the input sequence and generate the final input, as shown in equations (23) and (24): ; ; In the formula, It is the sequence length. Indicates the dimension of the vector; The input sequence is a high-dimensional vector matrix; It is the first A high-dimensional vector of input elements; This is the position encoding matrix; It is the first in the sequence Encoded vectors at each position; Indicates positional encoding The The position of the first Values ​​for each dimension; The encoder module consists of multiple identical stacked layers, each layer comprising two sub-layers. Each sub-layer includes a multi-head self-attention mechanism and a feedforward neural network. Residual connections are used in each sub-layer to add the sub-layer input and output, followed by layer normalization. For the input matrix... The output generated by the multi-head self-attention mechanism layer is represented as: ; In the formula, It is the weight matrix for learning; and These are the query and the key-value matrix, respectively; the feedforward neural network consists of fully connected layers and nonlinear activation functions, and its mathematical expression is: ; In the formula, and As the weight for learning, and For the bias term; after using residual connections in each of the sub-layers, the output of each sub-layer is represented as: ; The decoder module consists of multiple identical layers stacked together, each layer comprising three sub-layers: two multi-head self-attention mechanisms and a feedforward neural network; the first sub-layer is a masked multi-head self-attention mechanism used to process the decoder's own input sequence, specifically, given the query matrix of the decoder input. Key matrix Sum matrix The formula for calculating attention weights has been revised as follows: ; In the formula, It's a mask matrix used to mask attention weights for future positions; the second sub-layer of the decoder is a multi-head encoder-decoder attention mechanism, whose query matrix... The key matrix is ​​the output from the previous layer of the decoder. Sum matrix This process originates from the encoder output and is represented as follows: ; The feedforward neural network of the decoder has the same function as the encoder structure; each decoder layer produces an output after being processed by three sub-layers. After being processed layer by layer by multiple decoder layers, the final output is transformed by a linear layer and a softmax layer to generate the predicted target sequence. The model performance evaluation metrics are root mean square error, maximum absolute error, root mean square percentage error, and maximum absolute percentage error, which are used to evaluate the predictive performance of the model. The calculation formulas are shown in equations (28) to (31), respectively: ; ; 。 8. A battery pack degradation trajectory prediction system based on actual operating data, characterized in that, include: Module 1: Used to preprocess the mileage, cycle count, SOC, voltage, current of the battery pack, and voltage and temperature data of individual cells recorded by the battery management system under actual operating conditions; Module 2: Used to select current, voltage, SOC, and temperature as key factors affecting battery aging, and calculate the duration, ampere-hour throughput, and watt-hour throughput under the conditions of each type of factor interval; Module 3: Used to construct four levels of aging feature sets using the normalized charge-discharge usage intensity histogram sequence data; the aging feature sets include vehicle level, battery pack level, single cell level and inconsistency level; Module 4: Used to smooth the capacity trajectory of actual vehicles during their service life using locally weighted linear regression; Module 5: Used to perform feature filtering on the aging feature library obtained from Modules 3 and 4 using a three-step feature filtering method; Module 6: Used to predict the battery pack capacity decay trajectory by using the feature space obtained from Module 5 as input to the Transformer model.

9. The battery pack degradation trajectory prediction system based on actual operating data as described in claim 8, characterized in that, The preprocessing specifically includes, Data cleaning module: used to remove missing data from the collected battery voltage signals; Charge / discharge segment division module: used to obtain measurement data of charging and discharging segments based on the battery system charging status flag bits.

10. The battery pack degradation trajectory prediction system based on actual operating data as described in claim 8, characterized in that, The aforementioned three-step feature selection method specifically includes: The first step is to calculate the Pearson correlation coefficient. If the correlation coefficient is less than S1, it is considered an irrelevant feature and is deleted. After removing irrelevant features, the second step is to select the feature with the highest correlation score from the remaining feature library and designate it as the reference feature. Any feature whose Pearson correlation coefficient with the reference feature reaches or exceeds S2 is considered a redundant feature and is removed. The third step is to determine the best k features based on the maximum correlation-minimum redundancy algorithm. The method for determining the optimal k features based on the maximum correlation-minimum redundancy algorithm specifically includes, firstly, the mutual information between variables x and y, calculated as shown in equation (19): ; In the formula, Represents random variables and The joint probability density function; and Representing variables respectively and The marginal probability density function; Then, find the feature subset containing k features. As the final feature space, in (n is the total number of input features), feature set With target variable Mutual information between Calculated using equation (20): ; Secondly, it is necessary to ensure that the redundancy between features is minimized, and that the mutual information between features within the same set is sufficient. Calculated using equation (21): ; In the formula, and They represent the first in the set. The and the first One characteristic, Features and characteristics Mutual information between them; Furthermore, given mutual information, the correlation and redundancy between features are evaluated using two methods: additive integral and multiplicative integral, denoted as follows: and ; Finally, an incremental search method is typically used to determine the optimal feature set, that is, from the remaining features in the existing feature set. Search the The features are shown in equation (22): ; In the formula, It is the input feature set In but not existing Features in It is a feature set Features of [the text].