Battery system life prediction method based on pre-training architecture
Through the combination of Transformer and Gaussian regression network based on pre-training architecture, the problems of low data utilization efficiency and insufficient model generalization capabilities in battery life prediction are solved, and high-precision and efficient battery life prediction are achieved, which is suitable for different battery types and operating conditions.
Patent Information
- Application Number
- CN202510722260.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art has problems such as low data utilization efficiency, insufficient model generalization capability and low prediction accuracy in battery life prediction, especially in cross-scenario migration and uncertainty evaluation.
Using a pre-training architecture method, general features are extracted through Transformer pre-trained models, and battery capacity prediction is performed in combination with Gaussian regression network, unsupervised training is used for unsupervised training, and a two-stage framework is built to improve the generalization ability and prediction accuracy of the model.
It improves the accuracy and cross-scenario adaptability of battery life prediction, significantly improves the prediction efficiency and accuracy of the model, can effectively capture the complexity of the battery operating state and provide probability prediction of battery capacity.
Smart Images

Figure CN120470935A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of battery life prediction, and in particular relates to a battery system life prediction method based on a pre-training architecture. Background Art
[0002] With the rapid development of new energy technologies, battery systems, as core components in electric vehicles, energy storage devices, and other fields, have become crucial for predicting their lifespan, ensuring safe equipment operation, formulating maintenance strategies, and controlling costs. RUL (remaining useful life) prediction aims to accurately assess the remaining cycles of a battery from its current state to end of life (EoL) by analyzing historical operating data. This allows for optimized charging and discharging strategies and early warning of failure risks. Battery capacity is a core input parameter for RUL prediction. By monitoring capacity decay, RUL can be effectively estimated, accurately reflecting the battery's RUL status and optimizing battery maintenance or replacement strategies.
[0003] However, existing technologies face significant challenges in data utilization efficiency, model generalization, and prediction accuracy. Regarding data, battery life data requires a complete set of charge and discharge cycles (typically hundreds to thousands of times) to obtain true battery capacity labels, a time-consuming, labor-intensive, and costly process. Furthermore, battery capacity labels cannot be determined until the battery is scrapped, resulting in a severe shortage of labeled data. Existing technologies often rely on supervised learning, requiring a large number of labeled samples for model training. However, unlabeled data in real-world scenarios (such as massive amounts of real-time battery status data) is not effectively utilized. Data sources are limited and inefficient, leading to inaccurate battery life prediction results. Regarding models, traditional recurrent neural networks (RNNs) and their variants (such as LSTMs and GRUs) suffer from the vanishing gradient problem when processing long time series data, making it difficult to capture early, subtle aging characteristics. Existing models are typically trained for a single battery type (such as proton exchange membrane fuel cells or lithium batteries) or specific operating conditions, lacking cross-scenario transferability. Furthermore, these models often output a single battery capacity prediction value, lacking an assessment of prediction uncertainty. This makes them difficult to address the risk-based decision-making requirements of real-world applications, resulting in insufficient prediction accuracy and robustness. Summary of the Invention
[0004] The present invention provides a battery system life prediction method based on a pre-training architecture to solve the problem of low battery life prediction accuracy.
[0005] The basic solution provided by the present invention is a battery system life prediction method based on a pre-training architecture, comprising the following steps: S1: Acquire multi-dimensional raw time series data, perform data preprocessing, and determine an unlabeled training set, a labeled training set, and a test set. The preprocessing includes deleting abnormal characters and invalid data and identifying charge and discharge data with a complete charge and discharge cycle according to data extraction rules; the data extraction rules are that the voltage-current correlation is greater than 0.8 and the SOC difference is greater than 30%; S2: Split the preprocessed data into time windows and expand them into multiple dimensions through one-dimensional convolution to form an input matrix; S3: Build a Transformer pre-training model, input the input matrix of the unlabeled training set into the pre-training model for pre-training, and output representation features. The Transformer pre-training model is an encoder-only structure. S4: Construct a Gaussian regression network model, input the input matrix and representation features of the label training set into the Gaussian regression network for training, input the validation set into the trained Gaussian regression network model to output the battery capacity prediction value and standard deviation, and evaluate the model.
[0006] Preferably, in S3, the Transformer prediction model includes an input embedding layer, a masked multi-head self-attention layer, a point-by-point feedforward network layer, a normalization and residual connection layer, and an output prediction head layer.
[0007] Preferably, in S4, the Gaussian regression network model includes an input layer, a hidden layer, a mean prediction branch layer, a standard deviation prediction branch layer, and an output layer, wherein the input layer is used to receive representation features, the hidden layer is used to flatten the input matrix and fuse time series features, the mean prediction branch layer is used to output the battery capacity prediction value, the standard deviation prediction branch layer is used to output the standard deviation, and the output layer is used to output the joint distribution parameters.
[0008] Further preferably, in S4, the hyperparameters of the Gaussian regression network model are optimized by a Bayesian optimization algorithm, and the hyperparameters include learning rate, batch size, number of training rounds, number of hidden layer neurons, regularization parameter, optimizer, and weight of loss function.
[0009] Further preferably, the specific steps of optimizing the hyperparameters include: a. Define the hyperparameter space; b. Initial sampling generates initial hyperparameter combinations and evaluates the initial hyperparameter combinations; c. Build a random forest proxy model; d. Determine the next sampling point using the sampling function, train the Gaussian network regression model using the determined hyperparameter combination, evaluate the model, and update the random forest model based on the evaluation results; e. Repeat step d until the number of iterations meets the preset number or the hyperparameter optimization converges; f. Output the optimal hyperparameter combination.
[0010] Preferably, in S3, the constraint function of the Transformer pre-training model adopts MSELoss, and the loss function expression adopted is as follows:
[0011] Where, is the true label, is the predicted probability distribution.
[0012] Preferably, in S4, the expression of the loss function of the Gaussian regression network model is as follows:
[0013] Where y is the actual battery capacity, μ and σ are the predicted mean and standard deviation, respectively, and C is a constant term.
[0014] Preferably, in S1, the multi-dimensional original time series data includes voltage, current, mileage, timestamp, and battery SOC value.
[0015] The principles and advantages of the present invention are: 1. Acquiring multidimensional data can fully capture the complexity of the battery's operating state. When processing data, the voltage-current correlation is considered to avoid the problem of voltage-current mismatch during real-vehicle data collection. The SOC difference is also limited, so that the filtered data represents a complete charge and discharge cycle, providing complete information from low SOC to high SOC and vice versa. At the same time, it can reduce the impact of measurement errors caused by changes in the charge and discharge curves and potential nonlinear behavior, or other transient interference-induced noise on the model. In addition, block segmentation and one-dimensional convolution of the preprocessed data can effectively reduce the number of model parameters and inference time, improving model prediction efficiency.
[0016] 2. Pre-training is performed based on the Transformer architecture, and prediction values are output based on the Gaussian network regression model. The combination of pre-training and Gaussian regression forms a two-stage framework of "feature learning-probability prediction". The pre-training model extracts common features, and Gaussian regression is fine-tuned for specific tasks, improving the model's generalization ability for different battery types, operating conditions, or environments, and effectively improving the model's prediction accuracy.
[0017] 3. Set up labeled training sets and unlabeled training sets, and perform unsupervised training on the unlabeled training sets in the pre-trained model to learn common temporal features and significantly improve cross-scenario adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a logic block diagram of the present invention; Figure 2 is a flow chart of the present invention; Figure 3 This is a distribution diagram of the prediction results of the training set of the present invention; Figure 4 This is a distribution diagram of the error range of the prediction results of the training set of the present invention; Figure 5 This is a distribution diagram of the prediction results of the test set of the present invention; Figure 6 This is a distribution diagram of the error range of the prediction results of the test set of the present invention. DETAILED DESCRIPTION
[0019] The following is further described in detail through specific implementation methods: The specific implementation process is as follows: Figures 1 to 6 , a battery system life prediction method based on a pre-trained architecture, comprising the following steps: S1: Acquire multi-dimensional raw time series data, perform data preprocessing, and determine an unlabeled training set, a labeled training set, and a test set. The preprocessing includes deleting abnormal characters and invalid data and identifying charge and discharge data with a complete charge and discharge cycle according to data extraction rules; the data extraction rules are that the voltage-current correlation is greater than 0.8 and the SOC difference is greater than 30%; In S1, the multi-dimensional original time series data includes voltage, current, mileage, timestamp, and battery SOC value. Specifically, the operation data based on GBT32960 is downloaded from the cloud platform of the electric vehicle, and the multi-dimensional time series data is selected. ,in represents the voltage at time step t , current ,mileage , timestamp and battery SOC value st; The following table shows some of the time series data collected in this example:
[0020] S2: The preprocessed data is divided into time windows and expanded into multiple dimensions through one-dimensional convolution to form an input matrix; the length of the time window is 3000 to 4000, which is set according to the actual prediction target, battery characteristics, and collected charging data; the data dimension expansion range includes 128 to 512. In this embodiment, the time window length is 4000 and the data is expanded to 256 dimensions.
[0021] S3: Build a Transformer pre-training model, input the input matrix of the unlabeled training set into the pre-training model for pre-training, and output representation features. The Transformer pre-training model is an encoder-only structure, the constraint function uses MSELoss, and the loss function expression used is as follows:
[0022] Where, is the true label, is the predicted probability distribution.
[0023] In S3, the Transformer prediction model includes an input embedding layer, a masked multi-head self-attention layer, a point-by-point feedforward network layer, a normalization and residual connection layer, and an output prediction head layer.
[0024] S4: Construct a Gaussian regression network model, input the input matrix and representation features of the label training set into the Gaussian regression network for training, input the validation set into the trained Gaussian regression network model to output the battery capacity prediction value and standard deviation, and use the negative log-likelihood function as the loss function to evaluate the model.
[0025] In S4, the loss function of the Gaussian regression network model is expressed as follows:
[0026] Where y is the actual battery capacity, μ and σ are the predicted mean and standard deviation, respectively, and C is a constant term.
[0027] In S4, the Gaussian regression network model includes an input layer, a hidden layer, a mean prediction branch layer, a standard deviation prediction branch layer, and an output layer. The input layer is used to receive representation features, the hidden layer is used to flatten the input matrix and fuse time series features, the mean prediction branch layer is used to output the battery capacity prediction value, the standard deviation prediction branch layer is used to output the standard deviation, and the output layer is used to output the joint distribution parameters.
[0028] In S4, the hyperparameters of the Gaussian regression network model are optimized by the Bayesian optimization algorithm, wherein the hyperparameters include the learning rate, batch size, number of training rounds, number of hidden layer neurons, regularization parameter, optimizer, and weight of the loss function.
[0029] The specific steps for optimizing hyperparameters include: a. Define the hyperparameter space; In this embodiment, the specific range is the learning rate 10 -5 ~10 -2 , batch size is 32, 64, 128, 256, number of training rounds is 50-200, number of hidden layer neurons is 32-256, regularization parameter is 10-6 ~10 -2 ,The optimizer uses Adam and SGD, and the weight of the loss function is 0.1 to 0.9.
[0030] b. Initial sampling generates initial hyperparameter combinations and evaluates the initial hyperparameter combinations; Specifically, Latin hypercube sampling is used to generate the initial hyperparameter combinations; c. Build a random forest proxy model; d. Determine the next sampling point (i.e., the new hyperparameter combination) through the sampling function, train the Gaussian network regression model with the newly determined hyperparameter combination, evaluate the model using evaluation metrics, and update the random forest model based on the evaluation results. e. Repeat step d until the number of iterations meets the preset number or the hyperparameter optimization converges; the number of iterations is set to 150 to 200 times, and the number of iterations in this example is set to 180 times.
[0031] f. Output the optimal hyperparameter combination.
[0032] In S4, the Gaussian network regression model is evaluated using the following evaluation metrics:
[0033] Where n is the number of test samples, is the actual battery capacity of the i-th sample, The battery capacity predicted for the i-th sample.
[0034] like Figure 3 The distribution of prediction results of the training set of this embodiment. In the figure, train_real represents the real label data of the capacity, train_gpr represents the model prediction data, and the gray represents the 95% confidence interval, that is, the interval <3σ; Figure 4 This is the distribution diagram of the error range of the prediction results of the training set. The mean absolute error MAPE of the training set is 1.65%.
[0035] Figure 5 This is the distribution diagram of the prediction results of the test set of this invention. train_real represents the real label data of the capacity, train_gpr represents the model prediction data, and the gray represents the 95% confidence interval. Figure 6 This is the distribution diagram of the prediction result error range of the test set of the present invention. The mean absolute error MAPE of the test set is 1.72%.
[0036] In this embodiment, the MAPE of the training set and the prediction set are both less than 2%, and the model prediction accuracy is high, which meets industrial requirements.
[0037] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme are not described in detail here. Ordinary technicians in the field are aware of all common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the field can improve and implement this scheme in combination with their own abilities under the inspiration given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A battery system life prediction method based on a pre-training architecture, characterized in that: The following steps are involved: S1: Acquire multi-dimensional raw time series data, perform data preprocessing, and determine an unlabeled training set, a labeled training set, and a test set. The preprocessing includes deleting abnormal characters and invalid data and identifying charge and discharge data with a complete charge and discharge cycle according to data extraction rules; the data extraction rules are that the voltage-current correlation is greater than 0.8 and the SOC difference is greater than 30%; S2: Split the preprocessed data into time windows and expand them into multiple dimensions through one-dimensional convolution to form an input matrix; S3: Build a Transformer pre-training model, input the input matrix of the unlabeled training set into the pre-training model for pre-training, and output representation features. The Transformer pre-training model is an encoder-only structure. S4: Construct a Gaussian regression network model, input the input matrix and representation features of the label training set into the Gaussian regression network for training, input the validation set into the trained Gaussian regression network model to output the battery capacity prediction value and standard deviation, and evaluate the model.
2. The battery system life prediction method based on a pre-training architecture according to claim 1, characterized in that: In S3, the Transformer prediction model includes an input embedding layer, a masked multi-head self-attention layer, a point-by-point feedforward network layer, a normalization and residual connection layer, and an output prediction head layer.
3. The battery system life prediction method based on a pre-training architecture according to claim 1, characterized in that: In S4, the Gaussian regression network model includes an input layer, a hidden layer, a mean prediction branch layer, a standard deviation prediction branch layer, and an output layer. The input layer is used to receive representation features, the hidden layer is used to flatten the input matrix and fuse time series features, the mean prediction branch layer is used to output the battery capacity prediction value, the standard deviation prediction branch layer is used to output the standard deviation, and the output layer is used to output the joint distribution parameters.
4. The battery system life prediction method based on a pre-training architecture according to claim 3 is characterized in that: In S4, the hyperparameters of the Gaussian regression network model are optimized by the Bayesian optimization algorithm, wherein the hyperparameters include the learning rate, batch size, number of training rounds, number of hidden layer neurons, regularization parameter, optimizer, and weight of the loss function.
5. The battery system life prediction method based on pre-training architecture according to claim 4 is characterized in that: The specific steps for optimizing hyperparameters include: a. Define the hyperparameter space; b. Initial sampling generates initial hyperparameter combinations and evaluates the initial hyperparameter combinations; c. Build a random forest proxy model; d. Determine the next sampling point using the sampling function, train the Gaussian network regression model using the determined hyperparameter combination, evaluate the model, and update the random forest model based on the evaluation results; e. Repeat step d until the number of iterations meets the preset number or the hyperparameter optimization converges; f. Output the optimal hyperparameter combination.
6. The battery system life prediction method based on pre-training architecture according to claim 1, characterized in that: In S3, the constraint function of the Transformer pre-training model uses MSELoss, and the loss function expression used is as follows: Where, is the true label, is the predicted probability distribution.
7. The battery system life prediction method based on pre-training architecture according to claim 1, characterized in that: In S4, the loss function of the Gaussian regression network model is expressed as follows: Where y is the actual battery capacity, μ and σ are the predicted mean and standard deviation, respectively, and C is a constant term.
8. The battery system life prediction method based on pre-training architecture according to claim 1, characterized in that: In S1, the multi-dimensional raw time series data includes voltage, current, mileage, timestamp, and battery SOC value.