Lithium battery health state and residual life combined prediction large model and prediction method

By generating dual-mode images and using a joint prediction model of Transformer block and attention mask strategy, the problems of insufficient data and noise impact in lithium battery health status and residual life prediction are solved, achieving higher accuracy and robust prediction.

CN120386998APending Publication Date: 2025-07-29GUANGDONG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510471380.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Existing methods for predicting health status and residual life of lithium batteries rely on a large amount of high-quality data, but in actual applications, insufficient data or a lot of noise, resulting in underfitting or overfitting the model, affecting the prediction accuracy and generalization ability.

Method used

Dual-mode images are generated using lithium battery charging and discharging data, and joint prediction of lithium battery health status and remaining life is performed through joint prediction models. Transformer block and attention mask strategy are used to enhance feature extraction and fusion, and the training process of the prediction model is optimized.

Benefits of technology

It significantly improves the accuracy, robustness and generalization ability of lithium battery health status prediction, and solves the problems of high data quality dependence, weak model generalization, and large long-term prediction errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386998A_ABST
    Figure CN120386998A_ABST
Patent Text Reader

Abstract

The invention discloses a lithium battery health state and residual life combined prediction large model and a prediction method, a bimodal image is generated through charging and discharging data of a lithium battery, and a lithium battery health state and residual life prediction result is output through a combined prediction model. In the training stage, an attention mask strategy is adopted, an original image and a mask image are predicted twice, key aging feature attention is dynamically enhanced, and noise is suppressed; the accuracy, robustness and generalization ability of lithium battery health state prediction are remarkably improved, and the problems that in the prior art, the data quality dependency degree is high, the model generalization is weak, and the long-term prediction error is large are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lithium batteries, and more specifically, to a large model for jointly predicting the state of health and remaining useful life of a lithium battery and a prediction method. Background Art

[0002] Lithium-ion batteries have become the core components of modern energy storage systems due to their high energy density, long cycle life, and low self-discharge characteristics. In a battery management system (BMS), the state of health (SOH) and remaining useful life (RUL), as key parameters characterizing battery aging, their accurate prediction is of great significance for ensuring the safe operation of the battery system. However, the non-linear aging characteristics presented by the battery during charge and discharge processes and the diversity of degradation modes under the action of multi-stress coupling make traditional prediction methods face severe challenges.

[0003] Current mainstream prediction methods can be divided into two categories: physics-based models and data-driven models. Physics-based models establish parameter mapping relationships by constructing electrochemical models or equivalent circuit models, but their modeling process is limited by the complete understanding of the internal reaction mechanism of the battery. In contrast, data-driven methods directly mine the aging characteristics in the operating data through machine learning algorithms, and among them, deep learning methods have attracted much attention due to their multi-modal data processing capabilities. Existing research has attempted to apply network structures such as CNN and LSTM to process time-series data and explore to alleviate the problem of insufficient data through transfer learning. However, practice has shown that such methods have defects such as the performance of the model highly depending on the quality of the health indicators extracted manually, the network architecture design needs to be customized for specific tasks, and transfer learning is easily affected by the data distribution differences between the source domain and the target domain, resulting in negative transfer phenomena; existing pre-trained models lack adaptive training for battery degradation characteristics. It is worth noting that large-scale vision models (LVMs) in the field of computer vision have demonstrated powerful feature extraction and cross-task generalization capabilities, but due to the one-dimensional time-series characteristics of battery data and the scale of public datasets, this technology has not been effectively applied to the field of battery health management. Although existing research has attempted to convert charge and discharge curves into two-dimensional images for feature extraction, there are still technical gaps in multi-modal image representation and degradation mode correlation analysis. Summary of the Invention

[0004] In order to overcome the defects of the prior art that rely on a large amount of high-quality battery aging data, but in actual applications, the data is insufficient or there is a lot of noise, resulting in model underfitting or overfitting, which affects the prediction accuracy, the present invention provides a large model for jointly predicting the state of health and remaining useful life of a lithium battery and a prediction method.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] The present invention provides a large model for jointly predicting the health state and remaining life of a lithium battery, including a data conversion unit, a feature extraction unit, a feature fusion unit, and a prediction unit connected in sequence;

[0007] The data conversion unit is used to convert the charge and discharge data of the lithium battery obtained into a bimodal image and transmit the bimodal image to the feature extraction unit;

[0008] The feature extraction unit is used to extract the bimodal features in the bimodal image and transmit the bimodal features to the feature fusion unit;

[0009] The feature fusion unit is used to fuse the bimodal features into deep fusion features and transmit the deep fusion features to the prediction unit;

[0010] The prediction unit outputs a joint prediction result of the health state and remaining life of the lithium battery according to the deep fusion features.

[0011] Preferably, the feature extraction unit includes a patch embedding layer, a first Transformer block, a second Transformer block, a third Transformer block, a fourth Transformer block, a fifth Transformer block, a sixth Transformer block, a seventh Transformer block, an eighth Transformer block, a ninth Transformer block, a tenth Transformer block, an eleventh Transformer block, and a twelfth Transformer block connected in sequence; and the fifth Transformer block, the seventh Transformer block, and the ninth Transformer block are fine-tuned;

[0012] The feature extraction unit outputs bimodal features to the feature fusion unit, and the bimodal features include first-modal features and second-modal features.

[0013] Preferably, the feature fusion unit includes a first branch subunit and a second branch subunit arranged in parallel; the first-modal features are correspondingly input into the first branch subunit, and the second-modal features are correspondingly input into the second branch subunit; both the first branch subunit and the second branch subunit include a multi-attention block and a feature aggregation block connected in sequence; the output end of the twelfth Transformer block is connected to the input ends of both the first branch subunit and the second branch subunit.

[0014] Preferably, the first branch sub-unit and the second branch sub-unit have the same structure. The multi-attention block of the first branch sub-unit includes a first normalization layer, a second normalization layer, a third normalization layer, a first self-attention layer, a first cross-attention layer, a first multi-layer perceptron, a first summation point, a second summation point, and a third summation point;

[0015] The second-modal features are correspondingly input into the cross-attention layer of the first branch sub-unit; the first-modal features are correspondingly input into the cross-attention layer of the second branch sub-unit;

[0016] The output end of the first normalization layer is connected to the input end of the first self-attention layer; the output end of the first self-attention layer is connected to the input end of the first summation point; the output end of the first summation point is connected to both the input ends of the second normalization layer and the second summation point; the output end of the second normalization layer is connected to the input end of the first cross-attention layer; the output end of the first cross-attention layer is connected to the input end of the second summation point; the output end of the second summation point is connected to both the input ends of the third normalization layer and the third summation point; the output end of the third normalization layer is connected to the input end of the first multi-layer perceptron; the output end of the first multi-layer perceptron is connected to the input end of the third summation point;

[0017] The output end of the third summation point is connected to the input end of the feature aggregation block;

[0018] Preferably, the feature aggregation block includes a first global average pooling layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a first activation function layer, a second activation function layer, a fourth summation point, a fifth summation point, a first multiplication point, and a second multiplication point;

[0019] The output end of the third summation point is connected to both the input ends of the fourth summation point and the first multiplication point; the output end of the multi-attention block of the second branch sub-unit is connected to both the input ends of the fourth summation point and the second multiplication point in the first branch sub-unit; the output end of the multi-attention block of the first branch sub-unit is connected to the feature aggregation block in the second branch sub-unit in the same way;

[0020] The output end of the fourth summation point is connected to the input end of the first global average pooling layer; the output end of the first global average pooling layer is connected to the input end of the first convolutional layer; the output end of the first convolutional layer is connected to the input end of the first activation function layer; the output end of the first activation function layer is connected to both the input ends of the second convolutional layer and the third convolutional layer; the output ends of the second convolutional layer and the third convolutional layer are both connected to the input end of the second activation function layer; the output end of the second activation function layer is connected to the input ends of the first multiplication point and the second multiplication point; the output ends of the first multiplication point and the second multiplication point are connected to the input end of the fifth summation point;

[0021] An output terminal of the fifth summing point is connected to an input terminal of the prediction unit.

[0022] The present invention also provides a prediction method based on the above-mentioned large model for joint prediction of lithium battery health status and remaining life, the method comprising:

[0023] Obtain the charge and discharge data of the lithium battery to be predicted;

[0024] Preprocessing the charge and discharge data of the lithium battery to be predicted to obtain the preprocessed charge and discharge data of the lithium battery;

[0025] The charge and discharge data of the lithium battery to be predicted after preprocessing is input into the joint prediction model to obtain the prediction results of the health status and remaining life of the lithium battery.

[0026] Preferably, the preprocessing of the charge and discharge data of the lithium battery to obtain the preprocessed charge and discharge data of the lithium battery includes:

[0027] Clean the charge and discharge data of lithium batteries and remove the charge and discharge data with abnormal life or temperature changes;

[0028] The charge and discharge data of the cleaned lithium battery are interpolated, low-pass filtered, and resampled at equal intervals to obtain the final resampling length;

[0029] The charge and discharge data of the lithium battery are smoothed according to the final resampling length to obtain the preprocessed charge and discharge data of the lithium battery.

[0030] Preferably, before using the joint prediction model, the joint prediction model should be trained, and an attention masking strategy is introduced during the training process. The attention masking strategy is:

[0031]

[0032] F′ CLS ~Categorical(F CLS ,Weights(F CLS ))

[0033] Weights(F CLS )=Norm(GAP(F CLS ))

[0034] Among them, Categorical(·) is the operation of sampling from the probability distribution; θ is the mask threshold; Weights(·) is the weighted operation; Norm(·) is the square root normalization operation.

[0035] Preferably, before using the joint prediction model, the joint prediction model should be trained, and the total loss function during training is:

[0036] J = L ori + αL mask

[0037] wherein, L ori is the first loss function; α is the first hyperparameter; L mask is the second loss function.

[0038] Preferably, the Huber loss is used to calculate the first loss function and the second loss function, and the calculation formula is:

[0039]

[0040] wherein, is the predicted value of the state of health and remaining life of the lithium battery; y n is the predicted value of the state of health and remaining life of the lithium battery; δ is the second hyperparameter.

[0041] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0042] The present invention proposes a large model and prediction method for jointly predicting the state of health and remaining life of a lithium battery. The dual-modal battery images generated from the charge and discharge data of the lithium battery are used to jointly predict the state of health and remaining life of the lithium battery through a joint prediction model, which significantly improves the accuracy, robustness and generalization ability of the lithium battery state of health prediction, and solves the problems of high dependence on data quality, weak model generalization, and large long-term prediction errors in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a schematic structural diagram of the large model for jointly predicting the state of health and remaining life of the lithium battery in Embodiment 1;

[0044] Figure 2 is a schematic structural diagram of the large model for jointly predicting the state of health and remaining life of the lithium battery in Embodiment 1;

[0045] Figure 3 is a schematic structural diagram of the feature fusion unit in Embodiment 1;

[0046] Figure 4 is a flowchart of the method for jointly predicting the state of health and remaining life of the lithium battery in Embodiment 2;

[0047] Figure 5 is a flowchart of the method for jointly predicting the state of health and remaining life of the lithium battery in Embodiment 3;

[0048] Figure 6 is a schematic structural diagram of the large model for jointly predicting the state of health and remaining life of the lithium battery in Embodiment 3;

[0049] Figure 7 It is a schematic structural diagram of the feature fusion unit described in Embodiment 3. Detailed implementation manners

[0050] The accompanying drawings are only for illustrative purposes and should not be construed as limitations on this patent;

[0051] To better illustrate this embodiment, some components in the accompanying drawings are omitted, enlarged or reduced, which do not represent the dimensions of the actual product;

[0052] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.

[0053] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0054] Embodiment 1

[0055] This embodiment provides a joint prediction large model for the health state and remaining life of a lithium battery, as Figure 1 shown, including a data conversion unit, a feature extraction unit, a feature fusion unit, and a prediction unit connected in sequence;

[0056] The data conversion unit is used to convert the charge and discharge data of the acquired lithium battery into a bimodal image and transmit the bimodal image to the feature extraction unit;

[0057] The feature extraction unit is used to extract the bimodal features in the bimodal image and transmit the bimodal features to the feature fusion unit;

[0058] The feature fusion unit is used to fuse the bimodal features into deep fusion features and transmit the deep fusion features to the prediction unit;

[0059] The prediction unit outputs a joint prediction result of the health state and remaining life of the lithium battery according to the deep fusion features.

[0060] As Figure 2As shown, the feature extraction unit includes a patch embedding layer, a first Transformer block, a second Transformer block, a third Transformer block, a fourth Transformer block, a fifth Transformer block, a sixth Transformer block, a seventh Transformer block, an eighth Transformer block, a ninth Transformer block, a tenth Transformer block, an eleventh Transformer block, and a twelfth Transformer block connected in sequence; and the fifth Transformer block, the seventh Transformer block, and the ninth Transformer block are fine-tuned;

[0061] The feature extraction unit outputs bimodal features to the feature fusion unit, and the bimodal features include first-modal features and second-modal features.

[0062] As Figure 3 shown, the feature fusion unit includes a first branch subunit and a second branch subunit arranged in parallel; the first-modal features are correspondingly input into the first branch subunit, and the second-modal features are correspondingly input into the second branch subunit; both the first branch subunit and the second branch subunit include a multi-attention block and a feature aggregation block connected in sequence; the output end of the twelfth Transformer block is connected to the input ends of both the first branch subunit and the second branch subunit.

[0063] The structures of the first branch subunit and the second branch subunit are the same. The multi-attention block of the first branch subunit includes a first normalization layer, a second normalization layer, a third normalization layer, a first self-attention layer, a first cross-attention layer, a first multi-layer perceptron, a first summation point, a second summation point, and a third summation point;

[0064] The second-modal features are correspondingly input into the cross-attention layer of the first branch subunit; the first-modal features are correspondingly input into the cross-attention layer of the second branch subunit;

[0065] The output end of the first normalization layer is connected to the input end of the first self-attention layer; the output end of the first self-attention layer is connected to the input end of the first summation point; the output end of the first summation point is connected to the input ends of both the second normalization layer and the second summation point; the output end of the second normalization layer is connected to the input end of the first cross-attention layer; the output end of the first cross-attention layer is connected to the input end of the second summation point; the output end of the second summation point is connected to the input ends of both the third normalization layer and the third summation point; the output end of the third normalization layer is connected to the input end of the first multi-layer perceptron; the output end of the first multi-layer perceptron is connected to the input end of the third summation point;

[0066] The output end of the third summing point is connected to the input end of the feature aggregation block;

[0067] The feature aggregation block includes a first global average pooling layer, a first convolution layer, a second convolution layer, a third convolution layer, a first activation function layer, a second activation function layer, a fourth summation point, a fifth summation point, a first multiplication point, and a second multiplication point;

[0068] The output end of the third summing point is connected to the input ends of the fourth summing point and the first multiplication point; the output end of the multi-attention block of the second branch sub-unit is connected to the input ends of the fourth summing point and the second multiplication point in the first branch sub-unit; the output end of the multi-attention block of the first branch sub-unit is connected to the feature aggregation block in the second branch sub-unit in the same way;

[0069] The output end of the fourth summing point is connected to the input end of the first global average pooling layer; the output end of the first global average pooling layer is connected to the input end of the first convolutional layer; the output end of the first convolutional layer is connected to the input end of the first activation function layer; the output end of the first activation function layer is connected to the input ends of the second convolutional layer and the third convolutional layer; the output ends of the second convolutional layer and the third convolutional layer are both connected to the input end of the second activation function layer; the output end of the second activation function layer is connected to the input ends of the first multiplication point and the second multiplication point; the output ends of the first multiplication point and the second multiplication point are connected to the input end of the fifth summing point;

[0070] An output terminal of the fifth summing point is connected to an input terminal of the prediction unit.

[0071] Example 2

[0072] This embodiment also provides a prediction method based on the large model for joint prediction of lithium battery health status and remaining life in embodiment 1, such as Figure 4 As shown, the method includes:

[0073] Obtain the charge and discharge data of the lithium battery to be predicted;

[0074] Preprocessing the charge and discharge data of the lithium battery to be predicted to obtain the preprocessed charge and discharge data of the lithium battery;

[0075] The charge and discharge data of the lithium battery to be predicted after preprocessing is input into the joint prediction model to obtain the prediction results of the health status and remaining life of the lithium battery.

[0076] The preprocessing of the charge and discharge data of the lithium battery to be predicted to obtain the preprocessed charge and discharge data of the lithium battery includes:

[0077] Clean the charge and discharge data of lithium batteries and remove the charge and discharge data with abnormal life or temperature changes;

[0078] Interpolate, perform low - pass filtering, and resample at equal intervals on the charge - discharge data of the lithium battery after cleaning to obtain the final resampled length;

[0079] Smooth the charge - discharge data of the lithium battery according to the final resampled length to obtain the charge - discharge data of the pre - processed lithium battery;

[0080] Before using the joint prediction model, the joint prediction model needs to be trained. The total loss function during training is:

[0081] J = L ori +αL nask

[0082] where L ori is the first loss function; α is the first hyper - parameter; L mask is the second loss function;

[0083] Calculate the first loss function and the second loss function using the Huber loss. The calculation formula is:

[0084]

[0085] where is the predicted value of the health state and remaining life of the lithium battery; y n is the predicted value of the health state and remaining life of the lithium battery; δ is the second hyper - parameter;

[0086] Before using the joint prediction model, the joint prediction model needs to be trained. During the training process, an attention mask strategy is introduced. The attention mask strategy is:

[0087]

[0088] F′ CLS ~Categorical(F CLS ,Weights(F CLS ))

[0089] Weights(F CLS )=Norm(GAP(F CLS ))

[0090] where Categorical(·) is the operation of sampling from a probability distribution; θ is the mask threshold; Weights(·) is the weighting operation; Norm(·) is the square - root normalization operation.<>

[0091] Example 3

[0092] This example provides a method for jointly predicting the health state and remaining life of a lithium battery, asFigure 5 As shown in, including:

[0093] Data preprocessing

[0094] Data such as voltage, current, temperature, and incremental capacity are widely used and easily obtained in BMS devices. Therefore, the present invention constructs a bimodal image based on the above data for subsequent joint prediction of SOH and RUL. The specific details are as follows:

[0095] 1) Clean the collected charge and discharge data. For batteries with significantly longer or shorter lifetimes compared to other batteries in the same batch, and batteries with extremely weak or abnormal temperature changes.

[0096] 2) Resample the cleaned data. First, perform interpolation by inserting a set number of zero values between each data point. Then, perform low-pass filtering on the interpolated data to remove the high-frequency components introduced by interpolation. Finally, perform equally spaced decimation on the filtered signal to obtain the final resampled length.

[0097] 3) Smooth the resampled data. Set the window size and use the moving average method to take the arithmetic mean of the data within the window.

[0098]

[0099] Among them, N is the window size, and X smooth is the smoothed data.

[0100] 4) Generate a curve graph and GAF as bimodal images. First, the image can be divided into 4 regions, with current, incremental capacity, temperature, and voltage occupying the upper left, upper right, lower left, and lower right regions respectively. The data in each region of the curve graph consists of the first cycle and the last cycle in L historical cycles, defined as I1. GAF then uses all input cycles to generate the Gram and angular field, and the formula is:

[0101]

[0102] Among them, is the result of normalizing the smoothed data.

[0103]

[0104] I2[i,j] = cos(φ[i] + φ[j]) (4)

[0105] Among them, I2[i,j] is the result of the pixel point corresponding to the i-th row and j-th column in the GAF image.

[0106] 5) Generate training set labels. The label of SOH consists of the SOH corresponding to L historical cycles and the SOH of the next cycle. The RUL label consists of the RUL of the last cycle among the L historical cycles and the RUL of the next cycle.

[0107] Joint prediction

[0108] The joint prediction model BLSCN is as Figure 6 shown, including feature extraction, feature fusion and prediction stages, and introducing an attention mask strategy in the training stage to enhance the ability of LVM to extract battery aging features.

[0109] Feature extraction

[0110] LVM utilizes the pre-trained DINOv2, which is an image encoder for self-supervised pre-training and fine-tuning on a large and lean dataset, and can effectively extract battery aging features. Therefore, DINOv2 is selected as the pre-trained LVM for feature extraction in BLSCN, which consists of a patch embedding layer and 12 transformer blocks. Specifically, the patch embedding consists of a CNN, and each transformer block consists of multiple layer normalizations (LN), a self-attention, and a multi-layer perceptron (MLP).

[0111] The attention mechanism used in DINOv2 is multi-head attention (MHA), and the formula is:

[0112] MHA(Q,K,V) = Concat(head1,…,head h )W o (5)

[0113] head i = Attention(Q i ,K i ,V i ) (6)

[0114]

[0115] Q = X1W1, K = X K W K , V = X V W V (8)

[0116] Among them, MHA(·) and Attention(·) represent the MHA operation and the attention operation respectively. Softmax(·) is the softmax function. represents transpose. X, h, and d represent the input data, the number of attention heads, and the channel dimension of the input data respectively. When X Q , XK and X V When they are equal, MHA is self-attention. Otherwise, MHA is cross-attention.

[0117] It is assumed that the input image F1 is first segmented into multiple 14×14 patches through patch embedding, and then processed by 12 transformer blocks in sequence to extract the corresponding modal features F3. The formula is:

[0118] F3 = Transformer 12 (…Transformer1(PE(F1))…) (9)

[0119] Transformer1(PE(F1)) = Z1 + MLP(LN(Z1)) (10)

[0120] Z1 = PE(F1) + MHA(LN(PE(F1)), LN(PE(F1)), LN(PE(F1))) (11)

[0121] MLP(LN(Z1)) = Linear(GELU(Linear(LN(Z1)))) (12)

[0122] PE(F1) = Conv2d(F1) (13)

[0123] Among them, Transformer(·), PE(·), MLP(·), Linear(·), GELU(·) and LN(·) represent transformer block operation, patch embedding, MLP operation, linear operation, Gaussian error linear unit operation and LN operation respectively. Conv2d(·) is a two-dimensional convolution operation of 14×14. Similarly, another modal feature F4 is obtained from F2.

[0124] Feature Fusion and Prediction

[0125] The bimodal image features extracted from the pre-trained DINOv2 are sequentially interactively fused and predicted through the cross-fusion module (CFM) and the fully connected (FC) layer in the small model (SM). To better adapt to the SOH and RUL prediction tasks, SM uses two structurally identical but independent branches for F3 and F4. It should be noted that CFM also has two branches, and the bimodal image features between them are fused with each other, which can well characterize the battery life. One of the branches is as Figure 7 shown.

[0126] As Figure 7 shown, a type of bimodal image feature is sequentially passed through residual self-attention, residual cross-attention and residual MLP to obtain the shallow fusion feature F5. The formula is:

[0127] F5 = Z2 + MLP(LN(Z2)) (14)

[0128] Z2 = Z1 + MHA2(LN(Z1), F4, F4) (15)

[0129] Z1 = F3 + MHA1(LN(F3), LN(F3), LN(F3)) (16)

[0130] It should be noted that the bimodal feature interaction occurs in the residual cross - attention, as shown in formula (15). Due to the residual structure, this shallow fusion can be stacked N times (usually set to 3).

[0131] Subsequently, the shallow - fused feature F5 passes through global average pooling (GAP), point convolution, softmax, weighted operation, and pixel - by - pixel addition in the feature aggregation block (FAB) in sequence to obtain the deep - fused feature F7. The formula is as follows:

[0132] F7 = F5W1 + F6W2 (17)

[0133] W1, W2 = Softmax(PC2(Z1)), Softmax(PC3(Z1)) (18)

[0134] Z1 = GELU(PC1(GAP(F5 + F6))) (19)

[0135] Where PC(·) and GAP(·) represent point convolution operation and GAP operation respectively. It should be noted that the bimodal features interact further at the beginning and end of the FAB, as shown in formulas (17) and (19) respectively. The feature map size of PC1(·) is Similarly, another deep - fused feature F8 is obtained from F4.

[0136] Finally, the deep - fused features F7 and F8 pass through an FC layer respectively to obtain the joint prediction results of SOH and RUL.

[0137] Attention mask strategy

[0138] During the training process, the original bimodal image obtains a classification token (i.e., the first token) with size through DINOv2, where h represents the number of attention heads in DINOv2 (usually set to 6). Then, the H ′ and W ′ dimensions of the classification token are adjusted to H and W through bilinear interpolation to obtain the attention map To enhance robustness, the weights of h attention maps are used as a probability distribution for randomly sampling an attention map. And further threshold it to obtain the corresponding attention mask. The formula is:

[0139]

[0140] F′ CLS ~Categorical(F CLS ,Weights(F CLS )) (21)

[0141] Weights(F CLS )=Norm(GAP(F CLS )) (22)

[0142] Among them, Categorical(·) represents the operation of sampling from a probability distribution, θ represents the mask threshold, which is generally set to 0.8 times the maximum value in F′ CLS . Weights(·) and Norm(·) represent the weighting operation and the square root normalization operation respectively.

[0143] Loss function

[0144] Since BLSCN needs to learn two tasks simultaneously, its loss function is defined as the sum of the losses of the two tasks. For the loss J of each task, since BLSCN makes two predictions under the attention mask strategy, the prediction losses caused by the original image and the masked image need to be considered. The formula is:

[0145] J=L ori +αL mask (23)

[0146] Among them, L ori and L mask represent the losses of the original image and the masked image respectively. ɑ is a hyperparameter, which is generally set to 0.6.

[0147] L ori and L mask are calculated using the Huber loss. The formula is:

[0148]

[0149] Among them, and y n represent the predicted values and measured values of n samples in a batch respectively. δ is a hyperparameter, which is generally set to 0.01.

[0150] Components with the same or similar labels correspond to the same or similar parts;

[0151] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0152] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A large model for jointly predicting the health state and remaining life of a lithium battery, characterized in that, It includes a data conversion unit, a feature extraction unit, a feature fusion unit, and a prediction unit connected in sequence; The data conversion unit is used to convert the charge and discharge data of the lithium battery into a bimodal image and transmit the bimodal image to the feature extraction unit; The feature extraction unit is used to extract the bimodal features in the bimodal image and transmit the bimodal features to the feature fusion unit; The feature fusion unit is used to fuse the bimodal features into deep fusion features and transmit the deep fusion features to the prediction unit; The prediction unit outputs a joint prediction result of the health state and remaining life of the lithium battery according to the deep fusion features.

2. The method for jointly predicting the health state and remaining life of a lithium battery according to claim 1, wherein The feature extraction unit includes a patch embedding layer, a first Transformer block, a second Transformer block, a third Transformer block, a fourth Transformer block, a fifth Transformer block, a sixth Transformer block, a seventh Transformer block, an eighth Transformer block, a ninth Transformer block, a tenth Transformer block, an eleventh Transformer block, and a twelfth Transformer block connected in sequence; and fine-tune the fifth Transformer block, the seventh Transformer block, and the ninth Transformer block; The feature extraction unit outputs bimodal features to the feature fusion unit, and the bimodal features include first-modal features and second-modal features.

3. The method for jointly predicting the health state and remaining life of a lithium battery according to claim 2, wherein, The feature fusion unit includes a first branch subunit and a second branch subunit arranged in parallel; the first-modal features are correspondingly input into the first branch subunit, and the second-modal features are correspondingly input into the second branch subunit; both the first branch subunit and the second branch subunit include a multi-attention block and a feature aggregation block connected in sequence; The output end of the twelfth Transformer block is connected to the input ends of both the first branch subunit and the second branch subunit.

4. The method for jointly predicting the health state and remaining life of a lithium battery according to claim 3, wherein The structures of the first branch subunit and the second branch subunit are the same. The multi-attention block of the first branch subunit includes a first normalization layer, a second normalization layer, a third normalization layer, a first self-attention layer, a first cross-attention layer, a first multi-layer perceptron, a first summation point, a second summation point, and a third summation point; The second-modal features are correspondingly input into the cross-attention layer of the first branch subunit; The first-modal features are correspondingly input into the cross-attention layer of the second branch subunit; The output end of the first normalization layer is connected to the input end of the first self-attention layer; the output end of the first self-attention layer is connected to the input end of the first summation point; the output end of the first summation point is connected to the input ends of the second normalization layer and the second summation point; the output end of the second normalization layer is connected to the input end of the first cross-attention layer; the output end of the first cross-attention layer is connected to the input end of the second summation point; the output end of the second summation point is connected to the input ends of the third normalization layer and the third summation point; the output end of the third normalization layer is connected to the input end of the first multi-layer perceptron; the output end of the first multi-layer perceptron is connected to the input end of the third summation point; The output end of the third summation point is connected to the input end of the feature aggregation block.

5. The method for jointly predicting the health state and remaining life of a lithium battery according to claim 4, wherein The feature aggregation block includes a first global average pooling layer, a first convolutional layer, a second convolutional layer, a third convolutional layer, a first activation function layer, a second activation function layer, a fourth summation point, a fifth summation point, a first multiplication point, and a second multiplication point; The output end of the third summation point is connected to the input ends of the fourth summation point and the first multiplication point; the output end of the multi-attention block of the second branch sub-unit is connected to the input ends of the fourth summation point and the second multiplication point in the first branch sub-unit; the output end of the multi-attention block of the first branch sub-unit is connected to the feature aggregation block in the second branch sub-unit in the same way; The output end of the fourth summation point is connected to the input end of the first global average pooling layer; the output end of the first global average pooling layer is connected to the input end of the first convolutional layer; the output end of the first convolutional layer is connected to the input end of the first activation function layer; the output end of the first activation function layer is connected to the input ends of the second convolutional layer and the third convolutional layer; the output ends of the second convolutional layer and the third convolutional layer are connected to the input end of the second activation function layer; the output end of the second activation function layer is connected to the input ends of the first multiplication point and the second multiplication point; the output ends of the first multiplication point and the second multiplication point are connected to the input end of the fifth summation point; The output end of the fifth summation point is connected to the input end of the prediction unit.

6. A prediction method for a large model for jointly predicting the health state and remaining life of a lithium battery according to any one of claims 1-5, characterized in that, The method includes: Obtaining the charge and discharge data of the lithium battery to be predicted; Preprocessing the charge and discharge data of the lithium battery to be predicted to obtain the preprocessed charge and discharge data of the lithium battery; Inputting the preprocessed charge and discharge data of the lithium battery to be predicted into the joint prediction model to obtain the prediction results of the health state and remaining life of the lithium battery.

7. The method for jointly predicting the health state and remaining life of a lithium battery according to claim 6, wherein The preprocessing of the charge and discharge data of the lithium battery to obtain the preprocessed charge and discharge data of the lithium battery includes: Cleaning the charge and discharge data of the lithium battery to remove the charge and discharge data with abnormal life or abnormal temperature change; Interpolating, low-pass filtering, and equally spaced resampling the charge and discharge data of the lithium battery after cleaning to obtain the final resampling length; Smoothing the charge and discharge data of the lithium battery according to the final resampling length to obtain the preprocessed charge and discharge data of the lithium battery.

8. The method for jointly predicting the health state and remaining life of a lithium battery according to claim 6, wherein, Before using the joint prediction model, the joint prediction model needs to be trained, and an attention mask strategy is introduced during the training process. The attention mask strategy is: F′ CLS ~Categorical(F CLS ,Weights(F CLS )) Weights(F CLS ) = Norm(GAP(F CLS )) Among them, Categorical(·) is the operation of sampling from a probability distribution; θ is the mask threshold; Weights(·) is the weighting operation; Norm(·) is the square root normalization operation.

9. The method for jointly predicting the health state and remaining life of a lithium battery according to claim 6, wherein Before using the joint prediction model, the joint prediction model needs to be trained. The total loss function during training is: J = L ori + αL mask Among them, L ori is the first loss function; α is the first hyperparameter; L mask is the second loss function.

10. The method for jointly predicting the health state and remaining life of a lithium battery according to claim 6, characterized in that The Huber loss is used to calculate the first loss function and the second loss function. The calculation formula is: Among them, is the predicted value of the health state and remaining life of the lithium battery; y n is the predicted value of the health state and remaining life of the lithium battery; δ is the second hyperparameter.

Citation Information

Cited By

  • Power lithium battery multi-parameter automatic sorting and conveying system based on visual identification

    CN121266850A