Early remaining service life prediction method for commercial lithium ion battery

By constructing a teacher model for feature-time step screening and a student model based on deep separable convolution, and using discriminator hierarchical decomposition and adversarial generator network for feature distillation, the problems of insufficient data and insufficient learning ability in the early residual service life prediction of lithium-ion batteries are solved, and high-precision and low-cost prediction effects are achieved.

CN120044391APending Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510116632.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art faces insufficient data, ineffective redundant information transmission, and insufficient learning ability when predicting the early residual service life of lithium-ion batteries.

Method used

The Mamba model used by feature-time step filtering is used as the teacher model, combining dynamic exit and time step reduction techniques to convert time series features into image features through affine transformation. The student model uses deep separable convolutions and structurally reparameterized branch convolutions based on residual links, and performs feature distillation through discriminator hierarchical decomposition and adversarial generator networks to reduce computational costs and improve prediction accuracy.

Benefits of technology

It effectively improves the accuracy of early life prediction of lithium-ion batteries, reduces calculation costs, and ensures that student models can efficiently learn and imitate the knowledge of teacher models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120044391A_ABST
    Figure CN120044391A_ABST
Patent Text Reader

Abstract

The invention relates to an early remaining service life prediction method for a commercial lithium ion battery, which belongs to the technical field of battery energy, and comprises the following steps: firstly, carrying out affine transformation on a capacity voltage curve difference value of the first 100 cycles of each battery into pixel data as input data; and then calculating reconstructed features and soft labels by adopting a Mama model screened by feature-time steps and a depth separable convolution model combined with a re-parameterization method. Secondly, designing a feature generator for hierarchically decomposing the discriminator into a model; and finally, knowledge distillation loss is calculated based on the reconstructed features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of battery energy, and relates to a method for predicting the early remaining service life of commercial lithium-ion batteries. Background Art

[0002] Batteries are an important part of the clean energy transition. Electric vehicles, as well as the infrastructure of renewable energy systems and smart grids, require long battery life to achieve economic feasibility. Effectively predicting the remaining battery life is crucial for the reliable operation and safety maintenance of battery management systems.

[0003] Generally, RUL prediction methods include model-based methods, data-driven methods, and hybrid methods. Model-based solutions aim to explicitly model the relationship between sensed data and RUL. Data-driven methods aim to directly learn the relationship from a large amount of data without knowledge of the physical model of the system. Although there have been a large number of studies on the remaining battery life, accurately predicting battery life is not an easy task considering the following challenges.

[0004] (1) The number of early trainable samples is limited and relevant important features are difficult to capture: Deep learning models rely heavily on data quality and quantity.

[0005] (2) Redundant information in the teacher model is not learned by the student and a simple student network leads to insufficient learning ability of the student: In knowledge distillation, if the knowledge of the teacher model is not screened, it may cause the student model to learn noise or irrelevant information, thus affecting learning efficiency and the final effect.

[0006] (3) The teacher and student networks in knowledge distillation usually adopt the same or similar network structures: When the performance gap between the teacher and student models is too large, the student model may have difficulty improving from the teacher model.

[0007] Therefore, there is an urgent need to design a new knowledge distillation model for the early remaining service life of batteries to solve the above problems. Summary of the Invention

[0008] In view of this, the purpose of the present invention is to provide a method for predicting the early remaining service life of commercial lithium-ion batteries

[0009] To achieve the above purpose, the present invention provides the following technical solutions:

[0010] A method for predicting the early remaining service life of commercial lithium-ion batteries, comprising the following steps:

[0011] S1: Construct a data set;

[0012] S2: Preprocess the time series data in the dataset, and convert the time series features into image features through affine transformation;

[0013] S3: Construct a teacher model and a student model; the teacher model uses the Mamba model with feature-time step screening, introduces dynamic exit and time step reduction techniques, converts the activation function layer into a FAN network, and calculates the reconstructed features and soft labels of the teacher; the student model uses depthwise separable convolution based on residual links and branch convolution with structure reparameterization to calculate the reconstructed features and soft labels of the student;

[0014] S4: Hierarchically decompose the discriminator and put it into the feature generator of the student model to perform feature distillation on the features of the student and the teacher layer by layer;

[0015] S5: Use the preprocessed dataset to train the teacher model and the student model to obtain the final early remaining useful life prediction model;

[0016] S6: Use the early remaining useful life prediction model to predict the early remaining useful life of commercial lithium-ion batteries.

[0017] Furthermore, the dataset constructed in step S1 is specifically: Use multiple lithium-ion batteries that are cycled to failure under fast charging conditions, cycle with various fast charging curves and constant discharge curves, and collect battery numbers, cycle life numbers, detailed data during cycling, simplified battery life cycle data, voltage sampling points, and charging strategy data.

[0018] Furthermore, the preprocessing in step S2 specifically includes:

[0019] Use the ampere-hour integration method to evaluate the discharge capacity;

[0020] Construct a relationship between voltage and discharge capacity to obtain the discharge curve of a single battery at different cycles;

[0021] Use the difference curve between the voltage-capacity curves of different cycles as a feature, that is, calculate the difference curve by the following formula:

[0022]

[0023] where Dif qv is the difference curve, QV_Value 100 and QV_Value i are the voltage-capacity curves of the 100th cycle and the ith cycle;

[0024] For each voltage-capacity curve, an image of 100x100 pixels is first constructed, where each pixel value represents the value of a given feature at a specific cycle and voltage; the pixel values are obtained through an affine transformation of the feature values, with black representing the minimum feature value and white pixels representing the maximum feature value. Eventually, each battery obtains a different grayscale image.

[0025] Furthermore, when constructing the teacher model, a feature-time step screening strategy layer is constructed to adaptively eliminate fewer feature time steps, specifically including:

[0026] First, the feature map is put into the attention mechanism to calculate the importance score, obtaining the attention map; calculate the importance score of the i-th feature time step x i of the l-th layer attention matrix, and the formula is as follows:

[0027]

[0028] where, A l (x i , x j ) represents the value of the i-th row and j-th column of the l-th layer attention matrix, representing the degree of attention of other feature time steps x j to the feature time step x i ;

[0029] Then, sort the importance of the feature time steps according to the importance score, and use the top-K strategy for feature screening.

[0030] Furthermore, when constructing the student model, first divide the number of input channels into two equal groups, corresponding to convolutional kernels of 3×3 and 5×5 respectively. The output representation form of the depthwise separable convolution is:

[0031]

[0032] where, Conv dw and Conv pc are the depth convolution and pointwise convolution, and the four inputs of Conv dw are the number of channels, convolutional kernel, and padding size of i c and o c ;

[0033] Then, reparameterize the grouped depthwise separable convolution block based on the structural reparameterization method, including the following steps:

[0034] Add a residual link to each depthwise separable convolution block to aggregate shallow detail features and deep semantic information;

[0035] Integrate batch normalization BN into the convolution;

[0036] Aggregated multi-branch depth convolution;

[0037] Convert a convolution kernel of size m into a convolution kernel of size k by padding with zeros;

[0038] Integrate the multi-branch depth convolution into a 5-way convolution.

[0039] Furthermore, step S4 specifically includes the following steps:

[0040] The teacher model uses a pre-trained generative model to transfer the knowledge of features and attention, and retains the feature maps at each intermediate step in the generative model; extract the semantics represented by the attention map from the trained teacher model and require the student model to imitate it. The overall loss expression is:

[0041]

[0042] Among them, y represents the real data, and x represents any random data with the same structure as the real data;

[0043] Feature distillation is used to calculate the distance between the student model and the teacher model. The distance calculation formula is:

[0044]

[0045] Among them, F T and F S represent the feature maps of the teacher and student models, f(·) is an explicit mapping function, and n is the time dimension;

[0046] Minimize the above formula to make the student feature map and the teacher's feature map more similar;

[0047] Both the teacher and student models learn a generative model, and optimize the attention by using the intermediate features F and the number of cycles, expressed as:

[0048]

[0049] Specifically include:

[0050] First, freeze the student model and the attention module, put the student's features into the teacher's generator to obtain the reconstructed feature map, in order to narrow the gap between the teacher and the student. The process of feature reconstruction is expressed as:

[0051]

[0052] Among them, z t is the latent variable, and Φ is the learnable model;

[0053] Adopt the generative loss to optimize the feature reconstruction:

[0054]

[0055] Integrate the discriminator into the generator, and the learning process of updating the feature reconstruction loss is as follows:

[0056]

[0057] where D K represents the output of the k-th layer of the discriminator; during the training process of the generator model, the goal is to minimize the evidence lower bound loss

[0058]

[0059] where α is an empirical hyperparameter, z t is the latent variable, and the calculation formula is:

[0060]

[0061] where is the FAN network, θ i is a 3D convolution that connects features and attention by reducing the number of channels and adding parameters of the learnable model, and f z is the mean and variance function of the intermediate variable z to generate the latent variable;

[0062] Use the Decoder to generate the reconstructed feature map, using the latent variable Z and the attention λ, as follows:

[0063]

[0064] Finally, update the generator model in the student model using the above three losses, so that the generator model of the student model can learn the ability of the teacher generator model:

[0065]

[0066] Furthermore, convert the output values of the teacher model and the student model into soft labels through Softmax activation, and the calculation of the loss is divided into the following steps:

[0067] First, calculate the predicted probabilities y T , y S ;

[0068] Then, use the soft loss to define the difference between the teacher model prediction and the student model prediction:

[0069]

[0070] Next, use the hard loss to define the difference between the student model prediction and the true value label:

[0071]

[0072] Define a boundary loss:

[0073]

[0074] Finally, backpropagation is performed using the three losses to update the parameters of the student model, enabling the student model to learn the knowledge in the teacher model.

[0075] The beneficial effects of the present invention are as follows: This method can effectively improve the accuracy of early life prediction of lithium-ion batteries and reduce the computational cost.

[0076] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0078] Figure 1 is the knowledge distillation framework for discriminator-decomposed adversarial;

[0079] Figure 2 In (a), it is the discharge voltage-capacity curve of Battery No. 1 from the 1st cycle to the 100th cycle, and in (b), it is the grayscale difference image between the 100th cycle and the 2nd cycle;

[0080] Figure 3 is the structural diagram of the improved teacher model;

[0081] Figure 4 is the structural diagram of the improved student model;

[0082] Figure 5 is the feature distillation of discriminator decomposition. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0084] It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0085] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0086] Please refer to Figures 1 to 5 , the present invention provides a method for predicting the early remaining useful life of a battery based on a discriminator-decomposed adversarial knowledge distillation framework. First, the time series features are converted into image features through affine transformation. Then, a knowledge distillation life prediction framework is constructed. The improved Mamba model is used as the teacher model, and depthwise separable convolution is used as the student model. The discriminator in the adversarial generator network is incorporated into the feature distillation stage, and a bounded boundary loss is added to rewrite the objective function in the knowledge distillation stage, ultimately enabling the student to learn more refined knowledge. The specific steps of this method are as follows:

[0087] Step 1: The present invention uses the MIT dataset publicly disclosed by Severson. This dataset is based on 124 samples of APR18650M1A lithium phosphate batteries that are cycled to failure under 124 fast charging conditions. Each lithium-ion battery is cycled in a horizontal cylindrical fixture on a 48-channel Arbin LBT electrochemical workstation in a forced convection 30°C greenhouse. The dataset contains information such as battery number, cycle life number, detailed data during the cycle, simplified battery life cycle data, voltage sampling points, and charging strategies, and can be used for predicting the early remaining useful life of lithium-ion batteries.

[0088] Step 2: Preprocess the time series data in the dataset: Convert the time series features into image features through affine transformation to maintain the common relationship between data points. Usually, the voltage curve can reflect the degradation of the battery capacity. In this invention, the ampere-hour integration method is used to evaluate the discharge capacity. By constructing a relationship between voltage and discharge capacity, the discharge curve of a single battery sample at different cycles is obtained. As the number of cycles increases, the difference in discharge capacity between 3.6V and 2.0V gradually decreases, and the discharge voltage-capacity curve gradually moves to the left. To construct the correlation between features, time, and space, this invention uses the difference curve between the voltage-capacity curve cycles as a feature, that is, the difference curve is calculated by the following formula:

[0089]

[0090] where Dif qv is the difference curve, QV_Value 100 and QV_Value i are the voltage-capacity curves of the 100th cycle and the ith cycle.

[0091] For each voltage-capacity curve, an image of 100x100 pixels is first constructed, where each pixel value represents the value of a given feature at a specific cycle and voltage. The pixel value is obtained by affine transformation from the feature value, where black represents the minimum feature value and white pixels represent the maximum feature value. Finally, different grayscale images are obtained for each battery, which is convenient for using the teacher model and student model in the subsequent knowledge distillation framework to extract new features from the data. Example results are shown in Figure 2 (a) and (b) in

[0092] Step 3: Model construction: The teacher model uses the Mamba model with feature-time step screening to calculate the reconstructed features of the teacher and the soft labels of the teacher. The model structure diagram is shown in Figure 3 . In this model, dynamic exit is introduced, combined with the time step reduction technology to reduce the number of parameters, and the activation function layer is converted into a FAN network. The student model uses depthwise separable convolution and reparameterization method to calculate the reconstructed features of the student and the soft labels of the student. The model structure diagram is shown in Figure 4 . This model structure includes depthwise separable convolution based on residual links and branch convolution with structural reparameterization to further reduce the number of parameters.

[0093] Considering that not all voltage-capacity value pairs contribute equally to the remaining useful life, this invention constructs a feature-time step screening strategy layer to adaptively eliminate fewer feature time steps to reduce the redundant computational cost and make the knowledge of the teacher more refined.

[0094] First, put the feature map into the attention mechanism to calculate the importance score and obtain the attention map. The formula for calculating the importance score of the \(i\)-th feature time step \(x_i\) of the attention matrix in the \(l\)-th layer is as follows:

[0095]

[0096] Among them, \(A\) l (x i , x j ) represents the value of the \(i\)-th row and \(j\)-th column of the attention matrix in the \(l\)-th layer, representing the degree of attention of other feature time steps \(x\) j to the \(x\) of the feature time step i . Since the attention has positive and negative values, the absolute value is taken to reflect the more real attention correlation.

[0097] Then, sort the importance of the feature time steps according to the importance score, and use the top-K strategy for feature screening. In the present invention, a larger convolution kernel is used to obtain high-resolution detailed features, and a smaller convolution kernel captures low-resolution semantic information.

[0098] Depthwise separable convolutions are used in the student model. The grouped depthwise separable convolutions achieve a balance between accuracy and efficiency. First, the number of input channels is evenly divided into two groups, corresponding to convolution kernels of \(3\times3\) and \(5\times5\) respectively, to fuse multi-scale convolution information and mix different receptive fields, which is beneficial to feature fusion. The output representation form of the depthwise separable convolution is:

[0099]

[0100] Among them, Conv dw and Conv pc are the depth convolution and the pointwise convolution. The four inputs of Conv dw are the number of channels of \(i\) c and \(o\) c , the convolution kernel, and the padding size.

[0101] To improve the efficiency of the multi-branch structure adopted by the grouped depthwise separable convolution block when deployed in the deep model, the depth convolutions of the same layer are integrated based on the structure reparameterization method, and the architecture of the grouped depthwise separable convolution block is further simplified through equivalent replacement. The reparameterized grouped depthwise separable convolution block is established through two main steps. First, a residual connection is added to each block to effectively aggregate shallow detailed features and deep semantic information. Second, the effectiveness of the reparameterization depends on the equivalence of the parameter transformation in the method. First, batch normalization (BN) is integrated into the convolution. Then, the multi-branch depth convolutions are aggregated. The convolution kernel of size \(m\) is converted to the convolution kernel of size \(k\) by the method of supplementing 0 elements. The multi-branch depth convolutions are integrated into a 5-way convolution to save resources.

[0102] Step 4: In the present invention, the discriminator is hierarchically decomposed and placed in the feature generator of the student model, and the features of the student and the teacher are distilled layer by layer.

[0103] (Ⅰ) The teacher model uses a pre-trained generation model to transfer the knowledge of features and attention, and retains the feature maps of each intermediate step in the generation model. Extract the semantics represented by the attention map from the trained teacher model, and require the student model to imitate it. The overall loss expression is:

[0104]

[0105] where y represents the real data, and x represents any random data with the same structure as the real data. Feature distillation mainly calculates the distance between the student model and the teacher model, mainly as Figure 5 shown, and the distance calculation formula is:

[0106]

[0107] where F T and F S represent the feature maps of the teacher and student models, f(·) is an explicit mapping function, and n is the time dimension. Minimizing the above formula means making the student feature map and the teacher feature map more similar. To more accurately predict the total number of battery cycles, both the teacher and student models learn a generation model, and optimize the attention by using the intermediate feature F and the number of cycles. The specific expression is:

[0108]

[0109] Specifically, first freeze the student model and the attention module, put the features of the student into the generator of the teacher to obtain the reconstructed feature map, so as to narrow the gap between the teacher and the student. The process of feature reconstruction is expressed as:

[0110]

[0111] where z t is the latent variable, and Φ is the learnable model. During the training process, in order to let the student focus on the teacher's attention-based feature reconstruction, and not be affected by the differences in the features and attention maps extracted by different models, fit the distribution of the teacher model and improve the reconstruction quality of the student model, put the features and attention of the student into the generator of the teacher. The generation loss is used to optimize the feature reconstruction:

[0112]

[0113] To further match the reconstruction feature gap between the teacher model and the student model, the present invention integrates the discriminator into the generator, and the learning process of updating the feature reconstruction loss is as follows:

[0114]

[0115] where D K represents the output of the k-th layer of the discriminator. During the training process of the generator model, the goal is to minimize the evidence lower bound loss

[0116]

[0117] where α is an empirical hyperparameter, generally set to 0.1, and z t is a latent variable, and its calculation formula is:

[0118]

[0119] where is the FAN network, and θ i is a 3D convolution that connects features and attention by reducing the number of channels and adding learnable model parameters. f z is the mean and variance function of the intermediate variable z to generate the latent variable. Then, the Decoder is used to generate the reconstructed feature map, which utilizes the latent variable Z and the attention λ as follows:

[0120]

[0121] Finally, using the three losses mentioned above, the generator model in the student model is updated so that the generation model of the student model can learn the capabilities of the teacher generator model:

[0122]

[0123] (Ⅱ) Based on the JS divergence, the present invention converts the output values of the teacher model and the student model into soft labels through Softmax activation. The calculation of the loss is mainly divided into four steps: First, calculate the predicted probabilities y T , y S of the outputs of the teacher and student models. Then, use the soft loss to define the difference between the teacher model prediction and the student model prediction:

[0124]

[0125] Furthermore, use the hard loss to define the difference between the student model prediction and the true value label:

[0126]

[0127] To ensure the effectiveness of the students' predictions, i.e., the prediction error of the subsequent time is less than that of the previous time, a boundary loss is defined as follows:

[0128]

[0129] Finally, the above-mentioned three losses are used for backpropagation to update the parameters of the student model, enabling the student model to learn the knowledge in the teacher model and improve its prediction ability for the remaining useful life.

[0130] Step 5: To effectively verify and evaluate this method, the present invention uses the MIT dataset. This dataset is based on 124 samples of APR18650M1A lithium phosphate batteries that were cycled to failure under 124 fast charging conditions. Data such as battery number voltage interpolation, capacity interpolation, and total number of cycle periods are used. Each batch of data among the three batches of data is divided into a training set and a test set according to a ratio of 8:2. Server configuration information: i7-6800K CPU, 16.0GB RAM, and NVIDIA GeForce GTX 1080 graphics card. Regarding parameter settings, the activation function in the teacher model uses the activation function in the FAN network, the optimizer uses adam and the learning rate is initialized, and the loss function uses cross-entropy. The number of iteration rounds is set to 50, and the batchsize is set to 64. The specific test steps are as follows: After training is completed, the method of the sklearn library is used to evaluate the model performance using the test dataset to obtain the test loss and test accuracy.

[0131] In the above embodiment, a discriminator-decomposed adversarial knowledge distillation framework is proposed for transmitting knowledge between different network architectures for RUL prediction. The overall framework is as Figure 1 shown.

[0132] In the above embodiment, the reference in the specification to "this embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment are included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.

[0133] In the above embodiment, although the present invention has been described in connection with specific embodiments of the present invention, many substitutions, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the previous description. For example, other storage structures (e.g., dynamic RAM (DRAM)) can be used in the embodiments discussed. The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims.

[0134] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, any one of the methods in this embodiment is implemented.

[0135] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0136] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes any one of the methods in this embodiment.

[0137] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk or optical disc that can store program codes.

[0138] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store a computer program, the communication interface is used for communication, and the processor and the transceiver are used to run the computer program to make the electronic terminal execute each step of the above method.

[0139] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0140] The above-mentioned processor may be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it may also be a digital signal processor (Digital Signal Processing, abbreviated as DSP), an application specific integrated circuit (Application SpecificIntegrated Circuit, abbreviated as ASIC), a field programmable gate array (Field-Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0141] The present invention can be used in numerous general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0142] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for predicting the remaining useful life of commercial lithium-ion batteries in the early stage, characterized in that: The following steps are involved: S1: Construct dataset; S2: Preprocess the time series data in the dataset and convert the time series features into image features through affine transformation; S3: Construct a teacher model and a student model; the teacher model adopts the Mamba model of feature-time step screening, introduces dynamic exit and time step reduction technology, converts the activation function layer into a FAN network, and calculates the teacher's reconstruction features and the teacher's soft labels; the student model adopts residual link-based depth-separable convolution and structural reparameterized branched convolution to calculate the student's reconstruction features and the student's soft labels; S4: Decompose the discriminator hierarchically and put it into the feature generator of the student model, and perform feature distillation on the features of the student and teacher layer by layer; S5: Use the preprocessed data set to train the teacher model and the student model to obtain the final early remaining useful life prediction model; S6: Predicting the early remaining useful life of commercial lithium-ion batteries using the early remaining useful life prediction model.

2. The method for predicting the early remaining useful life of a commercial lithium-ion battery according to claim 1, characterized in that: The data set constructed in step S1 is specifically: using lithium-ion batteries that are cycled to failure under multiple fast charging conditions, cycling with various fast charging curves and constant discharge curves, collecting battery numbers, cycle life numbers, detailed cycle process data, simplified battery life cycle data, voltage sampling points, and charging strategy data.

3. The method for predicting the early remaining useful life of a commercial lithium-ion battery according to claim 1, characterized in that: The pre-processing in step S2 specifically includes: The discharge capacity was evaluated using the ampere-hour integration method; The relationship between voltage and discharge capacity is established to obtain the discharge curve of a single battery under different cycles; The difference curve between the voltage-capacity curve cycles is used as a feature, that is, the difference curve is obtained by calculating the following formula: Among them, Dif qv is the difference curve, QV_Value 100 and QV_Value i is the voltage-capacity curve of the 100th cycle and the i-th cycle; For each voltage-capacity curve, a 100x100 pixel image is first constructed, where each pixel value represents the value of a given feature at a specific cycle and voltage; the pixel value is obtained from the eigenvalue through an affine transformation, where black represents the smallest eigenvalue and white pixels represent the largest eigenvalue. Finally, each battery obtains a different grayscale image.

4. The method for predicting the early remaining useful life of a commercial lithium-ion battery according to claim 1, characterized in that: When building the teacher model, a feature-time step screening strategy layer is constructed to adaptively eliminate fewer feature time steps, including: First, put the feature map into the attention mechanism to calculate the importance score and obtain the attention map; calculate the i-th feature time step x of the l-th layer attention matrix i The importance score is as follows: Among them, A l (x i ,x j ) represents the value of the i-th row and j-th column of the attention matrix of the l-th layer, representing other feature time steps x j For the characteristic time step x i The degree of attention; Then, the importance of feature time steps is ranked according to the importance score, and the top-K strategy is used for feature screening.

5. The method for predicting the early remaining useful life of a commercial lithium-ion battery according to claim 1, characterized in that: When constructing the student model, the number of input channels is first divided into two groups, corresponding to the convolution kernels 3×3 and 5×5, respectively. The output representation of the depthwise separable convolution is: Among them, Conv dw and Conv pc It is depth convolution and point-by-point convolution, Conv dw The four inputs are i c and c The number of channels, convolution kernels, and padding size; Then the grouped depth-wise separable convolutional blocks are reparameterized based on the structure reparameterization method, including the following steps: Each depth-wise separable convolutional block adds a residual link to aggregate shallow detail features and deep semantic information; Integrate batch normalization BN into convolution; Aggregate multi-branch deep convolution; The convolution kernel of size m is converted to a convolution kernel of size k by adding 0 elements; Integrate multi-branch depthwise convolution into 5-way convolution.

6. The method for predicting the early remaining useful life of a commercial lithium-ion battery according to claim 1, characterized in that: Step S4 The specific steps include: The teacher model uses the pre-trained generative model to transfer the knowledge of features and attention, and retains the feature maps of each intermediate step in the generative model; the semantics represented by the attention map is extracted from the trained teacher model, and the student model is required to imitate it. The overall loss expression is: Among them, y represents real data, and x represents any random data with the same structure as the real data; Feature distillation is used to calculate the distance between the student model and the teacher model. The distance calculation formula is: Among them, F T and F S represents the feature graph of the teacher and student models, f(·) is the explicit mapping function, and n is the time dimension; Minimize the above formula to make the student feature map more similar to the teacher feature map; Both the teacher and student models learn a generative model by optimizing the attention using the intermediate features F and the number of cycles, expressed as: Specifically include: First, freeze the student model and attention module, put the student's features into the teacher's generator to obtain the reconstructed feature map to narrow the gap between the teacher and the student. The feature reconstruction process is expressed as: Among them, z t is a latent variable, Φ is a learnable model; Using Generative Loss To optimize feature reconstruction: Integrate the discriminator into the generator and update the learning process of feature reconstruction loss as follows: Among them, D K Represents the output of the kth layer of the discriminator; during the training of the generator model, the goal is to minimize the evidence lower bound loss Among them, α is an empirical hyperparameter, z t is a latent variable, and the calculation formula is: in, is the FAN network, θ i is a 3D convolution that connects features and attention by reducing the number of channels and adding parameters of a learnable model, f z is the mean and variance function of the intermediate variable z, which generates the latent variable; Use Decoder to generate the reconstructed feature map, using the latent variable Z and attention λ, as follows: Finally, the above three losses are used to update the generator model in the student model so that the generator model of the student model can learn the capabilities of the teacher generator model:

7. The method for predicting the early remaining useful life of a commercial lithium-ion battery according to claim 1, characterized in that: The output values ​​of the teacher model and the student model are converted into soft labels through Softmax activation, and the calculation of the loss is divided into the following steps: First, calculate the predicted probability y of the teacher and student model outputs T ,y S ; Then, a soft loss is used to define the difference between the teacher model prediction and the student model prediction: Next, a hard loss is used to define the difference between the student model prediction and the true value label: Define a margin loss: Finally, the three losses are used for back propagation to update the parameters of the student model so that the student model can learn the knowledge in the teacher model.

Citation Information

Patent Citations

  • A computer-implemented method for training a neural network and an electronic system

    CN113435568A

  • Lightweight Internet of Things malicious traffic identification method based on knowledge distillation space-time neural network

    CN116260642A

  • Internet of Things malicious software family classification method based on lightweight convolutional neural network and multi-teacher knowledge distillation

    CN116541837A

  • Model distillation real-time target detection method and device based on pseudo label filtering

    CN117011640A

  • Semantic relation preserving knowledge distillation for image-to-image translation

    EP4150528A1