Battery state prediction method, system and device and storage medium
By constructing a lightweight CNN+LSTM architecture and using knowledge distillation techniques, the problem of inaccurate battery state prediction in existing technologies is solved, achieving high-precision battery state prediction on low-power chips and enhancing the model's adaptability and generalization ability.
Patent Information
- Application Number
- CN202511152766.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies lack a lightweight, high-performance CNN+LSTM architecture combined with knowledge distillation for battery state prediction, resulting in inaccurate battery state prediction and making it difficult to deploy on low-power chips.
We employ a teacher model based on convolutional neural networks and long short-term memory networks, compress the model using knowledge distillation techniques, and construct a student model adapted for embedded deployment. This includes data acquisition, sliding window processing, time alignment, multi-scale patch attention modules, and comprehensive loss function optimization.
It achieves high-precision, low-latency, and low-power battery state prediction, enhances the model's adaptability and generalization ability, and improves the prediction accuracy of SOC and SOE.
Smart Images

Figure CN120971979A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence and battery management systems, and in particular to a battery state prediction method, system, device, and storage medium. Background Technology
[0002] Currently, batteries play a crucial role in new energy vehicles, energy storage systems, and other fields. Accurate prediction of battery state (such as SOC and SOE) is of great significance for system safety and energy efficiency optimization. CNN (Convolutional Neural Network) and LSTM (Long Short-Term Memory Network) have been widely used in battery state modeling tasks, effectively extracting local features and time-series features.
[0003] However, these deep models typically contain a large number of parameters, consuming significant computational resources and making them difficult to deploy on edge embedded devices such as Cortex-M7. Knowledge distillation, as a model compression method, can transfer knowledge from high-performance "large models" to "lightweight models," significantly reducing resource overhead while maintaining accuracy.
[0004] Current technologies still lack a complete process that uses a lightweight, high-performance CNN+LSTM architecture, combines it with knowledge distillation, and optimizes its deployment for low-power chips, resulting in inaccurate battery state prediction. Summary of the Invention
[0005] To more accurately predict battery state, this application provides a battery state prediction method, system, device, and storage medium.
[0006] Firstly, this application provides a battery state prediction method, which is based on convolutional neural networks and long short-term memory networks, and adopts the following technical solution: Collect battery data and process the battery data by constructing a sliding window to obtain battery training data samples; A teacher model is constructed based on the convolutional neural network and the long short-term memory network; The battery training data samples are input into the teacher model for training, and the distillation target data of the teacher model is recorded. A student model is constructed based on the convolutional neural network and the long short-term memory network. The battery training data samples and the distillation target data are input into the student model to perform distillation training to obtain a trained target student model. The target student model is converted into an embeddable student model, and the embeddable student model is deployed. When real-time battery data is received, battery status is predicted using the embedded student model.
[0007] Through the above technical solutions, this application constructs a high-performance teacher model based on a CNN+LSTM architecture and further improves the performance of the lightweight student model using knowledge distillation model compression technology. The final result is a high-performance student model suitable for embedded deployment, achieving a high-precision, low-latency, and low-power battery state prediction system.
[0008] In one specific implementation scheme, the process of collecting battery data and obtaining battery training data samples by constructing a sliding window includes: Based on the battery discharge test records of the battery's rated capacity and rated energy, several battery data are collected according to the preset sampling frequency. Several battery label data are calculated based on several battery data, the battery rated capacity and the battery rated energy, and the battery label data includes battery capacity percentage and battery energy percentage; A battery data set is obtained by organizing several battery data and several battery tag data corresponding to the battery data, and the battery data set is then normalized. A sliding window is constructed, the sliding window having a preset window length, and battery training data samples are obtained by collecting the battery data set according to the window length using the sliding window.
[0009] Through the above technical solution, the data acquisition and preprocessing stage constructs time series samples through a sliding window and performs normalization processing to provide high-quality input data for subsequent models.
[0010] In one specific implementation scheme, the acquisition of several battery data points according to a preset sampling frequency includes: Several first battery data are collected according to a preset sampling frequency. The first battery data includes battery voltage, battery current, charging control voltage, charging control current, discharging control current, and battery temperature data. Based on several first battery data, several second battery data are calculated, including battery voltage change rate, battery current change rate, and battery temperature change rate. Several sets of first battery data and several sets of second battery data are combined to obtain several sets of battery data.
[0011] By employing the above technical solutions, dynamic information is added through the calculation of the rate of change, enabling the model to capture dynamic changes in battery state, rather than just static values. The rate of change information helps the model better understand the instantaneous behavior of the battery, thereby improving the prediction accuracy of SOC and SOE. The rate of change feature allows the model to better adapt to different charging and discharging conditions and battery aging states, enhancing the model's adaptability.
[0012] In one specific implementation, the step of inputting the battery training data samples into the teacher model for training, and recording the distillation target data of the teacher model, includes: The first battery training data sample is obtained by temporally aligning several features of the battery data using the Reshape layer, Transpose layer, and several depthwise separable two-dimensional convolutional layers of the teacher model. The first battery training data sample is intervalized using the temporal multi-scale Patch attention module of the teacher model to obtain the teacher model output result of the temporal multi-scale Patch attention module. The teacher model output result includes teacher model output attention features and teacher model output attention weights. The teacher model output attention features are obtained by extracting temporal features from several LSTM layers of the teacher model output attention features. The weighted output features of the teacher model LSTM layer are obtained by weighting the output teacher model's attention weights; The predicted output value of the teacher model is obtained by processing the weighted LSTM layer output features of the teacher model using several fully connected layers and Sigmoid activation function layers. The teacher model is trained using the predicted values output by the teacher model and the battery label data. Record the distillation target data of the teacher model, which includes the predicted output value of the teacher model, the attention features output by the teacher model, the attention weights output by the teacher model, and the output features of the LSTM layer of the teacher model.
[0013] In one specific implementation, training the teacher model using the output predictions of the teacher model and the battery tag data includes: The teacher model is trained using its output predictions and the battery tag data. The loss of the teacher model is calculated using the mean squared error formula:
[0014] Where n is the number of predicted categories, This represents the i-th value of the battery tag data; Let be the i-th predicted output value of the teacher model.
[0015] Through the above technical solutions, the teacher model adopts a complex CNN+LSTM structure, using deep two-dimensional convolutions to perform fine-grained temporal alignment of the input, avoiding errors caused by temporal misalignment when directly modeling time series. A temporal multi-scale patch attention module is used to extract and fuse features at different temporal granularities, enhancing the model's ability to perceive features at multiple time scales and improving its adaptability to short-term abrupt changes and long-term trends. Multi-layer LSTMs are used to perform global temporal modeling of patch features, enabling the modeling of long-term dependencies. Attention is applied at the patch level to focus on key temporal intervals. This application enhances the model's generalization ability, prediction accuracy, and inference speed through temporal alignment, multi-scale patch feature extraction and fusion, and adaptive dynamic weight calculation.
[0016] In one specific implementation, the step of inputting the battery training data samples and the distillation target data of the teacher model into the student model for distillation training to obtain a trained target student model includes: The battery training data samples are input into the student model; The student model's Reshape layer, Transpose layer, and several depthwise separable 2D convolutional layers are used to perform time alignment on several features of the battery data to obtain a second battery training data sample. The second battery training data samples are intervalized using the temporal multi-scale Patch attention module of the student model to obtain the student model output result of the temporal multi-scale Patch attention module. The student model output result includes student model output features and student model output attention weights. The output features of the student model are obtained by extracting temporal features from the output features of the student model using several LSTM layers. The weighted output features of the student model's LSTM layer are obtained by weighting the student model's output attention weights; The student model output prediction value is obtained by processing the weighted LSTM layer output features of the student model using the fully connected layer and the Sigmoid activation function layer. The student model is trained by distillation using the output prediction value of the student model, the battery label data, and the distillation target data of the teacher model. After training, a trained target student model is obtained.
[0017] In one specific implementation, the distillation training of the student model using the output predicted value of the student model, the battery label data, and the distillation target data of the teacher model includes: The student model is trained by distillation using the output predictions of the student model, the battery label data, and the distillation target data of the teacher model. The hard-label loss of the student model is calculated using the mean squared error, and the calculation formula is as follows:
[0018] in, It is a hard-label loss, where n is the number of predicted categories. This represents the i-th value of the battery tag data; This is the i-th predicted output value of the student model; The soft-label loss of the student model is calculated using the mean squared error, and the calculation formula is as follows:
[0019] in, This is a soft-label loss, where n is the number of predicted categories. This is the i-th predicted output value of the student model; This is the i-th predicted output value of the teacher model; The attention weight distillation loss of the student model is calculated using the mean squared error formula, which is as follows:
[0020] in, This is the attention weight distillation loss; n is the number of patches; This is the i-th value in the attention weights of the student model; This represents the i-th value in the teacher's attention weights model. The intermediate feature distillation loss of the student model is calculated using cosine similarity loss, and the calculation formula is as follows:
[0021] in, This is the intermediate feature distillation loss; n is the number of feature layers distilled during training; The intermediate features of the i-th layer of the student model; The intermediate features of the i-th layer of the teacher model; This indicates that dimensionality matching is performed on the intermediate features of the teacher model to align with the intermediate features of the student model. This represents the modulo operation; the intermediate features are the output features of the temporal multi-scale Patch attention module and the output features of the LSTM layer; The total loss of the student model is calculated based on the hard label loss, the soft label loss, the attention weight distillation loss, and the intermediate feature distillation loss, using the following formula:
[0022] in, It is hard label loss; It is soft label loss; It is the attention weight distillation loss; This is an intermediate characteristic of distillation loss; These are weighting coefficients that regulate the contribution of each component.
[0023] Through the above technical solutions, the student model adopts a simplified CNN+LSTM structure, reducing the number of parameters and computational complexity. Using knowledge distillation, the student model can learn the "knowledge" of the teacher model, significantly reducing model size while maintaining high prediction accuracy. The design of the comprehensive loss function ensures that the student model can both learn real label information and mimic the output and intermediate layer features of the teacher model, thus maintaining good performance while being lightweight.
[0024] Secondly, this application provides a battery state prediction system. The method is based on convolutional neural networks and long short-term memory networks, and adopts the following technical solution: The system includes: The training data acquisition module is used to collect battery data and process the battery data by constructing a sliding window to obtain battery training data samples; A teacher model construction module is used to construct a teacher model based on the convolutional neural network and the long short-term memory network. The teacher model training module is used to input the battery training data samples into the teacher model for training and to record the distillation target data of the teacher model. The student model distillation training module is used to construct a student model based on the convolutional neural network and the long short-term memory network. The battery training data samples and the distillation target data of the teacher model are input into the student model for distillation training to obtain a trained target student model. The student model deployment module is used to convert the format of the target student model to obtain an embeddable student model and deploy the embeddable student model. The battery status prediction module is used to predict the battery status using the embedded student model when real-time battery data is received.
[0025] Thirdly, this application provides a computer device that adopts the following technical solution: it includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as described above for a battery state prediction method.
[0026] Fourthly, this application provides a computer-readable storage medium that stores a computer program capable of being loaded by a processor and executing the aforementioned battery state prediction method.
[0027] In summary, this application has the following beneficial technical effects: (1) This application constructs a high-performance teacher model based on the CNN+LSTM architecture and further improves the performance of the lightweight student model by using knowledge distillation model compression technology. Finally, a high-performance student model that can be adapted to embedded deployment is obtained, realizing a high-precision, low-latency, and low-power battery state prediction system.
[0028] (2) Through knowledge distillation, the student model can learn the "knowledge" of the teacher model, which can significantly reduce the model size while maintaining high prediction accuracy. The design of the comprehensive loss function ensures that the student model can learn real label information and imitate the output and intermediate layer features of the teacher model, thus maintaining good performance while being lightweight.
[0029] (3) The training data adds dynamic information, and the model can capture the dynamic changes in the battery state, rather than just static values; the rate of change information can help the model better understand the instantaneous behavior of the battery, thereby improving the prediction accuracy of SOC and SOE; the rate of change feature enables the model to better adapt to different charging and discharging conditions and battery aging states, enhancing the model's adaptability.
[0030] (4) This application enhances the generalization ability, prediction accuracy and inference speed of the model by using time alignment, multi-scale patch feature extraction and fusion and adaptive dynamic weight calculation when constructing the model. Attached Figure Description
[0031] Figure 1 This is a flowchart of an embodiment of this application.
[0032] Figure 2 This is a schematic diagram of the teacher model structure.
[0033] Figure 3 This is a schematic diagram of the student model structure.
[0034] Figure 4 This is a flowchart of the knowledge distillation training process.
[0035] Figure 5 This is a block diagram of an embedded data acquisition and control system based on STM32H7.
[0036] Figure 6 This is the logic control flowchart for battery state prediction.
[0037] Figure 7This is a structural block diagram of an embodiment of this application.
[0038] Figure labels: 701, Training data acquisition module; 702, Teacher model construction; 703, Teacher model training module; 704, Student model distillation training module; 705, Student model deployment module; 706, Battery state prediction module. Detailed Implementation
[0039] The following is in conjunction with the appendix Figures 1-7 This application will be described in further detail.
[0040] This application discloses a battery state prediction method, which is used to predict battery state more accurately.
[0041] Currently, batteries play a crucial role in new energy vehicles, energy storage systems, and other fields. Accurate prediction of battery state (such as SOC and SOE) is of great significance for system safety and energy efficiency optimization. CNN (Convolutional Neural Network) and LSTM (Long Short-Term Memory Network) have been widely used in battery state modeling tasks, effectively extracting local features and time-series features.
[0042] However, these deep models typically contain a large number of parameters, consuming significant computational resources and making them difficult to deploy on edge embedded devices such as Cortex-M7. Knowledge distillation, as a model compression method, can transfer knowledge from high-performance "large models" to "lightweight models," significantly reducing resource overhead while maintaining accuracy.
[0043] Current technologies still lack a complete process that uses a lightweight, high-performance CNN+LSTM architecture, combines it with knowledge distillation, and optimizes its deployment for low-power chips, resulting in inaccurate battery state prediction.
[0044] Therefore, this application proposes a battery state prediction method, which can be used to predict battery state more accurately.
[0045] like Figure 1 As shown, the method includes: S10: Collect battery data and process the battery data by constructing a sliding window to obtain battery training data samples.
[0046] Specifically, battery data is collected to train the model. The battery training data samples are obtained by processing the battery data by moving a window of a fixed length over the time series using the sliding window method.
[0047] S20, a teacher model is constructed based on convolutional neural networks and long short-term memory networks.
[0048] Specifically, a teacher model with a large number of parameters is built based on convolutional neural networks (CNN) and long short-term memory networks (LSTM) to improve the performance of lightweight student models.
[0049] S30: Input the battery training data samples into the teacher model for training and record the distillation target data of the teacher model.
[0050] Specifically, battery training data samples are input into the teacher model for training. The teacher model structure is as follows: Figure 2 As shown, the teacher model is trained using battery training data samples, and the distillation target data required to improve the lightweight student model is recorded. The distillation target data includes the output prediction value (soft label) of the teacher model, the output features of the LSTM layer, and the output features and attention weights of the temporal multi-scale Patch attention module.
[0051] S40 constructs a student model based on convolutional neural networks and long short-term memory networks. The battery training data samples and distillation target data are input into the student model for distillation training to obtain a trained target student model.
[0052] Specifically, a student model with a small number of parameters is built based on convolutional neural networks (CNN) and long short-term memory networks (LSTM). The student model structure is as follows: Figure 3 As shown, battery training data samples are input into the student model. The distillation loss function of the student model is calculated using the distillation target data from the teacher model. After training, the trained target student model is obtained, as shown below. Figure 4 This is a knowledge distillation training flowchart. First, a training set, namely battery training data samples, is constructed. Then, a complex teacher model with a large number of parameters is trained using the battery training data samples to obtain the distillation target data of the teacher model. Next, a student model with a small number of parameters is trained using the distillation target data and the battery training data samples. Finally, the student model is deployed to an embedded device.
[0053] S50, convert the target student model to an embeddable student model, and deploy the embeddable student model.
[0054] Specifically, the trained student model is converted into a format suitable for embedded environments, enabling model deployment and real-time inference on low-power embedded chips such as the Cortex-M7. The SavedModel format student model is converted to .tflite format using TensorFlow Lite Converter, and then the .tflite model is further compiled into an array (C array) that can embed C language code using xxd commands or Python scripts.
[0055] The S60, upon receiving real-time battery data, predicts battery status using an embeddable student model.
[0056] Specifically, the model deployment platform is based on the STM32H7 series microcontroller, featuring high-performance, low-power peripheral interfaces. The platform structure is as follows: Figure 5 As shown, the platform includes a power supply module, a current sampling module, a voltage sampling module, a temperature sampling module, a communication module, and Flash memory. Each sampling module is responsible for sampling the corresponding data, and the communication module is used to obtain control data during battery charging and discharging and to interact with external systems.
[0057] First, the input preprocessing, model invocation, and output postprocessing logic are encapsulated into modular C functions for easy integration and debugging.
[0058] Then, the inference engine and interface driver are deployed, which includes initializing the MicroInterpreter inference instance, configuring the memory allocator and tensor buffer, registering the required operators (DepthwithConv2D, ReLU, UnidirectionalSequenceLSTM, Reshape, etc.), and setting up input and output data channels.
[0059] Finally, real-time data acquisition is performed, and inference is executed by calling Invoke() through MicroInterpreter.
[0060] Logic control flowchart as follows Figure 6 As shown, after powering on via the power supply module, model initialization, peripheral initialization, and parameter initialization are performed. Specifically, this includes initializing the MicroInterpreter inference instance, configuring the memory allocator and tensor buffer, registering the required operators (DepthwithConv2D, ReLU, UnidirectionalSequenceLSTM, Reshape, etc.), and setting up input and output data channels.
[0061] Then, it determines whether the ADC data acquisition is complete. If the acquisition is complete, it performs real-time data acquisition, processes the data, and performs SOX prediction based on the deployed student model. If the acquisition is not complete, it returns to continue acquisition.
[0062] Next, the system determines whether the state is abnormal based on the prediction. If it is abnormal, an alarm is triggered and the abnormal information is recorded. If there is no abnormality, the system returns to continue calculating and predicting based on the newly collected data.
[0063] In one embodiment, to more accurately predict battery state, the step of collecting battery data and processing the battery data through a sliding window to obtain battery training data samples can be specifically performed as follows: First, based on the battery discharge test record of the battery's rated capacity and rated energy, a number of battery data are collected according to the preset sampling frequency. Specifically, under standard charge and discharge conditions, a complete discharge test is completed, and the capacity and energy are recorded as the rated capacity and rated energy. During the battery's cyclic charge and discharge, the sampling frequency is set to, for example, 1 set / second, to collect battery data. Then, based on several battery data, battery rated capacity, and battery rated energy, several battery tag data are calculated. The battery tag data includes battery capacity percentage and battery energy percentage. Specifically, SOC (State of Charge, i.e., battery capacity percentage) reflects the battery's remaining capacity, which is the ratio of the battery's current remaining charge to its rated charge. SOE (State of Energy, i.e., battery energy percentage) reflects the battery's remaining energy, which is the ratio of the battery's current remaining energy (releaseable energy) to its rated energy. Using the collected battery data, battery rated capacity, and battery rated energy, the battery capacity percentage and battery energy percentage are calculated to obtain several battery tag data corresponding to the battery data. Next, a set of battery data and corresponding battery tag data are organized to obtain a set of battery data. The set of battery data is then normalized. Specifically, the set of battery data and corresponding battery tag data are organized to obtain a set of battery data. The set of battery data is then normalized to the range of [0,1] by performing maximum and minimum value normalization. Finally, a sliding window is constructed, which includes a preset window length. The battery training data samples are obtained by collecting battery data according to the window length using the sliding window. Specifically, a sliding window with a window length of N is constructed, and battery training data samples are formed by collecting data according to the window length using the sliding window.
[0064] In the data acquisition and preprocessing stage, time series samples are constructed using a sliding window and normalized to provide high-quality input data for subsequent models.
[0065] In one embodiment, to more accurately predict battery status, the step of collecting several battery data points according to a preset sampling frequency can be specifically performed as follows: First, several first battery data are collected according to the preset sampling frequency. The first battery data includes battery voltage, battery current, charging control voltage, charging control current, discharging control current, and battery temperature data. Specifically, during the battery's cyclic charging and discharging, the sampling frequency is set to, for example, 1 set / second, to collect the first battery data. The types of first battery data include battery voltage, battery current, charging control voltage, charging control current, discharging control current, and battery temperature data. Then, based on several first battery data, several second battery data are calculated. The second battery data includes the battery voltage change rate, the battery current change rate, and the battery temperature change rate. Specifically, the battery voltage change rate, the battery current change rate, and the battery temperature change rate are calculated based on the first battery data. Next, several sets of first battery data and several sets of second battery data are merged to obtain several sets of battery data. Specifically, the first battery data and several sets of second battery data are merged to obtain several sets of battery data. That is, the dimension of each set of input data of the model is 9 (battery voltage, battery current, charging control voltage, charging control current, discharging control current, battery temperature, battery voltage change rate, battery current change rate, battery temperature change rate), while the dimension of output data is 2 (battery capacity percentage, battery energy percentage).
[0066] By calculating the rate of change, dynamic information is added, allowing the model to capture dynamic changes in battery state, rather than just static values. The rate of change information helps the model better understand the instantaneous behavior of the battery, thereby improving the prediction accuracy of SOC and SOE. The rate of change feature enables the model to better adapt to different charging and discharging conditions and battery aging states, enhancing the model's adaptability.
[0067] In one embodiment, to more accurately predict battery state, the step of inputting battery training data samples into the teacher model for training and recording the distillation target data of the teacher model can be specifically performed as follows: like Figure 2 The teacher model structure is explained in detail below: Input is the input tensor; Reshape() is the tensor shape reshaping layer; Transpose() is the tensor dimension exchange layer; DepthwiseConv2D() is the depthwise separable 2D convolutional layer; Conv2D() is the ordinary 2D convolutional layer; Dense() is the fully connected layer; Sigmoid() is the Sigmoid activation function layer; LSTM() is the Long Short-Term Memory network layer; Mul() is the element-wise multiplication operation; Concat() is the dimension-wise concatenation operation; Add() is the element-wise addition operation; Output is the output tensor.
[0068] The following will combine Figure 2 Explanation: First, the teacher model's Reshape layer, Transpose layer, and several depthwise separable 2D convolutional layers are used to temporally align several features of the battery data to obtain the first battery training data sample. Specifically, the battery training data sample is input into the teacher model, represented by Input, with a shape of N×C, where N represents the number of time steps and C represents the feature dimension of each time step. The input is adjusted to a 1×C×N shape through Reshape and Transpose operations. Depthwise Conv2D (kernel size 1×3, stride 1, padding='same') is used to independently perform temporal convolution on each channel to achieve fine-grained temporal alignment, and the output tensor remains 1×C×N. Finally, the tensor shape is restored to N×C through Reshape and Transpose operations. Then, the training data samples of the first battery are intervalized using the temporal multi-scale Patch attention module of the teacher model to obtain the output of the teacher model of the temporal multi-scale Patch attention module. The output of the teacher model includes the teacher model output attention features and the teacher model output attention weights. Specifically, the data is further processed using the temporal multi-scale Patch attention module. The output of this module includes: Patch Feature, with shape n×C', representing the features extracted after dividing into n patches; Attention Weight, with shape n, representing the attention weight assigned to each patch. The temporal multi-scale Patch attention module in the teacher model takes a time-aligned sequence tensor N×C as input and performs patch partitioning and feature extraction at three different temporal granularities (fine-grained, medium-grained, and coarse-grained).
[0069] For medium-grained branches (with a partition length of l), the steps are as follows: First, the input is reshaped as l×C×n, where n=N / l is the number of patches; Then, DepthwiseConv2D (kernel size 3×3) is used to extract features within each patch, with an output shape of l×C×n; Next, the Transpose and Reshape operations yield a tensor of shape n×(C×l), which serves as one of the inputs for the three granularity branches Concat (concatenate by dimension) and Add (add element by element).
[0070] For both fine-grained (l_short) and coarse-grained (l_long) branches, each branch contains the following steps: First, reshape the input as l_*×C×n_*, where n_*=N / l_*; Then, local patch features are extracted using DepthwiseConv2D, and the output is l_*×C×n_*; Next, the patch features are fused using regular Conv2D (input channels n_*, output channels n), and the number of output channels is unified, with the output shape being l_*×C×n; Then, the Transpose operation yields a tensor of shape n×C×l_*, and subsequent processing is divided into two paths: Path 1: After the Reshape operation, a tensor with shape n×(C×l_*) is obtained, which is used for Concat and concatenated with the other two branches to obtain the Patch Feature with shape n×C'; Path 2: Align the features to a medium-granularity length l using Conv2D, then reshape them to n×(C×l). This reshape is then used to add features to other branches, resulting in fused features of shape n×(C×l). These features are then fed into Dense and Sigmoid layers to obtain the attention weights for each patch, with shape n.
[0071] Next, the teacher model output attention features are extracted using several LSTM layers to obtain the output features of the teacher model LSTM layers. Specifically, the Patch Feature is modeled through a 3-layer LSTM network to output the LSTM Feature (with an n×C'' shape). Then, the weighted output features of the teacher model LSTM layer are obtained by weighting the output features of the teacher model LSTM layer using the attention weights of the output teacher model. Specifically, the LSTM features are multiplied with the attention weights patch by patch to achieve weighted fusion. Next, the weighted output features of the teacher model LSTM layer are processed using several fully connected layers and a Sigmoid activation function layer to obtain the predicted output value of the teacher model. Specifically, the weighted result is fed into two fully connected layers for mapping, and the predicted output is output through the Sigmoid activation function. The output has a shape of 2, which corresponds to the predicted values of SOC and SOE, respectively. Next, the teacher model is trained using the output predictions of the teacher model and the battery label data. Specifically, the teacher model is trained using the real labels, i.e., the battery label data, based on the output predictions of the teacher model to obtain high-precision prediction capabilities. The mean squared error (MSE) is used as the loss of the teacher model. Finally, the distillation target data of the teacher model is recorded. The distillation target data includes the teacher model's output prediction value, the teacher model's output attention features, the teacher model's output attention weights, and the teacher model's LSTM layer output features. Specifically, the distillation target data of the teacher model includes the teacher model's output prediction value (soft label), the teacher model's LSTM layer output features, and the teacher model's temporal multi-scale Patch attention module output features and output attention weights.
[0072] In one embodiment, to more accurately predict battery state, the step of training a teacher model using the output predictions of the teacher model and battery label data can be specifically performed as follows: The teacher model is trained using its output predictions and battery tag data. The loss of the teacher model is calculated using the mean squared error formula:
[0073] Where n is the number of predicted categories, This represents the i-th value of the battery tag data; Let be the i-th predicted output value of the teacher model.
[0074] Specifically, the teacher model is trained using real-label (battery label) data based on its output predictions to achieve high-precision prediction capabilities. The loss of the teacher model is the mean squared error (MSE), calculated as follows:
[0075] Where n is the number of predicted categories, This represents the i-th value of the battery tag data; Let be the i-th predicted output value of the teacher model.
[0076] The teacher model employs a complex CNN+LSTM structure, using deep 2D convolutions to perform fine-grained temporal alignment of the input, avoiding errors caused by temporal misalignment when directly modeling time series. A temporal multi-scale patch attention module extracts and fuses features at different temporal granularities, enhancing the model's ability to perceive features at multiple time scales and improving its adaptability to short-term abrupt changes and long-term trends. Multi-layer LSTMs are used for global temporal modeling of patch features, enabling the modeling of long-term dependencies. Attention is applied at the patch level to focus on key temporal intervals. This application enhances the model's generalization ability, prediction accuracy, and inference speed through temporal alignment, multi-scale patch feature extraction and fusion, and adaptive dynamic weight calculation.
[0077] In one embodiment, to more accurately predict battery state, the step of inputting battery training data samples and distillation target data from the teacher model into the student model for distillation training to obtain a trained target student model can be specifically performed as follows: like Figure 3 The student model structure is explained in detail below: Input is the input tensor; Reshape() is the tensor shape reshaping layer; Transpose() is the tensor dimension exchange layer; DepthwiseConv2D() is the depthwise separable 2D convolutional layer; Conv2D() is the ordinary 2D convolutional layer; Dense() is the fully connected layer; Sigmoid() is the Sigmoid activation function layer; LSTM() is the Long Short-Term Memory network layer; Mul() is the element-wise multiplication operation; Concat() is the dimension-wise concatenation operation; Add() is the element-wise addition operation; Output is the output tensor. The following will combine... Figure 3 Explanation: First, the battery training data samples are input into the student model. Specifically, it is represented by Input, which has an N×C shape, where N represents the number of time steps and C represents the feature dimension of each time step.
[0078] Then, using the student model's Reshape layer, Transpose layer, and several depthwise separable 2D convolutional layers, several features of the battery data are temporally aligned to obtain the second battery training data samples. Specifically, the input is adjusted to a 1×C×N shape through Reshape and Transpose operations; a depthwise separable 2D convolution (DepthwiseConv2D, kernel size 1×3, stride 1, padding='same') is used to independently perform temporal convolution on each channel to achieve fine-grained temporal alignment, keeping the output tensor at 1×C×N; and then the tensor shape is restored to N×C through Reshape and Transpose operations. Next, the temporal multi-scale Patch attention module of the student model is used to intervalize the training data samples of the second battery to obtain the output results of the student model of the temporal multi-scale Patch attention module. The output results of the student model include the output features of the student model and the output attention weights of the student model. Specifically, the temporal multi-scale Patch attention module is used to further process the data. The output of this module includes: Patch Feature, with shape n×C', representing the features extracted after dividing into n patches; Attention Weight, with shape n, representing the attention weight assigned to each patch. The temporal multi-scale Patch attention module in the teacher model takes a time-aligned sequence tensor N×C as input and performs patch partitioning and feature extraction at three different temporal granularities (fine-grained, medium-grained, and coarse-grained).
[0079] For medium-grained branches (with a partition length of l), the steps are as follows: First, the input is reshaped as l×C×n, where n=N / l is the number of patches; Then, DepthwiseConv2D (kernel size 3×3) is used to extract features within each patch, with an output shape of l×C×n; Next, the Transpose and Reshape operations yield a tensor of shape n×(C×l), which serves as one of the inputs for the three granularity branches Concat (concatenate by dimension) and Add (add element by element).
[0080] For both fine-grained (l_short) and coarse-grained (l_long) branches, each branch contains the following steps: First, reshape the input as l_*×C×n_*, where n_*=N / l_*; Then, local patch features are extracted using DepthwiseConv2D, and the output is l_*×C×n_*; Next, the patch features are fused using regular Conv2D (input channels n_*, output channels n), and the number of output channels is unified, with the output shape being l_*×C×n; Then, the Transpose operation yields a tensor of shape n×C×l_*, and subsequent processing is divided into two paths: Path 1: After the Reshape operation, a tensor with shape n×(C×l_*) is obtained, which is used for Concat and concatenated with the other two branches to obtain the Patch Feature with shape n×C'; Path 2: Align the features to a medium-granularity length l using Conv2D, then reshape them to n×(C×l). This reshape is then used to add features to other branches, resulting in fused features of shape n×(C×l). These features are then fed into Dense and Sigmoid layers to obtain the attention weights for each patch, with shape n.
[0081] Then, the output features of the student model are extracted by time series feature extraction using several LSTM layers of the student model. Specifically, the Patch Feature is modeled by two layers of LSTM network and outputs LSTMFeature (with shape n×C''). Next, the student model's LSTM layer output features are weighted using the student model's output attention weights to obtain the weighted student model LSTM layer output features. Specifically, the LSTM features are multiplied with the attention weights patch by patch to achieve weighted fusion. Then, the weighted output features of the LSTM layer of the student model are processed using the fully connected layer and the Sigmoid activation function layer of the student model to obtain the predicted output value of the student model. Specifically, the weighted result is fed into the fully connected layer for mapping; the predicted output is output through the Sigmoid activation function, with a shape of 2, corresponding to the predicted values of SOC and SOE respectively.
[0082] Finally, the student model is trained using the student model's output predictions, battery label data, and the teacher model's distillation target data. The trained target student model is then obtained. Specifically, the student model is trained using the student model's output predictions, battery label data, and the teacher model's distillation target data. In one embodiment, to more accurately predict battery state, the step of distilling the student model using the student model's output predictions, battery label data, and the teacher model's distillation target data can be specifically performed as follows: The student model is trained by distillation using the output predicted values of the student model, battery label data, and distillation target data of the teacher model. The hard-label loss of the student model is calculated using the mean squared error, and the formula is as follows:
[0083] in, It is a hard-label loss, where n is the number of predicted categories. This represents the i-th value of the battery tag data; This is the i-th predicted output value of the student model; Specifically, the hard label loss uses the mean squared error (MSE), calculated as follows:
[0084] in, It is a hard-label loss, where n is the number of predicted categories. This represents the i-th value of the battery tag data; This is the i-th predicted output value of the student model; The soft-label loss of the student model is calculated using the mean squared error, and the formula is as follows:
[0085] in, This is a soft-label loss, where n is the number of predicted categories. This is the i-th predicted output value of the student model; This is the i-th predicted output value of the teacher model; Specifically, the soft label loss also uses the mean squared error (MSE), and the calculation formula is as follows:
[0086] in, This is a soft-label loss, where n is the number of predicted categories. This is the i-th predicted output value of the student model; This is the i-th predicted output value of the teacher model; The attention weight distillation loss of the student model is calculated using the mean squared error formula, which is as follows:
[0087] in, This is the attention weight distillation loss; n is the number of patches; This is the i-th value in the attention weights of the student model; This represents the i-th value in the teacher's attention weights model. Specifically, the attention weight distillation loss also uses mean squared error (MSE), calculated as follows:
[0088] in, This is the attention weight distillation loss; n is the number of patches; This is the i-th value in the attention weights of the student model; This represents the i-th value in the teacher's attention weights model. The intermediate feature distillation loss of the student model is calculated using cosine similarity loss, and the formula is as follows:
[0089] in, This is the intermediate feature distillation loss; n is the number of feature layers distilled during training; The intermediate features of the i-th layer of the student model; The intermediate features of the i-th layer of the teacher model; This indicates that dimensionality matching is performed on the intermediate features of the teacher model to align with the intermediate features of the student model. This represents the modulo operation; the intermediate features are the output features of the temporal multi-scale Patch attention module and the output features of the LSTM layer; Specifically, the intermediate feature distillation loss uses cosine similarity loss, calculated as follows:
[0090] in, This is the intermediate feature distillation loss; n is the number of feature layers distilled during training; The intermediate features of the i-th layer of the student model; The intermediate features of the i-th layer of the teacher model; This indicates that dimensionality matching is performed on the intermediate features of the teacher model to align with the intermediate features of the student model. This represents the modulo operation; the intermediate features are the output features of the temporal multi-scale Patch attention module and the output features of the LSTM layer; The student model is trained under the guidance of the teacher's model. The complete training process is attached. Figure 4 As shown, the intermediate layer features used for distillation are the output features of the deep 2D convolutional layers of the teacher model and the student model, as well as the output features of the LSTM layer.
[0091] The total loss of the student model is calculated based on the hard label loss, soft label loss, attention weight distillation loss, and intermediate feature distillation loss, using the following formula:
[0092] in, It is hard label loss; It is soft label loss; It is the attention weight distillation loss; This is an intermediate characteristic of distillation loss; These are weighting coefficients that regulate the contribution of each component.
[0093] Specifically, the total loss of the student model is a weighted sum of hard label loss, soft label loss, attention weight distillation loss, and intermediate feature distillation loss, calculated as follows:
[0094] in, It is hard label loss; It is soft label loss; It is the attention weight distillation loss; This is an intermediate characteristic of distillation loss; These are weighting coefficients that regulate the contribution of each component.
[0095] The student model employs a simplified CNN+LSTM structure, reducing the number of parameters and computational complexity. Through knowledge distillation, the student model learns the "knowledge" of the teacher model, significantly reducing model size while maintaining high prediction accuracy. The design of the comprehensive loss function ensures that the student model can learn both real label information and mimic the output and intermediate layer features of the teacher model, thus maintaining good performance while remaining lightweight.
[0096] By using a teacher model to guide the training of a student model, the student model can learn richer knowledge representations by imitating the output and intermediate layer features of the teacher model, thus maintaining high prediction accuracy while being lightweight.
[0097] Based on the above method, this application also discloses a battery state prediction system. For example... Figure 7 The system includes the following modules: Training data module 701 is used to collect battery data and obtain battery training data samples by processing the battery data through a sliding window. Teacher model building module 702 is used to build a teacher model based on convolutional neural networks and long short-term memory networks; The teacher model training module 703 is used to input battery training data samples into the teacher model for training and to record the distillation target data of the teacher model. The student model distillation training module 704 is used to build a student model based on a convolutional neural network and a long short-term memory network. It inputs battery training data samples and distillation target data into the student model to perform distillation training and obtain a trained target student model. The student model deployment module 705 is used to convert the format of the target student model to obtain an embeddable student model and deploy the embeddable student model. The battery state prediction module 706 is used to predict the battery state using an embeddable student model when real-time battery data is received.
[0098] In one embodiment, the training data module 701 is specifically used to collect several battery data points based on the battery rated capacity and rated energy recorded in the battery discharge test, according to a preset sampling frequency; calculate several battery tag data points based on the several battery data points, the battery rated capacity, and the battery rated energy, the battery tag data points including battery capacity percentage and battery energy percentage; organize the several battery data points and the several battery tag data points corresponding to the several battery data points to obtain a battery data set, and normalize the battery data set; construct a sliding window, the sliding window having a preset window length, and use the sliding window to collect battery data set according to the window length to obtain battery training data samples.
[0099] In one embodiment, the training data module 701 is specifically used to collect a plurality of first battery data according to a preset sampling frequency. The first battery data includes battery voltage, battery current, charging control voltage, charging control current, discharging control current, and battery temperature data. Based on the plurality of first battery data, a plurality of second battery data is calculated. The second battery data includes battery voltage change rate, battery current change rate, and battery temperature change rate. The plurality of first battery data and the plurality of second battery data are combined to obtain a plurality of battery data.
[0100] In one embodiment, the teacher model training module 703 is specifically used to: align several features of the battery data in time using the teacher model's Reshape layer, Transpose layer, and several depthwise separable 2D convolutional layers to obtain a first battery training data sample; intervalize the first battery training data sample using the teacher model's temporal multi-scale Patch attention module to obtain the teacher model output of the temporal multi-scale Patch attention module, the teacher model output including teacher model output attention features and teacher model output attention weights; extract temporal features from the teacher model's output attention features using several LSTM layers to obtain the teacher model LSTM layer output features; weight the teacher model LSTM layer output features using the output teacher model attention weights to obtain weighted teacher model LSTM layer output features; process the weighted teacher model LSTM layer output features using several fully connected layers and a Sigmoid activation function layer to obtain the teacher model output predicted value; train the teacher model using the teacher model output predicted value and battery label data; and record the teacher model's distillation target data, the distillation target data including the teacher model output predicted value, teacher model output attention features, teacher model output attention weights, and teacher model LSTM layer output features.
[0101] In one embodiment, the teacher model training module 703 is specifically used to train the teacher model using the output prediction values of the teacher model and battery tag data, and to calculate the teacher model loss using the mean squared error formula:
[0102] Where n is the number of predicted categories, This represents the i-th value of the battery tag data; Let be the i-th predicted output value of the teacher model.
[0103] In one embodiment, the student model distillation training module 704 is specifically used to input battery training data samples into the student model; to perform time alignment on several features of the battery data using the student model's Reshape layer, Transpose layer, and several depthwise separable 2D convolutional layers to obtain second battery training data samples; to perform intervalization on the second battery training data samples using the student model's temporal multi-scale Patch attention module to obtain the student model output result of the temporal multi-scale Patch attention module, the student model output result including student model output features and student model output attention weights; to perform temporal feature extraction on the student model output features using several LSTM layers of the student model to obtain the student model LSTM layer output features; to weight the student model LSTM layer output features using the student model output attention weights to obtain weighted student model LSTM layer output features; to process the weighted student model LSTM layer output features using the student model's fully connected layer and Sigmoid activation function layer to obtain the student model output prediction value; and to perform distillation training on the student model using the student model's output prediction value, battery label data, and the teacher model's distillation target data, resulting in a trained target student model.
[0104] In one embodiment, the student model distillation training module 704 is specifically used to perform distillation training on the student model using the output prediction value of the student model, battery label data, and distillation target data of the teacher model. The hard-label loss of the student model is calculated using the mean squared error, and the formula is as follows:
[0105] in, It is a hard-label loss, where n is the number of predicted categories. This represents the i-th value of the battery tag data; This is the i-th predicted output value of the student model; The soft-label loss of the student model is calculated using the mean squared error, and the formula is as follows:
[0106] in, This is a soft-label loss, where n is the number of predicted categories. This is the i-th predicted output value of the student model; This is the i-th predicted output value of the teacher model; The attention weight distillation loss of the student model is calculated using the mean squared error formula, which is as follows:
[0107] in, This is the attention weight distillation loss; n is the number of patches; This is the i-th value in the attention weights of the student model; This represents the i-th value in the teacher's attention weights model. The intermediate feature distillation loss of the student model is calculated using cosine similarity loss, and the formula is as follows:
[0108] in, This is the intermediate feature distillation loss; n is the number of feature layers distilled during training; The intermediate features of the i-th layer of the student model; The intermediate features of the i-th layer of the teacher model; This indicates that dimensionality matching is performed on the intermediate features of the teacher model to align with the intermediate features of the student model. This represents the modulo operation; the intermediate features are the output features of the temporal multi-scale Patch attention module and the output features of the LSTM layer; The total loss of the student model is calculated based on the hard label loss, soft label loss, attention weight distillation loss, and intermediate feature distillation loss, using the following formula:
[0109] in, It is hard label loss; It is soft label loss; It is the attention weight distillation loss; This is an intermediate characteristic of distillation loss; These are weighting coefficients that regulate the contribution of each component.
[0110] This application also discloses a computer device.
[0111] Specifically, the computer device includes a memory and a processor, the memory storing a computer program that can be loaded by the processor and executed as described above for predicting battery state.
[0112] This application also discloses a computer-readable storage medium.
[0113] Specifically, the computer-readable storage medium stores a computer program that can be loaded by a processor and executed as described above for a battery state prediction method. The computer-readable storage medium includes, for example, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0114] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of the present invention.
Claims
1. A method for predicting battery state, characterized in that, The method is based on convolutional neural networks and long short-term memory networks, and the method includes: Collect battery data and process the battery data by constructing a sliding window to obtain battery training data samples; A teacher model is constructed based on the convolutional neural network and the long short-term memory network; The battery training data samples are input into the teacher model for training, and the distillation target data of the teacher model is recorded. A student model is constructed based on the convolutional neural network and the long short-term memory network. The battery training data samples and the distillation target data are input into the student model to perform distillation training to obtain a trained target student model. The target student model is converted into an embeddable student model, and the embeddable student model is deployed. When real-time battery data is received, battery status is predicted using the embedded student model.
2. The method according to claim 1, characterized in that, The process of collecting battery data and processing it through a sliding window to obtain battery training data samples includes: Based on the battery discharge test records of the battery's rated capacity and rated energy, several battery data are collected according to the preset sampling frequency. Several battery label data are calculated based on several battery data, the battery rated capacity and the battery rated energy, and the battery label data includes battery capacity percentage and battery energy percentage; A battery data set is obtained by organizing several battery data and several battery tag data corresponding to the battery data, and the battery data set is then normalized. A sliding window is constructed, the sliding window having a preset window length, and battery training data samples are obtained by collecting the battery data set according to the window length using the sliding window.
3. The method according to claim 2, characterized in that, The battery data collected according to the preset sampling frequency includes: Several first battery data are collected according to a preset sampling frequency. The first battery data includes battery voltage, battery current, charging control voltage, charging control current, discharging control current, and battery temperature data. Based on several first battery data, several second battery data are calculated, including battery voltage change rate, battery current change rate, and battery temperature change rate. Several sets of first battery data and several sets of second battery data are combined to obtain several sets of battery data.
4. The method according to claim 2, characterized in that, The step of inputting the battery training data samples into the teacher model for training, and recording the distillation target data of the teacher model, includes: The first battery training data sample is obtained by temporally aligning several features of the battery data using the Reshape layer, Transpose layer, and several depthwise separable two-dimensional convolutional layers of the teacher model. The first battery training data sample is intervalized using the temporal multi-scale Patch attention module of the teacher model to obtain a first output result, which includes the first output feature and the first output attention weight of the temporal multi-scale Patch attention module of the teacher model. The teacher model output attention features are obtained by extracting temporal features from several LSTM layers of the teacher model output attention features. The weighted output features of the teacher model LSTM layer are obtained by weighting the output teacher model's attention weights; The predicted output value of the teacher model is obtained by processing the weighted LSTM layer output features of the teacher model using several fully connected layers and Sigmoid activation function layers. The teacher model is trained using the predicted values output by the teacher model and the battery label data. Record the distillation target data of the teacher model, which includes the predicted output value of the teacher model, the attention features output by the teacher model, the attention weights output by the teacher model, and the output features of the LSTM layer of the teacher model.
5. The method according to claim 4, characterized in that, The step of training the teacher model using the output prediction value of the teacher model and the battery tag data includes: The teacher model is trained using its output predictions and the battery tag data. The loss of the teacher model is calculated using the mean squared error formula: Where n is the number of predicted categories, This represents the i-th value of the battery tag data; Let be the i-th predicted output value of the teacher model.
6. The method according to claim 5, characterized in that, The step of inputting the battery training data samples and the distillation target data of the teacher model into the student model for distillation training to obtain the trained target student model includes: The battery training data samples are input into the student model; The student model's Reshape layer, Transpose layer, and several depthwise separable 2D convolutional layers are used to perform time alignment on several features of the battery data to obtain a second battery training data sample. The second battery training data sample is intervalized using the temporal multi-scale Patch attention module of the student model to obtain a second output result, which includes the second output feature and the second output attention weight of the temporal multi-scale Patch attention module of the student model. The output features of the student model are obtained by extracting temporal features from the output features of the student model using several LSTM layers. The weighted output features of the student model's LSTM layer are obtained by weighting the student model's output attention weights; The student model output prediction value is obtained by processing the weighted LSTM layer output features of the student model using the fully connected layer and the Sigmoid activation function layer. The student model is trained by distillation using the output prediction value of the student model, the battery label data, and the distillation target data of the teacher model. After training, a trained target student model is obtained.
7. The method according to claim 6, characterized in that, The process of training the student model by distillation using the output predictions of the student model, the battery label data, and the distillation target data of the teacher model includes: The student model is trained by distillation using the output predictions of the student model, the battery label data, and the distillation target data of the teacher model. The hard-label loss of the student model is calculated using the mean squared error, and the calculation formula is as follows: in, It is a hard-label loss, where n is the number of predicted categories. This represents the i-th value of the battery tag data; This is the i-th predicted output value of the student model; The soft-label loss of the student model is calculated using the mean squared error, and the calculation formula is as follows: in, This is a soft-label loss, where n is the number of predicted categories. This is the i-th predicted output value of the student model; This is the i-th predicted output value of the teacher model; The attention weight distillation loss of the student model is calculated using the mean squared error formula, which is as follows: in, This is the attention weight distillation loss; n is the number of patches; This is the i-th value in the attention weights of the student model; This represents the i-th value in the teacher's attention weights model. The intermediate feature distillation loss of the student model is calculated using cosine similarity loss, and the calculation formula is as follows: in, This is the intermediate feature distillation loss; n is the number of feature layers distilled during training; The intermediate features of the i-th layer of the student model; The intermediate features of the i-th layer of the teacher model; This indicates that dimensionality matching is performed on the intermediate features of the teacher model to align with the intermediate features of the student model. This represents the modulo operation; the intermediate features are the output features of the temporal multi-scale Patch attention module and the output features of the LSTM layer; The total loss of the student model is calculated based on the hard label loss, the soft label loss, the attention weight distillation loss, and the intermediate feature distillation loss, using the following formula: in, It is hard label loss; It is soft label loss; It is the attention weight distillation loss; This is an intermediate characteristic of distillation loss; These are weighting coefficients that regulate the contribution of each component.
8. A battery state prediction system, characterized in that, The method is based on convolutional neural networks and long short-term memory networks, and the system includes: The training data acquisition module (701) is used to acquire battery data and process the battery data by constructing a sliding window to obtain battery training data samples. The teacher model construction module (702) is used to construct a teacher model based on the convolutional neural network and the long short-term memory network. The teacher model training module (703) is used to input the battery training data samples into the teacher model for training and record the distillation target data of the teacher model; The student model distillation training module (704) is used to construct a student model based on the convolutional neural network and the long short-term memory network, and input the battery training data samples and the distillation target data of the teacher model into the student model for distillation training to obtain a trained target student model. The student model deployment module (705) is used to convert the format of the target student model to obtain an embeddable student model and deploy the embeddable student model. The battery state prediction module (706) is used to predict the battery state using the embedded student model when real-time battery data is received.
9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer program is stored that can be loaded by a processor and executed according to any one of claims 1 to 7.