Energy storage battery health state evaluation method considering coupling effect of power distribution and micro-grid

By extracting global aging features using differential thermal voltammetry and incremental singular value decomposition, and combining Transformer and GRU models, along with ensemble learning using knowledge distillation and gradient boosting decision trees, the problem of insufficient accuracy in energy storage battery health status assessment was solved, achieving high-precision and robust assessment results.

CN121633853APending Publication Date: 2026-03-10STATE GRID ZHEJIANG ELECTRIC POWER CO LTD NINGBO POWER SUPPLY CO
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing methods for assessing the health status of energy storage batteries have insufficient accuracy in microgrid coupling scenarios, making it difficult to meet practical application needs and affecting the operational stability and safety of energy storage systems.

Method used

Differential thermal voltammetry and incremental singular value decomposition were used to extract global aging features. A health status assessment model was constructed by combining the Transformer model and the GRU model. The student model was trained on the distribution network side and fine-tuned on the microgrid side using the knowledge distillation method. Gradient boosting decision tree was used for ensemble learning to improve the assessment accuracy.

Benefits of technology

It achieves high-precision health status assessment of energy storage batteries in microgrid scenarios, improves the accuracy and robustness of the assessment, reduces problems such as battery overcharging, over-discharging and accelerated aging, and ensures the safety and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121633853A_ABST
    Figure CN121633853A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of energy storage battery state-of-health assessment, in particular to an energy storage battery state-of-health assessment method considering a micro-grid coupling effect, which comprises the following steps: performing signal analysis and global aging feature extraction on obtained target micro-grid side energy storage battery circulation aging data based on a differential thermal voltammetry to obtain global aging features; constructing a health state evaluation model comprising a distribution network side teacher model and a micro-grid side student model; training the student model on the distribution network side based on a knowledge distillation method, and finely adjusting the student model on the micro-grid side by using local data; inputting the global aging features into a health state evaluation model for health state evaluation, and integrating a first health state evaluation result output by the teacher model and a second health state evaluation result output by the student model after fine adjustment updating based on an integrated learning method taking a gradient boosting decision tree as a core, therefore, the health state evaluation result of the energy storage battery is obtained, and the high-precision health state evaluation result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy storage battery health status assessment technology, and specifically to an energy storage battery health status assessment method that takes into account the coupling effect of distribution microgrids. Background Technology

[0002] With the continuous increase in the penetration rate of distributed energy, the coupled operation of distribution networks and microgrids has become an important operating mode of new power systems. As a core hub for smoothing the output fluctuations of distributed power sources, ensuring power quality, and realizing coordinated dispatch of power sources, grids, loads, and storage, the accurate assessment of the state of health (SOH) of energy storage batteries directly determines the safe, stable operation and economical maintenance of distribution microgrids. In the coupled operation scenario of distribution microgrids, energy storage batteries need to frequently cope with the switching between grid-connected and off-grid modes, the impact load caused by the output fluctuations of distributed power sources, and the flexible charging and discharging under the guidance of peak and valley electricity prices. This leads to the battery aging process exhibiting highly nonlinear and multi-factor coupled characteristics, which places extremely high demands on the accuracy of health status assessment.

[0003] However, existing energy storage battery health status assessment methods generally suffer from insufficient assessment accuracy in microgrid coupled scenarios, making it difficult to meet practical application needs. Direct measurement methods calculate battery capacity or test battery internal resistance by performing ampere-hour integration under experimental conditions, thereby directly obtaining the battery's state of health (SOH). Model-based methods estimate SOH by constructing equivalent circuit models and electrochemical mechanism models, offering strong interpretability. Data-driven methods eliminate reliance on complex battery mechanistic knowledge, using machine learning and deep learning algorithms to directly perform nonlinear fitting of the battery's SOH based on battery data and feature extraction, making it easier to achieve high-precision SOH estimation.

[0004] The disadvantages of direct measurement methods are that they can usually only be implemented in the laboratory under specific operating conditions and experiments, and are difficult to apply in the dynamic, complex, and real-time-critical vehicle environment. The disadvantages of model-based methods are that the complex side reaction mechanisms and multi-field coupling characteristics inside the battery make it extremely difficult to build a high-precision model. At the same time, it is difficult to achieve strong generalization of SOH estimation under different external conditions. The disadvantages of data-driven methods are that they are completely dependent on data, have poor interpretability, and have extremely high requirements for data volume and data quality. At the same time, a single model is difficult to take into account both the global long-term aging trend and the local short-term capacity fluctuation.

[0005] Existing methods have limitations that lead to significant discrepancies between assessment results and the actual health status of batteries. This insufficient assessment accuracy directly impacts fault warnings, lifespan predictions, and charge / discharge strategy optimization for energy storage batteries in microgrids. It can easily cause problems such as overcharging, over-discharging, and accelerated aging, not only reducing the operational stability and lifespan of the energy storage system but also potentially leading to safety risks such as power outages in the microgrid. Summary of the Invention

[0006] The purpose of this invention is to solve the technical problem that when performing State of Health (SOH) assessment on energy storage batteries in distributed microgrid application scenarios under distribution networks, the isolated algorithm inside the battery management system is difficult to adapt to the distribution-microgrid collaborative system, resulting in low accuracy of energy storage battery health status assessment.

[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0008] This invention provides a method for assessing the health status of energy storage batteries considering the coupling effect of distribution microgrids. The method includes: analyzing the acquired cyclic aging data of the target microgrid-side energy storage battery using differential current voltammetry (DCV) to obtain the DCV response signal; extracting global aging features from the DCV response signal using incremental singular value decomposition (SVD) to obtain global aging features; constructing a health status assessment model, which includes a teacher model on the distribution network side and a student model on the microgrid side; the student model includes a CNN layer and a GRU layer connected sequentially; training the student model on the distribution network side using knowledge distillation, and then distributing the trained student model to the microgrid side; fine-tuning the distributed student model on the microgrid side using local data, freezing the CNN layer, and fine-tuning and updating the GRU layer; inputting the global aging features into the health status assessment model for health status assessment, and integrating the first health status assessment result output by the teacher model with the second health status assessment result output by the fine-tuned and updated student model using an ensemble learning method centered on gradient boosting decision trees to obtain the energy storage battery health status assessment result.

[0009] Optionally, the differential thermal voltage-current response signal can be obtained by: monitoring the voltage and temperature changes during the charging and discharging process of the battery based on the cyclic aging data of the target microgrid-side energy storage battery, and utilizing the entropy change heat effect accompanying the phase transition during the electrode material insertion and extraction process to obtain the differential thermal voltage-current response signal.

[0010] Optionally, the global aging features can be obtained through the following methods: Singular value decomposition is performed on the differential current transformer (DCT) matrix corresponding to the DCT response signal to obtain the first decomposition result; when new cyclic data arrives, the new cyclic data is projected onto the right singular vector matrix corresponding to the DCT matrix, and the projection coefficients and projection residuals are calculated; Singular value decomposition is performed on the L2 norm matrix of the projection residuals to obtain the second decomposition result; the first and second decomposition results are merged to obtain the merged decomposition result, and the elements in the diagonal singular value matrix of the merged decomposition result corresponding to each cycle of the DCT matrix are extracted as the global aging features.

[0011] Optionally, the teacher model is built around the encoder in the Transformer model, including an input layer, a self-attention mechanism layer, a residual connection and normalization layer, a feedforward neural network layer, and an output layer connected in sequence. The input layer is used to encode the trigonometric function position of the global aging features, and the encoded result is superimposed with the global aging features and then input into the self-attention mechanism layer. The self-attention mechanism layer adopts a multi-head attention structure.

[0012] Optionally, the student model is built around the gated recurrent unit (GRU) in a convolutional neural network. The CNN layer of the student model includes an input layer, a hidden layer, and an output layer connected in sequence. The hidden layer includes a convolutional layer and a max pooling layer connected in sequence. The convolutional layer is used to extract features through convolution and process them with the ReLU activation function. The max pooling layer is used to reduce the dimensionality of the features. The GRU layer of the student model includes an update gate, a reset gate, a candidate hidden state layer, and a final hidden state layer connected in sequence. The update gate and the reset gate are used to control the updating and resetting of the input information.

[0013] Optionally, it also includes: setting the loss function of the student model as the superposition of error loss and distillation loss, and determining the distillation loss in the following way: discretizing the health state value intervals and modifying the model output layer, outputting the probability distribution of health state values ​​falling in each interval through the Softmax function to obtain the probability distribution of the teacher model and the student model; determining the KL divergence of the teacher model and the student model based on the probability distribution, and then determining the distillation loss of the student model.

[0014] Optionally, the health status value range can be discretized and the model output layer can be modified by dividing the health status value range into multiple continuous intervals and setting the number of neurons in the model output layer based on the number of continuous intervals.

[0015] Optionally, before performing signal analysis on the acquired target microgrid-side energy storage battery cycle aging data based on differential thermal voltammetry, the method further includes: acquiring the original target microgrid-side energy storage battery cycle aging data; performing data cleaning on the original target microgrid-side energy storage battery cycle aging data to obtain the cleaned target microgrid-side energy storage battery cycle aging data; data cleaning includes removing abnormal data, fixing the sampling interval, data interpolation, or data resampling.

[0016] Optionally, the student model can be trained by: using Bayesian optimization to train the student model for hyperparameter optimization, and using Adam gradient descent to update the parameters of the student model.

[0017] Optional ensemble learning methods centered on gradient boosting decision trees include histogram optimization and one-sided gradient sampling.

[0018] The beneficial effects of this invention are as follows: 1. Signal analysis of energy storage batteries is performed based on differential thermal voltammetry (DVT), and the global aging features of energy storage batteries are extracted based on incremental singular value decomposition (SVD), which differs from traditional methods that rely solely on peak and valley features as aging characteristics. DVT can capture the coupled response of electrical and thermal signals during battery cycle aging, while SVD can perform global dimension feature reduction and key information mining on time-series DVT response signals, effectively avoiding evaluation bias caused by local features. The extracted global aging features are more comprehensive and representative, improving the accuracy of health status assessment.

[0019] 2. Distribution network-side training based on knowledge distillation leverages the massive data and mature assessment experience of the distribution network-side teacher model to transfer complex health status assessment knowledge to the more lightweight microgrid-side student model, ensuring that the student model maintains high initial assessment performance despite its simplified structure. The microgrid-side local fine-tuning strategy, freezing the CNN layers and updating only the GRU layers, utilizes both the spatial feature extraction capabilities of the CNN layers for global aging characteristics and the temporal feature learning of the GRU layers to adapt to the personalized operating conditions of the microgrid-side batteries. This fully utilizes the high precision and global information of the teacher model, overcoming the performance loss issues caused by the limited computing resources and data of the student model on the microgrid side, and improving the accuracy of health status assessment.

[0020] 3. An ensemble learning method centered on gradient boosting decision trees is used to weightedly fuse the global evaluation results (first health status evaluation results) of the distribution network-side teacher model and the localized evaluation results (second health status evaluation results) of the microgrid-side student model, effectively complementing the advantages of the two models. The teacher model relies on global data, and its evaluation results have global universality; the student model is adapted to the microgrid-side operating conditions, and its evaluation results are more consistent with the actual state of local energy storage batteries. This overcomes the evaluation errors of a single model under extreme operating conditions or data noise, resulting in a more robust and accurate health status evaluation result. Attached Figure Description

[0021] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings.

[0022] Figure 1 This is a flowchart of a method for assessing the health status of energy storage batteries that takes into account the coupling effect of distribution microgrids in this invention;

[0023] Figure 2 This is a schematic diagram of a method for assessing the health status of energy storage batteries that takes into account the coupling effect of distribution microgrids in this invention.

[0024] Figure 3 This is a schematic diagram of DTV analysis for a method of assessing the health status of energy storage batteries that takes into account the coupling effect of distribution microgrids in this invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only one preferred embodiment of this invention and are only used to explain this invention. They do not limit the scope of protection of this invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0026] As one implementation method, the purpose of this invention is to solve the problem that when performing SOH evaluation on energy storage batteries in distributed microgrid application scenarios under distribution networks, the isolated algorithms inside the battery management system are difficult to adapt to the distribution-microgrid collaborative system, resulting in information silos and low evaluation accuracy, which in turn makes it difficult to achieve high-precision SOH evaluation of energy storage batteries with global collaboration in distribution-microgrid collaborative scenarios.

[0027] like Figure 1 As shown, this invention provides a method for assessing the state of health (SOH) of energy storage batteries considering the coupling effect of microgrids. This method is characterized by strong practicality, high assessment accuracy, strong anti-interference capability, and strong robustness. This invention can be used for SOH assessment of energy storage batteries in microgrid systems, and has diverse application prospects in energy storage and power system applications. This method can achieve high-precision and robust SOH assessment of energy storage batteries. Specifically, it includes:

[0028] S1. Data Acquisition and Cleaning: Raw target microgrid-side energy storage battery cycle aging data is acquired through various means, including power distribution terminals, edge computing gateways, cloud monitoring platforms, and public datasets. This raw data includes current, voltage, temperature data, and SOH (State of Health) tags. Data cleaning is then performed on the raw data to obtain cleaned target microgrid-side energy storage battery cycle aging data. Data cleaning includes removing outlier data, fixing sampling intervals, data interpolation, or data resampling.

[0029] S2. Feature Extraction: Signal analysis is performed on the acquired cycle aging data of the target microgrid-side energy storage battery using the differential current voltammetry (DCV) method to obtain the DCV response signal. Global aging features are then extracted from the DCV response signal using incremental singular value decomposition (SVD) to obtain global aging features.

[0030] S21. The differential thermal voltammetry (DTV) response signal is obtained through the following methods: monitoring voltage and temperature changes during battery charging and discharging based on the cyclic aging data of the target microgrid-side energy storage battery; utilizing the entropy change thermal effect accompanying phase transitions during electrode material insertion and extraction to obtain the DTV response signal. Signal analysis is performed using the Differential Thermal Voltammetry (DTV) method, a highly efficient electrochemical characterization technique widely used in battery aging analysis and characterization. The insertion and extraction processes in electrode materials trigger phase transitions accompanied by entropy changes, resulting in significant thermal effects. The DTV response signal can be obtained by monitoring voltage and temperature changes during battery charging and discharging. During battery aging, the decrease in peak value and the shift of peak position to higher potentials, along with the decrease in valley value and the shift of valley position to lower potentials, lead to the derived peak-valley values, peak-valley positions, and peak-valley areas. These characteristics effectively reflect changes such as increased resistance and uneven electrode performance during battery capacity decay, thus effectively reflecting the microscopic changes in the battery during aging on a macroscopic scale.

[0031] S22. Global aging features are obtained through the following method. Global aging feature extraction refers to feature extraction based on the entire DTV curve to fully utilize the overall information of the DTV curve, unlike traditional methods that rely solely on peak and valley features as aging features. This method treats the entire DTV curve as a high-dimensional signal and extracts its principal component features through matrix decomposition, thereby comprehensively utilizing the global morphological change information contained in the entire curve. This method not only utilizes peak and valley information but also fully leverages the curve offset information strongly correlated with battery aging contained in the entire curve, improving the utilization rate of DTV information and the quality of the extracted features.

[0032] Global aging feature extraction employs incremental singular value decomposition (I-SVD) to extract aging features from the DTV curve. I-SVD is a variant of singular value decomposition, offering higher efficiency, especially for online applications like SOH assessment. When new data is added or data changes, I-SVD utilizes the existing singular value decomposition results of the original matrix to quickly approximate the updated matrix's singular values ​​through incremental updates, thus avoiding a complete re-decomposition of the entire matrix and achieving higher efficiency and real-time performance.

[0033] S221. Perform singular value decomposition on the differential thermal voltammetry matrix corresponding to the differential thermal voltammetry response signal to obtain the first decomposition result.

[0034] S222. When new cyclic data arrives, the new cyclic data is projected onto the right singular vector matrix corresponding to the differential thermal voltammetry matrix. The projection coefficients and projection residuals are calculated, and the L2 norm matrix of the projection residuals is decomposed into singular values ​​to obtain the second decomposition result.

[0035] S223. Merge the first decomposition result and the second decomposition result to obtain the merged decomposition result. Extract the elements in the diagonal singular value matrix and the corresponding positions of the differential thermal current matrix in each cycle of the merged decomposition result as global aging features.

[0036] S3. Model Construction: Construct a health status assessment model, which includes a teacher model on the distribution network side and a student model on the microgrid side; the student model includes a CNN layer and a GRU layer connected in sequence.

[0037] The teacher model is built around the encoder in the Transformer model, and includes an input layer, a self-attention mechanism layer, a residual connection and normalization layer, a feedforward neural network layer and an output layer connected in sequence. The input layer is used to encode the trigonometric function position of the global aging features, and the encoded result is superimposed with the global aging features and then input into the self-attention mechanism layer. The self-attention mechanism layer adopts a multi-head attention structure.

[0038] The teacher model refers to a deep learning model deployed on the distribution network side for SOH (State of Health) assessment. SOH is a core indicator for measuring the performance degradation of electrochemical energy storage devices such as energy storage batteries and power batteries. The distribution network side has abundant computing, hardware, and data resources, thus the teacher model has high model complexity and computational consumption, making it easier to achieve globally collaborative feature capture and high-precision SOH assessment. The teacher model is built around the Transformer model, a deep learning model suitable for solving sequence problems. The core mechanism of the Transformer model is self-attention, which can model dependencies without considering distances in the input or output sequences, directly calculating the dependencies between any two positions. This allows for better capture of global information and higher parallelism and computational efficiency.

[0039] The student model is built around the gated recurrent unit (GRU) in a convolutional neural network. The CNN layer of the student model consists of an input layer, a hidden layer, and an output layer connected in sequence. The hidden layer consists of a convolutional layer and a max pooling layer connected in sequence. The convolutional layer is used to extract features through convolution and process them with the ReLU activation function, while the max pooling layer is used to reduce the dimensionality of the features. The GRU layer of the student model consists of an update gate, a reset gate, a candidate hidden state layer, and a final hidden state layer connected in sequence. The update gate and the reset gate are used to control the updating and resetting of the input information.

[0040] The student model refers to the deep learning model for SOH evaluation deployed on the microgrid side. Due to limited hardware and computing resources on the microgrid side, the SOH evaluation model deployed there is typically a performance-constrained, lightweight model. Its goal is to achieve high-accuracy, real-time SOH evaluation for its corresponding microgrid system. The student model is built around a Convolutional Neural Network (CNN)-Gated Recurrent Unit (GRU) model. CNNs possess powerful spatial feature recognition capabilities, effectively capturing spatial correlations between features. The CNN is followed by a GRU model to capture temporal relationships. The GRU model is a deep learning model suitable for solving sequence problems; it is a variant of the Recurrent Neural Network (RNN). By introducing a gating mechanism, it can selectively control the updating and resetting of information, allowing the network to learn long-term dependencies more efficiently. Facing scenarios with differentiated temporal feature distributions of multiple energy storage units on different microgrid sides, it can efficiently and automatically capture their differentiated temporal features.

[0041] S4. Micro-network Collaborative Training: The student model is trained on the distribution network side using knowledge distillation, and the trained student model is then distributed to the micro-network side. On the micro-network side, the distributed student model is fine-tuned using local data, the CNN layers are frozen, and the GRU layers are fine-tuned and updated. Bayesian optimization is used to train the student model for hyperparameter optimization, and Adam gradient descent is used to update the parameters of the student model.

[0042] Distribution-micro collaboration refers to a collaborative architecture where a complex model is trained on the distribution network side, while a lightweight model is deployed on the microgrid side. A teacher model is built on the distribution network side, and a student model is trained using a cross-model knowledge distillation strategy before being deployed to the microgrid side for fine-tuning. Cross-model knowledge distillation is a model compression and knowledge transfer method that leverages the high accuracy and global information of the teacher model to compensate for performance losses caused by the limited complexity of the student model and the limited data on the microgrid side. By modifying the network structures of the teacher and student models, the regression problem model is transformed into a pseudo-classification problem model, allowing for a comparison of the probability distributions output by the teacher and student models. This introduces distillation loss into the loss function during student model training, enabling knowledge distillation between models with different structures. The complex knowledge of the distribution network-side teacher model is transferred to the lightweight student model on the microgrid side, thus achieving efficient deployment of a lightweight student model with high accuracy and strong generalization capabilities, approaching that of the teacher model, on computationally limited microgrid-side devices.

[0043] Microgrid-side student model fine-tuning refers to the process of distributing the student model, distilled from the teacher model on the distribution network side, to each microgrid node. Each microgrid node then uses local data to train and update the student model. This allows the student model to learn general knowledge about battery aging and global information from the teacher model, enabling it to perform targeted updates to the microgrid nodes it is deployed with, thereby improving its evaluation accuracy. Since the differences between microgrid nodes mainly lie in the time-series characteristics under different operating conditions, the CNN layers of the student model are frozen during model fine-tuning, and only the GRU layers are fine-tuned.

[0044] The loss function for the student model is set as the superposition of error loss and distillation loss. The distillation loss is determined by: discretizing the health state value intervals and modifying the model output layer; using the Softmax function to output the probability distribution of health state values ​​falling within each interval to obtain the probability distributions of the teacher and student models; determining the KL divergence of the teacher and student models based on the probability distributions, and then determining the distillation loss of the student model. Discretizing the health state value intervals and modifying the model output layer includes: dividing the health state value range into multiple continuous intervals and setting the number of neurons in the model output layer accordingly based on the number of continuous intervals.

[0045] S5. Health Status Assessment: The aging characteristics of the entire domain are input into the health status assessment model for health status assessment. Based on the ensemble learning method with gradient boosting decision tree as the core, the first health status assessment result output by the teacher model is integrated with the second health status assessment result output by the student model after fine-tuning and updating to obtain the health status assessment result of the energy storage battery.

[0046] The ensemble learning method refers to integrating the evaluation results of the teacher model on the distribution network side and the student model on the microgrid side, thereby achieving a high-precision SOH evaluation that takes into account both global information and local characteristics. The ensemble learning method is built on machine learning algorithms, with LightGBM as the core. LightGBM significantly improves the model efficiency by using optimization strategies such as histogram optimization and one-sided gradient sampling on the basis of the traditional Gradient Boosting Decision Tree (GBDT).

[0047] As one implementation method, such as Figure 2 As shown, it is divided into the following four steps: data processing and feature extraction; deep learning model construction; micro-cooperative training; ensemble learning and model validation.

[0048] Step 1: Data processing and feature extraction: Taking two battery cycle aging datasets as examples, the two datasets have different battery material systems and experimental conditions to fully simulate the actual micro-cooperative scenario and thus verify the method proposed in this invention.

[0049] The specific information for Dataset 1 is as follows: Eight 0.74Ah pouch cells were used, and the cells were placed in a MK53 hot chamber at 40°C. The cells were subjected to constant current charge-discharge cycle aging tests at a rate of 1C using an 8-channel bioMPG205 battery tester.

[0050] The specific information for Dataset 2 is as follows: Four 18650 lithium batteries were used, and an aging experiment was conducted at room temperature using a constant current and constant voltage charging mode. During charging, the batteries were initially charged at a constant current of 1.5A. After reaching the upper cutoff voltage, the charging was switched to constant voltage charging. The charging process stopped when the charging current dropped below 20mA. During discharging, the batteries continued to discharge at a continuous current of 2A until they reached their lower cutoff voltage.

[0051] The same data processing and feature extraction were then performed on both datasets.

[0052] The collected voltage and temperature data undergo preprocessing by removing outliers, fixing sampling intervals, and performing data interpolation or resampling to ensure equal data length in each cycle. The differential voltage-temperature voltammetry (DCFT) method is then used to derive characteristics from the processed voltage and temperature data, and is widely applied in battery aging analysis and characterization. For example... Figure 3 As shown, during battery aging, the decrease in peak value and its shift towards higher potential, along with the decrease in valley value and its shift towards lower potential, give rise to various characteristics that can effectively reflect the battery capacity degradation process. The DTV response signal can be obtained by monitoring voltage and temperature changes during battery charging and discharging, as shown below:

[0053]

[0054] in, Represents temperature difference, Represents the voltage step size. Represents the time step.

[0055] The DTV is calculated using preprocessed data, followed by global aging feature extraction from the DTV curve. Global aging feature extraction refers to feature extraction based on the entire DTV curve to fully utilize its overall information, unlike traditional methods that rely solely on peak and valley features as aging characteristics. This method treats the entire DTV curve as a high-dimensional signal and extracts its principal component features through matrix factorization, thereby comprehensively utilizing the global morphological change information contained within the entire curve. This approach not only utilizes peak and valley information but also fully leverages the curve offset information strongly correlated with battery aging contained within the entire curve, improving the utilization rate of DTV information and the quality of the extracted features.

[0056] I-SVD technology is used to extract global aging features from DTV curves. I-SVD is a variant of Singular Value Decomposition (SVD) and is more efficient than traditional SVD, especially for online applications such as SOH assessment. When new data is added or data changes, I-SVD uses the existing SVD results of the original matrix to quickly approximate the updated matrix's singular values ​​through incremental updates, thus avoiding a complete SVD of the entire matrix again. This results in higher efficiency and real-time performance. The calculation of I-SVD is as follows:

[0057] For the DTV data matrix, the standard SVD decomposition is as follows:

[0058]

[0059] in, The data is a DTV matrix with a size of j rows. Column, where j is the cycle number, Let U be the DTV data volume, i.e., the dimension of the high-dimensional DTV signal. U is the left singular vector matrix, with size j rows and r columns, where r is the rank of the DTV data matrix. V is the right singular vector matrix, with size j rows and r columns. Rows and columns r. It is a diagonal singular value matrix, where the elements on the diagonal are singular values.

[0060] When new cyclic data arrives, a new matrix is ​​obtained:

[0061]

[0062] Decomposed and updated as follows:

[0063]

[0064] in, This is the new matrix obtained when new cyclic data arrives. For the new left singular vector matrix, For the new right singular vector matrix, For the new diagonal singular value matrix, For new cyclic data.

[0065] Update the decomposition using the following steps:

[0066] 1: Project the new row onto the original right singular vector matrix, and calculate the projection coefficient vector of the new row onto the original right singular vector matrix:

[0067]

[0068] 2: Calculate the projection residual, representing the portion of the new row not covered by the original space:

[0069]

[0070] 3: Construct a 1x1 matrix from the L2 norm of the residuals, and then perform singular value decomposition on the small matrix constructed from the residuals:

[0071]

[0072] 4: Merge the original decomposition with the new decomposition to obtain the updated left singular vector matrix, diagonal singular value matrix, and right singular vector matrix:

[0073]

[0074]

[0075]

[0076] In the formula, Let be the residual vector, representing the difference between the new row and the projection; This is the projection coefficient vector.

[0077] Finally, the singular value decomposition of the updated matrix was obtained. The elements at the corresponding positions of each row of the diagonal singular value matrix and the DTV matrix were extracted as features, which are the global aging features of the DTV curve. These features were then used as input to the deep learning model for SOH evaluation.

[0078] Step 2: Deep learning model construction: Construct teacher models deployed on the distribution network side and student models deployed on the microgrid side.

[0079] The teacher model refers to the deep learning model for SOH evaluation deployed on the distribution network side. The distribution network side has abundant computing, hardware, and data resources, thus the teacher model has high model complexity and high computational cost, making it easier to achieve globally collaborative feature capture and high-precision SOH evaluation. The teacher model is built around the Transformer model, which has excellent performance in solving time series problems. Furthermore, its core self-attention mechanism can model time dependencies without considering sequence distance, resulting in higher parallelism and computational efficiency. SOH evaluation is a many-to-one regression problem, which solves the regression problem based on feature sequences rather than sequence-to-sequence generation; therefore, only the encoder part of the Transformer is used. The Transformer encoder model includes an input layer, a self-attention mechanism layer, a residual connection and normalization layer, a feedforward neural network layer, and an output layer.

[0080] First, in the input layer, the input DTV global aging feature sequence is positionally encoded. Positional encoding uses trigonometric functions to inject positional information into the time series, generating a unique code for each vector at each position. This code is then added to the input time series data. The trigonometric function positional encoding can be described as follows:

[0081]

[0082]

[0083] in Indicates location, Representing feature dimension, Encoding the position of trigonometric functions, This indicates that the position code is for an even position. This indicates that the position code is for an odd-numbered position.

[0084] Then the input timing data and position code are added together:

[0085]

[0086]

[0087] in, The input is the global aging feature sequence of the DTV, and the input of the m-th encoder layer is... The output is , The length of the time series. The final output of the input layer, superscript This indicates that the variable is an input sequence. Encoding the position of trigonometric functions, For input feature time series data, subscript This indicates that the variable is the time series data of the input features of the input layer.

[0088] Then, the self-attention mechanism layer is entered. For a single head, the input sequence is transformed linearly to generate a query (Q), a key (K), and a value (V):

[0089]

[0090] in , and These are the linear transformation weight matrices, and are the learnable parameters, with superscripts... , and These indicate that the linear transformation weight matrix corresponds to the linear transformation weight matrix for the query, key, and value, respectively. This is the input to the m-th encoder layer.

[0091] Then, the scaling dot product attention is calculated, which involves calculating the dot product of each dot with all others, dividing each result by a normalization factor, and then processing it using the softmax activation function.

[0092]

[0093] in , , These are the query matrix, key matrix, and value matrix, respectively. This represents the softmax activation function, used to normalize values ​​to a range of weights between 0 and 1, where the sum of the weights equals 1. As the key dimension, Scaling factor This is the result of scaling the dot product attention calculation.

[0094] The multi-head attention mechanism is built upon the scaled dot product attention mechanism described above. It concatenates the attention outputs of each head and performs a linear transformation by multiplying them by the projection matrix, ultimately yielding the output of the multi-head attention mechanism.

[0095]

[0096]

[0097] in, This represents a splicing operation. Represents the projection matrix. To ultimately obtain the output of the multi-head attention mechanism, The result of splicing the attention outputs of each head.

[0098] Then, we proceed to the residual connection and normalization layer:

[0099]

[0100] The layer normalization formula is as follows:

[0101]

[0102] In the above formula, And β are learnable parameters. The mean, To prevent small constants with denominators of 0, The variance of the input. For the output of the residual connection and normalization layer, The input to the representation layer is normalized. The formula for calculating layer normalization is given.

[0103] Then, it enters the feedforward neural network for nonlinear transformation:

[0104]

[0105] in, This represents the activation function. This represents the output of the feedforward neural network layer. and These represent the weights and biases of the first layer of the feedforward neural network. and These are the weights and biases of the second layer of the feedforward neural network, respectively.

[0106] The feedforward neural network is followed by another residual connection layer and a normalization layer. The specific formulas are as described above and will not be repeated here. The final output is set to... .

[0107] at last, The input is fed into the output layer, which consists of fully connected layers. The output layer outputs the SOH evaluation results of the teacher model.

[0108]

[0109] in, The subscript for the SOH assessment results output by the teacher model. This indicates that the variable is a variable of the output layer in the Transformer. For the activation function, this application chooses the Sigmoid function. and These are the weights and biases of the output layer in the Transformer, respectively.

[0110] The student model refers to the deep learning model for SOH evaluation deployed on the microgrid side. Due to limited hardware and computing resources on the microgrid side, SOH evaluation models deployed there are typically lightweight models with performance constraints. Their goal is to achieve high-accuracy, real-time SOH evaluation for their corresponding microgrid systems. The student model is built around a Convolutional Neural Network (CNN)-Gated Recurrent Unit (GRU) model. CNNs have powerful feature extraction capabilities, effectively capturing spatial relationships between features. The CNN is followed by a GRU model to capture temporal relationships. The GRU model is a deep learning model suitable for solving sequence problems. By introducing a gating mechanism, it can selectively control the updating and resetting of information, allowing the network to learn long-term dependencies more efficiently. Faced with scenarios where multiple energy storage units have differentiated temporal feature distributions on different microgrid sides, it can efficiently and automatically capture their differentiated temporal features.

[0111] CNN consists of three layers: input layer, hidden layer, and output layer.

[0112] The extracted DTV global aging features are input into the CNN through the input layer;

[0113] The hidden layers consist of convolutional layers and max-pooling layers; the convolutional layer is the core of the CNN, responsible for extracting abstract features from the input. Max-pooling layers are added after the convolutional layers for further dimensionality reduction, thereby improving efficiency.

[0114] The input data is fed into the input layer and then passed through the convolutional layer for convolution calculation:

[0115]

[0116] Among them, for the first One input and one output, The input is the global aging feature sequence of the DTV. For convolution kernel weights, The kernel size is [size]. This is the index used for convolution calculation. This is the bias of the convolution kernel. This is the output of the convolutional layer.

[0117] Then, non-linear activation is performed using an activation function. In this application, the ReLU function is chosen as the activation function.

[0118]

[0119] in, The input to the ReLU function, i.e., the output of the convolutional layer. ; These are the features after processing by the activation function.

[0120] Then, the data enters the pooling layer for further dimensionality reduction:

[0121] First, zero values ​​are added to both ends of the input sequence to determine the length of the output sequence after pooling:

[0122]

[0123] in, The length of the output sequence after pooling. Given the sequence length of the input pooling layer, To determine the pooling window size, This is an adjustable padding parameter, representing the length of zero-value padding at both ends of the input sequence. This is the stride size of the pooling window. This indicates rounding down to the nearest integer.

[0124] The first output sequence element It is determined by the maximum value within the corresponding window in the input sequence, as shown in the formula:

[0125]

[0126] Finally, the data is output to the GRU model through the output layer, which is generally composed of fully connected layers and is responsible for generating the output for the next layer of the network.

[0127]

[0128] in, The abstract features output by the CNN layer of the student model, with subscripts This indicates that the variable is a variable in the output layer of the student model CNN. For the activation function, this application chooses the Sigmoid function. and These represent the weights and biases in the output layer of the student CNN model.

[0129] The CNN is followed by the GRU model to capture temporal relationships. The GRU mainly consists of update gate, reset gate, candidate hidden state layer, final hidden state layer and output layer.

[0130] First, the data is updated to reflect the degree to which historical information of the gate control is retained:

[0131]

[0132] in, For the activation function, this application chooses the Sigmoid function. The input to the GRU is the abstract feature output by the CNN. , and These are the weights and biases of the updated gates, respectively. The old state, To update the gate output, subscript This indicates that the variable is a variable in the update gate of the GRU layer in the student model.

[0133] Then, the impact of resetting the gate control history information on the candidate state is as follows:

[0134]

[0135] in, For the activation function, this application chooses the Sigmoid function. For the input of GRU, A value close to 0 indicates that historical states should be ignored. A value close to 1 indicates a combination of historical states. and These are the weights and biases for resetting the gate, respectively.

[0136] Then, a temporary new state is generated through the candidate hidden state layer:

[0137]

[0138] In the formula, The temporary new state generated for the candidate hidden state layer. The hyperbolic tangent activation function is used. and These represent the weights and biases of the candidate hidden state layer, respectively.

[0139] Then, the old and new states are merged through the final hidden state layer:

[0140]

[0141] in, This is the output of the final hidden state layer. This indicates element-wise multiplication.

[0142] Finally, the GRU output layer outputs the final SOH evaluation result of the student model:

[0143]

[0144] in, The subscript of the SOH assessment results output by the student model. This indicates that the variable is a variable in the output layer of the GRU in the student model. This is the output of the final hidden state layer. For the activation function, this application chooses the Sigmoid function. and These represent the weights and biases of the output layer in the GRU in the student model.

[0145] Step 3: Perform micro-cooperative training.

[0146] Distribution-micro collaboration refers to a collaborative architecture where a complex model is trained on the distribution network side, while a lightweight model is deployed on the microgrid side. A teacher model is built on the distribution network side, and a student model is trained using a cross-model knowledge distillation strategy before being deployed to the microgrid side for fine-tuning. Cross-model knowledge distillation is a model compression and knowledge transfer method that leverages the high accuracy and global information of the teacher model to compensate for performance losses caused by the limited complexity of the student model and the limited data on the microgrid side. By modifying the network structures of the teacher and student models, the regression problem model is transformed into a pseudo-classification problem model, allowing for a comparison of the probability distributions output by the teacher and student models. This introduces distillation loss into the loss function during student model training, enabling knowledge distillation between models with different structures. The complex knowledge of the distribution network-side teacher model is transferred to the lightweight student model on the microgrid side, thus achieving efficient deployment of a lightweight student model with high accuracy and strong generalization capabilities, approaching that of the teacher model, on computationally limited microgrid-side devices.

[0147] Both the microgrid and distribution network models were trained using Bayesian optimization techniques for automatic hyperparameter optimization and the Adam gradient descent algorithm. During the training of the student model on the microgrid side, cross-model knowledge distillation was also performed based on the teacher model on the distribution network side. Cross-model knowledge distillation is a model compression and knowledge transfer method.

[0148] 1. First, the network structures of the teacher and student models are modified, transforming the regression problem model into a pseudo-classification problem model, thereby enabling a comparison of the probability distributions output by the teacher and student models:

[0149] The continuous SOH values ​​are discretized because a battery is generally considered unusable when the SOH drops below 80%. Therefore, the range from 80% to 100% is divided into 200 intervals of 0.1%. Then, the output layers of both the teacher and student models are modified from one neuron to 200 neurons to output a 200-dimensional probability distribution vector. This vector is activated using a Softmax function, and each value represents the probability that the SOH falls within that interval. The median of the intervals with the highest probabilities is taken as the final SOH evaluation value. The Softmax function is:

[0150]

[0151] in, This is the input to the Softmax function.

[0152] 2. Then, in the loss function during student model training, in addition to the basic error loss between the SOH evaluation value and the actual SOH value, distillation loss is added based on the model modifications described above. Distillation loss is determined based on the KL divergence of the output probability distributions of the student and teacher models. KL divergence is an indicator that evaluates the degree of difference between two probability distributions, measuring the information loss or asymmetric distance of one probability distribution relative to another. Finally, the loss function during student model training is:

[0153] Error loss:

[0154]

[0155] Distillation loss:

[0156]

[0157]

[0158] Total loss of student model:

[0159]

[0160] In the above formula, For error loss, MSE is the sum of the student model's SOH evaluation value and the actual SOH value; This represents the actual SOH value. The SOH evaluation value for the student model; This represents the total number of cycles; Circular index; For distillation loss, KL divergence is given for the probability distributions of the student model SOH evaluation and the teacher model SOH evaluation. Here is the formula for the KL divergence. and These are two probability distributions for comparison; The probability distribution for the teacher model SOH evaluation; The probability distribution for evaluating the student model SOH; The total loss of the student model is the sum of the distillation loss and the error loss.

[0161] This enables knowledge distillation between different structural models, transferring the complex knowledge of the distribution network-side teacher model to the lightweight student model on the microgrid side. This allows for the efficient deployment of a lightweight student model with high accuracy and strong generalization ability, similar to the teacher model, in microgrid-side devices with limited computing resources.

[0162] The trained student model will be deployed to microgrid nodes and fine-tuned using local microgrid data. This allows the student model to learn general knowledge about battery aging and global information from the teacher model, enabling it to make targeted updates to the deployed microgrid nodes and improve its evaluation accuracy. Since the differences between microgrid nodes are mainly reflected in the time-series characteristics under different operating conditions, the CNN layers of the student model are frozen during model fine-tuning, and only the GRU layers are retrained.

[0163] Step 4: Integration of learning and model validation.

[0164] When conducting SOH assessment, the assessment results of the distribution network-side teacher model and the microgrid-side student model are integrated to obtain a high-precision SOH assessment result that takes into account both global information and local characteristics. The ensemble learning method is built on machine learning algorithms, with LightGBM as the core. LightGBM significantly improves model efficiency by using optimization strategies such as histogram optimization and one-sided gradient sampling on the basis of traditional gradient boosting decision tree (GBDT).

[0165] The SOH evaluation results of the distribution network-side teacher model and the microgrid-side student model are used as input features:

[0166]

[0167] LightGBM, as a gradient boosting tree, achieves its fusion rule by accumulating multiple decision trees:

[0168] The final fusion result of the meta-model is a weighted sum of the predictions from T decision trees:

[0169]

[0170] In the above formula, For the number of decision trees, For the first The output of each decision tree, index For decision tree indexing, The features of the input meta-model are the SOH evaluation values ​​output by the student model and the teacher model. For the first The weights of each decision tree; This represents the final fusion result of the meta-model; The representative meta-model is LightGBM.

[0171] Each decision tree is a piecewise constant function that divides the input space into several leaf nodes through feature splitting, with each leaf node corresponding to a fixed predicted value; for the input... , No. The output of the tree is:

[0172]

[0173] in, For the first The output of each tree; For the first The number of leaf nodes in the tree; For the first Tree No. The predicted value for the nth leaf node; for the nth leaf node. The first tree The input space region corresponding to each leaf node; This is an indicator function; it is 1 when the input belongs to the corresponding input space region, and 0 otherwise. Index for leaf nodes;

[0174] The weights of each tree are determined by minimizing the MSE. Each tree is updated using the gradient descent algorithm. The loss function for the regression task is:

[0175]

[0176] in, The loss function for LightGBM is... This represents the actual SOH value. This is the final SOH evaluation value output by the meta-model; This represents the total number of cycles; Circular index;

[0177] When performing model validation, the data from Dataset 1 and Dataset 2 are distributed to the distribution network side and the microgrid side to fully simulate the distribution-microgrid collaboration scenario: On the distribution network side, the four batteries in Dataset 1 and the two batteries in Dataset 2 are used as the data on the distribution network side. Cross-validation is performed during the model training process, that is, one battery is used as the test set and the data of the other batteries are used as the training set for training and testing.

[0178] On the microgrid side: The remaining battery data of dataset 1 and dataset 2 are used as data for two different microgrid nodes, with the remaining data of dataset 1 used as data for microgrid node 1 and the remaining data of dataset 2 used as data for microgrid node 2.

[0179] Use one of the other four batteries in Dataset 1 as the fine-tuning training set and the remaining batteries as the test set; use one of the other two batteries in Dataset 2 as the fine-tuning training set and the remaining batteries as the test set.

[0180] This invention provides a method for assessing the state of health (SOH) of energy storage batteries considering the coupling effect of distribution microgrids. It features high practicality, high assessment accuracy, strong anti-interference capability, and robustness. The method employs signal analysis to perform aging analysis and extract global aging features. Deep learning models for SOH assessment are constructed on both sides of the distribution microgrid, and model training is performed for distribution-microgrid collaboration based on cross-model knowledge distillation and fine-tuning strategies. The final SOH assessment result is obtained by fusing the SOH assessment results from both sides of the distribution microgrid using ensemble learning. Through this technical solution, this invention can be used for SOH assessment of energy storage batteries in microgrid systems, showing diverse application prospects in energy storage and power systems. This method can achieve high-precision and robust SOH assessment of energy storage batteries.

[0181] This invention proposes a health status assessment method for energy storage batteries that considers the coupling effect of microgrids. It uses differential current voltammetry (DCV) for signal analysis of the energy storage battery and incremental singular value decomposition (SVD) for global aging feature extraction. Unlike traditional methods that rely solely on peak-valley characteristics for aging features, this method treats the entire DCV curve as a high-dimensional signal and extracts its principal component features through SVD, thus comprehensively utilizing the global morphological change information contained in the entire curve. By analyzing macroscopic signals, this invention effectively reflects the microscopic changes in battery aging using only the cyclic aging data of the target microgrid-side energy storage battery. This solves the problems of low data and information utilization and low feature extraction quality in signal analysis-based battery aging characterization.

[0182] This application proposes a method for assessing the health status of energy storage batteries that considers the coupling effect of distribution and microgrids. It constructs a distribution-microgrid collaborative architecture by training a complex model on the distribution network side and deploying a lightweight model on the microgrid side. A teacher model is built on the distribution network side, and a student model is trained using a cross-model knowledge distillation strategy and then distributed to the microgrid side for fine-tuning. Based on this distribution-microgrid collaboration, a lightweight student model with high accuracy and strong generalization ability, approaching the teacher model, is efficiently deployed in microgrid-side devices with limited computing resources. This fully utilizes the high accuracy and global information of the teacher model, solving the performance loss problem caused by the limited computing resources and data on the microgrid side.

[0183] This application proposes a method for assessing the health status of energy storage batteries that considers the coupling effect of distribution microgrids. During knowledge distillation, the network structure of the teacher model and the student model is modified, and the regression problem model is transformed into a pseudo-classification problem model. This allows for a comparison of the probability distributions of the outputs of the teacher model and the student model. Consequently, distillation loss is introduced into the loss function during the training process of the student model. This solves the problem that heterogeneous models are difficult to distill through model compression pruning or loss functions when performing a many-to-one regression problem for state assessment.

[0184] This application proposes a method for assessing the health status of energy storage batteries that considers the coupling effect of distribution and microgrids. It constructs a high-precision teacher model on the distribution network side with the computationally expensive Transformer model as the core, and a lightweight student model on the microgrid side with the computationally inexpensive CNN-GRU model as the core, thus solving the adaptability problem of deploying SOH assessment data-driven models on the distribution network side and the microgrid side.

[0185] This application proposes a method for assessing the state of health (SOH) of energy storage batteries that considers the coupling effect of distribution microgrids. Based on LightGBM, it integrates the assessment results of the teacher model on the distribution network side and the student model on the microgrid side, thereby efficiently obtaining high-precision SOH assessment results that take into account both global information and local characteristics. This solves the problem of SOH assessment that is difficult to achieve in independent microgrid nodes while taking into account both global information and local characteristics, as well as the problem of low efficiency of traditional gradient boosting tree models.

[0186] Compared with the prior art, the present invention has the following beneficial effects based on the above embodiments:

[0187] 1. Signal analysis of energy storage batteries is performed based on differential current voltammetry (DCF), and global aging features are extracted using incremental singular value decomposition (SVD), which differs from traditional methods that rely solely on peak and valley features as aging characteristics. DCF can capture the coupled response of electrical and thermal signals during battery cycling, while SVD can perform global feature reduction and key information mining on time-series DCF response signals, effectively avoiding evaluation bias caused by local features. The extracted global aging features are more comprehensive and representative, improving the accuracy of health status assessment.

[0188] 2. Distribution network-side training based on knowledge distillation leverages the massive data and mature assessment experience of the distribution network-side teacher model to transfer complex health status assessment knowledge to the more lightweight microgrid-side student model, ensuring that the student model maintains high initial assessment performance despite its simplified structure. The microgrid-side local fine-tuning strategy, freezing the CNN layers and updating only the GRU layers, utilizes both the spatial feature extraction capabilities of the CNN layers for global aging characteristics and the temporal feature learning of the GRU layers to adapt to the personalized operating conditions of the microgrid-side batteries. This fully utilizes the high precision and global information of the teacher model, overcoming the performance loss issues caused by the limited computing resources and data of the student model on the microgrid side, and improving the accuracy of health status assessment.

[0189] 3. An ensemble learning method centered on gradient boosting decision trees is used to weightedly fuse the global evaluation results (first health status evaluation results) of the distribution network-side teacher model and the localized evaluation results (second health status evaluation results) of the microgrid-side student model, effectively complementing the advantages of the two models. The teacher model relies on global data, and its evaluation results have global universality; the student model is adapted to the microgrid-side operating conditions, and its evaluation results are more consistent with the actual state of local energy storage batteries. This overcomes the evaluation errors of a single model under extreme operating conditions or data noise, resulting in a more robust and accurate health status evaluation result.

Claims

1. A method for evaluating the state of health of an energy storage battery considering the coupling of a microgrid, characterized in that, The method comprises the following steps: Based on the differential heat voltammetry method, the obtained target micro-grid side energy storage battery cycle aging data is analyzed to obtain a differential heat voltammetry response signal; Based on the incremental singular value decomposition method, the differential heat voltammetry response signal is subjected to global aging feature extraction to obtain a global aging feature; A health state evaluation model is constructed, which includes a teacher model on the distribution network side and a student model on the micro-grid side; the student model includes a CNN layer and a GRU layer connected in sequence; Based on the knowledge distillation method, the student model is trained on the distribution network side, and the trained student model is distributed to the micro-grid side; the student model is fine-tuned on the micro-grid side using local data, the CNN layer is frozen, and the GRU layer is fine-tuned and updated; The global aging feature is input into the health state evaluation model for health state evaluation, and based on the ensemble learning method with the gradient boosting decision tree as the core, the first health state evaluation result output by the teacher model and the second health state evaluation result output by the fine-tuned and updated student model are integrated to obtain the energy storage battery health state evaluation result. 2.The method of claim 1, wherein, The differential heat voltammetry response signal is obtained by the following method: based on the target micro-grid side energy storage battery cycle aging data, the voltage change and temperature change during the battery charging and discharging process are monitored, and the entropy change heat effect accompanying the phase change during the electrode material insertion and extraction process is used to obtain the differential heat voltammetry response signal. 3.The method of claim 2, wherein, The global aging feature is obtained by the following method: The singular value decomposition is performed on the differential heat voltammetry matrix corresponding to the differential heat voltammetry response signal to obtain a first decomposition result; When new cycle data arrives, the right singular vector matrix corresponding to the differential heat voltammetry matrix is used to project the new cycle data, and the projection coefficient and the projection residual are calculated; The L2 norm matrix of the projection residual is subjected to singular value decomposition to obtain a second decomposition result; The first decomposition result and the second decomposition result are combined to obtain a combined decomposition result, and the elements corresponding to each cycle of the differential heat voltammetry matrix in the diagonal singular value matrix in the combined decomposition result are extracted as the global aging feature. 4.The method of claim 1, wherein, The teacher model is constructed based on the encoder in the Transformer model, including an input layer, a self-attention mechanism layer, a residual connection and a normalization layer, a feedforward neural network layer and an output layer connected in sequence; the input layer is used for triangular function position coding of the global aging feature, and the coding result is superimposed with the global aging feature and input into the self-attention mechanism layer; the self-attention mechanism layer adopts a multi-head attention structure. 5.The method of claim 1, wherein, The student model is constructed based on the gated recurrent unit in the convolutional neural network, and the CNN layer of the student model includes an input layer, a hidden layer and an output layer connected in sequence; the hidden layer includes a convolution layer and a max pooling layer connected in sequence, the convolution layer is used for extracting features through convolution calculation and processing through a ReLu activation function, and the max pooling layer is used for dimension reduction of the features; the GRU layer of the student model includes an update gate, a reset gate, a candidate hidden state layer and a final hidden state layer connected in sequence, and the update gate and the reset gate are used for controlling the update and reset of the input information. 6.The method of claim 1, wherein, The method further comprises the following steps: The loss function of the student model is set as the superposition of an error loss and a distillation loss, and the distillation loss is determined by discretizing the health state value interval and modifying the model output layer, outputting the probability distribution of the health state value falling in each interval through a Softmax function to obtain the probability distribution of the teacher model and the student model. The KL divergence of the teacher model and the student model is determined based on the probability distribution, and the distillation loss of the student model is further determined.

7. The method of claim 6, wherein the method further comprises: Discretizing the health state value interval and modifying the model output layer comprises: dividing the health state value range into a plurality of continuous intervals, and setting the number of neurons of the model output layer based on the number of continuous intervals. 8.The method of claim 1, wherein, Before signal analysis is performed on the obtained target micro-grid side energy storage battery cycle aging data based on differential heat voltammetry, the method further comprises: obtaining original target micro-grid side energy storage battery cycle aging data; performing data cleaning on the original target micro-grid side energy storage battery cycle aging data to obtain data cleaned target micro-grid side energy storage battery cycle aging data; the data cleaning comprises removing abnormal data, fixing the sampling interval, data interpolation or data resampling. 9.The method of claim 1, wherein, The student model is trained by: using the Bayesian optimization method to perform hyperparameter optimization training on the student model, and using the Adam gradient descent method to update the parameters of the student model. 10.The method of claim 1, wherein, The ensemble learning method with gradient boosting decision tree as the core includes histogram optimization method and one-sided gradient sampling method.

Citation Information

Cited By

  • Energy storage power station cloud-edge collaborative health management method and system based on knowledge distillation and transfer learning

    CN121965711A

  • A cloud-edge collaborative health management method and system for energy storage power stations based on knowledge distillation and transfer learning

    CN121965711B