Artificial Intelligence-Driven Method for Predicting the State of Health and Remaining Useful Life of Battery Energy Storage
Through artificial intelligence-driven multi-scale patch hybrid network and self-supervised learning method, multi-scale features of battery data are extracted, solving the shortcomings of traditional methods in battery energy storage health status monitoring and residual life prediction, and achieving higher prediction accuracy and generalization capabilities.
Patent Information
- Application Number
- CN202510213107.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional methods face problems such as lack of labeled data, high false alarm rate, poor generalization ability in battery energy storage health status monitoring and residual life prediction, and insufficient assumptions of comparison learning methods, and image analysis methods are not suitable for the timing characteristics and complex periodicity of battery data.
Using a multi-scale patch hybrid network driven by artificial intelligence, time series features are extracted through time mixers and spatial mixers, combined with self-supervised comparison representation learning and self-supervised classification, and the triple loss function and custom loss function are used to optimize feature extraction and classification effects.
It improves the accuracy, stability and generalization ability of battery energy storage health status monitoring and residual life prediction, reduces the false alarm rate, and enhances the model's adaptability to complex battery data.
Smart Images

Figure CN119716587B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an artificial intelligence-driven method for predicting the health state and remaining life of battery energy storage. Background Art
[0002] In the aspect of monitoring the health state and predicting the remaining life of battery energy storage, traditional methods face many challenges. In real-world scenarios, the labeled data is extremely scarce, which greatly limits the application of supervised learning methods. Most existing methods adopt an unsupervised way to learn the normal behavior of unlabeled time series. However, the definition of the normal boundary of these methods is often too narrow, resulting in minor deviations being misjudged as anomalies, thus generating a high false alarm rate. In addition, the generalization ability of the model is also greatly limited and it is difficult to adapt to complex and changing practical application scenarios.
[0003] When traditional contrastive learning methods are applied to monitor the health state and predict the remaining life of battery energy storage, their assumptions are seriously insufficient. For example, it is usually considered that the enhanced time series window is a positive sample, and the window that is far away in time is a negative sample. However, this assumption often does not hold in practical applications because time series enhancement may generate negative samples, and windows that are far away may also belong to positive samples. In addition, directly applying image analysis methods to monitor the health state and predict the remaining life of battery energy storage does not achieve good results because battery data has unique time series characteristics and complex periodicity, which are essentially different from image data.
[0004] These deficiencies in the prior art seriously affect the accuracy, stability, and generalization ability of monitoring the health state and predicting the remaining life of battery energy storage. To overcome these challenges, a new technical method is urgently needed to more accurately capture the time series characteristics and periodicity of battery data, improve the generalization ability of the model and the prediction accuracy. Summary of the Invention
[0005] To overcome the deficiencies of the prior art, the present invention provides an artificial intelligence-driven method for predicting the health state and remaining life of battery energy storage, aiming to improve the accuracy, stability, and generalization ability of monitoring the health state and predicting the remaining life of battery energy storage, improve the application defects of traditional contrastive learning methods, and achieve accurate monitoring of the health state and prediction of the remaining life of battery energy storage, so as to solve the deficiencies in anomaly detection, feature extraction, and model optimization of traditional methods.
[0006] In a first aspect, the present invention provides an artificial intelligence-driven method for predicting the health state and remaining life of battery energy storage, including:
[0007] Obtain the time series data of various characteristics and status information of the battery at historical time points and perform preprocessing; the various characteristics and status information include voltage data, current data, temperature data, capacity data, battery internal resistance, service life, and historical cycle times;
[0008] Input the preprocessed time series data into a feature extraction model based on a multi-scale patch hybrid network for feature extraction, where the multi-scale patch hybrid network consists of a temporal mixer and a spatial mixer;
[0009] In the self-supervised contrastive representation learning stage, use the triplet loss function to train the feature extraction model to optimize the time series features extracted by the network, and use the learned feature representations to find semantically meaningful nearest neighbors and farthest neighbors for each sample;
[0010] In the self-supervised classification stage, perform classification based on the nearest neighbor and farthest neighbor information of all feature samples, and train the classification model through a custom loss function to achieve the best classification effect;
[0011] Use the trained model to classify and predict the new time series data of the battery, and output the battery health status.
[0012] Furthermore, where the preprocessing of the time series data of various characteristics and status information of the battery at historical time points includes:
[0013] Adopt instance normalization to standardize the time series window;
[0014] Based on point anomalies and subsequence anomalies in the standardized time series data, randomly select dimensions and start times and perform anomaly injection to construct anomaly samples.
[0015] Furthermore, where the temporal mixer captures short-term and long-term dependencies using convolutional kernels of different sizes to extract variable time series information; the spatial mixer performs feature aggregation based on the maximum variance framework to extract multi-scale information on variable dimensions.
[0016] Furthermore, where the temporal mixer captures short-term and long-term dependencies using convolutional kernels of different sizes to extract variable time series information, including:
[0017] The temporal mixer extracts variable time series information by using linear layers and multi-layer perceptrons along the time dimension, and adds the input and output through short connections or residual connections. The specific formula is:
[0018] ,
[0019] where, is the input time series feature; is a linear layer for performing linear layer normalization on the input time series features; is a multi-layer perceptron for feature transformation and extraction, performing a non-linear mapping on the normalized data to learn deeper representations; is the final output feature, i.e., the extracted variable time series information.
[0020] Furthermore, among them, the spatial mixer performs feature aggregation based on the maximum variance framework, extracting multi-scale information on the variable dimension, including:
[0021] Analyze the data period using the maximum variance framework to identify periodic patterns in the data;
[0022] Based on the results of the data period analysis, perform multi-scale patch segmentation, including: determining the patch sizes of different scales according to the data period; dividing the original preprocessed time series data into patches of multiple scales, each patch containing data within a specific time range;
[0023] For each scale of patch, use a multi-layer perceptron with shared parameters to extract features along the variable dimension;
[0024] Determine the multi-scale feature aggregation weights based on the amplitude calculated by the maximum variance framework, fuse the multi-scale features according to the calculated multi-scale feature aggregation weights, and add them to the output of the time mixer through a short connection. The specific formula is:
[0025] ,
[0026] where, is the i-th scale patch determined according to the data period analysis; n represents the number of scales; represents the features extracted by applying the multi-layer perceptron to the i-th scale patch; is the aggregation weight of the i-th scale feature determined by the amplitude calculated by the maximum variance framework; is the output of the time mixer; S is the final output of the spatial mixer.
[0027] Furthermore, among them, the analysis of the data period using the maximum variance framework to identify periodic patterns in the data includes:
[0028] The maximum variance framework finds a modified circulant matrix by matching the modified circulant matrix, that is, finding a modified circulant matrix C such that the variance of the data CX obtained by transforming the original data X by this modified circulant matrix C is maximized. By maximizing the variance, the most important features in the time series data, that is, the data period patterns, are extracted.
[0029] Further, in the self-supervised contrastive representation learning stage, training the feature extraction model using a triplet loss function to optimize the time series features extracted by the network includes:
[0030] Training the feature extraction model using a triplet loss function to minimize the distance between the anchor sample and the positive sample, while maximizing the distance between the anchor sample and the negative sample, with a difference of at least a preset interval ;
[0031] If the distance between the negative sample and the anchor fails to exceed the distance between the positive sample and the anchor plus the preset interval , a loss will be generated, otherwise the loss is zero, thereby optimizing the time series features extracted by the feature extraction model to better distinguish normal and abnormal samples.
[0032] Further, in the self-supervised classification stage, classification is performed based on the nearest neighbor and farthest neighbor information of all feature samples, and the classification model is trained through a custom loss function to achieve the best classification effect, including:
[0033] Training the classification model through a custom loss function to enable the classification model to accurately classify the nearest neighbor and farthest neighbor information of all feature samples; among them, the loss function maximizes the similarity of the window with its nearest neighbor and minimizes the similarity of the window with its farthest neighbor. The loss function formula is:
[0034] ,
[0035] where represents the th window; represents the nearest neighbor; represents the farthest neighbor; represents the similarity metric; controls the weight for optimizing the nearest neighbor similarity; controls the weight for optimizing the farthest neighbor similarity; represents the th window and its nearest neighbor ; represents the th window and its farthest neighbor ; α is a hyperparameter used to control the weight of the regularization term; H(x) represents entropy.
[0036] Further, using the trained model to classify and predict the new time series data of the battery and output the battery health status, including:
[0037] After completing model training, determine the class assignment and majority class of the dataset through the inference process, including:
[0038] First, use the trained classification model to assign classes to the battery data, identify normal and abnormal battery data, and then count the proportion of the majority class, that is, normal battery data;
[0039] Calculate the probability that the new window belongs to the majority class, determine whether the new window is abnormal based on this probability, and output the abnormal score of the battery health status indicator.
[0040] Furthermore, the method further includes:
[0041] Use the trained model to extract features from the battery health status data of the battery, and input the battery health status data after feature extraction into the remaining life prediction model for prediction, output the remaining life prediction value of the battery. If the remaining life prediction value is less than a certain threshold, trigger an alarm or propose a suggestion to replace the battery.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] (1) Improve prediction accuracy: By introducing a multi-scale patch hybrid network, the present invention makes full use of the multi-scale characteristics of time series data, improving the prediction accuracy of the model for battery health status and remaining life; the design of the custom loss function enables the model to better distinguish normal and abnormal data, further enhancing the prediction accuracy.
[0044] (2) Enhance model stability: Through the application of the modified circulant matrix and the maximum variance framework, the present invention enables the model to more flexibly model non-uniform periodic changes, enhancing the stability of the model; through the centralized kernel alignment method, the similarity between different feature layers is measured, improving the stability and generalization ability of data features.
[0045] (3) Improve generalization ability: The method of the present invention does not rely on a large amount of labeled data. Through self-supervised learning and feature extraction, the generalization ability of the model is improved; the application of the anomaly injection technique enables the model to learn more comprehensive anomaly features, further enhancing the generalization ability of the model.
[0046] (4) Have important practical value: The present invention proposes a new multi-scale patch hybrid network structure, providing new ideas and methods for feature extraction of time series data; by combining self-supervised learning and feature extraction, accurate prediction of battery health status and remaining life is achieved, having important practical value.
[0047] In summary, the present invention overcomes the deficiencies of traditional methods in the health state monitoring and remaining life prediction of battery energy storage, significantly improves the accuracy, stability, and generalization ability of prediction, and has important practical value.
[0048] It should be understood that the content described in the summary of the invention is not intended to limit the key or important features of the embodiments of the present invention, nor to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present invention. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0050] Figure 1 is a flowchart of an artificial intelligence-driven method for predicting the health state and remaining life of battery energy storage according to an embodiment of the present invention;
[0051] Figure 2 is a schematic diagram of the specific steps of an artificial intelligence-driven method for predicting the health state and remaining life of battery energy storage according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] In addition, the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0054] The embodiments of the present invention provide an artificial intelligence-driven method for predicting the health state and remaining life of battery energy storage. The present invention includes the following innovative points:
[0055] 1. The present invention proposes a multi-scale patch hybrid network to make full use of the multi-scale characteristics of time series data. This network consists of a temporal mixer and a spatial mixer. The temporal mixer uses convolutional kernels of different sizes to capture short-term and long-term dependencies and combines residual connections to enhance the temporal modeling ability. The spatial mixer performs feature aggregation based on the maximum variance framework to extract multi-scale information in the variable dimension and improve the model's ability to characterize the battery health state. 2. Aiming at the complex periodicity and non-stationary characteristics of battery energy storage data, the present invention constructs a modified circulant matrix and a maximum variance framework, which can more flexibly model non-uniform periodic changes and improve the time series pattern recognition ability. 3. The present invention designs a custom loss function suitable for self-supervised contrastive learning to optimize the effect of time series feature extraction. Through the triplet loss function, the distance between the anchor sample and the positive sample is minimized, while the distance to the negative sample is increased, thereby enhancing the model's ability to distinguish normal and abnormal data; the loss function based on the nearest neighbor and the farthest neighbor emphasizes the similarity constraint between windows to ensure that the time series representation has stronger discriminability and robustness. 4. To further optimize the representation ability of time series, the present invention measures the similarity between different feature layers through the centralized kernel alignment method, improves the stability and generalization ability of data features, and enhances the credibility and interpretability of the model.
[0056] Figure 1 The flowchart of an artificial intelligence-driven method for predicting the health state and remaining life of battery energy storage is shown; Figure 2 The schematic diagram of the specific steps of an artificial intelligence-driven method for predicting the health state and remaining life of battery energy storage is shown. As Figure 1 and Figure 2 shown, an artificial intelligence-driven method 100 for predicting the health state and remaining life of battery energy storage, which is a method for predicting the health state and remaining life of battery energy storage based on self-supervised contrastive representation learning, is divided into two stages, namely the self-supervised contrastive representation learning stage and the self-supervised classification stage. This method 100 includes:
[0057] (I) Self-supervised contrastive representation learning stage
[0058] S110: Obtain the time series data of various features and state information at the historical time points of the battery and perform preprocessing; the various features and state information include voltage data, current data, temperature data, capacity data, battery internal resistance, service life, and historical cycle times;
[0059] In the monitoring of the health state and prediction of the remaining life of battery energy storage, time series data is a series of observed values arranged in chronological order, which reflects various characteristics and state information of the battery at different time points, including voltage data, current data, temperature data, capacity data, battery internal resistance, service life, historical cycle times, etc. In this step S110, it is used to preprocess the time series data, specifically including:
[0060] S111: Standardization of time series data
[0061] The instance normalization is used to standardize the time series window to eliminate the drift of the input data distribution. Since the distribution of time series data is prone to shift due to natural or human factors, and it is difficult for the training set to cover all distributions. This module connects the prediction timestamp with the historical sequence for synchronous normalization (using instance normalization), and removes the denormalization term, only retaining the instance normalization for input stability to help the model cope with data distribution changes. The instance normalization formula is:
[0062] ,
[0063] where, represents the th feature of the th sample, represents the mean of the th feature, represents the variance of the th feature, is a small constant to prevent the denominator from being zero, is the result after normalization.
[0064] By standardizing the time series data, the dimensional difference between different time series is eliminated, making the subsequent analysis more accurate.
[0065] S112: Anomaly injection into time series data
[0066] Anomaly injection is usually performed on the time series data after standardization, and point anomalies (single-point context deviation) and subsequence anomalies (periodic, trend anomalies, etc.) are used to construct anomaly samples. In multivariate time series data, dimensions and start times are randomly selected for anomaly injection:
[0067] ,
[0068] : The original value of the
[0069] th dimension in the multivariate time series at time
[0070] : Abnormal disturbance term (point anomaly or subsequence anomaly such as periodic anomaly or trend anomaly can be adopted).
[0071] In order to enrich the training samples by generating time series data similar to real anomalies in the case of insufficient data or high annotation costs, time series data anomaly injection can improve the model's ability to identify anomalies. By injecting different types of anomalies, various anomaly situations can be simulated, enabling the model to learn more comprehensive anomaly features. By training the model on data containing anomalies, the model's robustness to anomalies can be improved, reducing false positives and false negatives.
[0072] S120: Input the preprocessed time series data into the feature extraction model based on the multi-scale patch hybrid network for feature extraction, where the multi-scale patch hybrid network consists of a temporal mixer and a spatial mixer; among them, the temporal mixer captures short-term and long-term dependencies using convolutional kernels of different sizes and extracts variable time series information; the spatial mixer performs feature aggregation based on the maximum variance framework and extracts multi-scale information in the variable dimension.
[0073] In this step S120, when performing feature extraction using the multi-scale patch hybrid network, multi-scale window features of the time series are extracted using different channels and kernel sizes. Different channels correspond to different dimensions, and different kernel sizes are used to capture multi-time scale features. Finally, a multi-layer perceptron layer is added to generate a low-dimensional feature vector. The multi-scale patch hybrid network structure consists of a temporal mixer and a spatial mixer. Specifically:
[0074] S121: The temporal mixer captures short-term and long-term dependencies using convolutional kernels of different sizes and extracts variable time series information, including:
[0075] The temporal mixer is a key component of the multi-scale patch mixing network, responsible for extracting variable time series information along the temporal dimension. The input to the temporal mixer is typically a three-dimensional tensor, whose dimensions can be represented as (number of samples, number of time steps, number of features). This tensor contains time series data for multiple samples, where each sample records the values of multiple features at multiple time steps. The temporal mixer first uses a linear layer (also known as a fully connected layer or a dense layer) to process the input data. The linear layer maps the original feature space to a new feature space by applying a weight matrix and a bias term to the input data, for capturing the linear relationships in the input data. After the linear layer, the temporal mixer uses a multi-layer perceptron (MLP) to further extract complex non-linear features. The MLP consists of multiple hidden layers, each of which contains a linear layer. The number of neurons in each hidden layer can be the same or different. Typically, as the number of layers increases, the number of neurons gradually decreases to capture higher-level features. An activation function (such as ReLU, Tanh, etc.) is usually used after each hidden layer to increase non-linearity. To improve the training efficiency and performance of the network, the temporal mixer adds the input to the output using a short connection or a residual connection. This step helps to alleviate the vanishing gradient problem in deep networks and facilitates the learning of the identity mapping.
[0076] The temporal mixer extracts variable time series information by using a linear layer and a multi-layer perceptron along the temporal dimension, and adds the input to the output using a short connection or a residual connection. The specific formula is as follows:
[0077] ,
[0078] where, is the input time series feature, which can represent multi-dimensional time series data within a time window. After being normalized, it is input into the model; is the linear layer, used for linear layer normalization of the input time series feature to improve the stability and convergence speed of model training; is the multi-layer perceptron, used for feature transformation and extraction. It performs non-linear mapping on the normalized data to learn deeper representations; is the final output feature, i.e., the extracted variable time series information. It is an incremental adjustment after being transformed by the multi-layer perceptron based on the input feature , enabling the network to better learn the feature information in the time series.
[0079] S122: The spatial mixer performs feature aggregation based on the maximum variance framework to extract multi-scale information in the variable dimension, including:
[0080] The spatial mixer is responsible for extracting variable-dimensional features and performing multi-scale feature aggregation based on the maximum variance framework. This process includes data cycle analysis, multi-scale patch segmentation, feature extraction, weight calculation, and feature fusion, etc. Data cycle analysis: The spatial mixer first analyzes the data using the maximum variance framework to identify periodic patterns in the data. The maximum variance framework is a statistical method that can reveal the most significant components or cycles in the data. By analyzing the data cycle, the spatial mixer can determine the different scales present in the data (such as short-term fluctuations, long-term trends, etc.). Multi-scale patch segmentation: Based on the results of the data cycle analysis, the spatial mixer performs multi-scale patch segmentation. This step aims to divide the data into segments or "patches" of different scales for subsequent independent extraction and analysis of features at different scales. Feature extraction: For each scale of patch, the spatial mixer uses a parameter-sharing multi-layer perceptron (MLP) to extract features along the variable dimension. Parameter sharing means that patches of different scales use the same MLP structure for feature extraction, which helps reduce the number of model parameters and improve generalization ability. Among them, the multi-layer perceptron includes an input layer, multiple hidden layers, and an output layer; the MLP is applied to each scale of patch to extract features on the variable dimension; the number of hidden layers and neurons of the MLP can be adjusted according to the specific task and data characteristics. Weight calculation: The spatial mixer determines the weights for multi-scale feature aggregation based on the amplitude calculated by the maximum variance framework. The amplitude reflects the importance or significance of features at different scales and can therefore be used as a weight to fuse multi-scale features. Feature fusion and short connection: The spatial mixer fuses the multi-scale features according to the calculated weights and adds them to the output of the temporal mixer through a short connection. This step realizes multi-scale data fusion analysis and feature representation enhancement.
[0081] Through the above process, the spatial mixer can achieve multi-scale data fusion analysis and feature representation enhancement. Specifically, it includes:
[0082] S1221: Analyze the data cycle using the maximum variance framework to identify periodic patterns in the data;
[0083] Furthermore, step S1221: Analyze the data cycle using the maximum variance framework to identify periodic patterns in the data, specifically including:
[0084] The maximum variance framework extracts the data cycle pattern and enhances the time series modeling ability by matching a modified circulant matrix, that is, finding a modified circulant matrix C such that the variance of the data CX obtained by transforming the original data X by the matrix C is maximized. By maximizing the variance, the most important features in the time series data, namely the data cycle pattern, can be extracted. The specific formula is:
[0085] ,
[0086] Among them, is a modified circulant matrix used to capture periodic patterns in time series data. A modified circulant matrix refers to a matrix that, based on the standard circulant matrix, introduces additional adjustments (such as weighting, non-uniform shifting, etc.) to adapt to non-stationary time series, making it more flexible to adapt to real-world data. Through the modified circulant matrix, complex time-dependent relationships can be better modeled, especially capturing key features in non-strictly periodic data (such as battery health data). : The goal of this optimization problem is to find a modified circulant matrix such that the transformed data has the maximum variance, that is, emphasizing the change pattern of the periodic signal.
[0087] S1222: Based on the results of data period analysis, perform multi-scale patch segmentation, including: determining the patch sizes at different scales according to the data period; dividing the original preprocessed time series data into patches at multiple scales, with each patch containing data within a specific time range;
[0088] S1223: For each scale of patch, use a multi-layer perceptron (MLP) with parameter sharing to extract features along the variable dimension;
[0089] S1224: Determine the multi-scale feature aggregation weights based on the amplitude calculated by the maximum variance framework, fuse the multi-scale features according to the calculated multi-scale feature aggregation weights, and add them to the output of the temporal mixer through a short connection to achieve multi-scale data fusion analysis and enhanced feature representation. The specific formula is:
[0090] ,
[0091] Among them, is the patch at the i-th scale determined according to data period analysis; I represents the number of scales; represents the features extracted by applying the multi-layer perceptron to the patch at the i-th scale; is the aggregation weight of the features at the i-th scale determined by the amplitude calculated by the maximum variance framework; is the output of the temporal mixer; S is the final output of the spatial mixer.
[0092] Extract multi-scale features from time series data through the multi-scale patch mixing network feature extraction model, providing effective feature representations for subsequent contrastive representation learning and classification tasks.
[0093] S130: In the self-supervised contrastive representation learning stage, use the triplet loss function to train the feature extraction model to optimize the time series features extracted by the network, and find the semantically meaningful nearest neighbor and farthest neighbor for each sample using the learned feature representation;
[0094] S131: In the self-supervised contrastive representation learning stage, using the triplet loss function to train the feature extraction model to optimize the time series features extracted by the network includes:
[0095] Use the triplet loss function to train the feature extraction model, minimize the distance between the anchor sample and the positive sample, while maximizing the distance between the anchor sample and the negative sample, and at least differ by a preset interval ;
[0096] If the distance between the negative sample and the anchor fails to exceed the distance between the positive sample and the anchor plus the preset interval , a loss will be generated, otherwise the loss is zero, thereby optimizing the time series features extracted by the feature extraction model to better distinguish normal and abnormal samples. The specific formula is:
[0097] ,
[0098] where, is the anchor sample, i.e., the current sample 's data point; is the positive sample, i.e., the sample belonging to the same category or similar to the anchor sample; is the negative sample, i.e., the sample with a different category or dissimilar to the anchor sample; is the feature extraction function, used to map the input data to the embedding space; is the square of the Euclidean distance between the anchor and the positive sample, representing the distance between similar samples; is the square of the Euclidean distance between the anchor and the negative sample, representing the distance between dissimilar samples; Interval, used to ensure that the distance between the negative sample and the anchor is greater than the distance of the positive sample, usually a hyperparameter;: i.e., , ensuring that the loss value is non-negative.
[0099] In this stage, set the triplet loss function for contrastive learning, aiming to make similar samples closer in the feature space, while different samples are farther apart in the feature space. By setting the triplet loss function to train the model, the model can learn an effective feature representation for distinguishing similar and dissimilar samples.
[0100] S132: Use the learned feature representations to find semantically meaningful nearest neighbors and farthest neighbors for each sample, generating prior information that can capture the fine-grained fluctuations and overall trends of the time series, enhancing the richness of the feature representations to improve performance in the self-supervised classification stage.
[0101] Specifically, step S132 includes:
[0102] S1321: Process the time series data through a multi-scale patch mixing network, which outputs one or more feature vectors that contain the feature information of the time series data at different scales. Each sample corresponds to one or more feature vectors, which can be regarded as the representation of the sample in the feature space;
[0103] S1322: To compare the similarity between feature vectors, select a suitable metric method. Commonly used metric methods include Euclidean distance, cosine similarity, etc. For each pair of samples, calculate the distance or similarity between their feature vectors to provide an indicator for measuring the semantic similarity between samples;
[0104] S1323: Find the nearest neighbors and farthest neighbors:
[0105] (1) Finding the nearest neighbors: For each sample, calculate its distance or similarity to all other samples in the dataset. Then, sort the samples according to these distances or similarities. Among the sorted samples, the sample with the smallest distance (or the largest similarity) to the current sample is the nearest neighbor.
[0106] (2) For each sample, we calculate its distance or similarity to all other samples in the dataset. Then, sort the samples according to these distances or similarities. Among the sorted samples, the sample with the largest distance (or the smallest similarity) to the current sample is the farthest neighbor.
[0107] Through the above process, it is possible to use the learned feature representations to find semantically meaningful nearest neighbors and farthest neighbors for each sample.
[0108] (II) Self-supervised classification stage
[0109] S140: In the self-supervised classification stage, classify based on the nearest neighbor and farthest neighbor information of all feature samples, and train the classification model through a custom loss function to achieve the best classification effect;
[0110] In step S140, classification is performed based on the nearest neighbor and farthest neighbor information of all feature samples obtained in step S130. As a prior, it is incorporated into the learnable method to maximize the similarity between the window representation and the nearest neighbor and minimize the similarity between the window representation and the farthest neighbor. The model is trained through a custom loss function, and the formula of this loss function is:
[0111] ,
[0112] where, represents the th window; represents the nearest neighbor; represents the farthest neighbor; represents the similarity metric; controls the weight for optimizing the nearest neighbor similarity; controls the weight for optimizing the farthest neighbor similarity; represents the th window and the similarity between it and its nearest neighbor ; represents the th window and the similarity between it and its farthest neighbor ; α is a hyperparameter used to control the weight of the regularization term; H(x) represents entropy.
[0113] This loss function maximizes the similarity between the window representation and its nearest neighbor, minimizes the similarity between the window representation and its farthest neighbor, and applies the entropy loss to prevent overfitting. Eventually, a discriminative representation is learned to distinguish normal and abnormal windows, and the classification performance of the model is optimized.
[0114] In the classification stage, a suitable classification model can be selected or designed, such as a fully connected layer, a support vector machine, etc., to classify the extracted features.
[0115] S150: Use the trained model to classify and predict the new time series data of the battery, and output the battery health status.
[0116] In this step S150, for the time series data features after feature extraction and self-supervised learning, the trained model is used to predict the new time series data, and information such as the health status (normal, abnormal, etc.) is output. Specifically, it includes: After completing the model training, the class assignment and majority class of the data set are determined through the inference process, and further includes:
[0117] S151: First, use the trained classification model to assign categories to the battery data, identify normal and abnormal battery data, then count the proportion of the majority class (i.e., normal battery data), calculate the probability that the new window belongs to the majority class, and determine whether the new window is abnormal based on this probability, and output the abnormal score of the battery health status indicator;
[0118] During the self-supervised inference process, determine whether the window is abnormal by calculating the probability P(y = c|x) that the new window belongs to the majority class. The formula is:
[0119] ,
[0120] where x represents the sample; represents the category; represents the majority class; represents the given sample After that, the probability that it belongs to the category is used to determine whether a certain window is abnormal. When is low, it means that the sample has a low possibility of belonging to the majority class (normal class), so it can be determined as abnormal; is the prediction score of the model for the sample belonging to the category ; is the total number of categories; is the exponentiated prediction score of the category to make it always positive for easy normalization; is the exponential sum of the prediction scores of all categories, which plays a role in normalization to ensure that the sum of the probabilities of all categories is 1.
[0121] Determine whether the window is abnormal based on this probability, and an abnormal score can be generated for further analysis. The lower the score, the higher the possibility of abnormality. The inference process will finally give a judgment on whether the new window is abnormal. By calculating the probability that the new window belongs to the majority class, if the probability P(y = c|x) is lower than a certain preset threshold (set according to the specific application scenario and the tolerance for false positives and false negatives), then determine that the window is abnormal; if the probability P(y = c|x) is higher than the threshold, then determine it as normal.
[0122] For battery health status monitoring, the input is the current data of the battery (such as voltage, current, temperature, etc.), and the model outputs the abnormal score of the battery health status indicator. Assume that the threshold for health status monitoring is set to 0.8. If the predicted value of the current health status of the battery is 0.6, then it can be determined that the battery is in an abnormal state.
[0123] S152: Extract features from the battery health state data of the battery using the trained model. The battery health state data includes voltage, current, temperature, capacity data, battery internal resistance, service life, historical cycle count, etc. Then input the battery health state data after feature extraction into the remaining life prediction model for prediction, and output the predicted remaining life value of the battery. If the predicted remaining life value is less than a certain threshold, trigger an alarm or propose a suggestion to replace the battery.
[0124] Specifically, perform feature extraction on the input new time series data through a multi-scale patch hybrid network. Then, through self-supervised learning, that is, training in the contrastive representation learning and classification stages, enable the model to learn the complex relationships between the time series data, health state, and remaining life. Then input the extracted features into a remaining life prediction model such as a regressor (such as a fully connected layer or support vector regression SVR, etc.), and output the remaining life, expressed in time or charge-discharge cycle count.
[0125] For remaining life prediction, the input is the current health state data of the battery (including voltage, current, temperature, capacity data, battery internal resistance, service life, historical cycle count, etc.) into the remaining life prediction model. If the model predicts that the remaining life is less than a certain threshold (such as 50 hours), trigger an alarm or propose a suggestion to replace the battery.
[0126] Preferably, in some embodiments, to improve the classification performance, the feature extraction model can be fine-tuned. The fine-tuning process is to adjust the parameters of its last few layers to adapt to a specific classification task while keeping most of the parameters of the feature extraction model unchanged.
[0127] Preferably, in some embodiments, it also includes model evaluation and optimization: Use the validation set or test set to evaluate the trained model to measure its performance. The evaluation metrics can include accuracy, recall rate, F1 score, etc. Optimize the model according to the evaluation results, such as adjusting hyperparameters, changing the model architecture, increasing training data, etc. The optimized model will have better generalization ability and performance.
[0128] According to the above embodiments of the present invention, by introducing a multi-scale patch hybrid network, making full use of the multi-scale characteristics of time series data, the prediction accuracy of the model for the battery health state and remaining life is improved. Through the design of a custom loss function, the model can better distinguish normal and abnormal data, further improving the prediction accuracy. Through self-supervised learning and feature extraction, the generalization ability of the model is improved.
[0129] Experimental comparison:
[0130] Use the quarterly dataset and weekly dataset of the battery for comparison to prove the stability of the method of the present invention in long-term and short-term predictions.
[0131] Calculate the symmetric mean percentage error to measure the prediction error. The result ranges from to, and the smaller the value, the better the prediction effect:
[0132] ,
[0133] where : The total number of samples (the number of prediction points).
[0134] : The true value (the true battery health state or service life).
[0135] : The predicted value (the health state or service life predicted by the model).
[0136] : The absolute error between the predicted value and the true value.
[0137] : The average of the true value and the predicted value, used to normalize the error.
[0138] Calculate the mean absolute scaled error to measure whether the error of the model is lower than the simple average rate of change of the time series. A mean absolute scaled error less than 1 indicates that the model is better than directly using the average rate of change of the time series:
[0139] ,
[0140] where : The total number of samples in the test dataset.
[0141] : The reference sequence length of the seasonal or trend data.
[0142] : The true value.
[0143] : The predicted value.
[0144] : That is, the mean absolute error between the predicted value and the true value.
[0145] That is, the rate of change in the time series (calculate the mean absolute change of historical data).
[0146] Adopt centralized kernel alignment to evaluate the effect of different levels of representation learning, and calculate the centralized kernel alignment value between the hidden layer representation of the present invention and the distribution of the true battery service life , with a value range between 0 and 1. The higher the value, the more similar the two distributions are, indicating that the model has learned features close to the true distribution. By comparing the representation capabilities of other methods, the advantages of the method of the present invention in feature extraction are verified.
[0147] ,
[0148] G: The Gram matrix of the first set of feature representations (output of the neural network).
[0149] Q: The Gram matrix of the second set of feature representations (true battery health state or service life distribution).
[0150] : The trace of the matrix, representing the sum of the eigenvalues of the matrix.
[0151] Table 1 Experimental comparison results
[0152]
[0153] In this experiment on battery energy storage health state monitoring and remaining life prediction, we comprehensively compared and evaluated the method of the present invention with traditional convolutional neural networks. In terms of prediction error, measured by two key indicators, the symmetric mean percentage error and the mean absolute scaled error, the method of the present invention has shown significant advantages. Its prediction error has been effectively reduced compared to convolutional neural networks, indicating that the method of the present invention has significant advantages in prediction accuracy and can more accurately estimate the battery energy storage health state and remaining life.
[0154] The experimental comparison results are shown in Table 1. From the evaluation of representation capabilities, the centralized kernel alignment value of the method of the present invention reaches 0.812, while that of the convolutional neural network is only 0.693. This clearly shows that the method of the present invention has stronger representation capabilities and can better extract and express key features in the data, thus providing better feature representations for subsequent analysis and prediction.
[0155] It should be understood that various forms of the processes shown above can be reordered, steps can be added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. No limitations are imposed herein.
[0156] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An artificial intelligence driven method for predicting the health status and remaining life of battery energy storage, characterized in that: include: Obtain time series data of various characteristics and status information of the battery at historical time points and perform preprocessing; Various characteristics and status information including voltage data, current data, temperature data, capacity data, battery internal resistance, service life and historical cycle count; Inputting the preprocessed time series data into a feature extraction model based on a multi-scale patch mixing network for feature extraction, wherein the multi-scale patch mixing network is composed of a time mixer and a space mixer; Among them, the temporal mixer uses convolution kernels of different sizes to capture short-term and long-term dependencies and extract variable time series information; the spatial mixer performs feature aggregation based on the maximum variance framework to extract multi-scale information on the variable dimension; The time mixer uses convolution kernels of different sizes to capture short-term and long-term dependencies and extract variable time series information, including: The time mixer extracts variable time series information by using linear layers and multi-layer perceptrons along the time dimension, and adds the input and output through short connections or residual connections. The specific formula is: , in, is the input time series feature; It is a linear layer, which is used to perform linear layer normalization on the input time series features; It is a multi-layer perceptron used for feature transformation and extraction, and performs nonlinear mapping on the normalized data to learn deeper representations; is the final output feature, i.e. the extracted variable time series information; In the self-supervised contrastive representation learning stage, the feature extraction model is trained using a triplet loss function to optimize the time series features extracted by the network, and the learned feature representation is used to find semantically meaningful nearest neighbors and farthest neighbors for each sample; In the self-supervised classification stage, classification is performed based on the nearest neighbor and farthest neighbor information of all feature samples, and the classification model is trained through a custom loss function to achieve the best classification effect; Use the trained model to classify and predict the new time series data of the battery and output the battery health status.
2. The method according to claim 1, characterized in that in, Preprocessing of the time series data of various characteristics and status information of the battery at historical time points includes: Use instance normalization to normalize the time series window; Point anomalies and subsequence anomalies are used in the standardized time series data, and the dimensions and starting time are randomly selected and anomaly injected to construct abnormal samples.
3. The method according to claim 2, characterized in that in, The spatial mixer performs feature aggregation based on the maximum variance framework and extracts multi-scale information on the variable dimension including: Data cycles were analyzed using a maximum variance framework to identify periodic patterns in the data; Based on the results of data period analysis, multi-scale patch segmentation is performed, including: determining the patch sizes of different scales according to the data period; dividing the original pre-processed time series data into patches of multiple scales, each patch containing data within a specific time range; For each scale of patch, a parameter-sharing multilayer perceptron is used to extract features along the variable dimension; The multi-scale feature aggregation weight is determined based on the amplitude calculated based on the maximum variance framework, and the multi-scale features are fused according to the calculated multi-scale feature aggregation weight, and added to the output of the time mixer through a short connection. The specific formula is: , in, is the i-th scale patch determined based on the data periodicity analysis; n represents the number of scales; represents the features extracted by applying a multi-layer perceptron to the i-th scale patch; is the aggregation weight of the i-th scale feature determined by the magnitude calculated by the maximum variance framework; is the output of the time mixer; S is the final output of the space mixer.
4. The method according to claim 3, characterized in that in, The use of a maximum variance framework to analyze data cycles to identify periodic patterns in the data includes: Maximum variance framework via matching modified circulant matrices , that is, find a modified circulant matrix , so that the variance of the data CX obtained after the modified circulant matrix C transforms the original data X is maximized, and by maximizing the variance, the most important feature in the time series data, namely the data periodic pattern, is extracted.
5. The method according to claim 2, characterized in that: in, In the self-supervised contrastive representation learning stage, the feature extraction model is trained using a triplet loss function to optimize the time series features extracted by the network, including: The feature extraction model is trained using a triplet loss function to minimize the distance between the anchor sample and the positive sample, while maximizing the distance between the anchor sample and the negative sample, and the difference is at least a preset interval. ; If the distance between the negative sample and the anchor point does not exceed the distance between the positive sample and the anchor point plus , a loss will occur, otherwise the loss is zero, thereby optimizing the time series features extracted by the feature extraction model to better distinguish normal and abnormal samples.
6. The method according to claim 3, characterized in that in, In the self-supervised classification stage, classification is performed based on the nearest neighbor and farthest neighbor information of all feature samples, and the classification model is trained through a custom loss function to achieve the best classification effect, including: The classification model is trained by a custom loss function so that the classification model can accurately classify the nearest neighbor and farthest neighbor information of all feature samples; the loss function maximizes the similarity between the window representation and its nearest neighbor and minimizes the similarity between the window representation and its farthest neighbor. The loss function formula is: , in, Indicates windows; represents the nearest neighbor; represents the farthest neighbor; represents a similarity measure; Control the weight of nearest neighbor similarity optimization; Control the weight of the farthest neighbor similarity optimization; Indicates Windows Its nearest neighbors The similarity between Indicates Windows Its furthest neighbor ; α is a hyperparameter used to control the weight of the regularization term; H(x) represents entropy.
7. The method according to claim 5, characterized in that in, Use the trained model to classify and predict the new battery time series data and output the battery health status, including: After model training is complete, the class assignment and majority class of the dataset are determined through the inference process, including: First, use the trained classification model to classify the battery data, identify normal and abnormal battery data, and then count the proportion of the majority class, that is, normal battery data; Calculate the probability that the new window belongs to the majority class, determine whether the new window is abnormal based on the probability, and output the abnormal score of the battery health status indicator.
8. The method according to claim 1, characterized in that in, The method further comprises: The trained model is used to extract features from the battery health status data of the battery, and the battery health status data after feature extraction is input into the remaining life prediction model for prediction, and the remaining life prediction value of the battery is output. If the remaining life prediction value is less than a certain threshold, an alarm is triggered or a suggestion to replace the battery is made.
Citation Information
Patent Citations
Method for predicting state of charge of lithium battery of new energy electric vehicle
CN117540879A