Fault prediction method for lithium ion battery of energy storage power station
Through the combination of feature selection, fusion and extraction layers, high-quality features are screened out and attention mechanisms are used to improve the accuracy of lithium-ion battery failure prediction, solving the problem of low prediction accuracy in the prior art, and achieving more efficient fault detection.
Patent Information
- Application Number
- CN202510605374.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, the lithium-ion battery fault prediction method based on the deep learning model shows performance degradation when the time series increases, resulting in the prediction accuracy being insufficient and the failure of lithium-ion batteries in energy storage power stations cannot be effectively detected.
A failure prediction method for lithium-ion batteries of energy storage power stations is adopted, and the Pearson correlation coefficient screening features are calculated through the feature selection layer, and the feature fusion layer is used for dimensionality reduction and denoising. The feature extraction layer uses the attention mechanism to screen key features, and outputs the fault prediction results through the full connection layer and the output layer.
It improves the accuracy of lithium-ion battery failure prediction, saves computing resources, and significantly improves the prediction accuracy of the model.
Smart Images

Figure CN120470441A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of battery failure prediction, and in particular to a method for predicting failures of lithium-ion batteries in energy storage power stations. Background Art
[0002] Energy storage power station systems consist of thousands of lithium-ion battery cells, forming a highly complex system. These cells exhibit significant nonlinearity, temperature sensitivity, significant aging characteristics, and inconsistencies. These characteristics can lead to thermal runaway during use due to heat accumulation from chemical reactions within the battery pack or from external heat sources. In severe cases, this can threaten the overall safety of the energy storage station. As a crucial component of grid intelligence, the safety of energy storage systems is directly related not only to the grid's peak and frequency regulation capabilities, but also to the safety of the battery storage devices themselves. Battery cell arrays are connected in various configurations, including series and parallel. However, in actual use, individual battery cells or their components may fail due to battery aging or improper operation. If these failures are not detected and addressed promptly, they can adversely affect the safety of the battery system and, in extreme cases, even lead to thermal runaway.
[0003] In existing data-driven lithium battery fault prediction methods, commonly used deep learning models include recurrent neural networks, long-short-term memory networks, and gated recurrent units. These models directly analyze and process large amounts of offline and online lithium-ion battery operating data, establish a mapping mechanism between lithium-ion battery inputs and outputs, and extract corresponding features for fault diagnosis. This fault diagnosis method does not require the establishment of a precise lithium-ion power battery model. However, the deep learning models used in existing technologies all show performance degradation as the time series increases, resulting in inaccurate predictions.
[0004] Therefore, there is an urgent need for a battery fault prediction method to improve the accuracy of battery fault prediction. Summary of the Invention
[0005] Based on this, it is necessary to provide a method for predicting lithium-ion battery failures in energy storage power stations in response to the above technical problems.
[0006] The present invention adopts the following technical solutions:
[0007] The present invention provides a method for predicting lithium-ion battery failure in an energy storage power station, comprising:
[0008] Obtaining operating data of energy storage power stations;
[0009] Inputting the operating data into a trained lithium-ion battery fault prediction model; the lithium-ion battery fault prediction model includes: a feature selection layer, a feature fusion layer, a feature extraction layer, a fully connected layer and an output layer;
[0010] In the feature selection layer, the Pearson correlation coefficient between the battery voltage feature and other data features in the operating data is calculated to obtain the feature correlation value, and the other data features and their corresponding feature correlation values are mapped to the correlation space. Feature correlation values greater than a preset threshold are screened from the correlation space and the corresponding correlation features are retained.
[0011] In the feature fusion layer, the retained correlation features are reduced in dimension and denoised, and then encoded and decoded to output the fused features;
[0012] At the feature extraction layer, based on the position encoding and fusion features of the local timestamp, fused features with time information are calculated and converted into query vectors, key vectors, and value vectors. The correlation between each query vector and all key vectors is calculated separately, and features are filtered based on the correlation between each query vector and all key vectors. Based on the query vector, key vector, and value vector of the filtered features, the attention score of each filtered feature is calculated, and based on the attention score, each filtered feature is weighted to extract key features.
[0013] The key features are extracted using the ReLU activation function through the fully connected layer, and the battery failure prediction results are output through the output layer.
[0014] Preferably, obtaining the energy storage power station operation data specifically includes:
[0015] Collect the original operating data of the energy storage power station through the battery management system;
[0016] The original operating data is sequentially subjected to missing value filling, outlier detection and standardization to obtain the operating data of the energy storage power station.
[0017] Preferably, calculating the Pearson correlation coefficient between the battery voltage feature and other data features in the operating data to obtain a feature correlation value specifically includes:
[0018] According to the position coding of the running data involving the local timestamp, the time feature is extracted. The formula is:
[0019] PE(pos,2i)=sin(pos / 2L 2i / d );
[0020] PE(pos,2i+1)=cos(pos / 2L 2i / d );
[0021] Where PE is the position code, pos represents the position of the current running data in the entire input sequence, i is the dimension of the current calculation value, ranging from 1 to d, d represents the dimension of the input sequence, and L represents the length of the sequence;
[0022] The Pearson correlation coefficients, i.e., feature correlation values, between the battery voltage feature and the data feature in the operating data, and between the battery voltage feature and the time feature are calculated respectively.
[0023] Preferably, the feature fusion layer includes a first deep autoencoder and a second deep autoencoder; the dimensionality reduction and denoising of the retained correlation features and encoding and decoding to output fusion features specifically include:
[0024] Inputting the retained correlation features into the first deep autoencoder for encoding, and then decoding the encoded features into the retained correlation features to obtain the network parameters of the first deep autoencoder;
[0025] Updating the network parameters in the second deep autoencoder to the network parameters of the first deep autoencoder;
[0026] The retained correlation features are input into the second deep autoencoder, and the first-order fusion features are obtained through the first encoding;
[0027] The first-order fusion features are encoded a second time to obtain fusion features; the fusion features are second-order fusion features.
[0028] Preferably, the encoding and decoding formulas are:
[0029]
[0030]
[0031] In the formula, m represents the number of fusion features, n is the number of input and output nodes, f(.) is the activation function of the node, b1 and b2 are weights, and the n-dimensional correlation feature x is retained in the encoding stage. i Mapped to m-dimensional h through activation function j , in the decoding stage, h j Remapping to n-dimensional preserved correlation features
[0032] Preferably, the fusion feature with time information The calculation formula is:
[0033]
[0034] Where u iRepresents the fusion feature, i∈[1,2,…,L], L is the sequence length of the fusion feature, t is the sequence number, α is the factor of the size between the balance mapping vector and the position encoding, PE (L×(t-1)+i) The position code of the local timestamp, [SE (L×(t-1)+i) ] is a scaling factor, which is set to the sequence length or a fixed constant and is used to adjust the scale of the position encoding to adapt it to different sequence lengths and feature dimensions. P is the abbreviation of position.
[0035] Preferably, feature screening is performed based on the correlation between each query vector and all key vectors, specifically including:
[0036] Calculate the i-th query vector q i With the key vector k of all features j The correlation is calculated as follows:
[0037]
[0038] Where, represents the correlation coefficient between the i-th query vector and the key vectors of all features, d is the dimension of the input feature, L K is the length of the key sequence, Represents the query vector q i With all key vectors k j The maximum dot product of the query vector q i The strength of association with the most relevant key vector; is the average of all dot products, used to measure the query vector q i With the key vector k j Overall relevance of
[0039] Sort the features corresponding to all query vectors in descending order according to relevance;
[0040] Filter out the features corresponding to the first u query vectors, where u = clogL Q , c is the sampling factor, L Q The length of the query.
[0041] Preferably, the calculation formula of the attention score is:
[0042]
[0043] Where A(K,Q,V) is each attention score, where K is the key, Q is the query, V is the value, softmax(.) is the activation function, and d is the input dimension.
[0044] Preferably, weights are assigned to each filtered feature based on the attention score to extract key features, specifically including:
[0045] Through the linear activation function, the attention score of each filtered feature is normalized and converted into a probability distribution, that is, the weight of each feature;
[0046] According to the weight of each feature, all features are weighted and summed, and features with a weight greater than the threshold are determined as key features.
[0047] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects:
[0048] In the field of battery fault prediction, the present invention constructs a lithium-ion battery fault prediction model for an energy storage power station. The feature screening layer calculates the correlation between historical data and target data, and screens out high-quality data features, laying the foundation for improving the accuracy of model prediction. Secondly, the feature fusion layer reduces the dimension and denoises the high-quality feature data, and fuses the high-quality data features through encoding and decoding. Finally, the feature extraction layer filters the query vector of the fused features, significantly reducing the computing resources of the model. The attention score of the input feature is obtained and assigned a corresponding weight through the filtered query vector and its corresponding key and value. The key features in the energy storage power station operation data that are helpful for fault judgment are extracted according to the weight, and the prediction results of the battery fault are output through the fully connected layer and the output layer, which saves computing resources while effectively improving the accuracy of the model prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0050] Figure 1 A schematic diagram of a flow chart of a method for predicting lithium-ion battery failure in an energy storage power station provided by the present invention;
[0051] Figure 2 A diagram showing the arrangement of cells in a battery module of a method for predicting lithium-ion battery failures in an energy storage power station provided by the present invention;
[0052] Figure 3 A schematic diagram of a curve of collected data of an energy storage power station for a method for predicting lithium-ion battery failures in an energy storage power station provided by the present invention;
[0053] Figure 4 A schematic diagram of the correlation coefficients of features in historical data and target data of a method for predicting lithium-ion battery failures in energy storage power stations provided by the present invention;
[0054] Figure 5A schematic diagram of the structure of a feature fusion layer deep autoencoder for a method for predicting lithium-ion battery faults in energy storage power stations provided by the present invention;
[0055] Figure 6 A qualitative analysis diagram of the fusion characteristics of a method for predicting lithium-ion battery failures in an energy storage power station provided by the present invention;
[0056] Figure 7 A schematic diagram of a structure of a lithium-ion battery failure prediction model for an energy storage power station provided by the present invention;
[0057] Figure 8 Comparative model prediction results and prediction error diagram of a lithium-ion battery failure prediction method for an energy storage power station provided by the present invention;
[0058] Figure 9 The prediction results and prediction error diagram of the model of the present invention for the method for predicting lithium-ion battery failure in an energy storage power station provided by the present invention;
[0059] Figure 10 This is a graph showing the prediction results of a No. 60 battery in a method for predicting lithium-ion battery failure in an energy storage power station provided by the present invention on training sets of different sizes;
[0060] Figure 11 A diagram showing prediction results at different time intervals of a method for predicting lithium-ion battery failure in an energy storage power station provided by the present invention;
[0061] Figure 12 This is a result analysis chart of the average percentage error, RMSE and MAE of all data sets of a lithium-ion battery fault prediction method for energy storage power stations provided by the present invention. DETAILED DESCRIPTION
[0062] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in the specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0063] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0064] Figure 1 The figure is a flow chart of a method for predicting lithium-ion battery failure in an energy storage power station according to the present invention, which specifically includes the following steps:
[0065] S101: Obtaining energy storage power station operation data.
[0066] The original energy storage power station operation data is collected through the BMS system; missing values are filled, outliers are detected and standardized on the original energy storage power station operation data to obtain the energy storage power station operation data.
[0067] Specifically, if a certain energy storage power station consists of three battery compartments, see Figure 2 , is a diagram of the arrangement of cells in a battery module. Each compartment contains two stacks (1, 2), each stack contains three clusters, each cluster consists of 19 modules, and each module consists of two parallel cells and 12 series cells. In this study, the term "single cell" refers to "one series cell" in the two parallel and 12 series configuration. Figure 2 Figure 1 illustrates the arrangement of battery cells in the battery module. The rated voltage of each cell is 3.2V and the nominal capacity is 120Ah. The rated voltage of the battery pack is 38.4V and the capacity is 240Ah. A BMS is used to continuously record real-time operating data at 60-second intervals, including battery current, voltage, temperature, and SOC.
[0068] Alternatively, see Figure 3 Figure 2 shows the correlation coefficients of features in historical and target data. The variability and unpredictability of factors such as grid peak regulation and frequency regulation often lead to random fluctuations in most time-varying parameters, including battery voltage (normal and abnormal), cluster voltage, SOC, temperature, and current. As can be seen from the figure, the battery and cluster voltages fluctuate primarily in the range of 3.6V to 4.2V and 50.4V to 58.8V, respectively. Abnormal voltages range from 0V to 3.39V, and temperatures range from 22°C to 28°C. Current jumps are caused by the charge-discharge switching of the energy storage station. The SOC ranges from 17.5% to 86.6%. This phenomenon can be explained by the ancillary services provided by the energy storage station. The confidence level of SOC is relatively low, making it difficult to accurately estimate even using state-of-the-art methods. SOC significantly affects voltage variation, but the relationship is nonlinear and exhibits time lag. This limits the possibility of complete battery discharge. Other factors influencing battery voltage within the energy storage station include current, temperature, and unavoidable aging effects.
[0069] Alternatively, given the wide variety of information described above, there is a certain degree of redundancy. When the dimensionality of the training data is too high or too low, it can complicate the model training process and negatively impact prediction accuracy. Therefore, the goal of this study is to remove features irrelevant to the prediction results, reduce redundant information, and improve data validity through dimensionality reduction, thereby enhancing the prediction accuracy of the subsequent voltage prediction model.
[0070] Specifically, with the continuous advancement of intelligent metering technology, energy storage power station operation and management can monitor and record various data within the system in real time. This enables dynamic tracking of battery status, load conditions, and voltage changes. However, due to factors such as BMS failures and human error, some critical data may be lost or abnormal, affecting the normal operation of the system and the accuracy of decision-making. This requires effective preprocessing of the collected raw data before data analysis and prediction to ensure data integrity, accuracy, and reliability. Furthermore, the different types of data collected by the BMS, such as battery voltage, temperature, and SOC, often have inconsistent dimensions, and intelligent fusion algorithms are sensitive to the numerical range and scale of the input data. Therefore, all collected data must be standardized and quantized before input. Standardization ensures that features across different data dimensions are within a consistent magnitude range, preventing any feature from having an undue impact on the algorithm due to excessive or insufficient scale. Furthermore, effective data preprocessing can significantly improve model learning, optimize algorithm accuracy, and further enhance the reliability of prediction results. Therefore, reasonable data preprocessing is not only the basis for improving model performance, but also a key step in achieving accurate battery voltage prediction and fault diagnosis.
[0071] Alternatively, different evaluation metrics often have different scales and measurement units, and these differences may affect the results of data analysis. To eliminate the influence of these metrics, data standardization measures are required to ensure comparability between metrics. This method standardizes the raw data by scaling them proportionally. Linear transformation is considered an effective method to map the raw data to the range [0, 1]. The normalization formula is:
[0072]
[0073] Where D norm The data collected by the BMS after normalization in the previous section are collectively referred to as raw data. D represents the raw data; D max and D min are the maximum and minimum values of the original data set, respectively.
[0074] S102: Inputting the operating data into a trained lithium-ion battery fault prediction model; the lithium-ion battery fault prediction model includes: a feature selection layer, a feature fusion layer, a feature extraction layer, a fully connected layer and an output layer.
[0075] S103: In the feature selection layer, the Pearson correlation coefficient between the battery voltage feature and other data features in the operating data is calculated to obtain a feature correlation value, and other data features and corresponding feature correlation values are mapped to a correlation space; feature correlation values greater than a preset threshold are screened out from the correlation space and the corresponding correlation features are retained.
[0076] Optionally, the Pearson correlation coefficient between the battery voltage feature and other features in the energy storage power station operation data is calculated, specifically including:
[0077] Position coding involving local timestamps is introduced to extract the time characteristics of energy storage power station operation data. The formula is:
[0078] PE(pos,2i)=sin(pos / 2L 2i / d );
[0079] PE(pos,2i+1)=cos(pos / 2L 2i / d );
[0080] Where PE is the position code of the local timestamp, pos represents the position of the current word in the entire input sequence, i is the dimension of the current calculated value, ranging from 1 to d, and d represents the dimension of the input sequence. L Indicates the length of the sequence;
[0081] Based on multiple data features related to the battery voltage features of the energy storage power station operation data, the Pearson correlation coefficient between the battery voltage features and multiple data features and time features in the historical data and the target data is calculated, and the correlation features of the Pearson correlation coefficient with the voltage data features within a certain range are retained.
[0082] The calculation formula for the Pearson correlation coefficient between the battery voltage feature and multiple data features and the time feature is:
[0083]
[0084] Where ρ represents the Pearson correlation coefficient between the battery voltage feature and multiple data features and time features, n represents the number of multiple data features, and x i Indicates voltage, Represents x i The average value of y i The characteristic variable representing the voltage at time i; represents y i The average value of .
[0085] Specifically, online estimation of battery SOH and RUL remains a challenging problem. To more fully consider the temporal characteristics of battery voltage, three time series features of the voltage data of battery 60 from August 1 to August 31 were extracted based on the above equation. In addition, ten parameters related to battery voltage data were recorded using the BMS, including cell temperature, stack voltage, stack current, stack SOC, cluster voltage, cluster current, and cluster SOC. In order to extract the most valuable features from these parameters and quantify their correlation with voltage, the PCC method was adopted to eliminate harmful parameters and extract high-quality parameters as input to the voltage prediction model.
[0086] Specifically, the range of ρ is -1 to 1, and the absolute value of ρ reflects the strength of the correlation between the characteristic variable and the voltage. An absolute value close to 1 indicates that the linear correlation between the two parameters is strong. The PCC values of the selected 11 parameters are shown in Figure 4 , which shows the correlation coefficients of features in historical and target data. Selecting highly correlated parameters can reduce the computational burden while maintaining high predictive performance. The PCC values for the battery pack total voltage, cluster total voltage, cluster SOC, and daily information are all greater than 0.85, indicating strong linear correlation.
[0087] S104: In the feature fusion layer, the retained correlation features are subjected to dimensionality reduction and denoising, and are encoded and decoded, and the fused features are output.
[0088] In the feature fusion layer, the retained correlation features are reduced in dimension and denoised, and the features after dimension reduction and denoising are encoded and decoded to output the fused features.
[0089] The feature fusion layer is constructed by two deep autoencoders, including an input layer, a first hidden layer, and a second hidden layer. The retained correlation features are input to the input layer for encoding and decoding, and the network parameters between the input layer and the first hidden layer are determined. The first-order fusion features are output in the first hidden layer through the network parameters between the input layer and the first hidden layer. The first-order fusion features are encoded and decoded to determine the network parameters between the first hidden layer and the second hidden layer. The second-order fusion features, i.e., fusion features, are output based on the network parameters between the first hidden layer and the second hidden layer.
[0090] Among them, the encoding and decoding formulas are:
[0091]
[0092] In the formula, m represents the number of fusion features, n is the number of input and output nodes, f(.) is the activation function of the node, b1 and b2 are weights, and the encoding stage converts the n-dimensional input x i Mapped to m-dimensional h through activation functionj , in the decoding stage, h j Remap to n-dimensional
[0093] Specifically, the deep autoencoder is an unsupervised neural network that can be used for dimensionality reduction and feature extraction. The feature fusion layer is used to process the selected highly relevant features to reduce the dimensionality and noise. The specific structure of the feature fusion layer is as follows: Figure 5 (a) shows that the two feature fusion layers are combined into a stacked autoencoder for second-order feature fusion. The structure of the stacked autoencoder is as follows: Figure 5 (b) The fusion process is as follows: First, the retained correlation features are used as Figure 5 The input and output of the structure in (a) are encoded and decoded to obtain the input layer to h (1) The network parameters of the layer. Figure 5 (b) The parameters of the structure are updated to the network parameters, and the hidden layer output is the first-order fusion result. Then, the result is used as Figure 5 (b) Structure the input of the second hidden layer and set the number of nodes in the second hidden layer to 1. Finally, a sequence is obtained as the fusion result. The output of the hidden layer is the second-order fusion feature, which is recorded as HI. n ]and Represent the input and output of the network respectively; W ij and W jk is the weight of the connection. The node output of the hidden layer is h j , the output layer node is
[0094] Specifically, to further explore the impact of high-quality feature fusion HI on the quantitative characterization of battery voltage data, this paper conducted quantitative and qualitative analysis. The quantitative analysis results show that the PCC value between the fused features and the voltage data is 0.9914, which is greater than the PCC value between any selected features and voltage, indicating that there is a significant correlation between the two. The qualitative analysis results can be found in Figure 6 As can be seen from the figure, the curve after fusion of HI is highly consistent with the change trend of the original voltage curve. Combining the results of quantitative and qualitative analysis, it can be concluded that the fusion feature can effectively represent the various states of voltage data.
[0095] S105: In the feature extraction layer, based on the position encoding and fusion features of the local timestamp, the fusion features with time information are calculated and converted into query vectors, key vectors and value vectors; the correlation between each query vector and all key vectors is calculated respectively, and features are filtered based on the correlation between each query vector and all key vectors; based on the query vector, key vector and value vector of the filtered features, the attention score of each filtered feature is calculated, and according to the attention score, each filtered feature is weighted to extract key features.
[0096] In the feature extraction layer, the fused features are input into the embedding layer of the feature extraction layer. The embedding layer obtains the input features through calculation and obtains the query vector, key vector and value vector of the input features. The query vector, key vector and value vector of the input features are filtered by calculating the correlation between each query and all keys. The attention score of the input features is calculated by the query vector, key vector and value vector of the filtered input features. According to the attention score, different weights are assigned to the input features to capture the key features in the input features. The starting mark of the key feature is input into the decoder of the feature extraction layer for decoding, and the multi-head attention of the feature extraction layer is masked so that each position of the key feature is only associated with the information of the current position.
[0097] The fused features are input to the embedding layer of the feature extraction layer. The embedding layer obtains the input features through calculation, specifically including:
[0098] The position encoding and fusion features involving the local timestamp are input to the embedding layer of the feature extraction layer. The embedding layer obtains the input features of the encoding layer through calculation. The formula is:
[0099]
[0100] Where ui represents the fusion feature, i∈[1,2,…,L], L is the sequence length of the fusion feature, t is the sequence number, α is the factor of the size between the balance mapping vector and the position encoding, PE (L×(t-1)+i) The position code of the local timestamp, [SE (L×(t-1)+i) ] is a scaling factor, usually set to the sequence length or a fixed constant, used to adjust the scale of the position encoding to adapt it to different sequence lengths and feature dimensions. P is the abbreviation of position.
[0101] Calculate the attention score of the input features, specifically including:
[0102] Define the i-th query vector as q i , calculate q i With all key vectors k j The relevance of the query and the importance of all queries are obtained. The formula is:
[0103]
[0104] Where, Represents the query q i The importance of d is the dimension of the input feature, L K is the length of the key sequence, Represents the query q i With all keys k j The maximum dot product of the query q i the strength of the association with the most relevant key; is the average of all dot products, used to measure the query q i With key k j Overall relevance of
[0105] Sort all queries in descending order of importance, select the first u queries and their corresponding keys and values to calculate the attention score of each query, where u = clogL Q , c is the sampling factor, L Q is the length of the query;
[0106] By activating the softmax function, the attention scores of the first u queries and their corresponding keys and values are calculated as follows:
[0107]
[0108] Where A(K,Q,V) is the attention score of the query vector, key vector and value vector, where K is the key, Q is the query, V is the value, softmax(.) is the activation function, and d is the input dimension.
[0109] Based on the attention score, different weights are assigned to the input features to capture the key features in the input features, including:
[0110] The model converts the input data into three vectors: query, key, and value. The query can be thought of as a question, the key as the clue to answer the question, and the value as the answer itself.
[0111] The model then calculates the similarity between the query and the key, which simply means how closely the query matches each key (clue). Features with higher matching scores are considered more important.
[0112] To make these similarity values easier to understand, the model uses a method called Softmax to convert all similarities into a probability distribution. In this way, each feature will get a weight, and the higher the weight of the feature, the more important it is.
[0113] Based on these weights, the model performs a weighted summation of the actual information (value vectors) of all input features. Simply put, the most important feature information is amplified, while the unimportant ones are relatively weakened, ultimately resulting in a new representation.
[0114] (5) Efficient processing of long sequences:
[0115] To process very long time series data, the feature extraction layer uses a sparse attention mechanism that focuses only on key parts rather than every detail, thereby improving computational efficiency. This makes the model more efficient and accurate in long time series tasks.
[0116] The step of associating each position of the key feature with only the information of the current position specifically includes:
[0117] When performing self-attention calculations, instead of considering every position in the input sequence, it focuses only on the parts most relevant to the current time step. This allows the model to avoid excessive computation when processing long time series, saving computing resources while effectively capturing key information.
[0118] In traditional self-attention mechanisms, each position interacts with all other positions. By limiting each position to associate only with nearby key positions, or selecting the most important time steps for calculation based on certain rules, this approach not only improves computational efficiency but also enables the model to focus more on data at key time steps, reducing interference from irrelevant information.
[0119] By design, the features of each position are allowed to be associated with its local time step and some global time steps. This not only preserves the information of local features, but also captures the long-range dependencies in the global sequence. In short, although the model limits direct interaction with all positions, it can still identify which global dependencies are important and pay attention to them accordingly.
[0120] By learning the characteristics of the input sequence, it can effectively select features that have an important impact on the current prediction task and pay more attention to these features. In this way, it avoids the interference of irrelevant features and ensures that the model can focus on processing the most critical time step data.
[0121] Specifically, since the position information of the data cannot be directly recorded in the input, position embedding is required. The feature extraction layer incorporates the position encoding into the data input to ensure that the model captures the correct order of the input sequence. This position encoding is divided into a local timestamp and a global timestamp. The equations for the local timestamp are provided in equations (2) and (3).
[0122] Specifically, self-attention is an important feature extraction mechanism used in the encoding and decoding process. The original self-attention mechanism is designed to map the query vector, key vector, and value vector vector to the output. The weights are calculated by applying a specific formula to calculate the relevance of the query to the corresponding key, and then the weighted sum of these values is calculated according to these weights;
[0123] Specifically, according to previous experimental results, the score of self-attention follows a long-tail distribution, that is, a few dot products contribute most of the attention. Therefore, identifying attention is very important to reduce computational complexity, which is also the core strategy of probabilistic sparse self-attention. We define the i-th query as q i , and q i The probability of paying attention to all keys is defined as p(k j |q i ). When q i When the contribution to attention is not important, p(k j |q i ) will be close to uniform distribution u(kj|qi)=1 / L k ..
[0124] In order to reduce the computational complexity, M(q i ,K) requires traversing all dot products of the query and the key, and the complexity of this process is quadratic O(L Q L K In addition, the LSE operation (M(q i , the first part of K) is not stable enough for numerical calculation. Therefore, two methods are used to reduce the complexity: approximate simplification and sampling calculation. i ,K) has obvious upper and lower bounds, the formula is:
[0125]
[0126] Due to the existence of long-tail distribution, there is no need to evaluate for each query We can randomly sample N = ln K Ln Q Calculate for And set the residual value to zero. This method makes the numerical calculation more stable and reduces unnecessary traversal. K =L Q =L, then the time and space complexity is O(LlnL).
[0127] Specifically, the encoder is mainly responsible for extracting dependencies from lengthy data sequences. The probabilistic sparse self-attention mechanism processes sequences containing redundant V vectors. Therefore, in this case, the distillation operation gives higher weight to basic features. The distillation process formula from stage j to j+1 is:
[0128]
[0129] Among them, among them[.] AB Represents the key operation in the multi-head probabilistic sparse self-attention mechanism. ELU is an activation function. The effect of convolution and pooling is to halve the input sequence, thereby reducing memory utilization to O((2-λ)LlogL).
[0130] Optionally, the decoder receives the following input vector:
[0131]
[0132] in represents the input of the decoder, is the start marker of the sequence, A placeholder representing the target sequence. To maintain the consistency of the input dimension, the timestamp is padded with zeros. Next, masked multi-head attention is used to prevent autoregression to ensure that each position only focuses on information related to the current position. Finally, the final result is obtained. The voltage dataset we use has a high sampling frequency. Based on the correlation heat map, we selected data that is strongly correlated with the input data in order to train and predict the model for long sequence data. The prediction length was verified to be 4012 sample points. The results show that the model can effectively process and predict data sequences with long sequences and large data features. After multiple tests, convergence was achieved in the 10th round of the dataset. Given the characteristics of the battery voltage data of the energy storage power station, traditional methods cannot quickly complete model training when faced with newly generated data. Therefore, while ensuring prediction accuracy, reducing the model size and computing runtime, three encoder layers and one decoder layer are selected.
[0133] After the self-attention mechanism and other processing modules (such as positional encoding and sparse attention), the resulting feature information is passed to the fully connected layer. The fully connected layer maps the high-dimensional features of different layers in the network to the final output space. Through the fully connected layer, the model can synthesize the previously extracted features and extract a more compact and effective representation. This layer typically contains multiple neurons, each responsible for processing a portion of the information and passing it to the next layer.
[0134] S106: Extract key features using the ReLU activation function through the fully connected layer, and output the battery fault prediction results through the output layer.
[0135] The output of the fully connected layer generally passes through the activation function ReLU (Rectified Linear Unit) to introduce nonlinear relationships, allowing the network to capture more complex patterns.
[0136] The use of activation functions enables the network to better adapt to nonlinear voltage change patterns, especially in voltage prediction tasks where battery voltage fluctuations may have complex nonlinear characteristics.
[0137] After being processed by the fully connected layer, the final result is passed to the output layer. The function of the output layer is to output the final prediction result according to the task requirements.
[0138] For voltage prediction tasks, the output layer is typically a regression layer that outputs continuous voltage values. Specifically, the output layer generates a predicted value corresponding to the actual battery voltage.
[0139] The dimensionality of the output layer generally matches the number of prediction targets. For example, if the task is to predict voltage changes over a period of time in the future, the output layer may have multiple neurons, one for each predicted voltage value at a time step.
[0140] (4) Generation of prediction results:
[0141] Through the above steps, the model will generate a prediction result based on the input battery data. This prediction result is the model's estimate of the voltage at the future time step.
[0142] For different prediction tasks, such as one-step prediction and multi-step prediction, the design of the output layer may be different, but the overall process is similar: after the fully connected layer processes the feature information, the final prediction result is given through the output layer.
[0143] (5) Post-processing and evaluation:
[0144] The output prediction results usually undergo certain post-processing, such as smoothing or error correction, to further improve the accuracy of the prediction.
[0145] In addition, the performance of the model will be quantitatively evaluated through a variety of error evaluation indicators (such as RMSE, MSE, MAE, etc.) to verify the effectiveness of the model in actual tasks.
[0146] Specifically, during the model training process, the loss function plays a vital role, essentially reflecting the error of the network. The smaller the value of the loss function, the better the performance of the network in solving the problem. Therefore, choosing an appropriate loss function is crucial to ensure that the network parameters are optimized in a more reasonable direction. There are many loss functions to choose from, including absolute value loss function, mean square loss function, cross entropy loss function, etc. In this experiment, we used MSE. The specific expression of the mean square loss function is as follows:
[0147]
[0148] In the above formula, n is the amount of data of the energy storage power station operation data, Ui Represents the true value. Represents the predicted value
[0149] Optionally, key model parameters include the historical data length (h-len), initial learning rate, dropout rate, number of encoder layers, and batch size. The basic hyperparameters and the prediction results of the dataset discussed in the "Decoder" section of this article preliminarily determined the model structure and hyperparameters. The experimental model parameters are shown in Table 1:
[0150] Table 1 Parameters of voltage anomaly prediction model
[0151] Batch size 64 period 10 Activation Function ELU Learning rate 0.001 Encoder layer 3 Decoder layer 1 Temporal feature encoding hour Discard unit 0.1 Loss Function mse h-len 12
[0152] Specifically, BO neural networks contain several hyperparameters, including the loss function, the number of encoder layers, the number of decoder layers, h-len, the learning rate, the dropout rate, and the batch size. These parameters are typically set empirically, making optimal selection difficult. These hyperparameters significantly impact the runtime and prediction accuracy of the neural network, making their optimization crucial. This study focuses on the historical data length, the learning rate, the dropout rate, and the batch size as hyperparameters of the neural network. The BO algorithm is used to efficiently identify the optimal configuration of these hyperparameters. h-len determines the number of past time steps that the model considers during prediction.
[0153] An appropriate h-len enables the model to accurately capture long-term dependencies and trends within the time series, thereby enhancing forecast accuracy. Bayesian optimization is a method for optimizing black-box functions by constructing a Gaussian process model. Its fundamental principle is to select the parameter values most likely to lead to optimization at each iteration based on the current Gaussian process model. It leverages Bayes' theorem to update the model's prior probability distribution using information gained from function evaluation. This allows BO to select the next sampling point based on insights provided by the current Gaussian process model, iteratively enhancing the black-box function.
[0154] Additionally, model parameters were optimized using Bayesian optimization, including: Similar to the original model, the model optimized using the BO neural network also includes certain parameters from the original model. However, BO specifically optimizes parameters such as h-len, initial learning rate, dropout rate, and batch size. To ensure effective optimization and feasibility of the resulting parameters, it is crucial to set appropriate ranges for the hyperparameters, as shown in Table 2.
[0155] Table 2 Hyperparameter range and results
[0156] Hyperparameters Range domain Parameter results h-len [12,24] 16 Learning rate [0.0001,0.01] 0.0045 Packet loss rate [0.1,0.5] 0.2 Batch size [12,64] 30
[0157] The final lithium-ion battery failure prediction model structure can be found in Figure 7,First, the first layer is the data preparation layer.,The main task of this layer is to preprocess and integrate the data.
[0158] Data preprocessing: Before data is fed into the model, it must be preprocessed to meet the model's input requirements. The data is reshaped into a three-dimensional form: [sample size, time step, high-quality feature dimensions]. In this model, the time step is set to 24 (based on voltage predictions at 1-minute intervals). After determining the time step, each data entry is iterated over and divided into multiple overlapping windows with the time step as the window width. The dataset is then divided into training and test sets according to a specific ratio. Furthermore, the data is normalized and corrected for outliers to eliminate the influence of unit differences between features.
[0159] PCC feature selection: Preprocessed data is input into the PCC layer for feature selection. The PCC layer consists of two layers of correlation calculation and a feature selection mechanism. First, features are mapped into a correlation space by calculating the Pearson correlation coefficient between the input data and the target data. Next, the second layer further evaluates the feature correlation values generated by the first layer, selecting high-quality parameters with a correlation between 0.85 and 1 with the target variable. Finally, the feature selection mechanism retains these highly correlated parameters and removes irrelevant features, thereby improving the model's predictive ability and ensuring training effectiveness.
[0160] Feature Fusion: Following the PCC layer is the SAE layer, which consists of an input layer, a hidden layer, and an output layer. The SAE layer performs dimensionality reduction and denoising on the selected highly correlated parameters. The high-quality parameters processed by the PCC layer serve as the input and output of the SAE layer. After encoding and decoding through the network, the hidden layer ultimately outputs the fused parameters.
[0161] The second layer is the feature extraction layer. The core of this layer is the feature extraction layer model architecture, which consists of three encoding layers and one decoding layer. The feature extraction layer utilizes a self-attention mechanism to effectively capture the relationship between global and local information when processing long time series data, overcoming the limitations of traditional LSTM models in long-series learning. First, the feature maps fused by the feature fusion layer are input into the encoding layer of the feature extraction layer. After processing through the three encoding layers, intermediate outputs are generated. The decoding layer then processes these intermediate outputs and, incorporating a self-attention mechanism, assigns weights to different inputs, focusing resources on critical information. Finally, the fully connected layer and output layer generate the final prediction results. Within this layer, the feature extraction layer's hyperparameters are optimized using a Bayesian optimization algorithm to improve model training efficiency and prediction accuracy. This optimization process adjusts the values of multiple hyperparameters to better adapt the model to the data characteristics.
[0162] The third layer is the Bayesian optimization layer. To further enhance the model's predictive performance, a Bayesian optimization algorithm is used to optimize key hyperparameters in the feature extraction layer, including the learning rate, packet loss rate (dropout), historical data length (h-len), batch size (batch size), and the number of neurons in the dense layer (denseunit). Optimizing these hyperparameters effectively improves the model's generalization and stability.
[0163] Error calculation is a key step in evaluating model prediction accuracy, and the accuracy of the results directly reflects the model's performance. To comprehensively measure the deviation between the predicted and true values, this study used RMSE, MSE, MAE, and maximum absolute error as standard evaluation metrics. These metrics quantify the distribution characteristics of prediction errors from different perspectives, providing a multi-dimensional reference for model optimization.
[0164] RMSE calculates the square root of the mean square of the prediction error and can effectively reflect the impact of large errors. It is suitable for scenarios that are sensitive to outliers. MSE amplifies the error through squaring operations and is particularly suitable for loss functions used in regression tasks. MAE directly calculates the absolute difference between the predicted value and the true value. It is less sensitive to outliers and can more robustly reflect the average level of error. The calculation formulas for these indicators are as follows:
[0165]
[0166] MaxAE=Max|V True -V Predict |;
[0167] The BMS dataset for an energy storage power station was sampled for 60 seconds, with 11,671 data points from battery No. 60 between August 3, 2023, and August 11, 2023, serving as the dataset. Abnormal voltage faults were identified in the dataset from 22:18:48 on August 8, 2023, to 9:51:49 on August 9, 2023, and from 21:40:48 on August 9, 2023, to 10:24:49 on August 10, 2023. To validate the generalizability of the proposed model, a comparative experiment was conducted using a dataset from batteries No. 3, No. 5, and No. 6 at another energy storage power station. During the model training phase, the dataset was divided into training and test sets, with 70% of the data used for training and the remaining 30% for testing. During training, the network continuously adjusted parameters to minimize error, while the test data was used to evaluate the model's generalization and predictive performance.
[0168] Figure 8(a) shows the performance of a comparison neural network without hyperparameter optimization in predicting battery voltage anomalies. There is still a large deviation between the model's predictions and the actual observed data. Throughout the prediction process, as time passes, errors gradually accumulate, resulting in a continuous increase in the deviation between the later predicted values and the true values. This error accumulation phenomenon indicates that although there is a certain degree of prediction accuracy during the model training phase, the model's predictive ability decreases over time. Figure 8 (b) further reveals the absolute error between the predicted results and the actual data. Although the model was able to maintain a small error range in the early stages of training, the absolute error increased significantly as the prediction time continued, especially when the voltage fluctuated violently. This shows that when the voltage value suddenly changed, the model's predictive ability was greatly challenged. Overall, the absolute error of the prediction results ranged from 0 to 165mV, and the highest absolute error occurred at 7423 minutes, reaching 1030mV. Such an error range may reflect the overfitting phenomenon caused by improper hyperparameter configuration, resulting in the model's failure to generalize effectively. To overcome this problem, this section optimizes the hyperparameters through the Bayesian optimization algorithm, thereby improving the predictive performance of the feature extraction layer model.
[0169] Figure 9 The prediction results of the neural network model after training on the first 70% of the experimental data after adjusting the hyperparameters are shown. Compared with the comparison model, the model of the present invention has significantly improved in prediction accuracy. During the verification process, the optimized model showed a lower absolute error range, reduced to between 0 and 43mV. Although at some moments, the maximum absolute error of the prediction still reached 890mV, which occurred at a time point of 7423 minutes, this result was significantly more accurate and reliable than the unoptimized model. The introduction of BO effectively reduced the fluctuation of the error, improved the stability and accuracy of the model, and demonstrated the important role of the Bayesian algorithm in improving model performance and providing more accurate predictions.
[0170] In prediction tasks, data-driven methods are not only dependent on the construction of the prediction model itself, but are also significantly affected by the dataset partitioning strategy. The effective partitioning of the dataset is one of the key factors for successful prediction. In particular, the size of the training set has an important impact on the prediction accuracy of the model. Generally speaking, larger training sets help improve the prediction accuracy of neural network models, because more training data allows the model to better learn the patterns in the data. However, the relationship between the size of the training set and the performance of the model does not increase linearly. An excessively large training set may lead to a decrease in computational efficiency and may even cause problems such as overfitting. Therefore, a reasonable selection of the size of the training set is the key to ensuring prediction accuracy. In this study, training sets of 25%, 50%, and 75% were used, respectively. Figure 10 (a) to Figure 10 (c) As shown. It can be clearly seen that when the training set size is 50% and 75%, the prediction accuracy of the model is significantly better than that of the 25% training set. However, when the training set size is increased from 50% to 75%, the prediction accuracy of the model is not significantly improved. Specifically, the MSE is reduced by 866.79mV and 864.68mV, respectively, the RMSE is reduced by 727.568mV and 725.89mV, and the MAE is improved by 59.68% and 60.63%. These changes indicate that although increasing the size of the training set generally improves the learning ability of the model, especially learning more complex patterns, in this case, the accuracy improvement brought about by increasing the training set from 50% to 75% is relatively small, and the increased computational cost does not seem to be so cost-effective.
[0171] In order to further verify the robustness of the model, the present invention uses a No. 5 battery test data set from another energy storage power station, whose rated battery voltage is 3.1V. Figure 10 (d) to Figure 10 (f) shows the prediction results for this dataset. The neural network model provided by the present invention demonstrates extremely high accuracy when predicting voltage anomalies. Compared to the actual voltage values, the predicted values have a minimal error rate, with an average error sum (MSE) of approximately 1.9 mV and a maximum error sum (MAE) of approximately 9.06%. These results demonstrate the model's ability to maintain high-precision predictions despite variations in battery parameters, demonstrating its robustness and broad adaptability under diverse conditions.
[0172] Sampling the initial BMS data every minute has a significant impact on the training efficiency of the neural network model. The size of the dataset is directly related to the ability of the neural network to effectively learn and obtain the expected results. Generally, larger datasets can provide more samples, which helps the training and generalization ability of the model. However, overly large datasets may introduce additional noise and increase the complexity of model training, which may affect the accuracy of predictions. Therefore, this study explored data sampling strategies at different time intervals, including sampling intervals of 2, 3, 4, 5, 6, and 7 minutes. To improve operational efficiency, a strategy of model training based on the first 50% of the dataset was adopted, and a one-step prediction method was used. Figure 11 The prediction results at different time intervals are shown.
[0173] In the model's experiments, the sampling interval was set to 2 minutes and 3 minutes, which showed good results in predicting short-term battery anomalies in energy storage facilities. However, insufficient data sampling resulted in the neural network model showing low prediction accuracy when predicting abnormal feature changes, such as Figure 11 (c) to Figure 11(f) As shown. For training set ratios exceeding 25%, the MSE, RMSE, and MAE indicators initially show a gradual downward trend as the sampling interval increases. This indicates that when the proportion of training data is high, a larger sampling interval can improve the stability and prediction accuracy of the model. However, when the training set ratio is 25%, the error fluctuates greatly due to the exclusion of some abnormal voltage data, resulting in the accumulation of errors over time. Although the results of this study emphasize the advantages of shorter sampling intervals (such as 2-3 minutes) in anomaly detection, they also point out the factors that need to be weighed when selecting the optimal sampling interval. This is not only to improve the accuracy of the prediction, but also to consider the noise level, computing resource limitations, and the specific requirements of the energy storage system. By reasonably selecting the sampling interval, the computing efficiency and the accuracy of the model prediction can be effectively balanced, thereby optimizing the operating performance of the energy storage system.
[0174] For a more comprehensive model evaluation, a detailed dissection of the results from the perspective of error analysis is required. Figure 12 The performance of all datasets in terms of mean percentage error, RMSE, and MAE is shown. MAE more accurately reflects the true forecast deviation when measuring error because it avoids the influence of error drift, while MSE is more sensitive to increases in error and can therefore capture large changes in error. RMSE is consistent with the scale of the data, making its results more intuitive and easier to understand, facilitating the interpretation of changes in overall forecast accuracy. Table 3 lists the specific values of various error evaluation metrics.
[0175] As the training set ratio increased from 25% to 50%, the model's prediction accuracy improved significantly. However, when the training set ratio was further increased to 75%, the error evaluation indicators did not show significant improvement. It is worth noting that at a sampling interval of 4 minutes, the model's prediction performance reached its best, with MAE of 233.873mV and 31.997mV, RMSE of 515.683mV and 61.692mV, and MSE of 265.928mV and 3.806mV. These values showed the lowest error under different training settings. Figure 12 The results show that smaller data sets fail to effectively capture the true characteristics of the data, resulting in a decrease in prediction accuracy. In contrast, a sampling interval of 4 minutes seems to provide the best data distribution and quality balance for the model of the present invention, resulting in the model showing the lowest MAE when predicting voltage changes. From the overall error analysis, it can be seen that there is a complex relationship between the training set size, sampling interval and prediction accuracy. The analysis results show that choosing a sampling interval of 4 minutes, combined with a training set ratio of 50%, can lead to more accurate prediction performance. The prediction errors in different data amounts and training sets are shown in Table 3:
[0176] Table 3 Specific values of various error evaluation indicators
[0177]
[0178] The proposed model optimizes the hyperparameters of the feature extraction layer model using the BO algorithm and combines it with the SAE model to process input data. This hybrid model demonstrates excellent voltage prediction capabilities. Training and validation results on an earlier dataset demonstrate strong prediction accuracy and good generalization for battery voltage prediction. Compared to traditional models without hyperparameter optimization, experimental results show that the proposed hybrid model outperforms the individual models across multiple performance metrics. This demonstrates that the hybrid model not only performs well in voltage prediction but also validates its advantages and effectiveness in battery fault prediction. Furthermore, this chapter explores the impact of the training set ratio and sampling interval on prediction accuracy. Analysis results show that using a dataset with a 4-minute interval and setting the training set size to 50% significantly improves prediction accuracy. In real-world energy storage plant operations, tasks such as peak shaving and frequency regulation are common, leading to complex voltage fluctuations. Future research could investigate extending the model to diagnose specific types of voltage anomalies and enhance fault detection capabilities. Furthermore, consideration could be given to exploring the model's adaptability to voltage prediction for other battery systems.
[0179] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.
Claims
1. A method for predicting lithium-ion battery failure in an energy storage power station, characterized in that: include: Obtaining operating data of energy storage power stations; Inputting the operating data into the trained lithium-ion battery failure prediction model; The lithium-ion battery fault prediction model includes: a feature selection layer, a feature fusion layer, a feature extraction layer, a fully connected layer and an output layer; In the feature selection layer, the Pearson correlation coefficient between the battery voltage feature and other data features in the operating data is calculated to obtain the feature correlation value, and the other data features and their corresponding feature correlation values are mapped to the correlation space. Feature correlation values greater than a preset threshold are screened from the correlation space and the corresponding correlation features are retained. In the feature fusion layer, the retained correlation features are reduced in dimension and denoised, and then encoded and decoded to output the fused features; At the feature extraction layer, based on the position encoding and fusion features of the local timestamp, fused features with time information are calculated and converted into query vectors, key vectors, and value vectors. The correlation between each query vector and all key vectors is calculated separately, and features are filtered based on the correlation between each query vector and all key vectors. Based on the query vector, key vector, and value vector of the filtered features, the attention score of each filtered feature is calculated, and based on the attention score, each filtered feature is weighted to extract key features. The key features are extracted using the ReLU activation function through the fully connected layer, and the battery failure prediction results are output through the output layer.
2. A method for predicting lithium-ion battery failure in an energy storage power station according to claim 1, characterized in that: The obtaining of energy storage power station operation data specifically includes: Collect the original operating data of the energy storage power station through the battery management system; The original operating data is sequentially subjected to missing value filling, outlier detection and standardization to obtain the operating data of the energy storage power station.
3. The method for predicting lithium-ion battery failure in an energy storage power station according to claim 1, wherein: Calculating the Pearson correlation coefficient between the battery voltage feature and other data features in the operating data to obtain a feature correlation value specifically includes: According to the position coding of the running data involving the local timestamp, the time feature is extracted. The formula is: <h2 style=";text-align:left;direction:ltr">PE(pos,2i) = sin(pos / 2L)<h2 style=";text-align:left;direction:ltr"> 2i / d <h2 style=";text-align:left;direction:ltr"> ); PE(pos,2i+1)=cos(pos / 2L 2i / d ); Where PE is the position code, pos represents the position of the current running data in the entire input sequence, i is the dimension of the current calculation value, ranging from 1 to d, d represents the dimension of the input sequence, and L represents the length of the sequence; The Pearson correlation coefficients, i.e., feature correlation values, between the battery voltage feature and the data feature in the operating data, and between the battery voltage feature and the time feature are calculated respectively.
4. The method for predicting lithium-ion battery failure in an energy storage power station according to claim 1, wherein: The feature fusion layer includes a first deep autoencoder and a second deep autoencoder; the dimension reduction and denoising of the retained correlation features are performed and encoded and decoded, and the fusion features are output, specifically including: Inputting the retained correlation features into the first deep autoencoder for encoding, and then decoding the encoded features into the retained correlation features to obtain the network parameters of the first deep autoencoder; Updating the network parameters in the second deep autoencoder to the network parameters of the first deep autoencoder; The retained correlation features are input into the second deep autoencoder, and the first-order fusion features are obtained through the first encoding; The first-order fusion features are encoded a second time to obtain fusion features; the fusion features are second-order fusion features.
5. The method for predicting lithium-ion battery failure in an energy storage power station according to claim 4, wherein: The encoding and decoding formulas are: In the formula, m represents the number of fusion features, n is the number of input and output nodes, f(.) is the activation function of the node, b1 and b2 are weights, and the n-dimensional correlation feature x is retained in the encoding stage. i Mapped to m-dimensional h through activation function j , in the decoding stage, h j Remapping to n-dimensional preserved correlation features 6. The method for predicting lithium-ion battery failure in an energy storage power station according to claim 1, wherein: The fusion feature with time information The calculation formula is: Where u i Represents the fusion feature, i∈[1,2,…,L], L is the sequence length of the fusion feature, t is the sequence number, α is the factor of the size between the balance mapping vector and the position encoding, PE (L×(t-1)+i) The position code of the local timestamp, [SE (L×(t-1)+i) ] is a scaling factor, which is set to the sequence length or a fixed constant and is used to adjust the scale of the position encoding to adapt it to different sequence lengths and feature dimensions. P is the abbreviation of position.
7. The method for predicting lithium-ion battery failure in an energy storage power station according to claim 1, wherein: The feature screening is performed based on the correlation between each query vector and all key vectors, specifically including: Calculate the i-th query vector q i With the key vector k of all features j The correlation is calculated as follows: Where, represents the correlation coefficient between the i-th query vector and the key vectors of all features, d is the dimension of the input feature, L K is the length of the key sequence, Represents the query vector q i With all key vectors k j The maximum dot product of the query vector q i The strength of association with the most relevant key vector; is the average of all dot products, used to measure the query vector q i With key vector k j Overall relevance of Sort the features corresponding to all query vectors in descending order according to relevance; Filter out the features corresponding to the first u query vectors, where u = clogL Q , c is the sampling factor, L Q The length of the query.
8. The method for predicting lithium-ion battery failure in an energy storage power station according to claim 1, wherein: The calculation formula of attention score is: Where A(K,Q,V) is each attention score, where K is the key, Q is the query, V is the value, softmax(.) is the activation function, and d is the input dimension.
9. The method for predicting lithium-ion battery failure in an energy storage power station according to claim 1, wherein: According to the attention score, each filtered feature is weighted and key features are extracted, specifically including: Through the linear activation function, the attention score of each filtered feature is normalized and converted into a probability distribution, that is, the weight of each feature; According to the weight of each feature, all features are weighted and summed, and features with a weight greater than the threshold are determined as key features.