Energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention
By constructing a knowledge graph of energy storage battery failure and utilizing the graph self-attention mechanism, the problem of low efficiency of battery failure analysis in traditional methods is solved, and efficient and accurate battery failure prediction is achieved.
Patent Information
- Application Number
- CN202510669314.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-23
AI Technical Summary
Existing energy storage battery failure analysis methods rely on traditional experiments and experience, resulting in inefficient and costly data processing, and fail to fully consider the correlation between battery property information.
A method based on neighborhood aggregation and graph self-attention is adopted to construct a knowledge graph of energy storage battery failure, map entities and relationships into low-dimensional vectors, aggregate node information using the graph self-attention mechanism, generate a topological representation, and train a classifier to predict the battery failure probability.
It improves the efficiency and accuracy of battery failure analysis, reduces computational complexity and cost, and can comprehensively integrate battery failure-related information, capture complex topological relationships and long-distance dependencies, and generate more representative node representations.
Smart Images

Figure CN120670978A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text prediction technology, and in particular to a method for predicting energy storage battery failure based on neighborhood aggregation and graph self-attention. Background Art
[0002] As a key storage device for renewable energy, energy storage batteries play a vital role in balancing grid loads, increasing renewable energy utilization, serving as an emergency backup power source, improving power quality, and reducing electricity costs. However, battery technology still faces demands for improvement in energy density, cycle life, safety, and other areas, making research and optimization of battery materials particularly urgent.
[0003] Existing energy storage battery failure analysis mainly relies on traditional experimental and empirical methods, which usually require complex laboratory tests and long-term field inspections to analyze the failure conditions that occur during battery use and storage, such as capacity decay, short cycle life, increased internal resistance, abnormal voltage, lithium deposition, gas production, leakage, short circuit, deformation, thermal runaway, etc., thereby providing a basis for battery design, manufacturing, use and subsequent maintenance, and providing a reference for the development of future battery technology.
[0004] Traditional energy storage battery failure analysis methods require extensive physical testing and manual intervention, resulting in inefficient data processing and high costs. Complex laboratory testing and long-term field testing consume significant time and manpower, leading to high equipment and operating costs. This not only delays the development and market adoption of new battery technologies but also increases the economic burden of research and production. Therefore, how to combine existing research data to obtain correlations between energy storage battery properties and employ artificial intelligence methods to predict battery failure has become an urgent challenge. Summary of the Invention
[0005] The present invention proposes a method for energy storage battery failure prediction based on neighborhood aggregation and graph self-attention, which solves the problem of high cost and low efficiency of battery failure analysis caused by the existing methods not fully considering the correlation between battery attribute information.
[0006] To solve the above technical problems, the present invention provides a method for predicting energy storage battery failure based on neighborhood aggregation and graph self-attention, comprising the following steps:
[0007] Step S1: defining entity classes and inter-class relationships of energy storage battery failure, extracting entities and relationships from literature in the field of energy storage batteries based on the entity classes and inter-class relationships, obtaining triple data, and constructing a knowledge graph of energy storage battery failure based on the triple data;
[0008] Step S2: Constructing a battery failure prediction model, wherein the battery failure prediction model maps entities and relationships in the knowledge graph into low-dimensional vectors to obtain embedded representations of the entities and relationships;
[0009] Step S3: Taking the entities as nodes, an adjacency matrix is constructed based on the relationships between the nodes. The battery failure prediction model generates a topological representation of each node by aggregating the multi-hop neighbor information of each node, and fuses the topological representations of each node based on the graph self-attention mechanism to obtain a topological representation of the node.
[0010] Step S4: train the classifier of the battery failure model using the fused embedding representation and topological representation to predict the battery failure probability, input the battery information to be predicted into the trained battery failure prediction model to obtain the failure probability of the battery to be predicted.
[0011] Preferably, in step S1, extracting entities and relationships in the literature of energy storage batteries according to the entity classes and inter-class relationships includes the following steps:
[0012] Step S11: converting the text sequence of energy storage battery field literature into semantic representation;
[0013] Step S12: capturing sequence dependencies from the forward and backward directions of the semantic representation, generating a forward embedding representation and a backward embedding representation, respectively, and concatenating the forward embedding representation and the backward embedding representation to obtain a bidirectional feature representation of each word in the text sequence;
[0014] Step S13: Generate a label sequence based on the bidirectional feature representation of each word, calculate the score of each word corresponding to the candidate label and the transition probability between each label;
[0015] Step S14: Calculate the initial score of the first word in the text sequence corresponding to all labels based on the score of each word corresponding to the candidate label and the transition probability between each label. Calculate the maximum score when transferring from all predecessor labels for the subsequent words of the first word and each candidate label, and record the path.
[0016] Step S15: Backtrack the path matrix to obtain the global optimal label sequence and obtain all entities and relationships in the literature in the field of energy storage batteries.
[0017] Preferably, in step S11, the input text sequence H is converted into a context embedding vector including token embedding, segment embedding and position embedding by a multi-layer bidirectional Transformer encoder structure, and the semantic representation H of the text sequence is generated according to the context embedding vector. seq , the semantic representation H seq The expression is:
[0018] H seq=BERT({e 1 ,e 2 ,...,e M});
[0019]
[0020] In the above formula, the BERT() function is the encoder of the BERT model; i is the context embedding vector of the i-th word; M is the number of words in the text sequence; are the token embedding, segment embedding, and position embedding of the i-th word respectively.
[0021] Preferably, in step S12, a bidirectional feature representation of each word in the text sequence is obtained by a bidirectional long short-term memory network BiLSTM, and the expression of the bidirectional feature representation of each word in the text sequence is:
[0022]
[0023] In the above formula, H BiLSTM The final representation of BiLSTM output; It is the feature representation after aggregating the forward and backward feature representations at time t; are the forward and backward feature representations at time t; LSTM for () is the encoder structure of the LSTM model; is the semantic representation of the text sequence; are the forward and backward feature representations at time t-1 respectively; are the time step states at time t-1 respectively.
[0024] Preferably, the expression of the global optimal label sequence in step S15 is:
[0025]
[0026] In the above formula, Y * is the global optimal label sequence; argmax means finding the parameters of the function; H is the text sequence; Y is the label sequence; S(H,Y) is the total score of the label sequence Y; T is the length of the text sequence; s t [y t ] is the label y t The score of A[y t-1 ,y t ] is the label y t and y t-1 The transfer score between .
[0027] Preferably, in step S2, mapping the entities and relationships in the knowledge graph into low-dimensional vectors comprises the following steps: mapping the head entity X H and tail entity X T Map it to a dimensional vector representation, map the relationship R to a dimensional diagonal matrix, and construct the triple data <X H ,R,X T >, a scoring function is used to evaluate the rationality of the triple data, and a cross entropy loss function is used to optimize the triple data to obtain the entity embedding representation Z(X) and the relationship embedding representation Z(R). The expression of the scoring function is:
[0028]
[0029] Where, f(X H ,R,X T ) is the scoring function; diag() means extracting the elements of the main diagonal of the matrix; P is the number of entities in the knowledge graph.
[0030] Preferably, in step S3, a message passing graph neural network MPNN is used to generate a topological representation of the node, and an aggregation function is used to aggregate the multi-hop neighbor information of each node. The expression of the aggregation function is:
[0031]
[0032] In the above formula, Z(N c ,G) is node N c The topological representation of the K-hop neighbor nodes; L is the number of hidden layers of MPNN; K represents the number of neighbor hops; For node N c and the set of all neighboring nodes within K hops; α k is the coefficient of the neighbor node within hop number K; MPNN() is the graph neural network encoder.
[0033] Preferably, the expression for fusing the topological representation of each node based on the graph self-attention mechanism in step S3 is:
[0034]
[0035] Where Z(N c ) is node N c Topological representation of ATT G represents the graph attention mechanism; δ(,) is the topology perception function; γ(ε c ) is node N c Adjacency matrix optimization; For node N c The linear transformation function of the absolute position encoding; N is the total number of nodes.
[0036] Preferably, the expression for fusing the embedding representation and the topological representation in step S4 is:
[0037]
[0038] Where Z is the representation of the entity; Norm represents the normalization operation; Z(X) is the embedding representation of the entity; softmax is the activation function; ε c For node N c The adjacency matrix of Z(N c ) is node N c Topological representation of ; N is the number of nodes c Spend.
[0039] Preferably, in step S4, the battery information to be predicted is input into the trained battery failure prediction model, and the expression of the failure probability of the battery to be predicted is obtained as follows:
[0040]
[0041] Where, is the failure probability of the battery to be predicted; X H is the head entity; R is the relationship between nodes; X T is the tail entity; sigmoid is the activation function; Z H 、Z T They represent the feature representation of the head entity and the tail entity respectively; Z(R) is the relation embedding representation; ⊙ represents matrix multiplication.
[0042] The benefits of the present invention include at least:
[0043] 1. By defining entity classes and inter-class relationships for energy storage battery failure, extracting entities and relationships from literature in the field of energy storage batteries to construct triple data, and further building a knowledge graph, this method can comprehensively integrate massive amounts of information related to energy storage battery failure, including the cross-relationships between various failure modes and energy storage performance parameters. This provides a rich and systematic data foundation for subsequent analysis, overcoming the shortcomings of traditional methods that rely solely on limited experiments and experience and have low data processing efficiency.
[0044] 2. Mapping entities and relationships in the knowledge graph into low-dimensional vectors to obtain embedded representations of entities and relationships. This can preserve the semantic and structural information in the original knowledge graph in a lower-dimensional space, simplifying and condensing the originally complex knowledge graph data while retaining key information. This helps improve the computational efficiency and performance of the model and avoids the difficulties and excessive computational overhead of traditional methods in processing high-dimensional complex data.
[0045] 3. By constructing an adjacency matrix and utilizing the graph self-attention mechanism to aggregate the multi-hop neighbor information of each node, the complex topological relationships and long-distance dependencies between nodes can be fully considered, solving the limitation of traditional methods that only focus on local neighbor information or simple statistical features. This enables the model to capture richer and more comprehensive contextual information, thereby more accurately reflecting the mutual influence and action mechanism between various factors in the failure process of energy storage batteries, and generating a more representative and discriminative node topology representation. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;
[0047] Figure 2 Schematic diagram of the structure of a named entity recognition model according to an embodiment of the present invention;
[0048] Figure 3 A schematic diagram of a knowledge graph of energy storage battery failure according to an embodiment of the present invention;
[0049] Figure 4 This is a schematic diagram of the structure of the energy storage battery failure prediction model according to an embodiment of the present invention;
[0050] Figure 5 Schematic diagram of the system structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following is a clear and complete description of the technical solutions in the embodiments of the present invention, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are within the scope of protection of the present invention.
[0052] like Figure 1 As shown, an embodiment of the present invention provides a method for predicting energy storage battery failure based on neighborhood aggregation and graph self-attention, comprising the following steps:
[0053] Step S1: Define the entity classes and inter-class relationships of energy storage battery failure, extract entities and relationships in the literature on energy storage batteries based on the entity classes and inter-class relationships, obtain triple data, and construct a knowledge graph of energy storage battery failure based on the triple data.
[0054] Specifically, we crawled the battery field related research literature from the literature website and screened the literature containing energy storage battery information. We cleaned the obtained literature, removed the noise and redundant information, and obtained accurate and high-quality literature information to construct a battery failure literature dataset.
[0055] In the embodiment of the present invention, keywords such as "battery degradation" and "failure" are used to search for literature titles on literature websites such as Web of Science and PubChem to screen relevant literature in the battery field. Then, the searched literature was screened using keywords such as "Energy Storage Battery", "Battery Degradation", "Battery Failure Mechanism", "Lithium-ion Battery Aging", "Sodium-ion Battery Performance", "Battery Thermal Runaway", "Electrochemical Degradation", "Capacity Fade in Batteries", "Battery Lifetime Prediction", "Failure Mode and Effect Analysis for Batteries", "Battery State of Health", and "Battery Cycle Life" to obtain energy storage battery literature.
[0056] Use regular expressions to mark punctuation marks and special characters such as $, #, % in the energy storage battery literature to reduce noise interference. Use word segmentation tools to segment text data and normalize synonyms and antonyms to reduce vocabulary diversity.
[0057] 10% of the literature in the dataset was annotated based on rule matching, and the BIO annotation method was used to annotate the battery type, electrode material, battery properties, failure mode, etc. in the literature for subsequent training of the named entity recognition model.
[0058] Specifically, the rules determine whether the document contains entities such as battery type, electrode material, battery properties, and failure mode. If so, the BIO annotation method is used to annotate them. Among them, battery types include lithium-ion battery, sodium-ion battery, solid-state battery, flow battery, and lead-acid battery. Electrode materials include positive electrode materials, negative electrode materials, and electrolytes. Electrode materials are matched by identifying chemical formulas. Battery properties include rate effect (C-rateEffec), overcharge, overdischarge, temperature, and humidity. When battery properties are identified, the subsequent specific property values must be annotated. Failure models include capacity fade, increase in internal resistance, solid electrolyte interface layer growth, thermal runaway, and lithium metal precipitation.
[0059] To automatically label entities in battery failure literature, a named entity recognition model based on Bert-BiLSTM was constructed and trained using labeled data. The best-performing model was saved during training and applied to the remaining 90% of unlabeled battery failure literature to complete the entity labeling task.
[0060] like Figure 2 As shown in Figure 1, the named entity recognition model consists of three main parts: the encoding layer SciBERT, the bidirectional long short-term memory layer BiLSTM, and the probability output layer. The SciBERT encoding layer uses the BERT pre-training model to obtain the semantic representation H of the input text. seq , the BiLSTM layer further encodes the semantic representation, and the probability output layer is responsible for outputting the label sequence with the highest probability.
[0061] The SciBERT encoding layer converts the input text sequence H into a token embedding E through a multi-layer bidirectional Transformer encoder structure. tok , segment embedded E seg and position embedding E pos The contextual embedding vector of the joint representation effectively captures the semantic features of each word. The process can be expressed as:
[0062]
[0063]
[0064] In the above formula, M is the number of words in the text sequence; R d is the hidden layer size of the SciBERT model.
[0065] In SciBERT, the overall embedding e of the i-th word i and the semantic representation of the text sequence H seq The expression is:
[0066]
[0067] H seq =BERT({e 1 ,e 2 ,...,e M});
[0068] In the above formula, BERT() represents the encoder function of the BERT model.
[0069] The BiLSTM layer semantically represents the text sequence generated by the SciBERT encoding layer H seq As input, it relies on the contextual dependencies of text sequences to improve entity recognition. BiLSTM consists of forward and backward long short-term memory units LSTM, and each LSTM unit consists of an input gate, a forget gate, an output gate, and a memory unit. seq =[H1 seq ,H2 seq ,...,H T seq ], where T is the sequence length, which is represented in LSTM as:
[0070]
[0071] h t =o t ⊙tanh(c t );
[0072] In the above formula, i t 、f t and o t They are input gate, forget gate and output gate respectively; h t is hidden state; c t is the memory unit; W f 、W i 、W o 、W c represents the weight matrix; U f 、U i 、Uo 、U c is the weight matrix of the hidden state; b f 、b i 、b o 、b c is the bias term; σ() is the sigmoid function; ⊙ represents element-wise multiplication; and tanh is the activation function.
[0073] The final representation H of the BiLSTM layer BiLSTM It is represented jointly by the forward LSTM and the backward LSTM:
[0074]
[0075] In the above formula, It is the feature representation after aggregating the forward and backward feature representations at time t; are the forward and backward feature representations at time t; LSTM for () is the encoder structure of the LSTM model; are the forward and backward feature representations at time t-1 respectively; are the time step states at time t-1 respectively.
[0076] The probability output layer recognizes and labels entities in the sequence by considering context information. Specifically, first, the H BiLSTM Generate a label sequence, calculate the score of each word corresponding to the candidate label and the transition probability between each label, calculate the initial score of the first word in the text sequence corresponding to all labels based on the score of each word corresponding to the candidate label and the transition probability between each label, calculate the maximum score when transferring from all predecessor labels for the subsequent words of the first word and each candidate label, and record the path at the same time; backtrack the path matrix to obtain the global optimal label sequence, and obtain all entities and relationships in the literature in the field of energy storage batteries.
[0077] For a text sequence H and its label sequence Y = [y1,y2,...,y T ], the total score of the sequence is:
[0078]
[0079] Where S(H,Y) is the total score of the label sequence Y; s t [y t ] represents BiLSTM for label y t The score assigned; A[y t-1 ,y t ] represents the transfer score between labels.
[0080] Perform entity recognition on unlabeled data based on the saved named entity recognition model, and combine the Viterbi algorithm to decode the label sequence with the highest probability in the recognition sequence:
[0081]
[0082] Where Y * is the global optimal label sequence; argmax means finding the parameters of the function.
[0083] The embodiment of the present invention divides the labeled data into a training set and a test set in a ratio of 8:2. The model training is performed on an NVIDIA GeForce RTX 4090 graphics card equipped with 24GB RAM, and the Adam optimizer is used to update and optimize the model parameters. During the training process, the hyperparameters set include: training epoch is 300, activation function is ReLU, dropout ratio is 0.2, number of hidden units is 762, number of hidden layers is 3, maximum sequence length is 512, and optimal learning rate is 0.05. In model training, the log-likelihood maximization of the predicted label of the probabilistic output layer is used as the loss function to optimize the model performance. The expression of the constructed loss function L is:
[0084]
[0085] Where Q is the number of entities.
[0086] To train the model and optimize the loss function, the present embodiment sets a corresponding number of training rounds. After training, the model with the best results is selected from the test set, saved, and used for subsequent entity labeling tasks. Table 1 shows a performance comparison of the named entity recognition model proposed in this embodiment of the present invention with other existing methods.
[0087] Table 1 Comparison of named entity recognition performance
[0088]
[0089] As can be seen from Table 1, the named entity recognition model proposed in the embodiment of the present invention performs well in all indicators, especially in key indicators such as the area under the curve AUC, the average precision recall curve area AUPR and the F1 score, which is significantly better than other comparison models.
[0090] To construct a knowledge graph for energy storage battery failure, we first organized battery data extracted and annotated from literature into a CSV format. The construction of the knowledge graph involves four main steps: entity recognition, relationship extraction, entity-relationship model construction, and the final construction of the knowledge graph.
[0091] In this embodiment of the present invention, the entity types related to battery failure include Battery Type, Cathode Material, Anode Material, Electrolyte, Failure Mode, and Impact Factor. These entities are automatically identified using a named entity recognition model and saved in a CSV file.
[0092] Three relationship names are defined: Used Materials (USES_MATERIAL), Failure Modes (HAS_FAILURE_MODE), and Influencing Factors (AFFECTED_BY). Uses_material describes the materials used in the battery, HAS_FAILURE_MODE describes the possible failure modes of the battery, and AFFECTED_BY describes how the failure mode is affected by specific factors. Based on these relationships, the following matching rules are set: "X battery + use + Y material" corresponds to Uses_material, "X has Y failure" corresponds to HAS_FAILURE_MODE, and "Y failure + affected by Z" corresponds to AFFECTED_BY. Using these rules, the relationships between entities are matched and stored in a CSV file.
[0093] The entity-relationship model consists of four entities: Battery (battery type), Material (material), FailureMode (failure mode), and ImpactFactor (influencing factor). Materials include positive electrode materials, negative electrode materials, and electrolytes. Relationships consist of three types: Uses_Material, Has_Failure_Mode, and Affected_By. Combining these entities and relationships, the following triples are constructed and saved in a CSV file:<Battery,USES_MATERIAL,Material> 、<Battery,HAS_FAILURE_MODE,FailureMode> 、<FailureMode,AFFECTED_BY,ImpactFactor> .
[0094] Finally, import the entities and relationships saved in the CSV file into the Neo4j database and get the following Figure 3 The complete knowledge graph of energy storage battery failure is shown.
[0095] Step S2: Construct a battery failure prediction model. The battery failure prediction model maps the entities and relationships in the knowledge graph into low-dimensional vectors to obtain embedded representations of the entities and relationships.
[0096] Step S3: Take the entities as nodes and construct an adjacency matrix based on the relationship between nodes. The battery failure prediction model generates a topological representation of each node by aggregating the multi-hop neighbor information of each node. The topological representation of each node is fused based on the graph self-attention mechanism to obtain the topological representation of the node.
[0097] Step S4: The classifier of the battery failure model is trained by the fused embedding representation and topological representation to predict the battery failure probability, and the battery information to be predicted is input into the trained battery failure prediction model to obtain the failure probability of the battery to be predicted.
[0098] Specifically, if Figure 4 As shown in Figure 1, the battery failure prediction model consists of three main modules: a knowledge representation module, a graph embedding module, and a prediction module. The knowledge representation module is responsible for acquiring embedded representations of entities and relationships within the graph. The graph embedding module combines graph neural networks to learn the global topological representation of entities. The prediction module fuses the embedded and topological representations to predict battery failure.
[0099] The knowledge representation module uses the DistMult model to obtain the embedded representation of entities and relations. Specifically, DistMult transforms the head entity X H and tail entity X T Map to vector representation and map the relation R to a diagonal matrix. Calculate the rationality of the triple embedding representation through the scoring function and optimize the triple representation with the cross entropy loss function to obtain the final entity embedding representation Z(X) and relation embedding representation Z(R):
[0100]
[0101] Where, f(X H ,R,X T ) is the scoring function; diag() means extracting the elements of the main diagonal of the matrix; P is the number of knowledge graph entities.
[0102] The graph embedding module represents the knowledge graph as an undirected graph G = (N, ε), where N = {N1, N2, ..., N S} is an entity, i.e., a set of nodes; ε is an adjacency matrix, which is used to represent the association relationship between entities. The embodiment of the present invention uses a message passing neural network MPNN to generate an embedded representation of a node, and uses an aggregation function to aggregate the features of the node's multi-hop neighbor nodes for entity representation. For node N c , whose embedding representation consisting of k-hop neighbor nodes is:
[0103]
[0104] Where Z(N c ,G) is node N cThe topological representation of the K-hop neighbor nodes; L is the number of hidden layers of MPNN; K represents the number of neighbor hops; For node N c and the set of all neighboring nodes within K hops; α k is the coefficient of neighbor nodes within hop number K, which is based on the feature representation of MPNN of central atoms with different hop numbers. k (N i ); MPNN() is the graph neural network encoder.
[0105] The embodiment of the present invention further designs a graph self-attention mechanism to integrate the feature information of nodes:
[0106]
[0107] Where Z(N c ) is node N c Topological representation of ATT G represents the graph attention mechanism; δ(,) is the topology perception function; γ(ε c ) is node N c Adjacency matrix optimization; For node N c The linear transformation function of the absolute position encoding; N is the total number of nodes.
[0108] The prediction module fuses the entity embedding representation Z(X) and the node topology representation Z(N c ) to obtain the representation Z of the entity in the knowledge graph:
[0109]
[0110] Where Z is the representation of the entity; Norm represents the normalization operation; soft max is the activation function; ε c For node N c The adjacency matrix of N is the number of nodes c Spend.
[0111] The embodiment of the present invention uses Z as the input of the prediction model and reconstructs the failure relationship into a joint probability distribution represented by the sigmoid function
[0112]
[0113] The embodiment of the present invention uses a binary cross entropy loss function to optimize the energy storage battery failure prediction model, and selects the model with the best prediction performance in the test set and saves it for subsequent battery failure prediction tasks.
[0114] The experimental examples were trained and tested on an NVIDIA GeForce RTX 4090 graphics card with 24GB of video memory. The main training parameters were set as follows: 100 epochs, a learning rate of 0.001, a batch size of 128, three hidden layers in the graph neural network (GNN), two node clusters, and a dropout ratio of 0.2.
[0115] In the experiment, the constructed knowledge graph was used as the dataset. Referring to the experimental setup of previous related research, the dataset was divided into a training set: validation set: test set ratio of 8:1:1. Table 2 shows a comparison of the failure prediction model proposed in this embodiment of the present invention and other methods in the energy storage battery failure knowledge graph.
[0116] Table 2 Performance comparison in failure prediction
[0117]
[0118] The results show that the improved model of the embodiment of the present invention achieved the best performance in all indicators, especially in accuracy and F1 score, which significantly outperformed other comparison methods. This shows that the failure prediction model of the embodiment of the present invention has a clear advantage in processing complex energy storage battery failure knowledge graphs.
[0119] like Figure 5 As shown, an embodiment of the present invention further provides an energy storage battery failure prediction system based on neighborhood aggregation and graph self-attention, which is implemented based on the above-mentioned energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention. The system includes:
[0120] Data acquisition module: used to collect relevant scientific literature on energy storage batteries, perform preliminary noise removal and data cleaning on the collected literature, and provide data support for entity recognition and knowledge graph construction.
[0121] Entity recognition and relationship extraction module: Use the trained named entity recognition model to identify entities in the literature, combine the pre-set relationship matching rules to extract the relationships between entities, and provide triple data for knowledge graph construction.
[0122] Knowledge graph construction module: Use the extracted triple data to build a knowledge graph to provide data for subsequent battery failure prediction.
[0123] Battery failure prediction module: Loads the trained battery failure prediction module, predicts the provided battery data, determines whether the battery has failed, and displays the results.
[0124] The technical features of the above embodiments may be combined in any manner. To simplify the description, not all possible combinations of the technical features in the above embodiments are described. Only preferred embodiments of the present invention are presented. While the description is relatively specific and detailed, it should not be construed as limiting the scope of the present invention. As long as there are no conflicts in the combination of these technical features, they should be considered to be within the scope of this specification.
[0125] It should be noted that, for those skilled in the art, various modifications and improvements can be made without departing from the scope of the present invention, and these modifications and improvements fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for predicting energy storage battery failure based on neighborhood aggregation and graph self-attention, characterized in that: The following steps are involved: Step S1: defining entity classes and inter-class relationships of energy storage battery failure, extracting entities and relationships from literature in the field of energy storage batteries based on the entity classes and inter-class relationships, obtaining triple data, and constructing a knowledge graph of energy storage battery failure based on the triple data; Step S2: Constructing a battery failure prediction model, wherein the battery failure prediction model maps entities and relationships in the knowledge graph into low-dimensional vectors to obtain embedded representations of the entities and relationships; Step S3: Taking the entities as nodes, an adjacency matrix is constructed based on the relationships between the nodes. The battery failure prediction model generates a topological representation of each node by aggregating the multi-hop neighbor information of each node, and fuses the topological representations of each node based on the graph self-attention mechanism to obtain a topological representation of the node. Step S4: train the classifier of the battery failure model using the fused embedding representation and topological representation to predict the battery failure probability, input the battery information to be predicted into the trained battery failure prediction model to obtain the failure probability of the battery to be predicted.
2. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 1 is characterized by: In step S1, entities and relationships in the energy storage battery field literature are extracted based on the entity classes and inter-class relationships, including the following steps: Step S11: converting the text sequence of energy storage battery field literature into semantic representation; Step S12: capturing sequence dependencies from the forward and backward directions of the semantic representation, generating a forward embedding representation and a backward embedding representation, respectively, and concatenating the forward embedding representation and the backward embedding representation to obtain a bidirectional feature representation of each word in the text sequence; Step S13: Generate a label sequence based on the bidirectional feature representation of each word, calculate the score of each word corresponding to the candidate label and the transition probability between each label; Step S14: Calculate the initial score of the first word in the text sequence corresponding to all labels based on the score of each word corresponding to the candidate label and the transition probability between each label. Calculate the maximum score when transferring from all predecessor labels for the subsequent words of the first word and each candidate label, and record the path. Step S15: Backtrack the path matrix to obtain the global optimal label sequence and obtain all entities and relationships in the literature in the field of energy storage batteries.
3. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 2 is characterized by: In step S11, the input text sequence H is converted into a context embedding vector including token embedding, segment embedding and position embedding through a multi-layer bidirectional Transformer encoder structure, and the semantic representation H of the text sequence is generated according to the context embedding vector. seq , the semantic representation H seq The expression is: H seq =BERT({e 1 ,e 2 ,...,e M }); In the above formula, BERT() function is the encoder of the BERT model; e i is the context embedding vector of the i-th word; M is the number of words in the text sequence; are the token embedding, segment embedding, and position embedding of the i-th word respectively.
4. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 2 is characterized by: In step S12, a bidirectional feature representation of each word in the text sequence is obtained through a bidirectional long short-term memory network BiLSTM. The expression of the bidirectional feature representation of each word in the text sequence is: In the above formula, H BiLSTM The final representation of BiLSTM output; It is the feature representation after aggregating the forward and backward feature representations at time t; are the forward and backward feature representations at time t respectively; LSTM for ( ) is the encoder structure of the LSTM model; is the semantic representation of the text sequence; are the forward and backward feature representations at time t-1 respectively; are the time step states at time t-1 respectively.
5. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 2 is characterized by: The expression of the global optimal label sequence in step S15 is: In the above formula, Y * is the global optimal label sequence; argmax means finding the parameters of the function; H is the text sequence; Y is the label sequence; S(H,Y) is the total score of the label sequence Y; T is the length of the text sequence; s t [y t ] is the label y t The score of A[y t-1 ,y t ] is the label y t and y t-1 The transfer score between .
6. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 1 is characterized by: In step S2, the entities and relationships in the knowledge graph are mapped into low-dimensional vectors, including the following steps: H and tail entity X T Map it to a dimensional vector representation, map the relationship R to a dimensional diagonal matrix, and construct the triple data <X H ,R,X T >, a scoring function is used to evaluate the rationality of the triple data, and a cross entropy loss function is used to optimize the triple data to obtain the entity embedding representation Z(X) and the relationship embedding representation Z(R). The expression of the scoring function is: Where, f(X H ,R,X T ) is the scoring function; diag() means extracting the elements of the main diagonal of the matrix; P is the number of entities in the knowledge graph.
7. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 1 is characterized by: In step S3, the message passing graph neural network MPNN is used to generate the topological representation of the node, and the aggregation function is used to aggregate the multi-hop neighbor information of each node. The expression of the aggregation function is: In the above formula, Z(N c ,G) is node N c The topological representation of the K-hop neighbor nodes; L is the number of hidden layers of MPNN; K represents the number of neighbor hops; For node N c and the set of all neighboring nodes within K hops; α k is the coefficient of the neighbor node within hop number K; MPNN() is the graph neural network encoder.
8. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 1 is characterized by: The expression for fusing the topological representation of each node based on the graph self-attention mechanism in step S3 is: Where Z(N c ) is node N c Topological representation of ATT G represents the graph attention mechanism; δ(,) is the topology perception function; γ(ε c ) is node N c Adjacency matrix optimization; For node N c The linear transformation function of the absolute position encoding; N is the total number of nodes.
9. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 1 is characterized by: The expression for fusing the embedding representation and the topological representation in step S4 is: Where Z is the representation of the entity; Norm represents the normalization operation; Z(X) is the embedding representation of the entity; softmax is the activation function; ε c For node N c The adjacency matrix of Z(N c ) is node N c Topological representation of ; N is the number of nodes c Spend.
10. The energy storage battery failure prediction method based on neighborhood aggregation and graph self-attention according to claim 1 is characterized by: In step S4, the battery information to be predicted is input into the trained battery failure prediction model, and the expression of the failure probability of the battery to be predicted is obtained as follows: Where, is the failure probability of the battery to be predicted; X H is the head entity; R is the relationship between nodes; X T is the tail entity; sigmoid is the activation function; Z H 、Z T are the feature embedding representations of the head entity and the tail entity respectively; Z(R) is the relation embedding representation; ⊙ represents matrix multiplication.
Citation Information
Patent Citations
Knowledge graph construction method for fault diagnosis and analysis of lithium ion battery of energy storage station
CN116384487A
Power battery fault analysis method and device, computer equipment and storage medium
CN117192373A
Entity description and topological structure information enhanced knowledge graph completion method
CN117725228A
Battery fault prediction model training method and device, and prediction method and device
CN118395370A
Lithium battery abnormal state prediction method and system based on deep learning
CN118734233A