An operation and maintenance decision-making support method based on the knowledge graph of relay protection device defects
By constructing a defect knowledge graph of relay protection devices, using named entity recognition and knowledge integration, the problems of difficulty in operation and maintenance work and unused data are solved, efficient utilization of defective data is achieved, and the efficiency of power grid operation and maintenance is improved.
Patent Information
- Application Number
- CN202410066728.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-01-17
AI Technical Summary
In the prior art, the defects and accident statistical analysis of relay protection devices rely on manual professional capabilities, which makes operation and maintenance work difficult and defect data not effectively utilized.
Build a knowledge graph based on the defects of relay protection devices, and use the knowledge graph to assist in decision-making through naming entity recognition, knowledge extraction and integration, and provide operation, maintenance and maintenance references.
The full mining and utilization of defective data has been achieved, the difficulty of operation and maintenance work has been reduced, and the level of power grid operation and maintenance has been improved.
Smart Images

Figure CN117892814B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power operation and maintenance technology, and specifically to an operation and maintenance auxiliary decision-making method based on a knowledge graph of relay protection device defects. Background Art
[0002] With the rapid development and increasing intelligence of power grids, the number of relay protection devices has increased rapidly, and the management requirements for statistical analysis of their defects and accidents have also continued to increase, increasing the difficulty of power operation and maintenance. In existing technologies, statistical analysis of defects and accidents relies primarily on the professional capabilities of operators, which is a difficult task. Furthermore, the large amount of defect data accumulated by operators after they eliminate relay protection defects often remains idle and is not effectively mined and utilized. Therefore, there is an urgent need for intelligent technology to integrate and utilize relevant data, helping on-site operators to more efficiently handle defects and improve power grid operation and maintenance. Summary of the Invention
[0003] The purpose of the present invention is to provide an operation and maintenance auxiliary decision-making method based on the knowledge graph of relay protection device defects to solve the problems raised in the above background technology.
[0004] To achieve the above objectives, the present invention provides the following technical solution: an operation and maintenance decision-making support method based on a knowledge graph of relay protection device defects, comprising the following steps: step 1, preprocessing the relay protection device defect text; step 2, building a named entity recognition model; step 3, knowledge extraction; step 4, knowledge fusion; step 5, building a knowledge graph; step 6, using the knowledge graph to assist decision-making;
[0005] In the above step 1, the protection device defect report is selected as the original training data set, the record characteristics of the protection device defect text are analyzed, and on this basis, the text data is preprocessed to obtain the BIO annotation data set;
[0006] In step 2 above, a BERT-BiLSTM-CRF named entity recognition model is constructed and optimized to obtain a MacBERT-BiLSTM-CRF model.
[0007] In step 3 above, the MacBERT-BiLSTM-CRF model optimized in step 2 is used to extract entities from the unstructured defect text, and a rule-based approach is used to extract relationships.
[0008] In the above step 4, cosine similarity is used to perform knowledge fusion on the original structured data in the defect text and the combined data extracted from the knowledge in step 3;
[0009] In the above step 5, the knowledge graph of relay protection device defects is constructed based on the pattern layer, and the knowledge graph is stored using the Neo4j graph database;
[0010] In step six above, based on the defect phenomenon, the knowledge graph is used to infer diagnostic results and assist in decision-making, providing a reference for operation and maintenance work.
[0011] Preferably, the step 1 specifically includes the following steps:
[0012] 1.1 Analysis of the characteristics of relay protection device defect text: The defects recorded in the relay protection device defect report mainly include the following four categories:
[0013] 1) Defects in the installed relay protection device that has not yet been put into operation;
[0014] 2) Defects in the installed and operational relay protection device itself;
[0015] 3) Defects that occur during operation of a relay protection device that has been put into operation or trial operation;
[0016] 4) Defects found in the periodic inspection or other tests of the relay protection device put into operation or trial operation;
[0017] 1.2 Based on the text data format of the defect report, the data is divided into structured data and unstructured data. Structured data includes defect information, device information, substation information, and fault elimination information, which is standardized information automatically generated by the system; unstructured data includes protection anomalies or problems found during inspection, as well as their treatment and improvement measures, which are manually recorded in the form of short text descriptions;
[0018] 1.3 Use the following three methods to clean the unstructured data divided in step 1.2:
[0019] 1) Filter stop words: First, build a stop word dictionary based on the data characteristics, and then use the Jieba word segmentation package in Python combined with the built stop word dictionary to filter stop words in the defect text;
[0020] 2) Outlier processing: Remove extra spaces and abnormal line breaks in the text, delete voltage level information that is repeated in the structured data, and normalize the upper and lower case of English letters;
[0021] 3) Synonym merging: Construct a power industry dictionary based on power industry standard naming to merge and unify synonymous entities;
[0022] 1.4 Data annotation: Use the BIO annotation method to annotate the cleaned unstructured sample data content, match all characters and labels one by one, and generate a BIO annotation file;
[0023] 1.5 Data enhancement: A data enhancement method combining similar entity replacement and synonym replacement is used to expand the corpus of the annotated data. Specifically, first, the file data with BIO annotation is read, and entity information and non-entity information can be extracted according to the annotated label type; second, the label of the entity information is determined, and entity information with the same label is randomly selected from the data set for similar replacement; then, the Jieba word segmentation package is used to segment the remaining non-entity information, and synonym replacement is achieved by constructing a synonym dictionary; finally, the replaced entity information and non-entity information are combined to generate a new BIO annotation file.
[0024] Preferably, in step 1.5, in order to avoid semantic confusion and semantic repetition in the enhanced text, when replacing the same entity, it should be ensured that the entity replaced by the same entity that appears multiple times in the text is also the same, and when there are two or more defect phenomenon labels, it should be ensured that the replaced defect phenomenon entities are different; in addition, when performing synonym replacement, punctuation marks in the segmentation are ignored, and the replacement ratio is set to 20% for random replacement to avoid serious semantic deviation.
[0025] Preferably, the step 2 specifically includes the following steps:
[0026] 2.1 Constructing a BERT-BiLSTM-CRF named entity recognition model: First, pre-train the character sequence in the BERT layer to obtain dynamic word vectors. Then, the dynamic word vectors are input into the BiLSTM layer for bidirectional encoding to extract contextual semantic information. Finally, the semantic vector containing contextual information is input into the CRF layer for predicting label constraints. After Viterbi decoding, the globally optimal predicted label sequence is obtained, that is, the label with the maximum probability corresponding to each entity is obtained, completing the entity extraction task.
[0027] 2.2 The following five improvement methods are used to improve the BERT-BiLSTM-CRF model:
[0028] 1) Replace the BERT model with the MacBERT model;
[0029] 2) Concatenate the output vector of the MacBERT layer and the semantic vector output by the BiLSTM layer to fuse the features in series;
[0030] 3) Introduce the Dropout layer after the concatenated vector;
[0031] 4) Using the learning rate update strategy, the optimal local area is found through rapid training with the initial learning rate in the early stage. When the model monitoring indicator does not increase for two consecutive rounds, the model is trained using the decaying learning rate. The relationship between the decaying learning rate Learning_rateS and the initial learning rate Learning_rate is as follows:
[0032]
[0033] Where n is the number of times the learning rate is decayed;
[0034] 2.3 According to the confusion matrix, the precision rate P, recall rate R and F1 value are calculated as the evaluation indicators of the named entity recognition model. During model training, the F1 value is used as a monitoring indicator to judge the training effect of each round.
[0035] Preferably, the step 2.3 is specifically as follows:
[0036] In the confusion matrix, TP represents the proportion of entity samples whose actual value and predicted value are both 1; FN represents the proportion of entity samples whose actual value is 1 but predicted value is 0; FP represents the proportion of entity samples whose actual value is 0 but predicted value is 1; TN represents the proportion of entity samples whose actual value and predicted value are both 0;
[0037] The accuracy rate A is the percentage of sample data with correct prediction results to the total number of samples. It can reflect the extraction effect of the model to a certain extent. The calculation formula is as follows:
[0038]
[0039] Since the accuracy rate A is difficult to measure the entity extraction results when the samples are unbalanced, the F1 value is introduced as a monitoring indicator for model training. The F1 value is the harmonic mean of the precision rate P and the recall rate R, and the expression is as follows:
[0040]
[0041] The higher the F1 value, the better the overall recognition effect of the named entity recognition model. The precision rate is the percentage of samples with correct P entity label type judgment to all samples with positive prediction values for the classification. Its value is positively correlated with the classification accuracy. The calculation formula is as follows:
[0042]
[0043] The recall rate R is the percentage of samples with correct entity label type judgment among all actual samples of that type. Its value is negatively correlated with the number of samples missed during the prediction process. The calculation formula is as follows:
[0044]
[0045] Preferably, in the step three, entity extraction is specifically: using the constructed MacBERT-BiLSTM-CRF improved model to extract five types of key entity information, namely defective devices, defective parts, defective phenomena, defective causes and solutions, from the unstructured text of the data source; wherein, the defective phenomena are extracted separately, and when two or more defective phenomena are extracted, the defective phenomena are merged; relationship extraction is specifically: formulating relationship extraction rules to extract the semantic associations between the five types of entities.
[0046] Preferably, in step 4, the knowledge fusion is specifically as follows: using the cosine similarity algorithm to match and fuse entities, assuming that B1 and B2 are the entity word vectors in the electric power professional dictionary and the MacBERT model output word vectors of the extracted entities, respectively. The cosine similarity cosθ calculation formula of the two entities is as follows:
[0047]
[0048] The closer the cosine similarity is to 1, the more similar the two entities are. The entity with the highest similarity in the dictionary is selected for linking to complete knowledge fusion.
[0049] Preferably, the step 5 specifically includes the following steps:
[0050] 5.1 Constructing the model layer of the relay protection device defect knowledge graph. The model layer contains the entity information required to construct the relay protection device defect knowledge graph and establishes the connection relationship and attribute relationship between some entities;
[0051] 5.2 Build the data layer. The data sources of the data layer include the power company's relay protection device defect report and relay protection device defect management regulations. The original structured data in the defect text can be directly instantiated according to the pattern layer. For unstructured data, knowledge extraction and knowledge fusion are performed through steps 3 and 4, and then instantiated according to the pattern layer.
[0052] 5.3 Use the py2neo development framework to link to the Neo4j graph database to build a knowledge graph and realize the visualization of the knowledge graph in the field of relay protection device defects;
[0053] 5.4 Combined with the Neo4j graph database, the knowledge graph of relay protection device defects is updated in an incremental update manner.
[0054] Preferably, the step 5.4 specifically includes the following steps:
[0055] 5.4.1 For the newly appeared entity types in the relay protection device defect text, the experts summarized the relationship and subordination between the newly added entities and the original entities, and updated the model layer;
[0056] 5.4.2 Extract and integrate knowledge of newly added entities based on the model layer to complete the update of the data layer;
[0057] 5.4.3 Use the py2neo development framework to link to the Neo4j graph database, add and modify knowledge modules based on the original knowledge graph, and update the knowledge graph.
[0058] Preferably, the step six specifically includes the following steps:
[0059] 6.1 The defect content is identified through the improved MacBERT-BiLSTM-CRF model to obtain the defect device and defect phenomenon entities;
[0060] 6.2 Use cosine similarity to calculate the matching degree between the entities extracted by the model and the existing entities in the knowledge graph, and set the threshold to link the entities;
[0061] 6.3 For extracted entities greater than the threshold, the Cypher language is designed to match them with cases in the graph, infer the possible defect level, defect location and defect cause, and provide auxiliary strategies to improve the power grid operation and maintenance level; for extracted entities less than the threshold, they are analyzed and processed by experts, and the relay protection device defect knowledge graph is updated according to the knowledge update steps.
[0062] Compared with the existing technology, the beneficial effects of the present invention are: the present invention completes the construction of the relay protection device defect knowledge graph by performing knowledge extraction and knowledge fusion on the relay protection device defect text, and realizes the full mining and utilization of the defect data. The present invention proposes a process for using the knowledge graph to assist in decision-making and a method for updating the knowledge graph. The knowledge graph and defect phenomenon can be used to infer the defect level, defect location, defect cause and defect elimination suggestions, providing a reference for on-site operation and maintenance personnel, effectively reducing the difficulty of operation and maintenance work and improving the power grid operation and maintenance level. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 Construct an architecture diagram for the knowledge graph of relay protection device defects;
[0064] Figure 2 This is the data structure diagram of the defect report of the relay protection device;
[0065] Figure 3 This is a flow chart of the data enhancement method;
[0066] Figure 4 This is the BERT model architecture diagram;
[0067] Figure 5 This is the BERT-BiLSTM-CRF model architecture diagram;
[0068] Figure 6 Improved model architecture diagram for MacBERT-BiLSTM-CRF;
[0069] Figure 7 This is the structure diagram of the knowledge graph model layer of the relay protection device defect;
[0070] Figure 8 This is a flowchart for the operation and maintenance decision-making support application based on the knowledge graph of relay protection device defects;
[0071] Figure 9 This is a comparison chart of training monitoring indicators for different models;
[0072] Figure 10 The knowledge graph structure diagram instantiated for the pattern layer;
[0073] Figure 11 This is a visualization interface diagram of the knowledge graph of defects of some relay protection devices;
[0074] Figure 12 Flowchart of the method of the present invention. DETAILED DESCRIPTION
[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0076] See also Figure 1-12 , Table 1-11, an embodiment provided by the present invention: an operation and maintenance decision-making support method based on a knowledge graph of relay protection device defects, comprising the following steps: step 1, pre-processing of relay protection device defect text; step 2, building a named entity recognition model; step 3, knowledge extraction; step 4, knowledge fusion; step 5, building a knowledge graph; step 6, using the knowledge graph to assist decision-making;
[0077] In the above step 1, a power company's protection device defect report was selected as the original training data set. After preliminary screening, a total of 1578 valid historical defect data were obtained. The record characteristics of the protection device defect text were analyzed, and the text data was preprocessed on this basis. After 1x data enhancement, a total of 3156 defect text data were obtained. This was used as the data set for the named entity recognition model, and the training set, test set, and validation set were divided into 8:1:1 ratios. The specific steps include:
[0078] 1.1 Analysis of the characteristics of relay protection device defect text: The defects recorded in the relay protection device defect report mainly include the following four categories:
[0079] 1) Defects in the installed relay protection device that has not yet been put into operation;
[0080] 2) Defects in the installed and operational relay protection device itself;
[0081] 3) Defects that occur during operation of a relay protection device that has been put into operation or trial operation;
[0082] 4) Defects found in the periodic inspection or other tests of the relay protection device put into operation or trial operation;
[0083] 1.2 Based on the text data format of the defect report, the data is divided into structured data and unstructured data. Structured data includes defect information, device information, substation information, and fault elimination information, which is standardized information automatically generated by the system; unstructured data includes protection anomalies or problems found during inspection, as well as their treatment and improvement measures, which are manually recorded in the form of short text descriptions;
[0084] 1.3 Use the following three methods to clean the unstructured data divided in step 1.2:
[0085] 1) Filter stop words: First, build a stop word dictionary based on the data characteristics, and then use the Jieba word segmentation package in Python combined with the built stop word dictionary to filter stop words in the defect text;
[0086] 2) Outlier processing: Remove extra spaces and abnormal line breaks in the text, delete voltage level information that is repeated in the structured data, and normalize the upper and lower case of English letters;
[0087] 3) Synonym merging: Construct a power industry dictionary based on power industry standard naming to merge and unify synonymous entities;
[0088] 1.4 Data annotation: Use the BIO annotation method to annotate the cleaned unstructured sample data content, match all characters and labels one by one, and generate a BIO annotation file;
[0089] 1.5 Data enhancement: A data enhancement method combining similar entity replacement and synonym replacement is used to expand the corpus of the annotated data. Specifically, first, the BIO annotated file data is read, and entity information and non-entity information can be extracted according to the annotated label type; second, the label of the entity information is determined, and entity information with the same label is randomly selected from the data set for similar replacement; then, the Jieba word segmentation package is used to segment the remaining non-entity information, and synonym replacement is achieved by building a synonym dictionary; finally, the replaced entity information and non-entity information are combined to generate a new BIO annotated file; in order to avoid semantic confusion and semantic repetition in the enhanced text, when performing similar entity replacement, it should be ensured that the same entity that appears multiple times in the text is replaced by the same entity, and when there are two or more defect phenomenon labels, it should be ensured that the replaced defect phenomenon entities are different; in addition, punctuation marks in the segmentation are ignored when performing synonym replacement, and the replacement ratio is set to 20% for random replacement to avoid serious semantic deviation;
[0090] In step 2 above, a BERT-BiLSTM-CRF named entity recognition model is constructed and optimized to obtain a MacBERT-BiLSTM-CRF model. Specifically, the following steps are included:
[0091] 2.1 Constructing a BERT-BiLSTM-CRF named entity recognition model: First, pre-train the character sequence in the BERT layer to obtain dynamic word vectors. Then, the dynamic word vectors are input into the BiLSTM layer for bidirectional encoding to extract contextual semantic information. Finally, the semantic vector containing contextual information is input into the CRF layer for predicting label constraints. After Viterbi decoding, the globally optimal predicted label sequence is obtained, that is, the label with the maximum probability corresponding to each entity is obtained, completing the entity extraction task.
[0092] 2.2 The following five improvement methods are used to improve the BERT-BiLSTM-CRF model:
[0093] 1) Replace the BERT model with the MacBERT model;
[0094] 2) Concatenate the output vector of the MacBERT layer and the semantic vector output by the BiLSTM layer to fuse the features in series;
[0095] 3) Introduce the Dropout layer after the concatenated vector;
[0096] 4) Using the learning rate update strategy, the optimal local area is found through rapid training with the initial learning rate in the early stage. When the model monitoring indicator does not increase for two consecutive rounds, the model is trained using the decaying learning rate. The relationship between the decaying learning rate Learning_rateS and the initial learning rate Learning_rate is as follows:
[0097]
[0098] Where n is the number of times the learning rate is decayed;
[0099] 2.3 Based on the confusion matrix, the precision rate P, recall rate R and F1 value are calculated as the evaluation indicators of the named entity recognition model. During model training, the F1 value is used as a monitoring indicator to judge the training effect of each round; specifically:
[0100] In the confusion matrix, TP represents the proportion of entity samples whose actual value and predicted value are both 1; FN represents the proportion of entity samples whose actual value is 1 but predicted value is 0; FP represents the proportion of entity samples whose actual value is 0 but predicted value is 1; TN represents the proportion of entity samples whose actual value and predicted value are both 0;
[0101] The accuracy rate A is the percentage of sample data with correct prediction results to the total number of samples. It can reflect the extraction effect of the model to a certain extent. The calculation formula is as follows:
[0102]
[0103] Since the accuracy rate A is difficult to measure the entity extraction results when the samples are unbalanced, the F1 value is introduced as a monitoring indicator for model training. The F1 value is the harmonic mean of the precision rate P and the recall rate R, and the expression is as follows:
[0104]
[0105] The higher the F1 value, the better the overall recognition effect of the named entity recognition model. The precision rate is the percentage of samples with correct P entity label type judgment to all samples with positive prediction values for the classification. Its value is positively correlated with the classification accuracy. The calculation formula is as follows:
[0106]
[0107] The recall rate R is the percentage of samples with correct entity label type judgment among all actual samples of that type. Its value is negatively correlated with the number of samples missed during the prediction process. The calculation formula is as follows:
[0108]
[0109] 2.4 Using BERT-BiLSTM-CRF and MacBERT-BiLSTM-CRF 1 The named entity recognition experiment was conducted with the MacBERT-BiLSTM-CRF* model. The precision rate P, recall rate R and F1 value were used as evaluation indicators to verify the extraction effect of the MacBERT-BiLSTM-CRF improved model. 1 =(\textbf){\textbf} represents the model improved using only the first method, and MacBERT-BiLSTM-CRF* represents the final model improved using all four methods. The experimental results show that the BERT and MacBERT pre-trained models can achieve a high F1 value greater than 88% at the beginning of training, and the F1 value tends to increase and then stabilize with the increase in the number of training times. The MacBERT-BiLSTM-CRF* model has a higher entity extraction level than the MacBERT-BiLSTM-CRF model. 1 model and BERT-BiLSTM-CRF model, verifying the effectiveness of the model improvement method; MacBERT-BiLSTM-CRF 1 Compared with the BERT-BiLSTM-CRF model, the model's precision, recall, and F1 value increased by 1.418%, 1.446%, and 1.432%, respectively. This shows that for the sample data of this embodiment, the MacBERT preprocessing model is better than the BERT model; the value of the MacBERT-BiLSTM-CRF* model is 93.048%, and the precision, recall, and F1 value of this model are higher than those of the MacBERT-BiLSTM-CRF 1 The results are 0.138%, 0.262% and 0.201% higher than those of the previous model, respectively. This shows that the improved method proposed in this embodiment can improve the recognition effect of the named entity recognition model and can complete the entity extraction work of the relay protection device defect text;
[0110] In step 3, the MacBERT-BiLSTM-CRF model optimized in step 2 is used to extract entities from the unstructured defect text, and a rule-based approach is used to extract relationships. Specifically, entity extraction involves extracting five key entity types from the unstructured text of the data source using the constructed MacBERT-BiLSTM-CRF improved model: defective devices, defective locations, defective phenomena, defective causes, and solutions. Defective phenomena are extracted separately, and when two or more defective phenomena are extracted, they are merged. Specifically, relationship extraction involves formulating relationship extraction rules to extract semantic associations between the five types of entities.
[0111] In the above step 4, cosine similarity is used to perform knowledge fusion on the original structured data in the defect text and the combined data extracted in step 3. The knowledge fusion is specifically as follows: the cosine similarity algorithm is used to match and fuse entities. Let B1 and B2 be the entity word vector in the power professional dictionary and the MacBERT model output word vector of the extracted entity, respectively. The cosine similarity cosθ calculation formula of the two entities is as follows:
[0112]
[0113] The closer the cosine similarity is to 1, the more similar the two entities are. The entity with the highest similarity in the dictionary is selected for linking to complete knowledge fusion. To verify the effectiveness of the knowledge fusion method, this embodiment extracts 50 entities with irregular names from the relay protection device defect text, manually creates 50 entities with incomplete names, and matches and analyzes the 100 irregular entities with entities in the power professional dictionary. According to experimental statistics, 91 of the 100 irregular entities can be correctly matched, and entity alignment can be completed to a certain extent, realizing knowledge fusion of the knowledge graph in the field of relay protection device defects.
[0114] In the above step 5, the knowledge graph of relay protection device defects is constructed based on the pattern layer, and the knowledge graph is stored using the Neo4j graph database. Specifically, the following steps are included:
[0115] 5.1 Constructing the model layer of the relay protection device defect knowledge graph. The model layer contains the entity information required to construct the relay protection device defect knowledge graph and establishes the connection relationship and attribute relationship between some entities;
[0116] 5.2 Build the data layer. The data sources of the data layer include the power company's relay protection device defect report and relay protection device defect management regulations. The original structured data in the defect text can be directly instantiated according to the pattern layer. For unstructured data, knowledge extraction and knowledge fusion are performed through steps 3 and 4, and then instantiated according to the pattern layer.
[0117] 5.3 Use the py2neo development framework to link to the Neo4j graph database to build a knowledge graph and realize the visualization of the knowledge graph in the field of relay protection device defects;
[0118] 5.4 Combined with the Neo4j graph database, the knowledge graph of relay protection device defects is updated in an incremental update manner, which specifically includes the following steps:
[0119] 5.4.1 For the newly appeared entity types in the relay protection device defect text, the experts summarized the relationship and subordination between the newly added entities and the original entities, and updated the model layer;
[0120] 5.4.2 Extract and integrate knowledge of newly added entities based on the model layer to complete the update of the data layer;
[0121] 5.4.3 Use the py2neo development framework to connect to the Neo4j graph database, add and modify knowledge modules based on the original knowledge graph, and update the knowledge graph;
[0122] In step 6 above, based on the defect phenomenon, the knowledge graph is used to infer the diagnosis results and assist in decision-making, providing a reference for operation and maintenance work. Specifically, the following steps are included:
[0123] 6.1 The defect content is identified through the improved MacBERT-BiLSTM-CRF model to obtain the defect device and defect phenomenon entities;
[0124] 6.2 Use cosine similarity to calculate the matching degree between the entities extracted by the model and the existing entities in the knowledge graph, and set the threshold to link the entities;
[0125] 6.3 For extracted entities greater than the threshold, the Cypher language is designed to match them with cases in the graph, infer the possible defect level, defect location and defect cause, and provide auxiliary strategies to improve the power grid operation and maintenance level; for extracted entities less than the threshold, they are analyzed and processed by experts, and the relay protection device defect knowledge graph is updated according to the knowledge update steps.
[0126] Table 1 Unstructured sample data
[0127]
[0128] Table 2 Some professional dictionaries of electricity
[0129]
[0130] Table 3 Character label types
[0131]
[0132]
[0133] Table 4 BIO annotation results
[0134] character Label character Label character Label Ⅱ O now O bad I-Defect Cause part O field O , O mother B-Defective Device transport O right O Wire I-Defective Device dimension O C B-defective area Save I-Defective Device people O P I-Defective Site Protect I-Defective Device member O U I-Defective Site Pack I-Defective Device Inspection B-Solution plate I-Defective Site Place I-Defective Device check I-Solution Enter O transport B-defect phenomenon hair O OK O OK I-defect phenomenon now O Even B-Solution lamp I-defect phenomenon Save O Change I-Solution Destruction I-defect phenomenon Protect O , O , O C B-defective area Recovery O liquid B-defect phenomenon P I-Defective Site complex O crystal I-defect phenomenon U I-Defective Site just O black I-defect phenomenon plate I-Defective Site often O screen I-defect phenomenon damage B-Defect Cause 。 O
[0135] Table 5 Confusion Matrix
[0136]
[0137] Table 6 Relationship extraction rules
[0138] Entity Type relation Relationship attribute value Entity Type Defective device Occurrence Defect discovery method Defect phenomenon Defect phenomenon Occurrence site Defect phenomenon "+ location" Defective areas Defective areas Cause Defect phenomenon "+ cause" Cause of defect Cause of defect Solution Defect phenomenon "+ solution" Workaround
[0139] Table 7 Data augmentation results
[0140]
[0141] Table 8 Training parameters of the MacBERT-BiLSTM-CRF* model
[0142] Training parameters Setting Values Parameter meaning Epoch 20 Number of training sessions Batch_size 32 Batch size Learning_rate 0.00001 Initial learning rate LSTM_units 128 Number of neurons in the LSTM hidden layer Dropout_rate 0.1 Random discard rate Max_length 150 BERT reads maximum sequence length
[0143] Table 9 Experimental results of different models
[0144]
[0145] Table 10 Partial data relationship extraction results
[0146] entity relation Relationship attribute value entity Main transformer protection device Occurrence Run patrol Trip output circuit abnormality Trip output circuit abnormality Occurrence site Abnormal tripping output circuit + part Export plate Export plate Cause Trip output circuit abnormality + reason damage damage Solution Trip output circuit abnormality + solution replace
[0147] Table 11 Partial entity matching results
[0148]
[0149]
[0150] Based on the above, the advantages of the present invention are that, when the present invention is used, firstly, the recording characteristics of the defect text of the relay protection device are analyzed, and the unstructured text is cleaned, labeled and enhanced, thereby improving the quality and utilization rate of the defect text training data set; secondly, an improved method is proposed based on the BERT-BiLSTM-CRF model, and a MacBERT-BiLSTM-CRF improved model is constructed to perform entity extraction tasks. Compared with the BERT-BiLSTM-CRF model, the MacBERT-BiLSTM-CRF model has better entity extraction effect and can accurately identify the relay protection device. The entity information in the relay protection device defect text is extracted; then, the relation extraction rules of the relay protection device defect text are defined, and the relation extraction task is completed together with the entity extraction model. The semantic connection between entities is enhanced by defining the attribute values of the relationship, and the knowledge fusion is realized by calculating the cosine similarity of the entities to complete the entity alignment task; finally, the model layer of the relay protection device defect knowledge graph is constructed, and the Neo4j graph database is used to realize the storage of the knowledge graph data layer. The application process of relay protection device defect auxiliary decision-making and the knowledge graph update method are proposed, which realizes the full mining and utilization of defect data and effectively improves the power grid operation and maintenance level.
[0151] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. An operation and maintenance decision-making support method based on a knowledge graph of relay protection device defects, comprising the following steps: Step 1: Preprocessing of relay protection device defect text The protection device defect report is selected as the original training dataset. The record characteristics of the protection device defect text are analyzed. Based on this, the text data is preprocessed to obtain the BIO annotation dataset. The pretreatment comprises the following steps: S110. Analysis of defect text characteristics: Defects recorded in the relay protection device defect report mainly include the following four categories: S1101. Defects in installed relay protection devices that have not yet been put into operation; S1102. Defects in installed and operational relay protection devices; S1103. Defects occurring during operation of a relay protection device that has been put into operation or trial operation; S1104. Defects found in relay protection devices put into operation or trial operation during periodic inspection or other tests; S120. Data Classification: Based on the textual data format of defect reports, data is divided into structured data and unstructured data. Structured data includes defect information, device information, substation information, and fault elimination information, and is standardized information automatically generated by the system. Unstructured data includes protection anomalies or problems discovered during inspections, their treatment, and improvement measures, and is manually recorded in the form of short text descriptions. S130. Data cleaning: Use the following three methods to clean unstructured data: S1301. Filter stop words: First, build a stop word dictionary based on the data characteristics. Then, use the Jieba word segmentation package in Python combined with the built stop word dictionary to filter stop words in the defect text. S1302. Outlier processing: Delete extra spaces and abnormal line breaks in the text, delete voltage level information that is repeated in the structured data, and normalize the case of English letters; S1303. Synonym Merging: Construct a professional electric power dictionary based on electric power industry standard naming, and merge and unify synonymous entities; S140. Data annotation: Use the BIO annotation method to annotate the cleaned unstructured sample data content, match all characters to labels one by one, and generate a BIO annotation file; S150. Data Augmentation: This method uses a combination of similar entity replacement and synonym replacement to augment the annotated data. Specifically, the following steps are performed: First, the BIO-annotated file data is read and entity and non-entity information is extracted based on the annotated label type. Second, the entity information label is determined and entities with the same label are randomly selected from the dataset for similar replacement. Then, the remaining non-entity information is segmented using the Jieba word segmentation package and synonym replacement is achieved by constructing a synonym dictionary. Finally, the replaced entity information is combined with the non-entity information to generate a new BIO annotated file. S160. Data Augmentation Optimization: To avoid semantic confusion and duplication in the enhanced text, when replacing similar entities, ensure that the entities that appear multiple times in the text are replaced by the same entity. Furthermore, when there are two or more defect labels, ensure that the replaced defect entities are all different. Additionally, when performing synonym replacement, ignore punctuation marks in the segmentation, and set the replacement ratio to 20% for random replacement to avoid serious semantic deviation. Step 2: Build a named entity recognition model: Build a BERT-BiLSTM-CRF named entity recognition model and optimize it to obtain the MacBERT-BiLSTM-CRF model; The model construction and optimization includes the following steps: S210. Model Construction: First, character sequences are pre-trained in the BERT layer to generate dynamic word vectors. The dynamic word vectors are then fed into the BiLSTM layer for bidirectional encoding to extract contextual semantic information. Finally, the contextual semantic vectors are fed into the CRF layer for predicting label constraints. Viterbi decoding yields the globally optimal predicted label sequence, i.e., the maximum probability label for each entity, completing the entity extraction task. S220. Model Improvement: The following five improvement methods are used to improve the BERT-BiLSTM-CRF model: S2201. Replace the BERT model with the MacBERT model; S2202. Concatenate the output vector of the MacBERT layer and the semantic vector output by the BiLSTM layer; S2203. Introduce a Dropout layer after the concatenated vector; S2204. Using the learning rate update strategy, the optimal local area is found through rapid training with the initial learning rate in the early stage. When the model monitoring indicator does not increase for two consecutive rounds, the model is trained using the decaying learning rate. The relationship between the decaying learning rate Learning_rateS and the initial learning rate Learning_rate is as follows: Where n is the number of times the learning rate is decayed; S230. Model evaluation: Based on the confusion matrix, the precision rate P, the recall rate R, and the F1 value are calculated as evaluation indicators of the named entity recognition model. During model training, the F1 value is used as a monitoring indicator to judge the training effect of each round; Step 3: Knowledge extraction: Use the MacBERT-BiLSTM-CRF model optimized in step 2 to extract entities from unstructured defect text, and combine it with rule-based methods to extract relations; The entity extraction specifically involves extracting five key entity information categories, namely defective device, defective location, defective phenomenon, defective cause, and solution, from the unstructured text of the data source; wherein the defective phenomena are extracted separately, and when two or more defective phenomena are extracted, the defective phenomena are merged; The relationship extraction specifically includes: formulating relationship extraction rules to extract semantic associations between five types of entities; Step 4: Knowledge Fusion: Use cosine similarity to perform knowledge fusion on the original structured data in the defect text and the combined data extracted from step 3; The knowledge fusion is specifically as follows: using the cosine similarity algorithm to match and fuse entities, assuming that B1 and B2 are the entity word vectors in the electric power professional dictionary and the MacBERT model output word vectors of the extracted entities, respectively. The cosine similarity cosθ calculation formula of the two entities is as follows: The closer the cosine similarity is to 1, the more similar the two entities are. The entity with the highest similarity in the dictionary is selected for linking to complete knowledge fusion; Step 5: Build a knowledge graph: Construct a knowledge graph of relay protection device defects based on the pattern layer and use the Neo4j graph database to store the knowledge graph; The knowledge graph construction includes the following steps: S510.: Constructing the model layer: The model layer contains the entity information required to construct the knowledge graph of relay protection device defects, and establishes the connection relationship and attribute relationship between some entities; S520. Build the data layer: The data sources for the data layer include the power company's relay protection device defect reports and relay protection device defect management regulations. The existing structured data in the defect text can be directly instantiated according to the model layer. For unstructured data, knowledge extraction and knowledge fusion are performed through steps 3 and 4, and then instantiated according to the model layer. S530. Knowledge Graph Visualization: Utilize the py2neo development framework to link with the Neo4j graph database to construct a knowledge graph and visualize the knowledge graph in the field of relay protection device defects. S540. Knowledge Update: Integrate the Neo4j graph database to update the relay protection device defect knowledge graph in an incremental update manner. Specifically, the following steps are performed: S5401. For newly appeared entity types in the relay protection device defect text, experts summarize the relationships and subordinate relationships between the newly added entities and the original entities, and update the model layer; S5402. Extract and integrate knowledge of new entities according to the model layer to complete the update of the data layer; S5403. Use the py2neo development framework to connect to the Neo4j graph database, add and modify knowledge modules based on the original knowledge graph, and update the knowledge graph; Step 6: Use knowledge graphs to assist decision making Based on the defect phenomenon, the knowledge graph is used to infer the diagnosis results and assist in decision-making, providing a reference for operation and maintenance work; The auxiliary decision-making comprises the following steps: S610. Entity Recognition: Entity recognition is performed on the defect content using the improved MacBERT-BiLSTM-CRF model to obtain the defective device and defect phenomenon entities. S620. Entity Matching: Use cosine similarity to calculate the matching degree between the entities extracted by the model and the existing entities in the knowledge graph, and perform entity linking by setting a threshold; S630. Reasoning and decision-making: For extracted entities greater than the threshold, a Cypher language is designed to match them with cases in the graph, inferring the possible defect level, defect location, and defect cause, and providing auxiliary strategies to improve the power grid operation and maintenance level. For extracted entities less than the threshold, experts analyze and process them, and update the relay protection device defect knowledge graph according to the knowledge update steps.
Citation Information
Patent Citations
Shelter design knowledge graph construction method
CN116523043A
Power grid alarm information auxiliary decision-making method based on knowledge graph
CN116957236A