Wind power fault knowledge graph construction method based on improved event extraction technology
By improving the event extraction technology and knowledge graph construction method, the problem of inefficient processing of text data of wind turbine faults in traditional methods is solved, and efficient extraction of wind farm fault information is achieved and operational efficiency is improved.
Patent Information
- Application Number
- CN202510382538.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
AI Technical Summary
Traditional manual analysis methods cannot efficiently process and extract relevant information in the text data of wind turbine failures, resulting in low operation and maintenance efficiency and high cost of wind farms.
The improved event extraction technology is adopted to design the ACFGEE model, combine BERT and GCN for multi-source event encoding, optimize event extraction through joint comparison learning and adaptive loss weighting layer, build a wind power failure knowledge graph, and use the Neo4j graph database to store and visualize the failure events.
It realizes efficient extraction and structured storage of wind power failure information, supports rapid fault location and root cause analysis, reduces operation and maintenance costs, and improves the operational efficiency and economic benefits of wind farms.
Smart Images

Figure CN120354920A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and natural language processing, and particularly relates to a method for constructing a wind power fault knowledge graph based on an improved event extraction technology. Background Technique
[0002] As an important clean energy at present, wind energy is renewable, and its development has promoted the rapid progress of wind power generation technology. With the continuous increase in the global demand for renewable energy, the wind power industry has also entered a rapid development stage, the number of wind turbines has been increasing, and the scale of wind farms has been expanding day by day. However, as complex mechanical equipment, wind turbines will inevitably be affected by various factors during operation, resulting in various types of faults. These faults not only affect the normal operation of wind turbines, but also may cause economic losses and safety hazards, and may seriously affect the stability of the entire power grid in severe cases.
[0003] At present, most of the fault information of wind turbines is stored in unstructured text data, such as fault reports, maintenance records, monitoring logs, etc. These data often contain a large number of professional terms, industry knowledge, and fault event details, and the data volume is huge. Traditional manual analysis methods cannot efficiently process and extract valuable information from them. Manual analysis is not only inefficient, but also easily limited by the experience and subjective judgment of the analyst. Therefore, how to efficiently reuse this unstructured fault text data to provide technical guidance for practitioners has become a key issue in improving the operation and maintenance efficiency of wind farms.
[0004] Event extraction (EE) is a technology for extracting key information from unstructured or semi-structured data. Its goal is to identify and extract the dynamic behavior or state change of an event itself within a specific time and space range. Different from traditional entity, relationship, and attribute information extraction technologies, the core of event extraction lies in extracting structured information related to "events" and answering key questions such as "who, when, where, what was done, why", and "how" of the event. The application of event extraction technology can effectively improve the extraction efficiency and accuracy of fault information, and provide support for subsequent fault diagnosis, root cause analysis, and maintenance decision-making.
[0005] A knowledge graph is a knowledge representation method that represents and stores information through a graph structure, aiming to express things (entities) in the real world and their relationships through nodes and edges. Each node usually represents a type of entity, such as a location, an event, a concept, etc., while the edges represent the relationships between these entities, such as "belongs to", "is associated with", "owns", etc. Through the structure of the graph, the knowledge graph can clearly show the multi-level relationships between things and support efficient query and reasoning. Combining natural language processing technologies such as event extraction, the knowledge graph not only provides an intelligent way for information storage but also improves the efficiency of data query and analysis.
[0006] Therefore, by combining event extraction and knowledge graph technologies, automatically processing and analyzing these fault text data to construct a wind power fault knowledge graph can store and display information such as fault events, causes, and influencing factors of wind power equipment in a structured manner, and reveal the potential relationships between events through fault graph analysis, so as to achieve rapid fault location and root cause analysis, provide data support for maintenance decision-making, effectively reduce the operation and maintenance costs of wind farms, and improve the operation efficiency and economic benefits of wind farms. Summary of the Invention
[0007] Aiming at the problem that traditional manual analysis methods cannot efficiently process and extract relevant information from fault text data, the present invention provides a method for constructing a wind power fault knowledge graph based on improved event extraction technology.
[0008] A method for constructing a wind power fault knowledge graph based on improved event extraction technology of the present invention includes the following steps:
[0009] Step 1: Obtain the wind power fault case dataset windPFC and perform annotation and division.
[0010] Step 2: Design an event extraction algorithm model ACFGEE (adaptive contrastive fusion graph-based event extraction) including four core modules: a multi-source event encoding layer, an event graph information fusion layer with joint contrastive learning, a one-stage decoding layer, and an adaptive loss weighting layer.
[0011] Step 3: Use the event extraction dataset windPFC to train the ACFGEE event extraction model.
[0012] Step 4: Input the data of the event extraction test set into the trained event extraction model, record the extraction results of the model, and evaluate the model performance.
[0013] Step 5: Design a fault event knowledge ontology and store the extracted structured fault event knowledge to construct a wind power fault event knowledge graph.
[0014] Step 6: Implement the visual display and recommendation application of the fault knowledge graph.
[0015] Further, step 1 is specifically as follows: Obtain the windPFC dataset (wind power fault case dataset), use the label-studio text annotation tool to annotate the event structure therein, and divide the dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0016] Further, step 2 is specifically as follows: Design an improved event extraction model ACFGEE, which includes a multi-source event encoding layer that combines the pre-trained language model BERT (bidirectional encoder representation from transformers) and GCN (graph convolutional network), an event graph information fusion layer for joint contrast learning, a one-stage decoding layer that simultaneously realizes trigger word and argument recognition, and an adaptive loss weighting layer that dynamically adjusts the contrast fusion learning loss and the one-stage decoding loss.
[0017] Multi-source event encoding layer: Independently encode and integrate the information from two sources, namely sentences and event graphs. Use the pre-trained language model BERT to embed the sentence information to obtain an embedded representation containing context information; subsequently, use the embedding layer to encode the nodes (event types and argument roles) to generate a single semantic event graph composed of node embeddings and edge index matrices, thereby obtaining all semantic event graphs containing all event types. Subsequently, use GCN to encode the semantic event graphs and model the dependency relationships between event types and corresponding argument roles into the semantic event graph embedding information.
[0018] Event graph information fusion layer for joint contrast learning: First, introduce the multi-head self-attention mechanism to calculate the weighted context representation of the sentence semantics and the global event graph structure, so that the words in each sentence are weighted and adjusted according to the structure information of the global event graph; subsequently, through the gated fusion mechanism of fusion contrast learning, maximize the similarity between positive sample features (true event graphs) and minimize the similarity between negative sample features (irrelevant event graphs) to fully and efficiently fuse the event graph and sentence encoding information.
[0019] One-stage decoding layer: Introduce the rotation position encoding technology to calculate the scores between trigger word pairs and argument word pairs to represent valid relationships. After obtaining the word pair scores, use the cross-entropy loss function ZLPR (zero-bounded log-sum-exp & pairwise rank-based) of the multi-label classification task to minimize the relationship scores of trigger word pairs and argument pairs.
[0020] Adaptive Loss Weighting Layer: Introduce the optimization strategy of multi-task learning in computer vision, design an adaptive loss weighting layer to assign different weights to the contrastive learning loss and the decoding loss, dynamically adjust the optimization objective to minimize the losses of all tasks, and achieve the adaptive optimization of multi-task learning.
[0021] Further, step 3 is specifically as follows: After processing the windPFC dataset and designing the ACFGEE model, train it; the input is sentence information with a length less than or equal to 512 characters, and the output is all potential entity information, namely trigger words and arguments, as well as the corresponding event types and argument roles; the experiment uses an NVIDIA A30 graphics card with 24G of video memory for training, the programming language is Python 3.8.19, and the deep learning framework is PyTorch 1.9.1 + cuda 11.1; select AdamW as the optimizer, use Bert-Base-Chinese as the pre-trained language model for the three datasets, and set its learning rate to 0.00002, the hidden layer to 768, the learning rate of GCN to 0.0008, the learning rate of other modules to 0.001, randomly select the number of negative samples to be 5, the number of iterations to be 70 epochs, the batch size to be 4, and use the cross-entropy loss function ZLPR for multi-label classification tasks as the loss function.
[0022] Further, step 4 is specifically as follows:
[0023] The evaluation criteria include four subtasks: trigger word identification (TI), trigger word classification (TC), argument identification (AI), and argument classification (AC).
[0024] TI: If the predicted trigger word span matches the true label, the trigger word is correctly identified.
[0025] TC: If the trigger word is correctly identified and assigned to the correct event type, the trigger word is correctly classified.
[0026] AI: If the event type is correctly identified and the predicted argument range matches the true label, the argument is correctly identified.
[0027] AC: If the argument is correctly identified and the predicted role matches the true label, the argument is correctly classified.
[0028] For all four evaluation criteria, the precision (Precision), recall (Recall), and F1 value representing the harmonic mean of P and R are used to measure the model performance, which are respectively expressed as:
[0029] The definition of Precision is the ratio of the number of samples predicted as positive by the model to the number of samples that are actually positive, as shown in formula (1):
[0030]
[0031] Among them, TP (True Positive) represents the number of instances correctly predicted as the target, and FP (False Positive) is the number of instances wrongly predicted as the target.
[0032] The definition of Recall is the ratio of all samples that are actually positive classes to the number of samples accurately identified as positive classes by the model, as shown in formula (2):
[0033]
[0034] Among them, FN (False Negative) is the number of instances that are actually the target but not correctly predicted.
[0035] The F1 value represents the harmonic mean of Precision and Recall. The F1 value is a comprehensive indicator used to measure the performance of a classification model, especially when there is a trade-off between Precision and Recall. Its formula is as follows:
[0036]
[0037] Furthermore, step 5 is specifically as follows: Use the Neo4j graph database to store the structured fault information extracted from events and construct a knowledge graph of wind power fault events; the construction of the knowledge graph follows the ontology structure of fault events, which consists of two parts: nodes and relationships, specifically as follows:
[0038] Nodes represent different entities in the fault event ontology, including the following eight types of entities:
[0039] 1) Time: Represents the time when the fault occurred.
[0040] 2) Project: Represents the name or location of the wind farm project.
[0041] 3) Wind turbine: Includes attributes such as turbine number and model type, representing a specific wind turbine.
[0042] 4) Fault phenomenon: The manifestation of the fault event itself, including fault type, fault unit, fault status, etc.
[0043] 5) Fault troubleshooting: Represents the fault troubleshooting process, including fault unit, troubleshooting operations, and fault status.
[0044] 6) Fault location: The result of fault location, clarifying the specific fault unit and status.
[0045] 7) Maintenance measures: Indicates the maintenance measures that have been taken.
[0046] 8) Structure: The structure tree of the wind turbine represents each component of the wind turbine, including blades, pitch system, generator, etc.
[0047] Relationship description represents the mutual connections between nodes, including the following seven types of relationships:
[0048] 1) Belong to: Describes the relationship between the fault phenomenon and the wind turbine structure.
[0049] 2) Time: Connects the fault phenomenon with the fault occurrence time.
[0050] 3) Location: Connects the fault phenomenon with the item where the fault occurs.
[0051] 4) Wind turbine: Connects the fault phenomenon with a specific wind turbine.
[0052] 5) Troubleshooting: Connects the fault phenomenon with fault troubleshooting and the relationships between multiple fault troubleshootings.
[0053] 6) Location: Connects fault troubleshooting with fault location.
[0054] 7) Correspondence: Connects fault location with repair measures.
[0055] The node and relationship design of the fault event ontology ensures a comprehensive representation of the fault event, fully reflecting information in various dimensions such as structure, time, location, wind turbine, fault troubleshooting, fault location, repair measures, etc. Each node represents an important entity (such as time, wind turbine, fault phenomenon, etc.), and each relationship represents the association between entities, such as the relationship between the fault phenomenon and time, the relationship between the fault phenomenon and the wind turbine, etc. On this basis, according to the design requirements of the event knowledge ontology, the knowledge storage algorithm converts the structured data automatically extracted by the event extraction model from the fault text into the format of a graph database, creates nodes and relationships in sequence according to the structural constraints of the event ontology, and imports them into the Neo4j graph database. This process ensures that the data layer information stored in the graph database conforms to the designed standard knowledge and provides efficient data support for subsequent query, analysis, and repair recommendation.
[0056] Furthermore, step 6 is specifically as follows: Design an intelligent operation and maintenance algorithm engine for wind turbine faults to realize the visual display of the fault event knowledge graph, enabling operation and maintenance personnel to intuitively view fault events, root cause analysis, and relevant data, and further using the similarity of fault events to recommend repair plans, so that it can be applied to fault diagnosis, prediction, and decision support in wind farms.
[0057] The intelligent operation and maintenance algorithm engine system for wind turbine faults adopts a B / S (Browser / Server) architecture. Users interact with the system through a browser (Browser). The event extraction intelligent calculation, data storage, and processing of the system are all carried out on the server side (Server). The system is developed in a front-end and back-end separated manner. The front-end uses the Vue.js framework for interface development. The back-end uses the Spring Boot framework to handle business logic and data storage respectively, and uses the Flask framework to provide functions related to deep learning and data analysis. Finally, based on the constructed knowledge graph, the system further uses the similarity of fault events to recommend maintenance solutions. By performing clustering analysis, calculating the similarities in multiple dimensions of fault phenomena, fault units, involved structures, and taken maintenance measures, the system helps to automatically identify and match the most similar maintenance solutions according to historical fault cases, thus forming a complete wind farm operation and maintenance system to help operation and maintenance personnel efficiently manage fault events and optimize the operation of wind farms.
[0058] The beneficial technical effects of the present invention are as follows:
[0059] A method for constructing a wind power fault knowledge graph based on an improved event extraction technology. First, the present invention designs an improved event extraction model, which automatically extracts event information from unstructured texts (such as fault reports, device logs, etc.). This model combines a multi-source event encoding layer of BERT (Bidirectional Encoder Representations from Transformers) and GCN (Graph Convolutional Network), and obtains the semantic information of sentences and the event semantics integrating role information by encoding and embedding sentence and event graph information. Secondly, an event graph information fusion layer using joint contrast learning fully fuses event graph and sentence information to enhance the effect of event extraction. Then, a one-stage decoding layer is used to simultaneously identify trigger words and arguments, reduce error propagation in multi-stage decoding, and improve the accuracy of extraction results. Finally, an adaptive loss weighting layer dynamically adjusts the contrast fusion learning loss and the one-stage decoding loss to optimize the performance of multi-task learning and achieve a better event extraction effect. After the event extraction model is optimized, the present invention further designs a fault event ontology, stores structured event data in accordance with the event ontology using the Neo4j graph database, converts the extraction results into structured information, and constructs a wind power fault event knowledge graph based on this information. Finally, through the visualization display of the graph and using the similarity of fault events to recommend maintenance plans, operation and maintenance personnel can intuitively view the fault events, root cause analysis and related data of wind turbine generators, and thus quickly locate problems. The present invention can be widely applied to the fault diagnosis and decision support system of wind farms, help operation and maintenance personnel quickly locate problems, improve the operation efficiency and economic benefits of wind farms, and has been successfully applied to a certain wind farm for fault diagnosis and decision support. Description of the Drawings
[0060] Figure 1 The flow chart for constructing the wind power fault event knowledge graph proposed by the present invention.
[0061] Figure 2 The detailed illustration of the event structure of the label-studio annotated text.
[0062] Figure 3 The framework diagram of the ACFGEE event extraction model proposed by the present invention.
[0063] Figure 4 The event structure diagram.
[0064] Figure 5 The metric performance of different algorithm models on the windPFC dataset.
[0065] Figure 6The ACFGEE event extraction model designed for the present invention is in the event extraction model management interface diagram trained with data in the wind power fault field.
[0066] Figure 7 It is the extraction result of windPFC test data.
[0067] Figure 8 It is the wind power fault event knowledge ontology designed for the present invention.
[0068] Figure 9 It is the structural diagram of the intelligent operation and maintenance algorithm engine for wind turbine faults designed for the present invention.
[0069] Figure 10 It is the visual display of the wind power fault knowledge graph constructed for the present invention.
[0070] Figure 11 It is the display of the intelligent recommendation page for wind power fault knowledge designed for the present invention. Detailed implementation method
[0071] The following further elaborates on the present invention in conjunction with the attached drawings and specific implementation methods.
[0072] The process of a method for constructing a wind power fault knowledge graph based on improved event extraction technology of the present invention is as Figure 1 shown. First is the annotation of the fault event knowledge text: Obtain the text data of wind power fault cases, use the label-studio text annotation tool to perform event structure annotation, and divide the annotated text data into a training set, a validation set, and a test set according to a ratio of 8:1:1 for training and evaluation; Subsequently, design an improved event extraction model ACFGEE (adaptive contrastive fusion graph-based event extraction) and conduct training and evaluation; Secondly, deploy and test the evaluated event extraction model; Then model the fault event ontology and import it into the Neo4j graph database, and import the structured data obtained from event extraction into the Neo4j graph database according to the event ontology structure to achieve knowledge storage and construct a wind power fault event knowledge graph; Finally is the knowledge service, including the knowledge query service for the visual display of the wind power fault event knowledge graph and the knowledge recommendation service for fault repair recommendation of fault events based on cosine similarity. Specifically, it includes the following steps:
[0073] Step 1: Obtain the wind power fault case dataset windPFC and perform annotation and division.
[0074] Obtain the windPFC dataset (wind power fault case dataset), use the label-studio text annotation tool to annotate the event structure therein, and the annotation details are as Figure 2As shown in the figure, the dataset is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0075] Step 2: Design an event extraction algorithm model ACFGEE that includes four core modules: a multi-source event encoding layer, an event graph information fusion layer with joint contrast learning, a one-stage decoding layer, and an adaptive loss weighting layer.
[0076] The event extraction framework diagram is as Figure 3 shown. ACFGEE includes four modules: a multi-source event encoding layer that combines the pre-trained language model BERT (bidirectional encoder representation from transformers) and GCN (graph convolutional network), an event graph information fusion layer with joint contrast learning, a one-stage decoding layer that simultaneously realizes trigger word and argument recognition, and an adaptive loss weighting layer that dynamically adjusts the contrast fusion learning loss and the one-stage decoding loss.
[0077] Multi-source event encoding layer: Independently encode and integrate the information from two sources, namely sentences and event graphs. Use the pre-trained language model BERT to embed the sentence information to obtain an embedding representation containing context information; subsequently, use the embedding layer to encode the nodes (event types and argument roles) to generate a single semantic event graph composed of node embeddings and edge index matrices, thereby obtaining all semantic event graphs containing all event types. Then, use GCN to encode the semantic event graphs and model the dependency relationships between event types and corresponding argument roles into the semantic event graph embedding information.
[0078] Event graph information fusion layer with joint contrast learning: First, introduce the multi-head self-attention mechanism to calculate the weighted context representation of sentence semantics and the global event graph structure, so that the words in each sentence are weighted and adjusted according to the structure information of the global event graph; subsequently, through the gated fusion mechanism of fusion contrast learning, maximize the similarity between positive sample features (true event graphs) and minimize the similarity between negative sample features (irrelevant event graphs) to fully and efficiently fuse the event graph and sentence encoding information.
[0079] One-stage decoding layer: Introduce the rotational position encoding technology to calculate the scores between trigger word pairs and argument word pairs to represent valid relationships. After obtaining the word pair scores, use the cross-entropy loss function ZLPR (zero-bounded log-sum-exp & pairwise rank-based) of the multi-label classification task to minimize the relationship scores of trigger word pairs and argument pairs.
[0080] Adaptive Loss Weighting Layer: Introduce the optimization strategy of multi-task learning in computer vision, design an adaptive loss weighting layer to assign different weights to the contrastive learning loss and the decoding loss, dynamically adjust the optimization objective to minimize the losses of all tasks, and achieve the adaptive optimization of multi-task learning.
[0081] The optimization is specifically as follows:
[0082] (1) Introduce the event graph structure as shown in Figure 4 . Encode the event graph through GCN, model the semantic relationship between the event type and the corresponding argument role as graph embedding information, assist the model to capture global dependencies, and enhance the ability of event semantic representation.
[0083] (2) Adopt a contrastive learning strategy when fusing sentence semantic information and global event graph structure information. Randomly select event graph structure information unrelated to the current event type in the global event graph structure information as negative samples, and event graph structure information of the same event type as positive samples, ensuring that the fused event information representation is as far away from negative samples as possible and close to positive samples, improving the model's discrimination ability for different event graph structure types, deeply integrating sentence semantics and event graph structure information, and effectively enhancing the correlation modeling ability between event types and argument roles.
[0084] (3) Design an adaptive loss weighting layer to dynamically adjust the loss weights of the contrastive fusion learning task and the one-stage decoding task, balance the contributions of each module, and thus optimize the performance of the multi-task learning framework.
[0085] Step 3: Use the event extraction dataset windPFC to train the ACFGEE event extraction model.
[0086] After processing the windPFC dataset and designing the ACFGEE model, train it; the input is sentence information with a length less than or equal to 512 characters, and the output is all potential entity information, namely trigger words and arguments, as well as the corresponding event types and argument roles; the experiment uses an NVIDIA A30 graphics card with 24G of video memory for training, the programming language is Python 3.8.19, and the deep learning framework is PyTorch 1.9.1 + cuda 11.1; select AdamW as the optimizer, use Bert-Base-Chinese as the pre-trained language model for the three datasets, set its learning rate to 0.00002, the hidden layer to 768, the learning rate of GCN to 0.0008, the learning rate of other modules to 0.001, randomly select the number of negative samples to be 5, the number of iterations to be 70 epochs, the batch size to be 4, and use the cross-entropy loss function ZLPR for multi-label classification tasks as the loss function.
[0087] Step 4: Input the data of the event extraction test set into the trained event extraction model, record the extraction results of the model, and evaluate the performance of the model.
[0088] The evaluation criteria include four subtasks: trigger word identification (TI), trigger word classification (TC), argument identification (AI), and argument classification (AC).
[0089] TI: If the predicted trigger word span matches the true label, the trigger word is correctly identified.
[0090] TC: If the trigger word is correctly identified and assigned to the correct event type, the trigger word is correctly classified.
[0091] AI: If the event type is correctly identified and the predicted argument range matches the true label, the argument is correctly identified.
[0092] AC: If the argument is correctly identified and the predicted role matches the true label, the argument is correctly classified.
[0093] For all four evaluation criteria, the precision, recall, and F1 value, which represents the harmonic mean of P and R, are used to measure the model performance, respectively expressed as:
[0094] The definition of Precision is the ratio of the number of samples predicted as positive by the model to the number of samples that are actually positive among them, as shown in formula (1):
[0095]
[0096] Among them, TP (True Positive) represents the number of instances correctly predicted as the target, and FP (False Positive) is the number of instances incorrectly predicted as the target.
[0097] The definition of Recall is the ratio of the number of samples that are actually positive to the number of samples accurately identified as positive by the model among them, as shown in formula (2):
[0098]
[0099] Among them, FN (False Negative) is the number of instances that are truly the target but not correctly predicted;
[0100] The F1 value represents the harmonic mean of Precision and Recall. The F1 value is a comprehensive indicator used to measure the performance of a classification model, especially when there is a trade-off between precision and recall. Its formula is as follows:
[0101]
[0102] By testing on the windPFC dataset, the method proposed in the present invention is compared with a variety of advanced event extraction models, and the comparison results on the windPFC dataset are shown in Table 1. The bold numbers indicate the best results, and the underlined numbers indicate the sub-optimal results. Figure 5 It is a three-dimensional graph plotted for the F1 values of four metrics of the windPFC dataset to more intuitively represent the performance of different models.
[0103] Table 1 Comparative experiments on the windPFC dataset
[0104]
[0105] In the windPFC dataset, the results of all models in the TI and TC tasks are exactly the same. By analyzing this dataset, it is found that there is a high correlation between trigger words and event types. For example, the trigger words for the "fault phenomenon" event type include "report, occur failure, alarm, report out", etc.; the trigger words for the "fault troubleshooting" event type include "troubleshoot, boarding inspection, inspection and discovery, measurement", etc.; the trigger words for the "fault location" event type include "locate, determine", etc.; and the trigger words for the "repair measure" event type include "replace, coordinate spare parts, coordinate delivery", etc. In this case, the relationship between trigger words and event types is almost fixed, so that the model can automatically and accurately classify the event type while identifying the trigger words. Therefore, the results of the TI and TC tasks tend to be the same. Similarly, for the AI and AC tasks, there are mainly six arguments: "time, location, model type, fan number, fault unit, and fault status", and there is also a strong distinguishability between them. So when the model identifies the arguments, it can basically determine the classification of the arguments.
[0106] To verify the effectiveness and feasibility of the event extraction method in the present invention in actual application scenarios, the event extraction model trained with data in the wind power fault field is deployed in the server and called and tested on the front-end page. The relevant model management page is as Figure 6 shown. Figure 7 It shows the event extraction test result log on the windPFC dataset, in which the relevant items, model types, and fan numbers of wind power fault cases are blurred.
[0107] In Figure 7In the case of wind power fault events, the model extracts four types of events: fault phenomena, fault troubleshooting, fault location, and repair measures. In the first three types of events, two arguments, namely the faulty unit and the corresponding fault status, are mainly extracted. For repair measures, only the faulty unit is extracted as an argument. Fault phenomena also include four arguments: time, location, model type, and fan number. As can be seen from the diagram, the model can effectively perform event extraction in the field of wind power faults, verifying the effectiveness and feasibility of the model in actual application scenarios.
[0108] Step 5: Design the knowledge ontology of fault events and store the extracted structured fault event knowledge to construct a wind power fault event knowledge graph.
[0109] Use the Neo4j graph database to store the structured fault information extracted from events and construct a wind power fault event knowledge graph; the construction of the knowledge graph follows the Figure 8 fault event ontology structure shown below. This structure consists of two parts: nodes and relationships, as follows:
[0110] Nodes represent different entities in the fault event ontology, including the following eight types of entities:
[0111] 1) Time: Represents the time when the fault occurred.
[0112] 2) Project: Represents the name or location of the wind farm project.
[0113] 3) Fan: Includes attributes such as fan number and model type, representing a specific fan.
[0114] 4) Fault Phenomenon: The manifestation of the fault event itself, including fault type, faulty unit, fault status, etc.
[0115] 5) Fault Troubleshooting: Represents the process of fault troubleshooting, including the faulty unit, troubleshooting operations, and fault status.
[0116] 6) Fault Location: The result of fault location, specifying the specific faulty unit and status.
[0117] 7) Repair Measures: Indicates the repair measures that have been taken.
[0118] 8) Structure: The structure tree of the fan, representing the various components of the fan, including blades, pitch systems, generators, etc.
[0119] Relationships describe the mutual connections between nodes, including the following seven types of relationships:
[0120] 1) Belong to: Describes the relationship between the fault phenomenon and the fan structure.
[0121] 2) Time: Connects the fault phenomenon with the time when the fault occurred.
[0122] 3) Location: Connect the fault phenomenon to the item where the fault occurred.
[0123] 4) Fan: Connect the fault phenomenon to the specific fan.
[0124] 5) Troubleshooting: Connect the fault phenomenon to the fault troubleshooting and the relationships between multiple fault troubleshootings.
[0125] 6) Location: Connect the fault troubleshooting to the fault location.
[0126] 7) Correspondence: Connect the fault location to the repair measures.
[0127] The design of the nodes and relationships of the fault event ontology ensures the comprehensive representation of the fault event, fully reflecting the information in various dimensions such as structure, time, location, fan, fault troubleshooting, fault location, and repair measures. Each node represents an important entity (such as time, fan, fault phenomenon, etc.), and each relationship represents the association between entities, such as the relationship between the fault phenomenon and time, the relationship between the fault phenomenon and the fan, etc. On this basis, according to Figure 8 the design requirements of the event knowledge ontology shown, the knowledge storage algorithm converts the structured data automatically extracted by the event extraction model from the fault text into the format of a graph database, creates nodes and relationships in sequence according to the structural constraints of the event ontology, and imports them into the Neo4j graph database. This process ensures that the data layer information stored in the graph database conforms to the design standard knowledge and provides efficient data support for subsequent query, analysis, and repair recommendation.
[0128] The optimization is specifically as follows:
[0129] (1) Design the event knowledge ontology as shown in Figure 8 ;
[0130] (2) Design the knowledge storage algorithm so that the structured data of event extraction can be stored according to the designed event ontology structure for subsequent visual display and recommended application.
[0131] Step 6: Implement the visual display and recommended application of the fault knowledge graph.
[0132] Design a wind turbine fault intelligent operation and maintenance algorithm engine to realize the visual display of the fault event knowledge graph, enabling the operation and maintenance personnel to intuitively view the fault events, root cause analysis, and relevant data, and further using the similarity of the fault events to recommend repair solutions, so that it can be applied to the fault diagnosis, prediction, and decision support of the wind farm.
[0133] The intelligent operation and maintenance algorithm engine system for wind turbine failures adopts a B / S (Browser / Server) architecture. Users interact with the system through a browser (Browser). The event extraction intelligent calculation, data storage, and processing of the system are all carried out on the server side (Server). The system is developed in a front-end and back-end separated manner. The front-end uses the Vue.js framework for interface development, and the back-end uses the Spring Boot framework to handle business logic and data storage respectively, and the Flask framework to provide functions related to deep learning and data analysis. Finally, based on the constructed knowledge graph, the system further uses the similarity of failure events to recommend repair solutions. By performing clustering analysis, calculating the similarities in multiple dimensions of failure phenomena, failure units, involved structures, and taken repair measures, it helps the system automatically identify and match the most similar repair solutions according to historical failure cases, thus forming a complete wind farm operation and maintenance system to help operation and maintenance personnel efficiently manage failure events and optimize the operation of the wind farm.
[0134] The system adopts a Figure 9 hybrid technology architecture of layered and microservices as shown. It realizes the extraction and dynamic expansion of knowledge through the tool layer, model algorithm layer, and data persistence layer, and realizes knowledge services through the intelligent calculation layer, engine interface layer, and application page interaction layer.
[0135] 1) Application page interaction layer - Embedded and front-end page interactive applications, the entry for users to interact with the system.
[0136] 2) Engine interface layer - Ontology modeling of failure event knowledge, extraction of failure event texts, query of failure event knowledge graphs, intelligent recommendation of failure cases, intelligent Q&A of failure knowledge.
[0137] 3) Analysis and calculation layer - Including analysis and calculation services such as natural language processing, problem processing, and answer retrieval.
[0138] 4) Data persistence layer - Adopts Neo4j and MiniO to store data knowledge such as graph knowledge and document models respectively.
[0139] 5) Model algorithm layer - Provides D2R algorithm parsing to convert database structured knowledge into graph database knowledge, provides E2R algorithm parsing to convert Excel table structured knowledge into graph database knowledge, uses the ACFGEE event extraction model for text event extraction, and a knowledge storage algorithm to import the structured knowledge obtained from event extraction into the graph database according to the event ontology structure, and calculates cosine similarity to calculate the similarity after failure events for repair case recommendation.
[0140] 6) Tool layer - including data knowledge storage tools, text data annotation tool label-studio, ontology construction tool Protégé, OWL ontology parsing tool, etc.
[0141] The fault event knowledge graph after the structural event knowledge obtained by event extraction is stored in the graph database is as Figure 10 shown (the relevant items and fan entities of the wind power fault cases are hidden). After the event extraction model successfully extracts structured information from unstructured texts such as fault reports and device logs, the wind power fault event knowledge graph constructed based on the Neo4j graph database becomes the core for subsequent analysis and decision support. The graph contains multiple important nodes, such as the time node, project node, fan node, fault phenomenon node, fault troubleshooting node, fault location node, maintenance measure node, etc. of the fault event, and connects these nodes through defined relationships to form a clear wind turbine fault management system. Based on the knowledge graph, the system further uses the similarity of fault events to recommend maintenance plans. By performing clustering analysis and calculating the similarities in multiple dimensions such as fault phenomena, fault units, involved structures, and taken maintenance measures, it helps the system automatically identify and match the most similar maintenance plan according to historical fault cases. The maintenance plan recommendation page for similarity calculation based on the extracted fault events is as Figure 11 shown.
[0142] The optimization is specifically as follows:
[0143] (1) Design the system architecture of the intelligent operation and maintenance algorithm engine for wind turbines;
[0144] (2) Implement the deployment, testing, and invocation of the event extraction algorithm model in the algorithm engine;
[0145] (3) Use cosine similarity to recommend fault repairs based on fault event cases.
[0146] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for constructing a wind power fault knowledge graph based on improved event extraction technology, characterized in that It includes the following steps: Step 1: Obtain the wind power fault case dataset windPFC and perform annotation and partitioning; Step 2: Design an event extraction algorithm model ACFGEE that includes four core modules: a multi-source event encoding layer, an event graph information fusion layer with joint contrast learning, a one-stage decoding layer, and an adaptive loss weighting layer; Step 3: Use the event extraction dataset windPFC to train the ACFGEE event extraction model; Step 4: Input the data of the event extraction test set into the trained event extraction model, record the extraction results of the model, and evaluate the model performance; Step 5: Design a fault event knowledge ontology and store the extracted structured fault event knowledge to construct a wind power fault event knowledge graph; Step 6: Implement the visual display and recommendation application of the fault knowledge graph.
2. The method for constructing a wind power fault knowledge graph based on an improved event extraction technology according to claim 1, wherein The specific content of Step 1 is as follows: Obtain the windPFC dataset, use the label-studio text annotation tool to annotate the event structure therein, and divide the dataset into a training set, a validation set, and a test set according to the ratio of 8:1:
1.
3. A method for constructing a wind power fault knowledge graph based on an improved event extraction technology according to claim 1, characterized in that, The specific content of Step 2 is as follows: Design the improved event extraction model ACFGEE, which includes four modules: a multi-source event encoding layer that combines the pre-trained language model BERT and GCN, an event graph information fusion layer with joint contrast learning, a one-stage decoding layer that simultaneously realizes trigger word and argument recognition, and an adaptive loss weighting layer that dynamically adjusts the contrast fusion learning loss and the one-stage decoding loss; Multi-source event encoding layer: Independently encode and integrate the information from two sources, namely sentences and event graphs. Use the pre-trained language model BERT to embed the sentence information to obtain an embedding representation containing context information; Subsequently, use the embedding layer to encode the nodes, generate a single semantic event graph composed of node embeddings and edge index matrices, so as to obtain all semantic event graphs containing all event types. Then use GCN to encode the semantic event graphs and model the dependency relationship between event types and corresponding argument roles into the semantic event graph embedding information; Event graph information fusion layer with joint contrast learning: First, introduce a multi-head self-attention mechanism to calculate the weighted context representation of sentence semantics and the global event graph structure, so that the words in each sentence are weighted and adjusted according to the structure information of the global event graph; Subsequently, through the gated fusion mechanism of fusion contrast learning, maximize the similarity between positive sample features and minimize the similarity between negative sample features to fully and efficiently fuse the event graph and sentence encoding information; One-stage decoding layer: Introduce the rotation position encoding technology to calculate the scores between trigger word pairs and argument word pairs to represent valid relationships. After obtaining the word pair scores, use the cross-entropy loss function ZLPR of the multi-label classification task to minimize the relationship scores of trigger word pairs and argument pairs; Adaptive loss weighting layer: Introduce the multi-task learning optimization strategy in computer vision, design an adaptive loss weighting layer to assign different weights to the contrast learning loss and the decoding loss, dynamically adjust the optimization objective to minimize the losses of all tasks, and achieve the adaptive optimization of multi-task learning.
4. A method for constructing a wind power fault knowledge graph based on an improved event extraction technology according to claim 1, characterized in that Step 3 is specifically as follows: After processing the windPFC dataset and designing the ACFGEE model, train it; the input is sentence information with a length less than or equal to 512 characters, and the output is all potential entity information, namely trigger words and arguments, as well as the corresponding event types and argument roles; the experiment uses an NVIDIA A30 graphics card with 24G of video memory for training, the programming language is Python 3.8.19, and the deep learning framework is PyTorch 1.9.1 + cuda11.1; select AdamW as the optimizer, use Bert-Base-Chinese as the pre-trained language model for the three datasets, and set its learning rate to 0.00002, the hidden layer to 768, the learning rate of GCN to 0.0008, the learning rate of other modules to 0.001, randomly select the number of negative samples to be 5, the number of iterations to be 70 epochs, the batch size to be 4, and use the cross-entropy loss function ZLPR for multi-label classification tasks as the loss function.
5. A method for constructing a wind power fault knowledge graph based on an improved event extraction technology according to claim 1, characterized in that, Step 4 is specifically as follows: The evaluation criteria include four subtasks: trigger word identification TI, trigger word classification TC, argument identification AI, and argument classification AC; TI: If the predicted trigger word span matches the true label, the trigger word is correctly identified; TC: If the trigger word is correctly identified and assigned to the correct event type, the trigger word is correctly classified; AI: If the event type is correctly identified and the predicted argument range matches the true label, the argument is correctly identified; AC: If the argument is correctly identified and the predicted role matches the true label, the argument is correctly classified; For all four evaluation criteria, the precision, recall, and F1 value representing the harmonic mean of P and R are used to measure the model performance, which are respectively expressed as: The definition of Precision is the ratio of the number of samples predicted as positive by the model to the number of samples that are actually positive, as shown in formula (1): Among them, TP represents the number of instances correctly predicted as the target, and FP is the number of instances wrongly predicted as the target; The definition of Recall is the ratio of the number of samples that are actually positive to the number of samples accurately identified as positive by the model, as shown in formula (2): Among them, FN is the number of instances that are truly the target but not correctly predicted; The F1 value represents the harmonic mean of precision and recall. The F1 value is a comprehensive indicator used to measure the performance of a classification model, especially when there is a trade-off between precision and recall. Its formula is as follows:
6. The method for constructing a wind power fault knowledge graph based on an improved event extraction technology according to claim 1, wherein Step 5 is specifically as follows: Use the Neo4j graph database to store the structured fault information extracted from events and construct a wind power fault event knowledge graph; the construction of the knowledge graph follows the fault event ontology structure, which consists of two parts: nodes and relationships, specifically as follows: Nodes represent different entities in the fault event ontology, including the following eight types of entities: 1) Time: Represents the time when the fault occurred; 2) Project: Represents the name or location of the wind farm project; 3) Wind turbine: including attributes such as wind turbine number and model, representing a specific wind turbine; 4) Fault phenomenon: the manifestation of the fault event itself, including fault type, fault unit, and fault status; 5) Fault troubleshooting: representing the fault troubleshooting process, including fault unit, troubleshooting operations, and fault status; 6) Fault location: the result of fault location, specifying the specific fault unit and status; 7) Maintenance measures: indicating the maintenance measures that have been taken; 8) Structure: the structure tree of the wind turbine, representing the various components of the wind turbine, including blades, pitch system, and generator; The relationship describes the mutual connections between nodes, including the following seven types of relationships: 1) Belong to: describing the relationship between the fault phenomenon and the wind turbine structure; 2) Time: connecting the fault phenomenon and the fault occurrence time; 3) Location: connecting the fault phenomenon and the project where the fault occurs; 4) Wind turbine: connecting the fault phenomenon and the specific wind turbine; 5) Troubleshooting: connecting the fault phenomenon and fault troubleshooting, as well as the relationships between multiple fault troubleshootings; 6) Location: connecting fault troubleshooting and fault location; 7) Correspondence: connecting fault location and maintenance measures; On this basis, according to the design requirements of the event knowledge ontology, the knowledge storage algorithm converts the structured data automatically extracted by the event extraction model from the fault text into the format of a graph database, creates nodes and relationships in sequence according to the structural constraints of the event ontology, and imports them into the Neo4j graph database.
7. A method for constructing a wind power fault knowledge graph based on an improved event extraction technology according to claim 1, characterized in that, The specific content of step 6 is as follows: Design a fault intelligent operation and maintenance algorithm engine for wind turbines to realize the visual display of the fault event knowledge graph, enabling operation and maintenance personnel to intuitively view fault events, root cause analysis, and relevant data, and further using the similarity of fault events to recommend maintenance plans, so that it can be applied to fault diagnosis, prediction, and decision support in wind farms; The fault intelligent operation and maintenance algorithm engine system for wind turbines adopts a B / S architecture. Users interact with the system through a browser. The event extraction intelligent calculation, data storage, and processing of the system are all carried out on the server side. The system development adopts a front-end and back-end separation method. The front end uses the Vue.js framework for interface development, and the back end uses the Spring Boot framework to handle business logic and data storage respectively, and uses the Flask framework to provide functions related to deep learning and data analysis. Finally, based on the constructed knowledge graph, the system further uses the similarity of fault events to recommend maintenance plans. By performing clustering analysis, calculating the similarity in multiple dimensions of fault phenomena, fault units, involved structures, and maintenance measures that have been taken, it helps the system automatically identify and match the most similar maintenance plan according to historical fault cases, thus forming a complete wind farm operation and maintenance system to help operation and maintenance personnel efficiently manage fault events and optimize the operation of wind farms.
Citation Information
Cited By
Wind generating set fault processing method, device and equipment based on multi-modal information and medium
CN121167314A