Wind turbine generator abnormal knowledge association method and device based on knowledge graph and medium
By constructing an abnormal knowledge graph based on the knowledge graph, the problems of data silos and knowledge fragmentation in the operation of wind turbines are solved, intelligent fault diagnosis and predictive maintenance are realized, and the operating stability and equipment life of wind turbines are improved.
Patent Information
- Application Number
- CN202511220340.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-29
AI Technical Summary
During the operation of wind turbines, there are problems of data silos and knowledge fragmentation, which makes abnormal diagnosis and predictive maintenance difficult. Existing methods rely on manual experience or isolated systems and lack systematic and dynamic responses.
A knowledge graph-based method is adopted to construct an abnormal knowledge graph through distributed multi-source heterogeneous data collection, OPC UA information model establishment, graph convolutional network and long short-term memory network optimization, perform entity association and knowledge reasoning, and realize intelligent retrieval of fault knowledge and predictive maintenance.
It improves the accuracy and efficiency of wind turbine fault diagnosis, extends equipment life, provides a feasible intelligent operation and maintenance framework, and solves the problems of data silos and knowledge fragmentation.
Smart Images

Figure CN120763709A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a method, device and medium for associating abnormal knowledge of wind turbines based on a knowledge graph. Background Art
[0002] Wind turbines, as core equipment for renewable energy, are widely deployed in remote areas (such as offshore wind farms and high-altitude mountainous areas). Their operational stability directly impacts grid reliability and economic profitability. A single wind turbine can cost tens of millions of yuan, and downtime can incur significant losses. A single 5MW turbine can incur losses exceeding 100,000 yuan per day. Wind turbine operation involves heterogeneous data from multiple sources, including real-time SCADA system parameters (such as wind speed and power output), vibration sensor data, thermal imaging monitoring, lubricant metal particle counts, and maintenance logs. This data is complex and voluminous, and failure modes are highly correlated. For example, a gearbox anomaly can lead to a cascading generator failure. However, the harsh environments and dispersed equipment of wind farms pose challenges to data collection and integration, including communication delays and data loss. This makes anomaly diagnosis and predictive maintenance a major pain point in the industry. Existing methods for wind turbine fault management rely primarily on manual experience or isolated systems (such as SCADA threshold alarms). This results in fragmented and unsystematic knowledge about anomalies, resulting in data fragmentation, low knowledge utilization, and insufficient dynamic response. Therefore, more effective methods are needed to manage and apply fault data. To overcome these limitations, the introduction of knowledge graph technology has become a key solution. Knowledge graphs can construct semantic networks of multi-source heterogeneous data (e.g., transforming vibration data, maintenance records, and supply chain information into nodes) and, through graph structure learning (e.g., convolutional networks), mine implicit associations between faults (e.g., the causal chain between main shaft imbalance and generator current anomalies). Furthermore, knowledge graphs can be efficiently constructed with the aid of industrial multimodal large models, fully leveraging the advantages of knowledge graphs in building knowledge networks and displaying knowledge associations. This allows general knowledge graph methods to be domain-specific, targeting the unique characteristics of wind turbine scenarios. This addresses core issues such as data silos and knowledge fragmentation, providing a viable intelligent operation and maintenance framework for the wind power industry. Summary of the Invention
[0003] The purpose of the present invention is to address the deficiencies in the prior art and provide a method, device and medium for associating abnormal knowledge of wind turbines based on a knowledge graph.
[0004] The purpose of the present invention is achieved through the following technical solutions: A wind turbine anomaly knowledge association method based on knowledge graph includes: Step 1: Based on the operation and maintenance needs of the wind power industry, distributed multi-source heterogeneous data and knowledge collection is carried out; Step 2: Establish the OPC UA information model of the wind power equipment node and determine the node's common attributes and reference types; Step 3: Use prompt word engineering to guide the industrial multimodal industrial big model to standardize the description and identification of node data, complete entity extraction and attribute extraction, and use graph convolutional networks to extract and represent relationships, and then import them into the graph database to build an abnormal knowledge graph; Step 4: Use graph convolutional networks and long short-term memory networks to optimize and update the semantic connections between node entities, complete knowledge merging and processing, and dynamically update the abnormal knowledge graph; Step 5: Aiming at the application scenario of abnormal knowledge, perform knowledge reasoning by mining entity associations based on the abnormal knowledge graph.
[0005] Furthermore, the step 1 includes: Step 1.1: Multi-source knowledge data integration: Data sources include internal and external data sources of wind power equipment operation and maintenance enterprises. Internal data sources include real-time equipment operation data, system operation and maintenance logs, and internal documents. External data sources include industry standards and public data sets, academic literature, and patents. Step 1.2, data knowledge preprocessing: Based on data quality, construct feature engineering to perform mean value replacement, data fitting, and data elimination on abnormal data and missing data, conduct rationalization tests on knowledge sources, and remove low-quality information sources.
[0006] Furthermore, the step 2 includes: Step 2.1, define the OPC UA information model: The node types of the OPC UA information model include objects, object types, variables, variable types, data types, reference types, methods, and views. The common attributes of the OPC UA information model include node identifiers, node types, browse names, display names, descriptions, write masks, and user write masks. Additional attributes of the OPC UA information model include whether it is abstract, whether it is symmetric, and reverse names. Step 2.2, build the OPC UA information model: The establishment of the OPC UA information model includes four steps: obtaining requirements, defining type models, instantiating information models, and exporting XML and CSV documents.
[0007] Furthermore, the step 3 includes: Step 3.1: Standardize the OPC UA information model for wind turbine equipment: Use the multimodal industrial macro model fine-tuned with industrial data to preprocess the OPC UA information model from Step 2, clean up illogical structured information, and combine the collected knowledge data to enhance embedded reasoning capabilities in the form of thought chains, thereby expanding the relationships between different individuals. Step 3.2, Entity Extraction and Attribute Extraction: Based on a variety of prompt word engineering design methods, including zero-shot learning, small-shot learning, role prompting, and thought chaining, the multimodal industrial large model is guided to normalize the OPC UA information model into an ontology under the attribute graph model to complete entity extraction and attribute extraction; Step 3.3, forming knowledge representation: performing relationship extraction and representation on the pre-processed OPC UA information model data based on the relationship graph convolutional neural network.
[0008] Furthermore, the step 4 includes: Step 4.1, knowledge merging and processing: Complete the knowledge merging and processing process based on graph convolutional neural networks, including link prediction, knowledge reasoning, and entity disambiguation, to form a standard knowledge representation; Step 4.2, construct the dynamic evolution of the knowledge graph: Use the graph convolutional neural network to capture the dependencies of the graph, use the long short-term memory network to learn and evolve the parameter matrices of different layers of the graph convolutional network at each moment, and thus use knowledge embedding to construct the dynamic evolution of the knowledge graph.
[0009] Furthermore, the link prediction in step 4.1 is based on the framework of original knowledge graph - graph neural network encoder - decoder - knowledge graph link prediction; the graph convolutional neural network uses the graph neural network to implement the link prediction task, explicitly representing the process of updating the information of a single node in the knowledge graph, collecting information from adjacent nodes, aggregating and updating it into the representation of this node, and continuously iterating and updating through downstream tasks until the node vector reaches a fixed point. After learning the abnormal knowledge graph, the DisMult decoder is used to generate a possibility score for each potential edge in the graph, and finally selects the most likely edge prediction; The knowledge reasoning in step 4.1 includes: the graph convolutional neural network uses the learned entity embedding to predict new triples; by calculating the scores of all possible tail entities and selecting the entity with the highest score as the prediction result; In step 4.1, the graph convolutional neural network implements entity disambiguation through the following process: Each entity reference in the text is linked to the entity in the knowledge graph by calculating the similarity between the context vector of the entity reference and the entity embedding vector in the knowledge graph. Based on the similarity score, the most likely entity is selected as the disambiguation result of the entity reference.
[0010] Furthermore, the step 5 includes: Knowledge graph reasoning applications of abnormal knowledge association of wind turbines: Knowledge graph reasoning applications of abnormal knowledge association include intelligent retrieval of fault knowledge, repetitive fault identification, role-based fault data analysis, corrective action effectiveness analysis, and predictive fault maintenance.
[0011] The application also provides a wind turbine unit abnormal knowledge association device based on a knowledge graph, comprising one or more processors, and is used for realizing the wind turbine unit abnormal knowledge association method based on the knowledge graph.
[0012] The application also provides a readable storage medium, which stores a program, and the program is executed by a processor to realize the wind turbine unit abnormal knowledge association method based on the knowledge graph.
[0013] Compared with the prior art, the application has the following beneficial effects: The application realizes a multi-modal knowledge graph construction method based on graph structure learning and a multi-modal industrial large model, explores the correlation of different abnormalities in the wind turbine unit operation and maintenance scene, constructs a knowledge base rich in abnormal semantics, and proposes an abnormal feature expression technology based on knowledge graph reinforcement, increases the probability of correctly handling faults under abnormal conditions, prolongs the service life of wind power equipment, solves core problems such as data island and knowledge fragmentation, and provides a landable intelligent operation and maintenance framework for the wind power industry. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description only constitute some embodiments of the application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.
[0015] Figure 1 It is an OPC UA information model modeling flowchart in the wind turbine unit abnormal knowledge association method based on the knowledge graph of the application.
[0016] Figure 2 It is an abnormal knowledge graph construction flowchart in the wind turbine unit abnormal knowledge association method based on the knowledge graph of the application.
[0017] Figure 3 It is a structure schematic diagram of the wind turbine unit abnormal knowledge association device based on the knowledge graph of the application. DETAILED DESCRIPTION
[0018] In order to better understand the technical solutions of the present application, the embodiments of the present application will be described in detail below with reference to the drawings.
[0019] It should be clear that the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0020] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "an", "the" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0021] The embodiment of the present invention provides a wind turbine abnormality knowledge association method based on knowledge graph, such as Figure 2 As shown, the following steps are included: Step 1: Based on the operation and maintenance business needs of the wind power industry, distributed multi-source heterogeneous data collection is carried out. Specifically, it includes: Step 1.1, Multi-Source Knowledge Data Integration: Data support can be primarily divided into two aspects: internal data sources and external data sources within the wind turbine operation and maintenance enterprise. Internal data sources include real-time equipment operation data, system operation and maintenance logs, and internal documentation. Real-time equipment operation data primarily comes from sensors. The original SCADA dataset collected over 75 different variables. Based on the incentive knowledge of the direct-drive wind turbine, 25 valid variables were selected from this dataset, including hub speed, hub angle, blade 1 angle, blade 2 angle, blade 3 angle, pitch motor 1 current, pitch motor 2 current, pitch motor 3 current, x-axis vibration value, y-axis vibration value, nacelle weather station wind speed, absolute wind direction value, inverter grid-side active power, wind tower ambient temperature, nacelle temperature, blade 1 motor case temperature, blade 2 motor case temperature, blade 3 motor case temperature, blade 1 pitch motor temperature, blade 2 pitch motor temperature, blade 3 pitch motor temperature, blade 1 inverter case temperature, blade 2 inverter case temperature, blade 3 inverter case temperature, and wind tower ambient temperature. System operation and maintenance logs include fault type (electrical, control system, pitch, yaw, hydraulic, drive system, tower, generator, and nacelle), fault time, equipment operation data from the three days prior to the fault, fault resolution (part replacement, parameter calibration), and technician identification (maintenance team information), forming a fault attribution chain. Internal documentation includes technical documentation (instrument and equipment manuals and design drawings, providing equipment functional specifications and safety thresholds) and equipment maintenance guides (recording core operation and maintenance knowledge such as fault resolution paths to support compliance assessment of abnormal operating conditions). External data sources primarily include industry standards and public datasets, academic literature and patents, and supply chain data. Industry standards and public datasets, such as fault accident databases, serve as core data sources. They integrate fault report descriptions (such as abnormal vibration and sudden temperature rise), handling methods (downtime for maintenance, parameter adjustment), and zeroing reports (root cause analysis documents) to extract anomaly-related knowledge (fault-symptom-solution triples). In addition, open databases such as production safety accident case studies published by relevant organizations and equipment reliability statistics shared by industry associations can supplement macro-risk models. Academic literature and patents primarily extract advanced algorithms (such as deep learning-based anomaly detection models) and fault diagnosis logic (such as Bayesian network causal reasoning) from these papers and patents to enhance the intelligence capabilities of knowledge graphs. Supply chain data is linked to supplier qualifications (such as component quality inspection reports) and raw material quality data (such as metal hardness test values) to analyze how supply chain issues (such as inferior bearings) affect equipment anomalies.
[0022] Step 1.2, data knowledge preprocessing: Since many data, especially text information, are filled by artificial, there are a large number of non-standard filling and non-uniform negative signs, it is necessary to construct feature engineering to replace the average value of abnormal data, missing data, data fitting and data rejection and other conventional processing, to ensure that the data meet the requirements, to reasonably test the knowledge source, to remove low-quality information sources, the data preprocessing rules include: filtering out the points with power less than or equal to zero, wind turbine limited operation and affected by switching transition; filter out the power outliers. Feature engineering is a known technology for data processing in machine learning, and will not be described here.
[0023] Step 2, establish the OPC UA information model of the wind power equipment node, determine the general attributes and reference types of the node, and model the node instance facing the equipment, as shown in Figure 1 Specifically, it includes: Step 2.1, define the OPC UA information model: traditional OPC can only provide data information, and cannot reflect semantic association. While the OPC UA specification can not only provide data information, but also provide effective data semantics, such as the OPC UA information model not only displays the attribute information of the temperature sensor device, but also allows the type hierarchy supported by the device to be exposed. OPC UA information modeling uses object-oriented technology, including type hierarchy and inheritance; type information is exposed to the outside, and can be accessed in the same way as accessing instances; a full mesh node network is used, which allows information to be connected in various ways; the type hierarchy and the reference type between nodes are extensible; OPC UA has no restrictions on how to model information, in order to establish appropriate models for the data provided; OPC UA information modeling is always done on the server side.
[0024] Specifically, the node types of the OPC UA information model in the wind turbine operation and maintenance field include object (Object), object type (ObjectType), variable (Variabe), variable type (VariableType, DataType, ReferenceType, Method, and View). The seven common properties of an OPC UA node include node identifier (NodeId), node type (NodeClass), browse name (BrowseName), display name (DisplayName), description (Description), write mask (WriteMask), and user write mask (UserWriteMask). Additional properties of the OPC UA information model include whether it is abstract (IsAbstract), whether it is symmetric (Symmetric), and inverse name (InverseName). The main entity information includes fault (Fault), product (Product), processing method (Manufacture), and elimination measure (Solution); the main attribute information includes product name (ProductName), product drawing number (Product No), product category (Product Category), fault description (Fault Desc), reporter (Reporter), fault category (Fault Category), cause category (Fault Cause), process instruction number (Craft No), process instruction version number (Craft Version), processing method (Process Type), elimination method (Process Detail), and processor (Processor).
[0025] Step 2.2, build the OPC UA information model: Modeling mainly includes four steps: requirement acquisition, defining type model, instantiating information model, and exporting XML and CSV documents.
[0026] Taking a motor as an example, we first need to obtain relevant information about the motor, including basic motor information, fault information, troubleshooting and resolution methods, and related product information. Basic motor information includes the motor model, manufacturer, and serial number. Fault information includes the fault type, description, time of occurrence, and sensor parameter logs before and after the fault. Troubleshooting and resolution methods include solutions and steps for different faults. Related product information includes information about other devices used with the motor, such as sensors and controllers.
[0027] Define the object type (ObjectType) and variable type (VariableType) related to the motor and its fault monitoring.
[0028] Object Type: MotorType (basic properties and methods of the motor); Parameters (motor configuration parameters); Methods (motor operations, such as starting and stopping); Faults (motor fault information); Troubleshooting (troubleshooting methods and handling methods); AssociatedProducts (associated product information).
[0029] Variable type (VariableType): StatusVariableType (motor status variable type); ConfigurationVariableType (motor configuration parameter variable type); FaultInfoVariableType (fault information variable type); TroubleshootingVariableType (troubleshooting method and handling method variable type).
[0030] Next, the information model is instantiated, the specific motor object is instantiated, and the properties and methods related to fault monitoring are included.
[0031] MotorInstance (a specific motor instance) inherits from MotorType (a motor type). This instance includes Status (an instantiation of StatusVariableType, representing the current state of the motor); Configuration (an instantiation of ConfigurationVariableType, containing the specific motor configuration parameters); Faults (fault type, fault description, and fault time); and Troubleshooting (resolution and steps). It also includes AssociatedProducts (related product information).
[0032] Then use the tool to export XML and CSV documents. Use OPC UA modeling tools (such as UAmodeler or opcua-modeler) to design and export XML and CSV documents. These documents will serve as the data source for instantiation information.
[0033] It is also necessary to bind the data source and bind the exported XML and CSV documents to the actual motor data source to ensure that the OPCUA server can correctly read and write the motor status, configuration parameters, and fault information.
[0034] Finally, testing and verification are performed to test the interaction between the OPC UA client and the motor to ensure that all properties and methods work correctly and that the data interaction process meets expectations, especially the fault monitoring and handling parts.
[0035] In practical applications, the type model can be adjusted according to the device object. For variable pitch motor 1, basic properties can be obtained: variable pitch motor model (such as PITCH-5MW-V3), offshore corrosion protection level (IEC 61400-22), gearbox reduction ratio (1:98.7); fault monitoring parameters: blade angle command deviation (>2° alarm), winding temperature gradient (>5℃ / min), backup power supply voltage drop (<18VDC); environmental related data: real-time wind speed (0-25m / s), atmospheric humidity (>90% icing risk), salt spray concentration (exclusive to coastal wind farms); grid interaction indicators: low voltage ride-through status (LVRT Active) and active power change rate (dP / dt <-0.5MW / s).
[0036] Step 3: Use prompt word engineering to guide the industrial multimodal industrial big model to standardize the description and identification of node data, complete entity extraction and attribute extraction, use graph convolutional networks to extract and represent relationships, and then import them into the graph database to build an abnormal knowledge graph. Specifically, it includes: Step 3.1, Standardizing the Information Model: This involves cleaning and normalizing OPC UA modeling information, uncovering hidden connections between OPC UA modeling objects, and laying the foundation for subsequent steps such as entity and attribute extraction, relationship extraction and identification, and importing into a graph database. The following sections describe these specific implementation steps.
[0037] Specifically, OPC UA modeling information preprocessing is mainly based on the powerful natural language and multimodal processing capabilities of multimodal industrial large models. First, a set of specifications and standards should be defined. These specifications and standards should include: Data naming rules: unified naming conventions, such as camel case or underscore case.
[0038] Data type definition: Clarify the specific requirements of each data type, such as integer, floating-point number, string, etc.
[0039] Units and dimensions: Standardize the use of units and dimensions, such as using Celsius or Fahrenheit for temperature.
[0040] Encoding rules: For encoded data, such as enumeration values, define clear encoding rules.
[0041] Data format: Define a unified format for date, time and other data.
[0042] Next, we use the data cleaning capabilities of the multimodal industrial big model to clean non-standard data. Multimodal industrial big models can be developed in-house or sourced from commercial sources, such as the Qizhi Kongming Industrial Big Model, Cosmo-GPT, and Industry-GPT. Data cleaning primarily involves removing invalid data, formatting, and converting data types. Removing invalid data: Identifying and removing invalid data such as null values and error values. Formatting: Converting data in different formats to a unified format. Data type conversion: Converting data to the correct data type.
[0043] Data validation is then performed using a multimodal industrial model to ensure data accuracy and consistency. This primarily includes range checks, uniqueness checks, and integrity checks. Range checks verify that data falls within a reasonable range. Uniqueness checks ensure the uniqueness of key fields. Integrity checks ensure the integrity of data records.
[0044] Furthermore, based on a multimodal industrial big model, it can identify and correct irregular data in batches. For example, it can classify data to identify different types of irregular entries; predict reasonable values for data and correct outliers; and group similar data for batch processing.
[0045] Step 3.2, Entity and Attribute Extraction: Specifically, uncovering the hidden connections between objects in the OPC UA information model relies primarily on the reasoning capabilities of the large language model enhanced by Chain of Thought. After the aforementioned OPC UA information modeling and normalization, errors in the data itself are corrected. However, due to data quality limitations, the semantic connections between objects have not yet been fully clarified. Therefore, the reasoning capabilities of the large language model enhanced by Chain of Thought are needed to mine and expand these relationships.
[0046] Large language models inherently possess a certain level of reasoning capability. However, the thought chaining prompting method, as a computationally efficient method that eliminates the need for training and fine-tuning large language models, can improve their reasoning capabilities by breaking down complex problems into multi-step sub-problems through prompt engineering. For example, the thought chaining-enhanced GPT-3.5 model significantly outperforms the standard model on datasets such as mathematical problems, commonsense reasoning, and symbolic reasoning.
[0047] Next, we will use the maintenance of motors as an example to provide guidance on engineering design through thought chains: Q: The motor overheats and cannot continue to run. What part is related? A: Let's think step by step. Based on the collected data, list possible causes of the fault: high temperature due to cooling system failure, poor ventilation; abnormal current due to short circuit, winding damage... Logical reasoning: if the current temperature exceeds the normal range and the cooling system does not work properly, it can be inferred that the overheating is due to insufficient cooling.
[0048] A: Action taken: Recommend immediate inspection of cooling liquid level, cleaning of radiator, etc.
[0049] Thus, the object relationship of motor overheating - cooling system - cooling liquid level / radiator is established. New reference relationship between each other is created.
[0050] Specifically, entity extraction and attribute extraction mainly rely on the natural language processing capabilities of large language models based on prompt engineering. Entity extraction in a knowledge graph refers to automatically identifying named entities from raw corpus. Entities, as the most basic elements of a knowledge graph, their extraction accuracy directly affects the quality of the subsequent knowledge base. Traditional entity extraction methods can be roughly divided into three categories: (1) Rule-based method, which is characterized by performing under limited semantic and text field conditions. Under the defined rules, entity information is extracted, but this method relies heavily on expert experience and is difficult to adapt to new requirements of various changing data; (2) Statistical learning-based entity extraction method, which treats named entity recognition as a sequence labeling problem and uses partially labeled or fully labeled semantic information for training; (3) Deep learning-based method, which has good fitting ability for complex nonlinear problems and can learn complex features. The large language model method based on prompt engineering belongs to method (3). Prompt engineering consists of three parts: instruction (Instruction), content (Content), and input text (InputText). The following four methods are used for prompt word design in this invention, and specific examples are as follows: 1. Zero-shot (zero-shot learning): First, Instruction (instruction): "According to the following OPC UA information model, extract the main body and attribute of the knowledge graph." Content (content): NULL (null value).
[0051] InputText: MotorInstance: a specific motor instance, inherited from MotorType; Status: an instantiation of StatusVariableType, indicating the current state of the motor; Configuration: an instantiation of ConfigurationVariableType, containing the specific configuration parameters of the motor.
[0052] Faults: FaultType (fault type); FaultDescription (fault description); FaultTime (fault occurrence time).
[0053] Troubleshooting: Resolution; Steps; Processors.
[0054] AssociatedProducts: Associated product information.
[0055] The model gives an Answer (response): the main entity information includes fault (Fault) - Faults, product (Product) - MotorInstance, processing method (Manufacture) - Status / Configuration, and troubleshooting (Solution); the main attribute information includes product name (Product Name) - MotorInstance, product drawing number (Product No), product category (Product Category), fault description (FaultDesc) - FaultDescription, reporter (Reporter), fault category (FaultCategory) - FaultType, cause category (FaultCause), process instruction number (Craft No), process instruction version number (CraftVersion), processing method (Process Type) - Resolution, processing method (Process Detail) - Steps, and processor (Processor) - Processors.
[0056] 2. Few-Shot (small sample learning): Instruction: Based on the following OPC UA information model, extract the knowledge graph subject and attribute from it.
[0057] Content: MotorInstance-others (similar to MotorInstance).
[0058] InputText: Omitted (same as Zero-Shot).
[0059] Answer: Omitted.
[0060] 3.Role-Prompt Instruction: Based on the following OPC UA information model, extract the knowledge graph subject and attribute from it.
[0061] Content: Assume that you are a professional maintenance worker of xx wind turbine equipment.
[0062] InputText: Omitted (same as Zero-Shot).
[0063] Answer: Omitted.
[0064] 4. Chain-of-Thought Instruction: Based on the following OPC UA information model, extract the knowledge graph subject and attribute from it.
[0065] Content: Let's think step by step.
[0066] InputText: Omitted (same as Zero-Shot).
[0067] Answer: Omitted.
[0068] Step 3.3, forming knowledge representation: Specifically, the extraction and identification of knowledge graph relationships are based on the relational graph convolutional neural network (Relation-GCN), and can also be completed using other graph convolutional networks or graph attention networks. R-GCN divides relationships into three categories: one is the relationship pointing to the node itself, the second is the relationship pointing to external nodes, and the third is the relationship between the node itself. By aggregating the information vectors of these three types of relationships, the corresponding probability distribution is obtained according to the task requirements through the RELU function in turn. Compared with GCN, R-GCN separates the information of the edge and the information of the node itself in the hidden layer, and uses Wr to represent and adjust the information of the adjacent nodes on its edge, which makes up for the lack of consideration of the information of the nodes connected by the edge by GCN. Define the network G=(V,ε,R), where the node v i ∈V, edge(v i ,r,vj )∈ε, where r∈R denotes the type of the edge of the graph network, i.e., the relationship type. The information of a specific node is calculated as follows:
[0069] where denotes the node v i the hidden layer information of the first l layer, denotes the node v i the hidden layer information of the first l+ layer, denotes the node v j the hidden layer information of the first l layer, and σ denotes an activation function, denotes the neighbor node of the node i with the r-type edge, is a normalization constant for a specific problem, which can be learned or pre-set, denotes the weight matrix of the relationship r in the first l layer, denotes the feature transformation weight of the node itself in the first l layer.
[0070] R-GCN needs to specify a transformation function for each type of edge. In order to reduce the computational overhead, the transformation function parameters of different types of relationships are shared, and a base function decomposition is used:
[0071] denotes a set of base matrices of the first l layer, and there are B base matrices in total, is used to construct all relationship weights, is the linear combination coefficient of the relationship r to the base matrix .
[0072] For entity classification, a plurality of R-GCN layers are stacked, and a softmax function is further passed after the output of the last layer. Ignoring the unlabeled nodes, the cross-entropy loss Loss1 is minimized for all labeled nodes as follows:
[0073] where Y is a set of labeled nodes, K represents the output dimension, denotes the output of the i-th labeled node at the k-th position, L represents the output layer, and t ik denotes the label real value of the i-th labeled node at the k-th position.
[0074] Based on the entity extraction, attribute extraction, and relationship extraction implemented in the previous article, we import the Neo4j graph database to build a knowledge graph (KG), KG = {E, R, F}, where E represents the entity set {e1, e2, ..., e E}, R represents the relationship set between different entities {r1,r2,...,r R}, F represents the triple set {f1,f2,...,f T}, each triple f is defined as (s, r, o), where s, r, o represent the head entity, relation, and tail entity respectively, and T represents the total number of triples.
[0075] Step 4: Optimize and dynamically update the knowledge graph: Use graph convolutional networks and long short-term memory networks to optimize and update the semantic connections between node entities, complete knowledge merging and processing, and dynamically update the abnormal knowledge graph. Specifically, this includes: Step 4.1, Knowledge Merging and Processing: Although Step 3 has completed preliminary knowledge representation, the knowledge graph is still incomplete, resulting in slightly poor performance when executing downstream tasks. To avoid this, knowledge merging and processing of the knowledge graph are required, and a dynamic update mechanism must be established. Knowledge merging and processing primarily include link prediction, knowledge reasoning, and entity disambiguation. The dynamic update mechanism primarily relies on the LSTM (Long Short-Term Memory) network to update the graph convolutional network parameter matrix, thereby utilizing temporal information.
[0076] Link prediction is based on a general model prediction framework: original knowledge graph, graph neural network encoder, decoder, and knowledge graph link prediction. R-GCN uses graph neural networks to implement link prediction tasks, demonstrating the process of updating information for a single node in the knowledge graph representation. Information is collected from adjacent nodes, aggregated and updated into the representation of the node itself, and continuously updated through downstream tasks until the node vector reaches a fixed point. After the knowledge graph representation is learned, the DisMult decoder is used to generate a likelihood score for each potential edge in the graph, ultimately selecting the most likely edge prediction. The DisMult (Distributed Multiplicative Model) decoder is a common decoder model in the field of knowledge graph embedding (KGE) and will not be discussed in detail here.
[0077] Define a triple (s, r, o), where s and o represent two entities, relation r and diagonal matrix R respectively. r ∈R d×d One-to-one correspondence, the scoring function f of the triple (s, r, o) is calculated as follows:
[0078] Among them, e s Represents the vector mapped by the encoder composed of R-GCN, e o Represents the vector mapped by the encoder composed of R-GCN. The superscript T represents the transpose.
[0079] We use negative sampling training. For each positive sample (s, r, o), we randomly destroy its head entity or tail entity to generate ω negative samples. The total training set T includes positive samples (y=1) and negative samples (y=0). y is the sample type identifier. The cross entropy loss function Loss2 is designed as follows:
[0080] in is a subset of the edges E in the graph network G=(V,E,R), σ refers to the logistic sigmoid function, which maps the score to the interval [0,1], indicating the probability that the triple is true.
[0081] The DistMult score f(s, r, o) determines the ranking order of candidate triplets, and evaluation metrics (such as MRR and HITS@K) completely rely on this ranking result. Since global reasoning of anomaly associations requires a comprehensive evaluation of ranking quality, the average reciprocal ranking MRR is used here:
[0082] Among them, S represents the set of triples of the test set, that is, the set defined in step 3 F= {f1,f2,...,f T},rank i Indicates the ranking of the i-th triple in the prediction results.
[0083] Knowledge reasoning uses specific methods to infer new relationships or identify erroneous information based on existing data, addressing the problem of incomplete knowledge graphs. Graph Convolutional Networks (GCNs) can be used for knowledge reasoning. Their core approach is to learn embedded representations of entities through convolution operations within a graph structure. GCNs can capture information about an entity's neighboring nodes and update the node's feature representation layer by layer, effectively learning complex relationships between entities.
[0084] GCN aggregates information about neighboring nodes through graph convolution operations and updates the feature representation of nodes. Its basic graph convolution formula is:
[0085] Where A represents the original adjacency matrix of the graph, represents the adjacency matrix with self-connection, I is the identity matrix, and the degree matrix , Indicates whether nodes i and j are adjacent in the graph. If they are adjacent, it is 1; otherwise, it is 0. σ is the activation function. H (l) It is l The node feature matrix of the layer, H (l+1) is the node feature matrix of the l+1th layer, W (l) It is l The weight matrix of the layer.
[0086] In the knowledge reasoning stage, GCN can use the learned entity embeddings to predict new triples by calculating the scores of all possible tail entities and selecting the entity with the highest score as the prediction result.
[0087] Entity disambiguation is a challenging problem in knowledge graphs because real-world entities may appear in many different forms in the graph. Graph Convolutional Neural Networks (GCNs) can achieve entity disambiguation through the following process and formula: After using GCN to learn the embedding representation of each entity in the knowledge graph, context information aggregation is performed. For each entity reference in the text, it needs to be linked to the entity in the knowledge graph. This can be achieved by calculating the similarity between the context vector of the entity reference and the entity embedding vector in the knowledge graph. The similarity function can be cosine similarity, Euclidean distance, etc. Based on the similarity score, the most likely entity is selected as the disambiguation result of the entity reference. This is expressed by the following formula: Score(e,m)=Similarity(H e ,H m ) Among them, H e is the context vector referred to by entity e, H m is the embedding vector of entity m in the knowledge graph, Score represents the calculated score, and Similarity is the similarity calculation function.
[0088] The loss function Loss3 during training can be expressed as:
[0089] Where T is the set of triplets in the training set, m' is the negative entity used for contrastive learning during training, and γ is a hyperparameter used to control the boundary of the score.
[0090] Step 4.2, construct the dynamic evolution of the knowledge graph: In order to overcome the problem that the existing knowledge graph cannot update the node structure and semantic information in time due to the fact that most of the knowledge graphs are static, a dynamic knowledge graph update mechanism will be introduced. The specific implementation method is to use the graph convolutional neural network to capture the dependency relationship of the graph, and use the long short-term memory (LSTM) network to learn and evolve the parameter matrix of the lth layer of the GCN model at each time t. , thereby using knowledge embedding to construct the dynamic evolution of the knowledge graph.
[0091] First, rewrite the expression of the convolutional neural network into the expression at time t:
[0092] Where H (l) is the first l The node feature matrix of the layer, H (l+1) is the node feature matrix of the l+1th layer at time t, W (l) is the first l The weight matrix of the layer, is the degree matrix at time t, for , X t is the input at time t, which is the initial feature matrix of the graph. In this method, it represents the encoding matrix of each node in the graph.
[0093] The core of weight evolution is to use long short-term memory network (LSTM) to update the weight matrix at time t based on current and historical information. 。 In a typical graph convolutional network is constant, while in a dynamic knowledge graph, the graph structure at different times t may be different. In order to reflect the evolution of the graph structure over time, let It changes with time steps and convolution layers, that is:
[0094] In the formula, LSTM represents the long short-term memory network, and the long short-term memory network is used to update , the specific update formula is as follows:
[0095] in, represents the parameter matrix of the lth layer of the graph convolutional network at time t-1, Indicates the time t-1 l The node feature matrix of the layer, σ represents the activation function.
[0096] W f is the weight matrix of the forget gate, f tis the forget gate, which determines how many old states to retain; W i is the weight matrix of the input gate, i t is the input gate, which controls the inflow of new information; W o is the weight matrix of the output gate, O t is the output gate, which controls the output of the current state; W c is the weight matrix of the candidate state, Represents the candidate state, generates a new candidate value, c t Represents the update state; the final output parameter matrix , It represents the Hadamard product, which is the element-wise multiplication of two matrices or vectors, and tanh represents the hyperbolic tangent function.
[0097] Using long short-term memory network, use the t-1 time Go to update , which can analyze the impact of graph evolution at different times and dynamically update the knowledge graph.
[0098] Step 5: Application of knowledge graph reasoning related to abnormal knowledge of wind turbines: including intelligent retrieval of fault knowledge, identification of repetitive faults, role-based fault data analysis, analysis of the effectiveness of corrective measures, and predictive fault maintenance.
[0099] Intelligent retrieval based on the fault knowledge base primarily relies on entities, attributes, and associations designed within the knowledge graph. Natural language processing methods are used to automatically extract knowledge triples to form the core knowledge base. Intelligent retrieval can be divided into two forms: the question-and-answer form. This form defines templates for frequently asked questions. When a user asks a question, the question is structured. Using word segmentation tools, part-of-speech analysis, and syntactic analysis, the structure, sentence structure, and key entities of the question are extracted. The corresponding template is identified. Once the defined question template is matched, matching entities and attributes are found within the knowledge graph. The best matching result is taken as the answer to the question and assembled into a sentence that is returned to the user. The other form uses a multi-keyword combination mode. The user enters a keyword for the fault or problem they are interested in. The fault or keyword is identified. If it is a term, the term is entered into the graph to search for related fault information. If it is not a term, only the keyword combination is matched to find fault information containing the term in the fault description, thus achieving intelligent fault information retrieval.
[0100] Repetitive faults occur when two faults differ in their process and symptoms, but their underlying logic and causes are consistent. This type of fault is a key focus for quality improvement. A knowledge graph-based fault knowledge base can effectively identify repetitive faults. This is because the knowledge graph uses entity alignment to identify the similarity between two fault entities, enabling accurate queries using experimentally determined thresholds.
[0101] Role-based fault analysis involves leveraging the information in the fault knowledge base to improve product quality from different perspectives. Based on the actual business roles, designers, process engineers, and inspectors are more interested in fault knowledge. Different roles prioritize different fault types and attributes. Leveraging knowledge graphs to unlock the potential value of data helps practitioners quickly identify root causes and solutions.
[0102] Corrective action effectiveness analysis is a crucial component of fault management in the industrial sector. By applying a knowledge graph-based fault knowledge base, we can associate fault entities with corresponding corrective actions. The core idea is to use a recurring fault identification method to identify a specific fault. We search for the number of recurring faults in the past few months since the fault occurred, then search for the number of recurring faults in the future months after the fault occurred. Dividing these two values yields an effectiveness coefficient, which is then used to evaluate the effectiveness of the corrective action.
[0103] Predictive maintenance is an important goal to ensure the smooth operation of instruments and equipment and avoid major accidents and losses. The abnormal knowledge association method of wind turbines based on knowledge graph can model the fault information of devices from the historical knowledge base, comprehensively consider relevant semantic information, and construct a nonlinear least squares function to fit the equipment cycle life, thereby realizing predictive maintenance of equipment.
[0104] See also Figure 3 , an embodiment of the present invention provides a wind turbine abnormality knowledge association device based on a knowledge graph, including one or more processors for implementing a wind turbine abnormality knowledge association method based on a knowledge graph in the above embodiment.
[0105] The implementation of the wind turbine abnormality knowledge association device based on the knowledge graph of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for execution. From the hardware level, if Figure 3As shown, this is a hardware structure diagram of any device with data processing capability where the abnormal knowledge association device of a wind turbine generator system based on knowledge graph is located. Figure 3 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in the embodiment may also include other hardware according to the actual function of the device with data processing capabilities, which will not be described in detail.
[0106] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0107] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0108] An embodiment of the present invention also provides a readable storage medium on which a program is stored. When the program is executed by a processor, a wind turbine abnormality knowledge association method based on a knowledge graph in the above embodiment is implemented.
[0109] The readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0110] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "an," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0111] It should be understood that although the terms first, second, third, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when..." or "when..." or "in response to determining."
[0112] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included in the scope of protection of one or more embodiments of this specification.
Claims
1. A wind turbine anomaly knowledge association method based on knowledge graph, characterized in that: include: Step 1: Based on the operation and maintenance needs of the wind power industry, distributed multi-source heterogeneous data and knowledge collection is carried out; Step 2: Establish the OPC UA information model of the wind power equipment node and determine the node's common attributes and reference types; Step 3: Use prompt word engineering to guide the industrial multimodal industrial big model to standardize the description and identification of node data, complete entity extraction and attribute extraction, and use graph convolutional networks to extract and represent relationships, and then import them into the graph database to build an abnormal knowledge graph; Step 4: Use graph convolutional networks and long short-term memory networks to optimize and update the semantic connections between node entities, complete knowledge merging and processing, and dynamically update the abnormal knowledge graph; Step 5: Aiming at the application scenario of abnormal knowledge, perform knowledge reasoning by mining entity associations based on the abnormal knowledge graph.
2. The wind turbine anomaly knowledge association method based on knowledge graph according to claim 1 is characterized in that: The step 1 comprises: Step 1.1: Multi-source knowledge data integration: Data sources include internal and external data sources of wind power equipment operation and maintenance enterprises. Internal data sources include real-time equipment operation data, system operation and maintenance logs, and internal documents. External data sources include industry standards and public data sets, academic literature, and patents. Step 1.2, data knowledge preprocessing: Based on data quality, construct feature engineering to perform mean value replacement, data fitting, and data elimination on abnormal data and missing data, conduct rationalization tests on knowledge sources, and remove low-quality information sources.
3. The wind turbine anomaly knowledge association method based on knowledge graph according to claim 1 is characterized in that: The step 2 includes: Step 2.1, define the OPC UA information model: The node types of the OPC UA information model include objects, object types, variables, variable types, data types, reference types, methods, and views. The common attributes of the OPC UA information model include node identifiers, node types, browse names, display names, descriptions, write masks, and user write masks. Additional attributes of the OPC UA information model include whether it is abstract, whether it is symmetric, and reverse names. Step 2.2, build the OPC UA information model: The establishment of the OPC UA information model includes four steps: obtaining requirements, defining the type model, instantiating the information model, and exporting XML and CSV documents.
4. The wind turbine anomaly knowledge association method based on knowledge graph according to claim 1 is characterized in that: The step 3 includes: Step 3.1: Standardize the OPC UA information model for wind turbine equipment: Use the multimodal industrial macro model fine-tuned with industrial data to preprocess the OPC UA information model from Step 2, clean up illogical structured information, and combine the collected knowledge data to enhance embedded reasoning capabilities in the form of thought chains, thereby expanding the relationships between different individuals. Step 3.2, Entity Extraction and Attribute Extraction: Based on a variety of prompt word engineering design methods, including zero-shot learning, small-shot learning, role prompting, and thought chaining, the multimodal industrial large model is guided to normalize the OPC UA information model into an ontology under the attribute graph model to complete entity extraction and attribute extraction; Step 3.3, forming knowledge representation: performing relationship extraction and representation on the pre-processed OPC UA information model data based on the relationship graph convolutional neural network.
5. The wind turbine anomaly knowledge association method based on knowledge graph according to claim 1 is characterized in that: The step 4 comprises: Step 4.1, knowledge merging and processing: Complete the knowledge merging and processing process based on graph convolutional neural networks, including link prediction, knowledge reasoning, and entity disambiguation, to form a standard knowledge representation; Step 4.2, construct the dynamic evolution of the knowledge graph: Use the graph convolutional neural network to capture the dependencies of the graph, use the long short-term memory network to learn and evolve the parameter matrices of different layers of the graph convolutional network at each moment, and thus use knowledge embedding to construct the dynamic evolution of the knowledge graph.
6. The wind turbine anomaly knowledge association method based on knowledge graph according to claim 5 is characterized in that: The link prediction in step 4.1 is based on the framework of original knowledge graph - graph neural network encoder - decoder - knowledge graph link prediction; the graph convolutional neural network uses the graph neural network to implement the link prediction task, explicitly representing the process of updating the information of a single node in the knowledge graph, collecting information from adjacent nodes, aggregating and updating it into the representation of this node, and continuously iterating and updating through downstream tasks until the node vector reaches a fixed point. After learning the abnormal knowledge graph, the DisMult decoder is used to generate a likelihood score for each potential edge in the graph, and finally selects the most likely edge prediction; The knowledge reasoning in step 4.1 includes: the graph convolutional neural network uses the learned entity embedding to predict new triples; By calculating the scores of all possible tail entities and selecting the entity with the highest score as the prediction result; In step 4.1, the graph convolutional neural network implements entity disambiguation through the following process: Each entity reference in the text is linked to the entity in the knowledge graph by calculating the similarity between the context vector of the entity reference and the entity embedding vector in the knowledge graph. Based on the similarity score, the most likely entity is selected as the disambiguation result of the entity reference.
7. The wind turbine anomaly knowledge association method based on knowledge graph according to claim 1 is characterized in that: The step 5 comprises: Knowledge graph reasoning applications of abnormal knowledge association of wind turbines: Knowledge graph reasoning applications of abnormal knowledge association include intelligent retrieval of fault knowledge, repetitive fault identification, role-based fault data analysis, corrective action effectiveness analysis, and predictive fault maintenance.
8. A wind turbine abnormality knowledge association device based on knowledge graph, characterized in that: It includes one or more processors for implementing a wind turbine abnormality knowledge association method based on a knowledge graph as described in any one of claims 1-7.
9. A readable storage medium, characterized in that A program is stored thereon, and when the program is executed by a processor, a wind turbine abnormality knowledge association method based on a knowledge graph as described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Pediatric disease prediction system based on big data
CN118538400A
Aeronautical manufacturing field fault knowledge graph modeling method based on OPC UA protocol and improved CASREL model
CN119168035A
Wind turbine generator operation and maintenance knowledge base construction method based on large model and mechanism self-learning
CN120354922A
Document-level intelligent manufacturing process flow relation extraction method
CN120493930A
Cited By
Method, system and equipment for analyzing abnormal phenomenon of semiconductor structure
CN121542487A