Wind power fault diagnosis operation and maintenance method based on multi-source data and knowledge retrieval enhancement
By combining knowledge graphs and generating large models, dynamically analyzing the multi-source data flow of the SCADA system, the problem of data dependence and knowledge base disconnection in wind turbine fault diagnosis is solved, and efficient and accurate fault diagnosis and operation and maintenance support is achieved.
Patent Information
- Application Number
- CN202510362045.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
The existing wind turbine fault diagnosis technology has problems such as strong data dependence, insufficient generalization capability, inaccurate inference results, and disconnection of static knowledge base from dynamic operating data, resulting in insufficient real-time and adaptability of diagnosis.
Combining the structured domain knowledge of the knowledge graph and the independent learning ability to generate large models, we dynamically analyze the real-time multi-source data flow of the SCADA system, and realize the intelligence of wind power fault diagnosis through multimodal data acquisition and preprocessing, knowledge base construction, multi-grained problem decomposition and retrieval optimization, knowledge retrieval enhancement generation and result feedback optimization, and knowledge retrieval enhancement generation and result feedback optimization.
It significantly improves the accuracy and efficiency of fault diagnosis, reduces the "illusion" phenomenon of generating answers, improves the accuracy and comprehensiveness of fault recognition, shortens the response time of fault diagnosis, and enhances the interpretability and scientific operation and maintenance of the inference process.
Smart Images

Figure CN120297410A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wind turbine maintenance, and particularly to a wind power fault diagnosis and operation and maintenance method based on multi-source data and enhanced knowledge retrieval. Background Art
[0002] At present, the fault diagnosis technologies of wind turbines are mainly divided into data-driven diagnosis methods and knowledge-driven diagnosis methods.
[0003] The data-driven diagnosis method performs state detection and fault analysis by using the historical operation data of each component of the wind turbine. However, due to the complex and highly uncertain operating environment of the wind turbine, such methods face great challenges in applicability and accuracy in practical applications. The knowledge-driven diagnosis method has gradually become the mainstream choice for wind turbine fault diagnosis in recent years. Such methods rely on a deep understanding of the physical process of the system, but the complexity of the wind power system makes it very difficult to establish an accurate mathematical model. The inevitable simplifying assumptions in the model limit its prediction ability in complex environments and are difficult to comprehensively reflect the true behavior of the system. At the same time, the generalization ability of the model is poor.
[0004] Currently, most wind power enterprises still mainly rely on the SCADA (Supervisory Control and Data Acquisition) system and manual experience to diagnose the faults of wind turbines, and can only analyze the current operating state of the turbines, which is difficult to meet the urgent needs of the wind power industry for efficient and accurate operation and maintenance. Moreover, although the SCADA system can collect the real-time operation data of wind turbines (such as vibration, temperature, speed, etc.), its analysis function is mainly limited to threshold alarm and simple trend analysis, lacking the intelligent reasoning ability for complex faults. Existing fault diagnosis systems usually separate the SCADA multi-source data from the knowledge base, resulting in the inability to deeply integrate real-time data with domain knowledge and making it difficult to achieve accurate diagnosis in a dynamic environment. In addition, traditional methods cannot directly analyze the SCADA multi-source data stream through large models, resulting in insufficient real-time performance and adaptability of fault diagnosis.
[0005] At the same time, when existing pre-trained large models have insufficient data, they often fabricate answers, resulting in the "hallucination" phenomenon, making it difficult to guarantee the accuracy of the diagnosis results. In recent years, some studies have begun to use retrieval augmentation technology (RAG) to improve the "hallucination" problem of large models. However, traditional retrieval augmentation generation methods usually rely on retrieval mechanisms based on text or vector similarity, which may lead to incomplete or even duplicate retrieval results in the field of fault diagnosis that requires complex reasoning. Summary of the Invention
[0006] To solve the problems of strong data dependence, insufficient generalization ability, inaccurate inference results, and disconnection between static knowledge bases and dynamic operation data in the existing technologies, the present invention proposes a wind power fault diagnosis and operation and maintenance method enhanced by multi-source data and knowledge retrieval. By combining the structured domain knowledge of a knowledge graph with the autonomous learning ability of a generative large model, and dynamically parsing the real-time multi-source data stream of the SCADA system, the above problems are solved.
[0007] This application discloses a wind power fault diagnosis and operation and maintenance method enhanced by multi-source data and knowledge retrieval, including the following steps:
[0008] S1. Multi-modal data collection and preprocessing: Collect multivariate data related to wind turbine fault diagnosis and preprocess the data;
[0009] S2. Knowledge base construction: Extract knowledge from the preprocessed data and construct a corresponding wind turbine operation and maintenance knowledge graph database;
[0010] S3. Multi-granularity problem decomposition and retrieval optimization: After receiving the user's query, retrieve the user's recent conversation records through the conversation context for prompt word optimization processing, and use the natural language processing ability of the large model to split the user's problem into multiple simple sub-problems;
[0011] S4. Knowledge retrieval enhanced generation: Based on the sub-problem list, perform SPO triple entity recognition, and combine knowledge graph sub-graph retrieval with dynamic weight adjustment of SCADA real-time data to generate the final answer;
[0012] S5. Result generation and feedback optimization: Hand over the final answer to the large model for processing, realize the standardized output of the diagnosis result and the continuous optimization of the system, and generate the result required to be returned to the user.
[0013] Preferably, the S1 includes the following steps:
[0014] S11. Divide the fault diagnosis data of the wind turbine into structured data and unstructured data;
[0015] Among them, the structured data includes sensor time series data (vibration, temperature, speed, etc.), SCADA system logs, equipment parameter tables, fault code libraries, work order records, etc., which have clear field definitions and numerical characteristics; the unstructured data covers operation and maintenance manuals, expert diagnosis reports, text descriptions of repair work orders, PDF documents of fault cases, voice work order recordings (which need to be processed into text), etc., and deep semantic parsing is required;
[0016] S12. Multimodal data preprocessing, using a sliding window algorithm based on semantic segmentation, combined with the Sentence-BERT model to calculate the paragraph similarity, ensuring context coherence, performing different text extraction processes on different types of data, and after extracting the text, performing chunk cutting on the text;
[0017] S13. Real-time access to multi-source data streams of the SCADA system, including sensor time-series data, device status codes, alarm logs, etc., performing real-time cleaning, alignment, and feature extraction on the data through time-series analysis algorithms, generating dynamic data chunks and storing them in a cache queue for subsequent real-time retrieval and inference calls.
[0018] Preferably, the S2 includes the following steps:
[0019] S21. Construct an association constraint system in the field of wind turbine operation and maintenance, defining association constraint conditions such as core entity types and relationship types in the field of wind turbine operation and maintenance;
[0020] S22. Utilize the natural language processing ability of the locally deployed large wind power operation and maintenance model and the pre-constructed private domain knowledge Schema to perform knowledge extraction on the chunked text according to the prompt words of the Schema, including named entity extraction and relationship extraction, and construct the index relationship between the sub-graph and the text chunks;
[0021] S23. Use BGE-M3 and Sentence-BERT to perform vectorized embedding representation on the extracted entities, relationships, and triples;
[0022] S24. Establish a dynamic association mechanism between SCADA real-time data and the knowledge graph, bind the device parameters and historical alarm records in the SCADA system to the fault types and maintenance plan entities in the knowledge graph, and support dynamic knowledge update based on real-time data;
[0023] S25. Save the knowledge extraction results and the association binding information between SCADA and the knowledge base to the Neo4j graph database.
[0024] Preferably, the S23 includes the following steps:
[0025] S231. Vectorization of text data: For unstructured text data, use the Sentence-BERT model for semantic embedding to generate a vector representation of the text paragraph;
[0026] S232. Vectorization of structured data: For structured data, use the BGE-M3 model for feature extraction and vectorized representation;
[0027] S233. Multimodal Data Fusion: Align and fuse text data and structured data representations to generate a unified multimodal vector representation for subsequent retrieval and reasoning.
[0028] Preferably, the said S3 includes the following steps:
[0029] S31. Conduct a complexity assessment of the question, calculate the doubt degree and dependency parsing depth of the user query through a large model, and define multi-level complexity classification;
[0030] S32. Split questions of different complexities to generate the following strategies:
[0031] Explicit split: For queries containing coordinating conjunctions, directly split them into independent sub-questions;
[0032] Implicit reasoning split: Use chain-of-thought prompting to guide the large model to generate intermediate questions;
[0033] S33. Integrate the split sub-questions into a sub-question list in the order of precedence.
[0034] Preferably, the said S31 includes the following steps:
[0035] Doubt degree calculation: According to the user query Q, tokenize it by word to obtain the word sequence {w1, w2, …, w n}, and the large model predicts the conditional probability P(w i |w1, w2, …, w i-1 ) based on the context, and calculate the overall doubt degree through the cross-entropy loss function:
[0036]
[0037] where N is the number of words in the query, and the higher the PP(Q), the higher the semantic complexity of the query or the uncertainty of the large model;
[0038] Dependency parsing depth calculation: First, use the en_core_web_trf model of the dependency parsing tool spaCy to parse the user query Q to generate a dependency syntax tree, and the dependency depth Depth(Q) is defined as the path length from the root node to the deepest leaf node;
[0039] Complexity grading rule: Combine the doubt degree and the dependency parsing depth, and set thresholds through experiments to divide into three levels of complexity:
[0040] Simple query (single intent), such as "What is the solution to the E102 fault?":
[0041] PP(Q) < 50 and Depth(Q) ≤ 3;
[0042] Compound query (multiple sub-questions): such as "How to distinguish between gearbox overheating and generator bearing failure?":
[0043] 50 ≤ PP(Q) < 100 and 3 < Depth(Q) ≤ 6;
[0044] Complex reasoning query (requiring multi-hop reasoning): such as "E102 was reported multiple times at the same wind farm last month. What are the possible environmental factors?":
[0045] PP(Q) > 100 and Depth(Q) > 6.
[0046] Preferably, the S4 includes the following steps:
[0047] S41. Query and perform retrieval reasoning according to the sub-question list, use knowledge extraction to perform entity recognition on the sub-questions, and extract the SPO triples containing logical relationships in the sub-queries;
[0048] S42. Check the entity integrity, summarize the entity pairs, and perform retrieval queries in the constructed knowledge graph knowledge base according to the extracted triples;
[0049] S43. Generate a response, organize the sub-graphs retrieved by all sub-queries in S42 into a text answer summary, and determine whether the problem is solved. If it is solved, submit it to the large model to generate the final answer according to the summary of the sub-question results. If the problem is not solved, the large model generates a supplementary question and repeats the above steps;
[0050] S44. During the retrieval process, synchronously call the SCADA real-time data cache, dynamically adjust the weights of the retrieval results in combination with the current device status, and preferentially match the knowledge sub-graphs most relevant to the real-time working conditions to ensure the timeliness and pertinence of the diagnostic suggestions.
[0051] Preferably, the realization of the standardized output of the diagnostic results requires structured generation, and the following hierarchical response generation templates are set:
[0052] Emergency failure: Red mark, with "Check immediately for shutdown" displayed on the first line;
[0053] Early warning prompt: Yellow mark, suggesting "Arrange an inspection within 72 hours";
[0054] Knowledge base missing: Blue mark, triggering the manual expert access process.
[0055] Preferably, the feedback closed-loop mechanism is: update the knowledge base through the operation and maintenance personnel's adoption / modification / rejection operations on the diagnostic results, and calculate the KPI indicators of accuracy, recall rate, and response time monthly.
[0056] Preferably, the real-time verification mechanism is as follows: compare the generated maintenance plan with the SCADA real-time data. If the equipment status is not restored, trigger a secondary diagnosis and update the fault solution cases in the knowledge base.
[0057] Advantages of the present invention:
[0058] 1. Combination of knowledge graph and large model: By introducing a knowledge graph and combining the reasoning ability of the large model, reasoning is carried out using structured domain knowledge, significantly improving the accuracy of fault diagnosis. The knowledge graph provides the structure of the wind turbine, fault types and association rules, supporting the generation model to generate accurate and reasonable diagnostic results in different fault scenarios.
[0059] 2. Retrieval-Augmented Generation (RAG) technology optimizes the generation model: Through the Retrieval-Augmented Generation (RAG) technology, the system queries relevant knowledge and historical fault data in the knowledge graph before generating an answer, providing accurate reference information for the generation model and effectively reducing the "hallucination" phenomenon. Experiments show that the fault diagnosis accuracy has increased by 15%-20%, especially performing excellently in complex or rare fault scenarios.
[0060] 3. Intelligent reasoning of multi-modal data: The system integrates multi-modal data such as sensor data, operation logs, and fault cases, and improves the accuracy and comprehensiveness of fault recognition through joint reasoning. The question-answering system can provide targeted fault diagnosis suggestions and clarify the solution for the operation and maintenance personnel.
[0061] 4. Enhanced interpretability of the reasoning process: Combining the knowledge graph and the generation model makes the reasoning process more transparent and interpretable. It not only outputs the diagnostic conclusion but also shows the key factors and knowledge sources in the reasoning process, helping the operation and maintenance personnel understand the fault cause and repair plan, and improving the scientific nature of maintenance decisions.
[0062] 5. Optimization of fault diagnosis efficiency and response time: Through the retrieval-augmented generation large model, it can quickly respond to the fault description or sensor data input by the user, achieving efficient diagnosis. Experiments show that the fault diagnosis response time is shortened by about 25% compared with traditional methods, significantly improving the operation and maintenance efficiency. Description of the Drawings
[0063] Figure 1 It is a flowchart of the wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to an embodiment of the present invention;
[0064] Figure 2 It is a flowchart of multi-modal data acquisition and preprocessing according to an embodiment of the present invention;
[0065] Figure 3 It is a flowchart of knowledge base construction according to an embodiment of the present invention;
[0066] Figure 4 Flow chart of multi - granularity problem decomposition and retrieval optimization for embodiments of the present invention;
[0067] Figure 5 Flow chart of knowledge retrieval enhanced generation for embodiments of the present invention;
[0068] Figure 6 Flow chart of result generation and feedback optimization for embodiments of the present invention. Detailed implementation manners
[0069] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the following provides embodiments with reference to the accompanying drawings and further elaborates on the present application in detail.
[0070] Embodiments of the present application disclose a wind power fault diagnosis and operation and maintenance method based on multi - source data and knowledge retrieval enhancement. The process is as Figure 1 shown and includes the following steps:
[0071] S1. Multi - modal data collection and pre - processing: Collect multivariate data related to wind turbine fault diagnosis and pre - process the data. The process is as Figure 2 shown.
[0072] S11. Divide the fault diagnosis data of the wind turbine into structured data and unstructured data.
[0073] Among them, structured data includes sensor time - series data (vibration, temperature, speed, etc.), SCADA system logs, equipment parameter tables, fault code libraries, work order records, etc., which have clear field definitions and numerical characteristics; unstructured data covers operation and maintenance manuals, expert diagnosis reports, text descriptions of maintenance work orders, fault case PDF documents, voice work order recordings (which need to be processed into text), etc., and deep semantic parsing is required.
[0074] S12. Multi - modal data pre - processing: Adopt a sliding window algorithm based on semantic segmentation (window size 512 tokens, overlap rate 20%), combine with the Sentence - BERT model to calculate the paragraph similarity to ensure context coherence, perform different text extraction processes on different types of data, and after extracting the text, perform block cutting on the text.
[0075] S13. Real - time access to multi - source data streams of the SCADA system, including sensor time - series data (vibration, temperature, speed, etc.), equipment status codes, alarm logs, etc. Perform real - time cleaning, alignment and feature extraction on the data through a time - series analysis algorithm (LSTM is used in this embodiment), generate dynamic data blocks and store them in a cache queue for subsequent real - time retrieval and inference calls.
[0076] S2. Knowledge base construction: Extract knowledge from the preprocessed data and construct a corresponding wind turbine operation and maintenance knowledge graph database. The process is as follows Figure 3 as shown.
[0077] S21. Construct an association constraint system in the field of wind turbine operation and maintenance, and define association constraint conditions such as core entity types (fault types, equipment components, sensors, maintenance plans) and relationship types (causing, associated components, repair methods) in the field of wind turbine operation and maintenance.
[0078] S22. Utilize the natural language processing ability of the locally deployed large model for wind power operation and maintenance, as well as the pre-constructed private domain knowledge Schema, to extract knowledge from the segmented text according to the prompt words of the Schema, including named entity extraction and relationship extraction, and construct the index relationship between the subgraph and the text block. The large model for wind power operation and maintenance is a large model obtained by fine-tuning the general large language model (DeepSeek) for the field of wind power fault diagnosis
[0079] S23. Use BGE-M3 and Sentence-BERT to perform vectorized embedding representations on the extracted entities, relationships, and triples.
[0080] S231. Vectorization of text data: For unstructured text data (such as operation and maintenance manuals, fault case descriptions), use the Sentence-BERT model for semantic embedding to generate vector representations of text paragraphs.
[0081] S232. Vectorization of structured data: For structured data (such as sensor time series data, equipment parameter tables), use the BGE-M3 model for feature extraction and vectorized representation.
[0082] S233. Multi-modal data fusion: Align and fuse the representations of text data and structured data to generate a unified multi-modal vector representation for subsequent retrieval and reasoning.
[0083] S24. Establish a dynamic association mechanism between SCADA real-time data and the knowledge graph, bind the equipment parameters and historical alarm records in the SCADA system to the fault type and maintenance plan entities in the knowledge graph, and support dynamic knowledge update based on real-time data.
[0084] S25. Save the knowledge extraction results and the association binding information between SCADA and the knowledge base to the Neo4j graph database.
[0085] S3. Multi-granularity problem decomposition and retrieval optimization: After receiving the user's query, retrieve the user's recent conversation records through the conversation context for prompt word optimization processing, and use the natural language processing ability of the large model to split the user's problem into multiple simple sub-problems. The process is as follows Figure 4As shown. The large model used in the embodiments of this application is a general large language model, such as DeepSeek, ChatGpt, Llama, etc. In this embodiment, DeepSeek is used as the large model.
[0086] S31. Evaluate the complexity of the problem, calculate the doubt degree and the depth of dependency parsing of the user's query through the large model, and define multi-level complexity classification.
[0087] Calculation of perplexity. Perplexity is used to quantify the semantic complexity of the query and the uncertainty of the large model prediction. According to the user's query Q, tokenize it by words to obtain the word sequence {w1, w2, …, w n}, and the large model predicts the conditional probability P(w i |w1, w2, …, w i-1 ) based on the context. Calculate the overall perplexity through the cross-entropy loss function:
[0088]
[0089] where N is the number of words in the query. The higher the PP(Q), the higher the semantic complexity of the query or the uncertainty of the large model.
[0090] Calculation of dependency parsing depth. The dependency parsing depth reflects the complexity of the syntactic structure of the query. First, use the en_core_web_trf model of the dependency syntax analysis tool spaCy to parse the user's query Q to generate a dependency syntax tree. The dependency depth Depth(Q) is defined as the path length from the root node to the deepest leaf node (i.e., the maximum depth). For example, the dependency tree of the sentence "Check whether the vibration of the gearbox is abnormal" may contain the path "Check → Vibration → Abnormal", and the depth is 3.
[0091] Complexity classification rules. Combine the perplexity and the dependency parsing depth, and set thresholds through experiments to divide into three levels of complexity:
[0092] Simple query (single intention), such as "What is the solution to the E102 fault?":
[0093] PP(Q) < 50 and Depth(Q) ≤ 3;
[0094] Compound query (multiple sub-questions): such as "How to distinguish between overheating of the gearbox and generator bearing failure?":
[0095] 50 ≤ PP(Q) < 100 and 3 < Depth(Q) ≤ 6;
[0096] Complex inference query (requiring multi-hop inference):
[0097] PP(Q) > 100 and Depth(Q) > 6.
[0098] S32. Split problems of different complexities to generate the following strategies:
[0099] Explicit split: For queries containing coordinating conjunctions (such as "and", "or"), directly split them into independent sub-problems;
[0100] Implicit reasoning split: Use the Chain-of-Thought prompt to guide the large model to generate intermediate problems.
[0101] S33. Integrate the split sub-problems into a sub-problem list in the order of precedence.
[0102] S4. Knowledge retrieval enhanced generation. Based on the sub-problem list, perform SPO triple entity recognition, combine knowledge graph sub-graph retrieval with SCADA real-time data dynamic weight adjustment to generate the final answer. The process is as Figure 5 shown.
[0103] S41. Query and perform retrieval reasoning according to the sub-problem list generated by S3, use knowledge extraction to perform entity recognition on the sub-problems, and extract the SPO triples containing logical relationships in the sub-queries.
[0104] S42. Check entity integrity, summarize entity pairs, and perform retrieval queries in the constructed knowledge graph knowledge base according to the extracted triples.
[0105] S43. Generate a response, organize the sub-graphs retrieved by all sub-queries in S42 into a text answer summary, and determine whether the problem is solved. If it is solved, hand it over to the large model to generate the final answer based on the summary of the sub-problem results. If the problem is not solved, the large model generates a supplementary question and repeats the above steps.
[0106] S44. During the retrieval process, synchronously call the SCADA real-time data cache, combine the current device status (such as temperature anomaly, vibration amplitude) to perform dynamic weight adjustment on the retrieval results, and preferentially match the knowledge sub-graph most relevant to the real-time working condition to ensure the timeliness and pertinence of the diagnostic suggestions. The most relevant knowledge sub-graph is measured through the following dimensions:
[0107] Real-time data matching degree: Whether the retrieval result is highly relevant to the SCADA real-time data (such as temperature, vibration, rotational speed);
[0108] Fault type consistency: Whether the retrieval result is consistent with the current alarm or fault code;
[0109] Historical case similarity: Whether the retrieval result is highly similar to historical fault cases (such as the same equipment model, similar working conditions).
[0110] S5, Result generation and feedback optimization. The final answer is processed by the large model to achieve the standardized output of the diagnosis result and the continuous optimization of the system, generating the result required to be returned to the user. The process is as Figure 6 shown.
[0111] The steps for realizing the standardized output of the diagnosis result need to be generated structurally. The following grading response generation templates are set:
[0112] Emergency failure: Red mark, with "Check immediately for shutdown" displayed on the first line;
[0113] Early warning prompt: Yellow mark, suggesting "Arrange inspection within 72 hours";
[0114] Knowledge base missing: Blue mark, triggering the process of accessing human experts.
[0115] To realize the feedback optimization function, the system adopts a feedback closed-loop mechanism. The operation and maintenance personnel can perform "adopt / modify / reject" operations on the diagnosis results generated by the system to ensure the accuracy and practicality of the answers, and synchronize the newly added knowledge to the knowledge base in real time. At the same time, the system automatically calculates the key performance indicators (KPIs) every month, including accuracy (consistency with expert diagnosis), recall rate (fault coverage), and response time (<3 seconds), to quantitatively evaluate the system performance and continuously optimize the diagnosis effect, thus forming a closed-loop iterative process of "human feedback - knowledge update - effect evaluation" to continuously improve the intelligent level and engineering practicality of the system.
[0116] Real-time verification mechanism - The system compares the generated maintenance plan with the SCADA real-time data (such as the sensor readings after maintenance). If the diagnosis result does not restore the equipment status to normal, the secondary diagnosis process is automatically triggered, and the fault solution cases in the knowledge base are updated to form a dynamic iterative optimization process.
[0117] The present application aims to provide a fault diagnosis method for intelligent operation and maintenance of wind turbines that deeply integrates a knowledge retrieval enhanced generative large model with real-time data of the SCADA system, so as to solve the problems of strong data dependence, insufficient generalization ability, inaccurate reasoning results, and disconnection between the static knowledge base and dynamic operation data in the prior art. By combining the structured domain knowledge of the knowledge graph with the autonomous learning ability of the generative large model, and dynamically parsing the real-time multi-source data stream of the SCADA system (such as sensor time series data, device parameters, and fault codes), where the SCADA system, as the core source of multi-source data, provides key real-time data on the operation of wind turbines, including various parameters such as wind speed, power, temperature, and vibration. These data, together with other sources (such as historical maintenance records and meteorological data), constitute a multi-source data system. The solution proposed in the present application can significantly improve the accuracy, reliability, and real-time performance of fault diagnosis in a limited fault data environment. Optimize the retrieval enhanced reasoning process of the large model using the knowledge graph, effectively avoid the "hallucination" phenomenon of generating answers, and at the same time enhance the dynamic adaptability of diagnosis through the deep integration of real-time data and historical cases, providing efficient and accurate intelligent operation and maintenance support for wind turbines, thereby reducing operation and maintenance costs, improving the safety and stability of the system, and promoting the high-quality development of smart wind power.
[0118] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A wind power fault diagnosis and operation and maintenance method enhanced based on multi-source data and knowledge retrieval, characterized in that It includes the following steps: S1. Multimodal data collection and preprocessing: Collect wind turbine fault diagnosis data and preprocess the data; S2. Knowledge base construction: Extract knowledge from the preprocessed data and construct a corresponding wind turbine operation and maintenance knowledge graph database; S3. Multi-granularity problem decomposition and retrieval optimization: Evaluate the complexity of the user query through a large model, split the problem hierarchically by combining the doubt degree and the dependency parsing depth, and generate a list of sub-problems; S4. Knowledge retrieval enhanced generation: Based on the list of sub-problems, perform SPO triple entity recognition, combine knowledge graph sub-graph retrieval and dynamic weight adjustment of SCADA real-time data to generate the final answer; S5. Result generation and feedback optimization: Hand over the final answer to the large model for processing to achieve the standardized output of the diagnosis result, combine the feedback closed-loop mechanism for feedback optimization, and verify the effectiveness of the diagnosis result through the real-time verification mechanism.
2. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 1, wherein The S1 includes the following steps: S11. Divide the wind turbine fault diagnosis data into structured data and unstructured data; The structured data includes sensor time series data, SCADA system logs, equipment parameter tables, fault code libraries, and work order records, which have clear field definitions and numerical characteristics; The unstructured data includes operation and maintenance manuals, expert diagnosis reports, text descriptions of maintenance work orders, PDF documents of fault cases, and voice recordings of work orders, which require in-depth semantic parsing; S12. Multimodal data preprocessing: Adopt a sliding window algorithm based on semantic segmentation, combine with the Sentence-BERT model to calculate the paragraph similarity to ensure context coherence, perform different text extraction processes on different types of data, and after extracting the text, perform chunk cutting on the text; S13. Real-time access to multi-source data streams of the SCADA system, including sensor time series data, equipment status codes, and alarm logs, perform real-time cleaning, alignment, and feature extraction on the data through time series analysis algorithms, and generate dynamic data blocks and store them in the cache queue.
3. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 2, characterized in that The S2 includes the following steps: S21. Construct an association constraint system for the wind turbine operation and maintenance field, and define association constraint conditions including the core entity types and relationship types in the wind turbine operation and maintenance field; S22. Utilize the natural language processing ability of the locally deployed wind power operation and maintenance large model and the pre-constructed private domain knowledge Schema to extract knowledge from the chunked text according to the prompt words of the Schema, including named entity extraction and relationship extraction, and construct the index relationship between the sub-graph and the text block; S23. Use BGE-M3 and Sentence-BERT to perform vectorized embedding representation on the extracted entities, relationships, and triples; S24. Establish a dynamic association mechanism between SCADA real-time data and the knowledge graph, bind the equipment parameters and historical alarm records in the SCADA system to the fault types and maintenance plan entities in the knowledge graph, and support dynamic knowledge update based on real-time data; S25. Save the knowledge extraction results and the association binding information between SCADA and the knowledge base to the Neo4j graph database.
4. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 3, characterized in that The S23 includes the following steps: S231. Vectorize text data: For unstructured text data, use the Sentence-BERT model for semantic embedding to generate vector representations of text paragraphs. S232. Vectorize structured data: For structured data, use the BGE-M3 model for feature extraction and vector representation. S233. Multimodal data fusion: Align and fuse the text data and structured data representations to generate a unified multimodal vector representation for subsequent retrieval and reasoning.
5. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 4, wherein The above S3 includes the following steps: S31. Evaluate the complexity of the question, calculate the doubt degree and dependency parsing depth of the user query through a large model, and define multi-level complexity classification. S32. Split questions of different complexity levels to generate the following strategies: Explicit split: For queries containing coordinating conjunctions, directly split them into independent sub-questions. Implicit reasoning split: Use chain-of-thought prompts to guide the large model to generate intermediate questions. S33. Integrate the split sub-questions into a sub-question list in the order of precedence.
6. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 5, wherein The above S31 includes the following steps: Degree of doubt calculation. Given a user query Q, word segmentation is performed on a word-by-word basis to obtain a word sequence {w1, w2, …, w n}. The large model predicts the conditional probability P(w i | w1, w2, …, w i-1 ) for each word based on the context, and the overall degree of doubt is calculated through the cross-entropy loss function: Among them, N is the number of words in the query. The higher the PP(Q), the higher the semantic complexity of the query or the uncertainty of the large model. Dependency parsing depth calculation: First, use the en_core_web_trf model of the dependency syntax analysis tool spaCy to parse the user query Q to generate a dependency syntax tree. The dependency depth Depth(Q) is defined as the path length from the root node to the deepest leaf node. Complexity grading rules: Combine the doubt degree and dependency parsing depth, and set thresholds through experiments to divide into three levels of complexity: Simple query: PP(Q) < 50 and Depth(Q) ≤ 3. Compound query: 50 ≤ PP(Q) < 100 and 3 < Depth(Q) ≤ 6. Complex reasoning query: PP(Q) > 100 and Depth(Q) > 6.
7. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 6, characterized in that The above S4 includes the following steps: S41. Query and perform retrieval reasoning according to the sub-question list, use knowledge extraction to perform entity recognition on the sub-questions, and extract SPO triples containing logical relationships in the sub-queries. S42. Check entity integrity, summarize entity pairs, and perform retrieval queries in the constructed knowledge graph knowledge base according to the extracted triples. S43. Generate a reply, organize the sub-graphs retrieved by all sub-queries in S42 into a summary of text answers, and determine whether the question is solved. If it is solved, hand it over to the large model to generate the final answer according to the summary results of the sub-questions. If the question is not solved, the large model generates a supplementary question and repeats the above steps. S44. During the retrieval process, synchronously call the SCADA real-time data cache, dynamically adjust the weights of the retrieval results in combination with the current device status, and preferentially match the knowledge sub-graphs most relevant to the real-time working conditions to ensure the timeliness and pertinence of the diagnostic suggestions.
8. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 7, characterized in that, The realization of the standardized output of the diagnostic results requires structured generation, and the following hierarchical response generation templates are set: Emergency failure: Marked in red, prompt to stop the machine immediately for inspection. Early warning: Marked in yellow, it is recommended to arrange a patrol within 72 hours. Knowledge base missing: marked in blue, triggering the process of involving human experts.
9. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 8, wherein The feedback closed-loop mechanism is as follows: The knowledge base is updated through the operations of adoption / correction / rejection of the diagnosis results by the operation and maintenance personnel, and the accuracy rate, recall rate, and response time KPI indicators are calculated monthly.
10. The wind power fault diagnosis and operation and maintenance method based on multi-source data and knowledge retrieval enhancement according to claim 9, wherein The real-time verification mechanism is as follows: The generated maintenance plan is compared with the SCADA real-time data. If the device status is not restored, secondary diagnosis is triggered, and the fault solution cases in the knowledge base are updated.
Citation Information
Cited By
Intelligent operation and maintenance question-answering system for cable manufacturing equipment
CN120561253A
Intelligent operation and maintenance question and answer system for cable manufacturing equipment
CN120561253B
Expressway toll auxiliary question and answer method and system based on knowledge graph
CN120744141A
A highway tolling auxiliary question and answer method and system based on a knowledge graph
CN120744141B
Industrial fault diagnosis method and system based on natural language fault features and large model knowledge enhanced reasoning
CN120821995A