Fault diagnosis method for offshore wind turbine generator and related equipment
By constructing a fault knowledge graph in the data center and encrypting and deploying it on the operation and maintenance terminal, and using a large language model to parse and filter extended queries, structured triple data is generated. This solves the problem of low efficiency in fault diagnosis of offshore wind turbines in traditional methods and realizes efficient fault diagnosis on low computing power equipment.
Patent Information
- Application Number
- CN202511546114.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional methods for diagnosing faults in offshore wind turbines are inefficient and struggle to quickly and accurately acquire fault knowledge in complex scenarios. Furthermore, existing large language models and knowledge graphs are difficult to deploy on low-computing-power devices, affecting the real-time performance of fault diagnosis.
A fault diagnosis method for dual systems is proposed. A fault knowledge graph is built in the data center and deployed to the operation and maintenance terminal through encryption. The initial fault query is parsed using a large language model to generate extended queries. The optimal candidate document set is selected by combining the unit maintenance logs to generate structured triple data. Accurate structured knowledge is extracted through the fault knowledge graph to improve the efficiency of fault diagnosis.
It improves the accuracy and efficiency of fault diagnosis, adapts to low computing power environments, and meets the real-time fault diagnosis needs of offshore wind turbines in complex environments.
Smart Images

Figure CN121502008A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fault diagnosis technology, and in particular to a fault diagnosis method and related equipment for offshore wind turbines. Background Technology
[0002] Equipment fault diagnosis is of great value in the field of control. In the field operation and maintenance scenarios of large and complex equipment such as wind power, it is crucial to quickly and accurately obtain fault knowledge and auxiliary decision-making suggestions. At present, traditional operation and maintenance knowledge acquisition mainly relies on manually reviewing maintenance logs, consulting senior engineers, or using simple keyword search systems, which results in low fault diagnosis efficiency. Summary of the Invention
[0003] The main objective of this application is to propose a fault diagnosis method and related equipment for offshore wind turbines, which can improve fault diagnosis efficiency.
[0004] To achieve the above objectives, one aspect of this application proposes a fault diagnosis method for offshore wind turbines, the method comprising: Receive initial fault query, parse the initial fault query according to the preset language big model, and determine the extended query; The optimal candidate document set is determined by filtering based on the extended query and unit maintenance logs. The optimal candidate document set and the extended query are then processed according to the preset language big data model and the preset fault knowledge graph to generate structured triple data. The preset fault knowledge graph is encrypted and synchronized unidirectionally by the data center. The problem prompt words are determined by concatenating the structured triple data, the optimal candidate document set, and the initial fault query; the problem prompt words are then input into the preset language model for processing to obtain the diagnostic results.
[0005] In some embodiments, parsing the initial fault query according to a preset language model to determine the expanded query specifically includes: Based on the analysis of the preset language model, preset fault information is determined; wherein, the preset fault information includes general domain knowledge and fault diagnosis priors; The initial fault query is identified based on the preset fault information and the preset language model, and a hypothesis document is generated. The initial fault query is concatenated with the hypothetical document to obtain the expanded query.
[0006] In some embodiments, the step of filtering based on the extended query and unit maintenance logs to determine the optimal candidate document set specifically includes: The extended query is matched with the unit maintenance logs to determine a matching score; and the unit maintenance logs are sorted according to the matching score to determine a candidate set. Based on the candidate set and the extended query, vector encoding is performed to obtain several document vectors; Similarity is calculated based on several document vectors to determine cosine similarity, and the candidate set is sorted according to the cosine similarity to determine the optimal candidate document set.
[0007] In some embodiments, the step of processing the optimal candidate document set and the extended query based on the preset language big model and the preset fault knowledge graph to generate structured triple data specifically includes: Data is read from the preset fault knowledge graph to determine entity relationship constraints; The optimal candidate document set and the extended query are parsed to determine entity information; and the entity information and the preset fault knowledge graph are analyzed to determine query constraints. The entity relationship constraints and query constraints are processed according to the preset prompt template and the preset language model to determine the query code; the preset fault knowledge graph is queried according to the query code to determine the structured triple data.
[0008] In some embodiments, the preset fault knowledge graph is determined by the following method: The unit maintenance logs are processed based on a preset lightweight model, a multi-head self-attention network, and predefined relationship types to determine a candidate relationship set; The candidate relation set is convolved according to the preset lightweight model to obtain entity data; Based on the candidate relation set and the entity data, generate format triple data, perform synonym merging on the format triple data, and determine the target format triple data; The preset fault knowledge graph is generated based on the target format triple data.
[0009] In some embodiments, the step of processing the unit maintenance logs based on a preset lightweight model, a multi-head self-attention network, and predefined relationship types to determine a candidate relationship set specifically includes: The first associated data is determined by concatenating the predefined relationship type with the unit maintenance log; wherein, the first associated data represents the association between the textual semantics of the unit maintenance log and the predefined relationship type; The first associated data is processed according to the preset lightweight model to generate embedded data; and the embedded data is processed according to the multi-head self-attention network to generate several relation type probabilities. The candidate relationship set is determined by filtering the probabilities of several relationship types based on a first preset threshold.
[0010] In some embodiments, the step of performing convolution processing on the candidate relation set according to the preset lightweight model to obtain entity data specifically includes: Data is extracted from the candidate relationship set to determine entity type information, and the entity type information is then text-fused with the unit maintenance log to obtain the first fused data; The first fused data is processed according to the preset lightweight model to generate text context embedding data, and the text context embedding data is convolved to determine long-distance entity feature data. The long-distance entity feature data is fused using a preset improved self-attention network to obtain second fused data, and the second fused data is then corrected for label bias to obtain the entity data.
[0011] To achieve the above objectives, another aspect of this application proposes a fault diagnosis system for offshore wind turbines, the system comprising: The parsing module is used to receive initial fault queries, parse the initial fault queries according to a preset language model, and determine extended queries; The generation module is used to filter the extended query and unit maintenance logs to determine the optimal candidate document set, and process the optimal candidate document set and the extended query according to the preset language big model and the preset fault knowledge graph to generate structured triple data; wherein, the preset fault knowledge graph is encrypted and unidirectionally synchronized by the data center; The diagnostic module is used to concatenate the structured triplet data, the optimal candidate document set, and the initial fault query to determine the problem prompt words; and input the problem prompt words into the preset language big data model for processing to obtain the diagnostic results.
[0012] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0013] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0014] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.
[0015] The embodiments of this application include at least the following beneficial effects: This application provides a fault diagnosis method, system, electronic device, storage medium, and program product for offshore wind turbines. This solution parses the received initial fault query using a preset language model to determine extended queries; it filters the turbine maintenance logs and extended queries to determine the optimal candidate document set; it processes the determined optimal candidate document set and extended queries according to the preset language model and preset fault knowledge graph to obtain structured triplet data; it concatenates the obtained structured triplet data, the optimal candidate document set, and the received initial fault query into a problem prompt word, inputs the problem prompt word into the preset language model for processing, and outputs the corresponding fault diagnosis result. By generating extended queries, the accuracy of the semantic expression of the fault is improved; and by extracting precise structured knowledge from the fault knowledge graph based on the extended queries, the efficiency of fault diagnosis is improved. Attached Figure Description
[0016] Figure 1 This is a flowchart of a fault diagnosis method for offshore wind turbines provided in an embodiment of this application; Figure 2 yes Figure 1 The flowchart of step S101 in the text; Figure 3 yes Figure 1 The flowchart of step S102 in the document; Figure 4 yes Figure 1 The flowchart of step S102 in the document; Figure 5 This is a flowchart illustrating the construction of a fault knowledge graph in a fault diagnosis method for offshore wind turbines provided in this application embodiment; Figure 6 yes Figure 5 The flowchart of step S501 in the text; Figure 7 yes Figure 5 The flowchart of step S502 in the document; Figure 8 This is a flowchart of an intelligent fault diagnosis method provided in a specific embodiment of this application; Figure 9 This is a flowchart illustrating the construction of a dual-system architecture and lightweight graph modeling in a specific embodiment provided in this application. Figure 10This is a flowchart of graph retrieval enhanced question answering in a specific embodiment provided in this application; Figure 11 This is a flowchart illustrating performance verification in a specific embodiment provided in this application. Figure 12 This is a schematic diagram of the structure of a fault diagnosis system for offshore wind turbines provided in an embodiment of this application; Figure 13 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0018] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0019] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] In related technologies, traditional methods rely on manually querying paper maintenance logs, consulting senior engineers, or using simple keyword search systems to obtain maintenance knowledge for on-site maintenance of industrial equipment. However, these traditional methods struggle to handle complex scenarios, resulting in low fault diagnosis efficiency and significantly reduced decision reliability. Existing technical solutions apply large language models and knowledge graphs to industrial equipment maintenance to improve fault diagnosis efficiency. However, the large number of parameters in existing large language models and knowledge graphs makes them difficult to deploy on low-computing-power devices such as maintenance terminals, thus affecting the real-time performance of fault diagnosis.
[0022] In view of this, this application provides a fault diagnosis method and related equipment for offshore wind turbines. This solution constructs a dual-system fault diagnosis method, builds a fault knowledge graph in a data center, and deploys it to the operation and maintenance terminal through encryption. The initial fault query input by the user is expanded by combining it with the unit maintenance log through a language big model to obtain an expanded query, thereby improving the accuracy of the semantic expression of the fault query. Then, based on the obtained expanded query, knowledge is queried from the fault knowledge graph to extract accurate structured knowledge. The language big model processes the extracted structured knowledge and the initial fault query to improve the efficiency of fault diagnosis.
[0023] This application provides a fault diagnosis method for offshore wind turbines, relating to the field of information technology. This fault diagnosis method for offshore wind turbines can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle-mounted terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing a fault diagnosis method for offshore wind turbines, but is not limited to the above forms.
[0024] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0025] Figure 1 This is an optional flowchart of a fault diagnosis method for offshore wind turbines provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S103.
[0026] Step S101: Receive the initial fault query, parse the initial fault query according to the preset language big model, and determine the extended query; Step S102: Based on the extended query and unit maintenance log, the optimal candidate document set is determined. The optimal candidate document set and extended query are then processed according to the preset language big model and preset fault knowledge graph to generate structured triple data. The preset fault knowledge graph is encrypted and synchronized unidirectionally by the data center. Step S103: Based on the structured triplet data, the optimal candidate document set, and the initial fault query, the problem prompt words are concatenated to determine the problem prompt words; the problem prompt words are input into the preset language large model for processing to obtain the diagnostic results.
[0027] Steps S101 to S103 of this embodiment improve the data security and reliability of fault diagnosis data by deploying a fault knowledge graph with one-way encrypted push from the data center locally on the operation and maintenance terminal. A large language model is used to semantically parse the fault query input by the user, obtaining corresponding extended queries. Based on these extended queries and existing unit operation and maintenance log data, an optimal candidate document set is obtained to improve the accuracy of subsequent knowledge retrieval from the fault knowledge graph, thereby improving the accuracy and efficiency of fault diagnosis. The large language model extracts relevant fault diagnosis knowledge from the fault knowledge graph based on the optimal candidate document set and extended queries. This extracted fault diagnosis knowledge, the optimal candidate document set, and the user-input fault query are then combined to form a problem prompt. The large language model then generates structured fault diagnosis results based on the problem prompt, providing technical support for wind turbine operation and maintenance.
[0028] Please see Figure 2 In some embodiments, step S101 may include, but is not limited to, steps S201 to S203: Step S201: Analyze the preset language model to determine preset fault information; wherein, the preset fault information includes general domain knowledge and fault diagnosis priors; Step S202: Identify the initial fault query based on the preset fault information and preset language model, and generate a hypothesis document; Step S203: Combine the initial fault query with the hypothesis document to obtain the expanded query.
[0029] In step S201 of some embodiments, the initial fault query Q input by the user is input into a locally deployed language big model, such as the Deepseek model; the language big model analyzes the database and extracts general domain knowledge and fault diagnosis priors as preset fault information; the language big model parses the initial fault query Q input by the user based on the extracted preset fault information, and refines the user's query intent to improve the accuracy and efficiency of subsequent fault diagnosis.
[0030] In step S202 of some embodiments, the language big model analyzes the input initial fault query based on general domain knowledge and fault diagnosis priors, extracts the core intent in the initial fault query, generates a hypothetical answer document D_h, and then accurately locates the corresponding fault diagnosis knowledge from the fault knowledge graph based on the hypothetical answer document.
[0031] In step S203 of some embodiments, the generated hypothetical answer document D_h is concatenated with the initial fault query Q input by the user to obtain an extended query Q_ext=[Q;D_h]. This extended query includes the original fault query input by the user and the extracted semantic expression, which is used for locating knowledge data in the subsequent fault knowledge graph. In this embodiment, the obtained extended query is processed by a lightweight model to generate a multi-dimensional dense vector embedding to enhance the semantic matching ability of subsequent retrieval.
[0032] Please see Figure 3 In some embodiments, step S102 may include, but is not limited to, steps S301 to S303: Step S301: Perform matching calculations based on the extended query and the unit maintenance logs to determine the matching score; and sort the unit maintenance logs according to the matching score to determine the candidate set; Step S302: Perform vector encoding based on the candidate set and extended query to obtain several document vectors; Step S303: Calculate the similarity based on several document vectors, determine the cosine similarity, and sort the candidate set according to the cosine similarity to determine the optimal candidate document set.
[0033] In step S301 of some embodiments, due to the low computing power environment of the maintenance terminal, the efficiency of matching the fault knowledge graph based on the obtained extended query in the traditional method is low. Therefore, the efficiency and positioning accuracy of the matching query are improved by a series strategy of "coarse word frequency screening and fine semantic screening". Based on word frequency statistics and inverse document frequency, the matching score between the extended query and the unit maintenance log is calculated according to the following formula. , in, Inverse document frequency, Total number of documents For terms containing words The number of documents, For terms In the document The frequency of terms in the logs. After calculating the matching score between the extended query and the unit maintenance logs, the documents in the unit maintenance logs are sorted according to the matching score, and the top 100 logs with the highest scores are selected as the candidate set.
[0034] In step S302 of some embodiments, after obtaining the candidate set, since the number of documents in the candidate set is large, it has a certain impact on the efficiency of subsequent fault knowledge graph query matching. Therefore, it is necessary to further filter the candidate set and extended queries to balance the efficiency and accuracy of query matching. The obtained candidate set and extended queries are vector-encoded for further filtering. In this embodiment, the candidate set and extended queries are BERT vector-encoded to obtain a corresponding number of document vectors and query vectors.
[0035] In step S303 of some embodiments, after vector encoding of the candidate set and the extended query, the optimal candidate documents are selected as the optimal candidate document set by calculating the similarity between the vectors. In this embodiment, the FAISS approximate nearest neighbor algorithm is used to calculate the cosine similarity between the vector-encoded candidate set and the extended query. The algorithm formula is as follows: , in, For cosine similarity, To expand the query vector, The candidate set document vectors.
[0036] Please see Figure 4 In some embodiments, step S102 may also include, but is not limited to, steps S401 to S403: Step S401: Read data from the preset fault knowledge graph and determine entity relationship constraints; Step S402: Parse the optimal candidate document set and extended query to determine entity information; and analyze the entity information with the preset fault knowledge graph to determine query constraints. Step S403: Process entity relationship constraints and query constraints according to the preset prompt template and preset language big model to determine the query code; query the preset fault knowledge graph according to the query code to determine the structured triple data.
[0037] In step S401 of some embodiments, after determining the optimal candidate document, the optimal candidate document is parsed by the language big model to generate a corresponding query statement for the fault knowledge graph. Based on this query statement, precise structured knowledge is extracted from the fault knowledge graph to generate subsequent fault diagnosis results. In this embodiment, the fault knowledge graph is stored in the Neo4j database. The language big model connects to and accesses the Neo4j database, reads the fault diagnosis event ontology from the fault knowledge graph, including 12 types of entities and 8 types of relationships, such as fault types and fault phenomena, and the relationships between fault types and fault phenomena. The language big model generates an ontology concept hierarchy description based on the read fault diagnosis event ontology and stores it, obtaining the grammatical constraints between entities and relationships so that the language big model can subsequently generate structured diagnostic results.
[0038] In step S402 of some embodiments, the language big model parses the optimal candidate document and extended query, extracts the core entity information of the fault event, such as "generator bearing high temperature" corresponding to the "fault mode" entity; the language big model analyzes the extracted core entity information and fault events in the fault knowledge graph to determine the type and association relationship of the optimal candidate document and extended query in the graph, which serves as the constraint condition for the query.
[0039] In step S403 of some embodiments, the language big model processes the preset prompt template and the obtained constraints to generate query code that conforms to the syntax of the fault knowledge graph. For example, the prompt template is "MATCH(:fault mode{name:xxx})-[:relation type]->(o:entity type) RETURN o". The language big model queries the deployed fault knowledge graph according to the generated query code to obtain structured triple data.
[0040] Please see Figure 5 In some embodiments, the preset fault knowledge graph in the fault diagnosis method for offshore wind turbines provided in this application can be determined through steps S501 to S504: Step S501: Process the unit maintenance logs according to the preset lightweight model, multi-head self-attention network and predefined relationship types to determine the candidate relationship set; Step S502: Perform convolution processing on the candidate relation set according to the preset lightweight model to obtain entity data; Step S503: Generate format triplet data based on the candidate relation set and entity data, perform synonym merging on the format triplet data, and determine the target format triplet data; Step S504: Generate a preset fault knowledge graph based on the target format triplet data.
[0041] In step S501 of some embodiments, under the low computing power environment of the operation and maintenance terminal, the computing power requirement of traditional knowledge graph modeling is too high, so it is necessary to construct a lightweight fault knowledge graph. In this embodiment, the lightweight model deployed in the operation and maintenance terminal is used to identify the unit operation and maintenance logs, determine the predefined relationship that conforms to fault diagnosis, and provide semantic constraints for subsequent entity extraction.
[0042] In step S502 of some embodiments, after identifying the predefined relationship, the entity type corresponding to the predefined relationship in the unit operation and maintenance log is obtained by inference based on the predefined relationship. Then, the corresponding entity information is extracted by convolution processing, and the corresponding label bias is corrected to improve the accuracy of the structured data in the subsequent fault knowledge graph.
[0043] In step S503 of some embodiments, the identified relationships and extracted arrival entity information are processed to generate corresponding format triplet data; then the generated format triplet data is synonymized and merged to reduce data redundancy, thereby improving the efficiency of subsequent fault diagnosis.
[0044] In step S504 of some embodiments, a topological structure is constructed based on the subsequent formatted triplet data through synonym merging to form a fault diagnosis knowledge graph. In some embodiments, after identifying relationships and extracting entity data, the lightweight model is adjusted by jointly optimizing the relationship classification loss and sequence labeling loss to improve the extraction accuracy of triplet data in the lightweight model and ensure adaptability to low-computing-power environments. In this embodiment, the cross-entropy loss is used to measure the identification error between candidate relationships as the relationship classification loss, and the relationship classification loss is calculated using the following formula: , in, Input text length, To predict relationships, This is a real relationship.
[0045] Meanwhile, cross-entropy loss is used to measure the entity label prediction error as the sequence labeling loss, and the calculation formula is as follows: , in, To predict entity labels, For the real labels. Then through hyperparameters. Balancing relation identification and entity extraction yields the total loss: , Please see Figure 6 In some embodiments, step S501 may include, but is not limited to, steps S601 to S603: Step S601: Concatenate the predefined relationship type with the unit maintenance log to determine the first associated data; wherein, the first associated data represents the association between the textual semantics of the unit maintenance log and the predefined relationship type; Step S602: Process the first associated data according to the preset lightweight model to generate embedded data; and process the embedded data according to the multi-head self-attention network to generate several relation type probabilities. Step S603: Filter the probabilities of several relation types according to the first preset threshold to determine the candidate relation set.
[0046] In step S601 of some embodiments, the lightweight model concatenates the document data in the unit maintenance log with the predefined relationship types in BERT format, so that the language big model and the lightweight model can identify the text semantics and the associations in the relationship types in the concatenated data, so that the lightweight model can identify the corresponding associations for information extraction and construct a fault knowledge graph.
[0047] In step S602 of some embodiments, a lightweight model processes the concatenated data to generate initial text and embedded data incorporating relationships. The generated embedded data is then input into a constructed multi-head self-attention network for analysis to deeply mine the semantic relationships between text domain relationships. Finally, a classifier outputs the probability of the concatenated data for each relationship type. In this embodiment, the obtained embedded data is input into a three-layer multi-head attention network for processing, and a Softmax classifier outputs probability data. The expressions for the three-layer multi-head attention network and the Softmax classifier are as follows: , in, For the first The input of attention head, These are the learnable parameters. Calculate the scaling factor for attention.
[0048] In step S603 of some embodiments, the obtained probabilities are sorted and filtered according to a preset threshold to obtain a set of valid candidate relationships. ,in, , This is the predefined set of relations.
[0049] Please see Figure 7 In some embodiments, step S502 may include, but is not limited to, steps S701 to S703: Step S701: Extract data from the candidate relationship set, determine entity type information, and perform text fusion of entity type information with unit maintenance logs to obtain the first fused data; Step S702: Process the first fused data according to the preset lightweight model to generate text context embedding data, and perform convolution processing on the text context embedding data to determine long-distance entity feature data. Step S703: The long-distance entity feature data is fused according to the preset improved self-attention network to obtain the second fused data, and the label bias is corrected on the second fused data to obtain entity data.
[0050] In step S701 of some embodiments, after obtaining the candidate relation set, the candidate relation set is analyzed according to the abstract event ontology to determine the entity type information therein. For example, the abstract event ontology is "fault mode-has_Symptom-fault phenomenon". The lightweight model fuses the determined entity type information with the unit operation and maintenance log to obtain fused data, and adds location embedding according to the fused data to enhance the model's location awareness, so as to improve the accuracy and efficiency of subsequent fault diagnosis.
[0051] In step S702 of some embodiments, the lightweight model processes the obtained fused data to generate text context embedding data, and then processes the generated text context embedding data through a convolutional network to extract long-range entity features from the embedding data. In this embodiment, BERT is used to process the fused data to generate text context embedding data, and then the generated text context embedding data is input into a four-layer dilated gated convolutional network for convolution processing to capture long-range entity features from the text context embedding data. The expression of the four-layer dilated gated convolutional network is as follows: , in, and These are single-dimensional dilated convolutions with non-shared weights. It is the Hadamard product.
[0052] In step S703 of some embodiments, the captured long-distance entity features are input into an improved self-attention network for processing. The entity structure information in the long-distance entity features is fused, and the label bias of the fused data is corrected to obtain entity data. The entity data includes corresponding entity labels. The lightweight model and the language large model can determine the start and end positions of the entity in the unit operation and maintenance log based on the entity labels. In this embodiment, after fusing the entity structure information in the long-distance entity features through the improved self-attention network to obtain fused data, a CRF layer is introduced to correct the label sequence logical bias of the fused data and output the final entity data and its labels.
[0053] The following is a detailed description and explanation of the solutions in the embodiments of the present invention, using specific application examples: In a specific embodiment, according to Figure 8 The above describes a fault diagnosis platform for offshore wind turbines, constructed using a dual-system collaborative architecture platform consisting of a high-performance data center and a low-performance maintenance terminal. The data center handles large-scale model training and knowledge storage, as well as generating remote maintenance decisions. The maintenance terminal focuses on lightweight on-site inference and independent operation, providing fault knowledge retrieval and real-time decision support, and supports offline single-machine operation in network-free environments. This achieves functional division of labor and secure coupling, meeting the fault diagnosis needs of offshore wind turbines in complex environments. Figure 9 As shown, secure collaboration between the two systems is achieved through an encryption-synchronization-one-way transmission mechanism. The data center transmits the complete fault knowledge graph to the operation and maintenance terminal via encryption one-way transmission. The operation and maintenance terminal then deploys a lightweight terminal graph based on the received fault knowledge graph. The expression for this encryption one-way transmission is as follows: , in, A complete fault knowledge graph for data centers. This is an AES encryption operation; Sync is a periodic synchronization mechanism. For lightweight terminal map, The terminal receives knowledge but does not send back the original data, thus improving data security.
[0054] The data center leverages high-performance hardware and industrial-grade software to build a powerful computing platform, meeting the needs of parallel processing of tens of millions of sensor monitoring data points and training large-parameter models. At the software level, a multi-layered technical architecture is constructed, utilizing the Neo4j graph database to store fault diagnosis knowledge graphs and locally deploying the Deepseek model as a large-scale language model to query the fault diagnosis knowledge graph and generate fault diagnosis results. Lightweight dependency libraries, including a lightweight BERT model, are pre-installed on the maintenance terminals, and the fault knowledge graphs pushed from the data center are stored as Neo4j offline files for offline operation. Furthermore, the Deepseek model generates fault diagnosis results containing clear "fault phenomenon - repair steps - cause analysis" inference chains, facilitating understanding and verification by maintenance personnel.
[0055] The maintenance terminal uses the BERT model to concatenate maintenance log text with predefined relation types. The BERT model then processes the concatenated text data to generate 768-dimensional embedding data. This embedding data is input into a multi-head self-attention module for deep learning, yielding a candidate relation set. Predefined relations conforming to the abstract event ontology are selected from the maintenance logs, providing semantic constraints for subsequent entity extraction. The abstract event ontology, combined with the candidate relation set, extracts the corresponding entity types. These entity types are then fused with the maintenance log text, and the BERT model generates contextual embedding data. Dilated gated convolution and an improved self-attention module are then used to process the contextual embedding data. Feature extraction yields entity features; a CRF layer is then introduced to correct label bias in these features, outputting entity annotations conforming to the BIO specification. The extracted entities and identified relationships are integrated to generate triplet data. Synonym merging is performed on the generated triplet data, and a topological structure is constructed to obtain a searchable fault diagnosis knowledge graph, which is then imported into the Neo4j database for storage. Simultaneously, during the construction of the fault diagnosis knowledge graph, the Adam optimizer is used to iterate the lightweight model, setting the number of iterations to 300 and the batch size per round to 8. Model parameters are updated through gradient backpropagation, controlling the model parameter size to not exceed 17M to adapt to the low-computing-power environment of maintenance terminals. Please refer to [link / reference]. Figure 10 The user inputs a fault problem into the maintenance terminal. The terminal uses this fault problem as the initial query. The deployed Deepseek model, based on general domain knowledge and fault diagnosis priors, processes the initial query using the HyDE method to generate a hypothesis document and expand the semantics of the user's problem to obtain an expanded query. Then, a concatenated strategy of BM25 coarse screening and FAISS fine screening is used to filter maintenance logs to obtain the optimal candidate log documents. The Deepseek model parses the optimal candidate log documents and the expanded query to generate a Cypher query statement. Based on this query statement, it accesses the fault diagnosis knowledge graph in the Neo4j database to obtain structured triples. The optimal candidate log documents, triple data, and the initial query are integrated to obtain prompt words. The Deepseek model outputs a fault diagnosis result of "fault phenomenon - repair steps - cause analysis" based on the prompt words.
[0056] Please see Figure 11To verify the performance of the fault diagnosis method provided in this application embodiment, historical wind farm operation and maintenance log data was collected. After cleaning the log data, 1227 complete fault work orders were obtained as experimental data. The experimental data were stratified and sampled according to a 7:1:1 ratio to obtain a training set of 977 data points, a validation set of 100 data points, and a test set of 150 data points. The performance of the lightweight fault knowledge graph modeling core module RO-JTEM was evaluated using the divided training set, validation set, and test set to verify its adaptability to low computing power environments. The F1 score of entity extraction (Precision / P / R), relation recognition (Precision / P / R), and triple extraction (Precision / P / R) was used as the core indicator, and the calculation formula is as follows: , Meanwhile, mainstream entity relation joint extraction models such as TSM-JERHDE, TDEER, Tabel-Seq, and Know-Enhanced were selected as baselines. Experimental results show that RO-JTEM achieved an F1 score of 85.2% for triple extraction, only 2.9% lower than the best-performing TSM-JERHDE (88.1%), but significantly better than TDEER (84.2%) and Tabel-Seq (82.7%), verifying its graph modeling accuracy. RO-JTEM has only 17M parameters (a 63% reduction from TSM-JERHDE's 46M), a single-round training time of 4.7min (a 67.8% reduction from TSM-JERHDE's 14.6min), an inference time of 0.35s / sample, and a GPU memory usage of 4.1GB, all significantly lower than the baseline models, verifying its deployment feasibility in low-computing-power environments. Then, Hit@k (the proportion of the first k answers containing correct results) was used as an indicator to test the performance of the large language model in the system and quantify its accuracy. The calculation formula is as follows: , in, To test the number of questions, For the standard answer set, For the first k answers of the model, The indicator function is used. By comparison, under low computing power environment, DeepSeek-7B achieves 66% Hit@1, which is significantly better than LLaMA2-7B (23%) and LLaMA3-8B (48%). The Hit@1 (66%) and Hit@10 (96%) of this method are higher than those of large models only (33% / 84%), traditional RAG (47% / 89%) and GraphRAG (53% / 89%), verifying the effectiveness of graph retrieval enhancement strategy. Finally, 100 real wind farm operation and maintenance issues were selected to manually evaluate the system, assessing its operability, interpretability, and response efficiency, and assigning corresponding scores. The results showed that the system's average score was 4.2 points, with the highest score (4.5 points) in "interpretability." This was significantly better than the traditional black-box model (3.1 points) because the reasoning process could be traced back to knowledge graph triples (such as the causal chain of "bearing high temperature → insufficient lubrication"). This verifies that the system meets the needs of on-site operation and maintenance. Therefore, the RO-JTEM question-answering system can run stably in a low-computing-power environment, with inference latency ≤0.35s / sample and memory usage ≤4.1GB, meeting the "low latency, low resource" requirements of operation and maintenance terminals.
[0057] The embodiments of this application include at least the following beneficial effects: This application provides a fault diagnosis method, system, electronic device, storage medium, and program product for offshore wind turbines. This solution parses the received initial fault query using a preset language model to determine extended queries; it filters the turbine maintenance logs and extended queries to determine the optimal candidate document set; it processes the determined optimal candidate document set and extended queries according to the preset language model and preset fault knowledge graph to obtain structured triplet data; it concatenates the obtained structured triplet data, the optimal candidate document set, and the received initial fault query into a problem prompt word, inputs the problem prompt word into the preset language model for processing, and outputs the corresponding fault diagnosis result. By generating extended queries, the accuracy of the semantic expression of the fault is improved; and by extracting precise structured knowledge from the fault knowledge graph based on the extended queries, the efficiency of fault diagnosis is improved.
[0058] Please see Figure 12 This application also provides a fault diagnosis system for offshore wind turbines, which can implement the above-mentioned method. The system includes: The parsing module is used to receive initial fault queries, parse the initial fault queries according to a preset language model, and determine extended queries; The generation module is used to filter the extended query and unit maintenance logs to determine the optimal candidate document set, and process the optimal candidate document set and the extended query according to the preset language big model and the preset fault knowledge graph to generate structured triple data; wherein, the preset fault knowledge graph is encrypted and unidirectionally synchronized by the data center; The diagnostic module is used to concatenate the structured triplet data, the optimal candidate document set, and the initial fault query to determine the problem prompt words; and input the problem prompt words into the preset language big data model for processing to obtain the diagnostic results.
[0059] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0060] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0061] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0062] Please see Figure 13 , Figure 13 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 1301 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1302 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1302 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1302 and is called and executed by the processor 1301 using the methods described in the embodiments of this application. The input / output interface 1303 is used to implement information input and output; The communication interface 1304 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1305 transmits information between various components of the device (e.g., processor 1301, memory 1302, input / output interface 1303, and communication interface 1304); The processor 1301, memory 1302, input / output interface 1303 and communication interface 1304 are connected to each other within the device via bus 1305.
[0063] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0064] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0065] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0066] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0067] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0068] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0069] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0070] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0071] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0072] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0073] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0074] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0075] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0076] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0077] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0078] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A fault diagnosis method for offshore wind turbines, characterized in that, The method includes: Receive initial fault query, parse the initial fault query according to the preset language big model, and determine the extended query; The optimal candidate document set is determined by filtering based on the extended query and unit maintenance logs. The optimal candidate document set and the extended query are then processed according to the preset language big data model and the preset fault knowledge graph to generate structured triple data. The preset fault knowledge graph is encrypted and synchronized unidirectionally by the data center. The problem prompt words are determined by concatenating the structured triple data, the optimal candidate document set, and the initial fault query; the problem prompt words are then input into the preset language model for processing to obtain the diagnostic results.
2. The method according to claim 1, characterized in that, The step of parsing the initial fault query according to a preset language model to determine the expanded query specifically includes: Based on the analysis of the preset language model, preset fault information is determined; wherein, the preset fault information includes general domain knowledge and fault diagnosis priors; The initial fault query is identified based on the preset fault information and the preset language model, and a hypothesis document is generated. The initial fault query is concatenated with the hypothetical document to obtain the expanded query.
3. The method according to claim 1, characterized in that, The step of filtering based on the extended query and unit maintenance logs to determine the optimal candidate document set specifically includes: The extended query is matched with the unit maintenance logs to determine a matching score; and the unit maintenance logs are sorted according to the matching score to determine a candidate set. Based on the candidate set and the extended query, vector encoding is performed to obtain several document vectors; Similarity is calculated based on several document vectors to determine cosine similarity, and the candidate set is sorted according to the cosine similarity to determine the optimal candidate document set.
4. The method according to claim 1, characterized in that, The step of processing the optimal candidate document set and the extended query based on the preset language big model and the preset fault knowledge graph to generate structured triple data specifically includes: Data is read from the preset fault knowledge graph to determine entity relationship constraints; The optimal candidate document set and the extended query are parsed to determine entity information; and the entity information and the preset fault knowledge graph are analyzed to determine query constraints. The entity relationship constraints and query constraints are processed according to the preset prompt template and the preset language model to determine the query code; the preset fault knowledge graph is queried according to the query code to determine the structured triple data.
5. The method according to claim 1, characterized in that, The preset fault knowledge graph is determined by the following method: The unit maintenance logs are processed based on a preset lightweight model, a multi-head self-attention network, and predefined relationship types to determine a candidate relationship set; The candidate relation set is convolved according to the preset lightweight model to obtain entity data; Based on the candidate relation set and the entity data, generate format triple data, perform synonym merging on the format triple data, and determine the target format triple data; The preset fault knowledge graph is generated based on the target format triple data.
6. The method according to claim 5, characterized in that, The step of processing the unit maintenance logs based on a preset lightweight model, a multi-head self-attention network, and predefined relationship types to determine a candidate relationship set specifically includes: The first associated data is determined by concatenating the predefined relationship type with the unit maintenance log; wherein, the first associated data represents the association between the textual semantics of the unit maintenance log and the predefined relationship type; The first associated data is processed according to the preset lightweight model to generate embedded data; and the embedded data is processed according to the multi-head self-attention network to generate several relation type probabilities. The candidate relationship set is determined by filtering the probabilities of several relationship types based on a first preset threshold.
7. The method according to claim 5, characterized in that, The step of performing convolution processing on the candidate relation set according to the preset lightweight model to obtain entity data specifically includes: Data is extracted from the candidate relationship set to determine entity type information, and the entity type information is then text-fused with the unit maintenance log to obtain the first fused data; The first fused data is processed according to the preset lightweight model to generate text context embedding data, and the text context embedding data is convolved to determine long-distance entity feature data. The long-distance entity feature data is fused using a preset improved self-attention network to obtain second fused data, and the second fused data is then corrected for label bias to obtain the entity data.
8. A fault diagnosis system for offshore wind turbines, characterized in that, The system includes: The parsing module is used to receive initial fault queries, parse the initial fault queries according to a preset language model, and determine extended queries; The generation module is used to filter the extended query and unit maintenance logs to determine the optimal candidate document set, and process the optimal candidate document set and the extended query according to the preset language big model and the preset fault knowledge graph to generate structured triple data; wherein, the preset fault knowledge graph is encrypted and unidirectionally synchronized by the data center; The diagnostic module is used to concatenate the structured triplet data, the optimal candidate document set, and the initial fault query to determine the problem prompt words; and input the problem prompt words into the preset language big data model for processing to obtain the diagnostic results.
9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.