Wind power operation and maintenance knowledge management method, device, equipment and medium

CN121144534BActive Publication Date: 2026-09-11SANY ELECTRIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511374320.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-09-11
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

[0005]本申请实施例提供风电运维知识管理方法、装置、设备及介质,用以解决现有技术中存在无法有效指导实际运维工作,使用场景受限的问题

Benefits of technology

[0072] This application provides a method, apparatus, equipment, and medium for wind power operation and maintenance (O&M) knowledge management. This method significantly improves the accuracy and professionalism of the wind power O&M knowledge question-and-answer system by integrating large language model fine-tuning and retrieval enhancement algorithms with unified management of structured and unstructured data in the wind power O&M field. By generating a high-quality training dataset from unstructured data according to a preset annotation template, a large language model with knowledge of the wind power O&M field is trained, enabling it to summarize and answer common and fundamental questions. Simultaneously, for structured data, semantic vector mapping is used and stored in a vector database. Combined with retrieval enhancement algorithms, this allows for rapid retrieval and accurate access to real-time data such as equipment parameters, operating data, and operating procedures, ensuring the accuracy and timeliness of answers. This method effectively covers the diverse and complex question-and-answer needs of wind power O&M sites, not only improving the response speed and accuracy to O&M issues and reducing reliance on human experience and knowledge transfer costs, but also enabling continuous updating and optimization of the wind power O&M knowledge base, constructing an intelligent knowledge management platform with real-time performance, professionalism, and sustainability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121144534B_ABST
    Figure CN121144534B_ABST
Patent Text Reader

Abstract

The application provides a wind power operation and maintenance knowledge management method, device, equipment and medium. The method comprises the following steps: firstly, obtaining wind power operation and maintenance original data; then, labeling unstructured data according to a preset labeling template to generate a training data set; further, training a large language model to be trained based on the training data set to obtain a trained large language model; then, processing structured data based on a retrieval enhancement algorithm to obtain a mapping relationship between a semantic vector and semantic text, and storing the mapping relationship between the semantic vector and the semantic text in a vector database; finally, deploying a wind power operation and maintenance knowledge management platform with the trained large language model and the vector database to provide functional services to users. Through this method, the management of wind power operation and maintenance knowledge is improved, and the operation and maintenance efficiency and equipment reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing and information retrieval technology, and in particular to a method, device, equipment and medium for wind power operation and maintenance knowledge management. Background Technology

[0002] In wind power operation and maintenance (O&M) scenarios, knowledge management is crucial. Wind power equipment is diverse, has long operating cycles, and suffers from complex and widely distributed faults. A standardized and systematic knowledge management system can effectively accumulate and share O&M experience, standardize operating procedures, reduce human error and repeated trial and error, improve fault response and troubleshooting efficiency, and ensure the long-term stable operation of equipment. At the same time, knowledge management helps accelerate the training of new employees, reduces the risk of technical gaps caused by personnel turnover, and supports wind power companies in achieving their goals of safe, efficient, and sustainable O&M management.

[0003] In existing technologies, wind power operation and maintenance knowledge management methods generally construct entity-relation-entity triples through named entity recognition and relation extraction, forming a knowledge base similar to common questions and corresponding answers. When a user asks a question, the model vectorizes the user's input text, compares similarity, and outputs the answer corresponding to the most similar common question, thereby achieving the function of classifying and answering questions.

[0004] However, existing wind power operation and maintenance knowledge management methods have limitations in effectively guiding actual operation and maintenance work and are subject to application scenarios. Summary of the Invention

[0005] This application provides a wind power operation and maintenance knowledge management method, device, equipment, and medium to solve the problem that existing technologies cannot effectively guide actual operation and maintenance work and have limited application scenarios.

[0006] In a first aspect, embodiments of this application provide a wind power operation and maintenance knowledge management method, including:

[0007] Obtain raw wind power operation and maintenance data, which includes multiple structured data and multiple unstructured data.

[0008] The unstructured data is labeled according to a preset labeling template to generate a training dataset, which includes multiple question-answer pairs.

[0009] The large language model to be trained is trained based on the training dataset to obtain a trained large language model, which is a model with intelligent question answering capabilities in the field of wind power operation and maintenance.

[0010] The structured data is processed based on the retrieval enhancement algorithm to obtain the mapping relationship between semantic vectors and semantic text, and the mapping relationship between semantic vectors and semantic text is stored in the vector database.

[0011] By combining the trained large language model with the vector database, a wind power operation and maintenance knowledge management platform is constructed to provide functional services to users.

[0012] In one possible implementation, the step of labeling the unstructured data according to a preset labeling template to generate a training dataset includes:

[0013] The unstructured data is preprocessed to obtain preprocessed unstructured data. The preprocessing includes at least format conversion and removal of erroneous and redundant data.

[0014] Each piece of preprocessed unstructured data is labeled according to the preset labeling template to generate question-answer pairs;

[0015] The training dataset is obtained from multiple question-answer pairs.

[0016] In one possible implementation, training the large language model to be trained based on the training dataset to obtain a trained large language model includes:

[0017] A device-level LoRA adapter is inserted into the middle and lower-level conversion modules of the large language model to be trained, and a site-level LoRA adapter and a fusion adapter are inserted into the middle and higher-level conversion modules of the large language model to be trained.

[0018] The device-level LoRA adapter is trained based on the device-level wind power operation and maintenance data in the training dataset to adjust the parameters of the device-level LoRA adapter, and the site-level LoRA adapter is trained based on the site-level wind power operation and maintenance data in the training dataset to adjust the parameters of the site-level LoRA adapter.

[0019] Based on the correlation between the equipment-level wind power operation and maintenance data and the station-level wind power operation and maintenance data, a fusion test set is constructed.

[0020] The fusion adapter is trained based on the fusion test set to adjust the parameters of the fusion adapter;

[0021] The trained device-level LoRA adapter, the station-level LoRA adapter, and the fusion adapter are integrated into the large language model to be trained to obtain the trained large language model.

[0022] In one possible implementation, the step of processing the structured data based on the retrieval enhancement algorithm to obtain the mapping relationship between semantic vectors and semantic text, and storing the mapping relationship between semantic vectors and semantic text in a vector database, includes:

[0023] Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts, each of which contains complete contextual information.

[0024] For each semantic text, a semantic vector corresponding to the semantic text is generated using the trained large language model;

[0025] Based on the similarity between the semantic texts, a semantic knowledge graph is constructed;

[0026] The semantic text, its corresponding semantic vector, and the semantic knowledge graph are stored in the vector database.

[0027] In one possible implementation, the method further includes:

[0028] Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts. Then, the image data and table data in the structured data are descriptively embedded using a multimodal model to obtain embedding vectors.

[0029] For each semantic text, generating a semantic vector corresponding to the semantic text through a language model includes:

[0030] Generate a fusion vector based on the semantic vector and the embedding vector;

[0031] The step of storing the semantic text and its corresponding semantic vector into the vector database includes:

[0032] The semantic text and the corresponding fusion vector are stored in the vector database.

[0033] In one possible implementation, the method further includes:

[0034] In the first preset update cycle, the vector database is updated based on an incremental update mechanism.

[0035] In one possible implementation, the method further includes:

[0036] In the second preset update cycle, the vector database is updated based on a full update mechanism.

[0037] Secondly, embodiments of this application provide a wind power operation and maintenance knowledge management device, comprising:

[0038] The acquisition module is used to acquire raw wind power operation and maintenance data, which includes multiple structured data and multiple unstructured data.

[0039] The generation module is used to annotate the unstructured data according to a preset annotation template to generate a training dataset, which includes multiple question-answer pairs;

[0040] The training module is used to train the large language model to be trained based on the training dataset to obtain a trained large language model, which is a model with intelligent question answering capabilities in the field of wind power operation and maintenance.

[0041] The processing module is used to process the structured data based on the retrieval enhancement algorithm to obtain the mapping relationship between semantic vectors and semantic text, and to store the mapping relationship between semantic vectors and semantic text in the vector database.

[0042] The application module is used to combine the trained large language model and the vector database to build a wind power operation and maintenance knowledge management platform to provide functional services to users.

[0043] In one possible implementation, the generation module is specifically used for:

[0044] The unstructured data is preprocessed to obtain preprocessed unstructured data. The preprocessing includes at least format conversion and removal of erroneous and redundant data.

[0045] Each piece of preprocessed unstructured data is labeled according to the preset labeling template to generate question-answer pairs;

[0046] The training dataset is obtained from multiple question-answer pairs.

[0047] In one possible implementation, the training module is specifically used for:

[0048] A device-level LoRA adapter is inserted into the middle and lower-level conversion modules of the large language model to be trained, and a site-level LoRA adapter and a fusion adapter are inserted into the middle and higher-level conversion modules of the large language model to be trained.

[0049] The device-level LoRA adapter is trained based on the device-level wind power operation and maintenance data in the training dataset to adjust the parameters of the device-level LoRA adapter, and the site-level LoRA adapter is trained based on the site-level wind power operation and maintenance data in the training dataset to adjust the parameters of the site-level LoRA adapter.

[0050] Based on the correlation between the equipment-level wind power operation and maintenance data and the station-level wind power operation and maintenance data, a fusion test set is constructed.

[0051] The fusion adapter is trained based on the fusion test set to adjust the parameters of the fusion adapter;

[0052] The trained device-level LoRA adapter, the station-level LoRA adapter, and the fusion adapter are integrated into the large language model to be trained to obtain the trained large language model.

[0053] In one possible implementation, the processing module is specifically used for:

[0054] Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts, each of which contains complete contextual information.

[0055] For each semantic text, a semantic vector corresponding to the semantic text is generated using a language model;

[0056] Based on the similarity between the semantic texts, a semantic knowledge graph is constructed;

[0057] The semantic text, its corresponding semantic vector, and the semantic knowledge graph are stored in the vector database.

[0058] In one possible implementation, the processing module is further configured to:

[0059] Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts. Then, the image data and table data in the structured data are descriptively embedded using a multimodal model to obtain embedding vectors.

[0060] For each semantic text, generating a semantic vector corresponding to the semantic text through a language model includes:

[0061] Generate a fusion vector based on the semantic vector and the embedding vector;

[0062] The step of storing the semantic text and its corresponding semantic vector into the vector database includes:

[0063] The semantic text and the corresponding fusion vector are stored in the vector database.

[0064] In one possible implementation, the processing module is further configured to:

[0065] In the first preset update cycle, the vector database is updated based on an incremental update mechanism.

[0066] In one possible implementation, the processing module is further configured to:

[0067] In the second preset update cycle, the vector database is updated based on a full update mechanism.

[0068] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0069] The memory stores computer-executed instructions;

[0070] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0071] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions that, when executed on a computer, cause the computer to perform the first aspect and / or various possible implementations of the first aspect.

[0072] This application provides a method, apparatus, equipment, and medium for wind power operation and maintenance (O&M) knowledge management. This method significantly improves the accuracy and professionalism of the wind power O&M knowledge question-and-answer system by integrating large language model fine-tuning and retrieval enhancement algorithms with unified management of structured and unstructured data in the wind power O&M field. By generating a high-quality training dataset from unstructured data according to a preset annotation template, a large language model with knowledge of the wind power O&M field is trained, enabling it to summarize and answer common and fundamental questions. Simultaneously, for structured data, semantic vector mapping is used and stored in a vector database. Combined with retrieval enhancement algorithms, this allows for rapid retrieval and accurate access to real-time data such as equipment parameters, operating data, and operating procedures, ensuring the accuracy and timeliness of answers. This method effectively covers the diverse and complex question-and-answer needs of wind power O&M sites, not only improving the response speed and accuracy to O&M issues and reducing reliance on human experience and knowledge transfer costs, but also enabling continuous updating and optimization of the wind power O&M knowledge base, constructing an intelligent knowledge management platform with real-time performance, professionalism, and sustainability. Attached Figure Description

[0073] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0074] Figure 1 This is a schematic diagram of the structure of the wind power operation and maintenance knowledge management system provided in the embodiments of this application;

[0075] Figure 2A flowchart illustrating the wind power operation and maintenance knowledge management method provided in this application embodiment. Figure 1 ;

[0076] Figure 3 A flowchart illustrating the wind power operation and maintenance knowledge management method provided in this application embodiment. Figure 2 ;

[0077] Figure 4 A schematic diagram of the structure of the wind power operation and maintenance knowledge management device provided in the embodiments of this application;

[0078] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0079] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0080] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0081] In wind power operation and maintenance (O&M) scenarios, knowledge management is crucial. Wind power equipment is diverse, has long operating cycles, and suffers from complex and widely distributed faults. A standardized and systematic knowledge management system can effectively accumulate and share O&M experience, standardize operating procedures, reduce human error and repeated trial and error, improve fault response and troubleshooting efficiency, and ensure the long-term stable operation of equipment. At the same time, knowledge management helps accelerate the training of new employees, reduces the risk of technical gaps caused by personnel turnover, and supports wind power companies in achieving their goals of safe, efficient, and sustainable O&M management.

[0082] In existing technologies, wind power operation and maintenance knowledge management methods generally construct entity-relation-entity triples through named entity recognition and relation extraction, forming a knowledge base similar to common questions and corresponding answers. When a user asks a question, the model vectorizes the user's input text, compares similarity, and outputs the answer corresponding to the most similar common question, thereby achieving the function of classifying and answering questions.

[0083] However, existing wind power operation and maintenance knowledge management methods can only provide a limited number of non-targeted answers, which cannot cope with the large number of diverse and dynamically changing personalized questions in wind power operation and maintenance sites. They also lack the ability to understand context and reason about complex knowledge, making it difficult to meet the higher requirements for accuracy, timeliness and professionalism in scenarios such as equipment fault diagnosis and operation and maintenance strategy optimization.

[0084] Based on this, this application proposes a wind power operation and maintenance (O&M) knowledge management method. Traditional O&M knowledge management methods can only perform simple matching based on preset question-answer pairs, failing to provide targeted and accurate answers, and lacking effective utilization of unstructured data and tacit knowledge. To overcome these problems, this application introduces a technical approach based on the fusion of large language model fine-tuning and retrieval enhancement. First, raw wind power O&M data is acquired, combined with structured data (such as wind farm regulations and O&M training materials) and unstructured data (such as wind power books, papers, and wind farm O&M records). For unstructured data, manual or automatic annotation is performed using preset annotation templates to construct a high-quality question-answer pair training dataset. This dataset is then used to fine-tune the large language model to enable it to possess language understanding and question-answering capabilities in the wind power O&M domain. Simultaneously, for structured data, semantic vectorization is performed using retrieval enhancement algorithms, establishing a mapping relationship between semantic vectors and the original text and storing it in a vector database to achieve real-time retrieval and content traceability. Finally, the trained large language model and the constructed vector database are deployed together on the wind power operation and maintenance knowledge management platform. When users ask questions, the model's generalization ability and the database's real-time retrieval ability are combined to output more accurate, timely and professional answers, thereby effectively making up for the shortcomings of traditional methods and improving the effectiveness and practical value of knowledge management and intelligent question answering in the field of wind power operation and maintenance.

[0085] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0086] Figure 1 This is a schematic diagram of the structure of the wind power operation and maintenance knowledge management system provided in the embodiments of this application; as shown below. Figure 1 As shown, the system includes: a data layer, a model layer, an application layer, and an operations layer.

[0087] The system comprises several layers: a data layer, serving as the knowledge foundation, aggregates multi-source data resources including professional books, academic papers, wind farm operation and maintenance records, regulations, training materials, and equipment manuals in the wind power field, providing a comprehensive and rich knowledge base for model training and optimization; a model layer, the core of the system's technology, fine-tunes a large model using wind power expertise, while combining RAG retrieval enhancement generation technology to achieve accurate knowledge output; an application layer transforms technical capabilities into practical products, seamlessly integrating with other application software through workflow construction; and an operation layer ensures the system's continuous and efficient operation by regularly updating the database, collecting user feedback, and iteratively optimizing the knowledge base, enabling the wind power operation and maintenance knowledge management platform to continuously adapt to the ever-changing actual needs of wind farms and provide long-term, stable technical support.

[0088] Figure 2 A flowchart illustrating the wind power operation and maintenance knowledge management method provided in this application embodiment. Figure 1 The method includes:

[0089] S201. Obtain raw data for wind power operation and maintenance.

[0090] The raw data for wind power operation and maintenance includes multiple structured data and multiple unstructured data.

[0091] It should be noted that structured data includes wind farm regulations, operation and maintenance training materials, and equipment manuals; unstructured data includes wind power books, wind power papers, and wind power operation and maintenance records.

[0092] S202. Label the unstructured data according to the preset labeling template to generate a training dataset.

[0093] In one feasible approach, firstly, the unstructured data is preprocessed to obtain preprocessed unstructured data; then, each preprocessed unstructured data is labeled according to a preset labeling template to generate question-answer pairs; finally, a training dataset is obtained based on multiple question-answer pairs.

[0094] The training dataset includes multiple question-answer pairs; preprocessing includes at least format conversion and removal of erroneous and redundant data.

[0095] It should be noted that, for unstructured data to be easily processed and analyzed by machines, the unstructured text needs to be converted into a uniform format, and then errors and redundant information need to be removed to improve data quality and processing efficiency; then a training dataset suitable for model learning should be created from the cleaned data.

[0096] Understandably, classifying and processing the structured and unstructured data in the original wind power operation and maintenance data not only improves data quality and processing efficiency, but also lays the foundation for building a high-quality training dataset, significantly improving the model's learning effect and output accuracy in the field of wind power.

[0097] S203. Train the large language model to be trained based on the training dataset to obtain a trained large language model.

[0098] Among them, the large language model is a model with intelligent question-answering capabilities in the field of wind power operation and maintenance.

[0099] It should be understood that the multiple question-answer pairs obtained from the aforementioned annotations are used as training samples and imported into the training framework according to a predetermined format to ensure that the data is consistent with the input and output requirements of the model. Then, a large language model with good language understanding ability is selected as the base model (such as the Qwen3-8B model, Gemma3 12B, etc.). Through supervised fine-tuning or parameter efficient fine-tuning (such as LoRA), targeted training is carried out using the wind power field training dataset to continuously optimize the model parameters, so that it can better adapt to the corpus expression and question-answering logic of this field. After training is completed, a dedicated large language model with professional knowledge and question-answering ability in the field of wind power operation and maintenance is obtained, which can accurately and standardly output professional answers based on the wind power-related questions raised by users.

[0100] It should be noted that, in addition to QLoRA fine-tuning technology, other efficient parameter fine-tuning methods can also be used, such as P-Tuning, Prompt-Tuning, or Adapter-Tuning, etc., and the most suitable fine-tuning strategy should be selected according to the specific scenario.

[0101] Understandably, this method significantly enhances the model's understanding and mastery of terminology, knowledge systems, and question-and-answer logic in the wind power operation and maintenance field, enabling it to possess more professional, accurate, and stable intelligent question-and-answer capabilities for wind power operation and maintenance. This effectively supports relevant users in conducting knowledge queries and assisting in decision-making during actual business operations, thereby improving work efficiency and accuracy.

[0102] S204. Based on the retrieval enhancement algorithm, the structured data is processed to obtain the mapping relationship between semantic vectors and semantic text, and the mapping relationship between semantic vectors and semantic text is stored in the vector database.

[0103] Understandably, by semantically vectorizing structured data in the wind power sector and constructing mapping relationships, and then managing and storing it in a vector database, the accuracy and response speed of subsequent semantic-based retrieval can be significantly improved. This ensures that the model can quickly and accurately retrieve relevant normative original texts or data when users ask questions, thereby effectively supporting the Retrieval Augmentation (RAG) process and guaranteeing the authority, accuracy, and timeliness of the answers.

[0104] S205. Combine the trained large language model with the vector database to build a wind power operation and maintenance knowledge management platform to provide functional services to users.

[0105] The functional services include rules and regulations inquiry, equipment operation guidance, operation and maintenance problem answering, and training and support.

[0106] Understandably, by employing a dual-pipeline parallel processing approach of fine-tuning and retrieval enhancement, a professional knowledge graph in the field of wind power operation and maintenance can be constructed and combined with a large language model to achieve more accurate knowledge reasoning and correlation analysis.

[0107] It should be noted that, in order to balance the comprehensiveness, accuracy, timeliness and maintenance cost of the knowledge base, this application adopts a dual-track update mechanism of full update and incremental update for the training dataset to meet the continuous operation requirements of intelligent application systems based on large models.

[0108] In the first preset update cycle, the vector database is updated based on an incremental update mechanism; in the second preset update cycle, the vector database is updated based on a full update mechanism.

[0109] Specifically, full updates are performed every 3 months, matching the update frequency of the large model to maintain its technological advancement. If the large model is replaced, LoRA fine-tuning is re-executed; otherwise, fine-tuning is performed based on the last full update model using the new data. Incremental updates are performed every 2-4 weeks based on new data, with a single round of fine-tuning. After a full update, the latest incremental version is discarded, forming a complete update process.

[0110] Understandably, the dual-track mechanism of low-frequency full-volume maintenance and high-frequency incremental maintenance can balance daily maintenance efficiency with long-term effects, reducing costs while maintaining continuous optimization of results.

[0111] This application provides a wind power operation and maintenance (O&M) knowledge management method. This method significantly improves the accuracy and professionalism of the wind power O&M knowledge question-and-answer system by integrating large language model fine-tuning and retrieval enhancement algorithms with unified management of structured and unstructured data in the wind power O&M field. By generating a high-quality training dataset from unstructured data according to a preset annotation template, a large language model with knowledge of the wind power O&M field is trained, enabling it to summarize and answer common and fundamental questions. Simultaneously, for structured data, semantic vector mapping is used and stored in a vector database. Combined with retrieval enhancement algorithms, this allows for rapid retrieval and accurate access to real-time data such as equipment parameters, operating data, and operating procedures, ensuring the accuracy and timeliness of answers. This method effectively covers the diverse and complex question-and-answer needs of wind power O&M sites, not only improving the response speed and accuracy to O&M issues and reducing reliance on human experience and knowledge transfer costs, but also enabling continuous updating and optimization of the wind power O&M knowledge base, constructing an intelligent knowledge management platform with real-time performance, professionalism, and sustainability.

[0112] In one feasible approach, the structured data is first semantically divided based on a preset segmentation strategy to obtain multiple semantic texts, each containing complete contextual information. Then, for each semantic text, a semantic vector corresponding to the semantic text is generated using a trained large language model. Next, a semantic knowledge graph is constructed based on the similarity between the semantic texts. Finally, the semantic texts, their corresponding semantic vectors, and the semantic knowledge graph are stored in a vector database.

[0113] The preset segmentation strategy can be based on paragraph or chapter titles, sentence structure, or semantic boundary recognition using natural language models.

[0114] It should be understood that structured data requires appropriate methods to achieve semantic retrieval and reasoning that more closely resembles human thinking, significantly improving accuracy and relevance. Specifically, structured data in the wind power field (such as regulations, training materials, equipment manuals, etc.) should be rationally divided into logical units to obtain several semantic texts with complete contextual information, facilitating subsequent modeling and retrieval. Then, a language model is used to semantically encode each semantic text, obtaining a semantic vector representing its deep meaning, and the semantic vector similarity (such as cosine similarity) between all semantic texts is calculated. A semantic knowledge graph is then established based on the similarity relationship. Finally, the semantic texts, their corresponding semantic vectors, and the semantic knowledge graph are stored in a vector database. By managing the data through the vector database, efficient similarity retrieval and association reasoning can be used to achieve efficient retrieval and scalable management of structured knowledge, facilitating subsequent applications such as enhanced retrieval generation and intelligent question answering.

[0115] Optionally, the method also includes:

[0116] Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts. Then, the image data and table data in the structured data are descriptively embedded using a multimodal model to obtain embedding vectors.

[0117] For each semantic text, a semantic vector corresponding to the semantic text is generated through a language model, including generating a fusion vector based on the semantic vector and the embedding vector.

[0118] The process of storing semantic text and its corresponding semantic vectors in a vector database includes storing semantic text and its corresponding fused vectors in a vector database.

[0119] It should be understood that after dividing structured data into multiple semantic texts with complete context based on a preset segmentation strategy, a multimodal model is further used to embed image data (such as wind power equipment topology diagrams) and tabular data (such as operating parameter tables) within the structured data, generating embedding vectors that represent their semantic content. Simultaneously, a language model generates corresponding semantic vectors for each semantic text, and these text semantic vectors are further fused with the chart embedding vectors to obtain a more comprehensive and information-dense fused vector. Finally, the semantic texts and their corresponding fused vectors are uniformly stored in a vector database, forming a multimodal integrated knowledge vector index encompassing structured data, image data, and tabular data.

[0120] Understandably, by fusing multimodal information such as text and charts, the semantic representation capability of structured data in machine understanding can be comprehensively improved, solving the problem of insufficient representation by traditional single text vectors. This enables subsequent retrieval or question-answering systems not only to accurately understand text content but also to respond precisely by associating chart data. The stored fused vectors possess higher semantic relevance and retrieval performance, significantly enhancing the accuracy, comprehensiveness, and practicality of wind power-related knowledge bases in intelligent retrieval and question-answering reasoning.

[0121] Figure 3 A flowchart illustrating the wind power operation and maintenance knowledge management method provided in this application embodiment. Figure 2 ,like Figure 3 As shown, in this embodiment... Figure 2 Based on the examples, the process of training a large language model is described in detail, and the method includes:

[0122] S301. Insert a device-level LoRA adapter into the mid-to-low-level conversion module of the large language model to be trained, and insert a site-level LoRA adapter and a fusion adapter into the mid-to-high-level conversion module of the large language model to be trained.

[0123] It should be noted that Low-Rank Adaptation (LoRA) adjusts the model's expressive power for specific domain tasks by introducing trainable low-rank matrices into some layers of the transformation module, without changing the original weights. The insertion position of LoRA at different levels determines the level of abstraction it models.

[0124] The low-to-mid level conversion module is responsible for modeling and understanding lower-level language features, inserting equipment-level LoRA to focus on learning fine-grained information such as operation and maintenance terminology, principles, and common problems at the equipment level, including wind turbines, converters, and towers. The high-to-mid level conversion module is responsible for organizing higher-level abstract knowledge and semantic reasoning, inserting site-level LoRA to focus on learning macro-level knowledge such as overall wind farm operation management, scheduling strategies, and site regulations. The high-to-mid level fusion adapter undertakes the task of cross-layer fusion of knowledge at different levels (equipment vs. site), bridging the information gap caused by the hierarchical span and improving the consistency of overall contextual understanding.

[0125] Understandably, by inserting different LoRA adapters in layers, the model can be equipped with multi-level knowledge modeling capabilities at different granularities from equipment to site, thereby refining domain understanding, improving the accuracy of multi-level problem response, and avoiding information mixing or missing.

[0126] S302. Train the device-level LoRA adapter based on the device-level wind power operation and maintenance data in the training dataset to adjust the parameters of the device-level LoRA adapter, and train the site-level LoRA adapter based on the site-level wind power operation and maintenance data in the training dataset to adjust the parameters of the site-level LoRA adapter.

[0127] It should be understood that this step uses a supervised fine-tuning method to train the corresponding LoRA adapters independently using the training dataset, only updating the LoRA parameter matrix of the insertion layer and freezing other model layers. This ensures that different layers focus on their respective tasks and do not cause cross-contamination, allowing them to learn the knowledge representation capabilities of their own level and form clear knowledge boundaries and modular capabilities.

[0128] Understandably, this approach enables lower- and middle-level staff to have a strong ability to characterize equipment problems and provide more accurate and professional answers; it also enhances the understanding and summarization abilities of middle and senior management regarding station-level management issues, avoiding confusion and misuse of knowledge.

[0129] S303. Based on the correlation between equipment-level wind power operation and maintenance data and station-level wind power operation and maintenance data, construct a fusion test set.

[0130] It should be noted that although equipment-level and station-level data are at different levels, they are inherently related (such as equipment failure affecting station operation and maintenance strategies). It is necessary to merge test sets to simulate real business cross-level correlation problems, train the model to understand the linkage of data at different levels, and thus provide a data foundation for subsequent fusion adapter training, ensuring that the model can correctly associate knowledge at different levels and avoid local overfitting or logical gaps.

[0131] S304. Train the fusion adapter based on the fusion test set to adjust the parameters of the fusion adapter.

[0132] S305. Integrate the trained device-level LoRA adapter, site-level LoRA adapter, and fusion adapter into the large language model to be trained to obtain the trained large language model.

[0133] It should be noted that the Multilayer Perceptron (MLP) layer in the conversion module, as the core component for high-order feature transformation and nonlinear mapping in the model, has a large number of parameters and strong representational capabilities, making it a key area for model knowledge storage. By injecting device-level LoRA adapters and site-level LoRA adapters here, the feature representations within the model can be efficiently adjusted in a lightweight manner, enabling the model to accurately adapt to the specific scenario requirements of different levels while retaining its general language capabilities. Therefore, the device-level LoRA adapter and site-level LoRA adapter are set in the MLP layer of the conversion module.

[0134] Understandably, this approach enables the creation of a large-scale wind power model with hierarchical understanding, fine-grained modeling, and cross-layer linkage reasoning capabilities. This model can accurately cover equipment, stations, and integrate various operation and maintenance question-and-answer scenarios, thereby improving overall domain adaptability and response accuracy, and realizing hierarchical association and multi-granular representation of wind farm operation and maintenance knowledge.

[0135] Figure 4 This is a schematic diagram of the structure of the wind power operation and maintenance knowledge management device provided in the embodiments of this application; as shown below. Figure 4 As shown, the device includes:

[0136] The acquisition module 401 is used to acquire raw wind power operation and maintenance data, which includes multiple structured data and multiple unstructured data.

[0137] The generation module 402 is used to annotate unstructured data according to a preset annotation template to generate a training dataset, which includes multiple question-answer pairs;

[0138] Training module 403 is used to train the large language model to be trained based on the training dataset to obtain a trained large language model. The trained large language model is a model with intelligent question answering capabilities in the field of wind power operation and maintenance.

[0139] The processing module 404 is used to process structured data based on the retrieval enhancement algorithm, obtain the mapping relationship between semantic vectors and semantic text, and store the mapping relationship between semantic vectors and semantic text in the vector database.

[0140] Application module 405 is used to combine the trained large language model and vector database to build a wind power operation and maintenance knowledge management platform to provide functional services to users.

[0141] In one possible implementation, the generation module 402 is specifically used for:

[0142] Unstructured data is preprocessed to obtain preprocessed unstructured data. Preprocessing includes at least format conversion and removal of erroneous and redundant data.

[0143] Each preprocessed unstructured data point is labeled according to a preset labeling template, generating question-answer pairs;

[0144] The training dataset is obtained from multiple question-answer pairs.

[0145] In one possible implementation, the training module 403 is specifically used for:

[0146] Insert a device-level LoRA adapter into the mid-to-low-level conversion module of the large language model to be trained, and insert a site-level LoRA adapter and a fusion adapter into the mid-to-high-level conversion module of the large language model to be trained.

[0147] The device-level LoRA adapter is trained based on the device-level wind power operation and maintenance data in the training dataset to adjust the parameters of the device-level LoRA adapter, and the site-level LoRA adapter is trained based on the site-level wind power operation and maintenance data in the training dataset to adjust the parameters of the site-level LoRA adapter.

[0148] Based on the correlation between equipment-level wind power operation and maintenance data and site-level wind power operation and maintenance data, a fusion test set is constructed;

[0149] The fusion adapter is trained based on the fusion test set to adjust its parameters;

[0150] The trained device-level LoRA adapter, site-level LoRA adapter, and fusion adapter are integrated into the large language model to be trained to obtain the trained large language model.

[0151] In one possible implementation, the processing module 404 is specifically used for:

[0152] Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts, each of which contains complete contextual information.

[0153] For each semantic text, a semantic vector corresponding to the semantic text is generated through a language model;

[0154] Construct a semantic knowledge graph based on the similarity between semantic texts;

[0155] The semantic text, its corresponding semantic vectors, and the semantic knowledge graph are stored in the vector database.

[0156] In one possible implementation, the processing module 404 is further configured to:

[0157] Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts. Then, the image data and table data in the structured data are descriptively embedded using a multimodal model to obtain embedding vectors.

[0158] For each semantic text, a semantic vector corresponding to the semantic text is generated through a language model, including:

[0159] Generate a fused vector based on the semantic vector and the embedding vector;

[0160] The semantic text and its corresponding semantic vectors are stored in a vector database, including:

[0161] The semantic text and its corresponding fused vector are stored in the vector database.

[0162] In one possible implementation, the processing module 404 is further configured to:

[0163] In the first preset update cycle, the vector database is updated based on the incremental update mechanism.

[0164] In one possible implementation, the processing module 404 is further configured to:

[0165] In the second preset update cycle, the vector database is updated based on the full update mechanism.

[0166] The wind power operation and maintenance knowledge management device provided in this application embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0167] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5As shown, the electronic device 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus 504.

[0168] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.

[0169] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0170] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0171] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0172] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0173] This application also provides a computer-readable storage medium storing computer-executable instructions that, when executed on a computer, cause the computer to perform the above-described method.

[0174] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0175] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0176] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0177] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0178] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0179] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0180] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0181] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A wind power operation and maintenance knowledge management method, characterized in that, include: Obtain raw wind power operation and maintenance data, which includes multiple structured data and multiple unstructured data. The unstructured data is labeled according to a preset labeling template to generate a training dataset, which includes multiple question-answer pairs. The large language model to be trained is trained based on the training dataset to obtain a trained large language model, which is a model with intelligent question answering capabilities in the field of wind power operation and maintenance. The structured data is processed based on the retrieval enhancement algorithm to obtain the mapping relationship between semantic vectors and semantic text, and the mapping relationship between semantic vectors and semantic text is stored in the vector database. The trained large language model and the vector database are combined to build a wind power operation and maintenance knowledge management platform to provide functional services to users; The process of training the large language model to be trained based on the training dataset to obtain a trained large language model includes: A device-level low-rank adaptive LoRA adapter is inserted into the middle and lower-level conversion modules of the large language model to be trained, and a station-level LoRA adapter and a fusion adapter are inserted into the middle and higher-level conversion modules of the large language model to be trained. The device-level LoRA adapter is trained based on the device-level wind power operation and maintenance data in the training dataset to adjust the parameters of the device-level LoRA adapter, and the site-level LoRA adapter is trained based on the site-level wind power operation and maintenance data in the training dataset to adjust the parameters of the site-level LoRA adapter. Based on the correlation between the equipment-level wind power operation and maintenance data and the station-level wind power operation and maintenance data, a fusion test set is constructed. The fusion adapter is trained based on the fusion test set to adjust the parameters of the fusion adapter; The trained device-level LoRA adapter, the station-level LoRA adapter, and the fusion adapter are integrated into different positions in the large language model to be trained to obtain the trained large language model.

2. The method according to claim 1, characterized in that, The step of labeling the unstructured data according to a preset labeling template to generate a training dataset includes: The unstructured data is preprocessed to obtain preprocessed unstructured data. The preprocessing includes at least format conversion and removal of erroneous and redundant data. Each piece of preprocessed unstructured data is labeled according to the preset labeling template to generate question-answer pairs; The training dataset is obtained from multiple question-answer pairs.

3. The method according to claim 1, characterized in that, The process of processing the structured data based on the retrieval enhancement algorithm to obtain the mapping relationship between semantic vectors and semantic text, and storing the mapping relationship between semantic vectors and semantic text in a vector database, includes: Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts, each of which contains complete contextual information. For each semantic text, a semantic vector corresponding to the semantic text is generated using the trained large language model; Based on the similarity between the semantic texts, a semantic knowledge graph is constructed; The semantic text, its corresponding semantic vector, and the semantic knowledge graph are stored in the vector database.

4. The method according to claim 3, characterized in that, The method further includes: Based on a preset segmentation strategy, the structured data is semantically divided to obtain multiple semantic texts. Then, the image data and table data in the structured data are descriptively embedded using a multimodal model to obtain embedding vectors. For each semantic text, generating a semantic vector corresponding to the semantic text through a language model includes: Generate a fusion vector based on the semantic vector and the embedding vector; The step of storing the semantic text and its corresponding semantic vector into the vector database includes: The semantic text and the corresponding fusion vector are stored in the vector database.

5. The method according to claim 1, characterized in that, The method further includes: In the first preset update cycle, the vector database is updated based on an incremental update mechanism.

6. The method according to claim 1, characterized in that, The method further includes: In the second preset update cycle, the vector database is updated based on a full update mechanism.

7. A wind power operation and maintenance knowledge management device, characterized in that, include: The acquisition module is used to acquire raw wind power operation and maintenance data; the raw wind power operation and maintenance data includes multiple structured data and multiple unstructured data. The generation module is used to annotate the unstructured data according to a preset annotation template to generate a training dataset, which includes multiple question-answer pairs; The training module is used to train the large language model to be trained based on the training dataset to obtain a trained large language model, which is a model with intelligent question answering capabilities in the field of wind power operation and maintenance. The processing module is used to process the structured data based on the retrieval enhancement algorithm to obtain the mapping relationship between semantic vectors and semantic text, and to store the mapping relationship between semantic vectors and semantic text in the vector database. The application module is used to deploy the large language model and the vector database into a wind power operation and maintenance knowledge management platform to provide functional services to users; The training module is specifically used to insert a device-level low-rank adaptive LoRA adapter into the middle and lower-level conversion modules of the large language model to be trained, and to insert a site-level LoRA adapter and a fusion adapter into the middle and higher-level conversion modules of the large language model to be trained. The device-level LoRA adapter is trained based on the device-level wind power operation and maintenance data in the training dataset to adjust the parameters of the device-level LoRA adapter, and the site-level LoRA adapter is trained based on the site-level wind power operation and maintenance data in the training dataset to adjust the parameters of the site-level LoRA adapter. Based on the correlation between the equipment-level wind power operation and maintenance data and the station-level wind power operation and maintenance data, a fusion test set is constructed. The fusion adapter is trained based on the fusion test set to adjust the parameters of the fusion adapter; The trained device-level LoRA adapter, the station-level LoRA adapter, and the fusion adapter are integrated into different positions in the large language model to be trained to obtain the trained large language model.

8. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Knowledge-operation mapping fine-tuning LLM-based power system calibration tuning agent

    CN119357321A

  • Knowledge question-answering system based on large language model

    CN119396975A