Power plant equipment retrieval and question and answer method and system based on RAG

By preprocessing and slicing power plant equipment data using a RAG-based method, and combining natural language processing and the RAG question-answering engine model, the problems of low information retrieval efficiency and low answer credibility in power plant equipment management are solved, achieving efficient and accurate question-answering services.

CN120910194APending Publication Date: 2025-11-07YANCHI ZHONGYING CHUANGNENG NEW ENERGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510967005.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Traditional information retrieval methods are inefficient in power plants, generative models have low reliability, and retrieval-based question-answering systems struggle to understand complex questions, failing to meet the actual needs of power plant equipment management.

Method used

The RAG-based approach is used to preprocess and slice power plant equipment data, and semantic features are extracted through natural language processing to form a vector database. Combined with the RAG question-answering engine model, accurate retrieval and answer generation are performed to ensure the credibility and accuracy of the answers.

Benefits of technology

It improves the efficiency of power plant equipment management, reduces model illusion, provides more credible and accurate answers, can understand complex questions, and adapts to the diverse question-and-answer needs of power plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910194A_ABST
    Figure CN120910194A_ABST
Patent Text Reader

Abstract

The invention discloses an RAG-based power plant equipment retrieval and question-answering method, system, equipment, medium and program, and the method comprises the steps: obtaining the original data of power plant equipment, preprocessing the data of the power plant equipment, and obtaining the processed data of the power plant equipment; performing slicing processing on the processed power plant equipment data to obtain sliced power plant equipment data; extracting semantic features from the sliced power plant equipment data by using a natural language processing technology to form a vector database; searching power plant equipment data corresponding to the question in a vector database by using an RAG question and answer engine model; and inputting the power plant equipment data corresponding to the question and the question into the generative model for processing to obtain an answer corresponding to the question. The method is based on the RAG technology, deep fusion of professional documents and intelligent questions and answers is achieved, the model illusion phenomenon is reduced, the answer credibility is improved, and the power plant equipment management efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data electric digital data processing, and particularly relates to a power plant equipment retrieval and question-answering method, system, device, medium and program based on RAG. BACKGROUND

[0002] In the daily operation and management system of a power plant, ensuring the stable operation state of equipment is a core link for guaranteeing the continuous supply of electricity and maintaining the safe and orderly production. The work efficiency and accuracy of an operator, who is the direct executor of equipment operation monitoring and maintenance, are directly related to the overall operation quality of the power plant. To complete this key task, the operator needs to frequently consult various professional documents such as equipment manuals, maintenance manuals and DCS (Distributed Control System) operation guides in daily work. These documents cover a large amount of information such as technical parameters, operation processes, maintenance points and fault elimination methods of equipment, and are the "knowledge treasure house" for the operator to deal with various problems in the equipment operation process.

[0003] However, the traditional information retrieval method gradually exposes many drawbacks when dealing with the complex and massive knowledge system of a power plant. Keyword search, as the most common information retrieval means, has obvious limitations. Due to the rich and specific context of power plant professional terms, simple keyword matching often cannot accurately locate the required information, resulting in a large amount of irrelevant content in the search results. The operator needs to spend a lot of time screening useful parts from a large amount of information, which greatly reduces the query efficiency. Artificially reviewing documents is even more time-consuming and laborious. In the face of thick operation manuals and equipment manuals, the operator needs to quickly find the answer to a specific problem among numerous chapters, which is like finding a needle in a haystack. Not only is it inefficient, but it is also easy to miss key information due to human negligence.

[0004] With the rapid development of artificial intelligence technology, intelligent question answering systems have brought new ideas and solutions to information retrieval. Currently, generative models based on neural networks (such as the GPT series) have made significant progress in the field of intelligent question answering and are widely used in various scenarios. However, in the specific field of power plants, such models have obvious shortcomings. On the one hand, generative models often lack accurate references to professional documents when generating answers. They tend to generate seemingly reasonable answers based on training data, but these answers may lack actual basis and deviate from the actual situation of power plant equipment, resulting in a significant reduction in the credibility of the answers. For power plant operators, accurate and reliable information is crucial, and any minor error can cause serious equipment failure or even safety accidents, so such answers lacking accurate references cannot meet the actual needs. On the other hand, training large language models requires massive data as support, while the power plant field has high professionalism and specificity, with relatively limited and scattered relevant data. This makes it difficult to directly apply general large language models to power plant question answering needs, as the model is difficult to fully learn the professional knowledge of the power plant field and provide accurate and effective answers. In addition, although pure retrieval-based question answering systems can perform information retrieval based on existing document libraries, they have obvious shortcomings in understanding complex user questions. Power plant operators may use complex sentence patterns, include multiple conditions, or ask questions related to multiple devices, and pure retrieval-based question answering systems are difficult to accurately understand these complex semantics, making it difficult to provide in-depth analysis and comprehensive and accurate answers, thus limiting their application effectiveness in power plant operations. SUMMARY

[0005] To solve the problems of time-consuming and inefficient retrieval, low credibility of generative models, and difficulty in understanding complex questions for retrieval-based systems, which cannot meet the needs of power plant question answering, the present application provides a RAG-based power plant equipment retrieval and question answering method. Based on RAG technology, the method realizes the deep integration of professional documents and intelligent question answering, reduces model hallucination, improves answer credibility, and improves power plant equipment management efficiency.

[0006] To achieve the above-mentioned purposes, the present application provides the following technical solutions.

[0007] In a first aspect, the present application provides a RAG-based power plant equipment retrieval and question answering method, comprising: Obtaining power plant equipment raw data, preprocessing the power plant equipment data to obtain processed power plant equipment data; Slicing the processed power plant equipment data to obtain sliced power plant equipment data; Extracting semantic features from the sliced power plant equipment data using natural language processing technology to form a vector database; The RAG question and answer engine model is used to find the power plant equipment data corresponding to the question in the vector database. The power plant equipment data corresponding to the question and the question are input into the generation model for processing to obtain the answer corresponding to the question.

[0008] As a further improvement of the present application, the acquisition of the power plant equipment raw data, the preprocessing of the power plant equipment data, and the obtaining of the processed power plant equipment data, comprise: The power plant equipment raw data, including the equipment manual, the maintenance manual, the DCS operation guide and the maintenance record, are acquired. The power plant equipment raw data are subjected to data cleaning processing and format correction to obtain the processed power plant equipment data.

[0009] As a further improvement of the present application, the slice processing of the processed power plant equipment data, the sliced power plant equipment data, comprise: The granularity of the slice is determined according to the requirement of the set target. The semantic relationship in the power plant equipment data is analyzed to find the division point. The processed power plant equipment data are subjected to slice processing based on the granularity of the slice and the division point to obtain the sliced power plant equipment data.

[0010] As a further improvement of the present application, the extraction of the semantic features of the sliced power plant equipment data using the natural language processing technology to form the vector database, comprise: The semantic features of the sliced power plant equipment data are extracted using the natural language processing technology, and the extracted semantic features are converted into high-dimensional vectors to form the vector database.

[0011] As a further improvement of the present application, the use of the RAG question and answer engine model to find the power plant equipment data corresponding to the question in the vector database, comprise: The question text is converted into a query vector. The RAG question and answer engine model is used to perform similarity matching in the vector database according to the query vector to retrieve the power plant equipment data related to the question.

[0012] As a further improvement of the present application, the input of the power plant equipment data corresponding to the question and the question into the generation model for processing to obtain the answer corresponding to the question, comprise: The power plant equipment data corresponding to the question and the question are input into the generation model for processing to generate the answer. The RAG question and answer engine model is used to verify the generated answer and the power plant equipment raw data, and if the verification result meets the set requirement, the answer corresponding to the question is obtained.

[0013] In a second aspect, the present application provides a RAG-based power plant equipment retrieval and question-answering system, comprising: a data processing module configured to obtain raw data of the power plant equipment, pre-process the data of the power plant equipment, and obtain processed data of the power plant equipment; a data slicing module configured to slice the processed data of the power plant equipment, and obtain sliced data of the power plant equipment; a database module configured to extract semantic features of the sliced data of the power plant equipment using a natural language processing technology, and form a vector database; a data searching module configured to search for the power plant equipment data corresponding to a question in the vector database using a RAG question-answering engine model; an answer obtaining module configured to input the power plant equipment data corresponding to the question and the question into a generation model for processing, and obtain an answer corresponding to the question.

[0014] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the RAG-based power plant equipment retrieval and question-answering method.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the RAG-based power plant equipment retrieval and question-answering method.

[0016] In a fifth aspect, the present application provides a computer program product, comprising computer instructions, and the computer instructions are executed by a processor to implement the steps of the RAG-based power plant equipment retrieval and question-answering method.

[0017] Compared with the prior art, the present application has the following beneficial effects: The application forms a vector database by preprocessing and slicing the original data of power plant equipment and extracting semantic features using natural language processing technology. When searching, the RAG question and answer engine model can quickly locate the power plant equipment data matching the question in the vector database, greatly shortening the search time, improving the search efficiency, allowing staff to quickly obtain the required information, and gaining valuable time for timely maintenance and management of power plant equipment. Second, the application combines the advantages of search and generation. First, use RAG technology to accurately search for power plant equipment data highly related to the question, and then input these data together with the question into the generation model. This answer generation method based on reliable data effectively reduces the model illusion phenomenon and avoids the errors or misleading answers generated by the generative model due to the lack of effective information support, ensuring that the generated answers have higher credibility and accuracy, providing a solid basis for the correct operation and decision-making of power plant equipment. Finally, the method can better understand the true intention of the question by deeply extracting semantic features using natural language processing technology. Even in the face of complex, ambiguous or multi-semantic questions, it can accurately grasp the key information and find the truly matching power plant equipment data in the vector database, effectively solving the understanding problem of the retrieval model in complex question scenarios, widening the application range of the question and answer system, and making it adapt to the diversified questioning needs in the actual work of power plants. BRIEF DESCRIPTION OF DRAWINGS

[0018] The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way. In the drawings: Figure 1 A flowchart of a RAG-based power plant equipment search and question and answer method; Figure 2 A RAG knowledge base construction flowchart in a RAG-based power plant equipment search and question and answer method; Figure 3 A structure diagram of a RAG-based power plant equipment search and question and answer system; Figure 4 An electronic device schematic diagram in an embodiment of the application. DETAILED DESCRIPTION

[0019] In order to enable personnel skilled in the art to better understand the technical solutions in the present application, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application, and the described embodiments are only some of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor should fall within the scope of the present application.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0021] To address the problems of inefficient retrieval, low reliability of generative models, and complex, difficult-to-understand retrieval queries in existing technologies, which fail to meet the question-and-answer requirements of power plants, this invention provides a power plant equipment retrieval and question-and-answer method based on RAG (Research and Analysis Groups). Figure 1 As shown, it includes: S100: Obtain raw data of power plant equipment, preprocess the power plant equipment data, and obtain processed power plant equipment data. S200: The processed power plant equipment data is sliced; the sliced ​​power plant equipment data. S300: Extract semantic features from the sliced ​​power plant equipment data using natural language processing technology to form a vector database; S400: Uses the RAG question-answering engine model to find the power plant equipment data corresponding to the question in the vector database; S500: Input the power plant equipment data and the question into the generative model for processing to obtain the answer to the question.

[0022] This method, based on RAG technology, achieves deep integration of professional documents and intelligent question answering, reduces model illusion, improves answer credibility, and enhances the efficiency of power plant equipment management.

[0023] The present application will be further explained and illustrated below with reference to specific embodiments.

[0024] like Figure 1 As shown, the RAG-based power plant equipment retrieval and question-answering method of the present invention includes: S1: Data Collection Acquire relevant data from power plant equipment and organize it into raw data for power plant equipment.

[0025] Specifically, this involves acquiring and organizing various technical documents related to power plant equipment, including equipment manuals, maintenance manuals, and DCS operation guidelines. Data sources can include electronic documents provided by manufacturers, internal maintenance records, and other relevant materials to ensure the comprehensiveness and authority of the information.

[0026] S2: Data Cleaning The raw data of power plant equipment is cleaned and standardized to remove redundant information, correct format errors, and unify technical terms for subsequent data processing and retrieval. This improves the effectiveness of subsequent data analysis, retrieval, and generation. Technical documents related to power plants are diverse and often contain a lot of redundant information, format errors, and inconsistent terminology, which can affect data quality, leading to low retrieval efficiency and even affecting the accuracy of generated models.

[0027] Specifically, a text similarity algorithm is used to determine the similarity between documents or segments. If the similarity of two segments exceeds a predetermined threshold, they are considered duplicate content and are removed. Hash calculation is performed on the document content, and it is checked whether there are duplicate hash values to avoid storing duplicate data.

[0028] After deduplication, the power plant equipment data is corrected in format, including punctuation correction, paragraph structure correction, and handling of invalid characters. This ensures the readability and standardization of the document.

[0029] After format modification, the power plant data is unified in technical terms.

[0030] Under the premise of ensuring information integrity, the power plant equipment data after technical term unification is sorted for missing values. Under the premise of ensuring information integrity, missing parts are filled or deleted. If the missing content has no impact on subsequent data processing, the part can be directly deleted. When deleting, consider whether the deleted content is important to avoid deleting critical information. For some missing data, use appropriate methods to fill in. For example, use information from other parts of the document to infer and fill in, or fill in default values based on domain knowledge.

[0031] S3: Slice processing of data cleaned power plant equipment data Specifically, to improve retrieval efficiency and accuracy, the system divides the document into logical segments, each containing complete technical information. The slicing can be based on chapters, paragraphs, or specific technical keywords, making the retrieval process more targeted.

[0032] S31: Determine the requirements of the set target and decide the granularity of slicing based on the hierarchical structure in the power plant equipment data.

[0033] S32: Use machine learning and natural language processing techniques to analyze the semantic relationships in the power plant equipment data to determine the most suitable place for slicing. For example, use syntactic analysis to determine logically coherent blocks of information to ensure that each slice is self-contained, making it easier for subsequent retrieval and model generation.

[0034] S33: By pre-defined key automatic segmentation of documents, ensure that document slices contain relevant technical content. For example, when a document mentions a specific device or operation, it is automatically divided into a slice to provide detailed information related to it directly when queried later. For documents involving complex technical terms, keyword-driven slicing can be used. When slicing, not only does each segment have complete technical information, but automatic summarization techniques can also be used to compress and refine long text segments, providing a condensed version of the slice, reducing redundancy, and improving document retrieval efficiency.

[0035] S34: To ensure the quality of the text after slicing, a quality check mechanism can be established to evaluate the relevance, accuracy, and completeness of each segment, automatically marking or correcting potential problems such as irrelevant information, repetitive content, etc. This helps improve the accuracy of subsequent vectorization and retrieval.

[0036] S4: Vectorization of power plant equipment data after slicing Based on RAG technology, build RAG question and answer engine model, through question and answer engine combined with document retrieval and language model to generate answer. RAG question and answer engine model is composed of multiple core modules, including user interface layer, question and answer engine layer, document management layer and model training optimization layer. The user interface layer supports web and mobile queries, supports voice input and text input, and improves user query convenience. The question and answer engine layer is responsible for processing user input, calling the RAG question and answer engine model for retrieval and generation to provide accurate device question and answer services. The document management layer is responsible for storing, indexing and updating power plant related documents, and ensures that the system always obtains the latest device information. The model training optimization layer supports continuous learning, optimizes the question and answer effect based on user feedback, so that the system can continuously improve the response quality and accuracy.

[0037] S41: Natural language processing (NLP) semantic analysis of power plant equipment data, converting questions into machine-understandable query vectors. The sliced text is also converted into high-dimensional vector representation, using natural language processing (NLP) technology (such as BERT, Sentence Transformers) to extract semantic features, forming a vector database for efficient retrieval later.

[0038] S42: Based on the parsed query vector, the RAG question and answer engine model uses the vector database (such as FAISS, Milvus) for similarity matching to retrieve the most relevant document slices. The vector database finds the most matching document paragraphs based on pre-defined similarity measures such as cosine similarity.

[0039] S5: Get the search answer S51: The retrieved document fragments are input into the generation model along with the user question, which generates an answer by combining the retrieved information with the context, ensuring semantic accuracy and technical content coherence.

[0040] S52: The generated answer is automatically verified, such as compared with the original document, to ensure its compliance with the latest information of power plant equipment. If the RAG question and answer engine model automatically verifies failure or has doubts, the artificial intervention mechanism will be started, and the answer will be manually reviewed and corrected by experts or administrators. Artificial intervention can also modify, supplement or confirm the generated answer, ultimately providing higher quality output.

[0041] S53: The RAG question and answer engine model records user feedback. If the user is not satisfied with the answer or believes that some information is inaccurate, the system will learn from itself and further optimize the model through incremental learning strategies to ensure that subsequent query answers are more accurate and reliable.

[0042] RAG question and answer engine model. Supervised training using annotated question and answer datasets improves the model's understanding of power plant equipment-related questions. Incremental learning strategies enable the model to adapt to new documents and new equipment. Through user feedback mechanisms, continuously optimize answer quality to ensure the system can provide high-quality question and answer content. In addition, the RAG question and answer engine model prediction part, after the user input question, first performs semantic analysis, triggers the retrieval module to extract relevant documents from the knowledge base, and combines the retrieval content to provide the final answer by the generation model, and attaches relevant document references to enhance the credibility and explainability of the answer.

[0043] The RAG question and answer engine model uses an efficient data storage solution, such as a vector database (FAISS, Milvus), to store and index text vectors, and combines traditional database storage of original text data for subsequent query and management.

[0044] After the user inputs the query, the RAG question and answer engine model first uses vector search technology to retrieve the most relevant document fragments from the database, and then inputs the retrieved information into the generation model (such as DeepSeek, ChatGLM) to generate answers, ensuring accurate and reliable answers. To ensure the efficiency and reliability of the system, the invention uses multiple evaluation methods to optimize the intelligent question and answer system. The system uses BLEU, ROUGE and other text evaluation indicators to measure the accuracy of question and answer, and continuously optimizes system performance through manual annotation and user feedback. In addition, expert knowledge-based test sets are used for verification to ensure that the model-generated answers meet the actual needs of power plants and can effectively assist in equipment operation and management.

[0045] In summary, this application achieves a deep integration of professional documents and intelligent question answering by combining RAG technology with retrieval and generative methods. Compared to traditional methods, RAG technology improves accuracy by reducing model illusions and enhancing answer credibility when combined with professional documents; it also enhances interpretability by providing the original document source along with the answer, facilitating user tracing; and it improves response speed by employing an efficient index structure and optimized model inference strategies to increase query efficiency. This system can significantly improve the efficiency of power plant equipment management, reduce manual review time, and provide maintenance personnel with faster and more reliable technical support.

[0046] The second objective of this invention is to propose a power plant equipment retrieval and question-answering system based on RAG, such as... Figure 3 As shown, it includes: Data processing module 100: used to acquire raw data of power plant equipment, preprocess the power plant equipment data, and obtain processed power plant equipment data; Data slicing module 200: Used to slice the processed power plant equipment data; the sliced ​​power plant equipment data. Database module 300: Used to extract semantic features from the sliced ​​power plant equipment data using natural language processing technology to form a vector database; Data Search Module 400: Used to search for power plant equipment data corresponding to a question in a vector database using the RAG question answering engine model; The answer acquisition module 500 is used to input the power plant equipment data corresponding to the question and the question into the generative model for processing, and obtain the answer to the question.

[0047] like Figure 4 As shown, a third objective of this invention is to provide an electronic device comprising a processor 601, a memory 602, and a display screen 603. The memory 602 and the display screen 603 are both connected to the processor 601, such as via a bus 604. Optionally, the electronic device may further include a transceiver 605. It should be noted that in practical applications, the transceiver 605 is not limited to one type, and the structure of this electronic device does not constitute a limitation on the embodiments of this application.

[0048] The processor 601 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute the various exemplary logical blocks, modules and circuits described in connection with the disclosure. The processor 601 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0049] The bus 604 can include a path for transmitting information between the above-mentioned components. The bus 604 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 604 can be divided into an address bus, a data bus, a control bus, etc.

[0050] The memory 602 can be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto.

[0051] The memory 602 is used to store application program code for implementing the scheme of the present application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the application program code stored in the memory 602 to realize the content shown in the foregoing method embodiments.

[0052] Figure 4 The electronic device shown is merely an example and should not bring any limitation to the function and use range of the embodiments of the present application.

[0053] A fourth object of the present application is to provide a computer readable storage medium storing a computer program, the computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the processes of the method embodiments as described above. Figure 1 The memory including instructions executable by the processor of the electronic device to perform the method described above.

[0054] The computer readable storage medium can be a tangible device that maintains and stores instructions for use by an instruction execution device. The computer readable storage medium can be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof. Specifically, the computer readable storage medium can be a portable computer diskette, a hard disk, a USB flash drive, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, an optical disk, a magnetic disk, a mechanical encoding device, and any combination thereof.

[0055] A fifth object of the present application is to provide a computer program product comprising computer instructions, which, when executed by a processor, implement the processes of the method embodiments as described above and achieve the same technical effects. To avoid repetition, they will not be described here. Figure 1

[0056] Many embodiments and many applications other than those described herein will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. Therefore, it is intended that the scope of the application be limited only by the broadest interpretation of the appended claims to be accorded under 35 U.S.C. § 112. It is intended and it is specifically contemplated that all of the articles, references, and courses of action discussed above can be employed as a part of the present application and are hereby incorporated by reference in their entirety. Any part of the subject matter disclosed herein, which is not specifically covered by the claims, is not abandoned and should not be considered as being outside the scope of the application.

[0057] The above is a further detailed description of the present application, which cannot be deemed as limiting the specific embodiments of the present application to the above. For those skilled in the art, some simple deductions or replacements can be made without departing from the concept of the present application, which should be deemed as falling within the protection scope of the present application as defined by the claims.​

Claims

1. A RAG-based power plant equipment retrieval and question answering method, characterized in that, The method comprises the following steps: Obtaining raw data of power plant equipment, preprocessing the data of the power plant equipment, and obtaining processed data of the power plant equipment; Slicing the processed data of the power plant equipment, and obtaining sliced data of the power plant equipment; Extracting semantic features from the sliced data of the power plant equipment using natural language processing technology, and forming a vector database; Using a RAG question and answer engine model to find the power plant equipment data corresponding to the question in the vector database; Inputting the power plant equipment data corresponding to the question and the question into a generation model for processing, and obtaining an answer corresponding to the question.

2. A RAG-based power plant equipment retrieval and question answering method according to claim 1, characterized in that, The method comprises the following steps: Obtaining raw data of power plant equipment, preprocessing the data of the power plant equipment, and obtaining processed data of the power plant equipment; Obtaining raw data of power plant equipment, including equipment manuals, maintenance manuals, DCS operation guides and maintenance records; 3. The RAG-based power plant equipment retrieval and question answering method of claim 1, wherein, Cleaning and formatting the raw data of the power plant equipment, and obtaining processed data of the power plant equipment. The method comprises the following steps: Determining the slicing granularity according to the requirements of the set target; Analyzing the semantic relationship in the data of the power plant equipment, and finding the slicing point; 4. The RAG-based power plant equipment retrieval and question answering method of claim 1, wherein, Based on the slicing granularity and the slicing point, the processed data of the power plant equipment is sliced to obtain the sliced data of the power plant equipment. The method comprises the following steps:

5. The RAG-based power plant equipment retrieval and question answering method of claim 1, wherein, Extracting semantic features from the sliced data of the power plant equipment using natural language processing technology, and converting the extracted semantic features into high-dimensional vectors to form a vector database. The method comprises the following steps: Converting the question text into a query vector; 6. The RAG-based power plant equipment retrieval and question answering method of claim 1, wherein, Using the RAG question and answer engine model to perform similarity matching in the vector database according to the query vector, and retrieving the power plant equipment data related to the question. The method comprises the following steps: Inputting the power plant equipment data corresponding to the question and the question into a generation model for processing, and generating an answer; 7. A RAG-based power plant equipment retrieval and question answering system, characterized in that, Using the RAG question and answer engine model to verify the generated answer with the raw data of the power plant equipment, and if the verification result meets the set requirements, obtaining the answer corresponding to the question. The method comprises the following steps: Data processing module: used for obtaining raw data of power plant equipment, preprocessing the data of the power plant equipment, and obtaining processed data of the power plant equipment; Data slicing module: used for slicing the processed data of the power plant equipment, and obtaining sliced data of the power plant equipment; Database module: used for extracting semantic features from the sliced data of the power plant equipment using natural language processing technology, and forming a vector database; Finding data module: used for using a RAG question and answer engine model to find the power plant equipment data corresponding to the question in the vector database; Answer acquisition module: used for inputting the power plant equipment data corresponding to the question and the question into a generation model for processing, and obtaining an answer corresponding to the question.

8. An electronic device, comprising: A computer readable storage medium storing a computer program, the computer program, when executed by a processor, implements the steps of the RAG-based power plant equipment search and question answering method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, A computer readable storage medium storing a computer program, the computer program, when executed by a processor, implements the steps of the RAG-based power plant equipment search and question answering method according to any one of claims 1-6.

10. A computer program product, characterised in that, Computer instructions, when executed by a processor, implement the steps of the RAG-based power plant equipment search and question answering method according to any one of claims 1-6.