Data construction and fine tuning method for knowledge retrieval model in energy power field
By constructing the search fine-tuning data in the energy and power field and using comparative learning and LoRA fine-tuning technology, the lack of knowledge and diversified input processing problems in the application of existing search models in the energy and power field are solved, and the search accuracy and semantic understanding ability are significantly improved, achieving a more cost-effective information acquisition process.
Patent Information
- Application Number
- CN202510252833.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing search models face many challenges in the application of energy and electricity fields, including lack of domain-specific knowledge, difficulty in dealing with diverse user inputs and professional terms, resulting in insufficient search accuracy and semantic understanding.
By constructing search fine-tuning data for the energy and power field, using a diverse expression method to generate problems - the document uses positive sample sets and negative sample sets, and combines comparison learning and LoRA parameter fine-tuning to fine-tune the search model to optimize the similarity between positive and negative sample pairs, and improve the semantic understanding ability and retrieval accuracy of the model.
It significantly improves the search accuracy and semantic understanding ability of the search model in the energy and power field, can have a deeper understanding of the specific search needs in the energy and power field, generates highly reliable search suggestions, and reduces resource consumption and application costs.
Smart Images

Figure CN120179781A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of using artificial intelligence for knowledge retrieval in the energy and power field, and particularly to a data construction and fine-tuning method for a knowledge retrieval model oriented to the energy and power field. Background Art
[0002] Currently, the energy and power industry is facing a sharp expansion of data scale and increasing complexity of information needs. How to efficiently retrieve and obtain valuable information from massive data has become an important challenge. In the operation, management, and equipment maintenance processes of the power system, timely access to accurate information is crucial for supporting decision-making and fault analysis. However, traditional information retrieval methods often struggle to meet the rapidly growing demand for diverse and refined information in the energy and power industry.
[0003] In recent years, the rapid development of deep learning has enabled large language models (LLMs) to demonstrate excellent capabilities in the pre-training of large-scale corpora, especially in natural language understanding and text generation. However, although LLMs perform well in text generation in general domains, when it comes to vertical domains (such as energy and power), the generated content may lack factual consistency and even introduce irrelevant or false information, resulting in the "hallucination" phenomenon. To address this issue, the Retrieval-Augmented Generation (RAG) method has been proposed. This method provides relevant facts and context information for text generation tasks by introducing external knowledge sources, thereby improving the accuracy and reliability of the generated content.
[0004] In a RAG system, the recall performance of the retrieval model is a key factor determining the overall effectiveness of the system. An efficient retrieval model can accurately locate relevant knowledge fragments from large-scale data sources and provide high-quality input for subsequent text generation. However, the application of existing retrieval models in the energy and power field still faces many challenges. First, the professional knowledge such as technical manuals, equipment maintenance guides, and operation specifications in the power industry is usually confidential or proprietary information and is difficult to obtain publicly, resulting in the lack of domain-specific knowledge in the training corpus of existing retrieval models. Second, in the energy and power field, the mapping relationship between professional terms and complex question-document pairs poses higher requirements for the semantic understanding ability of the model. In a production environment, users may express their questions in various ways, including precise queries using professional terms, colloquial expressions, or keyword-based retrievals. However, existing models often perform inadequately when dealing with these diverse expressions and are difficult to accurately capture the specific professional vocabulary and context in the energy and power industry.
[0005] For example, the invention application with the application number 202410717394.6 discloses a large model knowledge question-answering method and device for the power field, belonging to the technical field of natural language understanding question-answering. Using the application solution, the large language model can output answers based on the content of all sibling nodes of the knowledge points, which can improve the efficiency of document processing and information retrieval in the power field. However, the solution also has the following problems: lack of diverse expression methods, insufficient performance, and difficulty in accurately capturing the professional vocabulary and context unique to the energy and power industry.
[0006] Therefore, in order to improve the applicability of the retrieval augmented generation system in the energy and power field, it is urgent to enhance the understanding ability of the retrieval model for specific knowledge in the energy and power field through methods such as domain adaptation, knowledge enhancement, and professional training. There is an urgent need for a more intelligent and efficient technical means to optimize the information acquisition process to cope with the challenges in data management and knowledge application in the energy and power industry. Summary of the Invention
[0007] In view of the above problems, the purpose of the present invention is to provide a data construction and fine-tuning method for a knowledge retrieval model for the energy and power field, empower the knowledge of the retrieval augmented generation system, adopt diverse expression methods, construct retrieval fine-tuning data for the energy and power field, and provide an economical and efficient intelligent retrieval solution for knowledge in the energy and power field.
[0008] An embodiment of the present invention provides a data construction and fine-tuning method for a knowledge retrieval model for the energy and power field.
[0009] First aspect: A data construction and fine-tuning method for a knowledge retrieval model for the energy and power field, comprising:
[0010] S1. Preprocess the document data in the energy and power field, and segment the document data into document segments suitable for input to the retrieval model;
[0011] S2. Execute question generation based on the large language model, generate a positive sample set of question-document pairs according to the document segments, and sample to generate a negative sample set of question-document pairs;
[0012] S3. Combine the positive sample set and the negative sample set for contrast learning and LoRA parameter fine-tuning, fine-tune the retrieval model, optimize the similarity between the positive and negative sample pairs, and improve the semantic understanding ability and retrieval accuracy of the retrieval model.
[0013] Further, the S1 includes the steps of:
[0014] S11. Read the document data from the documents and technical specification data sources in the energy and power field, perform formatting processing and in-depth cleaning on the document data, and remove the noise data;
[0015] S12. For document data with excessive length, use a semantic-based text segmentation method to segment the document data into document fragments suitable for input to the retrieval model.
[0016] Further, the semantic-based text segmentation method adopted in S12 includes:
[0017] S12a. Use punctuation marks or delimiters to perform chunking on the document;
[0018] S12b. Evaluate the semantic similarity between adjacent chunks. If the similarity is lower than a preset threshold, merge the adjacent chunks to ensure the semantic integrity and coherence of each document fragment;
[0019] S12c. Control the length of each document fragment not to exceed the maximum input length limit of the retrieval model to meet the model processing requirements.
[0020] Further, for the segmented document fragments, generate questions to construct positive samples of question-document pairs. The positive sample set of question-document pairs includes standardized questions, colloquial questions, and keyword generation. The question generation template includes role description, task requirements, generation rules, and input-output formats, where:
[0021] Standardized questions cover the core information of the text, meet the professional requirements of the energy and power field, and use industry terms and standard expressions;
[0022] Colloquial questions understand the user's daily language habits, adopt a natural and direct colloquial style, and maintain the accurate conveyance of the key information of the text;
[0023] Keyword generation captures the core information and themes of the text, understands the overall intention of the text, and extracts keyword information.
[0024] Further, the negative sample set includes random negative sample sampling, BM25-based negative sample sampling, and vector similarity-based negative sample sampling, where:
[0025] Random negative sample sampling means excluding the positive samples of the current query from the full-text paragraph library, and then randomly selecting K1 fragments as negative samples;
[0026] BM25-based negative sample sampling means using the bm25s library to query according to the current question using the BM25 algorithm, performing recall of word-level similarity on the text, obtaining the top N most relevant fragments, excluding the top 5 most relevant fragments, and randomly selecting K2 from the remaining fragments as negative samples;
[0027] Negative sample sampling based on vector similarity is to use a pre-trained vector representation model to encode all document fragments in the candidate text paragraph library to generate a high-dimensional vector space; encode the query question text to obtain a query vector; use an efficient vector retrieval tool to build an index, and perform a nearest neighbor search on the query vector in the vector space to retrieve the top N document fragments that are semantically most relevant to the query vector; exclude the top 5 fragments that are most relevant to the query vector from the retrieval results, and randomly select K3 from the remaining fragments as negative samples.
[0028] Furthermore, the LoRA parameter fine-tuning includes:
[0029] For the weight matrix W ∈ R of the pre-trained retrieval model m×n , reparameterize it using the parameter update matrix ΔW to obtain the updated weight matrix W′, which is expressed by the formula:
[0030] W′ = W + ΔW = W + AB
[0031] where A ∈ R m×r and B ∈ R r×n are two low-rank matrices obtained by decomposing the parameter update matrix ΔW, where
[0032]
[0033] Furthermore, the contrastive learning by combining the positive sample set and the negative sample set includes:
[0034] S31. Use the pre-trained retrieval model to encode the query question and the positive and negative sample documents into high-dimensional vectors respectively, which is expressed by the formula:
[0035] v q = f(q), v d = f(d)
[0036] where q is the query question, d is the positive and negative sample documents, and v d is the positive and negative sample document vector, and f(·) is the encoding function of the retrieval model;
[0037] S32. Calculate the similarity between the query vector and the positive and negative sample document vectors, which is expressed by the formula:
[0038]
[0039] where v q is the query vector, and v d is the positive and negative sample document vector;
[0040] S33. Calculate the contrastive learning loss function to make the similarity of the positive sample pairs higher than that of the negative sample pairs, which is expressed by the formula:
[0041]
[0042] Among them, represents the positive sample document vector, represents the negative sample document vector, τ represents the temperature hyperparameter, and K represents the number of all negative samples corresponding to a single question.
[0043] Second aspect: An electronic device, including a memory, a processor, and a computer program stored on the memory and operable on the processor. When the processor executes the program, it implements the steps of the method provided in the first aspect.
[0044] Third aspect: A non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method provided in the first aspect.
[0045] Advantages of the present invention:
[0046] 1. By generating high-quality positive samples of question-document pairs and challenging negative samples, and combining contrastive learning techniques, the present invention optimizes the retrieval model, enabling it to deeply master the professional knowledge in the field of energy and power. Thus, it significantly improves the retrieval accuracy of the vector model in this field. Using the method of the present invention, it is possible to deeply understand the specific retrieval needs in the field of energy and power, generate highly reliable retrieval suggestions that conform to the habits of this field, and thus promote the wide application of the retrieval-augmented generation system (RAG) in the field of energy and power.
[0047] 2. Through domain enhancement, the present invention realizes knowledge empowerment for the retrieval-augmented generation system, significantly improving the application effect of the retrieval model in the field of energy and power; not only strengthening the professional knowledge level of the model in the field, but also being able to intelligently retrieve decision-making suggestions that meet the field requirements. At the same time, it also has the ability to process diverse user inputs, is closer to user usage habits, greatly improves the user experience and satisfaction, and demonstrates excellent robustness. In addition, by adopting the LoRA method of parameter-efficient fine-tuning, it significantly reduces the resource requirements during the fine-tuning process, thereby reducing the resource consumption and application costs of the model in actual application scenarios, and bringing a more economical and efficient solution for users. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic flow chart of the data construction and fine-tuning method of the knowledge retrieval model for the energy and power field of the present invention;
[0049] Figure 2 is a schematic diagram of the standardized question generation template of the present invention;
[0050] Figure 3 is a schematic diagram of the colloquial question generation template of the present invention;
[0051] Figure 4 Schematic diagram of the keyword generation template for the present invention;
[0052] Figure 5 Flowchart of the negative sample sampling process for the present invention;
[0053] Figure 6 Architecture diagram of the fine-tuning of the retrieval model based on contrastive learning and LoRA for the present invention;
[0054] Figure 7 Schematic diagram of the structure of the data construction and fine-tuning device of the knowledge retrieval model for the energy and power field of the present invention;
[0055] Figure 8 Schematic diagram of the structure of the electronic device of the present invention. Detailed implementation manners
[0056] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0057] In order to improve the performance of the retrieval model in the energy and power field, this example proposes a data construction and fine-tuning method for the knowledge retrieval model in the energy and power field. Utilizing the powerful generation ability of large language models (LLMs), by constructing various types of questions and diverse negative samples on domain documents, combined with contrastive learning and parameter-efficient fine-tuning technology (LoRA), the retrieval model is fine-tuned to improve the retrieval accuracy and generalization ability of the model in the energy and power field.
[0058] The specific implementation manners of this embodiment will be given below. Figure 1 Flowchart of the data construction and fine-tuning method for the knowledge retrieval model in the energy and power field provided by the embodiment of the present invention, including the following steps:
[0059] S1. Preprocess the document data in the energy and power field, and segment the document data into document fragments suitable for input to the retrieval model.
[0060] First, read the document data from data sources such as documents and technical specifications in the energy and power field, and perform formatting processing to ensure data consistency. Subsequently, deeply clean the documents to remove noise data such as garbled characters, irrelevant symbols, and HTML tags, thereby ensuring the purity and high quality of the document data. For documents with too large a length, a semantic-based text segmentation method is used to segment the documents into fragments suitable for processing by the retrieval model.
[0061] Specifically, it includes: First, the document is chunked using punctuation marks or delimiters; then, the semantic similarity between adjacent chunks is evaluated. If the similarity is lower than a preset threshold, the adjacent chunks are merged to ensure the semantic integrity and coherence of each document fragment; finally, the length of each document fragment is controlled not to exceed the maximum input length limit of the retrieval model (e.g., 512), so as to meet the requirements of model processing.
[0062] S2. Perform question generation based on large language models, generate a positive sample set of question-document pairs according to the document fragments, and sample to generate a negative sample set.
[0063] Question generation based on large language models (LLMs) mainly generates corresponding questions for the list of preprocessed energy and power document fragments. The generation process includes three types of questions: standardized question generation, colloquial question generation, and keyword query generation, to handle different types of user inputs in the scenarios of energy and power and equipment maintenance. Different question generation steps are controlled by designing different large language model text formatting templates.
[0064] In the selection of large language models, API interfaces provided by existing research institutions or companies can be selected, such as Deepseek, ChatGPT, Qwen, etc., or local deployment of large models, such as vLLM, can be used for question generation requests. This embodiment uses the locally deployed open-source Qwen2.5-32B model to generate different types of questions and generate question-document pairs in a standardized format.
[0065] First is standardized question generation. Common questions are extracted from energy and power industry documents, technical specifications, and reports, and standardized queries are generated through large language models.
[0066] The generated questions are as detailed, standardized, and easy to understand as possible. The question generation template for standardized questions consists of four parts: role description, task requirements, generation rules, and input-output format. The detailed question generation template is as Figure 2 shown.
[0067] The main task requirements description for standardized questions is: The generated questions should comprehensively cover the core information of the document fragments; the questions should preferably meet the professional requirements of the energy and power field, using industry terms and standard expressions; handle text containing complex structures of data or HTML tags.
[0068] Then is colloquial question generation. Colloquial questions mainly generate colloquial expressions through the LLM to simulate the input of users under normal circumstances. The question generation template for colloquial questions consists of four parts: role description, task requirements, generation rules, and input-output format. The detailed question generation template is as Figure 3 shown.
[0069] The main task requirements of colloquial questions are described as follows: understand the user's daily language habits, and ensure that the generated questions are presented in a natural and direct colloquial style; colloquial questions should maintain the accurate conveyance of key information in the text and avoid distortion.
[0070] Then there is keyword generation. Keyword generation mainly summarizes the input text paragraph fragments and outputs summary statements or keywords that can express the key information of the fragments. It is used to simulate the keyword type queries of user input. The question generation template for keyword generation consists of four parts: role description, task requirements, generation rules, and input-output format. The complete template for keyword generation is as Figure 4 shown.
[0071] The main task requirements of keyword generation are described as follows: be able to quickly capture the core information and theme, and ensure the representativeness and retrieval function of the keywords. Understand the overall intention of the text and extract words or phrases with high efficient retrieval value from it. Have the ability to refine information to ensure that the keywords or phrases are highly recognizable and meaningful independently of the specific context.
[0072] Merge the questions and document fragments generated by the three methods of standardized questions, colloquial questions, and keyword generation to obtain the positive sample set for model training.
[0073] The negative sample sampling method proposed in this example aims to generate a challenging negative sample set for the documents in the energy and power field through various sampling strategies, so as to improve the model's ability to distinguish positive and negative samples, thereby enhancing the generalization performance and robustness of the retrieval model in complex energy and power scenarios.
[0074] As Figure 5 shown, the negative sample set mainly includes three parts of sample sampling, including: random negative sample sampling, BM25-based negative sample sampling, and vector similarity-based negative sample sampling.
[0075] Random negative sample sampling is to randomly select fragments from the full-text paragraph library in the energy and power field as negative samples, aiming to provide the model with diverse irrelevant samples and provide negative sample signals for the fine-tuning of the retrieval model.
[0076] Specifically, it includes first excluding the positive samples of the current query from the full-text paragraph library, and then randomly selecting K1 fragments as negative samples.
[0077] BM25-based negative sample sampling is a commonly used text similarity calculation method, aiming to select negative sample documents with a certain degree of relevance at the lexical level, increase sample diversity, thereby reducing the model's dependence on simple patterns, and helping the model to effectively distinguish positive and negative samples in the real semantic features of the energy and power field.
[0078] Specifically, it includes: First, use the BM25 algorithm to query according to the current question. Here, the bm25s library is adopted to recall the similarity at the word level of the text and obtain the top N most relevant segments. After that, exclude the top 5 most relevant segments to avoid selecting potential positive samples, and randomly select K2 from the remaining segments as negative samples. This strategy retains subtle semantic similarities at low relevance levels, helping the model to perform refined learning.
[0079] Negative sample sampling based on vector similarity uses a pre-trained vector representation model to calculate the vector similarity of text paragraphs. This method measures semantic similarity through the distance in a high-dimensional vector space and selects negative samples that are close but not the most relevant in the semantic space to increase the difficulty of negative samples.
[0080] Specifically, it includes: First, use a pre-trained vector representation model to encode all document segments in the candidate text paragraph library to generate high-dimensional vector representations; at the same time, also encode the query question text to obtain its corresponding query vector representation; then, use an efficient vector retrieval tool (such as Faiss) to build an index and perform a nearest neighbor search for the query vector in the vector space to retrieve the top N document segments that are most semantically relevant to the query vector.
[0081] To ensure the diversity and challenge of negative samples, this embodiment adopts a hierarchical sampling strategy: exclude the top 5 segments most relevant to the query from the retrieval results (to avoid misselecting positive samples), and then randomly select K3 from the remaining segments as negative samples. The negative samples generated by this strategy are semantically somewhat relevant to the query but not exactly the same, which can help the model learn more fine-grained semantic discrimination capabilities.
[0082] Combine the negative sample set generated by the sampling method with the positive sample set for the training of the model in the energy and power field.
[0083] S3. Combine the positive sample set and the negative sample set for contrastive learning and LoRA parameter fine-tuning, train the pre-trained retrieval model, and use the trained retrieval model in the energy and power field to find knowledge segments related to the question for question retrieval.
[0084] The training of the retrieval model in this embodiment, as Figure 6 shown, mainly conducts efficient fine-tuning of the retrieval model training on the energy and power field data by combining contrastive learning and LoRA technology.
[0085] The efficient fine-tuning of LoRA parameters aims to achieve efficient fine-tuning of the pre-trained retrieval model through low-rank matrix factorization, significantly reducing the consumption of computing resources while maintaining the performance of the model in the energy and power field.
[0086] Specifically, LoRA introduces low-rank matrix decomposition technology based on model fine-tuning. For the weight matrix W∈R m×n , reparameterize the parameter update matrix ΔW and decompose it into two low-rank matrices A∈R m×r and B∈R r×n ,in The updated weight matrix is expressed as follows:
[0087] W′=W+ΔW=W+AB
[0088] During the fine-tuning process, most of the parameters of the retrieval model are frozen and only the parameters of the low-rank matrices A and B are updated. This strategy significantly reduces the number of parameters that need to be trained, thereby reducing the consumption of computing resources.
[0089] Contrastive learning refers to improving the model's ability to understand the semantics of texts in the energy and power sector by optimizing the similarity between positive and negative sample pairs. For example, it can more accurately identify key information such as power equipment fault descriptions and power grid dispatch instructions.
[0090] The goal of contrastive learning is to optimize the model, shorten the semantic feature distance between the query vector features and the positive samples, and increase the distance with the negative sample features, so that it can better distinguish between positive and negative samples, thereby improving the semantic understanding ability and retrieval accuracy of the retrieval model in the field of energy and power.
[0091] Specifically, first, the query and document are encoded into high-dimensional vector representations using the pre-trained retrieval model; for the input query q and document d, their vector representations are obtained through the retrieval model f(·):
[0092] v q =f(q),v d =f(d)
[0093] After that, the similarity between the query vector and the positive sample document vector, as well as the similarity between the query vector and the negative sample document, is calculated; the similarity is calculated using cosine similarity, and the formula is expressed as:
[0094]
[0095] Finally, the model is optimized by contrastive loss function so that the similarity of positive sample pairs is higher than that of negative sample pairs. The contrastive learning loss function is defined as follows:
[0096]
[0097] in, represents the positive sample document vector, Denote the negative sample document vector, τ represents the temperature hyperparameter, and K represents the number of all negative samples corresponding to a single question.
[0098] The embodiment of the present application provides a method for constructing fine-tuning data for a retrieval model in a vertical domain, including two core parts: diversified positive sample generation of document-question pairs and diversified collection of negative samples of question-document pairs.
[0099] Diversified positive sample generation of question-document pairs. For the input list of vertical domain documents, relevant questions for the corresponding document segments are generated through three methods: standardized question generation, colloquial question generation, and keyword generation.
[0100] Diversified negative sample collection of question-document pairs. First, through random negative sample document sampling, negative samples are randomly selected from all document paragraph lists. Second, based on BM25 negative sample sampling, the top N samples with the highest similarity are recalled using the BM25 algorithm, and the top 5 samples with the highest similarity are excluded from them to avoid selecting potential positive samples, and sampling is performed in the remaining set, and documents with high word matching similarity are selected as negative samples. Finally, based on vector similarity for negative sample sampling, both the question and the document are represented as embedding vectors, and the top N samples with the highest similarity are retrieved through the vector similarity algorithm, and samples with similarity between 5-N are selected for sampling, and documents with high semantic matching similarity are selected as negative samples.
[0101] Through multi-channel negative sample sampling, a diversified negative sample set is generated to enhance the model's ability to recognize different semantic differences.
[0102] Finally, through contrastive learning technology, during the training process, the retrieval query information and the document are converted into vector form, maximizing the similarity between positive samples while minimizing the similarity between positive and negative samples, thereby optimizing the retrieval effect. Using LoRA technology, the parameters of the retrieval model are efficiently fine-tuned to optimize the resource usage efficiency and improve the retrieval performance of the model in the energy and power field.
[0103] Through the above method, the trained retrieval model obtained can significantly improve the model's performance on complex data in the energy and power field, and at the same time has high robustness and good scalability, providing strong support for tasks such as power system analysis and equipment maintenance.
[0104] The present invention provides a device for data construction and fine-tuning of a knowledge retrieval model in the energy and power field, as Figure 7 shown. Based on the retrieval model, the device includes:
[0105] A sample acquisition module for acquiring a positive sample set and a negative sample set of question-document pairs;
[0106] The optimized parameter module uses the positive sample set and the negative sample set, combines contrastive learning technology and LoRA technology to train the retrieval model, and optimizes the parameters of the retrieval model.
[0107] This device trains the retrieval model through the sample acquisition module and the optimized parameter module to obtain the retrieval model in the trained Retrieval-Augmented Generation (RAG) system, in response to the challenges of difficult access to professional knowledge in the energy and power fields and diverse query requirements.
[0108] The present invention also provides an electronic device. Figure 8 As shown in the structural schematic diagram of the electronic device provided by the embodiment of the present invention, Figure 8 the electronic device may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communications interface, and the memory communicate with each other through the communication bus. The processor can call the logical instructions in the memory, for example, to execute the following methods:
[0109] S1. Preprocess the document data in the energy and power fields, and segment the document data into document fragments suitable for input to the retrieval model;
[0110] S2. Perform question generation based on the large language model, generate a positive sample set of question-document pairs according to the document fragments, and sample to generate a negative sample set of question-document pairs;
[0111] S3. Combine contrastive learning of the positive sample set and the negative sample set and LoRA parameter fine-tuning to fine-tune the retrieval model, optimize the similarity between positive and negative sample pairs, and improve the semantic understanding ability and retrieval accuracy of the retrieval model.
[0112] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.
[0113] An embodiment of the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above embodiments. For example, it includes:
[0114] S1. Preprocess the document data in the energy and power field, and segment the document data into document fragments suitable for input to the retrieval model.
[0115] S2. Perform question generation based on the large language model, generate a positive sample set of question-document pairs according to the document fragments, and sample to generate a negative sample set of question-document pairs.
[0116] S3. Fine-tune the retrieval model by combining contrastive learning of the positive and negative sample sets and LoRA parameter fine-tuning, optimize the similarity between the positive and negative sample pairs, and improve the semantic understanding ability and retrieval accuracy of the retrieval model.
[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data construction and fine-tuning method for knowledge retrieval model in the field of energy and power, characterized in that: include: S1. Preprocess the document data in the energy and power field and divide the document data into document segments suitable for retrieval model input; S2. Execute question generation based on the large language model, generate a positive sample set of question-document pairs based on document fragments, and sample and generate a negative sample set of question-document pairs; S3. Combine the comparative learning of positive sample sets and negative sample sets and LoRA parameter fine-tuning to fine-tune the retrieval model, optimize the similarity between positive and negative sample pairs, and improve the semantic understanding ability and retrieval accuracy of the retrieval model.
2. According to claim 1, a data construction and fine-tuning method for a knowledge retrieval model in the field of energy and power, characterized in that: The S1 comprises the steps of: S11. Read document data from documents and technical specification data sources in the energy and power field, format and deeply clean the document data, and remove noise data; S12. For document data that is too long, a semantic-based text segmentation method is used to segment the document data into document segments that are suitable for retrieval model input.
3. The data construction and fine-tuning method for the knowledge retrieval model in the energy and power field according to claim 2 is characterized in that: The text segmentation method based on semantics is adopted in S12, including: S12a, dividing the document into blocks using punctuation marks or separators; S12b, evaluating the semantic similarity between adjacent blocks, and if the similarity is lower than a preset threshold, merging the adjacent blocks to ensure the semantic integrity and coherence of each document fragment; S12c. Control the length of each document fragment so that it does not exceed the maximum input length limit of the retrieval model to meet the model processing requirements.
4. The data construction and fine-tuning method for the knowledge retrieval model in the energy and power field according to claim 1 is characterized in that: The question-document matching sample set includes standardized questions, colloquial questions and keyword generation. The question generation template includes role description, task requirements, generation rules and input and output formats, where: Standardized questions cover the core information of the text, meet the professional requirements of the energy and power sector, and use industry terminology and standard expressions; Colloquial questions understand the user's daily language habits, adopt a natural and direct colloquial style, and maintain accurate communication of key text information; Keyword generation captures the core information and themes of the text, understands the overall intent of the text, and extracts keyword information.
5. The data construction and fine-tuning method for the knowledge retrieval model in the energy and power field according to claim 1 is characterized in that: The negative sample set includes random negative sample sampling, BM25-based negative sample sampling, and vector similarity-based negative sample sampling, where: Random negative sample sampling is to exclude the positive sample of the current query from the full text paragraph library, and then randomly select K1 fragments as negative samples; The negative sample sampling based on BM25 is to use the BM25 algorithm to query the current question using the bm25s library, recall the text at the word level similarity, obtain the top N most relevant fragments, exclude the top 5 most relevant fragments, and randomly select K2 from the remaining fragments as negative samples; The negative sample sampling based on vector similarity is to use the pre-trained vector representation model to encode all document fragments in the candidate text paragraph library to generate a high-dimensional vector space; encode the query question text to obtain the query vector; use an efficient vector retrieval tool to build an index, perform a neighbor search on the query vector in the vector space, and retrieve the top N document fragments that are most semantically relevant to the query vector; exclude the top 5 fragments most relevant to the query vector from the retrieval results, and randomly select K3 from the remaining fragments as negative samples.
6. The data construction and fine-tuning method for the knowledge retrieval model in the energy and power field according to claim 1 is characterized in that: The LoRA parameter fine-tuning includes: The weight matrix W∈R of the pre-trained retrieval model m×n , use the parameter update matrix ΔW to reparameterize and obtain the updated weight matrix W′, which is expressed as: W′=W+ΔW=W+AB Among them, A∈R m×r and B∈R r×n are two low-rank matrices decomposed into the parameter update matrix ΔW, where 7. The data construction and fine-tuning method for the knowledge retrieval model in the energy and power field according to claim 1 is characterized in that: The S3 combines the positive sample set and the negative sample set for comparative learning, including: S31. Use the pre-trained retrieval model to encode the query question and positive and negative sample documents into high-dimensional vectors respectively. The formula is expressed as: v q =f(q),v d =f(d) Among them, q is the query question, d is the positive and negative sample documents, and v d is the document vector of positive and negative samples, and f(·) is the encoding function of the retrieval model; S32. Calculate the similarity between the query vector and the positive and negative sample document vectors. The formula is expressed as: Among them, v q is the query vector, v d is the positive and negative sample document vector; S33, calculate the contrastive learning loss function so that the similarity of the positive sample pair is higher than that of the negative sample pair. The formula is expressed as: in, represents the positive sample document vector, represents the negative sample document vector, τ represents the temperature hyperparameter, and K represents the number of all negative samples corresponding to a single question.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of a data construction and fine-tuning method for a knowledge retrieval model in the energy and power field as described in any one of claims 1 to 7 are implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a data construction and fine-tuning method for a knowledge retrieval model in the energy and power field as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Medical text semantic recognition method, medical text semantic recognition device, electronic equipment and readable storage medium
CN111128394A
Multi-task large model fine tuning method based on adapters and low-rank adaptation
CN116822611A
Domain retrieval method based on meta-learning and knowledge enhancement
CN117609419A
Large model knowledge question-answering method and device for power field
CN118585626A
Cited By
Electric power knowledge retrieval system based on large language model
CN120910228A
Retrieval enhancement method and system for multi-round dialogue type questions and answers and application
CN121029952A
BGE model fine tuning method and system based on vector retrieval precision enhancement
CN121212333A
A BGE model fine-tuning method and system based on vector retrieval precision enhancement
CN121212333B