Optimization and evaluation method and device of retrieval enhancement generation system, equipment and medium

By constructing vectorized model training data, optimizing model parameters and configuration prompt words, combining supervision fine-tuning and reward feedback training, the generation model is optimized and the system performance is evaluated, which solves the shortcomings of the search-enhanced generation system in generation, retrieval and evaluation, and improves the system effect.

CN120387512APending Publication Date: 2025-07-29MINSHENG BANKING CORP
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510319049.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing search enhancement generation system has shortcomings in generation, search and evaluation, and it is difficult to comprehensively improve its effectiveness.

Method used

By constructing vectorized model training data, optimizing model parameters, configuring prompt words, optimizing the generative model with supervised fine-tuning and reward feedback training, and evaluating system performance in combination with the Q&A evaluation index set.

Benefits of technology

Overall, the accuracy of the search and enhancement generation system has been improved, from 85% to 89.2%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387512A_ABST
    Figure CN120387512A_ABST
Patent Text Reader

Abstract

The invention provides an optimization method and device for a retrieval enhancement generation system, equipment and a medium. Comprising the following steps: constructing model training data according to a document extracted from a knowledge base and text fragments obtained by segmenting the document; performing model parameter optimization on the vectorization model based on the model training data; processing the user question and the text fragment based on a vectorization model to obtain a retrieval fragment; according to the retrieval fragment, the user question and the session context, configuring a cue word; calling a generation model to process the user question and the cue word, and generating a question answer text; a supervised fine tuning and reward feedback training mode is adopted, and a generation model is optimized according to a question and answer data set formed by a standard answer text, a user question and a text retrieved according to the user question; and determining a target evaluation index according to the question answer text, the manual annotation data of the user question and the question and answer evaluation index set, so as to evaluate the performance of the retrieval enhancement generation system according to the target evaluation index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to an optimization and evaluation method, device, equipment and medium for a retrieval-augmented generation system. Background Art

[0002] RAG (Retrieval Augmented Generation): Retrieval Augmented Generation is a technical method that combines retrieval and generation, aiming to enhance the generation ability of large models by retrieving relevant information from external knowledge bases. The working principle of RAG consists of two main steps: data retrieval and generation. First, the system recalls relevant knowledge according to the user's input. Then, the retrieved information is injected into the prompt and provided to the large model as context to help it generate more accurate and comprehensive answers. In the actual enterprise scenarios, the accuracy rate of the RAG system often becomes a bottleneck for popularization. Existing solutions often focus on the optimization of models in a certain link of the RAG system, such as evaluation, retrieval, etc. For the entire RAG system, it is difficult to cover all dimensions. Summary of the Invention

[0003] The technical problem to be solved by the embodiments of this application is to provide an optimization and evaluation method, device, equipment and medium for a retrieval-augmented generation system, so as to optimize the RAG system from three dimensions of generation, retrieval and evaluation, and improve the effect of the RAG system.

[0004] In a first aspect, the embodiments of this application provide an optimization and evaluation method for a retrieval-augmented generation system, and the method includes:

[0005] Construct model training data for the vectorization model of the retrieval-augmented generation system according to the documents extracted from the knowledge base and the text fragments obtained by segmenting the documents;

[0006] Optimize the model parameters of the vectorization model based on the model training data;

[0007] Process the user question and the text fragments based on the vectorization model to obtain retrieval fragments;

[0008] Configure prompt words according to the retrieval fragments, the user question and the session context corresponding to the user question;

[0009] Call the generation model of the retrieval-augmented generation system to process the user question and the prompt words, and generate a question answer text corresponding to the user question;

[0010] Optimize the generation model by means of supervised fine-tuning and reward feedback training, according to the standard answer text corresponding to the user question, the user question, and the Q&A dataset formed by the text retrieved according to the user question.

[0011] According to the question answer text, the manually annotated data corresponding to the user question, and the pre-constructed Q&A evaluation index set, determine the target evaluation index in the Q&A evaluation index set, so as to evaluate the performance of the retrieval-augmented generation system according to the target evaluation index.

[0012] In a second aspect, an embodiment of the present application provides an optimization and evaluation device for a retrieval-augmented generation system, and the device includes:

[0013] A model training data construction module, configured to construct model training data for the vectorization model of the retrieval-augmented generation system according to the documents extracted from the knowledge base and the text segments obtained by segmenting the documents.

[0014] A model parameter optimization module, configured to optimize the model parameters of the vectorization model based on the model training data.

[0015] A retrieval segment acquisition module, configured to process the user question and the text segment based on the vectorization model to obtain retrieval segments.

[0016] A prompt configuration module, configured to configure prompts according to the retrieval segments, the user question, and the conversation context corresponding to the user question.

[0017] An answer text generation module, configured to call the generation model of the retrieval-augmented generation system to process the user question and the prompts to generate the question answer text corresponding to the user question.

[0018] A generation model optimization module, configured to optimize the generation model by means of supervised fine-tuning and reward feedback training, according to the standard answer text corresponding to the user question, the user question, and the Q&A dataset formed by the text retrieved according to the user question.

[0019] An evaluation module, configured to determine the target evaluation index in the Q&A evaluation index set according to the question answer text, the manually annotated data corresponding to the user question, and the pre-constructed Q&A evaluation index set, so as to evaluate the performance of the retrieval-augmented generation system according to the target evaluation index.

[0020] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0021] A processor, a memory, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the optimization and evaluation method of the retrieval augmented generation system described in any one of the above.

[0022] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the optimization and evaluation method of the retrieval augmented generation system described in any one of the above.

[0023] Compared with the prior art, the embodiments of the present application include the following advantages:

[0024] In the embodiments of the present application, model training data for a vectorization model of a retrieval augmented generation system is constructed based on documents extracted from a knowledge base and text fragments obtained by segmenting the documents. The model parameters of the vectorization model are optimized based on the model training data. The user question and the text fragments are processed based on the vectorization model to obtain retrieval fragments. Prompt words are configured according to the retrieval fragments, the user question, and the conversation context corresponding to the user question. The generation model of the retrieval augmented generation system is called to process the user question and the prompt words to generate a question answer text corresponding to the user question. In a manner of supervised fine-tuning and reward feedback training, the generation model is optimized according to a question answer text corresponding to the user question, a question-answer dataset formed by the user question and the text retrieved according to the user question, and a standard answer text. According to the question answer text, the manually annotated data corresponding to the user question, and a pre-constructed question-answer evaluation index set, a target evaluation index in the question-answer evaluation index set is determined to evaluate the performance of the retrieval augmented generation system according to the target evaluation index. In the embodiments of the present application, the RAG system is optimized from three dimensions: generation (i.e., optimization of the generation model), retrieval (segmentation of documents and optimization of the vectorization model), and evaluation (i.e., determination of evaluation indexes and automatic performance evaluation of the RAG system), which can overall improve the effect of the RAG system.

[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flowchart of steps of an optimization and evaluation method for a retrieval augmented generation system provided by an embodiment of the present application;

[0027] Figure 2 It is a flowchart of a RAG system provided by an embodiment of the present application;

[0028] Figure 3 It is a schematic structural diagram of an optimization and evaluation device for a retrieval augmented generation system provided by an embodiment of the present application;

[0029] Figure 4 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific implementation manners

[0030] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0031] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0032] Referring to Figure 1 , a flowchart of steps of an optimization and evaluation method for a retrieval augmented generation system provided by an embodiment of the present application is shown. As Figure 1 shown, the optimization and evaluation method for the retrieval augmented generation system may include: Step 101, Step 102, Step 103, Step 104, Step 105, Step 106, and Step 107.

[0033] Step 101: Construct model training data for the vectorization model of the retrieval augmented generation system according to the documents extracted from the knowledge base and the text fragments obtained by segmenting the documents.

[0034] In this embodiment, when optimizing the RAG (retrieval augmented generation system), documents can be extracted from the knowledge base, and these documents can be any form of text data, such as articles, reports, book chapters, etc.

[0035] After extracting the documents from the knowledge base, the documents can be segmented to obtain text fragments. These fragments can be sentences, paragraphs, or more complex structures divided according to specific rules (such as keyword occurrences, topic transitions, etc.). In this embodiment, the segmentation effect of each text block can be improved by means of recursive segmentation, model segmentation, etc.

[0036] Furthermore, model training data for the vectorization model of RAG can be constructed according to the extracted documents and the text fragments obtained by segmenting the documents. The construction process of the model training data will be described in detail in combination with the following specific implementation manners.

[0037] In a specific implementation of the present application, the above Step 101 may include:

[0038] Sub-step A1: Process the documents extracted from the knowledge base through a question-and-answer pair generation model to obtain a plurality of initial question-and-answer pairs.

[0039] In this embodiment, when training the vectorization model of RAG, the documents extracted from the knowledge base can be processed by a question-and-answer pair generation model to obtain multiple initial question-and-answer pairs. Specifically, this process may include:

[0040] 1. First, extract the documents from the knowledge base for which question-and-answer pairs are desired, and preprocess the extracted documents, such as removing unnecessary formats, standardizing the text, etc.

[0041] 2. Use a large model to generate initial question-and-answer pairs. 1) Select a large model: Select a large language model suitable for generating question-and-answer pairs, such as the Qwen series (e.g., Qwen-2.5, etc.), the GPT series (e.g., GPT-3, GPT-4), etc. 2) Input the document: Input the preprocessed document into the large model. 3) Generate question-and-answer pairs: Utilize the generation ability of the large model to generate multiple potential question-and-answer pairs for each document. This usually involves asking an open-ended question to the model (such as "What questions and answers can you come up with regarding this document?"), and then having the model output a series of initial question-and-answer pairs.

[0042] Of course, in practical applications, other methods can also be used to generate question-and-answer pairs. For example, use the model to pose some questions based on the document content, and then traverse these questions and use the model to generate corresponding answers to obtain question-and-answer pairs, etc.

[0043] After processing the documents extracted from the knowledge base by the question-and-answer pair generation model to obtain multiple initial question-and-answer pairs, perform sub-step A2.

[0044] Sub-step A2: Obtain an available question-and-answer pair dataset after manual annotation and modification of the multiple initial question-and-answer pairs.

[0045] After processing the documents extracted from the knowledge base by the question-and-answer pair generation model to obtain multiple initial question-and-answer pairs, an available question-and-answer pair dataset after manual annotation and modification of the multiple initial question-and-answer pairs can be obtained. Specifically, the generated question-and-answer pairs can be initially screened to remove those that are clearly illogical, have grammar errors, or are insignificant. Then, the remaining question-and-answer pairs can be carefully annotated, which may include: verifying the accuracy of the answers: ensuring that the answers correctly answer the questions, evaluating the quality of the questions: ensuring that the questions are clear, specific, and closely related to the document content, annotating additional information: such as the difficulty level of the questions, the type of answers (e.g., factual, explanatory, inferential, etc.). Finally, modification and optimization can be carried out: Based on the annotation results, modify and optimize the question-and-answer pairs, which may include adjusting the wording of the questions, improving the details of the answers, adding necessary context, etc.

[0046] Sub-step A3: Slice the document to obtain multiple document slices.

[0047] After extracting the document from the knowledge base, the extracted document can be sliced to obtain multiple document slices.

[0048] After slicing the document to obtain multiple document slices, execute sub-step A4.

[0049] Sub-step A4: Process the multiple document slices based on a preset method to obtain the retrieval segments corresponding to the question-and-answer pairs in the available question-and-answer pair dataset.

[0050] After slicing the document to obtain multiple document slices, the multiple document slices can be processed based on a preset method to obtain the retrieval segments corresponding to the question-and-answer pairs in the available question-and-answer pair dataset. Specifically, each extracted document can be sliced, and the segments corresponding to the questions can be obtained through 3 methods (W-RAG, HyDE, BGE), and the k (k positive integers) segments with the highest scores are respectively selected from the segments of the document to which the question belongs and the segments of the document set to which the document belongs. The document to which the question belongs refers to the source document of the question, and the document set contains multiple documents in the same field, including the document to which the question belongs.

[0051] After processing the multiple document slices based on a preset method to obtain the retrieval segments corresponding to the question-and-answer pairs in the available question-and-answer pair dataset, execute sub-step A5.

[0052] Sub-step A5: Randomly select a set number of question-and-answer pair samples from the available question-and-answer pair dataset and the retrieval segments, and obtain the relevance of the manually annotated question-and-answer pair samples to obtain test data.

[0053] After processing the multiple document slices based on a preset method to obtain the retrieval segments corresponding to the question-and-answer pairs in the available question-and-answer pair dataset, a set number of question-and-answer pair samples can be randomly selected from the available question-and-answer pair dataset and the retrieval segments. Specifically, a certain number of question-and-answer pairs and retrieval segments can be randomly selected from the available question-and-answer pair dataset and the retrieval segments, and manually annotated to determine whether the retrieved segments are helpful for answering the questions, thereby obtaining test data.

[0054] After randomly selecting a set number of question-and-answer pair samples from the available question-and-answer pair dataset and the retrieval segments, execute sub-step A6.

[0055] Sub-step A6: Construct the model training data of the vectorization model according to the extracted question-and-answer pair samples and the relevance of the question-and-answer pair samples predicted by the model.

[0056] After randomly extracting a set number of Q&A pair samples from the available Q&A pair dataset and retrieval fragments, the model training data for the vectorization model can be constructed based on the extracted Q&A pair samples and the relevance of the Q&A pair samples predicted by the large model. Among them, the model training data can include: a set proportion of positive samples, hard negative samples, and negative samples. Specifically, the extracted Q&A pairs and retrieval fragments can be analyzed to determine whether the retrieval fragments are helpful for answering questions. If the large model determines that a certain retrieval fragment is helpful, it is used as a positive sample, and the rest are used as negative samples. At the same time, hard negative samples and ordinary negative samples are distinguished according to the retrieval scores. And the test data is excluded from it.

[0057] After constructing the model training data for the vectorization model of the retrieval-augmented generation system according to the documents extracted from the knowledge base and the text fragments obtained by segmenting the documents, step 102 is executed.

[0058] Step 102: Optimize the model parameters of the vectorization model based on the model training data.

[0059] After constructing the model training data, the model parameters of the vectorization model can be optimized based on the model training data. Specifically, the vectorization model can be trained based on the model training data to obtain the trained vectorization model. The trained vectorization model is tested based on the test data to obtain the recall rate of the trained vectorization model. The recall rate can reflect the model's ability to identify positive examples. A high recall rate means that the model can identify more positive examples, but it may also be accompanied by a relatively high false positive rate (i.e., the proportion of misidentifying negative examples as positive examples). Therefore, in practical applications, the recall rate is usually considered together with the precision to comprehensively evaluate the performance of the model.

[0060] The vectorization model training process can include:

[0061] 1. Select a vectorization model: Through public leaderboards, select a vectorization model with better performance, such as BGE-M3, BGE-Large.

[0062] 2. Adjust the proportions of positive samples, hard negative samples, and ordinary negative samples as training data.

[0063] 3. Adjust the training parameters for training.

[0064] Model evaluation: Use the trained model to obtain retrieval fragments for test questions, and compare with the manually annotated results, then the recall rate can be calculated to judge the model effect before and after training.

[0065] After optimizing the model parameters of the vectorization model based on the model training data, step 103 is executed.

[0066] Step 103: Process the user question and the text fragment based on the vectorization model to obtain a retrieval fragment.

[0067] After optimizing the model parameters of the vectorization model based on the model training data, the user question and the text fragment can be processed based on the vectorization model to obtain a retrieval fragment. The implementation process can be described in detail in combination with the following specific implementation manners.

[0068] In a specific implementation of the present application, the above step 103 may include:

[0069] Sub-step B1: Invoke the vectorization model to process the text fragment to obtain a document vectorization index.

[0070] In this embodiment, after obtaining the text fragment, the vectorization model can be invoked to process the text fragment to obtain a document vectorization index. Specifically, it may include: 1. Text preprocessing. Preprocess the text fragment, including steps such as word segmentation, stop word removal, and stemming, to simplify the text and reduce noise. 2. Vectorization model: Use a vectorization model such as BGE to convert the preprocessed text into high-dimensional vectors. These vectors can capture the semantic information in the text. 3. Index construction: Store the obtained document vectors into an index structure for fast retrieval. Commonly used index structures include inverted index, hash table, etc. As Figure 2 shown, for the knowledge base document, document slicing can be performed first, and the sliced text can be input into the vectorization model for processing to obtain a document vectorization index.

[0071] Sub-step B2: Invoke the vectorization model to process the user question to obtain a user question vector.

[0072] After completing the training of the vectorization model, the vectorization model can be invoked to process the user question input by the user to obtain a user question vector. Specifically, it may include: 1. Question preprocessing: Perform similar preprocessing steps on the user question to ensure consistency with the processing method of the text fragment. 2. Vectorization: Use the same vectorization model as the text fragment to convert the preprocessed user question into a vector. As Figure 2 shown, the user question can be processed by the vectorization model to obtain a user question vector.

[0073] Sub-step B3: Retrieve a target text fragment associated with the user question from a preset database according to the user question vector and the document vectorization index.

[0074] After obtaining the user question vector and the document vectorized index, the target text fragment associated with the user question can be retrieved from the preset database based on the user question vector and the document vectorized index. Specifically, the similarity between the user question vector and each document vector (such as cosine similarity, Euclidean distance, etc.) can be calculated. Then, the documents can be sorted according to the similarity scores, and several with the highest scores can be selected as the target text fragments. A threshold can be set in this step to filter out the documents with lower similarities.

[0075] After retrieving the target text fragment associated with the user question from the preset database based on the user question vector and the document vectorized index, sub-step B4 is executed.

[0076] Sub-step B4: Call the re-ranking model of the retrieval-augmented generation system to perform re-ranking processing on the target text fragment to obtain the retrieved fragment.

[0077] After retrieving the target text fragment associated with the user question from the preset database based on the user question vector and the document vectorized index, the re-ranking model of the retrieval-augmented generation system can be called to perform re-ranking processing on the target text fragment to obtain the retrieved fragment. In this example, the re-ranking model is usually a deep learning model that takes the initially retrieved text fragment and the user question as inputs and outputs a more accurate ranking result. This model can further consider factors such as the semantic matching degree and context relevance between the text fragment and the user question. After the re-ranking processing, multiple text fragments with the highest scores are selected as the final retrieval results. This result usually answers the user's question more accurately. As Figure 2 shown, the document vectorized index and the user question vector after being processed by the two vectorization models can retrieve relevant fragments and then perform re-ranking on the relevant fragments.

[0078] In this embodiment, the vectorization processing flow of RAG includes the vectorization processing of text fragments, the vectorization processing of user questions, the retrieval of target text fragments, and the re-ranking processing. Each step aims to improve the accuracy and efficiency of the retrieval, thereby providing a better Q&A experience for users. By combining deep learning models and advanced indexing techniques, the RAG model can quickly find the content most relevant to the user question in a large amount of text data.

[0079] After obtaining the retrieved fragment by processing the user question and the text fragment based on the vectorization model, step 104 is executed.

[0080] Step 104: Configure the prompt words according to the retrieved fragment, the user question, and the session context corresponding to the user question.

[0081] After obtaining the retrieval fragments by processing the user's question and text fragments based on the vectorized model, prompt words can be configured according to the retrieval fragments, the user's question, and the conversation context corresponding to the user's question, such as Figure 2 As shown, configure the prompt words by combining the user's question, the user's historical questions (i.e., the conversation context), and the result of reordering relevant fragments. Specifically, the user's question can be analyzed, that is, the question intention can be clarified: clarify what problem the user wants to understand or solve, identify the key information in the user's question, such as time, place, person, etc., and judge what type the user's question belongs to, such as factual, definitional, suggestive, etc. Then, the conversation context can be obtained: historical questions: review the user's previous questions to understand the background and continuity of the questions; change in user intention: analyze whether the user's intention has changed and how this change affects the understanding and answer of the current question. Finally, prompt words can be configured, including: combining the retrieval fragments and the user's question: combine the retrieval fragments and the user's question according to the analysis result of the user's question to form preliminary prompt words; integrating the conversation context: integrate the keywords and information in the conversation context into the prompt words to ensure that the prompt words can comprehensively reflect the user's intention and conversation background; optimizing the prompt words: refine and optimize the prompt words to ensure that they are concise, clear, and easy to understand.

[0082] After configuring the prompt words according to the retrieval fragments, the user's question, and the conversation context corresponding to the user's question, perform step 105.

[0083] Step 105: Call the generation model of the retrieval-augmented generation system to process the user's question and the prompt words, and generate the question answer text corresponding to the user's question.

[0084] After configuring the prompt words according to the retrieval fragments, the user's question, and the conversation context corresponding to the user's question, the generation model of the retrieval-augmented generation system can be called to process the user's question and the prompt words, and generate the question answer text corresponding to the user's question. Specifically, a trained generation model (such as GPT series, Qwen series, etc.) can be used to generate the answer text. This model generates possible answers based on the encoded input representation. Decode the output of the generation model into the text form of natural language, which involves post-processing the generated sequence, such as removing redundancy and adjusting grammar.

[0085] After calling the generation model of the retrieval-augmented generation system to process the user's question and the prompt words, and generating the question answer text corresponding to the user's question, perform step 106.

[0086] Step 106: Optimize the generation model in the way of supervised fine-tuning and reward feedback training according to the standard answer text corresponding to the user's question, the question-answer dataset formed by the user's question and the text retrieved according to the user's question.

[0087] After the generation model of the retrieval-augmented generation system is called to process the user question and the prompt words to generate the question answer text corresponding to the user question, the generation model can be optimized by means of supervised fine-tuning and reward feedback training. Specifically, according to the question-answer data set formed by the question answer text and the user question, the parameters of the generation model can be adjusted by means of supervised fine-tuning, that is, the preprocessed question-answer data set is used to fine-tune the generation model. The model parameters are adjusted by minimizing the difference between the predicted answer and the true answer (such as cross-entropy loss). The reward feedback training mechanism is adopted, and according to the quality of the answer generated by the generation model and the user feedback information, the generation model is iteratively optimized through a reinforcement learning algorithm so that the generation model generates an answer that meets the user's expectations.

[0088] In practical applications, a reward function can be designed to evaluate the quality of the generated answer. The reward function can be based on multiple factors, such as the relevance, accuracy, fluency, diversity, etc. of the answer. In the initial stage, automatic evaluation tools (such as BLEU, ROUGE, etc.) can be used to preliminarily evaluate the generated answer. As the system runs, user feedback is gradually introduced as a reference for rewards. Users can provide feedback by means of scoring, liking, commenting, etc. According to the reward function and user feedback, the reward value of each generated answer is calculated.

[0089] The reinforcement learning iterative optimization process can be as follows: 1. Policy gradient method: The generation model is optimized by using the policy gradient method (such as the REINFORCE algorithm, etc.). In each iteration, the model parameters are adjusted according to the reward value, so that the probability of generating a higher-quality answer increases. 2. Value function estimation: In order to improve the optimization efficiency, a value function (such as the Q function or V function, etc.) can be used to estimate the long-term reward of each action (that is, the generated word or phrase). This can be achieved by methods such as the deep Q network (DQN) or the actor-critic. 3. Iterative optimization: Through multiple iterations, the parameters of the generation model are continuously adjusted until the model performance converges or reaches the predetermined number of iterations.

[0090] In this embodiment, when optimizing the generation model, supervised fine-tuning data and reward feedback training data are used to optimize the generation model. The implementation process can be described in detail in combination with the following specific implementation manners.

[0091] In a specific implementation of this application, the above step 106 may include:

[0092] Step C1: Screen out multiple types of data from the question-answer data set and the general data set, and the multiple types of data include: supervised fine-tuning data and reward feedback training data.

[0093] In this embodiment, when optimizing the generation model, various types of data related to the retrieval-augmented generation system can be obtained from the question-answering dataset and the general dataset, such as Figure 2 As shown, the results returned by the large model can be obtained for collecting various types of data. The various types of data can include: supervised fine-tuning data and reward feedback training data. Among them, the supervised fine-tuning data is sourced from existing high-quality question-answer pairs, knowledge bases, or manually annotated data, which contains clear questions and corresponding answers, and is used to fine-tune the generation model to accurately understand and answer questions. The reward feedback training data can be obtained by simulating user behavior, historical user feedback, or manual annotation, and contains questions, generated answers, and corresponding reward values (or evaluations), and is used to train the model to learn to adjust the generation strategy according to the reward signal.

[0094] After collecting various types of data related to the retrieval-augmented generation system, step C2 is executed.

[0095] Step C2: Preprocess the collected various types of data to obtain preprocessed data.

[0096] After collecting various types of data related to the retrieval-augmented generation system, the collected various types of data can be preprocessed to obtain preprocessed data. Specifically, the preprocessing methods can be: data cleaning (i.e., removing duplicate, invalid, or low-quality data), formatting (unifying the data format, such as text encoding, the structure of question-answer pairs, etc.), annotation (manually or automatically annotating the data that needs to be annotated, such as annotating the reward value for the reward feedback data), tokenization and vectorization (performing tokenization on the text data and converting it into a vector representation for easy model processing).

[0097] After preprocessing the collected various types of data to obtain preprocessed data, step C3 is executed.

[0098] Step C3: Integrate the preprocessed data into a training dataset to optimize the generation model.

[0099] After preprocessing various types of collected data to obtain preprocessed data, the preprocessed data can be integrated into a training dataset to optimize the generation model. Specifically, a data integration strategy can be formulated according to the training objective and data characteristics. For example, general supervised fine-tuning data and question-and-answer data can be mixed in a certain proportion to balance the accuracy and generalization ability of the model. The integrated data is divided into a training set, a validation set, and a test set. The training set is used for model training, the validation set is used for model selection and parameter adjustment, and the test set is used for evaluating model performance. The generalization ability of the model is further improved by data augmentation techniques (such as synonym replacement, sentence restructuring, etc.).

[0100] Through the above process, the embodiments of the present application can collect and integrate various types of data, providing rich and high-quality training resources for the training of the retrieval-augmented generation system. This will help improve the performance of the model, enabling it to more accurately understand user questions and generate answers that meet expectations.

[0101] After completing the optimization of the generation model, step 107 is executed.

[0102] Step 107: Determine the target evaluation metric in the question-and-answer evaluation metric set according to the question-answer text, the manually annotated data corresponding to the user question, and the pre-constructed question-and-answer evaluation metric set, so as to evaluate the performance of the retrieval-augmented generation system according to the target evaluation metric.

[0103] After completing the optimization of the generation model, the target evaluation metric in the question-and-answer evaluation metric set can be determined according to the question-answer text, the manually annotated data corresponding to the user question, and the pre-constructed question-and-answer evaluation metric set, so as to evaluate the performance of the retrieval-augmented generation system according to the target evaluation metric. Specifically, evaluation data can be collected, and a question-and-answer evaluation metric set can be constructed, where the question-and-answer evaluation metric set includes: integrity metric, hallucination metric, loyalty metric, and answer relevance metric. The retrieval-augmented generation system processes the evaluation data to obtain a dataset to be evaluated. Obtain the manually annotated results corresponding to the dataset to be evaluated, and the manually annotated results include: questions, standard answers, generated answers, and annotation labels. Use an automated evaluation system to evaluate the data to be evaluated to obtain an automated evaluation result. Compare the automated evaluation result with the manually annotated result, and when the comparison result indicates that the trends of the automated evaluation result and the manually annotated result are the same, determine the target evaluation metric in the question-and-answer evaluation metric set.

[0104] In this embodiment, the evaluation part of the RAG system may include:

[0105] 1. Build an automated evaluation system: To accelerate the iteration of the link, aiming at the problem that manual evaluation is time-consuming and laborious, an automated evaluation system for the integrity (Completeness) and hallucination (Hallucination) of large models is constructed, which is more in line with the trend of manual evaluation.

[0106] 2. Build a scenario Q&A evaluation set to initially form an evaluation system for the Q&A scenario.

[0107] The work of the evaluation part includes three parts: evaluation data construction, automated evaluation system construction, and evaluation system assessment.

[0108] (1) Evaluation data construction.

[0109] The collected scenario data is divided into 3 cases:

[0110] a. For this scenario, there are already knowledge base documents and standard Q&A pairs, and this kind of data can be directly used as evaluation data.

[0111] b. For this scenario, after trial use, there are already enough questions, but most of them have no standard answers. Such data needs to generate answers through a large model and then be modified by annotators to obtain the correct answers, which can then be used as evaluation data.

[0112] c. For this scenario, there are only a small number of questions. At this time, the questions will be expanded through a large model, and then answers will be generated and the questions and answers will be modified and annotated manually as evaluation data.

[0113] (2) Automated evaluation system construction: Design evaluation methods based on cutting-edge work and experience, such as:

[0114] a. Completeness: Measure the coverage of the generated answer to the key points of the standard answer. Split the standard answer into key points through a large model, and then use the large model to judge whether the generated answer contains these key points and calculate the proportion.

[0115] b. Hallucination: The contradiction ratio between the generated answer and the key points of the standard answer. Use the large model to judge whether the generated answer conflicts with the key points of the standard answer and calculate the proportion.

[0116] c. Faithfulness: Measure whether the generated answer is faithful to the provided text fragment. Decompose the generated answer into multiple key point statements through a large model, and then use the large model to judge whether the provided text fragment can support this statement and calculate the proportion.

[0117] d. Answer relevance: Measure whether the generated answer solves the problem. According to the generated answer, use the large model to generate multiple questions, and calculate the similarity between the original question and the generated multiple questions through a vectorization model to obtain the average value.

[0118] (3) Evaluation system assessment.

[0119] a. For the problems in (1), use the RAG system to generate answers to obtain the data to be evaluated.

[0120] b. Manual annotation: According to the data to be evaluated, provide questions, standard answers, generated answers, and the annotators conduct annotation. The labels for annotation are: correct, wrong, unrecognized, should be unrecognized but not, unrecognized wrongly.

[0121] c. Obtain the automated evaluation result: Evaluate the data to be evaluated through an automated evaluation method to obtain a score.

[0122] d. Compare the automated evaluation result and the manual annotation result, and judge whether the trends of the automated evaluation result and the manual annotation result are the same. If they are the same, it can be retained as an automated evaluation index (i.e., the target evaluation index).

[0123] In the embodiment of the present application, by optimizing the RAG system from the perspectives of retrieval, generation, and evaluation, the overall accuracy rate of the RAG can be increased from 85% to 89.2%.

[0124] The optimization and evaluation method of the retrieval-enhanced generation system provided by the embodiment of the present application constructs the model training data of the vectorization model of the retrieval-enhanced generation system according to the documents extracted from the knowledge base and the text segments obtained by segmenting the documents. Optimize the model parameters of the vectorization model based on the model training data. Process the user questions and text segments based on the vectorization model to obtain retrieval segments. Configure prompt words according to the retrieval segments, user questions, and the conversation context corresponding to the user questions. Call the generation model of the retrieval-enhanced generation system to process the user questions and prompt words to generate the question answer text corresponding to the user questions. Adopt the method of supervised fine-tuning and reward feedback training to optimize the generation model according to the standard answer text corresponding to the user questions, the question-answer dataset formed by the user questions and the texts retrieved according to the user questions. Determine the target evaluation index in the question-answer evaluation index set according to the question answer text, the manual annotation data corresponding to the user questions, and the pre-constructed question-answer evaluation index set, so as to evaluate the performance of the retrieval-enhanced generation system according to the target evaluation index. The embodiment of the present application optimizes the RAG system from three dimensions: generation (i.e., optimization of the generation model), retrieval (segmentation of documents and optimization of the vectorization model), and evaluation (i.e., determination of evaluation indexes and automated performance evaluation of the RAG system), and can overall improve the effect of the RAG system.

[0125] Refer to Figure 3 , which shows the structural schematic diagram of an optimization and evaluation device for a retrieval-enhanced generation system provided by the embodiment of the present application, as Figure 3As shown in the figure, the optimization and evaluation device 300 of the retrieval-augmented generation system may include the following modules:

[0126] The model training data construction module 310 is used to construct the model training data of the vectorization model of the retrieval-augmented generation system according to the documents extracted from the knowledge base and the text fragments obtained by segmenting the documents;

[0127] The model parameter optimization module 320 is used to optimize the model parameters of the vectorization model based on the model training data;

[0128] The retrieval fragment acquisition module 330 is used to process the user question and the text fragment based on the vectorization model to obtain retrieval fragments;

[0129] The prompt configuration module 340 is used to configure prompts according to the retrieval fragments, the user question, and the session context corresponding to the user question;

[0130] The answer text generation module 350 is used to call the generation model of the retrieval-augmented generation system to process the user question and the prompt to generate the question answer text corresponding to the user question;

[0131] The generation model optimization module 360 is used to optimize the generation model by means of supervised fine-tuning and reward feedback training according to the standard answer text corresponding to the user question, the user question, and the question-answer data set formed by the text retrieved according to the user question;

[0132] The evaluation module 370 is used to determine the target evaluation index in the question-answer evaluation index set according to the question answer text, the manually annotated data corresponding to the user question, and the pre-constructed question-answer evaluation index set, so as to evaluate the performance of the retrieval-augmented generation system according to the target evaluation index.

[0133] Optionally, the model training data construction module includes:

[0134] The initial question-answer pair acquisition unit is used to process the documents extracted from the knowledge base through a question-answer pair generation model to obtain a plurality of initial question-answer pairs;

[0135] The available question-answer pair acquisition unit is used to obtain an available question-answer pair data set obtained by manually annotating and modifying the plurality of initial question-answer pairs;

[0136] The document slice acquisition unit is used to slice the document to obtain a plurality of document slices;

[0137] A retrieval fragment acquisition unit, configured to process the multiple document slices based on a preset method to obtain retrieval fragments corresponding to the question-and-answer pairs in the available question-and-answer pair dataset;

[0138] A question-and-answer pair sample extraction unit, configured to randomly extract a set number of question-and-answer pair samples from the available question-and-answer pair dataset and the retrieval fragments;

[0139] A model training data construction unit, configured to construct model training data for the vectorization model according to the extracted question-and-answer pair samples;

[0140] Wherein, the model training data includes: a set proportion of positive samples, hard negative samples, and negative samples.

[0141] Optionally, the model training data includes: training data and test data,

[0142] The model parameter optimization module includes:

[0143] A vectorization model training unit, configured to perform model training on the vectorization model based on the model training data to obtain a trained vectorization model;

[0144] A recall rate acquisition unit, configured to test the trained vectorization model based on the test data to obtain the recall rate of the trained vectorization model;

[0145] A model parameter optimization unit, configured to optimize the model parameters of the trained vectorization model according to the recall rate.

[0146] Optionally, the retrieval fragment acquisition module includes:

[0147] A vectorized index acquisition unit, configured to call the vectorization model to process the text fragment to obtain a document vectorized index;

[0148] A question vector acquisition unit, configured to call the vectorization model to process the user question to obtain a user question vector;

[0149] A target text fragment acquisition unit, configured to retrieve a target text fragment associated with the user question from a preset database according to the user question vector and the document vectorized index;

[0150] A retrieval fragment acquisition unit, configured to call the re-ranking model of the retrieval enhancement generation system to perform re-ranking processing on the target text fragment to obtain the retrieval fragment.

[0151] Optionally, the generation model optimization module includes:

[0152] A generation model adjustment unit, configured to adjust the parameters of the generation model in a supervised fine-tuning manner according to the question-and-answer data set;

[0153] A model optimization unit, configured to iteratively optimize the generation model by using a reward feedback training mechanism and a reinforcement learning algorithm according to the quality of the answers generated by the generation model and the user feedback information, so that the generation model generates answers that meet the user's expectations.

[0154] Optionally, the generation model optimization module includes:

[0155] A data acquisition unit, configured to screen out multiple types of data from the question-and-answer data set and the general data set, where the multiple types of data include: supervised fine-tuning data and reward feedback training data;

[0156] A preprocessing data acquisition unit, configured to preprocess the multiple types of data collected to obtain preprocessed data;

[0157] A generation model optimization unit, configured to integrate the preprocessed data into a training data set to optimize the generation model.

[0158] Optionally, the evaluation module includes:

[0159] An evaluation data collection unit, configured to collect evaluation data;

[0160] An index set construction unit, configured to construct a question-and-answer evaluation index set, where the question-and-answer evaluation index set includes: integrity index, hallucination index, loyalty index, and answer relevance index;

[0161] A data set acquisition unit, configured to process the evaluation data based on the retrieval-augmented generation system to obtain a data set to be evaluated;

[0162] A labeled result acquisition unit, configured to obtain the manual labeling result corresponding to the data set to be evaluated, where the manual labeling result includes: questions, standard answers, generated answers, and labeling tags;

[0163] An evaluation result acquisition unit, configured to evaluate the data set to be evaluated by using an automated evaluation system to obtain an automated evaluation result;

[0164] A target evaluation index determination unit, configured to compare the automated evaluation result with the manual labeling result, and determine the target evaluation index in the question-and-answer evaluation index set when the comparison result indicates that the trends of the automated evaluation result and the manual labeling result are the same.

[0165] The optimization and evaluation device of the retrieval-augmented generation system provided by the embodiments of the present application constructs the model training data of the vectorization model of the retrieval-augmented generation system according to the documents extracted from the knowledge base and the text fragments obtained by segmenting the documents. Optimize the model parameters of the vectorization model based on the model training data. Process the user question and the text fragments based on the vectorization model to obtain retrieval fragments. Configure prompt words according to the retrieval fragments, the user question, and the session context corresponding to the user question. Call the generation model of the retrieval-augmented generation system to process the user question and the prompt words to generate the question answer text corresponding to the user question. Adopt the method of supervised fine-tuning and reward feedback training to optimize the generation model according to the standard answer text corresponding to the user question, the user question, and the Q&A dataset formed by the text retrieved according to the user question. Determine the target evaluation metrics in the Q&A evaluation metric set according to the question answer text, the manually annotated data corresponding to the user question, and the pre-constructed Q&A evaluation metric set, so as to evaluate the performance of the retrieval-augmented generation system according to the target evaluation metrics. The embodiments of the present application optimize the RAG system from three dimensions: generation (i.e., optimization of the generation model), retrieval (segmentation of documents and optimization of the vectorization model), and evaluation (i.e., determination of evaluation metrics and automatic performance evaluation of the RAG system), which can overall improve the effect of the RAG system.

[0166] The embodiments of the present application also provide an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the above-mentioned optimization and evaluation method of the retrieval-augmented generation system is implemented.

[0167] Figure 4 Fig. shows a schematic structural diagram of an electronic device 400 according to an embodiment of the present invention. As Figure 4 shown, the electronic device 400 includes a central processing unit (CPU) 401, which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) 402 or the computer program instructions loaded from the storage unit 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0168] Multiple components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, a microphone, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0169] Each of the processes and treatments described above may be executed by the processing unit 401. For example, the method of any of the above embodiments may be implemented as a computer software program, which is tangibly contained in a computer-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the CPU 401, one or more actions in the method described above may be executed.

[0170] Additionally, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the optimization and evaluation method of the above-mentioned retrieval enhanced generation system is implemented.

[0171] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other.

[0172] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0173] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminals (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, may be implemented by computer program instructions. These computer program instructions may be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminals to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminals generate for implementing the processesFigure 1 one or more processes and / or blocks Figure 1 means for the functions specified in one or more blocks

[0174] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions in the process Figure 1 one or more processes and / or blocks Figure 1 specified in one or more blocks

[0175] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal, such that a series of operational steps are performed on the computer or other programmable terminal to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal provide for implementing the functions in the process Figure 1 one or more processes and / or blocks Figure 1 steps of the functions specified in one or more blocks

[0176] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present application

[0177] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or terminal including the element

[0178] The above has introduced in detail an optimization and evaluation method for a retrieval-augmented generation system, an optimization and evaluation device for a retrieval-augmented generation system, an electronic device, and a computer-readable storage medium. In this article, specific examples are used to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. An optimization method for a retrieval-augmented generation system, characterized in that, The method includes: Constructing model training data for the vectorization model of the retrieval-augmented generation system based on the documents extracted from the knowledge base and the text fragments obtained by segmenting the documents; Optimizing the model parameters of the vectorization model based on the model training data; Processing the user question and the text fragments based on the vectorization model to obtain retrieval fragments; Configuring prompt words according to the retrieval fragments, the user question, and the session context corresponding to the user question; Invoking the generation model of the retrieval-augmented generation system to process the user question and the prompt words to generate a question answer text corresponding to the user question; Optimizing the generation model by means of supervised fine-tuning and reward feedback training according to the standard answer text corresponding to the user question, the user question, and the question-answer dataset formed by the text retrieved according to the user question; Determining the target evaluation metric in the question-answer evaluation metric set according to the question answer text, the manually annotated data corresponding to the user question, and the pre-constructed question-answer evaluation metric set, so as to evaluate the performance of the retrieval-augmented generation system according to the target evaluation metric.

2. The method according to claim 1, characterized in that, The constructing model training data for the vectorization model of the retrieval-augmented generation system based on the documents extracted from the knowledge base and the text fragments obtained by segmenting the documents includes: Processing the documents extracted from the knowledge base through a question-answer pair generation model to obtain a plurality of initial question-answer pairs; Obtaining an available question-answer pair dataset after manual annotation and modification of the plurality of initial question-answer pairs; Performing slicing processing on the documents to obtain a plurality of document slices; Processing the plurality of document slices based on a preset method to obtain retrieval fragments corresponding to the question-answer pairs in the available question-answer pair dataset; Randomly extracting a set number of question-answer pair samples from the available question-answer pair dataset and the retrieval fragments, and obtaining the relevance of the manually annotated question-answer pair samples to obtain test data; Constructing model training data for the vectorization model according to the extracted question-answer pair samples and the relevance of the question-answer pair samples predicted by the model; Among them, the model training data includes: positive samples, hard negative samples, and negative samples in a set proportion.

3. The method according to claim 1, wherein The model training data includes: training data and test data, The optimizing the model parameters of the vectorization model based on the model training data includes: Performing model training on the vectorization model based on the model training data to obtain a trained vectorization model; Testing the trained vectorization model based on the test data to obtain the recall rate of the trained vectorization model; Optimizing the model parameters of the trained vectorization model according to the recall rate.

4. The method according to claim 1, characterized in that, The processing the user question and the text fragments based on the vectorization model to obtain retrieval fragments includes: Invoking the vectorization model to process the text fragments to obtain a document vectorization index; Invoking the vectorization model to process the user question to obtain a user question vector; Retrieve a target text segment associated with the user question from a preset database according to the user question vector and the document vectorized index; Call the re-ranking model of the retrieval-augmented generation system to re-rank the target text segment to obtain the retrieved segment.

5. The method according to claim 1, wherein The method of optimizing the generation model by means of supervised fine-tuning and reward feedback training, according to the standard answer text corresponding to the user question, the user question, and the question-answer dataset formed by the text retrieved according to the user question, includes: Adjust the parameters of the generation model by means of supervised fine-tuning according to the question-answer dataset; Adopt a reward feedback training mechanism and a reinforcement learning algorithm to iteratively optimize the generation model according to the quality of the answer generated by the generation model and the user feedback information, so that the generation model generates an answer that meets the user's expectations.

6. The method according to claim 1, characterized in that The method of optimizing the generation model by means of supervised fine-tuning and reward feedback training, according to the standard answer text corresponding to the user question, the user question, and the question-answer dataset formed by the text retrieved according to the user question, includes: Screen out various types of data from the question-answer dataset and the general dataset, and the various types of data include: supervised fine-tuning data and reward feedback training data; Preprocess the collected various types of data to obtain preprocessed data; Integrate the preprocessed data into a training dataset to optimize the generation model.

7. The method according to claim 1, characterized in that The method of determining the target evaluation metric in the question-answer evaluation metric set according to the question-answer text, the manually annotated data corresponding to the user question, and the pre-constructed question-answer evaluation metric set includes: Collect evaluation data; Construct a question-answer evaluation metric set, and the question-answer evaluation metric set includes: integrity metric, hallucination metric, loyalty metric, and answer relevance metric; Process the evaluation data based on the retrieval-augmented generation system to obtain a dataset to be evaluated; Obtain the manually annotated result corresponding to the dataset to be evaluated, and the manually annotated result includes: question, standard answer, generated answer, and annotation label; Evaluate the dataset to be evaluated by means of an automated evaluation system to obtain an automated evaluation result; Compare the automated evaluation result with the manually annotated result, and determine the target evaluation metric in the question-answer evaluation metric set when the comparison result indicates that the trends of the automated evaluation result and the manually annotated result are the same.

8. An optimization device for a retrieval-augmented generation system, characterized in that, The device includes: A model training data construction module, configured to construct model training data for the vectorization model of the retrieval-augmented generation system according to the documents extracted from the knowledge base and the text segments obtained by splitting the documents; A model parameter optimization module, configured to optimize the model parameters of the vectorization model based on the model training data; A retrieved segment acquisition module, configured to process the user question and the text segment based on the vectorization model to obtain a retrieved segment; A prompt word configuration module, configured to configure prompt words according to the retrieved segment, the user question, and the session context corresponding to the user question; An answer text generation module, configured to call the generation model of the retrieval-augmented generation system to process the user question and the prompt words, and generate an answer text corresponding to the user question; A generation model optimization module, configured to optimize the generation model by means of supervised fine-tuning and reward feedback training according to the standard answer text corresponding to the user question, the user question, and a question-answer dataset formed by the text retrieved according to the user question; An evaluation module, configured to determine a target evaluation metric in the question-answer evaluation metric set according to the answer text, the manually annotated data corresponding to the user question, and a pre-constructed question-answer evaluation metric set, so as to evaluate the performance of the retrieval-augmented generation system according to the target evaluation metric.

9. An electronic device, characterized in that, including: A processor, a memory, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the optimization and evaluation method of the retrieval-augmented generation system according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the optimization and evaluation method of the retrieval-augmented generation system according to any one of claims 1 to 7.

Citation Information

Cited By

  • Data set determination method and device, medium and product

    CN120705281A

  • A dataset determination method, apparatus, medium, and product

    CN120705281B

  • Answer generation method, device and equipment for multi-hop question

    CN120832957A

  • Evaluation result determination method and device, storage medium and electronic equipment

    CN120873150A

  • Insurance clause knowledge base self-maintenance method, system and equipment based on large model

    CN120975207A