Retrieval enhancement generation method and device, equipment and storage medium
Through the method of filtering and adjusting the number of paragraphs, combined with semantic segmentation model and language model evaluation, the search enhancement generation technology is optimized, the noise and missing search problems are solved, and the accuracy and completeness of the answers are improved.
Patent Information
- Application Number
- CN202510246477.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-03
AI Technical Summary
There are noise interference and missing search problems in the existing search enhancement generation technology, resulting in a decrease in answer accuracy and limiting its performance in practical applications.
By filtering the target paragraphs whose similarity score changes in the database, and using the second largest language model to evaluate and adjust the number of paragraphs, combining the semantic segmentation model to segment the document into semantic complete paragraphs, constructing a vector representation database, and optimizing the search process.
提高了检索的准确性和生成质量,解决了噪声检索和缺失检索的问题,提升了生成答案的完整性和准确性。
Smart Images

Figure CN120296141A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computer technology, and in particular, to a retrieval-augmented generation method, apparatus, device, and storage medium. Background Art
[0002] The Retrieval-Augmented Generation (RAG) technology significantly improves the performance of question-answering systems by integrating two key processes: information retrieval and text generation. The core technical path of this method is that the system first retrieves document information semantically relevant to the user's query problem from a large-scale corpus, and then inputs these retrieval results as context into a pre-trained language model to generate accurate and information-rich answers.
[0003] However, the retrieval-augmented generation technology still faces several key technical challenges in practical applications. The most prominent one is the noise interference problem in the retrieval stage. Specifically, when the system retrieves relevant documents, it inevitably introduces redundant information unrelated to the query intention. The existence of this noise information may mislead the reasoning process of the language model, resulting in a decrease in the accuracy of the generated answers. On the other hand, when the retrieval system fails to obtain sufficient key context information, that is, the missing retrieval phenomenon occurs, this will directly limit the generation ability of the language model and make it unable to provide complete and accurate answers. These technical bottlenecks severely restrict the performance of the retrieval-augmented generation method in actual application scenarios. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide a retrieval-augmented generation method, apparatus, device, and storage medium to improve retrieval accuracy and thus enhance the generation quality.
[0005] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:
[0006] According to a first aspect of one or more embodiments of this specification, a retrieval-augmented generation method is proposed, including:
[0007] Retrieving in a database according to a question raised by a user to obtain a plurality of candidate paragraphs;
[0008] Determining a target paragraph from the plurality of candidate paragraphs whose change rate of similarity score with the question meets a set requirement;
[0009] Inputting the target paragraph and the question into a first large language model to generate a first answer;
[0010] Input the target paragraph, the question, and the first answer into a second large language model so that the second large language model outputs an evaluation result, where the evaluation result includes a paragraph evaluation result indicating that the number of target paragraphs is excessive or insufficient;
[0011] Adjust the number of target paragraphs input to the second large language model according to the paragraph evaluation result, and repeat the process of generating the first answer, outputting the evaluation result, and adjusting the number of target paragraphs until a preset termination condition is met;
[0012] Input the question and the target paragraphs with the adjusted number into a first large language model to obtain the target answer to the question.
[0013] In some embodiments, the evaluation result further includes an evaluation score of the first answer, and the preset termination condition includes that the evaluation score meets a set condition, or the number of repeated executions reaches a set number threshold.
[0014] The method further includes:
[0015] In the case where the evaluation score does not meet the set condition, obtain the paragraph evaluation result; otherwise, obtain the answer to the question.
[0016] In some embodiments, the multiple candidate paragraphs are arranged in descending order of similarity scores to the question. Adjusting the number of target paragraphs input to the second large language model according to the paragraph evaluation result includes:
[0017] In the case where the paragraph evaluation result indicates an excessive number, reduce the number of target paragraphs, and the retained target paragraphs are those ranked earlier;
[0018] In the case where the paragraph evaluation result indicates an insufficient number, increase the number of target paragraphs in order of ranking.
[0019] In some embodiments, inputting the target paragraph, the question, and the first answer into the second large language model includes:
[0020] Combine the target paragraph, the question, and the first answer according to a preset instruction template to generate a self-feedback prompt, where the self-feedback prompt is used to instruct the second large language model to generate an evaluation result in a specified manner;
[0021] Input the self-feedback prompt into the second large language model.
[0022] In some embodiments, the method further includes building a database, specifically including the following steps:
[0023] Using a semantic segmentation model, divide the documents in the corpus into multiple paragraphs with complete semantics;
[0024] Use an embedding model to convert each paragraph into a vector representation;
[0025] Construct the database according to each vector representation.
[0026] In some embodiments, the method further includes training the semantic segmentation model, specifically including the following steps:
[0027] Obtain a sample data set, the sample data set includes multiple sentence pairs composed of two adjacent sentences, and the sentence pairs have labels indicating whether the two sentences are in the same paragraph;
[0028] Train the semantic segmentation model according to the sample data set.
[0029] In some embodiments, the using a semantic segmentation model to divide the documents in the corpus into multiple paragraphs with complete semantics includes:
[0030] Input every two adjacent sentences in the corpus into the semantic segmentation model, and obtain the prediction result of the semantic segmentation model for the two adjacent sentences;
[0031] When the prediction result indicates that the two sentences are not in the same paragraph, divide the two sentences into different paragraphs;
[0032] When the prediction result indicates that the two sentences are in the same paragraph, put the two sentences into the same paragraph.
[0033] In some embodiments, the semantic segmentation model includes a sentence encoding model, a feature enhancement model, and a multi-layer perceptron, and the method further includes:
[0034] The sentence encoding model obtains a first encoding vector and a second encoding vector according to the two sentences;
[0035] The feature enhancement model performs subtraction and multiplication operations on the first encoding vector and the second encoding vector to obtain a difference vector and a product vector;
[0036] The multi-layer perceptron obtains the prediction result according to the first encoding vector, the second encoding vector, the product vector, and the difference vector.
[0037] In some embodiments, the determining the target paragraph whose change rate of the similarity score with the question meets the set requirements from the multiple candidate paragraphs includes:
[0038] In the case where the multiple candidate paragraphs are arranged in descending order of the similarity score with the question, determine the score decrease rate of each paragraph according to the score ratio of each paragraph to the previous paragraph;
[0039] Select the paragraph before the candidate paragraph whose score decrease rate reaches the set gradient threshold as the target paragraph.
[0040] According to the second aspect of one or more embodiments of the present specification, a retrieval enhanced generation device is proposed, including:
[0041] A retrieval unit for retrieving in a database according to a question proposed by a user to obtain multiple candidate paragraphs;
[0042] A determination unit for determining a target paragraph whose change rate of the similarity score with the question meets the set requirements from the multiple candidate paragraphs;
[0043] A generation unit for inputting the target paragraph and the question into a first large language model to generate a first answer;
[0044] An evaluation unit for inputting the target paragraph, the question, and the first answer into a second large language model so that the second large language model outputs an evaluation result, and the evaluation result includes a paragraph evaluation result indicating that the number of the target paragraphs is too large or too small;
[0045] An adjustment unit for adjusting the number of target paragraphs input to the second large language model according to the paragraph evaluation result, and repeatedly executing the process of generating the first answer, outputting the evaluation result, and adjusting the number of target paragraphs until a preset termination condition is met;
[0046] An obtaining unit for inputting the question and the target paragraphs with adjusted quantity into the first large language model to obtain the target answer to the question.
[0047] In some embodiments, the evaluation result further includes an evaluation score of the first answer, and the preset termination condition includes that the evaluation score meets the set condition, or the number of repeated executions reaches the set number threshold,
[0048] The device further includes a termination unit for:
[0049] In the case where the evaluation score does not meet the set condition, obtain the paragraph evaluation result, otherwise obtain the answer to the question.
[0050] In some embodiments, the adjustment unit is specifically used for:
[0051] In the case where the paragraph evaluation result indicates that the number is too large, reduce the number of the target paragraphs, and the retained target paragraphs are those ranked in the front;
[0052] When the number indicated by the paragraph evaluation result is insufficient, the number of the target paragraphs is increased according to the sorting.
[0053] In some embodiments, when the evaluation unit is used to input the target paragraph, the question, and the first answer into the second large language model, it is specifically used for:
[0054] Combining the target paragraph, the question, and the first answer according to a preset instruction template to generate a self-feedback prompt, where the self-feedback prompt is used to instruct the second large language model to generate an evaluation result in a specified manner;
[0055] Inputting the self-feedback prompt into the second large language model.
[0056] In some embodiments, the device further includes a construction unit for:
[0057] Using a semantic segmentation model to segment the documents in the corpus into multiple semantically complete paragraphs;
[0058] Using an embedding model to convert each paragraph into a vector representation;
[0059] Constructing the database according to each vector representation.
[0060] In some embodiments, the device further includes a training unit for:
[0061] Obtaining a sample data set, where the sample data set includes multiple sentence pairs composed of two adjacent sentences, and the sentence pairs have labels indicating whether the two sentences are in the same paragraph;
[0062] Training the semantic segmentation model according to the sample data set.
[0063] In some embodiments, when the construction unit is used to use a semantic segmentation model to segment the documents in the corpus into multiple semantically complete paragraphs, it is specifically used for:
[0064] Inputting every two adjacent sentences in the corpus into the semantic segmentation model and obtaining the prediction result of the semantic segmentation model for the two adjacent sentences;
[0065] When the prediction result indicates that the two sentences are not in the same paragraph, splitting the two sentences into different paragraphs;
[0066] When the prediction result indicates that the two sentences are in the same paragraph, putting the two sentences into the same paragraph.
[0067] In some embodiments, the semantic segmentation model includes a sentence encoding model, a feature enhancement model, and a multi-layer perceptron. The apparatus further includes a prediction unit configured to:
[0068] The sentence encoding model obtains a first encoding vector and a second encoding vector according to the two sentences;
[0069] The feature enhancement model performs a subtraction operation and a multiplication operation on the first encoding vector and the second encoding vector to obtain a difference vector and a product vector;
[0070] The multi-layer perceptron obtains the prediction result according to the first encoding vector, the second encoding vector, the product vector, and the difference vector.
[0071] In some embodiments, the determining unit is specifically configured to:
[0072] In the case where the multiple candidate paragraphs are arranged in descending order of the similarity score with the question, determine the score decrease rate of each paragraph according to the score ratio of each paragraph to the previous paragraph;
[0073] Select the paragraph before the candidate paragraph whose score decrease rate reaches the set gradient threshold as the target paragraph.
[0074] According to a third aspect of one or more embodiments of the present specification, an electronic device is provided, including:
[0075] A processor;
[0076] A memory for storing processor-executable instructions;
[0077] Wherein, the processor runs the executable instructions to implement the steps of the method proposed in the above embodiments.
[0078] According to a fourth aspect of one or more embodiments of the present specification, a computer-readable storage medium is provided, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method proposed in the above embodiments are implemented.
[0079] According to a fifth aspect of one or more embodiments of the present specification, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method proposed in the above embodiments are implemented.
[0080] The retrieval-enhanced generation method proposed in the embodiments of this specification first determines the target paragraphs for the relevance between multiple candidate paragraphs retrieved from the database and the user's question. Then, for the answers generated by a large language model, another large language model is used to evaluate the rationality of the number of target paragraphs, and the number of target paragraphs is adjusted according to the evaluation results. Finally, the target answer to the user's question is obtained based on the adjusted number of target paragraphs as the context. This solution effectively solves the problems of noisy retrieval and missing retrieval in the retrieval-enhanced generation task, improves the accuracy of retrieval, and enhances the generation quality. Description of the Drawings
[0081] Figure 1 It is an application scenario diagram of a retrieval-enhanced generation method provided by an exemplary embodiment.
[0082] Figure 2 It is a flowchart of a retrieval-enhanced generation method provided by an exemplary embodiment.
[0083] Figure 3 It is a schematic diagram of the paragraph relevance score.
[0084] Figure 4 It is a block diagram of a retrieval-enhanced generation device provided by an exemplary embodiment.
[0085] Figure 5 It is a schematic structural diagram of a device provided by an exemplary embodiment. Detailed Embodiments
[0086] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0087] It should be noted that: in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0088] When existing retrieval-augmented generation methods retrieve relevant documents, they inevitably introduce redundant information unrelated to the query intent. The existence of this noisy information may mislead the inference process of the language model, resulting in a decrease in the accuracy of the generated answers. On the other hand, when insufficient key context information is obtained, i.e., the missing retrieval phenomenon occurs, this will directly limit the generation ability of the language model and prevent it from providing a complete and accurate answer. These technical bottlenecks severely restrict the performance of retrieval-augmented generation methods in practical application scenarios.
[0089] To solve the above problems, the embodiments of this application propose a retrieval-augmented generation scheme.
[0090] As Figure 1 shown, it is an application scenario diagram of a retrieval-augmented generation method in the embodiments of this application. Figure 1 It includes: client 10, database 20, first large language model 30, and second large language model 40.
[0091] Client 10 is a device with network connection, data processing, and user interaction capabilities, responsible for receiving requests input by users, executing the request processing flow, and presenting the processing results to users. Client 10 can be the client of a smartphone, tablet computer, desktop computer, smart wearable device, in-vehicle system, smart home device, learning machine, etc.
[0092] Database 20 is pre-constructed and can be stored either inside client 10 or independently outside client 10. Client 10 can interact with database 20.
[0093] The first large language model 30 and the second large language model 40 are large-scale neural network models in the field of natural language processing (NLP). It usually consists of hundreds of millions to billions of parameters and learns complex patterns and structures through pre-training on a large amount of text data, capable of understanding and generating complex language data. The first large language model 30 and the second large language model 40 are usually deployed on a server to provide backend support for the applications or services of client 10. Client 10 can communicate with the first large language model 30 and the second large language model 40 on the server through the Internet, send requests, and receive processed responses.
[0094] It should be noted that if additional modules are added to or individual modules are removed from the illustrated environment, the underlying concept of the example embodiments of this application will not be changed; and the request processing method proposed in this application is not only applicable to Figure 1 the application scenario shown, but also applicable to any device with request processing requirements.
[0095] Figure 2 shows a flowchart of a retrieval enhanced generation method provided by an embodiment of the present application. This method is executed by Figure 1 the client in Figure 2 . As shown, the method includes:
[0096] Step 201, retrieve in the database according to the question raised by the user to obtain multiple candidate paragraphs.
[0097] In some embodiments, the database may be a pre - constructed vector retrieval database, which contains multiple candidate paragraph vectors. By using an embedding model to vectorize the user's question, the question is converted into a vector. Next, query the vector retrieval database to extract multiple candidate vectors that are closest to the question vector. Specifically, the similarity (correlation) between the question vector and the paragraph vectors in the database can be calculated, and then all paragraphs are sorted in descending order according to the similarity score (correlation score), and the top N paragraphs with the highest scores are selected as candidate results. Among them, the number N of the multiple candidate paragraphs can be preset.
[0098] Step 202, determine a target paragraph from the multiple candidate paragraphs whose change rate of similarity score with the question meets the set requirements.
[0099] Since the multiple candidate paragraphs obtained in step 201 may ignore useful paragraphs due to insufficient selection quantity, or may include irrelevant paragraphs due to excessive selection quantity, the unnecessary information in these paragraphs may confuse the large - language model and make it difficult to give the correct answer.
[0100] Figure 3 shows two cases of the paragraph correlation scores of two articles in the database. As Figure 3 can be seen, the scoring pattern of the paragraphs shows a sharp drop before gradually decreasing. In this example, if only the first three paragraphs of article 1 and one paragraph of article 2 are selected, it is easier to get the correct answer. This indicates that the paragraphs before the sharp drop are significantly more relevant to the question than the subsequent paragraphs.
[0101] Therefore, in the embodiments of the present disclosure, the change rate obtained from the similarity between the multiple retrieved candidate paragraphs and the question raised by the user is used to further screen these candidate paragraphs, and the candidate paragraphs whose change rate meets the set requirements are determined as the target paragraphs for subsequent processing steps.
[0102] Specifically, when the multiple candidate paragraphs are arranged in descending order of similarity scores to the problem, the score decrease rate of each paragraph can be determined according to the score ratio of each paragraph to the previous paragraph; and the paragraph before the candidate paragraph whose score decrease rate reaches the set gradient threshold is selected as the target paragraph. That is, from the N candidate paragraphs obtained in step 201, the top K candidate paragraphs (K < N) with the highest ranking (selecting the top by descending order) before the significant decrease in similarity scores are selected as the target paragraphs, ensuring that only the candidate paragraphs whose relevance to the problem meets the set requirements are selected for the next stage.
[0103] This method of dynamically selecting paragraphs based on the score gradient of sorted paragraphs instead of selecting a fixed number of paragraphs can more accurately identify the paragraphs most relevant to a given problem.
[0104] Step 203, input the target paragraph and the problem into the first large language model to generate a first answer.
[0105] Use the target paragraph filtered in step 202 as the context, and output it together with the question raised by the user to the first large language model. The model generates an initial answer in response to the question, which is called the first answer here.
[0106] Step 204, input the target paragraph, the problem and the first answer into the second large language model so that the second large language model outputs an evaluation result.
[0107] Among them, the evaluation result can include a paragraph evaluation result indicating that the number of target paragraphs is too large or too small. That is, use the second large language model to evaluate whether the number of target paragraphs as the context is too large or too small.
[0108] Step 205, adjust the number of target paragraphs input to the second large language model according to the paragraph evaluation result, and repeat the process of generating the first answer, outputting the evaluation result and adjusting the number of target paragraphs until the preset termination condition is met.
[0109] Specifically, when the paragraph evaluation result indicates that the number is too large, reduce the number of target paragraphs, and the retained ones are the target paragraphs ranked in the front; when the paragraph evaluation result indicates that the number is too small, increase the number of target paragraphs according to the ranking.
[0110] For example, assume that in step 201, N candidate paragraphs are retrieved according to the user's question, and in step 202, K target paragraphs are selected from these N candidate paragraphs whose change rates of similarity scores meet the set requirements, that is, the K paragraphs with the highest relevance scores. If the paragraph evaluation result indicates that the number of target paragraphs is too large, then K - 1 paragraphs are selected from the sorted paragraphs; if the paragraph evaluation result indicates that the number of target paragraphs is insufficient, then K + 1 paragraphs are selected from the sorted paragraphs. Then steps 203 to 205 are repeatedly executed, that is, the reselected paragraphs are input into the first large language model to generate the first answer, and then the second large language model is used to evaluate whether the number of paragraphs is too large or insufficient, and adjustments are made again according to the evaluation result until the preset termination condition is met. The termination condition may include that the evaluation score meets the set condition, or the number of repeated executions reaches the set number threshold M. That is, when steps 203 to 205 have been repeatedly executed M times, or the paragraph evaluation result is neither too large nor too small (the judgment result of the second large language model is that the number of paragraphs remains unchanged), the adjustment is stopped to obtain the final target paragraphs.
[0111] By using the second large language model to adjust the number of retrieved target paragraphs, the problems of noisy retrieval and missing retrieval are further solved.
[0112] Step 206, input the question and the target paragraphs with adjusted quantity into the first large language model to obtain the target answer to the question.
[0113] The retrieval enhanced generation method proposed in the embodiments of this specification first determines target paragraphs for the relevance between multiple candidate paragraphs retrieved from the database and the user's question, then for the answer generated by a large language model, uses another large language model to evaluate the rationality of the number of target paragraphs, and adjusts the number of target paragraphs according to the evaluation result, and finally uses the target paragraphs with adjusted quantity as the context to obtain the target answer to the user's question. This solution effectively solves the problems of noisy retrieval and missing retrieval in the retrieval enhanced generation task, improves the accuracy of retrieval, and enhances the generation quality.
[0114] In some embodiments, the evaluation result further includes the evaluation score of the first answer. For example, it may indicate that the second large language model comprehensively scores the first answer from multiple dimensions such as the accuracy, relevance, integrity, fluency, and information richness of the first answer. The value of this score may be, for example, from 1 to 10 points.
[0115] In the case where the evaluation result also includes the evaluation score of the first answer, first determine whether the evaluation score meets the set conditions, such as whether the score has reached the feedback score threshold fs, for example 9. If it meets, directly obtain the target answer to the question and return it to the user; if it does not meet, further obtain the paragraph evaluation result and adjust the number of target paragraphs according to the paragraph evaluation result.
[0116] In some embodiments, the target paragraph, the question, and the first answer can be combined according to a preset instruction template to generate a self-feedback prompt, and the self-feedback prompt is used to instruct the second large language model to generate an evaluation result in a specified manner; next, the self-feedback prompt is input into the second large language model.
[0117] This self-feedback prompt can be used for two purposes: 1) to evaluate the quality of the first answer, 2) to evaluate the suitability of the retrieved paragraphs, that is, to evaluate whether the target paragraphs as the context contain redundant paragraphs or lack paragraphs necessary to answer the question. This feedback process produces two feedback outputs. The first is an evaluation score from 1 to 10, and the second is a paragraph evaluation result used to represent context adjustment, which can be represented by -1 or 1. Among them, -1 indicates that there is redundant information in the target paragraph, while 1 indicates that additional information is needed to fully answer the question.
[0118] After receiving the feedback output, first check the evaluation score of the first answer. If this score reaches or exceeds the feedback score threshold fs, then the generated first answer is considered acceptable, and it is determined as the target answer and then presented to the user; if not, context adjustment is performed according to the paragraph evaluation result. If the paragraph evaluation result indicates -1, it means that the number of target paragraphs should be reduced, for example, one is reduced according to the sorting; on the contrary, if the paragraph evaluation result indicates 1, it means that the number of target paragraphs should be increased, for example, one is increased according to the sorting. After making the adjustment, repeat steps 203 to 205 until the evaluation score reaches or exceeds the feedback score threshold fs, or this loop has been executed the set number of times threshold M.
[0119] In current RAG systems, two main corpus segmentation methods are adopted. The first method is to segment the corpus into paragraphs according to a predetermined number of tokens. The second method, based on the first method, ensures that each paragraph contains complete sentences. However, the first method may result in paragraphs containing incomplete sentences, weakening the overall coherence and meaning of each paragraph and making the calculated similarity between the user's question and these paragraphs invalid. The second method is to segment consecutive sentences shorter than a fixed length into one paragraph. According to this method, if a paragraph exceeds the predetermined length limit, the last sentence will be completely transferred to the next paragraph to ensure that each paragraph contains complete sentences. However, this method may also distort the meaning of the paragraph.
[0120] To solve the above problems, embodiments of the present disclosure propose a method for constructing a database. A lightweight and effective semantic segmentation model is used to segment the documents in the corpus into multiple semantically complete paragraphs to ensure that each paragraph conveys a clear and complete meaning and does not contain too many tokens. Then, an embedding model is used to convert each paragraph into a vector representation, and the database is constructed based on these vector representations.
[0121] The semantic segmentation model constructed in the embodiments of the present disclosure mainly includes three parts, namely a sentence encoding model, a feature enhancement model, and a multi-layer perceptron.
[0122] Among them, the sentence encoding model is a lightweight embedding model that obtains a first encoding vector x1 and a second encoding vector x2 according to the two input sentences. The feature enhancement model receives x1 and x2, performs subtraction and multiplication operations to process these embeddings, and obtains their difference vector (x1 - x2) and product vector x1 * x2. The multi-layer perceptron receives x1, x2, as well as (x1 - x2) and (x1 * x2), and generates a score representing the prediction result. This score is used to determine whether the two input sentences should be segmented or remain as a continuous part.
[0123] The semantic segmentation model proposed in the embodiments of the present disclosure can better capture the semantic differences or similarities between sentences, thereby improving the semantic segmentation effect.
[0124] In some embodiments, the semantic segmentation model can be trained in the following manner.
[0125] First, a sample data set is obtained, that is, a large amount of training data for semantic segmentation is collected. The sample data set contains multiple sentence pairs composed of two adjacent sentences, and the sentence pairs have labels indicating whether the two sentences are in the same paragraph. For example, a label of 1 indicates that these two sentences should be grouped into the same paragraph; a label of 0 indicates that they should be segmented into different paragraphs.
[0126] In one example, the Wikipedia dataset can be used as the data source, and the paragraphs in this dataset have been semantically segmented into paragraphs.
[0127] Next, all sentence pairs are input into the semantic segmentation model. A loss function is constructed based on the output of the model and the labels of the sentence pairs themselves. The parameters of the semantic segmentation model are trained using this loss function until the loss function converges, obtaining the trained semantic segmentation model.
[0128] In some embodiments, every two adjacent sentences in the corpus can be input into the semantic segmentation model, and the prediction result of the semantic segmentation model for the two adjacent sentences can be obtained; when the prediction result indicates that the two sentences are not in the same paragraph, the two sentences are segmented into different paragraphs; when the prediction result indicates that the two sentences are in the same paragraph, the two sentences are placed in the same paragraph.
[0129] Among them, the prediction result can be a score. If this score is lower than a predetermined segmentation score threshold ss (this threshold is between 0 and 1, for example, 0.5), then the two adjacent sentences will be segmented into different paragraphs; on the contrary, if the score is higher than or equal to ss, the two adjacent sentences will remain in the same paragraph.
[0130] For any given corpus, this segmentation process can be executed quickly because this lightweight semantic segmentation model can run in parallel on the GPU. For example, all sentence pairs in the corpus can be collected and organized into multiple batches, and then the semantic segmentation model is called to perform inference in parallel.
[0131] Figure 4 is a block diagram of a retrieval-enhanced generation device provided by an exemplary embodiment. As Figure 4 shown, the device includes:
[0132] A retrieval unit 401, configured to perform a retrieval in a database according to a question raised by a user to obtain a plurality of candidate paragraphs;
[0133] A determination unit 402, configured to determine a target paragraph from the plurality of candidate paragraphs whose change rate of the similarity score with the question meets a set requirement;
[0134] A generation unit 403, configured to input the target paragraph and the question into a first large language model to generate a first answer;
[0135] An evaluation unit 404, configured to input the target paragraph, the question, and the first answer into a second large language model, so that the second large language model outputs an evaluation result, and the evaluation result includes a paragraph evaluation result indicating that the number of the target paragraphs is too large or too small;
[0136] An adjustment unit 405, configured to adjust the number of target paragraphs input to the second large language model according to the paragraph evaluation result, and repeatedly execute the processes of generating the first answer, outputting the evaluation result, and adjusting the number of target paragraphs until a preset termination condition is met;
[0137] An obtaining unit 406, configured to input the question and the target paragraphs with the adjusted number into the first large language model to obtain the target answer to the question.
[0138] In some embodiments, the evaluation result further includes an evaluation score of the first answer, and the preset termination condition includes that the evaluation score meets a set condition, or the number of repeated executions reaches a set number threshold.
[0139] The apparatus further includes a termination unit, configured to:
[0140] In the case where the evaluation score does not meet the set condition, obtain the paragraph evaluation result, otherwise obtain the answer to the question.
[0141] In some embodiments, the adjustment unit is specifically configured to:
[0142] In the case where the paragraph evaluation result indicates that the quantity is excessive, reduce the number of target paragraphs, and the retained target paragraphs are those sorted in the front;
[0143] In the case where the paragraph evaluation result indicates that the quantity is insufficient, increase the number of target paragraphs according to the sorting.
[0144] In some embodiments, when the evaluation unit is configured to input the target paragraphs, the question, and the first answer into the second large language model, it is specifically configured to:
[0145] Combine the target paragraphs, the question, and the first answer according to a preset instruction template to generate a self-feedback prompt, where the self-feedback prompt is used to instruct the second large language model to generate an evaluation result in a specified manner;
[0146] Input the self-feedback prompt into the second large language model.
[0147] In some embodiments, the apparatus further includes a construction unit, configured to:
[0148] Use a semantic segmentation model to segment the documents in the corpus into multiple semantically complete paragraphs;
[0149] Use an embedding model to convert each paragraph into a vector representation;
[0150] Construct the database according to each vector representation.
[0151] In some embodiments, the device further includes a training unit configured to:
[0152] Obtain a sample data set, where the sample data set includes a plurality of sentence pairs each composed of two adjacent sentences, and the sentence pairs have labels indicating whether the two sentences are in the same paragraph;
[0153] Train the semantic segmentation model according to the sample data set.
[0154] In some embodiments, when the building unit is configured to use the semantic segmentation model to segment a document in a corpus into a plurality of semantically complete paragraphs, it is specifically configured to:
[0155] Input every two adjacent sentences in the corpus into the semantic segmentation model, and obtain the prediction result of the semantic segmentation model for the two adjacent sentences;
[0156] When the prediction result indicates that the two sentences are not in the same paragraph, segment the two sentences into different paragraphs;
[0157] When the prediction result indicates that the two sentences are in the same paragraph, put the two sentences into the same paragraph.
[0158] In some embodiments, the semantic segmentation model includes a sentence encoding model, a feature enhancement model, and a multi-layer perceptron. The device further includes a prediction unit configured to:
[0159] The sentence encoding model obtains a first encoding vector and a second encoding vector according to the two sentences;
[0160] The feature enhancement model performs a subtraction operation and a multiplication operation on the first encoding vector and the second encoding vector to obtain a difference vector and a product vector;
[0161] The multi-layer perceptron obtains the prediction result according to the first encoding vector, the second encoding vector, the product vector, and the difference vector.
[0162] In some embodiments, the determining unit is specifically configured to:
[0163] In the case where the plurality of candidate paragraphs are arranged in descending order of the similarity score with the question, determine the score decrease rate of each paragraph according to the score ratio of each paragraph to the previous paragraph;
[0164] Select the paragraph before the candidate paragraph whose score decrease rate reaches the set gradient threshold as the target paragraph.
[0165] Figure 5 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 5, at the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510. Of course, it may also include other hardware required for other services. One or more embodiments of this specification can be implemented in software. For example, the processor 502 reads the corresponding computer program from the non-volatile memory 510 into the memory 508 and then runs it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logical devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logical unit, and can also be hardware or logical devices.
[0166] The systems, devices, modules, or units described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.
[0167] In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0168] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0169] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0170] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.
[0171] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0172] The terms used in one or more embodiments of this specification are for the purpose of describing particular embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the" and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0173] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0174] The above are only the preferred embodiments of one or more embodiments of this specification and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.
Claims
1. A retrieval-augmented generation method, characterized in that, Including: Retrieving in a database according to a question raised by a user to obtain multiple candidate paragraphs; Determining a target paragraph from the multiple candidate paragraphs whose change rate of similarity score with the question meets a set requirement; Inputting the target paragraph and the question into a first large language model to generate a first answer; Inputting the target paragraph, the question and the first answer into a second large language model so that the second large language model outputs an evaluation result, where the evaluation result includes a paragraph evaluation result indicating that the number of target paragraphs is too large or insufficient; Adjusting the number of target paragraphs input to the second large language model according to the paragraph evaluation result, and repeating the process of generating the first answer, outputting the evaluation result and adjusting the number of target paragraphs until a preset termination condition is met; Inputting the question and the target paragraphs with adjusted quantity into the first large language model to obtain the target answer to the question.
2. The method according to claim 1, wherein The evaluation result further includes an evaluation score of the first answer, and the preset termination condition includes that the evaluation score meets a set condition, or the number of repeated executions reaches a set number threshold. The method further includes: In the case where the evaluation score does not meet the set condition, obtaining the paragraph evaluation result, otherwise obtaining the answer to the question.
3. The method according to claim 1 or 2, characterized in that, The multiple candidate paragraphs are arranged in descending order of similarity score with the question. Adjusting the number of target paragraphs input to the second large language model according to the paragraph evaluation result includes: In the case where the paragraph evaluation result indicates that the number is too large, reducing the number of target paragraphs, and the retained ones are the target paragraphs ranked earlier; In the case where the paragraph evaluation result indicates that the number is insufficient, increasing the number of target paragraphs in order.
4. The method according to claim 1, characterized in that, Inputting the target paragraph, the question and the first answer into the second large language model includes: Combining the target paragraph, the question and the first answer according to a preset instruction template to generate a self-feedback prompt, where the self-feedback prompt is used to instruct the second large language model to generate an evaluation result in a specified manner; Inputting the self-feedback prompt into the second large language model.
5. The method according to claim 1, characterized in that The method further includes constructing a database, specifically including the following steps: Using a semantic segmentation model to segment the documents in the corpus into multiple semantically complete paragraphs; Using an embedding model to convert each paragraph into a vector representation; Constructing the database according to each vector representation.
6. The method according to claim 5, wherein The method further includes training the semantic segmentation model, specifically including the following steps: Obtaining a sample data set, where the sample data set contains multiple sentence pairs composed of two adjacent sentences, and the sentence pairs have labels indicating whether the two sentences are in the same paragraph; Training the semantic segmentation model according to the sample data set.
7. The method according to claim 6, wherein Using the semantic segmentation model to segment the documents in the corpus into multiple semantically complete paragraphs includes: Inputting every two adjacent sentences in the corpus into the semantic segmentation model and obtaining the prediction result of the semantic segmentation model for the two adjacent sentences; When the prediction result indicates that the two sentences are not in the same paragraph, split the two sentences into different paragraphs; When the prediction result indicates that the two sentences are in the same paragraph, put the two sentences in the same paragraph.
8. The method according to claim 7, characterized in that The semantic segmentation model includes a sentence encoding model, a feature enhancement model, and a multi-layer perceptron. The method further includes: The sentence encoding model obtains a first encoding vector and a second encoding vector according to the two sentences; The feature enhancement model performs subtraction and multiplication operations on the first encoding vector and the second encoding vector to obtain a difference vector and a product vector; The multi-layer perceptron obtains the prediction result according to the first encoding vector, the second encoding vector, the product vector, and the difference vector.
9. The method according to claim 1, wherein Determining the target paragraph from the multiple candidate paragraphs whose change rate of the similarity score with the question meets the set requirements includes: When the multiple candidate paragraphs are arranged in descending order of the similarity score with the question, determine the score decrease rate of each paragraph according to the score ratio of each paragraph to the previous paragraph; Select the paragraph before the candidate paragraph whose score decrease rate reaches the set gradient threshold as the target paragraph.
10. A retrieval-augmented generation device, characterized in that, including: A retrieval unit for retrieving in a database according to a question raised by a user to obtain multiple candidate paragraphs; A determination unit for determining a target paragraph from the multiple candidate paragraphs whose change rate of the similarity score with the question meets the set requirements; A generation unit for inputting the target paragraph and the question into a first large language model to generate a first answer; An evaluation unit for inputting the target paragraph, the question, and the first answer into a second large language model so that the second large language model outputs an evaluation result, and the evaluation result includes a paragraph evaluation result indicating that the number of the target paragraphs is too large or too small; An adjustment unit for adjusting the number of target paragraphs input to the second large language model according to the paragraph evaluation result, and repeating the process of generating the first answer, outputting the evaluation result, and adjusting the number of target paragraphs until a preset termination condition is met; An obtaining unit for inputting the question and the target paragraphs with adjusted quantity into the first large language model to obtain the target answer to the question.
11. An electronic device, comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor realizes the method according to any one of claims 1-9 by running the executable instructions.
12. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1-9 are realized.
13. A computer program product, comprising computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1-9 are realized.
Citation Information
Patent Citations
Response method and device of question
CN108509463A
Intelligent dialogue method and system based on reading understanding model
CN112163079A