Error location method, electronic device, and medium for retrieval enhancement generation system

By obtaining the intermediate execution results of the search enhancement generation system and using the large language model for reasoning, the normality of the searcher and generation module in the search enhancement generation system is solved, and the module-level error positioning and diagnostic accuracy are improved.

CN119474276BActive Publication Date: 2025-05-13ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510027118.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-13
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

Existing search-enhanced generation systems (RAGs) have difficulties in error diagnosis, including modular complexity, defectiveness of traditional metrics, and inability to trace error modules.

Method used

By obtaining the intermediate execution results of the search-enhanced generation system, extracting fact triplets of the original search-related documents, and using a large language model to reason, judging the normality of the searcher and the generation module, thereby achieving module-level error positioning.

Benefits of technology

This method can extract key information in the context, perform in-depth reasoning, evaluate the contribution of each module in the retrieval enhancement generation system to the final answer, and improve the accuracy of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474276B_ABST
    Figure CN119474276B_ABST
Patent Text Reader

Abstract

The invention discloses an error location method, electronic device and medium for a retrieval enhancement generation system, comprising: obtaining an intermediate execution result of the retrieval enhancement generation system, including: a user question, an original retrieval-related document, a model response and a standard answer; inserting the original retrieval-related document into a first prompt word template, inputting it into a first large language model, and extracting an original retrieval fact triple; inserting all the original retrieval fact triples into a second prompt word template, inputting it into a second large language model, and judging whether all the original retrieval fact triples can deduce an answer to answer the user question; if the answer can be derived, judging that the retriever in the retrieval enhancement generation system is normal; otherwise, judging that the retriever is abnormal; inputting the user question, the model response, the standard answer and the original retrieval fact triples into a third large language model, judging the accuracy and completeness of the model response, and thus judging whether the large language model in the retrieval enhancement generation system is abnormal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing and machine learning, and in particular relates to an error location method, electronic device, and medium for a retrieval enhancement generation system. Background Art

[0002] Retrieval Augmented Generation (RAG) is a hybrid system that combines information retrieval and natural language generation. It is widely used in open-domain question answering, knowledge question answering and other tasks. The RAG model obtains relevant documents from an external knowledge base through a retrieval module, and then the re-ranking module selects the most relevant documents from the relevant documents. The generation module then generates answers based on the retrieval results to improve the accuracy and richness of the generated content. However, the complexity of the RAG system also leads to multiple challenges in practical applications. For example, the retrieval module may not be able to find documents that are highly relevant to the query, the generation module may generate inaccurate answers, or the re-ranking module may fail to correctly rank the most relevant results. The existence of these problems makes it difficult to guarantee the output quality of the RAG system, and it becomes particularly difficult to locate specific problem modules.

[0003] At present, there are mainly the following difficulties in error diagnosis of RAG system:

[0004] (1) Modular complexity: The RAG system consists of a retriever and a generator. Each module interacts with each other and jointly affects the final performance of the entire system.

[0005] (2) The shortcomings of traditional indicators: Most of the existing methods for evaluating RAG systems are rule-based and rely on strict manual annotation results. Specifically, traditional indicators of retrievers (such as recall@k and MRR) require the correct retrieval of text blocks to be defined in advance, while coarse-grained text alignment ignores the semantic relevance of the retrieved text. For generators, indicators based on n-grams (such as BLEU, ROUGE) and embeddings (such as BERTScore) cannot capture the subtle differences between responses and standard answers.

[0006] (3) Unable to trace the faulty module: Due to the mutual influence between modules, the evaluation method of the traditional RAG system only provides quantitative statistical indicators of each module, which is difficult to effectively capture the complexity and overall quality of the retrieval and generation components in the RAG system. For a specific RAG production environment, quantitative indicators cannot accurately identify the specific module that causes the system performance to degrade.

[0007] For example, suppose a RAG system generates an answer that is inconsistent with the retrieved document content when answering a user's question about "the main cause of global warming". The generated answer may emphasize some non-main factors and ignore the key factors. In this case, it may be that the retrieval process is wrong, or the generation module produces hallucinations (i.e., generates content that is irrelevant to the retrieved information).

[0008] Therefore, there is an urgent need to provide a method that can accurately locate the source of errors in the retrieval enhancement generation system. It can be found whether the retrieval module in the retrieval enhancement generation system fails to find suitable relevant documents, or whether the generation module has deviations in understanding and summarizing the retrieval information, thereby helping to identify the specific source of the problem and identify and analyze potential problems in each module. Summary of the invention

[0009] In view of the deficiencies in the prior art, the present invention provides an error location method, electronic device, and medium for a retrieval enhancement generation system.

[0010] In a first aspect, an embodiment of the present invention provides an error location method for a retrieval enhancement generation system, the method comprising:

[0011] Obtaining an intermediate execution result of the retrieval enhancement generation system, wherein the intermediate execution result includes: user questions, original retrieval-related documents, model responses, and standard answers;

[0012] Insert the original search related documents into the first prompt word template, input into the first language model, and extract the original search fact triples; insert all the original search fact triples into the second prompt word template, input into the second language model, and judge whether all the original search fact triples can deduce answers to answer the user's questions; if the answers can be derived, it is determined that the retriever in the search enhancement generation system is normal; otherwise, it is determined that the retriever is abnormal;

[0013] The triples of user questions, model responses, standard answers, and original retrieval facts are input into the third largest language model to determine the accuracy, completeness, derivability, and fluency of the model responses, thereby determining whether the large language model in the retrieval enhancement generation system is abnormal.

[0014] In a second aspect, an embodiment of the present invention provides an error location system for a retrieval enhancement generation system, the system comprising:

[0015] An intermediate execution result acquisition module is used to acquire the intermediate execution result of the retrieval enhancement generation system, wherein the intermediate execution result includes: user questions, original retrieval-related documents, model responses, and standard answers;

[0016] A fact triple extractor, taking the first prompt word template inserted into the original search related document as input, extracts the original search fact triple;

[0017] The retrieval module verifier takes the second prompt word template inserted into all the original retrieval fact triples as input, and determines whether all the original retrieval fact triples can deduce answers to answer the user's questions; if the answers can be derived, the retriever in the retrieval enhancement generation system is determined to be normal; otherwise, the retriever is determined to be abnormal;

[0018] The generation module verifier takes the user question, model response, standard answer, and original retrieval fact triple as input to determine the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval enhancement generation system is abnormal.

[0019] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned error location method for the retrieval enhancement generation system.

[0020] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned error location method for a retrieval enhancement generation system.

[0021] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned error localization method for a retrieval enhancement generation system.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] The present invention provides an error location method for a retrieval enhancement generation system, which performs triple extraction on original retrieval-related documents and combines a large language model to reason about all original retrieval-related documents. It can extract key information in the context and perform deep reasoning, thereby evaluating the contribution of the retriever and large language model generation modules in the retrieval enhancement generation system to the final answer, achieving module-level error location, and improving the accuracy of diagnosis under the condition of multi-level information flow transmission. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0025] Figure 1 A flowchart of an error location method for a retrieval enhancement generation system provided by an embodiment of the present invention;

[0026] Figure 2 A schematic diagram of a search-oriented enhanced generation system provided by an embodiment of the present invention;

[0027] Figure 3 A schematic diagram of extracting fact triples provided by an embodiment of the present invention;

[0028] Figure 4 A schematic diagram of retrieval verification provided by an embodiment of the present invention;

[0029] Figure 5 A schematic diagram of generating module verification provided by an embodiment of the present invention;

[0030] Figure 6 A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] It should be noted that, in the absence of conflict, the features in the following embodiments and implementations may be combined with each other.

[0033] Example 1

[0034] In this example, the Retrieval Augmented Generation (RAG) model includes a retriever and a large language model; after the user question is retrieved by the retriever, the original retrieval-related documents are obtained, and the user question and the original retrieval-related documents are inserted into the prompt word template and input into the large language model to obtain a model response.

[0035] For the above retrieval enhancement generation model, such as Figure 1 As shown, this example provides an error location method for a retrieval enhancement generation system, the method comprising:

[0036] Step S1, obtaining the intermediate execution result of the retrieval enhancement generation system, wherein the intermediate execution result includes: user question, original retrieval related documents, model response, and standard answer.

[0037] Step S2, insert the original search related documents into the first prompt word template, input into the first largest language model, extract the original search fact triples; insert all the original search fact triples into the second prompt word template, input into the second largest language model, and judge whether all the original search fact triples can deduce answers to answer user questions; if the answers can be derived, it is judged that the retriever in the search enhancement generation system is normal; otherwise, it is judged that the retriever is abnormal.

[0038] Step S3, input the user question, model response, standard answer, and original search fact triples into the third language model to determine the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval enhancement generation system is abnormal.

[0039] Accordingly, an embodiment of the present invention provides an error location system for a retrieval enhancement generation system, the system comprising:

[0040] The intermediate execution result acquisition module is used to obtain the intermediate execution results of the retrieval enhancement generation system, and the intermediate execution results include: user questions, original retrieval related documents, model responses, and standard answers.

[0041] The fact triple extractor, i.e., the first language model, takes the first prompt word template inserted into the original search-related document as input to extract the original search fact triples.

[0042] The retrieval module verifier, that is, the second largest language model, takes the second prompt word template inserted into all the original retrieval fact triples as input, and determines whether all the original retrieval fact triples can derive answers to answer user questions; if the answers can be derived, the retriever in the retrieval enhancement generation system is judged to be normal; otherwise, the retriever is judged to be abnormal.

[0043] The generation module verifier, that is, the third largest language model, takes the user question, model response, standard answer, and original retrieval fact triple as input, and determines the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval enhanced generation system is abnormal.

[0044] Example 2

[0045] In this example, the Retrieval Augmented Generation (RAG) model includes a rewriter, a retriever, and a large language model; the user question is processed by the rewriter to obtain a rewritten question, the user question and the rewritten question are retrieved by the retriever to obtain original retrieval-related documents and rewritten retrieval-related documents, the user question, the original retrieval-related documents, and the rewritten retrieval-related documents are inserted into a prompt word template and input into the large language model to obtain a model response.

[0046] In view of the above-mentioned retrieval enhancement generation model, this example provides an error location method for a retrieval enhancement generation system, the method comprising:

[0047] Step S1, obtaining the intermediate execution result of the retrieval enhancement generation system, wherein the intermediate execution result includes: user question, original retrieval-related documents, rewritten retrieval-related documents, model response, and standard answer.

[0048] Step S2, insert the original search related documents into the first prompt word template, input into the first largest language model, extract the original search fact triples; insert all the original search fact triples into the second prompt word template, input into the second largest language model, and judge whether all the original search fact triples can deduce answers to answer user questions; if the answers can be derived, it is judged that the retriever in the search enhancement generation system is normal; otherwise, it is judged that the retriever is abnormal.

[0049] Step S3, insert the rewritten retrieval related documents into the first prompt word template, input it into the first largest language model, and extract the rewritten retrieval fact triples; insert all the rewritten retrieval fact triples into the second prompt word template, input it into the second largest language model, and judge whether all the rewritten retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, it is determined that the rewriter in the retrieval enhancement generation system is normal; otherwise, it is determined that the rewriter is abnormal.

[0050] Step S4, input the user question, model response, standard answer, and rewritten retrieval fact triples into the third large language model to determine the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval enhancement generation system is abnormal.

[0051] Accordingly, an embodiment of the present invention provides an error location system for a retrieval enhancement generation system, the system comprising:

[0052] The intermediate execution result acquisition module is used to obtain the intermediate execution results of the retrieval enhancement generation system, and the intermediate execution results include: user questions, original retrieval-related documents, rewritten retrieval-related documents, model responses, and standard answers.

[0053] The fact triple extractor, i.e., the first language model, takes the first prompt word template inserted into the original search-related document as input to extract the original search fact triple;

[0054] The first prompt word template of the inserted rewritten retrieval related document is taken as input to extract the rewritten retrieval fact triple.

[0055] The retrieval module verifier, i.e., the second language model, takes the second prompt word template inserted into all the original retrieval fact triples as input, and determines whether all the original retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, the retriever in the retrieval enhancement generation system is judged to be normal; otherwise, the retriever is judged to be abnormal;

[0056] Taking the second prompt word template into which all rewritten retrieval fact triples are inserted as input, it is determined whether all rewritten retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, the retriever in the retrieval enhancement generation system is judged to be normal; otherwise, the retriever is judged to be abnormal.

[0057] The generation module verifier takes the user question, model response, standard answer, and rewritten retrieval fact triple as input to determine the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval enhancement generation system is abnormal.

[0058] Example 3

[0059] In this example, if Figure 2 As shown, the Retrieval Augmented Generation (RAG) model includes a rewriter, a retriever, a reranker and a large language model; the user question is processed by the rewriter to obtain a rewritten question, the user question and the rewritten question are retrieved by the retriever to obtain original retrieval-related documents and rewritten retrieval-related documents, the original retrieval-related documents and the rewritten retrieval-related documents are processed by the reranker to obtain reranked-related documents, the user question and the reranked-related documents are inserted into the prompt word template, and input into the large language model to obtain a model response.

[0060] In view of the above-mentioned retrieval enhancement generation model, this example provides an error location method for a retrieval enhancement generation system, the method comprising:

[0061] Step S1, obtaining the intermediate execution result of the retrieval enhancement generation system, wherein the intermediate execution result includes: user question Q (Question), original retrieval related document RT_C (Retrieval Context), rewritten retrieval related document RW_C (Rewritten Context), reranked related document RR_C (Reranked Context), model response Rep (Response), and standard answer GT (Ground Truth).

[0062] Step S2, insert the original search related documents into the first prompt word template, input into the first language model, extract the original search fact triples, such as Figure 3As shown; insert all original retrieval fact triples into the second prompt word template and input them into the second largest language model to determine whether all original retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, the retriever in the retrieval enhancement generation system is determined to be normal; otherwise, the retriever is determined to be abnormal, such as Figure 4 shown.

[0063] The fact triples are in JSON format and are expressed in the form of (subject, predicate, object), for example: (Beijing is the capital of China).

[0064] Step S3, insert the rewritten retrieval related documents into the first prompt word template, input it into the first largest language model, and extract the rewritten retrieval fact triples; insert all the rewritten retrieval fact triples into the second prompt word template, input it into the second largest language model, and judge whether all the rewritten retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, it is determined that the rewriter in the retrieval enhancement generation system is normal; otherwise, it is determined that the rewriter is abnormal.

[0065] Step S4, inserting the reordering related documents into the first prompt word template, inputting into the first largest language model, extracting the reordering fact triples; inserting all the reordering fact triples into the second prompt word template, inputting into the second largest language model, judging whether all the reordering fact triples can deduce answers to answer the user's questions; if the answers can be derived, it is judged that the rewriter in the retrieval enhancement generation system is normal; otherwise, it is judged that the rewriter is abnormal;

[0066] Calculate the average relevance score of the original search related documents and the average relevance score of the re-ranked related documents; when the average relevance score of the original search related documents is less than the average relevance score of the re-ranked related documents, the re-ranker is judged to be normal; otherwise, the re-ranker is judged to be abnormal.

[0067] The expressions for calculating the average relevance score of the original search related documents and the average relevance score of the re-ranked related documents are as follows:

[0068]

[0069]

[0070] In the formula, represents the average relevance score of the original retrieved related documents, represents the i-th original search related document, represents the average relevance score of the re-ranked related documents, Represents the i-th re-ranked related document.

[0071] It should be noted that when the average relevance score of the original retrieved related documents is less than the average relevance score of the re-ranked related documents, , indicating that the re-ranked related documents have higher factual relevance, the re-ranker is judged to be normal; when the average relevance score of the original retrieved related documents is greater than the average relevance score of the re-ranked related documents, , indicating that the original retrieved relevant documents have better factual relevance than the re-ranked relevant documents, and the re-ranker is judged to have failed.

[0072] Step S5, input the user question, model response, standard answer, and re-ordered fact triples into the third language model to determine the accuracy, completeness, derivability, and fluency of the model response, so as to determine whether the large language model in the retrieval enhancement generation system is abnormal, such as Figure 5 shown.

[0073] Among them, accuracy, completeness, derivability and fluency have the following meanings:

[0074] Accuracy: Determine whether the generated answer is consistent with the standard answer and has no error messages.

[0075] Completeness: Whether the generated answer contains all the necessary information required to answer the standard answer.

[0076] Derivability: Whether the generated answer can be derived from the extracted evidence triples.

[0077] Fluency: Whether the generated response is linguistically fluent, understandable, and free of grammatical errors.

[0078] Furthermore, when the model responds inaccurately or incompletely, the reasons for the inaccuracy or incompleteness are marked, such as "insufficient retrieval" or "large language model bias in retrieval-augmented generation systems."

[0079] In summary, the present invention analyzes and detects each module in the RAG system (retrieval, rewriter, reorderer, large language model generation module) one by one, and accurately points out the specific module where the system performance is degraded, thereby achieving module-level error location. The present invention uses a large language model to comprehensively evaluate the information interaction of multiple modules in the RAG system and identify problems caused by information loss or information deviation.

[0080] Accordingly, an embodiment of the present invention provides an error location system for a retrieval enhancement generation system, the system comprising:

[0081] The intermediate execution result acquisition module is used to obtain the intermediate execution results of the retrieval enhancement generation system, and the intermediate execution results include: user questions, original retrieval related documents, rewritten retrieval related documents, re-ordered related documents, model responses, and standard answers.

[0082] The fact triple extractor, i.e., the first language model, takes the first prompt word template inserted into the original search-related document as input to extract the original search fact triple;

[0083] Taking the first prompt word template inserted into the rewriting retrieval related document as input, extracting the rewriting retrieval fact triple;

[0084] Taking the first hint word template inserted into the re-ranking related documents as input, re-ranking fact triples are extracted.

[0085] The retrieval module verifier, i.e., the second language model, takes the second prompt word template inserted into all the original retrieval fact triples as input, and determines whether all the original retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, the retriever in the retrieval enhancement generation system is judged to be normal; otherwise, the retriever is judged to be abnormal;

[0086] Taking the second prompt word template inserted with all rewritten retrieval fact triples as input, it is determined whether all rewritten retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, it is determined that the retriever in the retrieval enhancement generation system is normal; otherwise, it is determined that the retriever is abnormal;

[0087] Taking the second prompt word template inserted into all reordered fact triples as input, it is determined whether all reordered fact triples can deduce answers to answer user questions; if the answers can be derived, the retriever in the retrieval enhancement generation system is determined to be normal; otherwise, the retriever is determined to be abnormal.

[0088] The generation module verifier takes user questions, model responses, standard answers, and reordered fact triples as input to determine the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval-enhanced generation system is abnormal.

[0089] Example 4

[0090] In this example, the retrieval enhancement generation system RAG object ω to be tested includes a rewriter, a retriever and a large language model; after the user question is processed by the rewriter, a rewritten question is obtained, and the user question and the rewritten question are retrieved by the retriever to obtain original retrieval-related documents and rewritten retrieval-related documents, and the user question, original retrieval-related documents and rewritten retrieval-related documents are inserted into a prompt word template and input into the large language model to obtain a model response.

[0091] Rewriter: Use the Meta / llama3-70b-instruct model as a rewriter; rewrite the user query to make it clearer without changing the original meaning of the sentence, and return 2 different rewritten queries;

[0092] Retriever: Use sentence-transformers / all-mpnet-base-v2 embedding model +FAISS as a vector retrieval tool, select the 10 closest retrieval texts for 1 original query and 2 rewritten queries respectively, and return a total of 30 original retrieval texts as the initial screening retrieval texts;

[0093] Reranker: Use the BAAI / bge-reranker-large model as the reranker, and select the top five samples as refined retrieval texts based on the relevance scores returned by the reranker;

[0094] Generator: Use the Meta / llama3-70b-instruct model as the generator; in this example, set temperature=0.1; max_tokens=100; do_sample=False as the model sampling parameters.

[0095] The present invention provides an error location system for a retrieval enhancement generation system, the system comprising: a fact triple extractor (Extractor), a retrieval module checker (Retriever_checker), and a generation module checker (Generator_checker); in this example, the fact triple extractor, the retrieval module checker, and the generation module checker all use GPT-4-Turbo as a base large model.

[0096] The present invention provides an error location method for a retrieval enhancement generation system, the method comprising:

[0097] Step S1, record the log of the detection object retrieval enhancement generation RAG system ω, obtain the intermediate execution result of the retrieval enhancement generation system and retain the intermediate execution result in the dictionary, the intermediate execution result includes: user question Q (Question), original retrieval related document RT_C (Retrieval Context), rewritten retrieval related document RW_C (Rewritten Context), reranked related document RR_C (Reranked Context), model response Rep (Response), standard answer GT (Ground Truth).

[0098] Take the following sample as an example:

[0099] User Question Q (Question): "The actor who played Phileas Fogg in Around the World in Eighty Days co-starred with Gary Cooper in the 1939 Goldwyn Productions film based on which author's novel?";

[0100] Original retrieval document RT_C (Retrieval Context): The 10 most relevant documents retrieved for the original question;

[0101] Rewrite retrieval document RW_C (Rewritten Context): retrieve the 20 most relevant documents for the rewritten question;

[0102] Reranked Context (RR_C): The original and rewritten retrieval documents are merged and ranked to get the top five documents.

[0103] Model response Rep(Response): "Jules Verne"

[0104] Standard answer GT (Ground Truth): "Charles L. Clifford".

[0105] It should be noted that the retrieval module in the retrieval enhancement generation system RAG first retrieves the 30 most relevant documents from the external knowledge base as context. The content recorded in the log to the retrieval includes different versions of the movie "Around the World in 80 Days" and the description of the role of actor David Niven in the movie, but no information related to Gary Cooper or the 1939 Goldwyn Productions movie was found. After passing through the rewriter, the retrieval module first retrieves the 30 most relevant documents from the external knowledge base as context. The content recorded in the log to the retrieval includes different versions of the movie "Around the World in 80 Days" and the description of the role of actor David Niven in the movie, but no information related to Gary Cooper or the 1939 Goldwyn Productions movie was found. Since the retrieval module failed to retrieve enough information to answer the user's query, the context after being filtered by the reranker still cannot contain relevant information to answer the user's question. Therefore, the final response returned by the retrieval enhancement generation RAG system ω is "Jules Verne", rather than the actual standard answer "Charles L. Clifford".

[0106] Step S2, inserting the original search related document into the first prompt word template, inputting it into the first language model, and extracting the original search fact triples;

[0107] In this example, the original retrieval fact triples obtained are as follows:

[0108] ("David Niven," "Playing," "Phileas Fogg in Around the World in 80 Days")

[0109] ("Around the World in 80 Days", "A huge success", "At the box office")

[0110] (“Around the World in 80 Days”, “won”, “the Academy Award for Best Picture”.)

[0111] All original retrieval fact triples are inserted into the second prompt word template and input into the second largest language model to determine whether all original retrieval fact triples can deduce answers to answer user questions; if answers can be derived, the retriever in the retrieval enhancement generation system is judged to be normal; otherwise, the retriever is judged to be abnormal.

[0112] The results show that the original retrieval fact triple only mentions David Niven, the actor who played Phileas Fogg. However, there is no information linking these two actors to the 1939 Goldwyn Productions film or to Gary Cooper. The evidence also does not mention the 1939 film or its author, so the evidence does not contain enough information to determine the author of the novel on which the 1939 film starring Gary Cooper was based, and the final answer cannot be inferred based on the original retrieval fact triple extracted by the retriever. Therefore, the retriever failed to provide enough information to answer the user query, and the evaluation result is retriever anomaly.

[0113] Step S3, inserting the rewriting retrieval related document into the first prompt word template, inputting it into the first language model, and extracting the rewriting retrieval fact triple T_1;

[0114] In this example, the rewritten document fact triple T_1 is obtained as follows:

[0115] ("David Niven," "Playing," "Phileas Fogg in Around the World in 80 Days")

[0116] ("Philea Fogg," "Played by," "Steve Coogan")

[0117] ("Around the World in 80 Days", "Won", "Oscar for Best Picture")

[0118] All rewritten retrieval fact triples T_1 are inserted into the second prompt word template and input into the second largest language model to determine whether all rewritten retrieval fact triples can deduce answers to answer user questions; the results show that in this instance, the rewritten retrieval fact triples cannot deduce answers to answer user questions, the rewriter fails to provide additional information to answer user queries, and the evaluation result is that the rewriting module fails.

[0119] Step S4, inserting the reordering related documents into the first prompt word template, inputting into the first language model, and extracting the reordering fact triple T_2;

[0120] In this example, the rewritten document fact triple T_2 is obtained as follows:

[0121] ("Philea Fogg," "Played by," "Steve Coogan")

[0122] ("Around the World in 80 Days", "Adapted from", "Jules Verne's 1873 novel")

[0123] ("David Niven," "Playing," "Phileas Fogg in Around the World in 80 Days")

[0124] All reordered fact triples T_2 are inserted into the second prompt word template and input into the second largest language model to determine whether all reordered fact triples can deduce answers to answer the user's question; in this example, the reordered context contains more detailed information about "Around the World in 80 Days", such as the performances of Steve Coogan and David Niven in different versions, but the answer cannot be derived to answer the user's question, so the rewriter is determined to be abnormal.

[0125] Calculate the average relevance score of the original search related documents and the average relevance score of the re-ranked related documents; in this example, the average relevance score of the original search related documents is 0.2, and the average relevance score of the re-ranked related documents is 0.4. The average relevance score of the re-ranked related documents is higher than the average relevance score of the original search related documents, and the re-ranker is determined to be normal.

[0126] Step S5, input the user question Q, model response Rep, standard answer GT, and reordered fact triple T_2 into the third largest language model to determine the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval enhancement generation system is abnormal.

[0127] In this case, the generation module generated "Jules Verne" as the answer, while the standard answer is "Charles L. Clifford". Based on the extracted fact triples and contextual content, this answer is obviously wrong because the retrieved content does not mention anything related to Gary Cooper or the author of the 1939 Goldwyn Productions film. Therefore, the evaluation criteria of the generation module show:

[0128] Accuracy: 0 (answer incorrect)

[0129] Completeness: 0 (necessary information is not included)

[0130] Derivability: 0 (cannot be deduced from the evidence)

[0131] Fluency: 1 (fluent in language expression)

[0132] The reason for incompleteness was marked as: “Insufficient retrieval”.

[0133] Combined with the analysis results in the above steps, this example determines that the error mainly comes from the retrieval module failing to find content related to Gary Cooper and the 1939 movie in the query, resulting in the generation module generating an incorrect answer. Therefore, the present invention can effectively identify the error module and help developers improve specific modules to improve system performance and stability.

[0134] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the error location method for the retrieval enhancement generation system as described above. Figure 6 As shown, a hardware structure diagram of any device with data processing capability for the error location method for the retrieval enhancement generation system provided by the embodiment of the present invention is shown, except Figure 6 In addition to the processor, memory and network interface shown, any device with data processing capability in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capability, which will not be described in detail.

[0135] Correspondingly, the present application also provides a computer-readable storage medium on which computer instructions are stored, and when the instructions are executed by the processor, the error location method for the retrieval enhancement generation system as described above is implemented. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.

[0136] Those skilled in the art will readily appreciate other embodiments of the present application after considering the description and practicing the contents disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application. The description and examples are intended to be exemplary only.

[0137] It should be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. An error location method for a retrieval enhancement generation system, characterized in that: The method comprises: a retrieval-oriented enhanced generation system comprising a rewriter, a retriever, a reranker and a large language model; Obtaining an intermediate execution result of the retrieval enhancement generation system, wherein the intermediate execution result includes: user questions, original retrieval-related documents, rewritten retrieval-related documents, re-ranked related documents, model responses, and standard answers; Insert the original search related documents into the first prompt word template, input into the first language model, and extract the original search fact triples; insert all the original search fact triples into the second prompt word template, input into the second language model, and judge whether all the original search fact triples can deduce answers to answer the user's questions; if the answers can be derived, it is determined that the retriever in the search enhancement generation system is normal; otherwise, it is determined that the retriever is abnormal; Insert the rewritten retrieval related documents into the first prompt word template, input into the first language model, and extract the rewritten retrieval fact triples; insert all the rewritten retrieval fact triples into the second prompt word template, input into the second language model, and judge whether all the rewritten retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, it is determined that the rewriter in the retrieval enhancement generation system is normal; otherwise, it is determined that the rewriter is abnormal; Insert the re-ranking related documents into the first prompt word template, input into the first largest language model, and extract the re-ranking fact triples; insert all the re-ranking fact triples into the second prompt word template, input into the second largest language model, and judge whether all the re-ranking fact triples can deduce answers to answer the user's questions; if the answers can be derived, it is determined that the re-ranker in the retrieval enhancement generation system is normal; otherwise, it is determined that the re-ranker is abnormal; Calculate the average relevance score of the original search related documents and the average relevance score of the re-ranked related documents; when the average relevance score of the original search related documents is less than the average relevance score of the re-ranked related documents, the re-ranker is judged to be normal; otherwise, the re-ranker is judged to be abnormal; The user question, model response, standard answer, and re-ordered fact triples are input into the third largest language model to determine the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval enhancement generation system is abnormal.

2. The error location method for a retrieval enhancement generation system according to claim 1, characterized in that: Fact triples are expressed in the form of (subject, predicate, object).

3. The error location method for a retrieval enhancement generation system according to claim 1, characterized in that: The method further comprises: When the model response is inaccurate or incomplete, mark the reasons for the inaccuracy or incompleteness.

4. An error localization system for a retrieval enhancement generation system, characterized in that: To implement the error location method for a retrieval enhancement generation system as described in any one of claims 1 to 3, the system comprises: An intermediate execution result acquisition module is used to acquire the intermediate execution result of the retrieval enhancement generation system, wherein the intermediate execution result includes: user questions, original retrieval-related documents, rewritten retrieval-related documents, re-ranked related documents, model responses, and standard answers; The fact triple extractor, i.e., the first language model, takes the first prompt word template inserted into the original retrieval-related document as input to extract the original retrieval fact triple; takes the first prompt word template inserted into the rewritten retrieval-related document as input to extract the rewritten retrieval fact triple; takes the first prompt word template inserted into the reordering-related document as input to extract the reordering fact triple; The retrieval module verifier, i.e., the second language model, takes the second prompt word template inserted into all the original retrieval fact triples as input, and determines whether all the original retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, the retriever in the retrieval enhancement generation system is judged to be normal; otherwise, the retriever is judged to be abnormal; Taking the second prompt word template inserted into all rewritten retrieval fact triples as input, it is determined whether all rewritten retrieval fact triples can deduce answers to answer user questions; if the answers can be derived, it is determined that the retriever in the retrieval enhancement generation system is normal; otherwise, it is determined that the rewriter is abnormal; Taking the second prompt word template inserted into all re-ordered fact triples as input, it is determined whether all re-ordered fact triples can deduce answers to answer user questions; if the answers can be derived, it is determined that the retriever in the retrieval enhancement generation system is normal; otherwise, it is determined that the re-ranker is abnormal; The generation module verifier takes user questions, model responses, standard answers, and reordered fact triples as input to determine the accuracy, completeness, derivability, and fluency of the model response, thereby determining whether the large language model in the retrieval-enhanced generation system is abnormal.

5. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the error location method for the retrieval enhancement generation system as described in any one of claims 1-3 above.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the error localization method for a retrieval enhancement generation system as described in any one of claims 1 to 3 is implemented.

7. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the error localization method for a retrieval enhancement generation system described in any one of claims 1-3 is implemented.

Citation Information

Patent Citations

  • Universal adaptive question answering method and system, storage medium and electronic equipment

    CN118643144A