Dynamic and static combined retrieval enhancement generation method and equipment
Through the search enhancement generation method combining dynamic and static search, combined with static search and dynamic search, the problems of unreasonable search frequency and limited information coverage in the existing technology are solved, and more efficient and better generation quality is achieved, and performance improvements are achieved to adapt to complex tasks.
Patent Information
- Application Number
- CN202510653237.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
In the prior art, the static RAG technology has unreasonable search frequency, limited coverage of the search content, and the dynamic RAG technology relies on fixed rules to determine the search timing, and the ability to capture real-time information of the model is limited, resulting in poor generation results.
The search-enhanced generation method of dynamic and static combination is adopted to generate preliminary answers through static search, and the information needs are detected in real time during the generation process, and the search timing and query content are dynamically adjusted to ensure that the model can obtain supplementary information in a timely manner during the generation process.
Improves generation quality and efficiency, reduces the number of dynamic searches, ensures a wider knowledge coverage, improves performance to adapt to complex tasks, and reduces the "illusion" phenomenon.
Smart Images

Figure CN120179795A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and particularly relates to a retrieval-augmented generation method and device combining dynamic and static aspects. Background Art
[0002] With the wide application of large language models (LLMs) in the field of natural language processing (NLP), they have demonstrated powerful capabilities in scenarios such as dialogue generation, question-and-answer systems, and knowledge reasoning. However, they may generate some content that is coherent and reasonable but does not conform to reality.
[0003] Retrieval-augmented generation (RAG) is an effective solution to address this issue. It enhances the LLM by retrieving relevant information from an external database and integrating it into the LLM's input. Existing methods typically use static rules to determine when to trigger retrieval or generate queries based on the most recently generated content. Although this approach is effective for simple tasks, it is often not suitable for complex multi-step tasks and long-form generation tasks. Moreover, frequent triggering of retrieval may lead to the introduction of a large amount of redundant or irrelevant information, which instead interferes with the model's generation process and reduces the output quality. Dynamic RAG performs multiple retrievals during the LLM's generation process. It consists of two steps: determining the optimal timing to activate the retrieval module (deciding when to retrieve), and designing an appropriate query after triggering the retrieval (determining what to retrieve). Currently, FL-RAG is a multi-round retrieval-augmentation method that triggers the retrieval module every n tokens. The tokens generated in the previous token window are used as the query. FS-RAG is also a multi-round retrieval-augmentation method that triggers the retrieval module for each sentence. The last generated sentence is used as the query.
[0004] Existing dynamic RAG techniques rely on fixed rules to determine the retrieval timing. Most existing dynamic RAG techniques trigger retrieval when generating each sentence, a fixed number of tokens, or when the model's generation confidence is low. However, they can only respond to a fixed range of rules and lack in-depth dynamic analysis of the generated content, which may still result in inaccurate triggering. Most existing dynamic RAG techniques may still waste resources due to inaccurate triggering under simple rules. Unnecessary retrieval augmentation may introduce irrelevant or noisy data to the LLM, thereby endangering the quality of the output. Summary of the Invention
[0005] In view of the above analysis, the present invention aims to provide a retrieval-augmented generation method and device combining dynamic and static aspects, which are used to solve the problems in the prior art that the retrieval frequency of static RAG techniques is unreasonable, the coverage of retrieved content is limited, and dynamic RAG techniques rely on fixed rules to determine the retrieval timing and have limitations in capturing real-time information of the model, resulting in poor generation results.
[0006] The object of the present invention is mainly achieved by the following technical solutions: On the one hand, the present invention provides a retrieval augmented generation method combining dynamic and static, including: S1: Based on the initial question, generate a preliminary answer through static retrieval; S2: Start the answer output from the starting token of the preliminary answer, and perform real-time information demand detection; the real-time information demand detection includes determining whether to trigger dynamic retrieval based on the current trigger threshold and the retrieval trigger index of the current token to be output; S3: If it is necessary to trigger dynamic retrieval, generate a retrieval query Query for dynamic retrieval, and continue to output the answer based on the dynamic retrieval result and dynamically update the trigger threshold; S4: Based on the dynamically updated trigger threshold, continue to judge whether to trigger new dynamic retrieval through real-time information demand detection, and repeat S3 until each token does not need to trigger new retrieval, and complete the answer output.
[0007] Further, the real-time information demand detection is performed by the following method to determine whether to trigger dynamic retrieval: Obtain the current trigger threshold; Obtain the retrieval trigger index corresponding to the token based on the uncertainty, context influence value and semantic importance of the current token to be generated; If the retrieval trigger index is greater than the trigger threshold, trigger dynamic retrieval.
[0008] Further, the retrieval trigger index is expressed as: ; where is the retrieval trigger index of the token at position , represents the uncertainty of the token at position , represents the context influence value of the token at position , represents the semantic importance of the token at position .
[0009] Further, the uncertainty of the corresponding token is characterized by the entropy value of each token generation position, expressed as: ; where represents the probability distribution of the token at position , and V is the vocabulary; The maximum attention value of the current token for subsequent generation represents the context influence value, expressed as: ; where represents the attention value of the -th token to the -th token; The semantic importance is expressed as: ; where is the stop word set.
[0010] Furthermore, based on the normalized confidence of the current dynamic retrieval result, the trigger threshold is dynamically updated, expressed as: ; where is the updated trigger threshold, is the base threshold, i.e., the threshold updated in the previous round of retrieval, is the hyperparameter controlling the threshold change amplitude, and Confidence is the normalized confidence of the current dynamic retrieval result.
[0011] Furthermore, the confidence is obtained by the following method: Calculate the entropy value of each generation position in the dynamic retrieval result; and obtain the average entropy value based on the entropy value of each generation position; Normalize the average entropy value to obtain the normalized average entropy value , taking as the confidence of the generated answer, i.e.: .
[0012] Furthermore, the average entropy value is expressed as: ; The normalization of the average entropy value is expressed as: ; where is the average entropy value, is the normalized average entropy value.
[0013] Furthermore, the retrieval query Query is generated by the following method: Based on the attention weights between the tokens of the generated answer, obtain the attention weight vector of each generation position; Select the top n tokens most relevant to the current generation position according to the attention weight vector; Arrange the selected n tokens in the order in the original context, and generate a retrieval query Query together with the user's initial question.
[0014] Furthermore, after generating the retrieval query Query, submit the Query to a retrieval engine, retrieve document fragments related to the Query, and sort them according to the relevance scores of the document fragments, and filter the top k most relevant retrieval results; append the k retrieval results to the input sequence in order to obtain a dynamic retrieval result.
[0015] On the other hand, a computer device is also disclosed, including at least one processor and at least one memory communicatively connected to the processor; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the aforementioned retrieval enhanced generation method combining dynamic and static.
[0016] Advantages of this technical solution: 1. The method provided by the present invention has higher efficiency and better generation quality, through static retrieval to filter preliminary information and dynamic threshold adjustment strategy. Compared with general dynamic RAG, it can reduce the number of dynamic retrievals and improve the overall system efficiency. By providing basic information through static RAG, it ensures that the preliminary answers generated have a wide range of knowledge coverage; the dynamic retrieval mechanism of dynamic RAG can fill knowledge gaps during the generation process and prevent the model from generating unreliable answers due to insufficient information. The threshold adaptive adjustment strategy achieves a dynamic balance between efficiency and generation quality.
[0017] 2. The retrieval enhanced generation method of the present invention greatly improves the performance of complex tasks. In multi-step reasoning or long text generation tasks, dynamic retrieval can significantly reduce the generation uncertainty of the model and reduce the "hallucination" phenomenon.
[0018] 3. The retrieval enhanced generation method of the present invention has a wider information coverage. Through the combination of static RAG and dynamic RAG, static RAG provides basic background information, and dynamic RAG is responsible for filling knowledge gaps to adapt to the dynamic information needs of the LLM during the generation process. Description of the Drawings
[0019] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components; Figure 1 is a flowchart of the retrieval enhanced generation method combining dynamic and static of the embodiment of the present invention; Figure 2 is a schematic application diagram of the retrieval enhanced generation method combining dynamic and static of the embodiment of the present invention. Detailed Implementation Modes
[0020] The preferred embodiments of the present invention will be specifically described below with reference to the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, rather than to limit the scope of the present invention.
[0021] An embodiment of the present invention provides a method for retrieval-augmented generation that combines dynamic and static methods, as Figure 1 shown, including: S1: Based on the initial question, generate a preliminary answer through static retrieval; Specifically, in this embodiment, an efficient retrieval large language model (LLM) is constructed based on the Transformer model to achieve efficient retrieval-augmented generation that combines dynamic and static methods. First, through the static retrieval module of the efficient retrieval large language model, a traditional single-round static RAG (retrieval-augmented generation) method is used to generate a preliminary answer to provide preliminary knowledge background support for the large language model (LLM). In practical applications, taking the initial question input by the user as a query, relevant documents are retrieved from an external knowledge base (such as Wikipedia). The retrieved documents and the question are used as the input of the static retrieval module of the large language model to generate a preliminary answer.
[0022] In this embodiment, a preliminary answer is generated through static retrieval, providing basic background information for dynamic retrieval, and solving the problem that existing LLMs may generate guesses or content that does not conform to facts due to lack of sufficient knowledge background or inaccurate embedded knowledge bases during the generation process. This "hallucination" phenomenon results in the generated results not meeting the high reliability requirements, especially performing poorly in knowledge-intensive tasks that require precise answers.
[0023] S2: Start answer output from the start token of the preliminary answer and perform real-time information demand detection; the real-time information demand detection includes determining whether to trigger dynamic retrieval based on the current trigger threshold and the retrieval trigger index of the current token to be output; Specifically, the following method is used for real-time information demand detection to determine whether to trigger dynamic retrieval: Obtain the current trigger threshold; Obtain the retrieval trigger index corresponding to the current token to be generated based on the uncertainty, context influence value, and semantic importance of the current token; If the retrieval trigger index is greater than the trigger threshold, trigger dynamic retrieval.
[0024] Among them, the retrieval trigger index is expressed as: ; Among them, is the position The retrieval trigger index of the token, indicating the position of the uncertainty of the token, indicating the position of the context influence value of the token, indicating the position of the semantic importance of the token.
[0025] Preferably, the entropy value of each token generation position is used to characterize the uncertainty of the corresponding token, expressed as: ; wherein, represents the probability distribution of the token at the position , V is the vocabulary; the higher the entropy, the more uncertain the model is about the token; Using the self-attention mechanism of the large prediction model, the maximum attention value generated by the current token for the subsequent generation is calculated to characterize the context influence value, expressed as: ; wherein, represents the attention value of the th token to the th token; Assign a semantic importance label to each token , and eliminate tokens with low semantic contributions such as stop words. The semantic importance is expressed as: ; wherein, is the stop word set (for example: um, de, ba, etc., and tokens without obvious semantics such as English articles and prepositions).
[0026] In this embodiment, starting from the preliminary answer generated by the static RAG, a multi-round dynamic retrieval mechanism of the dynamic RAG is used to optimize the output answer, and the "real-time information requirement detection module" (DTIND: Dynamic threshold real-time information requirement detection) of the dynamic RAG module of the large language model is used to detect the uncertainty and information requirements during the generation process. Trigger the retrieval module to perform targeted queries to ensure that the model can obtain supplementary information in real time in complex, multi-step reasoning tasks.
[0027] The real-time information demand detection module DTIND is responsible for dynamically determining whether a new retrieval needs to be triggered, ensuring that the retrieval module is triggered only when it is truly necessary, thereby improving efficiency. DTIND comprehensively evaluates three factors: the uncertainty of the token to be generated, the context influence value, and the semantic importance, and calculates a comprehensive retrieval trigger index for each token to be generated. When the retrieval trigger index of a certain token exceeds the current threshold, the retrieval module is triggered to perform a retrieval. If the retrieval trigger index does not exceed the current threshold, the retrieval is not triggered, and the answer output continues.
[0028] In this embodiment, through the introduction of the real-time information demand detection (DTIND) module, the generation process is analyzed dynamically in multiple dimensions, overcoming the limitations of fixed rules. By combining the uncertainty (entropy value) of generation, semantic importance (removing stop words), and context influence (attention distribution), it comprehensively evaluates whether each generated token needs to trigger a retrieval. By setting a trigger threshold based on confidence, the retrieval mechanism is triggered for local keywords above the threshold, which can more accurately determine the retrieval timing and avoid the phenomena of inaccurate triggering or over-retrieval.
[0029] S3: If a dynamic retrieval needs to be triggered, a retrieval query Query is generated for dynamic retrieval, and based on the dynamic retrieval results, the answer is continued to be output and the trigger threshold is dynamically updated. Specifically, in this embodiment, a retrieval query generation module (RQG module) based on self-attention is constructed, and the retrieval query Query is generated by the following method: An attention weight vector for each generation position is obtained based on the attention weights between the tokens of the generated answer. According to the attention weight vector, the top n tokens most relevant to the current generation position are selected. The selected n tokens are arranged in the order in the original context and are used together with the user's initial question to generate the retrieval query Query.
[0030] After generating the retrieval query Query, the Query is submitted to the retrieval engine to retrieve document fragments related to the Query, and the retrieved results are sorted according to the relevance scores of the document fragments, and the top k most relevant retrieval results are selected; the k retrieval results are appended to the user question in order as the context input of the prompt to obtain the corresponding dynamic retrieval result as the prompt input.
[0031] It should be noted that the design goal of the Retrieval Query Generation module (RQG) is to ensure that each retrieval can generate a query highly relevant to the model's information needs, thereby improving the accuracy of retrieval and generation. This module extracts the most valuable information from the context through the self-attention mechanism to construct the retrieval query. If a certain token , it means that this token needs to be output separately, and after retrieval enhancement, the output continues. During retrieval, a retrieval query Query needs to be used for retrieval. The generation method of this Query is as follows: The Retrieval Query Generation RQG uses the self-attention mechanism to extract the tokens most relevant to the current information needs from the context and constructs the Query together with the initial question: First, extract the self-attention matrix A: The last layer of the large language model based on the Transformer model generates the self-attention matrix A, and its elements represent the attention weight of the th token to the th token; Calculate the attention weight vector at the generation position : ; Among them, is the attention score of the model at position for the attention position , indicating the degree of dependence of the content generated at position on the context at position .
[0032] After obtaining the attention weight vector, select the most relevant tokens, that is, sort in descending order according to the attention weight vector , and select the top n tokens most relevant to the current position to ensure that the query can accurately reflect the real-time information needs of the model; Furthermore, sort the attention weights in descending order to obtain the index set of the top n tokens with the highest weights, which is expressed as: ; Among them, represents sorting in descending order; Then, The selected tokens are arranged in the order in which they appear in the original context to maintain the semantic coherence of the Query.
[0033] Exemplarily, as Figure 2 shown, assume that after the user asks "Please introduce Newton", the model first performs a preliminary static retrieval based on the user's question "Please introduce Newton" as the Query, and uses the retrieved results as context prompt words prompt to the model to generate a preliminary answer. So far, it is the traditional RAG, that is, static RAG.
[0034] The model's preliminary answer: "Isaac Newton is a famous British physicist..., and his research laid the foundation for basic physics..." When the model generates "basic physics", it detects that external retrieval is needed (here the token ); In the attention extraction results, "Isaac Newton", "physicist", "laid", and "basic physics" are the top 4 tokens with the most concentrated attention; (the number of tokens can be set by yourself as a hyperparameter).
[0035] Then generate the Query: "Isaac Newton physicist laid basic physics" (not a sentence, just some keywords) + "Please introduce Newton" (the user's initial question), that is, "Isaac Newton physicist laid basic physics Please introduce Newton".
[0036] Submit the generated Query to the knowledge base or document database for retrieval, such as submitting the Query constructed by RQS to the retrieval system (such as the Elasticsearch retrieval engine); retrieve document fragments related to the Query based on traditional methods such as BM25, or use dense vector retrieval (such as DPR, ColBERT) to obtain semantically similar documents; sort according to the relevance scores of the document fragments and select the top k most relevant results; Dynamically fuse the retrieval results with the context during the generation process, so that the language model can continue to generate based on the newly obtained external information. Append the retrieval results to the input sequence in order and continue to generate. The following is an example of the overall process: User's question: Please help me introduce Einstein; The large language model performs static RAG. The model first uses the user's question "Please help me introduce Einstein" as the Query for preliminary static retrieval, and uses the retrieved results as context prompt words prompt to the large language model to generate the original sequence (that is, the preliminary answer): "Albert Einstein was born in Ulm and later moved to Switzerland. During his studies, Einstein's research results..." When generating "Einstein's research achievements" through the DTIND module, it is detected that an external search for the token "achievements" is required. , Then, through the RQS module, the attention extraction results "Einstein", "research", and "Switzerland" are the top 3 tokens with the most concentrated attention (the number of tokens is a hyperparameter that can be set by oneself); together with the original question, a Query is generated: "Einstein, research, Switzerland Please introduce Einstein to me" (not a complete sentence, just keyword tokens + the original question); Based on the Query, it is submitted to the knowledge base or document database for retrieval, and the retrieved content is added to the prompt (the prompt words of the large model); The updated prompt is input into the large model. At this time, the prompt contains the retrieved relevant content as knowledge, and starting from the position where the retrieval began ("Einstein's research achievements"), the answer continues to be generated.
[0037] As a specific embodiment: User's question: Please introduce Einstein to me; The model first uses the user's question "Please introduce Einstein to me" as a Query for preliminary static retrieval, and takes the retrieved result as the context prompt words prompt to the large language model to generate the original sequence (i.e., the preliminary answer): Albert Einstein was born in Ulm and later moved to Switzerland. During his studies, Einstein's research achievements... DTIND detects that the retrieval trigger index of the token "achievements" is greater than the current trigger threshold ( ), pauses the generation, and conducts a retrieval; Based on the retrieval result, the threshold is dynamically updated, and after "Einstein's research achievements", the new retrieval result "which made him later become a key figure in shaping 20th-century science, although he completely revolutionized modern science through his groundbreaking contributions to physics and the development of the theory of relativity. In 1905, his Annus Mirabilis papers introduced revolutionary ideas, including the special theory of relativity" is continued to be output (here the of "special theory of relativity"), (triggering retrieval again, dynamically updating the threshold with the above method, and continuing to output) , where E represents energy, m represents mass, and c represents the speed of light in a vacuum (about 3×108 m / s). This formula states that the total energy of an object is proportional to its mass, and the proportionality constant is the square of the speed of light. Reveals the profound unity between mass and energy, laying the theoretical foundation for nuclear physics and high-energy physics. Understanding its applicability and physical environment is crucial...", until each token is traversed and no new retrieval needs to be triggered (subsequent generation ) and complete the output.
[0038] Furthermore, based on the confidence of the current dynamic retrieval result, the triggering threshold is dynamically updated, expressed as: ; where is the updated triggering threshold, is the base threshold, i.e., the threshold updated in the previous round of retrieval, is the hyperparameter controlling the change amplitude of the threshold, and Confidence is the normalized confidence of the current dynamic retrieval result.
[0039] It should be noted that when confidence = 0.5, maintain the original value; The boundary conditions of the normalized confidence are: ; ; The hyperparameter controlling the change amplitude of the threshold : : No adjustment of the threshold, ; : ; : ; Specifically, the confidence is obtained by the following method: Calculate the entropy value of each generation position in the dynamic retrieval result; and obtain the average entropy value based on the entropy value of each generation position; Normalize the average entropy value to obtain the normalized average entropy value , taking as the confidence of generating the answer.
[0040] The average entropy value is expressed as: ; The normalization of the average entropy value is expressed as: .
[0041] The normalized entropy value is between [0,1], and the higher the value, the lower the confidence Confidence. In this embodiment, Indicates confidence, that is: 。
[0042] It should be specifically noted that in this embodiment, first, an initial basic threshold is set based on the question and the preliminary answer proposed by the user, and then each time a retrieval is triggered, the ; Among them, ; that is, the previous retrieval threshold as the in the dynamic threshold formula to calculate the new 。
[0043] In this embodiment, a dynamic threshold is set by combining the preliminary results of static RAG : Statistically analyze the generation probability or output confidence in the generated answer; if the preliminary answer has a high confidence, then increase the dynamic RAG retrieval trigger threshold to reduce unnecessary retrieval calls. If the preliminary answer has a low confidence, that is, it performs poorly in complex reasoning, then appropriately lower the threshold to trigger the retrieval more frequently, thereby enhancing the answer quality. That is, if the generated result has a high confidence and the preliminary answer has good quality, for example, answering common sense questions (such as "Which country was Einstein born in?"), then increase the initial basic threshold and reduce the dynamic retrieval frequency. If the preliminary answer performs poorly in complex multi-step reasoning tasks (such as questions that require multiple rounds of reasoning), then lower the initial basic threshold to make the dynamic RAG trigger the retrieval more frequently and obtain the necessary external knowledge.
[0044] The present invention uses the static RAG technology during the first generation. The threshold problem is not considered during the initial generation, and a retrieval-enhanced generation must be performed to ensure that the model has complete and relevant knowledge input from the beginning. The model has obtained the context information enhanced by retrieval before generation, avoiding initial errors caused by insufficient embedded knowledge and providing a more reliable basis for subsequent dynamic adjustment. Then, the uncertainty and information requirements during the generation process are detected through the real-time information requirement detection module (DTIND: Dynamic threshold real-time information requirement detection), that is, calculate the , it is determined whether it is greater than the threshold. If it is greater than the threshold, it enters the query generation module. Subsequently, the "Retrieval Query Generation Module" (RQG) dynamically analyzes the global context and optimizes the construction of the retrieval query. Utilizing the self-attention mechanism of the LLM, the RQS preferentially selects tokens with higher attention weights, excludes redundant information, constructs high-quality and closely related queries, ensures that the queries can comprehensively reflect the real-time information requirements of the model, and optimizes the problem of relying on the most recently generated content in existing models.
[0045] Moreover, through the combination of static and dynamic mechanisms, the present invention enables the model to not only handle the large-scale information requirements in the initial stage but also flexibly respond to the detailed changes in the subsequent generation process. Under the joint operation of DTIND and RQS, the ability to capture the real-time information requirements of the model is significantly improved. In terms of real-time dynamic evaluation, the DTIND module dynamically determines whether the model needs to retrieve external knowledge by calculating the uncertainty, semantic value, and context contribution of each generated token. In terms of cross-context information integration, the RQS can capture key information from the entire context and generate accurate retrieval queries based on the current focus of the model, avoiding the limitation of relying only on the most recently generated content and enhancing the ability to capture the real-time information of the model.
[0046] S4: Continue to perform real-time information requirement detection based on the dynamically updated trigger threshold to determine whether a new dynamic retrieval needs to be triggered, and repeat S3 until each token does not require triggering a new retrieval, and then complete the answer output.
[0047] Specifically, in this embodiment, through multiple rounds of dynamic retrieval, each generated token is traversed. When each token does not require triggering a new retrieval, the answer output is completed.
[0048] In actual application, a threshold is adaptively generated based on the confidence of the new retrieval result , and the generation is controlled, that is, if the uncertainty during the generation process continues to increase, the DTIND triggers a new retrieval and Query construction again, forming a cycle of multiple rounds of dynamic retrieval and generation.
[0049] Since the tokens before the token to be retrieved have all been determined, only the subsequent generation output of the current token needs to be updated. Therefore, when the last token is traversed, there is no token greater than the threshold , and the retrieval ends, generating the final result.
[0050] Another embodiment of the present invention also discloses a computer device, including at least one processor and at least one memory communicatively connected to the processor; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the aforementioned retrieval enhanced generation method combining dynamic and static aspects.
[0051] In summary, the retrieval enhanced generation method and device of the present invention use a hybrid strategy that combines the advantages of static RAG and dynamic RAG. Static RAG ensures global coverage for the first generation, while dynamic RAG optimizes the generation details through subsequent dynamic adjustments, thus achieving the combination of comprehensive coverage and efficient dynamic adjustment. On the one hand, unnecessary retrievals are avoided. DTIND triggers retrievals only when truly needed by precisely evaluating retrieval requirements, reducing ineffective operations, thereby reducing resource waste and time overhead. On the other hand, RQS constructs highly relevant retrieval queries to ensure that the introduced data is accurate and effective, avoiding interference from noisy information to the output quality, and greatly optimizing the retrieval generation effect.
[0052] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.
[0053] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A dynamic and static combined retrieval enhancement generation method, characterized in that: include: S1: Based on the initial question, generate preliminary answers through static retrieval; S2: Outputting the answer starting from the initial token of the preliminary answer and performing real-time information demand detection; the real-time information demand detection includes determining whether dynamic retrieval needs to be triggered based on the current trigger threshold and the retrieval trigger index of the current token to be output; S3: If dynamic retrieval needs to be triggered, a retrieval query is generated to perform dynamic retrieval, and the answer is continuously output based on the dynamic retrieval result and the trigger threshold is dynamically updated; S4: Based on the dynamically updated trigger threshold, continue to determine whether a new dynamic search needs to be triggered through real-time information demand detection, and repeat S3 until each token does not need to trigger a new search and the answer is output.
2. The dynamic and static combined search enhancement generation method according to claim 1 is characterized in that: The following methods are used to detect real-time information needs to determine whether dynamic retrieval needs to be triggered: Get the current trigger threshold; Based on the uncertainty, contextual impact value and semantic importance of the token to be generated, the retrieval trigger index of the corresponding token is obtained; If the retrieval trigger index is greater than the trigger threshold, dynamic retrieval is triggered.
3. The dynamic and static combined search enhancement generation method according to claim 2 is characterized in that: The retrieval trigger index is expressed as: ; in, For location The token retrieval trigger index, Indicates location The uncertainty of the token, Indicates location The contextual impact value of the token, Indicates location The semantic importance of the token.
4. The dynamic and static combined search enhancement generation method according to claim 3 is characterized in that: The entropy value of each token generation position represents the uncertainty of the corresponding token, expressed as: ; in, Indicates at location Token The probability distribution of , V is the vocabulary; The maximum attention value of the current token on the subsequent generation is used to represent the context influence value, which is expressed as: ; in, Indicates Token pair The attention value of a token; The semantic importance is expressed as: ; in, is the set of stop words.
5. The dynamic and static combined search enhancement generation method according to claim 2 is characterized in that: Based on the confidence of the current dynamic retrieval results, the trigger threshold is dynamically updated, expressed as: ; in, is the updated trigger threshold, is the basic threshold, that is, the threshold after the previous round of retrieval update, is a hyperparameter that controls the threshold change range, and Confidence is the confidence of the current dynamic retrieval result.
6. The dynamic and static combined search enhancement generation method according to claim 5 is characterized in that: The confidence level is obtained by the following method: Calculate the entropy value of each generated position in the dynamic retrieval result; and obtain an average entropy value based on the entropy value of each generated position; The average entropy value is normalized to obtain a normalized average entropy value ,by As the confidence of the generated answer, that is: Confidence = 1 - Normalized Entropy.
7. The dynamic and static combined search enhancement generation method according to claim 6 is characterized in that: The average entropy value is expressed as: ; The average entropy value is normalized and expressed as: ; in, is the average entropy value, is the normalized average entropy value.
8. The dynamic and static combined search enhancement generation method according to claim 1 is characterized in that: Generate a search query using the following method: Based on the attention weights between the tokens of the generated answer, the attention weight vector of each generated position is obtained; Select the first n tokens most relevant to the current generation position according to the attention weight vector; Arrange the selected n tokens in the order of the original context and generate a retrieval query together with the initial question.
9. The dynamic and static combined search enhancement generation method according to claim 8 is characterized in that: After generating the search query Query, the Query is submitted to the search engine to retrieve document fragments related to the Query, and the document fragments are sorted according to their relevance scores to filter out the top k most relevant search results; the k search results are sequentially appended to the input sequence to obtain dynamic search results.
10. A computer device, characterized in that: comprising at least one processor, and at least one memory communicatively connected to the processor; The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the dynamic and static combined retrieval enhancement generation method described in any one of claims 1-9.
Citation Information
Patent Citations
Multi-stage dynamic retrieval question-answering system and working method thereof
CN118656481A
Large model illusion relieving method and device, equipment and storage medium
CN118964583A
Intelligent question answering method and system based on multi-module collaborative optimization
CN119557409A
Large language model question and answer method, device and equipment based on retrieval enhancement and medium
CN119597874A
Fine-grained image retrieval method based on deep learning
CN119988667A
Cited By
Multi-round reasoning question answering method and system based on query graph driving
CN120687579A
A Query Graph-Driven Multi-Turn Reasoning Question Answering Method and System
CN120687579B