Retrieval generation optimization method, medium and system based on long-term enhancement
By introducing a long-term enhancement mechanism in the search enhancement generation technology, the similarity between text blocks and query problems is calculated and the generation model is injected into batches, the problem of insufficient utilization of useful information is solved, and the accuracy and consistency of generated answers is significantly improved.
Patent Information
- Application Number
- CN202510082759.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Although search-enhanced generation technology (RAG) improves the accuracy of answers by searching external knowledge bases and combining the inference ability of generative models, there is still insufficient utilization of useful information, resulting in inaccurate answers or hallucinatory problems.
A search generation optimization method based on long-term enhancement is adopted. By calculating the similarity between each text block and the query problem, the number of injection times of each text block is determined, and the generation model is iterated in batches until the maximum injection times ends, and the final generated answer is obtained.
It significantly improves the factual accuracy and consistency of generated answers, improves the model's utilization of high-quality information, and reduces the possibility of generating false information.
Smart Images

Figure CN119494410B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of large language generation models, and in particular to a retrieval generation optimization method, medium and system based on long-term enhancement. Background Art
[0002] In traditional language generation models, the model relies on intrinsic parameters and training data to generate answers, but this approach has obvious limitations, especially in specific fields, where the model often cannot provide accurate answers due to knowledge lags or limitations. In order to overcome this problem, Retrieval-augmented Generation (RAG) technology came into being. RAG solves the problem of insufficient knowledge caused by the model relying solely on internal training data by retrieving external knowledge bases and combining them with the reasoning ability of the generative model. The core goal of RAG technology is to provide users with accurate and reliable answers by retrieving relevant information from existing knowledge bases and combining them with generative models, thereby reducing the "hallucination" phenomenon of generated content and ensuring that the text is more factual and accurate. Figure 1 As shown, the RAG system usually adopts the following basic process:
[0003] Indexing: RAG will clean and segment the knowledge base during the initialization phase, cutting longer documents into smaller "chunks" according to semantic units, and then convert these chunks into semantic vectors through the embedding model and create corresponding indexes. This process is done offline and stored in the vector database to speed up subsequent online retrieval;
[0004] Retrieval: When a user asks a query, the system uses the same embedding model to convert the question into a vector, searches the vector block database, and calculates the similarity between the query vector and the document block vector in the database. The system selects the document block with the highest similarity as the enhanced context information for the subsequent generation stage;
[0005] Generate: In the generation phase, RAG merges the retrieved document chunks with the user question to generate a comprehensive prompt for the large language model to answer. If there is historical dialogue information, the system can also merge it into the prompt to support multi-round dialogue generation.
[0006] Based on RAG, Advanced RAG introduces multiple optimization strategies:
[0007] Pre-retrieval optimization: Improve the accuracy of retrieval by optimizing text segmentation, index construction and query rewriting. In particular, semantic segmentation technology can cut long text into smaller and more relevant blocks according to semantic cohesion to avoid information loss and semantic truncation.
[0008] Retrieval optimization: In advanced RAG, the matching degree between query questions and documents is improved by fine-tuning the embedding model or using dynamic embedding technology in the retrieval stage. In addition, hybrid search technology further improves the accuracy of retrieval by combining vector search and keyword search.
[0009] Post-retrieval optimization: Before generation, the retrieved context is optimized through prompt compression and re-ranking, ensuring that the generative model only uses the most relevant information and reducing the chance of generating wrong answers.
[0010] After the above optimization strategies, Advanced RAG has significantly better results than RAG. However, although the system can improve the accuracy of answers by retrieving external knowledge bases and combining the reasoning ability of the generative model and various optimization strategies, the use of useful information is still seriously insufficient. The generative model still fails to accurately grasp the importance of effective information when constructing answers, resulting in inaccurate or hallucinatory answers, that is, generating false and untrue information. How to optimize it to further improve the accuracy of answers is a technical problem that needs to be solved in this field. Summary of the invention
[0011] In order to solve at least one of the above technical problems, the present invention provides a retrieval generation optimization method based on long-term enhancement, comprising:
[0012] S1: According to the query question input by the user, several text blocks are retrieved;
[0013] S2: Calculate the similarity between each text block and the query question;
[0014] S3: Determine the injection times of each text block according to the similarity;
[0015] S4: According to the injection times of each text block, the input generation model is iterated in batches until the highest injection times are reached to obtain the final generated answer.
[0016] Furthermore, step S1 includes:
[0017] S11: Divide the original documents in the knowledge base into blocks and generate context information;
[0018] S12: construct a complete text block containing context information, convert it into a vector representation, and store it in a vector database;
[0019] S13: According to the user query question, convert it into a vector representation, search the vector database, and obtain several text blocks.
[0020] Further, step S2 includes:
[0021] S21: Calculate the similarity score between each text block and the query question;
[0022] S22: Count the sum of similarity scores between all text blocks and the query question;
[0023] S23: Determine the weight of each text block according to the ratio of the similarity score between each text block and the query question to the total similarity score, which is the similarity between each text block and the query question.
[0024] Further, step S3 includes:
[0025] S31: Divide the text blocks into several levels of groups according to the similarity;
[0026] S32: Determine the injection times of each text block in descending order according to the level of the grouping;
[0027] Further, step S3 includes:
[0028] According to the similarity, the injection times of each text block are calculated using formula (3);
[0029] (3)
[0030] in, Indicates the number of injections of the i-th text block; represents the weight of the i-th text block; represents the similarity score; Indicates the maximum similarity score; Represents the minimum similarity score; represents the single injection scale factor; C represents the total injection scale factor.
[0031] Further, step S4 includes:
[0032] S41: Determine an input set according to the document block to be injected;
[0033] S42: inject the input set,into the generative model to obtain the output;
[0034] S43: subtract 1 from the number of input times of each document block, and determine whether the number of input times of each document block is 0. If so, delete the document block from the input set; otherwise, do not process it, and obtain an updated input set.
[0035] S44: Determine whether the updated input set is empty, if so, proceed to step S45, if not, return to step S42;
[0036] S45: Fusion of each output to obtain the final generated answer.
[0037] Further, step S4 includes:
[0038] S41': Determine an input set according to the document block to be injected;
[0039] S42': Inject the input set into the generative model to obtain the output;
[0040] S43': subtract 1 from the input times of each document block, and determine whether the input times of each document block is 0. If yes, delete the document block from the input set, otherwise, do not process it, and obtain an updated input set;
[0041] S44': Determine whether the updated input set is empty, if so, end; if not, add the summary of the current output to the updated input set, dynamically update the input set, and return to step S42'.
[0042] Furthermore, step S4 further includes:
[0043] Calculate the updated similarity between the output of the previous iteration and each text block to be input in the next iteration;
[0044] The injection times of each text block are dynamically updated according to the update similarity of each text block.
[0045] On the other hand, the present invention also provides a computer-readable storage medium on which a computer program for any of the above-mentioned retrieval generation optimization methods is stored.
[0046] On the other hand, the present invention also provides a computer system, comprising the above-mentioned computer-readable storage medium and one or more processors;
[0047] The processor is configured to run the computer program.
[0048] The present invention provides a retrieval generation optimization method, medium and system based on long-term potentiation, which applies the long-term potentiation (LTP) principle in neuroscience to the retrieval enhancement generation (RAG) technology, and optimizes the generation process of the RAG system by introducing the long-term potentiation mechanism. Specifically, by drawing on the LTP principle, the model's utilization rate of high-quality information is improved through the similarity calculation of document blocks and the batch multiple injection strategy, thereby significantly improving the factual accuracy and consistency of the generated answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 A flowchart of one embodiment of a method for optimizing the generation of prior art searches;
[0050] Figure 2 A flowchart of an embodiment of a method for optimizing retrieval generation according to the present invention;
[0051] Figure 3 A flowchart of another embodiment of the search and generation optimization method of the present invention. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] It should be noted that if the embodiments of the present invention involve directional indications, such as up, down, left, right, front, back, etc., then the directional indication is only used to explain the relative position relationship, movement status, etc. between the components in a certain specific posture. If the specific posture changes, the directional indication will also change accordingly. In addition, if the embodiments of the present invention involve descriptions of "first, second", "S1, S2", "step one, step two", etc., then such descriptions are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of indicated technical features or indicating the execution order of the method, etc. Those skilled in the art can understand that anything that does not violate the gist of the invention under the technical concept of the invention should be included in the protection scope of the present invention.
[0054] like Figure 2-3 As shown, the present invention provides a retrieval generation optimization method based on long-term enhancement, comprising:
[0055] S1: According to the query question input by the user, several text blocks are retrieved;
[0056] S2: Calculate the similarity between each text block and the query question;
[0057] S3: Determine the injection times of each text block according to the similarity;
[0058] S4: According to the injection times of each text block, the input generation model is iterated in batches until the highest injection times are reached to obtain the final generated answer.
[0059] In this embodiment, a retrieval generation optimization method based on long-term enhancement of the present invention is provided. By introducing the long-term enhancement mechanism, the generation process of the RAG system is optimized. By calculating the similarity of document blocks and injecting them in batches multiple times, the utilization rate of the model for high-quality information is improved, thereby significantly improving the factual accuracy and consistency of the generated answers. It is worth noting that the key to the present invention is to propose an innovative strategy of determining the number of injections by similarity calculation and then injecting them in batches multiple times. As for how to retrieve several text blocks with high relevance in step S1, how to calculate the similarity in step S2, how to determine the number of injections of each text block according to the similarity in step S3, and how to iterate the input generation model in batches in step S4 to obtain the final generated answer, a variety of methods can be used. The following preferred embodiments are only for illustrative purposes. As long as the conception is under the batch multiple injection strategy of the present invention, it should be within the protection scope of the present invention.
[0060] S1: According to the query question input by the user, several text blocks are retrieved;
[0061] Specifically, it is optional but not limited to using vector retrieval, hybrid retrieval strategy, etc. Vector retrieval strategy: vectorize the query question input by the user and the text block in the knowledge base, and then retrieve the relevant text block by calculating the similarity between the vectors; Hybrid retrieval strategy: combine BM25 and vector retrieval methods to improve the accuracy and recall rate of retrieval.
[0062] Preferably, as the second invention of the present invention, Figure 3 The example of segmentation enhancement shown, step S1, may optionally include but is not limited to:
[0063] S11: Divide the original documents in the knowledge base into blocks and generate context information;
[0064] Specifically, the document is preprocessed and context is generated, including two steps: document segmentation and context generation. Document segmentation: The original document can be segmented into smaller text blocks according to specific rules, such as a fixed-size block strategy. For example, each block contains 512 tokens, and an overlapping area of 100 tokens is set. This step helps to reduce the complexity of retrieval and improve the accuracy of retrieval. Context generation: Large language processing models and prompt technology can be used to automatically generate explanatory context for each text block. This context information is designed to supplement the context information lost due to document segmentation and enhance the semantic expression ability of the text block. When generating context, factors such as the subject, keywords, and contextual relationships of the text block can be considered to ensure that the generated context information is both concise and targeted.
[0065] S12: construct a complete text block containing context information, convert it into a vector representation, and store it in a vector database;
[0066] Specifically, the document is contextually embedded and represented by a vector, which may include two steps: concatenated embedding and vector representation. Concatenated embedding: The generated context information is concatenated with the corresponding text block to form a complete text block containing context information. This step ensures that the text block can more accurately reflect its original meaning and contextual relationship during the retrieval process. Vector representation: An encoder model, such as the BERT model, may be used to encode the concatenated text blocks and generate vector representations containing context information. These vector representations will be used in subsequent retrieval and matching processes.
[0067] S13: converting the query question input by the user into a vector representation, searching a vector database, and obtaining a number of text blocks;
[0068] Specifically, vector retrieval can be optionally adopted, and the above vector representation can be used to construct an efficient index structure to achieve fast retrieval and query. During the retrieval process, the vector database can be efficiently indexed and queried according to the user query and context information to obtain several text blocks.
[0069] Preferably, step S1 may also optionally include but is not limited to: S14: reordering and interference filtering step. Specifically, a reordering model and interference filtering mechanism may be optionally introduced to further optimize and filter the search results. These models can use context information to perform refined sorting and filtering on the search results, thereby ensuring that the most relevant and accurate document blocks are returned to the user.
[0070] In this embodiment, a preferred embodiment of step S1 is given, which is combined with segmentation enhancement to solve the problem of low relevance of retrieved documents in traditional RAG systems. This technology effectively compensates for the context information lost due to document segmentation by pre-adding specific explanatory context to each document block. These context information are usually automatically generated using advanced natural language processing models (such as GPT-4o, etc.) combined with Prompt technology to ensure conciseness and pertinence. After the context information is spliced with the document block, a vector representation containing the context information is generated through the Encoder model, making the text block more complete and easy to understand during the retrieval process. The application of this technical principle significantly improves the relevance and accuracy of the retrieval results and enhances the semantic understanding ability of the system.
[0071] S2: Calculate the similarity between each text block and the query question;
[0072] Specifically, the similarity can be calculated using a re-ranking model, multi-modal similarity calculation, context enhancement calculation, etc. For example, calculating the similarity using a re-ranking model provides a basis for the injection strategy and weight calculation in the next stage; multi-modal similarity calculation: in addition to traditional text similarity calculation, information from multiple modalities such as images and audio can also be considered to more comprehensively evaluate the relevance of text blocks; context enhancement calculation: when calculating similarity, consider the context information of the text block, such as the semantic association of the context before and after, to improve the accuracy of similarity calculation.
[0073] Specifically, step S2 may optionally but not limited to include:
[0074] S21: Calculate the similarity score of each text block with the query question ; Optionally, calculate the similarity score between each text block and the query question through the Cohere ranking model.
[0075] S22: Statistically calculate the total sum of the similarity scores of all text blocks with the query question ;
[0076] Specifically, formula (1) can be optionally used to calculate the total sum of the similarity scores of all blocks:
[0077] (1)
[0078] S23: Determine the weight of each text block according to the ratio of the similarity score of each text block with the query question to the total sum of the similarity scores , which is the similarity of each text block with the query question.
[0079] Specifically, the weight of each text block can be optionally calculated using formula (2):
[0080] (2)
[0081] Where, represents the weight of the th text block.
[0082] For example, assume there are 3 text blocks, and their similarity scores are S 1 = 0.8, S 2 = 0.6, S 3 = 0.2;
[0083] Calculate the total similarity:
[0084] ;
[0085] Calculate the weight of each text block:
[0086] ;
[0087] Therefore, the weight of the three text blocks is =0.5, =0.375, =0.125, and their sum is 1.
[0088] In this embodiment, a preferred embodiment of step S2 is given, and the exemplary system assigns initial weights to the retrieved document blocks according to the similarity scores obtained by re-ranking the Cohere re-ranking model. The document blocks with higher scores are more likely to contain accurate answers to the questions, and thus are given higher weights as the basis for subsequently determining the number of injections for each text block.
[0089] S3: Determine the injection times of each text block according to the similarity;
[0090] Specifically, the key of the present invention is to adopt a dynamic injection strategy to dynamically adjust the injection times according to the similarity between the text block and the query question. The text blocks with high similarity are injected multiple times to enhance their influence in the generation model.
[0091] Preferably, step S3 may optionally include but is not limited to:
[0092] S31: Divide the text blocks into several levels of groups according to the similarity;
[0093] S32: Determine the injection times of each text block in descending order according to the level of the grouping.
[0094] Specifically, it is possible to select a threshold value set at a certain level according to the similarity between each text block currently retrieved and the query question, and divide the text block into a certain level of groups. For example, when the similarity of a text block exceeds a first threshold, it is divided into a first group, and its injection frequency is increased for a high similarity, and its injection frequency is reduced for a later group; for a group with a particularly low similarity, it is possible to select not to inject it.
[0095] Or, in another embodiment, step S3 may optionally include but is not limited to:
[0096] According to the similarity, the injection times of each text block are calculated using formula (3);
[0097] (3)
[0098] in, Indicates the number of injections of the i-th text block; represents the weight of the i-th text block; represents the similarity score; Indicates the maximum similarity score; Represents the minimum similarity score; represents the single injection scale factor; C represents the total injection scale factor. and C are parameters for controlling the total number of injections, which provide an adjustable scale to ensure the number of injections. The overall range of meets the task requirements. The inventor creatively discovered through experimental data that if there is no injection scale factor, the number of injections may be directly determined by the weight, similarity score, similarity, etc., which will cause the generated result to be too heavy or too light on a single block, which is not conducive to the subsequent generation of the final answer.
[0099] S4: According to the injection times of each text block, the model is inputted iteratively in batches until the highest injection times are reached to obtain the final generated answer. Specifically, there are many ways to implement the iterative input of text and how to integrate the final generated answer.
[0100] In one embodiment, each text block can be optionally input into the generation model in batches according to the number of injections, and the answer obtained from each input can be integrated with the information of text blocks with different injection times using a fusion generation strategy when the answer is finally generated to generate a more comprehensive and accurate answer. Specifically: During the generation process, document blocks with higher weights will appear multiple times in the input sequence of the generation model to ensure that the generation model can see high-quality information more frequently. This process is similar to synaptic strengthening, which improves the generation accuracy of the model by repeatedly learning important information. The number of batches is determined by the number of injections with the highest weight. Therefore, it is particularly important to control the number of injections. It is recommended that the maximum number does not exceed 5 times, which can be optionally determined by the above-mentioned grouping, injection scale factor, etc. Specifically, step S4 optionally includes but is not limited to:
[0101] S41: Determine an input set according to the document block to be injected;
[0102] S42: inject the input set,into the generative model to obtain the output;
[0103] S43: subtract 1 from the number of input times of each document block, and determine whether the number of input times of each document block is 0. If so, delete the document block from the input set; otherwise, do not process it, and obtain an updated input set.
[0104] S44: Determine whether the updated input set is empty, if so, proceed to step S45, if not, return to step S42;
[0105] S45: Fusion of each output to obtain the final generated answer.
[0106] In another embodiment, the answer obtained from each input can be optionally dynamically adjusted to the next input, and the iterative input generation model is used, that is, in each iteration, a preliminary answer is generated based on the injected text block, and the subsequent injected text block is adjusted based on the generated result feedback. Specifically, step S4 can optionally include but is not limited to:
[0107] S41': Determine an input set according to the document block to be injected;
[0108] S42': Inject the input set into the generative model to obtain the output;
[0109] S43': subtract 1 from the input times of each document block, and determine whether the input times of each document block is 0. If yes, delete the document block from the input set, otherwise, do not process it, and obtain an updated input set;
[0110] S44': Determine whether the updated input set is empty, if so, end; if not, add the summary of the current output to the updated input set, dynamically update the input set, and return to step S42'.
[0111] Example:
[0112] Assume that the document blocks and their injection times are as follows:
[0113] Document block 1 (high weight): injected 3 times; Document block 2 (medium weight): injected 2 times; Document block 3 (low weight): injected 1 time.
[0114] 1. Initialize injection.
[0115] Input set: Inject all document chunks {document chunk1, document chunk2, document chunk3} into the model along with the user input.
[0116] formula:
[0117] ;
[0118] Output: Model generates preliminary results and summary .in, Represents the complete output generated by the model. Indicates from A summary of the key information extracted from the
[0119] 2. Recursively inject the second round. At this time, the injection count of document block 3 is 0, so no further injection is required.
[0120] Input set: Use document block 1 (high weight block) and document block 2 (medium weight block), combined with the first generated summary , injected into the model.
[0121] formula:
[0122] ;
[0123] Output: Model generates improved results and new abstract .
[0124] 3. Recursive injection for the third round.
[0125] Input set: Only document chunk 1 (high weight chunk) and the second generated summary are used , injected into the model to obtain the final result.
[0126] formula:
[0127] ;
[0128] Output: The model generates the final result .
[0129] Formula summary:
[0130] Round 1 input:
[0131] ;
[0132] Output:
[0133] ;
[0134] Round 2 input:
[0135] ;
[0136] Output:
[0137] ;
[0138] Round 3 input:
[0139] ;
[0140] Output:
[0141] ;
[0142] In this embodiment, after receiving the enhanced information, the system takes the user's question and the information obtained through the batch multiple injection strategy as input to the Chatgpt-o1 generation model. The generation model outputs the corresponding answer based on these prompts. Through multiple injections, the generated answer is more accurate and coherent, thereby improving the accuracy and consistency of the generated answer.
[0143] In another preferred embodiment, to further improve the accuracy and consistency of the generated answers, step S4, between the previous iteration and the next iteration, may optionally include but is not limited to:
[0144] Calculate the updated similarity between the output of the previous iteration and each text block to be input in the next iteration; specifically, the updated similarity can be calculated in the same way as step 2 to calculate the similarity of each text block.
[0145] The injection times of each text block are dynamically updated according to the update similarity of each text block.
[0146] In this embodiment, a further optimization of step S4 is given, and the number of injections of the next time is further dynamically updated through the previous result, so that the number of injections is dynamically adjusted following the previous result, gradually approaching the result similarity, further increasing the number of injections of texts with high similarity and reducing the number of injections of texts with low similarity. This dynamic adjustment can, on the one hand, further improve the utilization value of texts with high similarity and repeat injection; on the other hand, it can further reduce the interference of irrelevant information and the loss of time, and further improve the retrieval accuracy and efficiency.
[0147] In this embodiment, many preferred embodiments of step S4 are given. The principle of long-term potentiation (LTP) of segmentation enhancement technology is a phenomenon in neuroscience that describes the persistent enhancement of signal transmission between neurons. The present invention draws on it and applies it to the generation stage of the RAG system. In traditional RAG systems, retrieved document blocks are often injected into the generation model only once, resulting in low information utilization and inaccurate generation results. In order to overcome this defect, the present invention introduces a long-term enhancement mechanism to improve the model's utilization of high-quality information through weight calculation and batch multiple injection strategies. Specifically, highly relevant document blocks will be injected into the model input sequence multiple times during the generation stage, and combined with the previous summary, they will be input again according to the number of injections to ensure that the generation model can see and use these more relevant and useful information more frequently. This mechanism enhances the model's "memory" of key information, thereby improving the accuracy and coherence of generated answers.
[0148] In summary, the present invention aims to draw on the principle of long-term potentiation (LTP) in neuroscience to solve the two core problems existing in traditional retrieval-augmented generation (RAG) technology: low relevance of retrieved documents and low information utilization, and inaccurate generation results. By introducing the segmentation enhancement method and the long-term enhancement mechanism, the present invention aims to optimize the retrieval and generation process of the RAG system, improve the accuracy and relevance of the retrieval results, and enhance the utilization of high-quality information by the generation model, thereby significantly improving the factual accuracy and consistency of the generated answers. Specifically, the beneficial effects of the present invention include:
[0149] 1. Improve the relevance and accuracy of retrieved documents: By using segmentation enhancement methods, we add explanatory context to document blocks, enhance the semantic expression ability of document blocks, and make the retrieval results closer to the real intention of the user's query. At the same time, we use advanced re-ranking models and interference filtering mechanisms to further improve the accuracy and relevance of retrieval results and reduce the interference of irrelevant or low-relevance document blocks.
[0150] 2. Enhance information utilization and improve the accuracy of generated answers: Drawing on the LTP principle, through weight calculation and batch multiple injection strategies, highly relevant document blocks receive more attention and utilization in the generation stage. This mechanism ensures that the generation model sees and utilizes key information more frequently, thereby improving the accuracy and coherence of generated answers. At the same time, by optimizing the generation process, the possibility of generating false and untrue information is reduced, and the factuality and reliability of the generated content are improved.
[0151] The main problems to be solved are:
[0152] 1. Inaccurate retrieval caused by loss of context information: In traditional RAG (retrieval-augmented generation) systems, in order to improve retrieval efficiency, documents are usually divided into smaller text blocks. However, while this segmentation strategy improves efficiency, it also brings about the problem of context information loss. The lack of context information causes the text block to become incomplete when presented independently, and it is difficult to accurately reflect its original meaning and contextual relationship. This directly affects the relevance and accuracy of the retrieval results, making it difficult for the system to accurately find the document block that best matches the user's query. Cause: The main reason for the loss of context information lies in the document segmentation strategy. In order to quickly process and retrieve a large number of documents, the system usually cuts the document into smaller text blocks. However, this cutting method destroys the original structure and contextual relationship of the document, making the text block difficult to understand and use after being separated from the context. In addition, traditional retrieval methods often only focus on the semantic information of the text block itself, while ignoring the importance of contextual information to the retrieval results. Therefore, during the retrieval process, it is difficult for the system to accurately capture the text block that is most relevant to the user's query, thus affecting the accuracy and efficiency of the retrieval.
[0153] 2. Inaccurate generation due to insufficient information utilization:
[0154] Causes: The problem of insufficient information utilization mainly stems from the information processing mechanism in the generation phase of the RAG system. In traditional RAG systems, retrieved document blocks are regarded as static knowledge sources and are only injected once during the generation process. However, this approach ignores the dynamic nature and importance differences of information in the generation process. Highly relevant document blocks may contain information that is crucial to the answer to the question, but since they are only injected once, this information has limited influence in the generation model. In addition, the generation model may be disturbed when processing large amounts of information, resulting in the failure to effectively utilize the most relevant pieces of information to construct the answer.
[0155] Through experimental verification, the retrieval generation optimization method based on long-term enhancement proposed in the present invention has shown remarkable results. The inventor conducted experiments in the agricultural intelligent question-answering scenario. Compared with the traditional retrieval enhancement generation (RAG) system, the long-term retrieval enhancement generation algorithm of the present invention has greatly improved the accuracy of the answers. Specifically, the answer accuracy of the traditional RAG system is only 45%, but after adopting the algorithm of the present invention, the accuracy rate is increased to 95%, which is a very considerable improvement.
[0156] Compared with traditional RAG-related technologies, such as cosine similarity search, prompt engineering, query rewriting, and generation of summary documents, the present invention has more prominent advantages. Although these traditional technologies have a certain effect in improving retrieval and generation effects, they are limited by their inherent technical and method bottlenecks and it is difficult to achieve the accuracy level achieved by the present invention. The present invention innovatively applies the principle of long-term potentiation in neuroscience to RAG technology, effectively solving the core problems of low relevance of retrieval documents, low information utilization, and inaccurate generation results in traditional RAG systems, thereby achieving a significant improvement in the accuracy of answers.
[0157] On the other hand, the present invention further provides a computer storage medium storing executable program code; the executable program code is used to execute any of the above methods.
[0158] On the other hand, the present invention further provides a terminal device, comprising a memory and a processor; the memory stores a program code executable by the processor; the program code is used to execute any of the above methods.
[0159] Exemplarily, the program code may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the program code in the terminal device.
[0160] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that the terminal device may also include an input / output device, a network access device, a bus, etc.
[0161] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0162] The memory may be an internal storage unit of the terminal device, such as a hard disk or a memory. The memory may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal device. Further, the memory may also include both an internal storage unit of the terminal device and an external storage device. The memory is used to store the program code and other programs and data required by the terminal device. The memory may also be used to temporarily store data that has been output or is to be output.
[0163] The above-mentioned computer storage medium and terminal device are created based on the above-mentioned method, and their technical functions and beneficial effects are no longer repeated here. The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the various technical features in the above-mentioned embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0164] The above-mentioned embodiments only express several implementation methods of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A retrieval generation optimization method based on long-term enhancement, characterized in that: include: S1: According to the query question input by the user, several text blocks are retrieved; S2: Calculate the similarity between each text block and the query question; S3: Determine the injection times of each text block according to the similarity; S4: According to the injection times of each text block, the input generation model is iterated in batches until the highest injection times are reached to obtain the final generated answer; including: S41: Determine an input set according to the document block to be injected; S42: inject the input set,into the generative model to obtain the output; S43: subtract 1 from the number of input times of each document block, and determine whether the number of input times of each document block is 0. If so, delete the document block from the input set; otherwise, do not process it, and obtain an updated input set. S44: Determine whether the updated input set is empty, if so, proceed to step S45, if not, return to step S42; S45: Fusion of each output to obtain the final generated answer; Or, S41': determining an input set according to the document block to be injected; S42': Inject the input set into the generative model to obtain the output; S43': subtract 1 from the input times of each document block, and determine whether the input times of each document block is 0. If yes, delete the document block from the input set, otherwise, do not process it, and obtain an updated input set; S44': Determine whether the updated input set is empty, if so, end; if not, add the summary of the current output to the updated input set, dynamically update the input set, and return to step S42'.
2. The retrieval generation optimization method according to claim 1, characterized in that: Step S1 includes: S11: Divide the original documents in the knowledge base into blocks and generate context information; S12: construct a complete text block containing context information, convert it into a vector representation, and store it in a vector database; S13: According to the query question input by the user, convert it into a vector representation, search the vector database, and obtain several text blocks.
3. The retrieval generation optimization method according to claim 1, characterized in that: Step S2 comprises: S21: Calculate the similarity score between each text block and the query question; S22: Count the sum of similarity scores between all text blocks and the query question; S23: Determine the weight of each text block according to the ratio of the similarity score between each text block and the query question to the total similarity score, which is the similarity between each text block and the query question.
4. The retrieval generation optimization method according to claim 1, characterized in that: Step S3 includes: S31: Divide the text blocks into several levels of groups according to the similarity; S32: Determine the injection times of each text block in descending order according to the level of the grouping.
5. The retrieval generation optimization method according to claim 1, characterized in that: Step S3 includes: According to the similarity, the injection times of each text block are calculated using formula (3); (3) in, Indicates the number of times the i-th text block is injected; represents the weight of the i-th text block; represents the similarity score; Indicates the maximum similarity score; Represents the minimum similarity score; represents the single injection scale factor; C represents the total injection scale factor.
6. The retrieval generation optimization method according to any one of claims 1 to 5, characterized in that: Step S4 further includes: Calculate the updated similarity between the output of the previous iteration and each text block to be input in the next iteration; The injection times of each text block are dynamically updated according to the update similarity of each text block.
7. A computer-readable storage medium, characterized in that: A computer program for executing the retrieval generation optimization method described in any one of claims 1 to 6 is stored thereon.
8. A computer system, characterized in that: comprising the computer-readable storage medium of claim 7 and one or more processors; The processor is configured to run the computer program.