Retrieval enhancement generation method and system based on long text

By performing text analysis and vectorization on long text, combining vector database and search algorithms, we determine the priority of text blocks and merge adjacent text blocks, the problem of low quality of RAG system generated answers in long text scenarios is solved, and higher accuracy and completeness are achieved.

CN120011489APending Publication Date: 2025-05-16云鼎科技股份有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411831411.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the long text scenario, the existing search enhancement generation (RAG) system reduces information attention when processing long text, affects the quality of generated answers, and has a large amount of external knowledge, so the text length cannot be expanded without limit, resulting in limited RAG effect.

Method used

By parsing and splitting the text of the local knowledge base, using a vectorized model to convert text into vectors, and fine-tuning it according to the scene, using a vector database and search algorithm to find relevant text blocks, determine the priority of text blocks and add context information, merge adjacent text blocks to ensure the integrity of the output results, and finally input the merged text blocks into a large language model to generate a response.

Benefits of technology

It improves the accuracy of the RAG system in long text scenarios, enhances the quality of generated answers, ensures the integrity of the output results, and overcomes the problem of reduced information attention in long text processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011489A_ABST
    Figure CN120011489A_ABST
Patent Text Reader

Abstract

The invention relates to a retrieval enhancement generation method based on a long text. The method comprises the following steps: S1, analyzing and splitting the text; s2, converting the text into a vector by using a vectorization model; s3, searching an algorithm; s4, determining priorities of the text blocks, and adding more context information; s5, adding adjacent text blocks; s6, using the combined text blocks as input by the large language model, and generating final response or content by using the large language model; and S7, the large language model generates answers according to the input text blocks and the adjacent text blocks. By improving the existing RAG system, the accuracy of the RAG system in the long text scene is improved. The RAG capability is enhanced, and the purpose of improving the answer quality generated by the RAG system is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a retrieval enhancement generation method and system based on long text. Background Art

[0002] Retrieval-augmented generation, or RAG, provides a way for large language models to access recent and specific information by incorporating external knowledge as context. This approach significantly reduces hallucinations in model-generated content and improves the factual accuracy of generated content.

[0003] At present, the implementation of large language models mainly relies on RAG, and the effect of RAG mainly depends on retrieval. Usually, semantic retrieval is the most used, which has high requirements for vectorized model embedding. In specific scenarios, the embedding model needs to be fine-tuned, and the construction of the fine-tuning data set is difficult. Therefore, many people have focused on the processing capabilities of large models for long texts. At present, many models have even expanded the input length to 1M to solve the problem of insufficient input, but long text also has its own disadvantages. When the text becomes longer, the attention to effective information will decrease, which will affect the quality of the final generated answer. In addition, the amount of external knowledge is large, and the text length cannot be expanded indefinitely. Therefore, how to enhance RAG is the direction that should be explored most. Summary of the invention

[0004] (I) Purpose of the invention

[0005] In order to overcome the above shortcomings, the purpose of the present invention is to provide a retrieval enhancement generation method and system based on long text to solve the above technical problems.

[0006] (II) Technical solution

[0007] To achieve the above objectives, the technical solutions provided by this application are as follows:

[0008] A retrieval enhancement generation method based on long text includes the following steps:

[0009] S1 is the text parsing and splitting of the local knowledge base;

[0010] S2 uses a vectorization model to convert text into vectors and fine-tune the model according to the scenario, so that the model can better adapt to the distribution of text data in the current scenario;

[0011] S3 uses a vector database and search algorithms to find relevant text blocks based on input requests;

[0012] S4 prioritizes the text blocks and adds more contextual information;

[0013] S5 merges the text blocks before and after each selected text block to ensure the integrity of the output result;

[0014] The S6 large language model takes the merged text block as input and generates the final response or content using the large language model;

[0015] The S7 large language model generates answers based on the input text block and adjacent text blocks.

[0016] Preferably, the text formats include pdf, docx, excel, txt, and markdown.

[0017] Preferably, the text parsing tool in S1 is one or more of pdfminer, pyMuPDF, and python-dox.

[0018] Preferably, S4 specifically includes the following steps:

[0019] Assume that the length of the text is d, and the text is divided into N text blocks using Represents a text block, and index i represents the order of the text block in the long text d, i.e. c i-1 Indicates c i The previous text block, c i+1 Indicates c i For the following text block, given an input request q, the simplest similarity judgment method is used to calculate the similarity between the input request q and c i The distance between them is used to obtain the correlation between them.

[0020] s i =cos(emb(q),emb(c i ))

[0021] Where cos(·,·) represents the cosine similarity function and emb(·) represents the vectorized function.

[0022] Assume that based on the similarity score, we get the k text blocks with the highest scores. In order to ensure the integrity of the output results, we merge the text blocks before and after each text block together, that is, the number of text blocks is still k, but the size of each text block has changed. The new k text blocks are expressed as:

[0023]

[0024] Since the original order of the text blocks is retained, the following relationship is satisfied between the text blocks:

[0025]

[0026] A retrieval enhancement generation system based on long text, including the following main components:

[0027] The text parsing and splitting module is responsible for parsing long texts in various formats and splitting them into text blocks;

[0028] Vectorization module, which is used to convert text blocks into vector form for similarity comparison and retrieval. Commonly used models include BGE and m3e, which can also be fine-tuned as needed;

[0029] A search algorithm for retrieving the text block most relevant to the request in the vector database according to the input request;

[0030] The text block sorting and merging module determines the priority of the retrieved text blocks and merges them with adjacent text blocks to increase context information and ensure the integrity of the output results;

[0031] The large language model takes the merged text chunks as input and generates the final response or content;

[0032] The hardware system includes selecting a server with an AI chip corresponding to the number of parameters of the large model;

[0033] Software and hardware environment, including operating system, database management system, and programming language environment;

[0034] A user interface that allows the user to enter a request and displays the generated answer;

[0035] Storage system, used to store long text data in the local knowledge base and the vectorized vector database.

[0036] Beneficial effects:

[0037] The present invention improves the accuracy of the RAG system in the long text scenario by improving the existing RAG system. The main idea is to retain the position information of the retrieved text fragment in the text, and output the adjacent upper and lower fragments of the retrieved fragment as the retrieval result. Through this method, the ability of RAG is enhanced, and the purpose of improving the quality of the answer generated by the RAG system is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a system logic diagram of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present invention more clear, the following is a detailed description of the present invention in conjunction with the specific implementation methods and with reference to the attached drawings. Figure 1 , the present invention is further described in detail. It should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.

[0040] The present invention provides a long text-based retrieval enhancement generation method, comprising the following steps:

[0041] S1: text parsing and splitting of the local knowledge base, wherein the text formats include pdf, docx, excel, txt, and markdown, and the text parsing tools are one or more of pdfminer, pyMuPDF, and python-dox;

[0042] S2 uses a vectorization model to convert text into vectors and fine-tune the model according to the scenario, so that the model can better adapt to the distribution of text data in the current scenario;

[0043] S3 uses a vector database and search algorithms to find relevant text blocks based on input requests;

[0044] S4 prioritizes the text blocks and adds more contextual information;

[0045] Assume that the length of the text is d, and the text is divided into N text blocks using Represents a text block, and index i represents the order of the text block in the long text d, i.e. c i-1 Indicates c i The previous text block, c i+1 Indicates c i For the following text block, given an input request q, the simplest similarity judgment method is used to calculate the similarity between the input request q and c i The distance between them is used to obtain the correlation between them.

[0046] s i =cos(emb(q),emb(c i ))

[0047] Where cos(·,·) represents the cosine similarity function and emb(·) represents the vectorized function.

[0048] Assume that based on the similarity score, we get the k text blocks with the highest scores. In order to ensure the integrity of the output results, we merge the text blocks before and after each text block together, that is, the number of text blocks is still k, but the size of each text block has changed. The new k text blocks are expressed as:

[0049]

[0050] Since the original order of the text blocks is retained, the following relationship is satisfied between the text blocks:

[0051]

[0052] S5 merges the text blocks before and after each selected text block to ensure the integrity of the output result;

[0053] The S6 large language model takes the merged text block as input and generates the final response or content using the large language model;

[0054] The S7 large language model generates answers based on the input text block and adjacent text blocks.

[0055] A retrieval enhancement generation system based on long text, including the following main components:

[0056] The text parsing and splitting module is responsible for parsing long texts in various formats and splitting them into text blocks;

[0057] Vectorization module, which is used to convert text blocks into vector form for similarity comparison and retrieval. Commonly used models include BGE and m3e, which can also be fine-tuned as needed;

[0058] A search algorithm for retrieving the text block most relevant to the request in the vector database according to the input request;

[0059] The text block sorting and merging module determines the priority of the retrieved text blocks and merges them with adjacent text blocks to increase context information and ensure the integrity of the output results;

[0060] The large language model takes the merged text chunks as input and generates the final response or content;

[0061] The hardware system includes selecting a server with an AI chip corresponding to the number of parameters of the large model;

[0062] Software and hardware environment, including operating system, database management system, and programming language environment;

[0063] A user interface that allows the user to enter a request and displays the generated answer;

[0064] Storage system, used to store long text data in the local knowledge base and the vectorized vector database.

[0065] The present invention will be further described below by a specific embodiment:

[0066] The following is a diagram of dividing a long text into 12 text blocks, and calculating the similarity values ​​of each text block based on the user input.

[0067]

[0068] Generally, we select the top 3 as the search results. According to the traditional RAG method, the order of the selected text blocks is

[0069]

[0070] According to the present invention, the order of the selected text boxes is

[0071]

[0072] Here, we have omitted the step of combining adjacent text blocks, that is, c4 above actually represents the result of combining c3, ​​c4, and c5. It should also be noted that the large language model is highly dependent on the input. Even with the same input content, changing the text information before and after will result in different output answers. This is also the main idea proposed by the present invention. By controlling the order of text blocks and understanding and parsing the document according to the sequential logic of the document itself, better results will be achieved.

[0073] Comparing the long-text-based RAG system proposed in the present invention with the long-context large prediction model without RAG, the long-text LLM without RAG requires the input of a large number of tokens, which is both inefficient and expensive. In contrast, the RAG system proposed in the present invention not only significantly reduces the number of tokens, but also significantly improves the quality of the answers. In addition, compared with the traditional RAG system, the long-text-based RAG system proposed in the present invention not only retains the position of the retrieved content in the original text, but also expands the content, ensuring the integrity of the generated content and improving the accuracy of the RAG system.

[0074] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0075] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A retrieval enhancement generation method based on long text, characterized in that: The following steps are involved: S1 is the text parsing and splitting of the local knowledge base; S2 uses a vectorization model to convert text into vectors and fine-tune the model according to the scenario, so that the model can better adapt to the distribution of text data in the current scenario; S3 uses a vector database and search algorithms to find relevant text blocks based on input requests; S4 prioritizes the text blocks and adds more contextual information; S5 merges the text blocks before and after each selected text block to ensure the integrity of the output result; The S6 large language model takes the merged text block as input and generates the final response or content using the large language model; The S7 large language model generates answers based on the input text block and adjacent text blocks.

2. The retrieval enhancement generation method based on long text according to claim 1, characterized in that: The text formats include pdf, docx, excel, txt, and markdown.

3. The long text-based retrieval enhancement generation method according to claim 1, characterized in that: The text parsing tools in S1 are one or more of pdfminer, pyMuPDF, and python-dox.

4. The method for generating retrieval enhancement based on long text according to claim 1, characterized in that: S4 specifically includes the following steps: Assume that the length of the text is d, and the text is divided into N text blocks using Represents a text block, and index i represents the order of the text block in the long text d, i.e. c i-1 Indicates c i The previous text block, c i+1 Indicates c i For the following text block, given an input request q, the simplest similarity judgment method is used to calculate the similarity between the input request q and c i The distance between them is used to obtain the correlation between them. s i =cos(emb(q),emb(c i )) where cos(·,·) represents the cosine similarity function, emb(·) represents the vectorized function, Assume that based on the similarity score, we get the k text blocks with the highest scores. In order to ensure the integrity of the output results, we merge the text blocks before and after each text block together, that is, the number of text blocks is still k, but the size of each text block has changed. The new k text blocks are expressed as: Since the original order of the text blocks is retained, the following relationship is satisfied between the text blocks:

5. A retrieval enhancement generation system based on long text, characterized in that: The main components include the following: The text parsing and splitting module is responsible for parsing long texts in various formats and splitting them into text blocks; Vectorization module, which is used to convert text blocks into vector form for similarity comparison and retrieval. Commonly used models include BGE and m3e, which can also be fine-tuned as needed; A search algorithm for retrieving the text block most relevant to the request in the vector database according to the input request; The text block sorting and merging module determines the priority of the retrieved text blocks and merges them with adjacent text blocks to increase context information and ensure the integrity of the output results; The large language model takes the merged text chunks as input and generates the final response or content; The hardware system includes selecting a server with an AI chip corresponding to the number of parameters of the large model; Software and hardware environment, including operating system, database management system, and programming language environment; A user interface that allows the user to enter a request and displays the generated answer; Storage system, used to store long text data in the local knowledge base and the vectorized vector database.

Citation Information

Cited By

  • Streaming data-based retrieval enhancement generation method and electronic equipment

    CN121388144A