Retrieval enhancement generation method based on dynamic document block segment optimization

By dynamically adjusting the size of document blocks and fusing context information, the problem of insufficient context in the RAG model is solved, significantly improving the accuracy and efficiency of the model.

CN120045696AInactive Publication Date: 2025-05-27ZHEJIANG PRECE TECH CO LTD

Patent Information

Application Number
CN202510527022.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing RAG models often encounter insufficient context when processing document fragments, resulting in the inability to answer questions accurately or generate error messages.

Method used

The search-enhanced generation method based on dynamic document block segment optimization is adopted, and the document is divided into paragraphs and blocks through natural language processing technology, and the block size is dynamically adjusted according to query requests, and context information is fused to enhance the context understanding of paragraphs and blocks.

Benefits of technology

Reduces error messages and 'illusions' phenomena caused by insufficient context, and improves the accuracy and efficiency of RAG models when processing complex queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045696A_ABST
    Figure CN120045696A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval enhancement generation method based on dynamic document block segment optimization, which comprises the following steps: S1, acquiring an input document, and preprocessing the input document; s2, the document is segmented into a plurality of paragraphs through a natural language processing technology, each paragraph comprises a plurality of smaller blocks, and an index is constructed according to the paragraphs and the blocks; s3, after a query request is received, related paragraphs and blocks are retrieved through indexes; s4, according to the requirement of the query request, the size of the blocks is dynamically adjusted, segments are constructed through the adjusted blocks, and each segment comprises a plurality of blocks; s5, fusing the block obtained by retrieval with the corresponding context information, and promoting the organic combination of the segment, the block content and the context information through a fusion algorithm; and S6, generating answers based on the fused segments through a language model, and performing inspection confirmation and content optimization on the generated answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a retrieval-augmented generation method based on dynamic document block optimization. Background Art

[0002] In the field of natural language processing (NLP), the Retrieval-Augmented Generation (RAG) model is a deep learning method that combines retrieval and generation. The RAG model retrieves relevant information by retrieving a large number of documents and then uses this information to generate answers. However, the RAG model often encounters the problem of insufficient context when processing document fragments, resulting in inaccurate answers to questions or the generation of incorrect information.

[0003] Existing RAG models usually divide documents into multiple fragments and then use these fragments as the basis for retrieval. These fragments may not be correctly understood and used due to lack of sufficient context information.

[0004] The deficiencies of the prior art are as follows: 1. The fragments may contain implicit references and pronouns, resulting in the retrieval system being unable to retrieve or understand correctly.

[0005] 2. A single fragment may not contain all the answers to the question, and the answers may be scattered in multiple adjacent fragments.

[0006] 3. Presenting adjacent fragments to the language model (LLM) in the wrong order may lead to confusion and the generation of incorrect information.

[0007] 4. Simple fragment segmentation may cause the text to be segmented at the "thought interruption" point, making the fragments lose useful context.

[0008] 5. A single fragment usually only makes sense in the context of an entire section or the whole document, and reading it alone may be misleading.

[0009] Therefore, there is an urgent need for a retrieval-augmented generation method based on dynamic document block optimization. Summary of the Invention

[0010] The purpose of the present invention is to provide a retrieval-augmented generation method based on dynamic document block optimization to overcome the deficiencies in the prior art.

[0011] To achieve the above purpose, the present invention provides the following technical solutions: This application discloses a retrieval-augmented generation method based on dynamic document block optimization, including the following steps: S1: Obtain the input document and preprocess the input document; S2: Segment the document into several paragraphs through natural language processing technology. Each paragraph includes several smaller chunks, and construct an index based on the paragraphs and chunks. S3: After receiving a query request, retrieve relevant paragraphs and chunks through the index. S4: According to the requirements of the query request, dynamically adjust the size of the chunks, and construct segments from the adjusted chunks. Each segment includes several chunks. S5: Integrate the retrieved chunks with the corresponding context information, and promote the organic combination of the content of the segments and chunks with the context information through an integration algorithm. S6: Generate an answer based on the integrated segments through a language model, and verify, confirm and optimize the generated answer.

[0012] Preferably, the division of paragraphs in S2 includes the following: Through natural language processing technology, split the input document into several semantically coherent paragraphs; Determine the starting and ending points of the paragraphs, and ensure that each paragraph includes a complete semantic unit. The natural language processing technology includes sentence boundary detection and topic modeling.

[0013] Preferably, S2 includes the following: Further split each paragraph into several chunks, and each chunk includes sufficient information to support retrieval and generation.

[0014] Preferably, S2 includes the following: Generate a context fragment title for each paragraph to provide context information, and integrate the chunks divided from the paragraph with the context title to construct an index.

[0015] Preferably, S3 further includes the following: When retrieving paragraphs and chunks related to the query request, obtain the paragraphs and chunks that are most likely to contain the answer to the query request through a relevance scoring mechanism.

[0016] Preferably, S4 includes the following sub-steps: S41: Obtain initial chunks. Each chunk includes several sentences and paragraphs, including several context information. S42: Dynamically adjust the size of the chunks according to the complexity of the query request and the depth of the required information. S43: If the query request is a complex query that requires a wider context, merge the corresponding adjacent chunks to form larger chunks; if the query request is a simple query, keep the size of the chunks for precise retrieval. S44: Construct segments from the adjusted chunks. Each segment includes several chunks to provide richer context information.

[0017] Preferably, the construction of the middle section of S4 further includes the following: based on semantic coherence, query relevance, and information integrity, perform the construction of segments.

[0018] Preferably, S5 includes the following: fuse the retrieved segments, blocks with the corresponding context fragment titles, thereby enhancing the context information of the segments and blocks, and through the fusion algorithm, promote the organic combination of the context information and the block content.

[0019] Preferably, S6 includes the following: verify the answers generated by the language model through generation control technology to ensure that the generated answers conform to grammar specifications and are accurate in information.

[0020] Preferably, S6 includes the following: perform grammar checking and content optimization on the generated answers to improve the accuracy and readability of the answers, and ensure that the finally output answers meet the predetermined performance standards.

[0021] This application also discloses a retrieval enhanced generation device based on dynamic document block segment optimization, including a memory and one or more processors. Executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above-mentioned retrieval enhanced generation method based on dynamic document block segment optimization.

[0022] This application also discloses a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above-mentioned retrieval enhanced generation method based on dynamic document block segment optimization.

[0023] Advantages of the present invention: (1) Through the transformation from block to segment in the present invention, dynamically adjust the size of the block and the content of the context information, reduce the error information and "hallucination" phenomenon caused by insufficient context, and through dynamically adjusting the size of the block, provide more flexible and adaptable retrieval and generation capabilities; (2) Significantly improve the accuracy and efficiency of the RAG retrieval enhanced generation model when processing complex queries.

[0024] The features and advantages of the present invention will be described in detail through embodiments in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a step flow chart of a retrieval enhanced generation method based on dynamic document block segment optimization of the present invention; Figure 2 is a schematic diagram of the device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. However, it should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0027] Referring to Figure 1 , an embodiment of the present invention provides a retrieval-augmented generation method based on dynamic document block segment optimization, including the following steps, aiming to solve the problem of segment processing caused by insufficient context in the RAG retrieval-augmented generation model: 1. Document preprocessing: Receive the input document and perform format standardization.

[0028] Clean the document and remove irrelevant content such as headers, footers, comments, etc.

[0029] 2. Semantic segmentation: Utilize natural language processing techniques, such as sentence boundary detection and topic modeling, to segment the document into multiple semantically coherent paragraphs.

[0030] Determine the start and end points of the paragraphs to ensure that each paragraph contains a complete semantic unit.

[0031] 3. Context fragment title generation: Generate context fragment titles for each paragraph, which can be the document title, section title, or automatically generated abstract.

[0032] The titles are used to provide additional context information to help the language model better understand the paragraph content.

[0033] 4. Text segmentation: Further segment each paragraph into smaller blocks, where each block contains sufficient information to support retrieval and generation.

[0034] 5. Text block indexing: Fuse the blocks with the context fragment titles to build an index.

[0035] 6. Retrieval: For a given query, use the retrieval system to retrieve relevant paragraphs or blocks from the database.

[0036] Apply a relevance scoring mechanism to determine which paragraphs or blocks are most likely to contain the answer to the query.

[0037] The relevance scoring mechanism includes: A specific method is: (1) Vectorize the block to be queried using a vectorization model and persistently store it; (2) Vectorize the query using a vectorization model; (3) Calculate the distance between the vector of the block to be queried and the vector of the query. Generally, cosine distance is not used; (4) Sort by vector distance and select the top N blocks to be queried with the highest scores.

[0038] The method is not limited to vector distance. Common methods include: Vector distance, BM25 (a very famous algorithm in the field of information retrieval), and a hybrid distance metric of vector distance and BM25 7. Conversion from block to segment: Introduce a "block-to-segment" conversion mechanism to dynamically adjust the size of the block according to the query requirements to provide appropriate context.

[0039] Among them: The "block-to-segment" conversion process is described in detail The "block-to-segment" conversion process is one of the key innovations of this application. Its detailed steps are as follows: (1) Initial segmentation of blocks: Divide the document into initial blocks, each block containing a certain number of sentences or paragraphs to ensure sufficient context information.

[0040] (2) Dynamically adjust the block size: Dynamically adjust the block size according to the complexity of the query and the depth of the required information.

[0041] For complex queries that require a wider context, merge adjacent blocks to form larger segments.

[0042] For simple queries, the blocks may be kept small for precise retrieval.

[0043] For example: A simple query may only require one block to have enough information to answer the user's question.

[0044] While complex queries require multiple blocks; multiple blocks may include blocks that are relatively far apart in the original document, such as at the beginning and end of the document, that is, answering a question may require the first and last paragraphs of the document; and may also require blocks from multiple documents.

[0045] (3) Construction of segments: Construct the adjusted blocks into segments, each segment containing one or more blocks to provide richer context information.

[0046] Ensure that the construction of segments is not only based on semantic coherence but also takes into account query relevance and information integrity.

[0047] 8. Context Fusion: Fuse the retrieved paragraphs or chunks with the corresponding context fragment titles to enhance the context information of the paragraphs or chunks.

[0048] Use a specific fusion algorithm, such as weighted merging, to ensure the organic combination of context information and chunk content.

[0049] 9. Generate Answers: Use a language model, such as the Transformer deep learning model, to generate answers based on the fused paragraphs or chunks.

[0050] Apply generation control techniques to ensure that the generated answers conform to grammar rules and are accurate in information.

[0051] The described generation control technique refers to the method of controlling the output of the large model using prompt words.

[0052] The prompt words in this technique include the original question, relevant context, and instructions (such as "if the context does not contain information related to the original question, then answer 'I don't know'"). 10. Post - processing: Conduct grammar checking and content optimization on the generated answers to improve the accuracy and readability of the answers, and ensure that the finally output answers meet the predetermined performance standards.

[0053] Example 1: Suppose there is a research field of scientific papers containing multiple papers. The method of the present invention will show how to improve the performance of the RAG retrieval - augmented generation model through the method of the present invention.

[0054] 1. Document Pre - processing: Clean the scientific papers to remove irrelevant information, such as headers, footers, references, etc.

[0055] 2. Semantic Segmentation: Use natural language processing techniques, such as sentence boundary detection and topic modeling, to segment the papers into multiple semantically coherent paragraphs.

[0056] 3. Context Fragment Title Generation: Generate context fragment titles for each paragraph, such as using section titles or automatically generated abstracts.

[0057] 4. Initial Chunk Segmentation: Further segment each paragraph into smaller chunks, with each chunk having a maximum length of 256 characters and each chunk containing sufficient information to support retrieval and generation.

[0058] 5. Text block indexing: Integrate the blocks with the context fragment titles and build an index. The indexing method uses a vectorization model, such as the bge-large-zh-v1.5 model of BAAI Beijing Academy of Artificial Intelligence. Store the original text blocks and vectors in the vector database of the Elastic Search server.

[0059] 6. Retrieval: For a given query statement, use a vectorization model to generate the vector of the query statement, and use a retrieval system to retrieve relevant blocks from the database.

[0060] 7. Segment construction: Based on the blocks relevant to the query retrieved, filter out the blocks that are consecutive in the original text with the retrieved blocks. Calculate the vector distance between each block and the query, and use the cosine distance as the distance metric function to obtain the score of each block. Subtract a constant threshold of 0.1 from each score. Use the maximum subsequence sum algorithm to find the best segment, that is, a set of consecutive blocks that can maximize the score.

[0061] 8. Context fusion: Integrate the retrieved paragraphs or blocks with the corresponding context fragment titles to enhance the context information of the paragraphs or blocks.

[0062] 9. Generate answer: Use a large language model, such as GPT-4o, to generate an answer based on the fused paragraphs or blocks.

[0063] 10. Post-processing: Check the grammar and optimize the content of the generated answer to improve the accuracy and readability of the answer.

[0064] In a feasible embodiment, the distinguishing technical feature from Embodiment 1 is that different methods are used to generate the context fragment titles, such as using document summaries or keywords.

[0065] In a feasible embodiment, the distinguishing technical feature from Embodiment 1 is that different semantic segmentation methods are adopted to divide the document paragraphs.

[0066] In a feasible embodiment, the distinguishing technical feature from Embodiment 1 is that different algorithms are applied to identify and construct the most relevant text paragraphs.

[0067] An embodiment of the retrieval enhanced generation device based on dynamic document block segment optimization of the present invention can be applied to any device with data processing capabilities, and such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 2 shown, it is a hardware structure diagram of any device with data processing capabilities where the retrieval enhanced generation device based on dynamic document block segment optimization of the present invention is located. In addition to Figure 2 the processor, memory, network interface, and non-volatile memory shown, the any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here. The implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, which will not be elaborated here.

[0068] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0069] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a retrieval enhanced generation device based on dynamic document block segment optimization in the above embodiment.

[0070] The computer-readable storage medium may be an internal storage unit of any data processing-capable device described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any data processing-capable device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any data processing-capable device. The computer-readable storage medium is used to store the computer program and other programs and data required by any data processing-capable device, and may also be used to temporarily store data that has been output or is to be output.

[0071] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, or improvement made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A retrieval enhancement generation method based on dynamic document block optimization, characterized in that: The steps include: S1: Obtain input documents and preprocess the input documents; S2: Use natural language processing technology to divide the document into several paragraphs, each paragraph includes several smaller blocks, and build an index based on the paragraphs and blocks; S3: After receiving the query request, relevant paragraphs and blocks are retrieved through the index; S4: dynamically adjust the size of the block according to the query request, and construct segments with the adjusted blocks, each segment including several blocks; S5: Fusing the retrieved blocks with the corresponding context information, and promoting the organic combination of segment and block content with context information through a fusion algorithm; S6: Generate answers based on the fused segments through the language model, and verify, confirm and optimize the generated answers.

2. A retrieval enhancement generation method based on dynamic document block optimization as claimed in claim 1, characterized in that: The paragraph division in S2 includes the following contents: segmenting the input document into a number of semantically coherent paragraphs by natural language processing technology; determining the starting point and the ending point of the paragraph, and ensuring that each paragraph includes a complete semantic unit; The natural language processing techniques include sentence boundary detection and topic modeling.

3. The retrieval enhancement generation method based on dynamic document block optimization according to claim 1, characterized in that: The S2 includes the following contents: each paragraph is further divided into several blocks, each block includes information sufficient to support retrieval and generation.

4. The retrieval enhancement generation method based on dynamic document block optimization according to claim 1, characterized in that: The S2 includes the following contents: generating a context segment title for each paragraph for providing context information, and fusing the blocks divided by the paragraph with the context title to construct an index.

5. The retrieval enhancement generation method based on dynamic document block optimization according to claim 1, characterized in that: The S3 also includes the following content: when retrieving paragraphs and blocks related to the query request, the paragraphs and blocks that are most likely to contain the answer to the query request are obtained through a relevance scoring mechanism.

6. The retrieval enhancement generation method based on dynamic document block optimization according to claim 1, characterized in that: The S4 comprises the following sub-steps: S41: obtaining initial blocks, each block including a number of sentences, paragraphs, and a number of context information; S42: dynamically adjust the size of the block according to the complexity of the query request and the depth of the required information; S43: if the query request requires a complex query with a wider context, merge the corresponding adjacent blocks to form a larger block; if the query request is a simple query, keep the size of the block to facilitate accurate retrieval; S44: construct the adjusted blocks into segments, each segment including several blocks to provide richer context information.

7. A retrieval enhancement generation method based on dynamic document block optimization as claimed in claim 6, characterized in that: The construction of the segments in S4 also includes the following contents: constructing the segments based on semantic coherence, query relevance and information completeness.

8. The retrieval enhancement generation method based on dynamic document block optimization according to claim 1, characterized in that: The S5 includes the following contents: fusing the retrieved segments and blocks with the corresponding context segment titles, thereby enhancing the context information of the segments and blocks, and improving the organic combination of the context information and the block content through a fusion algorithm.

9. The retrieval enhancement generation method based on dynamic document block optimization according to claim 1, characterized in that: The S6 includes the following contents: checking the answer generated by the language model through generation control technology to ensure that the generated answer complies with grammatical specifications and the information is accurate.

10. The retrieval enhancement generation method based on dynamic document block optimization according to claim 1, characterized in that: The S6 includes the following contents: performing grammar checking and content optimization on the generated answers to improve the accuracy and readability of the answers and ensure that the answers finally outputted meet the predetermined performance standards.

Citation Information

Patent Citations

  • Advanced retrieval enhanced blocking and vectoring method based on large model

    CN118312579A

  • Text processing method and device for improving information retrieval and generation quality and computer system

    CN118535710A

  • Method for constructing vectorized knowledge base based on dynamic fields for RAG system

    CN119377418A

Cited By

  • Retrieval response method and device, storage medium, electronic equipment and program product

    CN120256546A

  • Tobacco professional knowledge document processing and analyzing method based on multistage feature aggregation

    CN122263849A