RAG method based on semantic precise blocking and precise background information generation
Through the RAG method generated by semantic precision chunking and accurate background information, the problems of inaccurate text chunking and poor retrieval effects in RAG technology are solved, and the accuracy and comprehensiveness of the intelligent question-and-answer system are improved.
Patent Information
- Application Number
- CN202510369383.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-29
AI Technical Summary
The existing RAG technology has problems such as inaccurate text chunking, lack of background information and poor retrieval effects, which affects the accuracy and comprehensiveness of the intelligent question-and-answer system.
Using the method of semantic precise chunking and precise background information generation, we use the text to split the text into fine-grained text blocks according to semantics, and generate context abstracts and association problems, and use the text embedding model for vector storage, combining the two-source retrieval and reordering model to improve retrieval accuracy.
It realizes accurate and reasonable chunking of text blocks, enriches background knowledge, improves the accuracy of problem retrieval and the performance of intelligent question-answer system.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a RAG method based on precise semantic chunking and precise background information generation. Background Art
[0002] With the rise of large language models, intelligent question answering systems have experienced significant development. There are mainly the following three technical routes for building intelligent question answering systems using large language models.
[0003] 1. Only relying on the pre-training stage of the large language model: This method is simple and easy to implement, without additional data annotation or training, and is suitable for rapid deployment. However, the knowledge of the model is limited by the time range and content breadth of the pre-training data, and it is easy to generate inaccurate answers or "hallucinations".
[0004] 2. Adopting the method of supervised fine-tuning: Fine-tuning the model with domain-specific annotated data to improve the accuracy of domain-specific questions. However, it requires a large amount of high-quality annotated data, with high costs, and may cause overfitting and "catastrophic forgetting".
[0005] 3. RAG (Retrieval-Augmented Generation) retrieval-augmented generation technology. The RAG technology can be roughly divided into two stages: pre-inference preprocessing preparation and retrieval generation during inference. In the preprocessing stage, various documents related to the question are collected, and these documents are segmented into several text chunks. Then the text chunks are vectorized and stored in the database as the external knowledge base of the large language model. In the inference stage, first, the vector retrieval method is used to find the most relevant text fragment to the question from the vector database as the background information. Then the large language model uses this background information as the enhanced supplementary information to answer the user's question. The advantages of the RAG technology are that the RAG system can obtain the latest information from external knowledge sources in real time, avoiding the problem of outdated model knowledge; by introducing external knowledge, the RAG system can reduce the information generated by the model that does not conform to the facts or fictional details; since there is no need to perform additional training on the model, it can flexibly handle questions in different fields and there is no risk of overfitting and "catastrophic forgetting". However, there are also some obvious defects restricting the performance of the RAG system, including,
[0006] Inaccurate text chunking: When the RAG system performs text chunking, it often adopts the equal fixed-word-count chunking method. For example, every 512 characters are used as a chunk. This method will lead to information loss or ambiguity, affecting the accuracy of the final answer.
[0007] Lack of background information: The retrieved document fragments usually lack necessary background information, which may lead to incomplete or inaccurate answers generated by the model, especially in scenarios that require comprehensive analysis.
[0008] Poor retrieval effect: When the question has a large semantic difference from the text blocks available for answering, relevant text blocks cannot be effectively retrieved.
[0009] To address these deficiencies, this patent proposes a RAG method based on precise semantic chunking and precise background information generation, aiming to further improve the performance and user experience of intelligent question-and-answer systems by optimizing the combination of text chunking technology, background information generation technology, and retrieval technology. Summary of the Invention
[0010] The present invention aims to provide a RAG method based on precise semantic chunking and precise background information generation to improve the problems of missing text block information and poor retrieval effect in existing RAG technologies. The present invention realizes the accurate and reasonable chunking of documents by optimizing the combination method of text chunking technology and background information generation technology, enriches the background knowledge of text blocks, and improves the accuracy of question retrieval by optimizing the retrieval technology. The technical problems to be solved by the present invention are achieved through the following technical solutions.
[0011] In the first aspect of the present invention, a RAG method based on precise semantic chunking and precise background information generation is proposed. In the pre-inference preparation stage, it includes:
[0012] Step S1: In the text data collection stage, documents related to the question are collected, and the text content therein is extracted. The collected documents include but are not limited to professional documents, high-quality Q&A, Internet-related text information, etc.
[0013] Step S2: The purpose of this step is to roughly split the text into coarse-grained text blocks with a relatively large number of characters but within the acceptable range of the maximum input window of the large language model.
[0014] Step S3: The purpose of this step is to split the coarse-grained text blocks obtained in Step S2 into finer-grained text blocks with fewer characters but clear meanings for subsequent retrieval of relevant text blocks. For each coarse-grained text block, the coarse-grained text block is further chunked using prompt words and the large language model to obtain several finer-grained text blocks with fewer characters. These finer-grained text blocks are composed of several sentences in the coarse-grained text block. On the premise of keeping the sentence order unchanged, sentences with closer semantics will be divided into the same finer-grained text block, and the large language model will rewrite and supplement each finer-grained text block to make its meaning clear. The prompt words used for text fine-grained chunking should include but are not limited to the following requirements: (1) The maximum number of sentences contained in each finer-grained text block; (2) If there are words with unclear meanings such as pronouns in the finer-grained text block, it should be rewritten and supplemented in combination with its corresponding coarsely chunked document to make its meaning clear; (3) If the fine-grained text is inconsistent with the meaning expressed in the original text, it should be explained to eliminate ambiguity; (4) On the premise of keeping the sentence order unchanged, sentences that are semantically closer should be grouped into the same fine-grained text block as much as possible.
[0015] This process can be formally expressed as:
[0016] where represents the coarse-grained text block, is the -th fine-grained text block split from , is the fine-grained chunking prompt template, represents the large language model.
[0017] Step S4: In this step, for each fine-grained text block, generate a context summary and several related questions that can be answered using this fine-grained text block. In the prompt for the large language model to process the current task, require the large language model to combine the information of the coarse-grained text block to generate a context summary and several related questions that can be answered using one or more fine-grained text blocks from this coarse-grained text block. In the prompt for this step, it should include but not be limited to the following requirements: (1) The maximum length of the generated context summary; (2) The generated context summary should be related to the current fine-grained text block and be able to supplement, explain, and expand it; (3) Combine the coarse-grained chunking and fine-grained chunking to generate n possible questions that can be answered using the fine-grained chunk.
[0018] Finally, generate a context summary and n related questions that can be answered using this fine-grained text block for each fine-grained text block. Concatenate a fine-grained text block with its corresponding context summary to form the final text block, which is used as the chunk containing accurate background information. The n possible questions are used as the related questions corresponding to this final text block.
[0019] The process of generating related questions and context summaries can be formally expressed as: and
[0020] where represents the coarse-grained text block, is the -th fine-grained text block split from , represent the corresponding context summary represent the th associated problem For generating the prompt template of the associated problem and the context summary represent the large language model represent the final text block represent to and are concatenated. Each associated problem and correspond. Record the mapping relationship between the described associated problem and the final text block containing the corresponding fine-grained text block, that is, the final text block can be found through each associated problem
[0021] Step S5: In this step, vectorize the text information in the final text block and the associated problems generated in step S4. Use a text embedding model such as bge, m3e, etc. to encode the text information in the final text block into text vectors, and store each vector together with the corresponding final text block as a vector index in the text knowledge vector database; extract each associated problem and the corresponding final text block from the mapping relationship recorded in S4, vectorize each associated problem, and store each associated problem vector as a vector index and the corresponding final text block vector in the associated problem vector database at the same time. Note: The text knowledge vector database and the associated problem vector database are two different vector databases
[0022] In the inference stage, the present invention provides a dual-source retrieval strategy, including the following steps
[0023] Step T1: First, vectorize the question text input by the user, and then calculate its cosine similarity with each vector in the text knowledge vector database. The calculation formula is as follows
[0024] where represents the question text vector represents a text block vector in the text block database. Select the top text blocks as the retrieval results of this database ) represents the set of relevant text blocks retrieved from the text knowledge vector database. In a similar way, calculate the cosine similarity between the user question vector and each vector in the associated problem vector database, select the set of text block vectors with the top similarities, and then retrieve the corresponding text blocks from these text block vectors in the text knowledge vector library to obtain the text block set , the union of the retrieval results of the two databases is obtained to get the dual-source retrieval result: , represents the set of relevant text blocks screened from the two vector databases.
[0025] Step T2: The set of text blocks retrieved in step T1 is fed into a ReRanker (reranking) model (such as bge-reranker, bce-reranker, etc.) for reranking, and several text blocks with the top ranking results are selected as the finally retrieved relevant text blocks.
[0026] Step T3: The retrieved relevant text and the user's question are combined into a prompt to prompt the large model to answer the user's question. Description of the Drawings
[0027] Figure 1 is the flowchart of the pre-inference preparation stage steps of the RAG method based on semantic precise chunking and precise background information generation of the present invention.
[0028] Figure 2 is the flowchart of the inference stage steps of the RAG method based on semantic precise chunking and precise background information generation of the present invention.
[0029] Figure 3 is the schematic diagram of the fine-grained chunking strategy process in the RAG method based on semantic precise chunking and precise background information generation of the present invention.
[0030] Figure 4 is the schematic diagram of the principle of context summary and related question generation in the RAG method based on semantic precise chunking and precise background information generation of the present invention.
[0031] Figure 5 is the example diagram of the embodiment of generating context summary and related questions in the RAG method based on semantic precise chunking and precise background information generation of the present invention. Detailed Embodiments
[0032] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described below with reference to the drawings and implementation examples. It should be understood that the specific implementation examples described here are only used to explain the present invention and are not used to limit the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present invention discloses a RAG method based on semantic precise chunking and precise background information generation. For the convenience of understanding and implementation, the key concepts involved in the present invention are first explained, and then the implementation methods are described in detail.
[0033] RAG (Retrieval-Augmented Generation): RAG is a hybrid model framework that combines retrieval and generation to improve the performance of natural language processing tasks. It first uses a retrieval module to find the most relevant fragments from a large number of documents, and then uses these fragments as additional information sources to assist large language models in generating more accurate and relevant outputs.
[0034] Large Language Model: A large language model refers to a language model with a large number of parameters. They are usually trained through deep learning techniques and can generate or understand natural language. These models are widely used in many fields such as text generation, translation, and question answering due to their powerful generalization ability. Examples include models like gpt3.5 and llama3.
[0035] Text chunking: Text chunking is a technique that divides long texts into multiple shorter parts (or chunks). This approach helps improve processing efficiency, especially in information retrieval or text analysis, where specific text chunks can be operated on instead of the entire document.
[0036] Prompt: In natural language processing, a prompt refers to words or phrases used to guide a model to generate a specific type of response. Through carefully designed prompts, users can effectively control the behavior of the generation model and make its output more in line with expectations.
[0037] Text embedding model: A text embedding model is a technique that converts text into numerical vectors that can capture the semantic information of the text in a multi-dimensional space. Text embedding is crucial for many NLP tasks such as similarity search, classification, and clustering. Examples include models like word2vec and Bert.
[0038] Vector database: A vector database is a database system specifically designed to store and query high-dimensional vector data. It optimizes similarity search for large-scale vector data and is commonly used in fields such as recommendation systems, image recognition, and natural language processing.
[0039] Reranker model: A Reranker model is a model that re-ranks the candidate results after initial retrieval to improve the quality of the final results. It usually relies on more complex evaluation criteria, such as considering context information or user's personalized preferences, to make the top-ranked results more in line with the actual needs of users. Examples include models like the reciprocal fusion ranking model and bge-reranker.
[0040] such as Figure 1As shown in the figure, this embodiment provides a RAG method based on semantic precise chunking and precise background information generation, including the following steps:
[0041] Step S1: In the text data collection stage, documents related to the question are collected, and the text content therein is extracted. The collected documents include but are not limited to professional documents, high-quality Q&A, Internet-related text information, etc. In addition, according to user feedback (such as the number of likes and dislikes), answers that are answered well or poorly by the intelligent Q&A system can be screened out. The better Q&A texts are adopted into the collected texts, and the poor Q&A are added to the collected texts after manual correction of the answer results.
[0042] Step S2: The purpose of this step is to divide the text into text chunks with a relatively large number of characters, but within the acceptable range of the large language model. Specifically, as Figure 2 shown, different segmentation methods are selected for different types of texts. For example, for documents with a table of contents structure, recursive segmentation is performed according to chapters. Texts that do not meet the character number requirements are split into several chapters until the character number is met; for texts without a table of contents structure, recursive segmentation can be performed according to paragraphs and punctuation marks until the character number is met. Through the above methods, different types of texts can be effectively segmented into coarse-grained text chunks suitable for processing by the large language model, which not only ensures that the coarse-grained text chunks contain more text information, but also maintains the coherence and logic of the text content, facilitating the generation of context summary information in subsequent steps. An example of a coarse-grained text chunk is shown in the Figure 5 coarse-grained text chunk part shown.
[0043] Step S3: The purpose of this step is to split the coarse-grained text chunks obtained in Step S2 into finer-grained text chunks with fewer words but clear meanings, facilitating the retrieval of relevant text chunks in subsequent steps. As Figure 3 shown, for each coarse-grained text chunk, the coarse-grained text chunk is further chunked using prompt words and the large language model to obtain several finer-grained text chunks with fewer words. These finer-grained text chunks are composed of several sentences in the coarse-grained text chunk. Under the premise of keeping the sentence order unchanged, sentences with closer semantics will be divided into the same finer-grained text chunk. At the same time, the large language model will rewrite and supplement each finer-grained text chunk to make its meaning clear. The prompt words used for text fine-grained chunking should include but are not limited to the following requirements: (1) The maximum number of sentences contained in each finer-grained text chunk; (2) If there are words with unclear meanings such as pronouns in the finer-grained text chunk, it should be rewritten and supplemented in combination with its corresponding coarsely chunked document to make its meaning clear; (3) If the fine-grained text is inconsistent with the meaning expressed by the text in the original text, it should be explained to eliminate ambiguity; (4) On the premise of keeping the sentence order unchanged, sentences that are semantically closer should be grouped into the same fine-grained text block as much as possible.
[0044] This process can be formally expressed as:
[0045] where, represents the coarse-grained text block, is the th fine-grained text block split from is the fine-grained chunking prompt template, represents the large language model.
[0046] Step S4: In this step, for each fine-grained text block, generate a context summary and several related questions that can be answered using this fine-grained text block. As Figure 4 shown, in the prompt for the large language model to process the current task, it is required that the large language model combines the information graph of the coarse-grained text block. In the prompt for this step, it should include but not be limited to the following requirements: (1) The maximum length of the generated context summary; (2) The generated context summary should be related to the current fine-grained text block and be able to supplement, explain, and expand it; (3) Combine the coarse-grained chunk and the fine-grained chunk to generate n possible questions that can be answered using the fine-grained chunk.
[0047] Each fine-grained text block finally obtains a context summary and n related questions that can be answered using this fine-grained text block. Concatenate a certain fine-grained text block with its corresponding context summary to form the final text block, and the n possible questions as the related questions corresponding to this final text block. Generate a context summary and several related questions that can be answered using this fine-grained text block for one or more fine-grained text blocks from this coarse-grained text block. Examples of the context summary, related questions, and final text blocks generated for the coarse-grained text block are as Figure 5 shown.
[0048] The process of generating related questions and context summaries can be formally expressed as: and
[0049] where, represents the coarse-grained text block, is the The th fine-grained text block split out, represents the corresponding context summary, represents the th associated problem, is to generate a prompt template for associated problems and context summaries, represents a large language model. represents the final text block, represents and are concatenated. Each associated problem and correspond. Examples of generating context summaries and associated problems are as Figure 5 shown.
[0050] Step S5: In this step, the text information in the final text block and the associated problems generated in step S4 is vectorized. Using a text embedding model such as bge, m3e, etc., the text information in the final text block is encoded into text vectors and stored in the text block vector database, and the associated problems corresponding to the final text block are encoded into text vectors and stored in the associated problem vector database.
[0051] In the inference stage, as Figure 2 shown, it includes the following steps:
[0052] Step T1: Vectorize the problem text. First, calculate the cosine similarity between it and each vector in the text knowledge vector database. The calculation formula is as follows:
[0053] represents the problem text vector, represents a text block vector in the text knowledge vector database. Select the text blocks with the top similarities as the retrieval results of this database: ), represents the set of relevant text blocks retrieved from the text knowledge vector database. Using a similar method, calculate the cosine similarity between the user problem vector and each vector in the associated problem vector database, select the set of text block vectors with the top similarities, and then retrieve the corresponding text blocks from these text block vectors in the text knowledge vector library to obtain the text block set . Take the union of the retrieval results from the two data sources to obtain the dual-source retrieval result: ), represents the set of relevant text blocks retrieved from the two vector databases.
[0054] Step T2: The set of text blocks retrieved in step T1 is fed into the bge-reranker model for re-ranking, and several text blocks with the top ranking results are selected as the finally retrieved relevant text blocks.
[0055] Step T3: The retrieved relevant text and the user's question are combined into a prompt to prompt the large model to answer the user's question.
Claims
1. A RAG method based on semantic precise chunking and precise background information generation, characterized in that, It includes a pre-inference preparation stage and an inference stage. The pre-inference preparation stage includes the following steps: Step S1: Collect multiple documents to be retrieved; Step S2: Cut each of the above documents into multiple coarse-grained text blocks according to a fixed block length, where the fixed block length refers to the maximum number of characters contained in each block set in advance; Step S3: Split each coarse-grained text block into multiple fine-grained text blocks semantically; specifically, construct prompt words and use a large language model to split each coarse-grained text block into several semantically coherent fine-grained text blocks with fewer words; the fine-grained text block is a text block composed of sentences with close semantics, and a coarse-grained text block can be split into multiple fine-grained text blocks; the sentences with close semantics refer to the sentences in the same natural paragraph or the relevant natural paragraphs describing the same thing; Step S4: Generate a context summary and multiple associated questions for each fine-grained text block; specifically, take each coarse-grained text block as the context background information corresponding to the several fine-grained text blocks split therefrom, form pairs of the context background information and each corresponding fine-grained text block and input them into the large language model, construct prompt words, and let the large language model generate a context summary and N associated questions that can use this fine-grained text block as an answer for each fine-grained text block. One fine-grained text block corresponds to one context summary and multiple associated questions, one associated question corresponds to one fine-grained text block, and N is an integer parameter greater than or equal to 1; splice the fine-grained text block and the corresponding generated context summary information into a final text block; record the mapping relationship between the associated question and the final text block containing the corresponding fine-grained text block, that is, the final text block can be found through each associated question; Step S5: Vectorize each final text block generated in Step S4, and store each vector as a vector index and the corresponding final text block in a text knowledge vector database; extract each associated question and the corresponding final text block from the mapping relationship in S4, vectorize each associated question, and store each associated question vector as a vector index and the corresponding final text block vector in an associated question vector database at the same time.
2. The inference stage according to claim 1, characterized in that, It includes the following steps: Step T1: For the question input by the user, vectorize the question text and retrieve the text knowledge vector database to obtain the N text blocks with the highest similarity to the question; then use the question text vector to retrieve the associated question vector database to obtain the M associated questions with the highest similarity and the corresponding final text block vectors, and then use the final text block vectors to retrieve the final text blocks from the text knowledge vector database respectively; take the union of the text blocks retrieved from the two databases to form a text block set S; N and M are integer parameters greater than or equal to 1 that can be set; Step T2: Use the Reranker re-ranking model to re-rank the set of text chunks S retrieved in Step T1, and select the top r text chunks as relevant texts, where r is an integer parameter greater than 1 that can be set; the Reranker model refers to a model for text re-ranking, which is used to evaluate and rank the retrieved documents for their relevance to the user's question. Step T3: Combine the r most relevant texts selected in Step T2 with the user's input question to form a prompt, and input it into the large language model to prompt the large language model to output an answer to the user's question.
3. According to step S1 described in claim 1, it is characterized in that: The documents include but are not limited to professional knowledge documents, high-quality Q&A, and Internet-related text information.
4. According to step S3 described in claim 1, it is characterized in that: In the prompt, it should include but is not limited to the following requirements: (1) The maximum number of sentences contained in each fine-grained text chunk; (2) If there are words with unclear meanings such as pronouns in the fine-grained text chunk, it should be rewritten and supplemented in combination with its corresponding coarse-grained document to make its meaning clear; (3) If the sentences in the fine-grained text chunk have inconsistent meanings with the corresponding sentences in the coarse-grained chunk, it should be explained to eliminate ambiguity; (4) On the premise of keeping the sentence order unchanged, sentences that are semantically closer should be grouped into the same fine-grained text chunk as much as possible.
5. According to step S4 described in claim 1, it is characterized in that: In the prompt, it should include but is not limited to the following requirements: (1) The maximum length of the generated context summary; (2) The generated context summary should be related to the current fine-grained text chunk and be able to supplement, explain, and expand it; (3) The number of generated related questions.
6. According to step T1 described in claim 2, it is characterized in that: The method for calculating the similarity between the user's question text vector and each vector in the text chunk vector database uses cosine similarity; the method for calculating the similarity between the user's question text vector and each vector in the related question vector database also uses cosine similarity.
7. According to step T2 described in claim 2, it is characterized in that: The ReRanker model can adopt bge-reranker or bce-reranker.
Citation Information
Cited By
Knowledge question-answering method based on improved RAG and agent workflow
CN120929577A
Method and system for reducing illusion information generated by power data agent
CN121117272A
Customs declaration information auditing method, device and equipment and storage medium
CN121684071A
Scene text video question answering method and system based on selection and focusing mechanism
CN122157282A