Method and device for improving recall capability of RAG system and computer equipment

By combining parsing tools and multi-scale semantic chunking strategies with query rewriting of large language models, the problems of inaccurate information extraction and insufficient multi-question handling capabilities of the RAG system in complex documents are solved, achieving efficient and accurate information retrieval and recall.

CN122045392APending Publication Date: 2026-05-15CHINA NANHU ACAD OF ELECTRONICS & INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610086287.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-22
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing RAG systems struggle to effectively address formatting issues when processing complex documents, especially PDF files, leading to inaccurate information extraction. Chunking strategies also struggle to balance the completeness and conciseness of answers. Traditional question rewriting techniques have limited capacity to handle multi-question questions, impacting the relevance and accuracy of retrieval.

Method used

By optimizing data reading, multi-scale semantic chunking, and question rewriting strategies, MinerU and RapidOCR tools are used to jointly parse PDF documents. Combined with a large language model, query rewriting and multi-channel recall are performed to generate multi-granularity index units, achieving efficient parsing and accurate retrieval.

Benefits of technology

It significantly improves the recall capability and query accuracy of the RAG system in complex document scenarios, ensures the integrity of document content and semantic alignment accuracy, adapts to the semantic features of different types of documents, and improves the recall integrity and relevance of multi-intent queries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045392A_ABST
    Figure CN122045392A_ABST
Patent Text Reader

Abstract

The invention provides a method and a device for improving recall capability of an RAG system. The method comprises the following steps: S1, carrying out structured analysis on a document; s2, performing semantic hierarchical segmentation on the document content by adopting a multi-scale semantic partitioning strategy to generate a multi-granularity index unit; s3, performing semantic enhancement and optimization on the query based on a query rewriting strategy of a large language model; and S4, combining the optimized query with the multi-scale semantic index, and executing a retrieval and generation task. According to the multi-scale semantic chunk strategy, the similarity is calculated by applying a vector model, knowledge is divided into a plurality of blocks with different semantic information, and the vector representation capability is enhanced. An LLM-based question rewriting strategy can effectively identify intention words and keywords in user query, and particularly for a multi-question question, the multi-question question can be split into a plurality of independent sub-questions for multi-path recall. According to the method, the accuracy and efficiency of intelligent query in the vertical field can be remarkably improved, and the query experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information retrieval technology, and in particular to a method and apparatus for improving the recall capability of a RAG system. Background Technology

[0002] Information retrieval technology, as an important component of modern information technology, has been widely applied in fields such as search engines, database queries, and enterprise knowledge management systems. Especially in vertical sectors, such as law, healthcare, and education, information retrieval systems require specific knowledge and terminology understanding capabilities to provide more accurate search results. RAG, as a technology combining information retrieval and generative models, significantly improves the accuracy and relevance of responses by retrieving relevant information from a large number of documents and then using a generative model to generate the final response.

[0003] However, existing information retrieval technologies still have some shortcomings when handling complex documents in specific vertical domains. For example, traditional RAG technology often fails to effectively solve layout problems when processing PDF files, leading to inaccurate information extraction. Regarding chunking strategies, existing single-scale methods struggle to balance the completeness and conciseness of answers. While these methods can quickly segment text, they cannot guarantee the semantic independence and completeness of each segment, especially when processing long texts, resulting in poor retrieval performance. Traditional question rewriting techniques mainly rely on keyword matching and synonym replacement, which have limited capabilities for handling multi-question questions, easily leading to incomplete recall and affecting the targeting and accuracy of retrieval, failing to fully cover users' query needs. These problems limit the application effectiveness of existing technologies in vertical domains, especially in scenarios requiring high accuracy and efficiency. Summary of the Invention

[0004] To address the problems of existing technologies, this invention provides a method and apparatus for improving the recall capability of a RAG system. By optimizing data reading, multi-scale semantic chunking, and question rewriting strategies, the accuracy and efficiency of queries are significantly improved.

[0005] One of the objectives of this invention is to provide a method for enhancing the recall capability of a RAG system, enabling efficient parsing and accurate retrieval of complex document content, thereby significantly improving the recall capability, query accuracy, and overall response efficiency of the RAG system.

[0006] A method for improving the recall capability of a RAG system includes the following steps:

[0007] S1. Perform structured parsing of the document;

[0008] S2. Employ a multi-scale semantic segmentation strategy to perform semantic hierarchical segmentation of document content, generating multi-granularity index units;

[0009] S3. A query rewriting strategy based on a large language model is used to enhance and optimize the semantics of queries;

[0010] S4. Combine the optimized query with the multi-scale semantic index to perform retrieval and generation tasks.

[0011] Preferably, in step S1, a document parsing tool is used to perform structured parsing on multi-format documents (including but not limited to PDF, Word, HTML, txt, and md) to extract their content, which includes text, titles, tables, and images.

[0012] Preferably, for PDF documents with complex structures and irregular layouts, it is preferable to use MinerU and RapidOCR tools in combination for parsing. MinerU is used to extract text content, hierarchical structure and paragraph relationships in the PDF document, while RapidOCR is used to identify and parse table content in the PDF document. The two work together to ensure the integrity and accuracy of the document structure and content.

[0013] Preferably, the multi-scale semantic segmentation strategy in step S2 adopts different segmentation methods based on the semantic density, logical structure and chapter level of the document, including sentence-level, paragraph-level and topic-level segmentation, to adapt to different types of query needs.

[0014] Preferably, the breakpoints of multi-scale semantic blocks are determined based on vector semantic similarity, wherein the vector representations of any two adjacent semantic blocks are respectively and When the cosine similarity is lower than a preset threshold, a semantic breakpoint is determined.

[0015] Preferably, after performing multi-scale semantic segmentation, semantic vector embedding is calculated for each chunk, and a vector index library based on Faiss or Milvus is constructed to support efficient similarity retrieval.

[0016] Preferably, the query rewriting strategy based on the large language model in step S3 involves semantic parsing, keyword extraction, and rewriting of the user query.

[0017] Preferably, step S4 includes a retrieval stage and a generation stage, wherein the retrieval stage is based on hybrid retrieval for document retrieval, and the generation stage is based on a large language model for context understanding and response generation.

[0018] Preferably, the retrieval stage employs a multi-channel recall mechanism, including preliminary recall based on keyword matching and semantic recall based on vector retrieval, and improves recall quality through a fusion ranking mechanism.

[0019] Compared with the prior art, the method of the present invention for improving the recall capability of the RAG system has the following beneficial technical effects:

[0020] 1) This invention, through the joint analysis of multiple tools, can effectively address the problem of complex layout structures and table content extraction, ensuring that the text and table information in PDF documents are accurately identified and preserved, providing a high-quality corpus foundation for subsequent retrieval.

[0021] 2) This invention uses a multi-scale semantic chunking strategy to divide document content into chunks at multiple levels of semantic granularity, such as sentence level, paragraph level, and topic level. Combined with chapter index and meta information, the index structure is more in line with human understanding logic, significantly improving the semantic alignment accuracy in the retrieval and generation stages.

[0022] 3) This invention employs a dynamic breakpoint algorithm based on vector similarity, combined with statistical methods such as percentiles, standard deviation, interquartile range, and gradient for chunk clustering, enabling the system to flexibly adjust the block boundaries according to changes in text semantics and adapt to the semantic features of different types of documents.

[0023] 4) This invention is based on a question rewriting strategy using a large language model (LLM). It uses few-shot learning technology to identify multiple questions and automatically generate subqueries, thereby decomposing and expanding the semantics of the query and significantly improving the recall completeness and relevance in multi-intent and complex query scenarios.

[0024] A second objective of this invention is to provide an apparatus for improving the recall capability of a RAG system, comprising:

[0025] The parsing module is used for structured parsing of various document types (including but not limited to PDF, Word, HTML, txt, md, etc.);

[0026] The chunking module is used to semantically hierarchically segment the document content and generate multi-granularity index units;

[0027] The query optimization module is used to semantically enhance and optimize queries in order to improve the accuracy and recall of search results.

[0028] The retrieval and generation module performs retrieval and generation tasks, thereby enabling the RAG system to achieve efficient retrieval and contextual understanding in complex document scenarios.

[0029] Preferably, the chunking module employs different chunking methods based on the semantic density, logical structure, and chapter level of the document, including sentence-level, paragraph-level, and topic-level chunking, to adapt to different types of query needs; it calculates semantic vector embedding for each chunk and constructs a vector index library based on Faiss or Milvus to support efficient similarity retrieval;

[0030] The query optimization module performs semantic parsing, keyword extraction, and rewriting of user queries;

[0031] The retrieval generation module performs document retrieval based on hybrid retrieval, and the generation sub-module performs contextual understanding and response generation based on a large language model. It adopts a multi-channel retrieval mechanism, including preliminary retrieval based on keyword matching and semantic retrieval based on vector retrieval, and improves the retrieval quality through a fusion ranking mechanism.

[0032] The present invention also provides a computer device for improving the recall capability of a RAG system, including a processor and a memory storing a plurality of computer instructions, characterized in that the computer instructions, when executed by the processor, implement the steps of the method for improving the recall capability of the RAG system. Attached Figure Description

[0033] Figure 1 This is a flowchart of a method for improving the recall capability of a RAG system according to an embodiment of the present invention;

[0034] Figure 2 This is a schematic diagram of the multi-scale semantic chunk strategy according to an embodiment of the present invention;

[0035] Figure 3 This is a schematic diagram of a device for improving the recall capability of a RAG system according to an embodiment of the present invention. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention.

[0038] Terminology Explanation:

[0039] Retrieval-Augmented Generation (RAG): Retrieval enhancement refers to the process of improving the efficiency and accuracy of information retrieval systems by applying advanced information retrieval technologies and artificial intelligence methods. In this field, technological innovation aims to enable retrieval systems to better understand the intent of user queries. Through technologies such as automatic learning, natural language processing, and data mining, it conducts deeper and more precise analysis of text and multimedia content to provide more relevant and personalized search results. The goal of retrieval enhancement is to optimize the information retrieval process, enabling users to obtain the information they need more quickly and accurately. It is applicable to various application scenarios, including search engines, document management systems, and knowledge base retrieval.

[0040] Example 1

[0041] like Figure 1 As shown, a method for improving the recall capability of a RAG system according to this embodiment includes the following steps:

[0042] S1. Use a document parsing tool to perform structured parsing on documents of various formats (including but not limited to PDF, Word, HTML, txt, md, etc.) to extract the main text, titles, tables, and images.

[0043] For PDF documents with complex structures and irregular layouts, it is preferable to use a combination of MinerU and RapidOCR tools for parsing to achieve accurate layout restoration and table content recognition. MinerU is used to perform structured parsing of the PDF file, extracting text content and document hierarchy. MinerU can recognize elements such as titles, paragraphs, lists, headers, and footers in PDFs, accurately restoring the document's logical structure and handling complex layouts, ensuring the completeness and accuracy of text extraction. For table or image areas detected in the document, MinerU marks them as structured blocks and outputs their corresponding coordinates. The table areas identified by MinerU are then processed by RapidOCR. RapidOCR uses OCR algorithms to perform text recognition and structure reconstruction of the tables, extracting row, column, and cell content to generate structured table data. This process ensures the complete extraction and semantic alignment of table information in the PDF, providing high-quality input for subsequent semantic segmentation and index construction.

[0044] Furthermore, other parsing tools can be selected, such as LayoutParser, PaddleOCR, DocTR, pdfplumber, Camelot, Tabula, etc., to enhance the system's versatility and robustness in multi-source document scenarios.

[0045] Furthermore, the data is processed using post-processing optimization methods such as rule-based pagination text merging strategies, meta information extraction, and text cleaning. The rule-based pagination text merging strategy combines text scattered across different pages into continuous paragraphs, ensuring text coherence and avoiding information breaks. Meta information extraction captures and stores document metadata such as author, title, and chapter index. Text cleaning removes irrelevant symbols and formatting, ensuring text purity.

[0046] S2. Employ a multi-scale semantic chunking strategy to perform semantic hierarchical segmentation of document content, generating multi-granularity index units.

[0047] S21. Segment the text according to symbols such as periods, semicolons, question marks, exclamation marks, and line breaks to generate multiple clauses. For example, a text containing multiple sentences will be segmented into multiple independent clauses, with each clause serving as a processing unit.

[0048] S22. Calculate the embedding for each clause using a vector model. The vector model can be a pre-trained model such as BERT or BGE. Calculate the vector representation of each clause, encoding each clause as a d-dimensional vector, denoted as:

[0049]

[0050] in This represents the vector representation of the i-th clause.

[0051] S23. Calculate the similarity between different clauses. Similarity calculation can use methods such as cosine similarity and Euclidean distance. For example, the formula for calculating the similarity between two clauses using cosine similarity is as follows:

[0052]

[0053] in, and These are the vector representations of the two clauses.

[0054] S24. Based on similarity, four breakpoint methods—based on percentiles, based on standard deviation, based on interquartile range, and based on gradient—are used to cluster the clauses, ultimately forming multiple chunks with independent semantics.

[0055] S25. Add document-related metadata to the chunks to enhance vector representation capabilities. For example, add chapter indexes, document categories, and other metadata as additional features to the vector representation of each chunk, thereby enhancing the vector representation capabilities and improving retrieval performance.

[0056] S3. A question rewriting strategy based on Large Language Model (LLM) is used to semantically enhance and optimize the input query in order to improve the accuracy and recall of the retrieval results.

[0057] The LLM (Limited Least Meaning) technique is used to determine if a question is a multi-question question. If so, the question is rewritten into multiple sub-questions, and multi-path recall is performed on each sub-question to improve the targeting and accuracy of the search. For example, for each sub-question, an independent search is performed using the RAG (Rapid Algorithm for Querying and Retrieving) model to generate a response for each sub-question. The search results for all sub-questions are then aggregated and deduplicated to ensure the accuracy and readability of the final response.

[0058] S4. Combine the optimized query with the multi-scale semantic index to perform retrieval and generation tasks, thereby enabling the RAG system to achieve efficient retrieval and contextual understanding in complex document scenarios.

[0059] The system weights and fuses recall results from different semantic scales (such as sentence-level, paragraph-level, and topic-level), dynamically adjusting the weights based on query intent and contextual relevance to ensure both comprehensive coverage and semantic accuracy. Combining the context matching and text semantic understanding capabilities of a large language model, candidate documents are re-ranked to generate a context-aware retrieval result set. This results in a final response that is semantically consistent, logically complete, and strongly contextually relevant.

[0060] Example 2

[0061] like Figure 2 As shown in the figure, this embodiment provides a multi-scale semantic chunking strategy in a method to improve the recall capability of a RAG system. The specific steps are as follows:

[0062] First, the input document is segmented using symbols such as periods, semicolons, question marks, exclamation marks, and line breaks to generate multiple clauses. For a text containing multiple sentences, the system divides it into several clauses, each treated as an independent processing unit. This process ensures the independence and integrity of each clause, providing a foundation for subsequent vector calculations and clustering.

[0063] Furthermore, clause embedding is calculated using a vector model to obtain a vector representation of each clause. By converting each clause into a high-dimensional vector, the system can further calculate the similarity between clauses, providing data support for subsequent breakpoint methods.

[0064] Furthermore, each clause is encoded into a d-dimensional vector using a vector model, denoted as:

[0065]

[0066] in This represents the vector representation of the i-th clause.

[0067] Furthermore, similarity is calculated based on the generated clause vectors, determining the similarity between different clauses. Similarity calculation can employ methods such as cosine similarity and Euclidean distance. Cosine similarity measures similarity by calculating the cosine of the angle between two vectors; a value closer to 1 indicates greater similarity. Euclidean distance measures similarity by calculating the straight-line distance between two vectors; a smaller value indicates greater similarity. For example, for the vectors of two clauses, the cosine similarity formula can be used:

[0068]

[0069] in, and These are the vector representations of the two clauses, It is the dot product of vectors. and It is the magnitude of the vector.

[0070] Furthermore, based on similarity, four breakpoint methods—percentile-based, standard deviation-based, interquartile range-based, and gradient-based—are employed to cluster clauses, ultimately forming multiple chunks with independent semantics. The percentile-based breakpoint method determines the breakpoint location by calculating the percentile of similarity. For example, calculating the 90th percentile of similarity and using clauses with similarity below this value as breakpoints to form new chunks. The standard deviation-based breakpoint method determines the breakpoint location by calculating the standard deviation of similarity. For example, calculating the standard deviation of similarity and using clauses with similarity below the average minus one standard deviation as breakpoints to form new chunks. The interquartile range-based breakpoint method determines the breakpoint location by calculating the interquartile range of similarity. For example, calculating the upper and lower quartiles of similarity and using clauses with similarity below the lower quartile as breakpoints to form new chunks. The gradient-based breakpoint method determines the breakpoint location by calculating the gradient of similarity changes. For example, the gradient of the similarity between adjacent clauses can be calculated, and the positions with larger gradient changes can be used as breakpoints to form new chunks.

[0071] Furthermore, a gradient-based breakpoint method determines breakpoint locations between clauses by calculating the gradient of similarity changes, thus forming multiple chunks with different semantic information. Specifically, the system calculates the rate of change of similarity between adjacent clauses; when the rate of change exceeds a preset threshold, that location is considered a breakpoint. The similarity gradient calculation is as follows:

[0072]

[0073] in, and They are the first The and the first The similarity of clauses This is the interval between clauses. If the gradient change exceeds a preset threshold (e.g., 0.1), then at the 1st... The and the first A breakpoint is set between each clause. This method can flexibly adapt to the semantic structure of different documents, improving the accuracy and diversity of chunks.

[0074] Furthermore, semantic vector embeddings are calculated for each chunk, and a vector index library based on Faiss or Milvus is built to support efficient similarity retrieval. Faiss is a high-performance vector retrieval engine that supports diverse index structures and can perform fast similarity retrieval in massive high-dimensional vectors, while Milvus is a distributed vector database suitable for long-term storage, distributed management, and horizontal scaling of massive semantic vectors.

[0075] Example 3

[0076] This embodiment provides a device for improving the recall capability of a RAG system, such as... Figure 3 As shown, it includes:

[0077] The parsing module is used for structured parsing of various document types (including but not limited to PDF, Word, HTML, txt, md, etc.);

[0078] The chunking module is used to semantically hierarchically segment the document content and generate multi-granularity index units;

[0079] The query optimization module is used to semantically enhance and optimize queries in order to improve the accuracy and recall of search results.

[0080] The retrieval and generation module performs retrieval and generation tasks, thereby enabling the RAG system to achieve efficient retrieval and contextual understanding in complex document scenarios.

[0081] Furthermore, the device's segmentation module employs different segmentation methods based on the document's semantic density, logical structure, and chapter hierarchy, including sentence-level, paragraph-level, and topic-level segmentation, to adapt to different types of query needs.

[0082] Furthermore, the device's chunking module calculates semantic vector embeddings for each chunk and builds a vector index library based on Faiss or Milvus to support efficient similarity retrieval.

[0083] Furthermore, the device's query optimization module performs semantic parsing, keyword extraction, and rewriting of user queries.

[0084] Furthermore, the device's retrieval and generation module performs document retrieval based on hybrid retrieval, while the generation sub-module performs contextual understanding and response generation based on a large language model.

[0085] Furthermore, the retrieval generation module of the device adopts a multi-channel recall mechanism, including preliminary recall based on keyword matching and semantic recall based on vector retrieval, and improves the recall quality through a fusion ranking mechanism.

[0086] It is understood that the device embodiments provided above correspond to the method embodiments described above, and the specific details can be referred to each other, which will not be repeated here.

[0087] Example 4

[0088] The present invention also provides a computer device, including a processor and a memory storing a plurality of computer instructions, wherein the computer instructions, when executed by the processor, implement the steps of the method for improving the recall capability of the RAG system.

[0089] For specific limitations on devices that enhance the recall capability of the RAG system, please refer to the limitations on methods for enhancing the recall capability of the RAG system mentioned above, which will not be repeated here.

[0090] The memory and processor are electrically connected directly or indirectly to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The memory stores a computer program that can run on the processor, which implements the method in the embodiments of the present invention by running the computer program stored in the memory.

[0091] The memory may be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory stores the program, and the processor executes the program upon receiving an execution instruction.

[0092] The processor may be an integrated circuit chip with data processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.

[0093] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for improving the recall capability of a RAG system, characterized in that, Includes the following steps: S1. Perform structured parsing of the document; S2. Employ a multi-scale semantic segmentation strategy to perform semantic hierarchical segmentation of document content, generating multi-granularity index units; S3. A query rewriting strategy based on a large language model is used to enhance and optimize the semantics of queries; S4. Combine the optimized query with the multi-scale semantic index to perform retrieval and generation tasks.

2. The method according to claim 1, characterized in that, In step S1, a document parsing tool is used to perform structured parsing on documents of various formats (including but not limited to PDF, Word, HTML, txt, and md) to extract their content, which includes text, titles, tables, and images.

3. The method according to claim 2, characterized in that, For PDF documents with complex structures and irregular layouts, it is preferable to use MinerU and RapidOCR together for parsing. MinerU is used to extract text content, hierarchical structure and paragraph relationships from PDF documents, while RapidOCR is used to identify and parse table content in PDF documents. The two work together to ensure the integrity and accuracy of the document structure and content.

4. The method according to claim 1, characterized in that, The multi-scale semantic segmentation strategy in step S2 adopts different segmentation methods based on the semantic density, logical structure and chapter level of the document, including sentence-level, paragraph-level and topic-level segmentation, to adapt to different types of query needs.

5. The method according to claim 4, characterized in that, The breakpoints for multi-scale semantic segmentation are determined based on vector semantic similarity, where the vector representations of any two adjacent semantic blocks are respectively... and When the cosine similarity is lower than a preset threshold, a semantic breakpoint is determined.

6. The method according to claim 1, characterized in that, After performing multi-scale semantic segmentation, semantic vector embeddings are calculated for each chunk, and a vector index library based on Faiss or Milvus is built to support efficient similarity retrieval.

7. The method according to claim 1, characterized in that, The query rewriting strategy based on the large language model in step S3 involves semantic parsing, keyword extraction, and rewriting of the user query.

8. The method according to claim 1, characterized in that, Step S4 includes a retrieval stage and a generation stage. The retrieval stage is based on hybrid retrieval for document recall, and the generation stage is based on a large language model for context understanding and response generation.

9. The method according to claim 8, characterized in that, The retrieval phase employs a multi-channel recall mechanism, including preliminary recall based on keyword matching and semantic recall based on vector retrieval, and improves recall quality through a fusion ranking mechanism.

10. An apparatus for enhancing the recall capability of a RAG system, comprising: The parsing module is used for structured parsing of various document types (including but not limited to PDF, Word, HTML, txt, md, etc.); The chunking module is used to semantically hierarchically segment the document content and generate multi-granularity index units; The query optimization module is used to semantically enhance and optimize queries in order to improve the accuracy and recall of search results. The retrieval and generation module performs retrieval and generation tasks, thereby enabling the RAG system to achieve efficient retrieval and contextual understanding in complex document scenarios.

11. The method according to claim 10, characterized in that, The segmentation module employs segmentation methods at different scales, including sentence-level, paragraph-level, and topic-level segmentation, based on the semantic density, logical structure, and chapter level of the document, to adapt to different types of query needs. Calculate semantic vector embeddings for each chunk and build a vector index library based on Faiss or Milvus to support efficient similarity retrieval; The query optimization module performs semantic parsing, keyword extraction, and rewriting of user queries; The retrieval generation module performs document retrieval based on hybrid retrieval, and the generation sub-module performs contextual understanding and response generation based on a large language model. It adopts a multi-channel retrieval mechanism, including preliminary retrieval based on keyword matching and semantic retrieval based on vector retrieval, and improves the retrieval quality through a fusion ranking mechanism.

12. A computer device for improving the recall capability of a RAG system, comprising a processor and a memory storing a plurality of computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method for improving the recall capability of the RAG system as described in any one of claims 1 to 9.