Method and system for improving RAG generation effect

By constructing a graph database and community summaries, and combining local and global indexing techniques, the accuracy and efficiency of RAG in complex queries have been improved, solving the problem of poor information integration in existing technologies and achieving more efficient information retrieval and generation.

CN120873202APending Publication Date: 2025-10-31DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510962577.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

When the knowledge base is of poor quality, user requests are ambiguous, and information integration is inadequate, the generated answers lack coherence and accuracy, the retrieval recall is low, and the information is redundant, making it difficult to handle complex queries.

Method used

By collecting and parsing documents to generate text blocks, constructing a knowledge graph database and community summaries, and using the knowledge graph for retrieval and response generation, combined with local and global indexing technologies, the ability to understand and respond to information fragment relationships is improved.

Benefits of technology

It improves the accuracy and efficiency of RAG in complex queries, enables faster location of relevant information, simplifies application development, provides better interpretability and traceability, and is suitable for multi-level analysis and reasoning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873202A_ABST
    Figure CN120873202A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for improving the RAG generation effect, and relates to the technical field of information retrieval, and the method specifically comprises the steps: 1, collecting and analyzing a document, and generating a text block; 2, extracting a ternary body from the text block, and constructing an atlas database and a community abstract; and 3, receiving a user question, performing question judgment on the user question, determining a question type, performing retrieval by using a graph database or a community abstract according to the question type, and generating an answer. According to the method, based on the knowledge graph, the relationship between different information fragments is captured and utilized, richer context reasoning is provided, and complex queries are more accurately understood and responded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information retrieval technology, and more specifically to a method and system for improving the performance of RAG generation. Background Technology

[0002] RAG technology, or Retrieval-Augmented Generation, is an artificial intelligence technique that combines information retrieval and text generation. By combining the inherent knowledge of large language models (LLMs) with non-parametric data from external databases, RAG technology improves the accuracy and reliability of models in knowledge-intensive tasks. Its core lies in enhancing information processing efficiency and accuracy by combining retrieval and generation. RAG technology can not only integrate massive amounts of data but also generate more precise content. It plays a crucial role in various natural language processing tasks, including question-answering systems, document generation, and intelligent assistants. It has demonstrated significant application value in areas such as enterprise knowledge management systems, online question-answering systems, and information retrieval systems.

[0003] RAG technology improves the accuracy, relevance, and timeliness of generated content. It allows AI systems to generate more context-aware responses by retrieving specific information relevant to the request. Implementing RAG is more efficient than continuously retraining LLM using new data.

[0004] However, RAG technology also has some limitations:

[0005] (1) The performance of RAG technology is highly dependent on the quality of the knowledge base. If there are errors or outdated information in the knowledge base, the fragments retrieved by RAG may also be inaccurate or irrelevant.

[0006] (2) The ambiguity of user requests, the limitations of the expressive power of the embedded model, and the poor quality of external knowledge bases may result in low relevance between the retrieved information and the user's question. In addition, low retrieval recall and information redundancy are also common problems, which limit the generative model from obtaining enough background information to construct a complete answer.

[0007] (3) During the generation phase, RAG needs to effectively integrate user input with retrieved information. Poor integration may result in a lack of coherence, insight, or comprehensive information in the generated answers.

[0008] Therefore, how to improve the effectiveness of RAG retrieval is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0009] In view of this, the present invention provides a method and system for improving the performance of RAG generation, which captures and utilizes the relationships between different information fragments based on knowledge graphs, provides richer contextual reasoning, and more accurately understands and responds to complex queries.

[0010] To achieve the above objectives, the present invention adopts the following technical solution:

[0011] A method to improve RAG generation quality includes the following steps:

[0012] Step 1: Collect and parse the document to generate text blocks;

[0013] Step 2: Extract ternaries from text blocks to construct a graph database and community summary;

[0014] Step 3: Receive user questions, classify the questions, determine the question type, and use the graph database or community summary to search based on the question type to generate answers.

[0015] Preferably, the document parsing process in step 1 includes:

[0016] Step 11: Process the document layout and format;

[0017] Step 12: Recognize the text within the images in the document and convert it into editable text;

[0018] Step 13: Using natural language processing techniques and contextual information, combined with the document's layout structure and element relationships, generate the text reading order;

[0019] Step 14: Divide the text into blocks to obtain several text blocks.

[0020] Preferably, the ternary body includes entities, relations, and claims. In step 2, the entities, relations, and claims of each text block are extracted to construct a local index and obtain a graph database. Based on the graph database, a global index is constructed to generate a community summary.

[0021] Preferably, the process of constructing a local index includes:

[0022] Step 211: Use the large extraction model to extract the entities, relationships, and claims of each text block;

[0023] Step 212: Based on the preset types of entities and relationships, generate entity descriptions for the extracted entities and relationship descriptions for the extracted relationships;

[0024] Step 213: Merge the extracted entities and relations with the same name respectively, retain one entity and one relation, and merge all entity descriptions with the same name using the large model to generate a new entity description that retains the entity, and merge all relation descriptions with the same name using the large model to generate a new relation description that retains the relation.

[0025] Step 214: Use the detection model to determine if there are any missing entities in the text block. If so, modify the prompt words extracted from the large model and return to step 211 until the preset number of iterations is reached and proceed to step 215; otherwise, proceed to step 215.

[0026] Step 215: Generate entity covariates and community reports based on the extracted entities and corresponding descriptions, construct a knowledge graph, and store it as a graph database.

[0027] Preferably, the global index construction process includes:

[0028] Step 221: Use the Leiden algorithm to recursively divide the knowledge graph in the graph database into multiple communities;

[0029] Step 222: Generate a corresponding community summary for each community using the summary big data model. Input the text segment corresponding to the community into the summary big data model to generate the community summary.

[0030] Preferably, in step 3, the user question is classified into a specific question or a summary question. If the user question is a specific question, a local index search is performed; if the user question is a summary question, a global index search is performed.

[0031] Preferably, the process of performing a local index retrieval includes:

[0032] Step 311: Identify semantically relevant entities from the knowledge graph based on the user's question;

[0033] Step 312: Extract connected entities, relationships, entity covariates, and community reports related to the identified entities from the knowledge graph;

[0034] Step 313: Extract text blocks related to the identified entities from the original input document, and use the identified entities, connected entities, relations, entity covariates, community reports and text blocks as candidate data sources;

[0035] Step 314: Prioritize and filter candidate data sources according to a predefined context window.

[0036] A single context window of a predefined size is used to generate a response to a user query by sorting and filtering.

[0037] Step 315: Input the sorted and filtered candidate data sources into the large model to obtain the answer and return it to the user.

[0038] Preferably, the process of performing a global index retrieval includes:

[0039] Step 321: Load all community summaries and all associated entities, use the sum of the occurrence frequencies of all entities associated with each community summary as the corresponding weight, randomly shuffle all community summaries, sort all community summaries according to the weight and divide them into different batches;

[0040] Step 322: Divide each batch of community summaries into several text blocks, and input each text block and user question into the large generation model. Each text block corresponds to one answer.

[0041] Step 323: Calculate the contribution of all answers to the user's question, and filter out high-scoring answers based on the set threshold;

[0042] Step 324: Populate the context window with high-scoring answers according to their contribution, as the context for high-scoring answers;

[0043] Step 325: Input the context of the high-scoring answer and the user's question into the large-scale model to generate an answer. Prioritize the context window based on contribution to avoid excessively long inputs. Simultaneously, aggregate the high-scoring answer as context and call the LLM to generate a coherent and comprehensive final response.

[0044] Preferably, all large models can use the Qwen family of large language models.

[0045] A system for improving RAG generation performance includes a file parsing module, an index building module, and a knowledge question answering module. The file parsing module collects and parses documents to generate text blocks. The index building module extracts ternaries from the text blocks to build a graph database and a community summary. The knowledge question answering module receives user questions, classifies them, determines the question type, and uses the graph database or community summary to retrieve answers based on the question type.

[0046] Preferably, the file parsing module includes a layout analysis unit, a text recognition unit, a table structure recognition unit, and a splitting unit; the layout analysis unit processes the document layout and format; the text recognition unit recognizes the text within images in the document and converts it into editable text; the table structure recognition unit uses natural language processing technology and contextual information, combined with the document's layout structure and element relationships, to generate a text reading order; and the splitting unit divides the text into blocks to obtain several text blocks.

[0047] Preferably, the index building module includes a local index building unit and a global index building unit; the local index building unit extracts the entities, relations and claims of each text block, builds a knowledge graph and stores it as a graph database; the global index building unit divides the knowledge graph into communities and generates a corresponding community summary for each community.

[0048] As can be seen from the above technical solutions, compared with the prior art, this invention discloses a method and system for improving RAG generation performance. Knowledge graph-based RAGs can capture and utilize the relationships between different information fragments, providing richer context and more accurately understanding and responding to complex queries. The graph structure of the knowledge graph enables knowledge graph-based RAGs to follow relational chains, facilitating more complex reasoning. This multi-hop reasoning capability allows the system to connect different information fragments through multiple intermediate steps to reach conclusions or generate answers. Compared to document structures, graph structures can more naturally represent hierarchical and non-hierarchical relationships between entities, making knowledge graph-based RAGs more closely resemble real-world knowledge organization methods when representing knowledge. Specifically, the beneficial effects of this invention include:

[0049] (1) For query types involving relational traversal, graph structures can significantly improve processing efficiency. The graph traversal algorithm of knowledge graph-based RAG enables faster location of relevant information during retrieval. (2) Once the knowledge graph is created, it is easier to build and maintain RAG applications. The application development of knowledge graph-based RAG is simpler because the data is visible and the graph has good iterability and scalability. (3) Knowledge graph-based RAG shows obvious advantages in handling tasks that require deep contextual understanding and complex relational analysis, especially in cases that require multi-level analysis and reasoning. (4) Knowledge graph-based RAG also shows advantages in governance, including better interpretability, traceability and access control, which are crucial for many industries. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0051] Figure 1 This invention provides a flowchart illustrating a method for improving RAG generation performance.

[0052] Figure 2 This is a schematic diagram of the local index retrieval answer process provided by the present invention;

[0053] Figure 3 This is a schematic diagram of the global index retrieval and response process provided by the present invention. Detailed Implementation

[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] This invention discloses a method for improving RAG generation quality, such as... Figure 1 As shown, it includes the following steps:

[0056] S1: Collect and parse the document to generate text blocks;

[0057] S2: Extract ternaries from text blocks to construct a graph database and community summary;

[0058] S3: Receive user questions, classify the questions, determine the question type, and use the graph database or community summary to search based on the question type to generate an answer.

[0059] Furthermore, the document parsing process in S1 includes:

[0060] S11: Process document layout and format; documents include various types such as docx and pdf; handle complex layouts and formats in documents, including various elements such as text, images, graphics, and tables, and solve problems such as element overlapping, diverse elements and complex layouts, to ensure the accuracy and completeness of the parsing results;

[0061] S12: Recognize the text within images in the document and convert it into editable text; accurately recognize text in images using Optical Character Recognition (OCR) technology and convert it into editable text;

[0062] S13: Using natural language processing technology and contextual information, combined with the document's layout structure and element relationships, a text reading order is generated;

[0063] S14: Divide the text into blocks to obtain several text blocks. Dividing the document into text blocks facilitates subsequent index construction and control over the number of characters input to large models.

[0064] Furthermore, the ternary structure includes entities, relations, and claims. In S2, the entities, relations, and claims of each text block are extracted to construct a local index, resulting in a graph database. Based on the graph database, a global index is constructed to generate a community summary. The local index is used to retrieve a specific question in the knowledge base, while the global index is used to retrieve summary questions.

[0065] Furthermore, the process of building a local index includes:

[0066] S211: Use a large model to extract the entities, relationships, and claims of each text block;

[0067] S212: Based on the preset types of entities and relationships, generate entity descriptions for the extracted entities and relationship descriptions for the extracted relationships;

[0068] S213: Merge the extracted entities and relations with the same name separately, retain one entity and one relation, and merge all entity descriptions with the same name using the large model to generate a new entity description that retains the entity, and merge all relation descriptions with the same name using the large model to generate a new relation description that retains the relation.

[0069] S214: Use the large model to determine if there are any missing entities in the text block. If so, modify the prompt words in the large model in S211 and return to S211 until the preset number of iterations is reached and enter S215; otherwise, enter S215; enter S215 after the number of iterations reaches 5.

[0070] S215: Generate entity covariates and community reports based on the extracted entities and corresponding descriptions, construct a knowledge graph, and store it as a graph database.

[0071] Furthermore, the local index constructed in the previous steps can be viewed as a uniformly undirected weighted graph. The process of constructing the global index includes:

[0072] S221: The Leiden algorithm is used to recursively divide the knowledge graph in the graph database into multiple communities. The connections between entities within a community are closer than those with other entities outside the community. The Leiden algorithm can effectively mine the hierarchical community structure of large-scale graphs.

[0073] S222: Utilizing a large model to generate a corresponding community summary for each community helps to gain a macro-level understanding at different levels of detail in the graph. Input the text segment corresponding to the community into the summary large model to generate the community summary.

[0074] Furthermore, S3 performs problem identification on user questions, determining whether the user question belongs to a specific question or a summary question. When the user question belongs to a specific question, a local index search is performed; when the user question belongs to a summary question, a global index search is performed.

[0075] Furthermore, during local index retrieval, knowledge graphs and unstructured text of user questions are used for enhancement, enabling entity-based reasoning. This leverages structured data from the knowledge graph and unstructured data from the input document to provide query-relevant entity information to the large language model (LLM) during the query process; for example... Figure 2 As shown, the process of performing a local index retrieval includes:

[0076] S311: Based on the user's question or in conjunction with the dialogue history, identify semantically relevant entities from the knowledge graph; the identified entities serve as entry points into the knowledge graph, thereby further extracting relevant details;

[0077] S312: Extract connected entities, relationships, entity covariates, and community reports related to the identified entities from the knowledge graph;

[0078] S313: Extract text chunks related to the identified entities from the original input document, and use the identified entities, connected entities, relations, entity covariates, community reports and text chunks as candidate data sources;

[0079] S314: Prioritize and filter candidate data sources to fit a single context window of a predefined size, which is used to generate a response to the user query;

[0080] S315: Input the sorted and filtered candidate data sources into a large language model, obtain the answer, and return it to the user.

[0081] Furthermore, community summaries are used to perform overall question reasoning on the entire corpus. A knowledge graph generated by LLM is used to organize and aggregate information. The community summary set generated by LLM is used as context data to generate responses in a map-reduce manner. In the map step, community reports are divided into predefined-sized text blocks. Each text block is used to generate an intermediate response containing a list of points, each point having a corresponding numerical score indicating its importance. In the reduce step, the most important points are selected from the intermediate responses and aggregated as the context for generating the final response to answer queries requiring cross-dataset information aggregation. Figure 3 As shown, the process of performing a global index retrieval includes:

[0082] S321: Load all community summaries and all associated entities, use the sum of the occurrence frequencies of all entities associated with each community summary as the corresponding weight, randomly shuffle all community summaries, sort all community summaries according to the weight and divide them into different batches;

[0083] S322: Divide each batch of community summaries into several text blocks, and input each text block and user question into the large generation model. Each text block corresponds to one answer.

[0084] S323: Calculate the contribution of all answers to the user's question and filter out high-scoring answers based on a set threshold;

[0085] S324: Fill the context window with high-scoring answers according to their contribution, and use it as the context for high-scoring answers;

[0086] S325: Input the context of high-scoring answers and the user's question into the large-scale generation model to generate an answer. Prioritize the context window based on contribution to avoid excessively long inputs, while aggregating high-scoring answers as context, and calling LLM to generate a coherent and comprehensive final response.

[0087] Furthermore, all large models can be selected from the Qwen family of large language models.

[0088] On the other hand, in one specific embodiment, a system for improving RAG generation includes a file parsing module, an index building module, and a knowledge question answering module; the file parsing module collects and parses documents to generate text blocks; the index building module extracts ternaries from the text blocks to build a graph database and a community summary; the knowledge question answering module receives user questions, classifies the user questions, determines the question type, and uses the graph database or community summary to retrieve answers based on the question type.

[0089] Furthermore, the document parsing module includes a layout analysis unit, a text recognition unit, a table structure recognition unit, and a splitting unit; the layout analysis unit processes the document layout and format; the text recognition unit recognizes the text within images in the document and converts it into editable text; the table structure recognition unit uses natural language processing technology and contextual information, combined with the document's layout structure and element relationships, to generate a text reading order; and the splitting unit divides the text into blocks to obtain several text blocks.

[0090] Furthermore, the index building module includes a local index building unit and a global index building unit; the local index building unit extracts the entities, relations and claims of each text block, builds a knowledge graph and stores it as a graph database; the global index building unit divides the knowledge graph into communities and generates a corresponding community summary for each community.

[0091] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0092] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for improving RAG generation quality, characterized in that, Includes the following steps: Step 1: Collect and parse the document to generate text blocks; Step 2: Extract ternaries from text blocks to construct a graph database and community summary; Step 3: Receive user questions, classify the questions, determine the question type, and use the graph database or community summary to search based on the question type to generate answers.

2. The method for improving RAG generation effect according to claim 1, characterized in that, The document parsing process in step 1 includes: Step 11: Process the document layout and format; Step 12: Recognize the text within the images in the document and convert it into editable text; Step 13: Using natural language processing techniques and contextual information, combined with the document's layout structure and element relationships, generate the text reading order; Step 14: Divide the text into blocks to obtain several text blocks.

3. The method for improving RAG generation effect according to claim 1, characterized in that, The ternary structure includes entities, relations, and claims. In step 2, the entities, relations, and claims of each text block are extracted to construct a local index and obtain a graph database. Based on the graph database, a global index is constructed to generate a community summary.

4. The method for improving RAG generation effect according to claim 3, characterized in that, The process of building a local index includes: Step 211: Use the large extraction model to extract the entities, relationships, and claims of each text block; Step 212: Based on the preset types of entities and relationships, generate entity descriptions for the extracted entities and relationship descriptions for the extracted relationships; Step 213: Merge the extracted entities and relations with the same name, and rewrite the entity description of the merged entity and the relation description of the merged relation. Step 214: Use the detection model to determine if there are any missing entities in the text block. If so, modify the prompt words extracted from the large model and return to step 211 until the preset number of iterations is reached and proceed to step 215; otherwise, proceed to step 215. Step 215: Generate entity covariates and community reports based on the extracted entities and corresponding descriptions, construct a knowledge graph, and store it as a graph database.

5. The method for improving RAG generation effect according to claim 4, characterized in that, The process of building a global index includes: Step 221: Use the Leiden algorithm to recursively divide the knowledge graph in the graph database into multiple communities; Step 222: Use the summary big model to generate a corresponding community summary for each community.

6. The method for improving RAG generation effect according to claim 5, characterized in that, In step 3, the user's question is classified as either a specific question or a summary question. If the user's question is a specific question, a local index search is performed; if the user's question is a summary question, a global index search is performed.

7. The method for improving RAG generation effect according to claim 6, characterized in that, The process of performing a local index retrieval includes: Step 311: Identify semantically relevant entities from the knowledge graph based on the user's question; Step 312: Extract connected entities, relationships, entity covariates, and community reports related to the identified entities from the knowledge graph; Step 313: Extract text blocks related to the identified entities from the original input document, and use the identified entities, connected entities, relations, entity covariates, community reports and text blocks as candidate data sources; Step 314: Prioritize and filter candidate data sources according to a predefined context window; Step 315: Input the sorted and filtered candidate data sources into the large model to obtain the answer.

8. The method for improving RAG generation effect according to claim 6, characterized in that, The process of performing a global index retrieval includes: Step 321: Load all community summaries and all associated entities, use the sum of the occurrence frequencies of all entities associated with each community summary as the corresponding weight, randomly shuffle all community summaries, sort all community summaries according to the weight and divide them into different batches; Step 322: Divide each batch of community summaries into several text blocks, and input each text block and user question into the large generation model. Each text block corresponds to one answer. Step 323: Calculate the contribution of all answers to the user's question, and filter out high-scoring answers based on the set threshold; Step 324: Populate the context window with high-scoring answers according to their contribution, as the context for high-scoring answers; Step 325: Input the context of the high-scoring answer and the user's question into the large-scale generation model to generate the answer.

9. A system for improving RAG generation quality, characterized in that, A method for improving RAG generation performance according to any one of claims 1-8 includes a file parsing module, an index building module, and a knowledge question answering module; The file parsing module collects and parses documents, generating text blocks. The index building module extracts ternary elements from text blocks, constructs a graph database, divides the graph database into communities, and summarizes a corresponding community summary for each community. The knowledge Q&A module receives user questions, classifies them, determines the question type, and uses a graph database or community summary to search for answers based on the question type.

10. The system for improving RAG generation effect according to claim 9, characterized in that, The file parsing module includes a layout analysis unit, a text recognition unit, a table structure recognition unit, and a splitting unit; The layout analysis unit processes document layout and formatting; The text recognition unit identifies the text within images in a document and converts it into editable text. The table structure recognition unit uses natural language processing technology and contextual information, combined with the document's layout structure and element relationships, to generate the text reading order; The text is divided into blocks to obtain several text blocks.

Citation Information

Cited By

  • Information processing method and device, computer equipment, computer readable storage medium and computer program product

    CN121979966A