Graph retrieval enhancement generation method simultaneously utilizing graph and tree structure
By combining the graph and tree structure method, using large language model chunking processing and SpaCy entity extraction, a two-way index is established, and an adaptive search strategy is adopted to solve the problem of inefficiency in the existing methods, and efficient graph retrieval enhancement generation is achieved, which is suitable for intelligent question-and-answer in long documents.
Patent Information
- Application Number
- CN202510713367.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-15
AI Technical Summary
The existing graph retrieval enhancement generation method is time-consuming in the indexing stage, lacks flexible search mode, is difficult to take into account query efficiency and semantic coverage, and the search process that relies on large language models is inefficient.
Using a method of combining graphs and tree structures, documents are processed in blocks through large language models and abstract trees and entity diagrams are built, SpaCy is used for entity extraction, bidirectional index is established, and a retrieval path is selected using adaptive search strategies to reduce calls to large language models.
It has achieved a 10x speed increase in the indexing stage and a 100x speed increase in the search stage, while maintaining the accuracy of Q&A, which is suitable for long documents and intelligent Q&A.
Smart Images

Figure CN120492592A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing, information mining and retrieval enhancement generation, and in particular to a high-efficiency and high-effective graph retrieval enhancement generation method. Background Art
[0002] With the rapid development of natural language processing (NLP) technology, large language models (LLMs) have become a core pillar of the field, demonstrating outstanding performance in a variety of tasks, including text summarization, machine translation, and question-answering systems. However, current large language models still have significant limitations, such as being prone to hallucinations and often performing poorly when processing domain-specific knowledge or very long texts. To address these issues, researchers have proposed "Retrieval-Augmented Generation" (RAG) technology, which provides large language models with relevant content from external knowledge sources during the generation phase, combining it with contextual learning capabilities to improve the accuracy and professionalism of responses.
[0003] Traditional RAG methods typically retrieve a small number of fragments from the original document as supplementary knowledge. However, due to the limited context, these methods struggle to provide a global understanding of the entire knowledge base. For example, in a novel question-answering task, when summarizing a character's personality changes, traditional methods can only retrieve partial fragments and are unable to model and judge the character's complete behavioral chain.
[0004] To make up for this shortcoming, researchers have proposed graph-structured RAG methods, such as GraphRAG, RAPTOR, and LightRAG. This type of method usually adopts an index-retrieval two-stage architecture: first, a graph structure index is constructed through a large language model or clustering method, and then global or local retrieval is performed in combination with the graph structure. Among them, GraphRAG attempts to extract entities and relationships directly from documents to build a multi-level knowledge graph; RAPTOR focuses on the construction of a document summary tree and focuses on the semantic extraction of the entire document; LightRAG attempts to complete multi-granularity entity and relationship extraction in a single large language model call in order to reduce indexing costs. However, the above methods all have problems such as insufficient efficiency, high cost of graph structure construction, or lack of flexible retrieval mechanisms.
[0005] In addition, existing methods generally rely on manually specified retrieval modes (such as "global retrieval" or "local retrieval"), lack adaptability and query granularity adaptation capabilities, and it is difficult to simultaneously take into account query efficiency and semantic coverage.
[0006] Therefore, the current key technical bottlenecks include:
[0007] 1) Low retrieval efficiency: Existing graph-structured RAG methods rely heavily on LLM calls during the indexing phase. Building entity graphs or clustering structures is time-consuming and difficult to apply to large-scale text.
[0008] 2) Single structure utilization: Most methods only build tree structures or graph structures, lacking in-depth exploration of how to combine the two;
[0009] 3) Fixed retrieval mode: It relies heavily on preset retrieval paths or strategies and lacks a retrieval mechanism that automatically adapts to entity structures.
[0010] To address the above problems, there is an urgent need for a graph retrieval enhancement generation method with efficient indexing capabilities, flexible retrieval path selection, and both local and global understanding capabilities. Summary of the Invention
[0011] This paper proposes a graph-based augmented generation (RAG) method that leverages both graphs and trees, aiming to improve the efficiency and effectiveness of retrieval-augmented generation (RAG). This method achieves a 10x speed increase in indexing and a 100x speed increase in retrieval without sacrificing question-answering accuracy, making it suitable for scenarios such as intelligent question-answering of long documents.
[0012] The specific technical solution for achieving the purpose of the present invention is:
[0013] A graph retrieval enhancement generation method utilizing both graph and tree structures, the method comprising the following steps:
[0014] Step 1: Document segmentation: Use the tokenizer corresponding to the large language model to convert the long document serving as the knowledge base into tokens that the large language model can understand. Then, the long text document is divided into multiple document blocks with overlapping areas based on the number of tokens. Specifically, each document block has 1200 tokens, and adjacent document blocks overlap by 100 tokens.
[0015] Step 2: Use the large language model to recursively summarize the document blocks obtained in step 1. Specifically, five adjacent document blocks are fed into the large language model, and then summarized using the leaf-level summary prompt words to obtain summary information of the document block content, namely the summary block. After all original document blocks are summarized, the resulting summary blocks are fed into the large language model in groups of five, and the summary prompt words for the summary blocks are used to obtain further summary information of the input. This process continues until only one document block that summarizes the entire document remains. A document summary tree is constructed, where the leaf layer is the original document content. The closer to the root node, the more summarized the node content is, and the more global information it contains.
[0016] Step 3: Use natural language processing tools to extract entities from each document block obtained in Step 1. Two entities appearing in the same sentence are treated as having an undirected edge. An entity graph is extracted from each document block. The entity graphs corresponding to all document blocks are then merged together to obtain an entity graph for the entire long document. Steps 2 and 3 are performed in parallel to save computing time.
[0017] Step 4: Build a bidirectional index, including an entity-to-document-block mapping index and a document-block-to-entity mapping index. These two indexes allow us to locate the corresponding entity from the document block and the corresponding document block from the entity, thus bridging the document summary tree and the entity graph. Specifically, for each entity, multiple document blocks are extracted for that entity, and for each document block, multiple entities are extracted.
[0018] Step 5: After receiving the query, use the natural language processing tool Spacy to extract the entities in the query, and use this entity to directly match the nodes in the extracted entity graph;
[0019] Step 6: Based on the structural relationships between query entities in the entity graph, an adaptive retrieval strategy is used. Depending on the degree of association between entities on the graph, a global search method based on vector retrieval or a method based on the mapping index positioning of entities to document blocks constructed in step 4 is used to search for the corresponding candidate document blocks.
[0020] Step 7: Sort the candidate document blocks based on entity coverage and frequency of occurrence, and prioritize document blocks that cover the query intent.
[0021] Step 8: Integrate and format the selected document blocks and entity information, and input them into the large language model as prompt information in a structured form to enhance the large language model's understanding of relevant information and generate answers to questions.
[0022] Furthermore, the leaf layer summary prompt words in step 2 specifically include:
[0023] 2.1.1 Task Definition: The large language model is required to summarize the input text content and ensure that the summary length does not exceed 1200 tokens;
[0024] 2.1.2 Output Format: The large language model is required to output the summary directly in text format without outputting other information;
[0025] The summary words for the summary block described in step 2 specifically include:
[0026] 2.2.1 Task Definition: The large language model is required to further summarize the input text content, find the relevance between these summaries, grasp the main content, and ensure that the summary length does not exceed 1200 tokens;
[0027] 2.2.2 Output Format: The large language model is required to output the summary directly in text format without outputting other information.
[0028] Furthermore, the adaptive search strategy described in step 6 includes the following judgment and execution processes:
[0029] 6.1 Determine whether an entity was extracted from the question. If no entity was found, directly use the pre-trained SentenceBERT model to obtain the text representation. Then, use similarity retrieval to retrieve the corresponding n document blocks or summary blocks from the document summary tree. n is a positive integer, ranging from 5 to 20. Step 6 ends. If an entity was extracted from the question, i.e., the question entity, proceed to 6.2.
[0030] 6.2 Consider all unordered question entity pairs among all question entities. Then, check the distance between the entities in the question entity pairs on the constructed entity graph for the entire long document. If the two entities are neighbors within h hops, where h hops means that from one entity, there are at most h edges to reach the other, then the two entities are considered related and continue with step 6.3.1 for these two entities. If the two entities are neighbors within h hops, then the two entities are considered unrelated and removed from the candidate question entity pairs. If no candidate question entity pairs meet the requirements after this step, proceed to step 6.3.2.
[0031] 6.3.1. Index each entity in the entity pair using the entity-to-document block mapping constructed in step 4 to obtain the document block sets corresponding to the two entities. Then, intersect these two document block sets to obtain the document blocks corresponding to both entities as candidate document blocks. If the number of candidate document blocks is not greater than k, step 6 ends. If the number of candidate document blocks is greater than k, proceed to step 6.4. k is a positive integer that depends on the input length limit of the large language model used and is at most 25.
[0032] 6.3.2 For candidate question entity pairs that do not meet the requirements, first use the pre-trained SentenceBERT model to obtain text representations. Then, use similarity retrieval to retrieve the 2n document blocks or summary blocks with the highest semantic similarity to the question from the document summary tree. For each retrieved document block, use the frequency of the question entity in the document block as its weight. For each retrieved summary block, use the sum of the weights of its child nodes as its weight. For each summary block node, if its child nodes are also summary blocks, recursively calculate the sum of the weights of their child nodes until a document block is encountered. Then, sort the retrieved document blocks and summary blocks in descending order of weight, select the first n as candidate document blocks, and end in step 6.
[0033] 6.4 If the number of candidate document blocks is greater than k, further filtering is performed; let h = h - 1, and then check the distance between the entities in the problem entity pair on the constructed entity graph of the entire long document. If the two entities are neighbors within h hops, they are considered related and continue to perform step 6.3.1 on these two entities; if the two entities are neighbors more than h hops, they are considered unrelated and are deleted from the candidate problem entity pairs; if after this step, there are no candidate problem entity pairs that meet the requirements, perform 6.5; if after this step, there are still more than k candidate document blocks, then repeat 6.4;
[0034] 6.5 If no relevant entities can be retrieved after further filtering in step 6.4, then the candidate document blocks obtained before executing step 6.4 are sorted and filtered as follows:
[0035] First, use the document block to entity mapping index built in step 4 to obtain the type of question entity contained in each document block as the weight of this document block 1;
[0036] Then execute 6.3.2 to obtain the ranking based on the frequency of the problem entity in the document block, which is used as the weight 2 of this document block;
[0037] Finally, all candidate document blocks are sorted, first in descending order of weight 1. If weight 1 is the same, they are sorted in descending order of weight 2. Finally, the first n document blocks are selected as candidate document blocks, and step 6 ends.
[0038] Furthermore, the selected document blocks and entity information are integrated and formatted as described in step 8, as follows:
[0039] 8.1: Organize the search results according to the structure of "entity 1-entity 2: document block content", use the extracted entity pairs as topics, and the related document blocks as supplementary information, and construct prompts in a structured format to input to the large language model;
[0040] 8.2: To reduce redundant input, if a document block is referenced by multiple entity pairs, these entities will be merged into a unified representation in the form of "entity 1-entity 2-...-entity n" to avoid adding the same document block repeatedly;
[0041] 8.3: In the candidate document block set, identify adjacent continuous document blocks and merge them to reduce cross-block duplication and context fragment overlap, thereby further compressing the prompt word length;
[0042] 8.4: Calculate entity coverage for all entity pairs, prioritize entity pairs that cover more entities in the query, and sort based on this;
[0043] 8.5: After the entity pairs are sorted, the corresponding document blocks are arranged in order according to the order of the original document blocks in the document to form the final large language model input prompt information, thereby enhancing the quality of the answers generated by the large language model.
[0044] Advantages compared to existing methods:
[0045] 1) Significantly Improved Indexing Efficiency: This method uses traditional NLP tools (such as SpaCy) instead of large language models for entity extraction. It also employs a concise, recursive summary tree construction strategy, requiring only a minimal number of LLM calls to complete indexing. Compared to methods like GraphRAG, this method achieves up to a 10x improvement in indexing speed, significantly reducing computational resources and time costs.
[0046] 2) Support for flexible adaptive retrieval mechanism: This paper proposes an adaptive retrieval strategy based on the entity graph structure, which can automatically select "local retrieval" or "global retrieval" paths according to the connectivity between entities in the query. This avoids the limitations of existing methods that rely on manually set retrieval modes and achieves higher query adaptability and semantic coverage capabilities.
[0047] 3) Reducing the reliance on large language models during retrieval: This method primarily uses the large language model in the summary tree construction phase. During the retrieval phase, structured indexes (entity graph, summary tree, index mapping) are fully used for query processing. This avoids the problem of multiple rounds of LLM calls during the retrieval process in traditional graph structures (RAGs), making the overall method more controllable and practical.
[0048] 4) Efficient construction of multi-granularity semantic structures: The summary tree captures global semantic hierarchical information, while the entity graph captures fine-grained relationship information. The two work together to significantly improve the semantic integrity and interpretability of the generation process, making up for the shortcomings of single-structure methods (such as building only trees or graphs) in semantic modeling.
[0049] 5) Multiple compression strategies reduce model input redundancy: This paper introduces compression strategies such as "entity pair merging" and "continuous block merging" in prompt word construction to avoid repeated input, reduce the number of input tokens for large language models, improve inference efficiency, and retain sufficient semantic context.
[0050] 6) Excellent performance-efficiency trade-off: Experimental results show that this method achieves a 100× retrieval speed improvement in multiple long-document question answering tasks (such as NovelQA and InfiniteQA) while maintaining accuracy comparable to the best baselines (such as GraphRAG), demonstrating its practical potential in real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 The figure is a flow chart and effect diagram of the pretreatment stage of the present invention;
[0052] Figure 2 An index effect diagram constructed in the pre-processing stage of the present invention;
[0053] Figure 3 Flowchart of the search phase of the present invention. DETAILED DESCRIPTION
[0054] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0055] The present invention provides a graph retrieval enhancement generation method based on graph and tree structure, the overall process of which includes document preprocessing, summary tree construction, entity graph construction, bidirectional index establishment, query processing, adaptive retrieval and result prompt information construction and other steps.
[0056] Example
[0057] This example uses the book Harry Potter and the Prisoner of Azkaban and the question "Did Slytherin win the Hogwarts House Cup?" as an example. Initially, h=3. Specifically, the following is the answer:
[0058] (1) Document block processing: The entire input long document is divided into tokens. Each document block is 1200 tokens long. There is an overlapping area of 100 tokens between adjacent document blocks to alleviate the semantic truncation problem.
[0059] (2) Constructing a document summary tree: Using a large language model, we perform hierarchical and recursive summarization of document blocks. First, we perform a preliminary summary of every five document blocks as a summary block. Then, we continue to summarize the summary results, i.e., the summary blocks, until a root node containing global semantic information is generated, and finally a document summary tree is formed.
[0060] (3) Construct a document entity graph: Use natural language processing tools (such as SpaCy) to extract entities from each document block and construct an entity subgraph based on entity co-occurrence relationships. Merge the entity graphs of all document blocks into an entity graph for the entire document to capture global entity relationships. For example, Slytherin and Hogwarts may be extracted from the first document block; Hogwarts and the College Cup may be extracted from the second document block. During integration, since both entity graphs have the common entity Hogwarts, Hogwarts is merged together, and then the entity graph corresponding to the entire long document of Slytherin-Hogwarts-College Cup is obtained. Specific examples of document summary trees and document entity graphs are as follows: Figure 1 shown.
[0061] (4) Construct a bidirectional index: Construct bidirectional indexes from entity to document block and document block to entity, i.e. Slytherin corresponds to the first document block, Hogwarts corresponds to the first and second document blocks, and the College Cup corresponds to the second document block; the first document block corresponds to Slytherin and Hogwarts, and the second document block corresponds to Hogwarts and the College Cup; this bidirectional index connects the entity graph and the document summary tree, providing a fast positioning mechanism for the subsequent retrieval stage. The constructed bidirectional index is as follows: Figure 2 shown.
[0062] (5) Extract query entities and match them: After receiving a natural language query, extract the entity information and locate the corresponding nodes in the entity graph. For example, in the example question, extract Slytherin, Hogwarts, and the House Cup, and then find the corresponding three nodes in the entity graph.
[0063] (6) Adaptive retrieval strategy: Dynamically select the retrieval path based on the structural relationship of the query entities in the graph: if the entities are closely related, perform local retrieval through the index; if the entities are sparse or have no structural connection, fall back to global vector retrieval on the summary tree. When the number of candidate document blocks exceeds the limit, use hop count constraints and entity coverage. Specifically, in the example question, the combination of these three entities will be exhausted to obtain the three problem entity pairs (Slytherin, Hogwarts), (Slytherin, College Cup), and (Hogwarts, College Cup). Then, it is found in the graph that Slytherin and Hogwarts are 1-hop neighbors, Slytherin and College Cup are 2-hop neighbors, and Hogwarts and College Cup are 1-hop neighbors, all of which meet the requirement of h<3, so they are all retained. Then the corresponding document blocks are found, and the document block corresponding to Slytherin and Hogwarts is the first document block. At the same time, the document block corresponding to Slytherin and College Cup does not exist, and the document block corresponding to Hogwarts and College Cup is the second document block. The document blocks finally retrieved are the first and second document blocks. After that, these two document blocks will be input into the large language model as supplementary knowledge to enhance the quality of the answers generated by the large language model. The method ends here. This whole process is as follows Figure 3 shown.
[0064] (7) Method evaluation
[0065] Mainstream graph retrieval enhancement generation methods were selected to evaluate both efficiency and effectiveness. Different indicators were used for the evaluation of effectiveness for different data sets. For multiple-choice questions, accuracy (Acc.) was used, and for short-answer questions, the Rouge-L score between the answer and the standard answer was used. In order to evaluate the efficiency, the running time (seconds) was calculated for the preprocessing stage and the retrieval stage respectively. Qwen2.5-7B-Instruct and Llama3.1-8B-Instruct were selected as the base models. Finally, the experimental results obtained by testing on the three data sets are shown in Table 1. The best results are displayed in bold, the suboptimal results are underlined, and if the two methods are tied for first place at this accuracy, both are indicated in bold. The method of the present invention is more efficient than the baseline method in various scenarios, while also retaining a strong retrieval enhancement generation capability.
[0066] Table 1 Comparison of the proposed method with other graph retrieval-based enhanced generation methods on three datasets
[0067]
Claims
1. A graph retrieval enhancement generation method that utilizes both graph and tree structures, characterized in that: The method comprises the following steps: Step 1: Document segmentation: Use the tokenizer corresponding to the large language model to convert the long document serving as the knowledge base into tokens that the large language model can understand. Then, the long text document is divided into multiple document blocks with overlapping areas based on the number of tokens. Specifically, each document block has 1200 tokens, and adjacent document blocks overlap by 100 tokens. Step 2: Use the large language model to recursively summarize the document blocks obtained in step 1. Specifically, five adjacent document blocks are fed into the large language model, and then summarized using the leaf-level summary prompt words to obtain summary information of the document block content, namely the summary block. After all original document blocks are summarized, the resulting summary blocks are fed into the large language model in groups of five, and the summary prompt words for the summary blocks are used to obtain further summary information of the input. This process continues until only one document block that summarizes the entire document remains. A document summary tree is constructed, where the leaf layer is the original document content. The closer to the root node, the more summarized the node content is, and the more global information it contains. Step 3: Use natural language processing tools to extract entities from each document block obtained in Step 1. Two entities appearing in the same sentence are treated as having an undirected edge. An entity graph is extracted from each document block. The entity graphs corresponding to all document blocks are then merged together to obtain an entity graph for the entire long document. Steps 2 and 3 are performed in parallel to save computing time. Step 4: Build a bidirectional index, including an entity-to-document-block mapping index and a document-block-to-entity mapping index. These two indexes allow us to locate the corresponding entity from the document block and the corresponding document block from the entity, thus bridging the document summary tree and the entity graph. Specifically, for each entity, multiple document blocks are extracted for that entity, and for each document block, multiple entities are extracted. Step 5: After receiving the query, use the natural language processing tool Spacy to extract the entities in the query, and use this entity to directly match the nodes in the extracted entity graph; Step 6: Based on the structural relationships between query entities in the entity graph, an adaptive retrieval strategy is used. Depending on the degree of association between entities on the graph, a global search method based on vector retrieval or a method based on the mapping index positioning of entities to document blocks constructed in step 4 is used to search for the corresponding candidate document blocks. Step 7: Sort the candidate document blocks based on entity coverage and frequency of occurrence, and prioritize document blocks that cover the query intent. Step 8: Integrate and format the selected document blocks and entity information, and input them into the large language model in a structured form as prompt information to enhance the large language model's understanding of relevant information and generate answers to the questions; in: The adaptive search strategy described in step 6 includes the following judgment and execution processes: 6.1 Determine whether an entity was extracted from the question. If no entity was found, directly use the pre-trained SentenceBERT model to obtain the text representation. Then, use similarity retrieval to retrieve the corresponding n document blocks or summary blocks from the document summary tree. n is a positive integer, ranging from 5 to 20. Step 6 ends. If an entity was extracted from the question, i.e., the question entity, proceed to 6.
2. 6.2 Consider all unordered question entity pairs among all question entities. Then, check the distance between the entities in the question entity pairs on the constructed entity graph for the entire long document. If the two entities are neighbors within h hops, where h hops means that from one entity, there are at most h edges to reach the other, then the two entities are considered related and continue with step 6.3.1 for these two entities. If the two entities are neighbors within h hops, then the two entities are considered unrelated and removed from the candidate question entity pairs. If no candidate question entity pairs meet the requirements after this step, proceed to step 6.3.
2. 6.3.
1. Index each entity in the entity pair using the entity-to-document block mapping constructed in step 4 to obtain the document block sets corresponding to the two entities. Then, intersect these two document block sets to obtain the document blocks corresponding to both entities as candidate document blocks. If the number of candidate document blocks is not greater than k, step 6 ends. If the number of candidate document blocks is greater than k, proceed to step 6.
4. k is a positive integer that depends on the input length limit of the large language model used and is at most 25. 6.3.2 For candidate question entity pairs that do not meet the requirements, first use the pre-trained SentenceBERT model to obtain text representations. Then, use similarity retrieval to retrieve the 2n document blocks or summary blocks with the highest semantic similarity to the question from the document summary tree. For each retrieved document block, use the frequency of the question entity in the document block as its weight. For each retrieved summary block, use the sum of the weights of its child nodes as its weight. For each summary block node, if its child nodes are also summary blocks, recursively calculate the sum of the weights of their child nodes until a document block is encountered. Then, sort the retrieved document blocks and summary blocks in descending order of weight, select the first n as candidate document blocks, and end in step 6. 6.4 If the number of candidate document blocks is greater than k, further filtering is performed; let h = h - 1, and then check the distance between the entities in the problem entity pair on the constructed entity graph of the entire long document. If the two entities are neighbors within h hops, they are considered related and continue to perform step 6.3.1 on these two entities; if the two entities are neighbors more than h hops, they are considered unrelated and are deleted from the candidate problem entity pairs; if after this step, there are no candidate problem entity pairs that meet the requirements, perform 6.5; if after this step, there are still more than k candidate document blocks, then repeat 6.4; 6.5 If no relevant entities can be retrieved after further filtering in step 6.4, then the candidate document blocks obtained before executing step 6.4 are sorted and filtered as follows: First, use the document block to entity mapping index built in step 4 to obtain the type of question entity contained in each document block as the weight of this document block 1; Then execute 6.3.2 to obtain the ranking based on the frequency of the problem entity in the document block, which is used as the weight 2 of this document block; Finally, all candidate document blocks are sorted, first in descending order of weight 1. If weight 1 is the same, they are sorted in descending order of weight 2. Finally, the first n document blocks are selected as candidate document blocks, and step 6 ends.
2. The graph retrieval enhancement generation method according to claim 1, characterized in that: The leaf layer summary prompt words in step 2 specifically include: 2.1.1 Task Definition: The large language model is required to summarize the input text content and ensure that the summary length does not exceed 1200 tokens; 2.1.2 Output Format: The large language model is required to output the summary directly in text format without outputting other information; The summary words for the summary block described in step 2 specifically include: 2.2.1 Task Definition: The large language model is required to further summarize the input text content, find the relevance between these summaries, grasp the main content, and ensure that the summary length does not exceed 1200 tokens; 2.2.2 Output Format: The large language model is required to output the summary directly in text format without outputting other information.
3. The graph retrieval enhancement generation method according to claim 1, characterized in that Step 8 integrates and formats the selected document blocks and entity information as follows: 8.1: Organize the search results according to the "entity 1-entity 2: document block content" structure, use the extracted entity pairs as topics, and the associated document blocks as supplementary information, and construct prompts in a structured format to input to the large language model; 8.2: To reduce redundant input, if a document block is referenced by multiple entity pairs, these entities will be merged into a unified representation of "entity 1-entity 2-...-entity n" to avoid repeated addition of the same document block; 8.3: In the candidate document block set, identify adjacent continuous document blocks and merge them to reduce cross-block duplication and context fragment overlap, thereby further compressing the prompt word length; 8.4: Calculate entity coverage for all entity pairs, prioritize entity pairs that cover more entities in the query, and sort based on this; 8.5: After the entity pairs are sorted, the corresponding document blocks are arranged in order according to the order of the original document blocks in the document to form the final large language model input prompt information, thereby enhancing the quality of the answers generated by the large language model.