Retrieval enhancement generation method and system based on structured document
By integrating structured information of structured documents in the search enhancement generation technology, the problem that traditional RAG technology is difficult to answer abstract questions when dealing with structured documents is solved, and the answer accuracy and rationality of community structure are improved.
Patent Information
- Application Number
- CN202510126899.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-27
AI Technical Summary
Traditional search-enhanced generation (RAG) technology is difficult to effectively answer abstract questions when processing structured documents, and it is easy to lead to sparse relationships and loss of structural information when building knowledge graphs, resulting in poor answering results.
By integrating the structured information of structured documents into the knowledge graph, increasing the number of relationships and adding structured information, a more reasonable and rich community structure is built and the construction method of the knowledge graph is improved.
It improves the accuracy of the final answer, solves the problems of sparse knowledge graph relationships and loss of structural information, making the community structure more reasonable and information richer.
Smart Images

Figure CN120067340A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of retrieval enhancement, and specifically relates to a retrieval enhancement generation method and system based on structured documents. Background Art
[0002] Currently, there are many documents with obvious chapters, numbers, or other similar structures on the Internet, that is, structured documents. The characteristics of such documents are that the content is organized into ordered parts, such as chapters, sections, subsections, lists, etc. For example, books and long documents with chapter relationships, legal documents with paragraph and sub-paragraph relationships, and standards and technical documents with clause and sub-clause hierarchical relationships.
[0003] As the capabilities of large language models become more and more powerful, when answering questions about such documents, the technology of retrieval augmented generation (RAG) can be used. First, according to a question, relevant text fragments are retrieved from the document, and then these text fragments and the question are input into the large language model to obtain an answer. However, there is a problem with traditional retrieval augmented generation (RAG): it is designed for answers contained within text blocks and is not good at answering summary questions. Therefore, the proposed GraphRAG solves this problem by constructing text data into a knowledge graph, partitioning communities, and then summarizing community summaries to answer summary questions through community summaries.
[0004] However, GraphRAG will have the problem of sparse relationships when constructing a knowledge graph, resulting in some entity relationship information being lost in the constructed communities; secondly, when facing structured documents, the constructed graph will lose the structure information of the document, such as which subheadings are included in the title and which entities are under the same title, resulting in the loss of the structure information of the document in the constructed communities. These will all lead to poor final answer effects. Summary of the Invention
[0005] To solve the above-mentioned problems, the present method proposes a retrieval enhancement generation method and system based on structured documents. By integrating the structured information of structured documents into the knowledge graph, it not only increases the number of relationships to solve the problem of sparse relationships in the knowledge graph, but also adds structured information to solve the problem of missing structured information in the knowledge graph, thereby making the constructed community structure more reasonable and the information more abundant, so the accuracy of the final answer is higher.
[0006] The present invention provides a retrieval enhancement generation method based on structured documents, including the following steps:
[0007] Step 1: Obtain a set of structured documents and extract the structured information of the documents;
[0008] Identify the structured document, extract the text of element nodes and text nodes, and the positions of text nodes and element nodes in the document. The element node refers to the basic unit in the document for organizing and hierarchically classifying specific content, and the text node refers to the specific text content at the lowest content level in the document, without involving the document structure.
[0009] For the extracted element nodes, use a large language model to construct the hierarchical structure of the element nodes as the document structured information HS = (V1, R), where V1 is the set of element nodes and R is the set of adjacency relationships R = {(v j , v j ) | v i ∈V1, v j ∈V1};
[0010] Step 2: Divide text blocks according to the document structure;
[0011] Based on the positions of element nodes and text nodes in the document, obtain the lowest-level element node to which the text node belongs; divide the text of the text node according to the maximum token number of the large language model to obtain text blocks and the lowest-level element node to which the text blocks belong, as text block information;
[0012] Step 3: Extract entities, relationships, and their descriptions from text blocks;
[0013] According to the text blocks and the lowest-level element nodes to which the text blocks belong, construct an entity relationship extraction prompt, use the large language model to identify entities and entity descriptions, as well as relationships and relationship descriptions between entities, and finally obtain entity information and relationship information;
[0014] Step 4: Construct an enhanced knowledge graph
[0015] According to the entity information and relationship information, align the entities and relationships, and then construct a knowledge graph G, G = {(E, V)}, where E is the set of aligned entities E = {(e, a1)}, where e is the entity name and a1 is the entity attribute, including the entity description and the associated text blocks, and V is the set of relationships, V = {(e 1 , a, e 2 ) | e 1 ∈E, e 2 ∈E}, a is the attribute of the relationship edge between entity e 1 and e 2 , including the relationship description.
[0016] Using the document structure information obtained in step 1 and the text block information obtained in step 2, enhance the knowledge graph to obtain an enhanced knowledge graph G′={(E′,V′)}, where E′ = E ∪ EH, EH represents the element nodes, and V′ = V ∪ VH, where VH represents the relationships between the lowest-level element nodes and entities and the hierarchical relationships between element nodes.
[0017] Step 5: Community division and community summary construction;
[0018] Invoke the community discovery algorithm for the enhanced knowledge graph for community division. For each divided community, construct a summary community summary prompt, and concatenate the summary community summary prompt with the entity name e, entity description, and relationship description included in the community, and input them together into the large language model to obtain the community summary.
[0019] Use the embedding model to vectorize the community summary to obtain community summary information, where the community summary information includes the mapping relationship between the community summary and the community summary embedding.
[0020] Step 6: Retrieval and answer generation;
[0021] Obtain the summary-type user question to be answered, where the summary-type user question refers to a question that aims to obtain the core information or brief overview related to a specific topic or requirement, without the need for detailed background information or complex analysis.
[0022] Vectorize the summary-type user question using the embedding model, calculate the similarity with the community summary embedding, and filter out the community summaries with similarity lower than a certain threshold to obtain several community summaries with high similarity;
[0023] Based on the summary-type user question to be answered, construct a prompt for generating intermediate answers and scoring the intermediate answers, and use the large language model to generate an intermediate answer and the score of the intermediate answer for each community summary with high similarity. Sort these intermediate answers in descending order of score, add them to the large language model context until the maximum context window limit is met, and then construct a prompt for summarizing the intermediate answers, and use the large language model to return the answer to the user.
[0024] Preferably, in step 2, the lowest-level element node to which the text block belongs is obtained as follows; according to the belonging relationship between the text block and the text node, and the belonging relationship between the text node and the lowest-level element node, obtain the lowest-level element node to which the text block belongs;
[0025] Preferably, in step 3, the entity information includes: entity name, entity description, and the text block where the entity is located; the relationship information includes: head entity name, tail entity name, and relationship description.
[0026] Preferably, in step 1, the element node is specifically a title, and the text node is specifically a text segment under the lowest-level title.
[0027] Preferably, step 1 specifically includes the following steps:
[0028] Step 1.1 Use a layout analysis model to identify the positions of element nodes and text nodes in the structured document, and use an optical character recognition model to extract the text in the element nodes and text nodes;
[0029] Step 1.2 Construct a prompt for identifying the hierarchical structure of element nodes, arrange the element nodes from top to bottom according to their positions in the document, and input them together to a large language model to obtain the hierarchical structure of the element nodes as the structured information of the document, denoted as HS=(V1,R), where V1 is the set of element nodes, and R is the adjacency relationship set R={(v i ,v j )|v i ∈V1,v j ∈V1}; The structured information of the document stores the mapping relationship between element nodes and their sub-element nodes.
[0030] As a preferred method, the entity relationship alignment method mentioned in step 4 is specifically as follows: According to the entity information and relationship information, align the entities according to their names, align the relationships according to the head and tail entity names, merge the text blocks where the entities appear into a list, and use a large language model to summarize the multiple descriptions of the aligned entities and relationships into one description.
[0031] Preferably, the enhanced knowledge graph in step 4 specifically includes the following steps:
[0032] Step 4.1 Add the element nodes to the knowledge graph G according to the structured information of the document;
[0033] Step 4.2 Establish a relationship between each element node and its parent element node and add it to the knowledge graph G;
[0034] Step 4.3 According to the text block information, find the text blocks contained in the lowest-level element node, find all the entities corresponding to the text blocks from the knowledge graph G to obtain the entities contained in the lowest-level element node, and establish a relationship between the lowest-level element node and all the entities it contains and add it to the knowledge graph G.
[0035] The present invention also provides a retrieval enhanced generation system based on structured documents, including:
[0036] A document structured information extraction module for identifying the structured document, extracting the text of element nodes and text nodes, and the positions of text nodes in the document. The element nodes refer to the basic units used to organize and hierarchically classify specific content in the document, and the text nodes refer to the specific text content at the lowest content level in the document; for the extracted element nodes, use a large language model to construct the hierarchical structure of the element nodes as the document structured information HS = (V1, R), where V1 is the set of element nodes and R is the set of adjacency relations R = {(v i , v j ) | v i ∈ V1, v j ∈ V1};
[0037] A text block division module for: obtaining the lowest-level element node to which the text node belongs according to the positions of the element node and the text node in the document; splitting the text of the text node according to the maximum token number of the large language model to obtain text blocks and the lowest-level element node to which the text block belongs as text block information;
[0038] An entity relationship extraction module for constructing an entity relationship extraction prompt according to the text block and the lowest-level element node to which the text block belongs, using the large language model to identify entities and entity descriptions, as well as relationships and relationship descriptions between entities, and finally obtaining entity information and relationship information;
[0039] An enhanced knowledge graph construction module for: aligning entities and relationships according to entity information and relationship information, and then constructing a knowledge graph G, G = {(E, V)}, where E is the set of aligned entities E = {(e, a1)}, where e is the entity name and a1 is the entity attribute, including the entity description and the associated text block, and V is the set of relationships, V = {(e 1 , a, e 2 ) | e 1 ∈ E, e 2 ∈ E}, a is the attribute of the relationship edge between entity e 1 and e 2 , including the relationship description; using the document structured information obtained in step 1 and the text block information obtained in step 2 to enhance the knowledge graph to obtain an enhanced knowledge graph G' = {(E', V')}, E' = E ∪ EH, where EH represents element nodes, and V' = V ∪ VH, where VH represents the relationship between the lowest-level element nodes and entities and the hierarchical relationship between element nodes;
[0040] The community division and community summary construction module is used to: call the community discovery algorithm to divide the enhanced knowledge graph into communities, construct a summary community summary prompt for each community obtained by the division, and concatenate the summary community summary prompt and the entity name e, entity description, and relationship description contained in the community and input them into the large language model to obtain a community summary; use the embedding model to vectorize the community summary to obtain community summary information, and the community summary information includes a mapping relationship between the community summary and the community summary embedding;
[0041] The retrieval and answer generation module is used to: obtain summary user questions to be answered, wherein the summary user questions refer to questions that are intended to obtain core information or brief overviews related to specific topics or needs, without the need for detailed background information or complex analysis. The summary user questions are vectorized using the embedding model, similarity is calculated with the community summary embedding, and community summaries with similarities below a certain threshold are screened out to obtain several community summaries with high similarity; based on the summary user questions to be answered, a prompt for generating intermediate answers and scoring intermediate answers is constructed, and an intermediate answer and a score for the intermediate answer are generated for each community summary with high similarity using a large language model. These intermediate answers are sorted in descending order of score, added to the large language model context, and then a prompt summarizing the intermediate answers is constructed, and the answer is returned to the user using the large language model.
[0042] Compared with the prior art, the above-mentioned retrieval enhancement generation method based on structured documents has higher accuracy.
[0043] Compared with existing methods, since the structural information of structured documents is integrated into the knowledge graph, it not only increases the number of relationships to solve the problem of sparse relationships in the knowledge graph, but also adds structured information to solve the problem of missing structured information in the knowledge graph, making the constructed community structure more reasonable and the information richer, so the final answer is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the specific implementation or the prior art description. Some specific embodiments of the present invention will be described in detail in an exemplary but not restrictive manner with reference to the drawings. The same reference numerals in the drawings indicate the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:
[0045] Figure 1 It is a flow chart of a retrieval enhancement generation method based on structured documents of the present invention.
[0046] Figure 2 It is a sample diagram of the aviation industry engineering design specification document of the present invention.
[0047] Figure 3 It is a sample diagram of the enhanced knowledge graph constructed according to the aviation industry engineering design specification document of the present invention. Detailed implementation manners
[0048] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0049] In the present invention, embedding represents vectorized embedding; prompt represents a prompt word; Specific embodiments
[0051] In order to make the technical solutions and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.
[0052] A retrieval-enhanced generation method based on structured documents mainly includes the following steps:
[0053] First, execute step S1 to import the aviation industry engineering design specification document set and extract the document structured information.
[0054] The aviation industry engineering design specification document set contains multiple aviation industry engineering design specification documents. The aviation industry engineering design specification documents are in PDF format.
[0055] Generate an image for each page of the aviation industry engineering design specification document, then use the layout analysis model 360LayoutAnalysis to identify the positions of each aviation industry engineering design specification document element node and text node in the document according to the image, and then use the PP-OCRv3 character recognition model to extract the characters in the element node and text node.
[0056] Construct a recognition element node hierarchy prompt, arrange the element nodes from top to bottom according to the position in the document, and input them together to the GPT-4o model to obtain the hierarchy of the element nodes, that is, the document structured information, and the document structured information stores the mapping relationship between the element nodes and their sub-element nodes.
[0057] In this embodiment, the element node refers to the title, and the text node refers to the text fragment under the lowest-level title. The document structured information is a JSON file, the key is the title, the data type is a string, and the value is the sub-directory title of the title, and the data type is an object.
[0058] The prompt for identifying the element node hierarchy is:
[0059] Please organize the following element node information into a hierarchical JSON structure. The key of the JSON is of string type, representing the element node, and the value is of object type, representing the sub-element node.
[0060] # Example
[0061] {Examples}
[0062] # Input
[0063] {input_text}
[0064] Then execute step 2. Based on the positions of the element nodes and text nodes in the document, obtain the lowest-level element node to which the text node belongs; split the text of the text node according to the maximum number of tokens to obtain text chunks; based on the belonging relationship between the text chunks and the text nodes, and the belonging relationship between the text nodes and the lowest-level element nodes, obtain the lowest-level element node to which the text chunks belong; finally, obtain the text chunk information, where the text chunk information includes the text chunk content and the lowest-level element node to which it belongs.
[0065] In this embodiment, the text chunk information is a CSV file, and the fixed length for splitting is set to 512, and the length of the overlapping part is set to 256.
[0066] Then execute step 3. Based on the text chunk information, construct an entity relationship extraction prompt, use the GPT-4o model to identify the entities and their descriptions, as well as the relationships between the entities and the descriptions of the relationships, and finally obtain the entity information and relationship information. The entity information includes: entity name, entity description, and the text chunk where the entity is located; the relationship information includes: head entity name, tail entity name, and relationship description.
[0067] In this embodiment, the entity relationship extraction prompt is:
[0068] - Goal --
[0069] Given a text and a list of entity types, identify all entities of these types in the text and all relationships between the entities.
[0070] - Steps -
[0071] 1. Identify all entities. For each identified entity, extract the following information:
[0072] - Entity name: The name of the entity
[0073] - Entity type: One of the following types [{entity_types}]
[0074] - Entity description: Description related to the entity
[0075] The format of each entity is set as ("entity"{tuple_delimiter}<entity name>{tuple_delimiter}<entity type>{tuple_delimiter}<entity description>).
[0076] 2. From the entities identified in step 1, identify all (head entity, tail entity) pairs that are significantly related.
[0077] For each pair of entities, extract the following information:
[0078] - Head entity: The name of the head entity
[0079] - Tail entity: The name of the tail entity
[0080] - Relationship description: The reason for considering the head entity and the tail entity to be related
[0081] The format of each relationship is set as ("relationship"{tuple_delimiter}<head entity>{tuple_delimiter}<tail entity>{tuple_delimiter}<relationship description>).
[0082] 3. All entities and relationships identified in steps 1 and 2. Use {record_delimiter} as the list separator.
[0083] 4. After completion, output {completion_delimiter}
[0084] #
[0085] - Example -
[0086] #
[0087] {Examples}
[0088] ##
[0089] - Real data -
[0090] #
[0091] Entity type: {entity_types}
[0092] Text: {input_text}
[0093] #
[0094] Answer:
[0095] Then perform step 4 to construct an enhanced knowledge graph.
[0096] First, construct a knowledge graph. According to the entity information and relationship information, align the entities by name, align the relationships by the head and tail entity names, merge the text blocks where the entities are located into a list, and use the GPT-4o model to summarize multiple descriptions of the aligned entities and relationships into one description.
[0097] Use the aligned entities as nodes and the aligned relationships as edges to construct a knowledge graph. Add the description of the entity and the list of text blocks where the entity is located as attributes of the entity, and add the description of the relationship as an attribute of the relationship to the knowledge graph.
[0098] Then enhance the knowledge graph. According to the document structured information, add the element nodes as nodes to the knowledge graph, establish a relationship between each element node and its parent element node, and add it to the knowledge graph; according to the text block information, find the text blocks included under the lowest-level element node, find all the entities corresponding to these text blocks from the knowledge graph to obtain the entities included in the lowest-level element node, and establish a relationship between the lowest-level element node and all the entities it includes, and add it to the knowledge graph.
[0099] Finally, obtain the enhanced knowledge graph and store the enhanced knowledge graph in a graph database at the same time.
[0100] In this embodiment, when merging the text blocks where the entities are located, it is necessary to use set to remove duplicates; the graph database uses the Neo4j graph database.
[0101] The prompt for summarizing the descriptions of entities and relationships is:
[0102] You are a helpful assistant responsible for generating a comprehensive summary of the data provided below.
[0103] Given one or two entities and a list of descriptions, all descriptions correspond to one entity or a group of entities.
[0104] Please summarize all of these into a comprehensive description. Ensure that all information in the descriptions is included.
[0105] If there are contradictions between the provided descriptions, resolve these contradictions and provide a coherent summary.
[0106] Ensure that it is written in the third person and includes entity names for complete context.
[0107] #
[0108] - Data -
[0109] Entity: {entity_name}
[0110] Description list: {description_list}
[0111] #
[0112] Output:
[0113] """
[0114] Then execute step 5, call the community discovery algorithm on the enhanced knowledge graph for community partitioning, construct a summary community abstract prompt, and concatenate the summary community abstract prompt with the entities and entity descriptions, relationships and relationship descriptions included in the community, and input them together into the GPT-4o model to obtain the community abstract.
[0115] Use the embedding model to vectorize the community abstract using the embedding model, and together with the community abstract, form community abstract information, which stores the mapping relationship between the community abstract and the community abstract embedding.
[0116] In this embodiment, the community abstract information is a CSV file; the maximum community size of the Leiden community discovery algorithm is set to 10, and the random seed is set to 0xDEADBEEF; the community abstract vector dimension is 1536;
[0117] The community abstract prompt is:
[0118] # Goal
[0119] Generate a community report based on the given list of entities and their relationships in the community. The community report contains information related to the community and its importance.
[0120] # Community report structure
[0121] The community report should include the following parts:
[0122] - Title: The name of the community, which needs to include its key entities. The title should be short but specific. If possible, include a representative entity in the title.
[0123] - Abstract: A summary reflecting the overall structure of the community, including how the entities are interconnected and important information related to the entities.
[0124] - Details: A list of 5 - 10 key pieces of information about the community. Each key piece of information should have a short summary followed by an explanation of the summary, which is based on the basic rules below and should be comprehensive.
[0125] Return the output as a properly formatted JSON string in the following format:
[0126]
[0127] # Regulations
[0128] Please do not provide irrelevant information without supporting evidence.
[0129] # Examples
[0130] {Examples}
[0131] -----------
[0132] # Real data
[0133] Use the following text as your answer. Do not fabricate anything in your answer.
[0134] Input:
[0135] {input_text}
[0136] The community report should include the following sections:
[0137] - Title: The name of the community, which needs to include its key entities. The title should be short but specific. If possible, include a representative entity in the title.
[0138] - Abstract: An abstract that reflects the overall structure of the community, including how the entities are interconnected and important information related to those entities.
[0139] - Details: A list of 5 - 10 key pieces of information about the community. Each key piece of information should have a short summary followed by an explanation of the summary, which is based on the basic rules below and should be comprehensive.
[0140] Return the output as a properly formatted JSON string in the following format:
[0141]
[0142] # Regulations
[0143] Please do not provide irrelevant information without supporting evidence.
[0144] Output:
[0145] Then, perform step 6 to obtain the summary-type user question to be answered, vectorize the summary-type user question using the embedding model, calculate the similarity with the community summary embedding, and filter out the community summaries with similarity lower than a certain threshold.
[0146] Construct a prompt for answering intermediate questions, use the GPT-4o model to generate an intermediate answer for each community summary with similarity higher than the threshold, and at the same time score the intermediate answers. Sort these intermediate answers in descending order of scores, add them to the context of the GPT-4o model until the maximum context window limit is reached, then construct a summary intermediate answer prompt, input it to the GPT-4o model, and then return the answer to the user.
[0147] In this embodiment, cosine similarity is used to represent the similarity between two vectors, and the similarity threshold is set to 0.77;
[0148] The prompt for answering intermediate questions is:
[0149] ---Role---
[0150] You are an assistant who can answer relevant questions based on the data in the table.
[0151] ---Goal---
[0152] Summarize all relevant information in the input table data to generate an answer, and the answer is a list of key points.
[0153] You should use the data provided in the following table as the main context for generating the answer.
[0154] If you don't know the answer, or the data in the input table does not contain enough information to provide an answer, please say so. Don't make anything up.
[0155] Each key point in the answer should contain the following parts:
[0156] - Description: A comprehensive description of the key point.
[0157] - Importance Score: An integer score between 0 - 100, indicating the importance of the key point when answering the user's question. The score for a "I don't know" type response is 0.
[0158] The response is in JSON format as follows:
[0159] {{
[0160] "Key Point List":
[0161] {{"description":"Description of key point 1","score":"score of key point 1}},
[0162] {{"description":"Description of key point 2","score":the score of key point 2}}, ]
[0164] }}
[0165] The original meaning and modal verbs (such as "maybe") should be retained in the answer.
[0166] Please do not provide information without supporting evidence.
[0167] Example:
[0168] {Examples}
[0169] ---Data Sheet---
[0170] {input_data}
[0171] ---Target---
[0172] Summarize all relevant information from the data in the input table to generate an answer, which is a list of key points. You should use the data provided in the following table as the primary context for generating your answer.
[0173] If you don't know the answer, or the data in the table you entered doesn't contain enough information to provide an answer, say so. Don't make anything up.
[0174] Each key point in your answer should contain the following parts:
[0175] - Description: A comprehensive description of the key points.
[0176] - Importance Score: An integer score between 0-100 indicating how important this key point is in answering the user's question. "I don't know" type responses should have a score of 0.
[0177] The original meaning and modal verbs (such as "maybe") should be retained in the answer.
[0178] Please do not provide information without supporting evidence.
[0179] The response should be in JSON format, like this:
[0180] {{
[0181] "Keypoint list":[
[0182] {{"description":"Description of key point 1","score":"score of key point 1}},
[0183] {{"Description": "Description of Key Point 2", "Score": Score of Key Point 2}},
[0185] }}
[0186] The summary of the intermediate answer prompt is:
[0187] ---Role---
[0188] You are an assistant who can answer questions about a dataset by synthesizing viewpoints from multiple different perspectives.
[0189] ---Goal---
[0190] Generate an answer that can solve the user's problem, summarizing answers from multiple different perspectives, each focusing on a different part of the dataset. The answers from different perspectives provided below are sorted in descending order of importance.
[0191] If you don't know the answer, or the answer provided doesn't contain enough information to answer, say so. Don't make anything up.
[0192] The final answer should remove all irrelevant information from the answers provided and combine the processed information into a comprehensive answer that provides explanations for all key points.
[0193] The answer should preserve the original meaning and modal verbs (such as "might").
[0194] Please don't provide information without supporting evidence.
[0195] ---Length and Format of the Target Answer---
[0196] {response_type}
[0197] ---Data Report---
[0198] {input_data}
[0199] ---Goal---
[0200] Generate an answer that can solve the user's problem, summarizing answers from multiple different perspectives, each focusing on a different part of the dataset. The answers from different perspectives provided below are sorted in descending order of importance.
[0201] If you don't know the answer, or the answer provided doesn't contain enough information to answer, say so. Don't make anything up.
[0202] The final answer should remove all irrelevant information from the answers provided and combine the processed information into a comprehensive answer that provides explanations for all key points.
[0203] The answer should preserve the original meaning and modal verbs (such as "may").
[0204] Please do not provide information without supporting evidence.
[0205] Corresponding to the embodiment of the foregoing retrieval enhanced generation method based on structured documents, the present invention further provides a retrieval enhanced generation system based on structured documents. It includes a document structured information extraction module, a text block division module, an entity relationship extraction module, an enhanced knowledge graph construction module, a community division and community summary construction module, and a retrieval and answer generation module.
[0206] The document structured information extraction module is used to identify the structured document, extract the text of the element nodes and text nodes, and the positions of the text nodes in the document. The element nodes refer to the basic units in the document for organizing and hierarchically classifying specific content, and the text nodes refer to the specific text content at the lowest content level in the document; for the extracted element nodes, use a large language model to construct the hierarchical structure of the element nodes as the document structured information HS=(V1,R), where V1 is the set of element nodes, and R is the set of adjacency relationships R={(v i ,v j )|v i ∈V1,v j ∈V1};
[0207] The text block division module is used to: obtain the lowest-level element node to which the text node belongs according to the positions of the element nodes and text nodes in the document; divide the text of the text node according to the maximum number of tokens of the large language model to obtain text blocks and the lowest-level element node to which the text block belongs as text block information;
[0208] The entity relationship extraction module is used to construct an entity relationship extraction prompt according to the text block and the lowest-level element node to which the text block belongs, use the large language model to identify entities and entity descriptions, as well as relationships and relationship descriptions between entities, and finally obtain entity information and relationship information;
[0209] The enhanced knowledge graph construction module is used to: align entities and relationships according to entity information and relationship information, and then construct a knowledge graph G, G={(E,V)}, where E is the set of aligned entities E={(e,a1)}, where e is the entity name, a1 is the entity attribute, including the description of the entity and the associated text block, and V is the set of relationships, V={(e 1 ,a,e 2 )|e 1 ∈E,e 2 ∈E}, a is the entity e 1and e 2 The attributes of the relationship edges between them, including relationship descriptions; using the document structure information obtained in step 1 and the text block information obtained in step 2 to enhance the knowledge graph, obtaining an enhanced knowledge graph G′={(E′,V′)}, where E′ = E ∪ EH, EH represents the element nodes, V′ = V ∪ VH, and VH represents the relationship between the lowest-level element nodes and entities and the hierarchical relationship between element nodes;
[0210] Community division and community summary construction module, which is used for: calling a community discovery algorithm to divide the enhanced knowledge graph into communities. For each divided community, construct a summary community abstract prompt, and input the summary community abstract prompt, the entity name e included in this community, the description of the entity, and the relationship description together into a large language model to obtain a community summary; using an embedding model to vectorize the community summary to obtain community summary information, and the community summary information includes the mapping relationship between the community summary and the community summary embedding;
[0211] Retrieval and answer generation module, which is used for: obtaining a summary-type user question to be answered, where the summary-type user question refers to a question aimed at obtaining core information or a brief overview related to a specific topic or requirement, without the need for detailed background information or complex analysis. Vectorize the summary-type user question using the embedding model, calculate the similarity with the community summary embedding, filter out community summaries with similarity lower than a certain threshold, and obtain several community summaries with high similarity; based on the summary-type user question to be answered, construct a prompt for generating an intermediate answer and scoring the intermediate answer, and use a large language model to generate an intermediate answer and an intermediate answer score for each community summary with high similarity. Sort these intermediate answers in descending order of score, add them to the context of the large language model, and then construct a prompt for summarizing the intermediate answers, and use the large language model to return the answer to the user.
[0212] As mentioned above, this is only a part of the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those familiar with the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A retrieval enhancement generation method based on structured documents, characterized in that: The following steps are involved: Step 1: Obtain a structured document, identify the structured document, extract element nodes, text nodes, and the positions of text nodes and element nodes in the document. The element node refers to the basic unit used to organize and hierarchize specific content in the document, and the text node refers to the specific text content at the lowest content level in the document; For the extracted element nodes, a large language model is used to construct a hierarchical structure of the element nodes as document structured information; Step 2: According to the positions of the element node and the text node in the document, the lowest-level element node to which the text node belongs is obtained; the text of the text node is segmented according to the maximum number of tokens of the large language model to obtain the text block and the lowest-level element node to which the text block belongs as the text block information; Step 3: construct an entity relationship extraction prompt according to the text block and the lowest-level element node to which the text block belongs, and use the large language model to identify entities and entity descriptions, as well as relationships between entities and relationship descriptions, to finally obtain entity information and relationship information; Step 4: Based on the entity information and relationship information, align the entities and relationships, and then construct the knowledge graph G, G = {(E, V)}, where E is the aligned entity set E = {(e, a1)}, where e is the entity name, a1 is the entity attribute, including the entity description and the associated text block, V is the relationship set, V = {(e1, a, e2) | e1∈E, e2∈E}, a is the attribute of the relationship edge between entities e1 and e2, including the relationship description; Using the document structure information obtained in step 1 and the text block information obtained in step 2, the knowledge graph is enhanced to obtain the enhanced knowledge graph G. ′ ={(E ′ ,V ′ )},E ′ =E∪EH, EH represents the element node, V ′ =V∪VH, where VH represents the relationship between the lowest level element node and the entity, and the hierarchical relationship between element nodes; Step 5: Call the community discovery algorithm to divide the enhanced knowledge graph into communities. For each community obtained by the division, a summary community summary prompt is constructed. The summary community summary prompt and the entity name e, entity description, and relationship description contained in the community are concatenated and input into the large language model to obtain a community summary. The community summary is vectorized using the embedding model to obtain community summary information, which includes a mapping relationship between the community summary and the community summary embedding. Step 6: Get the summary user questions to be answered. The summary-type user question is vectorized using the embedding model, and similarity is calculated with the community summary embedding, and community summaries with similarities below a certain threshold are screened out to obtain several community summaries with high similarity; based on the summary-type user question to be answered, a prompt for generating an intermediate answer and scoring the intermediate answer is constructed, and an intermediate answer and a score for the intermediate answer are generated for each community summary with high similarity using a large language model; these intermediate answers are sorted in descending order according to the score, added to the large language model context, and then a prompt summarizing the intermediate answers is constructed, and the answer is returned to the user using the large language model.
2. A retrieval enhancement generation method based on structured documents as claimed in claim 1, characterized in that: The method for constructing the enhanced knowledge graph described in step 4 specifically includes the following steps: Step 4.1: adding element nodes as nodes to the knowledge graph G according to the document structured information; Step 4.2 establishes a relationship between each element node and its parent element node, and adds them to the knowledge graph G; Step 4.3: According to the text block information, find the text block contained in the lowest-level element node, find all entities corresponding to the text block from the knowledge graph G, obtain the entities contained in the lowest-level element node, establish a relationship between the lowest-level element node and all entities contained therein, and add them to the knowledge graph G; The entity and relationship alignment described in step 4 specifically includes the following steps: based on the entity information and relationship information, entities are aligned according to names, relationships are aligned according to head and tail entity names, text blocks where entities appear are merged into a list, and multiple descriptions of aligned entities and relationships are summarized into one description using a large language model.
3. A retrieval enhancement generation method based on structured documents as claimed in claim 1, characterized in that: In step 2, the lowest level element node to which the text block belongs is obtained in the following manner: the lowest level element node to which the text block belongs is obtained according to the relationship between the text block and the text node, and the relationship between the text node and the lowest level element node.
4. A retrieval enhancement generation method based on structured documents as claimed in claim 1, characterized in that: In step 3, the entity information includes: entity name, entity description, and text block where the entity is located; the relationship information includes: head entity name, tail entity name, and relationship description.
5. A retrieval enhancement generation method based on structured documents as claimed in claim 1, characterized in that: In step 1, the element node is specifically a title, and the text node is specifically a text segment under the lowest level title.
6. A retrieval enhancement generation method based on structured documents as claimed in claim 1, characterized in that: Step 1 specifically includes the following steps: Step 1.1 uses a layout analysis model to identify the positions of element nodes and text nodes in the structured document, and uses a text recognition model to extract text in the element nodes and text nodes; Step 1.2 constructs a prompt to identify the hierarchical structure of element nodes, arranges the element nodes from top to bottom according to their positions in the document, and inputs them into the large language model together to obtain the hierarchical structure of the element nodes as the document structured information, denoted as HS = (V1, R), where V1 is the element node set, R is the adjacency relationship set R = {(v i ,v j )|v i ∈V1,v j ∈V1}; the document structured information stores the mapping relationship between element nodes and their sub-element nodes.
7. A retrieval enhancement generation system based on structured documents, characterized in that: include: The document structured information extraction module is used to identify the structured document, extract the element nodes, the text of the text nodes, and the position of the text nodes in the document. The element nodes refer to the basic units used to organize and hierarchize specific content in the document, and the text nodes refer to the specific text content at the lowest content level in the document. For the extracted element nodes, a hierarchical structure of the element nodes is constructed using a large language model as the document structured information HS = (V1, R), where V1 is the element node set, and R is the adjacency relationship set R = {(v i ,v j )|v i ∈V1,v j ∈V1}; The text block segmentation module is used to: obtain the lowest-level element node to which the text node belongs according to the positions of the element node and the text node in the document; segment the text of the text node according to the maximum number of tokens of the large language model to obtain the text block and the lowest-level element node to which the text block belongs as the text block information; An entity relationship extraction module, used to construct an entity relationship extraction prompt according to the text block and the lowest-level element node to which the text block belongs, and use the large language model to identify entities and entity descriptions, as well as relationships between entities and relationship descriptions, and finally obtain entity information and relationship information; The enhanced knowledge graph construction module is used to: align entities and relationships according to entity information and relationship information, and then construct a knowledge graph G, G = {(E, V)}, where E is the aligned entity set E = {(e, a1)}, where e is the entity name, a1 is the entity attribute, including the entity description and the associated text block, V is the relationship set, V = {(e1, a, e2) | e1∈E, e2∈E}, a is the attribute of the relationship edge between entities e1 and e2, including the relationship description; using the document structured information obtained in step 1 and the text block information obtained in step 2, the knowledge graph is enhanced to obtain the enhanced knowledge graph G ′ ={(E ′ ,V ′ )},E ′ =E∪EH, EH represents the element node, V ′ =V∪VH, where VH represents the relationship between the lowest level element node and the entity, and the hierarchical relationship between element nodes; The community division and community summary construction module is used to: call the community discovery algorithm to divide the enhanced knowledge graph into communities, construct a summary community summary prompt for each community obtained by the division, and concatenate the summary community summary prompt and the entity name e, entity description, and relationship description contained in the community and input them into the large language model to obtain a community summary; use the embedding model to vectorize the community summary to obtain community summary information, and the community summary information includes a mapping relationship between the community summary and the community summary embedding; The retrieval and answer generation module is used to: obtain summary user questions to be answered, vectorize the summary user questions using the embedding model, calculate similarity with the community summary embedding, filter out community summaries with similarity below a certain threshold, and obtain several community summaries with high similarity; based on the summary user questions to be answered, construct a prompt for generating intermediate answers and scoring the intermediate answers, and use the large language model to generate an intermediate answer and a score for the intermediate answer for each community summary with high similarity; sort the intermediate answers in descending order according to the scores, add them to the large language model context, and then construct a prompt summarizing the intermediate answers, and use the large language model to return the answers to the user.
Citation Information
Patent Citations
Construction method of RAG system based on Graph
CN118503407A
RAG question and answer method and system based on knowledge graph and medium
CN118673126A
Dynamic correlation enhancement retrieval generation system and method driven by intelligent knowledge graph
CN118839021A
Domain intelligent question-answering system and method based on knowledge graph library and text vector library
CN119128095A
Graph-based information question and answer method and device, storage medium and electronic device
CN119250202A
Cited By
Knowledge base construction method and device, equipment and medium
CN120297399A
DOM-based enterprise knowledge network construction method and system
CN120596685A
Question and answer method and device based on large and small model fusion and retrieval enhancement generation
CN120849564A
Question and answer method and device based on size model fusion and retrieval enhancement generation
CN120849564B
Ultra-large file deep analysis method and system based on dynamic segmentation and knowledge graph
CN120850991A