Multi-document question answering method and system based on long-term memory knowledge graph enhancement
Patent Information
- Application Number
- CN202610893310.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-06-22
AI Technical Summary
[0005]为了解决现有方案中多文档问答缺乏长期记忆机制、上下文连贯性不足、知识复用率低的问题,本发明提供了一种基于长期记忆知识图谱增强的多文档问答方法及系统,通过构建文档解析、记忆知识图谱构建、知识嵌入、问答生成及记忆完善的一体化流程,实现了文档内容、会话信息及任务经验的持续写入、管理与复用,增强了跨文档、跨轮次、跨会话的上下文连贯性,提升了问答结果一致性与准确性,减少了重复检索开销,能够有效支撑长周期复杂问答场景,提升了系统稳定性与用户交互体验
本发明通过构建文档解析、记忆知识图谱构建、知识嵌入、问答生成及记忆完善的一体化流程,有效解决了现有多文档问答缺乏长期记忆机制、上下文连贯性不足、知识复用率低的问题;具体的,先解析文档生成结构化内容单元,再基于会话构建记忆知识图谱,实现实体与关系的结构化沉淀;通过知识嵌入将文档与记忆信息向量化存储,为高效语义检索提供支撑;经过问题重写、图谱与文档联合检索,融合多源证据生成精准答案;最后通过价值评估筛选有效记忆并完善长期知识图谱,形成可持续迭代的知识体系;本发明实现了文档内容、会话信息及任务经验的持续写入、管理与复用,增强了跨文档、跨轮次、跨会话的上下文连贯性,提升了问答结果一致性与准确性,减少了重复检索开销,能够有效支撑长周期复杂问答场景,提升了系统稳定性与用户交互体验。
Smart Images

Figure CN122432304B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document question answering technology, specifically to a multi-document question answering method and system based on long-term memory knowledge graph enhancement. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the continuous development of artificial intelligence and large language model technology, multi-document question answering has been widely used in intelligent retrieval, knowledge services, and intelligent office scenarios. It can extract and integrate information from multiple documents and generate answers, adapting to complex interaction needs across multiple turns, conversations, and time spans. Current technologies mostly adopt text segmentation, vector retrieval, and one-time generation methods. Although research on long-term memory has been explored, it generally remains at the level of simple vector storage or basic graphs, and a complete knowledge accumulation, management, and retrieval system has not yet been formed. The overall technology still needs further improvement.
[0004] Existing multi-document question answering systems generally lack an integrated long-term memory mechanism for writing, management, and reading. They lack systematic accumulation of document content, historical questions and answers, entity relationships, and task experience, treating historical data merely as a passive retrieval library, failing to transform it into reusable long-term knowledge. Retrieval relies on fixed chunks, resulting in fragmented information and making it difficult to establish deep semantic connections across documents, rounds, and sessions, severely lacking contextual coherence. At the same time, the lack of memory update, conflict resolution, and value filtering mechanisms makes it impossible to dynamically adapt to the evolution of user intent, ultimately leading to poor question-and-answer consistency, low information reuse rate, and weak long-cycle interaction capabilities, making it difficult to meet the needs of deep semantic understanding and continuous knowledge reuse for complex questions and answers. Summary of the Invention
[0005] To address the shortcomings of existing multi-document question answering solutions, such as the lack of long-term memory mechanisms, insufficient contextual coherence, and low knowledge reuse rates, this invention provides a multi-document question answering method and system based on long-term memory knowledge graph enhancement. By constructing an integrated process of document parsing, memory knowledge graph construction, knowledge embedding, question and answer generation, and memory improvement, it achieves continuous writing, management, and reuse of document content, conversation information, and task experience. This enhances contextual coherence across documents, rounds, and conversations, improves the consistency and accuracy of question-and-answer results, reduces redundant retrieval overhead, effectively supports long-cycle complex question-and-answer scenarios, and improves system stability and user interaction experience.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a multi-document question answering method based on long-term memory knowledge graph enhancement.
[0007] A multi-document question answering method based on long-term memory knowledge graph enhancement includes the following process: It receives multiple types of input documents to be processed and performs document parsing operations, identifies the internal heading hierarchy of the document and decomposes the document content according to the heading boundaries, and generates various content units with complete contextual attributes; For each round of conversation content, entity extraction and semantic relationship generation operations are performed to construct a memory knowledge graph specific to the corresponding conversation; Knowledge embedding processing is performed on the content units and the entities and relations contained in the memory knowledge graph, respectively, to generate semantic vectors of a unified dimension and write them into the vector database for storage; Perform semantic completion and question rewriting operations on the original question input by the user to generate a more explicit question; Based on the explicit question retrieval memory knowledge graph and document knowledge vector, the final question answering result is generated by integrating multi-source evidence. Candidate memories are generated based on information from the question-and-answer process, the value of candidate memories is evaluated, and the long-term memory knowledge graph is improved.
[0008] In one implementation of the first aspect of the present invention, multiple types of input documents to be processed are received and a document parsing operation is performed to identify the internal title hierarchy structure of the document and decompose the document content according to the title boundaries, generating various content units carrying complete contextual attributes, including: It identifies headings, body paragraphs, tables, and charts in the input document and constructs the hierarchical relationships between headings within the document. Based on the heading hierarchy boundaries, it breaks down the overall content of the input document and generates structured content units that include text content, content type, the path of the heading, the page number, adjacent identifiers, and document version information.
[0009] In one implementation of the first aspect of the present invention, entity extraction and semantic relation generation operations are performed for each round of conversation content to construct a memory knowledge graph specific to the corresponding conversation, including: Scan the session content and identify key entities with memory value, and generate an entity set containing entity name, entity type, embedding vector, timestamp, source identifier and confidence information; Calculate the semantic similarity between different entities in the entity set, and perform normalization merging on semantically consistent entities; Analyze the semantic relationships between entities in the entity set and generate relation triples to construct a memory knowledge graph containing entity nodes and relation edges.
[0010] As a further limitation of the first aspect of the present invention, the semantic similarity between different entities in the entity set is calculated, and normalization merging processing is performed on semantically consistent entities, including: Extract the embedding vectors corresponding to the two entities to be compared, and calculate the cosine similarity between the two embedding vectors; When the cosine similarity reaches the preset similarity threshold, the two entities are determined to be the same semantic entity and a node merging operation is performed. When the cosine similarity does not reach the preset similarity threshold, an independent node is created for the new entity and added to the entity set.
[0011] In one implementation of the first aspect of the present invention, knowledge embedding processing is performed on the content unit and the entities and relations contained in the memory knowledge graph, respectively, to generate a semantic vector of a unified dimension and write it into a vector database for storage, including: Vectorization processing is performed on text blocks, table descriptions, and chart descriptions in content units to generate document knowledge vectors; The entities in the memory knowledge graph are directly vectorized, and the relation triples in the memory knowledge graph are transcribed into natural language and then vectorized to generate memory knowledge vectors. Document knowledge vectors and memory knowledge vectors are written into a unified vector database and associated with corresponding metadata information.
[0012] In one implementation of the first aspect of the present invention, semantic completion and question rewriting operations are performed on the original question input by the user to generate a defined question, including: Obtain the original question input by the user, and retrieve recent historical dialogue information and dialogue summary content from the conversation content; Complete the omitted information, referential information, and context-dependent information in the original problem to generate a semantically clear and intentionally explicit problem.
[0013] As a further limitation of the first aspect of the present invention, based on the explicit question retrieval memory knowledge graph and document knowledge vector, multi-source evidence is fused to generate the final question-answering result, including: Extract keywords from the explicit question, match them with entity nodes in the memory knowledge graph, and obtain the corresponding first-order subgraph; After transcribing the relation triples in the first-order subgraph into natural language, semantic matching is performed with the explicitation problem, and the triples with the highest matching degree are selected as graph-side evidence. Match the document title node corresponding to the document knowledge vector and obtain the associated text, tables and charts as document-side evidence; By integrating graph-based evidence and document-based evidence, the final question-and-answer results are generated by inputting the language model.
[0014] In one implementation of the first aspect of the present invention, generating candidate memories based on question-and-answer process information includes: Analyze the retrieval path, hit document structure, graph subgraph, and user feedback information during the question-and-answer result generation process; extract event-type, semantic-type, and program-type candidate memories, and generate a candidate memory set containing candidate type, candidate content, source information, timestamp, and confidence information; Evaluating the value of candidate memories includes: The value score is calculated by weighting candidate memories in the candidate memory set based on five dimensions: persistence, reusability, task relevance, evidence strength, and conflict risk. When the value score reaches the preset value threshold, the candidate memory is determined to be a valid memory and included in the writing scope. When the value score does not reach the preset value threshold, the candidate memory is temporarily stored in a short-term cache.
[0015] In one implementation of the first aspect of the present invention, improving the long-term memory knowledge graph includes: Based on the type of valid candidate memory, corresponding nodes are generated: event-type candidate memories generate event nodes, semantic-type candidate memories generate semantic nodes, and program-type candidate memories generate program nodes. Establish connections between new nodes and existing entity nodes, event nodes, and program nodes; update the node set and edge set of the long-term memory knowledge graph to complete the graph improvement operation.
[0016] Secondly, the present invention provides a multi-document question-answering system based on long-term memory knowledge graph enhancement.
[0017] A multi-document question answering system based on long-term memory knowledge graph enhancement includes: The document parsing unit is configured to: receive multiple types of input documents to be processed and perform document parsing operations, identify the internal heading hierarchy structure of the document and decompose the document content according to the heading boundaries, and generate various content units carrying complete contextual attributes; The graph construction unit is configured to perform entity extraction and semantic relationship generation operations for each round of conversation content, and construct a memory knowledge graph specific to the corresponding conversation. The knowledge embedding unit is configured to perform knowledge embedding processing on the content unit and the entities and relations contained in the memory knowledge graph, respectively, generate a semantic vector of a unified dimension and write it into the vector database for storage; The question rewriting unit is configured to perform semantic completion and question rewriting operations on the original question input by the user to generate a clear question; The question-answering generation unit is configured to generate the final question-answering result by fusing multi-source evidence based on a defined question retrieval memory knowledge graph and document knowledge vectors. The memory enhancement unit is configured to: generate candidate memories based on information from the question-and-answer process, evaluate the value of candidate memories, and improve the long-term memory knowledge graph.
[0018] Compared with the prior art, the beneficial effects of the present invention are: This invention effectively addresses the problems of existing multi-document question-and-answer systems, such as lack of long-term memory mechanisms, insufficient contextual coherence, and low knowledge reuse rates, by constructing an integrated process for document parsing, memory knowledge graph construction, knowledge embedding, question-and-answer generation, and memory improvement. Specifically, it first parses documents to generate structured content units, then constructs a memory knowledge graph based on the conversation to achieve structured accumulation of entities and relationships. Through knowledge embedding, document and memory information are vectorized and stored, providing support for efficient semantic retrieval. After question rewriting, graph and document joint retrieval, and the fusion of multi-source evidence, accurate answers are generated. Finally, through value assessment, effective memories are selected and the long-term knowledge graph is improved, forming a sustainable and iterative knowledge system. This invention enables the continuous writing, management, and reuse of document content, conversation information, and task experience, enhances contextual coherence across documents, rounds, and conversations, improves the consistency and accuracy of question-and-answer results, reduces redundant retrieval overhead, effectively supports long-cycle complex question-and-answer scenarios, and improves system stability and user interaction experience.
[0019] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0020] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0021] Figure 1 A flowchart illustrating a multi-document question-answering method based on long-term memory knowledge graph enhancement, provided as an exemplary embodiment of the present invention; Figure 2 A schematic diagram of a document parsing process provided as an exemplary embodiment of the present invention; Figure 3 A schematic diagram of the memory knowledge graph construction process provided as an exemplary embodiment of the present invention; Figure 4 A schematic diagram of a knowledge embedding process is provided as an exemplary embodiment of the present invention; Figure 5 A schematic diagram of a knowledge question-and-answer process is provided as an exemplary embodiment of the present invention; Figure 6 A schematic diagram of the process for improving a memory knowledge graph is provided as an exemplary embodiment of the present invention; Figure 7 This is a schematic diagram of a multi-document question-answering system based on long-term memory knowledge graph enhancement, provided as an exemplary embodiment of the present invention. Detailed Implementation
[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0023] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0024] Existing multi-document question-answering schemes still have the following shortcomings: Most systems simply treat historical documents or interaction records as external retrieval libraries, lacking a complete "write-manage-read" memory mechanism. Secondly, existing graph memory or vector memory schemes often only emphasize retrieval capabilities, but do not adequately consider memory promotion, conflict resolution, forgetting, and rollback. In addition, in document scenarios, systems often cannot distinguish between different types of information such as "original materials uploaded by users," "long-term knowledge summarized from documents," and "task experience formed during historical question-answering processes." Some researchers have proposed a three-stage memory approach: transient, working, and durable. Others have further used typified memory such as Core, Episodic, Semantic, and Procedural. This indicates that current cutting-edge research no longer views memory merely as a vector library, but rather as a system with a lifecycle and governance structure.
[0025] Given the problems with existing solutions, such as Figure 1 As shown, this invention proposes a multi-document question answering method based on long-term memory knowledge graph enhancement. The technical terms and related concepts involved in this processing scheme are briefly introduced below: Long-term memory refers to the memory mechanism by which a system continuously writes, manages, and retrieves document content, historical question-and-answer information, entity relationships, and task experience during multi-round, multi-session, and time-spanning interactions, in order to support knowledge reuse and context preservation in subsequent tasks. Multi-document question answering refers to a task format that retrieves, integrates, and infers relevant information from multiple documents in response to a user's question, and generates an answer. This task typically involves cross-document information association, cross-round contextual understanding, and multi-source evidence fusion.
[0026] The multi-document question-answering method of the present invention, such as Figure 1 As shown, the specific process includes the following: S101: Document parsing.
[0027] like Figure 2As shown, the system first inputs a document, then extracts the heading structure, further breaks down the document based on the headings, and then extracts text blocks, tables, and chart descriptions. Contextual attributes are then added, and finally, structured document content is output. Input documents can be PDF, Word, HTML, Markdown, web page text, or other parsable formats. The system first selects the corresponding parser based on the document format, identifies the headings, body text, tables, charts, and their adjacent descriptive text, and recovers the hierarchical heading relationships of the document. Let the set of extracted heading nodes in the document be: (1); in, These represent the first heading node, the second heading node, ..., the ... Each title node, any title node It should include at least the heading name, heading level, parent node identifier, and its position index in the document. The system constructs a heading tree structure based on the hierarchical relationship between headings and determines the logical organization of the document accordingly.
[0028] After obtaining the heading structure, the system breaks down the document according to the heading boundaries, forming a set of content units. (2); in, These represent the first content unit, the second content unit, ..., the ... Each content unit This can be a body paragraph, table, chart description, or other text segment related to the title. For any content unit... The system adds contextual attributes to it and represents them as follows: (3); in, Represents the content text. Indicates the content type. Indicates the path to the title. Indicates page number or position information. and These represent the identifiers of adjacent content blocks, respectively. This indicates the document version number. For tables, the system extracts the table header fields, table title, table body content, and footnotes, and preferably transcribes them into structured text. For charts, the system extracts the chart title, caption, and surrounding explanatory text, and uses a large language model to generate chart description text to improve the searchability of chart information. Thus, the system outputs document parsing results with title structure constraints and contextual attributes, providing a foundation for subsequent knowledge embedding and question-answering retrieval.
[0029] S102: Construction of a memory knowledge graph.
[0030] like Figure 3 As shown, each round of conversation input is processed first, and entity extraction is performed using a large language model. Then, entity types, embedding vectors, and timestamps are labeled. After entity normalization, relationships in the form of triples are generated, a memory subgraph is constructed, and finally, entity node pairing relationship triples are output. For each round of conversation, the system constructs a corresponding memory knowledge graph, denoted as the first... The content of the round-robin session is recorded as follows: (4); in, This indicates a user input question in this round. This indicates the system's response. This indicates user feedback or correction information. This represents the historical context and hit document content related to this round of conversation. The system first performs entity extraction on the conversation content. Specifically, it uses a large language model to scan the dialogue text and identify key entities, including people, places, items, preferences, events, attributes, document objects, method modules, and other concepts with memory value. Let the first... The entity set extracted by the round-robin session is: (5); in, They represent the first entity, the second entity, ..., the ... An entity, any entity It is represented as: (6); in, For entity name, For entity type, For entity embedding vectors, For timestamp metadata, For source identification, To extract confidence scores and avoid duplicate inclusion of synonymous or near-synonymous entities in the graph, the system calculates the semantic similarity between new entities and existing entity nodes: (7); in, Representing the first The entity and the first An entity, when If the two entities belong to the same semantic entity, the system will determine that they belong to the same semantic entity and perform entity normalization and node merging; otherwise, it will create a new entity node for the new entity.
[0031] After entity extraction, the system proceeds to the relation generation stage. The system uses a large language model to analyze the semantic relationships between entities and automatically generates relation triples. Let the relation set be: (8); in, For the source entity, For relational types, The target entity is defined as follows. Relationship types include, but are not limited to, "preference," "prohibit," "contain," "belong to," "depend on," "follow," "modify to," and "apply to." For example, the system can generate relationships such as ("user," "dietary preference," "vegetarian") and ("user," "prohibited from consumption," "dairy products"). Furthermore, the system constructs a memory subgraph corresponding to this round of sessions. (9); in, For a set of entity nodes, Let be the set of relation edges, and and Each entry includes a timestamp, source round, and confidence level. The output is structured knowledge of entity nodes and relation triples for that round of the session, which is then fed into the subsequent knowledge graph update stage.
[0032] S103: Knowledge Embedding.
[0033] After completing steps S101 and S102, the text elements obtained in the document parsing stage, as well as the entities, relations, and subgraph descriptions in the memory layer knowledge graph, are embedded to facilitate subsequent knowledge retrieval based on semantic similarity. Specifically, such as... Figure 4 As shown, the system first performs document element embedding, then graph element embedding, followed by vector writing, and finally outputs a unified vector representation space. The system vectorizes title nodes, body text blocks, table descriptions, chart descriptions, entity nodes, and relation triples. Let arbitrary knowledge units... The embedding is represented as: (10); in, Represents an embedded model. This is the corresponding vector representation.
[0034] For relation triples The system preferentially first transcribes the text into a natural language description, such as "the user's research direction is multi-document question answering," and then performs vectorization processing to improve its matching ability in natural language question answering scenarios. For the text elements in step 1, the system constructs a document knowledge vector set. (11); in, For embedding vectors in titles, text blocks, table descriptions, or chart descriptions. They represent the 1st embedding vector, the 2nd embedding vector, ..., the mth embedding vector, respectively.
[0035] For the session memory entities and relationships in step S102, the system constructs a set of memory knowledge vectors: (12); in, Describe the corresponding embedding vector for an entity node or triple. They represent the 1st embedding vector, the 2nd embedding vector, ..., the pth embedding vector, respectively.
[0036] The system writes the aforementioned vectors into a vector database and records the corresponding metadata, including object type, title path, source round, entity identifier, page number, and document version. During subsequent retrieval, the system performs matching using a semantic similarity function.
[0037] S104: Knowledge Q&A.
[0038] When a user raises a new question, the system first combines recent dialogue information and conversation summaries to semantically complete and rewrite the user's question to clarify the user's true needs. For example... Figure 5 As shown, the process first integrates recent dialogues and conversation summaries to rewrite the questions, resulting in rewritten questions. Then, the processing flow of the memory graph branch and the document branch is started simultaneously. The memory graph branch performs keyword extraction, initial node matching, first-order subgraph acquisition, triple reordering, and selection of Top-K triples in sequence. The document branch completes the steps of matching document titles, acquiring the corresponding text blocks, acquiring tables, and acquiring chart descriptions in sequence. After the two branches are processed, multi-source context fusion is performed, and the fused content is then input into the large language model to generate answers, finally outputting the question-and-answer results.
[0039] More specifically, let the user's original question be... ,recent The set of dialogue messages is The conversation summary is The system is based on the user's original question. Recent conversation information and conversation summary The execution problem is rewritten to obtain a more explicit problem: (13); in, This represents the problem rewriting function. Through this step, the system can complete the omitted information, referential information, and contextual dependencies in the user's problem, making the rewritten problem... To more accurately express the current user's intent and serve as a unified input for subsequent searches.
[0040] The problem after the rewrite Then, the system first performs keyword extraction to obtain a keyword set. Subsequently, the system uses keywords to match relevant entity nodes in the memory knowledge graph to locate the initial node most relevant to the current problem. Let the entity nodes in the memory knowledge graph be... Its vector representation is Keywords The vector representation is The initial node matching score can then be expressed as: (14); The system selects the highest-scoring entity nodes based on the matching scores to form an initial node set. Then, starting from these initial nodes, it retrieves a first-order subgraph from the memory knowledge graph and extracts all triples from this subgraph as a candidate triple set. To facilitate consistent matching with user questions, the system transcribes the candidate triples into natural language descriptions and then matches them to the rewritten question. Reorder the candidate triples. Let the candidate triples be... Its vector representation is The triple reordering score is expressed as (15); Based on the reordering score, the system selects the highest-scoring candidate triples from the set of candidates in the first-order subgraph. These triples form the final set of evidence for the memory knowledge graph: (16); At the same time, the system also utilizes the rewritten issues Semantic similarity matching is performed on document title nodes to locate the document titles most relevant to the question. Let the title node be... Its vector representation is The title matching score is then expressed as (17); The system selects the top few title nodes with the highest titles based on their matching scores, and then retrieves the corresponding text blocks, tables, and chart descriptions based on the title paths associated with these nodes, forming a document-side evidence set. Further filtering of the most relevant text blocks, tables, and chart descriptions within the matched title range is possible here, but the specific formulas are no longer listed separately.
[0041] Obtain the final evidence set on the memory knowledge graph side. After obtaining the document-side evidence set, the system organizes the two search results uniformly and inputs them into the large language model. Let the multi-source context set composed of the two types of evidence be denoted as . The final answer generation process is represented as follows: (18); in, This represents the question-answering function of a large language model. This represents the final generated answer.
[0042] S105: Improved memory knowledge graph.
[0043] After completing the question-and-answer session, the system does not stop at outputting the answer, but rather periodically refines the memory knowledge graph based on the entire dialogue process (refining it once every 10 rounds of dialogue). For example... Figure 6 As shown, successful retrieval patterns are first recorded, and event candidates, semantic candidates, and program candidates are generated in sequence. Then, value assessment is carried out, and qualified content is written into the long-term memory knowledge graph. Finally, the updated long-term memory knowledge graph is output.
[0044] The refinement process involves recording and summarizing successful retrieval patterns, effective formula call patterns, stable semantic knowledge, key events, and reusable procedures, thereby continuously expanding and optimizing the long-term memory structure. Specifically, the system analyzes the retrieval path, hit document structure, used graph subgraphs, key evidence used to generate the final answer, and user feedback on the answer in the current round of question answering, and generates event candidates, semantic candidates, and procedure candidates from these analyses.
[0045] Let the candidate memory set be: (19); in, These represent the first candidate memory, the second candidate memory, ..., the ... Each candidate memory has 10 candidate memories. It can be represented as: (20); in, Indicates the candidate type, which can be event candidate, semantic candidate, or program candidate. Indicates candidate content. Indicate the source, Represents a timestamp. The system indicates the confidence level. For event candidates, the system focuses on recording significant state changes that occur during the dialogue, such as changes in the user's diet, adjustments in learning direction, and shifts in the focus of the question. For semantic candidates, the system extracts knowledge that appears repeatedly and is semantically stable in multiple rounds of conversation, such as the definition of a module, the core logic of a method, and the purpose of a retrieval strategy. For program candidates, the system records processing flows that have been verified as effective in multiple rounds of question and answer, such as the execution mode of "first locating the definition clauses in the contract, then retrieving the rights and obligations clauses, then identifying risk points based on historical review experience, and finally generating review opinions."
[0046] To determine whether candidate memories are worth including in the long-term memory map, a large language model is used to conduct a multi-dimensional value assessment of candidate memories, ultimately setting candidate memories as... The value score is: (twenty one); in, Indicates persistence score, Indicates the reusability score. Indicates the task relevance score. The score indicates the strength of the evidence. Indicates the conflict risk score. For the corresponding weight parameters. When When a candidate memory is selected, the system writes it into the long-term memory knowledge graph; otherwise, it is only retained in the short-term cache or the current session context. For candidate memories that need to be written into the graph, the system generates new nodes and relationships based on their type. For example, if the candidate is an event candidate, an event node is generated and associated with related entities, session sessions, and timestamps; if the candidate is a semantic candidate, a semantic node is generated or existing stable relationships between entities are supplemented; if the candidate is a program candidate, a program node is generated and dependencies are established with the applicable task, related document structure, and successful retrieval path.
[0047] Let the improved long-term memory knowledge graph be: (twenty two); in, It includes not only the original entity nodes, but also event nodes, semantic nodes, and program nodes. This includes entity relationship edges, event association edges, process dependency edges, and evidence pointing edges. The system continues to use this graph as the basis for retrieval and updates in subsequent sessions, enabling the document agent to continuously accumulate knowledge, understand user preferences, remember effective processes, and gradually form long-term memory.
[0048] In summary, this invention integrates document title structure, text blocks under titles, tables, charts, and entity relationship knowledge accumulated in historical sessions to construct a long-term memory-enhanced question-answering system for multi-turn, multi-session, and cross-time-span tasks. Compared to existing multi-document question-answering methods that rely on fixed segmentation, vector retrieval, and one-time generation, this invention can more accurately understand user questions and achieve joint retrieval of document knowledge and historical memory knowledge, thereby improving contextual continuity, cross-session consistency, and answer generation quality in multi-document question-answering scenarios.
[0049] Moreover, instead of simply treating historical documents and interaction records as external retrieval libraries, it constructs a memory-layer knowledge graph to organize key entities, entity relationships, and subsequent event, semantic, and program candidates in the session into a continuously updated long-term memory structure. This enables the continuous writing, dynamic management, and subsequent retrieval of historical knowledge, gradually transforming previously processed document content and historical Q&A experience into reusable long-term knowledge, thereby improving the stability and adaptability of the system during long-term operation.
[0050] This invention improves the accuracy of question understanding and the degree of matching between search results and users' actual needs through a question-answering process of "question rewriting - memory map retrieval - document title retrieval - multi-source evidence fusion". By continuously recording and improving successful retrieval patterns and multiple candidate memories, the system has the ability to continuously iterate and enhance memory. This method can not only improve the retrieval efficiency and answer accuracy of subsequent question-answering tasks, but also reduce the additional consumption caused by repeated retrieval and repeated reasoning.
[0051] Figure 7 This paper presents a multi-document question-answering system based on long-term memory knowledge graph enhancement, including: The document parsing unit is configured to: receive multiple types of input documents to be processed and perform document parsing operations, identify the internal heading hierarchy structure of the document and decompose the document content according to the heading boundaries, and generate various content units carrying complete contextual attributes; The graph construction unit is configured to perform entity extraction and semantic relationship generation operations for each round of conversation content, and construct a memory knowledge graph specific to the corresponding conversation. The knowledge embedding unit is configured to perform knowledge embedding processing on the content unit and the entities and relations contained in the memory knowledge graph, respectively, generate a semantic vector of a unified dimension and write it into the vector database for storage; The question rewriting unit is configured to perform semantic completion and question rewriting operations on the original question input by the user to generate a clear question; The question-answering generation unit is configured to generate the final question-answering result by fusing multi-source evidence based on a defined question retrieval memory knowledge graph and document knowledge vectors. The memory enhancement unit is configured to: generate candidate memories based on information from the question-and-answer process, evaluate the value of candidate memories, and improve the long-term memory knowledge graph.
[0052] It is understood that the aforementioned units can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of the present invention. The aforementioned units are based on logical functional division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of the present invention, the system may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0053] According to another embodiment of the present invention, the system of this embodiment can be constructed by running a computer program (including program code) capable of performing the steps involved in the corresponding method of the present invention on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the aforementioned computing device through the computer-readable recording medium, and run therein.
[0054] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A multi-document question answering method based on long-term memory knowledge graph enhancement, characterized in that, Includes the following processes: It receives multiple types of input documents to be processed and performs document parsing operations, identifies the internal heading hierarchy of the document and decomposes the document content according to the heading boundaries, and generates various content units with complete contextual attributes; For each round of conversation content, entity extraction and semantic relationship generation operations are performed to construct a memory knowledge graph specific to the corresponding conversation; Knowledge embedding processing is performed on the entities and relations contained in the content units and memory knowledge graphs respectively to generate semantic vectors of a unified dimension and store them in a vector database. Specifically, this includes: performing vectorization processing on text blocks, table descriptions, and chart descriptions in the content units to generate document knowledge vectors; directly vectorizing the entities in the memory knowledge graphs; transcribing the relation triples in the memory knowledge graphs into natural language and then performing vectorization processing to generate memory knowledge vectors; and uniformly writing the document knowledge vectors and memory knowledge vectors into the vector database and associating them with corresponding metadata information. Perform semantic completion and question rewriting operations on the original question input by the user to generate a more explicit question; Based on the explicit question retrieval memory knowledge graph and document knowledge vector, the final question answering result is generated by integrating multi-source evidence. Candidate memories are generated based on the question-and-answer process information, the value of the candidate memories is evaluated, and the long-term memory knowledge graph is improved. The generation of candidate memories based on the question-and-answer process information includes: analyzing the retrieval path, hit document structure, graph subgraphs, and user feedback information in the process of generating question-and-answer results; extracting event-type, semantic-type, and program-type candidate memories, and generating a candidate memory set containing candidate type, candidate content, source information, timestamp, and confidence information. The evaluation of candidate memory value includes: a weighted score is calculated by considering five dimensions of candidate memory in the candidate memory set: persistence, reusability, task relevance, evidence strength, and conflict risk. When the value score reaches a preset value threshold, the candidate memory is determined to be a valid memory and included in the writing scope. When the value score does not reach the preset value threshold, the candidate memory is temporarily stored in a short-term cache.
2. The multi-document question answering method based on long-term memory knowledge graph enhancement as described in claim 1, characterized in that, It receives multiple types of input documents to be processed and performs document parsing operations, identifies the internal heading hierarchy structure of the document, and breaks down the document content according to the heading boundaries to generate various content units with complete contextual attributes, including: It identifies headings, body paragraphs, tables, and charts in the input document and constructs the hierarchical relationships between headings within the document. Based on the heading hierarchy boundaries, it breaks down the overall content of the input document and generates structured content units that include text content, content type, the path of the heading, the page number, adjacent identifiers, and document version information.
3. The multi-document question answering method based on long-term memory knowledge graph enhancement as described in claim 1, characterized in that, For each round of conversation content, entity extraction and semantic relationship generation operations are performed to construct a memory knowledge graph specific to the corresponding conversation, including: Scan the session content and identify key entities with memory value, and generate an entity set containing entity name, entity type, embedding vector, timestamp, source identifier and confidence information; Calculate the semantic similarity between different entities in the entity set, and perform normalization merging on semantically consistent entities; Analyze the semantic relationships between entities in the entity set and generate relation triples to construct a memory knowledge graph containing entity nodes and relation edges.
4. The multi-document question answering method based on long-term memory knowledge graph enhancement as described in claim 3, characterized in that, Calculate the semantic similarity between different entities in the entity set, and perform normalization merging on semantically consistent entities, including: Extract the embedding vectors corresponding to the two entities to be compared, and calculate the cosine similarity between the two embedding vectors; When the cosine similarity reaches the preset similarity threshold, the two entities are determined to be the same semantic entity and a node merging operation is performed. When the cosine similarity does not reach the preset similarity threshold, an independent node is created for the new entity and added to the entity set.
5. The multi-document question answering method based on long-term memory knowledge graph enhancement as described in claim 1, characterized in that, Perform semantic completion and question rewriting operations on the original question input by the user to generate a more explicit question, including: Obtain the original question input by the user, and retrieve recent historical dialogue information and dialogue summary content from the conversation content; Complete the omitted information, referential information, and context-dependent information in the original problem to generate a semantically clear and intentionally explicit problem.
6. The multi-document question answering method based on long-term memory knowledge graph enhancement as described in claim 5, characterized in that, Based on a defined question retrieval memory knowledge graph and document knowledge vectors, multi-source evidence is fused to generate the final question-answering result, including: Extract keywords from the explicit question, match them with entity nodes in the memory knowledge graph, and obtain the corresponding first-order subgraph; After transcribing the relation triples in the first-order subgraph into natural language, semantic matching is performed with the explicitation problem, and the triples with the highest matching degree are selected as graph-side evidence. Match the document title node corresponding to the document knowledge vector and obtain the associated text, tables and charts as document-side evidence; By integrating graph-based evidence and document-based evidence, the final question-and-answer results are generated by inputting the language model.
7. The multi-document question answering method based on long-term memory knowledge graph enhancement as described in claim 1, characterized in that, Improve the long-term memory knowledge graph, including: Based on the type of valid candidate memory, corresponding nodes are generated: event-type candidate memories generate event nodes, semantic-type candidate memories generate semantic nodes, and program-type candidate memories generate program nodes. Establish connections between new nodes and existing entity nodes, event nodes, and program nodes; update the node set and edge set of the long-term memory knowledge graph to complete the graph improvement operation.
8. A multi-document question answering system based on long-term memory knowledge graph enhancement, employing the multi-document question answering method based on long-term memory knowledge graph enhancement as described in any one of claims 1-7, characterized in that, include: The document parsing unit is configured to: receive multiple types of input documents to be processed and perform document parsing operations, identify the internal heading hierarchy structure of the document and decompose the document content according to the heading boundaries, and generate various content units carrying complete contextual attributes; The graph construction unit is configured to perform entity extraction and semantic relationship generation operations for each round of conversation content, and construct a memory knowledge graph specific to the corresponding conversation. The knowledge embedding unit is configured to perform knowledge embedding processing on the content unit and the entities and relations contained in the memory knowledge graph, respectively, generate a semantic vector of a unified dimension and write it into the vector database for storage; The question rewriting unit is configured to perform semantic completion and question rewriting operations on the original question input by the user to generate a clear question; The question-answering generation unit is configured to generate the final question-answering result by fusing multi-source evidence based on a defined question retrieval memory knowledge graph and document knowledge vectors. The memory enhancement unit is configured to: generate candidate memories based on information from the question-and-answer process, evaluate the value of candidate memories, and improve the long-term memory knowledge graph.
Citation Information
Patent Citations
Method and system for enhancing RAG questions and answers through mixed retrieval method
CN118627625A
Multi-document industry knowledge assistant system based on multi-modal knowledge graph
CN121350265A
Smart middle-screen voice assistant family scene personalized reply system
CN121743437A