Domain heterogeneous document question and answer enhancement method based on distant knowledge join reasoning, electronic equipment and computer readable storage medium
By constructing a knowledge graph that connects knowledge of distant relatives and close relatives and combining it with a large language model for intelligent question and answering, we solve the problems of completeness and accuracy of question and answer systems in complex fields and achieve high-quality answer generation.
Patent Information
- Application Number
- CN202510669928.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-12
AI Technical Summary
Existing intelligent question-answering systems have problems with insufficient question-answering completeness and low accuracy in complex fields. Especially when dealing with questions that require reasoning, traditional methods find it difficult to generate high-quality answers.
By defining distant and close relatives, building a knowledge graph, extracting multi-heterogeneous document data and performing vectorized processing, combining a large language model for intelligent answering and retrieval enhancement, and using a self-reflective learning mechanism for answer generation, the integrity and accuracy of knowledge are ensured.
It significantly improves the completeness and accuracy of question and answering in complex fields, especially in professional scenarios, and can generate more complete and accurate answers.
Smart Images

Figure CN120632019A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and natural language processing, and in particular to the technical field of intelligent question answering. It specifically proposes a domain heterogeneous document question answering enhancement method based on distant relative knowledge connection reasoning, an electronic device and a computer-readable storage medium. Background Art
[0002] With the rapid development of information technology and the emergence of massive amounts of data, automated knowledge question-answering systems have garnered widespread attention and research in various fields. In particular, technological advances based on large language models (LLMs) have made the Retrieval-Augmented Generation (RAG) method a crucial tool for complex document processing. The industry has proposed the Graph RAG method, which achieves more accurate question-answering of private text corpora by constructing a knowledge graph and expanding it based on the prevalence of user questions and the amount of source text to be indexed. However, in complex domains and complex reasoning-based questions, the current field of intelligent question-answering technology still suffers from insufficient completeness and accuracy in question-answering processing. Summary of the Invention
[0003] Technical Solution: To solve the above technical problems, the present invention defines distant and close kinship relationships. Based on the relationship settings, accurate intelligent question-answering results are generated by building a knowledge base and enhancing search. The specific technical solution provides:
[0004] A method for enhancing question answering of domain heterogeneous documents based on distant relative knowledge connection reasoning, the method comprising:
[0005] Define distant and close relationships in heterogeneous documents. Distant relationships refer to the knowledge about the spatial distance between the inference paths of entity nodes in the document within a set threshold. For example, the threshold can be set to 2-5 hops between the inference paths of entity nodes. Close relationships refer to the knowledge about the current text block in the document.
[0006] Extract data from multiple heterogeneous documents in the field and build a vector database;
[0007] Construct distant and close kinship relationships between different spatial locations in the knowledge graph and associated documents;
[0008] The final question-and-answer answer is obtained through intelligent answering, retrieval enhancement and verification.
[0009] As an improvement, when extracting data from multiple heterogeneous documents within a domain, we first perform preliminary data cleaning and content extraction on the heterogeneous documents, then perform semantic segmentation and divide them into semantic layers, and set text blocks with an upper threshold for the number of tokens; finally, we vectorize the text blocks through an embedding model and store them in a vector database.
[0010] As an improvement, it also includes preliminary recognition and professional field division of the text content in the vector database.
[0011] As an improvement, the method of constructing a knowledge graph is to use a large language model to extract data from the vector database, and according to different professional fields, realize the adaptation of the entity terminology library combined with different knowledge extraction prompt words; in the process of knowledge disambiguation, the timestamps and data lineage of different data sources must be retained, and when necessary, they can be traced back to the source document material.
[0012] As an improvement, the intelligent answering process includes:
[0013] (1) Perform basic analysis on the user input question through the large language model (LLM) to generate instructions for knowledge graph reasoning; then perform knowledge reasoning on close-kin knowledge connections;
[0014] (2) Based on the entities and relationships involved in the reasoning chain, vector retrieval is performed on the source material to match relevant source material paragraphs;
[0015] (3) Generate answers based on the graph reasoning chain and related TopP source material segments, where the TopP source material segments are a conventional and well-known sampling strategy based on cumulative probability thresholds, which is used to screen the range of candidate words when generating text).
[0016] As an improvement, the retrieval enhancement and inspection method is to first perform a rationality check and a completeness check. If it fails, it returns to the intelligent answer (2) and performs distant relative knowledge reasoning until the recalled knowledge is sufficient to support the question answer. When the distant relative distance reaches the upper limit and still cannot recall enough knowledge to answer the question, it returns failure. Based on the recalled knowledge, the final question-answering answer is obtained through the large language model LLM.
[0017] In the present invention, another specific embodiment provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned method.
[0018] In another specific embodiment of the present invention, a computer-readable storage medium is provided, which stores a computer program, wherein the computer program implements any of the above-mentioned methods when executed by a processor.
[0019] Beneficial Effects: The proposed method offers significant improvements over traditional RAG and GraphRAG intelligent question answering, most notably in terms of completeness and avoidance of hallucinations. This improvement is particularly evident in scenarios involving high-quality internal document repositories, where the effectiveness of answers is significantly improved. Therefore, the proposed method is particularly well-suited for professional domains, answering questions requiring specific reasoning logic. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flowchart of the heterogeneous document question-answering enhancement method based on distant relative and close relative knowledge connection reasoning of the present invention.
[0021] Figure 2 This is a logical diagram of Example 1 of the present invention. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present invention will be described clearly and completely below so that those skilled in the art can better understand the advantages and features of the present invention and thus more clearly define the scope of protection of the present invention. The embodiments described in the present invention are only some of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without making any creative work shall fall within the scope of protection of the present invention.
[0023] This paper proposes a method for enhancing heterogeneous document question answering based on the inference of distant and close relative knowledge connections, called Knowledge Interleaved Retrieval-Augmented Generation (KiRAG). Two relationships are defined in this paper: distant relatives refer to related knowledge within the document whose spatial distance exceeds a set threshold, while close relatives refer to related knowledge within the current text block of the document. The threshold is the spatial distance between the inference paths between entity nodes in the document, for example, 2 to 5 hops.
[0024] See Figure 1 As shown, it is a flow chart of the present invention, which includes two parts: one is the knowledge base construction process, and the other is the intelligent question-answering process of retrieval enhancement generation.
[0025] Part 1: The knowledge base construction process.
[0026] The steps of this part include: Step 1. Perform preliminary data cleaning and content extraction on heterogeneous documents, and perform semantic segmentation to divide them into semantic layers and text blocks with a certain upper limit on the number of tokens.
[0027] Step 2. Vectorize the text block using the embedding model and store it in the vector database.
[0028] Step 3. Conduct preliminary identification of text content and divide it into professional fields.
[0029] Step 4. Based on different professional fields, the entity terminology library is adapted to different knowledge extraction prompts and extracted into a knowledge graph using a large language model. During the knowledge disambiguation process, the timestamps and data lineage of different data sources are preserved, and when necessary, the source document can be traced back to the source material.
[0030] In step 4, the knowledge extraction method is adapted according to different special professional fields, and the data governance method in the graph ablation process can achieve complete and accurate labeling of data lineage.
[0031] The second part is the intelligent question answering process of retrieval enhancement generation.
[0032] The steps of this part include: Step 1. Use LLM to perform basic analysis on the questions input by the user and generate instructions for knowledge graph reasoning.
[0033] Step 2. Conduct knowledge reasoning of close relative knowledge connections.
[0034] Step 3. Based on the entities and relationships involved in the reasoning chain, perform vector retrieval on the source material and match it to the relevant source material paragraphs.
[0035] Step 4. Generate answers based on the graph reasoning chain and related TopP source material clips.
[0036] Step 5. Perform a rationality check and a completeness check. If they fail, return to step 2 and perform distant relative knowledge reasoning until the recalled knowledge is sufficient to support the answer to the question. When the distant relative distance reaches the upper limit (such as more than 3 hops) and sufficient knowledge cannot be recalled to answer the question, the process returns failure.
[0037] In this step 5, a self-referencing method is used to implement layer-by-layer reasoning of the knowledge graph, mine complex knowledge associations, discover implicit knowledge, and generate high-quality answers.
[0038] Step 6. Based on the recall knowledge from step 5, the final answer is generated through the large language model LLM.
[0039] The technical solutions and technical effects of the present invention are further described and illustrated below through specific embodiments.
[0040] Example 1
[0041] See Figure 2 As shown, the present invention raises the question: How big is the reconnaissance range of the U.S. fifth-generation aircraft? In order to solve the problem of the reconnaissance range of the U.S. fifth-generation aircraft, the following processing will be performed: jump from the node of "U.S. fifth-generation aircraft" to "F-22", and then jump to "AN / APG-77". With the attribute of "AN / APG-77", "detection distance 400 kilometers" is deduced to have a detection range of 400 kilometers for the U.S. fifth-generation aircraft. The reconnaissance capability of the U.S. fifth-generation aircraft mainly depends on its AN / APG-77 active electronically scanned array radar, and the maximum detection range of non-stealth conventional targets is about 400 kilometers.
[0042] Example 2
[0043] The test data used in this example was selected from online military news articles. This dataset contains military-related news articles published between January 2020 and December 2024, with content in both Chinese and English, and approximately 1.2 million tokens. Using the Rageval evaluation framework, we tested completeness, hallucination detection, irrelevance, and text generation quality (Rouge-L). We compared the GraphRag and LightRag frameworks within the knowledge graph Rag framework, with the results shown in Table 1.
[0044] Table 1 Test results
[0045]
[0046] Conclusion: This method aims to improve the accuracy and completeness of RAG's answer generation for complex reasoning questions. By dynamically constructing a domain knowledge graph for documents, associating the knowledge of distant and close relatives at different spatial locations of documents, and adopting a self-reflective learning mechanism for iterative evolution in the answer generation process, the reasoning ability of the RAG method in complex domains is enhanced, ensuring the completeness and accuracy of question and answer.
[0047] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A method for enhancing question answering of domain heterogeneous documents based on distant relative knowledge connection reasoning, characterized by: The method includes Define distant and close relationships in heterogeneous documents, where distant relationships refer to the knowledge of spatial distances between inference paths between entity nodes in the document outside a set threshold, and close relationships refer to the knowledge in the current text block of the document; Extract data from multiple heterogeneous documents in the field and build a vector database; Construct distant and close kinship relationships between different spatial locations in the knowledge graph and associated documents; The final question-and-answer answer is obtained through intelligent answering, retrieval enhancement and verification.
2. The method for enhancing question answering of heterogeneous documents based on distant relative knowledge connection reasoning according to claim 1 is characterized by: When extracting data from multiple heterogeneous documents within a domain, we first perform preliminary data cleaning and content extraction on the heterogeneous documents, then perform semantic segmentation and divide them into semantic layers, and set text blocks with an upper threshold for the number of tokens; finally, we vectorize the text blocks using an embedding model and store them in a vector database.
3. The method for enhancing question answering of heterogeneous documents based on distant relative knowledge connection reasoning according to claim 1 is characterized by: It also includes preliminary recognition and professional field division of text content in the vector database.
4. The method for enhancing question answering of heterogeneous documents based on distant relative knowledge connection reasoning according to claim 1 is characterized by: The method of constructing the knowledge graph is to extract the data in the vector database through a large language model, and realize the adaptation of the entity terminology library and different knowledge extraction prompt words according to different professional fields. In the process of knowledge disambiguation, the timestamps and data lineage of different data sources must be retained, and when necessary, they can be traced back to the source document material.
5. The method for enhancing question answering of domain heterogeneous documents based on distant relative knowledge connection reasoning according to claim 1 is characterized by: The process of intelligent answering includes (1) Perform basic analysis on the user input question through the large language model (LLM) to generate instructions for knowledge graph reasoning; then perform knowledge reasoning on close-kin knowledge connections; (2) Based on the entities and relationships involved in the reasoning chain, vector retrieval is performed on the source material to match relevant source material paragraphs; (3) Generate answers based on the graph reasoning chain and related TopP source material fragments.
6. The method for enhancing question answering of heterogeneous documents based on distant relative knowledge connection reasoning according to claim 5 is characterized by: The method of retrieval enhancement and checking is to first conduct a rationality check and a completeness check. If they fail, the system returns to (2) of the intelligent answer and performs distant relative knowledge reasoning until the recalled knowledge is sufficient to support the answer to the question; When the distance between distant relatives reaches the upper limit and sufficient knowledge cannot be recalled to answer the question, a failure is returned; Based on the recalled knowledge, the final question-answering answer is generated through the large language model LLM.
7. An electronic device, characterized in that: include: at least one processor; And, a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as claimed in any one of claims 1 to 6.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.