Government affair question and answer method and system based on knowledge graph and retrieval enhancement

By extracting and integrating multiple information from government documents in the government Q&A system, building a knowledge graph, and using search enhancement technology to generate and optimize answers, the existing system's poor accuracy and potential risk problems in dealing with multimodal information and complex problems are solved, and higher Q&A accuracy and user experience are achieved.

CN120179797AActive Publication Date: 2025-06-20北京大学长沙计算与数字经济研究院 +1

Patent Information

Application Number
CN202510670423.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing government Q&A system has problems of poor accuracy and potential risks when dealing with the "illusion" phenomena of multimodal information, complex problems and large language models.

Method used

By extracting text information, multimodal information and metadata from government documents, fuse this information to build a knowledge graph, and generate query answers in combination with search enhancement techniques, optimize answers to improve accuracy and security.

Benefits of technology

It significantly improves the accuracy of government affairs Q&A, reduces the possible "illusion" phenomena in large language models, ensures the safety, accuracy and readability of the answers, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179797A_ABST
    Figure CN120179797A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large models, and discloses a knowledge graph and retrieval enhancement-based government affair question-answering method and system, and the method comprises the steps: carrying out the information extraction of each government affair document, and obtaining text information, multi-modal information and metadata; fusing the information to obtain document information; performing entity relationship extraction based on the document information to obtain entity, relationship and text key value pairs; constructing a government affair knowledge graph based on the entity, the relationship and the text key value pair corresponding to each government affair document; based on the type of the user query statement and the government affair knowledge graph, adopting a retrieval enhancement technology to generate a query answer; and optimizing the query answer to generate a target answer. According to the method, through the multi-modal information and the knowledge graph, the document content and the query intention are comprehensively understood, the question and answer accuracy is improved, the knowledge graph and a retrieval enhancement technology are fused, the illusion phenomenon of a large language model is reduced, the answer accuracy is improved, the safety, accuracy and readability of the answer are ensured through optimization, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large models, and specifically relates to a government affairs question and answer method and system based on a knowledge graph and retrieval enhancement. Background Art

[0002] Currently, government affairs question and answer is mainly carried out in the following ways: one is query-based question and answer based on keyword matching, which matches the question input by the user with a pre-set keyword library and queries relevant content from the document library as the answer to return; the other is to introduce natural language processing technologies, such as using large language models, knowledge graphs and other technologies, to try to understand the semantics of the user's question and provide a more intelligent answer.

[0003] However, the problems existing in the above two methods are as follows: (1) Weak multi-modal information processing ability: It seriously relies on text information and is unable to process corresponding information when facing government affairs documents containing multi-modal information such as pictures, tables, and formulas, resulting in poor context understanding ability and thus affecting the accuracy of question and answer.

[0004] (2) Unable to handle complex problems: Facing complex government affairs problems, the keyword matching-based method can only search for keywords in isolation and is unable to logically disassemble and deeply analyze the problem. The query answers provided may be one-sided and fragmented, unable to meet the user's needs.

[0005] (3) "Hallucination" problem of large language models: When using large language models for government affairs question and answer, the "hallucination" phenomenon occurs from time to time. Although seemingly reasonable answers can be generated, there may be factual errors.

[0006] (4) Imperfect answer processing mechanism: After generating the query answer, it is directly passed to the user, resulting in poor accuracy and may disclose sensitive information, triggering potential risks. Summary of the Invention

[0007] In view of this, the present invention provides a government affairs question and answer method and system based on a knowledge graph and retrieval enhancement to solve the problems of poor accuracy and potential risks existing in the existing government affairs question and answer.

[0008] In the first aspect, the present invention provides a government affairs question and answer method based on a knowledge graph and retrieval enhancement, and the method includes: For each government affairs document, extract information from the government affairs document to obtain text information, multi-modal information, and metadata; Fuse the text information, multi-modal information, and metadata to obtain the document information of the government affairs document; Based on the document information of the government affairs document, perform entity relationship extraction to obtain entities, relationships, and text key-value pairs corresponding to the government affairs document; Construct a government affairs knowledge graph based on the entities, relationships, and text key-value pairs corresponding to each government affairs document; Generate query answers for the user's query statement using retrieval enhancement technology based on the type of the user's query statement and the government affairs knowledge graph; Optimize the query answers for the user's query statement to generate the target answers for the user's query statement.

[0009] The government affairs question-answering method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention ensures a comprehensive understanding of government affairs documents by extracting text information, multimodal information, and metadata, fuses different types of information to obtain document information, ensures the ability to fully understand the context, extracts entities and relationships from the document information, and generates text key-value pairs, enabling a better understanding of the semantic information in the document. By using entities, relationships, and text key-value pairs to construct a government affairs knowledge graph, logical reasoning and in-depth analysis can be performed based on the government affairs knowledge graph. At the same time, an efficient query mechanism is provided, which can quickly locate relevant information in large-scale data. When performing government affairs question-answering, combined with the user's query type and the knowledge graph, retrieval enhancement technology can more accurately match relevant answers, and through answer optimization, ensure that the generated answers are more accurate. Through multimodal information processing and the introduction of the knowledge graph, the content of the document and the user's query intention can be more comprehensively understood, significantly improving the accuracy of question-answering. By fusing the knowledge graph and retrieval enhancement technology, the "hallucination" phenomenon that may occur in large language models is reduced, the accuracy of the answers is improved, and through the optimization step, the security, accuracy, and readability of the answers are further ensured, enhancing the user experience.

[0010] In an alternative implementation manner, entity relationship extraction is performed based on the document information of the government affairs document to obtain the entities, relationships, and text key-value pairs corresponding to the government affairs document, including: Chunk the document information of the government affairs document to obtain multiple text chunks; For each text chunk, use a large language model to identify entities and relationships from the text chunk; For each entity, use the entity as the key and the relevant content of the entity in the text chunk as the value to generate the text key-value pair corresponding to the entity; For each relationship, use the relationship and the entities constituting the relationship as the key and the relevant content of the relationship in the text chunk as the value to generate the text key-value pair corresponding to the relationship.

[0011] The government affairs question - answering method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention can process these text chunks in parallel by splitting document information into multiple smaller text chunks. Within the smaller text chunks, the large - language model can better perceive the context, ensuring the accuracy of recognition results. The recognized entities, relationships, and their related content are stored in the form of key - value pairs, facilitating subsequent efficient querying and use, and improving the response speed.

[0012] In an alternative embodiment, based on the entities, relationships, and text key - value pairs corresponding to each government affairs document, a government affairs knowledge graph is constructed, including: Pre - process the text key - value pairs based on the text key - value pairs corresponding to each government affairs document to obtain target key - value pairs; Create nodes based on the entities in each government affairs document and create edges based on the relationships in each government affairs document; Based on the target key - value pairs corresponding to each government affairs document, connect the nodes and edges, and add attributes to the nodes and edges respectively to generate a government affairs knowledge graph, where the attribute refers to the value corresponding to the node or the value corresponding to the edge.

[0013] The government affairs question - answering method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention can improve the quality of text key - value pairs through pre - processing, convert the entities, relationships, and text key - value pairs in government affairs documents into nodes, edges, and their connection relationships in the graph, and add attributes to the nodes and edges at the same time, which can represent the information in government affairs documents more comprehensively and accurately. And the knowledge graph supports graph - based querying and reasoning, enabling more complex query problems to be processed and more complete answers to be provided.

[0014] In an alternative embodiment, before generating the query answer for the user query statement based on the type of the user query statement and the government affairs knowledge graph using retrieval enhancement technology, the method further includes: For each node in the government affairs knowledge graph, construct a vector corresponding to the node based on the node attributes of the node, the edges directly connected to the node, and the node and its attributes; Construct a vector database based on the vectors corresponding to all nodes.

[0015] The government affairs question - answering method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention extracts feature information from the attributes of the node, the edges directly connected to the node, and the node and its attributes, converts it into a vector, and thus obtains a vector database constructed based on the vectors corresponding to all nodes. Through vectorization processing, complex structured information can be converted into a form convenient for calculation, and at the same time, more complex query and analysis tasks are supported.

[0016] In an alternative embodiment, generating the query answer for the user query statement based on the type of the user query statement and the government affairs knowledge graph using retrieval enhancement technology includes: When the type of the user query statement is a concept type, a large language model is used to identify query entities from the user query statement; The query entities are converted into first query vectors, and at least one first target vector whose similarity to the first query vectors in the vector database is greater than a preset similarity threshold is determined; Based on at least one first target vector, corresponding first target nodes, first target edges, and first adjacent nodes are determined from the government affairs knowledge graph. The first target edges refer to the edges directly connected to the first target nodes, and the first adjacent nodes refer to the other nodes directly connected to the first target edges; Based on the attributes of the first target nodes, the attributes of the first adjacent nodes, the first target edges, and the attributes of the first target edges, a large language model is used to generate query answers to the user query statement.

[0017] The government affairs question-answering method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention, when the query statement input by the user is of the concept type, identifies and extracts entities in the query statement, converts them into first query vectors, and thus queries first target vectors from the vector database based on vector similarity. Then, based on the first target vectors, corresponding first target nodes, first target edges, and first adjacent nodes are determined from the government affairs knowledge graph. Not only are the target nodes found, but also the associated edges and nodes are extended, providing more comprehensive information. Finally, a large language model is used to generate query answers to ensure that the query answers are accurate, complete, and conform to the user's query intent, and the combination of the vector database and the knowledge graph significantly improves the query speed and efficiency.

[0018] In an alternative embodiment, when generating query answers to the user query statement by using retrieval enhancement technology based on the type of the user query statement and the government affairs knowledge graph, it further includes: When the type of the user query statement is a process type, a large language model is used to identify query entities, query relationships, and abstract concepts from the user query statement; The query entities, query relationships, and abstract concepts are respectively converted into second query vectors, and at least one second target vector whose similarity to each second query vector in the vector database is greater than a preset similarity threshold is respectively determined; Based on at least one second target vector, corresponding second target nodes, second target edges, and second adjacent nodes are determined from the government affairs knowledge graph. The second target edges include the edges directly connected to the second target nodes and the edges whose edge attributes are consistent with the query relationships, and the second adjacent nodes refer to the other nodes directly connected to the second target edges; The thinking chain technology is used to process based on the attributes of the second target nodes, the attributes of the second adjacent nodes, the second target edges, and the attributes of the second target edges to obtain query results; Use a large language model to integrate based on the query results and generate the query answer for the user's query statement.

[0019] The government affairs Q&A method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention, when the query statement input by the user is of the process type, identifies and extracts the entities, relationships and abstract concepts in the query statement, converts them into a second query vector, and thus queries the second target vector from the vector database based on the vector similarity. Then, based on the second target vector, determines the corresponding second target node, second target edge and second adjacent node from the government affairs knowledge graph, provides more comprehensive information, and finally uses a large language model to generate the query answer, ensuring that the query answer is accurate, complete and conforms to the user's query intention. Moreover, the combination of the vector database and the knowledge graph significantly improves the query speed and efficiency.

[0020] In an alternative embodiment, optimizing the query answer of the user's query statement to generate the target answer of the user's query statement includes: Performing rule filtering and model filtering on the query answer to obtain an initial answer; Using a large language model to score the initial answer to obtain a quality score; When the quality score is greater than the preset score threshold, taking the initial answer as the target answer; or, When the quality score is not greater than the preset score threshold, taking the initial answer as a new query answer and returning it to the step of performing rule filtering and model filtering on the query answer to obtain an initial answer until the quality score is not less than the preset score threshold, and taking the initial answer as the target answer.

[0021] The government affairs Q&A method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention improves the overall accuracy of the answer through a dual filtering mechanism by performing rule filtering and model filtering on the query answer, scores the initial answer after rule and model filtering to ensure that the quality score of the generated answer meets the requirements, improves the accuracy and reliability of the final answer, and at the same time avoids potential sensitivity risks.

[0022] In an alternative embodiment, for each government affairs document, information extraction is performed on the government affairs document to obtain text information, multimodal information and metadata, including: Converting the government affairs document into a standard format; Using a parsing tool to extract the text information and metadata of the government affairs document in the standard format; Using a computer vision model to perform non-text recognition on the government affairs document in the standard format to obtain non-text content; Performing preprocessing on the non-text content to obtain multimodal information.

[0023] The government affairs Q&A method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention unifies government affairs documents in different formats into a standard format for subsequent processing, and then uses parsing tools to extract the plain text content and metadata in the government affairs documents. At the same time, computer vision technology is used to extract non-text elements, so that when processing government affairs documents, it is not limited to understanding text, but can also interpret the semantics in multi-modal information, so as to comprehensively grasp the content of the documents.

[0024] In an alternative embodiment, text information, multi-modal information, and metadata are fused to obtain the document information of the government affairs document, including: Preprocess the text information and multi-modal information to obtain the preprocessed text information and multi-modal information; Adopt a large language model to convert metadata, preprocessed text information, and multi-modal information into vectors respectively; Adopt a fusion layer to perform feature fusion on multiple vectors to obtain a fusion vector; Decode the fusion vector to obtain the document information of the government affairs document.

[0025] The government affairs Q&A method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention improves the quality of text and multi-modal information by preprocessing the extracted text information, and uses a large language model to encode the preprocessed text information and convert it into a high-dimensional vector representation, which can capture the semantic information of the text. Therefore, a fusion layer is used to perform feature fusion on multiple vectors to obtain a fusion vector, and finally the fusion vector is decoded into readable document information, which can better reflect the semantic information of the government affairs document.

[0026] In a second aspect, the present invention provides a government affairs Q&A system based on knowledge graph and retrieval enhancement, the system includes: A document information processing module, which is used for each government affairs document, to extract information from the government affairs document to obtain text information, multi-modal information, and metadata, and fuse the text information, multi-modal information, and metadata to obtain the document information of the government affairs document; A knowledge graph construction module, which is used to extract entity relationships based on the document information of the government affairs document to obtain entities, relationships, and text key-value pairs corresponding to the government affairs document, and construct a government affairs knowledge graph based on the entities, relationships, and text key-value pairs corresponding to each government affairs document; A content query and generation module, which is used to generate query answers to user query statements based on the type of user query statements and the government affairs knowledge graph by using retrieval enhancement technology; A data quality control module, which is used to optimize the query answers to user query statements to generate target answers to user query statements.

[0027] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the government affairs Q&A method based on knowledge graph and retrieval enhancement according to the first aspect or any corresponding embodiment thereof.

[0028] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the government affairs Q&A method based on knowledge graph and retrieval enhancement according to the first aspect or any corresponding embodiment thereof.

[0029] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, and the computer instructions are used to cause a computer to execute the government affairs Q&A method based on knowledge graph and retrieval enhancement according to the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0031] Figure 1 is a schematic diagram of a government affairs Q&A system based on knowledge graph and retrieval enhancement according to an embodiment of the present invention; Figure 2 is a flowchart of a government affairs Q&A method based on knowledge graph and retrieval enhancement according to an embodiment of the present invention; Figure 3 is a schematic flowchart of a document information processing module according to an embodiment of the present invention; Figure 4 is a schematic diagram of a government affairs knowledge graph according to an embodiment of the present invention; Figure 5 is a schematic flowchart of a data quality control module according to an embodiment of the present invention; Figure 6 is a schematic hardware structure diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0033] Existing government affairs Q&A is mainly carried out in the following ways: one is query-based Q&A based on keyword matching; the other is the introduction of natural language processing technology. However, the above two methods have problems such as poor answer accuracy and potential risks. In the prior art, multi-strategy switching is used to process government affairs questions, focusing on multi-source data integration and intention inheritance. In addition, the prior art uses a general Q&A system to support the generation of query statements by large language models + matching of similar questions. The government affairs Q&A method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention can more comprehensively understand the document content and user query intention through multi-modal information processing and the introduction of knowledge graphs, significantly improving the accuracy of Q&A. By integrating knowledge graph and retrieval enhancement technologies, it reduces the "hallucination" phenomenon that may occur in large language models, improves the accuracy of answers, and further ensures the security, accuracy, and readability of answers through optimization steps, enhancing the user experience. At the same time, the present invention focuses on the combination of logical reasoning and retrieval enhancement of government affairs knowledge graphs, designs dual-mode retrieval for conceptual / process questions, and improves the parsing accuracy of complex government affairs processes, forming a unique technical path in three aspects: knowledge graph construction, Q&A mechanism, and answer optimization, solving the problems of insufficient generalization, weak logical reasoning, and security and compliance in the prior art in the government affairs field.

[0034] According to the embodiments of the present invention, an embodiment of a government affairs Q&A method based on knowledge graph and retrieval enhancement is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0035] Figure 1 is a schematic diagram of a government affairs Q&A system based on knowledge graph and retrieval enhancement according to the embodiments of the present invention, as Figure 1As shown in the figure, the system includes: a document information processing module, which is used to extract information from each government document to obtain text information, multimodal information, and metadata, and fuse the text information, multimodal information, and metadata to obtain the document information of the government document; a knowledge graph construction module, which is used to extract entity relationships based on the document information of the government document to obtain the entities, relationships, and text key-value pairs corresponding to the government document, and construct a government knowledge graph based on the entities, relationships, and text key-value pairs corresponding to each government document; a content query and generation module, which is used to generate query answers for the user query statement based on the type of the user query statement and the government knowledge graph by using retrieval enhancement technology; and a data quality control module, which is used to optimize the query answers for the user query statement to generate the target answers for the user query statement.

[0036] Specifically, the document information processing module is used to extract information from each government document to obtain text information, multimodal information, and metadata, and fuse the above three types of data to obtain document information that fully mines semantics and understands the context. The knowledge graph construction module is used to extract entities and relationships from the document information, and construct text key-value pairs based on this, so as to construct a government knowledge graph according to the entities, relationships, and text key-value pairs, to intuitively display the information contained in the government document and serve as the basis for efficient query of government question answering. The content query and generation module is used to efficiently and accurately query the corresponding information from the government knowledge graph based on the user query statement and generate query answers. The data quality control module is used to optimize the query answers to avoid potential risks.

[0037] In this embodiment, a government question answering method based on a knowledge graph and retrieval enhancement is provided, which can be used in the above-mentioned government question answering system based on a knowledge graph and retrieval enhancement. Figure 2 is a flowchart of the government question answering method based on a knowledge graph and retrieval enhancement according to an embodiment of the present invention. As Figure 2 shown, the process includes the following steps: Step S201: For each government document, extract information from the government document to obtain text information, multimodal information, and metadata. Specifically, a government document refers to a document used to record, convey, and process various government information, with legal effect and administrative binding force, and is the basis for government question answering, including policy documents, meeting minutes, work reports, application approval documents, etc. For each government document, extract the text information, multimodal information, and metadata therein. Among them, the text information refers to pure text content; the multimodal information refers to non-pure text content, including pictures, tables, and formulas, etc.; the metadata is used to describe the attributes of the government document, including basic attributes: title, author, subject, keywords, creation / modification date, etc., technical attributes: PDF version, software for generating PDF, encryption status, etc., extended attributes: XMP (Extensible Metadata Platform) metadata, custom metadata, etc. By extracting the above multiple types of information, compared with the related technology that only processes single text information, it can comprehensively extract and understand the content of the government document, providing a basis for conducting government question answering.

[0038] Step S202: Integrate the text information, multimodal information, and metadata to obtain the document information of the government document. Specifically, integrate multiple types of information, comprehensively integrate various aspects of information in the government document, fully explore the semantic information in the government document, form the document information of the government document, and make different types of information complement each other, providing a more complete and accurate data basis for subsequent analysis.

[0039] Step S203: Perform entity relationship extraction based on the document information of the government document to obtain the entities, relationships, and text key-value pairs corresponding to the government document. Specifically, identify different entities (such as names, times, locations, events, etc.) and their relationships from the integrated document information. For example, from the text "For more than 3 owners in the same unit, as well as the owners' committee or the property service enterprise entrusted by it, can submit an application to the Municipal Special Maintenance Fund Management Center for Elevator Maintenance Insurance with the funds in the value-added income pooling account", extract entities such as "owners", "owners' committee", "property service enterprise", "Municipal Special Maintenance Fund Management Center for Elevator Maintenance Insurance", and relationships such as "the owners' committee submits an application to the Municipal Special Maintenance Fund Management Center for Elevator Maintenance Insurance" and "the owners entrust the property service enterprise to purchase Elevator Maintenance Insurance from the Municipal Special Maintenance Fund Management Center". Based on the entities and relationships, text key-value pairs can also be generated, laying a foundation for constructing a government knowledge graph.

[0040] Step S204: Construct a government affairs knowledge graph based on the entities, relationships, and text key-value pairs corresponding to each government affairs document. Specifically, organize the entities, relationships, and text key-value pairs in the form of a graph to more intuitively display the information in the government affairs document, facilitate quick query and acquisition of information, help improve the accuracy and efficiency of data query, and provide support for government affairs Q&A.

[0041] Step S205: Generate a query answer for the user query statement based on the type of the user query statement and the government affairs knowledge graph. Specifically, the Retrieval-Augmented Generation (RAG) is a natural language processing technology that combines query and generation, which can improve the accuracy and reliability of complex queries. According to the type of the user query statement, combined with the government affairs knowledge graph, use the retrieval-augmented technology to generate accurate and comprehensive query answers to meet the user's needs.

[0042] Step S206: Optimize the query answer for the user query statement to generate the target answer for the user query statement. Specifically, since the query answer may contain factors that affect security and accuracy, the query answer can be optimized to obtain a more accurate, reliable, and user-demand-compliant target answer.

[0043] The government affairs Q&A method based on knowledge graph and retrieval augmentation provided by the embodiments of the present invention ensures a comprehensive understanding of government affairs documents by extracting text information, multimodal information, and metadata, fuses different types of information to obtain document information, ensures that the context can be fully understood, extracts entities and relationships from the document information and generates text key-value pairs, can better understand the semantic information in the document, constructs a government affairs knowledge graph through entities, relationships, and text key-value pairs, enables logical reasoning and in-depth analysis based on the government affairs knowledge graph, and at the same time provides an efficient query mechanism, which can quickly locate relevant information in large-scale data. When performing government affairs Q&A, combined with the user's query type and the knowledge graph, the retrieval-augmented technology can more accurately match relevant answers, and through answer optimization, ensure that the generated answers are more accurate. Through the introduction of multimodal information processing and knowledge graph, it is possible to more comprehensively understand the document content and the user's query intention, significantly improve the accuracy of Q&A, fuse the knowledge graph and retrieval augmentation technology, reduce the "hallucination" phenomenon that may occur in large language models, improve the accuracy of answers, and further ensure the security, accuracy, and readability of answers through the optimization step, thus enhancing the user experience.

[0044] In this embodiment, a government affairs Q&A method based on knowledge graph and retrieval augmentation is provided, which can be used in the above-mentioned government affairs Q&A system based on knowledge graph and retrieval augmentation. The method specifically includes the following steps: Step S301: For each government document, perform information extraction on the government document to obtain text information, multimodal information, and metadata.

[0045] Specifically, the above-mentioned step S301 includes: Step S3011: Convert the government document into a standard format. Specifically, since there are various formats for government documents, to facilitate subsequent information extraction and processing, government documents in different formats are uniformly converted into a standard format, such as PDF format.

[0046] Step S3012: Use a parsing tool to extract the text information and metadata of the government document in the standard format. Specifically, use a PDF parsing tool, such as MinerU, Apache PDFBox, PyPDF2, etc., to read the pure text content in the PDF government document, that is, the text information, and at the same time extract the metadata, which are the key to understanding the core content and background of the document.

[0047] Step S3013: Use a computer vision model to perform non-text recognition on the government document in the standard format to obtain non-text content. Specifically, use a computer vision model to recognize non-pure text content such as pictures, tables, and formulas in the government document to obtain non-text content. First, perform preprocessing operations such as layout recognition, paragraph sorting, removal of useless blocks, and filtering of useless layouts on the government document. Then, perform corresponding processing on different types of non-text content. For example, perform coordinate positioning (i.e., determine its context location information), content extraction and storage on pictures, and table merging and storage on tables.

[0048] Step S3014: Perform preprocessing on the non-text content to obtain multimodal information. Specifically, store the non-pure text content in a structured data format such as JSON or Excel to obtain multimodal information containing rich non-pure text content.

[0049] Step S302: Integrate the text information, multimodal information, and metadata to obtain the document information of the government document.

[0050] Specifically, the above-mentioned step S302 includes: Step S3021: Perform preprocessing on the text information and multimodal information to obtain the preprocessed text information and multimodal information. Specifically, perform preprocessing operations such as removing duplicate information and processing OCR (Optical Character Recognition) garbled codes on the text information and multimodal information to improve data availability.

[0051] Step S3022: Use a large language model to convert metadata, preprocessed text information, and multimodal information into vectors respectively. Specifically, a large language model (LLM) is an artificial intelligence technology based on deep learning, which is specifically used to process and generate natural language text. By training on a large-scale dataset, the large language model learns the complex structures and patterns of language and can understand and generate grammatically correct and semantically rich text. By using a large language model, the text information, multimodal information, and metadata in government documents are converted into vectors, facilitating subsequent processing by the model.

[0052] Step S3023: Use a fusion layer to perform feature fusion on multiple vectors to obtain a fused vector. Specifically, use the fusion layer in the large language model. Based on its semantic understanding ability and correlation mining ability obtained through pre-training, analyze the vectors of different types of information, and finally fuse this information in a logically coherent manner, so that the fused vector contains more comprehensive government information, reflecting the correlation between various types of information, which helps to more comprehensively understand the content of government documents.

[0053] Step S3024: Decode the fused vector to obtain the document information of the government document. Specifically, use a decoding algorithm to restore the fused vector to understandable and usable document information, obtain the document information of the integrated government document, realize the fusion of multi-source information, and improve the accuracy and reliability of government documents compared with only considering single text information. Optionally, the document information can be converted into a lightweight Markdown format for subsequent processing.

[0054] In some alternative embodiments, Figure 3 is a schematic flowchart of a document information processing module according to an embodiment of the present invention. As Figure 3 shown, the document information processing module includes a document preprocessing layer, a multimodal data processing layer, and a fusion layer. The government document is converted into a standard format through the document preprocessing layer, and then text information and metadata are extracted. Through the multimodal data processing layer, multimodal data recognition is performed on the government document in standard format, such as image recognition, table recognition, formula detection, and text block detection, so as to extract multimodal information and uniformly convert it into formats such as JSOM and Excel. Through the fusion layer, the text information and metadata extracted by the document preprocessing layer and the multimodal information extracted by the multimodal data processing layer are processed and finally fused to obtain document information.

[0055] Step S303: Perform entity relationship extraction based on the document information of the government document to obtain the entities, relationships, and text key-value pairs corresponding to the government document.

[0056] Specifically, the above Step S303 includes: Step S3031: Cut the document information of the government affairs document into pieces to obtain multiple text pieces. Specifically, as shown in the following formula (1), the document information is divided into smaller text pieces that are convenient for processing, improving the efficiency and accuracy in entity relationship extraction.

[0057] (1) Where, represents the set of text pieces of the government affairs document; represents the document information; represents the segmentation module; represents the index value of the text piece, usually represented by a hash value; represents the th text piece.

[0058] Step S3032: For each text piece, use a large language model to identify entities and relationships from the text piece. Specifically, utilize the natural language processing ability of the large language model to identify entities such as names, times, locations, events, etc. from the text piece, as well as the relationships between them, providing the basic data for subsequent generation of text key-value pairs and construction of the knowledge graph.

[0059] Step S3033: For each entity, use the entity as the key and the relevant content of the entity in the text piece as the value to generate the text key-value pair corresponding to the entity. Specifically, use the identified entity as the index key and the text description related to the entity in the text piece as the value to generate the text key-value pair. For example, assume the text piece is "For more than 3 owners in the same unit, as well as the owners' committee or the property service enterprise entrusted by it, they can all submit an application to the Municipal Property Special Maintenance Fund Management Center to purchase elevator maintenance insurance with the funds from the value-added income pooling account". Use the entity "owner" as the key and the relevant content "For more than 3 owners in the same unit, as well as the owners' committee or the property service enterprise entrusted by it, they can all submit an application to the Municipal Property Special Maintenance Fund Management Center to purchase elevator maintenance insurance with the funds from the value-added income pooling account" as the value to generate the text key-value pair. By generating the text key-value pair corresponding to the entity, the entity information is structured, facilitating the subsequent construction of the knowledge graph and information query.

[0060] Step S3034: For each relationship, use the relationship and the entities constituting the relationship as the key and the relevant content of the relationship in the text piece as the value to generate the text key-value pair corresponding to the relationship. Specifically, use the relationship and the entities constituting the relationship as the index key and the description of the relationship in the text piece as the value to generate the corresponding text key-value pair, which helps to accurately present the relationship between entities.

[0061] Step S304: Construct a government affairs knowledge graph based on the entities, relationships, and text key-value pairs corresponding to each government affairs document.

[0062] Specifically, the above Step S304 includes: Step S3041: Preprocess the text key-value pairs based on the text key-value pairs corresponding to each government affairs document to obtain target key-value pairs. Specifically, all the text key-value pairs corresponding to the government affairs documents can be imported into an open-source graph database to complete the construction of the knowledge graph in this graph database. After the import, first perform preprocessing such as deduplication, denoising, and conversion to a standard format on the text key-value pairs to improve the quality of the text key-value pairs and obtain the corresponding target key-value pairs.

[0063] Step S3042: Create nodes based on the entities in each government affairs document and create edges based on the relationships in each government affairs document. Specifically, each entity in the government affairs document is used as a node, and each relationship is used as an edge.

[0064] Step S3043: Connect the nodes and edges based on the target key-value pairs corresponding to each government affairs document, and add attributes to the nodes and edges respectively to generate a government affairs knowledge graph. The attribute refers to the value corresponding to the node or the value corresponding to the edge. Specifically, according to the key-value correspondence in the target key-value pairs, the relationships between entities can be obtained, so as to connect the entities with relationships, that is, connect the nodes and edges. At the same time, if the target key-value pair is generated based on an entity, the value in the key-value pair is used as the attribute of the node corresponding to the entity. If the target key-value pair is generated based on a relationship, the value in the key-value pair is used as the attribute of the edge corresponding to the relationship. Based on the above process, a complete government affairs knowledge graph is generated, which contains rich entities, relationships, and attributes, can intuitively display the relationships between government affairs information, and provides strong support for government affairs question answering.

[0065] In some alternative embodiments, the government affairs knowledge graph can be represented by the following formula (2).

[0066] (2) Wherein, represents the government affairs knowledge graph; represents the relationship index generation large language model; represents the entity; represents the relationship; represents the entity relationship extraction large language model; represents the text chunking.

[0067] In some alternative embodiments, Figure 4 is a schematic diagram of the government affairs knowledge graph according to the embodiment of the present invention, as shown in Figure 4As shown, the government affairs knowledge graph uses multiple entities as nodes and relationships as edges, connecting nodes and edges based on the relationships between entities, and showing the association relationships between government affairs information. Optionally, Figure 4 The node attributes, the relationships corresponding to the edges, and the attributes are not shown.

[0068] Step S305: For each node in the government affairs knowledge graph, construct a vector corresponding to the node based on the node attributes of the node, the edges directly connected to the node, and the node and its attributes. Specifically, comprehensively consider the attributes of the node, the edges directly connected to it, and the node and its attributes, and convert this information into a vector corresponding to the node. Optionally, the dimension and calculation method of the vector are determined according to specific requirements. Generate corresponding vectors for each node, so that the information in the knowledge graph can be stored and processed in vector form, saving storage resources and facilitating query.

[0069] Step S306: Construct a vector database based on the vectors corresponding to all nodes. Specifically, construct a vector database based on the vectors generated by all nodes in the government affairs knowledge graph for quick query.

[0070] Step S307: Based on the type of the user query statement and the government affairs knowledge graph, use retrieval enhancement technology to generate a query answer for the user query statement.

[0071] Specifically, the above step S307 includes: Step S3071: When the type of the user query statement is a concept type, use a large language model to identify the query entity from the user query statement. Specifically, the query statements input by users during querying generally belong to specific concept type queries or general process type queries. For different query types, the process of querying and generating answers is different. If the user query statement is of the concept type, such as "What is special deduction?", corresponding query answers are generated through steps S3071 to S3074. More specifically, use a large language model to identify the query entity in the user query statement. For example, the query entity in "What is special deduction?" is "special deduction".

[0072] Step S3072: Convert the query entity into a first query vector, and determine at least one first target vector in the vector database whose similarity to the first query vector is greater than a preset similarity threshold. Specifically, convert the identified query entity into vector form, that is, the first query vector. By calculating the similarity between vectors, vectors with similarity greater than the preset threshold are selected as the first target vectors from the vector database. By combining knowledge graph query and vector database query, the efficiency and accuracy of government affairs Q&A are improved.

[0073] Step S3073: Based on at least one first target vector, determine the corresponding first target node, first target edge, and first adjacent node in the government affairs knowledge graph. The first target edge refers to the edge directly connected to the first target node, and the first adjacent node refers to the other node directly connected to the first target edge. Specifically, finding the first target node corresponding to the first target vector, the first target edge directly connected to the first target node, and the other adjacent nodes directly connected to the first target edge (excluding the first target node) in the government affairs knowledge graph can not only accurately query the detailed information related to a specific query entity, but also obtain more extensive relevant knowledge, and thus be able to provide a comprehensive answer for the user to meet the user's needs.

[0074] Step S3074: Based on the attributes of the first target node, the attributes of the first adjacent node, the first target edge, and the attributes of the first target edge, use a large language model to generate the query answer for the user's query statement. Specifically, input the obtained first target node, first adjacent node, first target edge, and their attribute information into the large language model. As shown in the following formula (3), the large language model generates a query answer in natural language form based on this information to meet the user's need for specific concept information.

[0075] (3) Where, represents the query answer; represents the large language model; represents the user's query statement; represents the query engine; represents the government affairs knowledge graph.

[0076] Step S3075: When the type of the user's query statement is a process type, use a large language model to identify the query entity, query relationship, and abstract concept from the user's query statement. Specifically, if the user's query statement is of the process type, such as "How to fill in the special additional deductions for individual income tax?", through steps S3075 to S3079, generate the corresponding query answer. More specifically, use a large language model to identify the query entity in the user's query statement, such as the query entity "special additional deductions for individual income tax" and the query relationship "fill in" in "How to fill in the special additional deductions for individual income tax?". Among them, the abstract concept is the word with specific meaning in the user's query statement except for the entity and the relationship, such as specific rules and other elements.

[0077] Step S3076: Convert the query entity, query relationship, and abstract concept into second query vectors respectively, and determine at least one second target vector in the vector database whose similarity to each second query vector is greater than a preset similarity threshold. Specifically, convert the identified query entity, query relationship, and abstract concept into vectors, that is, second query vectors. For each second query vector, filter out the vectors in the vector database whose similarity is greater than the preset similarity threshold as its corresponding second target vector, providing clues for obtaining relevant information from the knowledge graph.

[0078] Step S3077: Based on at least one second target vector, determine the corresponding second target nodes, second target edges, and second adjacent nodes in the government affairs knowledge graph. The second target edges include the edges directly connected to the second target nodes and the edges whose edge attributes are consistent with the query relationship. The second adjacent nodes refer to the other nodes directly connected to the second target edges. Specifically, refer to Step S3073, which will not be elaborated here.

[0079] Step S3078: Use the chain of thought technique to process based on the attributes of the second target nodes, the attributes of the second adjacent nodes, the second target edges, and the attributes of the second target edges to obtain the query result. Specifically, use the chain of thought technique to sort out the logical process of solving the process problem proposed by the user, such as the steps of filling in the special additional deductions for individual income tax, according to the information obtained from the government affairs knowledge graph, to obtain the query result.

[0080] Step S3079: Use the large language model to integrate based on the query result to generate the query answer to the user's query statement. Specifically, input the query result obtained by the chain of thought technique into the large language model, and the large language model integrates it into a smooth and easy-to-understand natural language text as the query answer.

[0081] Step S308: Optimize the query answer to the user's query statement to generate the target answer to the user's query statement.

[0082] Specifically, the above Step S308 includes: Step S3081, perform rule filtering and model filtering on the query answer to obtain the initial answer. Specifically, use length rules, sensitive word libraries, etc. for rule filtering. For example, set a reasonable length range for the answer. If the query answer is too short, it may not be able to fully answer the question, and if it is too long, it may contain redundant information. Process the query answers that do not meet the length requirements. At the same time, check the query answer through the sensitive word library to remove content containing sensitive words and other illegal information, such as words related to personal privacy, confidential information, or inappropriate remarks. Then, perform model filtering, that is, perform question relevance filtering and answer accuracy filtering, and process incomplete, fabricated, or ambiguous query answers. More specifically, for query answers that do not meet rule filtering and model filtering, manual intervention will be carried out or retrieval enhancement techniques will be used to re-query relevant information and generate query answers. After rule filtering and model filtering, a relatively more compliant and accurate initial answer is obtained, providing a basis for subsequent scoring and further optimization.

[0083] Step S3082, use a large language model to score the initial answer to obtain a quality score. Specifically, utilize the large language model and combine multi-dimensional scoring metrics such as relevance, accuracy, and expression clarity to comprehensively evaluate the initial answer. The large language model will analyze the relevance between the initial answer and the user's query statement, for example, determine whether the initial answer accurately answers the user's query statement; check the accuracy of the initial answer to ensure there are no factual errors; evaluate the expression clarity of the initial answer to judge whether it is logically coherent and easy to understand. Based on the evaluation results of these dimensions, a score that can reflect the quality of the initial answer is obtained.

[0084] Step S3083, when the quality score is greater than the preset score threshold, use the initial answer as the target answer. Specifically, compare the quality score of the initial answer with the preset score threshold. If the quality score is greater than the preset score threshold, it means that the initial answer performs well in terms of relevance, accuracy, and expression clarity and can meet the user's needs. Return the initial answer as the target answer to accurately meet the user's query requirements.

[0085] Step S3084, when the quality score is not greater than the preset score threshold, the initial answer is used as the new query answer and returned to the step of obtaining the initial answer by performing rule filtering and model filtering on the query answer until the quality score is not less than the preset score threshold, and the initial answer is used as the target answer. Specifically, when the quality score is not greater than the preset score threshold, the current initial answer is re-used as the new query answer, and rule filtering, model filtering, and scoring are performed again, and optimization is performed again to improve the answer quality, and it is compared with the preset score threshold again, and this loop process is continued until the quality score is not less than the preset score threshold. After multiple iterations of optimization, the target answer that meets the quality requirements is finally obtained, improving the overall quality of government affairs Q&A.

[0086] In some alternative embodiments, Figure 5 is a schematic flowchart of the data quality control module according to an embodiment of the present invention. As Figure 5 shown, the data quality control module includes a rule filtering layer, a model filtering layer, and a model scoring layer. The query answer is processed by the rule filtering layer for length rules and sensitive word rules. Optionally, the sensitive word library can be expanded according to actual needs. The answer filtered by the rule filtering layer is further filtered for question relevance and answer accuracy by the model filtering layer to obtain the initial answer. The initial answer is evaluated by the model scoring layer in multiple dimensions to obtain the quality score. By comparing the quality score with the preset score threshold, if it is not less than the threshold, it is output, and if it is less than the threshold, it is returned to the rule filtering layer for re-filtering.

[0087] The government affairs Q&A method based on knowledge graph and retrieval enhancement provided by the embodiments of the present invention ensures a comprehensive understanding of government affairs documents by extracting text information, multi-modal information, and metadata, fuses different types of information together to obtain document information, ensures that the context can be fully understood, extracts entities and relationships from the document information, and generates text key-value pairs, which can better understand the semantic information in the document. Through entities, relationships, and text key-value pairs, a government affairs knowledge graph is constructed, enabling logical reasoning and in-depth analysis based on the government affairs knowledge graph. At the same time, an efficient query mechanism is provided, which can quickly locate relevant information in large-scale data. When performing government affairs Q&A, combined with the user's query type and the knowledge graph, the retrieval enhancement technology can be used to more accurately match relevant answers, and through answer optimization, it is ensured that the generated answers are more accurate. Through multi-modal information processing and the introduction of the knowledge graph, the content of the document and the user's query intention can be more comprehensively understood, significantly improving the accuracy of Q&A. By integrating the knowledge graph and retrieval enhancement technology, the "hallucination" phenomenon that may occur in large language models is reduced, the accuracy of the answer is improved, and through the optimization steps, the security, accuracy, and readability of the answer are further ensured, enhancing the user experience.

[0088] An embodiment of the present invention further provides a computer device. Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 6 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 In

[0089] FIG.

[0090] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device may be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0091] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0092] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.

[0093] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0094] An embodiment of the present invention further provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading through a network and originally stored in a remote storage medium or a non-transitory machine-readable storage medium and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiment is implemented.

[0095] A part of the present invention can be applied as a computer program product, such as computer program instructions. When executed by a computer, through the operation of the computer, the method and / or technical solution according to the present invention can be called or provided. Those skilled in the art should be able to understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways for computer program instructions to be executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.

[0096] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A government affairs Q&A method based on a knowledge graph and retrieval enhancement, characterized in that, The method includes: For each government document, extract information from the government document to obtain text information, multimodal information, and metadata; Fuse the text information, the multimodal information, and the metadata to obtain the document information of the government document; Perform entity relationship extraction based on the document information of the government document to obtain the entities, relationships, and text key-value pairs corresponding to the government document; Construct a government knowledge graph based on the entities, relationships, and text key-value pairs corresponding to each government document; Based on the type of the user query statement and the government knowledge graph, use retrieval enhancement technology to generate the query answer for the user query statement; Optimize the query answer for the user query statement to generate the target answer for the user query statement; Among them, the optimizing the query answer for the user query statement to generate the target answer for the user query statement includes: Perform rule filtering and model filtering on the query answer to obtain an initial answer; Use a large language model to score the initial answer to obtain a quality score; In the case where the quality score is greater than a preset score threshold, use the initial answer as the target answer; or, In the case where the quality score is not greater than the preset score threshold, use the initial answer as a new query answer and return it to the step of performing rule filtering and model filtering on the query answer to obtain an initial answer until the quality score is not less than the preset score threshold, and use the initial answer as the target answer.

2. The method according to claim 1, characterized in that, The performing entity relationship extraction based on the document information of the government document to obtain the entities, relationships, and text key-value pairs corresponding to the government document includes: Chunk the document information of the government document to obtain multiple text chunks; For each text chunk, use a large language model to identify entities and relationships from the text chunk; For each entity, use the entity as the key and the relevant content of the entity in the text chunk as the value to generate the text key-value pair corresponding to the entity; For each relationship, use the relationship and the entities constituting the relationship as the key and the relevant content of the relationship in the text chunk as the value to generate the text key-value pair corresponding to the relationship.

3. The method according to claim 2, characterized in that, The constructing a government knowledge graph based on the entities, relationships, and text key-value pairs corresponding to each government document includes: Preprocess the text key-value pairs based on the text key-value pairs corresponding to each government document to obtain target key-value pairs; Create nodes based on the entities in each government document and create edges based on the relationships in each government document; Based on the target key-value pairs corresponding to each government document, connect the nodes and edges, and add attributes to the nodes and edges respectively to generate the government knowledge graph, where the attribute refers to the value corresponding to the node or the value corresponding to the edge.

4. The method according to claim 3, characterized in that, Before generating the query answer for the user query statement based on the type of the user query statement and the government knowledge graph using retrieval enhancement technology, the method further includes: For each node in the government knowledge graph, construct a vector corresponding to the node based on the node attribute of the node, the edges directly connected to the node, and the node and its attributes; Construct a vector database based on the vectors corresponding to all nodes.

5. The method according to claim 4, characterized in that, Generating a query answer for the user query statement by using retrieval enhancement technology based on the type of the user query statement and the government affairs knowledge graph includes: When the type of the user query statement is a concept type, use a large language model to identify a query entity from the user query statement; Convert the query entity into a first query vector, and determine at least one first target vector in the vector database whose similarity to the first query vector is greater than a preset similarity threshold; Based on at least one first target vector, determine corresponding first target nodes, first target edges, and first adjacent nodes in the government affairs knowledge graph, where the first target edge refers to an edge directly connected to the first target node, and the first adjacent node refers to other nodes directly connected to the first target edge; Generate a query answer for the user query statement by using a large language model based on the attributes of the first target node, the attributes of the first adjacent node, the first target edge, and the attributes of the first target edge.

6. The method according to claim 4, characterized in that, Generating a query answer for the user query statement by using retrieval enhancement technology based on the type of the user query statement and the government affairs knowledge graph further includes: When the type of the user query statement is a process type, use a large language model to identify a query entity, a query relationship, and an abstract concept from the user query statement; Convert the query entity, the query relationship, and the abstract concept into second query vectors respectively, and determine at least one second target vector in the vector database whose similarity to each second query vector is greater than a preset similarity threshold; Based on at least one second target vector, determine corresponding second target nodes, second target edges, and second adjacent nodes in the government affairs knowledge graph, where the second target edge includes an edge directly connected to the second target node and an edge whose edge attribute is consistent with the query relationship, and the second adjacent node refers to other nodes directly connected to the second target edge; Adopt the chain of thought technology to process based on the attributes of the second target node, the attributes of the second adjacent node, the second target edge, and the attributes of the second target edge to obtain a query result; Use a large language model to integrate based on the query result to generate a query answer for the user query statement.

7. The method according to claim 1, wherein For each government affairs document, perform information extraction on the government affairs document to obtain text information, multimodal information, and metadata, including: Convert the government affairs document into a standard format; Use a parsing tool to extract the text information and metadata of the government affairs document in the standard format; Use a computer vision model to perform non-text recognition on the government affairs document in the standard format to obtain non-text content; Preprocess the non-text content to obtain the multimodal information.

8. The method according to claim 1, wherein Fusing the text information, the multimodal information, and the metadata to obtain the document information of the government affairs document includes: Preprocess the text information and the multimodal information to obtain preprocessed text information and multimodal information; Using a large language model, convert the metadata, preprocessed text information, and multimodal information into vectors respectively; Use a fusion layer to perform feature fusion on multiple vectors to obtain a fused vector; Decode the fused vector to obtain the document information of the government affairs document.

9. A government affairs question-answering system based on a knowledge graph and retrieval enhancement, wherein The system includes: A document information processing module for, for each government affairs document, extracting information from the government affairs document to obtain text information, multimodal information, and metadata, and fusing the text information, the multimodal information, and the metadata to obtain the document information of the government affairs document; A knowledge graph construction module for performing entity relationship extraction based on the document information of the government affairs document to obtain the entities, relationships, and text key-value pairs corresponding to the government affairs document, and constructing a government affairs knowledge graph based on the entities, relationships, and text key-value pairs corresponding to each government affairs document; A content query and generation module for generating a query answer to the user query statement based on the type of the user query statement and the government affairs knowledge graph by using retrieval enhancement technology; A data quality control module for optimizing the query answer to the user query statement to generate a target answer to the user query statement; Among them, the data quality control module is specifically used for: Performing rule filtering and model filtering on the query answer to obtain an initial answer; Using a large language model to score the initial answer to obtain a quality score; In the case where the quality score is greater than a preset score threshold, taking the initial answer as the target answer; or, In the case where the quality score is not greater than the preset score threshold, taking the initial answer as a new query answer and returning it to the step of performing rule filtering and model filtering on the query answer to obtain an initial answer until the quality score is not less than the preset score threshold, and taking the initial answer as the target answer.

Citation Information

Patent Citations

  • Government affair service field multi-strategy fusion dialogue method based on knowledge graph

    CN116628172A

  • Knowledge graph auxiliary question and answer method suitable for electric power field

    CN119322853A

  • Knowledge retrieval generation-based open domain question and answer method and system

    CN119577064A

  • System and method for extracting project information based on multi-modal large model

    CN119739907A

  • Knowledge graph question-answer method and apparatus based on deep learning technology, and device

    WO2021139283A1

Cited By

  • Government affair information recommendation method and device based on knowledge graph and multi-mode fusion

    CN120353924A

  • Modularized knowledge graph and retrieval enhanced large model fusion interaction method and system oriented to financial branch mechanism

    CN120448510A

  • A modular knowledge graph and retrieval-enhanced large model fusion and interaction method and system for financial branches

    CN120448510B

  • Retrieval enhancement generation method based on document knowledge base and knowledge graph

    CN120780849A

  • Business planning intelligent decomposition method and device based on large model fine tuning and medium

    CN120875272A