Government affair question and answer method and product based on knowledge enhancement large model

Through multi-stage semantic enhanced knowledge retrieval and networked query supplement, the problems of low efficiency and poor knowledge quality in the government Q&A system are solved, accurate response and service closed loop are achieved, and the overall performance of the government Q&A system is improved.

CN120256598AActive Publication Date: 2025-07-04SOUTHWEST JIAOTONG UNIV

Patent Information

Application Number
CN202510358669.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-04
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing government Q&A system is inefficient, cannot identify intention differences, lack dialogue context, affects the quality of retrieval knowledge, and is difficult to integrate multi-source policy information.

Method used

Adopt the government affairs question-and-answer method based on the knowledge-enhancing model, and supplement cross-domain knowledge through multi-stage semantic enhanced knowledge retrieval process and networked query, design intelligent problem diversion and structured feedback frameworks to achieve accurate response and service closed loop.

Benefits of technology

The manual evaluation efficiency was improved by 27.3%, the answer recall rate was 35.6%, the search quality was improved by 15% and 24.6%, and the professionalism and accuracy of multiple rounds of dialogue exceeded 30%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256598A_ABST
    Figure CN120256598A_ABST
Patent Text Reader

Abstract

The invention provides a government affair question-answering method and product based on a knowledge-enhanced large language model, and relates to the technical field of artificial intelligence. A current government affair question-answering system mainly adopts a manual reply mode and is low in efficiency, and a system based on RAG lacks a query and discrimination mechanism and cannot recognize intention differences and utilize dialogue contexts, so that the quality of retrieval knowledge is influenced. Aiming at the defects existing in the related technologies, the invention provides a government affair question answering method based on a knowledge enhancement large model, and particularly, aiming at the problem of low manual reply efficiency, a complete government affair hotline service framework containing intelligent question shunting and structured feedback is designed; aiming at the problem of poor knowledge retrieval quality of a traditional RAG system, a multi-stage semantic enhancement retrieval method is designed, cross-domain knowledge is supplemented through networking query, and accurate response and service closed loop of government affair consultation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a government affairs question answering method and product based on a knowledge-enhanced large model. Background Art

[0002] With the development of large language models, the field of intelligent question answering has undergone a major transformation. Intelligent question answering systems based on large language models have made great progress, especially in vertical fields such as finance and law. Intelligent question answering systems such as FinGPT and Chatlaw have promoted the practical application of technology and provided strong support for the intelligent processing of professional knowledge.

[0003] As an important bridge for communication between the government and citizens, the importance of the government service convenience hotline is becoming increasingly prominent. Facing the increasing volume of incoming calls for consultation, the existing government affairs question answering systems have certain limitations in dealing with a large number of similar questions, realizing real-time query and efficient transfer and handling, etc. Therefore, it is particularly urgent to build a professional government affairs question answering system to meet the complex and changeable question needs. Summary of the Invention

[0004] The present invention provides a government affairs question answering method and product based on a knowledge-enhanced large model to at least partially solve the above problems.

[0005] In a first aspect of the present invention, a government affairs question answering method based on a knowledge-enhanced large model is provided. The method includes: Receiving first text information in the form of natural language sent by a user; Based on the historical text information sent by the user and the first text information, obtaining a query statement corresponding to the first text information, where the query statement represents the true intention of the user; Retrieving the first N knowledge fragments from a government affairs knowledge base based on the query statement; Sorting the semantic matching degrees of the first N knowledge fragments and the query statement to obtain the first K knowledge fragments with the highest semantic matching degrees; Evaluating the availability of the first K knowledge fragments with the highest semantic matching degrees; Introducing the knowledge fragments with availability as external knowledge into a large language model; Obtaining an answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge.

[0006] Optionally, sorting the semantic matching degrees of the first N knowledge fragments and the query statement to obtain the first K knowledge fragments with the highest semantic matching degrees includes: For any one of the first K knowledge fragments with the highest semantic matching degrees: Input the query statement and the knowledge fragment as a whole into a large language model, and perform in-depth semantic interaction on the query and the document in combination with prompt words to output the semantic matching degree between the query statement and the knowledge fragment.

[0007] Optionally, the method further includes: In the case where none of the top K knowledge fragments with the highest semantic matching degree are available, trigger an online query tool to obtain supplementary knowledge; Obtain an answer text for the first text information based on the endogenous knowledge of the large language model and the supplementary knowledge.

[0008] Optionally, the method further includes: Perform intent recognition on the first text information through a large language model to determine the text classification corresponding to the first text information; In the case where the first text information belongs to the consultation category, retrieve external knowledge from the government affairs knowledge base based on the query statement and introduce it into the large language model, and obtain an answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge; In the case where the first text information belongs to the complaint category, help-seeking category, and praise category, summarize and record the first text information through a large language model, and generate an answer text for the first text information; In the case where the first text information belongs to other categories, generate an answer text for the first text information based on the large language model.

[0009] Optionally, the method further includes: Collect multi-source government affairs data; Perform data processing on the multi-source government affairs data to obtain processed government affairs data; Process the processed government affairs data based on a text encoder model to obtain vectorized data, and add the vectorized data to the government affairs knowledge base.

[0010] Optionally, the method further includes: Generate a service work order through a large language model based on the first text information and the answer text of the first text information, and send work order verification information based on the service work order; In the case where the service work order is verified, process the service work order through a text encoder model to obtain vectorized data, and supplement the vectorized data to the government affairs knowledge base.

[0011] A second aspect of the present invention provides a government affairs question-answering device based on a knowledge-enhanced large model, and the device includes: A receiving module, configured to receive the first text information in the form of natural language sent by a user; An intention analysis module, configured to obtain a query statement corresponding to the first text information based on the historical text information sent by the user and the first text information, where the query statement represents the true intention of the user; A retrieval module, configured to retrieve the top N knowledge fragments from a government affairs knowledge base based on the query statement; sort the semantic matching degrees between the top N knowledge fragments and the query statement to obtain the top K knowledge fragments with the highest semantic matching degrees; An evaluation module, configured to evaluate the availability of the top K knowledge fragments with the highest semantic matching degrees; An answer module, configured to introduce the knowledge fragments with availability as external knowledge into a large language model; obtain an answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge.

[0012] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the government affairs question-answering method based on a knowledge-enhanced large model as described in the first aspect of the present invention.

[0013] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the government affairs question-answering method based on a knowledge-enhanced large model as described in the first aspect of the present invention.

[0014] A fifth aspect of the present invention provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, it implements the steps in the government affairs question-answering method based on a knowledge-enhanced large model as described in the first aspect of the present invention.

[0015] The current government affairs question-answering system mainly adopts the method of manual reply, which is inefficient. And the RAG-based system lacks a query screening mechanism and cannot identify intention differences and utilize conversation context, affecting the quality of retrieved knowledge. In view of the defects existing in these related technologies, the embodiments of the present invention propose a government affairs question-answering method based on a knowledge-enhanced large model. Specifically, aiming at the problem of low efficiency of manual reply, a complete government affairs hotline service framework including intelligent question diversion and structured feedback is designed; aiming at the problem of poor quality of retrieved knowledge in traditional RAG systems, a multi-stage semantic enhancement retrieval method is designed, and cross-domain knowledge is supplemented through online query, realizing accurate response and service closed-loop for government affairs consultation.

[0016] The results of the comparative experiments show that the technical solution provided by the embodiment of the present invention has improved by 27.3% in terms of manual evaluation compared with the commercial dialogue model, and has improved by 35.6% in terms of answer recall rate compared with the fine-tuned model (GLM4-9b-chat); in terms of retrieval quality, compared with the naive RAG system, the query-knowledge relevance and knowledge support degree of the technical solution provided by the embodiment of the present invention have increased by 15% and 24.6% respectively; in terms of the overall performance of multi-round dialogue, the technical solution provided by the embodiment of the present invention has an average improvement of more than 30% in core indicators such as professionalism and accuracy compared with the naive RAG system (based on GLM4-9b-chat), the fine-tuned model (GLM4-9b-chat), and the commercial large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of the steps of the government affairs question-answering method based on the knowledge-enhanced large model provided by the present invention; Figure 2 is a schematic flowchart of the stage semantic enhancement-based knowledge retrieval process of the government affairs question-answering method based on the knowledge-enhanced large model provided by the present invention; Figure 3 is a schematic flowchart of the reply shunting mechanism process of the government affairs question-answering method based on the knowledge-enhanced large model provided by the present invention; Figure 4 is a schematic framework diagram of the government affairs question-answering system based on the knowledge-enhanced large model provided by the present invention; Figure 5 is a structural block diagram of the government affairs question-answering device based on the knowledge-enhanced large model provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] Compared with other professional fields, the application of large models in government affairs scenarios faces unique challenges: the polysemy of policy texts, the strong normativity of service processes, and the high frequency of knowledge updates make it difficult for existing general models to meet the requirements. Government service hotlines often involve the interpretation of policies that require multi-department collaboration and real-time updated livelihood information. Directly using large language models (LLMs) may lead to "hallucinations" during processing. Although question-and-answer systems based on knowledge graphs can achieve precise matching through structured knowledge bases, they face limitations such as high construction costs and slow dynamic updates. In contrast, the Retrieval-Augmented Generation (RAG) framework combines unstructured text retrieval and generation technologies, significantly expanding the information coverage while ensuring professionalism. A typical RAG system consists of three major modules: data processing, knowledge base construction and query, and answer generation, effectively enhancing the application ability of large models in vertical fields by introducing external professional knowledge. However, there are two key problems in the actual application of existing RAG systems: one is the lack of discrimination of user inputs, and the knowledge base is retrieved indiscriminately for each query; the other is that the query only depends on the current user input, neither effectively filtering redundant information that may affect the query accuracy nor easily missing key information in the conversation history. These problems result in less-than-ideal quality of the retrieved knowledge and sometimes even have a negative impact on answer generation.

[0021] Based on this, the present invention proposes a government affairs question-and-answer method based on a knowledge-enhanced large model. Facing problems such as low efficiency of manual responses and difficulties in integrating multi-source policy information in the government affairs field, this method innovatively applies large model technology to the government affairs field and proposes and designs a new government affairs question-and-answer system based on a knowledge-enhanced large model. Specifically, in the face of problems such as low-quality text backpropagation and omission of multi-round dialogue information in existing question-and-answer systems, the embodiments of the present invention introduce a multi-stage semantic-enhanced knowledge retrieval method in the knowledge base construction and query stage to improve the retrieval accuracy, and introduce online query as a means of cross-domain knowledge supplementation, thereby providing more accurate and comprehensive professional knowledge. In the answer generation stage, an intention recognition and question diversion model based on a large model is designed, which can automatically identify the question type - questions that require professional knowledge support will trigger the knowledge retrieval enhancement module, while questions that the model can handle independently will directly generate answers. This design not only reduces the "hallucinations" that may be brought by the introduction of external knowledge but also improves the response efficiency of the system. In addition, according to the requirements of the government service hotline scenario, we have also added a summary and feedback module, which structures the conversation content through intelligent information extraction, realizes a service closed-loop, and provides data support for continuous optimization.

[0022] Specifically, as Figure 1 shown, it shows a flowchart of the steps of a government affairs question-and-answer method based on a knowledge-enhanced large model provided by an embodiment of the present invention, and the method includes the following steps: S101. Receive the first text information in natural language sent by the user.

[0023] The government affairs Q&A method provided by the embodiments of the present invention is applied to a government affairs Q&A system, which can provide a user interface to receive the first text information in natural language sent by the user.

[0024] In the embodiments of the present invention, the first text information in natural language is a natural language text sent by the user, expressing the user's government affairs Q&A needs.

[0025] S102. Based on the historical text information sent by the user and the first text information, obtain a query statement corresponding to the first text information, where the query statement represents the true intention of the user.

[0026] S103. Retrieve the first N knowledge fragments from the government affairs knowledge base based on the query statement; sort the semantic matching degrees between the first N knowledge fragments and the query statement to obtain the first K knowledge fragments with the highest semantic matching degrees.

[0027] S104. Evaluate the availability of the first K knowledge fragments with the highest semantic matching degrees.

[0028] In the embodiments of the present invention, the above steps S102 to S104 are the multi-stage semantic enhanced knowledge retrieval processes provided by the embodiments of the present invention. Based on this process, problems such as low-quality text backpropagation and multi-round dialogue information omission in related technologies can be solved. Specifically, as Figure 2 shown, it shows a schematic flow diagram of the multi-stage semantic enhanced knowledge retrieval process in the embodiments of the present invention.

[0029] In the embodiments of the present invention, after receiving the current dialogue (i.e., the first text information) sent by the user, the historical text information sent by the user can be further considered to obtain a query statement corresponding to the first text information, specifically as shown in the historical dialogue summary retrieval stage in Figure 2 .

[0030] In the embodiments of the present invention, the historical text information refers to the historical dialogue sent by the current user before sending the current dialogue.

[0031] Traditional sparse text retrieval techniques based on word frequency statistics (such as BM25 and TF-IDF) mainly rely on lexical frequency for retrieval, but have weak fault tolerance when dealing with synonyms and near-synonyms; content-based dense retrieval techniques reduce information loss through complete Transformer inference, but the time cost is too high and not suitable for large-scale retrieval; while vector-based dense retrieval techniques optimize computational efficiency, but compressing documents into a single vector inevitably leads to information loss. Therefore, in order to balance query accuracy and time cost, the embodiments of the present invention propose a multi-stage semantic enhancement-based knowledge retrieval process for government affairs knowledge base retrieval, such as Figure 2 shown, which includes three core stages: historical dialogue summary retrieval stage, semantic re-ranking stage, and self-reflection decision-making stage.

[0032] Specifically, key entities can be extracted from historical dialogues, and the true intention of the user can be extracted from the current dialogue. Taking the key entities as supplements, the user's input can be rewritten based on the true intention of the user (such as Figure 2 in the historical dialogue summary retrieval stage), to obtain a query statement, and based on this query statement, the content in the knowledge base can be more accurately matched.

[0033] In the embodiments of the present invention, a rewritten large model can be constructed based on a large language model + prompt words to identify and remove redundant information in the user's input, extract core semantics and key intentions, and convert complex and colloquial expressions into standardized retrieval statements. Specifically, in terms of multi-round dialogue enhancement, by analyzing historical dialogues, the key information omitted in the user's input is restored, pronouns and references are parsed, and a coherent and complete query semantic representation is constructed, laying a solid foundation for accurate retrieval of the knowledge base in the follow-up, thereby improving the accuracy and relevance of the retrieval. This rewriting process can be represented by the following formula: (1) wherein, is the rewritten query statement, LLM represents the large language model, is the prompt word template dedicated to the rewriting task, is the first text information, H is the historical text information.

[0034] Furthermore, in the embodiments of the present invention, the query statement can be retrieved in the government affairs knowledge base.

[0035] In the embodiments of the present invention, the retrieval step is implemented based on step S103 (such as Figure 2(in the semantic rearrangement stage). In the embodiments of the present invention, a multi-stage semantic enhanced knowledge retrieval is adopted in the retrieval step. First, the query statement is encoded to obtain a query vector, and N semantic vectors are retrieved from the government affairs knowledge base based on the query vector. Specifically, first, the rewritten user query is vectorized to obtain the query q , and then the top N knowledge fragments (such as Figure 2 knowledge indexes 1 to N in the semantic rearrangement stage) are quickly retrieved from the government affairs knowledge base using vector-based dense retrieval technology. This process can be expressed by the following formula: (2) where D represents the knowledge fragments in the government affairs knowledge base, and Distance is the distance function.

[0036] In the embodiments of the present invention, a more complex semantic matching algorithm is further introduced to perform in-depth context analysis and fine evaluation on the preliminary retrieval results (the top N knowledge fragments), and screen out the top-K knowledge fragments with the most matching semantics and the richest information.

[0037] In the embodiments of the present invention, the semantic matching degrees of the top N knowledge fragments and the query statement are sorted to obtain the top K knowledge fragments with the highest semantic matching degrees, including: For any one of the top K knowledge fragments with the highest semantic matching degrees: The query statement and this knowledge fragment are taken as a whole and input into the large language model, and in combination with the prompt words, in-depth semantic interaction is performed on the query and the document, and the semantic matching degree of the query statement and the knowledge fragment is output.

[0038] In the embodiments of the present invention, a cross-encoder model can also be used to perform in-depth semantic interaction on the query and the document through the self-attention mechanism, and output the semantic matching degree of the query statement and the knowledge fragment.

[0039] Specifically, in the embodiments of the present invention, the screening of the top N knowledge fragments can be expressed by the following formula: (3) where It represents a cross-encoder model that takes the query q and the knowledge fragment d among the top N knowledge fragments as a whole input. Using the Transformer architecture, it conducts in-depth semantic interaction between the query and the document through the self-attention mechanism and outputs a score between 0 and 1, representing the semantic matching degree between the query and the document. The TopK() function means sorting the scores of all document-query pairs output by the CrossEncoder and selecting the top K documents with the highest scores as the final result. K is an adjustable hyperparameter, usually much smaller than the size of the candidate set N in the first stage.

[0040] In the embodiments of the present invention, the aforementioned semantic rearrangement stage realizes multi-dimensional screening of knowledge fragments. However, considering the capacity limitation of the government affairs knowledge base, its coverage still has certain limitations. For queries beyond the coverage of the knowledge base, even if the most relevant knowledge fragments are obtained through the first two stages, these retrieval results may still not fully support the response to the original query, and may even cause misleading outputs of the large language model, resulting in negative impacts. Therefore, in the embodiments of the present invention, an availability evaluation of the retrieved knowledge fragments is further proposed.

[0041] Specifically, as Figure 2 shown in the self-reflection decision-making stage in , in the embodiments of the present invention, after obtaining the retrieval results, a self-reflection evaluation module is constructed through prompt engineering, and this module determines the actual value of the knowledge fragment based on the large language model. The evaluation module assigns a binary label (4) where C represents the knowledge fragment retrieved from the knowledge base (i.e., the top K knowledge fragments with the highest semantic matching degree). Based on the evaluation results, in the embodiments of the present invention, the following decision-making mechanism is adopted for knowledge selection: (5) S105, introduce the available knowledge fragments as external knowledge into the large language model; obtain the answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge.

[0042] In the embodiments of the present invention, when it is determined that the finally retrieved top K knowledge fragments with the highest semantic matching degree are available, the knowledge fragment is injected into the large language model as external knowledge, so that the large language model generates the answer text for the first text information to reply to the current conversation sent by the user.

[0043] In an optional implementation manner, the method further includes: In the case where none of the top K knowledge fragments with the highest semantic matching degree are available, trigger the network query tool to obtain supplementary knowledge; obtain the answer text for the first text information based on the endogenous knowledge of the large language model and the supplementary knowledge.

[0044] In the embodiments of the present invention, when the matching degree of the internal knowledge base does not meet the threshold requirement, the external tool call interface Tool() can also be automatically triggered to realize the dynamic supplement of cross-domain knowledge, thereby improving the overall response quality of the system.

[0045] In an optional implementation manner, the method further includes: S1, perform intent recognition on the first text information through a large language model to determine the text classification corresponding to the first text information.

[0046] S2, in the case where the first text information belongs to the consultation category, retrieve external knowledge from the government affairs knowledge base based on the query statement and introduce it into the large language model, and obtain the answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge; in the case where the first text information belongs to the complaint category, help-seeking category, and praise category, summarize and record the first text information based on the large language model, and generate the answer text for the first text information; in the case where the first text information belongs to other categories, generate the answer text for the first text information based on the large language model.

[0047] In the embodiments of the present invention, classify the user's first text information through an intent recognition model based on prompt engineering and a large language model: (6) where respectively correspond to consultation category, complaint and suggestion category, and other category questions. For complaint and suggestion category and other category questions, the large language model will answer according to the set role and workflow based on its own capabilities.

[0048] For consultation category questions, the system determines whether cross-domain knowledge enhancement is required through a multi-stage semantic enhanced knowledge retrieval method: →{0,1}, where =1 indicates that knowledge support is required. When knowledge support is required, trigger the network query tool Tool to obtain supplementary knowledge: (7) Finally, the system integrates the original query q and the final knowledge knowledge and inputs them into the large language model to generate an answer: (8) Specifically, such as Figure 3As shown in the figure, it shows a schematic flowchart of the reply shunting mechanism provided by the embodiments of the present invention. In the embodiments of the present invention, an intent recognition model based on a large language model can classify the first text information sent by the user. When the category of the first text information is a consultation category, knowledge fragments can be obtained based on a multi-stage semantic enhancement-based knowledge retrieval process, and it is determined whether it is necessary to perform an online query to obtain supplementary knowledge based on the availability of the knowledge fragments. Thus, the retrieved knowledge fragments are used as external knowledge, or the supplementary knowledge retrieved through the online query is injected into the large language model for knowledge-based answering to generate a reply. When the category of the first text information is other categories, the large language model can autonomously respond to the first text information based on prompt words to generate a reply. When the category of the first text information is a complaint category, a help-seeking category, and a praise category, the large language model summarizes and records the first text information based on prompt words to generate a reply.

[0049] The embodiments of the present invention propose to classify user conversations and perform shunting processing. This innovative shunting mechanism and dynamic knowledge supplementation strategy not only avoid redundant retrieval of simple questions but also make up for the deficiencies of the knowledge base through online queries, significantly improving the answer quality and knowledge coverage of the system.

[0050] In an alternative embodiment, the method further includes: S11, collecting multi-source government affairs data.

[0051] S12, performing data processing on the multi-source government affairs data to obtain processed government affairs data.

[0052] S13, processing the processed government affairs data based on a text encoder model to obtain vectorized data, and adding the vectorized data to the government affairs knowledge base.

[0053] In the embodiments of the present invention, data such as 12345 government service platform Q&A data, information of government agencies at all levels in provinces and cities, and local laws and regulations can be collected as multi-source government affairs data.

[0054] Furthermore, in the embodiments of the present invention, data processing can be performed on the multi-source government affairs data. Specifically, the multi-source government affairs data can be parsed and sliced to obtain processed government affairs data, which can include: Q&A blocks and text blocks. These Q&A blocks and text blocks are further processed to obtain vectorized data and further obtain a government affairs knowledge base.

[0055] In the embodiments of the present invention, this government affairs knowledge base can be used as an offline database to provide external knowledge for subsequent retrieval.

[0056] In an alternative embodiment, the method further includes: S21, generate a service work order based on the first text information and the answer text of the first text information by a large language model, and send work order verification information based on the service work order.

[0057] S22, in the case where the service work order is verified and passed, process the service work order through a text encoder model to obtain vectorized data, and supplement the vectorized data to the government affairs knowledge base.

[0058] In the embodiments of the present invention, a large language model can summarize the answer text obtained based on the prompt engineering for the first text information of the current user to generate a service work order. The service work order may include: an event title, an event content, a processing result, etc. Among them, the event title and the event content may be summarized from the first text information, and the processing result may be summarized from the answer text.

[0059] In the embodiments of the present invention, after generating the service work order, work order verification information can be sent to the corresponding verification terminal based on the service work order, and the verification terminal can confirm whether there is an error in the service work order. In the case where the service work order is verified and passed, the content corresponding to the service work order can be supplemented to the government affairs knowledge base as a new knowledge fragment. Specifically, the service work order can be processed through a text encoder model to obtain vectorized data, and the vectorized data can be supplemented to the government affairs knowledge base.

[0060] In the embodiments of the present invention, by summarizing the processed government affairs question-and-answer process through a large language model to obtain new knowledge fragments and supplementing them to the government affairs knowledge base, the government affairs knowledge base can be continuously supplemented.

[0061] In the embodiments of the present invention, aiming at problems such as low efficiency of manual replies and difficulty in integrating multi-source policy information in the government affairs field, the large model technology is innovatively applied to the government affairs field, and a new government affairs question-and-answer method based on a knowledge-enhanced large model is proposed and designed. In the embodiments of the present invention, aiming at problems such as low-quality text backpropagation and omission of multi-round dialogue information in the existing question-and-answer systems, a multi-stage semantic enhancement-based knowledge retrieval method is also designed. And an automatic evaluation framework for the question-and-answer system based on the large model and prompt engineering is constructed to achieve a comprehensive evaluation of the system performance.

[0062] Based on the same inventive concept, the present invention also provides a government affairs question-and-answer system based on a knowledge-enhanced large model, as Figure 4 shown, which shows a framework schematic diagram of the government affairs question-and-answer system based on a knowledge-enhanced large model. The government affairs question-and-answer system based on a knowledge-enhanced large model includes: a data processing module, a knowledge base construction and query module, an answer generation module, and a summary feedback module.

[0063] Among them, the data processing module is used to collect data such as question-and-answer data from the 12345 government affairs platform, information of government agencies at all levels, and local laws and regulations as multi-source government affairs data, parse and slice the multi-source government affairs data to obtain question-and-answer blocks and text blocks. Further process these question-and-answer blocks and text blocks to obtain vectorized data, and further obtain a government affairs knowledge base. The government affairs knowledge base is used as an offline database to provide external knowledge for subsequent retrieval.

[0064] The knowledge base construction and query module is used to introduce a multi-stage semantic-enhanced knowledge retrieval process and introduce online query as a cross-domain knowledge supplement means, so as to provide more accurate and comprehensive professional knowledge. The specific process is similar to the above steps S102~S104 and will not be elaborated here.

[0065] The answer generation module is used to receive the first text information sent by the user and generate answer texts according to the corresponding reply strategies based on question classification. The specific process is similar to the above steps S1~S2 and will not be elaborated here.

[0066] The summary and feedback module is used to summarize the answer text generated by the current large language model in combination with the corresponding first text information to obtain a service work order, and supplement the government affairs knowledge base based on the verified service work order. The specific process is similar to the above steps S21~S22 and will not be elaborated here.

[0067] In the embodiment of the present invention, a comparative verification embodiment is also provided. This embodiment conducts experimental verification on the Ubuntu 22.04 Linux operating system, and the hardware configuration is as follows: Processor: 12th Gen Intel(R) Core(TM) i7-12700; Graphics processor: NVIDIA GeForce RTX 4090 (24GB VRAM); System memory: 1024GB. Based on the comprehensive consideration of the performance of the model in Chinese tasks, inference efficiency, and security, the GLM-9b-chat developed by Zhipu AI is selected as the basic language model in the embodiment of the present invention.

[0068] In the embodiment of the present invention, a test data set for government affairs knowledge Q&A is constructed according to the documents listed in Table 1.

[0069] Table 1 Statistical information of knowledge base documents

[0070] Among them, it includes two categories: single-round Q&A conversations and multi-round Q&A conversations. Among them, the single-round conversation data set Q&A is divided into unfiled data and knowledge base data, and the multi-round Q&A includes a multi-round conversation data set rewritten from the 12345 government affairs platform Q&A data. The specific test data set information is shown in Table 2.

[0071] Table 2 Dataset Problem Statistics

[0072] Traditional RAG evaluation methods face a triple contradiction of semantic quantization bias, evaluation cost, and result reliability. Rule-based metrics (such as ROUGE-L) rely on text overlap to quantify performance and are difficult to capture semantic consistency; manual annotation frameworks (such as TruLens) are highly reliable but costly; pure LLM automated evaluation (such as RAGAS) is low-cost but significantly different from manual evaluation. Therefore, the embodiments of the present invention combine the authority of manual annotation with the scalability of LLM evaluation, and use the Chain of Thought (CoT) method to decompose the evaluation object into independent units for fine-grained analysis, and calculate the comprehensive performance through the average unit score, constructing a comprehensive and systematic question-answering system evaluation method.

[0073] The evaluation framework starts from two core dimensions of retrieval quality and generation quality, and adopts a semantic association quantization scoring mechanism of 0-10 points. The evaluation objects are defined as: Q: query; C: retrieved knowledge; A: system answer; T: true answer.

[0074] The retrieval quality evaluation includes three core dimensions: 1) Query-Knowledge Relevance (QC_relevance) evaluates the support degree by calculating the semantic matching degree between the retrieved knowledge fragment and the query; 2) True Answer-Knowledge Relevance (TC_relevance) analyzes the coverage of the retrieved knowledge and the standard answer; 3) Knowledge Support (Support) detects the knowledge support situation of each sentence in the generated answer to verify the factuality and reliability.

[0075] (9) (10) (11) Where K, M, and N respectively represent the total number of retrieved knowledge fragments, the total number of sentences in the true answer T, and the total number of sentences in the system answer A. The prompt large model representing the calculation of the semantic similarity between evaluation objects.

[0076] The generation quality evaluation focuses on two dimensions: 1) Answer Similarity (TA_similarity) quantifies the semantic coverage of the system answer and the standard answer; 2) User Evaluation (User_acceptance) measures user satisfaction from dimensions such as ease of understanding and interaction experience.

[0077] (12) The prompt large model that calculates the semantic similarity between the true answer T and the system answer A.

[0078] Through this multi-dimensional and fine-grained evaluation method, a comprehensive and objective performance evaluation framework is provided for the RAG question-answering system, aiming to accurately quantify the semantic association degree between evaluation objects.

[0079] In the embodiments of the present invention, the performance of the ChatGovt system is comprehensively evaluated through rigorous comparative experiments. We selected three representative benchmark systems: (1) a naive RAG system based on GLM4-9b-chat, which only uses the dense vector retrieval method for knowledge acquisition; (2) an enhanced model obtained by fine-tuning GLM4-9b-chat with the constructed knowledge base data under the 4-bit quantization condition through the LoRA method; (3) a commercial large model with an online query function, serving as a reference standard for the current state of the art.

[0080] Based on a carefully constructed single-round dialogue dataset and continuous follow-up scenarios, a test set is constructed, and the system generation performance is systematically evaluated from two key dimensions of retrieval quality and generation quality. The evaluation results are shown in Table 3.

[0081] Table 3 System Reply Evaluation Results

[0082] It can be seen that in terms of answer generation quality, the government affairs question-answering system based on the knowledge-enhanced large model provided by the embodiments of the present invention achieved excellent scores of 9.19 points and 6.34 points in terms of human evaluation and answer recall rate respectively, and all generation quality indicators were significantly better than the benchmark systems. Specifically, its answer recall rate increased by 55.4% compared with the fine-tuned model, and the human evaluation score exceeded that of the commercial large model by 27.3%, demonstrating the effectiveness of the external knowledge base. Through comparative analysis, it was found that although the fine-tuned model was better than the naive RAG system (3.08 points) in terms of answer recall rate (4.08 points), its human evaluation (5.72 points) was lower than that of the latter (7.97 points). This may be because the mechanical reproduction of policy texts was overemphasized during the fine-tuning process, resulting in the generated results conforming to the standard expressions but lacking the connecting words and scene-based explanations of natural conversations, indicating that simply relying on model fine-tuning may reduce the practicality of answers due to overfitting.

[0083] In terms of the quality of knowledge retrieval, the government affairs Q&A system based on the knowledge-enhanced large model demonstrates significant technical advantages. As shown in Table 3, the three core indicators of query-knowledge relevance (7.20), true answer-knowledge relevance (7.00), and knowledge support (7.25) are all better than those of the naive RAG system, with increases of 15.0%, 7.4%, and 24.6% respectively. This improvement stems from the effective integration of the hybrid retrieval strategy: through the synergistic effect of the multi-stage semantic-enhanced knowledge retrieval method, the system can not only capture the deep associations at the semantic level but also accurately locate the professional terms in the policy text, thus achieving "double guarantee" for knowledge positioning in complex government affairs scenarios.

[0084] Based on the same inventive concept, the present invention also provides a government affairs Q&A device based on a knowledge-enhanced large model, as Figure 5 shown, which shows the structural block diagram of the government affairs Q&A device based on the knowledge-enhanced large model. The government affairs Q&A device 500 based on the knowledge-enhanced large model includes: A receiving module 501, configured to receive the first text information in the form of natural language sent by a user; An intention analysis module 502, configured to obtain a query statement corresponding to the first text information based on the historical text information and the first text information sent by the user, where the query statement represents the true intention of the user; A retrieval module 503, configured to retrieve the first N knowledge fragments from the government affairs knowledge base based on the query statement; sort the semantic matching degrees of the first N knowledge fragments and the query statement to obtain the first K knowledge fragments with the highest semantic matching degrees; An evaluation module 504, configured to evaluate the availability of the first K knowledge fragments with the highest semantic matching degrees; An answering module 505, configured to introduce the knowledge fragments with availability as external knowledge into the large language model; obtain an answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge.

[0085] Optionally, sorting the semantic matching degrees of the first N knowledge fragments and the query statement to obtain the first K knowledge fragments with the highest semantic matching degrees includes: For any one of the first K knowledge fragments with the highest semantic matching degrees: Input the query statement and this knowledge fragment as a whole into the large language model, perform in-depth semantic interaction on the query and the document in combination with the prompt words, and output the semantic matching degree of the query statement and the knowledge fragment.

[0086] Optionally, the device further includes: The network query module is used to trigger the network query tool to obtain supplementary knowledge when none of the top K knowledge fragments with the highest semantic matching degree are available; The answering module is further used to obtain the answer text for the first text information based on the endogenous knowledge of the large language model and the supplementary knowledge.

[0087] Optionally, the device further includes: The classification module is used to identify the intention of the first text information through the large language model and determine the text classification corresponding to the first text information; The retrieval module 503 is used to retrieve external knowledge from the government affairs knowledge base based on the query statement and introduce it into the large language model when the first text information belongs to the consultation category, and obtain the answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge; The answering module 505 is further used to summarize and record the first text information based on the large language model and generate the answer text for the first text information when the first text information belongs to the complaint category, help-seeking category, and praise category; The answering module 505 is further used to generate the answer text for the first text information based on the large language model when the first text information belongs to other categories.

[0088] Optionally, the device further includes: The collection module is used to collect multi-source government affairs data; The data processing module is used to process the multi-source government affairs data to obtain the processed government affairs data; The knowledge base construction module is used to process the processed government affairs data based on the text encoder model to obtain vectorized data, and add the vectorized data to the government affairs knowledge base.

[0089] Optionally, the device further includes: The summary module is used to generate a service work order based on the first text information and the answer text of the first text information through the large language model, and send work order verification information based on the service work order; The feedback module is used to process the service work order through the text encoder model to obtain vectorized data and supplement the vectorized data to the government affairs knowledge base when the service work order is verified.

[0090] Based on the same inventive concept, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the government affairs question-answering method based on the knowledge-enhanced large language model described in any of the above embodiments.

[0091] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the government affairs question-answering method based on a knowledge-enhanced large language model described in any one of the above embodiments are implemented.

[0092] Based on the same inventive concept, the present invention provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps in the government affairs question-answering method based on a knowledge-enhanced large language model described in any one of the above embodiments are implemented.

[0093] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0094] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0095] The present invention is described with reference to the flowcharts and / or block diagrams of methods, terminal devices (devices), and computer program products according to the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0097] These computer program instructions can also be loaded onto a computer or other programmable terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process or multiple processes and / or blocks. Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the functions specified in one block or multiple blocks.

[0098] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0099] Finally, it should also be noted that in the present invention, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0100] The above has introduced in detail a government affairs Q&A method based on a knowledge-enhanced large language model provided by the present invention. Specific examples are used in the present invention to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A government affairs Q&A method based on a knowledge-enhanced large model, characterized in that, The method includes: Receiving the first text information in the form of natural language sent by the user; Based on the historical text information sent by the user and the first text information, obtaining a query statement corresponding to the first text information, where the query statement represents the true intention of the user; Retrieving the top N knowledge fragments from the government affairs knowledge base based on the query statement; Sorting the semantic matching degrees of the top N knowledge fragments and the query statement to obtain the top K knowledge fragments with the highest semantic matching degrees; Evaluating the availability of the top K knowledge fragments with the highest semantic matching degrees; Introducing the available knowledge fragments as external knowledge into the large language model; Obtaining an answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge.

2. The government affairs Q&A method based on the knowledge-enhanced large model according to claim 1, wherein Sorting the semantic matching degrees of the top N knowledge fragments and the query statement to obtain the top K knowledge fragments with the highest semantic matching degrees, including: For any one of the top K knowledge fragments with the highest semantic matching degrees: Taking the query statement and this knowledge fragment as a whole and inputting them into the large language model, combining the prompt words to perform in-depth semantic interaction on the query and the document, and outputting the semantic matching degree of the query statement and the knowledge fragment.

3. The government affairs Q&A method based on the knowledge-enhanced large model according to claim 1, characterized in that, The method further includes: In the case where none of the top K knowledge fragments with the highest semantic matching degrees are available, triggering the network query tool to obtain cross-domain knowledge; Obtaining an answer text for the first text information based on the endogenous knowledge of the large language model and the supplementary knowledge.

4. The government affairs Q&A method based on a knowledge-enhanced large model according to claim 1, wherein, The method further includes: Performing intention recognition on the first text information through the large language model to determine the text classification corresponding to the first text information; In the case where the first text information belongs to the consultation category, retrieving external knowledge from the government affairs knowledge base based on the query statement and introducing it into the large language model, and obtaining an answer text for the first text information based on the endogenous knowledge of the large language model and the external knowledge; In the case where the first text information belongs to the complaint category, help-seeking category, and praise category, summarizing and recording the first text information through the large language model, and generating an answer text for the first text information; In the case where the first text information belongs to other categories, generating an answer text for the first text information based on the large language model.

5. The government affairs Q&A method based on a knowledge-enhanced large model according to any one of claims 1 to 4, characterized in that, The method further includes: Collecting multi-source government affairs data; Performing data processing on the multi-source government affairs data to obtain processed government affairs data; Processing the processed government affairs data based on the text encoder model to obtain vectorized data, and adding the vectorized data to the government affairs knowledge base.

6. The government affairs Q&A method based on a knowledge-enhanced large model according to claim 5, characterized in that, The method further includes: Generating a service work order through the large language model based on the first text information and the answer text of the first text information, and sending work order verification information based on the service work order; In the case where the service work order is verified, processing the service work order through the text encoder model to obtain vectorized data, and supplementing the vectorized data to the government affairs knowledge base.

7. A government affairs question-answering device based on a knowledge-enhanced large model, the device includes: A receiving module, configured to receive a first text message in natural language sent by a user; An intent analysis module, configured to obtain a query statement corresponding to the first text message based on the historical text message sent by the user and the first text message, where the query statement represents the true intent of the user; A retrieval module, configured to retrieve the top N knowledge fragments from a government affairs knowledge base based on the query statement; Sort the semantic matching degrees between the top N knowledge fragments and the query statement to obtain the top K knowledge fragments with the highest semantic matching degrees; An evaluation module, configured to evaluate the availability of the top K knowledge fragments with the highest semantic matching degrees; An answering module, configured to introduce the knowledge fragments with availability as external knowledge into a large language model; and obtain an answer text for the first text message based on the endogenous knowledge of the large language model and the external knowledge.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the government affairs question-answering method based on a knowledge-enhanced large model according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the government affairs question-answering method based on a knowledge-enhanced large model according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by a processor, the steps in the government affairs question-answering method based on a knowledge-enhanced large model according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Government affair service field multi-strategy fusion dialogue method based on knowledge graph

    CN116628172A

  • Multi-feature fusion semantic understanding method and device for government affair questions and answers

    CN117313748A

  • Policy question and answer method and system based on large language model and knowledge graph technology

    CN117725170A

  • Government service question and answer method based on knowledge graph enhanced large language model

    CN117851577A

  • Government affair intelligent response device and method based on intention recognition and large language model

    CN118035419A

Cited By

  • Device and method for realizing telephone question and answer customer service agent based on large model RAG

    CN121071099A

  • Method and system for realizing query rewriting based on llm and context

    CN121166890A

  • Large language model knowledge question-answering method and system fused with multi-modal knowledge graph

    CN121235124A