Civil administration service question and answer method based on large model and knowledge graph retrieval enhancement

By integrating knowledge graphs and large language models into the civil affairs service question-answering system and adopting question decomposition and explicit reasoning chain generation technology, the problems of insufficient knowledge accuracy and multi-hop reasoning capabilities of traditional question-answering systems are solved, and efficient and transparent civil affairs service question-answering is achieved, which can adapt to the rapid deployment and updating of different civil affairs fields.

CN120780798APending Publication Date: 2025-10-14SHIJIAZHUANG TIEDAO UNIV
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
CN202510823006.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Traditional civil affairs question-answering systems rely on large language models (LLMs) and suffer from knowledge cutoff and hallucination problems, making it difficult to ensure the accuracy and timeliness of civil affairs information. They also lack effective utilization of structured knowledge, resulting in insufficient multi-hop reasoning capabilities for complex multi-condition and cross-departmental problems.

Method used

By integrating knowledge graphs with large language models, adopting question decomposition modules and explicit reasoning chain generation technology, multi-hop information retrieval and answer explainability are achieved. The structured knowledge of knowledge graphs and the natural language processing capabilities of large language models are used to perform question parsing and entity recognition, and through lightweight deployment and dynamic update mechanisms, it can quickly adapt to different civil affairs service scenarios.

Benefits of technology

It improves the accuracy, efficiency and transparency of civil affairs service question and answering, solves the bottlenecks of traditional question and answer systems in knowledge accuracy and multi-hop reasoning capabilities, realizes accurate and transparent civil affairs services, and reduces system deployment and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780798A_ABST
    Figure CN120780798A_ABST
Patent Text Reader

Abstract

The invention discloses a civil administration service question and answer method based on a large model and knowledge graph retrieval enhancement, and the method comprises the steps: S10, inputting a user question, and carrying out the question analysis and entity recognition; s20, the problem complexity is judged, if the problem is a single-hop problem, knowledge graph single-hop retrieval is carried out, and if the problem is a multi-hop problem, knowledge graph multi-hop retrieval is carried out; s30, performing semantic matching sorting to generate sub-answers; and S40, based on the sub-answers and the user question, performing synthesis to generate a final answer. According to the method, the structured knowledge of the knowledge graph and the natural language processing capability of the large language model are fused; a question decomposition module is used to enhance the interpretability of multi-hop information retrieval and answers; and using contextual learning (ICL) and thinking chain (CoT) prompts to generate an individually processed explicit inference chain to improve authenticity; the defects of traditional civil administration service questions and answers in the aspects of knowledge accuracy, reasoning ability and interpretability are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing technology, and in particular relates to a civil affairs service question-answering method based on a large model and knowledge graph retrieval enhancement. Background Art

[0002] Against the backdrop of the rapid development of digital government, intelligent question-answering systems for civil affairs services have become a key tool for improving government service efficiency and public satisfaction. However, traditional civil affairs question-answering systems primarily rely on large language models (LLMs), which exhibit numerous limitations when handling complex civil affairs issues. On the one hand, LLMs suffer from knowledge cutoffs and hallucinations, making it difficult to ensure the accuracy and timeliness of civil affairs information. On the other hand, civil affairs knowledge is highly structured and dynamically updated, encompassing multiple business areas such as social assistance, marriage registration, funerals and cremations, disability benefits, child welfare, and elderly care services. Traditional models lack effective utilization of structured knowledge, resulting in insufficient multi-hop reasoning capabilities when handling complex civil affairs issues involving multiple conditions and across departments.

[0003] Knowledge graphs (KGs), as a structured knowledge representation method, can store entities, relationships, and attributes in the civil affairs field in a graph-like format, supporting efficient knowledge retrieval and complex reasoning. Large language models have advantages in natural language understanding and generation, but they suffer from knowledge cutoffs and hallucinations, and they underutilize structured knowledge in the civil affairs field.

[0004] Existing technologies often require extensive training data or complex fine-tuning to integrate knowledge graphs with large models, making it difficult to quickly adapt to the needs of diverse civil affairs service scenarios. Achieving a deep integration of knowledge graphs and large models without training, and thus improving the accuracy, interpretability, and domain adaptability of civil affairs question-and-answer services, has become a pressing technical challenge. Summary of the Invention

[0005] In order to solve the above problems, the present invention proposes a civil service question-answering method based on a large model and knowledge graph retrieval enhancement, which integrates the structured knowledge of the knowledge graph with the natural language processing capability of the large language model; uses a question decomposition module to enhance multi-hop information retrieval and the interpretability of answers; and uses in-context learning (ICL) and chain of thought (CoT) prompts to generate separately processed explicit reasoning chains to improve authenticity; and solves the shortcomings of traditional civil service question-answering in terms of knowledge accuracy, reasoning ability and interpretability.

[0006] To achieve the above-mentioned purpose, the technical solution adopted by the present invention is: a civil affairs service question-answering method based on a large model and knowledge graph retrieval enhancement, comprising the following steps:

[0007] S10, user question input, question parsing and entity recognition;

[0008] S20, determine the complexity of the problem. If it is a single-hop problem, perform a single-hop search on the knowledge graph. If it is a multi-hop problem, perform a multi-hop search on the knowledge graph.

[0009] S30, perform semantic matching sorting and generate sub-answers;

[0010] S40, synthesize and generate the final answer based on the sub-answers and the user question.

[0011] Furthermore, through the collaboration of the entity relationship network of the knowledge graph and natural language processing technology, question parsing and entity recognition are achieved, including:

[0012] When a user enters a question, the pre-trained named entity recognition model is used to extract the key entities in the question, and the complexity of the question is determined based on the pre-defined entity types and relationship patterns in the knowledge graph.

[0013] The complexity of the problem is divided into: single-hop problem: only involves the direct relationship between two entities; multi-hop problem: the answer must be obtained through reasoning across multiple entities and relationships;

[0014] For single-hop questions, directly trigger the single-hop search of the knowledge graph and generate answers through the direct association between entity nodes;

[0015] For multi-hop questions, we design ICL examples that include reasoning chains and sub-questions to guide the large language model to learn decomposition patterns; we decompose the question Q into a logically coherent sequence of sub-questions: Q→{q1,q2,...,q k}, and then perform multi-hop traversal in the knowledge graph starting from the core entities extracted from the question.

[0016] Furthermore, a reasoning chain generation mechanism is established. Through manually designed multi-hop question-reasoning chain-sub-problem examples, the large language model is taught how to decompose complex problems into ordered reasoning steps. The large language model is required to first output the reasoning chain, that is, to describe the logical relationship from the problem to the sub-problem; and then generate the sub-problem to ensure that each step of the reasoning is explainable.

[0017] Furthermore, during the knowledge graph retrieval process, candidate triples are retrieved:

[0018] For single-hop questions, there is no need to decompose the question and directly use the knowledge graph single-hop retrieval strategy;

[0019] For multi-hop questions, we first decompose the question into a logically coherent sequence of sub-questions. Then, we perform a multi-hop range search in the civil affairs knowledge graph to obtain hierarchical knowledge related to the question. The search strategy is dynamically selected based on the semantic relationship direction of the question:

[0020] Bidirectional breadth-first search strategy: Based on the breadth-first search algorithm, the knowledge graph is traversed, starting from the starting entity and the end entity or the reverse relationship, and expanding to the middle layer to form a bidirectional search path; for problems with uncertain direction or complex conditions, the outbound and inbound edges of the entity are analyzed at the same time, supporting reverse positioning of the organization from the service, and then superimposing the geographic location for cross-screening.

[0021] Furthermore, semantic matching and sorting include the following steps:

[0022] The pre-trained sentence embedding model is used to construct the semantic vector space and the sub-question q i and candidate triplet t i Encoded as high-dimensional vectors h i and g i , where the vector dimension matches the model architecture; through cosine similarity, the calculation formula is:

[0023]

[0024] This formula measures the cosine of the angle between vectors by normalizing the dot product. The value range is [-1, 1]. The larger the value, the higher the semantic association.

[0025] Based on the score, the candidate triples are sorted in descending order, and the top K most relevant triples are selected as the retrieval results T using the Top-K screening mechanism. K is a dynamically adjustable parameter with a default value of 30, which ensures a balance between retrieval coverage and computational efficiency. The screened triple set T contains policy clauses, process nodes, and departmental responsibility knowledge that are highly relevant to the sub-problem.

[0026] Furthermore, sub-answer generation and answer synthesis include:

[0027] Sub-question answer generation: Based on the candidate triple set T obtained by knowledge graph retrieval, the zero-sample generation capability of the pre-trained large model is used to generate the answer for the sub-question q i Generate natural language answer a i ; Among them, q i is the sub-question text output by the question decomposition module, T is the Top-K triples returned by the knowledge graph retrieval module, and the sub-answer a is generated by the large language model. i , the generated sub-answer a i Stored in the context cache as input conditions for subsequent sub-problems.

[0028] Answer synthesis and reasoning chain integration: According to the logical thread of the problem decomposition, the sub-answers a1, a2, ..., a nIt is hierarchically integrated with the reasoning chain c; the reasoning chain c is generated by the problem decomposition stage, reflecting the logical deduction process from the original problem to the sub-problem, and its content directly corresponds to the relationship path in the knowledge graph; the generated reasoning chain, sub-problem and sub-answer are used to synthesize the answer to the original question; the system uses the reasoning chain as the logical main line, connects the sub-problem and the answer in sequence, uses natural language transition words to connect each link, and clarifies the hierarchical structure through numbering or segmentation.

[0029] Furthermore, data collection is carried out to build a knowledge graph. Based on the characteristics of structured and unstructured data in the civil affairs field, differentiated processing paths are established to achieve structured modeling of the two types of data and form a complete knowledge network in the civil affairs field.

[0030] Data collection includes:

[0031] Government public data: serves as the core data source for constructing knowledge graphs in the civil affairs field;

[0032] User consultation logs: A key data source reflecting real civil affairs service needs, used to optimize the practicality and problem coverage of the knowledge graph;

[0033] Third-party data: As a supplementary source of civil affairs knowledge, it is used to expand the domain associations and dynamic information of the knowledge graph.

[0034] Furthermore, in constructing the knowledge graph, for structured data, there are clear field definitions and relationship models, and efficient conversion is achieved through a templated parsing engine;

[0035] During the parsing process, the cleaning rules are established as follows:

[0036] R c ={r1: remove duplicate records, r2: unify unit format, r3: fix missing values};

[0037] Through data cleaning rules, duplicate records are removed, coding formats are unified, and based on business logic, relationship types are predefined to build hierarchical associations between entities. Each structured data is converted into a triple T through mapping rules. i ={(s j ,r k ,o l )},s j As the main body, r k For contact, l for the object;

[0038] Dynamic updates of structured data are achieved through field change monitors, which are set to:

[0039]

[0040] in, Indicates the current value of the i-th field in the business system at timestamp t. represents the historical value of the same field i at timestamp (t-1);

[0041] When the field value in the business system changes When , the revision of the corresponding attribute triple in the graph is automatically triggered;

[0042] In constructing the knowledge graph, for unstructured data, natural language processing technology is used to achieve semantic extraction and structural conversion; the BERT-NER model is used for entity recognition, input text, and output entity set E = {e1, e2, ..., e n}, integrating domain dictionaries to enhance recognition capabilities; then using sequence labeling models or graph neural networks to determine relationship types based on contextual semantics; pre-defining the core relationship set R in the civil affairs field = {application conditions, handling procedures, material list, department responsibilities}; for entity pairs (e i , e j ), the output probability distribution of the relation extraction model is:

[0043]

[0044] Among them, s(r ij |e i ,e j ,Y) means that in the context text Y, the entity pair (e i ,e j ) belongs to the relation r ij The non-normalized score of , the larger the score value, the more likely the model believes that the entity pair belongs to the relationship r ij The higher the probability; ij To predict the relationship, r ′ is any relation in the relation set R;

[0045] After the entities and relationships are extracted, knowledge fusion is performed to disambiguate and align the entities and relationships in multi-source data, and a unified knowledge graph G in the civil affairs field is constructed. It is stored in the form of triples (s, r, o), where s is the subject, r is the relationship, and o is the object.

[0046] Furthermore, a lightweight deployment strategy and a training-free prompt learning framework enable rapid adaptation, avoiding the traditional model's reliance on large amounts of labeled data or complex fine-tuning. Specific implementations include:

[0047] (1) Parameter quantization and compression: Advanced quantization technology is used to compress large language model parameters;

[0048] (2) Multi-task parallel reasoning and prompt learning: For the sub-problem retrieval and generation process of multi-hop problems, the parallel computing capabilities of multi-core CPUs / GPUs are utilized to distribute the processing tasks of different sub-problems to multiple computing units for simultaneous execution;

[0049] (3) Cross-domain adaptation without training: Adapt to different civil affairs business scenarios through dynamic prompt engineering, predefine common entity types and relationship sets, and quickly map to the knowledge graph of the new domain by adjusting the prompt words.

[0050] Furthermore, service quality is continuously optimized through dynamic updates of the knowledge graph and a closed loop of user feedback, without the need for model training or parameter adjustment.

[0051] (1) Incremental update mechanism of knowledge graph: Establish a special information collection and update mechanism to obtain the latest civil affairs information in real time from authoritative channels including the official website of the civil affairs department and the policy document library; automatically identify new triples through unsupervised entity relationship extraction technology to trigger incremental updates of the knowledge graph;

[0052] (2) User feedback closed-loop optimization mechanism: Collect user feedback data on answers, including user evaluation of the accuracy, completeness, and logic of the answers, as well as suggestions for improving system functions. Use feedback data to optimize question decomposition strategies and retrieval parameters, and continuously improve system performance and user satisfaction.

[0053] The beneficial effects of adopting this technical solution are:

[0054] This invention promotes the upgrade of the RAG framework from "text-level retrieval" to "knowledge-level reasoning" by introducing knowledge graph-enhanced retrieval strategies and question decomposition mechanisms, solving the bottlenecks of knowledge accuracy, multi-hop reasoning capabilities and domain adaptability exposed by traditional civil service question and answering under the large language model technology path, and achieving accuracy and efficiency in civil service question and answering.

[0055] This invention breaks through the traditional RAG framework's reliance on text retrieval through a collaborative architecture of untrained knowledge graphs and large language models, dynamic question decomposition and multi-hop retrieval mechanisms, and explicit reasoning chain generation technology. Through technological innovation and adaptation to civil affairs scenarios, this method comprehensively improves the accuracy, efficiency, and transparency of civil affairs service Q&A, providing a feasible solution for the development of intelligent and precise digital government services, with significant social value and application prospects.

[0056] The present invention can create an accurate and reliable civil affairs knowledge supply system: the large language model that traditional civil affairs service questions and answers rely on has knowledge cutoff and hallucination problems, such as the possibility of citing outdated social assistance policies or fictitious elderly care service processing conditions. The present invention models civil affairs knowledge as (subject, relationship, object) triples through the structured retrieval mechanism of the knowledge graph, and links it with the real-time updated civil affairs knowledge base. When a user consults about civil affairs information, this method can directly retrieve the latest social assistance, elderly care services and other policy triples in the graph, providing real-time and accurate knowledge support for the large model, thereby reducing the error rate of reasoning to generate answers, ensuring the accuracy and timeliness of information such as policy interpretation and process description, and avoiding service misleading due to knowledge lag or fiction.

[0057] The present invention can enhance the reasoning ability of multi-hop problems: for multi-condition, cross-departmental complex problems commonly seen in the civil affairs field, such as "What materials and procedures are needed to apply for two subsidies for disabled people in other places", the present invention breaks down complex problems into a sequence of sub-problems through a problem decomposition module, such as "What are the application conditions for two subsidies for disabled people", "What additional supporting materials are required for applications in other places", and "What are the specific handling departments and procedures", and uses the multi-hop retrieval and reasoning capabilities of the knowledge graph to establish association paths between entities, realize the logical chain deduction from problem to answer, effectively improve the system's ability to handle multi-level and cross-link civil affairs issues, shorten the problem-solving path, and improve service efficiency.

[0058] This invention can enhance the transparency and user trust of civil affairs services: by generating an explicit reasoning chain, linking the problem-solving process with specific knowledge nodes in the knowledge graph, and clearly demonstrating the logical basis from user questions to answer generation. For example, when answering the question "How to apply for admission to a nursing home," the reasoning chain is displayed: "Determine the nursing home selection criteria → Search for qualified nursing homes → Understand the admission application process → Submit application materials", and the knowledge graph node corresponding to each link is labeled. This allows users to intuitively understand the rationality of system decisions, enhances the traceability and credibility of civil affairs services, and promotes the transformation of civil affairs question-answering from "black box generation" to "transparent reasoning."

[0059] This invention enables rapid adaptation in the civil affairs sector: It designs a training-free knowledge graph-large model fusion architecture. Through contextual learning and prompt engineering, it can quickly adapt knowledge graphs in different civil affairs fields, such as social assistance, elderly care services, and social organization management, without relying on large amounts of labeled data or complex model fine-tuning. This reduces the system's dependence on specific fields and deployment costs, improves the scalability and responsiveness of civil affairs services across different regions and departments, and meets the needs of diversified and dynamic civil affairs scenarios.

[0060] In summary, the present invention aims to construct an intelligent question-answering method for civil affairs services that integrates the structured knowledge of knowledge graphs and the natural language processing capabilities of large language models. It fills the knowledge gaps of large language models through real-time structured retrieval of knowledge graphs, and improves the accuracy and timeliness of civil affairs service questions and answers; with the help of problem decomposition and multi-hop retrieval of knowledge graphs, it strengthens the reasoning ability of complex civil affairs issues with multiple conditions and cross-links; by generating explicit reasoning chains, it associates the answer derivation process with knowledge graph nodes, and enhances the transparency of civil affairs services and user trust; designs a training-free fusion architecture to achieve rapid adaptation to knowledge graphs in different civil affairs fields, reduce system deployment and maintenance costs, and ultimately build an efficient, accurate, explainable and flexibly scalable intelligent question-answering system for civil affairs services, and promote the development of digital government affairs towards intelligence and precision. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 This is a flowchart of a civil affairs service question-answering method based on a large model and knowledge graph retrieval enhancement according to the present invention;

[0062] Figure 2 This is a schematic diagram of the overall framework in an embodiment of the present invention;

[0063] Figure 3 Schematic diagram of the steps for generating a sub-answer in an embodiment of the present invention. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described below with reference to the accompanying drawings.

[0065] In this embodiment, see Figure 1 and Figure 2 As shown, the present invention proposes a civil affairs service question answering method based on a large model and knowledge graph retrieval enhancement, comprising the following steps:

[0066] S10, user question input, question parsing and entity recognition;

[0067] S20, determine the complexity of the problem. If it is a single-hop problem, perform a single-hop search on the knowledge graph. If it is a multi-hop problem, perform a multi-hop search on the knowledge graph.

[0068] S30, perform semantic matching sorting and generate sub-answers;

[0069] S40, synthesize and generate the final answer based on the sub-answers and the user question.

[0070] As the above embodiment, the entity relationship network of the knowledge graph and the natural language processing technology are used in collaboration to achieve question parsing and entity recognition, including:

[0071] When a user enters a question, the pre-trained named entity recognition model is first used to extract key entities in the question, such as "nursing home," "disability subsidy," and "social assistance." The complexity of the question is then determined based on the pre-defined entity types and relationship patterns in the knowledge graph.

[0072] The complexity of the problem is divided into: single-hop problem: only involves the direct relationship between two entities; multi-hop problem: the answer must be obtained through reasoning across multiple entities and relationships;

[0073] For single-hop questions, directly trigger the single-hop search of the knowledge graph and generate answers through the direct association between entity nodes;

[0074] For multi-hop questions, we design ICL examples that include reasoning chains and sub-questions to guide the large language model to learn decomposition patterns; we decompose the question Q into a logically coherent sequence of sub-questions: Q→{q1,q2,...,q k}, and then perform multi-hop traversal in the knowledge graph starting from the core entities extracted from the question.

[0075] Preferably, a reasoning chain generation mechanism is established. Through manually designed multi-hop problem-reasoning chain-sub-problem examples, the large language model is taught how to decompose complex problems into ordered reasoning steps. The large language model is required to first output the reasoning chain, that is, to describe the logical relationship from the problem to the sub-problem; and then generate the sub-problem to ensure that each step of the reasoning is explainable.

[0076] Example format:

[0077] Original question: "Does a family meet the minimum living security requirements for urban and rural residents?"

[0078] Chain of reasoning: "In order to determine whether the family meets the minimum living security conditions, it is first necessary to determine whether the family's per capita income and property status are lower than the local minimum living security standard, and then verify whether there are any statutory support / maintenance obligations and other eligibility conditions."

[0079] Sub-questions:

[0080] 1. "What is the average monthly income per capita for this household? What are the sources of income?";

[0081] 2. "What is the family's financial situation? Does it exceed the local financial limit?"

[0082] 3. “Is there a legal caregiver / raiser in this family, and does he / she have the ability to provide support / raise the child?”

[0083] As an optimization solution of the above embodiment, Figure 2 As shown in the figure, during the knowledge graph retrieval process, candidate triples are retrieved:

[0084] For single-hop questions, there is no need to decompose the question and directly use the knowledge graph single-hop retrieval strategy;

[0085] For multi-hop questions, we first decompose the question into a logically coherent sequence of sub-questions. Then, we perform a multi-hop range search in the civil affairs knowledge graph to obtain hierarchical knowledge related to the question. The search strategy is dynamically selected based on the semantic relationship direction of the question:

[0086] Bidirectional breadth-first search strategy: Based on the breadth-first search algorithm, the knowledge graph is traversed, starting from the starting entity and the end entity or the reverse relationship, and expanding to the middle layer to form a bidirectional search path. For questions with uncertain directions or complex conditions, such as "Which nursing homes provide care services for disabled elderly people and are located in a certain area", the outbound and inbound edges of the entities are analyzed at the same time, supporting the reverse positioning of institutions from the service, and then superimposing the geographical location for cross-screening to improve the search coverage capability.

[0087] While unidirectional search may miss entities connected by reverse relationships, bidirectional search ensures that all possible entity-relationship paths in multi-hop questions are searched. Bidirectional breadth-first search ensures a comprehensive search of the knowledge graph in complex questions by simultaneously exploring both forward and reverse relationship paths. Combined with the filtering mechanism of text embedding models, this strategy improves answer accuracy while avoiding computational overhead, supporting the interpretability and robustness of multi-hop reasoning.

[0088] As an optimization solution of the above embodiment, semantic matching and sorting includes the following steps:

[0089] In order to achieve accurate association between sub-questions and candidate triples, a pre-trained sentence embedding model is used to construct a semantic vector space, and the sub-question q i and candidate triplet t i Encoded as high-dimensional vectors h i and g i , where the vector dimension matches the model architecture; through cosine similarity, the calculation formula is:

[0090]

[0091] This formula measures the cosine of the angle between vectors by normalizing the dot product. The value range is [-1, 1]. The larger the value, the higher the semantic association.

[0092] Based on the score, the candidate triples are sorted in descending order, and the top K most relevant triples are selected as the search results T using the Top-K screening mechanism. K is a dynamically adjustable parameter with a default value of 30, which ensures a balance between search coverage and computational efficiency. The screened triple set T contains policy clauses, process nodes, and departmental responsibility knowledge that are highly relevant to the sub-problem, such as the review department for social assistance, the contact information of nursing homes, and the payment time of disability subsidies. These are directly used to guide the large model to generate accurate answers, while also providing structured evidence support for the traceability of the reasoning chain.

[0093] As an optimization solution of the above embodiment, Figure 3 As shown, for sub-answer generation and answer synthesis, it includes:

[0094] Sub-question answer generation: Based on the candidate triple set T obtained by knowledge graph retrieval, the zero-sample generation capability of the pre-trained large model is used to generate the answer for the sub-question q i Generate natural language answer a i ; Among them, q i is the sub-question text output by the question decomposition module, T is the Top-K triples returned by the knowledge graph retrieval module, and the sub-answer a is generated by the large language model. i , the generated sub-answer a i Stored in the context cache as input conditions for subsequent sub-problems.

[0095] For example, when processing the "process of applying for two subsidies for disabled people in other places", the first sub-answer "household registration review requires the original household registration book" can be used as the reasoning basis for the subsequent "extra materials for off-site review" sub-question. If the current sub-answer fails to fully cover the problem requirements, the subsequent sub-questions are reconstructed through LLM. For example, when the answer to the sub-question "conditions for admission to nursing homes" only includes age restrictions, the reconstructed sub-question is "whether a health certificate is required for admission to nursing homes" to supplement the missing information. By forcing the large model to generate sub-answers a i This mechanism effectively avoids the hallucination problem of traditional large language models by binding the generation process of large language models with the structured triples of knowledge graphs, while providing clear knowledge source annotations for answers to ensure the accuracy of information and compliance of policy references.

[0096] Answer synthesis and reasoning chain integration: According to the logical thread of the problem decomposition, the sub-answers a1, a2, ..., a nThe system then hierarchically integrates the reasoning chain c, generated during the problem decomposition phase, with the logical deduction process from the original question to the subquestions. Its content directly corresponds to the relational paths in the knowledge graph. The generated reasoning chain, subquestions, and subanswers are used to synthesize the answer to the original question. The system uses the reasoning chain as the logical thread, sequentially linking the subquestions and answers. Natural language transitions are used to connect each link, and the hierarchical structure is clarified through numbering or segmentation. For example, for the question "How to apply for assistance for the extremely poor?" the integrated answer first presents the reasoning chain: "Applying for assistance for the extremely poor requires confirming household registration and income requirements, preparing application materials, and submitting for review." It then lists the answers to the subquestions in point-by-point format, with each answer item annotated with a triple reference from the knowledge graph. Through this mechanism, the generated answers exhibit both natural language fluency and logical traceability through explicit knowledge graph associations, enhancing the transparency and user trust of civil affairs services and meeting the requirements of civil affairs Q&A for process clarity and policy compliance.

[0097] As an optimization solution to the above embodiment, data collection is carried out to build a knowledge graph. Based on the characteristics of structured and unstructured data in the civil affairs field, differentiated processing paths are established to achieve structured modeling of the two types of data and form a complete knowledge network in the civil affairs field.

[0098] Data collection includes:

[0099] Government open data: As the core data source for constructing knowledge graphs in the civil affairs field, it is authoritative, standardized, and timely. It provides access to administrative regulations, local regulations, and departmental rules issued by civil affairs departments, as well as standardized service guidelines issued by provincial government platforms. These cover procedures and lists of materials for social assistance, marriage registration, funeral services, and cremation services, as well as frequently asked questions compiled by various government service halls, in various formats including text, PDF, and charts.

[0100] User consultation logs: A key data source reflecting real-world civil affairs service needs, used to optimize the practicality and coverage of the knowledge graph. This log collects user questions and system responses from civil affairs service hotlines and online consultation platforms, such as inquiries about disability subsidies and child welfare applications, as well as user-submitted work orders, materials submitted during service processing, and review comments.

[0101] Third-party data: Serves as a supplementary source of civil affairs knowledge, expanding the domain relevance and dynamic information of the knowledge graph. This data includes reports on elderly care services published by industry associations, social assistance policy assessments from research institutions, and publicly available information on the internet, such as articles on civil affairs policies from authoritative media outlets and transcripts of expert interview videos, to enrich the knowledge graph.

[0102] In building the knowledge graph, structured data, such as household registration information for social assistance recipients, the number of beds in nursing homes, and the standards for disbursement of disability subsidies, has clear field definitions and relationship models, enabling efficient conversion through a templated parsing engine.

[0103] During the parsing process, the cleaning rules are established as follows:

[0104] R c ={r1: remove duplicate records, r2: unify unit format, r3: fix missing values};

[0105] Through data cleaning rules, duplicate records are removed, coding formats are unified, and based on business logic, relationship types are predefined to build hierarchical associations between entities. Each structured data is converted into a triple T through mapping rules. i ={(s j ,r k ,o l )},s j As the main body, r k For contact, l for the object;

[0106] Dynamic updates of structured data are achieved through field change monitors, which are set to:

[0107]

[0108] in, Indicates the current value of the i-th field in the business system at timestamp t. represents the historical value of the same field i at timestamp (t-1);

[0109] When the field value in the business system changes When a new attribute is generated, the revision of the corresponding attribute triples in the graph is automatically triggered to ensure the timeliness of knowledge;

[0110] In constructing the knowledge graph, for unstructured data, such as policy documents, user consultation texts, and third-party reports, natural language processing technology is used to achieve semantic extraction and structured conversion; the BERT-NER model is used for entity recognition, input text, and output entity set E = {e1, e2, ..., e n}, and integrate domain dictionaries to enhance recognition capabilities; then use sequence labeling models (such as BiLSTM-CRF) or graph neural networks (GNN) to determine the relationship type based on contextual semantics; predefine the core relationship set R in the civil affairs field = {application conditions, handling procedures, material list, department responsibilities}; for entity pairs (e i , e j ), the output probability distribution of the relation extraction model is:

[0111]

[0112] Among them, s(r ij |e i ,e j ,Y) means that in the context text Y, the entity pair (e i ,e j ) belongs to the relation r ij The non-normalized score of , the larger the score value, the more likely the model believes that the entity pair belongs to the relationship r ij The higher the probability; ij To predict the relationship, r ′ is any relation in the relation set R.

[0113] After the entities and relationships are extracted, knowledge fusion is performed to disambiguate and align the entities and relationships in multi-source data, and a unified knowledge graph G in the civil affairs field is constructed. It is stored in the form of triples (s, r, o), where s is the subject, r is the relationship, and o is the object.

[0114] As an optimization solution for the above embodiment, the model deployment and reasoning of the method are as follows:

[0115] In actual civil affairs service scenarios, considering the differences in hardware resources among civil affairs departments in different regions, a lightweight deployment strategy and a training-free prompt learning framework are used to achieve rapid adaptation, avoiding the traditional model's reliance on large amounts of labeled data or complex fine-tuning. Specific implementation methods include:

[0116] (1) Parameter quantization and compression: Advanced quantization technology is used to compress large language model parameters; this quantization technology can significantly reduce the model storage requirements while preserving the model performance to the greatest extent. Through 4-bit quantization of large language models, its parameter storage space can be compressed to about a quarter of its original size, allowing the model to run efficiently on edge devices or government cloud servers. This process does not require retraining the model, and only hardware adaptation is achieved through quantization technology. This reduces the requirements for hardware resources, improves the scalability and deployment flexibility of the system, and adapts to the equipment conditions of grassroots civil affairs departments.

[0117] (2) Multi-task parallel reasoning and prompt learning: For the sub-problem retrieval and generation process of multi-hop problems, the parallel computing capabilities of multi-core CPUs / GPUs are utilized to distribute the processing tasks of different sub-problems to multiple computing units for simultaneous execution. For example, through context-based learning (ICL) and chain of thought (CoT) prompt templates, large models are guided to learn problem decomposition patterns. Without relying on training data in specific fields, the model can master the logical decomposition capabilities of complex problems through manually designed multi-hop problem-reasoning chain examples.

[0118] (3) Cross-domain adaptation without training: Adapt to different civil affairs business scenarios through dynamic prompt engineering, predefine common entity types and relationship sets, and quickly map them to the knowledge graph of the new domain by adjusting the prompt words; when expanding or changing the business, only relevant entity keywords need to be added or modified in the prompt, without the need to retrain the model or adjust the architecture.

[0119] As an optimization solution for the above embodiment, the method performance optimization mechanism is as follows:

[0120] Policies, regulations, and business processes are constantly changing. To ensure the information provided by the system is always accurate and effective, the system continuously optimizes service quality through dynamic updates of the knowledge graph and a closed loop of user feedback, without requiring model training or parameter adjustments.

[0121] (1) Incremental update mechanism of knowledge graph: Establish a special information collection and update mechanism to obtain the latest civil affairs information in real time from authoritative channels including the official website of the civil affairs department and the policy document library, such as the newly added pension service subsidy policy and the adjusted social assistance standards. Automatically identify new triples through unsupervised entity relationship extraction technology to trigger incremental updates of the knowledge graph; when new policies are introduced, the system can quickly identify and add relevant new entities and relationships to the knowledge graph, and update or delete existing triples. Through incremental updates, the tedious process of rebuilding the entire knowledge graph is avoided, the update efficiency is improved, and the timeliness and accuracy of the knowledge graph are guaranteed.

[0122] (2) User feedback closed-loop optimization mechanism: Collect user feedback data on answers, including user evaluation of the accuracy, completeness, and logic of the answers, as well as suggestions for improving system functions. Use feedback data to optimize question decomposition strategies and retrieval parameters, and continuously improve system performance and user satisfaction.

[0123] This paper proposes an intelligent question-answering method for civil affairs services based on a large model and knowledge graph retrieval enhancement. By constructing a deep collaborative architecture of knowledge graphs and large language models, it achieves the organic unity of structured management of civil affairs knowledge and natural language interaction. This method models civil affairs domain knowledge through the (subject, relationship, object) triples of the knowledge graph, and combines a dynamic retrieval mechanism to provide real-time and accurate structured knowledge support for the large language model, solving the knowledge gap and illusion problems of traditional large models; a problem decomposition module is designed to transform complex civil affairs problems into a sequence of sub-problems, and the multi-hop reasoning capability of the knowledge graph is relied upon to plan logical paths, thereby improving the efficiency of solving cross-departmental and multi-process problems; an inference chain generation mechanism is introduced to explicitly associate the answer derivation process with the knowledge graph nodes, thereby enhancing the traceability of civil affairs services and user trust; a training-free fusion architecture is constructed to achieve rapid adaptation to knowledge graphs in different civil affairs fields through contextual learning and prompt engineering, thereby reducing system deployment and maintenance costs. The technical solution of this invention aims to break through the performance bottleneck of traditional systems and build an intelligent question-and-answer system for civil affairs services that integrates accuracy, explainability, efficient reasoning capabilities and flexible scalability, providing technical support for the precision, transparency and agility of civil affairs services, and helping to improve government service efficiency and public satisfaction.

[0124] At the technical implementation level, the system adopts a lightweight deployment strategy to compress model parameters, combines parallel processing to optimize response speed, and ensures efficient operation in edge devices and government cloud environments; through incremental update mechanisms and user feedback closed loops, it continuously optimizes knowledge graphs and retrieval strategies to ensure the system's real-time adaptation to dynamic civil affairs knowledge.

[0125] In summary, this method is significantly superior to the traditional RAG model in terms of multi-hop problem processing accuracy, and can be quickly adapted to different civil affairs fields without a large amount of training data, reducing deployment and maintenance costs. It provides strong technical support for the intelligent upgrade of civil affairs services and has broad application prospects and social value.

[0126] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A civil affairs service question answering method based on a large model and knowledge graph retrieval enhancement, characterized by: Including steps: S10, user question input, question parsing and entity recognition; S20, determine the complexity of the problem. If it is a single-hop problem, perform a single-hop search on the knowledge graph. If it is a multi-hop problem, perform a multi-hop search on the knowledge graph. S30, perform semantic matching sorting and generate sub-answers; S40, synthesize and generate the final answer based on the sub-answers and the user question.

2. A civil affairs service question-answering method based on a large model and knowledge graph retrieval enhancement according to claim 1, characterized in that: Through the collaboration of the entity relationship network of the knowledge graph and natural language processing technology, question parsing and entity recognition are achieved, including: When a user enters a question, the pre-trained named entity recognition model is used to extract the key entities in the question, and the complexity of the question is determined based on the pre-defined entity types and relationship patterns in the knowledge graph. The complexity of the problem is divided into: single-hop problem: only involves the direct relationship between two entities; multi-hop problem: the answer must be obtained through reasoning across multiple entities and relationships; For single-hop questions, directly trigger the single-hop search of the knowledge graph and generate answers through the direct association between entity nodes; For multi-hop questions, we design ICL examples that include reasoning chains and sub-questions to guide the large language model to learn decomposition patterns; we decompose the question Q into a logically coherent sequence of sub-questions: Q→{q1,q2,...,q k }, and then perform multi-hop traversal in the knowledge graph starting from the core entities extracted from the question.

3. The civil affairs service question-answering method based on large model and knowledge graph retrieval enhancement according to claim 2 is characterized in that: Establish an inference chain generation mechanism. Through manually designed multi-hop question-inference chain-sub-problem examples, teach the large language model how to decompose complex problems into orderly reasoning steps. The large language model is required to first output the inference chain, that is, describe the logical relationship from the problem to the sub-problem; then generate the sub-problem to ensure that each step of the reasoning is explainable.

4. The civil affairs service question-answering method based on large model and knowledge graph retrieval enhancement according to claim 1 is characterized in that: During the knowledge graph retrieval process, candidate triples are retrieved: For single-hop questions, there is no need to decompose the question and directly use the knowledge graph single-hop retrieval strategy; For multi-hop questions, we first decompose the question into a logically coherent sequence of sub-questions. Then, we perform a multi-hop range search in the civil affairs knowledge graph to obtain hierarchical knowledge related to the question. The search strategy is dynamically selected based on the semantic relationship direction of the question: Bidirectional breadth-first search strategy: Based on the breadth-first search algorithm, the knowledge graph is traversed, starting from the starting entity and the end entity or the reverse relationship, and expanding to the middle layer to form a bidirectional search path; for problems with uncertain direction or complex conditions, the outbound and inbound edges of the entity are analyzed at the same time, supporting reverse positioning of the organization from the service, and then superimposing the geographic location for cross-screening.

5. The civil affairs service question answering method based on large model and knowledge graph retrieval enhancement according to claim 1 is characterized in that ,Semantic matching and sorting, including the following steps: The pre-trained sentence embedding model is used to construct the semantic vector space and the sub-question q i and candidate triplet t i Encoded as high-dimensional vectors h i and g i , where the vector dimension matches the model architecture; through cosine similarity, the calculation formula is: This formula measures the cosine of the angle between vectors by normalizing the dot product. The value range is [-1, 1]. The larger the value, the higher the semantic association. Based on the score, the candidate triples are sorted in descending order, and the top K most relevant triples are selected as the retrieval results T using the Top-K screening mechanism. K is a dynamically adjustable parameter with a default value of 30, which ensures a balance between retrieval coverage and computational efficiency. The screened triple set T contains policy clauses, process nodes, and departmental responsibility knowledge that are highly relevant to the sub-problem.

6. The civil affairs service question-answering method based on large model and knowledge graph retrieval enhancement according to claim 1 is characterized in that: For sub-answer generation and answer synthesis, including: Sub-question answer generation: Based on the candidate triple set T obtained by knowledge graph retrieval, the zero-sample generation capability of the pre-trained large model is used to generate the answer for the sub-question q i Generate natural language answer a i ; Among them, q i is the sub-question text output by the question decomposition module, T is the Top-K triples returned by the knowledge graph retrieval module, and the sub-answer a is generated by the large language model. i , the generated sub-answer a i Stored in the context cache as input conditions for subsequent sub-problems; Answer synthesis and reasoning chain integration: According to the logical thread of the problem decomposition, the sub-answers a1, a2, ..., a n It is hierarchically integrated with the reasoning chain c; the reasoning chain c is generated by the problem decomposition stage, reflecting the logical deduction process from the original problem to the sub-problem, and its content directly corresponds to the relationship path in the knowledge graph; the generated reasoning chain, sub-problem and sub-answer are used to synthesize the answer to the original question; the system uses the reasoning chain as the logical main line, connects the sub-problem and the answer in sequence, uses natural language transition words to connect each link, and clarifies the hierarchical structure through numbering or segmentation.

7. The civil affairs service question-answering method based on large model and knowledge graph retrieval enhancement according to claim 1 is characterized in that: Conduct data collection and build knowledge graphs. Targeting the characteristics of structured and unstructured data in the civil affairs field, establish differentiated processing paths, implement structured modeling of the two types of data, and form a complete knowledge network in the civil affairs field. Data collection includes: Government public data: serves as the core data source for constructing the knowledge graph in the civil affairs field; User consultation logs: A key data source reflecting real civil affairs service needs, used to optimize the practicality and problem coverage of the knowledge graph; Third-party data: As a supplementary source of civil affairs knowledge, it is used to expand the domain associations and dynamic information of the knowledge graph.

8. The civil affairs service question-answering method based on large model and knowledge graph retrieval enhancement according to claim 7 is characterized in that: In building a knowledge graph, for structured data, there are clear field definitions and relationship models, and efficient conversion is achieved through a templated parsing engine; During the parsing process, the cleaning rules are established as follows: R c ={r1: remove duplicate records, r2: unify unit format, r3: fix missing values}; Through data cleaning rules, duplicate records are removed, coding formats are unified, and based on business logic, relationship types are predefined to build hierarchical associations between entities. Each structured data is converted into a triple T through mapping rules. i ={(s j ,r k ,o l )},s j As the main body, r k For contact, l for the object; Dynamic updates of structured data are achieved through field change monitors, which are set to: in, Indicates the current value of the i-th field in the business system at timestamp t. represents the historical value of the same field i at timestamp (t-1); When the field value in the business system changes When , the revision of the corresponding attribute triple in the graph is automatically triggered; In constructing the knowledge graph, for unstructured data, natural language processing technology is used to achieve semantic extraction and structural conversion; the BERT-NER model is used for entity recognition, input text, and output entity set E = {e1, e2, ..., e n }, integrating domain dictionaries to enhance recognition capabilities; then using sequence labeling models or graph neural networks to determine relationship types based on contextual semantics; pre-defining the core relationship set R in the civil affairs field = {application conditions, handling procedures, material list, department responsibilities}; for entity pairs (e i , e j ), the output probability distribution of the relation extraction model is: Among them, s(r ij |e i ,e j ,Y) means that in the context text Y, the entity pair (e i ,e j ) belongs to the relation r ij The non-normalized score of , the larger the score value, the more likely the model believes that the entity pair belongs to the relationship r ij The higher the probability; ij is the predicted relation, r′ is any relation in the relation set R; After the entities and relationships are extracted, knowledge fusion is performed to disambiguate and align the entities and relationships in multi-source data, and a unified knowledge graph G in the civil affairs field is constructed. It is stored in the form of triples (s, r, o), where s is the subject, r is the relationship, and o is the object.

9. The civil affairs service question-answering method based on large model and knowledge graph retrieval enhancement according to claim 1 is characterized in that: Fast adaptation is achieved through a lightweight deployment strategy and a training-free prompt learning framework, avoiding the traditional model's reliance on large amounts of labeled data or complex fine-tuning. Specific implementations include: (1) Parameter quantization and compression: Advanced quantization technology is used to compress large language model parameters; (2) Multi-task parallel reasoning and prompt learning: For the sub-problem retrieval and generation process of multi-hop problems, the parallel computing capabilities of multi-core CPUs / GPUs are utilized to distribute the processing tasks of different sub-problems to multiple computing units for simultaneous execution; (3) Cross-domain adaptation without training: Adapt to different civil affairs business scenarios through dynamic prompt engineering, predefine common entity types and relationship sets, and quickly map to the knowledge graph of the new domain by adjusting the prompt words.

10. The civil affairs service question-answering method based on large model and knowledge graph retrieval enhancement according to claim 1 is characterized in that: Continuously optimize service quality through dynamic updates of the knowledge graph and a closed loop of user feedback, without requiring model training or parameter adjustments. (1) Incremental update mechanism of knowledge graph: Establish a special information collection and update mechanism to obtain the latest civil affairs information in real time from authoritative channels including the official website of the civil affairs department and the policy document library; automatically identify new triples through unsupervised entity relationship extraction technology to trigger incremental updates of the knowledge graph; (2) User feedback closed-loop optimization mechanism: Collect user feedback data on answers, including user evaluation of the accuracy, completeness, and logic of the answers, as well as suggestions for improving system functions. Use feedback data to optimize question decomposition strategies and retrieval parameters, and continuously improve system performance and user satisfaction.

Citation Information

Cited By

  • Content generation method and electronic equipment

    CN121071104A

  • Knowledge question and answer method based on multiple agents and heterogeneous data sources

    CN121092679A

  • BRT maintenance question and answer method and device based on expert feedback

    CN121119178A

  • Policy field-oriented reasoning generation method, apparatus and device, and storage medium

    CN121279459A

  • Policy field-oriented reasoning generation method, device and equipment and storage medium

    CN121279459B