Government affair hotline service knowledge graph construction method and system based on large language model

Through the semantic fusion and entity relationship extraction of government hotline Q&A data through the large language model, an efficient government affairs knowledge graph is built, the semantic implicitness and complex entity relationship problems in unstructured data processing are solved, and the intelligent upgrade of government services is realized.

CN120354923APending Publication Date: 2025-07-22NANJING YUNSHE INTELLIGENT TECH CO LTD +1

Patent Information

Application Number
CN202510816357.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate unstructured government hotline Q&A data, resulting in difficulty in extracting information, complex semantic implicitness and entity relationships, and low level of intelligence.

Method used

A chain thinking mechanism based on a large language model performs semantic fusion and logical refinement of unstructured government hotline Q&A data, uses step-by-step prompt strategy to identify knowledge entities, and extracts entity semantic relationships through semantic association analysis methods, constructs triple-tuple knowledge data and stores them in the graph database for dynamic updates and multi-dimensional retrieval.

Benefits of technology

It significantly improves the intelligence level of government services, improves the accuracy and efficiency of entity extraction and relationship extraction, realizes dynamic updates and multi-dimensional retrieval of knowledge graphs, and supports the construction of cross-field and multi-scenario government knowledge networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354923A_ABST
    Figure CN120354923A_ABST
Patent Text Reader

Abstract

The invention provides a government affair hotline service knowledge graph construction method and system based on a large language model, and relates to the technical field of natural language processing and government affair informationization, and the method comprises the steps: obtaining a government affair hotline service scene, and defining a mode layer of a multi-scene knowledge graph; unstructured government affair hotline question and answer data are collected, semantic fusion, logic extraction and ambiguity elimination are carried out on the data based on a chained thinking mechanism of a large language model, and a fused government affair hotline question and answer set is generated; driving a large language model to identify knowledge entities in the set by adopting a step-by-step prompt strategy, extracting semantic relationships among the knowledge entities by adopting a semantic association analysis method, constructing triple knowledge data based on the knowledge entities and the semantic relationships, storing the triple knowledge data into a graph database, and performing dynamic updating and multi-dimensional retrieval on the multi-scene knowledge graph to obtain a multi-scene knowledge graph set; therefore, a large language model and a knowledge graph technology can be fused, government affair data are efficiently integrated, and the intelligent level of government affair service is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of natural language processing and e-government informatization, and particularly to a method and system for constructing a knowledge graph of government service hotline based on a large language model. Background Art

[0002] With the digital transformation of government services, government service hotlines have become an important channel for the public to communicate with the government.

[0003] However, since the questions from citizens and the responses from the government often appear in the form of natural language, the data of government service hotlines lack clear entity and relationship annotations, which may lead to difficulties in information extraction. And traditional methods are difficult to accurately handle problems such as implicit logic and double attribution in question-and-answer pairs. Most of the existing knowledge graph construction technologies rely on structured data, such as policy documents, etc., and are difficult to adapt to unstructured question-and-answer data, thus resulting in low data utilization rate.

[0004] In the prior art, methods such as the BiLSTM-CRF model and RNN relation extraction are effective in some government service hotline scenarios, but their performance significantly degrades when facing complex semantics and implicit relationships. Large language models, such as ChatGPT, LLaMA, etc., although perform excellently in natural language processing, have not been systematically applied in constructing government knowledge graphs, and it is difficult to effectively integrate government data and improve service capabilities.

[0005] Therefore, it is necessary to provide a method and system for constructing a knowledge graph of government service hotline based on a large language model to solve the above technical problems. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method and system for constructing a knowledge graph of government service hotline based on a large language model, which are used to solve the problems existing in the prior art, such as implicit natural language semantics, complex entity relationships, low information extraction efficiency, and low intelligent level of government services.

[0007] The method for constructing a knowledge graph of government service hotline based on a large language model provided by the present invention, the construction method includes: Obtain the government service hotline scenario, and define the schema layer of the multi-scenario knowledge graph according to the government service hotline scenario; Collect unstructured government service hotline question-and-answer data, and perform semantic fusion, logic refinement, and ambiguity elimination on the unstructured government service hotline question-and-answer data based on the chain-of-thought mechanism of the large language model to generate a fused government service hotline question-and-answer set; Adopt a step-by-step prompting strategy to drive the large language model to identify knowledge entities in the integrated government service hotline Q&A set, and use a semantic association analysis method to extract the entity semantic relationships between the knowledge entities, and construct triple knowledge data based on the knowledge entities and the corresponding entity semantic relationships; Store the triple knowledge data in a graph database, and perform dynamic update and multi-dimensional retrieval on the multi-scenario knowledge graph based on the graph database.

[0008] Preferably, the schema layer includes 12 types of the knowledge entities and the entity semantic relationships between the knowledge entities.

[0009] Preferably, the 12 types of knowledge entities include person, submission time, type, matter, reply time, department, attribution, administrative region, policy, materials, contact phone number and other relevant entities; The entity semantic relationships between the knowledge entities include but are not limited to: The calling relationship between the person and the contact phone number; The belonging relationship between the matter and the department; The involved relationship between the matter and the policy; The required relationship between the matter and the materials; The located relationship between the department and the administrative region.

[0010] Preferably, the unstructured government service hotline Q&A data is sourced from multi-source heterogeneous government data sources, and the multi-source heterogeneous government data sources include public Q&A data platforms.

[0011] Preferably, the large language model is a pre-trained language model based on the Transformer architecture with a model parameter scale of not less than 6B, supporting multi-round dialogue and text generation functions, for performing natural language understanding, adapting to the domain requirements corresponding to the unstructured government service hotline Q&A data through fine-tuning, and performing knowledge extraction on the unstructured government service hotline Q&A data through local deployment.

[0012] Preferably, the chain-of-thought mechanism based on the large language model performs semantic fusion, logical refinement and ambiguity elimination on the unstructured government service hotline Q&A data to generate an integrated government service hotline Q&A set, specifically including: Use an ETL tool to perform data cleaning and format conversion on the collected unstructured government service hotline Q&A data; Perform semantic fusion on the unstructured government service hotline Q&A data based on the chain-of-thought mechanism of the large language model, determine the Q&A logical associations in the unstructured government service hotline Q&A data through a preset prompt template based on logical refinement, and generate an initial government service hotline Q&A set; Using the context understanding mechanism of the large language model, eliminate the double attribution or ambiguous expressions in the initial government hotline Q&A set, and generate the integrated government hotline Q&A set.

[0013] Preferably, the step-by-step prompting strategy is adopted to drive the large language model to identify the knowledge entities in the integrated government hotline Q&A set, specifically including: Adopt the step-by-step prompting strategy, and drive the large language model to identify the entity types in the integrated government hotline Q&A set through the first-round prompting template based on task decomposition; Drive the large language model to identify the entity attributes in the integrated government hotline Q&A set through the second-round prompting template based on task decomposition; Determine the knowledge entities in the integrated government hotline Q&A set based on the entity types and the entity attributes.

[0014] Preferably, the semantic association analysis method is adopted to extract the entity semantic relationships between the knowledge entities, specifically including: Obtain the context semantic association behaviors of the knowledge entities in the unstructured government hotline Q&A data; Adopt the semantic association analysis method, generate a candidate semantic relationship set through a preset prompting template based on relationship orientation, combine the context semantic association behaviors, screen the candidate semantic relationship set based on confidence to determine the entity semantic relationship, and optimize the entity semantic relationship based on the semantic similarity calculation method.

[0015] Preferably, the graph database adopts a native graph storage structure, supports dynamic update and multi-dimensional retrieval of the multi-scenario knowledge graph based on a preset path pattern through the attributes of nodes and edges.

[0016] A government hotline service knowledge graph construction system based on a large language model, the construction system includes: A knowledge graph definition module, used to obtain government hotline service scenarios and define the schema layer of the multi-scenario knowledge graph according to the government hotline service scenarios; A Q&A set generation module, used to collect unstructured government hotline Q&A data, perform semantic fusion, logical refinement and ambiguity elimination on the unstructured government hotline Q&A data based on the chain-of-thought mechanism of the large language model, and generate an integrated government hotline Q&A set; A triple construction module, used to adopt a step-by-step prompting strategy to drive the large language model to identify the knowledge entities in the integrated government hotline Q&A set, and adopt a semantic association analysis method to extract the entity semantic relationships between the knowledge entities, and construct triple knowledge data based on the knowledge entities and the corresponding entity semantic relationships; A knowledge graph update module for storing the triple knowledge data in a graph database and dynamically updating and multi-dimensionally retrieving the multi-scenario knowledge graph based on the graph database.

[0017] Compared with related technologies, the method and system for constructing a government affairs hotline service knowledge graph based on a large language model provided by the present invention have the following beneficial effects: The present invention obtains government affairs hotline service scenarios and defines the schema layer of the multi-scenario knowledge graph according to the government affairs hotline service scenarios; collects unstructured government affairs hotline Q&A data, performs semantic fusion, logical refinement, and ambiguity elimination on the unstructured government affairs hotline Q&A data based on the chain-of-thought mechanism of the large language model to generate a fused government affairs hotline Q&A set; adopts a step-by-step prompting strategy to drive the large language model to identify knowledge entities in the fused government affairs hotline Q&A set, and uses a semantic association analysis method to extract the entity semantic relationships between the knowledge entities, constructs triple knowledge data based on the knowledge entities and the corresponding entity semantic relationships; stores the triple knowledge data in the graph database, and dynamically updates and multi-dimensionally retrieves the multi-scenario knowledge graph based on the graph database, thereby enabling the integration of the large language model and knowledge graph technologies, efficiently integrating government affairs data, improving the intelligent level of government affairs services, and solving the problems of implicit natural language semantics, complex entity relationships, and low information extraction efficiency in traditional government affairs data processing methods.

[0018] The present invention utilizes the chain-of-thought ability of the large language model to perform semantic integration on unstructured government affairs hotline Q&A data, refine the Q&A interaction logic and eliminate ambiguity, improve semantic accuracy, ensure data consistency and accuracy, and effectively solve the problem of implicit natural language semantics in traditional government affairs data processing methods. The present invention identifies entities by adopting a step-by-step prompting strategy, extracts entity semantic relationships by using a semantic association analysis method, constructs triple data, and supports the dynamic update and multi-dimensional retrieval of the knowledge graph, thereby improving the accuracy and efficiency of entity recognition, ensuring the accuracy and reliability of relationship extraction, guaranteeing the timeliness and accuracy of the knowledge graph, enabling users to quickly obtain government affairs data according to different needs, and effectively solving the problems of complex entity relationships and low information extraction efficiency in the prior art. The present invention can significantly improve the construction efficiency of the knowledge graph through the large language model and has achieved remarkable results in entity extraction and relationship extraction tasks. Among them, the F1 score of entity extraction reaches 92.3%, and the F1 score of relationship extraction reaches 88.7%, which is about 15% higher than the F1 score of the traditional BiLSTM-CRF model. The present invention is adapted to various government affairs hotline service scenarios, supports the integration of multi-source data such as the mayor's mailbox and department hotlines, and constructs a cross-domain and multi-scenario government affairs knowledge network, which can effectively support the intelligent service upgrade of government affairs hotlines, provide efficient knowledge support for policy analysis, cross-departmental collaboration, and citizen consultation, and has broad application prospects. Description of the Drawings

[0019] Figure 1 The flowchart of the method for constructing a knowledge graph of government service hotline based on a large language model provided by an embodiment of the present invention; Figure 2 The system block diagram of the system for constructing a knowledge graph of government service hotline based on a large language model provided by an embodiment of the present invention; Figure 3 The schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] As Figure 1 shown, it is the flowchart of the method for constructing a knowledge graph of government service hotline based on a large language model provided by an embodiment of the present invention. Figure 1 The execution subject of the method shown can be a software and / or hardware device. The execution subject of the present application may include, but is not limited to, at least one of the following: user equipment, network equipment, etc. Among them, the user equipment may include, but is not limited to, a computer, a smart phone, a personal digital assistant (Personal Digital Assistant, abbreviated as: PDA), and the above-mentioned electronic equipment, etc. The network equipment may include, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing. Among them, cloud computing is a type of distributed computing, which consists of a group of loosely coupled computers forming a super virtual computer. This embodiment does not make any restrictions. It includes steps S1 to S4, specifically as follows: S1, obtain the government service hotline service scenario, and define the schema layer of the multi-scenario knowledge graph according to the government service hotline service scenario; Among them, the government service hotline service scenario refers to various government service interaction scenarios such as citizens' inquiries, complaints, and suggestions in the government service hotline. The schema layer is the top-level architecture design of the multi-scenario knowledge graph, which is used to standardize the data organization form.

[0022] In the field of government hotline services, citizens' interactive activities such as consultation, complaints, and suggestions through the government hotline platform constitute specific service scenarios, covering multiple business scenarios such as provident fund business handling, garbage classification policy consultation, and cross-departmental affairs coordination. As the top-level framework of knowledge organization, the model layer of the multi-scenario knowledge graph can systematically define 12 core entities such as people, submission time, matters, departments, policies, and materials based on the business logic of government services, and clarify the semantic associations between entities, such as the attribution relationship between matter entities and department entities, the involvement relationship between matter entities and policy entities, and the location relationship between department entities and administrative division entities. The model layer constructs a logical framework in the structured form of triples, providing a paradigm for data standardization in different business scenarios, ensuring that the knowledge graph architecture can adapt to the complexity and diversity of government hotline services.

[0023] S2, collecting unstructured government hotline question and answer data, and performing semantic fusion, logic extraction and ambiguity elimination on the unstructured government hotline question and answer data based on the chain thinking mechanism of the large language model to generate a fused government hotline question and answer set; It is understandable that unstructured government hotline Q&A data refers to citizen questions and government responses in the form of natural language, such as text Q&A pairs in the mayor's mailbox and hotline platform, which have the characteristics of semantic implicitness and logical fuzziness. The chain thinking mechanism of the large language model can make the implicit semantics in the unstructured Q&A data explicit through step-by-step logical deduction by using contextual reasoning ability, and refine the Q&A interaction logic to eliminate double attribution or ambiguous expressions, among which the Q&A interaction logic is such as the cause of the problem, the processing flow, etc., and double attribution is such as the ambiguity of the responsible party. The fused government hotline Q&A collection refers to the structured Q&A data set after cleaning and semantic integration, which provides standardized input for subsequent knowledge extraction.

[0024] In practical applications, most of the question and answer data generated by the government hotline exist in natural language text, which is unstructured data. It lacks a predefined data model and has problems such as semantic ambiguity, logical fuzziness, and double attribution. For example, when citizens consult about the status of housing provident fund sealing, they may involve ambiguous statements about the responsibility of the new and old units. Based on the chain thinking mechanism of the large language model, this type of data is deeply processed through step-by-step logical reasoning: first, the context understanding ability of the large language model is used to integrate the scattered question and answer pairs into semantically coherent texts, and clarify the entity association between questions and answers, such as associating the "housing provident fund sealing" problem with the reason that "the new unit has not completed the remittance"; secondly, through the preset logic extraction prompt template, the model is guided to extract deep logical relationships such as cause and effect and attribution in the question and answer data; finally, with the help of the model's context analysis ability, double attribution or ambiguous statements are eliminated, such as clarifying the responsible subject of the housing provident fund sealing status, generating a standardized fusion government hotline question and answer set, and providing high-quality structured input for subsequent knowledge extraction.

[0025] S3. Adopt a step-by-step prompting strategy to drive the large language model to identify the knowledge entities in the integrated government service hotline Q&A set, and use a semantic association analysis method to extract the entity semantic relationships between the knowledge entities, and construct triple knowledge data based on the knowledge entities and the corresponding entity semantic relationships; It should be noted that the step-by-step prompting strategy refers to a large language model interaction strategy based on task decomposition. A knowledge entity refers to a key information unit extracted from Q&A data. An entity semantic relationship refers to the semantic logical connection relationship between entities.

[0026] It can be understood that during entity recognition, the step-by-step prompting strategy interacts with the large language model in a task decomposition manner: the first-round prompt drives the large language model to identify the entities in the integrated Q&A set through a customized template. For example, it is determined that "provident fund seal status" in the text belongs to the "matter" entity, and "municipal provident fund management center" belongs to the "department" entity; the second-round prompt further refines the entity attributes, such as extracting the specific content of the matter, the contact information of the department, etc., so as to accurately identify the knowledge entities in stages.

[0027] During relationship extraction, the semantic association analysis method can combine the context semantic environment, generate candidate semantic relationship pairs through a relationship-oriented prompt template, and then determine the final entity semantic relationship based on a confidence screening mechanism. For example, the "call" relationship between "citizen" and "contact phone number", the "required" relationship between "matter" and "materials", etc. Based on the extracted knowledge entities and semantic relationships, triple knowledge data can be constructed and stored in a structured form of "head entity - relationship - tail entity" to ensure the standardization and computability of knowledge representation.

[0028] S4. Store the triple knowledge data in a graph database, and dynamically update and multi-dimensionally retrieve the multi-scenario knowledge graph based on the graph database.

[0029] It should be noted that the triple knowledge data is stored in a graph database, which adopts a native graph storage structure, using nodes to represent entities and edges to represent relationships, so as to efficiently process complex entity association relationships. The dynamic update mechanism supports adding new entity nodes and relationship edges in real time as new government affairs data is accessed. For example, when a new garbage classification policy document is added, the relationship connection of "matter - related - policy" is automatically updated to ensure that the knowledge graph keeps pace with the development of the government affairs service scenario. The multi - dimensional retrieval function relies on the path query ability of the graph database and supports users to retrieve associated knowledge according to different dimensions such as departments, administrative regions, policy types, etc. For example, query the "garbage classification" matters responsible by the "Urban Management Bureau" and their corresponding policy documents and processing procedures, providing efficient knowledge support for cross - departmental collaborative office and accurate citizen consultation, thus realizing the dynamic maintenance and flexible application of the knowledge graph.

[0030] In the specific implementation process, the schema layer includes 12 types of the knowledge entities and the entity semantic relationships between the knowledge entities.

[0031] The 12 types of the knowledge entities include person, submission time, type, matter, reply time, department, attribution, administrative region, policy, material, contact phone number and other related entities; The entity semantic relationships between the knowledge entities include but are not limited to: The calling relationship between the person and the contact phone number; The belonging relationship between the matter and the department; The related relationship between the matter and the policy; The required relationship between the matter and the material; The located relationship between the department and the administrative region.

[0032] In practical applications, the schema layer of the knowledge graph serves as the top - level framework for organizing government affairs hotline service knowledge, systematically defining 12 types of core knowledge entities and semantic associations, and constituting the basic paradigm for structuring government affairs data. Based on the logical characteristics of the government affairs hotline business scenario, this schema layer abstracts key information units in unstructured Q&A data into standardized entity types and constructs a logical network between entities through semantic relationships, providing a unified specification for subsequent knowledge extraction, storage and application.

[0033] In the scenario of government service hotline, the entity of person refers to the participants in government interactions, including specific objects such as citizens and government staff; the entities of submission time and response time respectively refer to the time points of question submission and official response, constituting the key attributes of the time sequence dimension; the entity of matter covers the specific contents of citizens' consultations and complaints, such as business matters like "querying the sealed state of provident fund" and "complaining about insufficient garbage classification facilities"; the entity of department corresponds to the administrative agencies handling government affairs, such as the Municipal Provident Fund Management Center, the Urban Management Bureau, etc.; the entity of attribution is used to describe the cause or responsible entity of the matter, such as "the new unit has not completed the remittance business"; the entity of administrative division limits the geographical scope of the department or matter, such as administrative regions at the provincial, municipal, and district levels; the entity of policy is associated with the regulatory documents involved in the matter; the entity of material refers to the certification documents required for handling the matter, such as ID cards, provident fund payment records, etc.; the entity of contact phone number refers to the information providing communication channels, such as the consultation phone number of the department; the entity of type refers to the classification identifier of the matter, such as consultation type, complaint type, suggestion type, etc.; other related entities, as an extended category, accommodate auxiliary information units outside the above classifications.

[0034] Furthermore, the semantic relationships between entities are constructed through the business logic of government services, forming a structured association network. Among them, the call relationship establishes the communication association between the entity of person and the entity of contact phone number, such as "citizens call the department consultation phone number", clarifying the channel of government interaction; the attribution relationship defines the responsibility association between the entity of matter and the entity of department, such as "garbage classification matters are attributed to the Urban Management Bureau", realizing the subject positioning of matter handling; the involvement relationship connects the entity of matter and the entity of policy, such as "the matter of provident fund sealing involves the provident fund management policy", supporting the query of policy basis; the need relationship represents the conditional association between the entity of matter and the entity of material, such as "handling provident fund matters requires submitting identity certification materials", clarifying the material requirements for business handling; the location relationship constructs the spatial association between the entity of department and the entity of administrative division, such as "the Municipal Provident Fund Management Center is located in a certain district of a certain city", realizing the knowledge organization in the geographical dimension.

[0035] This model layer transforms the natural language information in the government service hotline into structured knowledge units through standardized entity and relationship modeling, providing a clear goal for the knowledge extraction of large language models, and at the same time supporting the triple storage and multi-dimensional retrieval of the graph database. Experimental data shows that the knowledge graph constructed based on this model layer has high accuracy in entity extraction and relationship extraction tasks, and can effectively solve the problems of semantic implicit and complex relationships in traditional government data processing, providing underlying knowledge support for the intelligent upgrade of government services, such as cross-departmental collaboration, policy analysis, and accurate citizen consultation.

[0036] The unstructured government service hotline Q&A data comes from multi-source heterogeneous government data sources, and the multi-source heterogeneous government data sources include public Q&A data platforms.

[0037] In practical applications, the unstructured Q&A data in the government service hotline scenario originates from a multi-source heterogeneous government data source system, which covers various types and structures of government data collection channels. Among them, the public Q&A data platform, as the core data source, includes interactive carriers such as the mayor's mailbox, the exclusive hotline platforms of various government departments, and the public Q&A sections of government service websites. The citizen consultation texts and government reply contents stored in it exist in the form of natural language, lacking a predefined data model and structured format.

[0038] Due to the characteristics of semantic implicit and logical dispersion of such multi-source heterogeneous data, it is necessary to perform semantic fusion and logical extraction on it through the chain-of-thought mechanism of the large language model, realizing the transformation from unstructured text to a structured knowledge graph, providing basic data support for subsequent entity extraction and relationship construction.

[0039] The large language model is a pre-trained language model based on the Transformer architecture with a model parameter scale of not less than 6B, supporting multi-round dialogue and text generation functions, used for performing natural language understanding, adapting to the domain requirements corresponding to the unstructured government service hotline Q&A data through fine-tuning, and performing knowledge extraction on the unstructured government service hotline Q&A data through local deployment.

[0040] It should be noted that this large language model adopts the Transformer architecture and has a pre-trained basis with a parameter scale of not less than 6B (6 billion). It can integrate multi-round dialogue interaction and text generation functions, and perform semantic parsing on the unstructured Q&A data in the government service hotline scenario through natural language understanding technology.

[0041] In order to adapt to the special needs of the government affairs field, this model needs to carry out domain fine-tuning based on the government service hotline Q&A data, and conduct special training on professional terms, administrative processes, and Q&A logics in business scenarios such as provident fund handling and policy consultation to improve the accuracy of entity recognition and relationship extraction.

[0042] This model adopts the local deployment mode, builds an inference engine in the government affairs intranet environment to realize the knowledge extraction and processing of unstructured data. This deployment method not only meets the security management requirements of government affairs data, but also can improve the data processing efficiency through customized computing resource configuration, providing accurate knowledge extraction support for subsequent semantic fusion and triple construction.

[0043] Through the combination of pre-training capabilities and government affairs domain fine-tuning, this model effectively solves the problem of natural language semantic implication in traditional government affairs data processing methods, providing core technical support for the construction of the government service hotline service knowledge graph.

[0044] The chain thinking mechanism based on the large language model performs semantic fusion, logic extraction and ambiguity elimination on the unstructured government hotline question and answer data to generate a fused government hotline question and answer set, which specifically includes: Use ETL tools to clean and format the collected unstructured government hotline Q&A data; The chain thinking mechanism based on the large language model performs semantic fusion on the unstructured government hotline question and answer data, determines the question and answer logical association in the unstructured government hotline question and answer data through a preset prompt template based on logic extraction, and generates an initial government hotline question and answer set; The context understanding mechanism of the large language model is used to eliminate double attribution or ambiguous expressions in the initial government hotline question and answer set, thereby generating the fused government hotline question and answer set.

[0045] For the collected government hotline Q&A data, we first use ETL (extraction, transformation, loading) tools for preprocessing. This process uses data cleaning technology to remove noise information, such as repeated text, format error content, etc., and converts unstructured text into a standardized format. For example, the Q&A data on different platforms is unified into a binary structure of questions and answers to provide standardized input for subsequent semantic processing. This stage solves the consistency problem of multi-source heterogeneous data and ensures that the large language model can effectively process various types of government interaction texts.

[0046] Then, the large language model based on the Transformer architecture can perform deep semantic processing on the cleaned question and answer data through a chain thinking mechanism. The model uses preset logic extraction prompt templates, such as "Please integrate the core logic of the following question and answer data and clarify the entity associations", to drive the step-by-step reasoning process: first identify the key entities in the questions and answers, and then extract the logical relationship between the entities through contextual association analysis, and integrate the scattered question and answer pairs into semantically coherent text paragraphs. For example, when dealing with provident fund sealing consultation, the model will associate the reply "the new unit has not completed the remittance" with the citizen's question, generate an integrated text containing complete causal logic, and form an initial government hotline question and answer set. This process uses the sequence generation ability of the large language model to transform the implicit semantics in natural language into an explicit logical structure.

[0047] Regarding the possible double attributions in the initial Q&A set, such as ambiguous responsible entities or semantic ambiguity issues, the large language model optimizes them through the context understanding mechanism. Based on the semantic representation ability pre-trained with a large amount of government data, the model can identify semantic conflict points in the text. For example, in the expression "the provident fund seal is attributed to the new unit or the old unit", by analyzing the payment process and time nodes, it clarifies that the responsible entity is the new unit to eliminate the expression ambiguity. This process undergoes multiple rounds of semantic reasoning and logical verification to ensure the uniqueness and accuracy of each entity relationship in the integrated Q&A data, and finally generates a non-ambiguous integrated government hotline Q&A set.

[0048] Through the collaboration of the ETL tool and the large language model, a conversion channel from unstructured data to structured knowledge is constructed. The step-by-step reasoning feature of the chain of thought mechanism enables it to effectively handle complex semantics in government data, such as cross-departmental business logics and policy clause references, while the context understanding ability solves the ambiguity problems that are difficult to handle by traditional methods. Experimental data shows that the integrated Q&A set generated after this process provides high-quality input for subsequent entity extraction and relationship extraction, effectively solving the technical bottlenecks of semantic implication and complex logic in government hotline data, and laying a data foundation for the accurate construction of the knowledge graph.

[0049] The step-by-step prompting strategy is adopted to drive the large language model to identify the knowledge entities in the integrated government hotline Q&A set, specifically including: Adopting the step-by-step prompting strategy, through the first-round prompting template based on task decomposition, drive the large language model to identify the entity types in the integrated government hotline Q&A set; Through the second-round prompting template based on task decomposition, drive the large language model to identify the entity attributes in the integrated government hotline Q&A set; Determine the knowledge entities in the integrated government hotline Q&A set based on the entity types and the entity attributes.

[0050] In practical applications, the step-by-step prompting strategy can guide the large language model to extract the knowledge entities in the integrated Q&A set in stages through the interactive design of task decomposition.

[0051] Among them, the first-round prompt template aims at entity type recognition. Through a predefined domain knowledge framework, it guides the large language model to perform semantic segmentation and category mapping on the text in the fusion Q&A set. For example, for the provident fund business Q&A text "When a unit processes the addition of a new provident fund account, it is prompted that 'there is an unfinished contribution settlement'.", the first-round template will be guided by "Please identify the government entities in the following text", driving the model to classify "new provident fund account addition" as an "event entity" and "unit" as a "person entity" based on the semantic representation ability of the Transformer architecture. This process utilizes the entity generalization and recognition ability formed by the large language model during pre-training, combined with the type mapping rules fine-tuned in the government affairs domain, to achieve the preliminary positioning of entity categories and lay a classification foundation for subsequent attribute extraction.

[0052] The second-round prompt template focuses on the in-depth exploration of entity attributes. By decomposing tasks, it deconstructs complex entity features into quantifiable attribute fields. Taking the "event entity" as an example, the second-round template will be guided by "Please extract attributes such as the specific content, relevant policies, and handling materials of the above event", driving the large language model to analyze the detailed information in the text. Still taking the provident fund scenario as an example, the model will extract "unfinished contribution settlement business" as the specific content of the event and "T1" as the contact phone number attribute from the reply text "Need to contact the new unit to confirm the unfinished contribution settlement business, contact phone number T1".

[0053] Based on the entity types identified in the first round and the attribute information extracted in the second round, the final knowledge entities are determined through logical integration. This determination process first limits the attribute set according to the entity type, and then eliminates invalid recognition results through attribute integrity verification. For example, if the model identifies an entity as a "department" but does not extract the "administrative division" attribute, it will trigger a secondary prompt for completion; if the "specific content" attribute of the "event entity" is missing, the original text will be returned for re-parse. This mechanism can ensure that the structural degree of the knowledge entities meets the construction requirements of the knowledge graph. For example, it generates a "provident fund sealing status" entity containing complete attributes such as "event type", "specific content", and "involved department", providing a standardized data unit for subsequent triple construction.

[0054] The method of extracting the entity semantic relationship between the knowledge entities by using semantic association analysis specifically includes: Obtain the context semantic association behavior of the knowledge entities in the unstructured government affairs hotline Q&A data; Adopt the semantic association analysis method. Through a preset prompt template based on the relationship orientation, combined with the context semantic association behavior, generate a candidate semantic relationship set, screen the candidate semantic relationship set based on the confidence level to determine the entity semantic relationship, and optimize the entity semantic relationship based on the semantic similarity calculation method.

[0055] For the knowledge entities identified by the step-by-step prompting strategy, semantic association analysis first obtains the contextual semantic association behavior of these entities in the government question and answer text. This process is based on the contextual understanding ability of the large language model, analyzing the grammatical structure, semantic dependency and logical context of the text surrounding the entity. For example, in the text "Citizens consulted about the status of housing provident fund sealing, and the reply pointed out that it was caused by the failure of the new unit to complete the remittance", the positional relationship between "housing provident fund sealing status" and "new unit" in the causal logic is captured, as well as the association role of "remittance business" as an intermediate variable. Through the attention mechanism of the Transformer architecture, the large language model can quantify the strength of semantic associations between entities and provide contextual feature support for subsequent relationship extraction.

[0056] Based on predefined relationship types, a relationship-oriented preset prompt template is used to interact with the large language model, guiding the model to generate a set of candidate semantic relationships in combination with contextual semantic association behaviors. For example, if the input template is "analyze the logical relationship between 'housing fund sealed status' and 'new unit' in the text, and output possible relationship types", the model will generate candidate relationships such as "attributed to" and "caused by..." based on contextual features. This process uses the conditional generation capability of the large language model to map implicit relationships in natural language into standardized relationship types, forming a set of triples containing head entities, tail entities, and candidate relationships, providing multiple options for subsequent screening.

[0057] Furthermore, the generated candidate semantic relationship set can be initially filtered by the confidence threshold, such as setting 0.85 as the confidence screening standard to eliminate low-confidence relationship predictions. The confidence calculation integrates multi-dimensional features such as entity co-occurrence frequency, context matching and model generation probability. For example, the confidence of the "matter-attribution-department" relationship will be weighted in combination with government affairs domain knowledge such as departmental division of responsibilities and matter handling procedures. Subsequently, the semantic similarity calculation method is used to optimize the retained relationships: by comparing the semantic distance between the candidate relationship and the standard relationship library in the government affairs field, such as using cosine similarity for measurement, the accuracy of the relationship expression is adjusted, for example, "caused by..." is standardized as "attributed to", to ensure the consistency of the relationship expression and domain adaptability.

[0058] Through multi-layer processing of context perception, prompt guidance, confidence screening and semantic optimization, the problem of implicit relationships and variable expressions in government data can be effectively solved. For example, in the provident fund business scenario, the relationship of "provident fund sealed status-attributed to-new unit" can be accurately extracted, and the responsibility association of "garbage classification matters-attributed to-Urban Management Bureau" can be identified in the cross-departmental complaint scenario.

[0059] Through experimental verification of government service hotline Q&A data, the knowledge graph construction method based on large language models can achieve an F1 score of 92.3% in entity extraction tasks and 88.7% in relation extraction tasks. Compared with the traditional BiLSTM-CRF model, the improvement in the F1 score exceeds 15%. The experimental results show that the semantic fusion ability of large language models for unstructured government data through the chain of thought mechanism, and the precise drive of the step-by-step prompting strategy in entity and relation extraction, significantly optimize the accuracy and efficiency of knowledge graph construction.

[0060] Compared with traditional natural language processing methods, large language models show stronger context understanding ability when dealing with complex logics such as implicit semantics and dual attribution in government data. Their dynamic reasoning process effectively solves the recognition bottleneck of traditional models in semantic ambiguity scenarios. Experimental data further confirms that through the local deployment and domain fine-tuning of large language models, the structured processing efficiency of government service hotline data has been significantly improved, providing quantifiable technical support for the intelligent upgrade of government services. This method demonstrates high application value in scenarios such as policy analysis, cross-departmental collaboration, and citizen consultation.

[0061] The graph database adopts a native graph storage structure, supporting dynamic update and multi-dimensional retrieval of the multi-scenario knowledge graph based on a preset path pattern through the attributes of nodes and edges.

[0062] It should be noted that this graph database adopts a native graph storage structure, visualizing knowledge entities as nodes and modeling semantic relationships between entities as edges. Both nodes and edges carry attribute fields to record entity features and relationship features, such as department names, policy effective times, relationship confidence levels, etc.

[0063] Based on predefined knowledge association path patterns, such as the logical link of "government affairs item - attribution - administrative department - location - administrative division", this graph database supports multi-dimensional knowledge retrieval through the combination of node and edge attributes. For example, query all the items and related policies responsible by a certain department according to the "department name" attribute, or filter government affairs knowledge within a certain region through the "administrative division" attribute. The dynamic update mechanism allows real-time insertion of new nodes, such as adding new policy documents, or creating edges, such as establishing the attribution relationship between a new item and a department. Thus, through incremental data operations, it can maintain the synchronous iteration of the knowledge graph and government service scenarios, providing efficient graph-structured data access capabilities for cross-departmental collaborative governance, policy traceability analysis, and precise citizen consultation.

[0064] As Figure 2 shown, it is the system block diagram of the government service hotline service knowledge graph construction system provided by the embodiment of the present invention. The construction system includes: A knowledge graph definition module, configured to obtain government affairs hotline service scenarios and define the schema layer of a multi-scenario knowledge graph according to the government affairs hotline service scenarios; A question-and-answer set generation module, configured to collect unstructured government affairs hotline question-and-answer data, perform semantic fusion, logical refinement, and ambiguity elimination on the unstructured government affairs hotline question-and-answer data based on the chain-of-thought mechanism of a large language model, and generate a fused government affairs hotline question-and-answer set; A triple construction module, configured to adopt a step-by-step prompting strategy to drive the large language model to identify knowledge entities in the fused government affairs hotline question-and-answer set, and extract the entity semantic relationships between the knowledge entities by using a semantic association analysis method, and construct triple knowledge data based on the knowledge entities and the corresponding entity semantic relationships; A knowledge graph update module, configured to store the triple knowledge data in a graph database, and perform dynamic update and multi-dimensional retrieval on the multi-scenario knowledge graph based on the graph database.

[0065] Figure 2 The device in the illustrated embodiment can correspondingly be used to execute Figure 1 the steps in the method embodiment shown, and its implementation principle and technical effects are similar, and will not be elaborated here.

[0066] An electronic device includes a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the steps of the method for constructing a government affairs hotline service knowledge graph based on a large language model as described in any one of the above.

[0067] As Figure 3 shown, it is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. The electronic device 30 includes: a processor 31, a memory 32, and a computer program; wherein The memory 32 is configured to store the computer program, and the memory can also be a flash memory. The computer program is, for example, an application program, a functional module, etc. for implementing the above method.

[0068] The processor 31 is configured to execute the computer program stored in the memory to implement each step executed by the device in the above method. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.

[0069] Optionally, the memory 32 can be either independent or integrated with the processor 31.

[0070] When the memory 32 is a device independent of the processor 31, the device may further include: A bus 33, configured to connect the memory 32 and the processor 31.

[0071] A readable storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the steps of the method for constructing a government affairs hotline service knowledge graph based on a large language model as described in any one of the above.

[0072] Among them, the readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general or special-purpose computer. For example, the readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). In addition, the ASIC can be located in the user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0073] The present invention also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of the device can read the execution instructions from the readable storage medium, and the execution of the execution instructions by at least one processor causes the device to implement the methods provided by the above various embodiments.

[0074] In the embodiments of the above device, it should be understood that the processor can be a central processing unit (CPU for short), and can also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the present invention can be directly embodied as being completed by the execution of a hardware processor, or by a combination of hardware and software modules in the processor.

[0075] Through the introduction of the above embodiments, the present invention provides a method and system for constructing a knowledge graph of government affairs hotline services based on a large language model. By obtaining government affairs hotline service scenarios and defining the schema layer of a multi-scenario knowledge graph according to the government affairs hotline service scenarios; collecting unstructured government affairs hotline Q&A data, and performing semantic fusion, logic refinement, and ambiguity elimination on the unstructured government affairs hotline Q&A data based on the chain-of-thought mechanism of the large language model to generate a fused government affairs hotline Q&A set; adopting a step-by-step prompting strategy to drive the large language model to identify knowledge entities in the fused government affairs hotline Q&A set, and using a semantic association analysis method to extract the entity semantic relationships between the knowledge entities, constructing triple knowledge data based on the knowledge entities and the corresponding entity semantic relationships; storing the triple knowledge data in a graph database, and dynamically updating and multi-dimensionally retrieving the multi-scenario knowledge graph based on the graph database, so as to integrate the large language model and knowledge graph technologies, efficiently integrate government affairs data, improve the intelligent level of government affairs services, and solve the problems of implicit natural language semantics, complex entity relationships, and low information extraction efficiency in traditional government affairs data processing methods.

[0076] The present invention utilizes the chain-of-thought ability of the large language model to perform semantic integration on unstructured government affairs hotline Q&A data, refine the Q&A interaction logic and eliminate ambiguity, thereby improving semantic accuracy, ensuring data consistency and accuracy, and effectively solving the problem of implicit natural language semantics in traditional government affairs data processing methods. The present invention identifies entities by adopting a step-by-step prompting strategy, extracts entity semantic relationships by using a semantic association analysis method, constructs triple data, and supports the dynamic update and multi-dimensional retrieval of the knowledge graph, thereby improving the accuracy and efficiency of entity recognition, ensuring the accuracy and reliability of relationship extraction, guaranteeing the timeliness and accuracy of the knowledge graph, enabling users to quickly obtain government affairs data according to different needs, and effectively solving the problems of complex entity relationships and low information extraction efficiency in the prior art. The present invention can significantly improve the construction efficiency of the knowledge graph through the large language model and has achieved remarkable results in entity extraction and relationship extraction tasks. Among them, the F1 score of entity extraction reaches 92.3%, and the F1 score of relationship extraction reaches 88.7%, which is about 15% higher than the F1 score of the traditional BiLSTM-CRF model. The present invention is adapted to a variety of government affairs hotline service scenarios, supports the integration of multi-source data such as the mayor's mailbox and department hotlines, and then constructs a cross-domain and multi-scenario government affairs knowledge network, which can effectively support the intelligent service upgrade of government affairs hotlines, provide efficient knowledge support for policy analysis, cross-departmental collaboration, and citizen consultation, and has broad application prospects.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a knowledge graph of government hotline services based on large language models, characterized in that, The construction method includes: Obtain the government hotline service scenarios, and define the schema layer of the multi-scenario knowledge graph according to the government hotline service scenarios; Collect unstructured government hotline Q&A data, and perform semantic fusion, logical refinement, and ambiguity elimination on the unstructured government hotline Q&A data based on the chain-of-thought mechanism of the large language model to generate a fused government hotline Q&A set; Adopt a step-by-step prompting strategy to drive the large language model to identify the knowledge entities in the fused government hotline Q&A set, and use a semantic association analysis method to extract the entity semantic relationships between the knowledge entities, and construct triple knowledge data based on the knowledge entities and the corresponding entity semantic relationships; Store the triple knowledge data in a graph database, and perform dynamic update and multi-dimensional retrieval on the multi-scenario knowledge graph based on the graph database.

2. The method for constructing a knowledge graph of government affairs hotline services based on a large language model according to claim 1, wherein The schema layer includes 12 types of the knowledge entities and the entity semantic relationships between the knowledge entities.

3. The method for constructing a knowledge graph of government affairs hotline services based on a large language model according to claim 2, wherein, The 12 types of the knowledge entities include person, submission time, type, matter, reply time, department, attribution, administrative division, policy, materials, contact phone number, and other related entities; The entity semantic relationships between the knowledge entities include but are not limited to: The calling relationship between the person and the contact phone number; The belonging relationship between the matter and the department; The involved relationship between the matter and the policy; The required relationship between the matter and the materials; The located relationship between the department and the administrative division.

4. The method for constructing a government hotline service knowledge graph based on a large language model according to claim 1, wherein, The unstructured government hotline Q&A data is sourced from multi-source heterogeneous government data sources, and the multi-source heterogeneous government data sources include public Q&A data platforms.

5. The method for constructing a government affairs hotline service knowledge graph based on a large language model according to claim 1, wherein The large language model is a pre-trained language model based on the Transformer architecture with a model parameter scale of not less than 6B, supporting multi-turn conversations and text generation functions, used for performing natural language understanding, adapting to the domain requirements corresponding to the unstructured government hotline Q&A data through fine-tuning, and performing knowledge extraction on the unstructured government hotline Q&A data through local deployment.

6. The method for constructing a government affairs hotline service knowledge graph based on a large language model according to claim 1, wherein Performing semantic fusion, logical refinement, and ambiguity elimination on the unstructured government hotline Q&A data based on the chain-of-thought mechanism of the large language model to generate a fused government hotline Q&A set specifically includes: Use an ETL tool to perform data cleaning and format conversion on the collected unstructured government hotline Q&A data; Perform semantic fusion on the unstructured government hotline Q&A data based on the chain-of-thought mechanism of the large language model, and determine the Q&A logical associations in the unstructured government hotline Q&A data through a preset prompting template based on logical refinement to generate an initial government hotline Q&A set; Utilize the context understanding mechanism of the large language model to eliminate double attributions or ambiguous expressions in the initial government hotline Q&A set to generate the fused government hotline Q&A set.

7. The method for constructing a knowledge graph of government affairs hotline services based on a large language model according to claim 1, wherein, Adopting a step-by-step prompting strategy to drive the large language model to identify the knowledge entities in the fused government hotline Q&A set specifically includes: Adopt the step-by-step prompting strategy, and drive the large language model to identify the entity types in the integrated government service hotline Q&A set through the first-round prompting template based on task decomposition; Drive the large language model to identify the entity attributes in the integrated government service hotline Q&A set through the second-round prompting template based on task decomposition; Determine the knowledge entities in the integrated government service hotline Q&A set based on the entity types and the entity attributes.

8. The method for constructing a government affairs hotline service knowledge graph based on a large language model according to claim 1, characterized in that, The adoption of the semantic association analysis method to extract the entity semantic relationships between the knowledge entities specifically includes: Obtain the context semantic association behaviors of the knowledge entities in the unstructured government service hotline Q&A data; Adopt the semantic association analysis method, generate a candidate semantic relationship set through a preset prompting template based on relationship orientation, combine the context semantic association behaviors, screen the candidate semantic relationship set based on confidence to determine the entity semantic relationship, and optimize the entity semantic relationship based on the semantic similarity calculation method.

9. The method for constructing a government affairs hotline service knowledge graph based on a large language model according to claim 1, wherein, The graph database adopts a native graph storage structure, and supports dynamic update and multi-dimensional retrieval of the multi-scenario knowledge graph based on a preset path pattern through the attributes of nodes and edges.

10. A government affairs hotline service knowledge graph construction system based on a large language model, which is applied to the method for constructing a government affairs hotline service knowledge graph based on a large language model according to any one of claims 1-9, and is characterized in that, The construction system includes: A knowledge graph definition module, which is used to obtain the government service hotline service scenario and define the schema layer of the multi-scenario knowledge graph according to the government service hotline service scenario; A Q&A set generation module, which is used to collect unstructured government service hotline Q&A data, perform semantic fusion, logical refinement, and ambiguity elimination on the unstructured government service hotline Q&A data based on the chain-of-thought mechanism of the large language model, and generate an integrated government service hotline Q&A set; A triple construction module, which is used to adopt the step-by-step prompting strategy, drive the large language model to identify the knowledge entities in the integrated government service hotline Q&A set, adopt the semantic association analysis method to extract the entity semantic relationships between the knowledge entities, and construct triple knowledge data based on the knowledge entities and the corresponding entity semantic relationships; A knowledge graph update module, which is used to store the triple knowledge data in the graph database, and perform dynamic update and multi-dimensional retrieval of the multi-scenario knowledge graph based on the graph database.

Citation Information

Patent Citations

  • Medical insurance information question and answer method and device

    CN114153994A

  • Government affair service field multi-strategy fusion dialogue method based on knowledge graph

    CN116628172A

  • Method for providing high-quality data for multi-mode large model system

    CN117743315A

  • Government affair question and answer method based on knowledge graph and large language model

    CN118939761A

  • Knowledge graph question and answer method and system based on large language model

    CN119829720A

Cited By

  • Method and system for constructing Neo4j knowledge graph based on LLM natural language

    CN120744136A

  • Government affair policy question and answer method based on knowledge graph and related equipment

    CN120849569A

  • Government affair text auditing method and system based on knowledge graph reasoning

    CN120874852A

  • A government affair text auditing method and system based on knowledge graph reasoning

    CN120874852B

  • Government affair retrieval enhancement generation framework and method based on knowledge graph

    CN121255999A