A multi-source knowledge fusion administrative punishment intelligent question and answer model construction method and system
By constructing an intelligent question-and-answer model for administrative penalties that integrates multi-source knowledge, the problem of market supervision data integration has been solved, enabling precise and traceable question-and-answer support, improving the accuracy and interpretability of question-and-answer results, adapting to changes in laws and regulations, and providing efficient intelligent question-and-answer services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI MUNICIPAL ADMINISTRATION FOR MARKET REGULATION INFORMATION APPL RES CENT (SHANGHAI FOOD SAFETY TECH APPL CENT SHANGHAI MUNICIPAL ADMINISTRATION FOR MARKET REGULATION ARCHIVES)
- Filing Date
- 2026-03-04
- Publication Date
- 2026-06-05
AI Technical Summary
Existing technologies struggle to effectively integrate market supervision data scattered across different systems. Traditional search engines suffer from fragmented search results, untraceable answers, and a lack of semantic reasoning capabilities in the field of administrative penalties. Relying solely on large language models to generate answers in specialized fields lacks accuracy and interpretability.
This paper constructs an intelligent question-answering model for administrative penalties by integrating multi-source knowledge. It builds a multi-source heterogeneous data corpus covering the field of administrative penalties, extracts knowledge graphs based on domain ontology guidance and large language model, designs an intelligent question-answering architecture that coordinates retrieval and generation channels, and introduces knowledge evolution and optimization mechanisms to achieve continuous data updates and adaptability.
It achieves semantic unification and deep correlation of market supervision data, improves the accuracy and interpretability of question and answer results, possesses a complete chain of evidence and legal provisions, and can continuously adapt to changes in laws and regulations, providing high-precision, traceable and timely intelligent question and answer support.
Smart Images

Figure CN122154922A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and market supervision technology, and in particular relates to a method and system for constructing an intelligent question-and-answer model for administrative penalties that integrates multi-source knowledge. Background Technology
[0002] With the improvement of market supervision informatization, platforms such as the State Administration for Market Regulation, the National Standards Information Public Service Platform, and Credit China have accumulated massive amounts of structured and unstructured regulatory data. However, this data is scattered across different systems, with heterogeneous formats, making it difficult for non-professional users to efficiently retrieve and understand it. Traditional search engines, when faced with questions such as "Has a company been penalized?" or "What standards should a product meet?", suffer from fragmented search results, untraceable answers, and a lack of semantic reasoning capabilities.
[0003] In recent years, large language models (LMs) have demonstrated powerful capabilities in natural language understanding and generation. However, they suffer from limitations in specialized fields, such as "illusions," lack of factual basis, and inability to trace information sources. Relying solely on large models to generate answers is insufficient to meet the high requirements of accuracy, interpretability, and compliance in market regulation.
[0004] While existing technologies combine knowledge graphs with large-scale models, most are limited to general domains and lack ontology modeling, multi-source data fusion mechanisms, and dynamic update capabilities for administrative penalty scenarios, resulting in limited applicability. Therefore, there is an urgent need for an intelligent model construction method that can integrate multi-source market supervision knowledge and supervisory regulations to achieve accurate and traceable question answering. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention proposes a method and system for constructing an intelligent question-and-answer model for administrative penalties based on multi-source knowledge fusion, thereby resolving the issues present in the existing technologies.
[0006] Firstly, to achieve the above objectives, this invention provides a method for constructing an intelligent question-answering model for administrative penalties based on multi-source knowledge fusion, comprising the following steps: Construct a multi-source heterogeneous data corpus covering the field of administrative penalties; Based on domain ontology guidance and large language model extraction, an administrative penalty knowledge graph is constructed from the corpus. Design an intelligent question-answering architecture that coordinates retrieval and generation channels, and responds to user questions based on the knowledge graph; A knowledge evolution and optimization mechanism is established to continuously update the knowledge graph and the intelligent question-answering architecture.
[0007] Optionally, the process of constructing the corpus includes: We collect structured and unstructured data from regulatory announcements, standards information, credit information disclosures, and public opinion monitoring channels, including laws and regulations, enterprise information, product sampling results, penalty information, and standard documents. The collected data is formatted and semantically cleaned, text and table content are extracted, semantic deduplication and standardization are performed, and a domain corpus with a unified format is formed.
[0008] Optionally, the process of constructing the administrative penalty knowledge graph includes: Based on the business logic of the administrative penalty field, we define the core entity categories and their semantic relationships, and establish a formalized domain semantic framework. Based on the domain semantic framework, a prompt template is generated to guide the large language model to extract entities and relations from the unstructured text of the corpus and form knowledge triples. The extracted knowledge triples are subjected to cross-source entity semantic alignment and fusion, and the fused knowledge is stored in a graph database in the form of nodes and relations.
[0009] Optionally, the cross-source entity semantic alignment and fusion process includes: Semantic vector models are used to encode entity names from different sources into vectors; Calculate the semantic similarity between entity vectors. When the similarity exceeds a set threshold, they are determined to be the same semantic entity and aggregated. Consistency checks are performed on the aggregation results based on predefined semantic conflict rules.
[0010] Optionally, the process by which the intelligent question-answering architecture responds to user questions includes: In the retrieval channel, the semantics and intent of the user's question are parsed and mapped to a structured query of the knowledge graph, and relevant evidence chains are retrieved and extracted; In the generation channel, the user question, the chain of evidence, and relevant legal provisions are input into the large language model to generate a natural language answer with citations. The semantic consistency and confidence of the evidence chain and the natural language response are evaluated, and the final answer is output based on the evaluation results.
[0011] Optionally, after the process of generating a natural language answer, the method further includes: using a fact-checking module to perform semantic consistency verification on each claim in the generated answer and its corresponding cited content.
[0012] Optionally, the knowledge evolution and optimization mechanism includes: An automatic incremental update mechanism is established to periodically monitor changes in data sources, identify new content, and synchronously update it to the knowledge graph. Establish a timeliness management mechanism for regulations, create version chains and time sequence associations for regulatory nodes in the knowledge graph, and select the applicable version based on the time context when answering questions.
[0013] Optionally, the knowledge evolution and optimization mechanism further includes: Set up a user feedback interface to collect feedback on the question-and-answer results; Based on the evaluation, the prompt templates and semantic similarity thresholds in the intelligent question-answering architecture are automatically adjusted.
[0014] Secondly, the present invention also provides a system for constructing an intelligent question-and-answer model for administrative penalties based on multi-source knowledge fusion, for implementing a method for constructing an intelligent question-and-answer model for administrative penalties based on multi-source knowledge fusion, the system comprising: The data corpus construction module is used to build a multi-source heterogeneous data corpus covering the field of administrative penalties. The knowledge graph construction module is used to construct an administrative penalty knowledge graph from the corpus based on domain ontology guidance and large language model extraction. The intelligent question-answering module is used to respond to user questions in a coordinated manner through a retrieval channel and a generation channel based on the knowledge graph. The knowledge evolution module is used to continuously update and optimize the knowledge graph and the intelligent question answering module.
[0015] Compared with the prior art, the present invention has the following advantages and technical effects: This invention provides a method and system for constructing an intelligent question-answering model for administrative penalties based on multi-source knowledge fusion. It effectively integrates market supervision data scattered across different sources and formats, achieving semantic unity and deep correlation of multi-source knowledge through the construction of a domain knowledge graph. The dual-channel question-answering architecture designed based on this graph combines precise structured retrieval with the natural language generation capabilities of a large language model, ensuring that answers are not only accurate and semantically logical but also possess a complete chain of evidence and legal citations, significantly improving the credibility and interpretability of the question-answering results. Simultaneously, this invention introduces knowledge evolution and optimization mechanisms, including automatic incremental updates, regulatory timeliness management, and a closed-loop user feedback mechanism. This enables the system to continuously adapt to changes in laws and regulations and the growth of business knowledge, achieving dynamic knowledge improvement and real-time updating capabilities for the question-answering system. Therefore, it provides high-precision, traceable, and timely intelligent question-answering support for scenarios such as administrative penalties and corporate compliance. Attached Figure Description
[0016] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1This is a flowchart illustrating a method for constructing an intelligent question-answering model for administrative penalties based on multi-source knowledge fusion, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the dual-channel question-and-answer process according to an embodiment of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0019] Example 1 like Figure 1 As shown, this embodiment provides a method for constructing an intelligent question-answering model for administrative penalties based on multi-source knowledge fusion, including: Construct a multi-source heterogeneous data corpus covering the field of administrative penalties; Based on domain ontology guidance and large language model extraction, an administrative penalty knowledge graph is constructed from the corpus. Design an intelligent question-answering architecture that coordinates retrieval and generation channels, and responds to user questions based on the knowledge graph; A knowledge evolution and optimization mechanism is established to continuously update the knowledge graph and the intelligent question-answering architecture.
[0020] Furthermore, the process of constructing the corpus includes: We collect structured and unstructured data from regulatory announcements, standards information, credit information disclosures, and public opinion monitoring channels, including laws and regulations, enterprise information, product sampling results, penalty information, and standard documents. The collected data is formatted and semantically cleaned, text and table content are extracted, semantic deduplication and standardization are performed, and a domain corpus with a unified format is formed.
[0021] Specifically, the implementation process of this embodiment includes: Step S1: Data collection and preprocessing of multi-source administrative penalty laws, regulations, rules, normative documents, etc. This involves acquiring relevant structured and unstructured data from multiple channels, and constructing a unified, high-quality corpus through format normalization and noise removal. This includes the following steps: Step S11: Data collection. Structured and unstructured data are obtained from multiple channels, including government supervision, standards information, credit information, complaints and reports, and public opinion monitoring. The types of data collected cover laws and regulations, enterprise registration, product sampling inspection, penalty results, quality standards, policies and regulations, and social public opinion.
[0022] Step S12: Data preprocessing. Based on semantic cleaning and format standardization processes, deep text extraction, structured parsing, and semantic deduplication are performed on the collected multi-source data. Regular expressions and HTML parsing methods are used to remove noisy text. For scanned documents and image-based announcements, the LayoutLM model is used for end-to-end text and layout analysis. Complex table structures are accurately parsed using deep learning-based table recognition models such as TableNet or PaddleOCR Table. In the deduplication stage, a document embedding model (such as Sentence-BERT) is used to calculate semantic similarity, combined with timestamps for comprehensive judgment. At the data cleaning level, a large language model is used to achieve context-aware noise filtering and entity standardization. Finally, after unified encoding and time format normalization, a standardized administrative penalty corpus stored in JSON format is generated.
[0023] Furthermore, the process of constructing the administrative penalty knowledge graph includes: Based on the business logic of the administrative penalty field, we define the core entity categories and their semantic relationships, and establish a formalized domain semantic framework. Based on the domain semantic framework, a prompt template is generated to guide the large language model to extract entities and relations from the unstructured text of the corpus and form knowledge triples. The extracted knowledge triples are subjected to cross-source entity semantic alignment and fusion, and the fused knowledge is stored in a graph database in the form of nodes and relations.
[0024] Furthermore, the cross-source entity semantic alignment and fusion process includes: Semantic vector models are used to encode entity names from different sources into vectors; Calculate the semantic similarity between entity vectors. When the similarity exceeds a set threshold, they are determined to be the same semantic entity and aggregated. Consistency checks are performed on the aggregation results based on predefined semantic conflict rules.
[0025] Specifically, the implementation process of this embodiment includes: Step S2, knowledge graph construction guided by domain ontology, defines concepts, relationships, and attribute constraints through domain ontology to achieve semantic alignment and unified representation of multi-source heterogeneous data. Simultaneously, it rapidly generates prompt templates for entities and relationships using the ontology, providing prompts for knowledge extraction from large models, thus forming a reasonable and scalable administrative penalty knowledge graph. This includes the following steps: Step S21: Domain Ontology Design. Based on the business logic of administrative penalty laws, regulations, rules, and normative documents, and their applications in various fields, a semantic ontology model is established. Core entity classes are defined. Taking the product quality domain as an example, these include Law, Regulation, Regulatory Document, Enterprise, Product, Standard, Inspection Report, Nonconformity Item, and Administrative Penalty. Semantic relationship types are also defined, such as producers, basedOnStandard, hasDefect, issued by, violates regulations, and effective / repeal dates. The hierarchy, attribute constraints, and relationship rules between entities are established using OWL to construct the semantic framework for the product quality domain in the administrative penalty scenario.
[0026] Step S22, ontology-driven prompt generation, aims to accurately extract entities and relationships within the market supervision and management domain. This involves leveraging the domain ontology as a foundation to construct prompts for large-scale models. The method first structurally injects predefined core entity types (such as 'financial institutions', 'financial products', 'violations', and 'regulatory regulations') and relationship types (such as 'violation', 'penalty', and 'belonging to') from the market supervision ontology into the prompt instructions. By employing prompt engineering strategies of "role-playing" and "few-sample examples," the large model is guided to deeply understand specific terms and business contexts. For example, the model can be instructed to extract the semantic relationship of 'a securities company' being 'fined' for 'information disclosure violations' from a regulatory announcement and output it strictly following a predefined JSON schema. This ontology-based prompt generation mechanism effectively constrains the generation space of large models, significantly improving the accuracy and structuring level of knowledge extraction within professional domains, and providing a reliable data source for subsequent construction of high-quality knowledge graphs.
[0027] Step S23: Knowledge Extraction and Alignment. Based on a prompt-driven large language model, entities and relations are extracted from unstructured text to obtain a preliminary set of knowledge triples. The Sentence-BERT model is used to semantically vectorize the entity text. Cross-source entity matching is achieved through cosine similarity calculation. When the similarity exceeds a threshold, they are determined to be the same semantic entity and aggregation is performed. Then, semantic conflict detection is performed on the fusion result according to consistency constraint rules to ensure the logical correctness and semantic clarity of the knowledge structure.
[0028] Step S24: Knowledge Graph Construction and Storage. The cleaned and fused triples are imported into a graph database and stored as nodes and relationships, forming a structured administrative penalty knowledge graph. This graph supports multi-level semantic retrieval and path reasoning, enabling semantic association of information such as enterprises, products, standards, and penalties.
[0029] Furthermore, the process by which the intelligent question-answering architecture responds to user questions includes: In the retrieval channel, the semantics and intent of the user's question are parsed and mapped to a structured query of the knowledge graph, and relevant evidence chains are retrieved and extracted; In the generation channel, the user question, the chain of evidence, and relevant legal provisions are input into the large language model to generate a natural language answer with citations. The semantic consistency and confidence of the evidence chain and the natural language response are evaluated, and the final answer is output based on the evaluation results.
[0030] Furthermore, after the process of generating a natural language answer, the method further includes: using a fact-checking module to perform semantic consistency verification on each claim in the generated answer and its corresponding cited content.
[0031] Specifically, the implementation process of this embodiment includes: like Figure 2 Step S3, as shown, involves the design of a multi-source knowledge fusion question-answering architecture. By combining retrieval and generation channels, it achieves accurate retrieval of legal knowledge, semantic understanding, and natural language generation, thereby improving the accuracy and interpretability of the question-answering results. This includes the following steps: Step S31 involves constructing and executing the retrieval channel. The BERT-BiLSTM-CRF model is used for named entity recognition, combined with the BERT intent classification model to complete semantic parsing and determine the question type (e.g., legal and regulatory queries, standard basis queries, penalty result queries, sampling inspection information queries, etc.). The parsing results are then mapped to Cypher query statements, and a structured query is executed in the Neo4j knowledge graph. The semantic similarity between the question and node description is calculated using the Sentence-BERT model to achieve high-relevance retrieval and evidence chain extraction.
[0032] Step S31: Retrieval Channel Construction and Execution. Utilizing a large language model fine-tuned for market supervision (such as the Qwen series models), end-to-end query understanding is performed, completing named entity recognition, intent classification, and determining the question type (e.g., legal and regulatory queries, standard basis queries, penalty result queries, sampling inspection information queries, etc.) in one go. The parsed results are mapped to a standard graph query language, and structured queries are executed in the graph database. Simultaneously, a hybrid retrieval mechanism is employed, combining dense vector retrieval and graph neural networks for multi-hop relationship discovery, achieving high-relevance retrieval and evidence chain extraction.
[0033] Step S32: Generation Channel Design and Output Control. User questions, search results, and relevant legal content are input into the instruction-tuned large language model. A natural language answer is generated using a citation constraint prompt template. A lightweight "fact-checking module" is introduced. This module performs semantic consistency checks on each generated "claim-citation" pair through a cross-encoder model or a secondary LLM call, ensuring that the cited content truly supports the corresponding assertions in the answer, fundamentally eliminating "illusionary citations," and guaranteeing the traceability and factual accuracy of the output results.
[0034] Step S33: Result Fusion and Reliable Output Mechanism. Based on semantic consistency measurement and a multi-dimensional dynamic confidence assessment algorithm, the retrieved and generated results are fused. This confidence level considers not only semantic matching degree but also comprehensively evaluates source authority, regulatory timeliness, and the strength of the evidence chain. When the consistency and confidence levels are higher than preset thresholds, the fused natural language answer is output, along with an attached structured evidence chain; otherwise, it degenerates into a structured retrieval result, ensuring the accuracy, stability, and interpretability of the question-and-answer output.
[0035] Furthermore, the knowledge evolution and optimization mechanism includes: An automatic incremental update mechanism is established to periodically monitor changes in data sources, identify new content, and synchronously update it to the knowledge graph. Establish a timeliness management mechanism for regulations, create version chains and time sequence associations for regulatory nodes in the knowledge graph, and select the applicable version based on the time context when answering questions.
[0036] Furthermore, the knowledge evolution and optimization mechanism also includes: Set up a user feedback interface to collect feedback on the question-and-answer results; Based on the evaluation, the prompt templates and semantic similarity thresholds in the intelligent question-answering architecture are automatically adjusted.
[0037] Specifically, the implementation process of this embodiment includes: Step S4, Knowledge Evolution and Optimization Mechanism, utilizes an automatic incremental update mechanism, regulatory timeliness management, and user feedback mechanism to achieve dynamic evolution and continuous optimization of the knowledge graph and model parameters. This includes the following steps: Step S41: Automatic incremental update. A periodic task mechanism is established, utilizing Playwright in conjunction with scheduled tasks to automatically monitor changes in government announcements, standard revisions, and regulatory information. Semantic similarity is calculated through document embedding using SentenceTransformers to accurately identify substantive content changes. The entire process is implemented in real-time using stream processing frameworks such as Apache Flink, and Neo4J's MERGE operation is used to incrementally write nodes and relationships to the graph database, ensuring the timeliness of the knowledge graph.
[0038] Step S42, Regulation Timeliness Management: Addressing the issues of regulation revision and version replacement in administrative penalty scenarios, each regulation node is assigned time attributes such as "Publication Date," "Implementation Date," and "Repeal Date," and a version chain and temporal association are established in the knowledge graph. When a new regulation is issued or an old regulation is repealed, the semantic similarity between the new and old regulations is automatically calculated, and a replacement relationship is established. During the question-and-answer phase, the corresponding version of the regulation is automatically selected based on the temporal context of the user's question (such as the date of the event or the date of the sampling report), and the "Regulation Effective Status" and "Applicable Period" are marked in the generated results to ensure the legal validity and timeliness of the answers.
[0039] Step S43: Set up a user feedback interface to collect evaluations of the accuracy and citation standardization of the question-and-answer results. Based on the feedback, automatically adjust the prompt template structure, similarity threshold, and recognition parameters to achieve continuous optimization and adaptive adjustment of the model, forming a closed-loop learning mechanism of "data update—knowledge evolution—feedback optimization".
[0040] Example 2 Based on the same general inventive concept, this invention also provides a system for constructing an intelligent question-and-answer model for administrative penalties based on multi-source knowledge fusion. The system is described below, and the method described above can be referred to in conjunction with the system described below. The system includes: The data corpus construction module is used to build a multi-source heterogeneous data corpus covering the field of administrative penalties. The knowledge graph construction module is used to construct an administrative penalty knowledge graph from the corpus based on domain ontology guidance and large language model extraction. The intelligent question-answering module is used to respond to user questions in a coordinated manner through a retrieval channel and a generation channel based on the knowledge graph. The knowledge evolution module is used to continuously update and optimize the knowledge graph and the intelligent question answering module.
[0041] An application example of this invention is as follows: Information obtained from relevant announcements states: "Recently, market supervision departments conducted random inspections of electric water heaters sold in the market. Testing revealed that some samples had grounding resistance that did not meet the national standard GB 4706.1-2005. The relevant companies have been required to rectify the issues within a specified period and publicly announce the results of the investigation." After preprocessing, the announcement text is converted into a unified JSON format, containing the following fields: Release Date: 2024-03-18, Product Category: Electric Water Heater, Non-conforming Item: Excessive Grounding Resistance, Testing Standard: GB 4706.1-2005, Testing Institution: Market Supervision Department, Handling Measures: Ordered to rectify within a time limit.
[0042] The model input is automatically generated based on the Few-shot prompt template. Fields such as "detection conclusion" and "product description" in the JSON corpus generated in step S12 are concatenated into natural language fragments and input into the large language model (Qwen-Max) to perform entity and relation extraction.
[0043] Input text: "An electric water heater manufactured by a certain company was tested and found that its grounding resistance does not meet the GB 4706.1-2005 standard." Output triple: (A manufacturing company, produces, electric water heaters); (Electric water heaters, do not comply with GB 4706.1-2005).
[0044] Enterprise names from different sources (such as "×× Electric Appliance Co., Ltd." and "×× Home Appliance Manufacturer") are encoded into Sentence-BERT vectors, and cosine similarity is calculated. When the similarity is higher than a threshold (such as 0.92), they are identified as the same enterprise entity and aggregated.
[0045] The merged triples are imported in batches into the Neo4j graph database to form a semantic association network.
[0046] Suppose a user asks: "Which national standard should the grounding resistance of an electric water heater comply with?" Leveraging a finely tuned LLM (Limited Memory Management) model adapted for market regulation, end-to-end query understanding is performed, enabling named entity recognition (e.g., "electric water heater": Product, "grounding resistance": NonConformityItem), intent classification, and key information extraction in a single operation. The parsed results are mapped to graph query statements, and structured queries are executed in the graph database. Simultaneously, a hybrid retrieval mechanism is employed, combining vector retrieval and graph neural networks for multi-hop relationship discovery, achieving highly relevant retrieval and evidence chain extraction.
[0047] After execution, the returned result is: "GB 4706.1-2005 Safety of household and similar electrical appliances - Part 1: General requirements". User questions, search results, and relevant regulations are input into a fine-tuned large language model (Qwen-Max), which uses citation constraint prompts to generate a natural language answer. The generated answer is then passed through a semantic verification model based on a cross-encoder to fact-check the cited content, ensuring that the citations actually support the arguments.
[0048] The final output reads: "According to Article 8.2 of the national standard GB 4706.1-2005 'Safety of household and similar electrical appliances - Part 1: General requirements,' the grounding resistance of a Class I electric water heater shall not exceed 0.1Ω. [Source: Standard GB 4706.1-2005]." The output also indicates whether the standard is currently in effect and its applicable period, ensuring the legal timeliness of the answer.
[0049] Based on the retrieval and generation results, semantic consistency and confidence scores are calculated. If both scores are higher than the set threshold, the fused natural language answer is output, along with a "References" link to the original standard text and the sampling report.
[0050] If the monitoring program detects a new sampling result of "poor grounding of a certain brand of electric water heater", it uses SentenceTransformers to calculate document embeddings to identify content changes. The update process is triggered by the Apache Flink stream processing framework, and incremental writing of nodes and relationships is completed through MERGE statements, without manual intervention.
[0051] If a user reports that an answer does not display its source, the sample is recorded and template defects are analyzed. If a missing placeholder for a citation format is found in the prompt template, the template structure is automatically completed and the prompt parameters are updated. Simultaneously, the similarity thresholds in steps S23 and S33 are dynamically adjusted based on the statistical results of the feedback samples, thereby optimizing the fusion effect.
[0052] This invention provides a method and system for constructing an intelligent question-answering model for administrative penalties based on multi-source knowledge fusion. It effectively integrates market supervision data scattered across different sources and formats, achieving semantic unity and deep correlation of multi-source knowledge through the construction of a domain knowledge graph. The dual-channel question-answering architecture designed based on this graph combines precise structured retrieval with the natural language generation capabilities of a large language model, ensuring that answers are not only accurate and semantically logical but also possess a complete chain of evidence and legal citations, significantly improving the credibility and interpretability of the question-answering results. Simultaneously, this invention introduces knowledge evolution and optimization mechanisms, including automatic incremental updates, regulatory timeliness management, and a closed-loop user feedback mechanism. This enables the system to continuously adapt to changes in laws and regulations and the growth of business knowledge, achieving dynamic knowledge improvement and real-time updating capabilities for the question-answering system. Therefore, it provides high-precision, traceable, and timely intelligent question-answering support for scenarios such as administrative penalties and corporate compliance.
[0053] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for constructing an intelligent question-answering model for administrative penalties that integrates multi-source knowledge, characterized in that, Includes the following steps: Construct a multi-source heterogeneous data corpus covering the field of administrative penalties; Based on domain ontology guidance and large language model extraction, an administrative penalty knowledge graph is constructed from the corpus. Design an intelligent question-answering architecture that coordinates retrieval and generation channels, and responds to user questions based on the knowledge graph; A knowledge evolution and optimization mechanism is established to continuously update the knowledge graph and the intelligent question-answering architecture.
2. The method according to claim 1, characterized in that, The process of constructing the corpus includes: We collect structured and unstructured data from regulatory announcements, standards information, credit information disclosures, and public opinion monitoring channels, including laws and regulations, enterprise information, product sampling results, penalty information, and standard documents. The collected data is formatted and semantically cleaned, text and table content are extracted, semantic deduplication and standardization are performed, and a domain corpus with a unified format is formed.
3. The method according to claim 1, characterized in that, The process of constructing the administrative penalty knowledge graph includes: Based on the business logic of the administrative penalty field, we define the core entity categories and their semantic relationships, and establish a formalized domain semantic framework. Based on the domain semantic framework, a prompt template is generated to guide the large language model to extract entities and relations from the unstructured text of the corpus and form knowledge triples. The extracted knowledge triples are subjected to cross-source entity semantic alignment and fusion, and the fused knowledge is stored in a graph database in the form of nodes and relations.
4. The method according to claim 3, characterized in that, The process of cross-source entity semantic alignment and fusion includes: Semantic vector models are used to encode entity names from different sources into vectors; Calculate the semantic similarity between entity vectors. When the similarity exceeds a set threshold, they are determined to be the same semantic entity and aggregated. Consistency checks are performed on the aggregation results based on predefined semantic conflict rules.
5. The method according to claim 1, characterized in that, The intelligent question-answering architecture responds to user questions by including: In the retrieval channel, the semantics and intent of the user's question are parsed and mapped to a structured query of the knowledge graph, and relevant evidence chains are retrieved and extracted; In the generation channel, the user question, the chain of evidence, and relevant legal provisions are input into the large language model to generate a natural language answer with citations. The semantic consistency and confidence of the evidence chain and the natural language response are evaluated, and the final answer is output based on the evaluation results.
6. The method according to claim 5, characterized in that, After generating the natural language answer, the process further includes: using a fact-checking module to perform semantic consistency verification on each claim in the generated answer and its corresponding cited content.
7. The method according to claim 1, characterized in that, The knowledge evolution and optimization mechanism includes: An automatic incremental update mechanism is established to periodically monitor changes in data sources, identify new content, and synchronously update it to the knowledge graph. Establish a timeliness management mechanism for regulations, create version chains and time sequence associations for regulatory nodes in the knowledge graph, and select the applicable version based on the time context when answering questions.
8. The method according to claim 7, characterized in that, The knowledge evolution and optimization mechanism also includes: Set up a user feedback interface to collect feedback on the question-and-answer results; Based on the evaluation, the prompt templates and semantic similarity thresholds in the intelligent question-answering architecture are automatically adjusted.
9. A system for constructing an intelligent question-and-answer model for administrative penalties that integrates multi-source knowledge, characterized in that... The system for implementing the method of any one of claims 1-8 comprises: The data corpus construction module is used to build a multi-source heterogeneous data corpus covering the field of administrative penalties. The knowledge graph construction module is used to construct an administrative penalty knowledge graph from the corpus based on domain ontology guidance and large language model extraction. The intelligent question-answering module is used to respond to user questions in a coordinated manner through a retrieval channel and a generation channel based on the knowledge graph. The knowledge evolution module is used to continuously update and optimize the knowledge graph and the intelligent question answering module.