Government affair knowledge graph ontology construction and optimization method, device, equipment and medium

By constructing a domain knowledge base and adaptively selecting ontology analysis modes, the problems of weak adaptability and single mode in knowledge graph construction are solved, achieving efficient and accurate ontology construction, improving processing efficiency and exploration capabilities, and reducing manual intervention.

CN120930760BActive Publication Date: 2026-02-10INSPUR SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511476432.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-02-10
Estimated Expiration
2045-10-16

AI Technical Summary

Technical Problem

Existing knowledge graph construction technologies cannot meet the specialized needs of different domains. They have weak adaptability, a single construction mode, and lack a unified knowledge benchmark, resulting in low efficiency and unstable quality in ontology construction, making it difficult to cope with dynamic changes.

Method used

It constructs a domain knowledge base, automatically loads basic configuration parameters of the target domain, generates a hierarchical structured knowledge base, obtains data through API interfaces, adaptively selects ontology analysis modes, automatically matches the optimal analysis mode, generates a knowledge graph ontology, including fixed, scope and exploratory analysis modes, autonomously perceives ontology change characteristics, and reduces manual intervention.

Benefits of technology

It achieves high efficiency and accuracy in ontology construction, significantly improves terminology accuracy and concept association accuracy, enhances ontology reliability, improves processing efficiency and the ability to mine knowledge from exploration scenarios, and reduces anomaly response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930760B_ABST
    Figure CN120930760B_ABST
Patent Text Reader

Abstract

The application discloses a government affair knowledge graph ontology construction and optimization method and device, equipment and medium, belongs to the cross technical field of artificial intelligence and knowledge engineering, and the technical problem to be solved by the application is how to provide a unified and standardized knowledge benchmark for ontology construction, realize autonomous perception of ontology change characteristics of a system, accurately match an optimal analysis mode, and guarantee the efficiency and accuracy of ontology construction without manual intervention, and the technical scheme is that: a field knowledge base is constructed; an ontology construction system is started, basic configuration parameters of a target field are automatically loaded, core knowledge elements and unstructured documents of the target field are collected, the unstructured documents are classified, semantically marked and uniformly processed in format, a hierarchical structured field knowledge base is generated, and the field knowledge base is used as a unified knowledge benchmark for ontology construction; a processing data acquisition and analysis process is triggered; and an ontology analysis mode is adaptively selected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and knowledge engineering, specifically to a method, apparatus, equipment, and medium for constructing and optimizing a government knowledge graph ontology. Background Technology

[0002] Current knowledge graph construction technologies struggle to meet the specialized needs of different domains, with the following key shortcomings:

[0003] ① Weak domain adaptability: General construction solutions cannot accurately match the conceptual system and business rules of specific domains. For example, the semantic associations of professional terms and processes such as "cross-departmental approval" in the government sector are difficult to capture effectively, resulting in the disconnect between the graph and actual business.

[0004] ② The ontology construction mode is singular: it mostly adopts a fixed mode of "preset framework + knowledge filling", which cannot respond to the temporary needs of mature fields, and also limits the potential knowledge mining of immature fields. It is difficult to adapt to differentiated scenarios such as standardization, limited scope, and dynamic exploration.

[0005] ③ Lack of adaptation mechanism: The lack of a unified knowledge benchmark leads to inconsistencies in ontology concepts and relationships; the absence of a scientific model decision-making mechanism and reliance on manual selection of analysis models make it difficult to cope with dynamic changes in ontology structure and data, resulting in low efficiency and unstable quality in map construction.

[0006] Therefore, how to provide a unified and standardized knowledge benchmark for ontology construction, enable the system to autonomously perceive the characteristics of ontology changes, accurately match the optimal analysis mode, and ensure the efficiency and accuracy of ontology construction without human intervention is a technical problem that urgently needs to be solved. Summary of the Invention

[0007] The technical objective of this invention is to provide a method, apparatus, device, and medium for constructing and optimizing a government knowledge graph ontology, in order to address the problem of how to provide a unified and standardized knowledge benchmark for ontology construction, enable the system to autonomously perceive ontology change characteristics, accurately match the optimal analysis mode, and ensure the efficiency and accuracy of ontology construction without manual intervention.

[0008] The technical objective of this invention is achieved as follows: a method for constructing and optimizing a government knowledge graph ontology, the specific method of which is as follows:

[0009] Building a domain knowledge base: The ontology building system is launched, automatically loading the basic configuration parameters of the target domain, collecting the core knowledge elements and unstructured documents of the target domain, and classifying, semantically annotating and unifying the format of the unstructured documents to generate a hierarchical structured domain knowledge base as a unified knowledge benchmark for ontology building; among which, the core knowledge elements include standardized terminology, conceptual relationships and basic business rules.

[0010] Data Acquisition and Analysis Process Triggering: Acquire data to be processed in the target domain via API interface or file upload, and automatically trigger subsequent adaptive analysis processes;

[0011] Adaptive selection of ontology analysis mode: Automatically collect historical ontology structure change data, entity data fluctuation data, and semantic consistency data, combine the characteristics of the data to be processed to calculate the comprehensive ontology stability score, determine the stability level, and map the corresponding ontology analysis mode based on the stability level. Simultaneously verify the compatibility between the data to be processed and the mapping mode: If there is a compatibility conflict, the conflict is arbitrated according to the priority of "business rules → historical data comparison → manual intervention" to determine the final ontology analysis mode.

[0012] A knowledge graph ontology is generated based on the selected ontology analysis mode, and entity extraction, data processing and ontology construction are automatically performed according to the final determined mode to obtain a structured domain knowledge graph ontology; among which, the ontology analysis mode includes fixed analysis mode, scope analysis mode and exploratory analysis mode.

[0013] As a preferred approach, standardized terminology covers industry-wide terms and discipline-specific terms in the target field; conceptual relationships include synonyms, near-synonyms, and core concepts.

[0014] As a preferred option, the stability levels include high stability, medium stability, and low stability.

[0015] High stability refers to a stability score ≥ 85 points, corresponding to a fixed mapping analysis mode; medium stability refers to a stability score 65-84 points, corresponding to a range mapping analysis mode; low stability refers to a stability score < 65 points, corresponding to an exploratory mapping analysis mode.

[0016] More preferably, the fixed analysis mode is as follows:

[0017] Automatically loads a pre-defined, unmodifiable ontology framework, which contains the core concepts and fixed relationships of the target domain;

[0018] Only extract specific entity information from the data to be processed, without generating new concepts or adjusting existing relationships;

[0019] The extracted entities are mounted to the corresponding nodes according to the relationships within the ontology framework, thereby enabling incremental data supplementation within the framework.

[0020] More preferably, the range analysis mode is as follows:

[0021] Automatically identify the business sub-scope associated with the data to be processed, and determine the conceptual analysis boundary of the ontology;

[0022] Only extract the concepts and relationships directly related to the boundary of the concept analysis;

[0023] It automatically identifies and filters irrelevant information that exceeds the boundaries of conceptual analysis, and generates an ontology that focuses on a specific business sub-scope.

[0024] More specifically, the exploration and analysis model is as follows:

[0025] Automatically set the concept value threshold, which is based on the frequency of text occurrence, semantic relevance, and domain relevance.

[0026] Perform a full scan of the data to be processed to extract explicit concepts, implicit concepts, and potential relationships;

[0027] The comprehensive value of each concept is calculated using the DG-CW concept value algorithm, filtering out low-value concepts below the concept value threshold. The DG-CW concept value algorithm formula is as follows:

[0028] ;

[0029] ;

[0030] in, Indicates the adaptation parameters; This indicates the frequency of occurrence of concept c; This indicates the highest frequency of all concepts; This represents the value of each position in the title, core paragraph, and non-core paragraph; n represents the number of positions. This represents the similarity between concept c and core concept k; This indicates the value of the core concept k;

[0031] A preliminary ontology framework is generated by merging synonymous concepts and unifying terminology using the DG-TY domain synonymy algorithm; the formula for the DG-TY domain synonymy algorithm is as follows:

[0032] ;

[0033] in, Representation of terms and semantic similarity; Indicates the strength of the association between terms in the domain knowledge base; This represents the domain adaptation coefficient.

[0034] As a preferred option, the conflict arbitration process is as follows:

[0035] If the conflict stems from the pattern not conforming to the mandatory business rules of the target domain, then the business rule base in the domain knowledge base is called for verification, and the pattern is switched to conform to the mandatory rules.

[0036] If the conflict stems from a significant difference between the current matching mode and the historical matching mode for the same scenario, then retrieve the performance data of the matching mode for the same scenario over the past 3 months and use the mode with the better performance.

[0037] If the corresponding arbitration cannot resolve the conflict, a conflict arbitration work order will be automatically generated and pushed to the domain administrator. The administrator will then execute the specified mode according to the decision and record the decision result in the decision knowledge base.

[0038] A device for constructing and optimizing a government knowledge graph ontology, which implements the aforementioned method for constructing and optimizing a government knowledge graph ontology; the device is specifically as follows:

[0039] The knowledge base construction module is used to start the ontology construction system, automatically load the basic configuration parameters of the target domain, collect the core knowledge elements and unstructured documents of the target domain, and classify, semantically annotate and unify the format of the unstructured documents to generate a hierarchical structured domain knowledge base as a unified knowledge benchmark for ontology construction. Among them, the core knowledge elements include standardized terminology, conceptual relationships and basic business rules.

[0040] The data acquisition module is used to acquire data to be processed in the target domain through API interface or file upload, and automatically trigger subsequent adaptive analysis processes.

[0041] Adaptive mode selection module: used to automatically collect historical ontology structure change data, entity data fluctuation data and semantic consistency data, calculate the comprehensive ontology stability score by combining the characteristics of the data to be processed, determine the stability level, and map the corresponding ontology analysis mode based on the stability level, and simultaneously verify the compatibility between the data to be processed and the mapping mode: if there is a compatibility conflict, the conflict is arbitrated according to the priority of "business rules → historical data comparison → manual intervention" to determine the final ontology analysis mode.

[0042] The ontology generation module is used to generate a knowledge graph ontology based on the selected ontology analysis mode, and automatically perform entity extraction, data processing and ontology construction according to the finally determined mode to obtain a structured domain knowledge graph ontology; wherein, the ontology analysis mode includes fixed analysis mode, scope analysis mode and exploratory analysis mode;

[0043] The storage module is used to store the domain knowledge base, historical ontology change data, data to be processed, and the generated knowledge graph ontology;

[0044] The anomaly monitoring module is used to monitor the semantic conflict rate and the proportion of external entities during the ontology construction process in real time. When the abnormal indicators continue to exceed the preset threshold, it triggers emergency mode switching and abnormal data isolation, and simultaneously notifies the domain administrator.

[0045] An electronic device includes: a memory and at least one processor;

[0046] The memory contains computer programs;

[0047] The at least one processor executes the computer program stored in the memory, causing the at least one processor to execute the above-described method for constructing and optimizing a government knowledge graph ontology.

[0048] A computer-readable storage medium storing a computer program that can be executed by a processor to implement the above-described method for constructing and optimizing a government knowledge graph ontology.

[0049] The method, apparatus, equipment, and medium for constructing and optimizing government knowledge graph ontology of the present invention have the following advantages:

[0050] (i) This invention provides a unified and standardized knowledge benchmark for ontology construction by constructing a structured and highly available domain knowledge base, thereby solving the problem of insufficient support caused by the dispersion of domain knowledge and semantic ambiguity;

[0051] (ii) This invention designs multiple ontology analysis modes that are adaptable to different business scenarios (standardized processes, defined boundaries, dynamic mining), meeting the differentiated needs of the domain for ontology structure stability, scenario focus, and potential knowledge mining;

[0052] (III) This invention establishes an adaptive analysis mode mechanism, enabling the system to autonomously perceive the characteristics of ontology changes and accurately match the optimal analysis mode. It can ensure the efficiency and accuracy of ontology construction without human intervention, and ultimately form a practical domain-level knowledge graph ontology construction technology system.

[0053] (iv) In response to the pain points of weak specialization adaptability and single ontology construction mode in the existing technology, this invention proposes a domain-level knowledge graph ontology construction solution of "domain knowledge base construction - multi-scenario ontology analysis mode adaptation", forming a feasible technical system;

[0054] (v) This invention transforms scattered unstructured / semi-structured knowledge into a hierarchically clear structured knowledge base, significantly improving the accuracy of terminology and the correctness of concept association, effectively solving the problems of knowledge dispersion and semantic ambiguity, and greatly optimizing the level of knowledge structuring and standardization;

[0055] (vi) This invention relies on a standardized knowledge base to construct the ontology, which significantly reduces ontology bias compared to existing technologies, ensures the accuracy of ontology concept definitions and relational logic, and significantly enhances ontology reliability;

[0056] (vii) This invention adapts to batch standardized business through a fixed analysis mode, supplements only the entity data within the preset framework, and greatly improves processing efficiency while ensuring the stability and order of the ontology structure, thus significantly improving the processing efficiency of standardized scenarios.

[0057] (viii) This invention focuses on specific business sub-scenarios through range analysis mode, automatically filters out irrelevant information that exceeds the range, significantly reduces interference, makes ontology analysis more in line with target needs, and greatly improves the accuracy of limited scenario analysis;

[0058] (ix) This invention breaks through the limitations of traditional preset frameworks by exploring the analysis mode, and can discover explicit and implicit concepts and potential relationships in immature fields, providing strong support for innovative business needs and effectively strengthening the ability to mine knowledge in exploration scenarios;

[0059] (x) Existing technologies require manual selection of analysis modes. This invention adds an autonomous decision-making mechanism based on multi-dimensional ontology change data, namely, an autonomous mode matching function, which can accurately match the optimal mode without manual intervention and flexibly respond to dynamic changes in the ontology.

[0060] (xi) Existing technologies rely on manual inspection of anomalies. This invention adds an automatic anomaly detection and rapid processing function, which can promptly complete stability reassessment, mode switching and anomaly data isolation, and significantly shorten response time.

[0061] (xii) Existing technologies have difficulty balancing construction quality and efficiency. This invention achieves dual optimization of the consistency of the ontology structure and the efficiency of entity extraction by adding intelligent protection throughout the entire process, and forms a reusable technical system, which completely solves the core pain points of existing technologies such as "weak adaptability, single mode and low efficiency". Attached Figure Description

[0062] The invention will be further described below with reference to the accompanying drawings.

[0063] Appendix Figure 1 A flowchart illustrating the methods for constructing and optimizing government knowledge graph ontology;

[0064] Appendix Figure 2 A schematic diagram of the three major ontology analysis models;

[0065] Appendix Figure 3 This is a flowchart illustrating the dynamic fusion process. Detailed Implementation

[0066] The following detailed description of the method, apparatus, equipment, and medium for constructing and optimizing the government knowledge graph ontology of the present invention is based on the accompanying drawings and specific embodiments.

[0067] Example 1: As shown in the attached document Figure 1As shown in the figure, this embodiment provides a method for constructing and optimizing a government knowledge graph ontology, which is as follows:

[0068] S1. Constructing a Domain Knowledge Base: Start the ontology construction system, automatically load the basic configuration parameters of the target domain, collect the core knowledge elements and unstructured documents of the target domain, and classify, semantically annotate and unify the format of the unstructured documents to generate a hierarchical structured domain knowledge base as a unified knowledge benchmark for ontology construction; among which, the core knowledge elements include standardized terminology, conceptual relationships and basic business rules.

[0069] S2. Data Acquisition and Analysis Process Trigger: Acquire data to be processed in the target domain through API interface or file upload, and automatically trigger subsequent adaptive analysis processes;

[0070] S3. Adaptive selection of ontology analysis mode: Automatically collect historical ontology structure change data, entity data fluctuation data, and semantic consistency data. Combine the characteristics of the data to be processed to calculate the comprehensive ontology stability score, determine the stability level, and map the corresponding ontology analysis mode based on the stability level. Simultaneously verify the compatibility between the data to be processed and the mapping mode: If there is a compatibility conflict, the conflict is arbitrated according to the priority of "business rules → historical data comparison → manual intervention" to determine the final ontology analysis mode.

[0071] S4. Generate a knowledge graph ontology based on the selected ontology analysis mode, and automatically perform entity extraction, data processing and ontology construction according to the final determined mode to obtain a structured domain knowledge graph ontology; wherein, the ontology analysis mode includes fixed analysis mode, scope analysis mode and exploratory analysis mode.

[0072] The standardized terminology in step S1 of this embodiment covers industry-wide terms and discipline-specific terms in the target field; the conceptual relationships include synonyms, near-synonyms, and core concepts.

[0073] The stability levels in step S3 of this embodiment include high stability, medium stability, and low stability.

[0074] High stability refers to a stability score ≥ 85 points, corresponding to a fixed mapping analysis mode; medium stability refers to a stability score 65-84 points, corresponding to a range mapping analysis mode; low stability refers to a stability score < 65 points, corresponding to an exploratory mapping analysis mode.

[0075] The fixed analysis mode in this embodiment is as follows:

[0076] ① Automatically loads a pre-defined, unmodifiable ontology framework, which contains the core concepts and fixed relationships of the target domain;

[0077] ② Only extract specific entity information from the data to be processed, without generating new concepts or adjusting existing relationships;

[0078] ③ The extracted entities are mounted to the corresponding nodes according to the relationships within the ontology framework, thereby enabling incremental data supplementation within the framework.

[0079] The specific range analysis mode in this embodiment is as follows:

[0080] ① Automatically lock the business sub-scope associated with the data to be processed and determine the conceptual analysis boundary of the ontology;

[0081] ② Only extract the concepts and relationships directly related to the boundary of the concept analysis;

[0082] ③ Automatically identify and filter irrelevant information that exceeds the boundaries of concept analysis, and generate an ontology that focuses on a specific business sub-scope.

[0083] The specific exploratory analysis mode in this embodiment is as follows:

[0084] ① Automatically set the concept value threshold, which is based on the frequency of text occurrence, semantic relevance, and domain relevance strength;

[0085] ② Perform a full scan of the data to be processed to extract explicit concepts, implicit concepts, and potential relationships;

[0086] ③ Calculate the comprehensive value of each concept using the DG-CW concept value algorithm, filtering out low-value concepts below the concept value threshold; the DG-CW concept value algorithm formula is as follows:

[0087] ;

[0088] ;

[0089] in, Indicates the adaptation parameters; This indicates the frequency of occurrence of concept c; This indicates the highest frequency of all concepts; This represents the value of each position in the title, core paragraph, and non-core paragraph; n represents the number of positions. This represents the similarity between concept c and core concept k; This indicates the value of the core concept k;

[0090] ④ By merging synonymous concepts and unifying terminology through the DG-TY domain synonym algorithm, a preliminary ontology framework is generated; the formula for the DG-TY domain synonym algorithm is as follows:

[0091] ;

[0092] in, Representation of terms and semantic similarity; Indicates the strength of the association between terms in the domain knowledge base; This represents the domain adaptation coefficient.

[0093] The conflict arbitration situation in step S3 of this embodiment is as follows:

[0094] If the conflict stems from the pattern not conforming to the mandatory business rules of the target domain, then the business rule base in the domain knowledge base is called for verification, and the pattern is switched to conform to the mandatory rules.

[0095] If the conflict stems from a significant difference between the current matching mode and the historical matching mode for the same scenario, then retrieve the performance data of the matching mode for the same scenario over the past 3 months and use the mode with the better performance.

[0096] If the corresponding arbitration cannot resolve the conflict, a conflict arbitration work order will be automatically generated and pushed to the domain administrator. The administrator will then execute the specified mode according to the decision and record the decision result in the decision knowledge base.

[0097] Example 2: This example provides a device for constructing and optimizing a government knowledge graph ontology. This device is used to implement the government knowledge graph ontology construction and optimization method as described in Example 1. The device is specifically as follows:

[0098] The knowledge base construction module is used to start the ontology construction system, automatically load the basic configuration parameters of the target domain, collect the core knowledge elements and unstructured documents of the target domain, and classify, semantically annotate and unify the format of the unstructured documents to generate a hierarchical structured domain knowledge base as a unified knowledge benchmark for ontology construction.

[0099] The data acquisition module is used to acquire data to be processed in the target domain through API interface or file upload, and automatically trigger subsequent adaptive analysis processes.

[0100] Adaptive mode selection module: used to automatically collect historical ontology structure change data, entity data fluctuation data and semantic consistency data, calculate the comprehensive ontology stability score by combining the characteristics of the data to be processed, determine the stability level, and map the corresponding ontology analysis mode based on the stability level, and simultaneously verify the compatibility between the data to be processed and the mapping mode: if there is a compatibility conflict, the conflict is arbitrated according to the priority of "business rules → historical data comparison → manual intervention" to determine the final ontology analysis mode.

[0101] The ontology generation module is used to generate a knowledge graph ontology based on the selected ontology analysis mode, and automatically perform entity extraction, data processing and ontology construction according to the finally determined mode to obtain a structured domain knowledge graph ontology; wherein, the ontology analysis mode includes fixed analysis mode, scope analysis mode and exploratory analysis mode;

[0102] The storage module is used to store the domain knowledge base, historical ontology change data, data to be processed, and the generated knowledge graph ontology;

[0103] The anomaly monitoring module is used to monitor the semantic conflict rate and the proportion of external entities during the ontology construction process in real time. When the abnormal indicators continue to exceed the preset threshold, it triggers emergency mode switching and abnormal data isolation, and simultaneously notifies the domain administrator.

[0104] In this embodiment, the knowledge base construction module imports key knowledge elements from the target domain to form a structured pre-built domain knowledge base, defining core concepts within the domain. These core knowledge elements include standardized terminology, conceptual relationships, and basic business rules. Standardized professional terminology covers industry-wide terms and discipline-specific terms. A set of similar concepts clarifies similar concepts. Basic business rules include logical judgment rules and data association criteria within the domain.

[0105] The fixed analysis mode in this embodiment is suitable for scenarios where the domain ontology is highly mature and the ontology structure must be stable; its characteristics are as follows:

[0106] ① Framework preset locking: "Concepts" (nodes) and "relationships" (edges) are fixed, forming an unmodifiable ontology framework, ensuring that the framework is fully aligned with business standards;

[0107] ② Targeted entity parsing: Extracts only specific entity information from the text without changing the ontology;

[0108] ③ Structured data supplementation: The extracted entities are mounted to the corresponding nodes according to the ontology framework.

[0109] The scope analysis mode in this embodiment is applicable to scenarios where a specific business sub-scope needs to be focused on, as follows:

[0110] ① Define the boundaries of concept analysis: Specify only the "concepts" (nodes) of the ontology;

[0111] ② Targeted text parsing and extraction: Extracting concepts directly related to the defined scope and the relationships (edges) between these concepts.

[0112] The exploratory analysis mode in this embodiment is suitable for scenarios where the domain ontology is not yet mature. It does not presuppose any "concepts" or "relationships," but only explores potential "concepts" and "relationships" through the analysis of the dataset, thereby realizing the dynamic construction of the ontology framework. The process is as follows:

[0113] ① Parameter initialization: Set the "concept value threshold" according to the characteristics of the domain business. This threshold is the minimum standard for judging the concept value.

[0114] ② Full scan: The large model performs a full ontology scan on the original dataset, extracting all concepts and relationships in the text to form an initial knowledge set;

[0115] ③ Value calculation and cleaning, as detailed below:

[0116] Step 1: Calculate the comprehensive value of each concept using the DG-CW (Domain Graph - Concept Worth) algorithm; the DG-CW algorithm formula is as follows:

[0117] ;

[0118] ;

[0119] in, To adapt the parameters; It refers to the frequency of occurrence of concept c; It is the most frequent of all concepts; It represents the value of each position (heading, core paragraph, non-core paragraph, etc.); n is the number of positions. It is the similarity between concept c and core concept k; The value lies in the core concept k;

[0120] Step 2: Filter low-value concepts based on the preset "concept value threshold";

[0121] Step 3: Using the DG-TY (Domain Graph-Term Synonym) algorithm, synonymous concepts are merged to ultimately generate the domain ontology framework. The DG-TY algorithm formula is as follows:

[0122] ;

[0123] in, It is the semantic similarity between terms t1 and t2; It refers to the strength of the association between terms in the domain knowledge base; It is the domain adaptation coefficient.

[0124] The analysis mode adaptive mechanism in this embodiment is as follows:

[0125] (I) The data acquisition system for changes in the entity is as follows:

[0126] (1) Core data collection dimensions, as follows:

[0127] ① Structural Change Dimension: Records changes in the core structure of the knowledge graph ontology, specifically including:

[0128] Core concept changes: The number of new and deleted core concepts of the ontology during the statistical period;

[0129] Changes in relationships between concepts: The number of new additions, modifications, and deletions of relationships between core concepts within the statistical period;

[0130] Core framework adjustment: The frequency of adjustments to the core framework of the ontology within the statistical period;

[0131] ② Entity data dimension: Assess the pressure on ontology stability (the larger the amount of entity data, the stronger the impact on ontology structural stability). Specific statistical indicators include:

[0132] Total number of new entities added during the period: The total number of new entities added during the statistical period;

[0133] Entity mount location distribution: Statistics on the proportion of newly added entities belonging to the original ontology and the new ontology;

[0134] Duplicate Entity Frequency: The number of times the same entity is repeatedly entered within a statistical period, associated with different attributes / relationships;

[0135] ③ Semantic consistency dimension: Monitor the matching degree between newly added data and existing ontology semantics to help verify the stability of the ontology structure, specifically including:

[0136] Relationship rule compliance: Determines whether the relationships between newly added concepts comply with domain business rules;

[0137] Semantic conflict frequency: The number of times new data is inconsistent with the semantics of the existing ontology within the statistical period.

[0138] (2) Acquisition frequency and storage format, as detailed below:

[0139] ① Data Acquisition Frequency: The acquisition frequency is dynamically adjusted based on the update frequency of the ontological structure and entity data to ensure that stability changes are captured in a timely manner.

[0140] High-frequency stress scenarios: For scenarios with large daily increases in entity data and real-time updates to core business data, the "real-time acquisition (structural changes + semantic conflicts)" mode is adopted;

[0141] Low-frequency stress scenarios: For scenarios where the daily increase in entity data is small and the core structure is updated weekly, the "daily collection (structural changes + semantic conflicts)" mode is adopted.

[0142] ② Storage format: A hybrid storage architecture combining time-series databases and relational databases is adopted.

[0143] Time-series databases (such as InfluxDB): store ontological structure changes (time and frequency of core concept / relationship changes) and entity data (time series of total new additions and mounting percentages);

[0144] Relational databases (such as MySQL): Store semantically consistent data (semantic similarity results, conflict rule details) and entity data association information (approval business scenarios corresponding to duplicate entities), and establish an association index of "structural change - entity pressure - semantic conflict".

[0145] (II) Adaptive evaluation dimensions and quantitative standards, as detailed below:

[0146] (1) Core evaluation dimensions: The comprehensive score of ontology stability is calculated by weighting “structural stability score”, “entity stress adaptability score” and “semantic consistency score”. The three dimensions reflect ontology stability from the perspectives of “core structural state”, “external stress withstand capability” and “semantic adaptability”, respectively. The specific quantitative standards are shown in Table 1.

[0147] Table 1 Specific Quantitative Standards

[0148]

[0149] (2) Stability level judgment criteria: Based on the comprehensive stability score of the ontology, the ontology stability is divided into three levels, which serve as the direct basis for the matching analysis mode:

[0150] High stability: Overall score ≥ 85 points. Minimal changes in the body structure.

[0151] Medium stability: Overall score 65-84. Occasional changes in the body structure.

[0152] Low stability: Overall score < 65 points. Frequent changes in the structure.

[0153] (III) Adaptive decision-making logic and execution process, as detailed below:

[0154] (1) Stability - Pattern matching logic, as follows:

[0155] ① Basic matching rules: Based on the stability level as the core judgment criterion, directly map the default analysis mode - high stability → fixed analysis mode, medium stability → range analysis mode, low stability → exploratory analysis mode;

[0156] ②Scenario Supplement Rules: If the basic matching result conflicts with the current business scenario requirements, the system will automatically trigger scenario priority determination:

[0157] If the business scenario is marked as "standardized batch processing", then "stability first" will be maintained, and the fixed analysis mode of basic matching will be used.

[0158] If the business scenario is marked as "focusing on a specific range", then "scenario priority" will be triggered, and the mode will be switched to range analysis mode;

[0159] If the business scenario is labeled "Emerging Business Exploration", then regardless of the stability level, the exploration analysis mode will be forcibly applied.

[0160] (2) Conflict Arbitration Mechanism: When the stability assessment results conflict with business scenario requirements and historical pattern selection records, the system initiates a multi-dimensional conflict arbitration mechanism, judging the cases according to priority to ensure the rationality of the decision.

[0161] ① Priority 1: Business rule conflict arbitration: Call the "business rule base" in the domain knowledge base to verify whether the mode meets the mandatory rules - if not, switch to the mode that meets the rules, and record the rule conflict log at the same time;

[0162] ② Priority 2: Historical data conflict arbitration: If the conflict originates from "the current stability level matching mode is significantly different from the historical mode in the same scenario", the system retrieves the stability data and mode execution effect of the same scenario in the past 3 months. If the historical fixed mode execution effect is better than the current matching range mode, the historical fixed mode will be used, and the current data fluctuation will be marked as "temporary anomaly".

[0163] ③ Priority 3: Manual intervention arbitration: If the above two levels of arbitration still cannot resolve the conflict, the system will automatically notify the administrator, and after confirmation, the system will execute the specified mode.

[0164] (3) Mode switching triggering process: The dynamic switching of analysis mode is realized through the dual mechanism of "periodic evaluation triggering + real-time anomaly triggering" to ensure that the ontology construction always adapts to the current state. The specific process is as follows:

[0165] ① Periodic evaluation trigger process (routine switchover), details are as follows:

[0166] Step 1: The system automatically retrieves the collected body change data according to a preset cycle and calculates the overall stability score and level;

[0167] Step 2: Compare the current execution mode with the mode that matches the new stability level. If they do not match, generate a "Mode Switching Assessment Report".

[0168] Step 3: The system automatically verifies the feasibility of switching. If it is feasible, the switch is executed, and the mode configuration parameters are updated synchronously. If it is not feasible, the "parameter completion process" is triggered, which automatically extracts relevant sub-scope information from the domain knowledge base, completes the parameters, and then switches.

[0169] Step 4: After the switch is completed, the system will verify the effect of the first batch of built data. If it meets the standard, the new mode will run stably; if it does not meet the standard, it will revert to the original mode and trigger anomaly investigation.

[0170] ② Real-time anomaly triggering process (emergency switchover), details are as follows:

[0171] Step 1: The system monitors abnormal indicators in real time during the ontology construction process. When the abnormality continues to exceed the preset threshold, an emergency assessment is automatically triggered.

[0172] Step 2: Quickly calculate the "real-time stability score". If the score drops by more than 20 points compared to the previous period, it is judged as a "sudden drop in stability".

[0173] Step 3: The system automatically switches the current mode to the exploratory analysis mode adapted to low stability, and at the same time pauses batch data processing;

[0174] Step 4: Notify the domain administrator to handle the anomaly. After the administrator investigates the cause, the ontology will be reconstructed.

[0175] Example 3: This example provides a method for constructing and optimizing a government knowledge graph ontology. The method is as follows:

[0176] Step 1: System startup and knowledge base initialization (corresponding to "System startup + Knowledge base construction"), as follows:

[0177] (1) Start the ontology building system and automatically load the domain configuration (by default, load the basic parameters of the government administration approval domain).

[0178] (2) Construction of the domain knowledge base, as detailed below:

[0179] ① Collect core knowledge in the field (standardized terms such as "approval items" and "application materials", conceptual associations such as "government affairs APP - government affairs mobile application", business rules such as "approval items need to be associated with the accepting department") and unstructured documents (policy documents, service guides);

[0180] ② Automatically standardize documents (categorize, semantically annotate, and unify the format to JSON-LD), generate a well-structured knowledge base with clear hierarchy, store it in the specified database, complete the knowledge base initialization, and provide a unified knowledge benchmark for subsequent analysis.

[0181] ③ Algorithm Parameter Preset: Based on the characteristics of the government administration approval field, the core algorithm's basic parameters are initialized—among which, the "concept value threshold" is set to 0.66, for the DG-CW algorithm. The adaptation parameters are set to 0.4, 0.6, and 0.7; the adaptation parameters for the DG-TY algorithm. Set the values ​​to 0.7 and 0.3 to complete the knowledge base initialization, providing a unified knowledge benchmark and algorithm parameter support for subsequent analysis.

[0182] Step 2: Input of data to be processed (corresponding to "New file input"): The system receives data to be processed in the field of government administrative approval through API interface or file upload module, supporting multiple formats (such as batch approval data table, user feedback text, policy document). After the data is input, the next adaptive analysis process is automatically triggered without manual intervention.

[0183] Step 3: Adaptive Mechanism for Fully Automatic Mode Selection (corresponding to "Analysis Mode Adaptive Mechanism + Ontology Analysis Mode Selection"), this is the core connecting link in the process. The system automatically completes mode determination and selection based on "knowledge base benchmark + input data features":

[0184] (1) Data collection and stability assessment: Automatically collect historical ontology structure changes, entity data fluctuations, semantic consistency and other data, and combine them with input data characteristics (such as data volume and business relevance) to calculate the comprehensive stability score and level (high / medium / low stability).

[0185] (2) Automatic pattern matching: Map the analysis pattern according to the stability level (high stability → fixed pattern, medium stability → range pattern, low stability → exploration pattern), and simultaneously verify the compatibility between the input data and the pattern (such as standardized data to fit the fixed pattern).

[0186] (3) Conflict and anomaly preprocessing: If the data characteristics conflict with the stability matching mode (such as high stability but new business data is input), the arbitration will be automatically carried out by “business rule verification → historical data comparison”, and a work order will be pushed to the administrator if necessary to ensure accurate mode selection.

[0187] Step 4: Multi-scene ontology analysis and generation (corresponding to "Generate graph ontology"): Based on the adaptive mode selected in Step 3, the system automatically performs ontology analysis and construction, and outputs the graph ontology.

[0188] Example 4: This embodiment of the invention also provides an electronic device, including: a memory and at least one processor;

[0189] The memory stores computer-executed instructions;

[0190] The at least one processor executes the computer execution instructions stored in the memory, causing the at least one processor to execute the government knowledge graph ontology construction and optimization method in any embodiment of the present invention.

[0191] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.

[0192] Memory can be used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0193] Example 5: This example also provides a computer-readable storage medium storing multiple instructions, which are loaded by a processor to cause the processor to execute the government knowledge graph ontology construction and optimization method in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium may be provided, on which software program code implementing the functions of any of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device may read and execute the program code stored in the storage medium.

[0194] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of the present invention.

[0195] Storage media embodiments for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.

[0196] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0197] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion unit connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion unit execute some and all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0198] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing and optimizing a government knowledge graph ontology, characterized in that, The method is as follows: Building a domain knowledge base: The ontology building system is launched, automatically loading the basic configuration parameters of the target domain, collecting the core knowledge elements and unstructured documents of the target domain, and classifying, semantically annotating and unifying the format of the unstructured documents to generate a hierarchical structured domain knowledge base as a unified knowledge benchmark for ontology building; among which, the core knowledge elements include standardized terminology, conceptual relationships and basic business rules. Data Acquisition and Analysis Process Triggering: Acquire data to be processed in the target domain via API interface or file upload, and automatically trigger subsequent adaptive analysis processes; Adaptive selection of ontology analysis mode: Automatically collect historical ontology structure change data, entity data fluctuation data, and semantic consistency data, combine the characteristics of the data to be processed to calculate the comprehensive ontology stability score, determine the stability level, and map the corresponding ontology analysis mode based on the stability level. Simultaneously verify the compatibility between the data to be processed and the mapping mode: If there is a compatibility conflict, the conflict is arbitrated according to the priority of "business rules → historical data comparison → manual intervention" to determine the final ontology analysis mode. A knowledge graph ontology is generated based on the selected ontology analysis mode, and entity extraction, data processing and ontology construction are automatically performed according to the final determined mode to obtain a structured domain knowledge graph ontology; among which, the ontology analysis mode includes fixed analysis mode, scope analysis mode and exploratory analysis mode; The stability levels include high stability, medium stability, and low stability. High stability refers to a comprehensive stability score ≥ 85 points, corresponding to a fixed mapping analysis mode; medium stability refers to a comprehensive stability score of 65-84 points, corresponding to a range mapping analysis mode; low stability refers to a comprehensive stability score < 65 points, corresponding to an exploratory mapping analysis mode. The fixed analysis mode is as follows: Automatically loads a pre-defined, unmodifiable ontology framework, which contains the core concepts and fixed relationships of the target domain; Only extract specific entity information from the data to be processed, without generating new concepts or adjusting existing relationships; The extracted entities are mounted to the corresponding nodes according to the relationships within the ontology framework, thereby enabling incremental data supplementation within the framework. The specific range analysis mode is as follows: Automatically identify the business sub-scope associated with the data to be processed, and determine the conceptual analysis boundary of the ontology; Only extract the concepts and relationships directly related to the boundary of the concept analysis; Automatically identify and filter irrelevant information that exceeds the boundaries of conceptual analysis, and generate an ontology that focuses on a specific business sub-scope; The specific exploration and analysis modes are as follows: Automatically set the concept value threshold, which is based on the frequency of text occurrence, semantic relevance, and domain relevance. Perform a full scan of the data to be processed to extract explicit concepts, implicit concepts, and potential relationships; The comprehensive value of each concept is calculated using the DG-CW concept value algorithm, filtering out low-value concepts below the concept value threshold. The DG-CW concept value algorithm formula is as follows: ; ; in, Indicates the adaptation parameters; This indicates the frequency of occurrence of concept c; This indicates the highest frequency of all concepts; This represents the value of each position in the title, core paragraph, and non-core paragraph; n represents the number of positions. This represents the similarity between concept c and core concept k; This indicates the value of the core concept k; A preliminary ontology framework is generated by merging synonymous concepts and unifying terminology using the DG-TY domain synonymy algorithm; the formula for the DG-TY domain synonymy algorithm is as follows: ; in, Representation of terms and semantic similarity; Indicates the strength of the association between terms in the domain knowledge base; Indicates the domain adaptation coefficient; The conflict arbitration proceedings are as follows: If the conflict stems from the pattern not conforming to the mandatory business rules of the target domain, then the business rule base in the domain knowledge base is called for verification, and the pattern is switched to conform to the mandatory rules. If the conflict stems from a significant difference between the current matching mode and the historical matching mode for the same scenario, then retrieve the performance data of the matching mode for the same scenario over the past 3 months and use the mode with the better performance. If the corresponding arbitration cannot resolve the conflict, a conflict arbitration work order will be automatically generated and pushed to the domain administrator. The administrator will then execute the specified mode according to the decision and record the decision result in the decision knowledge base.

2. The method for constructing and optimizing a government knowledge graph ontology according to claim 1, characterized in that, Standardized terminology covers industry-wide terms and discipline-specific terms in the target field; conceptual relationships include synonyms, near-synonyms, and core concepts.

3. A device for constructing and optimizing a government knowledge graph ontology, characterized in that, This device is used to implement the method for constructing and optimizing a government knowledge graph ontology as described in claim 1 or 2; the device is specifically as follows: The knowledge base construction module is used to start the ontology construction system, automatically load the basic configuration parameters of the target domain, collect the core knowledge elements and unstructured documents of the target domain, and classify, semantically annotate and unify the format of the unstructured documents to generate a hierarchical structured domain knowledge base as a unified knowledge benchmark for ontology construction. Among them, the core knowledge elements include standardized terminology, conceptual relationships and basic business rules. The data acquisition module is used to acquire data to be processed in the target domain through API interface or file upload, and automatically trigger subsequent adaptive analysis processes. Adaptive mode selection module: used to automatically collect historical ontology structure change data, entity data fluctuation data and semantic consistency data, calculate the comprehensive ontology stability score by combining the characteristics of the data to be processed, determine the stability level, and map the corresponding ontology analysis mode based on the stability level, and simultaneously verify the compatibility between the data to be processed and the mapping mode: if there is a compatibility conflict, the conflict is arbitrated according to the priority of "business rules → historical data comparison → manual intervention" to determine the final ontology analysis mode. The ontology generation module is used to generate a knowledge graph ontology based on the selected ontology analysis mode, and automatically perform entity extraction, data processing and ontology construction according to the finally determined mode to obtain a structured domain knowledge graph ontology; wherein, the ontology analysis mode includes fixed analysis mode, scope analysis mode and exploratory analysis mode; The storage module is used to store the domain knowledge base, historical ontology change data, data to be processed, and the generated knowledge graph ontology; The anomaly monitoring module is used to monitor the semantic conflict rate and the proportion of external entities during the ontology construction process in real time. When the abnormal indicators continue to exceed the preset threshold, it triggers emergency mode switching and abnormal data isolation, and simultaneously notifies the domain administrator.

4. An electronic device, characterized in that, include: Memory and at least one processor; The memory contains computer programs; The at least one processor executes the computer program stored in the memory, causing the at least one processor to perform the government knowledge graph ontology construction and optimization method as described in claim 1 or 2.

5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method for constructing and optimizing a government knowledge graph ontology as described in claim 1 or 2.

Citation Information

Patent Citations

  • Knowledge management and control method and device based on knowledge graph

    CN115640403A

  • Military software defect multi-modal knowledge graph construction method, device and system

    CN116860986A