A government affair data automatic classification and grading method

By standardizing government data formats, constructing ontology and knowledge graphs, and combining machine learning and ontology reasoning technologies, the problems of data format diversity and ambiguous category boundaries have been solved, enabling efficient and secure data management and classification decisions.

CN120448890BActive Publication Date: 2025-12-05SCI CITY (GUANGZHOU) INFORMATION TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510451565.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-12-05
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The collection and processing of government data suffers from diverse data formats and blurred category boundaries, leading to uncertainty in classification decisions and affecting management efficiency and accuracy.

Method used

By standardizing data formats, we construct government domain ontology and knowledge graphs, use machine learning algorithms for data correction and classification, combine ontology reasoning and knowledge fusion technologies, formulate data grading standards, and establish dynamic adjustment and update mechanisms to ensure the flexibility and security of data management.

Benefits of technology

It has improved the efficiency and accuracy of government data management, ensured the robustness and reliability of data classification, realized the intelligence and security of data management, and provided a transparent accountability tracking mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448890B_ABST
    Figure CN120448890B_ABST
Patent Text Reader

Abstract

The application discloses a kind of government affair data automatic classification grading method, it is related to data management and information processing technical field.The application uses regular expression and pattern matching technology to convert data of different sources and formats into structured format, and improves the consistency and accuracy of data through machine learning algorithm, data quality detection, abnormal correction and other means, in the semantic enhancement and correlation analysis stage of data, by constructing government domain ontology and knowledge graph, combined with natural language processing technology, effectively classifies the semantics of government data, solves the overlap and uncertainty between categories, further improves the robustness and accuracy of classification decision through ensemble learning algorithm, ensures the efficiency and reliability of classification, by establishing data grading dynamic adjustment and updating mechanism, combined with real-time data monitoring, sensitivity evaluation and automatic adjustment of access control rules, ensures the flexibility and security of data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management and information processing technology, specifically a method for automatic classification and grading of government data. Background Technology

[0002] The collection and organization of government data faces the challenge of diverse data formats. Different types of documents, such as files, reports, and tables, vary significantly in their content organization and information presentation, making unified data processing difficult. When automatically classifying government data, the correlation and overlap between categories such as policy, law, finance, and administration lead to blurred category boundaries and uncertainty in classification decisions. Summary of the Invention

[0003] The purpose of this invention is to provide an automatic classification and grading method for government data. By standardizing data formats, constructing government domain ontology and knowledge graphs, and implementing dynamic data sensitivity assessment and update management strategies, this method solves the problems of data format diversity and classification uncertainty, thereby improving the efficiency and accuracy of government data management.

[0004] The objective of this invention can be achieved through the following technical solutions:

[0005] This application provides a method for automatically classifying and grading government data, including the following steps:

[0006] By defining data quality rules and constraints, statistical methods and machine learning algorithms are used to automatically correct and standardize data.

[0007] Based on automatically corrected and standardized data, a government domain ontology and knowledge graph are constructed to perform semantic enhancement and correlation analysis on government data. By constructing a domain concept hierarchy and relational network model, ontology reasoning and knowledge fusion technologies are used to classify data and make decisions. At the same time, a multi-classifier ensemble learning algorithm is used to integrate the classification results of different models.

[0008] Based on the characteristics and application needs of government data, a government domain ontology and knowledge graph are constructed, defining the core concepts, attributes and relationships of government data, and using the knowledge graph to express the relationships between data to form a conceptual hierarchy and relationship network model.

[0009] For the automatically corrected and standardized government data, natural language processing technology is used to perform semantic analysis on policy documents and announcements. Entity recognition technology is used to extract key policy measures, scope of impact and expected goals from policy documents, and semantic enhancement of the data is performed. Combined with ontology reasoning and knowledge graphs, cross-dataset association analysis is conducted.

[0010] Based on domain ontology and knowledge graph, we use ontology reasoning rules and knowledge fusion algorithms to perform correlation analysis on government data.

[0011] Based on the results of semantic enhancement and association analysis, a data classification model is trained using machine learning algorithms to automatically classify and support government data for decision-making.

[0012] In the process of data classification and decision-making, ontology reasoning technology is used to infer knowledge and derive implicit classification rules and decision-making basis based on the concept hierarchy and relationship network in the domain ontology.

[0013] By using knowledge fusion technology, the classification rules and decision criteria obtained from ontology reasoning are fused with the prediction results of machine learning models, and semantic information and data features are integrated to obtain the final classification and decision results.

[0014] Based on the classification results, a unified standard and specification for the hierarchical classification of government data is formulated. Based on the confidentiality dimension of the data, a data sensitivity assessment model and hierarchical decision-making mechanism are established to determine the data level. Through consultation and review by multiple stakeholders, a consensus-based data hierarchical classification scheme is formed.

[0015] After reaching a consensus on a data classification and grading scheme, a knowledge base and rule base for the government domain are constructed to determine knowledge reasoning and rule matching, and data is mapped to concepts and entities in the domain knowledge base to perform semantic association and contextual understanding of the data.

[0016] Establish a dynamic adjustment and update mechanism for data classification. By monitoring data usage in real time, dynamically assess the sensitivity and importance of data, and automatically trigger the reassessment and adjustment process of data classification based on preset adjustment rules and thresholds, update data management strategies and access control rules.

[0017] Before defining data quality rules and constraints, the process also includes: collecting data from different sources and types, converting it into a standardized structured format, and extracting key information using regular expressions based on the content and characteristics of the data, and mapping it to a predefined data model to form a unified data representation.

[0018] Furthermore, key information is extracted using regular expressions and mapped to a predefined data model to form a unified data representation, specifically including:

[0019] Acquire multi-source heterogeneous data, employ appropriate data cleaning and preprocessing methods to convert it into a standardized structured format, and determine the key information types and features to be extracted based on a predefined data model;

[0020] For standardized structured data, regular expression matching methods are used to extract key information that conforms to specific characteristics;

[0021] The extracted key information is mapped to the corresponding fields and attributes in the data model according to predefined mapping rules, and the mapped data is then validated and its integrity is checked.

[0022] The mapped data is stored in a unified format and structure to form a standardized data representation.

[0023] Furthermore, statistical methods and machine learning algorithms are used to automatically correct and standardize the data, specifically including:

[0024] Based on predefined data quality rules and constraints, a quality check is performed on the unified data to identify abnormal data that does not meet the quality requirements. Statistical methods are used to analyze the abnormal data, calculate the statistical characteristics of the abnormal data, and determine the distribution of the abnormal data.

[0025] By using anomaly detection algorithms in machine learning algorithms, abnormal data is automatically detected and identified, the location and range of abnormal data are determined, and then a data imputation algorithm is used to automatically estimate and fill in missing or abnormal data values.

[0026] According to the preset data standardization rules, the filled data is automatically standardized to convert the data into a unified format, unit and representation. Then the standardized data is quality verified and judged to see if the data meets the quality requirements according to the data quality rules and constraints.

[0027] The processed, high-quality data is stored in the database.

[0028] Furthermore, based on the confidentiality dimension of the data, a data sensitivity assessment model and a hierarchical decision-making mechanism are established to determine the data sensitivity levels, specifically including:

[0029] The government data set is acquired, and the metadata information of the data set is parsed using a natural language processing algorithm based on a pre-established rule base to obtain the data type and data source. It is then determined whether the data type belongs to a sensitive type. If it belongs to a sensitive type, the data set is classified into the dataset to be evaluated.

[0030] Data content is obtained from the dataset to be evaluated. Based on a pre-established keyword list, sensitive words are extracted from the data content using a text matching algorithm. The number and types of sensitive words are obtained, the weights of the sensitive words are determined, and the sensitivity score is calculated. Based on the sensitivity score and a preset threshold range, the data security level is determined.

[0031] Furthermore, when the data security level exceeds a preset threshold, the dataset is desensitized to obtain a desensitized dataset. Desensitization methods include data masking and data replacement.

[0032] User access records are obtained from the anonymized dataset. Based on user roles and permission matrices, the range of data that users can access is determined, and user access permissions to specific datasets are identified. When a user attempts to access data beyond their authorized access, a violation log is recorded.

[0033] Based on the data security level and sharing protocol, the sharing scope of the dataset is determined. When the data security level is high, the sharing scope is limited to a specific department. Data is transmitted through an interface to obtain the shared dataset.

[0034] Data update information is obtained from the shared dataset. The data update frequency is determined according to a pre-established schedule. When the data update frequency is high, the data synchronization mechanism is triggered to obtain the latest data version.

[0035] Based on the data storage period and destruction strategy, the data retention time is determined. When the data reaches the storage period, the data destruction process is triggered to delete the data from the storage medium and obtain the destroyed data record.

[0036] Furthermore, by constructing a knowledge base and rule base for the government domain, knowledge reasoning and rule matching are determined, and data is mapped to concepts and entities in the domain knowledge base to perform semantic association and contextual understanding of the data. Specifically, this includes:

[0037] By comparing structured data with data from the concept hierarchy and relational network, if the matching degree is higher than a preset threshold, the concept to which the data belongs is determined; if it is lower than the preset threshold, the structured data is compared with the entity layer to obtain the entity with the highest matching degree, thus obtaining the set of entities corresponding to the data.

[0038] By calling association rules in the rule base from the entity set, the relationships between entities are obtained. Based on the relationships between entities, an entity relationship graph is constructed. By traversing the graph, the data context information is determined, and the final semantic understanding result is obtained.

[0039] Furthermore, a data tiered dynamic adjustment and update mechanism will be established, specifically including:

[0040] Data collection and recording: Collect data operation logs, obtain operation user, operation type, operation time, and operation data object identification information through log parsing technology, and record them in the data usage record database;

[0041] Using frequency statistics and sensitivity assessment, the database information of data usage records is used to count the frequency of data object identifiers being used within a preset time range. When the frequency exceeds a preset frequency threshold, the data sensitivity and criticality assessment process is triggered. A pre-established assessment model is used to calculate the data object identifier set to obtain the data sensitivity set and the data criticality set.

[0042] When data classification adjustment is triggered, the data sensitivity set and data criticality set are obtained. When an indicator in the set exceeds the preset indicator threshold, the data classification adjustment process is triggered, the data classification rule base is called, and the corresponding classification rules are matched according to the data sensitivity set and data criticality set to generate a new classification result for the data object identifier set.

[0043] The control strategy is updated by obtaining the new classification result of the data object identifier set, determining the control strategy set corresponding to the new classification result according to the preset data classification and control strategy comparison table, comparing the original control strategy with the new control strategy using a difference analysis algorithm, and generating a control strategy update instruction set.

[0044] When adjusting access control rules, the system obtains a set of management policy update instructions, parses the instruction set through the rule engine, and triggers the access control rule adjustment process when the instruction type is access control rule adjustment. This process calls the access control rule generation module to generate a new set of data object identifier access control rules.

[0045] Verify the new access control rules. Based on the new data object identifier set access control rules, use a machine learning algorithm based on the Transformer model to verify the new access control rules and determine the degree of impact of the new access control rules on data security risk factors. If the degree of impact is lower than the preset threshold, the new access control rules are published to the access control system.

[0046] Data access behavior monitoring and dynamic adjustment: Based on the data object identifier set access logs fed back by the access control system, clustering machine learning algorithm combined with K-nearest neighbor machine learning algorithm is used to identify data access behavior patterns, determine the degree of deviation between the data access behavior pattern and the preset behavior pattern, and when the deviation exceeds the preset threshold, a data hierarchical dynamic adjustment process is carried out, and data operation logs are re-collected.

[0047] Furthermore, after updating data management strategies and access control rules, it also includes: establishing an audit and traceability mechanism for data classification changes, and recording the history and reasons for data classification adjustments.

[0048] Furthermore, establish an audit and traceability mechanism for data classification changes, recording the history and reasons for data classification adjustments, specifically including:

[0049] According to the preset periodic data collection hierarchy directory, the hierarchy information of each data item in the data hierarchy directory is obtained to obtain the hierarchy set. It is determined whether the data hierarchy set has changed within the period. If it has changed, the data hierarchy change record process is triggered. The changed data item set is determined by data comparison.

[0050] Then, a text analysis algorithm is used to parse the business log associated with the changed data item to obtain the text describing the reason for the change. The text describing the reason for the change is then parsed using an information extraction algorithm to obtain the phrase describing the reason for the change. Based on the phrase describing the reason for the change, the tag for the reason for the change is determined.

[0051] The system uses preset rules to match each data item in the set of changed data items, obtains the operator information of the changed data items, obtains the operator set and the identity information of each operator, and uses an information encryption algorithm to encrypt the identity information of each operator to obtain encrypted identity information.

[0052] Data change records are generated based on the set of changed data items, change reason tags, and encrypted identity information. These records are then stored on the blockchain to obtain a unique identifier for each data change record.

[0053] The beneficial effects of this invention are as follows:

[0054] By standardizing the processing of government data, extracting key information, and automatically classifying and correcting it, the problems of data format diversity and ambiguous category boundaries in the background technology are solved. First, regular expressions and pattern matching techniques are used to convert data from different sources and formats into a structured format. Then, machine learning algorithms, data quality detection, and anomaly correction are used to improve the consistency and accuracy of the data. In the semantic enhancement and association analysis stage, by constructing a government domain ontology and knowledge graph, combined with natural language processing technology, the government data is effectively semantically classified, which solves the overlap and uncertainty between categories. Ensemble learning algorithms are used to further improve the robustness and accuracy of classification decisions, ensuring the efficiency and reliability of classification.

[0055] By establishing a dynamic adjustment and update mechanism for data classification, combined with real-time data monitoring, sensitivity assessment, and automatic adjustment of access control rules, the flexibility and security of data management are ensured. By continuously assessing the sensitivity and usage of data, a reassessment of data classification is triggered, and corresponding management strategies and access controls are updated to ensure the security and compliance of government data. At the same time, the audit and traceability mechanism for data classification changes utilizes blockchain technology to store change records immutably, providing transparency and accountability for data management, thereby effectively improving the efficiency, security, and intelligence of government data management. Attached Figure Description

[0056] To better understand and implement this application, the technical solution is described in detail below with reference to the accompanying drawings.

[0057] Figure 1 A flowchart illustrating an automatic classification and grading method for government data provided in this application;

[0058] Figure 2 A flowchart illustrating the process of forming a unified data representation for an automatic classification and grading method for government data provided in this application;

[0059] Figure 3 This application provides a flowchart illustrating the automatic correction and standardization process for an automatic classification and grading method for government data. Detailed Implementation

[0060] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, exemplary embodiments will be described in detail below, examples of which are illustrated in the accompanying drawings. In the following description relating to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods and systems consistent with some aspects of this application as detailed in the appended claims.

[0061] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0062] The following detailed description of the specific implementation methods, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided in detail.

[0063] Please see Figures 1-3 This embodiment provides a method for automatically classifying and grading government data, including the following steps:

[0064] S1. Collect data from different sources and types, convert it into a standardized structured format, and extract key information using regular expressions based on the content and characteristics of the data, then map it into a predefined data model to form a unified data representation;

[0065] Furthermore, key information is extracted using regular expressions and mapped to a predefined data model to form a unified data representation, specifically including:

[0066] S11. Acquire multi-source heterogeneous data. For data from different sources and in different formats, adopt appropriate data cleaning and preprocessing methods to convert them into a standardized structured format. Based on the predefined data model, determine the key information types and features to be extracted.

[0067] S12. For standardized structured data, use pattern matching methods such as regular expressions to extract key information that conforms to specific characteristics;

[0068] S13. Map the extracted key information to the corresponding fields and attributes in the data model according to the predefined mapping rules, and perform verification and integrity checks on the mapped data to ensure the accuracy and consistency of the data.

[0069] S14. Store the mapped data in a unified format and structure to form a standardized data representation. Based on the standardized data representation, use machine learning algorithms, such as clustering, classification, and association rule mining, to discover patterns and rules in the data, supporting business decisions and optimization.

[0070] Specifically, by standardizing and extracting key information from multi-source heterogeneous data, and using techniques such as regular expressions to transform the data into a unified structured format, data consistency and accuracy are ensured. Based on this, key information is mapped according to a predefined data model, and data quality is guaranteed through validation and integrity checks. Ultimately, data stored in a unified format can support further analysis based on machine learning algorithms, such as pattern recognition and rule mining, providing intelligent support for business decisions, improving data utilization efficiency and value, and driving business optimization and intelligent decision-making.

[0071] S2. Based on unified data, by defining data quality rules and constraints, and using statistical methods and machine learning algorithms, such as anomaly detection and data imputation, the data is automatically corrected and standardized to improve its accuracy and usability.

[0072] Furthermore, statistical methods and machine learning algorithms, such as anomaly detection and data imputation, are used to automatically correct and standardize the data, specifically including:

[0073] S21. Based on predefined data quality rules and constraints, perform quality checks on the unified data, identify anomalous data that does not meet quality requirements, and analyze the anomalous data using statistical methods to calculate its statistical characteristics, such as mean, variance, and quantiles, to determine the distribution of the anomalous data. The quality rules and constraints include requirements for data format, data type, data range, data integrity, and data consistency, defining the standards that must be followed during data collection, storage, processing, and analysis to ensure that the data meets specific business needs and analytical objectives. Through quality rules and constraints, errors, omissions, and inconsistencies in the data can be identified and corrected, thereby improving the overall quality of the data and making it more suitable for subsequent analysis and decision-making processes.

[0074] S22. Through anomaly detection algorithms in machine learning algorithms, such as isolated forest and local anomaly factor, abnormal data is automatically detected and identified to determine the location and range of abnormal data. Then, data imputation algorithms, such as mean imputation and nearest neighbor imputation, are used to automatically estimate and fill in missing or abnormal data values ​​to ensure data integrity.

[0075] S23. According to the preset data standardization rules, the filled data is automatically standardized to convert the data into a unified format, unit and representation method to improve data consistency. Then, the standardized data is quality verified. According to the data quality rules and constraints, it is determined whether the data meets the quality requirements to ensure the accuracy and usability of the processed data.

[0076] S24. The processed high-quality data is stored in the database for subsequent data analysis and application. The data processing flow is completed, and the data quality and availability are significantly improved.

[0077] Specifically, by defining data quality rules and constraints, and combining statistical methods and machine learning algorithms, a comprehensive improvement in data quality was achieved. Anomaly detection algorithms (such as isolated forest and local outlier detection) accurately identify anomalous data, while data imputation algorithms (such as mean imputation and nearest neighbor imputation) effectively repair missing or outliers, ensuring data integrity. Furthermore, unified data standardization improves consistency, while quality verification further guarantees data accuracy and usability. Ultimately, the processed, high-quality data storage provides reliable support for subsequent analysis and applications, significantly improving the efficiency of data management and utilization.

[0078] S3. Based on automatically corrected and standardized data, construct a government domain ontology and knowledge graph, perform semantic enhancement and correlation analysis on government data, construct a domain concept hierarchy and relational network model, utilize ontology reasoning and knowledge fusion technology to classify data and make decisions, and use a multi-classifier ensemble learning algorithm to integrate the classification results of different models to improve the robustness and accuracy of classification; solve the problems of fuzzy categories and classification uncertainty of government data.

[0079] Furthermore, based on the characteristics and application needs of government data, an ontology and knowledge graph for the government domain are constructed, defining the core concepts, attributes, and relationships of government data, and using the knowledge graph to express the relationships between data to form a conceptual hierarchy and relationship network model;

[0080] For automatically corrected and standardized government data, natural language processing technology is used to perform semantic analysis on policy documents and announcements. Entity recognition technology is used to extract key policy measures, scope of impact, and expected goals from a large number of policy documents, such as the relationship between "policy" entities and "industry" entities, to enhance the semantics of the data. This helps to more effectively identify key entities, attributes, and relationships in the automatically corrected and standardized government data, thereby enhancing the semantics of the data. Combined with ontology reasoning and knowledge graphs, cross-dataset association analysis is carried out.

[0081] Based on domain ontology and knowledge graph, we use ontology reasoning rules and knowledge fusion algorithms to conduct correlation analysis on government data, discover implicit semantic relationships and patterns, and enrich the semantic representation of data. For example, by analyzing the data in the "Government Overview" module, we can discover the potential link between government service efficiency and public satisfaction. The Government Overview includes a display of the general situation of urban management, emergency response, human resources and social security, education, medical and health care, housing and construction, and civil affairs.

[0082] Based on the results of semantic enhancement and association analysis, machine learning algorithms such as decision trees and support vector machines are used to train a data classification model to automatically classify and support decision-making for the data in the "Government Overview" module of government data.

[0083] In the process of data classification and decision-making, ontology reasoning technology is used to perform knowledge deduction. Based on the concept hierarchy and relationship network in the domain ontology, implicit classification rules and decision-making basis are derived to enhance the interpretability of classification and decision-making. Specifically, by analyzing the data in the "Overview of Ecological Civilization" module, a balance strategy between environmental governance and economic development is derived. The "Overview of Ecological Civilization" includes displaying meteorological microservices, natural resources, and ecological environment overview.

[0084] By using knowledge fusion technology, the classification rules and decision criteria obtained from ontology reasoning are fused with the prediction results of machine learning models, and semantic information and data features are integrated to obtain the final classification and decision results, thereby improving the overall accuracy and reliability.

[0085] Suppose we are building a smart city management system that needs to process and analyze a large amount of government data, including traffic flow, public safety incidents, and environmental monitoring data. First, through automatically corrected and standardized data, we construct an ontology and knowledge graph containing core concepts, attributes, and relationships of urban management, such as entities like "traffic flow," "safety incidents," and "environmental indicators," as well as the relationships between them, such as "traffic flow" affecting "public safety." Next, we use natural language processing technology to perform semantic analysis on policy documents and announcements, extracting key policy measures, their scope of impact, and expected goals through entity recognition technology, such as identifying the relationship between "traffic control policy" and "traffic flow," thus performing semantic enhancement on the data. Then, based on the domain ontology and knowledge graph, the system uses ontology reasoning rules and knowledge fusion algorithms to perform correlation analysis on government data, discovering potential links such as "increased traffic flow" and "public safety incidents." Based on the results of these semantic enhancements and correlation analyses, machine learning algorithms, such as decision trees and support vector machines, are used to train a data classification model to automatically classify and support the data in the "Government Overview" module. Through ensemble learning and multi-classifier algorithms, the system can automatically classify urban traffic flow data into "peak hours" and "off-peak hours," identify "traffic congestion" events related to high traffic flow, and predict potential "public safety risks," such as areas with high traffic accident rates. It can also identify areas with air quality index exceeding standards based on environmental monitoring data and classify them as environmental problems that "require urgent intervention."

[0086] In the process of data classification and decision-making, ontology reasoning technology is used to deduce knowledge inference and derive strategies for balancing environmental governance and economic development. For example, by analyzing data in the "Overview of Ecological Civilization" module, it is possible to deduce how to reduce environmental pollution while maintaining economic growth. Finally, through knowledge fusion technology, the classification rules and decision-making basis obtained from ontology reasoning are fused with the prediction results of machine learning models, and semantic information and data features are integrated to obtain the final classification and decision-making results, thereby improving the overall accuracy and reliability and ensuring the scientific nature and effectiveness of urban management decisions.

[0087] Specifically, by constructing a government domain ontology and knowledge graph, and combining natural language processing (NLP) techniques with machine learning algorithms, semantic enhancement and accurate classification of government data can be effectively achieved, solving the problems of diverse data formats and ambiguous category boundaries. Specifically, the government data is first automatically corrected and standardized. Then, NLP techniques are used to extract key entities and attributes. Next, ontology reasoning and knowledge fusion techniques are used to discover implicit semantic relationships between data. Finally, ensemble learning algorithms are used to integrate the results of multiple classification models, improving the robustness and accuracy of classification. This process not only enhances the semantic representation of the data but also improves the interpretability of classification decisions, thereby solving the problems of uncertainty and category ambiguity in government data classification mentioned in the background technology.

[0088] S4. Based on the classification results, formulate unified standards and norms for the hierarchical classification of government data. Based on the confidentiality dimension of the data, establish a data sensitivity assessment model and hierarchical decision-making mechanism, determine the data level, and form a consensus on the data hierarchical classification scheme through consultation and review by multiple stakeholders, and embed it into the data lifecycle management process.

[0089] Furthermore, based on the confidentiality dimension of the data, a data sensitivity assessment model and a hierarchical decision-making mechanism are established to determine the data sensitivity levels, specifically including:

[0090] The government data set is acquired, and the metadata information of the data set is parsed using natural language processing algorithms according to the pre-established rule base to obtain the data type and data source. It is then determined whether the data type belongs to a sensitive type. If it belongs to a sensitive type, the data set is divided into the dataset to be evaluated. The government data set includes: population data, population statistics information, including registered population, floating population, population structure (age, gender, education level, etc.).

[0091] Economic data, including GDP, fiscal revenue and expenditure, tax data, industrial and agricultural output, commercial and service sector data, investment and consumption data;

[0092] Social management data, public safety data (such as crime rates and traffic accident statistics), public health data (such as disease control and medical services), and education data (such as the number of schools, the number of students, and educational outcomes);

[0093] Urban management data, including urban planning, construction, traffic flow, municipal facilities management, environmental protection, and pollution control;

[0094] Government service data includes information on the acceptance, processing, and feedback of government service matters, including administrative approvals, public services, and social affairs management.

[0095] Public resource data, including land resources, water resources, and energy consumption and allocation;

[0096] Social security data, including data on social insurance, social assistance, and social welfare.

[0097] Data content is obtained from the dataset to be evaluated. Based on a pre-established keyword list, sensitive words are extracted from the data content using a text matching algorithm. The number and types of sensitive words are obtained, the weights of the sensitive words are determined, and the sensitivity score is calculated. Based on the sensitivity score and a preset threshold range, the data security level is determined.

[0098] Furthermore, when the data security level exceeds a preset threshold, the dataset is desensitized to obtain a desensitized dataset. Desensitization methods include data masking and data replacement.

[0099] User access records are obtained from the anonymized dataset. Based on user roles and permission matrices, the range of data that users can access is determined, and user access permissions to specific datasets are identified. When a user attempts to access data beyond their authorized access, a violation log is recorded.

[0100] Based on the data security level and sharing protocol, the sharing scope of the dataset is determined. When the data security level is high, the sharing scope is limited to a specific department. Data is transmitted through an interface to obtain the shared dataset.

[0101] Data update information is obtained from the shared dataset. The data update frequency is determined according to a pre-established schedule. When the data update frequency is high, the data synchronization mechanism is triggered to obtain the latest data version.

[0102] Based on the data storage period and destruction strategy, the data retention time is determined. When the data reaches the storage period, the data destruction process is triggered to delete the data from the storage medium and obtain the destroyed data record.

[0103] Specifically, this process can identify and process sensitive information in government data. Through natural language processing and a pre-defined rule base, it automatically classifies datasets, assesses their sensitivity, determines their security level, and performs anonymization based on sensitivity scores. It also manages user access permissions to ensure data access compliance and controls the scope of data sharing based on security levels and sharing protocols. Furthermore, it monitors data update frequency and enforces data storage expiration policies, triggering a destruction process when data reaches its retention period, ensuring data security and compliance. Overall, this process improves the automation and accuracy of government data management, reduces the risk of data breaches, and optimizes data lifecycle management.

[0104] S5. After reaching a consensus on the data classification scheme, a knowledge base and rule base for the government domain are constructed to determine knowledge reasoning and rule matching, and the data is mapped to the concepts and entities in the domain knowledge base to perform semantic association and contextual understanding of the data.

[0105] Furthermore, by constructing a knowledge base and rule base for the government domain, knowledge reasoning and rule matching are determined, and data is mapped to concepts and entities in the domain knowledge base to perform semantic association and contextual understanding of the data. Specifically, this includes:

[0106] By comparing structured data with data from conceptual hierarchies and relational networks, if the matching degree is higher than a preset threshold, the concept to which the data belongs is determined. If it is lower than the preset threshold, the structured data is compared with the entity layer to obtain the entity with the highest matching degree, thus obtaining the set of entities corresponding to the data. For example, matching "GDP" data with the concept of "GDP" in the knowledge base, for data with a matching degree lower than the preset threshold, it is further compared with the entity layer to determine the specific entity corresponding to the data, such as "a company name" and the "company" entity in the knowledge base.

[0107] By using entity sets and calling association rules from the rule base, the relationships between entities are obtained, such as the relationship between "Company A" and "Industry B". Based on the relationships between entities, an entity relationship graph is constructed. By traversing the graph, data context information is determined, such as the status and influence of "Company A" in "Industry B", and the final semantic understanding result is obtained.

[0108] For example, in the tax field, a knowledge base containing concepts and entities such as "taxpayer," "tax type," and "tax rate" is constructed. When a company's tax return data is received, it is first matched against the "taxpayer" entity in the knowledge base. If the match exceeds a preset threshold, the data is determined to belong to the "taxpayer" entity. For data with a low match, it is further compared with the entity layer, such as matching "company name" with the "company" entity in the knowledge base to find the corresponding entity set. Subsequently, by calling association rules in the rule base, a tax relationship is discovered between "company A" and "industry B," constructing an entity relationship graph. The graph is then traversed to determine the status and influence of "company A" in "industry B," ultimately obtaining the semantic understanding result of the tax return data, namely, the tax contribution and economic activity importance of "company A" in "industry B." This entire process not only improves the semantic association and contextual understanding capabilities of the data but also provides strong data support for tax decision-making.

[0109] Specifically, by constructing a knowledge base and rule base for the government domain, and combining knowledge reasoning and rule matching technologies, semantic association and contextual understanding of government data were achieved. Based on a data hierarchical classification scheme, natural language processing technology was used to transform raw data into a structured form, and the semantic attribution of the data was accurately determined by matching it with concepts and entities in the knowledge base. Furthermore, association rules in the rule base were used to construct an entity relationship graph, revealing the relationships and semantic context between data. Ultimately, the deep understanding of data semantics significantly improved the intelligent processing capabilities of government data, providing strong support for optimizing government services and making scientific decisions.

[0110] S6. Establish a dynamic adjustment and update mechanism for data classification. By monitoring data usage in real time, dynamically assess the sensitivity and importance of data, and automatically trigger the reassessment and adjustment process of data classification according to preset adjustment rules and thresholds, and update data management strategies and access control rules.

[0111] Furthermore, a data tiered dynamic adjustment and update mechanism will be established, specifically including:

[0112] Data collection and recording: Collect data operation logs, obtain operation user, operation type, operation time, and operation data object identification information through log parsing technology, and record them in the data usage record database;

[0113] Using frequency statistics and sensitivity assessment, the database information of data usage records is used to count the frequency of data object identifiers being used within a preset time range. When the frequency exceeds a preset frequency threshold, the data sensitivity and criticality assessment process is triggered. A pre-established assessment model is used to calculate the data object identifier set to obtain the data sensitivity set and the data criticality set.

[0114] When data classification adjustment is triggered, the data sensitivity set and data criticality set are obtained. When an indicator in the set exceeds the preset indicator threshold, the data classification adjustment process is triggered, the data classification rule base is called, and the corresponding classification rules are matched according to the data sensitivity set and data criticality set to generate a new classification result for the data object identifier set.

[0115] The control strategy is updated by obtaining the new classification result of the data object identifier set, determining the control strategy set corresponding to the new classification result according to the preset data classification and control strategy comparison table, comparing the original control strategy with the new control strategy using a difference analysis algorithm, and generating a control strategy update instruction set.

[0116] When adjusting access control rules, the system obtains a set of management policy update instructions, parses the instruction set through the rule engine, and triggers the access control rule adjustment process when the instruction type is access control rule adjustment. This process calls the access control rule generation module to generate a new set of data object identifier access control rules.

[0117] Verify the new access control rules. Based on the new data object identifier set access control rules, use a machine learning algorithm based on the Transformer model to verify the new access control rules and determine the degree of impact of the new access control rules on data security risk factors. If the degree of impact is lower than the preset threshold, the new access control rules are published to the access control system.

[0118] Data access behavior monitoring and dynamic adjustment: Based on the data object identifier set access logs fed back by the access control system, clustering machine learning algorithm combined with K-nearest neighbor machine learning algorithm is used to identify data access behavior patterns, determine the degree of deviation between the data access behavior pattern and the preset behavior pattern, and when the deviation exceeds the preset threshold, a data hierarchical dynamic adjustment process is carried out, and data operation logs are re-collected.

[0119] Specifically, by establishing a dynamic adjustment and update mechanism for data grading, real-time monitoring and sensitivity assessment of government data usage are achieved. This mechanism automatically triggers reassessment and adjustment of data grading based on usage frequency and sensitivity indicators, and promptly updates data management strategies and access control rules. This mechanism not only improves the flexibility and responsiveness of data management but also verifies the effectiveness of new rules through machine learning algorithms, ensuring effective control of data security risks. This achieves intelligent monitoring and dynamic adjustment of data access behavior, enhancing the overall security and compliance of government data.

[0120] Furthermore, after updating data management strategies and access control rules, it also includes: establishing an audit and traceability mechanism for data classification changes, recording the history and reasons for data classification adjustments, and meeting compliance requirements.

[0121] Furthermore, establish an audit and traceability mechanism for data classification changes, recording the history and reasons for data classification adjustments, specifically including:

[0122] According to the preset periodic data collection hierarchy directory, the hierarchy information of each data item in the data hierarchy directory is obtained to obtain the hierarchy set. It is determined whether the data hierarchy set has changed within the period. If it has changed, the data hierarchy change record process is triggered. The changed data item set is determined by data comparison.

[0123] Then, a text analysis algorithm is used to parse the business log associated with the changed data item to obtain the text describing the reason for the change. The text describing the reason for the change is then parsed using an information extraction algorithm to obtain the phrase describing the reason for the change. Based on the phrase describing the reason for the change, the tag for the reason for the change is determined.

[0124] The system uses preset rules to match each data item in the set of changed data items, obtains the operator information of the changed data items, obtains the operator set and the identity information of each operator, and uses an information encryption algorithm to encrypt the identity information of each operator to obtain encrypted identity information.

[0125] Data change records are generated based on the set of changed data items, change reason tags, and encrypted identity information. These records are then stored on the blockchain to obtain a unique identifier for each data change record.

[0126] Specifically, by establishing an audit and traceability mechanism for data classification changes, the historical records and reasons for data classification adjustments can be recorded and tracked in detail, ensuring transparency and accountability in data management. Specifically, this mechanism automatically detects changes in data classification by periodically collecting and comparing data classification directories, and uses text analysis and information extraction technologies to parse the reasons for changes from business logs. Simultaneously, information about personnel involved in data changes is encrypted to ensure privacy and security. All these change records are securely stored on the blockchain, ensuring the immutability and traceability of data change records, meeting compliance requirements, and enhancing the security and credibility of data management.

[0127] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for automatically classifying and grading government data, characterized in that: Includes the following steps: By defining data quality rules and constraints, statistical methods and machine learning algorithms are used to automatically correct and standardize data. Based on automatically corrected and standardized data, a government domain ontology and knowledge graph are constructed to perform semantic enhancement and correlation analysis on government data. By constructing a domain concept hierarchy and relational network model, ontology reasoning and knowledge fusion technologies are used to classify data and make decisions. At the same time, a multi-classifier ensemble learning algorithm is used to integrate the classification results of different models. Based on the characteristics and application needs of government data, a government domain ontology and knowledge graph are constructed, defining the core concepts, attributes and relationships of government data, and using the knowledge graph to express the relationships between data to form a conceptual hierarchy and relationship network model. For the automatically corrected and standardized government data, natural language processing technology is used to perform semantic analysis on policy documents and announcements. Entity recognition technology is used to extract key policy measures, scope of impact and expected goals from policy documents, and semantic enhancement of the data is performed. Combined with ontology reasoning and knowledge graphs, cross-dataset association analysis is conducted. Based on domain ontology and knowledge graph, we use ontology reasoning rules and knowledge fusion algorithms to perform correlation analysis on government data. Based on the results of semantic enhancement and association analysis, a data classification model is trained using machine learning algorithms to automatically classify and support government data for decision-making. In the process of data classification and decision-making, ontology reasoning technology is used to infer knowledge and derive implicit classification rules and decision-making basis based on the concept hierarchy and relationship network in the domain ontology. By using knowledge fusion technology, the classification rules and decision criteria obtained from ontology reasoning are fused with the prediction results of machine learning models, and semantic information and data features are integrated to obtain the final classification and decision results. Based on the classification results, a unified standard and specification for the hierarchical classification of government data is formulated. Based on the confidentiality dimension of the data, a data sensitivity assessment model and hierarchical decision-making mechanism are established to determine the data level. Through consultation and review by multiple stakeholders, a consensus-based data hierarchical classification scheme is formed. After reaching a consensus on a data classification and grading scheme, a knowledge base and rule base for the government domain are constructed to determine knowledge reasoning and rule matching, and data is mapped to concepts and entities in the domain knowledge base to perform semantic association and contextual understanding of the data. Establish a dynamic adjustment and update mechanism for data classification. By monitoring data usage in real time, dynamically assess the sensitivity and importance of data, and automatically trigger the reassessment and adjustment process of data classification based on preset adjustment rules and thresholds, update data management strategies and access control rules.

2. The method for automatic classification and grading of government data according to claim 1, characterized in that: Before defining data quality rules and constraints, the process also includes: collecting data from different sources and types, converting it into a standardized structured format, and extracting key information using regular expressions based on the content and characteristics of the data, and mapping it to a predefined data model to form a unified data representation.

3. The method for automatic classification and grading of government data according to claim 2, characterized in that: Key information is extracted using regular expressions and mapped to a predefined data model to form a unified data representation, specifically including: Acquire multi-source heterogeneous data, employ appropriate data cleaning and preprocessing methods to convert it into a standardized structured format, and determine the key information types and features to be extracted based on a predefined data model; For standardized structured data, regular expression matching methods are used to extract key information that conforms to specific characteristics; The extracted key information is mapped to the corresponding fields and attributes in the data model according to predefined mapping rules, and the mapped data is then validated and its integrity is checked. The mapped data is stored in a unified format and structure to form a standardized data representation.

4. The method for automatic classification and grading of government data according to claim 1, characterized in that: Automatically correcting and standardizing data using statistical methods and machine learning algorithms, specifically including: Based on predefined data quality rules and constraints, a quality check is performed on the unified data to identify abnormal data that does not meet the quality requirements. Statistical methods are used to analyze the abnormal data, calculate the statistical characteristics of the abnormal data, and determine the distribution of the abnormal data. By using anomaly detection algorithms in machine learning algorithms, abnormal data is automatically detected and identified, the location and range of abnormal data are determined, and then a data imputation algorithm is used to automatically estimate and fill in missing or abnormal data values. According to the preset data standardization rules, the filled data is automatically standardized to convert the data into a unified format, unit and representation. Then the standardized data is quality verified and judged to see if the data meets the quality requirements according to the data quality rules and constraints. The processed, high-quality data is stored in the database.

5. The method for automatic classification and grading of government data according to claim 1, characterized in that: Based on the confidentiality dimension of the data, a data sensitivity assessment model and a hierarchical decision-making mechanism are established to determine the data sensitivity levels, specifically including: The government data set is acquired, and the metadata information of the data set is parsed using natural language processing algorithms according to the pre-established rule base to obtain the data type and data source. It is then determined whether the data type belongs to a sensitive type. If it belongs to a sensitive type, the data set is classified into the dataset to be evaluated. The government data set includes population data, economic data, social management data, urban management data, government service data, public resource data, and social security data. Data content is obtained from the dataset to be evaluated. Based on a pre-established keyword list, sensitive words are extracted from the data content using a text matching algorithm. The number and types of sensitive words are obtained, the weights of the sensitive words are determined, and the sensitivity score is calculated. Based on the sensitivity score and a preset threshold range, the data security level is determined.

6. The method for automatic classification and grading of government data according to claim 5, characterized in that: When the data security level exceeds the preset threshold, the dataset is de-identified to obtain the de-identified dataset. De-identification methods include data masking and data replacement. User access records are obtained from the anonymized dataset. Based on the user roles and permission matrix, the range of data accessed by the user is obtained, and the user's access permissions to a specific dataset are determined. When a user attempts to access data beyond their authority, a violation log is recorded. Based on the data security level and sharing protocol, the sharing scope of the dataset is determined. When the data security level is high, the sharing scope is limited to a specific department. Data is transmitted through an interface to obtain the shared dataset. Data update information is obtained from the shared dataset. The data update frequency is determined according to a pre-established schedule. When the data update frequency is high, the data synchronization mechanism is triggered to obtain the latest data version. Based on the data storage period and destruction strategy, the data retention time is determined. When the data reaches the storage period, the data destruction process is triggered to delete the data from the storage medium and obtain the destroyed data record.

7. The method for automatic classification and grading of government data according to claim 1, characterized in that: By constructing a knowledge base and rule base for the government domain, knowledge reasoning and rule matching are determined, data is mapped to concepts and entities in the domain knowledge base, and semantic association and contextual understanding of the data are performed. Specifically, this includes: By comparing structured data with data from the concept hierarchy and relational network, if the matching degree is higher than a preset threshold, the concept to which the data belongs is determined; if it is lower than the preset threshold, the structured data is compared with the entity layer to obtain the entity with the highest matching degree, thus obtaining the set of entities corresponding to the data. By calling association rules in the rule base from the entity set, the relationships between entities are obtained. Based on the relationships between entities, an entity relationship graph is constructed. By traversing the graph, the data context information is determined, and the final semantic understanding result is obtained.

8. The method for automatic classification and grading of government data according to claim 1, characterized in that: Establish a data hierarchical dynamic adjustment and update mechanism, specifically including: Data collection and recording: Collect data operation logs, obtain operation user, operation type, operation time, and operation data object identification information through log parsing technology, and record them in the data usage record database; Using frequency statistics and sensitivity assessment, the database information of data usage records is used to count the frequency of data object identifiers being used within a preset time range. When the frequency exceeds a preset frequency threshold, the data sensitivity and criticality assessment process is triggered. A pre-established assessment model is used to calculate the data object identifier set to obtain the data sensitivity set and the data criticality set. When data classification adjustment is triggered, the data sensitivity set and data criticality set are obtained. When an indicator in the set exceeds the preset indicator threshold, the data classification adjustment process is triggered, the data classification rule base is called, and the corresponding classification rules are matched according to the data sensitivity set and data criticality set to generate a new classification result for the data object identifier set. The control strategy is updated by obtaining the new classification result of the data object identifier set, determining the control strategy set corresponding to the new classification result according to the preset data classification and control strategy comparison table, comparing the original control strategy with the new control strategy using a difference analysis algorithm, and generating a control strategy update instruction set. When adjusting access control rules, the system obtains a set of management policy update instructions, parses the instruction set through the rule engine, and triggers the access control rule adjustment process when the instruction type is access control rule adjustment. This process calls the access control rule generation module to generate a new set of data object identifier access control rules. Verify the new access control rules. Based on the new data object identifier set access control rules, use a machine learning algorithm based on the Transformer model to verify the new access control rules and determine the degree of impact of the new access control rules on data security risk factors. If the degree of impact is lower than the preset threshold, the new access control rules are published to the access control system. Data access behavior monitoring and dynamic adjustment: Based on the data object identifier set access logs fed back by the access control system, clustering machine learning algorithm combined with K-nearest neighbor machine learning algorithm is used to identify data access behavior patterns, determine the degree of deviation between the data access behavior pattern and the preset behavior pattern, and when the deviation exceeds the preset threshold, a data hierarchical dynamic adjustment process is carried out, and data operation logs are re-collected.

9. The method for automatic classification and grading of government data according to claim 1, characterized in that: After updating data management strategies and access control rules, the following also applies: establishing an audit and traceability mechanism for data classification changes, and recording the history and reasons for data classification adjustments.

10. The method for automatic classification and grading of government data according to claim 9, characterized in that: Establish an audit and traceability mechanism for data classification changes, recording the history and reasons for data classification adjustments, specifically including: According to the preset periodic data collection hierarchy directory, the hierarchy information of each data item in the data hierarchy directory is obtained to obtain the hierarchy set. It is determined whether the data hierarchy set has changed within the period. If it has changed, the data hierarchy change record process is triggered. The changed data item set is determined by data comparison. Then, a text analysis algorithm is used to parse the business log associated with the changed data item to obtain the text describing the reason for the change. The text describing the reason for the change is then parsed using an information extraction algorithm to obtain the phrase describing the reason for the change. Based on the phrase describing the reason for the change, the tag for the reason for the change is determined. The system uses preset rules to match each data item in the set of changed data items, obtains the operator information of the changed data items, obtains the operator set and the identity information of each operator, and uses an information encryption algorithm to encrypt the identity information of each operator to obtain encrypted identity information. Data change records are generated based on the set of changed data items, change reason tags, and encrypted identity information. These records are then stored on the blockchain to obtain a unique identifier for each data change record.

Citation Information

Patent Citations

  • Method and system for improving data security of temporary office local area network

    CN118157996A

  • DCS intelligent decision-making method and system fusing large language model and knowledge graph

    CN118820778A