Government affair data automatic classification and grading method
By building the ontology and knowledge graph of the government field, combining natural language processing and machine learning algorithms, the problems of diversified government data formats and blurred category boundaries are solved, and efficient and accurate classification and management of government data are achieved, ensuring the security and compliance of data management.
Patent Information
- Application Number
- CN202510451565.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-11
AI Technical Summary
During the collection and sorting of government data, there are problems such as diversified data formats and blurred category boundaries, which leads to uncertainty in classification decisions and is difficult to achieve efficient and accurate automatic classification and management.
By building an ontology and knowledge graph in the government field, combining natural language processing technology and machine learning algorithms, semantic enhancement and correlation analysis of government data, a multi-classifier integrated learning algorithm is used to integrate classification results, and a dynamic adjustment and update mechanism for data hierarchy is established to ensure the flexibility and security of data management.
It realizes efficient and accurate classification and management of government data, improves data consistency and reliability, ensures the security and compliance of data management, and provides a transparent responsibility tracking mechanism.
Smart Images

Figure CN120448890A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data management and information processing technology, and in particular to a method for automatically classifying and grading government data. Background Art
[0002] The collection and organization of government data presents a diverse array of data formats. Documents, reports, forms, and other types of information vary widely in how they are organized and presented, making unified data processing difficult. The automatic classification of government data is often fuzzy due to the correlation and overlap between categories like policy, law, finance, and administration, leading to fuzzy category boundaries and uncertainty in classification decisions. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for automatic classification and grading of government data. By standardizing data formats, constructing government domain ontologies and knowledge graphs, and implementing dynamic data sensitivity assessments and update management strategies, the problems of data format diversity and classification uncertainty are solved, thereby improving the efficiency and accuracy of government data management.
[0004] The purpose of the present invention can be achieved through the following technical solutions:
[0005] This application provides a method for automatically classifying and grading government data, including the following steps:
[0006] Automatically correct and standardize data using statistical methods and machine learning algorithms by defining data quality rules and constraints;
[0007] Based on automatically corrected and standardized data, we construct government domain ontologies and knowledge graphs, conduct semantic enhancement and association analysis on government data, build domain concept hierarchies and relationship network models, and utilize ontology reasoning and knowledge fusion technologies to perform data classification and decision-making. We also employ multi-classifier ensemble learning algorithms to integrate the classification results of different models.
[0008] Among them, according to the characteristics and application requirements of government data, a government domain ontology and knowledge graph are constructed to define the core concepts, attributes and relationships of government data. The knowledge graph is used to express the relationship between data to form a conceptual hierarchy and relationship network model.
[0009] For automatically corrected and standardized government data, natural language processing technology is used to perform semantic analysis on policy documents and announcements. Entity recognition technology is used to extract key policy measures, impact areas, and expected goals from policy documents. Semantic enhancement of the data is then performed, and cross-dataset correlation analysis is conducted by combining ontology reasoning and knowledge graphs.
[0010] Based on domain ontology and knowledge graph, the association analysis of government data is carried out using ontology reasoning rules and knowledge fusion algorithms;
[0011] Based on the results of semantic enhancement and association analysis, machine learning algorithms are used to train data classification models to automatically classify and support decision-making for government data;
[0012] In the process of data classification and decision-making, ontology reasoning technology is used to perform knowledge deduction, and implicit classification rules and decision-making basis are derived based on the concept hierarchy and relationship network in the domain ontology;
[0013] Through knowledge fusion technology, the classification rules and decision-making basis obtained from ontology reasoning are integrated with the prediction results of the machine learning model, and semantic information and data features are integrated to obtain the final classification and decision results;
[0014] Based on the classification results, formulate unified standards and specifications for the grading and classification of government data. Based on the confidentiality dimension of the data, establish a data sensitivity assessment model and a grading decision-making mechanism to determine the level of data. Through consultation and review by multiple stakeholders, form a consensus data grading and classification plan.
[0015] After forming a consensus on the data classification scheme, we build a knowledge base and rule base for the government affairs field, determine knowledge reasoning and rule matching, map the data to concepts and entities in the domain knowledge base, and conduct semantic association and contextual understanding of the data.
[0016] Establish a dynamic adjustment and update mechanism for data classification, monitor data usage in real time, dynamically evaluate the sensitivity and importance of data, and automatically trigger the re-evaluation and adjustment process of data classification based on preset adjustment rules and thresholds, and update data management policies and access control rules.
[0017] Before defining data quality rules and constraints, it also includes: collecting data from different sources and types and converting them into a standardized structured format, and using regular expressions to extract key information based on the content and characteristics of the data, and mapping it to a predefined data model to form a unified data representation.
[0018] Furthermore, a regular expression method is used to extract key information and map it to a predefined data model to form a unified data representation, including:
[0019] Acquire multi-source heterogeneous data, convert it into a standardized structured format using appropriate data cleaning and preprocessing methods, and determine the key information types and features to be extracted based on predefined data models;
[0020] For standardized structured data, regular expression matching method is used to extract key information that meets specific characteristics;
[0021] Map the extracted key information to the corresponding fields and attributes in the data model according to predefined mapping rules, and perform verification and integrity checks on the mapped data;
[0022] The mapped data is stored in a unified format and structure to form a standardized data representation.
[0023] Furthermore, statistical methods and machine learning algorithms are used to automatically correct and standardize data, including:
[0024] According to pre-defined data quality rules and constraints, the unified data is quality checked to identify abnormal data that does not meet the quality requirements. The abnormal data is analyzed using statistical methods to calculate the statistical characteristics of the abnormal data and determine the distribution of the abnormal data.
[0025] Through the anomaly detection algorithm in the machine learning algorithm, abnormal data is automatically detected and identified to determine the location and range of abnormal data. Then, the data filling algorithm is used to automatically estimate and fill in missing or abnormal data values.
[0026] According to the preset data standardization rules, the filled data is automatically standardized and converted into a unified format, unit and representation. The standardized data is then quality verified to determine whether the data meets the quality requirements based on the data quality rules and constraints.
[0027] The processed high-quality data is stored in the database.
[0028] Furthermore, based on the confidentiality dimension of the data, a data sensitivity assessment model and hierarchical decision-making mechanism are established to determine the level of the data, including:
[0029] Obtain a government data set, and parse the data set metadata using a natural language processing algorithm based on a pre-established rule base to obtain the data type and data source, and determine whether the data type is sensitive. If it is sensitive, the data set is classified as a dataset to be evaluated.
[0030] Obtain data content from the dataset to be evaluated, extract sensitive words from the data content through a text matching algorithm based on a pre-established keyword table, obtain the number and types of sensitive words, determine the weights of sensitive words, and calculate the sensitivity score; determine the data confidentiality level based on the sensitivity score and the preset threshold range.
[0031] Furthermore, when the data confidentiality level exceeds a preset threshold, the data set is desensitized to obtain a desensitized data set. The desensitization methods include data masking and data replacement.
[0032] Obtain user access records from the desensitized dataset. Based on the user role and permission matrix, determine the data range that the user can access and determine the user's access rights to specific datasets. If a user attempts to access beyond their permission, a violation log is recorded.
[0033] Determine the sharing scope of the data set based on the data confidentiality level and sharing agreement. If the data confidentiality level is high, the sharing scope is limited to specific departments. Data is transmitted through the interface to obtain a shared data set.
[0034] Obtain data update information from shared datasets and determine the data update frequency based on a pre-established schedule. When the data update frequency is high, the data synchronization mechanism is triggered to obtain the latest data version.
[0035] The data retention time is determined based on the data storage period and destruction policy. When the data reaches the storage period, the data destruction process is triggered, the data is deleted from the storage medium, and a record of the destroyed data is obtained.
[0036] Furthermore, by building a knowledge base and rule base for the government affairs field, determining knowledge reasoning and rule matching, mapping data to concepts and entities in the domain knowledge base, and performing semantic association and contextual understanding of data, specifically including:
[0037] By comparing structured data with concept hierarchies and relational networks, if the matching degree is higher than a preset threshold, the concept to which the data belongs is determined. If it is lower than the preset threshold, the structured data is compared with the entity layer to obtain the entity with the highest matching degree, and the entity set corresponding to the data is obtained.
[0038] Through the entity set, the association rules in the rule library are called to obtain the association relationship between entities. Based on the association relationship between entities, an entity relationship graph is constructed. By traversing the graph, the data context information is determined to obtain the final semantic understanding result.
[0039] Furthermore, a data classification dynamic adjustment and update mechanism is established, specifically including:
[0040] Data collection and recording: Collect data operation logs, obtain the operation user, operation type, operation time, and operation data object identification information through log analysis technology, and record them in the data usage record database;
[0041] Frequency of use statistics and sensitivity assessment: By recording the database information of data usage, the frequency of use of data object identifiers within a preset time range is counted. When the frequency exceeds the preset frequency threshold, the data sensitivity and criticality assessment process is triggered. The pre-established assessment model is used to calculate the data object identifier set to obtain the data sensitivity set and data criticality set;
[0042] The data classification adjustment trigger obtains the data sensitivity set and the data criticality set. When an indicator in the set exceeds the preset indicator threshold, the data classification adjustment process is triggered, the data classification rule library is called, and the corresponding classification rules are matched according to the data sensitivity set and the data criticality set to generate a new classification result for the data object identification set;
[0043] Update the control strategy, obtain the new classification result of the data object identifier set, determine the control strategy set corresponding to the new classification result based on the preset data classification and control strategy comparison table, use the difference analysis algorithm to compare the original control strategy with the new control strategy, and generate a control strategy update instruction set;
[0044] Access control rule adjustment: obtain the control policy update instruction set, parse the instruction set through the rule engine, and when the instruction type is access control rule adjustment, trigger the access control rule adjustment process, call the access control rule generation module, and generate a new data object identifier set access control rule;
[0045] Verify the new access control rules. Based on the new data object identifier set access control rules, use the machine learning algorithm based on the Transformer model to verify the new access control rules and determine the impact of the new access control rules on data security risk factors. If the impact is lower than the preset threshold, the new access control rules will be published to the access control system.
[0046] Data access behavior monitoring and dynamic adjustment: Based on the access log of the data object identifier set fed back by the access control system, a clustering machine learning algorithm combined with a K-nearest neighbor machine learning algorithm is used to identify data access behavior patterns, and the degree of deviation between the data access behavior pattern and the preset behavior pattern is determined. When the degree of deviation exceeds the preset threshold, a data classification dynamic adjustment process is performed, and data operation logs are collected again.
[0047] Furthermore, after updating the data management strategy and access control rules, it also includes: establishing an audit and traceability mechanism for data classification changes to record the history and reasons for data classification adjustments.
[0048] Furthermore, an audit and traceability mechanism for data classification changes should be established to record the history and reasons for data classification adjustments, including:
[0049] Collect data hierarchical catalogs according to preset periods, obtain hierarchical information of each data item in the data hierarchical catalog, obtain hierarchical sets, determine whether the data hierarchical sets have changed within the period, and if so, trigger the data hierarchical change record process, and determine the changed data item set through data comparison;
[0050] Then, a text analysis algorithm is used to parse the business log associated with the changed data item to obtain a description of the change reason. The information extraction algorithm is used to parse the description of the change reason to obtain a change reason phrase. Based on the change reason phrase, a change reason label is determined.
[0051] Using preset rules to match each data item in the changed data item set, obtain the operator information of the changed data item, obtain the operator set and the identity information of each operator, and use the information encryption algorithm to encrypt the identity information of each operator to obtain encrypted identity information;
[0052] Generate a data change record based on the set of changed data items, the change reason label, and the encrypted identity information, store the data change record through the blockchain, and obtain a unique identifier for the data change record.
[0053] The beneficial effects of the present invention are:
[0054] By standardizing the processing of government data, extracting key information, and automatically classifying and correcting it, the problems of data format diversity and blurred category boundaries in background technologies are resolved. First, regular expressions and pattern matching techniques are used to convert data from different sources and formats into a structured format. Machine learning algorithms, data quality detection, anomaly correction, and other means are used to improve the consistency and accuracy of the data. During the semantic enhancement and association analysis phase of the data, the government domain ontology and knowledge graph are constructed, combined with natural language processing technology, to effectively semantically classify government data, resolving overlap and uncertainty between categories. The ensemble learning algorithm is used to further improve the robustness and accuracy of classification decisions, ensuring the efficiency and reliability of classification.
[0055] By establishing a dynamic adjustment and update mechanism for data classification, combined with real-time data monitoring, sensitivity assessment and automatic adjustment of access control rules, the flexibility and security of data management are ensured. By continuously evaluating the sensitivity and usage of data, re-evaluation of data classification is triggered, and corresponding management policies and access controls are updated to ensure the security and compliance of government data. At the same time, the audit and traceability mechanism for data classification changes uses blockchain technology to store change records in an unalterable manner, providing transparency and accountability tracking for data management, thereby effectively improving the efficiency, security and intelligence of government data management. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] For better understanding and implementation, the technical solution of the present application is described in detail below with reference to the accompanying drawings.
[0057] Figure 1 A flowchart of a method for automatic classification and grading of government data provided in this application;
[0058] Figure 2 A schematic diagram of the process of forming a unified data representation for the automatic classification and grading method of government data provided in this application;
[0059] Figure 3 A flowchart of an automatic classification and grading method for government data provided in this application for automatically correcting and standardizing data. DETAILED DESCRIPTION
[0060] To further illustrate the technical means and effectiveness of the present invention in achieving its intended purpose, exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present application. Rather, they are merely examples of methods and systems consistent with certain aspects of the present application, as detailed in the appended claims.
[0061] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0062] The following describes in detail the specific implementation methods, features and effects of the present invention in conjunction with the accompanying drawings and preferred embodiments.
[0063] See also Figure 1-Figure 3 This embodiment provides a method for automatically classifying and grading government data, including the following steps:
[0064] S1. Collect data from different sources and types and convert them into a standardized structured format. Based on the content and characteristics of the data, use regular expressions to extract key information and map it to a predefined data model to form a unified data representation.
[0065] Furthermore, a regular expression method is used to extract key information and map it to a predefined data model to form a unified data representation, including:
[0066] S11. Acquire multi-source heterogeneous data. Apply appropriate data cleaning and preprocessing methods to data from different sources and formats, converting them into standardized structured formats. Based on predefined data models, determine the key information types and features to be extracted.
[0067] S12. For standardized structured data, use pattern matching methods such as regular expressions to extract key information that meets specific characteristics;
[0068] S13. Map the extracted key information to the corresponding fields and attributes in the data model according to predefined mapping rules, and perform verification and integrity checks on the mapped data to ensure data accuracy and consistency;
[0069] S14. Store the mapped data in a unified format and structure to form a standardized data representation. Based on the standardized data representation, machine learning algorithms such as clustering, classification, and association rule mining are used to discover patterns and regularities in the data to support business decision-making and optimization.
[0070] Specifically, by standardizing and extracting key information from multi-source, heterogeneous data, and using techniques like regular expressions to convert the data into a unified structured format, data consistency and accuracy are ensured. Furthermore, key information is mapped according to predefined data models, and data quality is ensured through validation and integrity checks. Ultimately, data stored in a unified format supports further analysis based on machine learning algorithms, such as pattern recognition and pattern mining, providing intelligent support for business decision-making, improving data utilization efficiency and value, and promoting business optimization and intelligent decision-making.
[0071] S2. Based on unified data, by defining data quality rules and constraints, and using statistical methods and machine learning algorithms, such as anomaly detection and data filling, to automatically correct and standardize data to improve data accuracy and usability;
[0072] Furthermore, statistical methods and machine learning algorithms, such as anomaly detection and data filling, are used to automatically correct and standardize data, including:
[0073] S21. Based on predefined data quality rules and constraints, perform quality checks on unified data, identify abnormal data that does not meet quality requirements, analyze the abnormal data using statistical methods, calculate the statistical characteristics of the abnormal data, such as mean, variance, quantiles, etc., and determine the distribution of the abnormal data. Quality rules and constraints include requirements for data format, data type, data range, data integrity, data consistency, etc., and define the standards that must be followed during data collection, storage, processing, and analysis to ensure that the data meets specific business needs and analytical purposes. Quality rules and constraints can be used to identify and correct errors, omissions, and inconsistencies in the data, thereby improving the overall quality of the data and making it more suitable for subsequent analysis and decision-making processes.
[0074] S22. Automatically detect and identify abnormal data using anomaly detection algorithms within machine learning algorithms, such as Isolation Forest and Local Anomaly Factor, to determine the location and range of abnormal data. Then, use data filling algorithms, such as mean filling and nearest neighbor filling, to automatically estimate and fill missing or abnormal data values to ensure data integrity.
[0075] S23. Automatically standardize the populated data according to pre-set data standardization rules, converting the data into a unified format, unit, and representation to improve data consistency. Quality verification is then performed on the standardized data to determine whether the data meets quality requirements based on data quality rules and constraints, ensuring the accuracy and usability of the processed data.
[0076] S24. The processed high-quality data is stored in the database for subsequent data analysis and application. The data processing process is completed, and the data quality and availability are significantly improved.
[0077] Specifically, by defining data quality rules and constraints, combined with statistical methods and machine learning algorithms, we have achieved a comprehensive improvement in data quality. Anomaly detection algorithms (such as isolation forests and local anomaly factors) accurately identify anomalous data, and data filling algorithms (such as mean filling and nearest neighbor filling) effectively repair missing or outliers to ensure data integrity. On this basis, unified data standardization improves consistency, while quality verification further ensures data accuracy and availability. Ultimately, the processed, high-quality data storage provides reliable support for subsequent analysis and application, significantly improving the efficiency of data management and utilization.
[0078] S3. Based on the automatically corrected and standardized data, construct the government domain ontology and knowledge graph, perform semantic enhancement and association analysis on government data, and conduct data classification and decision-making by building a domain concept hierarchy and relationship network model, using ontology reasoning and knowledge fusion technology. At the same time, use a multi-classifier ensemble learning algorithm to integrate the classification results of different models to improve the robustness and accuracy of classification; solve the problems of fuzzy categories and classification uncertainty of government data.
[0079] Furthermore, based on the characteristics and application requirements of government data, we construct a government domain ontology and knowledge graph, define the core concepts, attributes, and relationships of government data, and use the knowledge graph to express the associations between data to form a conceptual hierarchy and relationship network model.
[0080] For automatically corrected and standardized government data, natural language processing technology is used to perform semantic analysis on policy documents and announcements. Entity recognition technology is used to extract key policy measures, impact areas, and expected goals from a large number of policy documents, such as the relationship between "policy" entities and "industry" entities, to perform semantic enhancement of the data. This helps to more effectively identify key entities, attributes, and relationships in automatically corrected and standardized government data, thereby enhancing the data semantics. Combined with ontology reasoning and knowledge graphs, cross-dataset association analysis can be performed.
[0081] Based on domain ontologies and knowledge graphs, we use ontology reasoning rules and knowledge fusion algorithms to conduct correlation analysis on government data, discover implicit semantic relationships and patterns, and enrich the semantic representation of data. For example, by analyzing data in the "Government Overview" module, we discovered the potential connection between government service efficiency and public satisfaction. The government overview includes overviews of urban management, emergency response, human resources and social security, education, medical care, housing and construction, and civil affairs.
[0082] Based on the results of semantic enhancement and association analysis, machine learning algorithms such as decision trees and support vector machines are used to train data classification models to automatically classify and provide decision support for the data in the "Government Overview" module of government data.
[0083] In the data classification and decision-making process, ontology reasoning technology is used to perform knowledge deduction. Based on the concept hierarchy and relationship network in the domain ontology, implicit classification rules and decision-making basis are derived to enhance the interpretability of classification and decision-making. Specifically, by analyzing the data in the "Ecological Civilization Overview" module, a balance strategy between environmental governance and economic development is derived. The Ecological Civilization Overview includes a display of meteorological microservices, natural resources, and ecological environment overviews.
[0084] Through knowledge fusion technology, the classification rules and decision-making basis obtained by ontology reasoning are fused with the prediction results of the machine learning model, and semantic information and data features are integrated to obtain the final classification and decision-making results, thereby improving the overall accuracy and reliability.
[0085] Imagine building a smart city management system that needs to process and analyze large amounts of government data, including traffic flow, public safety incidents, and environmental monitoring data. First, by automatically correcting and standardizing the data, an ontology and knowledge graph are constructed that encompasses core concepts, attributes, and relationships in urban management. These entities include "traffic flow," "safety incidents," and "environmental indicators," as well as the relationships between them, such as how "traffic flow" affects "public safety." Natural language processing techniques are then used to perform semantic analysis on policy documents and announcements. Entity recognition techniques are then used to extract key policy measures, their impact, and their intended goals. For example, the relationship between "traffic control policy" and "traffic flow" can be identified, enhancing the data's semantics. Then, based on the domain ontology and knowledge graph, the ontology reasoning rules and knowledge fusion algorithms are used to conduct association analysis on government data, and potential connections between "increased traffic flow" and "public safety incidents" are discovered. Based on the results of these semantic enhancements and association analyses, machine learning algorithms, such as decision trees and support vector machines, are used to train data classification models to automatically classify and provide decision support for the data in the "Government Overview" module. Through data processed by ensemble learning and multi-classifier algorithms, the system can automatically classify urban traffic flow data into "peak hours" and "off-peak hours", while identifying "traffic congestion" events related to high traffic flow and predicting possible "public safety risks", such as areas with a high incidence of traffic accidents. It can also identify areas where the air quality index exceeds the standard based on environmental monitoring data and classify them as environmental problems that "require urgent intervention".
[0086] In the process of data classification and decision-making, ontology reasoning technology is used to perform knowledge deduction and deduce the balance strategy between environmental governance and economic development. For example, by analyzing the data in the "Ecological Civilization Overview" module, it is deduced how to reduce environmental pollution while maintaining economic growth. Finally, through knowledge fusion technology, the classification rules and decision-making basis obtained from ontology reasoning are integrated with the prediction results of the machine learning model, and semantic information and data features are integrated to obtain the final classification and decision-making results, thereby improving the overall accuracy and reliability and ensuring the scientific nature and effectiveness of urban management decisions.
[0087] Specifically, by constructing a government domain ontology and knowledge graph, combined with natural language processing technology and machine learning algorithms, it is possible to effectively perform semantic enhancement and precise classification of government data, solving the problems of diverse data formats and blurred category boundaries. Specifically, government data is first automatically corrected and standardized, then natural language processing technology is used to extract key entities and attributes, followed by ontology reasoning and knowledge fusion technology to discover implicit semantic relationships between data, and finally, an ensemble learning algorithm is used to integrate the results of multiple classification models to improve the robustness and accuracy of classification. This process not only enhances the semantic representation of the data, but also improves the interpretability of classification decisions, thereby solving the problems of uncertainty in government data classification and category ambiguity mentioned in the background technology.
[0088] S4. Based on the classification results, formulate unified standards and specifications for the grading and classification of government data. Based on the confidentiality dimension of the data, establish a data sensitivity assessment model and a grading decision-making mechanism to determine the level of data. Through consultation and review by multiple stakeholders, form a consensus data grading and classification scheme and embed it into the data lifecycle management process.
[0089] Furthermore, based on the confidentiality dimension of the data, a data sensitivity assessment model and hierarchical decision-making mechanism are established to determine the level of the data, including:
[0090] Obtain government data sets and, based on a pre-established rule base, use natural language processing algorithms to parse the data set metadata to obtain the data type and source. Determine whether the data type is sensitive. If so, classify the data set as a dataset to be evaluated. Government data sets include: population data, demographic information, including registered population, floating population, and population structure (age, gender, education level, etc.);
[0091] Economic data, including GDP, fiscal revenue and expenditure, tax data, industrial and agricultural output, commercial and service industry data, investment and consumption data;
[0092] Social management data, public safety data (such as crime rates and traffic accident statistics), public health data (such as disease control and medical services), and education data (such as the number of schools, student numbers, and educational outcomes);
[0093] Urban management data, urban planning, construction, traffic flow, municipal facilities management, environmental protection and pollution control;
[0094] Government service data, including information on the acceptance, processing, and feedback of government service matters, including administrative approval, public services, and social affairs management;
[0095] Public resource data, including land resources, water resources, and energy consumption and distribution;
[0096] Social security data, including data on social insurance, social assistance, social welfare, etc.
[0097] Obtain data content from the dataset to be evaluated, extract sensitive words from the data content through a text matching algorithm based on a pre-established keyword table, obtain the number and types of sensitive words, determine the weights of sensitive words, and calculate the sensitivity score; determine the data confidentiality level based on the sensitivity score and the preset threshold range.
[0098] Furthermore, when the data confidentiality level exceeds a preset threshold, the data set is desensitized to obtain a desensitized data set. The desensitization methods include data masking and data replacement.
[0099] Obtain user access records from the desensitized dataset. Based on the user role and permission matrix, determine the data range that the user can access and determine the user's access rights to specific datasets. If a user attempts to access beyond their permission, a violation log is recorded.
[0100] Determine the sharing scope of the data set based on the data confidentiality level and sharing agreement. If the data confidentiality level is high, the sharing scope is limited to specific departments. Data is transmitted through the interface to obtain a shared data set.
[0101] Obtain data update information from shared datasets and determine the data update frequency based on a pre-established schedule. When the data update frequency is high, the data synchronization mechanism is triggered to obtain the latest data version.
[0102] The data retention time is determined based on the data storage period and destruction policy. When the data reaches the storage period, the data destruction process is triggered, the data is deleted from the storage medium, and a record of the destroyed data is obtained.
[0103] Specifically, it can identify and process sensitive information in government data. Through natural language processing and a preset rule base, the process automatically classifies data sets and assesses their sensitivity, determines the confidentiality level of the data, and performs desensitization based on the sensitivity score. It also manages user access rights to ensure the compliance of data access and controls the scope of data sharing based on the confidentiality level and sharing agreement of the data. It also includes monitoring the frequency of data updates and enforcing data storage period policies, thereby triggering the destruction process when the data reaches the retention period to ensure data security and compliance. Overall, this process improves the automation and accuracy of government data management, reduces the risk of data leakage, and optimizes data lifecycle management.
[0104] S5. After reaching a consensus on the data classification scheme, we will build a knowledge base and rule base for the government affairs domain, determine knowledge reasoning and rule matching, map the data to concepts and entities in the domain knowledge base, and conduct semantic association and contextual understanding of the data.
[0105] Furthermore, by building a knowledge base and rule base for the government affairs field, determining knowledge reasoning and rule matching, mapping data to concepts and entities in the domain knowledge base, and performing semantic association and contextual understanding of data, specifically including:
[0106] By comparing structured data with conceptual hierarchies and relationship networks, when the matching degree is higher than the preset threshold, the concept to which the data belongs is determined; when it is lower than the preset threshold, the structured data is compared with the entity layer to obtain the entity with the highest matching degree, and the entity set corresponding to the data is obtained; for example, the "GDP" data is matched with the "GDP" concept in the knowledge base. For data with a matching degree lower than the preset threshold, it is further compared with the entity layer to determine the specific entity corresponding to the data, such as "a certain company name" and the "company" entity in the knowledge base.
[0107] Through the entity set, the association rules in the rule library are called to obtain the association relationship between entities, such as the relationship between "Enterprise A" and "Industry B". Based on the association relationship between entities, an entity relationship graph is constructed. By traversing the graph, the data context information is determined, such as the status and influence of "Enterprise A" in "Industry B", to obtain the final semantic understanding result.
[0108] For example, in the tax field, a knowledge base containing concepts and entities such as "taxpayer," "tax type," and "tax rate" is constructed. Upon receiving a company's tax return data, it is first matched against the "taxpayer" entity in the knowledge base. If the match exceeds a preset threshold, the data is determined to belong to the "taxpayer" entity. For data with a lower match, it is further compared against the entity layer, for example, matching "a company name" with the "company" entity in the knowledge base to find the corresponding entity set. Subsequently, by invoking association rules from the rule base, the tax relationship between "company A" and "industry B" is discovered. An entity relationship graph is constructed, and the graph is traversed to determine the position and influence of "company A" in "industry B." Ultimately, the semantic understanding of the tax return data is obtained, namely, the tax contribution and economic significance of "company A" in "industry B." This entire process not only improves the semantic association and contextual understanding of the data but also provides strong data support for tax decision-making.
[0109] Specifically, by building a knowledge base and rule base for the government sector, combined with knowledge reasoning and rule matching techniques, we achieve semantic association and contextual understanding of government data. Based on a data classification scheme, we use natural language processing technology to transform raw data into a structured form. By matching it with concepts and entities in the knowledge base, we accurately determine the semantic attribution of the data. We further use the association rules in the rule base to construct an entity relationship graph, revealing the associations and semantic context between data. Ultimately, this deep understanding of data semantics significantly enhances the intelligent processing capabilities of government data, providing strong support for optimizing government services and scientific decision-making.
[0110] S6. Establish a dynamic adjustment and update mechanism for data classification. By real-time monitoring of data usage, dynamically assess the sensitivity and importance of data, and automatically trigger the reassessment and adjustment process of data classification based on preset adjustment rules and thresholds, and update data management policies and access control rules.
[0111] Furthermore, a data classification dynamic adjustment and update mechanism is established, specifically including:
[0112] Data collection and recording: Collect data operation logs, obtain the operation user, operation type, operation time, and operation data object identification information through log analysis technology, and record them in the data usage record database;
[0113] Frequency of use statistics and sensitivity assessment: By recording the database information of data usage, the frequency of use of data object identifiers within a preset time range is counted. When the frequency exceeds the preset frequency threshold, the data sensitivity and criticality assessment process is triggered. The pre-established assessment model is used to calculate the data object identifier set to obtain the data sensitivity set and data criticality set;
[0114] The data classification adjustment trigger obtains the data sensitivity set and the data criticality set. When an indicator in the set exceeds the preset indicator threshold, the data classification adjustment process is triggered, the data classification rule library is called, and the corresponding classification rules are matched according to the data sensitivity set and the data criticality set to generate a new classification result for the data object identification set;
[0115] Update the control strategy, obtain the new classification result of the data object identifier set, determine the control strategy set corresponding to the new classification result based on the preset data classification and control strategy comparison table, use the difference analysis algorithm to compare the original control strategy with the new control strategy, and generate a control strategy update instruction set;
[0116] Access control rule adjustment: obtain the control policy update instruction set, parse the instruction set through the rule engine, and when the instruction type is access control rule adjustment, trigger the access control rule adjustment process, call the access control rule generation module, and generate a new data object identifier set access control rule;
[0117] Verify the new access control rules. Based on the new data object identifier set access control rules, use the machine learning algorithm based on the Transformer model to verify the new access control rules and determine the impact of the new access control rules on data security risk factors. If the impact is lower than the preset threshold, the new access control rules will be published to the access control system.
[0118] Data access behavior monitoring and dynamic adjustment: Based on the access log of the data object identifier set fed back by the access control system, a clustering machine learning algorithm combined with a K-nearest neighbor machine learning algorithm is used to identify data access behavior patterns, and the degree of deviation between the data access behavior pattern and the preset behavior pattern is determined. When the degree of deviation exceeds the preset threshold, a data classification dynamic adjustment process is performed, and data operation logs are collected again.
[0119] Specifically, by establishing a dynamic data classification adjustment and update mechanism, real-time monitoring and sensitivity assessment of government data usage are achieved. This automatically triggers reassessment and adjustment of data classification based on data usage frequency and sensitivity indicators, allowing for timely updates to data management policies and access control rules. This mechanism not only improves the flexibility and responsiveness of data management, but also verifies the effectiveness of new rules through machine learning algorithms, ensuring that data security risks are effectively controlled. This enables intelligent monitoring and dynamic adjustment of data access behavior, enhancing the overall security and compliance of government data.
[0120] Furthermore, after updating the data management strategy and access control rules, it also includes: establishing an audit and traceability mechanism for data classification changes, recording the history and reasons for data classification adjustments, and meeting compliance requirements.
[0121] Furthermore, an audit and traceability mechanism for data classification changes should be established to record the history and reasons for data classification adjustments, including:
[0122] Collect data hierarchical catalogs according to preset periods, obtain hierarchical information of each data item in the data hierarchical catalog, obtain hierarchical sets, determine whether the data hierarchical sets have changed within the period, and if so, trigger the data hierarchical change record process, and determine the changed data item set through data comparison;
[0123] Then, a text analysis algorithm is used to parse the business log associated with the changed data item to obtain a description of the change reason. The information extraction algorithm is used to parse the description of the change reason to obtain a change reason phrase. Based on the change reason phrase, a change reason label is determined.
[0124] Using preset rules to match each data item in the changed data item set, obtain the operator information of the changed data item, obtain the operator set and the identity information of each operator, and use the information encryption algorithm to encrypt the identity information of each operator to obtain encrypted identity information;
[0125] Generate a data change record based on the set of changed data items, the change reason label, and the encrypted identity information, store the data change record through the blockchain, and obtain a unique identifier for the data change record.
[0126] Specifically, by establishing an audit and traceability mechanism for data classification changes, the historical records and reasons for data classification adjustments can be recorded and tracked in detail, ensuring transparency and accountability in data management. Specifically, this mechanism automatically detects changes in data classifications by periodically collecting and comparing data classification directories. It also utilizes text analysis and information extraction techniques to parse the reasons for changes from business logs. Furthermore, encryption is used to ensure the privacy and security of operator information involved in data changes. All these change records are securely stored on the blockchain, ensuring immutability and traceability, meeting compliance requirements, and enhancing the security and credibility of data management.
[0127] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for automatically classifying and grading government data, characterized by: The steps include: Automatically correct and standardize data using statistical methods and machine learning algorithms by defining data quality rules and constraints; Based on automatically corrected and standardized data, we construct government domain ontologies and knowledge graphs, conduct semantic enhancement and association analysis on government data, build domain concept hierarchies and relationship network models, and utilize ontology reasoning and knowledge fusion technologies to perform data classification and decision-making. We also employ multi-classifier ensemble learning algorithms to integrate the classification results of different models. Among them, according to the characteristics and application requirements of government data, a government domain ontology and knowledge graph are constructed to define the core concepts, attributes and relationships of government data. The knowledge graph is used to express the relationship between data to form a conceptual hierarchy and relationship network model. For automatically corrected and standardized government data, natural language processing technology is used to perform semantic analysis on policy documents and announcements. Entity recognition technology is used to extract key policy measures, impact areas, and expected goals from policy documents. Semantic enhancement of the data is then performed, and cross-dataset correlation analysis is conducted by combining ontology reasoning and knowledge graphs. Based on domain ontology and knowledge graph, the association analysis of government data is carried out using ontology reasoning rules and knowledge fusion algorithms; Based on the results of semantic enhancement and association analysis, machine learning algorithms are used to train data classification models to automatically classify and support decision-making for government data; In the process of data classification and decision-making, ontology reasoning technology is used to perform knowledge deduction, and implicit classification rules and decision-making basis are derived based on the concept hierarchy and relationship network in the domain ontology; Through knowledge fusion technology, the classification rules and decision-making basis obtained from ontology reasoning are integrated with the prediction results of the machine learning model, and semantic information and data features are integrated to obtain the final classification and decision results; Based on the classification results, formulate unified standards and specifications for the grading and classification of government data. Based on the confidentiality dimension of the data, establish a data sensitivity assessment model and a grading decision-making mechanism to determine the level of data. Through consultation and review by multiple stakeholders, form a consensus data grading and classification plan. After forming a consensus on the data classification scheme, we build a knowledge base and rule base for the government affairs field, determine knowledge reasoning and rule matching, map the data to concepts and entities in the domain knowledge base, and conduct semantic association and contextual understanding of the data. Establish a dynamic adjustment and update mechanism for data classification, monitor data usage in real time, dynamically evaluate the sensitivity and importance of data, and automatically trigger the re-evaluation and adjustment process of data classification based on preset adjustment rules and thresholds, and update data management policies and access control rules.
2. A method for automatically classifying and grading government data according to claim 1, characterized in that: Before defining data quality rules and constraints, it also includes: collecting data from different sources and types and converting them into a standardized structured format, and using regular expressions to extract key information based on the content and characteristics of the data, and mapping it to a predefined data model to form a unified data representation.
3. The method for automatically classifying and grading government data according to claim 2, characterized in that: Regular expressions are used to extract key information and map it to a predefined data model to form a unified data representation, including: Acquire multi-source heterogeneous data, convert it into a standardized structured format using appropriate data cleaning and preprocessing methods, and determine the key information types and features to be extracted based on predefined data models; For standardized structured data, regular expression matching method is used to extract key information that meets specific characteristics; Map the extracted key information to the corresponding fields and attributes in the data model according to predefined mapping rules, and perform verification and integrity checks on the mapped data; The mapped data is stored in a unified format and structure to form a standardized data representation.
4. The method for automatically classifying and grading government data according to claim 1, characterized in that: Automatically correct and standardize data using statistical methods and machine learning algorithms, including: According to pre-defined data quality rules and constraints, the unified data is quality checked to identify abnormal data that does not meet the quality requirements. The abnormal data is analyzed using statistical methods to calculate the statistical characteristics of the abnormal data and determine the distribution of the abnormal data. Through the anomaly detection algorithm in the machine learning algorithm, abnormal data is automatically detected and identified to determine the location and range of abnormal data. Then, the data filling algorithm is used to automatically estimate and fill in missing or abnormal data values. According to the preset data standardization rules, the filled data is automatically standardized and converted into a unified format, unit and representation. The standardized data is then quality verified to determine whether the data meets the quality requirements based on the data quality rules and constraints. The processed high-quality data is stored in the database.
5. The method for automatically classifying and grading government data according to claim 1, characterized in that: Based on the confidentiality dimension of the data, a data sensitivity assessment model and hierarchical decision-making mechanism are established to determine the level of the data, including: Obtain government data sets, parse the data set metadata using a natural language processing algorithm based on a pre-established rule base, obtain the data type and data source, and determine whether the data type is sensitive. If it is sensitive, the data set is classified as a dataset to be evaluated. Government data sets include population data, economic data, social management data, urban management data, government service data, public resource data, and social security data. Obtain data content from the dataset to be evaluated, extract sensitive words from the data content through a text matching algorithm based on a pre-established keyword table, obtain the number and types of sensitive words, determine the weights of sensitive words, and calculate the sensitivity score; determine the data confidentiality level based on the sensitivity score and the preset threshold range.
6. A method for automatically classifying and grading government data according to claim 5, characterized in that: When the data confidentiality level exceeds the preset threshold, the data set is desensitized to obtain a desensitized data set. Desensitization methods include data masking and data replacement. Obtain user access records from the desensitized dataset. Based on the user role and permission matrix, determine the data scope accessed by the user and determine the user's access rights to specific datasets. If a user attempts to access beyond their permission, a violation log is recorded. Determine the sharing scope of the data set based on the data confidentiality level and sharing agreement. If the data confidentiality level is high, the sharing scope is limited to specific departments. Data is transmitted through the interface to obtain a shared data set. Obtain data update information from shared datasets and determine the data update frequency based on a pre-established schedule. When the data update frequency is high, the data synchronization mechanism is triggered to obtain the latest data version. The data retention time is determined based on the data storage period and destruction policy. When the data reaches the storage period, the data destruction process is triggered, the data is deleted from the storage medium, and a record of the destroyed data is obtained.
7. The method for automatically classifying and grading government data according to claim 1, characterized in that: By building a knowledge base and rule base for the government affairs field, determining knowledge reasoning and rule matching, mapping data to concepts and entities in the domain knowledge base, and performing semantic association and contextual understanding of data, the following are specifically included: By comparing structured data with concept hierarchies and relational networks, if the matching degree is higher than a preset threshold, the concept to which the data belongs is determined. If it is lower than the preset threshold, the structured data is compared with the entity layer to obtain the entity with the highest matching degree, and the entity set corresponding to the data is obtained. Through the entity set, the association rules in the rule library are called to obtain the association relationship between entities. Based on the association relationship between entities, an entity relationship graph is constructed. By traversing the graph, the data context information is determined to obtain the final semantic understanding result.
8. The method for automatically classifying and grading government data according to claim 1, characterized in that: Establish a data classification dynamic adjustment and update mechanism, including: Data collection and recording: Collect data operation logs, obtain the operation user, operation type, operation time, and operation data object identification information through log analysis technology, and record them in the data usage record database; Frequency of use statistics and sensitivity assessment: By recording the database information of data usage, the frequency of use of data object identifiers within a preset time range is counted. When the frequency exceeds the preset frequency threshold, the data sensitivity and criticality assessment process is triggered. The pre-established assessment model is used to calculate the data object identifier set to obtain the data sensitivity set and data criticality set; The data classification adjustment trigger obtains the data sensitivity set and the data criticality set. When an indicator in the set exceeds the preset indicator threshold, the data classification adjustment process is triggered, the data classification rule library is called, and the corresponding classification rules are matched according to the data sensitivity set and the data criticality set to generate a new classification result for the data object identification set; Update the control strategy, obtain the new classification result of the data object identifier set, determine the control strategy set corresponding to the new classification result based on the preset data classification and control strategy comparison table, use the difference analysis algorithm to compare the original control strategy with the new control strategy, and generate a control strategy update instruction set; Access control rule adjustment: obtain the control policy update instruction set, parse the instruction set through the rule engine, and when the instruction type is access control rule adjustment, trigger the access control rule adjustment process, call the access control rule generation module, and generate a new data object identifier set access control rule; Verify the new access control rules. Based on the new data object identifier set access control rules, use the machine learning algorithm based on the Transformer model to verify the new access control rules and determine the impact of the new access control rules on data security risk factors. If the impact is lower than the preset threshold, the new access control rules will be published to the access control system. Data access behavior monitoring and dynamic adjustment: Based on the access log of the data object identifier set fed back by the access control system, a clustering machine learning algorithm combined with a K-nearest neighbor machine learning algorithm is used to identify data access behavior patterns, and the degree of deviation between the data access behavior pattern and the preset behavior pattern is determined. When the degree of deviation exceeds the preset threshold, a data classification dynamic adjustment process is performed, and data operation logs are collected again.
9. The method for automatically classifying and grading government data according to claim 1, characterized in that: After updating the data management strategy and access control rules, it also includes: establishing an audit and traceability mechanism for data classification changes, and recording the history and reasons for data classification adjustments.
10. A method for automatically classifying and grading government data according to claim 9, characterized in that: Establish an audit and traceability mechanism for data classification changes to record the history and reasons for data classification adjustments, including: Collect data hierarchical catalogs according to preset periods, obtain hierarchical information of each data item in the data hierarchical catalog, obtain hierarchical sets, determine whether the data hierarchical sets have changed within the period, and if so, trigger the data hierarchical change record process, and determine the changed data item set through data comparison; Then, a text analysis algorithm is used to parse the business log associated with the changed data item to obtain a description of the change reason. The change reason description text is parsed using an information extraction algorithm to obtain a change reason phrase. Based on the change reason phrase, a change reason label is determined. Using preset rules to match each data item in the changed data item set, obtain the operator information of the changed data item, obtain the operator set and the identity information of each operator, and use the information encryption algorithm to encrypt the identity information of each operator to obtain encrypted identity information; Generate a data change record based on the set of changed data items, the change reason label, and the encrypted identity information, store the data change record through the blockchain, and obtain a unique identifier for the data change record.
Citation Information
Patent Citations
Government affair service field multi-strategy fusion dialogue method based on knowledge graph
CN116628172A
Method and system for improving data security of temporary office local area network
CN118157996A
DCS intelligent decision-making method and system fusing large language model and knowledge graph
CN118820778A
Government affair data extraction and analysis method and system based on big data
CN119003780A
Mirror image type industry linkage engine system and method based on knowledge graph
CN119537864A
Cited By
Government affair knowledge graph ontology construction and optimization method and device, equipment and medium
CN120930760A
Data processing method and device and storage medium
CN120994671A
Government affair data processing system and method
CN121073402A
Foreign trade data intelligent classification system based on natural language processing
CN121233772A
Intelligent data security classification and grading method based on multi-modal large model
CN121234138A