Government affair data asset atlas construction method and device

By building a government data asset map, the problems of data heterogeneity, blood tracing, value evaluation and update in government data management are solved, efficient data sharing and cross-departmental data correlation are achieved, and data security compliance is met, and government decision-making is supported.

CN120386905APending Publication Date: 2025-07-29CICC DATA (WUHAN) SUPERCOMPUTING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510471926.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing data heterogeneity in the asset management of government data has led to difficulties in integration, insufficient credibility of data tracing, strong subjectivity of data value assessment, rigid map update mechanism, and contradiction between privacy protection and data utilization, making it difficult to achieve efficient data sharing and cross-departmental data correlation.

Method used

Multi-source heterogeneous data preprocessing, government data ontology modeling, asset value evaluation model, entropy weight method, blockchain technology and graph neural network are used to build a government data asset map to realize data visualization and dynamic updates, and combine federated learning and blockchain to ensure data security.

Benefits of technology

It improves data sharing efficiency, reduces the policy formulation cycle, improves the accuracy of cross-departmental data correlation rules, meets data security compliance requirements, and supports government decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386905A_ABST
    Figure CN120386905A_ABST
Patent Text Reader

Abstract

The invention provides a government affair data asset atlas construction method and device, and relates to the technical field of data processing. The method comprises the following steps: acquiring multi-source heterogeneous data; preprocessing the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data; inputting the preprocessed multi-source heterogeneous data into an ontology modeling tool, and performing ontology construction on the preprocessed multi-source heterogeneous data by adopting the ontology modeling tool to obtain a government affair data ontology model; determining a multi-dimensional index of the government affair data ontology model by adopting an asset value evaluation model; determining an asset value index according to the multi-dimensional index by adopting an entropy weight method; and visualizing the government affair data ontology model and the asset value index to obtain a data asset map. The data sharing efficiency can be improved, the policy making period is shortened, and the accuracy of the cross-department data association rule is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method and device for constructing a government affairs data asset graph. Background Art

[0002] With the in-depth promotion of the construction of digital government, government affairs data, as a core production factor, its efficient governance and value mining have become the key to improving the government's governance ability. However, there are still significant defects in the existing technology for the asset management of government affairs data, which restricts data sharing and in-depth application: data heterogeneity leads to integration difficulties:

[0003] Government affairs data is scattered in multi-department systems and exists in various forms such as structured (such as database tables), semi-structured (such as XML files), and unstructured (policy texts, GIS maps). Traditional data modeling methods (such as relational database Schema mapping) are difficult to achieve semantic unity. For example, the "personal number" of a certain city's social security bureau and the "resident ID" of the civil affairs bureau point to the same entity, but due to naming differences, they cannot be automatically associated, and relying on manual mapping is inefficient.

[0004] The credibility of data lineage tracing is insufficient:

[0005] Existing data lineage tracing technologies are mostly based on centralized logs (such as database trigger records of operation flowsheets), which have the risk of tampering and are difficult to verify across departments.

[0006] The subjectivity of data value evaluation is strong:

[0007] Current evaluation models mostly adopt fixed weights (such as the analytic hierarchy process) or single-dimensional indicators (such as data volume), ignoring the impact of dynamic business scenarios; in addition, traditional static models cannot adjust weights in real time, resulting in distorted value evaluation.

[0008] The graph update mechanism is rigid:

[0009] Government affairs data changes frequently with policy adjustments, while existing knowledge graphs mostly rely on full-scale updates (such as periodic ETL reruns), which are time-consuming and energy-consuming.

[0010] The contradiction between privacy protection and data utilization:

[0011] Existing federated learning solutions only support simple feature exchange and are difficult to support complex graph relationship reasoning, resulting in difficulty in balancing privacy protection and data utility. Summary of the Invention

[0012] The purpose of the present invention is to provide a method and device for constructing a government affairs data asset graph in view of the above deficiencies in the existing technology, so as to improve data sharing efficiency, reduce the policy-making cycle, and improve the accuracy of cross-department data association rules.

[0013] To achieve the above object, the technical solutions adopted in the embodiments of the present application are as follows:

[0014] In a first aspect, an embodiment of the present application provides a method for constructing a government affairs data asset map, including: obtaining multi-source heterogeneous data; the multi-source heterogeneous data includes: structured data, unstructured data; the structured data includes spatial data; preprocessing the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data; inputting the preprocessed multi-source heterogeneous data into an ontology modeling tool, and using the ontology modeling tool to construct an ontology for the preprocessed multi-source heterogeneous data to obtain a government affairs data ontology model; the government affairs data ontology model includes: an ontology, attributes of the ontology, and relationships between ontologies; using an asset value evaluation model to determine multi-dimensional indicators of the government affairs data ontology model; the multi-dimensional indicators include: a data integrity index, a data call frequency, and a data social benefit value; using the entropy weight method to determine an asset value index according to the multi-dimensional indicators; visualizing the government affairs data ontology model and the asset value index to obtain a data asset map.

[0015] In one implementation, the preprocessing the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data includes: de-duplicating, standardizing, and handling missing values of the multi-source heterogeneous data to obtain the preprocessed multi-source heterogeneous data.

[0016] In one implementation, the using the ontology modeling tool to construct an ontology for the preprocessed multi-source heterogeneous data to obtain a government affairs data ontology model includes: for the structured data other than the spatial data, using a Schema Mapping tool to map fields to the government affairs data ontology model; for the unstructured data, using an NLP model to extract relationships between ontologies and binding the relationships between ontologies to the government affairs data ontology model; for the spatial data, using the GeoJSON specification to convert the spatial data into standard spatial coordinate data and associating the standard spatial coordinate data to the government affairs data ontology model.

[0017] In one implementation, using an asset value evaluation model to determine multi-dimensional indicators of the government affairs data ontology model includes: using an integrity index calculation formula to determine the data integrity index; the integrity index calculation formula is as follows:

[0018]

[0019] Among them, μ is the field weight, which is determined by expert scoring or the information gain method; D is the existence score, where the field is non-empty as 1 and the null value is 0; E is the rule verification score, where the format verification passes as 1, fails as 0, the business rule verification passes as 1, and fails as 0.5.

[0020] The data call frequency is determined by using the call popularity calculation formula; the call popularity calculation formula is as follows:

[0021] χ = ∑(C t ×e (-λ×(T-t) )

[0022] Among them, C is the number of accesses; t is the call occurrence event, calculated by day; T is the current time; λ is the time decay factor.

[0023] The data social benefit value is determined by using the social benefit value calculation formula; the social benefit value calculation formula is as follows:

[0024] δ = β × A + (1 - β) × B

[0025] Among them, β is the weight of the objective index, and the objective index includes: the number of policy citations, the coverage of people's livelihood impact, and the economic value; A is the objective index; (1 - β) is the weight of the expert score; B is the expert score.

[0026] In one implementation manner, using the entropy weight method to determine the asset value index according to the multi-dimensional index includes: using the entropy weight method to determine the weights of different dimension indexes in the multi-dimensional index; the entropy weight method is as follows:

[0027]

[0028] Among them, w k is the principal component weight; F k is the contribution rate of the kth principal component; p k is the entropy weight of the kth principal component; F i is the contribution rate of the ith principal component; p i is the entropy weight of the ith principal component;

[0029] According to the weights of the different dimension indexes and the multi-dimensional index, the asset value index is determined.

[0030] In one implementation, after obtaining the data asset graph, the method further includes: using blockchain technology to record the data flow path, and when the data is called, triggering a smart contract to upload the data user's permissions and operation traces to the blockchain; obtaining the blockchain transaction log, extracting the data flow path from the blockchain transaction log, and storing it in the Neo4j graph database to obtain the data lineage graph; the nodes in the data lineage graph include: data tables, files, APIs; the edge relationships include: source, derivation, influence.

[0031] In one implementation, after obtaining the data lineage graph, the method further includes: constructing a graph neural network according to the data asset graph; the node features in the graph neural network include: the asset value index and sensitivity level of the data; the edge weights include: the data flow frequency and time decay factor; inputting the data propagation path into the graph neural network, and using the graph neural network to identify the data propagation path to output the abnormal data propagation path.

[0032] In one implementation, after obtaining the data asset graph, the method further includes: listening to an external event source; when detecting data updates, extracting keywords in the updated data and associating the keywords with the affected ontology.

[0033] In one implementation, after associating the keywords with the affected ontology, the method further includes: each department locally trains a GNN model; inputting the local graph into the GNN model, and using the GNN model to extract features from the local graph to obtain local graph features; the central server receives the corresponding local graph features uploaded by each department, aggregates the local graph features corresponding to each department to generate a global graph; the central server sends the global graph to each department so that each department can update the local graph according to the global graph.

[0034] Second aspect, an embodiment of the present application further provides a government affairs data asset graph construction device 900, including: an acquisition module configured to acquire multi-source heterogeneous data; the multi-source heterogeneous data includes: structured data, unstructured data; the structured data includes spatial data; a preprocessing module configured to preprocess the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data; an ontology construction module configured to input the preprocessed multi-source heterogeneous data into an ontology modeling tool, and use the ontology modeling tool to perform ontology construction on the preprocessed multi-source heterogeneous data to obtain a government affairs data ontology model; the government affairs data ontology model includes: an ontology, attributes of the ontology, and relationships between ontologies; a first determination module configured to use an asset value evaluation model to determine multi-dimensional indicators of the government affairs data ontology model; the multi-dimensional indicators include: a data integrity index, a data call frequency, and a data social benefit value; a second determination module configured to use the entropy weight method to determine an asset value index according to the multi-dimensional indicators; a visualization module configured to visualize the government affairs data ontology model and the asset value index to obtain a data asset graph.

[0035] In one implementation, the preprocessing module is configured to: perform deduplication, standardization, and missing value processing on the multi-source heterogeneous data to obtain the preprocessed multi-source heterogeneous data.

[0036] In one implementation, the ontology construction module is configured to: for the structured data other than the spatial data, use a Schema Mapping tool to map fields to the government affairs data ontology model; for the unstructured data, use an NLP model to extract relationships between ontologies and bind the relationships between ontologies to the government affairs data ontology model; for the spatial data, use the GeoJSON specification to convert the spatial data into standard spatial coordinate data and associate the standard spatial coordinate data with the government affairs data ontology model.

[0037] In one implementation, the first determination module is configured to: use an integrity index calculation formula to determine the data integrity index; the integrity index calculation formula is as follows:

[0038]

[0039] Among them, μ is the field weight, which is determined by expert scoring or the information gain method; D is the existence score, 1 for non-empty fields and 0 for null values; E is the rule verification score, 1 for passing the format verification, 0 for failure, 1 for passing the business rule verification, and 0.5 for failure.

[0040] Use a call heat calculation formula to determine the data call frequency; the call heat calculation formula is as follows:

[0041] χ = ∑(C t × e (-λ×(T-t) )

[0042] where C is the number of accesses; t is the event of call occurrence, calculated by day; T is the current time; λ is the time decay factor;

[0043] The social benefit value of the data is determined by using the social benefit value calculation formula; the social benefit value calculation formula is as follows:

[0044] δ = β × A + (1 - β) × B

[0045] where β is the weight of the objective indicators, and the objective indicators include: the number of policy citations, the coverage of people's livelihood impact, and economic value; A is the objective indicator; (1 - β) is the weight of the expert score; B is the expert score.

[0046] In one implementation, the second determination module is configured to: use the entropy weight method to determine the weights of different dimension indicators in the multi - dimension indicators; the entropy weight method is as follows:

[0047]

[0048] where w k is the principal component weight; F k is the contribution rate of the k - th principal component; p k is the entropy weight of the k - th principal component; F i is the contribution rate of the i - th principal component; p i is the entropy weight of the i - th principal component;

[0049] Determine the asset value index according to the weights of the different dimension indicators and the multi - dimension indicators.

[0050] In one implementation, the government affairs data asset graph construction device 900 further includes a data traceability module, and the data traceability module is configured to: use blockchain technology to record the data flow path, and when the data is called, trigger a smart contract to chain the data user's permissions and operation traces; obtain the blockchain transaction log, extract the data flow path from the blockchain transaction log, and store it in the Neo4j graph database to obtain a data lineage graph; the nodes in the data lineage graph include: data tables, files, APIs; the edge relationships include: source, derivation, and impact.

[0051] In one embodiment, the data traceability module is configured to: construct a graph neural network according to the data asset graph; the node features in the graph neural network include: the asset value index and sensitivity level of the data; the edge weights include: the data transfer frequency and time decay factor; input the data propagation path into the graph neural network, and use the graph neural network to identify the data propagation path and output the abnormal data propagation path.

[0052] In one embodiment, the government data asset graph construction device 900 further includes a data update module, and the data update module is configured to: monitor an external event source; when detecting data update, extract keywords in the updated data and associate the keywords with the affected ontology.

[0053] In one embodiment, the data update module is configured to: locally train a GNN model for each department; input the local graph into the GNN model, and use the GNN model to extract features of the local graph to obtain local graph features; the central server receives the corresponding local graph features uploaded by each department, aggregates the local graph features corresponding to each department to generate a global graph; the central server sends the global graph to each department so that each department can update the local graph according to the global graph.

[0054] In a third aspect, an embodiment of the present application provides a computer device, including: a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the computer device runs, the processor communicates with the storage medium through the bus, and the processor executes the program instructions to perform the steps of any of the above methods.

[0055] In a fourth aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of any of the above methods.

[0056] The beneficial effects of the present application are:

[0057] (1) It improves data sharing efficiency, reduces the policy-making cycle, and improves the accuracy of cross-department data association rules;

[0058] (2) By forming a closed loop of data access → modeling → traceability → update, it solves the problem of traditional graph staticization;

[0059] (3) Through federated learning + blockchain, it realizes "data can be used but not visible", meeting the compliance requirements of the Data Security Law;

[0060] (4) Through policy deduction, asset valuation, etc., it can directly support government decision-making, exceeding the positioning of pure technical tools. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 It is a schematic flowchart of a method for constructing a government affairs data asset map provided by an embodiment of the present application;

[0063] Figure 2 It is a schematic flowchart of a method for constructing a government affairs data asset map provided by an embodiment of the present application;

[0064] Figure 3 It is a schematic flowchart of a method for constructing a government affairs data asset map provided by an embodiment of the present application;

[0065] Figure 4 It is a schematic flowchart of a method for constructing a government affairs data asset map provided by an embodiment of the present application;

[0066] Figure 5 It is a schematic flowchart of a method for constructing a government affairs data asset map provided by an embodiment of the present application;

[0067] Figure 6 It is a schematic flowchart of a method for constructing a government affairs data asset map provided by an embodiment of the present application;

[0068] Figure 7 It is a schematic flowchart of a method for constructing a government affairs data asset map provided by an embodiment of the present application;

[0069] Figure 8 It is a schematic flowchart of a method for constructing a government affairs data asset map provided by an embodiment of the present application;

[0070] Figure 9 It is a schematic structural diagram of a device for constructing a government affairs data asset map provided by an embodiment of the present application;

[0071] Figure 10 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention.

[0073] Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.

[0074] In the description of the present application, it should be noted that if terms such as "upper", "lower", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship when the product of this application is usually placed, it is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.

[0075] In addition, terms such as "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0076] It should be noted that the features in the embodiments of the present application can be combined with each other without conflict.

[0077] Figure 1 A flowchart showing the process of a method for constructing a government affairs data asset map provided for the embodiments of the present application; as Figure 1 shown, the method includes the following steps 110 to step 160:

[0078] Step 110, obtaining multi-source heterogeneous data.

[0079] Among them, the multi-source heterogeneous data includes: structured data, unstructured data.

[0080] Structured data includes the database table structures of each department and spatial data; the database table structures of each department, such as: the MySQL tables of the Social Security Bureau and the Oracle tables of the Market Supervision Bureau; spatial data includes: GIS map files (Shapefile, GeoJSON), and the coordinates of Internet of Things sensors.

[0081] Unstructured data includes policy documents (PDF / Word), report texts, web data, and reference standards; reference standards include: the national government affairs data resource catalog and industry standards.

[0082] Step 120: Preprocess the multi-source heterogeneous data to obtain the preprocessed multi-source heterogeneous data.

[0083] Among them, in actual operations, the data of each department may not be interoperable, and there is a problem of data islands. In order to integrate the data of each department, it is necessary to preprocess the data of each department to standardize the data. Specifically, the above step 120 may further include the following steps:

[0084] Deduplicate, standardize, and handle missing values in the multi-source heterogeneous data to obtain the preprocessed multi-source heterogeneous data.

[0085] Among them, deduplication means deleting completely duplicate data or partially duplicate data and retaining key data.

[0086] Standardization means unifying the data format to ensure the consistency of formats such as dates, units, and categorical variables. For example: using regular expressions or date conversion tools to convert "2023 / 01 / 01" to "2023-01-01".

[0087] Missing value handling means identifying and handling null value fields in the data; here, specifically, the null value fields in the data can be marked and an alarm can be triggered.

[0088] Step 130: Input the preprocessed multi-source heterogeneous data into an ontology modeling tool, and use the ontology modeling tool to construct an ontology for the preprocessed multi-source heterogeneous data to obtain a government affairs data ontology model.

[0089] Among them, the government affairs data ontology model includes: ontology, the attributes of the ontology, and the relationships between ontologies.

[0090] The ontology modeling tool can adopt the Protégé ontology modeling tool. The Protégé ontology modeling tool is mainly used for the construction of ontologies in the semantic web and ontology-based knowledge applications, and is the core development tool for ontology construction.

[0091] For different types of data, the processing methods of the ontology modeling tool are different; specifically, such as Figure 2As shown, the above-mentioned capture 130 may further include the following steps 210 to 230:

[0092] Step 210: For structured data other than spatial data, use a Schema Mapping tool to map fields to the government data ontology model.

[0093] Among them, the Schema Mapping tool can use Altova MapForce.

[0094] The SchemaMapper tool can convert the data model to a new structure according to the mapping in an external lookup table (such as a database or CSV). It processes complex schema mappings and is suitable for situations that need to be maintained by non-FME users. The converter has two output ports, mapped and unmapped, and supports mapping types such as Filters, Feature Type Map, AttributeMap, and New Attribute. Examples are shown to illustrate how to configure and use these mapping types for attribute and feature type conversion.

[0095] Exemplarily, the SchemaMapper tool can map "Social Security Database. Personal Number" to "GDOM: Citizen. Social Security ID"; here, GDOM represents the above-mentioned government data ontology model.

[0096] Step 220: For unstructured data, use an NLP model to extract the relationships between ontologies and bind the relationships between ontologies to the government data ontology model.

[0097] Among them, the NLP model can use the BERT model, the BiLSTM-CRF model, etc., which are not limited here.

[0098] The NLP model is an important branch in the field of artificial intelligence, aiming to enable computers to understand, analyze, and generate human language. It processes a large amount of text data, learns the structure, grammar, and semantic information of language, and is widely used in tasks such as text classification, sentiment analysis, and machine translation.

[0099] The application scenarios of the NLP model are numerous. One of them is the application in named entity recognition and question answering systems, such as the application of the BERT model in Google Search.

[0100] The BERT (Bidirectional Encoder Representations from Transformers) model is a pre-trained language model based on the Transformer architecture. It uses a bidirectional Transformer encoder, which can capture the context information before and after words simultaneously, so as to understand the text semantics more comprehensively.

[0101] The BiLSTM-CRF model is a model that combines a bidirectional long short-term memory network (BiLSTM) and a conditional random field (CRF), and is mainly used for sequence labeling tasks such as named entity recognition (NER).

[0102] Exemplarily, <policy document, release date, 2023-01-01> can be bound to GDOM: policy.effective time.

[0103] Step 230: For spatial data, convert the spatial data into standard spatial coordinate data using the GeoJSON specification, and associate the standard spatial coordinate data with the government affairs data ontology model.

[0104] Among them, the GeoJSON specification defines how to uniformly convert different coordinate systems (such as: WGS84, CGCS2000) into the standard GeoJSON format.

[0105] Exemplarily, use the GeoJSON specification to define the geographic fence association rules, such as: the grid of community A → the affiliated street B → the administrative division C.

[0106] Step 140: Use the asset value evaluation model to determine the multi-dimensional indicators of the government affairs data ontology model.

[0107] Among them, the multi-dimensional indicators include: data integrity index, data call frequency, data social benefit value.

[0108] The asset value evaluation model determines data in different dimensions (i.e., the above multi-dimensional indicators) in different ways; specifically, as Figure 3 shown, the above step 140 may further include the following steps 310 to 330:

[0109] Step 310: Use the integrity index calculation formula to determine the data integrity index; the integrity index calculation formula is shown as formula (1) below:

[0110]

[0111] Among them, μ is the field weight, which is determined by expert scoring or information gain method; D is the existence score, 1 for non-empty fields and 0 for null values; E is the rule verification score, 1 for passing the format verification, 0 for failure, 1 for passing the business rule verification, and 0.5 for failure.

[0112] Exemplarily, a certain social security form contains 10 fields, the ID number (weight 0.3) exists and the format is correct, but the payment base exceeds the limit (score 0.5).

[0113] Step 320: Determine the data call frequency using the call heat calculation formula. The call heat calculation formula is as shown in formula (2) below:

[0114] χ = ∑(C t × e (-λ×(T-t) )(2)

[0115] where C is the number of accesses; t is the call occurrence event, calculated by day; T is the current time; and λ is the time decay factor.

[0116] λ being 0.1 means a 10% decay per day. λ can be set according to the actual situation and is not limited here.

[0117] Exemplarily, the call times of a certain API in the recent 3 days are: 10 times (today), 20 times (yesterday), and 30 times (the day before yesterday). Then, according to the above formula, the call heat is calculated to be 52.7.

[0118] Step 330: Determine the data social benefit value using the social benefit value calculation formula. The social benefit value calculation formula is as shown in formula (3) below:

[0119] δ = β × A + (1 - β) × B(3)

[0120] where β is the weight of the objective index, and the objective index includes: the number of policy citations, the coverage of people's livelihood impact, and economic value; A is the objective index; (1 - β) is the weight of the expert score; and B is the expert score.

[0121] The number of policy citations refers to the number of times the data is cited in policy documents.

[0122] The coverage of people's livelihood impact refers to the proportion of the affected population. For example, the medical insurance data covers 80% of the city's residents.

[0123] The economic value refers to the estimated economic benefits supported by the data. For example, the fuel cost saved after a certain traffic data optimizes the route.

[0124] Exemplarily, the objective index: cited by 3 policies (score 0.7 after standardization), covering 60% of the population (score 0.6), and the estimated economic value is 1 million yuan (score 0.5); the objective score = (0.7 + 0.6 + 0.5) / 3 = 0.6. The expert score = 0.8. The final social benefit value (β = 0.6): 0.6 × 0.6 + 0.4 × 0.8 = 0.68.

[0125] Step 150: Use the entropy weight method to determine the asset value index based on multi-dimensional indicators.

[0126] Among them, the entropy weight method is an objective weighting method based on the information entropy theory. Its core principle is to determine the weights of each index in the comprehensive evaluation by calculating the degree of dispersion of the index data. The smaller the information entropy value of an index, the greater the degree of dispersion and the higher the influence weight on the evaluation result.

[0127] In this step, the weights of the indexes in different dimensions can be determined first, and then the asset value index can be determined. Specifically, as Figure 4 shown, the above step 150 can further include the following steps 410 and 420:

[0128] Step 410: Use the entropy weight method to determine the weights of the indexes in different dimensions among the multi-dimensional indexes; the entropy weight method is shown in the following formula (4):

[0129]

[0130] where w k is the principal component weight; F k is the contribution rate of the k-th principal component; p k is the entropy weight of the k-th principal component; F i is the contribution rate of the i-th principal component; p i is the entropy weight of the i-th principal component.

[0131] The above entropy weight method is an improved and optimized entropy weight method. The existing entropy weight method directly calculates the information entropy of each index, ignoring the correlation between indexes, resulting in the problem that the weights of highly correlated indexes may be calculated repeatedly. For example, the call frequency is highly correlated with the number of users, and the weights may be calculated repeatedly. The improved and optimized entropy weight method can be understood as the principal component entropy weight method, which can use principal component analysis (PCA) to reduce the dimension of the data, eliminate multicollinearity (i.e., eliminate the correlation interference), and make the weight allocation more scientific. Its core idea is to allocate weights based on the contribution rate of the principal components.

[0132] Step 420: Determine the asset value index according to the weights of the indexes in different dimensions and the multi-dimensional indexes.

[0133] Among them, in this step, the weights of the indexes in different dimensions and the corresponding indexes can be weighted; for example: asset value index = 0.4 × data integrity index + 0.3 × data call popularity + 0.3 × social benefit value.

[0134] Step 160: Visualize the government affairs data ontology model and the asset value index to obtain the data asset atlas.

[0135] Among them, in this step, the data ontology model (i.e., the ontology, the attributes of the ontology, and the relationships between the ontologies) and the asset value index of the data can be visualized in the form of an atlas.

[0136] In actual operation, a data traceability function can be added based on modeling; specifically, as Figure 5 shown, after the above step 160, the method for constructing a government data asset map provided by the embodiments of the present application may further include the following steps 510 and 520:

[0137] Step 510: Use blockchain technology to record the data flow path. When the data is called, trigger a smart contract to record the data user's permissions and operation traces on the chain.

[0138] Among them, blockchain technology is a revolutionary distributed database technology that reshapes the trust architecture of the digital world through its decentralized, immutable, and transparent characteristics.

[0139] A smart contract is a computer protocol designed to spread, verify, or execute a contract in an informatized manner. Smart contracts allow for trusted transactions without a third party, and these transactions are traceable and irreversible.

[0140] In this step, when the data is called, trigger the smart contract to record, and then use the Hyperledger Fabric consortium chain to set the government department as a Peer node, and the data user's permissions and operation traces can be recorded on the chain.

[0141] Step 520: Obtain the blockchain transaction log, extract the data flow path from the blockchain transaction log, and store it in the Neo4j graph database to obtain a data lineage map.

[0142] Among them, the nodes in the data lineage map include: data tables, files, APIs; the edge relationships include: source, derivative, influence.

[0143] Exemplarily, the data flow path is the tax data of enterprise A → the financial bureau report → the statistical bureau economic analysis report.

[0144] After obtaining the data lineage map, causal inference anomaly propagation analysis can be implemented based on a graph neural network; specifically, as Figure 6 shown, after the above step 520, the method for constructing a government data asset map provided by the embodiments of the present application may further include the following steps 610 and 620:

[0145] Step 610: Construct a graph neural network according to the data asset map.

[0146] Among them, the node features in the graph neural network include: the asset value index and sensitivity level of the data; the edge weights include: the data flow frequency and time decay factor.

[0147] A Graph Neural Network (GNN) refers to a general term for algorithms that use neural networks to learn graph-structured data, extract and discover features and patterns in graph-structured data, and meet the requirements of graph learning tasks such as clustering, classification, prediction, segmentation, and generation.

[0148] Step 620: Input the data propagation path into the graph neural network, and use the graph neural network to identify the data propagation path, and output the abnormal data propagation path.

[0149] Among them, the above-mentioned causal inference abnormal propagation analysis can be realized based on the graph neural network; specifically, the graph neural network can identify the input data propagation path, thereby determining the propagation range of abnormal data (such as: missing fields), thereby locating the data abnormal propagation path, and outputting the data abnormal propagation path.

[0150] Exemplarily, if there is a field error in the "demographic data" of a certain city input, the graph neural network can output the influence chain population data → education bureau degree planning → transportation bureau bus line adjustment. Furthermore, the visualization tool Gephi is used to generate a red-highlighted warning path.

[0151] In actual operation, a data update function can be further added on the basis of data tracing, and then a closed loop of data access → modeling → tracing → update is formed, thereby solving the problem of static traditional graphs; specifically, as Figure 7 shown, the method for constructing a government affairs data asset graph provided by the embodiment of the present application can further include the following steps 710 and 720:

[0152] Step 710: Listen to external event sources.

[0153] Among them, external events can be listened to through the API, for example: relevant events of the government gazette are listened to through the government gazette API.

[0154] Step 720: When data update is detected, extract keywords in the updated data and associate the keywords with the affected ontology.

[0155] Exemplarily, when a policy document update is detected, extract keywords (such as: a 5% increase in pensions), and then associate the keywords with the affected ontology (such as: GDOM: social security. pension standard).

[0156] In actual operation, after data update, multi-domain collaboration supported by federated learning can also be performed to reduce the cost of data update; specifically, as Figure 8 shown, the method for constructing a government affairs data asset graph provided by the embodiment of the present application can further include the following steps 810 to 840:

[0157] Step 810: Each department trains the graph neural network locally.

[0158] Graph Neural Network (GNN) is a deep learning model specifically designed to process graph structured data and is widely used in tasks such as node classification, link prediction, and graph classification.

[0159] For example, the Civil Affairs Bureau, Medical Insurance Bureau, and Tax Bureau (the data does not leave the local area) each train a graph neural network locally.

[0160] Step 820: Input the local graph into the graph neural network, use the graph neural network to extract features of the local graph, and obtain local graph features.

[0161] Graph neural networks capture the structural information and node features of a graph and learn low-dimensional embedding representations of nodes. Their core concept is to transform graph data into a format that neural networks can process, typically through node and edge feature propagation. Graph neural networks can uncover deep patterns and semantic features underlying nodes and edges in graph data.

[0162] Step 830: The central server receives the corresponding local graph features uploaded by each department, aggregates the local graph features corresponding to each department, and generates a global graph.

[0163] Among them, the central server can use the FedAvg algorithm to aggregate the corresponding local graph features uploaded by each department, and then generate a global graph.

[0164] Step 840: The central server sends the global map to each department so that each department can update the local map based on the global map.

[0165] In this step, the central server simply sends the global graph to each department, and each department can then update its local graph based on the global graph. This reduces data update costs through multi-domain collaboration supported by federated learning.

[0166] The method for constructing a government affairs data asset map provided by the embodiments of the present application includes: First, obtaining multi-source heterogeneous data; the multi-source heterogeneous data includes: structured data, unstructured data; Secondly, preprocessing the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data; Thirdly, inputting the preprocessed multi-source heterogeneous data into an ontology modeling tool, and using the ontology modeling tool to construct an ontology for the preprocessed multi-source heterogeneous data to obtain a government affairs data ontology model; the government affairs data ontology model includes: ontology, attributes of the ontology, relationships between ontologies; Then, using an asset value evaluation model to determine multi-dimensional indicators of the government affairs data ontology model; the multi-dimensional indicators include: data integrity index, data call frequency, data social benefit value; Then, using the entropy weight method to determine the asset value index according to the multi-dimensional indicators; Finally, visualizing the government affairs data ontology model and the asset value index to obtain a data asset map. In this way, a closed loop is formed from data access → modeling → traceability → update, solving the problem of static traditional maps; federated learning + blockchain realizes "data can be used but not visible", meeting the compliance requirements of the Data Security Law; policy deduction, asset valuation, etc. can directly support government decision-making, exceeding the positioning of pure technical tools; improving data sharing efficiency, reducing the policy formulation cycle, and improving the accuracy of cross-departmental data association rules.

[0167] After introducing the method for constructing a government affairs data asset map according to the exemplary embodiments of the present disclosure, next, reference is made to Figure 9 to describe the government affairs data asset map construction device 900 according to the exemplary embodiments of the present disclosure.

[0168] Reference is made to Figure 9 , the government affairs data asset map construction device 900 includes: an acquisition module 910 configured to acquire multi-source heterogeneous data; the multi-source heterogeneous data includes: structured data, unstructured data; the structured data includes spatial data; a preprocessing module 920 configured to preprocess the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data; an ontology construction module 930 configured to input the preprocessed multi-source heterogeneous data into an ontology modeling tool, and use the ontology modeling tool to construct an ontology for the preprocessed multi-source heterogeneous data to obtain a government affairs data ontology model; the government affairs data ontology model includes: ontology, attributes of the ontology, relationships between ontologies; a first determination module 940 configured to use an asset value evaluation model to determine multi-dimensional indicators of the government affairs data ontology model; the multi-dimensional indicators include: data integrity index, data call frequency, data social benefit value; a second determination module 950 configured to use the entropy weight method to determine the asset value index according to the multi-dimensional indicators; a visualization module 960 configured to visualize the government affairs data ontology model and the asset value index to obtain a data asset map.

[0169] In one embodiment, the preprocessing module 920 is configured to: deduplicate, standardize, and handle missing values of multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data.

[0170] In one embodiment, the ontology construction module 930 is configured to: for structured data, use the SchemaMapping tool to map fields to the government affairs data ontology model; for unstructured data, use the NLP model to extract the relationships between ontologies and bind the relationships between ontologies to the government affairs data ontology model; for spatial data, use the GeoJSON specification to convert spatial data into standard spatial coordinate data and associate the standard spatial coordinate data with the government affairs data ontology model.

[0171] In one embodiment, the first determination module 940 is configured to: determine the data integrity index using the integrity index calculation formula; the integrity index calculation formula is as follows:

[0172]

[0173] where μ is the field weight, determined by expert scoring or the information gain method; D is the existence score, 1 for non-empty fields and 0 for null values; E is the rule verification score, 1 for passing the format verification, 0 for failure, 1 for passing the business rule verification, and 0.5 for failure.

[0174] Determine the data call frequency using the call heat calculation formula; the call heat calculation formula is as follows:

[0175] χ=∑(C t ×e (-λ×(T-t) )

[0176] where C is the number of accesses; t is the call occurrence event, calculated by day; T is the current time; λ is the time decay factor.

[0177] Determine the data social benefit value using the social benefit value calculation formula; the social benefit value calculation formula is as follows:

[0178] δ=β×A+(1-β)×B

[0179] where β is the weight of the objective index, and the objective index includes: the number of policy references, the coverage of people's livelihood impacts, and economic value; A is the objective index; (1-β) is the weight of the expert score; B is the expert score.

[0180] In one embodiment, the second determination module 950 is configured to: determine the weights of different dimension indicators in the multi-dimensional indicators using the entropy weight method; the entropy weight method is as follows:

[0181]

[0182] Among them, w k is the weight of the main component; F k is the contribution rate of the k-th principal component; p k is the entropy weight of the k-th principal component; F i is the contribution rate of the i-th principal component; p i is the entropy weight of the i-th principal component;

[0183] Determine the asset value index according to the weights of different dimensional indicators and multi-dimensional indicators.

[0184] In one implementation, the government affairs data asset graph construction device 900 further includes a data traceability module, and the data traceability module is configured to: use blockchain technology to record the data flow path, and when the data is called, trigger a smart contract to chain the data user's permissions and operation traces; obtain the blockchain transaction log, extract the data flow path from the blockchain transaction log, and store it in the Neo4j graph database to obtain the data lineage graph; the nodes in the data lineage graph include: data tables, files, APIs; the edge relationships include: source, derivation, influence.

[0185] In one implementation, the data traceability module is configured to: construct a graph neural network according to the data asset graph; the node features in the graph neural network include: the asset value index and sensitivity level of the data; the edge weights include: the data flow frequency and the time decay factor; input the data propagation path into the graph neural network, and use the graph neural network to identify the data propagation path and output the abnormal data propagation path.

[0186] In one implementation, the government affairs data asset graph construction device 900 further includes a data update module, and the data update module is configured to: monitor the external event source; when detecting data update, extract the keywords in the updated data and associate the keywords with the affected ontology.

[0187] In one implementation, the data update module is configured to: each department locally trains a GNN model; input the local graph into the GNN model, use the GNN model to extract the features of the local graph to obtain the local graph features; the central server receives the corresponding local graph features uploaded by each department, aggregates the local graph features corresponding to each department to generate a global graph; the central server sends the global graph to each part so that each department can update the local graph according to the global graph.

[0188] The above device is used to execute the method provided by the foregoing embodiment, and its implementation principle and technical effects are similar and will not be elaborated here.

[0189] The above-mentioned modules may be one or more integrated circuits configured to implement the above methods. For example, one or more Application Specific Integrated Circuits (ASICs), or one or more microprocessors, or one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduler code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0190] Figure 10 Schematic diagram of the computer device provided by the embodiment of the present application. This device may be integrated into a terminal device or a chip of a terminal device, and the terminal may be a computing device with data processing capabilities.

[0191] The device includes: a processor 1001, a storage medium 1002, and a bus 1003.

[0192] The storage medium 1002 stores program instructions executable by the processor 1001. When the computer device 1000 runs, the processor 1001 communicates with the storage medium 1002 through the bus 1003, and the processor 1001 executes the program instructions to execute the above method embodiments. The specific implementation manners and technical effects are similar and will not be elaborated here.

[0193] Optionally, the present invention also provides a program product, such as a computer-readable storage medium, including a program that is used to execute the above method embodiments when executed by a processor.

[0194] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces. The indirect coupling or communication connection of the devices or units may be in an electrical, mechanical or other form.

[0195] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0196] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0197] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit stored in a storage medium includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (abbreviation: ROM), random access memories (abbreviation: RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0198] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for constructing a government affairs data asset graph, characterized in that include: Acquire multi-source heterogeneous data; The multi-source heterogeneous data includes: structured data and unstructured data; the structured data includes spatial data; Preprocessing the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data; Inputting the pre-processed multi-source heterogeneous data into an ontology modeling tool, and using the ontology modeling tool to construct an ontology for the pre-processed multi-source heterogeneous data to obtain a government data ontology model; the government data ontology model includes: an ontology, ontology attributes, and relationships between ontologies; Using an asset value assessment model to determine the multi-dimensional indicators of the government data ontology model; the multi-dimensional indicators include: data integrity index, data call frequency, and data social benefit value; Using the entropy weight method, the asset value index is determined based on the multi-dimensional indicators; The government data ontology model and the asset value index are visualized to obtain a data asset map.

2. The method according to claim 1, wherein The preprocessing of the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data includes: Deduplication, standardization, and missing value processing are performed on the multi-source heterogeneous data to obtain the preprocessed multi-source heterogeneous data.

3. The method according to claim 1, wherein The ontology modeling tool is used to construct an ontology for the pre-processed multi-source heterogeneous data to obtain a government data ontology model, including: For the structured data other than the spatial data, use a Schema Mapping tool to map fields to the government data ontology model; For the unstructured data, an NLP model is used to extract the relationships between ontologies, and the relationships between the ontologies are bound to the government data ontology model; For the spatial data, the GeoJSON specification is used to convert the spatial data into standard spatial coordinate data, and the standard spatial coordinate data is associated with the government data ontology model.

4. The method according to claim 1, characterized in that: An asset value assessment model is used to determine the multi-dimensional indicators of the government data ontology model, including: The data integrity index is determined using an integrity index calculation formula; the integrity index calculation formula is as follows: Where μ is the field weight, determined by expert scoring or information gain method; D is the existence score, where a non-empty field is 1 and an empty field is 0; E is the rule verification score, where a format verification pass is 1 and a failure is 0; a business rule verification pass is 1 and a failure is 0.5; The data call frequency is determined using a call heat calculation formula; the call heat calculation formula is as follows: x=∑(C t ×e (-λ×(T-t) ) Where C is the number of visits; t is the call occurrence event, calculated in days; T is the current time; λ is the time decay factor; The social benefit value calculation formula is used to determine the social benefit value of the data; the social benefit value calculation formula is as follows: δ=β×A+(1-β)×B Among them, β is the weight of objective indicators, including: number of policy citations, coverage of people's livelihood impact, and economic value; A is the objective indicator; (1-β) is the weight of expert score; B is the expert score.

5. The method according to claim 1, characterized in that, The entropy weight method is used to determine the asset value index based on the multi-dimensional indicators, including: Determine the weights of different dimension indicators in the multi-dimensional indicators by using the entropy weight method; the entropy weight method is as follows: Among them, w k is the weight of the main component; F k is the contribution rate of the k-th main component; p k is the entropy weight of the k-th main component; F i is the contribution rate of the i-th main component; p i is the entropy weight of the i-th main component; Determine the asset value index according to the weights of the different dimension indicators and the multi-dimensional indicators.

6. The method according to claim 1, characterized in that After obtaining the data asset map, the method further includes: Use blockchain technology to record the data flow path. When the data is called, trigger a smart contract to record the data user's permissions and operation traces on the chain. Obtain the blockchain transaction log, extract the data flow path from the blockchain transaction log, and store it in the Neo4j graph database to obtain the data lineage map; the nodes in the data lineage map include: data tables, files, APIs; the edge relationships include: source, derivation, influence.

7. The method according to claim 6, characterized in that, After obtaining the data lineage map, the method further includes: Construct a graph neural network according to the data asset map; the node features in the graph neural network include: the asset value index and sensitivity level of the data; the edge weights include: the data flow frequency and time decay factor. Input the data propagation path into the graph neural network, and use the graph neural network to identify the data propagation path and output the abnormal data propagation path.

8. The method according to claim 1, wherein After obtaining the data asset map, the method further includes: Monitor the external event source; When data update is detected, extract the keywords in the updated data and associate the keywords with the affected ontology.

9. The method according to claim 8, wherein After associating the keywords with the affected ontology, the method further includes: Each department locally trains a graph neural network; Input the local map into the graph neural network, and use the graph neural network to extract features from the local map to obtain local map features; The central server receives the corresponding local map features uploaded by each department, aggregates the local map features corresponding to each department, and generates a global map; The central server distributes the global map to each part so that each department can update the local map according to the global map.

10. An apparatus for constructing a government affairs data asset graph, characterized in that, Includes: An acquisition module configured to acquire multi-source heterogeneous data; The multi-source heterogeneous data includes: structured data, unstructured data; the structured data includes spatial data; A preprocessing module configured to preprocess the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data; An ontology construction module configured to input the preprocessed multi-source heterogeneous data into an ontology modeling tool, and use the ontology modeling tool to construct an ontology for the preprocessed multi-source heterogeneous data to obtain a government affairs data ontology model; the government affairs data ontology model includes: ontology, attributes of the ontology, relationships between ontologies; A first determination module configured to use an asset value evaluation model to determine the multi-dimensional indicators of the government affairs data ontology model; the multi-dimensional indicators include: data integrity index, data call frequency, data social benefit value; A second determination module configured to use the entropy weight method to determine the asset value index according to the multi-dimensional indicators; A visualization module configured to visualize the government affairs data ontology model and the asset value index to obtain a data asset map.