A government affair data circulation management method and system based on a blockchain
By adopting a blockchain-based approach to the management of government data circulation, the problems of data anonymization ambiguity and weakened targeting in existing technologies have been solved, achieving efficient and accurate data circulation and privacy protection, and improving the adaptability and reliability of data.
Patent Information
- Application Number
- CN202511079499.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-04
AI Technical Summary
Existing data anonymization technologies rely on preset rules or human experience, resulting in blurred and weakened information after anonymization, which fails to meet high-standard guidance requirements and reduces data usability and departmental collaboration efficiency.
By using a blockchain-based approach to manage the flow of government data, requests for government data flow are obtained, the requests are parsed to determine the processing parameters, and the data is anonymized through parameterized configuration. Combined with historical flow records and data scenario coding stored on the blockchain, the flow topology and consistency are optimized to ensure that the anonymized data complies with privacy regulations and actual needs.
It improves the availability and integrity of data, enhances the accuracy and efficiency of data circulation, and optimizes the accuracy of data services and the overall circulation effect while ensuring privacy protection.
Smart Images

Figure CN120579219B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management, in particular to a government data circulation management method and system based on a blockchain. BACKGROUND
[0002] In the field of government data circulation management, with the increasing demand for data sharing among government departments, data desensitization technology is widely used to protect personal privacy and sensitive information. Desensitization processing removes or obscures identifying elements in the data to ensure that the data meets privacy regulations during circulation.
[0003] However, existing data desensitization technology mainly relies on preset rules or manual experience, resulting in blurred data information and weakened directionality after desensitization, which cannot meet the high-standard directional needs and reduces data practicality and department collaboration efficiency. SUMMARY
[0004] The present application provides a government data circulation management method and system based on a blockchain to solve the above problems.
[0005] In a first aspect, the present application provides a government data circulation management method based on a blockchain, comprising:
[0006] Obtaining a government data circulation request, analyzing the government data circulation request, and determining processing parameters;
[0007] According to the processing parameters, desensitizing the original government data to obtain desensitized data;
[0008] Analyzing the desensitized data to determine a data scenario code;
[0009] Obtaining historical circulation records stored in a blockchain; determining a circulation topology structure according to the historical circulation records and the data scenario code;
[0010] According to the circulation topology structure, determining circulation consistency, and outputting the desensitized data based on the circulation consistency.
[0011] By the scheme, the government affair data circulation request is acquired, the government affair data circulation request is analyzed, the processing parameter is determined, it is ensured that the desensitization processing can be parameterized according to the specific content of the request Configuration, avoid the problem that desensitization parameter setting is arbitrary, thereby providing input basis for the desensitization process, improving the adaptability and reliability of parameter setting. According to the processing parameter, the original government affair data is desensitized to obtain desensitized data, realizing the core of privacy protection, ensuring that the data meets the requirements of privacy regulations; at the same time, reducing the risk of excessive information blurring or detail loss, alleviating the defects of information blurring and weak directionality of desensitized data, thereby enhancing the usability and integrity of desensitized data. Analyzing the desensitized data, determining the data scene coding, eliminating the problem of lacking detailed scene evaluation mechanism, thereby improving the directionality and matching degree of data, ensuring that the desensitized data accurately serves the target demand in actual circulation. Obtain the historical circulation record stored in the block chain, provide tamper-proof traceability information. According to the historical circulation record and the data scene coding, the circulation topology structure is determined, the problem that the characteristics of the block chain cannot be effectively integrated is eliminated, thereby optimizing the accuracy of data-oriented and reducing the risk of desensitized information deviating from the actual demand. According to the circulation topology structure, the circulation consistency is determined, and based on the circulation consistency, the desensitized data is output, improving the efficiency and reliability of data circulation, eliminating the problem that desensitized information deviates from the actual demand, thereby enhancing the accuracy of data service and the overall circulation effect under the premise of ensuring privacy.
[0012] Optionally, the analyzing the government affair data circulation request and determining the processing parameter comprises:
[0013] Analyzing the government affair data circulation request, determining the data application scene and user access permission;
[0014] According to the data application scene and user access permission, determining the desensitization parameter and guiding metadata;
[0015] According to the desensitization parameter and guiding metadata, determining the processing parameter.
[0016] By the scheme, the government affair data circulation request is analyzed, the data application scene and user access permission are determined, and the desensitization processing is ensured to be combined with the real-time application scene, thereby eliminating the problems of information blurring and weak directionality; at the same time, it is ensured that the desensitization processing is dynamically adjusted according to the user access level, thereby eliminating the problem that the user permission is not deeply integrated into the desensitization decision. According to the data application scene and user access permission, the desensitization parameter and guiding metadata are determined, the desensitization process is dynamically matched with the scene demand and permission level, the problem of arbitrary desensitization parameter setting is eliminated, and the data value is improved; at the same time, it is ensured that the desensitization processing utilizes the inherent attributes of data, the problem that guiding metadata is not fully utilized is eliminated, and the guiding precision is optimized. According to the desensitization parameter and guiding metadata, the processing parameter is determined, realizing the goal of protecting privacy while improving data-oriented precision and circulation efficiency.
[0017] Optionally, determining the de-identification parameters based on the data application scenario and user access permissions includes:
[0018] Get the type of sensitive field;
[0019] Determine the initial desensitization level based on the sensitive field types and the data application scenarios;
[0020] Based on the user's access permissions, the user's security level is obtained;
[0021] Desensitization parameters are determined based on the initial desensitization level and the user security level.
[0022] This solution identifies sensitive field types, avoids misprocessing of non-sensitive fields, reduces the risk of data obfuscation, and provides basic input for setting de-identification parameters, ensuring that the de-identification operation focuses on core elements of privacy protection. Based on the sensitive field type and data application scenario, an initial de-identification level is determined, allowing the de-identification strength to be automatically adjusted for different scenarios, thus initially balancing privacy protection and data usability. Based on user access permissions, user security levels are obtained, optimizing data guidance accuracy while maintaining privacy and security. De-identification parameters are determined based on the initial de-identification level and user security levels, reducing excessive or insufficient obfuscation and improving data flow efficiency.
[0023] Optionally, obtaining the sensitive field type includes:
[0024] Analyze the government data circulation request to determine the data to be circulated;
[0025] Obtain the data field tags of the data to be circulated, match the data field tags with a preset sensitive field library, and determine the sensitive field type.
[0026] This solution analyzes government data circulation requests, identifies the data to be circulated, avoids interference from irrelevant data, and provides an accurate data foundation for sensitive field identification. It obtains the data field tags of the data to be circulated, matches these tags against a pre-defined sensitive field database, determines the sensitive field types, and ensures that de-identification decisions are based on precise field classification, thereby improving the targeting of the de-identification process.
[0027] Optionally, determining the circulation topology based on the historical circulation records and the data scenario encoding includes:
[0028] Analyze the historical circulation records to determine departmental nodes and departmental interaction relationships;
[0029] Construct a department topology graph based on the department interaction relationships and department nodes;
[0030] Based on the data scenario coding, the circulation target is determined;
[0031] Based on the departmental topology map, the distribution topology is determined according to the distribution objectives.
[0032] This solution analyzes historical data flow records to identify departmental nodes and their interactions, ensuring that these reflect the actual data flow history. This provides structured input for constructing departmental topology maps, avoiding reliance on static rules or assumptions and enhancing the traceability of flow decisions. Based on departmental interactions and nodes, a departmental topology map is constructed, ensuring the flow structure is built on dynamic historical interactions rather than a fixed template. This improves the accuracy and scalability of the departmental topology and supports dynamic optimization of flow paths. By encoding data scenarios, flow objectives are determined, eliminating guidance biases caused by static rules and improving the accuracy and scenario adaptability of data flow. Based on the departmental topology map and flow objectives, the flow topology structure is determined, optimizing data guidance accuracy and ensuring efficient and reliable connection between source and target departments, providing dynamic path support for data anonymization.
[0033] Optionally, determining the desensitization parameters based on the initial desensitization level and the user security level includes:
[0034] Based on the data application scenario, determine the fuzziness tolerance threshold;
[0035] Calculate the desensitization intensity correction coefficient based on the fuzziness tolerance threshold and the initial desensitization level;
[0036] The desensitization parameters are determined based on the desensitization intensity correction coefficient and the user confidentiality level.
[0037] This solution determines a tolerance threshold for ambiguity based on data application scenarios, reducing issues of information ambiguity and weak targeting. This prevents excessive ambiguity in anonymization parameter settings, which could lead to decreased data usability and improve data applicability in circulation. Based on the ambiguity tolerance threshold and the initial anonymization level, an anonymization intensity correction coefficient is calculated to optimize the rigidity of the initial anonymization level setting. This ensures that the anonymization parameters balance high-standard anonymization requirements with information guidance effects, providing intermediate variables for parameter determination and reducing arbitrariness in parameter setting. Finally, based on the anonymization intensity correction coefficient and user confidentiality level, the anonymization parameters are determined to prevent the anonymized information from deviating from actual guidance needs, thereby enhancing the usability of data in government decision-making.
[0038] Optionally, determining circulation consistency based on the circulation topology includes:
[0039] Based on the departmental topology map, the centrality index is determined according to the circulation objectives;
[0040] Based on the centrality index and the data scenario coding, key department nodes are determined;
[0041] Based on the processing parameters, analyze the de-identified data and determine the field retention rate of the de-identified data;
[0042] Analyze the key department nodes to determine field requirements;
[0043] Based on the retention rate of the field and the requirements of the field, determine the consistency of circulation.
[0044] This solution, based on the departmental topology map and circulation objectives, determines the centrality index, eliminating the problem of failing to use the centrality index to determine whether data conforms to the circulation structure. Based on the centrality index and data scenario coding, key departmental nodes are identified, achieving the goal of dynamically optimizing guidance accuracy in conjunction with application scenarios. Based on processing parameters, anonymized data is analyzed to determine the field retention rate, reflecting the degree of ambiguity after anonymization and addressing the defect of reduced data usability caused by ambiguous information. Key departmental nodes are analyzed to determine field requirements, fulfilling the need to optimize anonymization using guidance metadata. Based on the field retention rate and field requirements, circulation consistency is ensured, eliminating the problem of information deviating from actual guidance requirements after anonymization.
[0045] Optionally, the method further includes:
[0046] If the circulation consistency is lower than the preset circulation threshold, then based on the government data circulation request, the key department nodes are analyzed to determine the priority of field requirements.
[0047] Based on the priority of the field requirements and the consistency of circulation, generate a desensitization parameter correction instruction;
[0048] Update the processing parameters according to the desensitization parameter correction instruction.
[0049] This solution addresses the issue of data flow consistency falling below a preset threshold. Based on government data flow requests, it analyzes key department nodes to determine field priority and clarifies the minimum retention rate requirements for different fields from the target department, providing a basis for parameter correction. According to field priority and flow consistency, it generates de-identified parameter correction instructions to ensure the de-identified parameters better meet actual needs. Based on these instructions, it updates processing parameters, improving data guidance accuracy and flow efficiency while maintaining privacy protection standards.
[0050] Optionally, determining circulation consistency based on the field retention rate and the field requirements includes:
[0051] Obtain the metadata description information of the de-identified data, and analyze the data availability indicators based on the metadata description information;
[0052] Based on the department topology diagram, analyze the data flow requirements of adjacent department nodes and determine the collaboration consistency index;
[0053] Based on the field retention rate, data availability index, and collaborative consistency index, a multidimensional evaluation matrix is constructed;
[0054] Based on the multidimensional evaluation matrix, circulation consistency is determined according to the field requirements.
[0055] This solution acquires metadata descriptions of anonymized data, analyzes data availability metrics based on this information, and ensures that metadata details are effectively utilized to optimize anonymization decisions. This provides numerical data availability dimensions as a basis for constructing a multi-dimensional evaluation matrix. Based on departmental topology diagrams, it analyzes the data flow needs of adjacent departmental nodes, determines collaboration consistency metrics, and provides a collaboration dimension as a basis for constructing the multi-dimensional evaluation matrix, ensuring that data flow conforms to organizational structure dependencies. Based on field retention rates, data availability metrics, and collaboration consistency metrics, a multi-dimensional evaluation matrix is constructed, providing an overall evaluation basis for calculating flow consistency, facilitating quantitative comparison and decision-making. Based on the multi-dimensional evaluation matrix, flow consistency is determined according to field requirements, supporting dynamic adjustment of anonymization parameters to optimize government data services.
[0056] Secondly, this application provides a blockchain-based government data circulation management system, the system comprising:
[0057] The request parsing module is used to obtain government data circulation requests, parse the government data circulation requests, and determine the processing parameters;
[0058] The data processing module is used to perform desensitization processing on the original government data according to the processing parameters to obtain desensitized data;
[0059] The data analysis module is used to analyze the de-identified data and determine the data scenario encoding;
[0060] The structure determination module is used to acquire historical circulation records stored in the blockchain; and determine the circulation topology based on the historical circulation records and the data scenario encoding.
[0061] The data output module is used to determine the circulation consistency based on the circulation topology and output the de-identified data based on the circulation consistency.
[0062] Optionally, when the request parsing module parses the government data circulation request and determines the processing parameters, it is used for:
[0063] Analyze the government data circulation request to determine the data application scenario and user access permissions;
[0064] Based on the data application scenario and user access permissions, determine the de-identification parameters and guiding metadata;
[0065] Based on the desensitization parameters and the guiding metadata, the processing parameters are determined.
[0066] Optionally, when the request parsing module determines the de-identification parameters based on the data application scenario and user access permissions, it is used to:
[0067] Get the type of sensitive field;
[0068] Determine the initial desensitization level based on the sensitive field types and the data application scenarios;
[0069] Based on the user's access permissions, the user's security level is obtained;
[0070] Desensitization parameters are determined based on the initial desensitization level and the user security level.
[0071] Optionally, when the request parsing module obtains the sensitive field type, it is used for:
[0072] Analyze the government data circulation request to determine the data to be circulated;
[0073] Obtain the data field tags of the data to be circulated, match the data field tags with a preset sensitive field library, and determine the sensitive field type.
[0074] Optionally, when the structure determination module determines the circulation topology based on the historical circulation records and the data scenario encoding, it is used to:
[0075] Analyze the historical circulation records to determine departmental nodes and departmental interaction relationships;
[0076] Construct a department topology graph based on the department interaction relationships and department nodes;
[0077] Based on the data scenario coding, the circulation target is determined;
[0078] Based on the departmental topology map, the distribution topology is determined according to the distribution objectives.
[0079] Optionally, when the request parsing module determines the de-identification parameters based on the initial de-identification level and the user security level, it is used to:
[0080] Based on the data application scenario, determine the fuzziness tolerance threshold;
[0081] Calculate the desensitization intensity correction coefficient based on the fuzziness tolerance threshold and the initial desensitization level;
[0082] The desensitization parameters are determined based on the desensitization intensity correction coefficient and the user confidentiality level.
[0083] Optionally, when the data output module determines circulation consistency based on the circulation topology, it is used to:
[0084] Based on the departmental topology map, the centrality index is determined according to the circulation objectives;
[0085] Based on the centrality index and the data scenario coding, key department nodes are determined;
[0086] Based on the processing parameters, analyze the de-identified data and determine the field retention rate of the de-identified data;
[0087] Analyze the key department nodes to determine field requirements;
[0088] Based on the retention rate of the field and the requirements of the field, determine the consistency of circulation.
[0089] Optionally, the blockchain-based government data circulation management system further includes a parameter update module, used for:
[0090] If the circulation consistency is lower than the preset circulation threshold, then based on the government data circulation request, the key department nodes are analyzed to determine the priority of field requirements.
[0091] Based on the priority of the field requirements and the consistency of circulation, generate a desensitization parameter correction instruction;
[0092] Update the processing parameters according to the desensitization parameter correction instruction.
[0093] Optionally, when determining flow consistency based on the field retention rate and the field requirements, the data output module is used to:
[0094] Obtain the metadata description information of the de-identified data, and analyze the data availability indicators based on the metadata description information;
[0095] Based on the department topology diagram, analyze the data flow requirements of adjacent department nodes and determine the collaboration consistency index;
[0096] Based on the field retention rate, data availability index, and collaborative consistency index, a multidimensional evaluation matrix is constructed;
[0097] Based on the multidimensional evaluation matrix, circulation consistency is determined according to the field requirements. Attached Figure Description
[0098] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0099] Figure 1 This is a schematic diagram illustrating an application scenario provided in one embodiment of this application;
[0100] Figure 2 A flowchart illustrating a blockchain-based government data circulation management method provided in one embodiment of this application;
[0101] Figure 3 This is a schematic diagram of the structure of a blockchain-based government data circulation management system, provided as an embodiment of this application. Detailed Implementation
[0102] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0103] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0104] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0105] Existing data anonymization technologies mainly rely on preset rules or human experience, resulting in blurred and weakened information after anonymization, which fails to meet high-standard guidance requirements and reduces data usability and departmental collaboration efficiency.
[0106] Based on this, this application provides a blockchain-based method and system for managing the circulation of government data. It acquires and parses government data circulation requests, determines processing parameters, and ensures that the de-identification process can be parameterized according to the specific content of the request, avoiding arbitrary parameter settings. This provides input for the de-identification process and improves the adaptability and reliability of parameter settings. Based on the processing parameters, the original government data is de-identified to obtain de-identified data, achieving the core of privacy protection and ensuring that the data complies with privacy regulations. Simultaneously, it reduces the risk of excessive information ambiguity or loss of details, mitigating the defects of information ambiguity and weak directionality that de-identified data can easily produce, thereby enhancing the usability and integrity of the de-identified data. Analyzing the de-identified data, data scenario coding is determined, eliminating the problem of lacking a detailed scenario evaluation mechanism, thereby improving the directionality and matching degree of the data, ensuring that the de-identified data accurately serves the target needs in actual circulation. Historical circulation records stored on the blockchain are obtained, providing tamper-proof traceability information. Based on the historical circulation records and data scenario coding, the circulation topology is determined, eliminating the problem of failing to effectively integrate the characteristics of the blockchain, thereby optimizing the accuracy of data guidance and reducing the risk of de-identified information deviating from actual needs. Based on the circulation topology, circulation consistency is determined, and based on circulation consistency, anonymized data is output to improve the efficiency and reliability of data circulation, eliminate the problem of information deviating from actual guiding needs after anonymization, and thus enhance the accuracy of data services and the overall circulation effect while protecting privacy.
[0107] Figure 1 This is a schematic diagram illustrating an application scenario provided by this application, showing the application of the method provided in this application when managing the circulation of government data.
[0108] Specifically, the method provided in this application is applied to any server. The server interacts with the device of the user requesting data, obtains government data circulation requests through the user's device, parses the government data circulation requests, and determines processing parameters. Based on the processing parameters, the original government data is anonymized to obtain anonymized data. The anonymized data is analyzed to determine the data scenario code. Historical circulation records stored in the blockchain are obtained to provide tamper-proof traceability information. Based on the historical circulation records and data scenario code, the circulation topology is determined to eliminate the problem of failing to effectively integrate the characteristics of the blockchain, thereby optimizing the accuracy of data guidance and reducing the risk of anonymized information deviating from actual needs. Based on the circulation topology, circulation consistency is determined, and based on circulation consistency, anonymized data is output to improve the efficiency and reliability of data circulation, eliminate the problem of information deviating from actual guidance needs after anonymization, and thus enhance the accuracy of data services and the overall circulation effect while ensuring privacy. Specific implementation methods can be found in the following embodiments.
[0109] Figure 2This is a flowchart illustrating a blockchain-based government data circulation management method according to an embodiment of this application. The method of this embodiment can be applied to servers in the above scenarios. Figure 2 As shown, the method includes:
[0110] S201. Obtain government data circulation requests, parse government data circulation requests, and determine processing parameters;
[0111] A request for the flow of government data can be a request initiated by a government department to share government data.
[0112] The processing parameters can be the configuration parameters required for the desensitization process.
[0113] Specifically, when a user has a need for government data, the system receives the government data circulation request submitted by the user's device through the API interface; then, it parses the text fields in the government data circulation request; subsequently, based on the parsing results, it determines the desensitization level and the fields to be retained, and other processing parameters.
[0114] S202. Based on the processing parameters, the original government data is de-identified to obtain de-identified data;
[0115] Raw government data can be government data that has not undergone anonymization processing.
[0116] De-identified data can be government data that has undergone de-identification processing.
[0117] Specifically, the process involves retrieving raw government data from the government database; then, applying a desensitization algorithm to desensitize the processing parameters to determine the desensitized data: first, if the processing parameters specify a high level of fuzziness, data generalization techniques are used to blur the addresses; second, if the processing parameters specify fields to be retained, only identifiable elements are removed; finally, the desensitized data is output.
[0118] S203. Analyze the de-identified data and determine the data scenario coding;
[0119] Data scenario encoding can be a coded representation of the application scenario.
[0120] Specifically, NLP tools are used to scan the anonymized data to extract data features; then, based on manually labeled government data samples, machine learning is used to establish the correspondence between data features and application scenarios, generating a scenario coding mapping table; subsequently, based on the data features, the scenario coding mapping table is queried to determine the data scenario code.
[0121] S204. Obtain historical circulation records stored in the blockchain; determine the circulation topology based on the historical circulation records and data scenario coding;
[0122] Blockchain can be a distributed ledger technology used to store data.
[0123] Historical circulation records can be the data flow paths and modification history stored in the blockchain.
[0124] The circulation topology can be a graph of departmental dependencies in data circulation.
[0125] Specifically, the blockchain network is accessed through the blockchain node interface to query and obtain historical circulation records stored in the blockchain. Then, the historical circulation records are parsed to extract departmental nodes and circulation edges; subsequently, the topology is adjusted based on the data scenario encoding; finally, a graph database tool is used to apply path analysis algorithms to generate the circulation topology structure.
[0126] S205. Based on the circulation topology, determine circulation consistency and output de-identified data based on circulation consistency.
[0127] Circulation consistency can be defined as the degree of matching between data and circulation topology.
[0128] Specifically, by traversing the circulation topology, the requester information extracted from the government data circulation request is compared with the topology nodes to determine circulation consistency. If the circulation consistency is high, the output is allowed; if the circulation consistency is low, the output is rejected or an anomaly is recorded. Finally, based on the circulation consistency result, the final de-identified data is generated.
[0129] This solution acquires and analyzes government data circulation requests, determines processing parameters, and ensures that the anonymization process can be parameterized according to the specific content of the request. This avoids arbitrary parameter settings and provides input for the anonymization process, improving the adaptability and reliability of parameter settings. Based on the processing parameters, the original government data is anonymized to obtain anonymized data, achieving the core of privacy protection and ensuring data compliance with privacy regulations. Simultaneously, it reduces the risk of excessive information obscurity or loss of detail, mitigating the potential for information ambiguity and weak targeting in anonymized data, thereby enhancing the usability and integrity of the anonymized data. Analyzing the anonymized data determines data scenario coding, eliminating the lack of detailed scenario evaluation mechanisms, thus improving the data's targeting and matching accuracy, ensuring that the anonymized data accurately serves target needs in actual circulation. Historical circulation records stored on the blockchain provide tamper-proof traceability information. Based on historical circulation records and data scenario coding, the circulation topology is determined, eliminating the problem of ineffective integration of blockchain characteristics, thereby optimizing the accuracy of data guidance and reducing the risk of anonymized information deviating from actual needs. Based on the circulation topology, circulation consistency is determined, and based on circulation consistency, anonymized data is output to improve the efficiency and reliability of data circulation, eliminate the problem of information deviating from actual guiding needs after anonymization, and thus enhance the accuracy of data services and the overall circulation effect while protecting privacy.
[0130] In some embodiments, the government data circulation request is parsed to determine the data application scenario and user access permissions; based on the data application scenario and user access permissions, the de-identification parameters and guiding metadata are determined; based on the de-identification parameters and guiding metadata, the processing parameters are determined.
[0131] Data application scenarios can be the data usage purposes and conditions specified in the government data circulation request.
[0132] User access permissions can be determined based on the user identifier field in the government data circulation request.
[0133] The de-identification parameters can be dynamically set according to the data application scenario and user access permissions.
[0134] Guiding metadata can be an inherent attribute of government data.
[0135] Specifically, the process involves parsing government data circulation requests, querying the scenario description field in a predefined scenario coding mapping table, and converting the scenario description field value into a standardized data application scenario. Simultaneously, the user identifier field in the government data circulation request is read, and the user permission database is queried; then, user access permissions are obtained through the permission service interface. Next, based on the data application scenario and user access permissions, a rule engine built using expert experience and historical data-based statically preset de-identification parameter rules queries the configuration rule library to determine de-identification parameters and guiding metadata. Subsequently, based on the degree of fuzziness in the original data of the guiding metadata and the degree of fuzzification in the de-identification parameters, the final fuzziness strength is calculated; simultaneously, based on the directionality of the original data of the guiding metadata and the list of retained fields in the de-identification parameters, the guiding priority is determined; finally, the final fuzziness strength and guiding priority are integrated to generate processing parameters.
[0136] This solution analyzes government data circulation requests, identifies data application scenarios and user access permissions, and ensures that data anonymization is specifically tailored to real-time application scenarios, thereby eliminating information ambiguity and lack of clear direction. Simultaneously, it ensures that data anonymization is dynamically adjusted based on user access levels, thus addressing the issue of user permissions not being deeply integrated into the anonymization decision-making process. Based on data application scenarios and user access permissions, it determines anonymization parameters and guiding metadata, ensuring that the anonymization process dynamically matches scenario requirements and permission levels, eliminating arbitrary parameter settings and enhancing data value. Furthermore, it ensures that the anonymization process utilizes the inherent attributes of the data, eliminating the underutilization of guiding metadata and optimizing guidance accuracy. Based on the anonymization parameters and guiding metadata, it determines processing parameters to achieve the goal of protecting privacy while improving data guidance accuracy and circulation efficiency.
[0137] In some embodiments, the sensitive field type is obtained; the initial desensitization level is determined based on the sensitive field type and the data application scenario; the user security level is obtained based on the user's access permissions; and the desensitization parameters are determined based on the initial desensitization level and the user security level.
[0138] Sensitive field types can be the categories of fields in government data that need to be desensitized.
[0139] The initial desensitization level can be the preliminary level of desensitization intensity.
[0140] User security level can be a quantified level of a user's access permissions.
[0141] Specifically, historical data records are analyzed, the field list is traversed, and the sensitive category definitions in the predefined sensitive field mapping table are matched to extract the sensitive field types. Then, based on the sensitive field types and data application scenarios, the rule engine configuration rule base is queried to determine the initial de-identification level. Subsequently, the user identifier field in the government data circulation request is parsed, and the access level is obtained by querying the user permission database; then, the user's viewing permission is mapped to the user's security level. Furthermore, based on the initial de-identification level and the user's security level, the degree of obfuscation is determined by calling the calculation formula through the rule engine; simultaneously, based on the sensitive field types and user security levels, the retention rule base is queried to determine the field whitelist; finally, the degree of obfuscation and the field whitelist are combined to generate de-identification parameters.
[0142] This solution identifies sensitive field types, avoids misprocessing of non-sensitive fields, reduces the risk of data obfuscation, and provides basic input for setting de-identification parameters, ensuring that the de-identification operation focuses on core elements of privacy protection. Based on the sensitive field type and data application scenario, an initial de-identification level is determined, allowing the de-identification strength to be automatically adjusted for different scenarios, thus initially balancing privacy protection and data usability. Based on user access permissions, user security levels are obtained, optimizing data guidance accuracy while maintaining privacy and security. De-identification parameters are determined based on the initial de-identification level and user security levels, reducing excessive or insufficient obfuscation and improving data flow efficiency.
[0143] In some embodiments, the government data circulation request is parsed to determine the data to be circulated; the data field tags of the data to be circulated are obtained, and the data field tags are matched with a preset sensitive field library to determine the sensitive field type.
[0144] The data to be circulated can be the government data instance specified in the government data circulation request that needs to be circulated.
[0145] Data field labels can be names or descriptors that identify field attributes in the metadata of the data to be circulated.
[0146] The preset sensitive field library can be a collection of pre-stored sensitive field types. It is stored in the server and invoked when needed.
[0147] Specifically, the process involves analyzing requests for government data circulation to determine data identification information. Then, based on this identification information, data instances in the government database are queried and located to identify the data to be circulated. Next, the metadata storage of the data to be circulated is accessed to read the tag information for each field. Then, data field tags are obtained through a metadata query interface. Furthermore, a pre-defined sensitive field library is established based on the theoretical basis of pre-defining a set of sensitive field types using static rules. Then, a string matching algorithm is used to compare each data field tag with entries in the pre-defined sensitive field library. If a data field tag matches any sensitive field type in the pre-defined sensitive field library, it is marked as a successful match. Finally, sensitive field types are generated based on the matching results.
[0148] This solution analyzes government data circulation requests, identifies the data to be circulated, avoids interference from irrelevant data, and provides an accurate data foundation for sensitive field identification. It obtains the data field tags of the data to be circulated, matches these tags against a pre-defined sensitive field database, determines the sensitive field types, and ensures that de-identification decisions are based on precise field classification, thereby improving the targeting of the de-identification process.
[0149] In some embodiments, historical circulation records are parsed to determine departmental nodes and departmental interaction relationships; a departmental topology graph is constructed based on the departmental interaction relationships and departmental nodes; circulation targets are determined based on data scenario coding; and the circulation topology structure is determined based on the departmental topology graph and the circulation targets.
[0150] Department nodes can be government department entities.
[0151] Departmental interaction relationships can be defined as data flow relationships between departments.
[0152] A department topology diagram can be a graph structure consisting of department nodes and departmental interactions.
[0153] The circulation target can be the expected destination department for data circulation.
[0154] Specifically, historical circulation records are scanned to extract department identifiers as department nodes. Simultaneously, interaction events in the historical circulation records are analyzed to extract source and target department nodes, and department interaction relationships are defined based on event types. Then, the department interaction relationships are traversed, creating an edge for each relationship: the edge starts at the source department node and ends at the target department node. Several vertices and edges are then integrated to form a department topology graph. Subsequently, data scenario encoding is performed, and the encoded fields are parsed. Then, according to a predefined mapping table, the encoded fields are mapped to department identifiers to determine the circulation target. Finally, based on the department topology graph and the circulation target, a graph traversal algorithm is applied to determine the circulation topology structure.
[0155] This solution analyzes historical data flow records to identify departmental nodes and their interactions, ensuring that these reflect the actual data flow history. This provides structured input for constructing departmental topology maps, avoiding reliance on static rules or assumptions and enhancing the traceability of flow decisions. Based on departmental interactions and nodes, a departmental topology map is constructed, ensuring the flow structure is built on dynamic historical interactions rather than a fixed template. This improves the accuracy and scalability of the departmental topology and supports dynamic optimization of flow paths. By encoding data scenarios, flow objectives are determined, eliminating guidance biases caused by static rules and improving the accuracy and scenario adaptability of data flow. Based on the departmental topology map and flow objectives, the flow topology structure is determined, optimizing data guidance accuracy and ensuring efficient and reliable connection between source and target departments, providing dynamic path support for data anonymization.
[0156] In some embodiments, a fuzziness tolerance threshold is determined based on the data application scenario; a desensitization intensity correction coefficient is calculated based on the fuzziness tolerance threshold and the initial desensitization level; and desensitization parameters are determined based on the desensitization intensity correction coefficient and the user's security level.
[0157] The fuzziness tolerance threshold can be considered as the upper limit of the acceptable level of fuzziness in a data application scenario.
[0158] The desensitization intensity correction factor can be an adjustment factor used to modify the initial desensitization level to adapt to the needs of the scenario.
[0159] Specifically, based on the data application scenario, the scenario encoding field is extracted. Then, a predefined scenario-threshold mapping table is queried to determine the fuzziness tolerance threshold. Next, the relative relationship between the fuzziness tolerance threshold and the initial desensitization level is compared to calculate the desensitization intensity correction coefficient. Then, the impact of adjusting the desensitization intensity correction coefficient based on the user's security level is considered: if the user's security level is "high security level," less fuzziness is allowed, and the correction coefficient is further adjusted in the direction of reducing fuzziness; conversely, if the user's security level is "low security level," the original correction coefficient is maintained. Finally, based on the adjusted desensitization intensity correction coefficient and the user's security level, the final desensitization parameters are generated.
[0160] This solution determines a tolerance threshold for ambiguity based on data application scenarios, reducing issues of information ambiguity and weak targeting. This prevents excessive ambiguity in anonymization parameter settings, which could lead to decreased data usability and improve data applicability in circulation. Based on the ambiguity tolerance threshold and the initial anonymization level, an anonymization intensity correction coefficient is calculated to optimize the rigidity of the initial anonymization level setting. This ensures that the anonymization parameters balance high-standard anonymization requirements with information guidance effects, providing intermediate variables for parameter determination and reducing arbitrariness in parameter setting. Finally, based on the anonymization intensity correction coefficient and user confidentiality level, the anonymization parameters are determined to prevent the anonymized information from deviating from actual guidance needs, thereby enhancing the usability of data in government decision-making.
[0161] In some embodiments, based on the department topology map and according to the circulation objectives, the centrality index is determined; based on the centrality index and data scenario coding, key department nodes are determined; based on processing parameters, the anonymized data is analyzed to determine the field retention rate of the anonymized data; the key department nodes are analyzed to determine the field requirements; and based on the field retention rate and field requirements, circulation consistency is determined.
[0162] Centrality can be a numerical indicator that reflects the strength of association between nodes in a topological structure.
[0163] Key department nodes can be the core nodes in the department topology diagram that are most relevant to the current circulation scenario.
[0164] The field retention rate can be the proportion of retained fields to the total number of original fields.
[0165] Field requirements can be a set of minimum retention rates for different fields in the de-identified data required by the target department.
[0166] Specifically, the process involves traversing the nodes and edges of the department topology graph. Then, a degree centrality algorithm is used to count the number of connections between departments, and combined with the weights of the circulation targets to determine the centrality index. Next, the centrality index is compared with the data scenario encoding; based on the comparison results, key department nodes are identified. Then, based on processing parameters, the structure of the anonymized data is parsed, and the ratio of retained fields to the total number of original fields is calculated to determine the field retention rate of the anonymized data. Subsequently, a predefined department requirement table is queried, and combined with the attributes of key department nodes, the field requirements are determined. Finally, the field retention rate is compared with the field requirements; if several fields in the field requirements are retained in the anonymized data, circulation consistency is determined.
[0167] This solution, based on the departmental topology map and circulation objectives, determines the centrality index, eliminating the problem of failing to use the centrality index to determine whether data conforms to the circulation structure. Based on the centrality index and data scenario coding, key departmental nodes are identified, achieving the goal of dynamically optimizing guidance accuracy in conjunction with application scenarios. Based on processing parameters, anonymized data is analyzed to determine the field retention rate, reflecting the degree of ambiguity after anonymization and addressing the defect of reduced data usability caused by ambiguous information. Key departmental nodes are analyzed to determine field requirements, fulfilling the need to optimize anonymization using guidance metadata. Based on the field retention rate and field requirements, circulation consistency is ensured, eliminating the problem of information deviating from actual guidance requirements after anonymization.
[0168] In some embodiments, if the consistency of data flow is lower than a preset flow threshold, the key department nodes are analyzed based on the government data flow request to determine the priority of field requirements; a desensitization parameter correction instruction is generated according to the field requirement priority and the consistency of data flow; and the processing parameters are updated according to the desensitization parameter correction instruction.
[0169] The preset circulation threshold can be a predefined threshold parameter used to determine whether circulation consistency meets the minimum requirements. It is stored in the server in advance and invoked when needed.
[0170] Field requirement priority can be a quantified weight value that represents the importance of a field requirement.
[0171] The desensitization parameter modification command can be a structured operation command used to specify the updated content of the desensitization parameters.
[0172] Specifically, a circulation threshold is pre-set based on static rules and historical experience. Then, circulation consistency is compared with the pre-set threshold. If circulation consistency is lower than the pre-set threshold, a predefined department-field requirement priority mapping table is queried based on the government data circulation request. Next, based on a subset of key department nodes, a priority is assigned to each field requirement, thus generating a field requirement priority. Then, based on the field requirement priority and the insufficient information in circulation consistency, a de-identification parameter correction instruction is generated. Finally, according to the de-identification parameter correction instruction, the corresponding values in the processing parameters are modified, and the updated processing parameters are saved.
[0173] This solution addresses the issue of data flow consistency falling below a preset threshold. Based on government data flow requests, it analyzes key department nodes to determine field priority and clarifies the minimum retention rate requirements for different fields from the target department, providing a basis for parameter correction. According to field priority and flow consistency, it generates de-identified parameter correction instructions to ensure the de-identified parameters better meet actual needs. Based on these instructions, it updates processing parameters, improving data guidance accuracy and flow efficiency while maintaining privacy protection standards.
[0174] In some embodiments, metadata description information of the de-identified data is obtained, and data availability indicators are analyzed based on the metadata description information; based on the department topology diagram, the data flow requirements of adjacent department nodes are analyzed, and collaborative consistency indicators are determined; a multi-dimensional evaluation matrix is constructed based on field retention rate, data availability indicators, and collaborative consistency indicators; based on the multi-dimensional evaluation matrix, flow consistency is determined according to field requirements.
[0175] Metadata description information can be metadata attributes of de-identified data.
[0176] Data availability metrics can be a quantitative indicator that reflects the degree of ambiguity and directionality of anonymized data.
[0177] Adjacent department nodes can be pairs of departments that are directly connected in the department topology graph.
[0178] Data flow requirements can be based on field requirement priority values.
[0179] Collaboration consistency index can be a metric value used to assess the degree of collaboration and matching in data flow between departments.
[0180] A multidimensional evaluation matrix can be a structured data table that integrates three dimensions: field retention rate, data availability index, and collaborative consistency index.
[0181] Specifically, the process involves parsing the identifiers of the anonymized data, accessing the metadata database to match corresponding entries, and reading metadata descriptions containing ambiguity and directionality. Then, based on the metadata descriptions, the application scenario is matched. Subsequently, data availability metrics are determined according to a predefined rule table. Next, the department topology is traversed to identify adjacent department nodes associated with the current anonymized data. Then, the department-field requirement priority mapping table is queried to determine the data flow requirements of adjacent department nodes. Subsequently, by comparing the coverage of the current data fields with the required fields, the collaboration consistency metrics are determined. Then, the field retention rate, data availability metrics, and collaboration consistency metrics are input into the corresponding positions of the multidimensional evaluation matrix to generate the multidimensional evaluation matrix. Finally, based on the multidimensional evaluation matrix and combined with field requirements, a weighted sum is applied using predefined decision rules to determine flow consistency.
[0182] This solution acquires metadata descriptions of anonymized data, analyzes data availability metrics based on this information, and ensures that metadata details are effectively utilized to optimize anonymization decisions. This provides numerical data availability dimensions as a basis for constructing a multi-dimensional evaluation matrix. Based on departmental topology diagrams, it analyzes the data flow needs of adjacent departmental nodes, determines collaboration consistency metrics, and provides a collaboration dimension as a basis for constructing the multi-dimensional evaluation matrix, ensuring that data flow conforms to organizational structure dependencies. Based on field retention rates, data availability metrics, and collaboration consistency metrics, a multi-dimensional evaluation matrix is constructed, providing an overall evaluation basis for calculating flow consistency, facilitating quantitative comparison and decision-making. Based on the multi-dimensional evaluation matrix, flow consistency is determined according to field requirements, supporting dynamic adjustment of anonymization parameters to optimize government data services.
[0183] Figure 3 A schematic diagram of the structure of a blockchain-based government data circulation management system provided in one embodiment of this application is shown below. Figure 3 As shown, the blockchain-based government data circulation management system 300 of this embodiment includes: a request parsing module 301, a data processing module 302, a data analysis module 303, a structure determination module 304, and a data output module 305.
[0184] Request parsing module 301 is used to obtain government data circulation requests, parse the government data circulation requests, and determine processing parameters;
[0185] Data processing module 302 is used to perform desensitization processing on the original government data according to the processing parameters to obtain desensitized data;
[0186] Data analysis module 303 is used to analyze the de-identified data and determine the data scenario encoding;
[0187] The structure determination module 304 is used to obtain historical circulation records stored in the blockchain; and determine the circulation topology based on the historical circulation records and the data scenario encoding.
[0188] The data output module 305 is used to determine the circulation consistency according to the circulation topology and output the de-identified data based on the circulation consistency.
[0189] Optionally, when the request parsing module 301 parses the government data circulation request and determines the processing parameters, it is used for:
[0190] Analyze the government data circulation request to determine the data application scenario and user access permissions;
[0191] Based on the data application scenario and user access permissions, determine the de-identification parameters and guiding metadata;
[0192] Based on the desensitization parameters and the guiding metadata, the processing parameters are determined.
[0193] Optionally, when the request parsing module 301 determines the de-identification parameters based on the data application scenario and user access permissions, it is used to:
[0194] Get the type of sensitive field;
[0195] Determine the initial desensitization level based on the sensitive field types and the data application scenarios;
[0196] Based on the user's access permissions, the user's security level is obtained;
[0197] Desensitization parameters are determined based on the initial desensitization level and the user security level.
[0198] Optionally, when the request parsing module 301 obtains the sensitive field type, it is used for:
[0199] Analyze the government data circulation request to determine the data to be circulated;
[0200] Obtain the data field tags of the data to be circulated, match the data field tags with a preset sensitive field library, and determine the sensitive field type.
[0201] Optionally, when the structure determination module 304 determines the circulation topology based on the historical circulation records and the data scenario encoding, it is used to:
[0202] Analyze the historical circulation records to determine departmental nodes and departmental interaction relationships;
[0203] Construct a department topology graph based on the department interaction relationships and department nodes;
[0204] Based on the data scenario coding, the circulation target is determined;
[0205] Based on the departmental topology map, the distribution topology is determined according to the distribution objectives.
[0206] Optionally, when the request parsing module 301 determines the de-identification parameters based on the initial de-identification level and the user security level, it is used to:
[0207] Based on the data application scenario, determine the fuzziness tolerance threshold;
[0208] Calculate the desensitization intensity correction coefficient based on the fuzziness tolerance threshold and the initial desensitization level;
[0209] The desensitization parameters are determined based on the desensitization intensity correction coefficient and the user confidentiality level.
[0210] Optionally, when the data output module 305 determines the flow consistency based on the flow topology, it is used to:
[0211] Based on the departmental topology map, the centrality index is determined according to the circulation objectives;
[0212] Based on the centrality index and the data scenario coding, key department nodes are determined;
[0213] Based on the processing parameters, analyze the de-identified data and determine the field retention rate of the de-identified data;
[0214] Analyze the key department nodes to determine field requirements;
[0215] Based on the retention rate of the field and the requirements of the field, determine the consistency of circulation.
[0216] Optionally, the blockchain-based government data circulation management system further includes a parameter update module 306, used for:
[0217] If the circulation consistency is lower than the preset circulation threshold, then based on the government data circulation request, the key department nodes are analyzed to determine the priority of field requirements.
[0218] Based on the priority of the field requirements and the consistency of circulation, generate a desensitization parameter correction instruction;
[0219] Update the processing parameters according to the desensitization parameter correction instruction.
[0220] Optionally, when the data output module 305 determines circulation consistency based on the field retention rate and the field requirements, it is used to:
[0221] Obtain the metadata description information of the de-identified data, and analyze the data availability indicators based on the metadata description information;
[0222] Based on the department topology diagram, analyze the data flow requirements of adjacent department nodes and determine the collaboration consistency index;
[0223] Based on the field retention rate, data availability index, and collaborative consistency index, a multidimensional evaluation matrix is constructed;
[0224] Based on the multidimensional evaluation matrix, circulation consistency is determined according to the field requirements.
[0225] The system in this embodiment can be used to execute the methods of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
Claims
1. A blockchain-based method for managing the circulation of government data, characterized in that, include: Obtain government data circulation requests, parse the government data circulation requests, and determine the processing parameters; Based on the processing parameters, the original government data is de-identified to obtain de-identified data; Analyze the de-identified data to determine the data scenario encoding; Obtain historical circulation records stored on the blockchain; Based on the historical circulation records and the data scenario encoding, the circulation topology is determined, including: Analyze the historical circulation records to determine departmental nodes and departmental interaction relationships; Construct a department topology graph based on the department interaction relationships and department nodes; Based on the data scenario coding, the circulation target is determined; Based on the departmental topology map, the distribution topology structure is determined according to the distribution objectives; Based on the circulation topology, circulation consistency is determined, and based on the circulation consistency, the de-identified data is output; Determining circulation consistency based on the circulation topology includes: Based on the departmental topology map, the centrality index is determined according to the circulation objectives; Based on the centrality index and the data scenario coding, key department nodes are determined; Based on the processing parameters, analyze the de-identified data and determine the field retention rate of the de-identified data; Analyze the key department nodes to determine field requirements; Based on the field retention rate and the field requirements, determine circulation consistency, including: Obtain the metadata description information of the de-identified data, and analyze the data availability indicators based on the metadata description information; Based on the department topology diagram, analyze the data flow requirements of adjacent department nodes and determine the collaboration consistency index; Based on the field retention rate, data availability index, and collaborative consistency index, a multidimensional evaluation matrix is constructed; Based on the multidimensional evaluation matrix, circulation consistency is determined according to the field requirements.
2. The method according to claim 1, characterized in that, The process of parsing the government data flow request and determining the processing parameters includes: Analyze the government data circulation request to determine the data application scenario and user access permissions; Based on the data application scenario and user access permissions, determine the de-identification parameters and guiding metadata; Based on the desensitization parameters and the guiding metadata, the processing parameters are determined.
3. The method according to claim 2, characterized in that, The step of determining the de-identification parameters based on the data application scenario and user access permissions includes: Get the type of sensitive field; Determine the initial desensitization level based on the sensitive field types and the data application scenarios; Based on the user's access permissions, the user's security level is obtained; Desensitization parameters are determined based on the initial desensitization level and the user security level.
4. The method according to claim 3, characterized in that, The method for obtaining sensitive field types includes: Analyze the government data circulation request to determine the data to be circulated; Obtain the data field tags of the data to be circulated, match the data field tags with a preset sensitive field library, and determine the sensitive field type.
5. The method according to claim 3, characterized in that, The step of determining the desensitization parameters based on the initial desensitization level and the user security level includes: Based on the data application scenario, determine the fuzziness tolerance threshold; Calculate the desensitization intensity correction coefficient based on the fuzziness tolerance threshold and the initial desensitization level; The desensitization parameters are determined based on the desensitization intensity correction coefficient and the user confidentiality level.
6. The method according to claim 5, characterized in that, The method further includes: If the circulation consistency is lower than the preset circulation threshold, then based on the government data circulation request, the key department nodes are analyzed to determine the priority of field requirements. Based on the priority of the field requirements and the consistency of circulation, generate a desensitization parameter correction instruction; Update the processing parameters according to the desensitization parameter correction instruction.
7. A blockchain-based government data circulation management system, characterized in that, The method applied to any one of claims 1-6 includes: The request parsing module is used to obtain government data circulation requests, parse the government data circulation requests, and determine the processing parameters; The data processing module is used to perform desensitization processing on the original government data according to the processing parameters to obtain desensitized data; The data analysis module is used to analyze the de-identified data and determine the data scenario encoding; The structure determination module is used to acquire historical circulation records stored in the blockchain; and determine the circulation topology based on the historical circulation records and the data scenario encoding. The data output module is used to determine the circulation consistency based on the circulation topology and output the de-identified data based on the circulation consistency.
Citation Information
Patent Citations
Government affair data sharing and exchanging system based on data lake
CN114756622A
Government affair service affair information storage and circulation method and system
CN118626484A