Data de-identification privacy protection method and system
Through the sensitivity evaluation model, the sensitivity of data fields is divided into dynamic generation strategies for scenario requirements, and the multi-layer trusted execution architecture is combined with the multi-layer trusted execution architecture for layered processing, which solves the problem that data de-identification in the existing technology is difficult to balance privacy protection and data availability, and achieves the effect of dynamic balance.
Patent Information
- Application Number
- CN202510534091.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing technology lacks dynamic strategy adjustment capabilities and a complete execution architecture in data de-identification, making it difficult to balance privacy protection and data availability.
The sensitivity of data fields is divided by the preset sensitivity evaluation model, and a de-identification strategy is dynamically generated in combination with data usage scenarios and participant needs. Build a multi-layer trusted execution architecture, including lightweight layers, enhancement layers and security layers, and perform layered processing of different sensitive fields.
A dynamic balance between privacy protection and data availability is achieved, ensuring that data maintains sufficient availability while meeting privacy security requirements.
Smart Images

Figure CN120068157A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a privacy protection method and system for data de-identification. Background Art
[0002] With the rapid development of big data technology, data has become an important resource for promoting social progress and economic growth. However, the wide use of data also brings serious risks of privacy leakage. As an important privacy protection technology, data de-identification reduces the risk of data being re-identified by removing or replacing direct identifiers in the data, thereby supporting the legal use and sharing of data while protecting personal privacy.
[0003] Traditional data de-identification methods usually adopt a single technical means (such as data masking, generalization, or noise addition), which are difficult to cope with complex and changing actual application scenarios. For example, the sensitivity differences of different data fields are relatively large, and unified processing may lead to insufficient protection of low-sensitivity fields or over-processing of high-sensitivity fields, affecting the usability of data and the privacy protection effect. In addition, the diversity of data usage scenarios and the requirements of participants also pose higher requirements for the flexibility and adaptability of de-identification strategies. Summary of the Invention
[0004] This application provides a privacy protection method and system for data de-identification, which is used to solve the technical problem that the existing technology lacks the ability of dynamic policy adjustment and a perfect execution architecture, and it is difficult to balance privacy protection and data availability.
[0005] In the first aspect of this application, a privacy protection method for data de-identification is provided. The method includes: receiving original data, and based on a preset sensitivity assessment model, dividing the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields; obtaining data usage scenarios and participant requirement information, and performing policy matching through a dynamic policy engine to generate a de-identification execution policy; constructing a multi-layer trusted execution architecture, where the multi-layer trusted execution architecture includes a lightweight layer, an enhanced layer, and a security layer; based on the multi-layer trusted execution architecture, and in accordance with the de-identification execution policy, performing hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and outputting an available data set that meets privacy and security requirements.
[0006] In a second aspect of the present application, a privacy protection system for data de-identification is provided. The system includes: a sensitivity division module, which is configured to receive original data and divide the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity assessment model; an execution policy matching module, which is configured to obtain data usage scenarios and participant requirement information, and perform policy matching through a dynamic policy engine to generate a de-identification execution policy; a trusted execution architecture construction module, which is configured to construct a multi-layer trusted execution architecture, and the multi-layer trusted execution architecture includes a lightweight layer, an enhanced layer, and a security layer; a hierarchical de-identification processing module, which is configured to perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on the multi-layer trusted execution architecture according to the de-identification execution policy, and output an available data set that meets privacy and security requirements.
[0007] One or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0008] A privacy protection method and system for data de-identification provided in the present application relate to the technical field of data processing. The data is divided into low-, medium-, and high-sensitivity fields through a sensitivity assessment model, a de-identification policy is generated in combination with scenarios and requirements, and a multi-layer architecture of a lightweight layer, an enhanced layer, and a security layer is constructed to perform hierarchical processing on the fields, and an available data set that meets privacy and security requirements is output. This solves the technical problem in the prior art that there is a lack of dynamic policy adjustment ability and a perfect execution architecture, and it is difficult to balance privacy protection and data availability, and realizes the technical effect of dynamically balancing privacy protection and data availability through dynamic policy adjustment and a multi-layer trusted execution architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0010] Figure 1 It is a schematic flowchart of a privacy protection method for data de-identification provided by an embodiment of the present application;
[0011] Figure 2 It is a schematic structural diagram of a privacy protection system for data de-identification provided by an embodiment of the present application.
[0012] Explanation of the reference numerals: sensitivity classification module 11, execution strategy matching module 12, trusted execution architecture building module 13, hierarchical de-identification processing module 14. DETAILED DESCRIPTION
[0013] The present application provides a privacy protection method and system for data de-identification, which is used to solve the technical problem that the existing technology lacks dynamic policy adjustment capabilities and a complete execution architecture, and it is difficult to balance privacy protection and data availability.
[0014] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0015] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products, or devices.
[0016] Embodiment 1, as Figure 1 As shown, the present application provides a privacy protection method for data de-identification, the method comprising:
[0017] P10: Receive original data, and divide the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity assessment model.
[0018] Furthermore, the embodiment of the present application further includes step P10a, and step P10a further includes:
[0019] P11a: Construct an annotated dataset, which includes field feature vectors and field sensitivity level labels; P12a: Based on the annotated dataset, a random forest classifier is used for training to generate an initial sensitivity assessment model; P13a: Fuzzy field samples are constructed by adding small perturbations; P14a: Based on the fuzzy field samples, adversarial training is performed on the initial sensitivity assessment model to generate the sensitivity assessment model.
[0020] It should be understood that after receiving the original data first, it is necessary to divide the data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity assessment model. The core of this step lies in accurately classifying data fields through the data sensitivity assessment model to ensure that subsequent de-identification processing can adopt differential privacy protection strategies for fields with different sensitivity levels.
[0021] Among them, in order to construct an efficient and robust sensitivity assessment model, a labeled dataset is first constructed. The labeled dataset includes field feature vectors and field sensitivity level labels. The field feature vector is a multi-dimensional feature extracted from the original data to describe the field attributes, such as field type (e.g., numeric, text), field length, field value distribution, etc.; the field sensitivity level label is an artificial annotation of the field sensitivity according to business rules or expert experience. For example, the identity ID is marked as a high-sensitivity field, the age is marked as a medium-sensitivity field, and the gender is marked as a low-sensitivity field.
[0022] Next, based on the labeled dataset, a random forest classifier is used for training to generate an initial sensitivity assessment model. The random forest classifier is an ensemble learning method. By constructing multiple decision trees and integrating their prediction results, it can effectively process high-dimensional feature data and reduce the risk of overfitting. During the training process, through bootstrap sampling, multiple sample subsets are randomly and repeatedly drawn from the labeled dataset with replacement, and a decision tree is independently trained on each subset. Finally, the classification results of all decision trees are aggregated through a majority voting mechanism to obtain the final classification result. The advantages of random forests include the ability to handle high-dimensional data, tolerance to missing values and outliers, and the ability to evaluate feature importance, which makes them very suitable for training sensitivity assessment models.
[0023] Next, in order to improve the robustness of the model, by adding small perturbations, fuzzy field samples are constructed. Fuzzy field samples refer to slightly adjusting the original field feature vector (for example, adding random noise to numeric fields and replacing text fields with synonyms) to simulate the fuzziness and uncertainty of field features in the actual scenario. This kind of perturbation can simulate the uncertainty of data in the real scenario and help the model learn a wider range of feature patterns.
[0024] Further, the initial sensitivity assessment model is adversarially trained based on the fuzzy field samples to generate a final sensitivity assessment model. Adversarial training is a technique for improving the generalization ability of a model by introducing adversarial samples (i.e., fuzzy field samples). Its core idea is to enable the model to continuously learn how to distinguish normal samples from adversarial samples during training, thereby enhancing the stability and accuracy of the model in practical applications. During training, the model not only learns from the normal samples in the labeled dataset but also from the perturbed fuzzy field samples. In this way, the model can better adapt to minor changes in the input data, thus improving the accuracy and stability of sensitivity assessment.
[0025] Through the above steps, the generated sensitivity assessment model can more accurately classify the sensitivity of data fields, providing a reliable basis for subsequent data de-identification processing.
[0026] P20: Obtain data usage scenarios and information on the requirements of involved parties, and perform policy matching through a dynamic policy engine to generate a de-identification execution policy.
[0027] Further, step P20 of the embodiment of the present application further includes:
[0028] P21: Invoke a pre-set policy template library according to the data usage scenario to match a target policy template, where the policy template library contains de-identification rules for different fields; P22: Obtain information on the requirements of involved parties through a smart contract interface, and the information on the requirements of involved parties includes privacy budget and minimum data precision; P23: Referring to the information on the requirements of involved parties, dynamically adjust the target policy template based on a multi-party voting mechanism to generate a de-identification execution policy, where the voting weights are allocated according to the historical compliance rates of involved parties, and the de-identification execution policy includes pseudonymization intensity, generalization granularity, and noise volume.
[0029] Optionally, according to specific data usage scenarios and the privacy requirements of involved parties, perform policy matching through a dynamic policy engine to generate an adapted de-identification policy to ensure that the data remains sufficiently usable while meeting privacy protection requirements.
[0030] Specifically, first, invoke a pre-set policy template library according to the data usage scenario to match a target policy template. The policy template library is a collection containing de-identification rules for different fields, covering various data types and application scenarios. For example, the de-identification rules in the medical field may focus on protecting patient privacy, while those in the financial field pay more attention to the accuracy of transaction data. By invoking the policy template library, the system can quickly match the target policy template most suitable for the current data usage scenario, providing a basic framework for subsequent policy adjustment.
[0031] Next, obtain the participating party's requirement information through the smart contract interface. The participating party's requirement information includes the privacy budget and the minimum data accuracy. The privacy budget refers to the resource cost that the participating party is willing to pay for data privacy protection, such as the amount of differential privacy noise that is allowed to be added; the minimum data accuracy refers to the minimum requirement of the participating party for data availability, such as the granularity after data generalization cannot be lower than a certain threshold. The smart contract interface is an automated protocol based on blockchain technology, which can ensure the transparency and immutability of the participating party's requirement information, thus providing a trustworthy input for policy generation.
[0032] Furthermore, with reference to the participating party's requirement information, dynamically adjust the target policy template based on the multi-party voting mechanism to generate a de-identified execution policy. The multi-party voting mechanism is a distributed decision-making method, and the final policy adjustment direction is determined by the joint voting of multiple participating parties. The voting weight is allocated according to the historical compliance rate of the participating party, that is, the higher the historical compliance rate of the participating party, the greater its voting weight, thus ensuring the fairness and reasonableness of policy adjustment. The generated de-identified execution policy includes key parameters such as the pseudonymization strength, generalization granularity, and noise amount. Among them, the pseudonymization strength determines the degree of field replacement, the generalization granularity controls the accuracy of the data, and the noise amount is used to achieve differential privacy protection.
[0033] Through the above steps, the system can dynamically generate a de-identified execution policy with strong adaptability and good privacy protection effect according to different data usage scenarios and participating party requirements.
[0034] P30: Construct a multi-layer trusted execution architecture, which includes a lightweight layer, an enhanced layer, and a security layer.
[0035] Furthermore, step P30 of the embodiment of the present application further includes:
[0036] P31: Deploy the lightweight layer in a general security environment to perform data masking of pseudonymization and regular expression matching; P32: Deploy the enhanced layer in a trusted execution environment, and integrate a differential privacy noise addition module and an attribute generalization engine; P33: Deploy the security layer in a hardware-level security area, run a feature extraction model, and perform an identifier deletion operation.
[0037] Specifically, construct a multi-layer trusted execution architecture, which includes a lightweight layer, an enhanced layer, and a security layer. Through the design of the hierarchical architecture, de-identified operations with different complexities and security levels are allocated to the corresponding execution environments, thus optimizing system performance and resource utilization while ensuring the privacy protection effect.
[0038] Exemplarily, the lightweight layer is deployed in a general security environment and is mainly responsible for performing data masking operations such as pseudonymization and regular expression matching. The lightweight layer mainly processes the de-identification operations of low-sensitivity fields, and is characterized by low computational complexity and fast execution speed. Pseudonymization is a de-identification technique that replaces the original identifier with a random or pseudo-random value, such as replacing a name with a random string; regular expression matching is used to identify and mask sensitive information in a specific format, such as a card number or a phone number. Since the lightweight layer is deployed in a general security environment, it is applicable to scenarios with relatively high performance requirements but relatively low security requirements.
[0039] Next, the enhanced layer is deployed in a trusted execution environment, integrating a differential privacy noise addition module and an attribute generalization engine. The enhanced layer mainly processes the de-identification operations of medium-sensitivity fields, and is characterized by taking into account both security and data availability. The differential privacy noise addition module ensures the privacy of individual data in the statistical results by adding random noise to the data; the attribute generalization engine further reduces the identification risk of the data by replacing specific values with broader categories (for example, generalizing the age "25 years old" to "20 - 30 years old"). The trusted execution environment (TEE) is a hardware-level security isolation environment that can ensure that the operations in the enhanced layer are not interfered by external attacks during runtime.
[0040] Furthermore, the security layer is deployed in a hardware-level security area, running a feature extraction model and implementing an identifier deletion operation. The security layer mainly processes the de-identification operations of high-sensitivity fields, and is characterized by the highest security level, but also relatively high computational complexity. The feature extraction model is used to extract key features from the data while removing identifiers that may disclose individual identities; the identifier deletion operation directly deletes or permanently masks high-sensitivity fields, such as deleting an identity ID or a social security number. The hardware-level security area is a higher-level security protection mechanism than the trusted execution environment, usually implemented based on dedicated hardware (such as a security chip), and can provide the highest level of security guarantee for the security layer.
[0041] Through the above hierarchical deployment, the multi-layer trusted execution architecture can flexibly process data with different sensitivity levels in different security environments, taking into account both privacy protection and data availability, providing comprehensive technical support for data de-identification.
[0042] P40: Based on the multi-layer trusted execution architecture, according to the de-identification execution strategy, the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields are processed hierarchically to output an available data set that meets the privacy and security requirements.
[0043] It should be understood that, based on the multi-layer trusted execution architecture and in accordance with the de-identification execution policy, different levels of sensitive fields are processed in layers to output an available data set that meets the privacy and security requirements. The core of this step lies in the collaborative action of the multi-layer architecture and the execution policy to perform differential de-identification operations on fields with different sensitivity levels, thereby maximizing data availability while protecting data privacy.
[0044] Among them, the processing of low-sensitive fields is completed in the lightweight layer. According to the pseudonymization intensity in the de-identification execution policy, pseudonymization operations are performed on low-sensitive fields. For example, the user nickname is replaced with a random string. At the same time, through regular expression matching technology, sensitive information fragments that may exist (such as the username part in the email address) are identified and masked. Since the lightweight layer is deployed in a general security environment, it has a fast processing speed and low resource consumption, and is suitable for scenarios with high performance requirements.
[0045] The processing of medium-sensitive fields is completed in the enhanced layer. According to the generalization granularity and noise amount in the de-identification execution policy, the system calls the attribute generalization engine to perform generalization operations on medium-sensitive fields. For example, the specific age value is replaced with an age range (such as "20 - 30 years old"). At the same time, the differential privacy noise addition module adds an appropriate amount of random noise to the data to ensure the protection of individual privacy in the statistical results. The enhanced layer is deployed in a trusted execution environment (TEE), which can effectively prevent external attacks and ensure the security and reliability of the processing process.
[0046] The processing of high-sensitive fields is completed in the security layer. According to the identifier deletion requirement in the de-identification execution policy, the system runs a feature extraction model to extract key features from the data and delete or permanently mask high-sensitive fields (such as identity ID, social security number, etc.). The security layer is deployed in a hardware-level security area, providing the highest level of security guarantee based on dedicated hardware (such as a security chip) to ensure that the processing process of high-sensitive fields is not interfered by any external factors.
[0047] After the layered processing, the system integrates the processing results of low-, medium-, and high-sensitive fields to generate an available data set that meets the privacy and security requirements. While protecting privacy, this data set still retains sufficient data availability.
[0048] Furthermore, the embodiment of the present application further includes step P50, and step P50 further includes:
[0049] P51: Evaluate the privacy risk value of the available dataset through a re-identification attack simulator. If the privacy risk value exceeds the risk threshold, trigger the dynamic policy engine to upgrade the de-identification strength and generate an enhanced de-identification execution policy; P52: Based on the enhanced de-identification execution policy, perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and output an updated available dataset.
[0050] Optionally, the embodiment of the present application further evaluates and optimizes the privacy risk of the available dataset through a re-identification attack simulator and a dynamic policy engine to ensure the continuous effectiveness of data privacy protection. The core of this step lies in timely discovering and fixing potential privacy leakage risks in the dataset through simulated attacks and dynamic policy adjustments, thereby enhancing data security.
[0051] Specifically, first, evaluate the privacy risk value of the available dataset through a re-identification attack simulator. A re-identification attack simulator is a tool used to detect whether a de-identified dataset can be re-associated with the original individual, and it can simulate the process of an attacker attempting to re-identify an individual's identity through de-identified data. Its working principle is to calculate the probability of an individual in the dataset being re-identified, that is, the privacy risk value, by analyzing the feature distribution and correlation in the dataset. If the privacy risk value exceeds the preset risk threshold, it indicates that the current de-identification strength is insufficient to effectively protect data privacy. At this time, the system triggers the dynamic policy engine to upgrade the de-identification execution policy and generate an enhanced de-identification execution policy. The dynamic policy engine further enhances the de-identification strength by adjusting key parameters such as the pseudonymization strength, generalization granularity, and noise volume to cope with potential privacy leakage risks.
[0052] Furthermore, based on the enhanced de-identification execution policy, re-perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and output an updated available dataset. Exemplarily, the low-sensitivity fields further reduce the recognition risk in the lightweight layer by enhancing the pseudonymization strength; the medium-sensitivity fields increase the data privacy protection level in the enhanced layer by increasing the noise volume and expanding the generalization granularity; the high-sensitivity fields ensure that high-sensitivity information cannot be re-identified in the security layer through more rigorous feature extraction and identifier deletion operations. The updated available dataset is re-evaluated by the re-identification attack simulator to ensure that its privacy risk value is lower than the risk threshold, thus meeting higher privacy and security requirements.
[0053] By introducing step P50, this solution can effectively address the privacy risk after de-identification and ensure the security and availability of data throughout its life cycle.
[0054] Furthermore, step P51 of the embodiment of the present application further includes:
[0055] P51-1: The re-identification attack simulator includes a background knowledge attack model and a re-identification probability calculation formula; P51-2: Associate the publicly available database with the quasi-identifiers of the de-identified data in the available dataset to construct the background knowledge attack model; P51-3: Through the background knowledge attack model, calculate the number of unique combinations of quasi-identifiers in the de-identified data, where the number of unique combinations is the number of combinations of quasi-identifiers that can uniquely identify an individual; P51-4: According to the number of unique combinations and in combination with the re-identification probability calculation formula, calculate the privacy risk value of the available dataset.
[0056] The re-identification probability calculation formula is: ; where is the re-identification probability, N is the number of unique combinations, δ is the noise impact factor, δ = 1 + ε, and ε is the differential privacy noise intensity.
[0057] It should be understood that the specific process of privacy risk assessment can be further refined. By using the re-identification attack simulator to evaluate the de-identified available dataset, the privacy security of the data can be ensured.
[0058] Specifically, the re-identification attack simulator includes a background knowledge attack model and a re-identification probability calculation formula. The background knowledge attack model is used to simulate the process of an attacker using information in the publicly available database for correlation analysis with the de-identified data; the re-identification probability calculation formula is used to quantify the probability of an individual being re-identified, that is, the privacy risk value.
[0059] Among them, the publicly available database can be associated with the quasi-identifiers of the de-identified data in the available dataset to construct the background knowledge attack model. Quasi-identifiers refer to fields that can identify an individual when combined with other data sources, such as age, gender, postal code, etc. By matching the quasi-identifiers in the publicly available database with the de-identified data, the background knowledge attack model can simulate the process of an attacker inferring an individual's identity using external information.
[0060] Furthermore, through the background knowledge attack model, calculate the number of unique combinations of quasi-identifiers in the de-identified data. The number of unique combinations is the number of combinations of quasi-identifiers that can uniquely identify an individual. For example, in a certain dataset, if the combination "age = 25 years old, gender = male, postal code = 100001" corresponds to only one individual, then this combination is a unique combination. The more the number of unique combinations, the higher the privacy risk of the dataset.
[0061] Next, based on the unique combination numbers and combined with the re-identification probability calculation formula, the simulator calculates the privacy risk value of the available dataset. By introducing a noise impact factor, this formula takes into account the inhibitory effect of differential privacy noise on the re-identification probability. The differential privacy noise intensity ε is a key parameter for controlling the degree of privacy protection. The smaller ε is, the stronger the privacy protection, but the availability of the data may be reduced. By calculating the privacy risk value, the system can determine whether the current de-identification strategy is effective enough. If the privacy risk value exceeds the preset risk threshold, the dynamic policy engine is triggered to upgrade the de-identification intensity.
[0062] Through the above steps, the embodiments of the present application can effectively evaluate the privacy risk of de-identified data and dynamically adjust the de-identification strategy according to the risk assessment results to ensure the best balance between privacy protection and availability.
[0063] Further, step P51-2 of the embodiments of the present application further includes:
[0064] P51-21: Associate the public database with the available dataset, obtain public data and de-identified data, extract quasi-identifier fields based on the public data and de-identified data, and perform normalization processing; P51-22: Based on the quasi-identifier fields, extract quasi-identifiers, encode categorical variables, discretize continuous variables, and identify highly correlated combinations; P51-23: According to the highly correlated combinations, construct a training dataset, and through the training dataset, train and generate the background knowledge attack model.
[0065] Specifically, the embodiments of the present application further refine the construction process of the background knowledge attack model. By associating the public database with the available dataset, extracting and processing quasi-identifier fields, and finally training and generating the background knowledge attack model. The core of this step is to construct a high-quality training dataset through technical means such as data normalization, variable encoding, and discretization, so as to ensure that the background knowledge attack model can accurately simulate the behavior of attackers and provide a reliable basis for privacy risk assessment.
[0066] First, associate the public database with the de-identified available dataset to obtain data samples of both. Based on these data, extract quasi-identifier fields, which may be used by attackers to re-identify individuals. Quasi-identifier fields refer to fields that can identify individuals when combined with other data sources, such as age, gender, zip code, etc. To ensure data consistency and comparability, perform normalization processing on the extracted quasi-identifier fields, such as converting age to a unified range or standardizing text fields.
[0067] Next, after extracting the quasi-identifier fields, encode the categorical variables, discretize the continuous variables, and identify highly correlated combinations. For example, convert categorical variables (such as gender, occupation, etc.) into numerical form through one-hot encoding or label encoding; discretize continuous variables (such as age, income, etc.) through binning or clustering methods. Through these preprocessing steps, identify the quasi-identifier fields with highly correlated combinations. For example, the combination of "age = 25 years old, gender = male, postal code = 100001" may have a relatively high correlation. Identify these highly correlated combinations through statistical analysis or machine learning algorithms to provide key features for subsequent model training.
[0068] Furthermore, according to the highly correlated combinations, construct a training dataset and use the training dataset to train the background knowledge attack model. The training dataset includes the quasi-identifier fields in the public data and the quasi-identifier fields in the de-identified data, as well as the matching relationships between them. Machine learning algorithms (such as decision trees, random forests, or neural networks) can be used to train in combination with the training dataset to learn the association rules between the quasi-identifier fields, thereby simulating the process of an attacker re-identifying individuals using background knowledge and generating a background knowledge attack model.
[0069] Through the above steps, effectively generate a background knowledge attack model. This model can accurately simulate the behavior of attackers, provide a scientific basis for privacy risk assessment, thereby supporting the optimization and adjustment of the dynamic policy engine, and further improving the level of data privacy protection.
[0070] Furthermore, the embodiment of the present application further includes step P60, and step P60 further includes:
[0071] P61: After performing hierarchical de-identification locally, use the processed local data for model training and upload the model gradients; P62: The dynamic policy engine adjusts the differential privacy noise intensity distribution of each participant according to the global model accuracy; P63: Detect anomalies based on the gradient value distribution, evaluate the leakage risk, and automatically trigger the security layer processing when the leakage risk exceeds the risk threshold.
[0072] In a possible embodiment of the present application, the embodiment of the present application further expands the application scenario of data privacy protection, and realizes privacy protection and model training in a distributed environment through the combination of the federated learning framework and differential privacy technology. The core of this step lies in effectively preventing the risk of data leakage while ensuring the model accuracy through dynamically adjusting the differential privacy noise intensity and the anomaly detection mechanism, and ensuring the comprehensiveness and dynamics of data privacy protection.
[0073] First, after performing hierarchical de-identification locally, each participating party uses the processed local data for model training and uploads the model gradients. Exemplarily, each participating party performs hierarchical de-identification on the data locally to ensure that the data has been processed for privacy protection before leaving the local environment. The processed local data is used to train the local model, and the model gradients are uploaded to the central server through the federated learning framework instead of uploading the original data, thereby further reducing the risk of privacy leakage.
[0074] Next, the dynamic policy engine adjusts the differential privacy noise intensity distribution of each participating party according to the global model accuracy. The dynamic policy engine dynamically adjusts the differential privacy noise intensity added by each participating party during the gradient upload process by monitoring the accuracy change of the global model. For example, when the global model accuracy is low, the noise intensity is appropriately reduced to improve the model performance; when the global model accuracy is high, the noise intensity is increased to enhance privacy protection. This dynamic adjustment mechanism can achieve a balance between model accuracy and privacy protection, ensuring the efficiency and security of the federated learning framework.
[0075] Furthermore, anomalies are detected based on the gradient value distribution, and the leakage risk is evaluated. When the leakage risk exceeds the risk threshold, the security layer processing is automatically triggered. Exemplarily, by analyzing the distribution of the uploaded gradient values, whether there are abnormal patterns is detected. For example, some gradient values may imply certain characteristics of the original data, thus increasing the risk of privacy leakage. Through statistical analysis or machine learning algorithms, the abnormal conditions of the gradient value distribution are detected, and the potential leakage risk is evaluated. When the leakage risk exceeds the preset risk threshold, the dynamic policy engine automatically triggers the security layer processing, such as increasing the noise intensity, restricting the gradient upload frequency, or isolating the abnormal participating party, to effectively contain the risk of privacy leakage. The security layer is deployed in the hardware-level security area and can execute more stringent privacy protection measures, such as feature extraction, identifier deletion, or adding stronger noise.
[0076] By introducing step P60, the embodiments of the present application can dynamically adjust the privacy protection policy in distributed model training, ensuring privacy security while taking into account model performance and data availability. This step is not only applicable to the local data processing scenario but also can be extended to the collaborative learning scenarios across institutions and regions, with broad application value.
[0077] In summary, the embodiments of the present application at least have the following technical effects:
[0078] This application first divides the original data into low, medium, and high-sensitivity fields based on a sensitivity assessment model; secondly, combines the data usage scenarios and the requirements of the involved parties, and generates a de-identification execution policy through a dynamic policy engine; then, constructs a multi-layer trusted execution architecture including a lightweight layer, an enhanced layer, and a security layer; finally, performs hierarchical processing on different sensitivity fields according to the policy, and outputs an available data set that meets the privacy and security requirements.
[0079] It achieves the technical effect of dynamically balancing privacy protection and data availability through precise sensitivity assessment, flexible policy matching, and a multi-layer trusted execution architecture.
[0080] Embodiment 2, based on the same inventive concept as a privacy protection method for data de-identification in the foregoing embodiment, as Figure 2 shown, this application provides a privacy protection system for data de-identification. The system in the embodiment of this application and the method embodiment are based on the same inventive concept. Among them, the system includes:
[0081] A sensitivity division module 11, which is used to receive the original data and divide the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity assessment model.
[0082] An execution policy matching module 12, which is used to obtain data usage scenarios and the requirements information of the involved parties, and perform policy matching through a dynamic policy engine to generate a de-identification execution policy.
[0083] A trusted execution architecture construction module 13, which is used to construct a multi-layer trusted execution architecture, and the multi-layer trusted execution architecture includes a lightweight layer, an enhanced layer, and a security layer.
[0084] A hierarchical de-identification processing module 14, which is used to perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on the multi-layer trusted execution architecture according to the de-identification execution policy, and output an available data set that meets the privacy and security requirements.
[0085] Further, the sensitivity division module 11 is also used to perform the following steps:
[0086] Construct an annotated data set, where the annotated data set includes field feature vectors and field sensitivity level labels; based on the annotated data set, use a random forest classifier for training to generate an initial sensitivity assessment model; by adding small perturbations, construct fuzzy field samples; based on the fuzzy field samples, perform adversarial training on the initial sensitivity assessment model to generate the sensitivity assessment model.
[0087] Further, the execution policy matching module 12 is further configured to perform the following steps:
[0088] Call a pre-set policy template library according to the data usage scenario to match the target policy template, where the policy template library contains de-identification rules for different fields; obtain the participating party requirement information through the smart contract interface, and the participating party requirement information includes privacy budget and minimum data accuracy; with reference to the participating party requirement information, dynamically adjust the number of the target policy templates based on a multi-party voting mechanism to generate a de-identification execution policy, where the voting weight is allocated according to the historical compliance rate of the participating parties, and the de-identification execution policy includes pseudonymization intensity, generalization granularity, and noise volume.
[0089] Further, the trusted execution architecture construction module 13 is further configured to perform the following steps:
[0090] Deploy the lightweight layer in a general security environment to perform data masking of pseudonymization and regular expression matching; deploy the enhanced layer in a trusted execution environment to integrate a differential privacy noise addition module and an attribute generalization engine; deploy the security layer in a hardware-level security area to run a feature extraction model and perform an identifier deletion operation.
[0091] Further, the system further includes a re-identification simulation module, which is configured to perform the following steps:
[0092] Evaluate the privacy risk value of the available data set through a re-identification attack simulator. If the privacy risk value exceeds the risk threshold, trigger the dynamic policy engine to upgrade the de-identification intensity to generate an enhanced de-identification execution policy; based on the enhanced de-identification execution policy, perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and output an updated available data set.
[0093] Further, the re-identification simulation module is further configured to perform the following steps:
[0094] The re-identification attack simulator includes a background knowledge attack model and a re-identification probability calculation formula; associate the public database with the quasi-identifiers of the de-identified data in the available data set to construct the background knowledge attack model; through the background knowledge attack model, calculate the number of unique combinations of quasi-identifier combinations in the de-identified data, and the number of unique combinations is the number of quasi-identifier combinations that can uniquely identify an individual; according to the number of unique combinations, combine the re-identification probability calculation formula to calculate the privacy risk value of the available data set. The re-identification probability calculation formula is: ; where is the re-identification probability, N is the number of unique combinations, δ is the noise impact factor, δ = 1 + ε, and ε is the differential privacy noise intensity.
[0095] Furthermore, the re-identification simulation module is further configured to perform the following steps:
[0096] Associate the public database with the available dataset, obtain public data and de-identified data, extract quasi-identifier fields based on the public data and de-identified data, and perform normalization processing; based on the quasi-identifier fields, extract quasi-identifiers, encode categorical variables, discretize continuous variables, and identify highly correlated combinations; according to the highly correlated combinations, construct a training dataset, and through the training dataset, train and generate the background knowledge attack model.
[0097] Furthermore, the system further includes a distributed privacy detection module, which is configured to perform the following steps:
[0098] After performing hierarchical de-identification locally, use the processed local data for model training and upload the model gradients; the dynamic policy engine adjusts the differential privacy noise intensity distribution of each participant according to the global model accuracy; detect anomalies based on the gradient value distribution, evaluate the leakage risk, and when the leakage risk exceeds the risk threshold, automatically trigger the security layer processing.
[0099] It should be noted that the above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of the present specification have been described. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0100] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0101] This specification and the drawings are only exemplary descriptions of the present application and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A privacy protection method for data de-identification, characterized in that: The method comprises: Receiving raw data, and dividing the raw data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity assessment model; Obtain data usage scenarios and participant demand information, and perform policy matching through a dynamic policy engine to generate a de-identified execution policy; Constructing a multi-layer trusted execution architecture, wherein the multi-layer trusted execution architecture includes a lightweight layer, an enhancement layer, and a security layer; Based on the multi-layer trusted execution architecture and in accordance with the de-identification execution strategy, the low-sensitivity fields, medium-sensitivity fields and high-sensitivity fields are processed in layers to output a usable data set that meets privacy and security requirements.
2. A privacy protection method for data de-identification as claimed in claim 1, characterized in that: The method further comprises: Evaluate the privacy risk value of the available data set through a re-identification attack simulator, and if the privacy risk value exceeds a risk threshold, trigger the dynamic policy engine to upgrade the de-identification strength and generate an enhanced de-identification execution policy; Based on the enhanced de-identification execution strategy, the low-sensitivity fields, medium-sensitivity fields and high-sensitivity fields are processed in layers, and an updated available data set is output.
3. A privacy protection method for data de-identification as claimed in claim 2, characterized in that: The privacy risk value of the available dataset is evaluated through a re-identification attack simulator, including: The re-identification attack simulator includes a background knowledge attack model and a re-identification probability calculation formula; Associating the public database with the quasi-identifier of the de-identified data in the available data set to construct the background knowledge attack model; Calculate the number of unique combinations of quasi-identifier combinations in the de-identified data by using the background knowledge attack model, where the number of unique combinations is the number of quasi-identifier combinations that can uniquely identify an individual; According to the number of unique combinations and in combination with the re-identification probability calculation formula, a privacy risk value of the available data set is calculated.
4. A privacy protection method for data de-identification as claimed in claim 3, characterized in that: Constructing the background knowledge attack model includes: Associating the public database with the available data set, obtaining the public data and the de-identified data, extracting the quasi-identifier field based on the public data and the de-identified data, and performing normalization processing; Based on the quasi-identifier field, extract the quasi-identifier, encode the categorical variables, discretize the continuous variables, and identify the high correlation combination; A training data set is constructed according to the high correlation combination, and the background knowledge attack model is trained and generated through the training data set.
5. A privacy protection method for data de-identification as claimed in claim 3, characterized in that: The re-identification probability calculation formula is: ; in, is the re-identification probability, N is the number of unique combinations, δ is the noise impact factor, δ=1+ε, and ε is the differential privacy noise intensity.
6. A privacy protection method for data de-identification as claimed in claim 1, characterized in that: The training steps of the sensitivity assessment model include: Constructing a labeled data set, wherein the labeled data set includes a field feature vector and a field sensitivity level label; Based on the labeled data set, a random forest classifier is used for training to generate an initial sensitivity assessment model; Construct fuzzy field samples by adding small perturbations; The initial sensitivity assessment model is adversarially trained based on the fuzzy field samples to generate the sensitivity assessment model.
7. A privacy protection method for data de-identification as claimed in claim 1, characterized in that: Obtain data usage scenarios and participant demand information, and perform policy matching through the dynamic policy engine to generate a de-identification execution policy, including: Calling a preset policy template library according to the data usage scenario to match the target policy template, wherein the policy template library contains de-identification rules in different fields; Obtaining participant demand information through the smart contract interface, wherein the participant demand information includes privacy budget and minimum data accuracy; Referring to the demand information of the participants, the target policy template is dynamically adjusted based on a multi-party voting mechanism to generate a de-identification execution strategy, wherein the voting weight is allocated according to the historical compliance rate of the participants, and the de-identification execution strategy includes pseudonymization strength, generalization granularity and noise amount.
8. A privacy protection method for data de-identification as claimed in claim 1, characterized in that: Build a multi-layer trusted execution architecture, including: Deploy the lightweight layer in a common security environment and perform pseudonymization and regular expression matching data masking; Deploy the enhancement layer in a trusted execution environment, integrating a differential privacy noise addition module and an attribute generalization engine; The security layer is deployed in a hardware-level security zone to run feature extraction models and implement identifier removal operations.
9. The privacy protection method for data de-identification according to claim 1, characterized in that: The method further comprises: After performing layered de-identification locally, the processed local data is used for model training and the model gradients are uploaded; The dynamic policy engine adjusts the differential privacy noise intensity distribution of each participant according to the global model accuracy; Anomalies are detected based on the gradient value distribution and leakage risks are assessed. When the leakage risk exceeds the risk threshold, the safety layer processing is automatically triggered.
10. A privacy protection system for data de-identification, characterized in that: The system comprises: A sensitivity classification module, the sensitivity classification module is used to receive raw data, and based on a preset sensitivity assessment model, divide the raw data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields; An execution strategy matching module, which is used to obtain data usage scenarios and participant demand information, and perform strategy matching through a dynamic strategy engine to generate a de-identified execution strategy; A trusted execution architecture building module, wherein the trusted execution architecture building module is used to build a multi-layer trusted execution architecture, wherein the multi-layer trusted execution architecture includes a lightweight layer, an enhancement layer, and a security layer; A layered de-identification processing module is used to perform layered processing on the low-sensitivity fields, medium-sensitivity fields and high-sensitivity fields based on the multi-layer trusted execution architecture and in accordance with the de-identification execution strategy, and output a usable data set that meets privacy and security requirements.
Citation Information
Patent Citations
Diagnosis and treatment data de-identification method and device and query system
CN113591154A
Data product release method and system and storage medium
CN115935421A
Data privacy protection method and device, equipment and medium
CN116579018A
Financial customer information de-identification method and related device
CN118761084A
User data intelligent protection method and system based on differential privacy
CN119720263A
Cited By
Trusted data space construction method and equipment based on Eclipse EDC and medium
CN120805194A
Data sharing privacy protection platform responding to tax digitization demand
CN120951387A
Clinical data asset management method and system based on data intelligent anonymization
CN121071916A
Clinical data asset management method and system based on data intelligent anonymization
CN121071916B
Self-adaptive de-identification strategy generation method based on multi-objective optimization
CN122413450A