A privacy protection method and system for data de-identification

Through the hierarchical processing of the sensitivity evaluation model and multi-layer trusted execution architecture, the problem of insufficient adjustment of data de-identification strategies in the prior art is solved, and a dynamic balance between privacy protection and data availability is achieved.

CN120068157BActive Publication Date: 2025-08-01LINGSHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510534091.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-01
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing technology lacks dynamic strategy adjustment capabilities and a complete execution architecture in the process of data de-identification, making it difficult to balance privacy protection and data availability.

Method used

Through the sensitivity evaluation model, the data is divided into low, medium and high sensitive fields, and a de-identification strategy is generated based on data usage scenarios and participants' needs, and a multi-layer trusted execution architecture of lightweight layers, enhancement layers and security layers is built to process the fields layered.

Benefits of technology

A dynamic balance between privacy protection and data availability is achieved, ensuring that data remains sufficiently available while meeting privacy security requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068157B_ABST
    Figure CN120068157B_ABST
Patent Text Reader

Abstract

The present invention discloses a privacy protection method and system for data de-identification, which relates to the technical field of data processing. The method includes: after receiving the original data, dividing the data into low, medium, and high-sensitive fields based on a preset sensitivity assessment model; generating a de-identification execution policy through a dynamic policy engine according to the data usage scenario and the needs of the participating parties; constructing a multi-layer trusted execution architecture including a lightweight layer, an enhanced layer, and a security layer, and performing hierarchical processing on fields with different sensitivity levels according to the policy, and finally outputting an available data set that meets the privacy and security requirements. The present invention solves the technical problem that the prior art lacks the ability of dynamic policy adjustment and a perfect execution architecture, and it is difficult to balance between privacy protection and data availability, and achieves the technical effect of realizing the dynamic balance between privacy protection and data availability through dynamic policy adjustment and a multi-layer trusted execution architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a privacy protection method and system for data de-identification. Background Art

[0002] With the rapid development of big data technology, data has become an important resource to promote social progress and economic growth. However, the wide use of data also brings severe privacy leakage risks. As an important privacy protection technology, data de-identification reduces the risk of data being re-identified by removing or replacing direct identifiers in the data, thus supporting the legitimate use and sharing of data while protecting personal privacy.

[0003] Traditional data de-identification methods usually adopt a single technical means (such as data masking, generalization or noise addition), and it is difficult to cope with complex and changeable actual application scenarios. For example, the sensitivity differences of different data fields are relatively large, and unified processing may lead to insufficient protection of low-sensitivity fields or over-processing of high-sensitivity fields, affecting the usability of data and the privacy protection effect. In addition, the diversity of data usage scenarios and the needs of participants also pose higher requirements for the flexibility and adaptability of de-identification strategies. Summary of the Invention

[0004] This application provides a privacy protection method and system for data de-identification, which is used to solve the technical problem that the prior art lacks the ability of dynamic policy adjustment and a perfect execution architecture, and it is difficult to balance privacy protection and data usability.

[0005] In the first aspect of this application, a privacy protection method for data de-identification is provided. The method includes: receiving original data, and based on a preset sensitivity assessment model, dividing the original data into low-sensitivity fields, medium-sensitivity fields and high-sensitivity fields; obtaining data usage scenarios and participant requirement information, and performing policy matching through a dynamic policy engine to generate a de-identification execution policy; constructing a multi-layer trusted execution architecture, where the multi-layer trusted execution architecture includes a lightweight layer, an enhanced layer and a security layer; based on the multi-layer trusted execution architecture, processing the low-sensitivity fields, medium-sensitivity fields and high-sensitivity fields in layers according to the de-identification execution policy, and outputting an available data set that meets the privacy and security requirements.

[0006] In a second aspect of the present application, a privacy protection system for data de-identification is provided. The system includes: a sensitivity classification module for receiving original data and classifying the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity evaluation model; an execution policy matching module for obtaining data usage scenarios and participating party requirement information and generating a de-identification execution policy through a dynamic policy engine; a trusted execution architecture construction module for constructing a multi-layer trusted execution architecture including a lightweight layer, an enhanced layer, and a security layer; and a hierarchical de-identification processing module for performing hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on the multi-layer trusted execution architecture according to the de-identification execution policy and outputting an available data set that meets privacy and security requirements.

[0007] One or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0008] A privacy protection method and system for data de-identification provided in the present application relate to the technical field of data processing. The data is classified into low, medium, and high-sensitivity fields through a sensitivity evaluation model, a de-identification policy is generated in combination with scenarios and requirements, and a multi-layer architecture of a lightweight layer, an enhanced layer, and a security layer is constructed to perform hierarchical processing on the fields and output an available data set that meets privacy and security requirements. This solves the technical problem in the prior art of lacking the ability of dynamic policy adjustment and a perfect execution architecture and being difficult to balance privacy protection and data availability, and achieves the technical effect of realizing the dynamic balance between privacy protection and data availability through dynamic policy adjustment and a multi-layer trusted execution architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0010] Figure 1 It is a schematic flowchart of a privacy protection method for data de-identification provided by an embodiment of the present application;

[0011] Figure 2 It is a schematic structural diagram of a privacy protection system for data de-identification provided by an embodiment of the present application.

[0012] Explanation of the accompanying drawings: sensitivity classification module 11, execution strategy matching module 12, trusted execution architecture construction module 13, layered de-identification processing module 14. DETAILED DESCRIPTION

[0013] This application provides a privacy protection method and system for data de-identification, which is used to solve the technical problem that the existing technology lacks dynamic policy adjustment capabilities and a complete execution architecture, and is difficult to balance between privacy protection and data availability.

[0014] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0015] It should be noted that the terms "first", "second", etc. in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices.

[0016] Example 1, as Figure 1 As shown, this application provides a privacy protection method for data de-identification, which includes:

[0017] P10: Receive original data, and divide the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity assessment model.

[0018] Furthermore, the embodiment of the present application further includes step P10a, which further includes:

[0019] P11a: Construct a labeled dataset, which includes field feature vectors and field sensitivity level labels; P12a: Based on the labeled dataset, use a random forest classifier for training to generate an initial sensitivity assessment model; P13a: Construct fuzzy field samples by adding small perturbations; P14a: Perform adversarial training on the initial sensitivity assessment model based on the fuzzy field samples to generate the sensitivity assessment model.

[0020] It should be understood that after receiving the original data first, it is necessary to divide the data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity assessment model. The core of this step lies in accurately classifying data fields through the data sensitivity assessment model to ensure that subsequent de-identification processing can adopt differential privacy protection strategies for fields with different sensitivity levels.

[0021] Among them, in order to construct an efficient and robust sensitivity assessment model, a labeled dataset is first constructed. The labeled dataset includes field feature vectors and field sensitivity level labels. The field feature vector is a multi-dimensional feature extracted from the original data to describe the field attributes, such as field type (e.g., numerical type, text type), field length, field value distribution, etc.; the field sensitivity level label is an artificial annotation of the field sensitivity according to business rules or expert experience. For example, the identity ID is marked as a high-sensitivity field, the age is marked as a medium-sensitivity field, and the gender is marked as a low-sensitivity field.

[0022] Next, based on the labeled dataset, a random forest classifier is used for training to generate an initial sensitivity assessment model. The random forest classifier is an ensemble learning method. By constructing multiple decision trees and integrating their prediction results, it can effectively process high-dimensional feature data and reduce the risk of overfitting. During the training process, through bootstrap sampling, multiple sample subsets are randomly drawn from the labeled dataset with replacement, and a decision tree is independently trained on each subset. Finally, the classification results of all decision trees are aggregated through a majority voting mechanism to obtain the final classification result. The advantages of the random forest include the ability to handle high-dimensional data, tolerance to missing values and outliers, and the ability to evaluate feature importance. These characteristics make it very suitable for the training of the sensitivity assessment model.

[0023] Next, in order to improve the robustness of the model, by adding small perturbations, fuzzy field samples are constructed. Fuzzy field samples refer to slightly adjusting the original field feature vector (e.g., adding random noise to numerical fields and replacing synonyms for text fields) to simulate the fuzziness and uncertainty of field features in actual scenarios. This kind of perturbation can simulate the uncertainty of data in the real scenario and help the model learn a wider range of feature patterns. [[ID=I1]]

[0024] Furthermore, based on the fuzzy field samples, perform adversarial training on the initial sensitivity evaluation model to generate a final sensitivity evaluation model. Adversarial training is a technique that improves the generalization ability of the model by introducing adversarial samples (i.e., fuzzy field samples). Its core idea is to enable the model to continuously learn how to distinguish normal samples from adversarial samples during the training process, thereby enhancing the stability and accuracy of the model in practical applications. During the training process, the model not only learns the normal samples in the labeled dataset but also learns the perturbed fuzzy field samples. In this way, the model can better adapt to the minor changes in the input data, thus improving the accuracy and stability of sensitivity evaluation.

[0025] Through the above steps, the generated sensitivity evaluation model can more accurately classify the sensitivity of data fields, providing a reliable basis for subsequent data de-identification processing.

[0026] P20: Obtain data usage scenarios and participant requirement information, and perform policy matching through a dynamic policy engine to generate a de-identification execution policy.

[0027] Furthermore, step P20 of the embodiment of the present application further includes:

[0028] P21: Invoke a pre-set policy template library according to the data usage scenario to match a target policy template, where the policy template library contains de-identification rules for different fields; P22: Obtain participant requirement information through a smart contract interface, and the participant requirement information includes privacy budget and minimum data accuracy; P23: Refer to the participant requirement information and dynamically adjust the target policy template based on a multi-party voting mechanism to generate a de-identification execution policy, where the voting weights are allocated according to the historical compliance rates of the participants, and the de-identification execution policy includes pseudonymization intensity, generalization granularity, and noise volume.

[0029] Optionally, according to specific data usage scenarios and the privacy requirements of participants, perform policy matching through a dynamic policy engine to generate an adapted de-identification policy to ensure that the data remains sufficiently usable while meeting privacy protection requirements.

[0030] Specifically, first invoke the pre-set policy template library according to the data usage scenario to match a target policy template. The policy template library is a collection that contains de-identification rules for different fields, covering various data types and application scenarios. For example, the de-identification rules in the medical field may focus on protecting patient privacy, while the rules in the financial field pay more attention to the accuracy of transaction data. By invoking the policy template library, the system can quickly match the target policy template that is most suitable for the current data usage scenario, providing a basic framework for subsequent policy adjustment.

[0031] Next, obtain the participant requirement information through the smart contract interface. The participant requirement information includes the privacy budget and the minimum data accuracy. The privacy budget refers to the resource cost that a participant is willing to pay for data privacy protection, such as the amount of differential privacy noise that is allowed to be added; the minimum data accuracy refers to the lowest requirement of a participant for data availability, such as the granularity after data generalization cannot be lower than a certain threshold. The smart contract interface is an automated protocol based on blockchain technology, which can ensure the transparency and immutability of the participant requirement information, thus providing a trustworthy input for policy generation.

[0032] Furthermore, with reference to the participant requirement information, dynamically adjust the target policy template based on a multi-party voting mechanism to generate a de-identified execution policy. The multi-party voting mechanism is a distributed decision-making method, which determines the final policy adjustment direction through the joint voting of multiple participants. The voting weights are allocated according to the historical compliance rates of the participants, that is, the higher the historical compliance rate of a participant, the greater its voting weight, so as to ensure the fairness and rationality of policy adjustment. The generated de-identified execution policy includes key parameters such as the pseudonymization strength, generalization granularity, and noise amount. Among them, the pseudonymization strength determines the degree to which fields are replaced, the generalization granularity controls the accuracy of the data, and the noise amount is used to achieve differential privacy protection.

[0033] Through the above steps, the system can dynamically generate a de-identified execution policy with strong adaptability and good privacy protection effect according to different data usage scenarios and participant requirements.

[0034] P30: Construct a multi-layer trusted execution architecture, which includes a lightweight layer, an enhanced layer, and a security layer.

[0035] Furthermore, step P30 of the embodiment of the present application further includes:

[0036] P31: Deploy the lightweight layer in a general security environment to perform data masking of pseudonymization and regular expression matching; P32: Deploy the enhanced layer in a trusted execution environment to integrate a differential privacy noise addition module and an attribute generalization engine; P33: Deploy the security layer in a hardware-level security area to run a feature extraction model and perform an identifier deletion operation.

[0037] Specifically, construct a multi-layer trusted execution architecture, which includes a lightweight layer, an enhanced layer, and a security layer. Through the design of the layered architecture, de-identified operations with different complexities and security levels are assigned to the corresponding execution environments, so as to optimize the system performance and resource utilization while ensuring the privacy protection effect.

[0038] Exemplarily, the lightweight layer is deployed in a general security environment and is mainly responsible for performing data masking operations such as pseudonymization and regular expression matching. The lightweight layer mainly processes the de-identification operations of low-sensitivity fields, and is characterized by low computational complexity and fast execution speed. Pseudonymization is a de-identification technique that replaces the original identifier with a random or pseudo-random value, such as replacing a name with a random string; regular expression matching is used to identify and mask sensitive information in a specific format, such as a card number or a phone number. Since the lightweight layer is deployed in a general security environment, it is applicable to scenarios with relatively high performance requirements but relatively low security requirements.

[0039] Next, the enhanced layer is deployed in a trusted execution environment, integrating a differential privacy noise addition module and an attribute generalization engine. The enhanced layer mainly processes the de-identification operations of medium-sensitivity fields, and is characterized by taking into account both security and data availability. The differential privacy noise addition module ensures the privacy of individual data in the statistical results by adding random noise to the data; the attribute generalization engine further reduces the identification risk of the data by replacing specific values with broader categories (for example, generalizing the age "25 years old" to "20 - 30 years old"). The trusted execution environment (TEE) is a hardware-level secure isolation environment that can ensure that the operations in the enhanced layer are not interfered with by external attacks during runtime.

[0040] Furthermore, the security layer is deployed in a hardware-level secure area, running a feature extraction model and implementing an identifier deletion operation. The security layer mainly processes the de-identification operations of high-sensitivity fields, and is characterized by the highest security level, but also relatively high computational complexity. The feature extraction model is used to extract key features from the data while removing identifiers that may reveal individual identities; the identifier deletion operation directly deletes or permanently masks high-sensitivity fields, such as deleting an identity ID or a social security number. The hardware-level secure area is a higher-level security protection mechanism than the trusted execution environment, usually implemented based on dedicated hardware (such as a security chip), and can provide the highest level of security guarantee for the security layer.

[0041] Through the above hierarchical deployment, the multi-layer trusted execution architecture can flexibly process data with different sensitivity levels in different security environments, while taking into account privacy protection and data availability, providing comprehensive technical support for data de-identification.

[0042] P40: Based on the multi-layer trusted execution architecture, according to the de-identification execution strategy, the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields are processed in layers, and an available data set that meets the privacy and security requirements is output.

[0043] It should be understood that, based on the multi-layer trusted execution architecture and in accordance with the de-identification execution policy, different levels of sensitive fields are processed in layers to output an available dataset that meets privacy and security requirements. The core of this step lies in the coordinated action of the multi-layer architecture and the execution policy to perform differential de-identification operations on fields with different sensitivity levels, thereby maximizing data availability while protecting data privacy.

[0044] Among them, the processing of low-sensitivity fields is completed in the lightweight layer. According to the pseudonymization intensity in the de-identification execution policy, pseudonymization operations are performed on low-sensitivity fields. For example, the user nickname is replaced with a random string. At the same time, through regular expression matching technology, possible sensitive information fragments (such as the username part in the email address) are identified and masked. Since the lightweight layer is deployed in a general security environment, it has a fast processing speed and low resource consumption, making it suitable for scenarios with high performance requirements.

[0045] The processing of medium-sensitivity fields is completed in the enhanced layer. According to the generalization granularity and noise amount in the de-identification execution policy, the system calls the attribute generalization engine to perform generalization operations on medium-sensitivity fields. For example, the specific age value is replaced with an age range (such as "20 - 30 years old"). At the same time, the differential privacy noise addition module adds an appropriate amount of random noise to the data to ensure the protection of individual privacy in the statistical results. The enhanced layer is deployed in a trusted execution environment (TEE), which can effectively prevent external attacks and ensure the security and reliability of the processing process.

[0046] The processing of high-sensitivity fields is completed in the security layer. According to the identifier deletion requirement in the de-identification execution policy, the system runs a feature extraction model to extract key features from the data and delete or permanently mask high-sensitivity fields (such as identity ID, social security number, etc.). The security layer is deployed in a hardware-level security area, providing the highest level of security guarantee based on dedicated hardware (such as a security chip) to ensure that the processing process of high-sensitivity fields is not interfered by any external factors.

[0047] After the layered processing, the system integrates the processing results of low-, medium-, and high-sensitivity fields to generate an available dataset that meets privacy and security requirements. While protecting privacy, this dataset still retains sufficient data availability.

[0048] Furthermore, the embodiment of the present application further includes step P50, and step P50 further includes:

[0049] P51: Evaluate the privacy risk value of the available dataset through a re-identification attack simulator. If the privacy risk value exceeds the risk threshold, trigger the dynamic policy engine to upgrade the de-identification strength and generate an enhanced de-identification execution policy; P52: Based on the enhanced de-identification execution policy, perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and output an updated available dataset.

[0050] Optionally, the embodiment of the present application further evaluates and optimizes the privacy risk of the available dataset through a re-identification attack simulator and a dynamic policy engine to ensure the continuous effectiveness of data privacy protection. The core of this step lies in timely discovering and fixing potential privacy leakage risks in the dataset through simulated attacks and dynamic policy adjustments, thereby enhancing data security.

[0051] Specifically, first, evaluate the privacy risk value of the available dataset through a re-identification attack simulator. A re-identification attack simulator is a tool used to detect whether a de-identified dataset can be re-associated with the original individual, and it can simulate the process of an attacker attempting to re-identify an individual's identity through de-identified data. Its working principle is to calculate the probability of an individual in the dataset being re-identified, that is, the privacy risk value, by analyzing the feature distribution and correlation in the dataset. If the privacy risk value exceeds the preset risk threshold, it indicates that the current de-identification strength is insufficient to effectively protect data privacy. At this time, the system triggers the dynamic policy engine to upgrade the de-identification execution policy and generate an enhanced de-identification execution policy. The dynamic policy engine further enhances the de-identification strength by adjusting key parameters such as the pseudonymization strength, generalization granularity, and noise amount to cope with potential privacy leakage risks.

[0052] Furthermore, based on the enhanced de-identification execution policy, re-perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and output an updated available dataset. Exemplarily, the low-sensitivity fields further reduce the recognition risk in the lightweight layer by enhancing the pseudonymization strength; the medium-sensitivity fields improve the data privacy protection level in the enhanced layer by increasing the noise amount and expanding the generalization granularity; the high-sensitivity fields ensure that high-sensitivity information cannot be re-identified in the security layer through more strict feature extraction and identifier deletion operations. The updated available dataset is re-evaluated by the re-identification attack simulator to ensure that its privacy risk value is lower than the risk threshold, thus meeting higher privacy and security requirements.

[0053] By introducing step P50, this solution can effectively address the privacy risk after de-identification and ensure the security and availability of data throughout the entire life cycle.

[0054] Furthermore, step P51 of the embodiment of the present application further includes:

[0055] P51-1: The re-identification attack simulator includes a background knowledge attack model and a re-identification probability calculation formula; P51-2: Associate the publicly available database with the quasi-identifiers of the de-identified data in the available dataset to construct the background knowledge attack model; P51-3: Through the background knowledge attack model, calculate the number of unique combinations of quasi-identifiers in the de-identified data, where the number of unique combinations is the number of combinations of quasi-identifiers that can uniquely identify an individual; P51-4: According to the number of unique combinations and in combination with the re-identification probability calculation formula, calculate the privacy risk value of the available dataset.

[0056] The re-identification probability calculation formula is as follows: ; where is the re-identification probability, N is the number of unique combinations, δ is the noise impact factor, δ = 1 + ε, and ε is the differential privacy noise intensity.

[0057] It should be understood that the specific process of privacy risk assessment can be further refined. By using the re-identification attack simulator to evaluate the de-identified available dataset, the privacy security of the data can be ensured.

[0058] Specifically, the re-identification attack simulator includes a background knowledge attack model and a re-identification probability calculation formula. The background knowledge attack model is used to simulate the process of an attacker using the information in the publicly available database for correlation analysis with the de-identified data; the re-identification probability calculation formula is used to quantify the probability of an individual being re-identified, that is, the privacy risk value.

[0059] Among them, the publicly available database can be associated with the quasi-identifiers of the de-identified data in the available dataset to construct the background knowledge attack model. Quasi-identifiers refer to fields that can identify an individual when combined with other data sources, such as age, gender, postal code, etc. By matching the quasi-identifiers in the publicly available database with the de-identified data, the background knowledge attack model can simulate the process of an attacker inferring an individual's identity using external information.

[0060] Furthermore, through the background knowledge attack model, calculate the number of unique combinations of quasi-identifiers in the de-identified data. The number of unique combinations is the number of combinations of quasi-identifiers that can uniquely identify an individual. For example, in a certain dataset, if the combination of "age = 25 years old, gender = male, postal code = 100001" corresponds to only one individual, then this combination is a unique combination. The more the number of unique combinations, the higher the privacy risk of the dataset.

[0061] Next, based on the unique combination number and combined with the re-identification probability calculation formula, the simulator calculates the privacy risk value of the available data set. By introducing a noise impact factor, this formula takes into account the inhibitory effect of differential privacy noise on the re-identification probability. The differential privacy noise intensity ε is a key parameter for controlling the degree of privacy protection. The smaller ε is, the stronger the privacy protection, but the availability of the data may decrease. By calculating the privacy risk value, the system can determine whether the current de-identification strategy is effective enough. If the privacy risk value exceeds the preset risk threshold, the dynamic policy engine is triggered to upgrade the de-identification intensity.

[0062] Through the above steps, the embodiments of the present application can effectively evaluate the privacy risk of de-identified data and dynamically adjust the de-identification strategy according to the risk assessment results to ensure the best balance between privacy protection and availability.

[0063] Further, step P51-2 of the embodiments of the present application further includes:

[0064] P51-21: Associate the public database with the available data set, obtain public data and de-identified data, extract quasi-identifier fields based on the public data and de-identified data, and perform normalization processing; P51-22: Based on the quasi-identifier fields, extract quasi-identifiers, encode categorical variables, discretize continuous variables, and identify highly correlated combinations; P51-23: According to the highly correlated combinations, construct a training data set, and through the training data set, train and generate the background knowledge attack model.

[0065] Specifically, the embodiments of the present application further refine the construction process of the background knowledge attack model. By associating the public database with the available data set, extracting and processing quasi-identifier fields, and finally training and generating the background knowledge attack model. The core of this step is to construct a high-quality training data set through technical means such as data normalization, variable encoding, and discretization, so as to ensure that the background knowledge attack model can accurately simulate the behavior of attackers and provide a reliable basis for privacy risk assessment.

[0066] First, associate the public database with the de-identified available data set to obtain data samples of both. Based on these data, extract quasi-identifier fields, which may be used by attackers to re-identify individuals. Quasi-identifier fields refer to fields that can identify individuals when combined with other data sources, such as age, gender, zip code, etc. To ensure the consistency and comparability of the data, perform normalization processing on the extracted quasi-identifier fields, such as converting age to a unified range or standardizing text fields.

[0067] Next, after extracting the quasi-identifier fields, encode the categorical variables, discretize the continuous variables, and identify highly correlated combinations. For example, convert categorical variables (such as gender, occupation, etc.) into numerical form through one-hot encoding or label encoding; discretize continuous variables (such as age, income, etc.) through binning or clustering methods. Through these preprocessing steps, identify the quasi-identifier fields with highly correlated combinations. For example, the combination of "age = 25 years old, gender = male, postal code = 100001" may have a high correlation. Identify these highly correlated combinations through statistical analysis or machine learning algorithms to provide key features for subsequent model training.

[0068] Furthermore, based on the highly correlated combinations, construct a training dataset and use the training dataset to train the background knowledge attack model. The training dataset includes the quasi-identifier fields in the public data and the quasi-identifier fields in the de-identified data, as well as the matching relationships between them. Machine learning algorithms (such as decision trees, random forests, or neural networks) can be used to train in combination with the training dataset to learn the association rules between the quasi-identifier fields, thereby simulating the process of an attacker re-identifying individuals using background knowledge and generating a background knowledge attack model.

[0069] Through the above steps, effectively generate a background knowledge attack model. This model can accurately simulate the behavior of attackers, provide a scientific basis for privacy risk assessment, thereby supporting the optimization and adjustment of the dynamic policy engine, and further improving the level of data privacy protection.

[0070] Furthermore, the embodiment of the present application further includes step P60, and step P60 further includes:

[0071] P61: After performing hierarchical de-identification locally, use the processed local data for model training and upload the model gradients; P62: The dynamic policy engine adjusts the differential privacy noise intensity allocation of each participant according to the global model accuracy; P63: Detect anomalies based on the gradient value distribution, evaluate the leakage risk, and automatically trigger the security layer processing when the leakage risk exceeds the risk threshold.

[0072] In a possible embodiment of the present application, the embodiment of the present application further expands the application scenario of data privacy protection, and realizes privacy protection and model training in a distributed environment through a federated learning framework combined with differential privacy technology. The core of this step is to effectively prevent the risk of data leakage while ensuring the model accuracy through dynamic adjustment of the differential privacy noise intensity and anomaly detection mechanism, and ensure the comprehensiveness and dynamicity of data privacy protection.

[0073] First, after performing hierarchical de-identification locally, each participating party uses the processed local data for model training and uploads the model gradients. Exemplarily, each participating party performs hierarchical de-identification processing on the data locally to ensure that the data has undergone privacy protection processing before leaving the local environment. The processed local data is used to train the local model, and the model gradients are uploaded to the central server through the federated learning framework instead of uploading the original data, thereby further reducing the risk of privacy leakage.

[0074] Next, the dynamic policy engine adjusts the differential privacy noise intensity distribution of each participating party according to the global model accuracy. The dynamic policy engine dynamically adjusts the differential privacy noise intensity added by each participating party during the gradient upload process by monitoring the accuracy change of the global model. For example, when the global model accuracy is low, the noise intensity is appropriately reduced to improve the model performance; when the global model accuracy is high, the noise intensity is increased to enhance privacy protection. This dynamic adjustment mechanism can achieve a balance between model accuracy and privacy protection, ensuring the efficiency and security of the federated learning framework.

[0075] Furthermore, anomalies are detected based on the gradient value distribution, and the leakage risk is evaluated. When the leakage risk exceeds the risk threshold, the security layer processing is automatically triggered. Exemplarily, by analyzing the distribution of the uploaded gradient values, whether there are abnormal patterns is detected. For example, some gradient values may imply certain features of the original data, thus increasing the risk of privacy leakage. Through statistical analysis or machine learning algorithms, the abnormal conditions of the gradient value distribution are detected, and the potential leakage risk is evaluated. When the leakage risk exceeds the preset risk threshold, the dynamic policy engine automatically triggers the security layer processing, such as increasing the noise intensity, restricting the gradient upload frequency, or isolating the abnormal participating party, to effectively contain the risk of privacy leakage. The security layer is deployed in the hardware-level security area and can execute more stringent privacy protection measures, such as feature extraction, identifier deletion, or adding stronger noise.

[0076] By introducing step P60, the embodiments of the present application can dynamically adjust the privacy protection policy in distributed model training, taking into account model performance and data availability while ensuring privacy security. This step is not only applicable to the local data processing scenario but also can be extended to the collaborative learning scenarios across institutions and regions, with broad application value.

[0077] In summary, the embodiments of the present application at least have the following technical effects:

[0078] First, this application divides the original data into low, medium, and high-sensitivity fields based on a sensitivity assessment model. Secondly, in combination with the data usage scenario and the needs of the parties involved, a de-identification execution policy is generated through a dynamic policy engine. Then, a multi-layer trusted execution architecture including a lightweight layer, an enhanced layer, and a security layer is constructed. Finally, different sensitivity fields are processed in layers according to the policy to output an available data set that meets the privacy and security requirements.

[0079] It achieves the technical effect of dynamically balancing privacy protection and data availability through precise sensitivity assessment, flexible policy matching, and a multi-layer trusted execution architecture.

[0080] Embodiment 2, based on the same inventive concept as the privacy protection method for data de-identification in the foregoing embodiment, as Figure 2 shown, this application provides a privacy protection system for data de-identification. The system in the embodiments of this application and the method embodiments are based on the same inventive concept. Among them, the system includes:

[0081] A sensitivity division module 11, which is used to receive the original data and divide the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity assessment model.

[0082] An execution policy matching module 12, which is used to obtain data usage scenarios and information on the needs of the parties involved, and perform policy matching through a dynamic policy engine to generate a de-identification execution policy.

[0083] A trusted execution architecture construction module 13, which is used to construct a multi-layer trusted execution architecture. The multi-layer trusted execution architecture includes a lightweight layer, an enhanced layer, and a security layer.

[0084] A hierarchical de-identification processing module 14, which is used to perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on the multi-layer trusted execution architecture according to the de-identification execution policy, and output an available data set that meets the privacy and security requirements.

[0085] Further, the sensitivity division module 11 is also used to perform the following steps:

[0086] Construct an annotated data set, where the annotated data set includes field feature vectors and field sensitivity level labels; based on the annotated data set, train using a random forest classifier to generate an initial sensitivity assessment model; by adding small perturbations, construct fuzzy field samples; and perform adversarial training on the initial sensitivity assessment model based on the fuzzy field samples to generate the sensitivity assessment model.

[0087] Further, the execution policy matching module 12 is further configured to perform the following steps:

[0088] Call a pre-set policy template library according to the data usage scenario to match a target policy template, where the policy template library contains de-identification rules in different fields; obtain the participating party's requirement information through a smart contract interface, and the participating party's requirement information includes a privacy budget and a minimum data accuracy; with reference to the participating party's requirement information, dynamically adjust the number of the target policy templates based on a multi-party voting mechanism to generate a de-identification execution policy, where the voting weight is allocated according to the historical compliance rate of the participating party, and the de-identification execution policy includes a pseudonymization intensity, a generalization granularity, and a noise amount.

[0089] Further, the trusted execution architecture construction module 13 is further configured to perform the following steps:

[0090] Deploy the lightweight layer in a general security environment to perform data masking for pseudonymization and regular expression matching; deploy the enhanced layer in a trusted execution environment to integrate a differential privacy noise addition module and an attribute generalization engine; deploy the security layer in a hardware-level security area to run a feature extraction model and implement an identifier deletion operation.

[0091] Further, the system further includes a re-identification simulation module, which is configured to perform the following steps:

[0092] Evaluate the privacy risk value of the available data set through a re-identification attack simulator. If the privacy risk value exceeds the risk threshold, trigger the dynamic policy engine to upgrade the de-identification intensity to generate an enhanced de-identification execution policy; based on the enhanced de-identification execution policy, perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and output an updated available data set.

[0093] Further, the re-identification simulation module is further configured to perform the following steps:

[0094] The re-identification attack simulator includes a background knowledge attack model and a re-identification probability calculation formula; associate the public database with the quasi-identifiers of the de-identified data in the available data set to construct the background knowledge attack model; through the background knowledge attack model, calculate the number of unique combinations of quasi-identifier combinations in the de-identified data, and the number of unique combinations is the number of quasi-identifier combinations that can uniquely identify an individual; according to the number of unique combinations, combine the re-identification probability calculation formula to calculate the privacy risk value of the available data set. The re-identification probability calculation formula is: ; where is the re-identification probability, N is the number of unique combinations, δ is the noise impact factor, δ = 1 + ε, and ε is the differential privacy noise intensity.

[0095] Further, the re-identification simulation module is further configured to perform the following steps:

[0096] Associate the public database with the available dataset, obtain public data and de-identified data, extract quasi-identifier fields based on the public data and de-identified data, and perform normalization processing; based on the quasi-identifier fields, extract quasi-identifiers, encode categorical variables, discretize continuous variables, and identify highly correlated combinations; according to the highly correlated combinations, construct a training dataset, and through the training dataset, train and generate the background knowledge attack model.

[0097] Further, the system further includes a distributed privacy detection module, which is configured to perform the following steps:

[0098] After performing hierarchical de-identification locally, use the processed local data for model training and upload the model gradients; the dynamic policy engine adjusts the differential privacy noise intensity distribution of each participant according to the global model accuracy; detect anomalies based on the gradient value distribution, evaluate the leakage risk, and automatically trigger the security layer processing when the leakage risk exceeds the risk threshold.

[0099] It should be noted that the above sequence of embodiments of the present application is only for description and does not represent the advantages or disadvantages of the embodiments. And the above describes specific embodiments of this specification. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0100] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.

[0101] This specification and the drawings are only exemplary descriptions of the present application and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.

Claims

1. A privacy protection method for data de-identification, characterized in that, The method includes: Receiving the original data, and dividing the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a preset sensitivity evaluation model; Obtaining data usage scenarios and participant requirement information, and performing policy matching through a dynamic policy engine to generate a de-identification execution policy; Constructing a multi-layer trusted execution architecture, where the multi-layer trusted execution architecture includes a lightweight layer, an enhanced layer, and a security layer; Based on the multi-layer trusted execution architecture, and according to the de-identification execution policy, performing hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and outputting an available data set that meets privacy and security requirements. Among them, the lightweight layer processes the de-identification operations of low-sensitivity fields, the enhanced layer processes the de-identification operations of medium-sensitivity fields, and the security layer processes the de-identification operations of high-sensitivity fields.

2. The privacy protection method for data de-identification according to claim 1, characterized in that The method further includes: Evaluating the privacy risk value of the available data set through a re-identification attack simulator. If the privacy risk value exceeds the risk threshold, triggering the dynamic policy engine to upgrade the de-identification intensity and generate an enhanced de-identification execution policy; Based on the enhanced de-identification execution policy, performing hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields, and outputting an updated available data set.

3. A privacy protection method for data de-identification according to claim 2, characterized in that Evaluating the privacy risk value of the available data set through a re-identification attack simulator, including: The re-identification attack simulator includes a background knowledge attack model and a re-identification probability calculation formula; Associating a public database with the quasi-identifiers of the de-identified data in the available data set to construct the background knowledge attack model; Through the background knowledge attack model, calculating the number of unique combinations of quasi-identifiers in the de-identified data, where the number of unique combinations is the number of combinations of quasi-identifiers that can uniquely identify an individual; According to the number of unique combinations, and combining the re-identification probability calculation formula, calculating to obtain the privacy risk value of the available data set.

4. The privacy protection method for data de-identification according to claim 3, characterized in that, Constructing the background knowledge attack model includes: Associating a public database with the available data set, obtaining public data and de-identified data, extracting quasi-identifier fields based on the public data and de-identified data, and performing normalization processing; Based on the quasi-identifier fields, extracting quasi-identifiers, encoding categorical variables, discretizing continuous variables, and identifying highly correlated combinations; According to the highly correlated combinations, constructing a training data set, and training to generate the background knowledge attack model through the training data set.

5. The privacy protection method for data de-identification according to claim 3, characterized in that The formula for calculating the re-identification probability is as follows: ; where is the re-identification probability, N is the number of unique combinations, δ is the noise impact factor, δ = 1 + ε, and ε is the differential privacy noise intensity.

6. A privacy protection method for data de-identification as described in claim 1, characterized in that, The training steps of the sensitivity evaluation model include: Constructing an annotated data set, where the annotated data set includes field feature vectors and field sensitivity level labels; Based on the annotated data set, training using a random forest classifier to generate an initial sensitivity evaluation model; Constructing fuzzy field samples by adding small perturbations; Based on the fuzzy field samples, performing adversarial training on the initial sensitivity evaluation model to generate the sensitivity evaluation model.

7. A privacy protection method for data de-identification as described in claim 1, characterized in that, Obtaining data usage scenarios and participant requirement information, and performing policy matching through a dynamic policy engine to generate a de-identification execution policy, including: Call the pre-set policy template library according to the data usage scenario to match the target policy template, where the policy template library contains de-identification rules in different fields; Obtain the participant requirement information through the smart contract interface, where the participant requirement information includes privacy budget and minimum data accuracy; With reference to the participant requirement information, dynamically adjust the target policy template based on a multi-party voting mechanism to generate a de-identification execution policy, where the voting weight is allocated according to the historical compliance rate of the participants, and the de-identification execution policy includes pseudonymization intensity, generalization granularity, and noise volume.

8. A privacy protection method for data de-identification as claimed in claim 1, wherein Construct a multi-layer trusted execution architecture, including: Deploy the lightweight layer in a general security environment to perform data masking for pseudonymization and regular expression matching; Deploy the enhanced layer in a trusted execution environment, integrating a differential privacy noise addition module and an attribute generalization engine; Deploy the security layer in a hardware-level security area, run a feature extraction model, and perform identifier deletion operations.

9. A privacy protection method for data de-identification as claimed in claim 1, wherein The method further includes: After performing hierarchical de-identification locally, use the processed local data for model training and upload the model gradients; The dynamic policy engine adjusts the differential privacy noise intensity distribution of each participant according to the global model accuracy; Detect anomalies based on the gradient value distribution, evaluate the leakage risk, and automatically trigger the security layer processing when the leakage risk exceeds the risk threshold.

10. A privacy protection system for data de-identification, characterized in that, The system includes: A sensitivity division module, which is used to receive the original data and divide the original data into low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on a pre-set sensitivity assessment model; An execution policy matching module, which is used to obtain the data usage scenario and participant requirement information, and perform policy matching through a dynamic policy engine to generate a de-identification execution policy; A trusted execution architecture construction module, which is used to construct a multi-layer trusted execution architecture, and the multi-layer trusted execution architecture includes a lightweight layer, an enhanced layer, and a security layer; A hierarchical de-identification processing module, which is used to perform hierarchical processing on the low-sensitivity fields, medium-sensitivity fields, and high-sensitivity fields based on the multi-layer trusted execution architecture according to the de-identification execution policy, and output an available dataset that meets the privacy and security requirements, where the lightweight layer processes the de-identification operations of low-sensitivity fields, the enhanced layer processes the de-identification operations of medium-sensitivity fields, and the security layer processes the de-identification operations of high-sensitivity fields.

Citation Information

Patent Citations

  • Financial customer information de-identification method and related device

    CN118761084A

  • Data privacy protection method

    CN119848936A