A data anonymization method, apparatus, and readable storage medium
By obtaining the contribution and sensitivity of the attribute fields of the dataset and performing de-identification processing through non-uniform allocation of the privacy budget, the problem of poor data evaluation effect in existing technologies is solved, and a de-identified dataset that reflects the true characteristics of the data is generated.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAJUN TECHNOLOGY (CHONGQING) CO LTD
- Filing Date
- 2026-06-01
- Publication Date
- 2026-07-03
AI Technical Summary
Existing differential privacy technology distributes the privacy budget evenly across all data, resulting in anonymized data that fails to reflect the true distribution characteristics of the original data, thus affecting the effectiveness of data evaluation.
By obtaining the contribution index and sensitivity of the attribute fields of the dataset to be de-identified, and distributing the privacy budget non-uniformly according to the value weight and field sensitivity, a refined de-identification process is carried out.
It enables the accurate allocation of privacy budgets based on the importance and sensitivity of data features, generating de-identified datasets that reflect the true characteristics of the data, thus preventing the leakage of important information.
Smart Images

Figure CN122333531A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a data desensitization method, apparatus and readable storage medium. Background Technology
[0002] In data evaluation scenarios, evaluation models are used to analyze the statistical characteristics of data during the evaluation process. However, this data often involves personal and corporate information, especially in vertical fields that require high-precision data analysis, such as medical diagnostic assistance, financial credit scoring, and open office data. This issue is particularly prominent there. For example, medical data includes patient medical records, and credit data includes user behavior information. Directly using this data poses a risk of privacy breaches.
[0003] To address the risk of privacy breaches, existing technologies offer a differential privacy technique. Differential privacy achieves strict privacy protection by adding controllable noise to the data. However, it distributes the privacy budget evenly across all data, resulting in uniformly added noise. This causes the generated anonymized data to fail to reflect the true distribution characteristics of the original data, leading to problems in data evaluation. Summary of the Invention
[0004] This application provides a data anonymization method, apparatus, and readable storage medium that can solve the problem of not being able to distribute the privacy budget unevenly across all data.
[0005] In a first aspect, embodiments of this application provide a data anonymization method, including: Obtain the dataset to be anonymized, wherein the data to be anonymized in the dataset includes attribute fields, which are used to describe the characteristics of the data to be anonymized. The corresponding evaluation model is determined based on the first task requirement, which is used to indicate the target evaluation task type of the evaluation model. The evaluation model is used to evaluate the value of the dataset to be de-identified according to the first task requirement. The dataset to be de-identified is input into the evaluation model to obtain the first contribution index of each attribute field to the first evaluation result output by the evaluation model; For each attribute field, the value weight of the attribute field is determined according to the first contribution index of the attribute field, and the value weight is used to characterize the correlation strength between the attribute field and the first evaluation result; Obtain the field sensitivity of the attribute field; The target privacy budget for the attribute field is determined based on the value weight and the field sensitivity of the attribute field. Based on the target privacy budget of each attribute field, the attribute fields are anonymized to obtain an anonymized dataset.
[0006] Secondly, embodiments of this application provide a data desensitization device, comprising: The first acquisition module is used to acquire the dataset to be de-identified, wherein the data to be de-identified in the dataset includes attribute fields, and the attribute fields are used to describe the characteristics of the data to be de-identified. It is also used to determine the corresponding evaluation model according to the first task requirements, the first task requirements are used to indicate the target evaluation task type of the evaluation model, and the evaluation model is used to evaluate the value of the dataset to be de-identified according to the first task requirements. The second acquisition module is used to input the dataset to be desensitized into the evaluation model and acquire the first contribution index of each attribute field to the first evaluation result output by the evaluation model. The weight calculation module is used to determine the value weight of each attribute field based on the first contribution index of the attribute field, wherein the value weight is used to characterize the correlation strength between the attribute field and the first evaluation result. A sensitivity calculation module is used to obtain the field sensitivity of the attribute field; A privacy budget calculation module is used to determine the target privacy budget for the attribute field based on the value weight and the field sensitivity of the attribute field. The data desensitization module is used to desensitize each attribute field according to the target privacy budget of each attribute field to obtain a desensitized dataset.
[0007] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described in any one of the first aspects above.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method as described in any one of the first aspects above.
[0009] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to perform the method described in any one of the first aspects above.
[0010] The beneficial effects of the embodiments in this application compared with the prior art are: This application embodiment inputs the dataset to be de-identified into an evaluation model to obtain the first contribution index of each attribute field to the first evaluation result output by the evaluation model, thereby obtaining the feature importance of each attribute field in the current evaluation; for each attribute field, the value weight is determined according to the first contribution index of the attribute field, so as to accurately allocate the corresponding value according to the feature importance of each attribute field and obtain an accurate value evaluation of each attribute field; the field sensitivity of the attribute field is obtained, which represents the sensitivity of the attribute field to changes in the dataset to be de-identified; based on the value weight and field sensitivity of the attribute field, the target privacy budget of the attribute field is determined, and the required privacy budget of the attribute field is analyzed more finely from two dimensions, that is, the privacy budget of the attribute field is different with different value weights and field sensitivity, so as to achieve non-uniform allocation of privacy budget according to the own situation of each attribute field, and can allocate more appropriate privacy budget to attribute fields with different values, solving the problem of not being able to adaptively allocate privacy budget to data; based on the target privacy budget of each attribute field, the attribute field is de-identified, which can perform different de-identification processing according to different privacy budgets, so as to achieve non-uniform de-identification processing of each attribute field, and obtain a de-identified dataset in which important information is not significantly disclosed and reflects its own characteristics.
[0011] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of the first type of data desensitization method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the second process of the data desensitization method provided in the embodiments of this application; Figure 3 This is a schematic diagram of the data desensitization device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0014] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0015] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0016] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0017] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0018] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0019] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0020] In data evaluation scenarios, evaluation models are used to analyze the statistical characteristics of data during the evaluation process. However, this data often involves personal and corporate information; for example, medical data includes patient medical records, and credit data includes user behavior information. Direct use of this data poses a risk of privacy breaches. For instance, in scenarios such as credit approval and fraud monitoring, financial institutions need to anonymize data containing sensitive attributes such as user income, transaction records, credit history, and debt ratios before using it to train credit scoring or fraud detection models. In medical data sharing and collaborative analysis scenarios, hospitals or research institutions need to anonymize patient electronic medical records (including diagnostic codes, examination indicators, medication records, etc.) before using them for training cross-institutional disease risk prediction models or clinical research. In scenarios where management agencies open their data to the public to promote public services (such as traffic flow prediction and public health monitoring), data containing citizens' personal information needs to be anonymized before release.
[0021] To address the risk of privacy breaches, existing technologies offer a differential privacy technique. Differential privacy achieves strict privacy protection by adding controllable noise to the data. Specifically, it employs a uniform privacy budget allocation model, injecting uniform noise into all data according to the same privacy budget. The change in the probability distribution of the data after noise injection does not exceed the privacy budget.
[0022] Privacy budgets are used to control the strength of privacy protection. A smaller privacy budget allows for greater injected noise and stronger privacy protection, but also lower data usability. Conversely, a larger privacy budget results in less injected noise, weaker privacy protection, and increased data usability, but also higher privacy risks. This can lead to a situation where, after uniformly adding noise to attribute fields with varying values and sensitivities in the evaluation results using the same privacy budget, high-value, high-sensitivity fields are excessively masked, while low-value, low-sensitivity fields receive unnecessary protection. In other words, differential privacy technology uniformly distributes the privacy budget across all data, thus uniformly adding noise, resulting in anonymized data that fails to reflect the true nature of the data, leading to problems in data evaluation.
[0023] To address the aforementioned issues, this application provides a data desensitization method that can allocate an appropriate privacy budget to attribute fields.
[0024] In some embodiments, the method is applied to an electronic device.
[0025] In some implementations, the method may be specifically executed by a data desensitization device, which may be implemented in software and / or hardware and may be configured in the aforementioned electronic equipment.
[0026] like Figure 1 The method shown includes steps S11 to S17, which will be described in detail below.
[0027] S11: Obtain the dataset to be anonymized.
[0028] The dataset to be anonymized includes attribute fields. The evaluation model is used to evaluate the value of the dataset to be anonymized according to the task requirements. The attribute fields are used to describe the characteristics of the data to be anonymized.
[0029] For example, the data to be anonymized is patient medical records, and the attribute fields include continuous, discrete, and ordinal fields. Continuous fields include age, blood glucose level, hospitalization costs, etc. Discrete fields include disease codes, etc. Ordinal fields include disease severity (1-5 levels), satisfaction rating (1-5 levels), etc.
[0030] In the application, electronic devices obtain a dataset to be anonymized by inputting data for value assessment from the user. All input data is obtained through legal and compliant means. The dataset to be anonymized is a collection of data containing important information that requires anonymization processing.
[0031] In some implementations, the user-input data may be incomplete. This can be addressed by preprocessing and feature engineering. For example, preprocessing includes missing value imputation, outlier handling, and deduplication, while feature engineering includes feature filtering and encoding.
[0032] For example, data processing: missing values in continuous fields are filled using multiple imputation, missing values in discrete fields are filled using the mode, and outliers are truncated using the IQR (Interquartile Range) method (interquartile range ± 1.5 times).
[0033] Feature engineering: Ordinal fields use label encoding, and discrete fields use one-hot encoding (target encoding is used when the number of categories is greater than 10).
[0034] S12: Determine the corresponding evaluation model based on the requirements of the first task.
[0035] The evaluation model is used to assess the value of the de-identified dataset according to the first task requirement. The first task requirement indicates the type of task the evaluation model is targeting, such as predicting risk or predicting disease.
[0036] Specifically, the required evaluation model is determined based on the requirements of the first task, and the evaluation model is used to evaluate the value of the dataset to be de-identified, and the first evaluation result is output.
[0037] For example, if the first task requirement is customer value assessment, select a customer lifetime value prediction model. If the first task requirement is customer credit assessment, select a customer credit scoring model. If the first task requirement is risk assessment, select a risk assessment model. If the first task requirement is disease assessment, select a disease prediction model.
[0038] S13: Input the dataset to be desensitized into the evaluation model and obtain the first contribution index of each attribute field to the first evaluation result output by the evaluation model.
[0039] In the application, the dataset to be de-identified is used to train and infer on the evaluation model to obtain the first contribution index of each attribute field to the first evaluation result.
[0040] In some implementations, the dataset to be de-identified can be pre-divided into training, testing, and validation sets. The training set is used to train and evaluate the model, while the validation set is used to monitor overfitting. When the RMSE (Root Mean Square Error) value of the validation set increases for several consecutive rounds, early stopping is triggered.
[0041] Calculate model performance metrics using the test set. For regression tasks, obtain the model performance metric by calculating the RMSE value. For classification tasks, obtain the model performance metric by calculating the AUC (Area Under the Curve) value.
[0042] Model training is complete when the difference between the performance on the test set and the performance on the training set is less than or equal to a preset performance value; otherwise, feature engineering or model hyperparameters are re-optimized.
[0043] In some implementations, when the evaluation model is a tree model, the degree of influence of the attribute field on the evaluation result can be obtained by calculating the SHAP value (Shapley Additive Explanations), and the SHAP value can be determined as the primary contribution indicator.
[0044] Specifically, when the evaluation model is a tree model or similar, the data to be anonymized for each attribute field is input into the trained evaluation model. Based on the first evaluation result, the SHAP value of each attribute field is calculated using the SHAP library. For each attribute field, the mean of the absolute values of the attribute field's SHAP values is determined as the primary contribution indicator.
[0045] For example, the data to be anonymized is divided into training, validation, and test sets according to a preset ratio (e.g., 7:2:1). The hyperparameters and validation methods of the tree model are set. An initial decision tree is fitted using the training set, and the difference between the first evaluation result and the true value for each data point to be anonymized is calculated. Then, a second decision tree is fitted based on the difference value to correct the prediction error of the first decision tree. That is, in the Mth iteration, the difference value between the first m-1 decision trees is fitted by the m-th decision tree.
[0046] In each iteration, during the construction of the decision tree, each tree node traverses all possible split points for each attribute field, selecting the feature and split threshold that maximizes the decrease in the tree model's loss function. After construction, the tree is validated using a validation set.
[0047] After reaching the required number of iterations, the tree model with the best validation performance is selected through a validation method. The model performance of the tree model with the best validation performance is evaluated using a test set. When the evaluation meets the expected results, the tree model is used as the evaluation model for completing the training.
[0048] Then, the data to be anonymized for each attribute field is input into the trained evaluation model to obtain the corresponding first evaluation result output by the evaluation model. When the evaluation model is a tree model (such as a gradient boosting tree or random forest), the first contribution index can be obtained in the following way: The SHAP value (Shapley Additive Explanations) of each attribute field is calculated using the SHAP library, and the mean of the absolute values of the SHAP values of each attribute field is determined as the first contribution index of that attribute field. The SHAP value is based on the Shapley value in game theory, which can fairly allocate the contribution of each attribute field to the prediction result. Its calculation process is known to those skilled in the art, and therefore will not be described in detail here.
[0049] When the evaluation model is a neural network, the degree of influence of the attribute field on the evaluation result can be obtained by calculating the integral gradient, that is, the integral gradient is determined as the primary contribution indicator.
[0050] For example, the dataset to be desensitized is divided into training, validation, and test sets according to a preset ratio (e.g., 7:2:1). A neural network structure is constructed, including an input layer (number of nodes equal to the number of attribute fields), several hidden layers (e.g., two layers, 64 neurons per layer, ReLU activation function), and an output layer (linear output for regression tasks, softmax output for classification tasks, depending on the task type). Backpropagation and an optimizer (e.g., Adam) are used for training, with the loss function selected according to the task type (mean squared error (MSE) for regression tasks, cross-entropy for classification tasks). During training, the validation set is used to monitor overfitting; early stopping is triggered when the validation set loss does not decrease for several consecutive rounds. After training, the model performance is evaluated on the test set to ensure the model achieves the expected results. For example, for a neural network with two attribute fields (blood glucose level, age), for a sample (blood glucose level = 7.5, age = 55), a baseline (blood glucose level = 5.0, age = 30) is selected, with an integration step count m = 100. 100 interpolation points are uniformly selected on the straight line from the baseline to the sample. The partial derivatives of the neural network output with respect to blood glucose level and age are calculated for each point. These derivatives are summed, multiplied by the difference, and then divided by 100 to obtain the integral gradient values of the two attribute fields for that sample. This process is repeated for all samples, and the mean of the integral gradients of each field is taken as the first contribution index.
[0051] Specifically, when the evaluation model is a neural network or similar, the data to be anonymized for each attribute field is input into the trained evaluation model. Based on the first evaluation result, the integral gradient of each attribute field is calculated. For each attribute field, the mean of the integral gradient is determined as the first contribution indicator.
[0052] S14: For each attribute field, determine the value weight of the attribute field based on the first contribution index of the attribute field.
[0053] Among them, the value weight is used to characterize the correlation strength between the attribute field and the first evaluation result.
[0054] In application, a weight calculation model is constructed based on the actual scenario. The value weight is then calculated using the weight calculation model based on the primary contribution indicator.
[0055] In some implementations, a single-factor weighting calculation model is constructed to weight the first contribution indicator and obtain the value weight.
[0056] In some implementations, a multi-factor weighting calculation model is constructed, which performs weighting based on attribute fields and the primary contribution indicator to obtain value weights. The weight coefficients of the factors can be set according to the actual scenario.
[0057] S15: Get the field sensitivity of the attribute field.
[0058] In applications, the corresponding calculation model can be determined in advance based on the type of the attribute field. Based on the corresponding calculation model, the field sensitivity of the attribute field is calculated. Then, the field sensitivity is obtained.
[0059] In some implementations, the sensitivity of a field is calculated by determining the maximum amount of change in the attribute field.
[0060] S16: Determine the target privacy budget for the attribute field based on its value weight and sensitivity.
[0061] In the application, a two-factor coupling model of value weight and field sensitivity is pre-established. Based on the two-factor coupling model, the value weight and field sensitivity of the attribute field are coupled and calculated to obtain the target privacy budget for the attribute field, realizing the non-uniform distribution of the privacy budget, which does not rely on a fixed budget or a single factor to uniformly distribute the privacy budget.
[0062] S17: Based on the target privacy budget for each attribute field, perform desensitization processing on each attribute field to obtain a desensitized dataset.
[0063] In the application, based on the target privacy budget of the attribute field, noise that meets the target privacy budget is identified and injected into the data to be de-identified in the attribute field to achieve de-identification of the attribute field, so as to generate an alternative dataset with statistical similarity but cannot be associated with the original individuals, i.e., a de-identified dataset, thus protecting the important information in the data to be de-identified.
[0064] This application embodiment inputs the dataset to be de-identified into an evaluation model to obtain the first contribution index of each attribute field to the first evaluation result output by the evaluation model, thereby obtaining the feature importance of each attribute field in the current evaluation; for each attribute field, the value weight is determined according to the first contribution index of the attribute field, so as to accurately allocate the corresponding value according to the feature importance of each attribute field and obtain an accurate value evaluation of each attribute field; the field sensitivity of the attribute field is obtained, which represents the sensitivity of the attribute field to changes in the dataset to be de-identified; based on the value weight and field sensitivity of the attribute field, the target privacy budget of the attribute field is determined, and the required privacy budget of the attribute field is analyzed more finely from two dimensions, that is, the privacy budget of the attribute field is different with different value weights and field sensitivity, so as to achieve non-uniform allocation of privacy budget according to the own situation of each attribute field, and can allocate more appropriate privacy budget to attribute fields with different values, solving the problem of not being able to adaptively allocate privacy budget to data; based on the target privacy budget of each attribute field, the attribute field is de-identified, which can perform different de-identification processing according to different privacy budgets, so as to achieve non-uniform de-identification processing of each attribute field, and obtain a de-identified dataset in which important information is not significantly disclosed and reflects its own characteristics.
[0065] Understandably, by determining the target privacy budget for each attribute field based on its value weight and sensitivity, and then performing non-uniform desensitization on each attribute field, a desensitized dataset that reflects its own characteristics without significant information leakage can be obtained. This eliminates the need for direct encryption computation, thus avoiding the high computational overhead associated with encryption technologies (such as homomorphic encryption and secure multi-party computation). Furthermore, by determining the required evaluation model based on task requirements and using the primary contribution of each attribute field to the first evaluation result as a factor in determining the privacy budget, non-uniform desensitization can be applied to various evaluation models, addressing the limitation of encryption technologies to machine learning-based evaluation models.
[0066] In some embodiments, step S14: For each attribute field, determining the value weight of the attribute field based on the first contribution index of the attribute field includes: S141: For each attribute field, normalize the first contribution index of the attribute field to obtain the third contribution index.
[0067] In application, in order to eliminate the difference in units, the first contribution index of the attribute field is normalized to obtain the third contribution index.
[0068] For example, normalization can be Min-Max normalization, Z-Score normalization, etc.
[0069] Taking Min-Max normalization as an example, its formula is: ×(U'-L')+L', where, This is the original value (first contribution indicator). Let [X] be the normalized value (third contribution index). min(X) and max(X) are the minimum and maximum values of the first contribution index among all attribute fields, respectively. L' and U' are the lower and upper bounds of the value range of the third contribution index. The range [L', U'] can be a fixed empirical value (e.g., [0, 1] or [0.01, 0.1]), or an expected range set based on prior estimates or historical experience of the performance impact index. For example, let the value range of the third contribution index be [L', U'] = [0.02, 0.32], and given that min(X) = 0.01 and max(X) = 0.5, =0.01, then = ×0.3+0.02=0.02.
[0070] S142: Metrics for evaluating the performance of the model based on computed attribute fields.
[0071] In some implementations, the performance degradation rate after the attribute field is missing or after the attribute field is replaced is calculated, and the performance degradation rate is determined as the performance impact indicator.
[0072] The calculation formula is , Let i be the performance degradation rate of the i-th attribute field. This refers to the RMSE value of the model after attribute fields are missing or after attribute fields have been replaced. This is the RMSE value of the model when the attribute fields are not missing or when the attribute fields are not replaced.
[0073] In some implementations, the performance increase rate after adding the attribute field is calculated, and the performance increase rate is determined as a performance impact indicator.
[0074] The calculation formula is , Let i be the performance improvement rate of the i-th attribute field. Add the model's RMSE value to the attribute field. This is the RMSE value of the model when the attribute field is not added.
[0075] S143: Determine the initial weight of the attribute field based on the third contribution index and performance impact index of the attribute field.
[0076] In some implementations, a two-factor weighting calculation model is constructed to sum the third contribution index and performance impact index of the attribute field in a weighted manner to obtain the initial weight of the attribute field.
[0077] The two-factor weighting calculation model is Wi=α×SHAPi′+(1 α)×Impi, where Wi is the initial weight, α is the balance coefficient, SHAPi′ is the third contribution index of the i-th attribute field, and Impi is the performance impact index of the i-th attribute field.
[0078] For example, the range of values for α can be determined through cross-validation, and the range is 0.4≤α≤0.6, with a default setting of 0.5.
[0079] S144: If the attribute field is a key field, then perform an enhancement operation on the initial weight of the attribute field to obtain the value weight of the attribute field.
[0080] S145: If an attribute field is not a key field, then the initial weight of the attribute field is determined as the value weight of the attribute field.
[0081] In applications, key fields can be marked according to predefined business rules. For example, in the medical field, diagnostic codes are marked as key fields; in the financial field, risk fraud patterns are marked as key fields; and in the industrial field, sensor readings exceeding the "safety threshold" are marked as key fields.
[0082] The enhancement operation involves weighting the initial weights based on the enhancement coefficient to obtain the value weight of the attribute field. For example, the enhancement coefficient is set to 1.1~2.0.
[0083] This application embodiment determines the initial weight of an attribute field based on its third contribution index and performance impact index. By constructing a two-factor model, the basic weight of the attribute field is determined more accurately, ensuring that the weight of the attribute field reflects the correlation strength with the first evaluation result. If the attribute field is a key field, the initial weight of the attribute field is enhanced to obtain the value weight of the attribute field. Attribute fields that are key fields have a higher correlation with the first evaluation result, and the weight of the key field is increased through weight enhancement to ensure that the weight of the key field reflects the correlation strength with the first evaluation result.
[0084] In some embodiments, step S15: obtaining the field sensitivity of the attribute field is determined based on the type of the attribute field, including: S151: If the attribute field is a continuous field, the maximum fluctuation value of the attribute field is obtained within the preset evaluation period, and the maximum fluctuation value is used as the field sensitivity of the attribute field.
[0085] For example, continuous fields include income, transaction amount, monthly consumption, etc.
[0086] In the application, the continuous time series data of the continuous field is traversed to obtain the actual maximum fluctuation value.
[0087] The evaluation period is set according to the business logic of the task. This evaluation period is then used to scan and traverse all time-series data to determine the maximum fluctuation value. This maximum fluctuation value is then used as the sensitivity of the attribute field. , For field sensitivity, This represents the maximum fluctuation value.
[0088] For example, the business logic is time-related, setting a sliding window (30-day consumption window) corresponding to a preset evaluation period. The sliding window is used to scan and traverse all time-series data, recording the maximum difference in the attribute field (total consumption) within each sliding window to determine the maximum fluctuation value.
[0089] S152: If the attribute field is a discrete field, then perform data transformation on the dataset to be de-identified to obtain multiple new probability distributions.
[0090] For example, discrete fields include occupational classification, disease code, customer type, etc.
[0091] In the application, the empirical probability distribution of the calculated attribute field in the dataset to be de-identified is obtained to obtain an initial probability distribution. Then, the operation of adding or deleting a single data entry with a certain value is simulated. By iterating through all the data entries for a certain value that have been added or deleted, multiple new probability distributions are obtained.
[0092] For example, simulate the operation of adding or deleting a record (with the attribute field value set to category c) to obtain a new empirical probability distribution, i.e., a new probability distribution.
[0093] S153: Calculate the distribution offset between each new probability distribution and the initial probability distribution.
[0094] The initial probability distribution is the probability distribution of the attribute field in the dataset to be de-identified.
[0095] In application, the KL divergence (Kullback-Leibler Divergence, relative entropy) between each new probability distribution and the initial probability distribution is calculated. The KL divergence is then defined as the distribution offset.
[0096] S154: Determine the field sensitivity of the attribute field based on the maximum offset among all distribution offsets.
[0097] In the application, all distribution offsets are traversed to determine the maximum offset. The maximum offset is then normalized to obtain the field sensitivity of the attribute field.
[0098] In some implementations, the normalization formula is: , For field sensitivity, denoted as KL divergence when a single data point changes, Dmin is the minimum offset among all distribution offsets, and Dmax is the maximum offset among all distribution offsets.
[0099] For example, simulating the operation of adding or deleting a record (with the attribute field value being category c), Dmax and Dmin are obtained after iterating through all the add or delete operations of category c.
[0100] S155: If the attribute field is an ordinal field, then perform data changes on the dataset to be de-identified to obtain multiple ordinal offsets.
[0101] For example, ordinal fields include education level, satisfaction rating, and credit rating. The ordinal values for education level are: Primary School = 1, Junior High School = 2, Senior High School = 3, Bachelor's Degree = 4, Master's Degree = 5, Doctoral Degree = 6. The ordinal values for satisfaction rating are 1-5.
[0102] In the application, the operation of adding or deleting a single data entry is simulated to obtain the ordinal offset caused by the change in that single data entry. By iterating through all the data entries and performing addition or deletion operations, multiple ordinal offsets are obtained.
[0103] For example, simulate adding or deleting a single piece of data, where the value of an attribute field changes from 'a' to 'b', and calculate the ordinal offset based on 'b' and 'a'.
[0104] The calculation formula is: Sequence offset = |Sequence (b) - Sequence (a)|.
[0105] S156: Determine the maximum sequence offset among all sequence offsets.
[0106] S157: Determine the field sensitivity of the attribute field based on the maximum ordinal offset and the corresponding ordinal offset weight.
[0107] In application, the corresponding ordinal offset weights can be pre-set according to the degree of influence of each ordinal offset on the first evaluation result. Then, the maximum ordinal offset and its corresponding ordinal offset weight are weighted and summed to obtain the field sensitivity of the attribute field.
[0108] The calculation formula is ; For field sensitivity, is the maximum ordinal offset, and k is the ordinal offset weight.
[0109] In some implementations, the order offset weight corresponding to the maximum order offset can be set directly. For example, in a satisfaction rating of 1-5, the satisfaction rating changes from 1 to 5. The maximum order offset is 4, so the order offset weight corresponding to the maximum order offset can be directly set to 5, which is 5 × 4 = 20.
[0110] In some implementations, the ordinal offset weight is a preset coefficient reflecting the change in importance between adjacent ordinal positions in an ordinal field. For ordinal fields, the cumulative offset weight from the original ordinal position to the target ordinal position is equal to the sum of the preset weights between adjacent ordinal positions along the path. This weight is used to quantify the impact of different ordinal changes on the evaluation results to calculate field sensitivity. This allows setting semantic distance weights between adjacent ordinal positions, and then calculating the path weights of the maximum ordinal offset based on these semantic distance weights, thus obtaining the maximum ordinal offset and its corresponding ordinal offset weight. In practical scenarios, the value of improving from 4 to 5 points is twice that of improving from 3 to 4 points. The semantic distance weights for satisfaction ratings changing from 1 to 2, from 2 to 3, from 3 to 4, and from 4 to 5 are [1.0, 1.0, 1.0, 2.0]. For a satisfaction rating changing from 1 to 5, the maximum ordinal offset is 4, and the path weight sum = 1.0 + 1.0 + 1.0 + 2.0 = 5.0, 5.0 × 4 = 20.
[0111] This application embodiment determines the corresponding sensitivity calculation model based on the field type of the attribute field, and calculates the field sensitivity that conforms to the field type based on the sensitivity calculation model, thus accurately calculating the field sensitivity of the attribute field.
[0112] In some embodiments, step S16: determining the target privacy budget for the attribute field based on the value weight and field sensitivity of the attribute field includes: S161: Calculate the sum of the products of the attribute field's value weight and the field's sensitivity to obtain the first calculation result.
[0113] In the application, the corresponding field sensitivity is obtained based on the field type of the attribute field. The value weight of the attribute field and the field sensitivity are multiplied to obtain the first calculation result.
[0114] S162: Calculate the sum of the products of the value weights and field sensitivity of all attribute fields to obtain the second calculation result.
[0115] In application, in order to eliminate the difference in units, the first operation result of the attribute field is normalized, and then the sum of the products of the value weights and field sensitivities of all attribute fields is calculated to provide a normalization factor.
[0116] In some implementations, basic privacy protection for attribute fields is ensured by calculating the sum of the products of the effective attribute field value weight and the field sensitivity to obtain the second calculation result.
[0117] S163: Determine the target privacy budget based on the preset total privacy budget, the first calculation result, and the second calculation result.
[0118] The preset total privacy budget is derived based on the preset privacy protection level.
[0119] In the application, basic privacy protection for attribute fields is ensured by setting a preset total privacy budget so that invalid attribute fields can be allocated a privacy budget. Invalid attribute fields are those with data quality issues, flawed desensitization rules, or broken field associations. Furthermore, a second calculation ensures that the sum of the privacy budgets for all fields equals the preset total privacy budget.
[0120] In some implementations, the calculation formula is: , For the target privacy budget of the i-th attribute field, To preset the total privacy budget, The value weight of the i-th attribute field. Let S be the field sensitivity of the i-th attribute field. , , , The result of the second operation is Ω, which represents the set of valid attribute fields. and Invalid attribute field ( or ).
[0121] For example, in medical data, rare pathological indicators are weighted with high value. High field sensitivity Due to the unique values of rare pathological indicators, they exhibit high sensitivity, and the corresponding first calculation result is 40. Common physiological indicators: low value weight. Low field sensitivity The first operation result is 1. The second operation result for all attribute fields is 100. Rare pathological indicators... Common physiological indicators The target privacy budget for attribute fields with high value weights and high field sensitivity is larger than that for attribute fields with low value weights and low field sensitivity. A larger privacy budget corresponds to a smaller amount of noise, reducing noise interference. As a result, after subsequent noise injection, attribute fields with high value weights and high field sensitivity are less affected by noise interference, better retaining the statistical information reflecting their own situation, while attribute fields with low value weights and low field sensitivity are adequately protected.
[0122] This application embodiment accurately distinguishes attribute fields of different values by calculating the sum of the products of the value weight and the sensitivity of the attribute fields; it calculates the sum of the products of the value weight and the sensitivity of all attribute fields, determines the target privacy budget based on the preset total privacy budget, the first calculation result, and the second calculation result, allocates privacy budgets to invalid attribute fields based on the preset total privacy budget, and allocates appropriate privacy budgets to each attribute field based on the first calculation result and the second calculation result, ensuring that the sum of the privacy budgets of all attribute fields does not exceed the preset total privacy budget. High-value attribute fields are allocated high privacy budgets, and low-value attribute fields are allocated low privacy budgets, providing a basis for better preserving statistical information reflecting their own situation for high-value attribute fields after noise injection, and for providing appropriate protection for low-value attribute fields.
[0123] In some embodiments, step S17: De-identifying each attribute field according to the target privacy budget of each attribute field to obtain a de-identified dataset, including: S171: Based on the target privacy budget of each attribute field, inject noise into each attribute field to obtain an intermediate dataset.
[0124] In the application, the corresponding noise generation mechanism is determined based on the requirements of the second task. Based on the noise generation mechanism, the corresponding noise amount is calculated according to the target privacy budget of the attribute field, and the noise is injected into the attribute field.
[0125] S172: Perform statistical feature calibration on the intermediate dataset to obtain the desensitized dataset.
[0126] In the application, the attribute fields of the intermediate dataset are obtained and their statistical features are analyzed. The corresponding calibration model is determined based on the field type of the attribute fields. Based on the corresponding calibration model, the attribute fields are statistically calibrated according to their statistical features to obtain anonymized data for the attribute fields.
[0127] This application embodiment obtains an intermediate dataset by injecting noise into each attribute field according to the target privacy budget of each attribute field; then, it performs statistical feature calibration on the intermediate dataset to obtain a de-identified dataset. Different noise amounts are injected according to different privacy budgets to achieve non-uniform noise injection into each attribute field. At the same time, combined with statistical feature calibration of the intermediate dataset, the overall distribution drift of the intermediate dataset is corrected to reduce the distribution deviation between the de-identified dataset and the dataset to be de-identified, thereby reducing the value error of each attribute field in the de-identified dataset. This reduces the performance loss of the de-identified data to the evaluation model and also ensures the performance of the de-identified dataset in complex evaluation models.
[0128] In some embodiments, step S171: injecting noise into each attribute field according to the target privacy budget of each attribute field, including: S21: Select a noise injection strategy based on the requirements of the second task.
[0129] The second task requirement is used to indicate the statistical characteristics of the data to be retained, including the distribution pattern of the data records or the statistical characteristics of data aggregation.
[0130] The second task requires determining a noise injection strategy to meet the requirement of preserving statistical properties.
[0131] S22: For each attribute field, if the attribute field is not a field in the strongly correlated field set, and the second task requires preserving the distribution pattern of data records, then for each original value of the attribute field, the sampling probability of each candidate value is determined based on the target privacy budget of the attribute field, the utility score between the original value and each candidate value in the candidate set, and the global sensitivity.
[0132] The strongly correlated field set includes at least one attribute field. The correlation value between each attribute field in the strongly correlated field set is greater than a preset correlation threshold. The original value is the value of the attribute field in the dataset to be de-identified. The candidate set is a subset of the global value space of the attribute field. The utility score is determined based on the utility function, according to the similarity between the original value and the candidate value and the value weight of the attribute field. The global sensitivity is the sensitivity of the utility function on the neighboring datasets of the dataset to be de-identified.
[0133] For example, the second task requirement for classification, clustering, and other tasks is to preserve the distribution pattern of data records.
[0134] In application, the relevant values of all numerical fields (after continuous and ordinal encoding) and quantifiable discrete fields can be pre-calculated. Among them, the quantifiable discrete fields are obtained through feature engineering, such as one-hot encoding.
[0135] When the absolute value of the correlation between any two attribute fields is greater than a preset correlation threshold, these attribute fields are determined to be fields in the set of strongly correlated fields.
[0136] In some implementations, correlation values are calculated using Pearson correlation analysis.
[0137] For example, for continuous fields, the Pearson correlation value is calculated directly with a preset correlation threshold of 0.7. For ordinal fields, their ordinal codes can be retained first, and then the Spearman rank correlation coefficient can be calculated. The preset correlation threshold can also be set to 0.7. For discrete fields, the mutual information value between two discrete fields can be calculated first, and then normalized by dividing by the entropy of a single field to obtain the normalized mutual information (value range [0,1]). The preset correlation threshold can be set to 0.7.
[0138] Among them, the numerical representation of multi-category discrete fields is achieved through category frequency encoding, which means replacing the category value with the proportion of the category in the dataset, thereby reducing the dimensionality explosion of one-hot encoding.
[0139] When an attribute field is not in the strongly correlated field set, noise can be injected separately. The second task requirement, preserving the data record distribution pattern, employs an exponential noise generation mechanism. A candidate set can be pre-set according to the actual scenario requirements, and the sampling probability between each candidate value and the original value can be determined based on the exponential mechanism.
[0140] In some implementations, the exponential mechanism: sampling probability and , For the target privacy budget of the i-th attribute field, For utility function, r is the candidate value, and x is the original value. Let be the value weight of the i-th attribute field. For continuous fields, the similarity function is: The similarity between discrete / ordinal number segments is determined based on the degree of encoding. If they match, the similarity is 1; if they do not match, the similarity is 0. For global sensitivity.
[0141] In some implementations, because The maximum change is 1, which can be set. .
[0142] pass This makes it more likely that attribute fields with high value weights will obtain candidate values similar to the original values.
[0143] S23: Based on the sampling probability of each value in the candidate set, determine the first noise corresponding to the original value, and replace the original value with the first noise.
[0144] In the application, sampling is performed from the candidate set based on the sampling probability of each value. The sampled candidate values are determined as the first noise, and the original values are replaced with the first noise to obtain intermediate values, thus injecting noise into the attribute field. The intermediate values are the values of the attribute field in the intermediate dataset.
[0145] S24: If the attribute field is not a field in the strongly correlated field set, and the second task requires preserving the data aggregation and statistical characteristics, then generate second noise based on the field sensitivity of the attribute field, the target privacy budget, and the data dimension, and inject the second noise into the attribute field.
[0146] For example, the second task requirement for regression tasks, prediction tasks, etc., is to preserve the aggregated statistical properties of the data.
[0147] In applications, when an attribute field is not a field in a strongly correlated set of fields, noise can be injected separately. To preserve the statistical characteristics of data aggregation, a second requirement is to employ a truncated Laplace noise generation mechanism. Based on the truncated Laplace mechanism, second noise is generated according to the attribute field's sensitivity, target privacy budget, and data dimension, and then injected into the attribute field.
[0148] In some implementations, the Laplace mechanism is truncated: noise scale parameter , Let S be the field sensitivity of the i-th attribute field. , , `dim` represents the number of data dimensions participating in the current aggregation calculation or requiring joint consideration. The specific rules for determining this number are as follows: If the second task requires preserving the statistical characteristics of the data aggregation and calculating the marginal statistics (such as mean, variance, quantiles) of each attribute field, then `dim=1`. If calculating the joint statistics (such as covariance, correlation coefficient) between attribute fields, then `dim` equals the number of attribute fields participating in the joint statistical calculation. This number is determined by the feature dependency structure of the evaluation model or the statistical requirements specified by the user. For example, if the number of joint fields is 3, then `dim=3`. Let i be the target privacy budget for the i-th attribute field.
[0149] pass This allows us to use the differential privacy sequence combination theorem. As a sensitivity correction item for high-dimensional scenarios, it ensures that the privacy budget does not exceed the preset total privacy budget when multiple fields are jointly statistically analyzed.
[0150] The noise scaling parameter λ determines the amplitude of the generated noise; the larger λ is, the larger the absolute value of the generated noise. Based on the noise scaling parameter λ, Laplace noise n with a mean of 0 and a scaling parameter λ (probability density function is...) is generated. The second noise is obtained. The original value x is added to the second noise n to obtain an intermediate value, thus achieving noise injection.
[0151] In some implementations, in order to obtain the intermediate value that conforms to the reasonable value range of the attribute field, the data after injecting noise is truncated to the reasonable value range [L, U] of the attribute field to obtain the intermediate value, x´_final=max(L, min (U, x´_prelim)), where x´_final is the intermediate value and x´_prelim is the data after injecting noise.
[0152] Where L and U are defined by domain knowledge or determined from data estimation. For example, domain knowledge definition: for age, L=0, U=120. Data estimation determination: L=historical minimum, U=historical maximum.
[0153] S25: If the attribute field is a field in the strongly correlated field set, then determine the combined sensitivity of the strongly correlated field set based on the field sensitivity of each attribute field in the strongly correlated field set; determine the combined privacy budget of the strongly correlated field set based on the target privacy budget of each attribute field in the strongly correlated field set; generate third noise based on the combined sensitivity, combined privacy budget, each attribute field in the strongly correlated field set and its corresponding correlation value, and inject the third noise into the strongly correlated field set.
[0154] In applications, when an attribute field is a field in a strongly correlated field set, noise injection is required. The sensitivity of each attribute field in the strongly correlated field set is obtained, and this sensitivity is then normalized using Min-Max to obtain the normalized sensitivity. Finally, the combined sensitivity is calculated based on the normalized sensitivity of each attribute field.
[0155] In some implementations, the L2 norm of the normalized sensitivity of each attribute field is calculated, and the L2 norm is determined as the combined sensitivity.
[0156] Obtain the minimum privacy budget from the target privacy budgets of each attribute field in the set of strongly correlated fields, and use the minimum privacy budget as the combined privacy budget.
[0157] Determine the standard deviation of each attribute field in the strongly correlated field set, and construct an original covariance matrix based on the standard deviation and correlation value of each attribute field. Then, adjust the original covariance matrix according to the combined sensitivity and combined privacy budget to obtain a noise covariance matrix that satisfies the differential privacy Gaussian mechanism. The multidimensional Gaussian noise generated based on this matrix is identified as the third noise, and this third noise is injected into the strongly correlated field set. For example, the strongly correlated field set includes two strongly correlated attribute fields, X and Y. The original covariance matrix is constructed based on the standard deviation and correlation value of each attribute field. , Let X be the standard deviation of the attribute field. Let be the standard deviation of attribute field Y, and ρ be the correlation value. After dimensional adaptation, we obtain the noise covariance matrix that satisfies the differential privacy Gaussian mechanism. ,in, For combined sensitivity, To incorporate privacy budgets. , , Let be the normalized sensitivity of the i-th attribute field. This is used to balance the sensitivity contribution of individual attribute fields in a set of strongly correlated fields, and to avoid covariance distortion caused by inconsistent dimensions.
[0158] It is understandable that a set of strongly correlated fields includes two strongly correlated attribute fields whose third noise conforms to a two-dimensional Gaussian distribution, while a set of strongly correlated fields includes multiple strongly correlated attribute fields whose third noise conforms to a multidimensional Gaussian distribution.
[0159] Understandably, through , , To determine the normalized sensitivity of the i-th attribute field, the covariance matrix of the initial joint noise is adjusted for dimensionality adaptation. The diagonal elements of the covariance matrix of the third noise are... This is proportional to the original covariance matrix of the attribute field, ensuring that the noise amplitude after injection matches the field sensitivity and is controlled by the field sensitivity, rather than by the original dimensions of the attribute field. And through... The total variance of noise controlling the third noise and It is proportional to the third noise so that the differential privacy Gaussian mechanism theory is satisfied.
[0160] Based on the above, a third type of noise is injected into the set of strongly correlated fields to ensure that the covariance structure of the set of strongly correlated fields is proportionally maintained with the original covariance structure. This ensures that the attribute fields in the set of strongly correlated fields after the noise injection maintain a linear correlation, and that the correlation of each attribute field in the set of strongly correlated fields after the noise injection is not lost. This solves the problem of loss of correlation between fields after injecting isotropic Gaussian noise into strongly correlated attribute fields, and makes the evaluation model trained on the data after the noise injection more closely resemble the effect of training on the data to be desensitized.
[0161] For example, in consumer data, the correlation between "purchase frequency" and "average order value" is greater than the preset correlation threshold. Adding noise to each attribute field independently may destroy this relationship, leading to misjudgment by the evaluation model. By jointly injecting noise into "purchase frequency" and "average order value", the linear correlation between "purchase frequency" and "average order value" is maintained, making the evaluation model trained on the data with injected noise closer to the effect of training on the data to be desensitized.
[0162] This application embodiment adapts a noise generation mechanism to the second task requirements when the attribute field is not in the strongly correlated field set. Appropriate noise is generated based on this mechanism to ensure that the intermediate value of the attribute field is similar to the original value without significantly revealing important information. This makes the evaluation model trained on the intermediate dataset closer to the effect of training on the data to be anonymized. When the attribute field is in the strongly correlated field set, a third noise is generated based on the combined sensitivity, combined privacy budget, each attribute field in the strongly correlated field set, and its corresponding correlation value. The covariance matrix of the third noise is proportional to the original covariance matrix of the attribute field, ensuring that the noise amplitude after injection matches the field sensitivity. It is not dominated by the original dimensions of the attribute field but controlled by the field sensitivity, maintaining a linear correlation between attribute fields in the strongly correlated field set after noise injection. This makes the evaluation model trained on the noise-injected data closer to the effect of training on the data to be anonymized, improving the accuracy of the evaluation model, especially the accuracy of the evaluation model dependent on the correlation between features.
[0163] In some embodiments, step S172: statistical feature calibration of the intermediate dataset to obtain a desensitized dataset includes: S31: For each attribute field, if the attribute field is a continuous field, then obtain the original value of the attribute field at the preset baseline quantile.
[0164] In this application, a quantile matching correction mechanism is used for continuous fields. A preset baseline quantile is set, and the corresponding raw value is obtained at the preset baseline quantile.
[0165] S32: Based on the original value of the preset baseline quantile, perform linear stretching mapping on the intermediate value of the attribute field to obtain the desensitized data of the attribute field.
[0166] The intermediate value is the value of the attribute field in the intermediate dataset.
[0167] In the application, the median value at the same preset baseline quantile is determined. A linear ratio is determined based on the original value and the median value at the same quantile, and the median value of the attribute field is linearly stretched and mapped according to the linear ratio to obtain the de-identified data.
[0168] For example, the preset baseline quantiles include 5%, 50%, 75%, and 95% quantiles. Obtain the original values for the 5%, 50%, 75%, and 95% quantiles. The original value for the 75% quantile is 100,000 yuan, and the median value for the 75% quantile is 85,000 yuan. The linear ratio is 85,000 / 10 = 0.85. Based on this linear ratio, the median value greater than 85,000 yuan is scaled proportionally to ensure that the quantiles of the anonymized data in the attribute field match the quantiles of the data to be anonymized in terms of distribution pattern, quantiles, and marginal proportions.
[0169] S33: If the attribute field is a discrete field, determine the initial marginal distribution of the attribute field in the dataset to be de-identified.
[0170] In this application, a marginal distribution-constrained resampling mechanism is employed for discrete fields. The proportion of each value of the attribute field in the dataset to be anonymized is obtained. Based on the proportion of each value of the attribute field, the initial marginal distribution of the attribute field in the dataset to be anonymized is determined.
[0171] S34: Based on the initial marginal distribution, resample the attribute fields of the intermediate dataset to obtain a resampled dataset of the attribute fields.
[0172] In the application, corresponding constraint resampling conditions are determined based on the initial marginal distribution, including: marginal distribution constraints and covariance constraints. The attribute fields of the intermediate dataset are then resampled based on these constraint resampling conditions to obtain a resampled dataset of the attribute fields.
[0173] S35: If the target value exists in the resampled data, select the target number of data from the resampled dataset to obtain the data to be replaced.
[0174] The target value is the value at which the proportional deviation after resampling exceeds the allowable error, and the target quantity is determined based on the proportional deviation after resampling.
[0175] In this application, the resampled proportion deviation of each attribute field value in the resampled data is calculated between the proportion of each value and the proportion of the corresponding original value. This resampled proportion deviation is then compared to the allowable error. When the resampled proportion deviation exceeds the allowable error, the target value is obtained. Based on the resampled proportion deviation of the target value, the number of data points to be selected is determined, resulting in the target quantity. Finally, the target quantity of data points whose current value is not equal to the target value is randomly selected from the resampled dataset to obtain the data to be replaced.
[0176] Among them, the values of the attribute fields differ from the target values for the target number of data.
[0177] In some implementations, the tolerance for error is determined based on the value weight of the attribute field. The tolerance for error is set to ≤1% for high-value attribute fields, ≤3% for medium-value attribute fields, and ≤5% for low-value attribute fields.
[0178] For example, the tolerance for the attribute field "diabetes" is 1%. The proportion of "diabetes" in the disease code is 12%, and the proportion of "diabetes" after resampling is 9%. The proportion deviation after resampling is 12%-9%=3%, 3%>1%, so the number of data points corresponding to 3% is determined as the target number. Randomly select 3% of "non-diabetes" data to obtain the data to be replaced.
[0179] S36: Replace the value of the attribute field in the data to be replaced with the target value to obtain the replaced data of the attribute field.
[0180] S37: If the distribution of other attribute fields in the replaced data conforms to the corresponding initial marginal distribution, then the desensitized data of the attribute fields is obtained.
[0181] Some implementations also include: S38: If the distribution of other attribute fields in the replaced data does not conform to the corresponding initial marginal distribution, return to step S35. If the target value exists in the resampled data, select the target number of data from the resampled dataset to obtain the data to be replaced and the subsequent steps, until the retry termination condition is met.
[0182] In some implementations, the iteration termination condition can be set as follows: the distribution of other attribute fields in the replaced data conforms to the corresponding initial marginal distribution, satisfying the retry termination condition, resampling ends, and the currently obtained de-identified attribute field data is output. Alternatively, the iteration can be terminated when the number of enrichments reaches the preset number of retries (H_max), i.e., H = H_max, satisfying the retry termination condition, resampling ends, and the currently obtained de-identified attribute field data is output.
[0183] In the application, the retry status is compared with the retry termination condition. If the iteration termination condition is not met, resampling continues, and the process returns to step S35 and subsequent steps. If the iteration termination condition is met, resampling ends, and step S37 is executed to obtain the de-identified data of the attribute fields.
[0184] In application, the distribution of other attribute fields in the replaced data is detected by the chi-square test model to ensure that the distribution of other attribute fields in the replaced data conforms to the corresponding initial marginal distribution, thereby reducing the distortion of the overall distribution of the de-identified attribute fields and making the overall distribution of the de-identified attribute fields match the overall distribution of the data to be de-identified.
[0185] In some implementations, the other attribute fields in the replaced data can be selected from those strongly correlated with the original attribute field, or those with moderate correlation. Then, an attribute field contingency table is constructed. The chi-square statistic of the other attribute fields is calculated based on this contingency table. If the chi-square statistic of the other attribute fields is less than a preset threshold, it is determined that the distribution of the other attribute fields in the replaced data conforms to the corresponding initial marginal distribution, thus obtaining the desensitized data for the attribute fields. If the chi-square statistic of the other attribute fields is greater than or equal to the preset threshold, it is determined that the distribution of the other attribute fields in the replaced data does not conform to the corresponding initial marginal distribution. In this case, the target number of data points needs to be reselected until the distribution of the other attribute fields in the replaced data conforms to the corresponding initial marginal distribution.
[0186] In some implementations, the chi-square statistic is calculated using the following formula: χ² is the chi-square statistic, O is the observed frequency of other attribute fields in the replaced data, E is the expected frequency of other attribute fields, and E = (total number of rows × total number of columns) / total number of records. The total number of rows, total number of columns, and total number of records are the statistical data of the attribute field contingency table.
[0187] The preset critical value is obtained based on the chi-square statistic of the degrees of freedom. df = (number of rows - 1) × (number of columns - 1), where df is the degrees of freedom, and the number of rows and columns are the statistical data of the contingency table of attribute fields. This is a preset critical value, specifically the critical value of a chi-square distribution with df degrees of freedom at a significance level of α=0.05.
[0188] For example, the attribute fields strongly correlated with diabetes include blood glucose level and age. A two-dimensional contingency table of blood glucose level and age is constructed, with 2 rows and 2 columns. The chi-square statistics of blood glucose level and age are calculated based on this two-dimensional contingency table. df = (2-1) × (2-1) = 1 3.2 < 3.841, the distribution of other attribute fields in the replaced data conforms to the corresponding initial marginal distribution.
[0189] This application embodiment selects an appropriate statistical feature calibration model based on the field type of the attribute field, and performs accurate statistical feature calibration on the attribute field based on the appropriate statistical feature calibration model. This corrects the distribution drift of the attribute field and ensures that the anonymized data after statistical feature calibration matches the original data in terms of distribution shape, quantiles, and marginal proportions. This allows the anonymized data after statistical feature calibration to satisfy differential privacy while effectively controlling the distortion of the overall distribution and reducing the evaluation error caused by noise accumulation.
[0190] It is understandable that by injecting joint noise into attribute fields based on fields in a strongly correlated set of attribute fields, the correlation between high-dimensional features caused by individual noise addition is reduced, thus ensuring the correlation between high-dimensional features. Furthermore, by calibrating the statistical features of attribute fields according to their field types, the de-identified data can possess the micro-statistical characteristics of the data to be de-identified. This can solve the problem of the destruction of the micro-statistical characteristics and the correlation between high-dimensional features caused by the generalization or suppression of original values by generalization / anonymization techniques (such as k-anonymity and L-diversity).
[0191] In some embodiments, such as Figure 2 After obtaining the de-identified dataset, as shown, the process also includes: S41: Determine whether the iteration termination condition is met.
[0192] In some implementations, the iteration termination condition can be set to the deviation being less than or equal to a preset deviation threshold, that is when τ, the iteration termination condition is satisfied, the iteration ends, and the currently obtained desensitized data set is output. Or when the number of iterations reaches the preset number of iterations (K_max), that is K = K_max, but when τ, the iteration termination condition is satisfied, the iteration ends, and the currently obtained desensitized data set is output.
[0193] In the application, compare the iteration situation with the iteration termination condition. If τ or K < K_max, it is determined that the iteration termination condition is not satisfied, and the iteration continues, and step S42 is continued to be executed. If τ or K = K_max, it is determined that the iteration termination condition is satisfied, the iteration ends, step S46 is executed, and the currently obtained desensitized data set is output.
[0194] S42: If not, input the desensitized data set into the evaluation model to obtain the second evaluation result output by the evaluation model.
[0195] S43: Calculate the deviation between the first evaluation result and the second evaluation result.
[0196] In the application, compare the first evaluation result with the second evaluation result, and calculate the deviation between the first evaluation result and the second evaluation result.
[0197] For example, in regression tasks such as sales prediction, etc., determine the deviation according to the relative RMSE value of the evaluation model. The calculation formula is , where δ is the deviation, is the RMSE value of the desensitized data set, is the RMSE value of the data set to be desensitized.
[0198] For classification tasks such as credit scoring, disease diagnosis, etc., determine the deviation according to the AUC (Area Under the Roc Curve) difference of the evaluation model. The calculation formula is , where δ is the deviation, is the AUC value of the desensitized data set, is the AUC value of the data set to be desensitized.
[0199] S44: If the deviation is greater than the preset deviation threshold, obtain the second contribution degree index of each attribute field to the second evaluation result.
[0200] In the application, a preset deviation threshold is set in advance so that the evaluation result output by the evaluation model meets the scenario requirements.
[0201] In some implementations, the Kendall correlation coefficient is used to set an acceptable preset deviation threshold. For example, setting a preset deviation threshold... .
[0202] If the deviation exceeds the preset deviation threshold, it indicates that the obtained anonymized dataset contains information that is leaked or fails to reflect its own characteristics, and the anonymized dataset needs to be regenerated. Obtain the second contribution index of each attribute field to the second evaluation result.
[0203] S45: For each attribute field, calculate the difference in contribution between the first contribution index and the second contribution index of the attribute field, as well as the deviation contribution to the deviation.
[0204] In the application, the first contribution index and the second contribution index of the attribute field are compared, and the contribution difference between the first contribution index and the second contribution index of the attribute field is calculated.
[0205] In some implementations, the formula for calculating the contribution difference is: , The difference in contribution. As the second contribution indicator, It is the primary contribution indicator.
[0206] Modify the values of the attribute fields in the de-identified dataset and calculate the contribution of the attribute fields to the bias.
[0207] In some implementations, the modification method can be to replace the value of the attribute field in the de-identified dataset with the statistical characteristic value of the attribute field in the dataset to be de-identified.
[0208] For example, for continuous fields, replace the value of the attribute field in the de-identified dataset with the mean of that attribute field in the dataset to be de-identified. For discrete fields, replace the value of the attribute field in the de-identified dataset with the mode of that attribute field in the dataset to be de-identified. For ordinal fields, replace the value of the attribute field in the de-identified dataset with the median of that attribute field in the dataset to be de-identified.
[0209] In some implementations, the formula for calculating the deviation contribution is Ci=δ δ´ and Ci represent the contribution of the deviation. δ is the deviation before modification, and δ´ is the deviation after modification. The larger Ci is, the greater the contribution of this field to the deviation.
[0210] S46: Determine the value weight of the attribute field based on the contribution difference and deviation contribution of the attribute field, and return to the execution step S16.
[0211] In the application, the weight adjustment amount is determined based on the difference in contribution of the attribute fields. Based on the weight adjustment amount and the bias contribution, the value weight of the attribute field is re-determined, and the process returns to the step of determining the target privacy budget of the attribute field based on the value weight and field sensitivity, as well as subsequent steps (S16-S45).
[0212] In some implementations, if the contribution difference of an attribute field is negative, and the decrease in the contribution difference exceeds a first preset level, it indicates that the budget allocated to the attribute field may be insufficient, requiring an increase in the attribute field's evaluation value weight, resulting in a corresponding weight adjustment. If the change in the contribution difference of an attribute field is small, the attribute field's evaluation value weight is maintained, resulting in a corresponding weight adjustment. If the increase in the contribution difference of an attribute field exceeds a second preset level, it indicates that the budget allocated to the attribute field may be excessive, requiring a decrease in the attribute field's evaluation value weight, resulting in a corresponding weight adjustment.
[0213] For example, the contribution difference can be set to within ±5% as a small change.
[0214] In some implementations, the formula for calculating the value weight of attribute fields is redefined. , For the new value weight of the i-th attribute field, Let be the adjusted value weight of the i-th attribute field. , The value weight of the i-th attribute field. η is the weight adjustment amount. η is the learning rate for weight adjustment. The sign function is +1 when Ci > 0 and -1 when Ci < 0, and α is the smoothing factor (0 < α ≤ 1, usually α = 0.5). α is used to balance the historical weights and the current deviation contribution to avoid abrupt changes in value weights.
[0215] Example, deviation When the value is >20%, η=0.2, enabling rapid adjustment; when the value is <10%, η=0.2. When ≤20%, η=0.1, achieving routine adjustment; When the percentage is ≤10%, η=0.05, achieving fine adjustment.
[0216] S47: If so, output the desensitized dataset.
[0217] In application, When τ is reached, the iteration termination condition is met, the iteration ends, and the current high-quality de-identified dataset is output. When K = K_max, the iteration termination condition is met, the iteration ends, and the current de-identified dataset with the smallest deviation is output.
[0218] In some implementations, after obtaining the desensitized dataset with the minimum current deviation, a prompt can be set for the user to further optimize by increasing the preset total privacy budget or adjusting the preset deviation threshold.
[0219] In the embodiments of the present application, when the iteration termination condition is not met, the value weight of the attribute field is determined according to the contribution degree difference of the attribute field and the deviation contribution degree, and the steps of determining the target privacy budget of the attribute field according to the value weight of the attribute field and the field sensitivity are returned and subsequent steps are executed, so as to automatically learn and adjust the optimal privacy protection strategy, adaptively generate desensitized datasets that meet different evaluation models and data distributions, and improve the practicability.
[0220] In the embodiments of the present application, The value range of is 0.001 ≤ ≤ 0.5. The value of S_cont is consistent with the field numerical magnitude. The value range of S_dist is 0 < S_dist ≤ 1 (after KL divergence normalization). The value range of S_ord is S_ord ≥ 1. The value range of is 0 < ≤ 10. The value range of is 0 < < . The value range of dim is dim ≥ 1. The value range of λ is λ > 0. The value range of is ≥ 0. The value range of ρ is 1 ≤ ρ ≤ 1. The value range of η is 0.05 ≤ η ≤ 0.2. The value range of K_max is K_max ≥ 1. The value range of is ≥ 0. The value range of each normalized value is [0, 1]. The value range of is > 0. The value range of is 0 < < . The value range of is 0 ≤ ≤ 1.
[0221] To better understand the data desensitization method described in the embodiments of the present application, a specific example is provided for illustration.
[0222] Using medical data asset assessment as a scenario, the dataset to be anonymized contains 100,000 patient records. Attribute fields include: continuous fields (age, blood glucose level, hospitalization costs), discrete fields (disease code, medical insurance type), and ordinal fields (disease severity: 1-5, satisfaction score: 1-5). The assessment model is a disease risk prediction model, specifically a chronic disease risk prediction model. This model, constructed based on a gradient boosting tree, is used to evaluate the utility value of disease risk prediction based on the 100,000 patient records to be anonymized.
[0223] Set a preset total privacy budget =1.0, preset deviation threshold τ=10% (RMSE≤10%), preset correlation threshold |ρ|=0.7, maximum number of iterations K_max=5, minimum value L'=0.02, maximum value U'=0.32 of the third contribution index range, [L',U']=[0.02,0.32], correspondingly, ×(0.32-0.02)+0.02= ×0.3+0.02.
[0224] The specific construction of the gradient boosting tree model: Data partitioning: The original dataset was divided into training set, validation set, and test set in a 7:2:1 ratio.
[0225] Model hyperparameters: learning rate = 0.1, tree depth = 5, number of leaf nodes = 30, number of iterations = 100, regularization coefficient = 0.01.
[0226] Validation method: 5-fold cross-validation, with the goal of minimizing the RMSE of the validation set to select the optimal model.
[0227] The disease risk prediction model was trained using a dataset to be anonymized. Specifically, the dataset was divided into training, validation, and test sets in a 7:2:1 ratio. Based on 5-fold cross-validation, the training set was divided into five non-overlapping subsets. Five rounds of training were performed using these five subsets. In each round, four subsets were randomly selected as the training data, and the remaining subset was used as the internal validation data, ensuring that each round had at least one subset as the internal validation set for evaluation.
[0228] In each training round of each fold cross-validation iteration, a new decision tree is constructed using four subsets to approximate the difference value obtained in the previous iteration. The construction of the decision tree is subject to model hyperparameter constraints. After 100 iterations, the remaining subset and the validation set are used to validate the decision tree obtained in that round, thus obtaining the validation results.
[0229] After five rounds of training, the tree model with the best validation effect is selected through validation. The model performance of the tree model with the best validation effect is evaluated using the test set. When the evaluation effect is good, the tree model is used as the evaluation model to complete the training, thus realizing the training of the disease risk prediction model using the dataset to be desensitized.
[0230] The first contribution indicators were obtained by calculating SHAP values: disease severity (0.50), blood glucose level (0.435), disease code (0.271), age (0.173), hospitalization cost (0.075), satisfaction score (0.026), and medical insurance type (0.01). Based on the first contribution indicators of all the above attribute fields, the minimum value min(X) = 0.01 and the maximum value max(X) = 0.50 were determined. Then, Min-Max normalization was performed on the first contribution indicators of each attribute field according to the formula... = The third contribution index value is calculated by multiplying by 0.3 and adding 0.02. The third contribution index values for each attribute field are as follows: severity of illness (0.32), blood glucose level (0.28), disease code (0.18), age (0.12), hospitalization cost (0.06), satisfaction score (0.03), and medical insurance type (0.02).
[0231] Meanwhile, the performance degradation rate (as a performance impact indicator) after deleting attribute fields was measured as follows: blood glucose level (0.36), disease severity (0.30), disease code (0.253), age (0.08), hospitalization cost (0.06), satisfaction score (0.03), and medical insurance type (0.02).
[0232] The initial weights are calculated using a two-factor weighted calculation model, with α=0.5: For example, the initial weight of blood glucose value = 0.5×0.28+0.5×0.36=0.32; the initial weight of disease severity = 0.5×0.32+0.5×0.30=0.31; the initial weight of disease code (including rare disease characteristics) = 0.5×0.18+0.5×0.253=0.2165; the initial weight of age = 0.5×0.12+0.5×0.08=0.10; the initial weight of hospitalization expenses = 0.5×0.06+0.5×0.06 =0.06; the initial weight of satisfaction score = 0.5×0.03+0.5×0.03 =0.03; and the initial weight of medical insurance type = 0.5×0.02 +0.5×0.02 =0.02.
[0233] Since the disease code is a key field, its initial weight is augmented (assuming an augmentation coefficient of 1.2), resulting in an augmented weight of 0.2165 × 1.2 ≈ 0.26. Other fields are not augmented; their initial weights are their value weights. The final value weights for each attribute field are: disease severity (0.31), blood glucose level (0.32), disease code (0.26), age (0.10), hospitalization costs (0.06), satisfaction rating (0.03), and medical insurance type (0.02).
[0234] For continuous fields: the sliding window for age is set to 5 years, with a maximum fluctuation of 45 (S_cont=45); the sliding window for blood glucose is set to 7 days, with a maximum fluctuation of 12.3 (S_cont=12.3); and the sliding window for hospitalization expenses is set to 30 days, with a maximum fluctuation of 50000 (S_cont=50000). Discrete fields: The maximum offset of KL divergence for disease codes is 0.8, and the field sensitivity is 0.8 (S_dist=0.8); the maximum offset of KL divergence for medical insurance types is 0.2, and the field sensitivity is 0.2 (S_dist=0.2).
[0235] Ordinal fields: Severity of illness is coded in order from 1 to 5, with a uniform semantic weight of 1.0, a maximum ordinal offset of 4, a sum of path weights of 4, and a field sensitivity of 16 (S_ord=4×4=16); Satisfaction rating is coded in order from 1 to 5, with a semantic distance weight of [1.0,1.0,1.0,2.0], a maximum ordinal offset of 4, a sum of path weights of 5.0, and a field sensitivity of 20 (S_ord=4×5.0=20).
[0236] The sensitivity of continuous, discrete, and ordinal fields is multiplied with their value weights to obtain the first result.
[0237] Calculate the first result for each attribute field: Severity of illness: 0.31 × 16 = 4.96; Blood glucose level: 0.32 × 12.3 ≈ 3.936; Disease code: 0.26 × 0.8 = 0.208; Age: 0.10 × 45 = 4.5; Hospitalization cost: 0.06 × 50000 = 3000; Satisfaction score: 0.03 × 20 = 0.6; Medical insurance type: 0.02 × 0.2 = 0.004. Calculate the second result: =4.96+3.936+0.208+4.5+3000+0.6+0.004≈3014.208; Calculate the target privacy budget for each attribute field: Hospitalization cost ε = 3000 / 3014.208 ≈ 0.995; Age ε ≈ 4.5 / 3014.208 ≈ 0.0015; Blood glucose level ε ≈ 3.936 / 3014.208 ≈ 0.0013; Severity of illness ε ≈ 4.96 / 3014.208 ≈ 0.0016; Disease code ε ≈ 0.208 / 3014.208 ≈ 0.00007; Satisfaction score ε ≈ 0.6 / 3014.208 ≈ 0.0002; Medical insurance type ε ≈ 0.004 / 3014.208 ≈ 0.0000013 (according to the rule, ε_min = 0.001).
[0238] The assessment task was disease risk prediction (regression task), using a truncated Laplace mechanism: Hospitalization costs: =50000, =0.995, dim=1, λ=(50000×√1) / 0.995≈50251; truncated to [0,500000] after noise injection. Age: =45, =0.0015, dim=1, λ=(45×√1) / 0.0015=30000; truncated to [0,120] after noise injection. Blood glucose value: =12.3, =0.0013, dim=1, λ=(12.3×√1) / 0.0013≈9461.5; after noise injection, it is truncated to [2.8, 20.0].
[0239] The correlation coefficient ρ between blood glucose levels and disease severity is 0.78 > 0.7, forming a strongly correlated set of data. (Blood glucose levels) =0.0013, severity of illness =0.0016, =min(0.0013,0.0016)=0.0013; blood glucose level =12.3, = (12.3-0.2) / (50000-0.2)≈0.000246; Severity of the condition =16, = (16-0.2) / (50000-0.2)≈0.00031; =√(0.000246²+0.000316²)≈0.0004; The initial joint noise covariance matrix Σ=[[2.56, 1.8], [1.8, 1.21]], =0.000246 / 0.0004≈0.615, = 0.000316 / 0.0004 ≈ 0.79; The covariance matrix Σ´ of the third noise = (0.0004² / 0.0013²) × diag([0.615, 0.79]) × Σ ≈ (1.6×10^-7 / 1.69×10^-6) × [[0.615×2.56, 0.615×1.8], [0.79×1.8, 0.79×1.21]] ≈ 0.0947 × [[1.5744, 1.107], [1.422, 0.9559]] ≈ [[0.149, 0.105], [0.135, 0.0906]]. Generate a two-dimensional Gaussian noise vector, that is, the third noise, and inject the third noise into the set of strongly correlated fields between blood glucose values and disease severity.
[0240] The preset benchmark quantiles of the dataset to be desensitized: the 5% quantile of age = 18, the 50% quantile = 52, the 95% quantile = 82; the median values of the same quantiles: the 5% quantile of age = 19, the 50% quantile = 55, the 95% quantile = 85; perform a linear stretching mapping on the median values of the attribute fields: adjust the 50% quantile to 52, and scale the other values proportionally.
[0241] The proportion of the disease code "diabetes" in the dataset to be desensitized is 12%, and the proportion after desensitization is 9%; randomly select 3% (3000 records) of non-diabetes records and replace them with diabetes codes, and perform a chi-square test: select blood glucose values (≥7.0 mmol / L) and age (40 - 70 years old) as associated fields, construct a two-dimensional contingency table, calculate χ² = 3.2, the degree of freedom df = (2 - 1)(2 - 1) = 1, χ²0. 05 (1) = 3.841, because 3.2 < 3.841, it conforms to the original joint distribution, and obtain the desensitized data of the attribute fields.
[0242] Input the desensitized dataset into the disease risk prediction model to obtain the second evaluation result = 12.5%. = 2.3%, the deviation δ = 12.5% - 2.3% = 10.2% > τ = 10%, and the number of iterations = 1 < K_max = 5, enter the weight adjustment.
[0243] Calculate the contribution difference of the attribute fields: the contribution difference of blood glucose value decreases by 25%, and increase its weight to 0.38. Calculate the contribution C of blood glucose value = 10.2% - 8.0% = 2.2% (the deviation after replacing the original blood glucose value is 8.0%), η = 0.1, α = 0.5, =0.38 + 0.1×1×2.2%^0.5≈0.38+0.1×0.148≈0.395. Similarly, other attribute fields: severity of illness (0.30), disease code (0.25), age (0.09), hospitalization costs (0.05), satisfaction rating (0.03), and medical insurance type (0.02). Return to steps S15-S45. After the second iteration, the newly generated de-identified dataset has an RMSE of 9.8%≤10%, meeting the accuracy requirements. The iteration ends, and the currently obtained high-quality de-identified dataset is output.
[0244] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. Furthermore, the data collection in the above embodiments is compliant, and its use or implementation does not involve any infringement upon public interests.
[0245] For ease of explanation, only the parts related to the embodiments of this application are shown in the description of the methods described in the above embodiments.
[0246] In some embodiments, such as Figure 3 As shown, the device includes: The first acquisition module 10 is used to acquire the dataset to be de-identified. The dataset to be de-identified includes attribute fields, which are used to describe the characteristics of the data to be de-identified. It is also used to determine the corresponding evaluation model according to the first task requirements, the first task requirements are used to indicate the target evaluation task type of the evaluation model, and the evaluation model is used to evaluate the value of the dataset to be de-identified according to the first task requirements. The second acquisition module 11 is used to input the dataset to be de-identified into the evaluation model and obtain the first contribution index of each attribute field to the first evaluation result output by the evaluation model.
[0247] The weight calculation module 12 is used to determine the value weight of each attribute field based on the first contribution index of the attribute field. The value weight is used to characterize the correlation strength between the attribute field and the first evaluation result.
[0248] Sensitivity calculation module 13 is used to obtain the field sensitivity of the attribute field.
[0249] Privacy budget calculation module 14 is used to determine the target privacy budget for an attribute field based on the value weight and sensitivity of the attribute field.
[0250] The data desensitization module 15 is used to desensitize each attribute field according to the target privacy budget of each attribute field to obtain a desensitized dataset.
[0251] In some embodiments, the data desensitization module includes a noise injection unit and a statistical feature calibration unit.
[0252] The noise injection unit is used to inject noise into each of the attribute fields according to the target privacy budget of each attribute field to obtain an intermediate dataset.
[0253] A statistical feature calibration unit is used to perform statistical feature calibration on the intermediate dataset to obtain the desensitized dataset.
[0254] In some embodiments, the noise injection unit is specifically used to select a noise injection strategy according to a second task requirement, the second task requirement being used to indicate the data statistical characteristics to be retained, including retaining the data record distribution pattern or retaining the data aggregation statistical characteristics. Specifically, for each attribute field, if the attribute field is not a field in the strongly correlated field set, and the second task requirement is to preserve the distribution pattern of data records, then for each original value of the attribute field, the sampling probability of each candidate value is determined based on the target privacy budget of the attribute field, the utility score between the original value and each candidate value in the candidate set, and the global sensitivity. The strongly correlated field set includes at least one attribute field, and the correlation value between each attribute field in the strongly correlated field set is greater than a preset correlation threshold. The original value is the value of the attribute field in the dataset to be de-identified. The candidate set is a subset of the global value space of the attribute field. The utility score is determined based on the utility function, according to the similarity between the original value and the candidate value and the value weight of the attribute field. The global sensitivity is the sensitivity of the utility function on the neighboring datasets of the dataset to be de-identified. The noise injection unit is specifically used to determine the first noise corresponding to the original value based on the sampling probability of each value in the candidate set, and replace the original value with the first noise. The noise injection unit is specifically used to generate second noise and inject it into the attribute field if the attribute field is not a field in the strongly correlated field set and the second task requirement is to preserve the data aggregation and statistical characteristics. The noise injection unit is specifically used to determine the combined sensitivity of the strongly correlated field set based on the field sensitivity of each attribute field in the strongly correlated field set if the attribute field is a field in the strongly correlated field set; determine the combined privacy budget of the strongly correlated field set based on the target privacy budget of each attribute field in the strongly correlated field set; generate third noise based on the combined sensitivity, combined privacy budget, each attribute field in the strongly correlated field set and its corresponding correlation value, and inject the third noise into the strongly correlated field set.
[0255] In some embodiments, the statistical feature calibration unit is specifically used to obtain the original value of the attribute field of the preset benchmark quantile for each attribute field if the attribute field is a continuous field. The statistical feature calibration unit is specifically used to perform linear stretching mapping on the intermediate value of the attribute field based on the original value of the preset benchmark quantile to obtain the desensitized data of the attribute field. The intermediate value is the value of the attribute field in the intermediate dataset. The statistical feature calibration unit is specifically used to determine the initial marginal distribution of the attribute field in the dataset to be desensitized if the attribute field is a discrete field. The statistical feature calibration unit is specifically used to resample the attribute fields of the intermediate dataset according to the initial marginal distribution to obtain a resampled dataset of the attribute fields. The statistical feature calibration unit is specifically used to select the target number of data from the resampled dataset if there is a target value in the resampled data, and obtain the data to be replaced. The target value is the value where the proportion deviation after resampling exceeds the allowable error. The target number is determined based on the proportion deviation after resampling. The statistical feature calibration unit is specifically used to replace the value of the attribute field in the data to be replaced with the target value, so as to obtain the replaced data of the attribute field; The statistical feature calibration unit is specifically used to obtain desensitized data for attribute fields if the distribution of other attribute fields in the replaced data conforms to the corresponding initial marginal distribution.
[0256] In some embodiments, the weight calculation module is specifically used to normalize the first contribution index of each attribute field to obtain a third contribution index. The weight calculation module is specifically used to calculate the performance impact of attribute fields on the evaluation model. The weight calculation module is specifically used to determine the initial weight of an attribute field based on its third contribution metric and performance impact metric. The weight calculation module is specifically used to enhance the initial weight of an attribute field if the attribute field is a key field, thereby obtaining the value weight of the attribute field. The weight calculation module is specifically used to determine the initial weight of an attribute field as its value weight if the attribute field is not a key field.
[0257] In some embodiments, the privacy budget calculation module is specifically used to calculate the sum of the products of the value weight of the attribute field and the sensitivity of the field to obtain a first calculation result; The privacy budget calculation module is specifically used to calculate the sum of the products of the value weights of all attribute fields and the field sensitivity, and obtain the second calculation result; The privacy budget calculation module is specifically used to determine the target privacy budget based on the preset total privacy budget, the first calculation result, and the second calculation result. The preset total privacy budget is obtained based on the preset privacy protection strength.
[0258] In some embodiments, the apparatus further includes an iterative optimization module.
[0259] The iterative optimization module is used to input the desensitized dataset into the evaluation model and obtain the second evaluation result output by the evaluation model; The iterative optimization module is used to calculate the deviation between the first evaluation result and the second evaluation result; The iterative optimization module is used to obtain the second contribution index of each attribute field to the second evaluation result if the deviation is greater than the preset deviation threshold. The iterative optimization module is used to calculate the difference in contribution between the first contribution index and the second contribution index of each attribute field, as well as the deviation contribution to the deviation, for each attribute field. The iterative optimization module is used to determine the value weight of the attribute field based on the contribution difference and deviation contribution of the attribute field, and return to execute the steps of determining the target privacy budget of the attribute field based on the value weight and field sensitivity of the attribute field, and subsequent steps, until the iteration termination condition is met.
[0260] In some embodiments, the sensitivity calculation module is specifically used to obtain the maximum fluctuation value of the attribute field within a preset evaluation period if the attribute field is a continuous field, and use the maximum fluctuation value as the field sensitivity of the attribute field. The sensitivity calculation module is specifically used to perform data transformations on the dataset to be desensitized if the attribute field is a discrete field, thereby obtaining multiple new probability distributions. The sensitivity calculation module is specifically used to calculate the distribution offset between each new probability distribution and the initial probability distribution. The initial probability distribution is the probability distribution of the attribute field in the dataset to be desensitized. The sensitivity calculation module is specifically used to determine the field sensitivity of an attribute field based on the maximum offset among all distribution offsets. The sensitivity calculation module is specifically used to perform data changes on the dataset to be desensitized if the attribute field is an ordinal field, and obtain multiple ordinal offsets. The sensitivity calculation module is specifically used to determine the maximum sequence offset among all sequence offsets; The sensitivity calculation module is specifically used to determine the field sensitivity of an attribute field based on the maximum ordinal offset and the corresponding ordinal offset weight.
[0261] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4As shown, the electronic device 2 of this embodiment includes: at least one processor 20 ( Figure 4 (Only one is shown in the diagram), memory 21, and computer program 22 stored in said memory 21 and executable on said at least one processor 20, wherein said processor 20 executes said computer program 22 to implement the steps in any of the above method embodiments.
[0262] The electronic device 2 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. The electronic device 2 may include, but is not limited to, a processor 20 and a memory 21. Those skilled in the art will understand that... Figure 4 This is merely an example of electronic device 2 and does not constitute a limitation on electronic device 2. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0263] The processor 20 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0264] In some embodiments, the memory 21 may be an internal storage unit of the electronic device 2, such as a hard disk or memory of the electronic device 2. In other embodiments, the memory 21 may be an external storage device of the electronic device 2, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 2. Furthermore, the memory 21 may include both internal and external storage units of the electronic device 2. The memory 21 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 21 can also be used to temporarily store data that has been output or will be output.
[0265] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0266] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0267] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps in the above-described method embodiments.
[0268] This application provides a computer program product that, when run on an electronic device, enables the electronic device to implement the steps described in the various method embodiments above.
[0269] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some cases, the computer-readable medium cannot be an electrical carrier signal or a telecommunication signal.
[0270] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0271] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0272] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0273] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0274] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A data de-sensitization method, characterized in that, include: Obtain the dataset to be anonymized, wherein the data to be anonymized in the dataset includes attribute fields, which are used to describe the characteristics of the data to be anonymized. The corresponding evaluation model is determined based on the first task requirement, which is used to indicate the target evaluation task type of the evaluation model. The evaluation model is used to evaluate the value of the dataset to be de-identified according to the first task requirement. The dataset to be de-identified is input into the evaluation model to obtain the first contribution index of each attribute field to the first evaluation result output by the evaluation model; For each attribute field, the value weight of the attribute field is determined according to the first contribution index of the attribute field, and the value weight is used to characterize the correlation strength between the attribute field and the first evaluation result; Obtain the field sensitivity of the attribute field; The target privacy budget for the attribute field is determined based on the value weight and the field sensitivity of the attribute field. Based on the target privacy budget of each attribute field, the attribute fields are anonymized to obtain an anonymized dataset; The step of determining the target privacy budget for the attribute field based on its value weight and sensitivity includes: Calculate the sum of the products of the value weight and the sensitivity of the attribute field to obtain the first calculation result; Calculate the sum of the products of the value weights and the field sensitivity of all the attribute fields to obtain the second calculation result; The target privacy budget is determined based on the preset total privacy budget, the first calculation result, and the second calculation result. The preset total privacy budget is obtained based on a preset privacy protection strength.
2. The method of claim 1, wherein, The step of de-identifying each attribute field according to the target privacy budget of each attribute field to obtain a de-identified dataset includes: Based on the target privacy budget of each attribute field, noise is injected into each attribute field to obtain an intermediate dataset; The intermediate dataset is subjected to statistical feature calibration to obtain the desensitized dataset.
3. The method of claim 2, wherein, Based on the target privacy budget for each attribute field, noise injection is performed on each attribute field, including: Select a noise injection strategy based on the second task requirement, which indicates the statistical characteristics of the data to be retained, including retaining the distribution pattern of data records or retaining the aggregated statistical characteristics of data. For each attribute field, if the attribute field is not a field in the strongly correlated field set, and the second task requirement is to preserve the distribution pattern of data records, then for each original value of the attribute field, the sampling probability of each candidate value is determined based on the target privacy budget of the attribute field, the utility score between the original value and each candidate value in the candidate set, and the global sensitivity. The strongly correlated field set includes at least one attribute field, the correlation value between each attribute field in the strongly correlated field set is greater than a preset correlation threshold, the original value is the value of the attribute field in the dataset to be de-identified, the candidate set is a subset of the global value space of the attribute field, the utility score is determined based on the utility function, according to the similarity between the original value and the candidate value and the value weight of the attribute field, and the global sensitivity is the sensitivity of the utility function on the neighboring datasets of the dataset to be de-identified. Based on the sampling probability of each value in the candidate set, determine the first noise corresponding to the original value, and replace the original value with the first noise; If the attribute field is not a field in the strongly correlated field set, and the second task requirement is to retain the data aggregation and statistical characteristics, then a second noise is generated based on the field sensitivity of the attribute field, the target privacy budget, and the data dimension, and the second noise is injected into the attribute field. If the attribute field is a field in the strongly correlated field set, then the combined sensitivity of the strongly correlated field set is determined based on the field sensitivity of each attribute field in the strongly correlated field set; the combined privacy budget of the strongly correlated field set is determined based on the target privacy budget of each attribute field in the strongly correlated field set; and a third noise is generated based on the combined sensitivity, the combined privacy budget, each attribute field in the strongly correlated field set, and the corresponding correlation value, and the third noise is injected into the strongly correlated field set.
4. The method of claim 2, wherein, The step of performing statistical feature calibration on the intermediate dataset to obtain the de-identified dataset includes: For each attribute field, if the attribute field is a continuous field, then the original value of the attribute field at the preset benchmark quantile is obtained; Based on the original value of the preset baseline quantile, the intermediate value of the attribute field is linearly stretched and mapped to obtain the desensitized data of the attribute field, wherein the intermediate value is the value of the attribute field in the intermediate dataset. If the attribute field is a discrete field, then determine the initial marginal distribution of the attribute field in the dataset to be de-identified; Based on the initial marginal distribution, the attribute fields of the intermediate dataset are resampled to obtain a resampled dataset of the attribute fields; If a target value exists in the resampled data, then a target number of data is selected from the resampled dataset to obtain the data to be replaced. The target value is the value where the ratio deviation after resampling exceeds the allowable error. The target number is determined based on the ratio deviation after resampling. The value of the attribute field in the data to be replaced is replaced with the target value to obtain the replaced data of the attribute field; If the distribution of other attribute fields in the replaced data conforms to the corresponding initial marginal distribution, then the de-identified data of the attribute fields is obtained.
5. The method according to any one of claims 1 to 4, characterized in that, The step of determining the value weight of each attribute field based on the first contribution index of the attribute field includes: For each of the attribute fields, the first contribution index of the attribute field is normalized to obtain the third contribution index; Calculate the performance impact index of the attribute field on the evaluation model; The initial weight of the attribute field is determined based on the third contribution index and the performance impact index of the attribute field; If the attribute field is a key field, then the initial weight of the attribute field is enhanced to obtain the value weight of the attribute field. If the attribute field does not belong to the key field, then the initial weight of the attribute field is determined as the value weight of the attribute field.
6. The method of claim 1, wherein, After obtaining the de-identified dataset, the process also includes: The desensitized dataset is input into the evaluation model to obtain the second evaluation result output by the evaluation model; Calculate the deviation between the first evaluation result and the second evaluation result; If the deviation is greater than a preset deviation threshold, then the second contribution index of each attribute field to the second evaluation result is obtained; For each of the attribute fields, calculate the contribution difference between the first contribution index and the second contribution index of the attribute field, as well as the deviation contribution to the deviation. Based on the contribution difference of the attribute field and the deviation contribution, determine the value weight of the attribute field, and return to execute the step of determining the target privacy budget of the attribute field based on the value weight of the attribute field and the field sensitivity, and subsequent steps, until the iteration termination condition is met.
7. The method of claim 6, wherein, The field sensitivity of the attribute field is determined based on the type of the attribute field, including: If the attribute field is a continuous field, the maximum fluctuation value of the attribute field is obtained within a preset evaluation period, and the maximum fluctuation value is used as the field sensitivity of the attribute field. If the attribute field is a discrete field, then the dataset to be de-identified is subjected to data transformation to obtain multiple new probability distributions; Calculate the distribution offset between each of the new probability distributions and the initial probability distribution, wherein the initial probability distribution is the probability distribution of the attribute field in the dataset to be de-identified; The field sensitivity of the attribute field is determined based on the maximum offset among all the distribution offsets; If the attribute field is an ordinal field, then the dataset to be de-identified is modified to obtain multiple ordinal offsets; Determine the largest sequence offset among all the aforementioned sequence offsets; The field sensitivity of the attribute field is determined based on the maximum ordinal offset and the corresponding ordinal offset weight.
8. A data de-identification apparatus, comprising: include: The first acquisition module is used to acquire the dataset to be de-identified, wherein the data to be de-identified in the dataset includes attribute fields, and the attribute fields are used to describe the characteristics of the data to be de-identified. It is also used to determine the corresponding evaluation model according to the first task requirements, the first task requirements are used to indicate the target evaluation task type of the evaluation model, and the evaluation model is used to evaluate the value of the dataset to be de-identified according to the first task requirements. The second acquisition module is used to input the dataset to be desensitized into the evaluation model and acquire the first contribution index of each attribute field to the first evaluation result output by the evaluation model. The weight calculation module is used to determine the value weight of each attribute field based on the first contribution index of the attribute field, wherein the value weight is used to characterize the correlation strength between the attribute field and the first evaluation result. A sensitivity acquisition module is used to acquire the field sensitivity of the attribute field; A privacy budget calculation module is used to determine the target privacy budget for the attribute field based on the value weight and the field sensitivity of the attribute field. The data desensitization module is used to desensitize each attribute field according to the target privacy budget of each attribute field to obtain a desensitized dataset; The privacy budget calculation module is specifically used to calculate the sum of the products of the value weight and the sensitivity of the attribute field to obtain a first calculation result. The privacy budget calculation module is specifically used to calculate the sum of the products of the value weights and the sensitivity of all the attribute fields to obtain a second calculation result; The privacy budget calculation module is specifically used to determine the target privacy budget based on the preset total privacy budget, the first calculation result, and the second calculation result. The preset total privacy budget is obtained based on a preset privacy protection strength.
9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.