Data sample privacy risk grading method and system

By using the member inference attack model to evaluate the privacy risks of data samples in the member inference attack scenario, the problem of incomplete and accurate evaluation in the existing technology is solved, targeted protection is achieved, and the security of the machine learning model is improved.

CN119961976APending Publication Date: 2025-05-09BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063592.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the case of member reasoning attack, it is difficult to comprehensively and accurately evaluate the privacy risks of data samples in the existing technology, resulting in the inability to targeted protection, and there is a risk of privacy leakage.

Method used

By using the member inference attack model to attack the target model, the gradient information, neighborhood density and output probability of the data sample are combined, and the privacy risk score of the data sample is calculated and the privacy risk level is divided according to the score.

Benefits of technology

It realizes comprehensive and accurate assessment of data sample privacy risks in member reasoning attack scenarios, and can formulate targeted protection strategies, which significantly improves the security of machine learning models in sensitive data application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961976A_ABST
    Figure CN119961976A_ABST
Patent Text Reader

Abstract

The invention discloses a data sample privacy risk grading method and system, and the method comprises the steps: selecting a plurality of member reasoning attack models for simulating the behaviors of an attacker to attack a target model when the target model is trained through a data sample; calculating gradient information and neighborhood density of a data sample of the target model in a training process and an output probability output by the member reasoning attack model under the attack of the member reasoning attack model; performing normalization processing on the gradient information, the neighborhood density and the output probability which are obtained through calculation; and calculating a privacy risk score of the data sample according to the normalized gradient information, the neighborhood density and the output probability, and dividing a privacy risk level of the data sample according to the privacy risk score of the data sample. According to the grading method, gradient information and neighborhood density generated in the training process of a target model and the output probability output by a member reasoning attack model are utilized to comprehensively evaluate the privacy risk of a data sample in a member reasoning attack scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of privacy risk assessment for machine learning, and in particular to a method and system for grading privacy risks of data samples. Background Art

[0002] Membership Inference Attack (MIA) allows attackers to determine whether a data sample is used for model training by observing the output of the model. This type of attack is likely to lead to data privacy leakage in sensitive fields such as medicine and finance. For example, this type of attack can infer whether the patient's relevant information is involved in the training of a disease prediction or analysis model, thereby exposing the patient's condition information.

[0003] The National Institute of Standards and Technology (NIST) of the United States pointed out that MIA violates the principle of data confidentiality. At the same time, according to the EU General Data Protection Regulation (GDPR), such attacks may lead to illegal leakage of personal data. Therefore, it is urgent to design a method that can quantify and classify the privacy risks of data samples for MIA scenarios and propose an effective data sample privacy protection strategy.

[0004] Currently, there are few studies on evaluating the privacy risk of data samples under membership inference attack scenarios, and related research focuses on the following aspects:

[0005] 1. Privacy risk assessment based on information theory: Use information theory indicators such as information entropy, mutual information, conditional entropy, etc. to quantify privacy risks, which are usually suitable for measuring the effectiveness of data anonymization or protection mechanisms (such as k-anonymity, l-diversity).

[0006] 2. Privacy risk assessment based on attack simulation: Simulate the behavior of real attackers, evaluate privacy risks based on the attacker's success rate, and then conduct a quantitative assessment of the privacy risks of the data set or model.

[0007] 3. Privacy risk assessment based on game theory: Use the game theory framework to model the interaction between attackers and defenders, analyze the effectiveness of the privacy protection mechanism, and evaluate the possibility of privacy leakage and the attacker's benefits by solving the Nash equilibrium.

[0008] Although the above methods can be used for privacy risk assessment, there is still a lack of privacy risk assessment methods for data samples in reasoning attack scenarios. At the same time, the existing privacy risk assessment methods based on attack simulation only rely on the success rate of the attacker to assess the privacy risk of the data sample, which will lead to the privacy risk assessment of the data sample not being comprehensive and accurate enough, and thus failing to protect the data sample in a targeted manner.

[0009] Therefore, a method is needed to comprehensively and accurately evaluate the privacy risk of data samples in membership inference attack scenarios. Summary of the invention

[0010] In view of the deficiencies in the prior art, the present invention provides a data sample privacy risk grading method and system, which uses a member reasoning attack model to attack a target model, and utilizes the gradient information and neighborhood density of data samples generated by the target model during the training process, as well as the output probability of the member reasoning attack model output, to comprehensively evaluate the privacy risk of data samples under the member reasoning attack scenario.

[0011] The first object of the present invention is to provide a method for grading privacy risks of data samples, comprising:

[0012] When the target model is trained with data samples, multiple member reasoning attack models for simulating attacker behaviors are selected to attack the target model;

[0013] Calculate the gradient information of the data samples of the target model during the training process, the neighborhood density, and the output probability of the member reasoning attack model output under the attack of the member reasoning attack model;

[0014] Normalize the calculated gradient information, neighborhood density, and output probability respectively;

[0015] According to the normalized gradient information, neighborhood density, and output probability, the privacy risk score of the data sample is calculated, and the privacy risk level of the data sample is divided according to the privacy risk score of the data sample.

[0016] As a further improvement of the present invention, the member reasoning attack model includes: a member reasoning attack model based on a binary classifier, a member reasoning attack model based on predicted entropy, and a member reasoning attack model based on confidence.

[0017] As a further improvement of the present invention, the gradient information of the data sample is:

[0018]

[0019] Among them, x i is the data sample, θ is the parameter of the target model, L(x i ,θ) is the loss function.

[0020] As a further improvement of the present invention, the neighborhood density of the data sample is calculated using the k-nearest neighbor concept combined with local density and global density, and the calculation formula is:

[0021]

[0022] ρ(x i )=α·ρ local (x i )+(1-α)·ρ global (xi )

[0023] Where k is the data sample x i The number of nearest neighbors in the dataset, v_k(x i ) is a set of k data samples x i The minimum sphere volume of the nearest neighbors, ρ local (x i ) is the data sample x i The local density in the data set, ρ global (x i ) is the data sample x i The global density in the data set, α is the balance coefficient, ρ(x i ) is the data sample x i The neighborhood density of .

[0024] As a further improvement of the present invention, the calculation formula of the output probability output by the member reasoning attack model is:

[0025]

[0026] Among them, w j The weight of the attack model for the jth member inference, The data sample x output by the j-th member inference attack model i The output probability of whether it belongs to the training set, p combined (x i ) is the data sample x considering the weight of each member's inference attack model i The output probability of whether it belongs to the training set, m is the total number of member reasoning attack models, The data samples x output by the attack model for all members i The mean of the output probabilities of whether they belong to the training set, σ 2 (x i ) is the data sample x output by different members’ reasoning attack models i The variance of the output probability of whether it belongs to the training set, a(x i ) is the stability weighting coefficient, p final (x i ) is the data sample x after considering stability i The output probability of whether it belongs to the training set.

[0027] As a further improvement of the present invention, the privacy risk score of the calculated data sample is obtained by weighted calculation of normalized gradient information, neighborhood density, and output probability.

[0028] As a further improvement of the present invention, an adaptive optimization algorithm is used to dynamically adjust the normalized gradient information, neighborhood density, and weight coefficient of the output probability.

[0029] As a further improvement of the present invention, the step of dividing the privacy risk levels of the data samples according to the privacy risk scores of the data samples includes:

[0030] Using clustering methods to set thresholds for privacy risk scores;

[0031] If the privacy risk score of a data sample is greater than or equal to the threshold, the data sample is classified as a high-risk sample;

[0032] If the privacy risk score of a data sample is less than the threshold, the data sample is classified as a low-risk sample.

[0033] As a further improvement of the present invention, after dividing the privacy risk levels of data samples according to their privacy risk scores, data samples classified as high-risk samples are replaced with similar data generated by a generative adversarial network, or data features of the data samples are dynamically perturbed.

[0034] The second object of the present invention is to provide a data sample privacy risk grading system for implementing the above-mentioned grading method, comprising:

[0035] An attack model selection module, used for selecting a plurality of member reasoning attack models for simulating attacker behaviors to attack the target model when the target model is trained with data samples;

[0036] The data processing module is used to calculate the gradient information and neighborhood density of the data samples of the target model during the training process under the attack of the member reasoning attack model, as well as the output probability of the member reasoning attack model output; and normalize the calculated gradient information, neighborhood density, and output probability respectively;

[0037] The risk level classification module is used to calculate the privacy risk score of the data sample based on the normalized gradient information, neighborhood density, and output probability, and to classify the privacy risk level of the data sample based on the privacy risk score of the data sample.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The target model is attacked by using the membership reasoning attack model. The gradient information and neighborhood density of the data samples generated by the target model during the training process, as well as the output probability of the membership reasoning attack model are comprehensively considered. A multi-dimensional evaluation method is used to calculate the privacy risk score of each data sample to comprehensively evaluate the privacy risk of data samples under the membership reasoning attack scenario.

[0040] The privacy risk level of data samples is divided according to their privacy risk scores, providing a basis for the subsequent formulation and implementation of privacy protection strategies, thereby achieving targeted protection of data samples with different privacy risks, significantly improving the security of machine learning models in sensitive data application scenarios, and providing strong guarantees for the compliant use of data and privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A flowchart of the grading method is provided;

[0042] Figure 2 The overall flow chart of the grading method is as follows;

[0043] Figure 3 A schematic diagram of the grading system. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0046] See also Figure 1 , Figure 2 This embodiment provides a data sample privacy risk classification method, including:

[0047] When the target model is trained with data samples, multiple member reasoning attack models for simulating attacker behaviors are selected to attack the target model;

[0048] Calculate the gradient information of the data samples of the target model during the training process, the neighborhood density, and the output probability of the member reasoning attack model output under the attack of the member reasoning attack model;

[0049] Normalize the calculated gradient information, neighborhood density, and output probability respectively;

[0050] According to the normalized gradient information, neighborhood density, and output probability, the privacy risk score of the data sample is calculated, and the privacy risk level of the data sample is divided according to the privacy risk score of the data sample.

[0051] Use the member reasoning attack model to attack the target model, comprehensively consider the gradient information, neighborhood density, and output probability of the data samples generated by the target model during the training process, so as to comprehensively evaluate the privacy risks of data samples under the member reasoning attack scenario and provide a basis for the subsequent formulation and implementation of privacy protection strategies.

[0052] By selecting multiple member reasoning attack models for simulating attacker behaviors to attack the target model, the attacker's behaviors can be considered from different attack perspectives by simulating member reasoning attack scenarios.

[0053] Commonly used types of member reasoning attack models include, but are not limited to: a member reasoning attack model based on a binary classifier, a member reasoning attack model based on predicted entropy, and a member reasoning attack model based on confidence; in this embodiment, a total of m member reasoning attack models are selected.

[0054] Based on the simulation of member reasoning attack scenarios, various dimensional factors related to the privacy risk of data samples are comprehensively calculated, including: the gradient information of data samples during the training process of the target model, the neighborhood density, and the output probability of the member reasoning attack model output.

[0055] The gradient information of the data sample is used to reflect the sensitivity of the data sample to the target model training.

[0056] For each data sample, in the member inference attack scenario, calculate the gradient of the loss function of the target model training process with respect to the model parameters ▽ θ L(x i ,θ), where x i is the data sample, θ is the parameter of the target model, L(x i ,θ) is the loss function, ▽ θ L(x i ,θ) is the gradient information of the data sample.

[0057] At the same time, for each data sample, the gradient value of the loss function is dynamically calculated at different stages of model training (such as the initial stage, the intermediate stage, and the convergence stage) to capture the evolution trend of the privacy risk of the data sample at different stages of model training; the larger the gradient value, the greater the impact of the data sample on the target model, and the higher the risk of privacy information leakage of the data sample.

[0058] The neighborhood density of the data sample is used to measure the distribution of the data sample in the feature space. The higher the neighborhood density, the higher the similarity between the data sample and other samples. If the data sample is leaked, the attacker can obtain more information that is more helpful for model training, and the risk of privacy information leakage of the data sample is higher.

[0059] The neighborhood density of the data sample is calculated using the k-nearest neighbor idea combined with local density and global density. The calculation formula is:

[0060]

[0061] ρ(x i )=α·ρ local (x i )+(1-α)·ρ global (x i )

[0062] Where k is the data sample x i The number of nearest neighbors in the dataset; v_k(x i ) is a set of k data samples x i The minimum spherical volume of the nearest neighbors, for d-dimensional space, r is the sample x i The maximum distance between it and its k nearest neighbors, Γ is the gamma function; ρ local (x i ) is the data sample x i The local density in the data set, ρ global (x i ) is the data sample x i The global density in the data set, α is the balance coefficient, ρ(x i ) is the data sample x i The neighborhood density of .

[0063] The output probability output by the member reasoning attack model is used to indicate the probability of whether the data sample inferred by the member reasoning attack model (attacker) belongs to the training set.

[0064] The calculation steps of the output probability of the member reasoning attack model output are as follows:

[0065] In this embodiment, a total of m member reasoning attack models are selected for data sample x i , each member inference attack model will speculate whether the data sample is a data training set, considering the weight of each member inference attack model, and calculate p combined (x i ), the calculation formula is:

[0066]

[0067] Among them, w j The weight of the attack model for the jth member inference, The data sample x output by the j-th member inference attack model i The output probability of whether it belongs to the training set, p combined (x i) is the data sample x considering the weight of each member's inference attack model i The output probability of whether it belongs to the training set.

[0068] During the training process of the target model, the performance of the same data sample varies when different members infer the attack model, so it is necessary to consider the stability performance of the same data sample under different attack perspectives.

[0069] The probability variance reflects the stability of the same data sample under different attack perspectives, that is, reflects the degree of fluctuation of the inference results of the same data sample under different member reasoning attack models. The smaller the probability variance, the more consistent the inference results of different attackers (different reasoning attack models) on whether the data sample belongs to the training set, and the higher the privacy risk of the data sample; the larger the probability variance, the greater the difference in the inference results of different attackers (different reasoning attack models) on whether the data sample belongs to the training set, and the lower the privacy risk of the data sample.

[0070] The probability variance of each data sample is calculated as:

[0071]

[0072] Where m is the total number of member reasoning attack models, The data samples x output by the attack model for all members i The mean of the output probabilities of whether they belong to the training set, σ 2 (x i ) is the data sample x output by different members’ reasoning attack models i The variance of the output probability of whether it belongs to the training set.

[0073] Finally, the stability coefficient is calculated based on the probability variance of each data sample, and the stability coefficient is introduced into the output probability calculation of the member reasoning attack model output:

[0074]

[0075] a(x i ) is the stability weighting coefficient. The smaller the probability variance of the data sample, the greater the weight corresponding to the data sample; p final (x i ) is the data sample x after considering stability i The output probability of whether it belongs to the training set.

[0076] The calculated gradient information, neighborhood density, and output probability are normalized respectively to ensure that these data of different scales can be uniformly compared and comprehensively evaluated.

[0077] Normalized gradient information:

[0078]

[0079] Where g(x i ) is the data sample x i Gradient information, g max is the maximum value of the gradient information of all data samples, g min is the minimum value of the gradient information of all data samples.

[0080] Normalized neighborhood density:

[0081]

[0082] Among them, ρ(xi) is the data sample x i The neighborhood density, ρ max is the maximum value of the neighborhood density of all data samples, ρ min is the minimum value of the neighborhood density of all data samples.

[0083] The output probability is usually in the range of [0,1] and is normalized according to actual needs. The output probability after normalization is norm_prob(x i ).

[0084] The privacy risk score R(x) of the calculated data sample i ) is calculated by weighting the normalized gradient information, neighborhood density, and output probability. The calculation formula is:

[0085] R(x i )=w1·norm_g(x i )+w2·norm_ρ(x i )+w3·norm_prob(x i )

[0086] Among them, w1, w2, and w3 are weight coefficients, indicating the degree of influence of each dimensional factor on privacy risk.

[0087] Use adaptive optimization algorithms (such as Adam) to dynamically adjust the weight coefficients of each dimensional factor according to the performance of data samples during training.

[0088] Then, according to the privacy risk score R(x i ), and divide the privacy risk level of data samples into:

[0089] Use clustering methods (such as K-means) to set the threshold ε of the privacy risk score;

[0090] If the data sample x iThe privacy risk score R(xi) is greater than or equal to the threshold ε, and the data sample x i Classified as high-risk samples;

[0091] If the data sample x i The privacy risk score R(x i ) is less than the threshold, then the data sample x i Classified as low risk sample.

[0092] The classification of privacy risk levels of data samples can be expressed by the following relationship:

[0093]

[0094] After classifying the privacy risk levels of data samples, a privacy protection strategy is formulated according to the privacy risk levels of the data samples, so as to achieve targeted protection of data samples with different privacy risks, thereby significantly improving the security of machine learning models in sensitive data application scenarios and providing strong guarantees for the compliant use of data and privacy protection.

[0095] The specific privacy protection policy is:

[0096] For data samples classified as high-risk samples, they are replaced with similar data generated by a generative adversarial network, or the data features of the data samples are dynamically perturbed.

[0097] For data samples classified as low risk, they are kept as is to ensure the integrity of the data and the utility of the target model.

[0098] After comprehensively evaluating the privacy risks of data samples in membership reasoning attack scenarios, a privacy protection strategy is designed to resist membership reasoning attacks in order to achieve a reasonable balance between the privacy of data samples and the utility of the target model, thereby significantly improving the security of machine learning models in sensitive data application scenarios.

[0099] The following example specifically illustrates the calculation process of each dimension factor under the member reasoning attack model attack:

[0100] 1. Calculation of gradient information:

[0101] Gradient information is calculated for a single data sample during the target model training process. The specific process is as follows:

[0102] (1) For sample x i and label y i , calculate the loss function through the target model, such as cross entropy loss Where C is the number of label classification categories k, p k is the predicted probability of the target model for the k-th class label.

[0103] (2) Gradient calculation:

[0104] Calculate the gradient of the loss function with respect to the sample input Used to indicate the sensitivity of the target model to the input sample.

[0105] (3) Gradient norm:

[0106] Extract the L2-norm of the gradient as a metric: Where n is the input sample x i feature dimension.

[0107] Assume that the gradient norms of samples x1 and x2 are:

[0108] If the gradient norm of all data samples is in the range of [0,10], the normalized result of two data samples is: It can be seen that the privacy risk of data sample x2 is higher than that of data sample x1.

[0109] The gradient information of a sample reflects its sensitivity to the training of the template model. If the gradient of a sample is large, it means that the target model's prediction of the sample is very dependent on its specific features, which means that the model may overfit the sample and thus remember it more easily. Remembering the characteristics of the sample will increase the risk of privacy leakage. Because the membership inference attack essentially exploits the difference between the sample in the training set and the test set, the more the sample is remembered during the training process, the easier it is to infer the sample and the easier it is to be identified as a training sample by the attack model.

[0110] 2. Calculation of neighborhood density:

[0111] Neighborhood density is used to measure the distribution of samples in the feature space. The higher the neighborhood density, the more "typical" the sample is in the feature space or the more similar it is to other samples. Such samples are more likely to be remembered by the target model during training, and the risk of privacy leakage is higher. The specific calculation process is as follows:

[0112] (1) Feature vector extraction:

[0113] The sample x i Input into the target model, extract the last layer embedding vector f(xi), assuming that the dimension of the feature vector is d (such as 128 dimensions), This feature vector can be viewed as a sample x i Representation in feature space.

[0114] (2) Calculate the distance between samples:

[0115] In the feature space, calculate the sample x iWith other data samples x j The Euclidean distance of:

[0116] dist(x i ,x j )=‖f(x i )-f(x j )2

[0117] Find the distance sample x i The most recent k samples.

[0118] (3) Calculate the minimum sphere volume:

[0119] (4) Calculate neighborhood density:

[0120] Neighborhood density ρ(x i ) is calculated by the sample distribution of k nearest neighbors:

[0121] Where k is the number of neighbors, v_k(x i ) is the minimum sphere volume that contains the k nearest neighbor samples.

[0122] Assume that the target model extracts the feature vectors of samples x1 and x2, and performs the following calculation on the neighborhood density of the two samples:

[0123] The feature vector of sample x1 is f(x1) = [0.5, 0.3, 0.8, ...],

[0124] The feature vector f(x2) of sample x2 = [0.2, 0.6, 0.1, ...], and the dimension d of both data samples is 128.

[0125] Assuming k = 5, calculate the maximum distance r between samples x1 and x2 and their nearest neighbors, r(x1) = 1.5, r(x2) = 2.1; calculate the neighborhood density of samples x1 and x2 according to the calculation formula of neighborhood density, assuming ρ(x1) = 3.5, ρ(x2) = 1.2; finally, perform normalization.

[0126] Samples with high neighborhood density are densely distributed, indicating that they are similar to the features of a large number of samples and are "typical samples" in the feature space. The target model is more likely to remember these "typical samples" because the target model will give priority to remembering key samples that can accurately represent the features of most samples.

[0127] 3. Calculation of output probability:

[0128] The member inference attack model is used to simulate the attacker's behavior and calculate the probability of whether the target model is a member of the training set. The specific calculation process is as follows:

[0129] (1) Training of multiple attack models (taking the classification task of handwritten digit images as an example)

[0130] 1) Attack Model 1: Confidence-based Attack

[0131] Features: Maximum confidence p of the target model max =max(p0,p1,...,p9).

[0132] Training method:

[0133] Input: Maximum confidence p of the target model output max .

[0134] Label: member sample (training data, y member = 1) and non-member samples (test data, y member =0).

[0135] Output: probability p1(x i ), represents the sample x i Possibility for members.

[0136] 2) Attack Model 2: Attack based on predicted entropy

[0137] Features: Prediction entropy H(x i ), which is used to measure the output uncertainty of the target model:

[0138]

[0139] Training method:

[0140] Input: Prediction entropy H(x i ).

[0141] Tags: member sample and non-member sample.

[0142] Output: probability p2(x i ), represents the sample x i Possibility for members.

[0143] 3) Attack Model 3: Binary Classifier Based on Full Softmax Output

[0144] Features: The complete Softmax output probability distribution of the target model [p0,p1,...,p9].

[0145] Training method:

[0146] Input: The complete Softmax probability distribution.

[0147] Tags: member sample and non-member sample.

[0148] Output: probability p3(x i ), represents the sample x i Possibility for members.

[0149] (2) Use the above attack model to attack the target model

[0150] 1) Assume that sample x i is a handwritten digit image, and the input sample x is i , and get the Softmax output of the target model:

[0151] Softmax(x i ) = [0.05, 0.10, 0.05, 0.70, 0.05, 0.01, 0.01, 0.01, 0.01, 0.01]

[0152] 2) Use the above attack models to attack the target model respectively, and get the output of each attack model (each attack model has a certain effect on sample x i Membership probability prediction):

[0153] Attack model 1 (confidence-based): p1(x i )=0.85

[0154] Attack Model 2 (based on predicted entropy): H(x i )=-[0.05log(0.05)+0.10log(0.10)+...]≈0.88, p2(x i )=0.75

[0155] Gongzhi Model 3 (based on Softmax output): p3(x i )=0.80

[0156] (3) Combining the output probabilities of multiple attack models

[0157] 1) Calculate the mean and variance based on the output of the above attack models

[0158] Mean:

[0159]

[0160] variance:

[0161]

[0162] 2) Calculate the maximum variance: Assume that the maximum variance among all samples is maxσ 2 (x i )=0.005

[0163] 3) Calculate the stability coefficient:

[0164] Substituting the data into this we get:

[0165] 4) Calculate the output probability:

[0166]

[0167] The stability coefficient is introduced, and the outputs of multiple attack models are combined to calculate the final membership probability of the sample:

[0168]

[0169] Finally, normalization is performed.

[0170] (4) After normalizing the calculated gradient information, neighborhood density, and output probability, the normalized values ​​of the three factors are obtained, and the comprehensive privacy risk score is calculated according to the weight formula.

[0171] During the entire calculation process, there are no special requirements for data samples. The data samples can be either tabular data or image data, because in the process of training the target model, data features will be extracted from the data samples and converted into vector form to facilitate model training and calculation.

[0172] In the entire calculation process, there are no specific requirements for the target model. The target model is determined by the specific task, so the above calculation process does not specifically describe the specific parameters of the target model. The above method is mainly to calculate the privacy risk score of the data sample in the member reasoning attack scenario under the specific model training task, and determine the privacy risk level of each data sample based on the score, so that the data owner can take targeted protection measures.

[0173] See also Figure 3 This embodiment provides a data sample privacy risk classification system, which is used to implement the above classification method, including:

[0174] An attack model selection module, used for selecting a plurality of member reasoning attack models for simulating attacker behaviors to attack the target model when the target model is trained with data samples;

[0175] The data processing module is used to calculate the gradient information and neighborhood density of the data samples of the target model during the training process under the attack of the member reasoning attack model, as well as the output probability of the member reasoning attack model output; and normalize the calculated gradient information, neighborhood density, and output probability respectively;

[0176] The risk level classification module is used to calculate the privacy risk score of the data sample based on the normalized gradient information, neighborhood density, and output probability, classify the privacy risk level of the data sample according to the privacy risk score of the data sample, and formulate corresponding privacy risk protection strategies according to the privacy risk level of the data sample.

[0177] The target model is attacked by using the membership reasoning attack model, and the gradient information, neighborhood density, and output probability of the data samples generated by the target model during the training process are comprehensively considered to comprehensively evaluate the privacy risk of data samples in the membership reasoning attack scenario.

[0178] The privacy risk level of data samples is divided according to their privacy risk scores, and privacy protection strategies are formulated and implemented to achieve targeted protection of data samples with different privacy risks, thereby significantly improving the security of machine learning models in sensitive data application scenarios and providing strong guarantees for the compliant use of data and privacy protection.

[0179] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data sample privacy risk classification method, characterized in that: include: When the target model is trained with data samples, multiple member reasoning attack models for simulating attacker behaviors are selected to attack the target model; Calculate the gradient information of the target model's data samples during training, the neighborhood density, and the output probability of the member reasoning attack model output under the attack of the member reasoning attack model; Normalize the calculated gradient information, neighborhood density, and output probability respectively; According to the normalized gradient information, neighborhood density, and output probability, the privacy risk score of the data sample is calculated, and the privacy risk level of the data sample is divided according to the privacy risk score of the data sample.

2. A data sample privacy risk classification method according to claim 1, characterized in that: The member reasoning attack model includes: a member reasoning attack model based on a binary classifier, a member reasoning attack model based on prediction entropy, and a member reasoning attack model based on confidence.

3. A data sample privacy risk classification method according to claim 1, characterized in that: The gradient information of the data sample is: Among them, x i is the data sample, θ is the parameter of the target model, L(x i ,θ) is the loss function.

4. A data sample privacy risk classification method according to claim 1, characterized in that: The neighborhood density of the data sample is calculated using the k-nearest neighbor concept combined with local density and global density, and the calculation formula is: p(x i )=a·r local (x i )+(1-a)·r global (x i ) Where k is the data sample x i The number of nearest neighbors in the dataset, v_k(x i ) is a set of k data samples x i The minimum sphere volume of the nearest neighbors, ρ local (x i ) is the data sample x i The local density in the data set, ρ global (x i ) is the data sample x i The global density in the data set, α is the balance coefficient, ρ(x i ) is the data sample x i The neighborhood density of .

5. A data sample privacy risk classification method according to claim 1, characterized in that: The calculation formula of the output probability of the member reasoning attack model output is: Among them, w j The weight of the attack model for the jth member inference, The data sample x output by the j-th member inference attack model i The output probability of whether it belongs to the training set, p combined (x i ) is the data sample x considering the weight of each member's inference attack model i The output probability of whether it belongs to the training set, m is the total number of member reasoning attack models, The data samples x output by the attack model for all members i The mean of the output probabilities of whether they belong to the training set, σ 2 (x i ) is the data sample x output by different members’ reasoning attack models i The variance of the output probability of whether it belongs to the training set, a(x i ) is the stability weighting coefficient, p final (x i ) is the data sample x after considering stability i The output probability of whether it belongs to the training set.

6. A data sample privacy risk classification method according to claim 1, characterized in that: The privacy risk score of the calculated data sample is obtained by weighted calculation of normalized gradient information, neighborhood density, and output probability.

7. A data sample privacy risk classification method according to claim 6, characterized in that: An adaptive optimization algorithm is used to dynamically adjust the normalized gradient information, neighborhood density, and output probability weight coefficients.

8. A data sample privacy risk classification method according to claim 1, characterized in that: The step of classifying the privacy risk levels of the data samples according to the privacy risk scores of the data samples includes: Using clustering methods to set thresholds for privacy risk scores; If the privacy risk score of a data sample is greater than or equal to the threshold, the data sample is classified as a high-risk sample; If the privacy risk score of a data sample is less than the threshold, the data sample is classified as a low-risk sample.

9. A data sample privacy risk classification method according to claim 8, characterized in that: After dividing the privacy risk levels of data samples according to their privacy risk scores, data samples classified as high-risk samples are replaced with similar data generated by a generative adversarial network, or the data features of the data samples are dynamically perturbed.

10. A data sample privacy risk grading system, used to implement a data sample privacy risk grading method as claimed in any one of claims 1 to 9, characterized in that: include: An attack model selection module, used for selecting a plurality of member reasoning attack models for simulating attacker behaviors to attack the target model when the target model is trained with data samples; The data processing module is used to calculate the gradient information and neighborhood density of the data samples of the target model during the training process under the attack of the member reasoning attack model, as well as the output probability of the member reasoning attack model output; and normalize the calculated gradient information, neighborhood density, and output probability respectively; The risk level classification module is used to calculate the privacy risk score of the data sample based on the normalized gradient information, neighborhood density, and output probability, and to classify the privacy risk level of the data sample based on the privacy risk score of the data sample.