A feature privacy security measurement method and device and a storage medium

By constructing a synthetic dataset and training a prediction model to evaluate the intermediate layer features of the machine learning model, this approach addresses the insufficient privacy and security issues under unknown attack modes in existing technologies, achieving effective defense against unknown attacks and broad applicability.

CN119513907BActive Publication Date: 2025-10-17PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411431127.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-10-17
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

Existing technologies cannot effectively protect the intermediate features of machine learning models when faced with unknown or changing attack patterns, resulting in insufficient privacy and security, especially in sensitive data processing where there is a threat of inversion attacks.

Method used

By constructing a synthetic dataset, analyzing the intermediate layer features of the machine learning model, training the prediction model to evaluate its defense capability against unknown attacks, and optimizing the model parameters using a multi-layer neural network and a binary cross-entropy loss function, the prediction and defense against unknown attacks are achieved.

Benefits of technology

It improves the model's ability to defend against unknown attacks, provides broader and more flexible privacy protection, is applicable to a variety of machine learning models and attack environments, and enhances the defense effect in dynamic security environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005083699190000021
    Figure BDA0005083699190000021
  • Figure BDA0005083699190000022
    Figure BDA0005083699190000022
  • Figure BDA0005083699190000043
    Figure BDA0005083699190000043
Patent Text Reader

Abstract

The application discloses a feature privacy security measurement method and device and a storage medium, and belongs to the field of information security. The application is aimed at the problem that the existing method has limited defense capability for unknown attacks and depends on specific attack instances. By pre-training a victim model, intermediate layer features are extracted and attack success probabilities are recorded. Multiple inversion attacks are performed and a defense strategy optimization model is used. A synthetic data set is constructed, a prediction model is trained to predict attack success probabilities, and a trained prediction model is used to evaluate the privacy security of features under different defense strategies. By analyzing the distribution properties of the intermediate layer features, the prediction model can predict the anti-attack capability in an unknown attack environment, improve the comprehensiveness and generalization capability of privacy protection, and be suitable for various models and attack scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of information security, and particularly relates to a feature privacy security measurement method and device and a storage medium. BACKGROUND

[0002] In the field of artificial intelligence and machine learning, with the wide application of deep learning models in various data processing tasks, how to protect the intermediate layer features of the model from being maliciously used has become an important privacy protection problem. Especially in the processing of sensitive data (such as medical records, personal biological features, etc.), the intermediate layer features of the model often contain rich information, which may be used by attackers to reconstruct or infer the original training data. This inversion attack by extracting and using the intermediate layer features of the model has become a significant security threat. Traditional defense strategies, such as data perturbation, model regularization, and application of homomorphic encryption technology, are often designed for pre-defined attack patterns. They improve the ability of the intermediate layer to resist known attack patterns by training or processing the model in a specific way.

[0003] The existing technology mainly focuses on countering several known attack types by training the model in a specific defensive way to improve its performance on these known attacks. For example, some methods reduce the mutual information between the features of the intermediate layer of the model and the input data by introducing regularization techniques, thereby reducing the possibility of inferring the input data from the model output as much as possible. Although these strategies perform well against specific attack types, they generally lack the ability to generalize to unknown attack patterns and fail to provide a method to predict the resistance of the model to unknown or new attacks. A common feature defense method is Improving Robustness to Model Inversion Attacks via Mutual Information Regularization, which is a method based on regularization techniques to reduce the mutual information between the features of the intermediate layer of the model and the input data. Common attack types include: knowledge-enhanced attack method KEDMI, attack method GMI based on generative adversarial network, attack method LOMMA based on Logit maximization and model enhancement.

[0004] In the existing technology, inversion attacks on features (here, inversion attacks use the output or intermediate layer features of a machine learning model to reconstruct the input data of the model, thereby infringing on user privacy) pose a serious threat to the privacy protection of machine learning models. In the past defense strategies, evaluating the effectiveness of the defense usually relies on known attack strategies. However, this method is greatly reduced in effectiveness when encountering unknown or changing attack patterns. Especially in a dynamically changing security environment, past methods cannot provide continuous and adaptive protection. SUMMARY

[0005] To solve the above technical problems, the present application provides a feature privacy security measurement method. The core of the method is to analyze the intermediate layer features of the machine learning model to evaluate the privacy security of the features when facing unknown attack strategies. By extracting the intermediate layer features of the neural network and analyzing their distribution properties, the present application can predict and measure the privacy protection capability of the model under unknown attack modes.

[0006] The technical solution of the present application to solve the above technical problems is as follows:

[0007] A feature privacy security measurement method, comprising the following steps:

[0008] 1) Pre-training a victim model on a private dataset, and performing an inversion attack on the victim model, recording the output intermediate layer features and the corresponding attack success probability; improving the victim model through defense strategies, repeating the attack and data collection, and integrating all intermediate layer features and attack success probabilities into a synthetic dataset;

[0009] 2) Using the synthetic dataset to train a prediction model that predicts the attack success probability based on the input intermediate layer features;

[0010] 3) Using the trained prediction model to process the intermediate layer features of the victim model trained with different defense strategies to predict the attack success probability.

[0011] Further, in step 1), the private dataset is composed of private images and corresponding ID labels, and the victim model is a neural network model for handling classification tasks.

[0012] Further, in step 1), the inversion attack uses one of the attack methods of KEDMI, GMI, and LOMMA.

[0013] Further, in step 1), the steps of the inversion attack include: obtaining the intermediate layer features from the victim model, and performing image reconstruction attack based on the intermediate layer features through the attack model, and the probability of attack success is:

[0014]

[0015] where ID(x i ) represents the sensitive information extracted from the input image x i of the victim model, is the reconstructed image of the attack model, N is the total number of test samples, and I is an indicator function.

[0016] Further, in steps 1) and 3), the defense strategies include at least one of data perturbation, model regularization, and homomorphic encryption.

[0017] Further, the prediction model in step 2) is a multi-layer neural network model.

[0018] Further, the training model in step 2) uses a binary cross-entropy loss function to measure the difference between the predicted probability and the true probability, which is:

[0019]

[0020] where p attack,i is the probability of attack success, is the probability of attack success predicted by the training model, N is the total number of test samples, and the base of the logarithm is the natural logarithm.

[0021] Further, the training model in step 2) uses gradient descent or its variants to minimize the binary cross-entropy loss function L(θ) to optimize the model parameters θ.

[0022] A feature privacy security measurement device, comprising a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the above method.

[0023] A computer-readable storage medium storing a computer program, which when executed implements the steps of the above method.

[0024] The advantages of the present application mainly lie in the following aspects:

[0025] 1. Strong generalization ability: Unlike existing technologies that mainly defend against known attack patterns, the present application is not limited to known attack patterns, but can predict the resistance of the model to unknown attack environments by analyzing the intermediate layer features of the model. Compared with existing technologies, the present application not only improves the defense ability against known attacks, but also has adaptability to new or changing attack strategies, thereby providing stronger generalization ability.

[0026] 2. Independent of specific attack instances: The present application does not need to rely on specific attack patterns to improve the anti-attack ability of the model. Compared with existing methods that usually need to be trained for specific attacks, the present application can directly evaluate the anti-attack ability of the model by comprehensive analysis of the intermediate layer features, making the defense strategy more flexible and applicable to a wider range.

[0027] 3. Comprehensive privacy protection solution: The present application is applicable to a variety of machine learning models and attack environments, providing a more comprehensive privacy protection solution. By predicting the performance of the model in different attack scenarios, the present application can effectively improve the privacy protection level of the model.

[0028] 4. Improve defense effect in unknown attack environment: By analyzing the distribution of intermediate layer features through machine learning and neural networks, the invention can effectively predict the defense effect of the model in unknown attack environment, providing a more comprehensive solution for privacy protection. DETAILED DESCRIPTION

[0029] In order to make the technical features and advantages or technical effects of the above technical solutions of the present invention more obvious and easy to understand, the following detailed description is made.

[0030] The present invention provides a feature privacy security measurement method, which covers the complete process from data set construction, model training to attack resistance prediction. The following are the specific implementation steps:

[0031] 1. Construct a synthetic data set

[0032] In order to evaluate the privacy security of the features, a synthetic data set containing the intermediate layer features of the model and their corresponding attack success probabilities needs to be constructed first.

[0033] 1) Pre-train a victim model T on a private dataset D priv consisting of private images (such as personal facial images) and corresponding ID labels. The victim model T is a neural network (such as VGG16 and ResNet-34) that handles classification tasks, which can input an image x and output T(x), i.e. the probability distribution of the image on the downstream classification task (such as face recognition task).

[0034] 2) Perform multiple known inversion attacks (such as KEDMI, GMI, and LOMMA attack methods) on the victim model T. Record the intermediate layer features z and the corresponding attack success probability p attack , where the attack success probability p attack is obtained by evaluating the reconstructed image by the attacker (usually using existing image recognition models to predict the identity of the reconstructed image) and calculating the accuracy.

[0035] 2-1) Goal of inversion attack:

[0036] The main goal of inversion attack is to use the intermediate layer features z obtained from the victim model T to reconstruct the original input x through the attack model M. Such attacks are particularly aimed at input data containing sensitive information (such as personal facial images, biometric features, or other private data), and attack success will directly threaten the security of data privacy.

[0037] 2-2) The input of the attack model M is the intermediate layer feature z = T mid (x) extracted from the victim model T, where T midThe function that extracts the intermediate layer features in the model T. The output of the attack model M is the image reconstructed using the input intermediate layer features z

[0038] The probability p of the attack model M being successful attack :

[0039]

[0040] where ID(x) represents sensitive information extracted from the image x, such as identity, etc., N is the total number of test samples, and I is the indicator function, i.e.

[0041] 3) Adjust the victim model T after various defense strategies (such as data perturbation, model regularization, homomorphic encryption, etc.) to obtain a model T ′ , which is more resistant to various attacks, and through this model T ′ , the model T ′ repeats the above steps to obtain more data.

[0042] 4) Integrate all collected binary tuples (z, p attack (x)) into a synthetic dataset .

[0043] 2. Train the model to predict the attack resistance

[0044] 1) Use the collected synthetic dataset to train a prediction model F that predicts the probability of an attack being successful in a downstream attack scenario based on the intermediate layer features z

[0045] 2) The input of the prediction model F is the intermediate layer feature where d is the dimension of the feature. The output of the prediction model F is the predicted attack success probability

[0046] 3) Model structure: Assuming F is a function defined by parameters θ, it can be represented as: F(z; θ) = σ(g(z; θ)), where g(z; θ) is a function that can be implemented by a multi-layer neural network, responsible for extracting relevant information from the input feature z and the probability p attack,i of being attacked successfully, and σ is the Sigmoid activation function, used to compress the output to the interval [0, 1] to represent the probability.

[0047] 4) Training process:

[0048] Loss function: To train the model F, a binary cross-entropy loss function is used to measure the difference between the predicted probability and the true probability:

[0049]

[0050] where The base of the logarithm is the natural logarithm.

[0051] Optimization algorithm: Use gradient descent or its variants (such as Adam or RMSprop) to minimize the loss function L(θ) to optimize the model parameters θ.

[0052] 5) Verification and model selection: Optimize the model parameters θ and verify the prediction performance of the model F through techniques such as cross-validation to ensure good generalization ability on new attack data.

[0053] 3. Privacy and security assessment of features

[0054] Use the trained prediction model F to evaluate the privacy and security of the intermediate layer features z of the victim model T trained under different defense strategies.

[0055] 1) Extract the intermediate layer features z of the victim model T after using different defense strategies and use the prediction model F to evaluate the anti-attack ability.

[0056] 2) The results can help understand the effect of various defense strategies on improving the security of the victim model, and guide further model design and defense strategy selection.

[0057] Experimental test:

[0058] I. Experimental setup

[0059] 1. Private dataset setting: Use a subset of CelebA, carefully select samples containing 1000 people, 5 pictures per person, a total of 5000 pictures. The ID of the person is numbered from 0 to 999.

[0060] 2. Victim model configuration: Choose VGG16 as the framework of the victim model. Pre-train on the private dataset F priv to obtain the victim model T.

[0061] II. Inversion attack implementation

[0062] 1. Apply the following three attack methods to the victim model T respectively:

[0063] a) KEDMI (knowledge-enhanced attack method)

[0064] b) GMI (attack method based on generative adversarial network)

[0065] c) LOMMA (attack method based on Logit maximization and model enhancement)

[0066] 2. Collect the intermediate layer features z and the corresponding attack success probability p when the attack is successful attack .

[0067] 3. Data Collection and Preprocessing

[0068] 1. Data collection: For each image, the attack is performed without implementing the defense strategy, and the attack is performed again after implementing different defense strategies. This means that for each attack strategy, three sets of {z,p attack Data (for different defense strategies):

[0069] a) No defense

[0070] b) Defense strategy A (e.g. regularization)

[0071] c) Defense Strategy B (e.g., data perturbation)

[0072] 2. Data organization: For 5,000 images, each attack strategy generates 15,000 sets of data. In total, for the three attack strategies, there are 45,000 sets of data.

[0073] 4. Dataset Division

[0074] The data needs to be divided into training, test, and validation sets. The recommended split ratio is 70% training, 20% test, and 10% validation.

[0075] 5. Anti-attack capability prediction model

[0076] 1. Model training: Use the collected data to train a prediction model F. This model should be able to predict the probability of attack success in downstream attack scenarios based on the intermediate layer features z.

[0077] 2. Model structure: The prediction model can adopt a deep learning framework, such as a deep neural network, and use activation functions and loss functions to adjust the output to ensure that the output is a probability value.

[0078] 6. Performance Evaluation Indicators and Experimental Results

[0079] The experiment uses KL divergence (Kullback-Leibler Divergence) to measure the predicted probability With the true probability p attack Experiments show that on the test set, the average KL divergence between the predicted probability and the true probability is only 0.00805. This shows that the present invention can effectively measure the privacy security of features.

[0080] Although the present application has been disclosed with reference to the examples as above, it is not intended to limit the present application, and any suitable modification or equivalent replacement made by those skilled in the art to the technical solutions of the present application should be covered within the protection scope of the present application, and the protection scope of the present application is defined by the claims.

Claims

1. A feature privacy security measurement method, characterized in that: The following steps are involved: 1) Pre-train the victim model on a private dataset and perform an inversion attack on the victim model, recording the output intermediate layer features and the corresponding attack success probability; Improve the victim model through defense strategies, repeat attacks and data collection, and integrate all intermediate layer features and attack success probabilities into a synthetic dataset; The input of the attack model is the intermediate layer feature z=T extracted from the victim model T mid (x), where T mid Represents the function of extracting intermediate layer features in model T. The output of the attack model is the image reconstructed using the input intermediate layer features z Where M represents the attack model. The steps of the inversion attack include: obtaining the intermediate layer features from the victim model and performing an image reconstruction attack based on the intermediate layer features through the attack model. The probability of successful attack is: Among them, ID(x i ) represents the input image x from the victim model i Sensitive information extracted from is the image reconstructed by the attack model, N is the total number of test samples, and I is the indicator function; 2) Use the synthetic dataset to train a prediction model that predicts the probability of attack success based on the input intermediate layer features; 3) Use the trained prediction model to process the intermediate layer features of the victim model trained with different defense strategies to predict the probability of attack success.

2. The method according to claim 1, wherein In step 1), the private dataset consists of private images and corresponding ID tags, and the victim model is a neural network model that processes classification tasks.

3. The method according to claim 1, wherein In step 1), the inversion attack adopts one of the attack methods among KEDMI, GMI, and LOMMA.

4. The method according to claim 1, wherein The defense strategies in steps 1) and 3) include at least one of data perturbation, model regularization, and homomorphic encryption.

5. The method according to claim 1, wherein The prediction model in step 2) is a multi-layer neural network model.

6. The method according to claim 1 or 5, wherein: The training model in step 2) uses a binary cross entropy loss function to measure the difference between the predicted probability and the true probability. The binary cross entropy loss function is: Among them, p attack,i is the probability of a successful attack, is the probability of successful attack predicted by the training model, N is the total number of test samples, and the log base is the natural logarithm.

7. The method according to claim 6, wherein In step 2), the training model uses gradient descent or its variant algorithm to minimize the binary cross entropy loss function L(θ) to optimize the model parameters θ.

8. A feature privacy security measurement device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, characterized in that A computer program is stored, and when the computer program is executed, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and system for training data privacy measurement in machine learning

    CN113051620A

  • Image processing method for targeted countermeasure attack in middle layer

    CN113344090A