Privacy image classification method for defending model inversion attack
The feature purification module removes redundant information, which solves the instability and redundancy problems of the existing defense model inversion attack methods, and achieves a more stable defense effect and classification effect.
Patent Information
- Application Number
- CN202510427036.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-11
AI Technical Summary
The existing defense model inversion attack methods rely on data-driven, resulting in unstable defense effects, and only consider the first-order correlation between input data and features, and fail to effectively remove redundant information.
A private image classification method for defense model inversion attacks was designed, and the redundant information in the features was removed through the feature purification module, including removing feature pair correlation and feature correlation, using hash networks and random mapping vectors, combining classifiers and feature extractors, and optimizing the loss function to improve defense effects.
It significantly improves the defense performance of model inversion attacks, reduces the error between reconstructed images and private images, ensures classification utility and stable defense effects, and reduces training time.
Smart Images

Figure CN120298796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification and privacy protection, and in particular to a privacy image classification method for defending against model inversion attacks. Background Art
[0002] In recent years, the wide application of deep neural networks has significantly promoted the development of artificial intelligence, but the risk of its privacy leakage has become increasingly severe. Model inversion attacks can reconstruct privacy information in training data, such as face images, by reverse reasoning the output of a deep neural network model, posing a direct threat to personal privacy. Therefore, defending against model inversion attacks has become an important research topic.
[0003] Among the existing defense methods, one type of method introduces a generative adversarial framework to achieve the purpose of defending against model inversion attacks. These methods aim to construct an inversion model to simulate a powerful attacker and enhance the defense ability through adversarial training. However, since the inversion model relies on data-driven, the instability of data-driven may lead to unstable defense effects. Another type of method believes that the redundant information contained in features is usually the key factor for the success of model inversion attacks. Therefore, it defends by minimizing the correlation metric between the input data and the features to reduce redundant hidden information. It should be noted that this type of method only considers the first-order correlation between the input data and the features. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and propose a privacy image classification method for defending against model inversion attacks, eliminating the need to design an inversion model to simulate the reconstruction ability of an attacker through a data-driven method, and achieving better defense effects by removing more redundant information contained in the features.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: a privacy image classification method for defending against model inversion attacks, comprising the following steps:
[0006] S1: Obtain privacy image data and preprocess the images therein to obtain images of a unified size;
[0007] S2: Input the preprocessed private image data into the trained defense model to obtain an image classification result. The defense model consists of a feature extractor, a feature purification module, and a classifier. First, the feature extractor extracts features from the private image data to obtain private image features. Second, the feature purification module processes the private image features to remove redundant information therein to obtain purified private image features. This module restricts the attacker's ability to build an inversion model to reconstruct the private image data using the output of the defense model by reducing the private information that the attacker can obtain, thereby achieving the goal of defending against inversion attacks of the defense model. The feature purification module includes a feature pair correlation removal sub-module and a feature correlation removal sub-module. Here, a feature pair refers to a data set composed of two private image features. The feature pair correlation removal sub-module is used to remove the correlation between the Euclidean distance between two private image features in the input feature pair and the Euclidean distance between two semi-purified private image features in the output feature pair. The feature correlation removal sub-module is used to remove the correlation between the input semi-purified private image features and the output purified private image features. Finally, classify the purified private image features through the classifier to obtain the image classification result.
[0008] Further, in step S1, given a private image data set, the private image data therein is scaled to a size of 64×64. Let \(x_i\) i represent the \(i\)-th private image, and \(N\) represent the number of private images.
[0009] Further, the specific operation steps of step S2 are as follows:
[0010] S21: Use the feature extractor to extract corresponding private image features from the preprocessed private image data wherein the feature extractor is a deep neural network model. Let \(s_i\) represent the corresponding private image features obtained by the \(i\)-th private image passing through the feature extractor, i and let \(s_i\) represent that the private image features are a real number vector with a dimension of \(n\).
[0011] S22: Input the private image features \(s_i\) i into the feature pair correlation removal sub-module. The features will be processed through a fully connected layer, a BN layer, a LeakyReLU activation function, and a fully connected layer in sequence to obtain semi-purified private image features \(q_i\) i :
[0012] \(q_i\) i = FC(BN(LeakyReLU(FC(\(s_i\) i ))))
[0013] In the formula, FC represents a fully connected layer, and BN represents a batch normalization layer;
[0014] Calculate the Euclidean distance between the \(i\)-th private image feature \(s\) i and the \(j\)-th private image feature as well as the Euclidean distance between the \(i\)-th semi-purified private image feature \(q\) i and the \(j\)-th semi-purified private image feature where \(d\) represents calculating the Euclidean distance, \((s i , s j ) represents the feature pair composed of the \(i\)-th private image feature \(s i and the \(j\)-th private image feature \(s j , and \((q i , q j ) represents the feature pair composed of the \(i\)-th semi-purified private image feature \(q i and the \(j\)-th semi-purified private image feature \(q j ; By minimizing \(L RFPC , the removal of the correlation of the feature pair is realized, and the calculation method is as follows:
[0015]
[0016] In the formula, represents the distribution of \(D s , where represents the set of Euclidean distances between all private image features and other private image features, represents the distribution of \(D q , where represents the set of Euclidean distances between all semi-purified private image features and other semi-purified private image features, represents the joint distribution of \(D s and \(D q , represents the expectation;
[0017] S23: Input the obtained semi-purified private image feature \(q i into the feature correlation removal sub-module. The feature will first pass through a hash network to obtain the semi-purified private image feature \(h i , and the structure of the hash network is as follows:
[0018] h i = Tanh(FC(FCNet(FCNet(q i ))))
[0019] In the formula, Tanh represents the Tanh activation function, and FCNet represents a network with the structure of BN(LeakyReLU(FC(*))); During training, the hash network is controlled by the hash loss \(L hash , and the output feature value tends to 1 and -1, and the calculation method is as follows:
[0020]
[0021] In the formula, 1 represents a vector with dimension n and all elements being 1, cosh represents the hyperbolic cosine function, and n represents the dimension of the feature; then the semi-purified private image feature h i is multiplied by a random mapping vector to obtain the final purified private image feature z i , and the calculation method is as follows:
[0022] z i = h i × M
[0023] In the formula, M represents a random mapping vector that flips the feature values with a 50% probability, that is, M = {-1, +1} n , where the number of -1 and 1 each accounts for 50%, and the positions are randomly generated; during training, it is generated by a preset random number seed and does not change anymore;
[0024] S24: The purified private image feature z i is input into the classifier to obtain the confidence vector
[0025]
[0026] In the formula, FC n->C represents a fully connected layer with an input dimension of n and an output dimension of C, and C represents the number of classifications; to ensure the classification utility, the utility loss L utility is defined as:
[0027]
[0028] In the formula, L CE represents the cross-entropy function, α is a hyperparameter, represents that if {i, j|y i = y j , i ≠ j} is satisfied, the value is 1, and the value is 0 if not satisfied, y i represents the true classification label of the i-th private image, and y j represents the true classification label of the j-th private image.
[0029] Furthermore, to simultaneously ensure the classification utility of the model and the performance of defending against model inversion attacks, the final loss function L is used to train the defense model:
[0030] L = L utility + βL RFPC + γL hash
[0031] In the formula, both β and γ represent hyperparameters.
[0032] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0033] 1. The designed feature purification module of the present invention not only removes the correlation between the input features and the output features, but also removes the correlation between the Euclidean distances of the input feature pairs and the Euclidean distances of the output feature pairs, further increasing the error between the reconstructed image and the private image.
[0034] 2. The present invention defines the upper limit of the attacker's reconstruction function based on the Principal Inertia Components theory (du Pin Calmon F, Makhdoumi A, Médard M, et al. Principal inertia components and applications[J]. IEEE Transactions on Information Theory, 2017, 63(8): 5011 - 5038.), designs a loss function and method by minimizing the upper limit of the attacker's reconstruction function, eliminates the need to design an inversion model, thereby reducing a large amount of training time and ensuring stable results.
[0035] 3. The designed feature purification module of the present invention is a pluggable module and is independent of the feature extractor and the classifier.
[0036] 4. Compared with other methods for defending against model inversion attacks, the method of the present invention has significantly improved performance, a greater error between the reconstructed image and the private image, and better interpretability of the overall method.
[0037] In summary, the present invention can greatly reduce the accuracy of model inversion attacks by using the feature purification module while ensuring classification utility. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a framework diagram of the method of the present invention.
[0039] Figure 2 is a schematic diagram of the feature purification module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be further described in detail below in conjunction with the embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0041] As Figure 1 and Figure 2 shown, this embodiment discloses a privacy image classification method for defending against model inversion attacks, and the specific situation is as follows:
[0042] 1) Given a privacy image dataset, scale each image in it to a size of 64×64. Let \(x_i\) i represent the \(i\)-th privacy image, and \(N\) represent the number of privacy images.
[0043] 2) Input the preprocessed privacy image data into the trained defense model to obtain the image classification result. The defense model consists of a feature extractor, a feature purification module, and a classifier. First, the feature extractor extracts features from the privacy image data to obtain privacy image features. Second, the feature purification module processes the privacy image features to remove redundant information and obtain purified privacy image features. This module limits the attacker's ability to construct an inversion model to reconstruct the privacy image data using the output of the defense model by reducing the privacy information that the attacker can obtain, thus achieving the goal of defending against inversion attacks. The feature purification module includes a feature pair correlation removal sub-module and a feature correlation removal sub-module. Here, a feature pair refers to a data set composed of two privacy image features. The feature pair correlation removal sub-module is used to remove the correlation between the Euclidean distance between two privacy image features in the input feature pair and the Euclidean distance between two semi-purified privacy image features in the output feature pair. The feature correlation removal sub-module is used to remove the correlation between the input semi-purified privacy image features and the output purified privacy image features. Finally, classify the purified privacy image features through the classifier to obtain the image classification result. Specifically, it includes the following steps:
[0044] 2.1) Input the feature extractor VGG16 to extract the corresponding privacy image features Let \(s_i\) i represent the corresponding privacy image features obtained by the \(i\)-th privacy image passing through the feature extractor, indicating that the privacy image features are a real-valued vector with dimension \(n\).
[0045] 2.2) Input the privacy image features \(s_i\) i into the feature pair correlation removal sub-module. The features will pass through a fully connected layer, a BN layer, a LeakyReLU activation function, and another fully connected layer in sequence to obtain the semi-purified privacy image features \(q_i\) i :
[0046] \(q_i\) i = FC(BN(LeakyReLU(FC(\(s_i\) i ))))
[0047] In the formula, FC represents the fully connected layer, and BN represents the batch normalization layer.
[0048] Calculate the Euclidean distance between the \(i\)-th privacy image feature \(s_i\) i and the \(j\)-th privacy image feature respectively and the Euclidean distance between the i-th semi-purified private image feature q i and the j-th semi-purified private image feature where d represents the calculation of the Euclidean distance, and (s i , s j ) represents the feature pair composed of the i-th private image feature s i and the j-th private image feature s j . Similarly, (q i , q j ) represents the feature pair composed of the i-th semi-purified private image feature q i and the j-th semi-purified private image feature q j . By minimizing L RFPC , the removal of the correlation of the feature pairs is achieved, and the calculation method is as follows:
[0049]
[0050] In the formula, represents the distribution of D s , where represents the set of Euclidean distances between all private image features and other private image features, represents the distribution of D q , where represents the set of Euclidean distances between all semi-purified private image features and other semi-purified private image features, represents the joint distribution of D s and D q , represents the expectation. Since the distribution is not directly computable, as shown in Figure 2 (a) in, a loss estimation network (Cheng P, Hao W, Dai S, et al. Club: A contrastive log-ratio upper bound of mutual information[C] / / International conference on machine learning. PMLR, 2020: 1779-1788.) is used for approximate estimation to obtain L RFPC .
[0051] 2.3) Input the obtained semi-purified private image feature q i into the feature correlation removal sub-module. The feature will first pass through a hash network to obtain the semi-purified private image feature h i , and the structure of the hash network is as follows:
[0052] h i = Tanh(FC(FCNet(FCNet(qi ))))
[0053] Wherein, Tanh represents the Tanh activation function, and FCNet represents a network with the structure of BN(LeakyReLU(FC(*)))). During training, the hash network is controlled by the hash loss L hash and the output eigenvalue tends to 1 and -1, and the calculation method is as follows:
[0054]
[0055] Wherein, 1 represents a vector with dimension n and all elements being 1, cosh represents the hyperbolic cosine function, and n represents the dimension of the feature. Then, the semi-purified private image feature h i is multiplied by the random mapping vector to obtain the final purified private image feature z i , and the calculation method is as follows:
[0056] z i = h i ×M
[0057] Wherein, M represents a random mapping vector that flips the eigenvalue with a 50% probability, that is, M = {-1, +1} n , where the number of -1 and 1 each accounts for 50%, and the positions are randomly generated. During training, it is generated by a preset random number seed and does not change anymore.
[0058] 2.4) Input the purified private image feature into the classifier to obtain the confidence vector
[0059]
[0060] Wherein, FC n->C represents a fully connected layer with an input dimension of n and an output dimension of C, and C represents the number of classifications. To ensure the classification utility, the utility loss L utility is defined as:
[0061]
[0062] Wherein, L CE represents the cross-entropy function, α is a hyperparameter, represents that if {i, j|y i = y j , i ≠ j} is satisfied, the value is 1, and the value is 0 if not satisfied. y i represents the classification label of the i-th private image, and y j represents the true classification label of the j-th private image.
[0063] To simultaneously ensure the classification utility of the model and the performance of defending against model inversion attacks, the defense model composed of a feature extractor, a feature purification module, and a classifier is trained using the final loss function L:
[0064] L = L utility + βL RFPC + γL hash
[0065] In the formula, both β and γ represent hyperparameters.
[0066] The above embodiments are preferred embodiments of the present invention. However, the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A privacy image classification method for defending against model inversion attacks, characterized in that, It includes the following steps: S1: Obtain privacy image data, and preprocess the images therein to obtain images of a unified size; S2: Input the preprocessed privacy image data into the trained defense model to obtain an image classification result; among them, the defense model consists of a feature extractor, a feature purification module, and a classifier; first, the feature extractor extracts features from the privacy image data to obtain privacy image features; secondly, the feature purification module processes the privacy image features to remove redundant information therein to obtain purified privacy image features. This module limits the ability of the attacker to construct an inversion model to reconstruct the privacy image data using the output of the defense model by reducing the privacy information that the attacker can obtain, thereby achieving the goal of defending against inversion attacks of the defense model. The feature purification module includes a sub-module for removing feature pair correlations and a sub-module for removing feature correlations. Among them, a feature pair refers to a data set composed of two privacy image features. The sub-module for removing feature pair correlations is used to remove the correlation between the Euclidean distance between two privacy image features in the input feature pair and the Euclidean distance between two semi-purified privacy image features in the output feature pair. The sub-module for removing feature correlations is used to remove the correlation between the input semi-purified privacy image features and the output purified privacy image features; finally, the purified privacy image features are classified by the classifier to obtain the image classification result.
2. A privacy image classification method for defense model inversion attack according to claim 1, characterized in that, In step S1, given a privacy image dataset, the privacy image data therein is scaled to a size of 64×64, where x i represents the i-th privacy image, and N represents the number of privacy images.
3. A privacy image classification method for defense model inversion attacks according to claim 2, characterized in that, The specific operation steps of step S2 are: S21: For the preprocessed privacy image data Use a feature extractor to extract the corresponding privacy image features where the feature extractor is a deep neural network model, s i represents the corresponding privacy image features obtained by the feature extractor for the i-th privacy image, indicating that the privacy image features are a real vector with a dimension of n. S22: Input the privacy image feature s i into the feature pair correlation removal sub-module. The feature will be processed through a fully connected layer, a BN layer, a LeakyReLU activation function, and a fully connected layer in sequence to obtain the semi-purified privacy image feature q i : q i = FC(BN(LeakyReLU(FC(s i )))) In the formula, FC represents the fully connected layer, and BN represents the batch normalization layer; Calculate the Euclidean distance between the $i$-th private image feature $s$ i and the $j$-th private image feature as well as the Euclidean distance between the $i$-th semi-purified private image feature $q$ i and the $j$-th semi-purified private image feature where $d$ represents calculating the Euclidean distance, $(s$ i , $s$ j ) represents the feature pair composed of the $i$-th private image feature $s$ i and the $j$-th private image feature $s$ j , and $(q$ i , $q$ j ) represents the feature pair composed of the $i$-th semi-purified private image feature $q$ i and the $j$-th semi-purified private image feature $q$ j ; the removal of the correlation of the feature pair is achieved by minimizing $L$ RFPC and the calculation method is as follows: In the formula, represents the distribution of D s , where represents the set of Euclidean distances between all private image features and other private image features, represents the distribution of D q , where represents the set of Euclidean distances between all semi-purified private image features and other semi-purified private image features, represents the distribution of D s and D q 's joint distribution, represents the expectation; S23: Input the obtained semi-purified private image feature q i into the feature correlation removal sub-module. The feature will first pass through a hash network to obtain the semi-purified private image feature h i , and the structure of the hash network is as follows: h i = Tanh(FC(FCNet(FCNet(q i )))) Wherein, Tanh represents the Tanh activation function, and FCNet represents a network with the structure of BN(LeakyReLU(FC(*))); during training, the hash network is controlled by the hash loss L hash and the output eigenvalue tends to 1 and -1. The calculation method is as follows: Wherein, 1 represents a vector with dimension n and all elements being 1, cosh represents the hyperbolic cosine function, and n represents the dimension of the feature; then multiply the semi-purified private image feature h i by the random mapping vector to obtain the final purified private image feature z i , and the calculation method is as follows: z i = h i × M Wherein, M represents a random mapping vector of eigenvalues with a 50% probability of flipping, that is, M = {-1, +1} n , where the number of -1 and 1 each accounts for 50%, and the positions are randomly generated; during training, it is generated by a preset random number seed and will not change anymore; S24: Input the purified privacy image feature z i into the classifier to obtain the confidence vector where FC n->C represents a fully connected layer with an input dimension of n and an output dimension of C, and C represents the number of categories; to ensure the classification utility, the utility loss L utility is defined as: where L CE represents the cross-entropy function, and α is a hyperparameter. represents that when {i, j|y i =y j , i≠j} is satisfied, the value is 1, and when it is not satisfied, the value is 0. y i represents the true classification label of the i-th private image, and y j represents the true classification label of the j-th private image.
4. A privacy image classification method for defense model inversion attack according to claim 3, characterized in that, In order to ensure both the classification utility of the model and the performance of defending against inversion attacks of the defense model, the final loss function L is used to train the defense model: L = L utility + βL RFPC + γL hash In the formula, both β and γ represent hyperparameters.