Adversarial purification method based on de-noising diffusion probability model
By gradually adding and removing noise in the denoising diffusion probability model, the problem that existing adversarial purification methods are difficult to remove adversarial perturbations introduced by uncertain attack methods is solved, and effective purification of multiple adversarial attacks and the robustness of deep learning models is achieved.
Patent Information
- Application Number
- CN202510066890.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The existing adversarial purification methods are difficult to effectively remove adversarial perturbations introduced by uncertain attack methods.
The adversarial purification method based on the denoising diffusion probability model is adopted. By gradually adding noise in the forward process of the denoising diffusion probability model, and gradually removing noise by using the reverse process to obtain the purified image.
This method can effectively remove perturbations introduced by multiple adversarial attacks, improve the robustness of deep learning models, and does not require modification of the target model structure, and is highly operable and scalable.
Smart Images

Figure CN120014283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of anti-attack technology, and in particular to an anti-attack purification method based on a denoising diffusion probability model. Background Art
[0002] In recent years, deep learning models have achieved remarkable results in the field of computer vision, covering multiple application scenarios such as image classification, object detection, face recognition, autonomous driving, and medical image analysis. These technologies have not only greatly improved the computer's ability to understand visual tasks, but also promoted the widespread application of artificial intelligence in all walks of life.
[0003] However, despite the excellent performance of deep learning models in multiple tasks, their high dependence on input data also makes them particularly vulnerable to adversarial attacks. Attackers can make deep learning-based classification models produce wrong predictions by adding carefully designed, imperceptible perturbations to the input data. With the continuous development of deep learning technology, people are paying more and more attention to this security issue, especially when these models are applied to key fields such as finance, medical care, autonomous driving, etc., adversarial attacks may lead to serious consequences.
[0004] Existing adversarial defense methods can be divided into two main categories:
[0005] 1) Adversarial training: The core idea is to introduce adversarial samples into the training data set and continue training on the target model so that it can learn the characteristics of the adversarial samples, thereby improving the robustness of the target model. The shortcomings of adversarial training methods are mainly in three aspects: first, both the production and training of adversarial samples require a lot of computing power and time costs; second, using adversarial training to change the parameters of the target model may reduce the classification accuracy of the model for non-adversarial samples; third, most adversarial training methods can only defend against attacks of known categories, but cannot handle unknown threats.
[0006] 2) Adversarial purification: This method removes adversarial perturbations from adversarial samples before they are input into the target model, so there is no need to introduce additional data or retrain the target model. The purification model used in this method is trained independently of the attack model and the target model, so it has the ability to prevent unknown types of adversarial attacks. However, how to obtain a purification model with excellent performance is still a task of research significance.
[0007] Through the analysis of existing defense against adversarial attacks, it can be seen that although the current defense measures against specific attack methods have achieved remarkable results and can alleviate the threat posed by adversarial attacks to the security of deep neural network models to a certain extent, due to the different implementation methods of different adversarial attacks, both at the data level and the model level, the diversity of attack methods makes it still a huge challenge to deal with multiple adversarial attacks at the same time. Summary of the invention
[0008] The present invention proposes an adversarial purification method based on a denoising diffusion probability model to solve the technical problem that the existing adversarial purification methods are difficult to purify uncertain attack modes.
[0009] In order to solve the above technical problems, the present invention provides an adversarial purification method based on a denoising diffusion probability model, comprising the following steps:
[0010] Step S1: In the forward process of the denoising diffusion probability model, noise is gradually added to the input image in T steps;
[0011] Step S2: Using the reverse process of the denoising diffusion probability model, the noise is gradually removed by adopting a method matching the reverse step of the noise adding process to obtain a purified image.
[0012] Preferably, the noise in step S1 is Gaussian mixture noise.
[0013] Preferably, in step S1, the Gaussian mixed noise is generated by a linear combination of several Gaussian distributions with different means and variances.
[0014] Preferably, in step S1, noise is gradually added to the input image in T steps until the image becomes a pure random noise image conforming to a Gaussian normal distribution.
[0015] Preferably, the denoising diffusion probability model is a probability model based on a deep neural network.
[0016] Preferably, the deep neural network includes a convolutional layer, a deconvolutional layer and a fully connected layer.
[0017] Preferably, the deep neural network performs denoising and denoising operations on the image through supervised learning or unsupervised learning.
[0018] Preferably, in step S1, the input image includes a clean image and / or an adversarial image with disturbance.
[0019] Preferably, the sources of the adversarial images include: white box attack based on gradient information, black box attack, fast gradient sign attack FGSM, projected gradient descent PGD and adversarial sample generation algorithm DeepFool.
[0020] Preferably, the image purified in step S2 is output to a classifier for classification to determine the anti-purification effect.
[0021] The beneficial effects of the present invention include at least: the present invention is a plug-and-play adversarial purification method that can be directly integrated with existing deep learning models. In practical applications, the present invention, as a preprocessing module, purifies images during the data input stage to help deep learning models better cope with adversarial attacks. The implementation of this solution does not require changes to the target model structure, so it has high operability in actual deployment; the technical solution of the present invention has good scalability and can be adjusted and optimized according to different application requirements, such as combined with other defense mechanisms (such as adversarial training, model regularization, etc.) to further enhance the overall defense capability; as an additional technical feature, the present invention combines the advantages of the Gaussian mixed noise model and the diffusion model, and can more accurately simulate and remove a variety of complex noises in the image. By adopting mixed noises of multiple Gaussian distributions, the present invention exhibits stronger adaptability and robustness when dealing with a variety of adversarial threats. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A schematic diagram of a method flow of an embodiment of the present invention;
[0023] Figure 2 A schematic diagram of visualization results of a purified FGSM adversarial image according to an embodiment of the present invention;
[0024] Figure 3 A schematic diagram of visualization results of a purified PGD adversarial image according to an embodiment of the present invention;
[0025] Figure 4 A schematic diagram of visualization results of a purified DeepFool adversarial image according to an embodiment of the present invention; DETAILED DESCRIPTION
[0026] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.
[0027] The embodiment of the present invention provides an adversarial purification method based on a denoising diffusion probability model, comprising the following steps:
[0028] Step S1: In the forward process of the denoising diffusion probability model, noise is gradually added to the input image in T steps;
[0029] Step S2: Using the reverse process of the denoising diffusion probability model, the noise is gradually removed by using a method that matches the reverse step of the noise adding process, and the image is gradually restored to the original image that has not been attacked by the adversarial attack to obtain a purified image, thereby ensuring the removal of the adversarial disturbance in the image.
[0030] Specifically, the input image in step S1 can be a clean image or an adversarial image with disturbance, and the image types include but are not limited to handwritten digital images, natural images, etc., as well as scenes that may contain different visual features and contents. This method is applicable to adversarial attack purification in various image types and can provide strong robustness in different application fields.
[0031] The adversarial images with perturbations come from adversarial attacks, which include but are not limited to white-box attacks based on gradient information, black-box attacks, fast gradient sign attacks FGSM, projected gradient descent PGD, and adversarial sample generation algorithm DeepFool. They can effectively remove various adversarial perturbations and restore the original visual features and semantic information of the image.
[0032] In this embodiment, the added noise is mixed Gaussian noise, wherein the mixed Gaussian noise is generated by a linear combination of multiple Gaussian distributions with different means, variances and mixing coefficients, and the multiple Gaussian distributions can be adaptively adjusted according to the content of the input image and the attack type. Specifically, different Gaussian distributions will have different noise intensities, and the distribution area of each noise in the image can be dynamically changed according to the local features of the input image. In this way, the noise characteristics in complex images can be fitted more flexibly. Compared with the traditional single Gaussian noise, the Gaussian mixed noise can better capture the data distribution of the image.
[0033] After the above-mentioned denoising and denoising process, the final output image can not only remove the disturbance introduced by the adversarial attack, but also the purified image is consistent with the original image that has not been attacked in terms of semantic information and visual quality.
[0034] When the method of this embodiment is used, Figure 1 As shown, the following steps are included:
[0035] First, we select a representative standard image dataset, the MNIST handwritten digit dataset, which contains 10 types of handwritten Arabic numerals from 0 to 9. Each image in the training set and the test set is resized to 28x28 pixels, and all images are ensured to be single-channel grayscale images.
[0036] Use the training set of the MNIST dataset to train a DNN-based classifier and a denoising diffusion probability model respectively: in the forward process of the denoising diffusion probability model, we gradually add Gaussian mixed noise to the image over T steps until it becomes a pure noise image that conforms to the Gaussian distribution, and remove the noise over T steps in the reverse process. The training model learns how to restore the clean original image from the noise.
[0037] Different adversarial attack methods are used to attack the MNIST test set to generate adversarial samples, such as FGSM, PGD, and DeepFool. The generated adversarial samples are input into the denoising diffusion probability model trained in step 101. Through the process of gradual denoising and denoising, the original distribution and semantic information of the image are gradually restored, and the purified image is output.
[0038] The output purified image is sent to a trained DNN classifier for classification test to verify the adversarial purification effect of the present invention.
[0039] The adversarial training (AT) defense method is selected as the baseline and the adversarial purification method (GMDMP) based on Gaussian mixture noise and denoising diffusion probability model proposed in this invention is compared. The experimental results are shown in Table 1 and the visualization results are shown in Figure 2 , Figure 3 and Figure 4 shown.
[0040] Table 1 Comparison results between the GMDMP solution of the present invention and AT
[0041]
[0042] There is no significant decrease in the classification accuracy of GMDMP on clean samples. Compared with the results of the adversarial training method based on FGSM adversarial samples (FGSM-AT), the classification accuracy of GMDMP is only slightly lower than that of FGSM-AT when defending against FGSM attacks, while the classification accuracy of GMDMP when dealing with PGD and DeepFool attacks is much higher than that of FGSM-AT.
[0043] According to the above experimental results, the adversarial purification method proposed in the present invention can defend against most types of adversarial attacks and effectively improve the robustness of the classifier.
[0044] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is more specific and detailed, but it cannot be understood as limiting the scope of the present invention. As long as there is no contradiction in the combination of these technical features, they should be considered as within the scope of this specification.
[0045] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.
Claims
1. An adversarial purification method based on a denoising diffusion probability model, characterized by: The following steps are involved: Step S1: In the forward process of the denoising diffusion probability model, noise is gradually added to the input image in T steps; Step S2: Using the reverse process of the denoising diffusion probability model, the noise is gradually removed by adopting a method matching the reverse step of the noise adding process to obtain a purified image.
2. The method for countermeasure purification based on a denoising diffusion probability model according to claim 1, characterized in that: The noise in step S1 is Gaussian mixture noise.
3. The method for countermeasure purification based on a denoising diffusion probability model according to claim 2, characterized in that: In step S1, the Gaussian mixed noise is generated by a linear combination of several Gaussian distributions with different means and variances.
4. The method for countermeasure purification based on a denoising diffusion probability model according to claim 3, characterized in that: In step S1, noise is gradually added to the input image in T steps until the image becomes a pure random noise image that conforms to the Gaussian normal distribution.
5. The method for countermeasure purification based on denoising diffusion probability model according to claim 1, characterized in that: The denoising diffusion probability model is a probability model based on deep neural network.
6. The method for countermeasure purification based on a denoising diffusion probability model according to claim 5, characterized in that: The deep neural network includes a convolutional layer, a deconvolutional layer and a fully connected layer.
7. The method for countermeasure purification based on denoising diffusion probability model according to claim 6, characterized in that: The deep neural network performs denoising and denoising operations on the image through supervised learning or unsupervised learning.
8. The method for countermeasure purification based on denoising diffusion probability model according to claim 1, characterized in that: In step S1, the input image includes a clean image and / or an adversarial image with disturbance.
9. The method for countermeasure purification based on denoising diffusion probability model according to claim 8, characterized in that: The sources of the adversarial images include: white-box attack based on gradient information, black-box attack, fast gradient sign attack FGSM, projected gradient descent PGD and adversarial sample generation algorithm DeepFool.
10. The method for countermeasure purification based on denoising diffusion probability model according to claim 1, characterized in that: The image purified in step S2 is output to the classifier for classification to determine the effect of the adversarial purification.
Citation Information
Cited By
Double-branch collaborative confrontation defense method for network flow intrusion detection
CN121792248A
A dual-branch cooperative adversarial defense method for network traffic intrusion detection
CN121792248B