Method for generating backdoor attack defense model based on object detection
Through the backdoor attack defense model generation method based on target detection, the SinGAN and Yolov5 models are used to generate enhanced data sets, which solves the defense problem of backdoor attacks under single label and small data sets, and improves the robustness and security of deep neural networks.
Patent Information
- Application Number
- CN202211245119.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-12
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-10-12
AI Technical Summary
The existing backdoor attack defense methods lack effective defense methods in single-label, small data set scenarios, making it difficult to detect and defend backdoor attacks without significantly reducing model performance.
The backdoor attack defense model generation method based on target detection is adopted, and the enhanced data set is generated through multi-scale learning through SinGAN model, and the backdoor defense model is trained and generated by Yolov5 model, combining target detection technology for defense.
It realizes efficient defense against backdoor attacks in single-label and small data set scenarios, reducing the pressure of resource recomputation, and improving the robustness and security of deep neural networks.
Smart Images

Figure CN115632843B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of backdoor attack defense, unsupervised learning technology, and object detection technology, and particularly relates to a method for generating a backdoor attack defense model based on object detection. Background Art
[0002] In recent years, with the continuous development of the Internet and the continuous expansion of its application scope, a large number of software have been designed and developed by people. However, the software is very likely to contain different types of adversarial attacks such as viruses, Trojans, adware, worms, etc. The emergence of adversarial attacks will cause serious security risks and lead to serious consequences such as virus attacks, privacy information leakage, and infringement of users' interests and privacy. On the one hand, network security technicians continuously improve and optimize the adversarial attack detection technology; on the other hand, in order to evade security detection, the producers of adversarial attacks also continuously change and hide the attack methods. With the continuous development of deep learning, adversarial interference attacks are widespread, and a variety of black-box, white-box, and gray-box adversarial attack algorithms have been proposed, posing a serious threat to the security of machine learning and its applications. On the basis of traditional adversarial attack algorithms, backdoor attack methods are introduced into common attack means. The backdoor attack on the deep neural network (DNN) can achieve high attack performance without reducing the original performance of the model (that is, it will not cause a rapid decrease in the accuracy of the original DNN model). Therefore, the traditional defense method of judging whether the model is damaged by detecting the reduction of the original performance of the model is ineffective for backdoor attacks, and the backdoor attacks are more concealed.
[0003] The related technologies for backdoor attack detection in network security have undergone multiple evolutions. However, with the continuous evolution of the attackers' technology, efficient and secure detection, as well as improving the robustness of deep neural networks, remain important tasks in the field of network space security. For the backdoor attack on the deep neural network for image classification, in the early stage, the backdoor defense was mainly carried out by detecting fixed positions and fixed forms in the image, such as determining whether there are white squares, round blocks, etc. in the lower right corner. With the gradual in-depth research, the existing defense methods are mainly divided into three scenarios: 1) Before / during training, by detecting the backdoor trigger in the training process or before training and filtering out the backdoor samples from the data set; 2) After training, after the model has been trained, using the damaged model and the clean data set to determine whether the model or the training set is poisoned; 3) During actual online operation, this stage is mainly to judge it when the project is commercially launched. At present, there are fewer defense measures in this stage and the difficulty is relatively high.
[0004] Although significant results have been achieved in the defense technology against backdoor attacks based on deep learning, related methods show that current defense mechanisms mainly focus on multi-label and multi-dataset scenarios. There has not been in-depth analysis and understanding of the defense against backdoor attacks in application scenarios with single-label and small datasets. As a result, existing defense methods are ineffective in this area. Therefore, there is an urgent need to design and implement a robust and secure backdoor attack defense model from the perspective of single-label and small datasets. Summary of the Invention
[0005] An object of an embodiment of the present invention is to provide a method for generating a backdoor attack defense model based on object detection, which has high defense capabilities in scenarios with single-label and small datasets, and can also alleviate problems such as resources and pressure for service recalculation.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is a method for generating a backdoor attack defense model based on object detection, including the following steps:
[0007] Step 1: Select reference data, and select a clear image X from a single label class y to be detected as the reference data;
[0008] Step 2: Use the SinGAN model to perform multi-scale learning on the reference data, randomly initialize the parameters of the SinGAN model, and select the data of the 0th layer with large differences as the enhanced clean dataset D c ;
[0009] Step 3: Given a training backdoor set D with a quantity of N and a size of (w, h), and place D t at any position in the clean dataset D t . The synthesized sample set is denoted as the poisoned training set D c ; p ;
[0010] Step 4: Use the trained object detection tool Yolov5 model to generate a backdoor defense model M p from the poisoned training set D (x,y,w,h) .
[0011] Advantages of the present invention:
[0012] 1. The present invention realizes a "plug-and-play" backdoor defense model, which can alleviate problems such as resources and pressure for service recalculation.
[0013] 2. The present invention expands the defense perspective from large samples with fine granularity to small samples through SinGAN, and uses the YOLOv5 model for object detection attacks to expand the defense target from single-label to multi-label.
[0014] 3. Facing the security vulnerability problem of backdoor attacks existing in current deep neural networks, defense experiments are carried out from multiple dimensions. Under the implementation of a single sample, the effectiveness and efficiency of this method are verified through multiple detection angles such as single-label, cross-label, single-size, cross-size, and reference images, thereby assisting the research on the robustness mechanism of deep learning models, indicating that the invention has an efficient defense ability in the scenario of single-label and small datasets.
[0015] 4. The invention combines post-defense and target detection technologies, which has important application value and research value for designing and implementing a robust and secure deep learning model for backdoor attack defense. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 is the flowchart of the embodiment of the present invention;
[0018] Figure 2 is the generation architecture diagram of the backdoor defense model implemented in the present invention;
[0019] Figure 3 is the internal implementation structure diagram G of a single node of the deep neural network corresponding to the SinGAN model N , discriminator D n has the same structure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0021] An embodiment of the present invention provides a method for generating a backdoor attack defense model based on object detection. Based on the internal correlation in the normal distribution of the same label dataset, the SinGAN model is used to perform multi-scale learning on the benchmark image to simulate different distributions in the same label and achieve the data augmentation effect. Then, the backdoor type to be detected and the augmented data are synthesized to generate a poisoned training set. Finally, an optimized and improved object detection Yolov5 model is used to generate a defense model against backdoor attacks, improving the robustness of the deep neural network model.
[0022] The specific steps are as follows, as Figure 1 shown:
[0023] Step 1: Select benchmark data. Select a clear image X from a single label class Y to be detected as the benchmark data;
[0024] Step 2: Use the SinGAN model to perform multi-scale learning on the benchmark data. Randomly initialize the parameters of the SinGAN model, and select the data of the 0th layer with large differences as the augmented clean dataset D c ;
[0025] Step 3: Given a training backdoor set D with a quantity of N and a size of (w, h) t , and place D t at any position in the clean dataset D c . The synthesized sample set is denoted as the poisoned training set D p ;
[0026] Step 4: Use the trained object detection tool Yolov5 model to generate a backdoor defense model M p for the poisoned training set D (x,y,w,h) . Thus, the final model has a high defense effect against backdoor attacks on deep neural networks, improving the robustness and security of the deep neural network.
[0027] Among them, in the above Step 1, the single picture X selected from a single label y to be detected is clearly visible. Experimental results show that the clarity of the benchmark picture is proportional to the defense effect. Among them, y ∈ [1, C], and C represents the total number of output label types of the deep neural network.
[0028] In the above Step 2, the SinGAN model captures different data distributions from the multi-scale of a single image to generate a set of internally correlated datasets. The model internally includes: a generator G for generating data distributions n and a discriminator D for discriminating data distributions n , where G n and D nThe internal structure of is a convolutional neural network with five layers of convolutional blocks. The input layer is random noise and the image upsampled to the previous scale. The output layer is the result of the fusion of the convolutional neural network with five layers of convolutional blocks and the upsampled image. Combined with the definition of GAN (Generative Adversarial Network), Figure 3 The generator G shown n Structure, generator G n Described as:
[0029]
[0030] Among them, Z n For noise, represents a convolutional neural network with five layers of convolutional blocks, represents the upsampled version of the image, n represents the current scale, X n It's G N The corresponding output results. The convolutional network consists of 5 convolutional blocks in the form of: Conv(3×3)-BatchNorm-LeakyReLU.
[0031] Discriminator D n The architecture and G n In Be consistent.
[0032] The SinGAN model converts G n The training loss is divided into adversarial loss and reconstruction loss. Define the loss function of the SinGAN model for
[0033]
[0034] in, Represents the adversarial loss function, used to adjust X n The distance between the patch distribution and the patch distribution in the generated sample, Represents the reconstruction loss function, which is used to ensure that each G n A specific set of noise mappings needs to exist in . As a nested optimization objective function, the inner discriminator aims to maximize the loss function value of the two; the outer generator aims to minimize the loss function value, with the aim of ensuring that the perturbation on the original image is searched within a certain "degree".
[0035] Implement data enhancement and record it as clean dataset D c The process is: use the SinGAN model to perform multi-scale learning on the benchmark data, randomly initialize the SinGAN model parameters, and select the 0th layer data with large differences as the enhanced clean data set Dc 。
[0036] Furthermore, briefly describe the principle of the SinGAN model in step 2 above.
[0037] For a single image X, the SinGAN model provides a generator G N ={G0, ……, G n} to represent pyramid training of X (a single image) at n scales. Let X N ={X0, ……, X n} represent the corresponding output results of G N , where X n is the down-sampled version of X multiplied by r n . For a certain r > 1, r represents the base of the down-sampling coefficient.
[0038] The size of the image samples gradually subdivides from coarse-grained to fine-grained, and noise needs to be injected for each size. All generators and discriminators have the same domain, and the size captured during the generation process decreases as the dimension increases. At the initial size, the result is directly output from the noise data, that is, the SinGAN generator G n maps the noise data Z n to which is the output result of G n . The formula is as follows:
[0039]
[0040] As the size decreases, each generator G n adds the details not generated by the previous generator G n-1 . Therefore, for each generator G n , in addition to adding random noise Z n , it also receives the up-sampled version data of the coarser-size image, that is:
[0041]
[0042] All generators have a similar architecture. Specifically, before being fed into the convolutional layer, the noise Z n is added to the image This can ensure that the generator does not ignore the influence of the noise, as often happens in conditional scenarios involving randomness. The role of the convolutional layer is to generate the missing details in . That is, G n performs the following operations:
[0043]
[0044] Among them, the convolutional network consists of 5 convolutional blocks, in the form of: Conv(3×3)-BatchNorm-LeakyReLU.
[0045] Each generator G n and a discriminator D n form a pair. The discriminator classifies each overlapping block of its input as true or false. Among them, D n has the same architecture as G n in SinGAN defines the training loss of G n as consisting of two parts: adversarial loss and reconstruction loss.
[0046] The adversarial loss function is defined as to penalize the distance between the normal distribution in X n and the normal distribution in the generated samples. Since the WGAN-GP (Wasserstein GAN with Gradient Penalty) loss function can enhance the stability of training, the present invention selects it as the adversarial loss function, and its definition is:
[0047]
[0048] Among them, represents the output of the WGAN generator G n , represents the sample after adding random noise to ; represents the "gradient penalty" added in the WGAN-GP model. D w means the output of the discriminator.
[0049] To ensure that there is a specific set of noise mappings for each G in SinGAN n , on the basis of the reconstruction loss function is introduced and expressed as:
[0050]
[0051] Among them, represents the image generated using the noise at the corresponding nth scale, and Z * is the initially fixed noise data and remains fixed during training.
[0052] Finally, the loss function of SinGAN can be obtained as:
[0053]
[0054] In step 3, the number of backdoor types T needs to be limited within 10 to 50 to achieve an efficient detection effect.
[0055] In step 4, it specifically includes:
[0056] S41: Utilize the poisoned training set D synthesized in step 3 p to generate xml and txt format files corresponding to the described training information;
[0057] S42: Combine the poisoned training set D p and the file information (including txt files and xml files) generated in step S41 into the object detection dataset D of the Yolov5 model y , and use it as input for training the Yolov5 model.
[0058] S43: Set the initial number of input samples batch_size and the number of training loops epoch of the deep neural network according to the environment; randomly initialize the parameters of the Yolov5 model for object detection, and use the Yolov5 model to train the object detection dataset D y until the model loss function no longer changes.
[0059] S44: Finally, the backdoor defense model M (x,y,w,h) can be obtained.
[0060] Preferably, for a method of generating a backdoor defense model based on object detection, this method adopts a backdoor attack detection method integrating object detection technology in backdoor attacks.
[0061] To better illustrate the technical effects of the present invention, a specific example is used to experimentally verify the present invention and compare the technical effects with existing algorithms. In this experiment, the GTSRB (German Traffic Sign Recognition Images) dataset close to the engineering scenario is selected. The task is to recognize 43 different traffic signs, simulating the application scenario of autonomous driving vehicles. The GTSRB dataset contains 39.2K color training images and 12.6K test images, and the image size is set to 32×32. For the model architecture, ResNet-18 is selected as the classifier.
[0062] Following the first random strategy method proposed in the dynamic attack theory, random data is obtained from the normal distribution and applied to the image (such as Figure 2) The trigger size is default set to 4×4. Randomly select a target class and modify a part of the data into adversarial inputs of the target class to modify the training dataset. For a given dataset, select a certain proportion of samples to complete the backdoor operation (between [0.1, 0.2]). After setting the trigger, an attack success rate of over 92% can be achieved, while maintaining a high classification accuracy, that is, ensuring that the classification performance of the model for "clean" pictures without the inserted backdoor trigger is not affected.
[0063] Following the description of the method architecture in this paper, we study the detection performance of the backdoor defense model from multiple dimensions that may occur in the actual scenario of single target attack. In the experiment, the high availability, high effectiveness, and scalability of this method are verified from multiple scenarios such as target label, attack label, different trigger styles, cross-size, same-size, and different reference images. The experimental test results are shown in Table 1.
[0064] Table 1 Multi-dimensional defense performance results
[0065]
[0066]
[0067] In the experiment, except that the reference images in the attack target scenario are obtained from the attack label, the backdoor defense models in other scenarios are obtained from the target labels to be detected. It can be found from Table 1 that:
[0068] (1) While maintaining the classification performance basically unchanged, the backdoor attack success rate is reduced from 99.58% to 2.36% by the defense model, and in some scenarios, the backdoor attack success rate can be reduced to 0%. Therefore, the experimental results prove the effectiveness and efficiency of the backdoor attack defense method proposed in the present invention.
[0069] (2) The defense models in the experiment are all formed based on the assumed backdoor size of the defender. The test results from the cross-size perspective in the image show that when the size of the real backdoor trigger is inconsistent with the size during the training of the defense model, the performance of the backdoor attack defense will show a step-by-step decline. At the same time, when the two sizes are consistent, we re-conduct the experiment from the cross-size perspective while keeping the size of the real backdoor trigger and the size during the training of the defense model consistent. The test results show that excellent backdoor attack detection effects can also be achieved when the two are kept consistent.
[0070] (3) It can be found from the results of different reference image scenarios that when using a blurred image as the reference image, the backdoor attack rate is 3.47% in the speed limit label; however, the backdoor attack rate in the no passing label is 58.89%, and the performance gap is about 16.9 times. Therefore, the clarity of the reference image can affect the detection performance of the backdoor defense model.
[0071] In summary, the backdoor attack defense model generated by the present invention can effectively detect backdoor triggers; compared with traditional detection methods, the present invention has many advantages such as better performance, fewer assumptions, and finer detection points; it can significantly identify the existence of backdoor triggers without significantly reducing the model classification performance. At the same time, the method of enhancing backdoor samples in the present invention assists in the research of the robustness mechanism of the deep neural network model, which is beneficial to training a better and safer backdoor attack defense model.
[0072] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0073] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for generating a backdoor attack defense model based on object detection, characterized in that, Including the following steps: Step 1: Select reference data. Select a clear image X from a single label class y to be detected as the reference data; Step 2: Use the SinGAN model to perform multi-scale learning on the benchmark data. Randomly initialize the parameters of the SinGAN model, and select the data of the 0th layer with relatively large differences as the enhanced clean dataset D c ; Step 3: Given a training backdoor set D with a quantity of N and a size of (w, h) t , and place D t at any position in the clean dataset D c . The synthesized sample set is denoted as the poisoned training set D p ; Step 4: Use the trained object detection tool Yolov5 model to process the poisoned training set D p Generate the backdoor defense model M (x,y,w,h) ; In the second step, the SinGAN model captures different data distributions from the multi-scales of a single image to generate a set of internally correlated datasets. The model internally includes: a generator G for generating data distributions n and a discriminator D for discriminating data distributions n , where G n and D n have an internal structure of a convolutional neural network with five convolutional blocks. The input layer is random noise and the upsampled image from the previous scale, and the output layer is composed of the result of fusing the convolutional neural network with five convolutional blocks and the upsampled image; the generator G n is described as: where Z n is noise, represents a convolutional neural network with five convolutional blocks, represents the upsampled version of the image, n represents the current scale, and X n is the output result corresponding to G N The convolutional network consists of 5 convolutional blocks, and its form is: Conv(3×3)-BatchNorm-LeakyReLU; Discriminator D n has the same architecture as G n in and remains consistent The SinGAN model divides the training loss of G n into adversarial loss and reconstruction loss, and defines the loss function of the SinGAN model as Among them, represents the adversarial loss function, which is used to adjust X n the distance between the patch distribution and the patch distribution in the generated samples, represents the reconstruction loss function, which is used to ensure that each G n needs to have a specific set of noise mappings, As a nested optimization objective function, the internal discriminator aims to maximize the values of both loss functions; the outer generator aims to minimize the loss function values.
2. The method for generating a backdoor attack defense model based on object detection according to claim 1, wherein, In the said Step 1, from a single label y to be detected, y ∈ [1, C], where C represents the total number of output label types of the deep neural network.
3. The method for generating a backdoor attack defense model based on object detection according to claim 1, wherein, In the said Step 3, the number of backdoor types T is between 10 and 50.
4. The method for generating a backdoor attack defense model based on object detection according to claim 1, wherein In the said Step 4, it specifically includes: S41: Utilize the poisoned training set D synthesized in Step 3 p to generate xml and txt format files corresponding to the description training information; S42: Combine the poisoned training set D p and the file information generated in step S41 into the object detection dataset D of the Yolov5 model y , and use it as input to train the Yolov5 model; S43: Set the initial input sample number batch_size and the number of epochs for the deep neural network training loop according to the environment; randomly initialize the parameters of the Yolov5 model for object detection, and use the Yolov5 model to detect the target detection dataset D y for training until the model loss function no longer changes; S44: Finally, the backdoor defense model M can be obtained (x,y,w,Z) .
Citation Information
Patent Citations
Clean tag neural network backdoor implantation system based on general adversarial trigger
CN113255909A
Generating unsupervised adversarial examples for machine learning
US20220253714A1