An Adversarial Patch Generation Method Based on a Dynamic Optimization Integration Model

Through dynamic optimization of the integration model and loss function optimization, the generated adversarial patches show efficient migration and attack effects on multiple object detection models, solving the problem of insufficient migration between different models in the existing technology, and achieving effective deception of adversarial patches in the digital and physical worlds.

CN117151207BActive Publication Date: 2025-07-25NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311197147.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-07-25
Estimated Expiration
2043-09-18

AI Technical Summary

Technical Problem

When facing unknown black box models, existing adversarial patch generation methods are difficult to achieve efficient migration between different data sets and models, and ignore the structural differences of the target detection model, resulting in poor attack effects of adversarial samples on multiple target models.

Method used

The dynamic optimization integration model is adopted, and the integration model with dynamic optimization of parameters is constructed, which is divided into internal maximization and external minimization processes. The weight of the object detection model is dynamically adjusted, and an adversarial patch is generated that is effective for all attacked target detection models. The transformation function is used to simulate physical environment transformation, and the pixel value and smoothness of the adversarial patch are optimized in combination with the loss function.

Benefits of technology

The generated adversarial patches show good migration and attack effects on multiple object detection models, and can effectively deceive detectors in the digital and physical worlds, reduce the accuracy of object detection, and achieve generalization and stability of adversarial patches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117151207B_ABST
    Figure CN117151207B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence security technology, and specifically relates to an adversarial patch generation method based on a dynamically optimized integrated model. Initialize the adversarial patch, and use a transformation function to paste the adversarial patch on the original image to obtain a sample adversarial image; input the sample adversarial image into multiple object detection networks to obtain the output of each object detection network; establish an adversarial patch integrated model and obtain the first loss function of the integrated model; obtain the second loss function according to the pixel values of the pixel points in the adversarial patch; obtain the third loss function according to the pixel values of two adjacent pixel points in the adversarial patch; obtain the total loss function according to the first loss function, the second loss function, and the third loss function; use the integrated model to perform a preset number of iterations to obtain the adversarial patch when the total loss function is minimized. The present invention trains multiple object detection models at the same time and can generate adversarial patches that are effective for all attacked object detection models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence security technology, and in particular to an adversarial patch generation method based on a dynamically optimized integrated model. Background Art

[0002] Deep neural networks are highly expressive models. In 2013, Christian Szegedy et al. first discovered the vulnerable side of neural networks. After a large number of experiments, it was proved that adding a negligible perturbation to the original sample would cause the classification model to make mistakes. The sample after adding the perturbation is an adversarial sample, and the adversarial patch is a special form of the adversarial sample, becoming a popular research object in the field of adversarial attacks. The adversarial patch covers a local area of the image with a sticker-like pattern. Just placing the patch on the object to be measured can successfully deceive the victim model. Therefore, how to generate adversarial patches with high attack ability and transfer ability is the most important problem at present.

[0003] In the prior art, Wu et al. analyzed the attack performance and transferability of adversarial patches based on different datasets, target models, and attack categories. Hao et al. balanced the gradients of two object detectors in training by adding adversarial patches at key pixel positions of the image, enabling the adversarial patch to attack both YOLOv4 and Faster R-CNN simultaneously. With the rapid development of the target detection field, a large number of related works on personnel detection in the physical world have also emerged. However, when attacking in the physical world, an unknown black-box model is often faced. To improve its transferability between different data and different models, most research works enhance transferability by improving the optimization process, data augmentation, or scrambling the feature maps of intermediate layers. Xie et al. proposed to improve the transferability of adversarial samples by creating different input patterns, randomly transforming the input image at each iteration instead of only using the original image to generate adversarial samples. The adversarial samples generated by this method transfer better to different networks than existing benchmarks. Huang et al. proposed a transferable attack algorithm based on integrated gradients (TAIG), which can find highly transferable adversarial samples for black-box attacks. Different from previous methods that used multiple computational terms or combined with other methods, TAIG integrates three methods into one process. Zhu et al. studied the transferability of adversarial samples from the perspective of data distribution, proposed to form adversarial samples by manipulating the distribution of images, and assumed that removing the image from its original distribution could enhance the transferability of the adversarial. In non-target attacks, removing the image from its original distribution makes it difficult for different models to correctly classify the image; in target attacks, moving the image into the target distribution will mislead the model to classify the image as the target class. However, these methods interfere with features without selection due to not considering the intrinsic features of the objects in the image, and are prone to falling into model-specific local optima.

[0004] At present, the method based on integrated model is widely used in the transfer problem. Based on the idea of integrated model, an adversarial sample that can deceive multiple white-box models with different structures at the same time is trained, so the attack success rate will be higher when facing unknown black-box models. For example, Liu et al. studied the transferability of large models and large-scale data sets and proposed a method based on the integration of multiple target models for the first time to generate transferable adversarial samples; Zheng et al. proposed the Meta Gradient Adversarial Attack (MGAA) architecture, which can be integrated with any existing gradient-based attack method to improve the transferability of adversarial samples across models. However, the above work is still in the field of image classification, and there is little research on integrated models for dense prediction models (such as object detection). At the same time, most of the existing integrated model work fuses the outputs of all models together, ignoring the differences in the structures of different models, resulting in the optimization of adversarial samples always moving in a single direction, and it is difficult for adversarial samples to converge to have an attack effect on all target models. Summary of the invention

[0005] In order to solve the problems existing in the prior art, the present invention proposes a universal adversarial patch generation framework based on a dynamic optimization integration model. A method for generating adversarial patches based on a dynamic optimization integrated model is provided. In order to improve the generalization of adversarial patches and enhance the migration ability of adversarial patches between different models, the idea of adversarial training is introduced to construct an integrated model with dynamic parameter optimization. The entire training process is divided into an internal maximization process, in which the total loss output is maximized by dynamically adjusting the weight of the attacked target detection model, and an external minimization process, in which the total loss output is minimized by updating the adversarial patches. Through continuous adversarial training, the weights of each model are dynamically adjusted until a balance is achieved, and adversarial patches that are effective for all attacked target detection models are generated. The scheme specifically includes: initializing the adversarial patch, using a transformation function to paste the adversarial patch on the original image to obtain a sample adversarial image; inputting the sample adversarial image into multiple target detection networks to obtain the output of each target detection network; establishing an adversarial patch integrated model and obtaining a first loss function of the integrated model; obtaining a second loss function according to the pixel values of pixels in the adversarial patch; obtaining a third loss function according to the pixel values of two adjacent pixels in the adversarial patch; obtaining a total loss function according to the first loss function, the second loss function and the third loss function; and using the integrated model to perform a preset number of iterations to obtain an adversarial patch when the total loss function is minimized. The present invention simultaneously trains multiple target detection models and can generate adversarial patches that are effective for all attacked target detection models.

[0006] The present invention adopts the following technical solution: a method for generating adversarial patches based on a dynamic optimization integrated model, comprising:

[0007] Initialize the adversarial patch, and use the transformation function to paste the adversarial patch on the original image to obtain the sample adversarial image;

[0008] Input the sample adversarial image into multiple object detection networks respectively, and obtain the output of each object detection network;

[0009] Establish an integrated model of the adversarial patch according to the output of each object detection network and the weight corresponding to each object detection network, and obtain the first loss function of the integrated model based on the output of each object detection network;

[0010] Obtain the second loss function of the integrated model according to the pixel values of the pixel points in the adversarial patch and the printable color range of the printer;

[0011] Obtain the third loss function of the integrated model according to the pixel values of two adjacent pixel points in the adversarial patch;

[0012] Obtain the total loss function of the integrated model according to the first loss function, the second loss function and the third loss function;

[0013] Use the integrated model to perform a preset number of iterations, and obtain the adversarial patch output when the total loss function is the smallest.

[0014] Further, the transformation function is specifically:

[0015] X adv =M(p,X)=t((1 - M p )·X + t d (M p ·t TPS (p)))

[0016] The transformation function mainly represents the process of covering the adversarial patch on the original image through the transformation function to obtain the adversarial sample. Among them, M(·) is the transformation function, and t TPS represents performing TPS transformation on the generated adversarial patch, and t d represents blurring the image to different degrees according to the distance. t is used to simulate common transformations in the physical environment, such as changes in light, angle, etc. p is the adversarial patch, X is the original image, and M p ∈[0,1] d represents the masked area of the adversarial patch on the target image.

[0017] Further, after obtaining the output of each object detection network, it further includes:

[0018] Obtain the confidence of the target candidate boxes output by each object detection network; the target candidate boxes include positive sample information and negative sample information;

[0019] Obtain target candidate boxes with confidence greater than a set threshold, and calculate the confidence scores of the target candidate boxes with confidence greater than the set threshold to construct the loss function corresponding to the target detection network.

[0020] Further, the integrated model of the adversarial patch is:

[0021]

[0022] This formula is a basic integrated model constructed to calculate the total loss output of all target detection networks. Among them, i represents the i-th target detection network, K represents the number of target detection networks, w represents the weight parameter of the target detection network, p is the adversarial patch, and f(·) is the output of the integrated model.

[0023] Further, the expression of the first loss function of the integrated model is:

[0024]

[0025] This formula represents the total loss output of all attacked target detection networks in the designed dynamic optimization integrated model. Among them, the weight parameter w is defined as a probability simplex, satisfying W = {w|1 T w = 1, w i ∈[0,1]}, is the output of the i-th target detection network, where f j (·) represents the information contained in the j-th candidate box output by the target detection network, and μ is used to filter candidate boxes with confidence scores higher than it for all outputs, is the regularization term, where γ > 0 is the regularization parameter.

[0026] Further, the second loss function of the integrated model is specifically:

[0027]

[0028] This formula represents the total unprintable loss, which is used to calculate the pixel error after printing the patch. Among them, c print is the printable color, and i patch is a pixel in the adversarial patch p.

[0029] Further, the third loss function of the integrated model is specifically:

[0030]

[0031] This formula represents the smoothness loss, which is used to calculate the smoothness error after the adversarial patch covers the image, can improve the smoothness of the perturbed image, and enhance the feasibility of physical attacks. Among them, p i,jis a pixel value at the coordinate (i, j) in the adversarial patch p.

[0032] The beneficial effects of the present invention are as follows: The present invention proposes a dynamic optimization integrated model for the transferability problem in existing adversarial attacks. At the same time, it trains multiple object detection models, and can stabilize the training results by dynamically adjusting the weight parameters of the attacked object detection model. By combining the generation of adversarial patches and the dynamic optimization integrated model, the training process of adversarial patches is divided into two parts through constructing the loss function of the integrated model. First is the internal maximization, dynamically adjusting the weight of the attacked object detection model to maximize the total loss output. Then is the external minimization, minimizing the total loss output by updating the adversarial patch. Finally, through continuous adversarial training, dynamically adjusting the weight parameters of each model until balance, it can effectively generate adversarial patches that are effective for all attacked object detection models. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 is a schematic flowchart of a method for generating adversarial patches based on a dynamic optimization integrated model according to an embodiment of the present invention;

[0035] Figure 2 is a schematic flowchart of the implementation of a transformation function according to an embodiment of the present invention;

[0036] Figure 3 is a schematic diagram for verifying the experimental results of digital image attacks according to an embodiment of the present invention;

[0037] Figure 4 is a schematic diagram of the experimental results of a comparative experiment for generating adversarial patches according to an embodiment of the present invention;

[0038] Figure 5 is a schematic diagram of an adversarial patch T-shirt in the physical world according to an embodiment of the present invention;

[0039] Figure 6 is a schematic diagram of the indoor object detection results according to an embodiment of the present invention;

[0040] Figure 7 is a schematic diagram of the outdoor object detection results according to an embodiment of the present invention;

[0041] Figure 8 is a schematic diagram of the misrecognition results of object detection according to an embodiment of the present invention;

[0042] Figure 9 Schematic diagram of a category change result according to an embodiment of the present invention;

[0043] Figure 10 Schematic diagram of the attention distribution of an image by a target detection network according to an embodiment of the present invention. Detailed implementation manners

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] In order to generate adversarial patches that have a certain attack effect on mainstream target detectors, and at the same time have good migration ability and can be printed on fabrics so that the attack effect can also be retained in the physical world, the present invention explores the generation mechanism of adversarial patches. The adversarial patches are first processed by a transformation function and then overlaid on clean samples, and then the obtained adversarial samples are used as the input of a dynamic optimization ensemble model. The loss output by the ensemble model is L patch . To improve the attack effect of the adversarial patches, two parts, namely patch pixel value constraint and smoothness adjustment, are added. The output unprintable score loss and smoothness loss are used as another part of the loss function. Finally, the above three losses are combined to obtain the total loss, and the adversarial patches are updated by gradient descent optimization to minimize the total loss, realizing the migration of the adversarial patches between different target detection models. A method for generating adversarial patches based on a dynamic optimization ensemble model according to an embodiment of the present invention is as Figure 1 shown, including:

[0046] Initialize the adversarial patch, and use the transformation function to paste the adversarial patch on the original image to obtain a sample adversarial image;

[0047] Different from traditional adversarial example generation methods, adversarial patches are local perturbations applied to images. To achieve the migration of adversarial patches from the digital world to the physical world, when adding an adversarial patch to the surface of an object, a set of physical transformations T is usually added according to the Expectation over Transformation (EoT) algorithm to simulate contrast, brightness transformation, angle transformation, etc. in the physical environment. However, in the real physical environment, the scale of the patch decreases as the distance increases, and the imaging becomes blurrier. The present invention adds a distance transformation function on the basis of EOT and uses a Gaussian kernel to blur the image, that is, the smaller the target object, the larger the convolution kernel, and the blurrier the image after the patch transformation. The goal of the present invention is to print an adversarial patch pattern on a T-shirt so that people can be invisible under the target detection algorithm. Therefore, Thin Plate Spin (TPS) is added to simulate the non-rigid changes on the fabric surface, such as wrinkles. According to the above transformation measures, the transformation function M(·) is designed to paste the adversarial patch onto the target image. The specific transformation function is as follows:

[0048] X adv = M(p,X) = t((1 - M p )·X + t d (M p ·t TPS (p)))

[0049] where p is the adversarial patch, X is the original image, M p ∈[0,1] d represents the masked area of the adversarial patch on the target image, which is determined by the target candidate box obtained from the label. The M(·) function mainly includes three transformation methods. First, perform TPS transformation (t TPS ∈T TPS ) on the generated adversarial patch to make the image distorted to simulate the non-rigid changes on the clothes surface. Then, judge the distance according to the size of the target candidate box in the input image, and realize different degrees of blurring of the patch image through the distance transformation function (t d ∈T d ). Finally, simulate the common transformations in the physical environment through conventional physical transformations (t∈T). The overall process is as Figure 2 shown.

[0050] Input the sample adversarial images into multiple target detection networks respectively, and obtain the outputs of each target detection network;

[0051] For mainstream object detectors, the present invention needs to construct an appropriate loss function. By reducing the loss, the object detector is made to fail in detecting the category of "person". Image X and the ground truth label Y are used as the inputs to the detector. Through the detection network, the information and confidence scores of the candidate bounding boxes of the target categories are output, and then the candidate bounding boxes with scores higher than a preset threshold are selected as the prediction results. Therefore, the present invention takes the expectation of the output confidence scores as part of the optimization of the adversarial patch loss function, and reduces the confidence scores of the targets by optimizing the adversarial patches.

[0052] In a specific embodiment, first, the adversarial patch is randomly initialized. Through the transformation function M(·), the patch is pasted on the original image as the input to the object detection network f(·). Conventionally, the maximum probability of the target category output by the detection network is directly taken as part of the loss function. In each iteration process of the present invention, all N candidate bounding boxes output by the detection network are taken as the optimization objects, and the confidence changes in different scenarios are comprehensively considered. At the same time, in order to reduce the calculation, the present invention selects the candidate bounding boxes with the confidence scores of the category of "person" higher than the set threshold μ as the attack objects, and finally continuously optimizes and updates the pixels of the patch through gradient descent.

[0053] After obtaining the output of each object detection network, it further includes: obtaining the confidence of the target candidate bounding boxes output by each object detection network; the target candidate bounding boxes include positive sample information and negative sample information; obtaining the target candidate bounding boxes with confidence greater than the set threshold, and calculating the confidence scores of the target candidate bounding boxes with confidence greater than the set threshold to construct the loss function corresponding to the object detection network.

[0054] Thus, the present invention calculates the confidence scores of all candidate bounding boxes with scores higher than the threshold μ in the output of the object detection network for the adversarial samples, and takes them as part of the loss function for training the adversarial patch. By iteratively optimizing the adversarial patch, the total loss is minimized, and the adversarial patch is optimized during the training process to reduce the loss and make the object detector fail in detecting the category of human. That is, by training adversarial patches with complex features, the probabilities of other categories are increased, thereby reducing the probability of the category of "person". The loss function is defined as:

[0055]

[0056] An integrated model of the adversarial patch is established according to the output of each object detection network and the weight corresponding to each object detection network, and a first loss function of the integrated model is obtained based on the output of each object detection network;

[0057] To achieve the "invisibility" of pedestrians under multiple object detection algorithms, the generated adversarial patch can affect the prediction information output by all detection algorithms and reduce the probability of the category "person" in the output. In the present invention, for K object detection networks, the output of each object detection network is f(·), which is used to adjust the weights of different network models and satisfies Then the integrated model of the generated adversarial patch is expressed as:

[0058]

[0059] Among them, i represents the i-th object detection network, K represents the number of object detection networks, w represents the weight parameter of the object detection network, p is the adversarial patch, and f(·) is the output of the integrated model.

[0060] However, for detection algorithms with large structural differences such as YOLO and Faster R-CNN, it is difficult to adjust the w weight to the optimal balance point, and overfitting is likely to occur during the training process. To solve the above problems, based on the idea of Min-Max adversarial training, the present invention constructs a dynamic optimization integrated model for multi-object detection algorithms. During the training process, the proportion of each detection model in the total loss output is dynamically adjusted through the w parameter to achieve the generality of the adversarial patch. The entire training is divided into an external minimization process, where the parameter w is fixed, and the adversarial patch is optimized through backpropagation to minimize the output loss of the integrated model; and an internal maximization process, where the adversarial patch is fixed, and w is optimized through projected gradient descent to balance the weights of different models and maximize the total loss output. However, this single-coding form will reduce the generalization ability among other models, that is, the phenomenon that the training always biases towards one model appears, resulting in unstable training. Therefore, a strong concave regularizer is introduced in the internal maximization step to alleviate this problem. Thus, the expression of the first loss function of the integrated model is calculated as:

[0061]

[0062] Among them, the weight parameter w is defined as a probability simplex, satisfying W = {w|1 T w = 1, w i ∈[0, 1]}, is the output of the i-th object detection network, where f j (·) represents the information contained in the j-th candidate box output by the object detection network, and μ is used to filter the candidate boxes whose confidence scores of all outputs are higher than it. is a regularization term, where γ > 0 is the regularization parameter. The adversarial patch is designed to be trained under the maximum attack loss (the worst-case attack scenario). If γ → ∞, the inner maximization means calculating the average of K attack losses. Therefore, the regularization parameter γ adjusts the balance between the maximum loss attack strategy and the average attack strategy, making the adversarial patch have a better attack effect on all targeted detection models.

[0063] Obtain the second loss function of the integrated model according to the pixel values of the pixel points in the adversarial patch;

[0064] The objective of the present invention is to generate a universal adversarial patch that is effective in both the digital world and the physical world, with a focus on optimizing the pixel values and smoothness of the patch. Therefore, in a specific embodiment, the present invention optimizes the color difference between the adversarial patch and a conventional printer by evaluating the unprintable score of the adversarial patch pixel values, thereby constructing the second loss function of the integrated model of the present invention as:

[0065]

[0066] where c print is a printable color, i patch is a pixel in the adversarial patch p, and the color range (color gamut) that devices such as printers and screens can reproduce is only [0,1] 3 a subset of the RGB color space.

[0067] Obtain the third loss function of the integrated model according to the pixel values of two adjacent pixel points in the adversarial patch;

[0068] To ensure the smoothness of the generated adversarial patch, the present invention calculates the smoothness loss of the adversarial patch as the third loss function of the integrated model, specifically:

[0069]

[0070] where p i,j is a pixel value at the coordinate (i,j) in the adversarial patch p. When the values of adjacent pixels are close, the smoothness loss L smooth is low (i.e., the perturbation is smooth). Therefore, by minimizing L smooth the smoothness of the perturbed image can be improved, enhancing the feasibility of the physical attack.

[0071] Obtain the total loss function of the integrated model according to the first loss function, the second loss function, and the third loss function;

[0072] Finally, generate an adversarial patch that is effective in both the physical world and the digital world. The total training loss is expressed as follows:

[0073] Lattack = L patch + α·L NPS + β·L smooth

[0074] where α and β are weight coefficients determined according to experience, and the ultimate goal of optimization is to minimize the loss function L attack , each part of the loss function will be minimized together, and finally the adversarial patch output when the total loss function is minimized is obtained by performing a preset number of iterations using the integrated model.

[0075] The embodiment of the present invention further provides an attack experiment on the adversarial patch to verify the attack effect of the adversarial patch of the present invention. In the digital image attack experiment, the present invention uses the Inria Person dataset, which includes a training set of 614 pedestrian images and a test set of 288 images. The goal of the adversarial patch generated by the present invention is to achieve the "invisibility" of pedestrians under the object detection algorithm and reduce the recognition accuracy of various algorithms. Therefore, in order to evaluate the attack effect and transferability of the adversarial patch, first, the Average Precision (AP) value of the image before and after placing the adversarial patch is calculated to measure. The greater the decrease in AP, the stronger the attack effect. The adversarial patch generated by the solution of the present invention is mainly used to deceive the mainstream first-order object detection algorithms, YOLOv2, YOLOv3, and YOLOv4, and has an attack effect in both large object and small target detection. For the more complex second-order algorithm, it can make the prediction of the Faster R-CNN detection model based on the backbone of Vgg16 and Resnet50 fail. All detection models are trained on the coco dataset, the confidence threshold in the output stage is set to 0.5, and the Non-Maximum Suppression (NMS) threshold is set to 0.4.

[0076] The present invention designs three different detection model sets to train adversarial patches, namely DOEPatch(YY): obtained by training an integrated model composed of YOLOv2 and YOLOv3; DOEPatch(YF): obtained by training an integrated model composed of YOLOv3 and Faster-Vgg16; DOEPatch(YYF): obtained by training an integrated model composed of YOLOv2, YOLOv3 and Faster-Vgg16. And an adversarial patch function redesigned based on the design, the adversarial patches EmPatch(Y1), EmPatch(Y2), EmPatch(F) trained for YOLOv2, YOLOv3, Faster R-CNN. The entire algorithm is designed based on pytorch, trained on a 3090 device, the adversarial patch is set to a size of 3*300*300, the Adam optimizer is used, the learning rate for external minimization optimization is set to 0.03, the weight coefficients α = 0.01, β = 0.165 determined according to experience, the regularization term coefficient γ = 0.1, and the internal dynamic adjustment rate parameter ν is obtained by analyzing the experimental results on DOEPatch(YF). The schematic diagram of the verification results is as Figure 3 shown. It can be seen that when ν = 0.78, the adversarial patches generated by the present invention achieve better attack effects on both YOLOv2 and YOLOv3.

[0077] In order to more fully reflect the superiority of the adversarial patches designed based on the dynamic optimization integrated model in terms of attack performance and transferability, the embodiments of the present invention provide six groups of comparative experiments by comparing with the related algorithms of adversarial patches in the prior art. The experimental results are as Figure 4 shown, where Figure 4 (a) is the adversarial patch Noise composed of random noise; Figure 4 (b) is the solid-color adversarial patch Blue; Figure 4 (c) is the adversarial patch AdvPatch designed by Thys et al. for the YOLOv2 algorithm; Figure 4 (d) is the adversarial patch TCEGA designed by Zhanhao Hu et al. for the YOLOv2 algorithm; Figure 4 (e) is the adversarial patch NaPatch designed based on StyleGAN to attack the YOLOv2 algorithm; Figure 4 (f) is the adversarial patch UpcPatch for the FasterR-CNN algorithm. The adversarial patches EmPatch(Y1), EmPatch(Y2), EmPatch(F) designed by the present solution for YOLOv2, YOLOv3, FasterR-CNN are respectively as Figure 4 (g), Figure 4 (h), Figure 4(i); The adversarial patches DOEPatch(YY), DOEPatch(YF), and DOEPatch(YYF) designed based on the dynamic optimization integration model are shown in Figure 4 (j), Figure 4 (k), Figure 4 (l) respectively.

[0078] It should be noted that Figure 4 AdvPatch and TCEGA in [reference] are trained using the same dataset as this solution, and the YOLOv2 models and parameters used are also the same, making them highly comparable. While NaPatch and UpcPatch focus more on the concealment of adversarial patches, there are certain differences in the dataset compared to the present invention.

[0079] In the embodiments of the present invention, all adversarial patches are attached to the images of the Inria test set, and the effectiveness of the attack is reflected by calculating the change in the AP value of the "person" category in the outputs of different detection algorithms. The experimental results in the digital world are shown in Table 1. Clean represents the AP value of the "person" category in the original test set images. At the same time, the AP values of the test sets with 12 groups of adversarial patches attached are respectively tested under 5 mainstream object detection algorithms. Under the detection of YOLOv2, DOEPatch(YY) reduces the AP value by 67.49%, slightly lower than AdvPatch. However, the attack performance of DOEPatch(YY) is better than that of other adversarial patches such as TCEGA. At the same time, it achieves the best attack effect under the detection of YOLOv3, and can reduce the AP value to 29.20%. The adversarial patch EmPatch(Y2) generated by the present invention can reduce the AP value by 54.11% under the detection of YOLOv3; the attack performance is also much higher than the test results of other patches. For the more complex Faster-Vgg detection algorithm, DOEPatch(YF) and DOEPatch(YYF) achieve the highest attack effects, reducing the AP values by 43.94% and 44.81% respectively; they also have strong attack effects under the detection of YOLOv3, reducing the AP values by 41.91% and 41.65% respectively; due to the differences in the data sets, UpcPatch for Faster R-CNN only reduces the AP value by 6.10%; in order to more comprehensively compare the transferability of adversarial patches, black-box attack tests are carried out on YOLOv4 and Faster-Res algorithms. DOEPatch(YY) trained by the present invention can reduce the AP value by 50.55% under the detection of YOLOv4, while AdvPatch and TCEGA can only reduce it by 29.14% and 28.21% respectively. Under Faster-Res, the attack performances are all poor, but DOEPatch(YY) and DOEPatch(YF) still achieve the best attack effects, reducing by 18.55% and 17.05% respectively. Generally speaking, the adversarial patches generated based on the dynamic optimization integration model can all produce good attack effects among the target models to be attacked, and can also cause certain interference to the output results when facing unknown black-box models, and have good transferability.

[0080] Table 1 Comparison table of experimental results in the digital world

[0081]

[0082]

[0083] In another specific embodiment, the present invention further conducts an experimental evaluation in the physical world. To verify the attack effect of adversarial patches in the physical world, the present invention prints the adversarial patches on a T-shirt to simulate the object detection process in a real environment. Specifically, the adversarial patches in the digital world are composed of 300*300 pixels, and the present invention magnifies them to 30 cm * 30 cm, and covers the pattern of the adversarial patches on a pure cotton white T-shirt by offset printing for testing. As Figure 5 shown, to simulate the detection scenario in the real environment, the present invention uses a c920e camera to record two videos of pedestrian movements in two scenarios, indoor and outdoor. One of the pedestrians is wearing an adversarial T-shirt, and the detection results are as Figure 6 and Figure 7 shown. It can be seen that the attack effect of the adversarial patches generated by the present invention in the physical world is that YOLOv2 and YOLOv3 can identify pedestrians not wearing adversarial T-shirts in all frames in different environments, while the adversarial T-shirts made by the present invention with DOEPatch(YY), DOEPatch(YF), and DOEPatch(YYF) enable another pedestrian to achieve "invisibility" in both indoor and outdoor scenarios. Therefore, the adversarial patches generated by the present invention can successfully affect the recognition of the object detection system in the physical world, enable people wearing adversarial T-shirts to avoid detector recognition, and have good transferability under multiple detection algorithms.

[0084] In another specific embodiment, the present invention conducts an interpretability analysis of adversarial patches based on class transformation. Although most current work has studied the interpretability of adversarial samples, there is no relevant work that gives a comprehensive explanation of the attack performance for adversarial patches. Therefore, the present invention analyzes from the perspective of class transformation and combines the visualization tool Grad-CAM to make a more comprehensive interpretability analysis of the phenomenon that adversarial patches affect object detection.

[0085] The adversarial patches generated based on the dynamic optimization integration model have complex and distinct texture features in terms of color and feature distribution, similar to traditional adversarial patches, as Figure 8 shown. Figure 8 (a) shows the detection result of the adversarial patch DOEPatch(YYF) under YOLOv2, which is misclassified as the "teddy bear" class. Figure 8 (b) shows the detection result of the adversarial patch DOEPatch(YF) under YOLOv3, which is misclassified as the "apple" class. Figure 8 (c) shows the detection result of the adversarial patch EmPatch(F) under Faster R-CNN, which is misclassified as the "pottedplant" class. Just from some of the detection results, it is difficult to analyze the internal mechanism of adversarial patches.

[0086] From the perspective of class transformation, the present invention describes the class change process by comparing the number of class candidate boxes after screening. Figure 9 (a) shows the change of candidate boxes output by Faster R-CNN after adding EmPatch(F). Figure 9 (b) shows the transformation of candidate boxes output by YOLOv3 after adding DOEPatch(YF) to the test samples. Figure 9 (c) shows the change of candidate boxes output by YOLOv2 after adding DOEPatch(YYF) to the test samples. The left half of the bar chart represents the number of candidate boxes output after the original samples are predicted by the detection network, and the right half represents the output after detecting the adversarial samples. Specifically, it shows that the number of candidate boxes representing the "person" category will decrease sharply after adding the adversarial patch, while the number of some other categories will increase to a certain extent. By training adversarial patches with complex features, the probability of other categories is increased, thereby reducing the probability of the "person" category to a certain extent.

[0087] The present invention further uses Vgg16 pre-trained on ImageNet as the basic model and combines the Grad-CAM algorithm to show the change of the attention map before and after attacking the target image. After experiments on the test set, mainly two change results are produced as Figure 10 shown in Figure 10 (a) row represents the original sample and the adversarial sample after adding the adversarial patch. Figure 10 (b) row represents the recognition result under the detector. Figure 10 (c) and Figure 10 (d) row represents the detection results under Gard-CAM. As Figure 10 (b) shows, in result 1, the person is not recognized, and the attention is dispersed by the adversarial patch, and the feature distribution of the person changes due to the adversarial patch. In result 2, the adversarial patch is misrecognized, and the attention is completely attracted to the adversarial patch, resulting in a change in the decision of the model. Generally speaking, the adversarial patch can affect the distribution of target features in the detection network and interfere with the prediction result.

[0088] The present invention proposes a dynamic optimization integration model for the transferability problem in existing adversarial attacks. At the same time, it trains multiple object detection models, and can stabilize the training results by dynamically adjusting the weight parameters of the attacked object detection model. By combining the generation of adversarial patches and the dynamic optimization integration model, the training process of adversarial patches is divided into two parts by constructing the loss function of the integration model. First is the internal maximization, which dynamically adjusts the weights of the attacked object detection model to maximize the total loss output. Then is the external minimization, which minimizes the total loss output by updating the adversarial patches. Finally, through continuous adversarial training, the weight parameters of each model are dynamically adjusted until balanced, and adversarial patches that are effective for all attacked object detection models can be effectively generated. The present invention further analyzes the impact of adversarial patches on the class distribution in the target image from the perspective of class transformation, and at the same time combines the Grad-CAM visualization tool to analyze the image feature distribution before and after the attack, thus proposing a complete interpretability analysis framework for adversarial patches.

[0089] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An adversarial patch generation method based on a dynamic optimization integration model, characterized in that Including: Initializing an adversarial patch, and pasting the adversarial patch on the original image by using a transformation function to obtain a sample adversarial image; Inputting the sample adversarial image into multiple object detection networks respectively to obtain the outputs of each object detection network; Establishing an integrated model of the adversarial patch according to the output of each object detection network and the weight corresponding to each object detection network, and obtaining a first loss function of the integrated model based on the output of each object detection network; Obtaining a second loss function of the integrated model according to the pixel values of the pixel points in the adversarial patch and the printable color range of the printer; Obtaining a third loss function of the integrated model according to the pixel values of two adjacent pixel points in the adversarial patch; Obtaining a total loss function of the integrated model according to the first loss function, the second loss function and the third loss function; Performing a preset number of iterations by using the integrated model to obtain the adversarial patch output when the total loss function is the smallest.

2. The adversarial patch generation method based on a dynamic optimization integration model according to claim 1, wherein: The transformation function is specifically: X adv = M(p, X) = t((1 - M p )·X + t d (M p ·t TPS (p))) where M(·) is a transformation function, t TPS represents performing TPS transformation on the generated adversarial patch, t d represents blurring the image to different degrees according to the distance, and t is used to simulate the illumination and angle changes in the physical environment, p is the adversarial patch, X is the original image, M p ∈[0,1] d represents the masked area of the adversarial patch on the target image.

3. The method for generating adversarial patches based on a dynamic optimization integration model according to claim 1, wherein: After obtaining the output of each object detection network, it further includes: Obtaining the confidence of the target candidate boxes output by each object detection network; the target candidate boxes include positive sample information and negative sample information; Obtaining the target candidate boxes with confidence greater than a set threshold, and calculating the confidence scores of the target candidate boxes with confidence greater than the set threshold to construct the loss function of the corresponding object detection network.

4. A method for generating adversarial patches based on a dynamic optimization integration model according to claim 1, characterized in that: The integrated model of the adversarial patch is: Where, i represents the i-th object detection network, K represents the number of object detection networks, w represents the weight parameter of the object detection network, p is the adversarial patch, and f(·) is the output of the object detection model.

5. A method for generating adversarial patches based on a dynamic optimization integration model according to claim 1, characterized in that: The expression of the first loss function of the integrated model is: Among them, the weight parameter w is defined as a probability simplex, satisfying W = {w|1 T w = 1, w i ∈ [0, 1]}, is the output of the i-th object detection network, where f j (·) represents the information contained in the j-th candidate box output by the object detection network, and μ is used to filter candidate boxes whose confidence scores are higher than it, is a regularization term, where γ > 0 is a regularization parameter.

6. The adversarial patch generation method based on a dynamic optimization integration model according to claim 1, wherein The specific expression of the second loss function of the integrated model is: where c print is the printable color range of the printer, and i patch is a pixel in the adversarial patch p.

7. A method for generating adversarial patches based on a dynamic optimization integration model according to claim 1, characterized in that: The specific expression of the third loss function of the integrated model is: where p i,j is a pixel value at the coordinate (i, j) in the adversarial patch p.

Citation Information

Patent Citations

  • Adversarial sample generation method for unmanned aerial vehicle image target detection

    CN113643278A

  • Self-adaptive adversarial patch generation method and device facing physical attack

    CN115829877A