Adversarial patch generation-based infrared image target detection algorithm adversarial attack method
By constructing an adversarial patch generation model based on loss function optimization and Bernoulli random dropout, the problems of insufficient algorithm generalization ability and high computational overhead of the infrared image target detection algorithm in adversarial attacks are solved. The generated adversarial patches can effectively attack multiple target detection models, improving the robustness and security of the infrared image target detection algorithm.
Patent Information
- Application Number
- CN202510758053.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-17
AI Technical Summary
Existing infrared image target detection algorithms have problems such as insufficient algorithm generalization ability, high computational overhead, and poor migration performance when facing adversarial attacks. In particular, the lack of dedicated datasets and effective attack methods in infrared scenes makes the detection model susceptible to interference, and the performance of existing methods is significantly degraded when processing high-resolution images.
An adversarial patch generation model based on loss function optimization and Bernoulli random dropout is constructed. The quality of adversarial patches is evaluated by optimizing the loss function. The Bernoulli random dropout method is introduced to simulate multiple target detectors to generate adversarial patches with better migration performance. The local masking strategy is combined to prevent overfitting.
The generated adversarial patches can significantly reduce the recognition accuracy of infrared image target detection algorithms, enhance the effectiveness and adaptability of adversarial patches, improve cross-model transfer performance, and effectively attack multiple target detection models.
Smart Images

Figure CN120807994A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of infrared target detection model attack, in particular to an infrared image target detection algorithm adversarial attack method based on adversarial patch generation. BACKGROUND
[0002] The adversarial attack research of infrared image target detection algorithm originates from the urgent needs of the military security field. Infrared imaging systems play a core role in battlefield reconnaissance, border monitoring and other scenarios due to their night vision capabilities and anti-camouflage characteristics. However, their images generally have inherent defects such as lack of detail, blurred texture, and low information dimension, making target detection systems based on deep learning more susceptible to interference. Attackers can implement interference on unmanned reconnaissance devices or intelligent security systems by designing adversarial patches and other physically realizable attack forms, such as adding a small disturbance pattern to the surface of a key target, which can cause the detection algorithm to fail. The severity of this threat has been verified in the field of autonomous driving - research shows that adding specific disturbances to traffic signs can cause the classifier's misclassification rate to reach 99%, and if similar attack methods are applied to infrared detection systems, it may cause catastrophic consequences. At the same time, theoretical research shows that deep neural networks have inherent vulnerability, even in a simple two-layer ReLU network, a single step of gradient descent can generate effective adversarial samples, which provides a theoretical basis for the design of attack methods.
[0003] Current adversarial attack methods mainly revolve around three major technical routes. The generative adversarial network method generates adversarial samples in bulk by constructing a generator-discriminator framework, with AdvGAN as a typical representative achieving an attack success rate of 92.76% on the MNIST dataset. The generated perturbations have strong transferability, but are limited by model complexity and perform significantly worse when processing high-resolution images. The decision boundary analysis method is represented by DeepFool and the universal perturbation algorithm, which calculates the shortest perturbation distance from the sample to the classification boundary to achieve efficient attack. DeepFool generates L2 norm perturbations faster than traditional methods, but it is difficult to deal with the complex multi-target decision boundary in the target detection task. The gradient optimization method has evolved from FGSM single-step attack to PGD multi-step iteration, with MI-FGSM introducing a momentum mechanism to improve optimization stability, and CW method breaking through various defense mechanisms through confidence adjustment, but there is a contradiction between computational efficiency and attack effectiveness.
[0004] In addition, the existing method faces multiple challenges when migrating to an infrared detection scene. First, a large number of studies focus on visible light images, and only a small number of attack methods are directed at infrared characteristics, and there is a lack of dedicated data sets. Second, the algorithm generalization ability is insufficient, such as Dpatch is effective for YOLOv2 but ineffective for YOLOv3 and YOLOv4, and the TC-EGA method enhances the physical robustness through ring clipping, but is only applicable to human body detection and has limited view angle adaptability. At the technical implementation level, the problem of difficult convergence of GAN framework training is more prominent in the infrared scene, and the binary search process required by the CW method produces a huge computational overhead when processing high-resolution infrared images. SUMMARY
[0005] In view of the problems existing in the prior art, the present application proposes an infrared image target detection algorithm adversarial attack method based on adversarial patch generation, and an adversarial patch generation model based on loss function optimization and Bernoulli random dropout is constructed.
[0006] The present application first optimizes the adversarial attack loss function in terms of target hiding, structural similarity, pixel value, etc., to better evaluate the quality and effect of the generated adversarial patch and guide the algorithm to generate more effective patch patterns; secondly, a Bernoulli random dropout method is designed, which simulates a variety of target detectors by randomly dropping part of the neural network layers, improving the migration performance of the adversarial patch to multiple networks.
[0007] The technical scheme of the present application is as follows:
[0008] An infrared image target detection algorithm adversarial attack method based on adversarial patch generation, comprising the following steps:
[0009] Step 1: Obtain an infrared target detection model D and original input sample images: x1,...,x N ;
[0010] Step 2: Initialize the adversarial patch patch0, and iteratively optimize the adversarial patch until the set iteration termination condition is reached, and output the trained adversarial patch; for a certain iteration, the specific process is as follows:
[0011] Step 2.1: input each batch of original infrared images into the model D to obtain the target box position and confidence;
[0012] Step 2.2: apply the adversarial patch of this iteration to the original input infrared image to obtain the adversarial sample x adv ;
[0013] Step 2.3: perform Bernoulli random dropout on part of the neural network layers of the model D and perform local masking of the adversarial patch to obtain the processed model D';
[0014] Step 2.4: Get the adversarial example x obtained in step 2.2 adv The target box confidence and location information in the model D' obtained in step 2.3;
[0015] Step 2.5: Calculate the loss function value based on the target box position and confidence obtained in step 2.1 and the target box position and confidence obtained in step 2.4;
[0016] Step 2.6: Update the adversarial patch using the Adam optimizer based on the loss function value obtained in step 2.5; then proceed to the next iteration;
[0017] Step 3: Use the trained adversarial patches to conduct adversarial attacks on the infrared image target detection algorithm.
[0018] Furthermore, in step 2.3, the Bernoulli random dropout process is:
[0019] During the forward propagation, the input Z l The stacked layer output F(Z l ) add to get the output Z L :
[0020] Z L =Z l +(b+aa·b)·F(Z l )
[0021] The parameter a is sampled from the continuous uniform distribution U(0,1), and b is sampled from the Bernoulli distribution;
[0022] During the back propagation process, the loss function is defined as J, then:
[0023]
[0024] c is a Bernoulli variable, distributed in the same way as b.
[0025] Furthermore, in step 2.3, the local masking process of the adversarial patch is:
[0026] For the width and height of W and H respectively, the normalized input image Patch x ,y∈[0,1], with probability ζ=0.8, perform the following operations:
[0027] (1) Randomly sample within the adversarial patch and obtain a point p = (x0, y0);
[0028] (2) Generate a λH×λW rectangular area with point p as the center on the adversarial patch, where λ is the scale and the filling value is k∈[0,1].
[0029] Further, in step 2.5, the loss function is:
[0030] l total = l det + a l s + b l ssim (x, x adv ) + g l c
[0031] where a, b, g are hyperparameters; l det is the target hidden loss, l s is the smoothing loss, l ssim (x, x adv ) is the structural similarity loss, l c is the adversarial patch pixel value loss.
[0032] Further, for a single-stage target detector, the target hidden loss is:
[0033]
[0034] where n is the number of target boxes output by the target detector for the target class, is the confidence of the target box output by the target detector, represents the true value of the target box confidence.
[0035] Further, for a two-stage Faster R-CNN infrared target detection model, the target hidden loss is:
[0036]
[0037] where the set class probability value corresponding to the i-th region proposal box is and x adv is the adversarial sample after the adversarial patch is applied to the original infrared image.
[0038] Further, the smoothing loss is
[0039]
[0040] Further, the structural similarity loss is
[0041]
[0042] where n is the adversarial sample, m is the original infrared image sample, C is a constant, s mn is the covariance of the infrared images m and n, s m , s n are the standard deviations of the infrared images m, n, respectively.
[0043] Further, the adversarial patch pixel value loss is
[0044]
[0045] p j S is the pixel value corresponding to any one pixel point in the adversarial patch. patch S is the pixel value set of the adversarial patch.
[0046] Advantages
[0047] The present application combines the actual reconnaissance monitoring application of the infrared system, studies the adversarial patch generation algorithm based on loss function optimization and Bernoulli random discard, and proposes an infrared image target detection algorithm adversarial attack method based on adversarial patch generation. The adversarial patch is generated to attack the infrared image target detection algorithm, so that the recognition accuracy is reduced, so as to test the accuracy change of the infrared image target detection algorithm model and analyze its vulnerability.
[0048] The main advantages of the present application are as follows:
[0049] (1) Reduce the target detection accuracy: by generating adversarial patches, the recognition ability of the infrared image target detection algorithm to the target is significantly reduced, effectively attacking the target detection model, so as to evaluate the robustness and security of the model in practical application.
[0050] (2) Improve the effectiveness and adaptability of the adversarial patch: the adversarial patch generation algorithm in the present application is optimized by the loss function, fully considers various factors such as target hiding, structural similarity, pixel value, etc., so that the generated adversarial patch has higher quality and can more effectively interfere with the discrimination ability of the target detection model.
[0051] (3) Enhance the cross-model migration: by introducing the Bernoulli random discard strategy, the generated adversarial patch not only targets a single detection model, but also has strong cross-model migration performance. This method simulates the characteristics of multiple detectors, so that the adversarial patch can also produce attack effects under different neural network structures, improving the universality of the patch.
[0052] Additional aspects and advantages of the present application will be partially given in the following description, partially will become obvious from the following description, or will be understood by the practice of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0053] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description of the embodiments, taken in conjunction with the accompanying drawings, in which:
[0054] Figure 1 : principle diagram of a traditional adversarial attack method;
[0055] Figure 2The framework diagram of the method for resisting attacks is provided in the application.
[0056] Figure 3 The principle diagram of Bernoulli random discard method is shown in (a) forward propagation and (b) backward propagation.
[0057] Figure 4 The P-R curve of the YOLOv5s target detector is shown.
[0058] Figure 5 The schematic diagram of the adversarial patch is shown.
[0059] Figure 6 The loss function value change condition is shown. DETAILED DESCRIPTION
[0060] The embodiments of the application are described in detail below, which are exemplary and intended to explain the application, and cannot be understood as a limitation of the application.
[0061] The principle of the method for attacking the infrared image target detection algorithm by generating an adversarial patch is shown in Figure 1 First, a target object is selected from an input image and an adversarial disturbance is added to obtain an adversarial sample. Then, the adversarial sample is sent to a target detection model to obtain a model output result. Then, according to a pre-defined adversarial attack algorithm loss function, the error between the model output result and the target is calculated. Next, a gradient descent or iterative optimization method is used to seek the best adversarial disturbance. Finally, the new adversarial disturbance is added to the input image to obtain an adversarial sample, and target detection is performed again.
[0062] When the infrared image target detector is affected by the adversarial patch, the adversarial attack loss will guide the detector to reduce the confidence of recognizing the target, so that the target confidence result is lower than the threshold and is filtered out, and finally the infrared image target detector misses the specified target.
[0063] However, the traditional method has defects in the loss function and the migration performance. Therefore, the embodiment constructs a generative adversarial patch adversarial attack algorithm (GADP) based on Bernoulli random discard and loss function optimization, and the algorithm framework is shown in Figure 2 By using the Bernoulli random discard method, the neural network part of the target detector is randomly discarded, which approximates to simulate multiple target detectors, so that the target detection network structure is more general, and the adversarial patches generated by these detectors can have attack effects on multiple target detection algorithms at the same time. At the same time, by optimizing the loss function of the adversarial attack, the generated adversarial patch can be more suitable for the infrared image scene, thereby improving the attack success rate of the algorithm.
[0064] (1) Overall steps:
[0065] Step 1: Obtain the infrared target detection model D and the original input sample images: x1,...,x N ;
[0066] Step 2: Initialize the adversarial patch patch0 and iteratively optimize the adversarial patch until the set iteration termination condition is reached, output the trained adversarial patch; for a certain iteration, the specific process is as follows:
[0067] Step 2.1: input each batch of original infrared images into model D to obtain the target box position and confidence;
[0068] Step 2.2: apply the adversarial patch of this iteration to the original input infrared image to obtain the adversarial sample x adv ;
[0069] Step 2.3: perform Bernoulli random dropout on the neural network part of model D and perform local masking of the adversarial patch to obtain the processed model D';
[0070] Step 2.4: obtain the adversarial sample x adv in the model D' obtained in step 2.3 target box confidence and position information;
[0071] Step 2.5: calculate the loss function value according to the target box position and confidence obtained in step 2.1 and the target box position and confidence obtained in step 2.4;
[0072] Step 2.6: update the adversarial patch using the Adam optimizer according to the loss function value obtained in step 2.5; then proceed to the next iteration;
[0073] Step 3: use the trained adversarial patch to perform adversarial attack on the infrared image target detection algorithm.
[0074] The pseudo code of the GADP algorithm is as follows:
[0075]
[0076]
[0077] (B) Bernoulli random dropout:
[0078] For the Bernoulli random dropout method in step 2.3, the specific description is as follows:
[0079] In infrared target detection tasks, there are often multiple alternative infrared target detection models to choose from. Therefore, it is difficult for attackers or testers to determine which model the other party is using, and it is also difficult to obtain the model's parameters and gradient information. In order to improve the portability of adversarial patches, there is a method that integrates multiple target detection models to train adversarial patches together, but this method is more difficult, requires knowing which target detection model the other party is using, and has a certain amount of time and resource costs. In order to make the adversarial attack method more portable and reduce training costs, the present invention proposes a Bernoulli random drop method based on the random depth algorithm and introduces a local masking strategy for adversarial patches.
[0080] The Random Depth algorithm is an algorithm designed to reduce the depth of a network during training but maintain it during testing. This algorithm randomly deletes entire residual blocks during model training and bypasses the deleted blocks through skip connections. The following formula illustrates how Random Depth works:
[0081] H l =RELU(T l f l (H l-1 ))+id(H l-1 )
[0082] Among them, the Bernoulli random variable T l ∈{0,1}, id(·) represents the identity mapping. When T l = 1, the lth residual block is in the active state; when T l = 0, the lth residual block will be deleted. The output before the lth residual block is: H l-1 , the output after the lth residual block is: H l . Through the function f l By multiplying with the Bernoulli random variable, this algorithm can skip the lth residual block in the neural network. However, discarding a large number of shallow networks will cause the model to lose input data features and reduce accuracy. The random depth algorithm combines the linear decay rule to define the survival probability p of the lth residual block l =P(T l =1) = 1-l / L(1-p L ). The survival probability of the input layer is p0=1, and it will not be discarded. The survival probability of the last residual block is: p L .
[0083] Inspired by the random depth algorithm, the present invention constructs the Bernoulli random dropout method. By randomly dropping some layers in the neural network model, it approximately integrates multiple variants of the target detection model to improve the generalization performance of generating adversarial patches. This method can make the neural network model randomly retain or drop a certain proportion of features, thereby enhancing the transferability of the adversarial attack algorithm. The principle diagram is as followsFigure 3 is shown.
[0084] In the forward propagation process, the Bernoulli random dropout method randomly discards the input Z l and adds the output F(Z l ) of the stacked residual block to obtain the output Z L . Wherein the parameter a is sampled from the continuous uniform distribution U(0, 1), and b is sampled from the Bernoulli distribution. The formula is as follows:
[0085] Z L = Z l + (b + a - a·b)·F(Z l )
[0086] In the backward propagation process, the loss function is defined as J, and according to the chain rule, the Bernoulli random dropout method can be expressed as:
[0087]
[0088] c is a Bernoulli variable, which is distributed with b. Using this method, the neural network part of the target detector is randomly discarded, avoiding the problems of overfitting and gradient disappearance of deep neural networks. This method trains multiple "small models" after random dropout, which are used to approximate multiple target detectors, speed up the training speed, and improve the attack ability of the adversarial patch on different target models.
[0089] The local masking strategy of the adversarial patch is: in the training stage, the local masking strategy of the adversarial patch is used to randomly mask a local area of the adversarial patch, preventing the patch from overfitting to the specific target detection model and the training data.
[0090] The specific process is: for an input image Patch x , y ∈ [0, 1], with a probability ζ = 0.8, the following operations are performed:
[0091] (1) Randomly sample a point p = (x0, y0) in the adversarial patch;
[0092] (2) Generate a λH×λW rectangular region centered at point p on the adversarial patch, λ is the proportion, and the filled value is k ∈ [0, 1].
[0093] Similar to the random erasing method in the training of convolutional neural networks, the local masking strategy of the adversarial patch prevents the adversarial patch from relying too much on the local area.
[0094] (Three) Loss function
[0095] In the adversarial patch generation process, the loss function is very important to measure the difference between the target detector prediction and the expected output. Considering the realizability in the real world, compared with visible light images, infrared images are more difficult to present complex textures in the physical world due to the imaging mechanism based on thermal radiation. Therefore, when generating patches, patches with relatively few patterns are generated as much as possible. In this embodiment, the adversarial attack loss function is optimized to make it more suitable for adversarial attacks on infrared image target detection algorithm models.
[0096] In the infrared image target detection task, the adversarial patch generation algorithm is designed for the target to be identified. In order to make the infrared target detection model unable to identify the target, the GADP adversarial attack algorithm constructed needs to ensure that there is a significant deviation between the expected output and the actual infrared image label, so as to ensure that the expected output does not contain the specified target category. In order to optimize the texture and image information of the adversarial patch, the loss function needs to be constrained to ensure that the target detection model cannot identify the target. Generally, the purpose of designing the loss function is to make the adversarial attack algorithm generate effective adversarial patches faster and make the target detector predict errors as much as possible. The total loss function formula used is as follows:
[0097] l total =l det +αl s +βl ssim (x,x adv )+γl c
[0098] Wherein, α, β, γ are hyperparameters. The following introduces the target hiding loss l det , the smoothing loss l s , the structural similarity loss l ssim (x,x adv ) and the adversarial patch pixel value loss l c .
[0099] For a single-stage target detector, the confidence of the target object will determine whether the predicted bounding box contains the specified target category. In order to make the target detector unable to identify the target, the adversarial attack algorithm reduces the confidence of the target object to 0 as much as possible.
[0100] The target hiding loss designed for the target in the infrared image is:
[0101]
[0102] Wherein, n is the number of target boxes output by the target detector about the target category, is the confidence of the target box output by the target detector, represents the true value of the target box confidence. In order to make the target detector ignore the target in the infrared image, is zero, then:
[0103]
[0104] For the two-stage Faster R-CNN infrared target detection model, when generating the proposed region, the probability distribution of all N proposed regions is: {p1,p2,...,p n To hide the target in the infrared image and prevent the Faster R-CNN model from recognizing it, the adversarial algorithm reduces the foreground probability of the target of interest, and the area in the image other than the background is considered the foreground. At this point, the target hiding loss for the Faster R-CNN infrared target detection model is:
[0105]
[0106] The set category probability value corresponding to the i-th region suggestion box is And there are x adv is the adversarial example after the adversarial patch is applied to the original infrared image. det The smaller it is, the lower the probability that the target box in the Faster R-CNN detector is of the set category, and the more likely the set category target in the infrared image is to be ignored by the target detection model.
[0107] In infrared images, adversarial patches in the form of noise points are not natural and smooth, and are not easy to physically implement. This embodiment combines a smoothness loss to limit the number of noise points in the texture image used to generate the adversarial patch.
[0108]
[0109] After introducing the smoothness loss, the adversarial patch image becomes more natural and smoother visually, and the number of noise points is greatly reduced. Where i and j represent the pixel coordinates in the adversarial patch p.
[0110] In addition, this embodiment also introduces a structural similarity loss, which can make the adversarial patch image more consistent with the infrared image characteristics while ensuring that the difference between the original sample and the adversarial sample is small. The SSIM structural similarity is expressed as:
[0111]
[0112] σ mn is the covariance of infrared images m and n, μ m 、μ n , σ m , σ nrespectively, C1, C2 are constants. The SSIM value is less than or equal to 1, and the index has symmetry, regardless of the front or back order of the input infrared images. When calculating the structural similarity, start from a small block of N*N size, move the small block to calculate the local structural similarity in units of pixels, until the entire infrared image is traversed, and finally take the mean value of the local structural similarity. The closer the SSIM value is to 1, the more similar the two infrared images are. The mean square error can also measure the similarity of the two infrared images, but it focuses on considering the average brightness of the two images. In the task, more attention is paid to the contrast and structural difference between the adversarial sample and the infrared input image, and the addition of small patch adversarial patch operation has little effect on the overall brightness of the image. Moreover, the structural similarity measurement method is similar to the judgment standard of the human visual system. The resolution of the infrared image is low, and the brightness difference between the adversarial sample and the original image is small. In order to improve the calculation speed of the neural network model and accelerate the convergence, the brightness evaluation in the SSIM structural similarity is omitted in the embodiment. The final structural similarity loss used is:
[0113]
[0114] The size of the small slider in the above formula is 14*14, and the image is traversed from left to right and from top to bottom. The constant C=9.4*10 -4 , n is the adversarial sample, and m is the original infrared image sample. In the adversarial attack task, the smaller the value of l SSIM (m, n), the more original infrared image target information the adversarial sample retains.
[0115] In the embodiment, pixel values are also randomly selected from the infrared image to simulate the infrared radiation intensity in the physical world, so that the generated adversarial patch has infrared image characteristics and is more likely to be reproduced in a real environment. Since the infrared image is a single-channel image, it has fewer available pixel values compared to visible light images. In order to simulate the infrared radiation intensity in the physical world, 50 pixel values are randomly selected from the infrared image in the embodiment, and a pixel value set S pixel is constructed.
[0116] The adversarial patch pixel value loss in the embodiment is:
[0117]
[0118] p j is the pixel value corresponding to any one pixel point in the adversarial patch, S patch is the adversarial patch pixel value set. The smaller l c , the closer the pixel value in the adversarial patch is to the pixel value in the preset set S pixel , wherein the pixel value P j ∈S patchIn practical applications, the set of acceptable pixel values S can be adjusted according to the infrared radiation intensity value of the material pixel .
[0119] (iv) Simulation experiment
[0120] In the simulation experiment, the FLIR dataset is used, and the FLIR dataset is taken from the street and the highway, and the shooting time is not limited to day and night, and contains relatively complex natural scenes. Most of the infrared images are sampled at a speed of 2 per second. In an environment with fewer target objects, the infrared device samples at a speed of 1 per second. Among them, the target categories include people, cars, and bicycles. In combat scenarios, individual infrared night vision helmets are widely used. Soldiers need to use night vision helmets to detect important targets such as people. In order to better train and test the interference effect of the adversarial patch, the public FLIR dataset is filtered in this experiment. The specific conditions are as follows:
[0121] (1) The infrared image contains important targets such as people;
[0122] (2) The body height of the person in the infrared image is greater than or equal to 100 pixels.
[0123] Then, the experiment on the image that meets the conditions is data enhanced, and the infrared image is batch processed. After image flipping, random cropping, and rotation operations, 2600 infrared image data are obtained. However, these are mostly outdoor scenes, in order to make the generated adversarial patch adapt to more scenes, 400 indoor scene infrared images are added. Finally, a total of 3000 infrared image multi-scene dataset AFLIR is constructed. Among them, 2400 infrared images are used as the training set, and 600 images are used as the test set, and the number ratio of the training set to the test set is 4:1.
[0124] The effect of the adversarial patch generated by the present application:
[0125] The goal of this experiment is to make the infrared target detection model unable to detect people. In order to evaluate the adversarial attack effect from multiple aspects, the average precision AP (Average Precision), attack success rate ASR (Attack Success Rate), and AP change value are used to comprehensively evaluate the experimental results.
[0126] This experiment pays more attention to the detection results of the "person" category, and the subsequent AP values all represent the average precision of the "person" category. A lower AP value indicates that the target detector is more likely to misjudge the "person" category after being attacked by the adversarial patch. The attack success rate ASR is commonly used to evaluate the effectiveness of the adversarial attack method, and the formula is as follows:
[0127]
[0128] wherein N is the number of real labels of the target frame of the "person" in the test set T. Natt is the number of target frames predicted by the target detector as "person" in the test set T after the adversarial attack. The higher the attack success rate, the better the effect of the adversarial attack algorithm, and the more difficult it is for the target detector to identify the "person" target.
[0129] The AP change value of the target detector is the absolute value of the difference between the AP value before the attack and the AP value after the attack, that is:
[0130] AP change value = |AP pre -AP last |
[0131] The AP change value facilitates comparison of the attack effects of various adversarial attack algorithms applied to the same target detector. The larger the AP change value, the better the attack effect of the adversarial attack algorithm on the specific target detector.
[0132] The detection model targeted by the experiment is a lightweight model YOLOv5s commonly used for infrared image target detection tasks. The model has fewer parameters and faster inference speed. The model is trained for 200 steps on the AFLIR dataset, and the training results are shown in Figure 4 , the average precision of the YOLOv5s target detector performs well. Among them, the initial average precision of "person" is 81.3%.
[0133] The YOLOv5s target detector uses a confidence threshold of 0.25 and an IoU intersection over union threshold of 0.5, which can avoid multiple repeated prediction frames for a single target. In order to ensure the typical representativeness of the experimental results, the threshold selected in the experiment is consistent with the model threshold commonly used in the industry. The results of detecting the original infrared image by the YOLOv5s target detector show that the target detector can normally detect "person", "car" and "bicycle" categories. The infrared image contains multiple people and single person scenes, and the background contains highways, forests, buildings, streets and other scenes.
[0134] For the YOLOv5s target detector, the adversarial patches generated by the present application are as shown in Figure 5 These adversarial patches will cause the infrared image target detection algorithm to make errors when detecting "person" targets. Figure 5 (a), (b), (c), (d) respectively represent the adversarial patch images when training for 250 steps, 500 steps, 750 steps and 1000 steps. The loss function value change is shown in Figure 6 . In the experiment, 1000 steps of training were performed, the horizontal axis represents the number of iterations, and the vertical axis represents the loss function value. From the loss function value change, it can be seen that after about 60000 iterations, the algorithm converges, and the loss function value is about 0.5 at this time.
[0135] After the adversarial patch is applied to the original image, the adversarial sample is formed. When the input infrared image is an adversarial sample, the YOLOv5s detector cannot identify the person in the adversarial sample, but can still correctly identify the car and bicycle categories. Since the infrared image target detection model is difficult to correctly abstract all the features of the target, especially in a complex background, its abstract expression ability for target objects is low. The interference of the adversarial patch further reduces the ability of the model to extract target features, because it destroys the key structural information of the infrared target, thereby causing the infrared target detection model to fail to identify important targets.
[0136] Comparison of effects with other adversarial attack methods:
[0137] This experiment compares four typical adversarial attack methods with the GADP algorithm constructed in this paper. The QRAttack method attacks the infrared target detector by designing an adversarial two-dimensional code patch pattern; the advYOLOPatch algorithm designs an adversarial patch for "people" so that the infrared target detection model cannot identify "people"; in order to verify whether occlusion and random noise have an attack effect on the infrared target detection algorithm model, this experiment introduces white square patches and random noise patches as controls.
[0138] Without attack, in the infrared image, there is one target box whose category is "person" with a confidence of 0.71. The white square patch, QRAttack, and advYOLOPatch adversarial attack algorithms respectively reduce the confidence of "person" to 0.68, 0.66, and 0.57. The GADP algorithm constructed in this paper has better effect, and after adding the adversarial patch generated by this method, the YOLOv5s infrared target detector cannot correctly detect "person", that is, there is no target box in the infrared image whose category is correctly labeled as "person".
[0139] The quantitative experimental results of this experiment are shown in Table 1.
[0140] Table 1 Comparison results of adversarial attack methods
[0141]
[0142] The attack success rate ASR of the GADP algorithm constructed in the application is the highest, reaching 75.9%, followed by the QRAttack and advYOLOPatch methods, and the success rates are 74.6% and 69.2%, respectively. The white square and random noise patch cannot effectively attack the infrared target detection algorithm model, and the success rate is less than 9%. This is because the influence of occlusion and random noise has been simulated during the training of the infrared target detection network, and the interference has little effect on the key structural information of the infrared image target. The QRAttack method uses a loss function to constrain the shape of the two-dimensional code patch, and the adversarial sample after applying the patch makes the infrared target detector unable to identify the person. The GADP algorithm constructed in the application uses a more optimal adversarial attack loss function, which strengthens the attack effect of the adversarial patch on the blurred infrared image. The loss function of the GADP algorithm considers the characteristics of the infrared image, optimizes the pattern texture structure of the adversarial patch, and further interferes with the key structural information of the infrared image target, so that the average precision AP of the infrared image target detection algorithm model decreases to 15.1%, and it is difficult to correctly detect the person in the infrared image.
[0143] Verification of the transferability of the adversarial attack method
[0144] In this experiment, the mainstream single-stage target detection models YOLOv5s, YOLOv7, SSD and the two-stage target detection model FasterR-CNN were selected to verify the transferability of the GADP adversarial attack algorithm. On the AFLIR dataset, the average precision of the YOLOv5s, YOLOv7, SSD and FasterR-CNN target detection models for the "person" category is 81.3%, 88.4%, 70.9% and 74.2%, respectively. After the adversarial patch is applied to the original image, the generated adversarial sample has transferability. That is, the generated adversarial sample can achieve attack effect on multiple target detection models, causing the detection accuracy of different models to decrease, and the model cannot accurately detect the "person" category in the adversarial sample.
[0145] The following four groups of comparative experiments were conducted, using three adversarial attack algorithms: GADP, advYOLOPatch and QRAttack. These three algorithms all have excellent attack effect and can generate practical adversarial patches.
[0146] In the first set of experiments, the adversarial attack algorithm used the YOLOv5s object detector to train adversarial patches. Given the YOLOv5s model gradients and network structure, the attack targets various object detectors. The results are shown in Table 2. The GADP algorithm achieved the best performance against the YOLOv5s detector, reducing its average precision (AP) for the "person" class to 14.8%. Faster R-CNN performed second, with an AP drop to 45.2%. The worst performance was achieved against the SSD detector, with an AP drop to 49.1%. Because the adversarial attack algorithm fully understands the YOLOv5s detector's network structure and parameters during training, it can quickly identify features that the detector is prone to misidentification based on the loss function. Different object detectors focus on the image region containing the "person" when extracting features from infrared images, sharing common features. The adversarial patch disrupts these common features, causing errors in the extraction of "person" features by various object detectors. The GADP algorithm developed in this chapter outperformed both advYOLOPatch and QRAttack, reducing the AP of infrared object detectors to a relatively low level. The adversarial examples generated using the GADP algorithm combined with the YOLOv5s object detector are transferable and can be used to attack YOLOv7, SSD, and Faster R-CNN. The GADP algorithm reduces the average precision of these three object detectors by up to 45.2%. It is worth noting that YOLOv7, SSD, and Faster R-CNN were not trained using the adversarial patches.
[0147] Table 2 Attack effects of adversarial attack algorithms combined with YOLOv5s target detector to generate adversarial samples
[0148]
[0149] The second group of experiments uses the YOLOv7 target detector to train the adversarial patch when the adversarial attack algorithm is trained, the model gradient and network structure of YOLOv7 are known during training, and the attack target is multiple target detectors, and the effect is shown in Table 3. When the GADP algorithm attacks the YOLOv7 detector, the effect is the best, making the average precision AP of the YOLOv7 detector for the "person" class drop to 11.5%. Second is the YOLOv5s detector, and the average precision drops to 35.8%. Finally, the average precision of the SSD detector drops to 38.2%. The network structure of the YOLO series detector is similar, especially the feature extraction and post-processing method. Under the constraint of the loss function, the original sample after adding the adversarial patch is migrated to the data space that is easy to misjudge by the detector, for example: the area near the model decision boundary. The adversarial samples generated by combining the GADP algorithm with the YOLOv7 target detector have migration, that is, they can be used to attack YOLOv5s, SSD and FasterR-CNN. The GADP algorithm makes the average precision of these three target detectors drop to at most 35.8%. It is worth mentioning that these three target detectors did not participate in the training process of the adversarial patch.
[0150] Table 3 Attack effect of adversarial attack algorithm combined with YOLOv7 target detector to generate adversarial samples
[0151]
[0152] The third group of experiments uses the SSD target detector to train the adversarial patch when the adversarial attack algorithm is trained, the model gradient and network structure of SSD are known, and the attack target is multiple target detectors, and the effect is shown in Table 4. When the GADP algorithm attacks the SSD detector, the effect is the best, making the average precision AP of the SSD detector for the "person" class drop to 13.8%. Second is the YOLOv5s detector, and the average precision drops to 44.9%. Finally, the average precision of the YOLOv7 detector drops to 47.3%. The adversarial samples generated by combining the GADP algorithm with the SSD target detector have migration, that is, they can be used to attack YOLOv5s, YOLOv7 and FasterR-CNN. The GADP algorithm makes the average precision of these three target detectors drop to at most 45.2%. Moreover, these three target detectors did not participate in the training process of the adversarial patch.
[0153] Table 4 Attack effect of adversarial attack algorithm combined with SSD target detector to generate adversarial samples
[0154]
[0155]
[0156] The adversarial attack algorithm of the fourth group of experiments uses a Faster R-CNN target detector when training the adversarial patch, the model gradient and network structure of the Faster R-CNN are known, the attack target is multiple target detectors, and the effect is shown in Table 5. The GADP algorithm of this chapter attacks the Faster R-CNN detector, and the effect is the best, making the average precision AP of the Faster R-CNN detector for the "person" category drop to 10.1%. The second is the SSD detector, and the average precision drops to 38.1%. The last is the YOLOv5s detector, and the average precision drops to 43.7%. The adversarial samples generated by the GADP algorithm combined with the Faster R-CNN target detector have migration, that is, they can be used to attack the YOLOv5s, YOLOv7 and SSD target detectors. The GADP algorithm makes the average precision of the three target detectors drop to a maximum of 38.1%. Moreover, the three target detectors do not participate in the training process of the adversarial patch.
[0157] Table 5 Attack effect of adversarial attack algorithm combined with Faster R-CNN target detector to generate adversarial samples
[0158]
[0159] The above four groups of experiments show that the white-box attack effect is better than the migration attack, because more detector network gradient and output information are known under the white-box attack condition. Moreover, the GADP algorithm constructed by the present application has migration, and after the generated adversarial patch is applied to the original image, it can make multiple infrared image target detectors difficult to normally recognize the "person" category.
[0160] Compared with advYOLOPatch and QRAttack, the migration attack effect of the GADP algorithm is better, and can reduce the average precision of common target detectors to about 40%. Due to the difference in detector network structure, different adversarial patches will be generated when using the GADP algorithm for training, resulting in differences in attack effect. In addition, the adversarial patch generated by the GADP algorithm also has good white-box attack effect, that is, it can reduce the average precision of the detector with known gradient and network structure to about 11%, making the target detector difficult to normally detect the "person" category. The adversarial patch generated by this algorithm applied to the original sample can significantly reduce the detection performance of multiple different detectors. Although different infrared image target detection algorithm models have different feature extractors, such as Mobilenet, Resnet, VGG, Darknet-53, etc., they all focus on the feature information of the target nearby area. The adversarial patch generated by the GADP algorithm proposed in the present application can interfere with the target nearby area features that these algorithm models commonly focus on, thereby causing these algorithm models to produce false prediction results.
[0161] The application constructs an adversarial patch generation algorithm model based on loss function optimization and Bernoulli random drop. The adversarial patch generated by using the algorithm has certain migration for a typical infrared image target detection model of multiple categories, so that multiple infrared target detection models are difficult to accurately identify 'people'. On the AFLIR data set, the attack success rate reaches 75.9%.
[0162] Although the embodiments of the present application have been shown and described above, it should be understood by those skilled in the art that the above embodiments are exemplary and cannot be construed as limiting the present application, and those skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application without departing from the principles and purposes of the present application.
Claims
1. A method for countering attacks on infrared image target detection algorithms based on adversarial patch generation, characterized by: The following steps are involved: Step 1: Get the infrared target detection model D and the original input sample image: x1,...,x N ; Step 2: Initialize the adversarial patch patch0 and iteratively optimize the adversarial patch until the set iteration termination condition is reached, and output the trained adversarial patch. For a certain iteration, the specific process is as follows: Step 2.1: Input each batch of original infrared images into model D to obtain the target box position and confidence; Step 2.2: Apply the adversarial patch of this iteration to the original input infrared image to obtain the adversarial sample x adv ; Step 2.3: Perform Bernoulli random dropout on some layers of the neural network of model D and perform local masking with adversarial patches to obtain the processed model D'; Step 2.4: Get the adversarial example x obtained in step 2.2 adv The target box confidence and location information in the model D' obtained in step 2.3; Step 2.5: Calculate the loss function value based on the target box position and confidence obtained in step 2.1 and the target box position and confidence obtained in step 2.4; Step 2.6: Update the adversarial patch using the Adam optimizer based on the loss function value obtained in step 2.5; then proceed to the next iteration; Step 3: Use the trained adversarial patches to conduct adversarial attacks on the infrared image target detection algorithm.
2. The method for countering attacks on infrared image target detection algorithms based on countermeasure patch generation according to claim 1, characterized in that: In step 2.3, the Bernoulli random dropout process is: During the forward propagation, the input Z l The stacked layer output F(Z l ) add to get the output Z L : Z L =Z l +(b+a-a·b)·F(Z l ) The parameter a is sampled from the continuous uniform distribution U(0,1), and b is sampled from the Bernoulli distribution; During the back propagation process, the loss function is defined as J, then: c is a Bernoulli variable, distributed in the same way as b.
3. The method for countering attacks on infrared image target detection algorithms based on countermeasure patch generation according to claim 1, characterized in that: In step 2.3, the local masking process of the adversarial patch is: For the width and height of W and H respectively, the normalized input image Patch x ,y∈[0,1], with probability ζ=0.8, perform the following operations: (1) Randomly sample within the adversarial patch and obtain a point p = (x0, y0); (2) Generate a λH×λW rectangular area with point p as the center on the adversarial patch, where λ is the scale and the filling value is k∈[0,1].
4. The method for countering attacks on infrared image target detection algorithms based on countermeasure patch generation according to claim 1, characterized in that: In step 2.5, the loss function is: l total =l det +αl s +βl ssim (x,x adv )+γl c Among them, α, β, and γ are hyperparameters; l det is the target hidden loss, l s is the smoothing loss, l ssim (x,x adv ) is the structural similarity loss, l c To counter the patch pixel value loss.
5. The method for countering attacks on infrared image target detection algorithms based on countermeasure patch generation according to claim 4, characterized in that: For a single-stage object detector, the object hiding loss is: Among them, n is the number of target boxes output by the target detector regarding the target category, is the confidence of the target box output by the target detector, Represents the true value of the target box confidence.
6. The method for countering attacks on infrared image target detection algorithms based on countermeasure patch generation according to claim 4, characterized in that: For the two-stage Faster R-CNN infrared target detection model, the target hiding loss is: The set category probability value corresponding to the i-th region proposal box is f i person (·), and there is f i person (·)=p i , x adv is the adversarial example after the adversarial patch is applied to the original infrared image.
7. The method for countering attacks on infrared image target detection algorithms based on countermeasure patch generation according to claim 4, characterized in that: The smoothing loss is 8. The method for countering attacks on infrared image target detection algorithms based on countermeasure patch generation according to claim 4, characterized in that: The structural similarity loss is Where n is the adversarial sample, m is the original infrared image sample, C is a constant, σ mn is the covariance of infrared images m and n, σ m , σ n are the standard deviations of infrared images m and n respectively.
9. The method for countering attacks on infrared image target detection algorithms based on countermeasure patch generation according to claim 4, characterized in that: The pixel value loss of the adversarial patch is p j is the pixel value corresponding to any pixel in the adversarial patch, S patch is the set of pixel values of the adversarial patch.