Black box adversarial patch generation method for aerial image multi-scale target detection
Through the differential evolution algorithm, we optimize the generation of adversarial patches in the low-dimensional decision space. Combined with the saliency loss and fitness function, we solve the problem of generating universal adversarial patches in black-box physical attacks and achieve efficient and robust multi-scale target detection attacks suitable for complex aerial remote sensing images.
Patent Information
- Application Number
- CN202510778344.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-26
AI Technical Summary
Existing black-box physical attacks on deep neural networks in multi-scale target detection in aerial remote sensing images face the problems of high cost, information distortion and poor robustness. Especially in complex remote sensing images and high-precision target models, the difficulty and cost of generating effective adversarial samples increase, and existing methods are difficult to generate universal adversarial patches in practical applications.
A differential evolution algorithm is used to generate adversarial patches in a low-dimensional decision space. Through initialization, mutation, crossover and selection operations, combined with significance loss and fitness function, high-dimensional candidate adversarial patches are optimized and generated. Physical robustness data enhancement is performed to ensure that digital images match real images and adapt to different detectors and categories.
It significantly improves the attack performance of adversarial patch attacks on black-box aerial images, improves the attack success rate, robustness and efficiency, is suitable for a variety of complex application scenarios, reduces computational costs and enhances the anti-detection ability of adversarial patches.
Smart Images

Figure CN120708037A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a black-box adversarial patch generation method for multi-scale target detection in aerial images. Background Art
[0002] In recent years, deep neural networks (DNNs) have made tremendous progress in aerospace imagery analysis. For example, change detection, which automatically identifies and analyzes changes in images, can be applied to urban expansion, deforestation, and other areas. Another example is vessel classification, which automatically identifies and classifies vessel types in images, which can be applied to cargo ships, fishing vessels, and other types. Furthermore, urban planning, which uses image data for urban planning, can be applied to road planning and land use planning. Deep neural networks (DNNs) leverage their powerful feature extraction and automated processing capabilities in aerospace imagery analysis to effectively improve analysis efficiency and accuracy.
[0003] Recent research has shown that existing deep neural networks (DNNs) are vulnerable to adversarial examples. Adversarial attacks pose a serious threat to the security of DNN models, potentially leading to misjudgments and incorrect decisions, resulting in potential safety hazards and economic losses. Importantly, existing research has paid less attention to physical attacks in target detection in drone aerial remote sensing imagery. Compared to digital attacks, physical attacks in the real world pose a greater security threat. However, most existing methods are designed for white-box settings, which is impractical in real-world scenarios. In practical applications, attackers often lack access to complete information about the target model, such as its internal structure, training parameters, and training methods. However, attackers can observe the outputs through arbitrary inputs and, through the correspondence between inputs and outputs, find adversarial examples to attack the victim model. Black-box attacks are more practical in this context, simulating the limited knowledge of the model that attackers possess in the real world and providing a better way to assess the security of the model in unknown environments. White-box attacks, on the other hand, require access to the model's internal information and gradients, potentially being detected by defense mechanisms.
[0004] In the field of aerial remote sensing, the construction of a physical universal black-box adversarial patch attack for multi-scale target detection faces three major challenges. First, black-box attacks typically require sending a large number of query requests to the target model to obtain the model's output information and guide the generation of updated adversarial samples. Especially when faced with complex remote sensing images and high-precision target models, more queries are required to generate effective adversarial samples, which further increases the difficulty and cost of the attack. Second, unlike digital image adversarial attacks, adversarial patch attacks involve projecting carefully designed perturbations into real-world scenes through physical operations. During the digital-to-physical conversion process, adversarial patches inevitably experience information distortion, and their adversarial patches are easily detected by the human eye or patch detectors. Third, due to the differences between the adversarial patches generated on a specified victim model and other detector models in terms of structure, parameters, training methods, etc., the generation of universal adversarial patches is very challenging in terms of robustness. Summary of the Invention
[0005] The purpose of this invention is to provide a black-box adversarial patch generation method for multi-scale target detection in aerial images, which can significantly improve the attack performance of adversarial patch attacks on black-box aerial images, effectively improve the attack success rate, attack efficiency and robustness, and is suitable for a variety of complex application scenarios.
[0006] The present invention adopts the following technical solutions:
[0007] A black-box adversarial patch generation method for multi-scale object detection in aerial images, comprising the following steps:
[0008] S1: Preprocess the acquired aerial images and divide them into training and test sets;
[0009] S2: Generate random low-dimensional candidate adversarial patches in the low-dimensional decision space using the low-dimensional decision space model, and provide the original low-dimensional candidate adversarial patches for upgrading to high-dimensional candidate adversarial patches through initialization, mutation, crossover, and selection operations;
[0010] S3: In the high-dimensional patch space, the original low-dimensional candidate adversarial patch and the new low-dimensional candidate adversarial patch obtained after the cross operation are respectively upgraded to obtain the corresponding high-dimensional candidate adversarial patch;
[0011] S4: In the high-dimensional patch space, the generated high-dimensional candidate adversarial patches are subjected to physical robustness data augmentation and applied to image generation adversarial samples;
[0012] S5: Input the adversarial sample and the corresponding saliency loss value into the object detector, and use the fitness function algorithm to evaluate the detection results of the object detector;
[0013] S6: Repeat steps S2 to S5 to obtain a high-dimensional candidate adversarial patch that can completely attack all images or a high-dimensional candidate adversarial patch with the highest attack success rate and the lowest saliency value as the final high-dimensional adversarial patch;
[0014] S7, applies the final high-dimensional adversarial patch on top of the attack target category and migrates it to different target detectors for physical adversarial attacks.
[0015] The low-dimensional decision space model includes a population initialization module, a mutation module, a crossover module, and a selection module;
[0016] A population initialization module, used to generate a random number of low-dimensional candidate adversarial patches;
[0017] A mutation module is used to perform a mutation operation on the low-dimensional candidate adversarial patch to obtain a mutated low-dimensional candidate adversarial patch;
[0018] The crossover module is used to generate new low-dimensional candidate adversarial patches by crossover operation between the mutated low-dimensional candidate adversarial patches in the population and the low-dimensional candidate adversarial patches in the original population;
[0019] The selection module is used to screen high-dimensional candidate adversarial patches with high attack success rate in the population, which are formed by dimensionality upgrading of low-dimensional candidate adversarial patches, and then obtain the low-dimensional candidate adversarial patches corresponding to the high-dimensional candidate adversarial patches, which are used as the original population individuals for the next mutation, crossover and selection operations, namely the low-dimensional candidate adversarial patches.
[0020] Step S2 includes the following steps:
[0021] S21: Initialize low-dimensional candidate adversarial patches from a random population in a low-dimensional decision space;
[0022] S22: performing a mutation operation on the low-dimensional candidate adversarial patch in the initialized random population to obtain a mutated low-dimensional candidate adversarial patch;
[0023] S23: performing a cross operation on the mutated low-dimensional candidate adversarial patch and the original low-dimensional candidate adversarial patch to obtain a new low-dimensional candidate adversarial patch;
[0024] S24: For the original low-dimensional candidate adversarial patches and the new low-dimensional candidate adversarial patches after crossing in the low-dimensional decision space, the next generation of initial low-dimensional candidate adversarial patches is selected through the fitness function, and high-dimensional candidate adversarial patches and their indexes with high attack success rate obtained from the population by upgrading the low-dimensional candidate adversarial patches.
[0025] In step S22, a difference vector is first calculated for two random low-dimensional candidate adversarial patches, and then the structure of the difference vector calculation is multiplied by a scaling factor and added to the remaining adversarial patch to generate a new mutated low-dimensional candidate adversarial patch.
[0026] In step S3, the spatial color saliency digital enhancement transformation is first performed on the high-dimensional candidate adversarial patch through quantized color measurement; then the high-dimensional adversarial patch after the enhancement transformation is added to the attacked image through the MASK matrix.
[0027] The high-dimensional adversarial patch after the enhancement transformation is added only to the image where the detector can detect the object in the original image through the MASK matrix.
[0028] In step S5, the high-dimensional candidate adversarial patch formed by dimensional upgrading the original low-dimensional candidate adversarial patch and the new low-dimensional candidate adversarial patch after crossing is applied to the training set image, and the saliency loss value is marked in combination with the saliency loss calculation to generate an adversarial sample; then, the high-dimensional candidate adversarial patch whose fitness function value is lower than the confidence threshold is selected from the adversarial sample through the fitness function, and the low-dimensional candidate adversarial patch corresponding to the high-dimensional candidate adversarial patch is used as the next generation initial low-dimensional candidate adversarial patch.
[0029] In step S5, for each individual in the population, the fitness function evaluates the effect of the parent and child individuals in reducing the confidence of target detection by calculating the difference between the confidence scores of the parent and child individuals and the confidence threshold; the fitness function also compares the number of targets that are still detected by the parent and child individuals after being detected by the target detector, and selects individuals with shorter detection lengths to be retained in the next generation; at the same time, the fitness function also monitors in real time whether there is an individual that can make the confidence scores of all target objects lower than the confidence threshold. If such an individual exists, the index of the individual is immediately returned and the optimization process is terminated.
[0030] In step S23, according to the aerial photography height and magnification factor of the image, the mutated low-dimensional candidate adversarial patches in the population are dynamically combined with the original low-dimensional candidate adversarial patches to generate high-dimensional candidate adversarial patches of corresponding sizes.
[0031] Step S1 includes the following steps:
[0032] S11: Acquire aerial images and build an aerial image dataset;
[0033] S12: preprocess each aerial image in the aerial image dataset;
[0034] S13: Use the preprocessed aerial images to construct training and test sets.
[0035] This paper uses a specially designed differential evolution algorithm that includes mutation, crossover, and selection operations, combined with an optimized dimensionality reduction strategy and fitness function, to implement adversarial patch attacks on aerial image object detectors in a black-box setting, with the following significant beneficial effects:
[0036] First, the present invention can optimize the adversarial patch in a black-box setting by initializing the population, performing mutation, crossover, and selecting individuals, achieving a balance between global and local search. The algorithm can explore new solutions globally while utilizing the excellent features of parent individuals for local optimization, thereby obtaining significant results. Since the optimization process is performed in a low-dimensional space, the proposed method can effectively find the optimal solution within a limited computational budget, reducing computational costs and improving attack efficiency.
[0037] Secondly, the dimensionality-increasing method proposed in this invention innovatively realizes the adaptive generation of adversarial patches of corresponding sizes through scale factors according to aerial images of different scales, solving the problem of inconsistent patch sizes when shooting at different altitudes.
[0038] In addition, the present invention also takes into account the problem that digital world adversarial patches and real-world adversarial patches will cause distortion and inconsistent visual effects after printing, and performs saliency feature processing on the generated adversarial patches to ensure the match between the digital image and the actual printing effect, and makes the adversarial patches less colorful, thereby achieving the possibility of being detected to a large extent.
[0039] At the same time, through the application of fitness function, the present invention can assist the DE algorithm in selecting better individuals within a limited computational budget, and effectively conduct adversarial patch attacks on different target detectors and categories.
[0040] Finally, based on comprehensive performance evaluation and optimization, the present invention ensures the success and efficiency of the attack effect, and provides reliable technical support for the analysis and application of UAV aerial image counterattacks.
[0041] In summary, the present invention can significantly improve the attack performance of anti-patch attacks on black-box aerial images, has a higher attack success rate, stronger robustness and better attack efficiency, is suitable for a variety of complex application scenarios, and has broad promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the process of the present invention;
[0043] Figure 2 Schematic diagram of the framework of the low-dimensional decision space in the present invention;
[0044] Figure 3 Schematic diagram of the patch space framework in the present invention;
[0045] Figure 4 Schematic diagram of the aerial image in the present invention. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solutions and advantages of the embodiments of this document clearer, the technical solutions in the embodiments of this document will be clearly and completely described below in conjunction with the drawings in the embodiments of this document. Obviously, the described embodiments are part of the embodiments of this document, not all of the embodiments. Based on the embodiments of this document, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this document. It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of this document can be combined with each other in any way.
[0047] The present invention is described in detail below with reference to the accompanying drawings and embodiments:
[0048] like Figure 1 As shown, the black box adversarial patch generation method for multi-scale target detection in aerial images according to the present invention comprises the following steps in sequence:
[0049] S1: Preprocess the acquired aerial images and divide them into training and test sets;
[0050] In the present invention, since the images in the acquired data set have different sizes, in order to facilitate the experiment and ensure the accuracy of the subsequent model output results, the data set is cropped and classified.
[0051] Step S1 includes the following specific steps:
[0052] S11: Acquire aerial images and build an aerial image dataset;
[0053] In order to generate adversarial patches that can avoid detection of specific categories by well-trained aerial image detectors, in this embodiment, aerial images are obtained from the existing DOTA-v1.0 dataset and an aerial image dataset is constructed. The DOTA-v1.0 dataset is a large-scale dataset containing 2,806 aerial images. By acquiring 15 common categories from different sensors and platforms, including airplanes, ships, large vehicles, small vehicles, storage tanks, baseball fields, basketball courts, etc., the DOTA-v1.0 dataset can effectively support adversarial attack tasks in various scenarios, and the diversity and complexity of the dataset provides rich scene information for model training, which is particularly suitable for the study of physical general black box adversarial patch generation methods for multi-scale target detection.
[0054] S12: Preprocess each aerial image in the aerial image dataset.
[0055] Since the sizes of the original aerial images in DOTA vary greatly, the present invention unifies the size of each aerial image in the aerial image dataset by image cropping, and crops them into aerial images with a size of 608×608.
[0056] S13: Use the pre-processed aerial images to construct training and test sets;
[0057] In this example, 39,905 preprocessed aerial images were selected to form the training set, and 13,603 preprocessed aerial images were selected to form the test set. For each image category (such as aircraft, ship, large vehicle, small vehicle, etc.), 100 images were randomly selected from the test set for the black-box attack.
[0058] S2: Generate random low-dimensional candidate adversarial patches in the low-dimensional decision space using the low-dimensional decision space model, and provide the original low-dimensional candidate adversarial patches for upgrading to high-dimensional candidate adversarial patches through initialization, mutation, crossover, and selection operations;
[0059] The low-dimensional decision space model includes a population initialization module, a mutation module, a crossover module and a selection module;
[0060] A population initialization module, used to generate a random number of low-dimensional candidate adversarial patches;
[0061] The mutation module is used to perform mutation operations on low-dimensional candidate adversarial patches to introduce new mutated low-dimensional candidate adversarial patches and increase the diversity of the population;
[0062] The crossover module is used to generate new low-dimensional candidate adversarial patches by crossover operations between the mutated low-dimensional candidate adversarial patches in the population and the low-dimensional candidate adversarial patches in the original population, so as to promote information exchange between individuals in the population;
[0063] The selection module is used to screen high-dimensional candidate adversarial patches with high attack success rates in the population, which are formed by upgrading the low-dimensional candidate adversarial patches. This module then obtains the low-dimensional candidate adversarial patches corresponding to the high-dimensional candidate adversarial patches with high attack success rates, thereby providing the original population individuals, i.e., the low-dimensional candidate adversarial patches, for the next mutation, crossover, and selection operations.
[0064] In the present invention, random low-dimensional candidate adversarial patches are generated in the low-dimensional decision space, which mainly includes four steps: initialization of population, mutation, crossover, and selection. The flowchart is shown in the following figure. Figure 2 As shown. Aiming at the frequent query transmission in black-box adversarial patches, the present invention specifically designs the above-mentioned method for generating low-dimensional candidate adversarial patches in a low-dimensional decision space. Using black-box attacks as the primary attack method, combined with a differential evolution algorithm, this method achieves efficient generation and precise selection of high-dimensional adversarial patches, thereby addressing challenges such as high resource consumption, individual globality, locality, and integrity in generating high-dimensional adversarial patches.
[0065] The step S2 comprises the following steps:
[0066] S21: Initialize low-dimensional candidate adversarial patches from a random population in a low-dimensional decision space;
[0067] In the process of randomly generating the initial population in the low-dimensional decision space, in order to ensure the efficiency of subsequent operations, this embodiment selects 5 random low-dimensional candidate adversarial patches from the population. The decision variable size of each random low-dimensional candidate adversarial patch is set to 6×6×3. The decision variable size is much smaller than the size of the high-dimensional candidate adversarial patch, which prepares for the subsequent dimensional upgrade to high-dimensional candidate adversarial patches. Among them, 6×6 represents the number of pixels in the horizontal and vertical directions of the image, that is, the width and height of the image, 3 represents the three color channels of the image, and the population P t It can be expressed as:
[0068]
[0069] In formula (1), P t (i) represents the i-th low-dimensional candidate adversarial patch in the t-th evolution, and all individuals in each iterative optimization are from the interval The superscripts L and U are the upper and lower bounds of the variable; N is the population size, which can be set to 5 in this embodiment; T represents the maximum allowed number of evolutions, which can be set to 200 in this embodiment.
[0070] In this embodiment, and Can be set to 0 and 1 respectively to optimize and overwrite selected pixels in the destination image.
[0071] S22: perform mutation operation on the low-dimensional candidate adversarial patches in the initialized random population;
[0072] By mutating low-dimensional candidate adversarial patches, the present invention leverages random perturbations to increase population diversity, thereby expanding the method's search space and improving global search capabilities, facilitating the rapid identification of superior candidate adversarial patches. In subsequent adversarial patch attacks, the mutation employed by the present invention makes the generated low-dimensional candidate adversarial patches more effective, significantly reducing the confidence level of the detected target and increasing the success rate of the attack.
[0073] In this embodiment, each mutated low-dimensional candidate adversarial patch in the random population is obtained from three randomly different low-dimensional candidate adversarial patches. In this embodiment, the specific steps of the mutation operation are:
[0074] First, the difference vector of two random low-dimensional candidate adversarial patches is calculated. Then, the structure of the difference vector calculation is multiplied by the scaling factor f and added to the remaining adversarial patch (i.e., the third low-dimensional candidate adversarial patch) to generate a new mutated low-dimensional candidate adversarial patch. The mutation operation expression is:
[0075] P M (i) = P t (r1)+F·(P t (r2)-P t (r3)) (2)
[0076] r1≠r2≠r3≠i
[0077] In formula (2), P M (i) represents the low-dimensional candidate adversarial patch after mutation, the subscript M is the first letter of the English word Mutation, P t (r1) represents the remaining adversarial patch, F represents the scaling factor, and P t (r2) and P t (r3) represents the two candidate adversarial patches for difference vector calculation, r1, r2 and r3 are indices randomly selected from N candidate adversarial patches without replacement;
[0078] In this embodiment, the scaling factor F may be set to 0.5.
[0079] S23: Perform a cross operation on the mutated low-dimensional candidate adversarial patch and the original low-dimensional candidate adversarial patch;
[0080] In this paper, a crossover operation is used to generate new low-dimensional candidate adversarial patches, enabling the population to gradually converge to a more optimal solution. By properly setting the crossover probability, the intensity of the crossover operation can be controlled, thereby gradually improving the overall quality of the population while maintaining diversity. In subsequent adversarial patch attacks, this crossover operation helps to more quickly find patches that effectively reduce the confidence level of target detection.
[0081] In this embodiment, the specific steps of the crossover operation are:
[0082] After the individual mutation of the population, the original low-dimensional candidate adversarial patch in the candidate population and the mutated low-dimensional candidate adversarial patch are cross-operated to generate a new low-dimensional candidate adversarial patch after the crossover. The crossover operation expression is:
[0083]
[0084] In formula (3), P C (i) represents the new low-dimensional candidate adversarial patch after crossover, and the subscript C is the first letter of Crossover in English. r represents the crossover probability, rand refers to a random number in the interval [0,1], C r The larger the value of , the more characteristics the parent population has. Similarly, C r The smaller the value of , the more characteristics the offspring population has. Crossover is the key to generating offspring, which can be understood as the parent population according to Cr Proportional recombination. Crossover simulates genetic recombination in biological evolution to explore new candidate adversarial patch structures.
[0085] In this embodiment, the crossover probability C r Can be set to 0.4.
[0086] S24: The original low-dimensional candidate adversarial patches and the new low-dimensional candidate adversarial patches after the crossover in the low-dimensional decision space are selected through the fitness function to select the next generation of initial low-dimensional candidate adversarial patches, and the high-dimensional candidate adversarial patches with high attack success rate are screened from the population, which are obtained by upgrading the low-dimensional candidate adversarial patches. Then, the low-dimensional candidate adversarial patches and their indexes corresponding to the high-dimensional candidate adversarial patches are selected.
[0087] In the present invention, the above-mentioned selection operation is used to select high-dimensional candidate adversarial patches and their indexes with high attack success rate from the population, which are obtained by upgrading the dimensions of low-dimensional candidate adversarial patches, thereby screening out low-dimensional candidate adversarial patches with high attack effect for updating the population quality, so that high-quality individuals in the population can be retained and propagated, thereby accelerating the convergence process of the algorithm.
[0088] In this invention, by using the fitness function evaluation value in the high-dimensional patch space, low-dimensional candidate adversarial patches corresponding to high-fitness candidate adversarial patches can be selected and placed into the low-dimensional decision space population. Since the lower the fitness function evaluation value, the more effective the high-dimensional candidate adversarial patch is in attacking the target detector, and the higher the corresponding fitness, the individuals with high fitness are more likely to be selected in subsequent update optimization mutation and crossover operations, thereby further optimizing the population. For example, in an attack detection scenario, adversarial patches that can successfully evade detection can be selected.
[0089] In the present invention, the operation of obtaining the best candidate adversarial patch from the population can be expressed as:
[0090]
[0091] In formula (4), P tt1 (i) represents the offspring low-dimensional candidate adversarial patch, i.e., the next generation initial low-dimensional candidate adversarial patch, which is used to provide the initial population individuals for the next update, and f() represents the fitness function, which is used to evaluate the quality of each individual in step S5. This embodiment is based on the screening of the number of surviving instances, by counting the number of instances with confidence higher than the threshold (i.e., the "detection length" length (S)). i (i)))Prefer individuals that reduce the detection length.
[0092] S3: In the high-dimensional patch space, the original low-dimensional candidate adversarial patch and the new low-dimensional candidate adversarial patch after intersection are respectively upgraded to obtain the corresponding high-dimensional candidate adversarial patch;
[0093] In the present invention, the process of upgrading the low-dimensional candidate adversarial patches is completed in the high-dimensional patch space. The schematic diagram of the patch space structure is shown in Figure 3 shown.
[0094] After performing a cross operation on the mutated low-dimensional candidate adversarial patch in the population and the original low-dimensional candidate adversarial patch, the magnification factor ε is introduced according to the aerial height h of the image. h , the mutated low-dimensional candidate adversarial patches in the population are dynamically generated with the original low-dimensional candidate adversarial patches (collectively referred to as low-dimensional candidate adversarial patches) to generate high-dimensional candidate adversarial patches of corresponding sizes, thereby reducing the optimization space dimension.
[0095] In this embodiment, Figure 4 As shown, take the aerial image of the airport at an altitude of 200m as an example: according to the magnification factor ε h , the low-dimensional candidate adversarial patches in the population are enlarged into high-dimensional candidate adversarial patches with a length and width of 5 times the size. The specific steps of dimensionality increase are: by copying and randomly rotating the low-dimensional candidate adversarial patches, and then tiling the low-dimensional candidate adversarial patches (6×6×3) to high-dimensional candidate adversarial patches (30×30×3) in the high-dimensional patch space, thereby covering the attack target area of the image. The dimensionality increase operation reduces the optimization space dimension (from 900 dimensions to 36 dimensions in this embodiment), avoids the high-dimensional candidate adversarial patches from falling into local optimality, and greatly reduces the optimization cost, to a certain extent enhances the robustness of the high-dimensional candidate adversarial patches to other detections, and the generated high-dimensional candidate adversarial patches have repeated adversarial textures, which greatly enhances the attack efficiency and effectiveness.
[0096] S4: In the high-dimensional patch space, the generated high-dimensional candidate adversarial patches are subjected to physical robustness data augmentation and applied to image generation adversarial samples.
[0097] In the invention, a significant digital enhancement transformation is first performed on the high-dimensional candidate adversarial patches, which is used to solve the distortion problem in the conversion process from digital images to physical images, and the robustness problem of the adversarial patches on other detector models is solved by the noise processing method.
[0098] In this embodiment, first, an sRGB color saliency loss is used to bias the adversarial patch toward less vibrant or saturated colors, using a quantitative color metric derived from psycho-physical color scaling research. Compared to saturation, the color metric reflects more of the color information of the adversarial patch, overemphasizing the dark areas of the adversarial patch, making it less noticeable to humans or automated patch detection systems. This also significantly reduces the inconsistency between the digital image and the physical image when the adversarial patch is subsequently printed out.
[0099] The expressions of color metric and saliency loss are:
[0100] rg = RG;
[0101] yb=0.5*(R+G)-B;
[0102]
[0103] In formula (5), rg and yb are variables derived from the color channel R, G and B values of the color block, Represents the significant loss, μ represents the mean, which is used to measure the average value of the image brightness or color channel, σ represents the standard deviation, which is used to measure the dispersion of the image brightness or color channel. and Represents the square of the standard deviation of the red and green and yellow and blue color channels respectively, and Represents the square of the mean of the red and green and yellow and blue color channels respectively. The final significance loss value is used for subsequent fitness function evaluation.
[0104] Then, the high-dimensional adversarial patch after saliency transformation is added to the attacked image through the MASK matrix, which determines the shape and position of the added adversarial patch.
[0105] Among them, the high-dimensional candidate adversarial patches can be further applied to the images in which the detector can detect objects in the original image through the MASK matrix, and the images of objects in which the detector cannot detect in the original image are not subjected to the application of high-dimensional candidate adversarial patches, so as to further ensure the rigor of the present invention.
[0106] In this embodiment, the detector used may be a YOLOv3 detector, and the adversarial sample expression is expressed as:
[0107]
[0108] In formula (6), x ′ is the generated adversarial sample, M is the position mask used to mark the placement of the patch, ⊙ represents element-wise multiplication, x0 is the original image, ε is the adversarial noise, and st represents the constraint condition, requiring the generated adversarial sample x ′ This will be different from the original image x0, leading to incorrect predictions by the YOLOv3 detector model. Represents the victim model, whose weight is fixed during the attack process.
[0109] S5: Input the generated adversarial examples and the corresponding saliency loss values into the target detector, and use the fitness function algorithm to evaluate the detection results of the target detector.
[0110] In step S5, the high-dimensional candidate adversarial patch formed by upgrading the original low-dimensional candidate adversarial patch and the new low-dimensional candidate adversarial patch after crossing is applied to the training set image, and its significance loss value is marked in combination with the significance loss calculation to generate an adversarial sample; then, the high-dimensional candidate adversarial patch whose fitness function value is lower than the confidence threshold is selected from the adversarial sample through the fitness function, and the low-dimensional candidate adversarial patch corresponding to the high-dimensional candidate adversarial patch is used as the next generation initial low-dimensional candidate adversarial patch.
[0111] In this embodiment, the target detector may use Yolov3.
[0112] The adversarial sample with the high-dimensional candidate adversarial patch and the corresponding saliency loss value are input into the object detector to obtain the output detection result. The detector must simultaneously output the objectness score, bounding box, and category probability of all objects in the scene. In order to implement the vanishing attack, the attacker needs to reduce the detection confidence score of all instances to below the predefined confidence threshold as much as possible after applying the high-dimensional candidate adversarial patch. Therefore, formula (6) can be further expressed as:
[0113]
[0114] In formula (7), is the confidence score of the i-th target instance, S0 is a predefined confidence threshold; st represents the constraint condition, requiring that the confidence scores of all target instances are less than S0, and n is the total number of instances in the image.
[0115] In this embodiment, the fitness function is further optimized, and the key logic of the specific algorithm is as follows:
[0116] a: For each individual in the population, the difference between the confidence score of the parent and child individuals of the high-dimensional candidate adversarial patch and the confidence threshold is calculated to evaluate the effectiveness of the parent and child high-dimensional candidate adversarial patches in reducing the confidence of target detection; wherein, the parent and child individuals of the high-dimensional candidate adversarial patch are formed by increasing the dimensionality of the corresponding low-dimensional candidate adversarial patch parent and child individuals;
[0117] b: Compare the number of targets still detected by the parent and offspring individuals after being detected by the YOLOv3 detector (i.e., detection length), and select the individuals with the shorter detection length to be retained in the next generation. A shorter detection length means that the individual more effectively reduces the confidence of the target, that is, the attack effect is better, making the target more difficult to detect;
[0118] At the same time, real-time monitoring is performed to determine whether there is an individual that can make the confidence scores of all target objects lower than the confidence threshold (i.e., the detection length is 0). If such an individual exists, the index of the individual is immediately returned and the optimization process is terminated; thereby achieving the effective generation of the optimal resistance patch for the final effect and stopping in time when the attack is successful, avoiding unnecessary calculations and improving attack efficiency.
[0119] In step b, the detection lengths of the parent and offspring individuals are compared to select the individuals that can more effectively reduce the target detection confidence score. The corresponding low-dimensional candidate adversarial patches are retained to the next generation, thereby gradually optimizing the low-dimensional candidate adversarial patches in the population and improving the attack success rate. The optimization is stopped immediately after finding the individual that has successfully attacked to save computing resources.
[0120] S6: Repeat steps S2 to S5, and use the training set to optimize and update the low-dimensional candidate adversarial patches until the maximum number of iterations is reached or the specified attack effect is achieved. Finally, a high-dimensional candidate adversarial patch that can completely attack all images and is upgraded from the corresponding low-dimensional candidate adversarial patch is obtained, or a high-dimensional candidate adversarial patch with the highest attack success rate and the lowest significance value is obtained as the final high-dimensional adversarial patch, and applied to the test set images to verify the effectiveness and versatility of its generated high-dimensional adversarial patches.
[0121] In the present invention, by setting a low-dimensional decision space and a high-dimensional patch space respectively, the low-dimensional candidate adversarial patches in the population are iteratively updated in the low-dimensional decision space, and through mutation, crossover, and selection operations, and fitness evaluation in the high-dimensional patch space, low-dimensional candidate adversarial patch options are provided for the low-dimensional decision space selection operation, and then iterative updates are performed until a high-dimensional candidate adversarial patch that can completely attack individuals in all images appears or the number of updates reaches a threshold. If the high-dimensional candidate adversarial patch that has not completely attacked all instances reaches the threshold number of updates, we select the high-dimensional candidate adversarial patch with the highest attack success rate and the lowest significance value as the final high-dimensional adversarial patch, so as to provide the best effect for printing out the high-dimensional adversarial patch and attacking in the real world. The final high-dimensional adversarial patch is applied to the test set image to verify the effectiveness of the generated adversarial patch.
[0122] In S7, the final high-dimensional adversarial patch is applied on top of the attack target category and migrated to different target detectors for physical adversarial attacks.
[0123] In the present invention, the obtained final high-dimensional adversarial patch can be printed out, applied on top of the attack target category, and migrated to different target detectors for physical adversarial attacks.
[0124] In this example, the target category is an airplane. We print the generated high-dimensional adversarial patch in equal proportions and post it on the top of the airplane. We then use other detectors, including the one-stage detection model Yolov5 and the two-stage detection model Faster-RCNN, to test the robustness of the adversarial patch, achieve transferability, and solve the universality problem of the adversarial patch.
[0125] The black-box adversarial patch generation method for multi-scale target detection in aerial images described in the present invention addresses the problems of high resource consumption, printing distortion, and anti-robustness of adversarial patches in general adversarial attacks on black-box physics of aerial images. Through the innovative design of differential evolution algorithm, which includes mutation, crossover, and selection operations, a new dimensionality reduction strategy and fitness function are proposed, thereby realizing adversarial patch attacks on aerial image target detectors under black-box settings, and having the following significant beneficial effects: First, in order to better improve the success rate and efficiency of black-box attacks, the present invention adopts an advanced differential evolution algorithm, and optimizes low-dimensional candidate adversarial patches under black-box settings by initializing the population in a low-dimensional decision space, performing mutation, crossover, and selecting low-dimensional candidate adversarial patches, thereby achieving a balance between global search and local search. The algorithm can explore new solutions globally, while utilizing the excellent features of parent individuals for local optimization, thereby obtaining significant results. Since the process is optimized in a low-dimensional decision space, the proposed method can effectively find the optimal solution within a limited computational budget, reducing computational costs and improving attack efficiency.
[0126] The present invention innovatively achieves the adaptive generation of high-dimensional candidate adversarial patches of corresponding sizes based on aerial images of different scales through a scale factor, solving the problem of inconsistent patch sizes when shooting at different altitudes. Furthermore, the present invention addresses the problem of distortion and inconsistent visual effects caused by digital and real-world adversarial patches after printing. The high-dimensional candidate adversarial patches after dimensionality upgrade are processed for saliency features, ensuring a match between the digital image and the actual printout. The high-dimensional candidate adversarial patches are also rendered less vivid, significantly reducing their detectability and further enhancing their robustness. Furthermore, the present invention designs a universal fitness function and combines evaluation metrics such as attack success rate (ASR), average precision (AP), and number of queries to fully implement the invention's selection of better low-dimensional candidate adversarial patches within a limited computational budget. Furthermore, the invention effectively attacks adversarial patches against different target detectors and categories. Finally, based on comprehensive performance evaluation and optimization, the present invention ensures the success and efficiency of the attack, providing reliable technical support for the analysis and application of adversarial attacks against drone aerial images. In summary, the present invention can significantly improve the attack performance of adversarial patch attacks on black-box aerial images, with a higher attack success rate, stronger robustness, and better attack efficiency. It provides an efficient and reliable solution for a physical universal black-box adversarial patch generation method for multi-scale target detection in aerial images.
[0127] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A black-box adversarial patch generation method for multi-scale object detection in aerial images, characterized by: The following steps are involved: S1: Preprocess the acquired aerial images and divide them into training and test sets; S2: Generate random low-dimensional candidate adversarial patches in the low-dimensional decision space using the low-dimensional decision space model, and provide the original low-dimensional candidate adversarial patches for upgrading to high-dimensional candidate adversarial patches through initialization, mutation, crossover, and selection operations; S3: In the high-dimensional patch space, the original low-dimensional candidate adversarial patch and the new low-dimensional candidate adversarial patch obtained after the cross operation are respectively upgraded to obtain the corresponding high-dimensional candidate adversarial patch; S4: In the high-dimensional patch space, the generated high-dimensional candidate adversarial patches are subjected to physical robustness data augmentation and applied to image generation adversarial samples; S5: Input the adversarial sample and the corresponding saliency loss value into the object detector, and use the fitness function algorithm to evaluate the detection results of the object detector; S6: Repeat steps S2 to S5 to obtain a high-dimensional candidate adversarial patch that can completely attack all images or a high-dimensional candidate adversarial patch with the highest attack success rate and the lowest saliency value as the final high-dimensional adversarial patch; S7, applies the final high-dimensional adversarial patch on top of the attack target category and migrates it to different target detectors for physical adversarial attacks.
2. The black-box adversarial patch generation method according to claim 1, characterized in that: The low-dimensional decision space model includes a population initialization module, a mutation module, a crossover module, and a selection module; A population initialization module, used to generate a random number of low-dimensional candidate adversarial patches; A mutation module is used to perform a mutation operation on the low-dimensional candidate adversarial patch to obtain a mutated low-dimensional candidate adversarial patch; The crossover module is used to generate new low-dimensional candidate adversarial patches by crossover operation between the mutated low-dimensional candidate adversarial patches in the population and the low-dimensional candidate adversarial patches in the original population; The selection module is used to screen high-dimensional candidate adversarial patches with high attack success rate in the population, which are formed by dimensionality upgrading of low-dimensional candidate adversarial patches, and then obtain the low-dimensional candidate adversarial patches corresponding to the high-dimensional candidate adversarial patches, which are used as the original population individuals for the next mutation, crossover and selection operations, namely the low-dimensional candidate adversarial patches.
3. The black box adversarial patch generation method according to claim 1, characterized in that: Step S2 includes the following steps: S21: Initialize low-dimensional candidate adversarial patches from a random population in a low-dimensional decision space; S22: performing a mutation operation on the low-dimensional candidate adversarial patch in the initialized random population to obtain a mutated low-dimensional candidate adversarial patch; S23: performing a cross operation on the mutated low-dimensional candidate adversarial patch and the original low-dimensional candidate adversarial patch to obtain a new low-dimensional candidate adversarial patch; S24: For the original low-dimensional candidate adversarial patches and the new low-dimensional candidate adversarial patches after crossing in the low-dimensional decision space, the next generation of initial low-dimensional candidate adversarial patches is selected through the fitness function, and high-dimensional candidate adversarial patches and their indexes with high attack success rate obtained from the population by upgrading the low-dimensional candidate adversarial patches.
4. The black-box adversarial patch generation method according to claim 3, characterized in that: In step S22, a difference vector is first calculated for two random low-dimensional candidate adversarial patches, and then the structure of the difference vector calculation is multiplied by a scaling factor and added to the remaining adversarial patch to generate a new mutated low-dimensional candidate adversarial patch.
5. The black box adversarial patch generation method according to claim 1, characterized in that: In step S3, the spatial color saliency digital enhancement transformation is first performed on the high-dimensional candidate adversarial patch through quantized color measurement; then the high-dimensional adversarial patch after the enhancement transformation is added to the attacked image through the MASK matrix.
6. The black box adversarial patch generation method according to claim 5, characterized in that: The high-dimensional adversarial patch after the enhancement transformation is added only to the image where the detector can detect the object in the original image through the MASK matrix.
7. The black box adversarial patch generation method according to claim 1, characterized in that: In step S5, the high-dimensional candidate adversarial patch formed by dimensional upgrading the original low-dimensional candidate adversarial patch and the new low-dimensional candidate adversarial patch after crossing is applied to the training set image, and the saliency loss value is marked in combination with the saliency loss calculation to generate an adversarial sample; then, the high-dimensional candidate adversarial patch whose fitness function value is lower than the confidence threshold is selected from the adversarial sample through the fitness function, and the low-dimensional candidate adversarial patch corresponding to the high-dimensional candidate adversarial patch is used as the next generation initial low-dimensional candidate adversarial patch.
8. The black box adversarial patch generation method according to claim 1, characterized in that: In step S5, for each individual in the population, the fitness function evaluates the effect of the parent and child individuals in reducing the confidence of target detection by calculating the difference between the confidence scores of the parent and child individuals and the confidence threshold; the fitness function also compares the number of targets that are still detected by the parent and child individuals after being detected by the target detector, and selects individuals with shorter detection lengths to be retained in the next generation; at the same time, the fitness function also monitors in real time whether there is an individual that can make the confidence scores of all target objects lower than the confidence threshold. If such an individual exists, the index of the individual is immediately returned and the optimization process is terminated.
9. The black box adversarial patch generation method according to claim 1, characterized in that: In step S23, according to the aerial photography height and magnification factor of the image, the mutated low-dimensional candidate adversarial patches in the population are dynamically combined with the original low-dimensional candidate adversarial patches to generate high-dimensional candidate adversarial patches of corresponding sizes.
10. The black box adversarial patch generation method according to claim 1, characterized in that: Step S1 includes the following steps: S11: Acquire aerial images and build an aerial image dataset; S12: preprocess each aerial image in the aerial image dataset; S13: Use the preprocessed aerial images to construct training and test sets.