Self-similarity-based adversarial patch generation method

By generating and optimizing self-similarity adversarial patches, the attack success rate and stability of adversarial patches at different distances are solved, achieving more efficient attack effects and model robustness.

CN120355969APending Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196701.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the attack success rate and stability of the adversarial patch at different distances are low, resulting in insufficient security and robustness to deep learning models.

Method used

By obtaining the initial adversarial patch and performing transformations, the target adversarial samples are generated on the training image, and the total loss is determined by combining the unprintable score, smoothness loss, adversarial detection loss and self-similarity to obtain the optimal adversarial patch, maintaining the self-similarity of local and global patches to improve attack stability.

Benefits of technology

It improves the attack success rate and stability of the adversarial patch at different distances, enhances the physical reproducibility and robustness, and improves the attack effect of the adversarial patch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355969A_ABST
    Figure CN120355969A_ABST
Patent Text Reader

Abstract

The invention discloses an adversarial patch generation method based on self-similarity, and the method comprises the steps: obtaining an initial adversarial patch, carrying out the patch change of the initial adversarial patch, and obtaining a transformed patch; rendering the transformed patch to the collected training image to obtain a target confrontation sample; acquiring adversarial detection loss based on the target adversarial sample and the target detection model, and respectively acquiring an unprintable score, a smoothness loss and a self-similarity loss based on the initial adversarial patch; the total loss is determined through the non-printable score, the smoothness loss, the confrontation detection loss and the self-similarity loss; when the total loss does not meet the preset condition, the pixel value of the initial adversarial patch is updated through the total loss, the updated adversarial patch is rendered to the collected image circularly again, and new total loss calculation is carried out again until the optimal adversarial patch is obtained; according to the method, the attack success rate and stability of the adversarial patch under different distances are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of adversarial attacks, and relates to, but is not limited to, a method for generating adversarial patches based on self-similarity. Background Art

[0002] As a powerful adversarial example attack method, adversarial patch attack interferes with the detection results of deep learning models by adding carefully designed patches to images. Its strong aggressiveness poses a severe challenge to the security and robustness of deep learning models. However, this attack method also brings a series of beneficial effects and inspirations. On the one hand, adversarial patch attack reveals the vulnerability of deep learning models when facing specific types of perturbations, which prompts researchers to deeply explore and improve the defense strategies of the models. By simulating such attacks, potential vulnerabilities of the models can be discovered, and targeted improvements can be made to enhance the robustness of the models against adversarial patch attacks. This research not only helps to enhance the security of the models in actual deployment but also provides new methods and tools for model security assessment. On the other hand, the research on adversarial patch attack also inspires new research directions. During the process of exploring such attacks, researchers constantly discover new attack methods and defense strategies, which further enriches the research field of adversarial example attacks. This research also promotes the cross-integration between different disciplines and provides new perspectives and methods for solving complex problems. That is, the stronger the adversarial patch attack effect, the more deeply it can reveal the model vulnerability, promote the improvement of model robustness, deepen the understanding, drive the development of defense technologies, and expand the research field.

[0003] In related technologies, considering the attack effect of adversarial patch images under relatively short distances, adversarial perturbations are added to the original images, and the obtained adversarial samples are used as the input of the target model. The output prediction scores are used as part of the loss, and the adversarial perturbations are updated by gradient descent, as well as scaling and nesting the patches or optimizing them in regions. The above methods are generated considering distance changes, but there will be mutual interference between different regions during the optimization process, and simple scaling and nesting will also cause problems such as unstable final attack effects.

[0004] Therefore, how to improve the attack success rate and stability of adversarial patches at different distances has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a method for generating adversarial patches based on self-similarity, which at least solves the problem of low attack success rate and stability of adversarial patches at different distances in related technologies.

[0006] According to the first aspect of the embodiment of the present invention, a method for generating adversarial patches based on self-similarity is provided, including:

[0007] Obtain an initial adversarial patch, and perform patch transformation on the initial adversarial patch to obtain a transformed patch;

[0008] Render the transformed patch onto the collected training image to obtain a target adversarial sample;

[0009] Based on the target adversarial sample and the target detection model, obtain an adversarial detection loss, and based on the initial adversarial patch, respectively obtain a non-printable score, a smoothness loss, and a self-similarity loss; and determine a total loss through the non-printable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss;

[0010] When the total loss does not meet the preset condition, update the pixel values of the initial adversarial patch through the total loss, and render the updated adversarial patch onto the collected image again and calculate a new total loss again until an optimal adversarial patch is obtained.

[0011] According to a second aspect of an embodiment of the present invention, there is provided an electronic device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the method described in the first aspect.

[0012] According to a third aspect of an embodiment of the present invention, there is provided a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.

[0013] According to the solution provided by the embodiment of the present invention, an initial adversarial patch is obtained, and the initial adversarial patch is patched and changed to obtain a transformed patch; the transformed patch is rendered onto the collected training image to obtain a target adversarial sample; an adversarial detection loss is obtained based on the target adversarial sample and the target detection model, and a non-printable score, a smoothness loss, and a self-similarity loss are respectively obtained based on the initial adversarial patch; and a total loss is determined through the non-printable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss; when the total loss does not meet the preset conditions, the pixel values of the initial adversarial patch are updated through the total loss, and the updated adversarial patch is rendered onto the collected image again and a new total loss is calculated again until an optimal adversarial patch is obtained. In this process, introducing the target adversarial sample into the self-similarity loss can generate adversarial patches with similar images at multiple scales. By maintaining the self-similarity of local patches and global patches at different scales, the attack stability of the patches in different distance scenarios is improved. At the same time, introducing the target adversarial sample into the non-printable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss to determine the total loss, and adjusting the pixel values of the adversarial patch through the total loss until an optimal adversarial patch is obtained, improves the attack effect, enhances the physical reproducibility, and improves the robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:

[0015] Figure 1 Schematic flowchart of a method for generating an adversarial patch based on self-similarity provided by an embodiment of the present invention Figure 1 ;

[0016] Figure 2 Schematic flowchart of a method for generating an adversarial patch based on self-similarity provided by an embodiment of the present invention Figure 2 ;

[0017] Figure 3 Effect diagram of the influence of different numbers of local adversarial patches on the training effect of global adversarial patches provided by an embodiment of the present invention;

[0018] Figure 4 Variation trend diagram of the center point offset distance of a local adversarial patch and the sum of the perpendicular distances of the local adversarial patch from the adjacent two sides of the global adversarial patch and the attack effect of the adversarial patch provided by an embodiment of the present invention;

[0019] Figure 5Schematic diagram of the effect of an optimal adversarial patch multi-stage optimization strategy provided by an embodiment of the present invention;

[0020] Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] In the following description, reference is made to "some embodiments", which describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subsets or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0023] It should be noted that the terms "first / second / third" involved in the embodiments of the present invention are only used to distinguish similar objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present invention described here can be implemented in an order other than that illustrated or described here.

[0024] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the embodiments of the present invention belong. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0025] Figure 1 Flowchart of a method for generating an adversarial patch based on self-similarity provided by an embodiment of the present invention. The method for generating an adversarial patch based on self-similarity provided by an embodiment of the present invention can be executed by an electronic device, and the electronic device can be, for example, a computer, a server, etc.

[0026] As Figure 1 shown, a method for generating an adversarial patch based on self-similarity includes:

[0027] S101. Obtain an initial adversarial patch, and perform patch transformation on the initial adversarial patch to obtain a transformed patch.

[0028] An adversarial patch is a specific type of adversarial sample, which is a small image patch that can change the classification result of a machine learning model for an image when placed or superimposed at a certain position on a normal image.

[0029] In an embodiment of the present invention, the initial adversarial patch can be randomly initialized, and a series of transformations (such as rotation, scaling, translation, etc.) are performed on the initial adversarial patch to ensure its effectiveness under different perspectives and conditions, and then the transformed patch is obtained.

[0030] S102. Render the transformed patch onto the collected training image to obtain a target adversarial sample.

[0031] In an embodiment of the present invention, the transformed adversarial patch is applied to the collected training image (pedestrian image and the label corresponding to the pedestrian image) to generate an adversarial sample. This step is usually achieved by rendering the adversarial patch at a certain position on the target image. Pixel value constraint refers to imposing a certain limit on the pixel values of the adversarial patch, such as restricting them within a certain range (e.g., 0 - 255 for 8-bit images), or maintaining a specific pixel value distribution. Smoothness constraint is used to ensure that the adversarial patch does not look too sharp or obtrusive visually, usually achieved by reducing the high-frequency components in the adversarial patch. Aggressiveness constraint adds a small perturbation to the target image to mislead the model, but at the same time keeps this perturbation visually imperceptible. Structural similarity constraint is used to ensure that the processed adversarial patch is similar to the target image in structure, which can be achieved by comparing information such as the texture and edges of the images.

[0032] Furthermore, in order to achieve continuous attacks of the adversarial patch at different distances, self-similarity constraint is added to guide the patch to generate repetitive image structures. At the same time, in order to enhance the robustness of the adversarial patch in the physical environment, two parts of patch pixel value constraint and smoothness adjustment are added. In order to add a small perturbation to the target image to mislead the model, aggressiveness constraint is added, and finally the target adversarial sample is obtained.

[0033] S103. Obtain an adversarial detection loss based on the target adversarial sample and the target detection model, and respectively obtain a non-printable score, a smoothness loss, and a self-similarity loss based on the initial adversarial patch; and determine the total loss through the non-printable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss.

[0034] In an embodiment of the present invention, an adversarial detection loss can be obtained through a target adversarial sample and a target detection model, and an unprintable score, a smoothness loss, and a self-similarity loss can be obtained through an initial adversarial patch in the first loop (in subsequent loops, the loss is further calculated on the updated adversarial patch). Finally, the total loss is calculated through the unprintable score, the total variation loss, the adversarial detection loss, the self-similarity loss, and the total loss function formula. The total loss function formula is as shown in (3) below:

[0035] L total = αL det + βL sim + γL tv + δL NPS (3)

[0036] In the above formula (3), L det is the adversarial detection loss, L sim is the self-similarity loss, L tv is the smoothness loss, and α, β, γ, and δ are hyperparameters used to scale these four losses.

[0037] S104. When the total loss does not meet the preset condition, update the pixel values of the initial adversarial patch according to the total loss, and render the updated adversarial patch onto the acquired image again and calculate a new total loss until the optimal adversarial patch is obtained.

[0038] In an embodiment of the present invention, a preset condition is set. The preset condition can be a specific value or range. When the total loss does not meet the preset condition, update the pixel values of the initial adversarial patch according to the total loss and render the updated adversarial patch onto the acquired image to obtain a target adversarial sample. Based on the target adversarial sample, the updated adversarial patch, and the target detection model, obtain the unprintable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss, and determine a new total loss through the unprintable score, the total variation loss, the adversarial detection loss, and the self-similarity loss. Further, compare the new total loss with the preset condition. When it meets the preset condition, use the updated adversarial patch as the optimal adversarial patch. When the new total loss does not meet the preset condition, perform operations such as patch transformation on the new total loss again until the optimal adversarial patch is obtained.

[0039] As Figure 2 shown, Figure 2 is a flowchart of a method for generating an adversarial patch based on self-similarity provided by an embodiment of the present invention Figure 2 . In Figure 2 , the initial anti-patch is first processed by a conversion function and then overlaid on a clean sample (the training dataset image and the corresponding label), and then the obtained adversarial sample is used as the input of the target detection model, and the loss output by the model is Ldet (Adversarial detection loss). To achieve continuous attacks of adversarial patches at different distances, a self-similarity constraint is added to guide the patch to generate repetitive image structures, and the average structural similarity loss of the output is used as L sim Meanwhile, to enhance the robustness of the patch in the physical environment, two parts, namely patch pixel value constraint and smoothness adjustment, are added, and the unprintable score loss L NPS and the smoothness loss L tv are used as another part of the loss function. Finally, the above four losses are summed to obtain the total loss, and the adversarial patch is updated by gradient descent optimization to minimize the total loss, thus achieving continuous attacks of the adversarial patch at different distances.

[0040] It can be understood that in some embodiments of the present invention, an initial adversarial patch is obtained, and the initial adversarial patch is subjected to patch transformation to obtain a transformed patch; the transformed patch is rendered onto the collected training image to obtain a target adversarial sample; based on the target adversarial sample, the initial adversarial patch, and the target detection model, an unprintable score, a smoothness loss, an adversarial detection loss, and a self-similarity loss are obtained, and the total loss is determined by the unprintable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss; when the total loss does not meet the preset threshold, the pixel value of the initial adversarial patch is updated through the total loss, and the updated adversarial patch is rendered onto the collected image again and the new total loss is calculated again until the optimal adversarial patch is obtained. In this process, introducing the target adversarial sample into the self-similarity loss can generate adversarial patches with similar images at multiple scales. By maintaining the self-similarity of local patches and global patches at different scales, the attack stability of the patch in different distance scenarios is improved. At the same time, the rendered image is introduced into the unprintable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss to determine the total loss, and the pixel value of the adversarial patch is adjusted through the total loss until the optimal adversarial patch is obtained, which improves the attack effect, enhances the physical reproducibility, and improves the robustness.

[0041] In some embodiments of the present invention, obtaining the adversarial detection loss based on the target adversarial sample and the target detection model in S103 can be achieved through S1031 to S1032, and the following steps are used for illustration.

[0042] S1031: Input the target adversarial sample into the target detection model to obtain the confidence of the detected object and the class confidence of the detected object.

[0043] S1032: Obtain the adversarial detection loss through the confidence, the class confidence, and the adversarial detection loss function.

[0044] In some embodiments of the present invention, the target adversarial sample is input into the target detection model to obtain the confidence of the detected object and the class confidence of the detected object, and the confidence of the detected object and the class confidence of the detected object are input into the adversarial detection loss function to obtain the adversarial detection loss. Among them, the adversarial detection loss function is shown in the following formulas (4) and (5):

[0045]

[0046] In the above formula (4), i is the identifier of the detection box, and Conf(x i ) is the confidence of the detected object in the i-th detection box of the input initial adversarial sample, is the class confidence of the target class (such as the "person" class) of the detected object corresponding to the i-th detection box.

[0047] In some embodiments of the present invention, obtaining the unprintable score, smoothness loss, and self-similarity loss based on the initial adversarial patch in S103 can be implemented through S103A to S103C, and the following steps are used for illustration.

[0048] S103A: Take the initial adversarial patch as the global adversarial patch, and perform a preset number of intercepts at preset positions on the initial adversarial patch to obtain a plurality of local adversarial patches.

[0049] In some embodiments of the present invention, through simulation experiments, a preset number of local adversarial patches are intercepted at preset positions on the adversarial patch in each cycle. When intercepting the local adversarial patches in the first cycle, the initial adversarial patch is taken as the global adversarial patch, and a preset number of intercepts are performed at preset positions on the initial adversarial patch to obtain a plurality of local adversarial patches.

[0050] S103B: Determine the self-similarity loss through a plurality of local adversarial patches, the global adversarial patch, and the self-similarity loss function.

[0051] In some embodiments of the present invention, the self-similarity loss function is shown in the following formula (1):

[0052]

[0053] In the above formula (1), represents the i-th local adversarial patch intercepted from the preset position of the initial adversarial patch, represents the scaling operation, P global represents the global adversarial patch, scale i represents the size of the i-th local adversarial patch accounting for the global adversarial patch, and the specific scaling ratio of the global adversarial patch P global is based on scale iAdjusted by w i represents the weight of each scale, and SSIM represents the structural similarity.

[0054] Furthermore, substituting multiple local adversarial patches and global adversarial patches into the above formula (1) yields the self-similarity loss.

[0055] Among them, the calculation of the structural similarity is shown as follows (2):

[0056]

[0057] In the above formula (2), μ x and μ y are respectively the pixel averages within the local regions of P global and ; σ x and σ y are respectively the pixel value variances within the local regions of P global and ; σ xy is the covariance of the pixel values within the local regions of P global and ; C1 and C2 are stability constants used to avoid division by zero.

[0058] S103C, obtaining the smoothness loss through the pixel coordinates of the initial adversarial patch and the smoothness loss function; and obtaining the unprintable score through the pixel sum of the initial adversarial patch and the unprintable score loss function.

[0059] In some embodiments of the present invention, the features of natural images include smooth and consistent patches, and the colors gradually change within each patch. Therefore, to increase the plausibility of physical attacks, smooth and consistent perturbations are preferred. In addition, due to sampling noise, the camera may not be able to accurately capture the extreme differences between adjacent pixels in the perturbation, and non-smooth perturbations may not be physically realizable. To solve these problems, the total variation (TV) loss is introduced to maintain the smoothness of the perturbation, and the smoothness loss function is as follows:

[0060]

[0061] In the above formula (5), i and j refer to the pixel coordinates of the adversarial patch P, which are the pixel coordinates of the initial adversarial patch in the first loop.

[0062] Furthermore, the Non-Printability Score (NPS) is an indicator used to measure the printability of adversarial patches in the real world. Specifically, it evaluates whether an ordinary printer can accurately reproduce the generated colors when printing a patch. If the colors of the patch exceed the printable color gamut of the printer, the printer may not be able to correctly present these colors, thereby affecting the effect of the patch in the actual application scenario. The non-printable score loss function is defined as follows:

[0063]

[0064] In the above formula (6), p patch is a pixel in the adversarial patch P, and c print is a color in the set of printable colors C, which is the pixel of the initial adversarial patch in the first loop.

[0065] Furthermore, substitute the pixel coordinates of the initial anti-patch into the above (5) to obtain the smoothness loss, and substitute the pixels of the initial adversarial patch into the above (6) to obtain the non-printable score.

[0066] In some embodiments of the present invention, S103A may be preceded by S10 to S13, which are described in the following steps.

[0067] S10. Intercept at fixed positions on the training adversarial patch respectively to obtain multiple groups of training local adversarial patches with different numbers, and determine the optimal number among the multiple groups of different numbers.

[0068] S11. Perform patch interception in different proportions at different positions on the training adversarial patch according to the optimal number to obtain multiple groups of training local adversarial patches with different proportions at different positions.

[0069] S12. Conduct experiments based on the training adversarial patch and the training local adversarial patches, and determine the optimal position according to the experimental results.

[0070] S13. Use the optimal number and the optimal position as the preset number and the preset position.

[0071] In some embodiments of the present invention, the distance between the patch image and the sensor directly affects the size, blur degree of the adversarial patch, and the range of the adversarial patch captured by the detection model. Therefore, it is considered to generate adversarial patches that have attack effects at different scales, that is, the adversarial patches should have similar images at different scales. Inspired by fractal patterns, a fractal is a rough or fragmented geometric shape that can be divided into several parts, each of which is a reduced version of the whole (at least approximately), and it has irregularity, hierarchical nesting, and self-similarity. Among them, self-similarity is an invariance under scale transformation, that is, when observing the adversarial patch at different scales, an approximate and identical image can be seen. If the local part of the entire adversarial patch is magnified, and then the local part of the local part is magnified, similar structural features can be seen, that is, a unified and repetitive structural pattern is shown at different scales. Research shows that compared with other evaluation metrics, structural similarity can more accurately reflect the visual quality of the adversarial patch. At the same time, through experiments with the adversarial patch represented as sub-adversarial patches at multiple resolutions (such as multi-scale wavelet decomposition), it is found that structural similarity can better capture the similarity of the adversarial patch at different resolutions and is more robust in dealing with local transformations and capturing the structural information of the adversarial patch. Therefore, structural similarity is more suitable as an indicator for evaluating the self-similarity of the adversarial patch.

[0072] To generate adversarial patches with high self-similarity characteristics, by intercepting different local adversarial patches, the structural similarity between the local adversarial patch and the global adversarial patch is used as a constraint. However, during the generation process of the optimal adversarial patch, affected by multiple factors such as the number of local adversarial patches, the nesting method of local adversarial patches, and the position of local adversarial patches, the self-similarity and attackability of the generated adversarial patch will also change. Therefore, by adjusting the number of local adversarial patches, the relative positions between local adversarial patches, and the relative positions between local adversarial patches and the global adversarial patch, observe their effects on the finally generated optimal adversarial patch and the object detection model, so as to obtain the optimal self-similarity patch generation method.

[0073] First, by intercepting different numbers of training local adversarial patches (when determining the optimal number and optimal position, the training local adversarial patches are local adversarial patches) at fixed positions of the training adversarial patch (when determining the optimal number and optimal position, the training adversarial patch is the global adversarial patch) to observe its impact on the optimal adversarial patch. To facilitate observing the impact of the number on the optimization result of the adversarial patch, the most special type of adversarial patch nesting method is selected, that is, central nesting (i.e., the centers of local adversarial patches of different sizes coincide, and both coincide with the center of the global adversarial patch) for experiments. The impacts on the optimization result are respectively compared under the conditions where the number of local adversarial patches is 2, 3, and 4. It is found that under the same training condition settings, when the number of local adversarial patches is 2 and 4, a shape similar to a light beam is likely to appear, affecting the aggressiveness of the optimal adversarial patch. Specifically, as Figure 3 shown Figure 3 is the schematic diagram of the impact of different numbers of local adversarial patches on the training effect of the global adversarial patch provided by the embodiment of the present invention.

[0074] To compare the impacts of different nesting methods of local adversarial patches and the positions of local adversarial patches on the optimization result, the determination method of local nested adversarial patches is defined as follows: First, a point is randomly selected in the global adversarial patch as the center point of the local patch with the smallest proportion among 3 local adversarial patches (the local adversarial patch with the largest proportion, the local adversarial patch with the middle proportion, and the local adversarial patch with the smallest proportion). At the same time, it is judged whether the local adversarial patch with the smallest proportion is completely located in the global adversarial patch. If the judgment is true, this point is the center point of the local adversarial patch with the smallest proportion, and this point is (x1, y1); based on the first center point, a random offset direction θ and a random offset distance d are given. The random offset direction θ takes the positive x-axis direction as the reference axis and increases in the clockwise direction, and the value range is (0, 2π). The value range of the random offset distance d is (0, d max ), and the value of d max depends on the ratio of the two local patches, and the specific calculation method is as shown in the following formula (7):

[0075]

[0076] In the above formula (7), S large is the area of the local adversarial patch with the middle proportion, and S small is the area of the local adversarial patch with the smallest proportion.

[0077] The center point of the local adversarial patch with the middle proportion is obtained through the following formula (8), and formula (8) is as follows:

[0078]

[0079] Meanwhile, it is determined whether the local adversarial patch with the intermediate ratio completely contains the local adversarial patch with the minimum ratio and whether it is completely located within the global adversarial patch. The specific discrimination formula is as shown in the following (9):

[0080]

[0081] In the above (8) and (9), w large is the width of the local adversarial patch, h large is the height of the local adversarial patch, (x small , y small ) represents the center point of the local adversarial patch with the minimum ratio, (x large , y large ) represents the candidate center point of the local adversarial patch with the intermediate ratio obtained according to the random distance and direction. If all four discriminations are true, then this point is the center point (x2, y2) of the local adversarial patch with the intermediate ratio, and at the same time, the first random distance d1 is recorded; based on the center point of the local adversarial patch with the intermediate ratio, any offset direction and distance are given to obtain the center point of the local adversarial patch with the maximum ratio. At the same time, it is determined whether the local adversarial patch with the maximum ratio completely contains the local adversarial patch with the intermediate ratio and whether it is completely located within the global adversarial patch. If all four judgments are true, then this point is the center point (x3, y3) of the local adversarial patch with the maximum ratio, and at the same time, the second random offset distance d2 is recorded. After obtaining the local adversarial patch with the maximum ratio, the sum of the two offset distances is obtained to get the sum of the two random offset distances d total . Then, according to the generated local adversarial patch with the maximum ratio, the minimum vertical distance sum d min from the adjacent two sides of the global adversarial patch is obtained. The specific calculation method is as shown in the following (10):

[0082]

[0083] d a = min(d left , d right )(11)

[0084] d b = min(d top , d bottom ), d min = d a + d b (12)

[0085] In the above formulas (10), (11) and (12), w3 is the width of the local adversarial patch, h3 is the height of the local adversarial patch, d left , d right , d top and d bottomrespectively represent the vertical distances of the local adversarial patch with the largest proportion from the left, right, top, and bottom boundaries of the global adversarial patch. The sum of the two random offset distances d total and the sum of the minimum vertical distances to the adjacent sides of the global adversarial patch d min As two factors affecting the attack effect of the self-similarity patch, it is found that the two offset distances d total of the central adversarial patch are positively correlated with the result of adversarial patch training, that is, as the sum of the two offset distances d total increases, the aggressiveness of the adversarial patch becomes stronger; however, the sum of the minimum vertical distances d min to the adjacent sides of the global adversarial patch is negatively correlated with the result of adversarial patch training, that is, as the sum of the minimum vertical distances of the local adversarial patch with the largest proportion from the adjacent sides of the global adversarial patch d min increases, the aggressiveness of the adversarial patch becomes weaker. Such a phenomenon may be due to when the sum of the two offset distances d total of the center of the local adversarial patch is large and the sum of the vertical distances d min from the global adversarial patch is small, the remaining area of the entire adversarial patch forms a continuous available space, avoiding unnecessary segmentation and discontinuity of the remaining space, enhancing the continuity of the texture generation space of the adversarial patch, thereby allowing larger and more consistent texture structures to be generated in these areas. Such texture structures help maintain visual consistency and coherence during training, and at the same time, increase the attack effect of the adversarial patch.

[0086] As Figure 4 shown, Figure 4 This is a trend graph of the change in the offset distance of the center point of the local adversarial patch and the sum of the vertical distances of the local adversarial patch from the adjacent sides of the global adversarial patch and the attack effect of the adversarial patch provided by the embodiment of the present invention. It is observed that as the sum of the two offset distances increases, the aggressiveness of the adversarial patch becomes stronger. At the same time, as the sum of the minimum vertical distances of the local adversarial patch with the largest proportion from the adjacent sides of the global adversarial patch increases, the aggressiveness of the adversarial patch becomes weaker. At the same time, it is found that during the training process, whether the trajectory formed after the two offsets of the center point of the local adversarial patch is a straight line will also affect the training result of the adversarial patch. The center points forming a broken line do not conform to the movement trajectory of normal devices and it is difficult to achieve continuous adversarial attacks on moving unmanned devices. Therefore, the distribution method of the center points forming a broken line is difficult to achieve the optimal in terms of continuous attack effect.

[0087] Through the above experiments, it is determined that the number of local adversarial patches (optimal number) is 3. The upper left corner of the three local adversarial patches coincides with the upper left corner of the global adversarial patch, that is, the two offset distances of the center point are the largest and form a straight line, and the sum of the offset distances to the adjacent sides of the global adversarial patch is 0 (optimal position).

[0088] It is understandable that in some embodiments of the present invention, by setting two variables, namely the minimum offset and the minimum vertical distance, and analyzing the influence of the local patch composition method on the patch attack effect during the training process, the composition method of the general optimal nested patch is obtained. By verifying the influence of the local patch composition method on the attack effect, the deficiency of how the local structure of the nested patch affects the attack effect is solved, providing the best design scheme for multi-scale adversarial patches, and effectively improving the attack success rate and stability of the patches at different distances.

[0089] In the embodiments of the present invention, as Figure 5 described, Figure 5 is a schematic diagram of the effect of an optimal adversarial patch multi-stage optimization strategy provided by an embodiment of the present invention. In Figure 5 , the training is divided into multiple stages. As Figure 5 shown, during the optimization process of the adversarial patch, the detection loss L det and the self-similarity loss L sim weights are dynamically adjusted according to the training requirements of different stages to gradually optimize the attack effect of the adversarial patch and the visual consistency at different scales. In the initial stage of training (in stage 1), the similarity loss weight β is mainly increased in formula (3) to ensure the consistency of the adversarial patch at different scales and distances, and improve the robustness and generalization ability of the patch. As the patch gradually learns visually stable images, in the later stage (when the total loss function oscillates without decreasing for 50 epochs in stage 2), the adversarial detection loss weight α will be gradually increased in formula (3) to ensure that the adversarial patch can effectively attack the target detection model. In this way, overfitting of the adversarial patch at a single scale is avoided, its adaptability in different distance detection scenarios is ensured, and its attack effect at different scales is maintained.

[0090] In the embodiments of the present invention, the results of the optimal adversarial patch are presented, including details of the experimental setup and digital simulation experiment evaluation, specifically:

[0091] (1) Details of the experimental setup

[0092] In the digital image attack experiment, we used the Inria Person dataset, which includes a training set of 614 pedestrian images and a test set of 288 images. The generated adversarial patch aims to achieve the "invisibility" of pedestrians under the target detection algorithm. Therefore, to evaluate the attack effect of the adversarial patch, our evaluation metric is the mean average precision (mAP), which is a commonly used performance metric in the target detection task. To calculate the mAP, the detection boxes generated by each detector on the clean dataset are used as the ground truth boxes, and then the same detector is used to detect the dataset with the adversarial patch added, and the mAP of the detector in this case is calculated. The greater the decrease in mAP, the stronger the attack effect. The adversarial patch designed in this invention is mainly used to deceive the mainstream first-order target detection algorithms, YOLOv2 and YOLOv5, and has an attack effect in both large object and small target detections, with a focus on the attack effect on small targets. All detection models are trained on the coco dataset, the confidence threshold in the output stage is set to 0.5, and the non-maximum suppression (NMS) threshold is set to 0.4. The entire algorithm is designed based on pytorch and trained on a device equipped with an NVIDIA RTX4090 graphics card. The adversarial patch is set to a size of 3*300*300, the Adam optimizer is used, and the learning rate is set to 0.02, β1 = 0.9, β2 = 0.999.

[0093] (2) Digital simulation experiment evaluation

[0094] To more fully evaluate the superiority of self-similar adversarial patches in overall attack performance and small-target attack performance, the present invention compares the related work of adversarial patches that has been made public in recent years and sets up seven groups of experiments to compare the attack effects of different adversarial patches. Among them, RPAU forms nested patches by stacking sub-patches with the same distribution, enabling the adversarial patch to capture effective perturbations at any distance; Nested-AE aims to optimize different parts of the adversarial patch according to different distances to achieve the effectiveness of the adversarial patch at multiple distances; the purpose of AdvPatch is to reduce the maximum object score in all predicted bounding boxes by using a rectangular adversarial patch at the center of the target object; UPC aims to reduce the sum of the class scores of the target class of the detectable bounding boxes that are the target objects; DAP utilizes the prior information of natural images and directly changes the pixel values to produce an adversarial effect; DOE uses the method of model integration to adjust the weight parameters of the attacked target model to find the balance point where the generated adversarial patch can effectively attack all target models. By applying all the adversarial patches to the images in the test set of the Inria pedestrian dataset, the mAP value of the "person" category is calculated for different object detection models. Among them, for the Yolov5 model, the attack success rate of the present invention can reach 91.91%. At the same time, good attack effects are maintained for small targets on different detection models. Table 1 shows the test results in the digital world, representing the attack performance with the change of the mAP value. The left axis lists the attacked object detection models, and the upper axis shows all the adversarial patches.

[0095] Furthermore, to calculate the mAP, in the embodiments of the present invention, the detection boxes generated by each detection model on the clean dataset are used as the ground truth boxes, and then the same detection model is used to detect the dataset with the adversarial patches added, and the mAP of the detection model in this case is calculated. According to Table 2, the attack performance of the self-similarity patches trained for the Yolov5 model is superior to the patches of other methods. At the same time, as can be seen from Table 3, for small targets (the results are evaluated using the mAP calculation module in the COCOAPI, where the definition of a small target is that the area of the target box is less than 32 square pixels), the self-similarity patches of the embodiments of the present invention all obtain lower mAP values.

[0096] Table 2 Comparison of the attack effects between self-similarity patches and advanced adversarial patch methods

[0097] Detection model ours RPAU nested-AE AdvPatch UPC DAP DOE Yolov4 21.80 29.80 46.50 32.40 44.40 35.20 32.10 Yolov5 8.09 10.68 10.99 8.68 29.20 14.88 8.92 Yolov7 56.50 57.10 60.90 58.80 65.70 61.90 59.30 Detr 48.10 53.20 51.90 49.20 55.70 57.10 48.90 Faster-RCNN 48.90 55.70 50.80 52.20 66.90 59.20 49.60

[0098] Table 3 Attack effects of self-similarity patches and advanced adversarial patch methods on small targets

[0099] Detection model ours RPAU nested-AE AdvPatch UPC DAP DOE Yolov4 43.20 52.10 78.50 61.40 76.60 75.40 72.40 Yolov5 20.80 26.70 36.90 35.80 57.10 57.40 56.20 Yolov7 29.50 32.20 40.20 36.50 44.70 40.90 32.60 Detr 22.80 28.90 30.20 24.80 36.70 31.50 26.20 Faster-RCNN 63.70 65.40 64.40 65.90 75.50 67.70 66.90

[0100] Refer to Figure 6, showing a schematic structural diagram of an electronic device according to an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the electronic device.

[0101] As Figure 6 shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communication bus 508.

[0102] Among them:

[0103] The processor 502, the communications interface 504, and the memory 506 communicate with each other through the communication bus 508.

[0104] The communications interface 504 is used to communicate with other electronic devices or servers.

[0105] The processor 502 is used to execute the program 510, and specifically may execute the relevant steps in the above method embodiments.

[0106] Specifically, the program 510 may include program code, and the program code includes computer operation instructions.

[0107] The processor 502 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0108] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0109] The program 510 is specifically used to cause the processor 502 to execute the operations corresponding to the methods described in the above method embodiments.

[0110] For the specific implementation of each step in the program 510, reference may be made to the corresponding steps and descriptions in the corresponding units in the above method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.

[0111] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present invention can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present invention.

[0112] The methods according to the embodiments of the present invention described above can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the methods described herein can be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0113] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present invention.

[0114] The above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention shall be defined by the claims.

Claims

1. A method for generating adversarial patches based on self - similarity, characterized in that, Including: Obtain an initial adversarial patch, and perform patch transformation on the initial adversarial patch to obtain a transformed patch; Render the transformed patch onto the collected training image to obtain a target adversarial sample; Obtain an adversarial detection loss based on the target adversarial sample and a target detection model, and respectively obtain a non-printable score, a smoothness loss, and a self-similarity loss based on the initial adversarial patch; and determine a total loss through the non-printable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss; When the total loss does not meet the preset condition, update the pixel values of the initial adversarial patch through the total loss, and loop again to render the updated adversarial patch onto the collected image and calculate a new total loss again until an optimal adversarial patch is obtained.

2. The method according to claim 1, wherein Obtaining an adversarial detection loss based on the target adversarial sample and a target detection model includes: Input the target adversarial sample into the target detection model to obtain the confidence of the detected object and the class confidence of the detected object; Obtain the adversarial detection loss through the confidence, the class confidence, and an adversarial detection loss function.

3. The method according to claim 1, wherein The obtaining the non-printable score, the smoothness loss, and the self-similarity loss respectively based on the initial adversarial patch includes: Take the initial adversarial patch as a global adversarial patch, and perform a preset number of intercepts at preset positions on the initial adversarial patch to obtain a plurality of local adversarial patches; Determine the self-similarity loss through the plurality of local adversarial patches, the global adversarial patch, and a self-similarity loss function; Obtain the smoothness loss through the pixel coordinates of the initial adversarial patch and a smoothness loss function; and obtain the non-printable score through the pixel sum of the initial adversarial patch and a non-printable score loss function.

4. The method according to claim 3, characterized in that, Before taking the initial adversarial patch as a global adversarial patch and performing a preset number of intercepts at preset positions on the initial adversarial patch to obtain a plurality of local adversarial patches, the method further includes: Perform intercepts at fixed positions on the training adversarial patch to obtain multiple groups of first training local adversarial patches with different numbers, and determine an optimal number among the multiple groups with different numbers; Perform patch intercepts with different ratios at different positions on the training adversarial patch according to the optimal number to obtain multiple groups of second training local adversarial patches with different ratios at different positions; Conduct experiments based on the training adversarial patch and the second training local adversarial patches, and determine an optimal position according to the experimental results; Take the optimal number and the optimal position as the preset number and the preset position.

5. The method according to claim 3, characterized in that, The self-similarity loss function is as follows in formula (1): In the above formula (1), represents the i-th local adversarial patch intercepted from the preset position of the initial adversarial patch, represents a scaling operation, P global represents the global adversarial patch, scale i represents the size of the i-th local adversarial patch relative to the global adversarial patch, and the specific scaling ratio of the global adversarial patch P global is adjusted according to the scale i ; w i represents the weight of each scale, and the SSIM (Structural Similarity) represents the structural similarity; In the above formula (2), x is the global adversarial patch, y represents a certain proportion of the local adversarial patch, and μ x is the average pixel value within the local region of P global , and μ y is the average pixel value within the local region of . σ x and σ y are the variances of the pixel values within the local regions of P global and respectively, and σ xy is the covariance of the pixel values within the local regions of P global and P i local . C1 and C2 are stability constants used to avoid division by zero.

6. The method according to claim 1, wherein The determining the total loss through the non-printable score, the smoothness loss, the adversarial detection loss, and the self-similarity loss includes: Calculate the total loss through the non-printable score, the smoothness loss, the adversarial detection loss, the self-similarity loss, and a total loss function, and the total loss function is as shown in (3) below: L total = αL det + βL sim + γL tv + δL NPS (3) In the above formula (3), L det is the adversarial detection loss, L sim is the self-similarity loss, L tv is the smoothness loss, and α, β, γ, and δ are hyperparameters used to scale these four losses.