Adversarial Attack Method, Defense Method and Device Based on Enhanced General Patch

Through the anti-attack and defense methods of enhanced universal patches, the problems of insufficient migration and inefficiency of defense methods in the existing technology in the black box environment are solved, and the anti-attack patch with strong migration in the black box environment is generated, and efficient defense effects are achieved.

CN114359653BActive Publication Date: 2025-07-08BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111465111.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-07-08
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

The existing adversarial patches tend to converge prematurely during the generation process in a white box environment, resulting in insufficient migration in a black box environment. At the same time, the existing defense methods have limitations in detection efficiency and time cost, and have poor adaptability, making it difficult to effectively defend against various types of adversarial patch attacks.

Method used

Through an adversarial attack method based on enhanced universal patches, adversarial retraining and integrated network models are used to generate adversarial patches, deeper features of target categories are mined, and the defense method of local details is used to pre-process the gradient differences in the image key areas to enhance the black box migration and defense capabilities of adversarial patches.

Benefits of technology

The generated enhanced universal patch has stronger migration in a black box environment and can effectively attack multiple types of adversarial patches. At the same time, the defense method achieves effective defense against universal anti-adversarial patches at low cost and high efficiency, improving the defense effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359653B_ABST
    Figure CN114359653B_ABST
Patent Text Reader

Abstract

The present invention provides an adversarial attack method, a defense method and a device based on an enhanced general patch. The adversarial attack method includes: obtaining an image sample set; randomly selecting a first image sample from the image sample set as a first image to be trained; obtaining an original patch image, and performing image fusion on the original patch image and the first image to be trained to obtain a first adversarial sample image; inputting the first adversarial sample image into a patch generation model, and updating the patch image based on the gradient descent method to generate an adversarial patch image; randomly selecting a second image sample from the image sample set as a second image to be trained; performing image fusion on the adversarial patch image and the second image to be trained to obtain a second adversarial sample image, inputting the second adversarial sample image into an integrated network model, and updating the adversarial patch image based on the gradient descent method to generate an enhanced general patch image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular, to an adversarial attack method, a defense method and a device based on an enhanced general patch. Background Art

[0002] Deep learning plays an important role in various fields of our lives, including computer vision (CV), pattern recognition (PR), natural language processing (NLP), and other aspects. At the same time, more and more studies have shown that deep neural networks are vulnerable to adversarial samples, which not only occur in laboratory environments, but can also attack various CNN image recognition systems in the real world.

[0003] For real-world attacks, the existence of adversarial patches proves that such perturbations can easily attack white-box neural networks both in laboratory environments and in real-world environments. However, there is still a distance between adversarial patches and the features of target classes in visual perception, and adversarial defense can also use this characteristic to make judgments on adversarial samples. The main reason is that the generation of adversarial patches depends on a large amount of data training, and in the training of white boxes, premature convergence often occurs in the early stage of patch generation, and the prematurely converged patches often cannot fully mine the deep features of the target class. For the defense of real-world attacks, most of the existing methods make judgments on adversarial samples by relying on detection means through the gradient difference between images and patches. On the one hand, the detection efficiency of these methods will significantly decrease when facing patches after special processing such as image smoothing, which will further cause the decrease of defense efficiency. On the other hand, their generality and time cost are also relatively high; while some conventional preprocessing defense means such as image blurring or compression are usually inefficient.

[0004] Currently, Brown et al. proposed a general patch block to achieve targeted attacks, and achieved the robustness of the patch to position and angle by continuously optimizing and adjusting the gradient; Karmon further optimized the loss function on this basis and further explored the impact of the relationship between position and category on the patch to achieve attacks. The existing generation of such general patches is to achieve attacks on neural networks through the attack fitting process of white boxes, and the patches obtained by such a patch generation process largely depend on the attack difficulty of the white boxes they rely on; if the attack difficulty is too low, it means premature convergence of the patches, and more semantic information of the target class cannot be generated; which will further lead to insufficient transferability of adversarial patches in the black-box environment.

[0005] In addition, the defense strategies of general patch blocks are mainly divided into two categories: detection and perturbation and detection and repair. Naseer et al. proposed a local gradient smoothing scheme to counter physical world patch attacks; before inputting the image into CNN, the gradient of the estimated patch area is regularized to eliminate the impact of the adversarial attack. Hayes et al. proposed a physical image adversarial attack defense method based on image repair; based on traditional image processing methods, they detected the location of the patch in the input image and further used image repair technology to remove the patch. Xu et al. proposed a similarity measurement method, which uses the CAM algorithm to obtain the area in the image that has the greatest impact on the result; intercept the area, and then compare the similarity of the area with the image area belonging to the predicted label in the training data set; if the degree of dissimilarity exceeds a threshold, the input sample is determined to be an attack sample. However, the existing adversarial patches rely on the detection of the perturbed image, and have the characteristics of low detection efficiency and high detection time cost in the detection process. Therefore, this type of detection-based defense method still has certain limitations in terms of universality and time cost. At the same time, since detection often relies on the difference between patches and normal images, mainly in terms of gradient and semantic information, the currently used defense methods have poorer adaptability and can often only defend against a single type of patch attack. Therefore, how to provide a universal patch with stronger black-box migration capabilities and how to effectively defend against universal adversarial patches are technical issues that need to be solved urgently. Summary of the invention

[0006] In view of this, the present invention provides an anti-attack method, a defense method and a device based on an enhanced universal patch to solve one or more problems existing in the prior art.

[0007] According to one aspect of the present invention, the present invention discloses an anti-attack method based on an enhanced universal patch, the method comprising:

[0008] Acquire an image sample set, and randomly select a first image sample from the image sample set as a first image to be trained;

[0009] Acquire an original patch image, and perform image fusion on the original patch image and the first image to be trained to obtain a first adversarial sample image;

[0010] Inputting the first adversarial sample image into a patch generation model, and updating the patch image based on a gradient descent method to generate an adversarial patch image;

[0011] Randomly select a second image sample from the image sample set as the second image to be trained. Perform image fusion on the adversarial patch image and the second image to be trained to obtain a second adversarial sample image. Input the second adversarial sample image into the integrated network model, and update the adversarial patch image based on the gradient descent method to generate a strengthened general patch image.

[0012] In some embodiments of the present invention, the method further includes:

[0013] Perform image fusion on the adversarial patch image and the first image to be trained to obtain a third adversarial sample image. Perform robustness training on the patch generation model based on the original label corresponding to the first image to be trained and the third adversarial sample image.

[0014] In some embodiments of the present invention, the integrated network model is a series model of multiple convolutional neural networks.

[0015] In some embodiments of the present invention, performing image fusion on the original patch image and the first image to be trained to obtain a first adversarial sample image includes:

[0016] Obtain the focus of the first image to be trained based on the CAM algorithm;

[0017] Fuse the original patch image with the region of the first image to be trained that is far from the focus.

[0018] According to another aspect of the present invention, a defense method based on a strengthened general patch is also disclosed. The method includes:

[0019] Obtain an image to be recognized. Input the image to be recognized into a network model to generate a heat map of the image to be recognized, and determine the key region of the image to be recognized based on the heat map;

[0020] Perform multi-scale Gaussian blur processing on the image to be recognized to obtain multiple first low-resolution images;

[0021] Calculate the differences between the first low-resolution images at different scales, and obtain the local detail information of the image to be recognized at different scales based on the differences between the first low-resolution images at different scales;

[0022] Fuse each local detail information with the image to be recognized to obtain an enhanced image;

[0023] Obtain the local detail information of the enhanced image, and perform weakening processing on the obtained local detail information of the enhanced image to obtain a restored image corresponding to the image to be recognized.

[0024] In some embodiments of the present invention, the network model is a CAM visualization model.

[0025] In some embodiments of the present invention, obtaining the local detail information of the enhanced image includes:

[0026] Performing multi-scale Gaussian blur processing on the enhanced image to obtain a plurality of second low-definition images;

[0027] Calculating the differences between the second low-definition images at different scales, and obtaining the local detail information of the enhanced image at different scales based on the differences between the second low-definition images at different scales.

[0028] In some embodiments of the present invention, the calculation formula for the restored image is:

[0029]

[0030] wherein, I out is the restored image, I en is the enhanced image, m ij is the fusion ratio coefficient, and G de is the Gaussian blur set under different selected Gaussian kernels.

[0031] Correspondingly, the present invention also discloses an adversarial attack and defense system based on an enhanced general patch. The system includes a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method described in any one of the above embodiments.

[0032] According to another aspect of the present invention, a computer-readable storage medium is also disclosed, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in any one of the above embodiments.

[0033] In the adversarial attack method based on the enhanced general patch disclosed by the present invention, in the initialization training stage, deeper features of the target category are mined through adversarial retraining; then, the details of the patch are enhanced by combining the integrated model training method, thereby improving the black-box transferability of the adversarial patch. The defense method based on the enhanced general patch disclosed by the present invention only focuses on the key attention areas of the image and performs local processing on the image. This method makes full use of the difference between the patch gradient and the normal image gradient, performs undifferentiated preprocessing on all images, maximizes the influence of the preprocessing on the patch, and minimizes its influence on the normal image. Therefore, this defense method can effectively defend against general adversarial patches.

[0034] Additional advantages, objects, and features of the present invention will be partly set forth in the description which follows, and will partly become obvious to those of ordinary skill in the art upon examination of the following, or may be learned by practice of the present invention. The objects and other advantages of the present invention may be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings.

[0035] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to those specifically described above, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application, but do not limit the present invention. The components in the drawings are not drawn to scale, but are only for showing the principles of the present invention. For the convenience of showing and describing some parts of the present invention, the corresponding parts in the drawings may be enlarged, that is, may become larger relative to other components in the exemplary device actually manufactured according to the present invention. In the drawings:

[0037] Figure 1 is a schematic flowchart of an adversarial attack method based on an enhanced universal patch according to an embodiment of the present invention.

[0038] Figure 2 is a schematic flowchart of a defense method based on an enhanced universal patch according to an embodiment of the present invention.

[0039] Figure 3 is a generation strategy diagram of an enhanced universal patch according to an embodiment of the present invention.

[0040] Figure 4 is a defense strategy diagram of a general adversarial patch according to an embodiment of the present invention.

[0041] Figure 5 is a diagram showing the attack effects of a general adversarial patch and an enhanced universal patch according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] To make the objects, technical solutions, and advantages of the embodiments of the present invention clearer, the following further describes the embodiments of the present invention in detail with reference to the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0043] Here, it should be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, and other details less related to the present invention are omitted.

[0044] It should be emphasized that the term "comprising / including / having" as used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0045] The countermeasure patches commonly used in the prior art are effective against attacks in the digital or physical world, but they are only for successfully attacking a specified white-box model, and their transferability in a black-box environment is often insufficient. Therefore, to solve this problem, the present invention proposes an adversarial attack method based on enhanced universal patches, that is, generating Tsinghua-type universal patches based on an adversarial retraining method. Additionally, due to the problem of insufficient defense ability in the current general patch defense method, to solve this problem, the present invention also proposes a defense method based on enhanced universal patches, that is, performing defense based on a data preprocessing-based defense method.

[0046] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0047] Figure 1 It is a flowchart of an adversarial attack method based on enhanced universal patches according to an embodiment of the present invention. As Figure 1 shown, the adversarial attack method at least includes steps S10 to S40.

[0048] Step S10: Obtain an image sample set, and randomly select a first image sample from the image sample set as the first image to be trained.

[0049] In this step, each image sample in the image sample set is a pure image, that is, an image that can be correctly recognized and classified. For example, the first image to be trained randomly selected from the image sample set is a cat image, and the image label recognized by the model when the first image to be trained is not adversarially attacked should be cat.

[0050] Step S20: Obtain an original patch image, and perform image fusion on the original patch image and the first image to be trained to obtain a first adversarial sample image.

[0051] Among them, the original patch image can be randomly selected or initialized and generated, and the pixels of the original patch image are random noise; in this step, the initialized original patch image is further fused with the first image to be trained selected in step S10, that is, the original patch image is added to the first image to be trained, thereby generating a first adversarial sample image with a patch. For example, as Figure 3 shown, in the generation strategy diagram of this enhanced universal patch, the first adversarial sample image is the image located at the upper left of the strategy diagram, and this image is the fusion result of the cat image and the original patch image.

[0052] Further, the original patch image and the first image to be trained are fused to obtain a first adversarial sample image, which specifically includes: obtaining the focus of the first image to be trained based on the CAM algorithm; fusing the original patch image with the region of the first image to be trained that is far from the focus. For example, for Figure 3 the first adversarial sample image in, the original patch image is specifically placed in the lower left corner region of the first-generation training image.

[0053] Step S30: Input the first adversarial sample image into the patch generation model, and update the patch image based on the gradient descent method to generate an adversarial patch image.

[0054] Among them, the patch generation model can specifically be a CNN convolutional neural network model. The adversarial patch image is an adversarial patch obtained by performing multiple update iterations on the original patch image. In this step, specifically, the first adversarial sample image is input into the CNN convolutional neural network model, and then the gradient descent method is used to iteratively update the original patch image, thereby obtaining the adversarial patch image. And, since there are multiple images in the image sample set, in order to obtain a general adversarial patch image, other image samples in the image sample set can be reselected as the first image to be trained, and then the adversarial patch image obtained in the first pre-training process is added to other image samples in the image sample set to pre-train the patch generation model; after multiple pre-trainings, if all the image samples in the image sample set have been trained, the final adversarial patch image is generated.

[0055] Step S40: Randomly select a second image sample from the image sample set as the second image to be trained, fuse the adversarial patch image with the second image to be trained to obtain a second adversarial sample image, input the second adversarial sample image into the integrated network model, and update the adversarial patch image based on the gradient descent method to generate a strengthened general patch image.

[0056] In this step, the second image to be trained is a pure image in the image sample set similar to the first image to be trained. Such as Figure 3As shown, the images with elephant patterns and horse patterns in the strategy diagram can both be used as the second image to be trained. In this step, the adversarial patch image generated in step S30 is further added to the second image to be trained to generate a second adversarial sample image. The integrated network model is composed of a series of convolutional neural network (CNN) models with different structures in series; specifically, such as ResNet network, Vgg network, and Google network model connected in series. In this step, the integrated network model is used to replace the patch generation model in step S30; and the second image to be trained added with the adversarial patch image is input into the integrated network model to attack the integrated network model using the gradient descent method, and continuously iteratively update the adversarial patch image based on the complete training set until the attack converges for multiple CNN models in the integrated network model; then the final enhanced general patch image can be output. The complete training set can be understood as the sample set generated by fusing each sample image in the image sample set with the adversarial patch image.

[0057] Furthermore, the adversarial attack method based on the enhanced general patch further includes: fusing the adversarial patch image with the first image to be trained to obtain a third adversarial sample image, and performing robustness training on the patch generation model based on the original label corresponding to the first image to be trained and the third adversarial sample image. In this step, the third adversarial sample image is an adversarial sample that is successfully misclassified as the target category by the classification model. In this embodiment, based on the adversarial sample that is successfully misclassified as the target category, robustness retraining is performed using the original label. This step avoids the solidification of the gradient direction and expands the patch update space.

[0058] The above-mentioned adversarial attack method based on the enhanced general patch is specifically divided into two stages: the adversarial pre-training enhancement stage and the integrated attack enhancement stage. It should be noted that for the adversarial pre-training enhancement, the original patch image is randomly placed at a position far from the focus area of the image to be trained, rather than any position in the image to be trained, so as to avoid the influence of occlusion on the retraining.

[0059] The adversarial attack method based on the enhanced general patch in the above embodiments mines deeper semantic information of the target category, so that the enhanced general adversarial patch image can be more realistic and closer to the target category. Specifically, the improvement of this adversarial attack method focuses on the early input images of the training set; for early training, on the basis of optimizing the objective function, an adversarial training strategy is adopted; and the original labels that have been successfully classified are used to retrain the images, and then the patch generation model is updated; and further, the updated patch generation model is replaced with an integrated network model, and the patch image is iteratively updated. In short, the update of the adversarial patch is enhanced by continuously strengthening the attacked model; through this method, the patch training process can continuously learn the enhanced target class features, thereby narrowing the distance between the adversarial patch and the target image, and solving the problem of insufficient transfer performance of the existing general adversarial patch in the black box, that is, improving the black box transferability of the general adversarial patch.

[0060] Figure 2 As shown in the flowchart of the defense method based on the enhanced general patch according to an embodiment of the present invention, Figure 2 as shown, this defense method at least includes steps S001 to S005.

[0061] Step S001: Obtain the image to be recognized, input the image to be recognized into a network model to generate a heat map of the image to be recognized, and determine the key area of the image to be recognized based on the heat map.

[0062] In this step, the image to be recognized is an adversarial sample image added with a general adversarial patch image or a pure image directly obtained from an image sample set. And after obtaining the heat map of the image to be recognized, the key area mask of the image to be recognized is further obtained. This step is the model strong attention area positioning step, where the network model can be a CAM visualization model.

[0063] Specifically, since the CNN model often focuses on the local features of the input image, compared with global preprocessing, preprocessing only the attention area of the model has less impact on the original image; therefore, in this step, the CAM algorithm is used to obtain the heat map of the model for the image to be recognized. Specifically, it is necessary to deduce the importance of different region features for the model decision; first, is the sensitivity of the t-th class to the k-th channel of the feature map A output by the last convolutional layer k ; further, is used as the weight to perform a weighted linear combination of the last layer feature map and sent into the activation function to obtain the required heat map M t .

[0064]

[0065]

[0066] where Z is a normalization constant such that k is the serial number of the channel dimension of the feature map; i and j are the serial numbers of the width and height dimensions respectively; t is the target category, and y t is the gradient of the score of class t; is the pixel value at the position (i, j) in the feature map corresponding to the k-th channel.

[0067] Step S002: Perform multi-scale Gaussian blur processing on the image to be recognized to obtain multiple first low-resolution images.

[0068] In this step, the Gaussian blur sets under different Gaussian kernels selected can be denoted as G. To calculate the details of the image to be recognized, different-scale blur transformations are performed on the image to be recognized using different Gaussian blur kernels.

[0069] Step S003: Calculate the differences between the first low-resolution images at different scales, and obtain the local detail information of the image to be recognized at different scales based on the differences between the first low-resolution images at different scales.

[0070] In this step, the Gaussian difference image can be obtained by subtracting two layers of images obtained according to different blur kernels. Furthermore, the detail description of the image to be recognized at different scales can be obtained through Gaussian difference. The detail description can be further superimposed on the image to be recognized to increase the image's ability to express details, thereby obtaining the local detail information of the image to be recognized at different scales. In this step, specifically, the differences between the first low-resolution images (blurred images) at different scales can be calculated first, and then proportionally fused and multiplied by the magnification factor to obtain the global detail information of the image. Then, the global detail information is multiplied by the key region mask obtained in step S001 to obtain the local detail information of the image.

[0071] Step S004: Perform detail fusion on each piece of the local detail information and the image to be recognized to obtain an enhanced image.

[0072] This step is to further perform detail fusion on the local detail information and the image to be recognized, that is, add the obtained local detail information to the original image to be recognized and then perform normalization processing to obtain the enhanced image, and the enhanced image shows enhanced local details.

[0073] Steps S003 and S004 are to achieve local multi-scale detail fusion and enhancement. Suppose I is a clean image, and I C is a sample image without adding a universal adversarial patch to the clean image, and I adv is an adversarial sample image with a universal adversarial patch added to the clean image. The present invention adopts an indiscriminate preprocessing method, aiming to make IC The distortion of I is small, while I adv Compared with I, the distortion is large; however, the study found that the universal adversarial patch is relatively C Its feature distribution is denser and more irregular, and the regional image gradient is also higher, so I adv The Gaussian blur difference at different scales will be higher than I C Therefore, the image to be identified is fused based on multiple Gaussian blur differences of different scales and then subjected to certain enhancement processing. C The target’s contour details can be obtained, and for adversarial sample I adv Then the acquired detail information is highly dense and high-intensity. The acquired detail information is further added to the pure original image I for normalization to obtain the enhanced image I. en , enhanced image I en It is manifested as regional detail enhancement; that is, when I is I C When the enhanced image I en To enhance the contour details; and if I is I adv When the enhanced image I en This is manifested as a high degree of distortion in the adversarial patch area. This step can be expressed by the following calculation formula:

[0074] max(D(I c ,I)-D(I adv ,I))

[0075]

[0076] Among them, D is the measure of image distance; G is the set of Gaussian blurs under different selected Gaussian kernels; t is the magnification coefficient (t>1), which is used to indicate the magnification of detail information; m is the proportional coefficient, which is used to control the weight of the difference under adjacent Gaussian kernel blurs; norm is the normalization operation function.

[0077] Step S005: obtaining local detail information of the enhanced image, and performing attenuation processing on the obtained local detail information of the enhanced image to obtain a restored image corresponding to the image to be identified.

[0078] In this step, the enhanced image is similar to the image to be identified, and is also subjected to multi-scale Gaussian blur processing to further extract local detail information of the enhanced image at different scales. Specifically, obtaining the local detail information of the enhanced image includes performing multi-scale Gaussian blur processing on the enhanced image to obtain multiple second low-definition images; calculating the difference between the second low-definition images at different scales, and obtaining the local detail information of the enhanced image at different scales based on the difference between the second low-definition images at different scales. The above steps can also be understood as repeatedly executing steps S002 and S003. After the local detail information of the enhanced image is extracted, the local detail information in the enhanced image is further removed to restore the image.

[0079] Step S005 is used to reduce the multi-scale details of the enhanced image. Specifically, in order to reduce the impact of the preprocessing on the original restored image and to further distort the adversarial sample image, the enhanced image is reduced. C The detail information obtained under different scale sets is not much different (mainly manifested as contour information). Therefore, the enhanced image I after detail enhancement is en As input, take another scale set G de ={g1,g2,...,g n}, and perform similar detail reduction processing. This step reduces I C The distance between the image and the pure original image I, and the adversarial sample has a large distortion after the local detail enhancement and its own large scale sensitivity, further in the scale set G de The detail reduction processing under Gaussian blur will further aggravate the distortion of the adversarial sample. Specifically, this step can be expressed by the following calculation formula:

[0080]

[0081] Among them, I out To restore the image, I en To enhance the image, m ij is the fusion ratio coefficient, G de is the Gaussian fuzzy set under different selected Gaussian kernels; and That is, the scale set G de It is a Gaussian fuzzy set different from the scale set G.

[0082] After the above steps, the processed image to be identified is input into the defense model; if the image to be identified is a normal image, its category confidence remains unchanged; if the image to be identified is an adversarial sample image, its category label is converted from the wrong category to the correct category; then the entire defense stage is completed.

[0083] The above-listed defense methods based on data preprocessing focus on local processing of images by only paying attention to the key regions of the images, compared with various global data compressions. Compared with other methods based on image gradient discrimination, the defense method of the present invention can make full use of the subtle gradient differences between normal images and patches and amplify their impact on preprocessing. Specifically, regarding the feature centrality of patches, the distortion information (i.e., detail information) of patches obtained under Gaussian blur is greater than that of normal images; at the same time, the distortion information under Gaussian blur at multiple scales can be fused to further amplify this distance; using this characteristic, the obtained fused detail information is multiplied by an amplification factor to obtain enhanced detail information, which is then added to the original image, and finally the image is normalized to achieve detail enhancement of the key regions of normal images and high distortion at the adversarial sample patches; at the same time, in order to further reduce the impact of preprocessing on normal images, weakening processing is further performed under Gaussian blur at different scales. This defense method can be deployed in more lightweight production and living environments to effectively defend against adversarial attacks in the real environment.

[0084] Furthermore, in order to better demonstrate the advantages of the adversarial attack method based on the enhanced general patch and the defense method based on the enhanced general patch of the present invention, the following will be illustrated by a specific example.

[0085] In this example, the ImageNet-2012 dataset is used as the validation set, and 10 categories of images consisting of 10,000 images are selected to generate general adversarial patches. The size of the adversarial patches is 70x70, covering 10% of the image. During the training process of the general patch, the number of iterations is set to 8,100 times, and images of each category are randomly selected for each iteration. In addition, in order to further generate enhanced general adversarial patches, the adversarial pre-training strategy of the present invention is adopted for the first 100 input images of the first iteration, and the learning rate is set to 0.001 to re-train the patch generation model. For the attack performance, the performance of the generated enhanced general adversarial patches is evaluated under white-box and black-box settings; for black-box attacks, adversarial patches are generated based on ResNet-50, and then they are used to attack other models with different architectures and unknown parameters (VGG-16, Inception-V3, and ResNet-152), and their average target attack success rate is recorded. Figure 5 FIG. is a diagram showing the attack effects of ordinary adversarial patches and enhanced general patches of an embodiment of the present invention. The adversarial patches are generated by ResNet-50, and the target attack success rate refers to the average attack success rate tested in a black-box environment; as Figure 5 shown, the enhanced general patches are closer to the target category and have stronger black-box transferability.

[0086] The defense method of the present invention was further used to evaluate the defense effects against ordinary adversarial patches and enhanced general adversarial patches. Among them, the adversarial patches were generated using the ResNet-50 network structure, and the defense method was executed under the white-box setting; for each type of patch, 3000 successfully attacked adversarial samples (added with adversarial patches and enhanced patches) were randomly selected from the test data, and then the defense performances of the defense method of the present invention and other defense methods were compared. For the defense method of the present invention, three Gaussian kernels (5, 9, 19) were selected as the Gaussian blur set G to obtain details with the same proportional coefficient, the amplification coefficient t was set to 10, and then another 3 Gaussian kernels (3, 5, 11) were selected as the Gaussian blur set G de The enhanced image was weakened, and the defense model also used ResNet-50. The comparison of the defense performances of the defense method of the present invention and other defense methods is shown in the following table.

[0087] Table 1. Comparison of the defense performances of the defense method of the present invention and other defense methods

[0088]

[0089] Among them, Accuracy(%) is the classification accuracy that can be restored after defense processing for the successfully misclassified adversarial samples; Defense refers to different defense methods, and "Our" represents the defense method of the present invention. It can be seen from the above table that whether it is the commonly used adversarial patches at present or the enhanced patches generated by the adversarial attack method of the present invention, the defense method of the present invention has the highest defense accuracy (93.6% and 93.2%) compared with other defense methods, that is, the defense method of the present invention has high superiority in defending various adversarial patches.

[0090] Through the above embodiments, it can be found that the adversarial attack method based on the enhanced general patch proposed by the present invention can generate a general patch with stronger black-box transfer ability, and the enhanced general patch generated by using the adversarial attack method of the present invention can be better deployed in the real environment to effectively attack various artificial intelligence models, solving the problem of insufficient black-box transferability of the existing general adversarial patches; in addition, the defense method proposed by the present invention adopts a defense means based on data preprocessing, which can effectively defend against the enhanced general patch of the present invention, and compared with the existing detection and repair-based defense means, the processing time and cost of a single picture of the defense method of the present invention are lower than those of the existing defense technologies, and the defense effect is also better than other defense methods of the existing technologies.

[0091] Correspondingly, the present invention also discloses a countermeasure attack and defense system based on an enhanced general patch. The system includes a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method described in any of the foregoing embodiments.

[0092] In addition, the present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in any of the foregoing embodiments.

[0093] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link. A "machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet or an intranet.

[0094] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps. That is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0095] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An adversarial attack method based on an enhanced general patch, characterized in that, The method includes: Obtain an image sample set, and randomly select a first image sample from the image sample set as the first image to be trained; Obtain an original patch image, and perform image fusion on the original patch image and the first image to be trained to obtain a first adversarial sample image; Input the first adversarial sample image into a patch generation model, and update the patch image based on the gradient descent method to generate an adversarial patch image; Randomly select a second image sample from the image sample set as the second image to be trained, perform image fusion on the adversarial patch image and the second image to be trained to obtain a second adversarial sample image, input the second adversarial sample image into an integrated network model, and update the adversarial patch image based on the gradient descent method to generate a strengthened general patch image; Perform image fusion on the adversarial patch image and the first image to be trained to obtain a third adversarial sample image, and perform robustness training on the patch generation model based on the original label corresponding to the first image to be trained and the third adversarial sample image; wherein, the third adversarial sample image is an adversarial sample that is successfully misclassified as the target class by a classification model, and the first image to be trained and the second image to be trained are different images.

2. The adversarial attack method based on the enhanced general patch according to claim 1, wherein The integrated network model is a series model of multiple convolutional neural networks.

3. The adversarial attack method based on an enhanced general patch according to claim 1, wherein Performing image fusion on the original patch image and the first image to be trained to obtain a first adversarial sample image includes: Obtain the focus of the first image to be trained based on the CAM algorithm; Fuse the original patch image with the region of the first image to be trained that is far from the focus.

4. A defense method based on an enhanced general patch, characterized in that, The method includes: Obtain an image to be recognized, input the image to be recognized into a network model to generate a heat map of the image to be recognized, and determine the key region of the image to be recognized based on the heat map; Perform multi-scale Gaussian blur processing on the image to be recognized to obtain multiple first low-resolution images; Calculate the differences between the first low-resolution images at different scales, and obtain the local detail information of the image to be recognized at different scales based on the differences between the first low-resolution images at different scales; wherein, the global detail information of the first low-resolution images at different scales is obtained by fusing the differences between the first low-resolution images at different scales in proportion and multiplying by a magnification factor, and each global detail information is multiplied by the mask of the key region to obtain the local detail information at different scales; Perform detail fusion on each local detail information and the image to be recognized to obtain an enhanced image; Obtain the local detail information of the enhanced image, and perform weakening processing on the obtained local detail information of the enhanced image to obtain a restored image corresponding to the image to be recognized; Input the processed image to be recognized into a defense model; if the image to be recognized is a normal image, its class confidence remains unchanged; if the image to be recognized is an adversarial sample image, its class label is converted from the wrong class to the correct class.

5. The defense method based on the enhanced general patch according to claim 4, characterized in that The network model is a CAM visualization model.

6. The defense method based on an enhanced general patch according to claim 4, characterized in that, Obtaining the local detail information of the enhanced image includes: Perform multi-scale Gaussian blur processing on the enhanced image to obtain multiple second low-resolution images; Calculate the differences of the second-lowest-resolution images at different scales, and obtain the local detail information of the enhanced image at different scales based on the differences of the second-lowest-resolution images at different scales.

7. The defense method based on the enhanced general patch according to claim 6, characterized in that, The calculation formula for the restored image is: Among them, I out is the restored image, I en is the enhanced image, m ij is the fusion ratio coefficient, G de is the Gaussian blur set under different selected Gaussian kernels.

8. An adversarial attack and defense system based on an enhanced general patch, the system comprising a processor and a memory, characterized in that, Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for generating zoom robustness adversarial patch

    CN113689338A