Regional independence target attack confrontation sample generation method, device and equipment
By dividing the perturbation region into independent sub-regions in the target attack and combining it with hybrid loss function optimization, the problems of large computational overhead and limited robustness in the existing technology are solved, and a more efficient and stable target attack effect is achieved.
Patent Information
- Application Number
- CN202511032185.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-07
AI Technical Summary
Existing target attack methods suffer from high computational overhead and limited robustness in terms of improving portability. In particular, lightweight methods have a lower success rate and poor process generality in challenging scenarios.
By dividing the perturbation region into two independent masks, two independent sub-adversarial examples are generated. Gradient updates are performed by combining translation invariance and input transformation diversity. A hybrid loss function is used to optimize the perturbation region and generate new adversarial examples to mislead the deep learning model.
It significantly improves the portability and robustness of target attacks, ensures consistency in target output across different regions, reduces interference between regions, and enhances the stability and efficiency of attacks.
Smart Images

Figure CN120913032A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of deep learning security, in particular to a regional independence target attack adversarial sample generation method, device, equipment and medium. BACKGROUND
[0002] Deep neural networks have triggered revolutionary changes in the fields of computer vision, natural language processing, etc., but their inherent vulnerability makes them extremely susceptible to adversarial attacks. Such attacks can mislead the model to output incorrect prediction results by adding tiny perturbations to the input data, and have been proven to have cross-model transferability, thus realizing black-box attacks. Although non-target attacks have shown strong transferability in black-box scenarios, target attacks have significantly increased in complexity and challenge due to the need to precisely control the model output to the target class. The successful implementation of target attacks poses a serious threat to the security of deep learning models, so it is of great significance to study how to improve the transferability of target attacks.
[0003] Current target attack methods are mainly divided into resource-intensive and lightweight types. Resource-intensive methods rely on auxiliary networks or complex generation processes to improve transferability. For example, Wang et al. proposed a generative adversarial framework that directly generates adversarial samples with high transferability by integrating label information and feature distribution. Although this method has significant effects, the training process of its generator and discriminator requires a large amount of computational resources, resulting in high actual deployment costs. Similarly, the feature destruction attack method introduces an auxiliary classifier to capture the relationship between the current features and the target class and guide the generation direction of the adversarial perturbation. Although this method significantly improves transferability, it still has a large computational overhead due to the need to train an additional auxiliary model. In addition, some resource-intensive methods use complex input transformations to improve transferability, such as the ODI method, which generates adversarial perturbations on the surface of three-dimensional objects, uses rendered two-dimensional images to mislead the model, or the Admix method, which mixes input images with additional images of other classes to enhance transferability. However, these methods usually require additional training data, model components, or complex preprocessing steps, limiting their scalability and practicality in large-scale application scenarios.
[0004] The core of the lightweight attack method is to improve the transferability of cross-model by enhancing input diversity and stabilizing gradient update. For example, the diversified input method reduces overfitting to a specific model and improves transferability by applying random scaling and padding operations to the input image; the shift-invariant method ensures that the perturbation remains effective under small spatial shift input by applying a convolution smoothing kernel to the gradient; the self-universal method explores the internal relationship between universality and transferability, and achieves better transferability by aligning the local and global features of the image. In terms of stabilizing gradient update, the momentum iteration method introduces a momentum term into the gradient update to smooth the oscillation and prevent perturbation overfitting. The momentum iteration method is expected to further stabilize the update trajectory and help escape from local suboptimal solutions by aggregating the average gradient of the sampled data points in the previous iteration gradient direction. These lightweight methods are computationally efficient and are very suitable for large-scale attack scenarios. However, although the lightweight method significantly reduces the computational overhead, its attack effect and robustness are still limited, especially in challenging target attack scenarios, where the success rate will decrease significantly. In addition, such methods often lack flexibility, for example, the self-universal method needs to manually adjust the feature extraction layer according to the source model, resulting in poor process universality and difficulty in adapting to different model architectures.
[0005] In summary, the prior art has significant deficiencies in improving the transferability of target attacks. Resource-intensive methods will cause significant computational overhead in large-scale application scenarios, while the attack effect and robustness of lightweight methods are limited. Therefore, there is an urgent need for a solution that can maintain lightweight design while significantly improving the transferability of target attacks to address the above technical problems.
[0006] Therefore, the present application is proposed. SUMMARY
[0007] The present application aims to provide a region independence target attack adversarial sample generation method, device, equipment and medium to solve the problems of large computational overhead, limited transferability and robustness in improving the transferability of target attacks of existing methods.
[0008] To solve the above technical problems, the present application realizes the following technical solutions: A region independence target attack adversarial sample generation method, comprising: S1, obtaining an original image input; S2, randomly and dynamically generating a first mask and a second mask based on the original image, the second mask being a complementary region of the first mask, and the two masks not overlapping in space; S3, applying the first mask and the second mask as two independent perturbation regions during perturbation optimization to the original image to generate two independent sub-adversarial samples; S4, inputting two sub-adversarial samples into a convolutional neural network for gradient calculation and optimization, and updating the gradient in combination with translation invariance and input transformation diversity to update the perturbation region; S5, adding the perturbation region after optimization to the original image to generate a new adversarial sample for misleading the deep learning model to output a specified error category.
[0009] Preferably, the first mask is a binary square mask, the starting coordinates of which are a random position on the original image, the side length of which is a preset proportion range of the size of the original image, and the random discard method is used to set part of the region of the first mask to 0.
[0010] The first mask is represented as: ; Wherein, is the first mask.
[0011] Preferably, the second mask is represented as: ; Wherein, is the second mask.
[0012] Preferably, two independent sub-adversarial samples are represented as: ; ; Wherein, , respectively represent two independent sub-adversarial samples; is the original image; , respectively are the first mask and the second mask; is the global perturbation; represents the diversity input transformation.
[0013] Preferably, when the gradient calculation and optimization are performed on the two sub-adversarial samples, a consistency target loss function is used to synchronize the output distribution of the two sub-adversarial samples and improve the confidence of the target category; Then, the calculation formula of the gradient at the i-th iteration is: ; Wherein, represents the gradient representation of the current iteration; is the gradient of the global perturbation ; represents a white-box model; represents the target category, i.e. the error category specified by the attacker; represents the target loss function; , respectively represent two independent sub-adversarial samples.
[0014] Preferably, the formula for updating the gradient in combination with translation invariance and input transformation diversity is: ; wherein, represents the gradient representation of the current iteration; represents the updated gradient representation; is the L1 norm; represents the momentum term, which is used to stabilize the gradient and determine the optimal direction of disturbance update at each step; T represents the convolution kernel in the translation invariance TI.
[0015] Preferably, when updating the disturbance region, it also includes clipping the disturbance region, and the formulas are respectively: ; ; wherein, is the disturbance region of the current iteration; is the updated disturbance region; represents the step size; represents the updated gradient representation; is the sign function, which is used to generate a disturbance vector with a fixed direction but a standardized amplitude; is the original image; is the disturbance budget; represents the clipping operation on the original image.
[0016] Preferably, it also includes dynamically optimizing the output of different disturbance regions by using a hybrid loss function; wherein the hybrid loss function is represented as: ; wherein, is the normalized probability of the target class ; The term is used to increase the probability of the target class , while implicitly suppressing the probability of the non-target class ; represents the logit value of the target class ; The gradient calculation formula of the hybrid loss function is: ; wherein, represents the logit value of the non-target class; is the non-target class normalized probability.
[0017] The application further provides a region independence target attack adversarial sample generation device, comprising: An acquisition unit is configured to acquire an input original image. A mask generation unit is configured to randomly and dynamically generate a first mask and a second mask based on the original image, the second mask being a complementary region of the first mask, and the two masks not overlapping in space. A sub-adversarial sample generation unit is configured to apply the first mask and the second mask as two independent perturbation regions during perturbation optimization to the original image respectively to generate two independent sub-adversarial samples. A perturbation optimization unit is configured to input the two sub-adversarial samples into a convolutional neural network to perform gradient calculation and optimization, and perform gradient update in combination with translation invariance and input transformation diversity to update the perturbation region. A new adversarial sample generation unit is configured to add the perturbation region after optimization to the original image to generate a new adversarial sample for misleading a deep learning model to output a specified error category.
[0018] The application further provides a region independence target attack adversarial sample generation device, comprising a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement the region independence target attack adversarial sample generation method.
[0019] The application further provides a computer readable storage medium, wherein the computer readable storage medium stores computer readable instructions, and the computer readable instructions are executed by a processor of a device where the computer readable storage medium is located to implement the region independence target attack adversarial sample generation method.
[0020] Compared with the prior art, the application has the following beneficial effects: The application significantly improves the transferability and robustness of target attacks by enhancing the independence between different perturbation regions and combining mixed loss function optimization.
[0021] The perturbation region division method adopted by the application ensures that the generated sub-adversarial samples have higher misleading ability, and the region division is more consistent with the model decision logic, thereby improving the stability of the attack.
[0022] The application enhances the independence of adversarial attacks between different segmentation regions of an input image, thereby reducing the interference and dependence between regions; at the same time, the consistency of these regions in target output is ensured, so that the application realizes more stable and efficient target attacks in a cross-model scenario, and ensures that different regions can independently mislead a deep neural network to output a target category.
[0023] The mixed loss function of the present application accelerates the convergence of the perturbation generation process and enhances the robustness of the attack by dynamically adjusting the logit value of the target class and implicitly suppressing the logit value of the non-target class.
[0024] The present application solves the problem of limited transferability of lightweight target adversarial attacks in the prior art by proposing a region independence target attack (RITA) framework, providing an important technical improvement for the field of deep learning security. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.
[0026] Figure 1 A schematic diagram of a region independence target attack adversarial sample generation method provided for embodiment one.
[0027] Figure 2 A schematic diagram of the framework of a region independence target attack adversarial sample generation method provided for embodiment one.
[0028] Figure 3 A success rate (%) result statistical diagram of using four source models to perform target attack on a black-box target model provided for embodiment one.
[0029] Figure 4 A target attack success rate (%) result statistical diagram of using an integrated model to perform target attack on a black-box target model provided for embodiment one. Figure 5 An ASR (%) improvement effect diagram of the method model RITA of the present application combined with Admix, SI and ODI provided for embodiment one. Figure 6 A result diagram of using RN-50 as a source model to reduce the influence of perturbation region on the attack success rate (%) of transferability on Inc-v3, DN-121 and VGG-16 respectively provided for embodiment one. Figure 7 A result diagram of using RN-50 as a source model to show the influence of different partition sizes on the attack success rate (%) provided for embodiment one. Figure 8 An example diagram of successful target adversarial samples generated by the method of the present application for attacking Google Cloud Vision provided for embodiment one. Figure 9 A schematic diagram of a region independence target attack adversarial sample generation device provided for the embodiment two.
[0030] The present application will be further described below in conjunction with the drawings and specific embodiments. DETAILED DESCRIPTION
[0031] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0032] Embodiment one The embodiment one of the present application provides a region independence target attack adversarial sample generation method, which can be realized by a region independence target attack adversarial sample generation device (hereinafter referred to as an adversarial attack device), in particular, by one or more processors in the adversarial attack device.
[0033] In the present embodiment, the adversarial attack device can be an electronic device equipped with a processor, the processor is provided with a computer program of the region independence target attack adversarial sample generation method and the computer program can be executed, for example, a computer, a smart phone, a smart tablet, a workstation, etc., which are not limited here.
[0034] The embodiment of the application proposes a region-independence targeted attack (RITA) method for generating an adversarial sample model. Specifically, the disturbance is divided into two random independent regions, and the disturbance of these regions is applied to the clean original image sample respectively to generate two independent sub-adversarial samples. However, due to the fact that the disturbance information of different regions cannot be shared, the output prediction of the sub-adversarial sample is difficult to keep consistent. To solve this problem, the application imposes a consistency constraint on the output of the sub-adversarial sample during optimization, and then synchronously optimizes the two regions to improve the target class confidence. Each region of the model can mislead the deep neural network to output the target class independently. This new type of targeted attack method significantly improves the transferability of targeted attacks by enhancing the region independence of adversarial disturbance.
[0035] In addition, since RITA optimizes the disturbance of different regions independently and does not share information between regions, these regions naturally exhibit different attack saturation levels. However, the commonly used logit loss applies uniform gradient updates to all regions, which cannot adapt to this imbalance. To solve this limitation, a hybrid loss function with a simplified process is proposed, that is, by dynamically optimizing the target logit value of different regions while implicitly suppressing the probability of non-target classes, not only the convergence is accelerated and the robustness is enhanced, but also the target class is ensured to be approached faster and more consistently.
[0036] As shown in Figures 1-2 A region-independence targeted attack adversarial sample generation method includes steps S1 to S5.
[0037] S1, obtaining an input original image.
[0038] The embodiment can use Python as the main programming language and use the pytorch deep learning framework for implementation. Of course, other suitable programming languages and deep learning frameworks can also be selected, which are not limited herein.
[0039] The original image obtained in this step is the original data used for generating adversarial samples. These original images can come from various data sets (such as CIFAR-10, ImageNet, etc. for image classification), and collected images in actual application scenarios (such as security monitoring images, mobile phone shooting images, etc.).
[0040] S2, randomly and dynamically generating a first mask and a second mask based on the original image, the second mask being a complementary region of the first mask, and the two not overlapping in space.
[0041] In this step, the spatial region of the original image is divided into two independent parts by generating two spatially complementary and non-overlapping masks, preparing for subsequent independent perturbation in different regions, increasing the flexibility and pertinence of adversarial sample generation, and using regional independence to improve attack effect.
[0042] Specifically, first, define two masks (the first mask) and (the second mask) for selective operation. In this embodiment, a simple square selection strategy is adopted, and in each iteration, the side length and starting coordinates are randomly generated.
[0043] Specifically, the first mask is a binary square mask, the starting coordinates (such as the top-left corner coordinates) are a random position in the original image, and the side length is a preset proportion range of the size of the original image, such as , times, such as [0.1, 0.9], and the random dropout (Random Dropout) method is used to set part of the region of the first mask to 0.
[0044] The first mask is represented as: ; For example, 20% of the probability sets part of the pixel value in the square mask to 0 to increase the randomness of the mask; or the probability can determine which pixels are 0 according to the randomly generated probability.
[0045] In order to minimize the dependence between different regions, and are not overlapped in space. Therefore, can be considered as the complementary region of , and the second mask is represented as: ; wherein is the second mask.
[0046] Of course, other region division methods can also be used to generate two masks, which are not limited here.
[0047] S3, the first mask and the second mask are used as two independent perturbation regions during perturbation optimization, and are applied to the original image respectively to generate two independent sub-adversarial samples.
[0048] In this step, using and , two independent sub-adversarial samples can be generated, represented as: ; ; wherein, , respectively represent two independent sub-adversarial samples; is the original image; , respectively represent the first mask and the second mask; is the global perturbation; represents the diversity input transformation, and represents random scaling or padding.
[0049] Before the iteration process starts, the global perturbation δ can be initialized as a zero matrix with the same size as the original image, or as a random noise matrix uniformly sampled in a certain range interval.
[0050] S4, input the two sub-adversarial samples into the convolutional neural network for gradient calculation and optimization, and combine translation invariance and input transformation diversity for gradient update to update the perturbation region.
[0051] In this step, when the two sub-adversarial samples are calculated and optimized, the target loss function is used to simultaneously approximate the output distribution of the two sub-adversarial samples and improve the confidence of the target class.
[0052] Then, at the i-th iteration, the calculation formula of the gradient is: ; wherein, represents the gradient representation of the current iteration; is the gradient of the global perturbation ; represents a white-box model; represents the target class, i.e., the error class specified by the attacker; represents a target loss function, which can be any loss function, such as mean square error loss, cross-entropy loss, KL divergence loss, etc. , respectively represent two independent sub-adversarial samples.
[0053] Due to the lack of diverse input transformation, the gradient is easy to become fixed, resulting in overfitting to the model . To solve this problem, translation invariance TI and input transformation diversity DI are combined for gradient update to enhance the diversity of inputs and avoid local optimum. The formula is: ; wherein, represents the gradient representation of the current iteration; denotes the updated gradient representation; is the L1 norm; denotes the momentum term, which is used to stabilize the gradient and determine the optimal direction of the perturbation update at each step; T denotes the convolution kernel in translation invariance (TI), which can be a Gaussian kernel with size 7x7, for example.
[0054] When updating the perturbation region, clipping of the perturbation region is also included, and the formulas are as follows: ; ; wherein, is the perturbation region of the current iteration; is the updated perturbation region; denotes the step size, which can be set to , wherein N is the total number of iterations (e.g., N = 10); denotes the updated gradient representation; is the sign function, which is used to generate a perturbation vector with a fixed direction but a standardized amplitude; is the original image; is the perturbation budget, which can be set to , for example, 16 / 255; denotes the clipping operation on the original image.
[0055] In terms of model optimization, the prior art generally uses a Logit loss function for optimization, which can effectively alleviate the gradient vanishing problem compared to the cross-entropy loss. Given the output vector z of a neural network (wherein, denotes the logit value of the target class, denotes the logit value of the non-target class), the loss function is optimized by minimizing the value of the target logit, and the formula is as follows: ; The gradient of the Logit loss function is given by the following formula: ; This optimization process only affects , and equal gradient updates are used for all regions. In addition, the loss function does not actively suppress the logit value of the non-target class ( ). However, classification is essentially a relative process, and its confidence depends on the difference between the logit values of the target class and the non-target class. Although increasing can increase the final confidence, the loss function does not actively suppress the non-target class.
[0056] Since the model of the application updates the perturbations of different regions independently and does not share the perturbation information between regions, the attack effects of each region are significantly different. In this case, using the traditional Logit loss function will make it more difficult for the attack optimization to converge. To solve these limitations, a hybrid loss function is proposed, which combines the advantages of probability-based optimization and Logit-based optimization. The core idea is to introduce the softmax probability in the cross-entropy loss (CE) into the Logit loss function, dynamically enhance the logit value of the target class ( ) and suppress the logit value of the non-target class ( ), thereby improving the class discrimination.
[0057] The hybrid loss function is expressed as: ; wherein, is the normalized probability of the target class ; The term is used to improve the probability of the target class , while implicitly suppressing the probability of the non-target class ; represents the logit value of the target class , The term is used to ensure stable updating of the logit value of the optimization output.
[0058] The gradient calculation formula of the hybrid loss function is: ; wherein, represents the logit value of the non-target class; is the normalized probability of the non-target class .
[0059] The hybrid loss function can provide stronger gradients when the predicted probability is small, which promotes the value of to rise quickly; when increases, the gradient converges smoothly to -1, which not only alleviates numerical instability, but also avoids the problem of gradient disappearance. This feature is highly consistent with the optimization process of the RITA model of the application, i.e., for regions with low confidence, the hybrid loss function provides stronger gradients during updating, thereby accelerating the convergence of these regions. For high-confidence regions, the gradient neither disappears nor stabilizes near -1, ensuring that all regions can obtain stable and efficient optimization. In addition, the logit value of the non-target class is dynamically updated based on the probability, further expanding the separation degree from the target class. This implicit adjustment does not require direct modification of the logit value of the non-target class, significantly improving the computational efficiency.
[0060] S5, adding the completed optimized perturbation region to the original image to generate a new adversarial sample for misleading the deep learning model to output the specified error category.
[0061] After all iterations are completed, the final perturbation is added to the original image to generate a new adversarial sample for misleading the deep learning model to output the specified error category.
[0062] The core purpose of generating a target attack adversarial sample is to induce the deep learning model to misclassify the input image as the target category specified by the attacker through a carefully designed small perturbation, rather than the original true category. This process reveals the vulnerability of the model in a specific scenario and serves multiple fields such as security research, model improvement, and defense strategy development. Target attacks can be used in face recognition systems, autonomous driving perception systems, medical image diagnosis, etc., such as misclassifying a face image as a specific identity (such as unlocking a phone or bypassing identity verification), using an iterative gradient attack combined with a target loss function to generate a perturbation that makes the model output the specified identity.
[0063] The present application solves the challenge of limited transferability of lightweight target adversarial attacks from a novel and insightful perspective. Specifically, we reduce the interference and dependence between regions by enhancing the independence of adversarial attacks between different partition regions of the input image. By ensuring the consistency of these regions in the target output, we achieve more stable and efficient target attacks in cross-model scenarios. Then, we also propose a hybrid loss function that can accelerate the convergence of regional perturbations and improve the robustness of attacks. Through extensive experiments, it is shown that while maintaining a lightweight design, the present method exhibits better transferability compared to existing baseline schemes, and can be combined with most existing methods to further improve the attack success rate.
[0064] In a preferred embodiment, in order to verify the effect of the method of the present application, RITA is applied to the Google Cloud Vision API to evaluate the feasibility of target attacks in real-world scenarios. Among the 100 samples, 45% were misclassified as semantically related categories (target attack successful), and 33% achieved non-target attacks. These results show that RITA can effectively exploit the vulnerabilities of the Google Cloud Vision API and pose a significant threat to deployed commercial systems.
[0065] In a preferred embodiment, based on the DI-TI-MI, DI-TI-EMI, DI-TI-MI-SU and the model of the present application combined with the logit loss function (RITA-logit), the model of the present application combined with the hybrid loss function (RITA-Hybird) target attack adversarial sample generation method, using four original models (RN-50, DN-121, VGG-16, Inc-v3) to respectively attack other nine kinds of black box target models, the target attack success rate result chart is shown in the following figure: Figure 3 It can be seen from Figure 3 that the success rate of target attack by the method of the present application is significantly higher than that of the other three methods, and the method combined with the hybrid loss function RITA-Hybird is slightly better than RITA-logit.
[0066] Referring to the following figure Figure 4 , using the integrated method to attack six kinds of black box target models, it can be seen from the target attack success rate that the success rate of target attack by the method of the present application is significantly higher than that of the other three methods, and the method combined with the hybrid loss function RITA-Hybird is slightly better than RITA-logit.
[0067] Referring to the following figure Figure 5 , RN-50 as the original model, the RITA proposed by the present application is combined with Admix, SI and ODI method, and the target model attack comparison result of the method combined with Admix, SI and ODI method without combination can be known. The improvement effect of the target attack method combined with RITA is obviously improved.
[0068] Referring to the following figure Figure 6 , Figure 7 , using RN-50 as the original model, based on DI-TI-MI, DI-TI-EMI, DI-TI-MI-SU and the model of the present application combined with the logit loss function (RITA-logit), the model of the present application combined with the hybrid loss function (RITA-Hybird) target attack adversarial sample generation method, on three target models (Inc-v3, DN-121 and VGG-16), the influence result of attack success rate with different partition size (100% area, 81% area, 64% area, 49% area) can be known. The closer the partition area is, the higher the attack success rate is, but the target attack adversarial sample generation method combined with the model of the present application is less affected by the partition size of the disturbance area than the other three methods, which shows that the robustness of the model of the present application is good. And from Figure 7 it can be known that when the preset proportion range of the first mask side length , , at this time, the partition area size is close, and the attack success rate is the highest.
[0069] Referring to the accompanying drawings Figure 8 This is a successful target adversarial sample generated using the RITA method of the present application. The accompanying Figure 8 A number of successful cases are demonstrated, including different types of input images and their corresponding adversarial samples. These adversarial samples are highly similar in appearance to the original images, but successfully mislead the model output target categories in the classification prediction of deep neural networks. For example, an image that was correctly classified as a "hickory tree" is misclassified as a "goose" after adding perturbations, indicating the high efficiency of the RITA method in practical applications.
[0070] In summary, compared with the prior art, the present application has the following beneficial effects: The present application solves the problem of limited transferability of lightweight target adversarial attacks in the prior art by proposing a regional independence target attack (RITA) framework. In the specific implementation process, by enhancing the independence between perturbation regions, applying uniformity constraints, combining mixed loss function optimization and other means, the stability and cross-model transferability of the target attack are significantly improved. Experimental results show that, while maintaining lightweight design, the present method exhibits better transferability compared to existing baseline schemes, and can be combined with most existing methods to further improve the attack success rate. In particular, the present method achieves a target attack success rate of 45% and a non-target attack success rate of 33% in practical application tests on Google Cloud Vision API, proving its threat in the real world.
[0071] In this embodiment, DI (the Diverse Inputs): Diverse input method, by applying random scaling and padding operations to input images, reduces overfitting to specific models and improves transferability.
[0072] TI (the Translation-Invariant): Translation-invariant method, applies a convolution smoothing kernel to the gradient to ensure that the perturbation remains effective under small spatial shifts in the input.
[0073] SI (the Scale-Invariant): Scale-invariant method, uses the scale-invariant properties of deep neural networks to scale the original image multiple times to avoid overfitting to the source model.
[0074] MI (the Momentum Iterative): Momentum iterative method, introduces a momentum term into the gradient update to smooth oscillations and prevent perturbation overfitting.
[0075] SU (Self-Universality): Self-universality method, explores the internal relationship between universality and transferability, realizes self-universality by aligning local and global features of images, and further improves transferability.
[0076] ODI (Object-based Diverse Input): Object-based diverse input method, generates adversarial perturbations on the surface of three-dimensional objects, uses rendered two-dimensional images to mislead the model, and improves the success rate of cross-model attacks through physical-level perturbations.
[0077] EMI (Expected Momentum Iteration): Expected momentum iteration method, aggregates the average gradient of the sampled data points in the previous iteration gradient direction, stabilizes the update trajectory and helps to escape from local suboptimal solution.
[0078] Admix: Improve transferability by mixing input images with additional images of other categories.
[0079] For example, DI-TI-MI is a target attack method that combines diverse input, translation invariance and momentum iteration, and DI-TI-EMI is a target attack method that combines diverse input, translation invariance and expected momentum iteration.
[0080] RN-50: ResNet-50, a classic deep learning model in the ResNet (Residual Network) family, with a network depth of 50 layers (including residual blocks and other structures). By introducing residual connection (Residual Connection), it effectively solves the problem of gradient disappearance and degradation during deep network training, and can extract multi-level and multi-scale features of images. It is widely used in computer vision tasks such as image classification and object detection, and is often used as one of the benchmark models for evaluating the effectiveness of adversarial sample attacks.
[0081] DN-121: DenseNet-121, a model in the DenseNet (Dense Connection Network) series, with a network depth of 121 layers. Its characteristic is to use dense connection, that is, each layer is connected with all previous layers in the channel dimension, promoting the reuse and propagation of features, making the network parameters more efficient. It performs well in image classification tasks and can be used as a target for adversarial sample attacks or a basic model for generating adversarial samples to test model robustness.
[0082] VGG-16: is a classic convolutional neural network model proposed by the Visual Geometry Group at the University of Oxford, with a network depth of 16 layers (including 13 convolutional layers and 3 fully connected layers). It mainly uses a simple convolutional layer stack structure, uses small convolutional kernels (3x3) and max-pooling operations, extracts image features through repeated convolution-pooling structures, and has an important position in image classification, target recognition and other fields. It is often used for adversarial sample related research and as an attack model to measure the effectiveness of attack methods.
[0083] Inc-v3: Inception-v3, the third generation of Inception series models, developed by the Google team, performs well in image classification and target detection tasks. It uses Inception modules that use different sizes of convolutional kernels (such as 1x1, 3x3, 5x5) and pooling operations in parallel, then concatenate the outputs, which can extract image features at different scales, improving the diversity and efficiency of feature extraction. It is often used as a target model for adversarial sample attacks to test the effectiveness of attack methods.
[0084] RN-34: ResNet-34, a model in the ResNet family with a depth of 34 layers, also solves deep network problems based on residual connections, with a relatively simpler structure than ResNet-50 and less computational complexity. It can be applied to some classification tasks that do not require high computational resources or have relatively simple image features, and can also be used as an attack model in adversarial sample attack research to test the effectiveness of different attack methods on this model.
[0085] RN-152: ResNet-152, a model in the ResNet family with a depth of 152 layers, has a large depth and strong feature extraction capability, and can learn more complex and detailed image features, achieving high accuracy in image classification tasks. However, it has a large computational complexity and requires a lot of resources, and is often used as a high-end image classification task or model robustness testing model. In adversarial sample attack research, it is used to test the attack effect of attack methods on deep and strong feature expression models.
[0086] VGG-19: is a model in the VGG series with a depth of 19 layers (including 16 convolutional layers and 3 fully connected layers), similar in structure to VGG-16, but with more convolutional layers. It uses a repeated structure of 3x3 small convolutional kernels and 2x2 max-pooling to further deepen the network to extract more rich image features, which may have higher accuracy than VGG-16 in image classification tasks (under certain data sets). It is often used as a target model for adversarial sample attacks to study the effectiveness of attack methods on deeper VGG models.
[0087] AlexNet: One of the classic models that marked a significant breakthrough in image classification using deep learning, proposed by Alex Krizhevsky et al. The network consists of 5 convolutional layers and 3 fully connected layers, using ReLU activation functions (instead of traditional Sigmoid, etc.), Dropout to prevent overfitting, data augmentation techniques, and other techniques to significantly improve classification accuracy on datasets such as ImageNet, opening up a wave of deep learning in the field of computer vision. It is often used as a basic model for adversarial sample-related research to test the effectiveness of attack methods on early classic deep models.
[0088] Mob-v2: MobileNet-v2, the second generation of MobileNet models, focuses on efficient image classification models for mobile devices and resource-constrained scenarios. It uses depthwise separable convolution to decompose convolution operations into depthwise convolution and pointwise convolution, significantly reducing the number of parameters and computational complexity. It also introduces linear bottleneck and inverted residual structures to improve feature extraction capabilities and model efficiency. It is widely used in mobile device image recognition, adversarial sample attack testing, and other scenarios, and can be used as an attacked model to test the effectiveness of attack methods on lightweight models.
[0089] Mob-v3: MobileNet-v3, the third generation of MobileNet models, further optimizes model structure and efficiency, combining neural architecture search (NAS) technology to automatically search for the optimal network structure configuration. It also introduces h-swish activation functions and other improvements to improve classification accuracy while maintaining efficient computation. It is suitable for mobile image classification, edge computing scenarios, and adversarial sample attack testing. As an attacked model, it can be used to study the effectiveness of attack methods on the latest lightweight and efficient models.
[0090] Embodiment Two As shown in Figure 9 The second embodiment of the present application also provides a regional independence target attack adversarial sample generation device, comprising: An acquisition unit for acquiring an input original image; A mask generation unit for randomly and dynamically generating a first mask and a second mask based on the original image, the second mask being the complementary region of the first mask, and the two not overlapping in space; The sub-adversarial sample generation unit is configured to apply the first mask and the second mask as two independent perturbation regions during perturbation optimization to the original image respectively to generate two independent sub-adversarial samples. The perturbation optimization unit is configured to input the two sub-adversarial samples into a convolutional neural network to perform gradient calculation and optimization, and perform gradient update in combination with translation invariance and input transformation diversity to update the perturbation region. The new adversarial sample generation unit is configured to add the perturbation region after optimization to the original image to generate a new adversarial sample for misleading a deep learning model to output a specified error category.
[0091] Embodiment three The third embodiment of the present application further provides a region independence target attack adversarial sample generation device, which comprises a memory and a processor, the memory stores a computer program, and the computer program can be executed by the processor to implement the region independence target attack adversarial sample generation method as described above.
[0092] Embodiment four The fourth embodiment of the present application further provides a computer readable storage medium, which stores computer readable instructions, and the computer readable instructions are executed by a processor of a device where the computer readable storage medium is located to implement the region independence target attack adversarial sample generation method as described above.
[0093] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus and method embodiments described above are only illustrative, for example, the flowchart in the drawings shows the possible implementation architecture, function and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementation manners, the functions annotated in the blocks can also occur in different order from that annotated in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for executing the specified function or action, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0094] In addition, each functional module in various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0095] If the functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes. It should be noted that in this paper, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "includes a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0096] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "an" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0097] It should be understood that the term "and / or" used herein is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in this paper generally represents that the front and rear associated objects are a "or" relationship.
[0098] Depending on the context, the word "if" as used herein can be interpreted to mean "when" or "while" or "in response to determining" or "in response to detecting." Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "when it is determined" or "in response to determining" or "when [a stated condition or event] is detected" or "in response to detecting [a stated condition or event]."
[0099] The terms "first" and "second" mentioned in the embodiments are only to distinguish similar objects, and do not represent a specific order of the objects. Understandably, the terms "first" and "second" can be interchanged in a specific order or sequence as permitted. It should be understood that the objects distinguished by the terms "first" and "second" can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0100] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for generating an adversarial sample of a region independence target attack, characterized in that, The method comprises the following steps: S1, obtaining an input original image; S2, randomly and dynamically generating a first mask and a second mask based on the original image, the second mask being a complementary region of the first mask, and the two masks not overlapping in space; S3, applying the first mask and the second mask as two independent perturbation regions during perturbation optimization to the original image respectively to generate two independent sub-adversarial samples; S4, inputting the two sub-adversarial samples into a convolutional neural network for gradient calculation and optimization, and combining translation invariance and input transformation diversity for gradient update to update the perturbation region; S5, adding the optimized perturbation region to the original image to generate a new adversarial sample for misleading a deep learning model to output a specified wrong category.
2. The method of claim 1, wherein The first mask is a binary square mask, the starting coordinates of which are random positions on the original image, the side length of which is in a preset proportion range of the size of the original image, and the random discard method is used to set part of the region of the first mask to 0. The first mask is represented as: ; wherein, is the first mask.
3. The method of claim 2, wherein The second mask is represented as: ; wherein, is the second mask.
4. The method of claim 1, wherein The two independent sub-adversarial samples are represented as: ; ; wherein, , respectively represent two independent sub-adversarial samples; is the original image; , are respectively a first mask, a second mask; is a global perturbation; represents a diversity input transformation.
5. The method of claim 1, wherein When performing gradient calculation and optimization on the two sub-adversarial samples, a consistency target loss function is used to synchronize the output distribution of the two sub-adversarial samples and improve the confidence of the target category. Then, the calculation formula of the gradient at the i-th iteration is: ; wherein, denotes the gradient representation of the current iteration; is the gradient of the global perturbation ; denotes a white-box model; denotes the target class, i.e. the erroneous class specified by the attacker; denotes the target loss function; , denote two independent sub-adversarial samples, respectively.
6. The method of claim 5, wherein The formula for updating the gradient combining translation invariance and input transformation diversity is: ; wherein, represents the gradient representation of the current iteration; represents the updated gradient representation; is the LI norm; represents the momentum term, which is used to stabilize the gradient and determine the optimal direction of the perturbation update at each step; T represents the convolution kernel in translational invariance, TI.
7. The method of claim 1, wherein When updating the perturbation region, clipping the perturbation region is also included, and the formulas are respectively: ; ; wherein, is the perturbed region for the current iteration; is the updated perturbed region; denotes the step size; denotes the updated gradient representation; is a sign function that generates a perturbation vector with a fixed direction but normalized magnitude; is the original image; is the perturbation budget; denotes a cropping operation on the original image.
8. The method of claim 1, wherein Further, the method comprises dynamically optimizing the outputs of different perturbation regions by using a hybrid loss function; wherein the hybrid loss function is represented as: ; wherein, is the normalized probability of the target class ; is used to increase the probability of the target class , while implicitly suppressing the probability of non-target classes ; represents the logit value of the target class ; The gradient calculation formula of the mixed loss function is: ; in, This represents the logit value for non-target categories; Non-target category The normalized probability.
9. An area-independent target attack adversarial sample generation apparatus characterized by comprising: The method comprises the following steps: An acquisition unit is configured to acquire an input original image; A mask generation unit is configured to randomly and dynamically generate a first mask and a second mask based on the original image, the second mask being a complementary region of the first mask, and the two masks not overlapping in space; A sub-adversarial sample generation unit is configured to apply the first mask and the second mask as two independent perturbation regions during perturbation optimization to the original image respectively to generate two independent sub-adversarial samples; A perturbation optimization unit is configured to input the two sub-adversarial samples into a convolutional neural network for gradient calculation and optimization, and combine translation invariance and input transformation diversity for gradient update to update the perturbation region; A new adversarial sample generation unit is configured to add the optimized perturbation region to the original image to generate a new adversarial sample for misleading a deep learning model to output a specified wrong category.
10. A region independence target attack counter sample generation device, comprising: The method comprises a processor and a memory, and the memory stores a computer program which can be executed by the processor to implement the region independence target attack adversarial sample generation method of any one of claims 1-8.