Image classification black box adversarial sample generation method and device based on input transformation
By selecting a random circular area in the image for rotation perturbation, the problem of information loss and local transformation potential in the existing adversarial sample generation method is solved, which significantly improves the attack effect and migration of adversarial samples.
Patent Information
- Application Number
- CN202510098555.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
AI Technical Summary
The existing adversarial sample generation methods have information loss problems in the geometric transformation of images, and the potential of local transformation has not been fully explored, resulting in insufficient attack effect and migration.
By selecting random circular areas in the image for rotational perturbation, a diverse input image is generated to avoid global information loss and make full use of the potential of local area transformation.
It significantly improves the attack effect and migration of adversarial samples, so that the generated adversarial samples can more effectively deceive neural network models and have good cross-model migration capabilities.
Smart Images

Figure CN120047730A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of black-box adversarial attacks for image classification, and relates to a method and device for generating black-box adversarial samples for image classification based on input transformation. Background Art
[0002] As an important technology in the field of artificial intelligence, neural networks have developed rapidly in recent years and have been widely applied in many fields. Whether it is in image recognition, speech recognition, natural language processing, or cutting-edge technologies such as autonomous driving, neural networks have demonstrated excellent performance. By mimicking the working principle of biological neurons, neural networks can self-learn and process complex data, greatly promoting the progress of machine learning technology.
[0003] The core function of a neural network is to learn complex patterns through a large amount of data training and make inferences and classifications on new inputs. However, the complexity of neural networks also means that they may overfit to certain inputs or produce misleading learning, and these problems are often difficult to discover through conventional testing means. Adversarial samples can guide the model to make incorrect classifications by adding tiny perturbations to the input data, thereby revealing the vulnerability of the model in certain boundary situations.
[0004] For example, neural networks usually rely on specific features for classification, but the design of adversarial samples can break this feature dependence, making the model produce confident outputs for obviously incorrect inputs. This phenomenon indicates that neural networks may not truly understand the essential features of the data, but only capture some surface patterns. In this case, adversarial samples not only demonstrate the misjudgment of neural networks but also expose the weaknesses of the model when dealing with irregular or abnormal inputs in advance. By analyzing these adversarial samples, we can deeply understand why neural networks make incorrect decisions, thereby discovering potential defects of the model. This analysis process helps to optimize the design of neural networks, enabling them to better handle complex inputs in the real world and improve the generalization ability of the model.
[0005] Although certain progress has been made in the current research on the transferability of adversarial samples, there are still some problems and deficiencies. The main challenges are reflected in two aspects:
[0006] Information loss during attacks caused by methods such as rotation and distortion: Many current adversarial sample generation methods rely on geometric transformations of images, such as rotation, scaling, and distortion. These transformations can increase the diversity of adversarial samples to a certain extent, but they may also lead to the loss of original image information, thereby reducing the effectiveness of the attack. For example, rotating an image may cause misalignment of important feature information, resulting in a reduced attack success rate of the generated adversarial samples on the target model.
[0007] The potential of local transformation has not been fully explored: Currently, most generation methods based on input transformation perform global image transformation. However, research shows that local transformation may generate more diverse samples and thus more aggressive adversarial samples. This research direction of local image transformation has not been fully explored and is expected to become an effective means to improve the transferability of adversarial samples in the future. Summary of the Invention
[0008] Aiming at the deficiencies in the prior art, the object of the present invention is to provide a method and device for generating black-box adversarial samples for image classification based on input transformation. By randomly selecting circular regions in the image for rotational perturbation, it not only avoids the risk of losing global information but also fully explores the potential of local region transformation, significantly improving the attack effect and transferability of adversarial samples.
[0009] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions:
[0010] A method for generating black-box adversarial samples for image classification based on input transformation, comprising the following steps:
[0011] Step 1: Input the original input image into the image transformation module, perform a series of diverse random transformation operations on the original image to generate diverse input images;
[0012] Step 2, determine a surrogate model in the neural network structure module, and the surrogate model performs forward propagation and backward propagation to complete the classification and gradient calculation of the input image;
[0013] Step 3, select an objective function: perform a weighted sum of the cross-entropy loss and LPIPS to obtain the final objective function;
[0014] Step 4, use the objective function selected in Step 3, add the perturbation to be optimized to the diverse input images obtained in Step 1, input them into the surrogate model determined in Step 2, optimize the perturbation to be optimized until the set number of iterations is reached, and add the final perturbation to the original input image to obtain an adversarial sample that can mislead the neural network classification.
[0015] The present invention also includes the following technical features:
[0016] Specifically, the said Step 1 includes:
[0017] Step 1.1: Generate a circular region mask: Randomly generate a circular region on the original input image, and generate two masks for this circular region, namely a positive mask and a negative mask;
[0018] Step 1.2: Assigning mask region pixels: Assign the pixels within the circular region of the positive mask as 1 and the pixels outside the circular region as 0; assign the pixels within the circular region of the negative mask as 0 and the pixels outside the circular region as 1;
[0019] Step 1.3: Mask image generation: Perform a bitwise AND operation between the positive mask with assigned pixels and the original input image to generate a positive mask image; perform a bitwise AND operation between the negative mask with assigned pixels and the original input image to generate a negative mask image;
[0020] Step 1.4: Image rotation transformation: Rotate the generated positive mask image at a random angle to increase the diversity of the image;
[0021] Step 1.5: Merging images: Add the rotated positive mask image and the negative mask image to generate a diverse input image.
[0022] Specifically, in the said Step 2, the proxy models include: AlexNet model, ResNet model, Transformer model.
[0023] Specifically, in the said Step 3, the final objective function is:
[0024]
[0025] where L is the objective function, C is the total number of categories, y i is the true label of the i-th category, taking values of 0 or 1, and p i is the probability of the i-th category predicted by the model; α is the weight factor, and LPIPS is an index used to measure the perceptual similarity between images.
[0026] Specifically, the said Step 4 includes:
[0027] Step 4.1: First, add the diverse input image obtained in Step 1 to the current perturbation, and input it into the determined proxy model for forward propagation; calculate the value of the objective function based on the model output and the target output; then, through the backpropagation algorithm, calculate the gradients at the positions of each pixel in each image;
[0028] Step 4.2: Perform a per-pixel averaging on the gradients of all images to calculate the current gradient magnitude of each pixel; perform a weighted summation of the current gradient and the gradients obtained in previous iterations to form a stable gradient update strategy;
[0029] Step 4.3: Take the final gradient sign after the weighted summation of each pixel to determine the update direction of each pixel;
[0030] Step 4.4: Multiply the update direction by the step size to generate a preliminary perturbation, and clip the perturbation within the perturbation norm;
[0031] Step 4.5: Repeat the process of steps 4.1 to 4.4 until the set number of iterations is reached. Add the final perturbation to the original image to obtain an adversarial sample that can mislead the neural network classification.
[0032] An image classification black-box adversarial sample generation device based on input transformation, comprising:
[0033] An image transformation module, configured to perform a series of diverse random transformation operations on the original image to generate diverse input images;
[0034] A neural network structure module, configured to determine a surrogate model in the neural network structure module. The surrogate model performs forward propagation and backward propagation to complete the classification and gradient calculation of the input image; the surrogate models include: AlexNet model, ResNet model, Transformer model;
[0035] A target function determination module, configured to perform weighted summation of the cross-entropy loss and LPIPS to obtain a final target function; the final target function is:
[0036]
[0037] where L is the target function, C is the total number of categories, y i is the true label of the i-th category, taking values of 0 or 1, and p i is the probability of the i-th category predicted by the model; α is a weight factor, and LPIPS is an index used to measure the perceptual similarity between images;
[0038] A gradient optimization and adversarial sample generation module, configured to use the selected target function to add a perturbation to be optimized to the diverse input images, input them into the determined surrogate model, optimize the perturbation to be optimized until the set number of iterations is reached, and add the final perturbation to the original input image to obtain an adversarial sample that can mislead the neural network classification.
[0039] Specifically, the image transformation module can be used to implement:
[0040] Generate a circular region mask: Randomly generate a circular region on the original input image, and generate two masks for this circular region, namely a positive mask and a negative mask;
[0041] Assign pixel values to the masked region: Assign the pixel points inside the circular region of the positive mask to 1, and the pixel points outside the circular region to 0; assign the pixel points inside the circular region of the negative mask to 0, and the pixel points outside the circular region to 1;
[0042] Mask image generation: Perform a bitwise AND operation on the positive mask after pixel assignment and the original input image to generate a positive mask image; perform a bitwise AND operation on the negative mask after pixel assignment and the original input image to generate a negative mask image;
[0043] Image rotation transformation: Rotate the generated positive mask image at a random angle to increase image diversity;
[0044] Merge images: Add the rotated positive mask image and the negative mask image to generate a diverse input image.
[0045] Specifically, the gradient optimization and adversarial sample generation module can be used to achieve:
[0046] Add the diverse input image and the current perturbation, and input it into a determined surrogate model for forward propagation; Calculate the value of the objective function based on the model output and the target output; Then, through the backpropagation algorithm, calculate the gradients at the positions of each pixel in each image;
[0047] Average the gradients of all images pixel by pixel to calculate the current gradient magnitude of each pixel; Perform a weighted sum of the current gradient and the gradients obtained in previous iterations to form a stable gradient update strategy;
[0048] Take the final gradient sign after weighted summation of each pixel to determine the update direction of each pixel;
[0049] Multiply the update direction by the step size to generate a preliminary perturbation, and clip the perturbation within the perturbation norm;
[0050] Repeat until the set number of iterations is reached, add the final perturbation to the original image to obtain an adversarial sample that can mislead the neural network classification.
[0051] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the image classification black-box adversarial sample generation method based on input transformation.
[0052] A computer-readable storage medium is used to store program instructions, and these program instructions can be executed by a processor to implement the steps of the image classification black-box adversarial sample generation method based on input transformation.
[0053] Compared with the prior art, the present invention has the following technical effects:
[0054] Through the effective combination of input transformation, gradient optimization, neural network structure, and objective function, the present invention proposes an efficient method for generating black-box adversarial samples. Compared with existing adversarial sample generation techniques, the significant advantages of this method are as follows: diverse input transformations greatly enhance the diversity and robustness of adversarial samples, making the attack more transferable; a stable gradient optimization strategy ensures the smoothness of perturbation updates, improves the attack efficiency of adversarial samples, and has good transfer ability; a reasonable objective function design combines classification accuracy and visual similarity, ensuring that the generated adversarial samples maintain visual similarity with the original image while misleading the model, making the attack more concealed. This method can not only successfully deceive deep learning models in black-box scenarios but also effectively improve the cross-model transferability and attack effect of adversarial samples, thus having significant advantages in practical applications. Description of the Drawings
[0055] Figure 1 It is the overall algorithm flow chart.
[0056] Figure 2 It is a schematic diagram of the random rotation method of a random circular area.
[0057] Figure 3 It is a test result diagram of adversarial samples generated on the ResNet-18 model.
[0058] Figure 4 It is a test result diagram of adversarial samples generated on the ResNet-101 model.
[0059] Figure 5 It is a test result diagram of adversarial samples generated on the Inception-v3 model. Detailed Implementation Manner
[0060] The generation of adversarial samples is achieved by using the gradient of the target model to find a small perturbation of the input image, thereby causing the model to misclassify it. Common white-box attack methods such as the Fast Gradient Sign Method (FGSM) and the Projected Gradient Descent (PGD) use similar principles. Here, FGSM is taken as an example to illustrate how to generate adversarial samples.
[0061] FGSM is a simple and efficient white-box attack method that uses the gradient of the loss function of the target model to generate adversarial samples. Its basic idea is that given an input sample x and the corresponding true label y, by applying a small perturbation along the direction of the loss function gradient on the sample, the model misclassifies the sample. The strategy for FGSM to generate adversarial samples is to add a perturbation along the direction of the loss function gradient to the input. The calculation formula for the perturbation is:
[0062]
[0063] where: δ is the added adversarial perturbation; ∈ is a coefficient that controls the size of the perturbation, usually a very small positive number;
[0064] represents the sign of the gradient of the loss function, that is, the positive or negative sign of the gradient of each pixel. Adding the perturbation δ to the original input sample x gives a new adversarial sample x adv :
[0065]
[0066] At this time, the adversarial sample x adv is almost indistinguishable from the original sample x to the human eye, but to the model, it has changed significantly and may lead to misclassification.
[0067] However, when actually attacking a target model, an attacker usually cannot obtain the specific parameters of the model and thus cannot directly perform backpropagation on the model to calculate the gradient. However, research has shown that even if an adversarial sample is generated on a certain known model, it may still have a misleading effect on different models. This phenomenon is called the transferability of adversarial samples. Transferability enables an attacker to successfully attack a target model with adversarial samples generated by other models even without knowing the internal information of the target model.
[0068] The present invention provides a method for generating black-box adversarial samples for image classification based on input transformation, as Figure 1As shown, the core idea of this method is to generate adversarial samples through four modules: an image transformation module, a gradient optimization module, a neural network structure, and an objective function, aiming to mislead the classification results of deep neural network models. Specifically, first, the input image is processed by the input image transformation module to perform diverse random transformations, generating a series of diverse transformed images; then, the gradient optimization module iteratively calculates the update direction and magnitude of the perturbation, generates the perturbation, and tests it in the neural network through forward propagation; in the neural network structure module, a specific image classification model, such as AlexNet, ResNet, or Transformer model, is selected to perform forward propagation and classification on the input image; finally, in the objective function module, the cross-entropy loss and LPIPS (Perceptual Image Patch Similarity) are combined to optimize the model output result. By gradually adding perturbations, the model produces incorrect classification outputs, thereby generating adversarial samples, enabling the generated adversarial samples to effectively deceive the model and improving the transferability of the adversarial samples. The main innovation is that in the input transformation module, a transformation method based on random rotation of a random circular region is innovated, and in the objective function module, a weighted sum of the traditional cross-entropy loss and LPIPS is innovated as a new objective function. This method enhances the diversity of adversarial samples through the masking and rotation operations of the input transformation module. The gradient optimization module continuously updates the perturbation iteratively, improving the attack effect, and finally generates adversarial samples with strong transfer ability.
[0069] The following are specific embodiments of the present invention. It should be noted that the present invention is not limited to the following specific embodiments, and all equivalent transformations based on the technical solutions of this application fall within the protection scope of the present invention.
[0070] Embodiment 1:
[0071] As Figure 1 shown, this embodiment provides a method for generating black-box adversarial samples for image classification based on input transformation, including the following steps:
[0072] Step 1: Input the original input image into the image transformation module, perform a series of diverse random transformation operations on the original image to generate diverse input images, aiming to increase the diversity and complexity of adversarial samples. As Figure 2 shown, the specific steps are as follows:
[0073] Step 1.1: Generate a circular region mask: Randomly generate a circular region on the original input image, and generate two masks for this circular region, namely a positive mask and a negative mask;
[0074] Step 1.2: Assign pixel values to the masked area: Assign the pixel points within the circular area of the positive mask as 1, and the pixel points outside the circular area as 0; assign the pixel points within the circular area of the negative mask as 0, and the pixel points outside the circular area as 1. This enables the mask to highlight different parts of the image respectively.
[0075] Step 1.3: Generate the masked images: Perform a bitwise AND operation between the positive mask with assigned pixel values and the original input image to generate the positive masked image; perform a bitwise AND operation between the negative mask with assigned pixel values and the original input image to generate the negative masked image; through such processing, certain parts of the image are masked, and other parts are highlighted.
[0076] Step 1.4: Image rotation transformation: Perform a rotation with a random angle (e.g., -180° and 180° in Example 1) on the generated positive masked image to further increase the diversity of the image; through the rotation operation, more visual changes are produced in the transformed image.
[0077] Step 1.5: Merge the images: Add the rotated positive masked image and the negative masked image to generate a diverse input image. These transformed images will be used as inputs to provide more samples for gradient optimization.
[0078] Step 2: Determine the surrogate model in the neural network structure module to generate adversarial samples on this known model, and the generated adversarial samples are also aggressive towards other unknown models; in the neural network structure module, select a specific surrogate model for forward propagation and backward propagation to complete the classification and gradient calculation of the input image.
[0079] The surrogate models include:
[0080] AlexNet model: Adopt the ReLU activation function, introduce the Dropout mechanism and the max-pooling strategy, which enhance the training speed and feature extraction ability of the network and are suitable for large-scale image classification tasks. The basic architecture of AlexNet consists of 8 layers of neural networks, including 5 convolutional layers and 3 fully connected layers. This network first adopted the ReLU activation function, replacing the traditional Sigmoid function, which significantly improved the training speed of the model on large-scale datasets; at the same time, the Dropout mechanism was introduced to effectively prevent overfitting, and the training time was significantly shortened through GPU parallel computing; in addition, AlexNet adopted an overlapping max-pooling strategy to improve the feature extraction ability of the network and enhance the spatial feature expression effect.
[0081] ResNet Model: ResNet introduces skip connections between convolutional layers, enabling the input of each layer to be directly passed to subsequent layers, thus avoiding the common vanishing gradient problem in deep networks; it allows for the construction of very deep networks while maintaining high accuracy, especially suitable for complex image classification tasks. This innovative design allows ResNet to maintain high training efficiency when building extremely deep neural networks and effectively improve the accuracy and robustness of classification tasks, especially performing well in complex tasks such as image classification.
[0082] Transformer Model: Through the multi-head self-attention mechanism and feed-forward neural network, it can capture the global and local features of the input image, and has powerful sequence processing and feature extraction capabilities. The Transformer architecture consists of an encoder and a decoder. The encoder is responsible for extracting the global features of the input sequence, and the decoder generates the output based on these features; each encoder and decoder unit is composed of a multi-head self-attention mechanism and a feed-forward neural network. The self-attention mechanism can capture the long-range dependencies in the input image, and the feed-forward neural network is responsible for further extracting local features. This modular structure makes Transformer have powerful feature extraction and sequence processing capabilities in image classification tasks.
[0083] Step 3, Select the objective function: Weighted sum of cross-entropy loss and LPIPS, with the weight factor α set to 0.9, thus obtaining the optimized objective function. This method can generate more effective adversarial samples. The final objective function is:
[0084]
[0085] where L is the objective function, C is the total number of classes, y i is the true label of the i-th class, taking values of 0 or 1 (one-hot encoding, only the y i of the true class is 1, and the rest are 0), p i is the probability of the i-th class predicted by the model; α is the weight factor, and LPIPS is an index used to measure the perceptual similarity between images.
[0086] In the supervised learning of image classification, cross-entropy is often used as one of the objective functions for optimization algorithms. The model minimizes the cross-entropy loss to make the predicted probabilities of its outputs close to the probability distribution of the true labels. The smaller the loss value, the closer the model's predictions are to the true labels. Compared with other loss functions, cross-entropy combines the logarithmic function and can effectively handle the problem of numerical overflow. Especially in deep learning, when the predicted values are close to 0 or 1, it can avoid the problems of vanishing gradients or exploding gradients. Cross-entropy loss is often used in combination with the Softmax layer. Softmax converts the original output of the model into a probability distribution, and cross-entropy can directly calculate the loss within the probability space, which is particularly useful for multi-classification problems. In adversarial attacks, adversarial attack algorithms need to maximize the cross-entropy loss to make the predictions of the model on adversarial samples deviate greatly from the true labels. In this way, the generated adversarial samples can effectively deceive the model and make it make incorrect classifications for the input. For multi-classification problems, the common cross-entropy loss after the Softmax output layer has the formula:
[0087]
[0088] where C is the total number of classes. y i is the true label of the i-th class, taking values of 0 or 1 (one-hot encoding, where only the y i of the true class is 1 and the rest are 0). p i is the probability of the i-th class predicted by the model, and the output is converted into a probability distribution through Softmax.
[0089] LPIPS (Learned Perceptual Image Patch Similarity) is a metric used to measure the perceptual similarity between images. Different from traditional similarity measurement methods (such as L2 distance, PSNR, etc.), LPIPS takes into account the sensitivity of the human visual system to image differences, so it can more accurately reflect the visual similarity of images. Research shows that the transferability of adversarial samples is positively correlated with their LPIPS values, that is, the higher the LPIPS value of an adversarial sample, the stronger its transferability. As shown in the following table:
[0090] Table 1: Transferability and LPIPS values under various methods
[0091] TIM DIM SIM SSA Admix Transferability 57.4 77.6 79.3 80.6 83.6 LPIPS 0.25 0.43 0.48 0.54 0.73
[0092] Therefore, when generating adversarial samples, the transferability of the samples can be enhanced by increasing the LPIPS value. In the design of the objective function, cross-entropy and LPIPS can be weighted and summed, and the weight factor α is set to 0.9 to obtain the optimized objective function. This method can generate more effective adversarial samples.
[0093] The final objective function is as follows: where L is the objective function, C is the total number of classes, and y i is the true label of the i-th class, taking values of 0 or 1 (one-hot encoding, where only the y of the true class i = 1 and the rest are 0), and p i is the probability of the i-th class predicted by the model; α is the weight factor, and LPIPS is an index used to measure the perceptual similarity between images.
[0094] Step 4: Using the objective function selected in Step 3, add the perturbation to be optimized to the diverse input images obtained in Step 1, input them into the surrogate model determined in Step 2, optimize the perturbation to be optimized until the set number of iterations is reached, and add the final perturbation to the original input image to obtain an adversarial sample that can mislead the neural network classification. The specific steps are as follows:
[0095] Step 4.1: First, add the diverse input images obtained in Step 1 to the current perturbation, and input them into the determined surrogate model for forward propagation; calculate the value of the objective function based on the model output and the target output; then, through the backpropagation algorithm, calculate the gradients at the positions of each pixel in each image, reflecting the sensitivity of the current perturbation at these pixels.
[0096] Step 4.2: Average the gradients of all images pixel by pixel to calculate the current gradient magnitude of each pixel; this average represents the sensitivity of this pixel to the perturbation. To improve the optimization effect, sum the current gradient and the gradients obtained in previous iterations with weights, and use a hyperparameter to balance the current and historical gradients to form a stable gradient update strategy.
[0097] Step 4.3: Take the sign of the final gradient after weighted summation of each pixel to determine the update direction of each pixel, with three update directions: +, -, 0; this direction indicates how the perturbation is adjusted to maximize the value of the objective function.
[0098] Step 4.4: Multiply the update direction by the step size (the step size is usually 1 / 8 of the maximum perturbation norm) to generate a preliminary perturbation, and clip the perturbation within the perturbation norm to ensure that the generated perturbation is within the specified limits and avoid the adversarial sample being too prominent.
[0099] Step 4.5: Repeat the process from Step 4.1 to Step 4.4 until the set number of iterations is reached, and add the final perturbation to the original image to obtain an adversarial sample that can mislead the neural network classification.
[0100] The termination condition of the iterative process is to reach the preset number of iterations. To improve the transferability of the generated adversarial samples, the number of images generated by the input transformation is set to 10. This number can balance the consumption of computing resources while ensuring the attack effect, and avoid overloading the CPU and memory. In addition, an appropriate number of images can improve the universality of adversarial samples on different models, enabling the attack to have better cross-model transfer ability.
[0101] The present invention also provides an image classification black-box adversarial sample generation device based on input transformation, including:
[0102] An image transformation module, configured to perform a series of diverse random transformation operations on the original image to generate diverse input images;
[0103] A neural network structure module, configured to determine a surrogate model in the neural network structure module, and the surrogate model performs forward propagation and backward propagation to complete the classification and gradient calculation of the input image; the surrogate models include: AlexNet model, ResNet model, Transformer model;
[0104] An objective function determination module, configured to perform weighted summation of the cross-entropy loss and LPIPS to obtain the final objective function; the final objective function is:
[0105]
[0106] where L is the objective function, C is the total number of categories, y i is the true label of the i-th category, taking values of 0 or 1, and p i is the probability of the i-th category predicted by the model; α is a weight factor, and LPIPS is an index used to measure the perceptual similarity between images;
[0107] A gradient optimization and adversarial sample generation module, configured to use the selected objective function, add the perturbation to be optimized to the diverse input images, input them into the determined surrogate model, optimize the perturbation to be optimized to reach the set number of iterations, and add the final perturbation to the original input image to obtain an adversarial sample that can mislead the neural network classification.
[0108] Specifically, the image transformation module can be used to implement:
[0109] Generate a circular region mask: randomly generate a circular region on the original input image, and generate two masks for this circular region, namely a positive mask and a negative mask;
[0110] Assign pixel values to the masked region: assign the pixel points inside the circular region of the positive mask to 1, and the pixel points outside the circular region to 0; assign the pixel points inside the circular region of the negative mask to 0, and the pixel points outside the circular region to 1;
[0111] Mask image generation: Perform a bitwise AND operation on the positive mask after pixel assignment and the original input image to generate a positive mask image; perform a bitwise AND operation on the negative mask after pixel assignment and the original input image to generate a negative mask image.
[0112] Image rotation transformation: Rotate the generated positive mask image at a random angle to increase image diversity.
[0113] Merge images: Add the rotated positive mask image and the negative mask image to generate a diverse input image.
[0114] Specifically, the gradient optimization and adversarial sample generation module can be used to implement:
[0115] Add the diverse input image and the current perturbation, and input it into a certain surrogate model for forward propagation; calculate the value of the objective function based on the model output and the target output; then, through the backpropagation algorithm, calculate the gradient at the position of each pixel in each image.
[0116] Perform per-pixel averaging on the gradients of all images to calculate the current gradient magnitude of each pixel; perform a weighted sum of the current gradient and the gradients obtained in previous iterations to form a stable gradient update strategy.
[0117] Take the sign of the final gradient after weighted summation of each pixel to determine the update direction of each pixel.
[0118] Multiply the update direction by the step size to generate a preliminary perturbation, and clip the perturbation within the perturbation norm.
[0119] Repeat until the set number of iterations is reached, add the final perturbation to the original image to obtain an adversarial sample that can mislead the neural network classification.
[0120] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-described embodiment of the method for generating black-box adversarial samples for image classification based on input transformation are implemented.
[0121] In one embodiment, a computer-readable storage medium is provided. This computer-readable storage medium is used to store program instructions, and these program instructions can be executed by a processor to implement the steps of the above-described embodiment of the method for generating black-box adversarial samples for image classification based on input transformation.
[0122] The above computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, etc.
[0123] In each embodiment, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a volatile or non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the essence of this solution, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0124] In the present invention, the algorithm pseudocode is summarized as follows:
[0125] Algorithm 1: Black-box Adversarial Sample Generation Algorithm
[0126] Input: Neural network parameters θ, neural network model f, objective function L, each input image x, corresponding label y, total number of training rounds N, number of generated input transformation images m, perturbation range ε, perturbation update step size α; weighted ratio μ of historical gradient and current gradient, random circular region random rotation transformation T(·)
[0127] Output: Adversarial sample x with good transferability adv
[0128] 1: α = ε / T; g 0 = 0
[0129] 2: for t = 1 to N do
[0130] 3: for i = 1 to m do
[0131] 4: Obtain the total gradient: / / Calculate the gradient after randomly transforming the image
[0132] 5: Obtain the average gradient: g = sum / m
[0133] 6: Perform weighted summation with the historical gradient:
[0134] 7: Obtain the gradient sign: sign = g t+1.sign()
[0135] 8: Update the current adversarial example:
[0136] 9: return
[0137] To verify the effectiveness of the method in this case, the following further explains with experimental data:
[0138] Experimental setup: Dataset: Randomly selected 1000 images from the ILSVRC 2012 validation set. These images belong to 1000 categories and were correctly classified by the adopted model. All images were pre-adjusted to 299×299×3.
[0139] Test models: Generate adversarial examples on ResNet-18, ResNet-101, and Inception-v3 respectively, and test the attack success rate on ResNet-18, ResNet-101, Inception-v3, ResNeXt-50, DenseNet-121, MobileNet, ViT, and Swin.
[0140] Baseline methods: Compare the method in this paper with attacks based on input transformation, namely DIM, TIM, DEM, Admix, and SSA.
[0141] Hyperparameter settings: The perturbation size ∈ is set to 16 / 255, the number of iterations T = 10, the step size α = ∈ / T = 1.6 / 255, the decay factor μ = 1, and the LPIPS weighting factor is 0.9.
[0142] Experimental results: As shown in Figure 3 , 4 , 5, where the abscissa represents each model and the ordinate represents the attack success rate of various attack methods on the model. Compared with the baseline methods, the method in this paper always outperforms the best baseline in terms of performance on all eight models with different architectures. An attack success rate higher than the best baseline by at least 2.8% is achieved on all models.
[0143] The preferred embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, and these simple modifications all fall within the protection scope of the present invention.
[0144] In addition, it should be noted that, among the various specific technical features described in the above specific embodiments, they can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the present invention will not separately describe various possible combination methods.
[0145] In addition, any combinations can also be made among various different embodiments of the present invention, as long as they do not violate the idea of the present invention, and they should equally be regarded as the content disclosed by the present invention.
Claims
1. A method for generating black-box adversarial samples for image classification based on input transformation, characterized in that: The following steps are involved: Step 1: Input the original input image into the image transformation module, and perform a series of diverse random transformation operations on the original image to generate diverse input images; Step 2, determining a proxy model in the neural network structure module, and performing forward propagation and back propagation on the proxy model to complete the classification and gradient calculation of the input image; Step 3, select the objective function: perform weighted summation of the cross entropy loss and LPIPS to obtain the final objective function; In step 4, using the objective function selected in step 3, the diversified input images obtained in step 1 are added with the perturbation to be optimized, and the perturbation to be optimized is passed into the proxy model determined in step 2. The perturbation to be optimized is optimized for the set number of iterations, and the final perturbation is added to the original input image to obtain an adversarial sample that can mislead the neural network classification.
2. The method for generating black-box adversarial samples for image classification based on input transformation according to claim 1, characterized in that: The step 1 comprises: Step 1.1: Generate circular area mask: randomly generate a circular area on the original input image, and generate two masks for the circular area, namely, a positive mask and a negative mask; Step 1.2: Assign values to pixels in the mask area: Assign values to pixels in the positive mask circular area as 1, and values to pixels outside the circular area as 0; Assign values to pixels in the negative mask circular area as 0, and values to pixels outside the circular area as 1; Step 1.3: Mask image generation: perform bitwise AND operation on the positive mask after pixel value assignment and the original input image to generate a positive mask image; perform bitwise AND operation on the negative mask after pixel value assignment and the original input image to generate a negative mask image; Step 1.4: Image rotation transformation: Rotate the generated positive mask image at random angles to increase the diversity of the image; Step 1.5: Merge images: Add the rotated positive mask image and the negative mask image to generate a diversified input image.
3. The method for generating black-box adversarial samples for image classification based on input transformation according to claim 1, characterized in that: In step 2, the proxy models include: AlexNet model, ResNet model, and Transformer model.
4. The method for generating black-box adversarial samples for image classification based on input transformation according to claim 1, characterized in that: In step 3, the final objective function is: Among them, L is the objective function, C is the total number of categories, and y i is the true label of the i-th category, which takes the value of 0 or 1, p i is the probability of the i-th class predicted by the model; α is the weight factor, and LPIPS is a metric used to measure the perceptual similarity between images.
5. The method for generating black-box adversarial samples for image classification based on input transformation according to claim 1, characterized in that: The step 4 comprises: Step 4.1: First, add the diversified input image obtained in step 1 to the current disturbance and input it into the determined proxy model for forward propagation; calculate the value of the objective function based on the model output and the target output; then calculate the gradient of each pixel position in each image through the back propagation algorithm; Step 4.2: Average the gradients of all images pixel by pixel to calculate the current gradient size of each pixel; perform weighted summation of the current gradient and the gradient obtained in the historical iteration to form a stable gradient update strategy; Step 4.3: Take the final gradient sign after weighted summation of each pixel point and determine the update direction of each pixel point; Step 4.4: Multiply the update direction by the step size to generate a preliminary perturbation, and clip the perturbation to within the perturbation norm; Step 4.5: Repeat the process from step 4.1 to step 4.4 until the set number of iterations is reached, and add the final perturbation to the original image to obtain an adversarial sample that can mislead the neural network classification.
6. A device for generating black-box adversarial samples for image classification based on input transformation, characterized in that: include: The image transformation module is used to perform a series of diverse random transformation operations on the original image to generate diverse input images; The neural network structure module is used to determine the proxy model in the neural network structure module. The proxy model performs forward propagation and back propagation to complete the classification and gradient calculation of the input image. The proxy models include: AlexNet model, ResNet model, and Transformer model. The objective function determination module is used to perform weighted summation of the cross entropy loss and LPIPS to obtain the final objective function; the final objective function is: Among them, L is the objective function, C is the total number of categories, and y i is the true label of the i-th category, which takes the value of 0 or 1, p i is the probability of the i-th class predicted by the model; α is the weight factor, and LPIPS is a metric used to measure the perceptual similarity between images; The gradient optimization and adversarial sample generation module is used to use the selected objective function to add the perturbation to be optimized to the diverse input images, pass it into the determined proxy model, optimize the perturbation to be optimized for a set number of iterations, and add the final perturbation to the original input image to obtain an adversarial sample that can mislead the neural network classification.
7. The device for generating black-box adversarial samples for image classification based on input transformation according to claim 6, characterized in that: The image transformation module can be used to achieve: Generate circular area mask: randomly generate a circular area on the original input image, and generate two masks for the circular area, namely, positive mask and negative mask; Assigning values to pixels in the masked area: Assigning values of 1 to pixels in the positive masked circular area and 0 to pixels outside the circular area; Assigning values of 0 to pixels in the negative masked circular area and 1 to pixels outside the circular area; Mask image generation: perform bitwise AND operation on the positive mask after pixel value assignment and the original input image to generate a positive mask image; perform bitwise AND operation on the negative mask after pixel value assignment and the original input image to generate a negative mask image; Image rotation transformation: Rotate the generated positive mask image at random angles to increase the diversity of the image; Merge images: Add the rotated positive mask image and the negative mask image to generate a diversified input image.
8. The device for generating black-box adversarial samples for image classification based on input transformation according to claim 6, characterized in that: The gradient optimization and adversarial sample generation module can be used to achieve: Add the diverse input images to the current disturbance and input the determined proxy model for forward propagation; calculate the value of the objective function based on the model output and the target output; Then, the gradient of each pixel position in each image is calculated through the back propagation algorithm; The gradients of all images are averaged pixel by pixel to calculate the current gradient size of each pixel; the current gradient is weighted summed with the gradient obtained in the historical iteration to form a stable gradient update strategy; Take the final gradient sign after weighted summation of each pixel point to determine the update direction of each pixel point; Multiply the update direction by the step size to generate a preliminary perturbation, and clip the perturbation to within the perturbation norm; Repeat until the set number of iterations is reached, and add the final perturbation to the original image to obtain an adversarial sample that can mislead the neural network classification.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for generating black-box adversarial samples for image classification based on input transformation according to any one of claims 1 to 5 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program instructions, which can be executed by a processor to implement the steps of the image classification black-box adversarial sample generation method based on input transformation described in any one of claims 1 to 5.