Confrontation sample generation method and device, equipment, storage medium and program product
By using improved momentum iterative fast gradient symbolic method and iterative processing technology in deep neural network models, target adversarial samples are generated, which solves the problems of low success rate of adversarial sample attacks and weak migration capabilities in the existing technology, and improves the model's resistance and quality of adversarial samples.
Patent Information
- Application Number
- CN202510111523.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, the success rate of adversarial sample attacks is low and the migration ability is weak, making it difficult to effectively improve the ability of deep neural network models to resist attacks.
By acquiring the input image and processing it to obtain the target perturbation area, the input image is input as an initial adversarial sample into the trained target model for iterative processing, using the improved momentum iterative fast gradient symbol method to accumulate the loss gradient and update the momentum gradient until the second loss result converges, and a target adversarial sample is generated.
The attack success rate and migration ability of the adversarial samples are improved, and the resistance to deep neural network models is enhanced, while ensuring the image quality and perturbation concealment of the adversarial samples are ensured.
Smart Images

Figure CN120219868A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to an adversarial sample generation method, apparatus, device, storage medium, and program product. Background Art
[0002] With the development of artificial intelligence technology, deep neural network models have achieved extremely remarkable achievements in many fields such as image recognition and image classification. However, deep neural network models are vulnerable to adversarial samples. An adversarial sample refers to a situation where, by making a small and intentional perturbation to the input data, the deep neural network model produces an incorrect result during classification or recognition. Therefore, generating threatening and highly transferable adversarial samples is of great significance for improving the attack resistance ability of deep neural network models.
[0003] In the prior art, the Momentum Iterative Fast Gradient Sign Method (MIFGSM), as a method for generating adversarial samples, integrates a momentum term into the iterative attack to generate more threatening adversarial samples. At the same time, it can also improve the stability of the attack, making the generated adversarial samples show stronger robustness under different models and input conditions. However, this method still has the technical problems of low success rate of adversarial sample attacks and weak transfer ability. Summary of the Invention
[0004] The present application provides an adversarial sample generation method, apparatus, device, storage medium, and program product to solve the technical problems of low success rate of adversarial sample attacks and weak transfer ability in the prior art.
[0005] In a first aspect, the present application provides an adversarial sample generation method, including:
[0006] Obtain an input image, process the input image to obtain a target perturbation region corresponding to the input image;
[0007] Obtain a trained target model, use the input image as an initial adversarial sample, and input it into the target model for iterative processing to obtain a first loss result and a second loss result of the adversarial sample in each iteration; wherein, the first loss result is used to guide the generation of the target adversarial sample; the second loss result is used to optimize the generation process of the target adversarial sample;
[0008] According to the first loss result, use the improved Momentum Iterative Fast Gradient Sign Method to accumulate the loss gradients corresponding to the first loss result in each iteration, and according to the accumulated loss gradients, predict the gradients in the next iteration process, and update the momentum gradients in the current iteration according to the gradients to obtain an update result;
[0009] According to the update result, add perturbations to the target perturbation region, traverse the iterative process until the second loss result converges, and obtain the target adversarial sample.
[0010] In a possible implementation, the input image is used as an initial adversarial sample and input into the target model for iterative processing to obtain the first loss result and the second loss result of the adversarial sample in each iteration, including:
[0011] The input image is used as an initial adversarial sample and input into the target model for iterative processing to obtain the first feature vector of the adversarial sample in each iteration, and the label corresponding to the adversarial sample and the input image in each iteration is calculated to obtain the first loss result corresponding to the adversarial sample in each iteration;
[0012] The second feature vector of the target image is obtained, and the feature loss is obtained according to the first feature vector and the second feature vector; wherein, the target image is used to guide the generation of the target adversarial sample;
[0013] The adversarial sample and the input image in each iteration are calculated to obtain the perturbation loss;
[0014] According to the feature loss and the perturbation loss, the second loss result corresponding to the adversarial sample in each iteration is obtained.
[0015] In a possible implementation, according to the first loss result, the loss gradient corresponding to the first loss result in each iteration is accumulated by using the improved momentum iterative fast gradient sign method, and according to the accumulated loss gradient, the gradient of the next iteration process is predicted, and the momentum gradient in the current iteration is updated according to the gradient to obtain the update result, including:
[0016] When the nth iteration is performed, according to the first loss result corresponding to the current iteration, the first loss result is calculated by using the backpropagation algorithm to obtain the loss gradient of the adversarial sample relative to the output of the target model in the current iteration;
[0017] The loss gradients corresponding to the first loss results in the previous n iterations are accumulated by using the improved momentum iterative fast gradient sign method;
[0018] According to the loss gradient, the gradient of the (n + 1)th iteration is predicted to obtain the prediction result;
[0019] If the first gradient direction corresponding to the prediction result is the same as the second gradient direction of the loss gradient in the nth iteration, the prediction result and the loss gradients in the previous n iterations are calculated to obtain the calculation result, and according to the calculation result, the momentum gradient in the current iteration is updated to obtain the update result;
[0020] If the first gradient direction corresponding to the prediction result is opposite to the second gradient direction of the loss gradient in the nth iteration, the momentum gradient in the current iteration is updated according to the prediction result to obtain the update result.
[0021] In a possible implementation, processing the input image to obtain a target perturbation region corresponding to the input image includes:
[0022] Processing the input image to obtain a perturbation region corresponding to the input image, and obtaining a target perturbation region at a target pixel position in the perturbation region according to the perturbation region.
[0023] In a possible implementation, processing the input image to obtain a perturbation region corresponding to the input image, and obtaining a target perturbation region at a target pixel position in the perturbation region according to the perturbation region includes:
[0024] Processing the input image to obtain a corresponding attention heatmap, and selecting any pixel position in the attention heatmap as the center to determine a square region with a preset side length;
[0025] Calculating multiple total weight values for the square regions corresponding to each pixel position in the attention heatmap; wherein, the square regions do not exceed the boundary of the attention heatmap;
[0026] Comparing the multiple total weight values to obtain a perturbation region corresponding to the input image;
[0027] Performing binarization processing on the perturbation region to obtain a target perturbation region at a target pixel position in the perturbation region.
[0028] In a possible implementation, performing binarization processing on the perturbation region to obtain a target perturbation region at a target pixel position in the perturbation region includes:
[0029] Calculating a corresponding weight mean according to the total weight value corresponding to the perturbation region;
[0030] Obtaining the weight value of each pixel position in the perturbation region, comparing the weight value of each pixel position with the weight mean, and obtaining a binarization processing result of each pixel position in the perturbation region;
[0031] Obtaining a target perturbation region at a target pixel position in the perturbation region according to the binarization processing result.
[0032] In a second aspect, the present application provides an adversarial sample generation device, including:
[0033] A first processing module, configured to obtain an input image, process the input image, and obtain a target perturbation region corresponding to the input image;
[0034] The second processing module is used to obtain the trained target model, take the input image as the initial adversarial sample, input it into the target model for iterative processing, and obtain the first loss result and the second loss result of the adversarial sample in each iteration; wherein, the first loss result is used to guide the generation of the target adversarial sample; the second loss result is used to optimize the generation process of the target adversarial sample;
[0035] The update module is used to, according to the first loss result, use the improved momentum iterative fast gradient sign method to accumulate the loss gradients corresponding to the first loss result in each iteration, and according to the accumulated loss gradients, predict the gradients in the next iteration process, and update the momentum gradients in the current iteration according to the gradients to obtain the update result;
[0036] The addition module is used to, according to the update result, add perturbations to the target perturbation region, traverse the iterative process until the second loss result converges, and obtain the target adversarial sample.
[0037] In a possible implementation manner, the second processing module is further used to:
[0038] Take the input image as the initial adversarial sample, input it into the target model for iterative processing, obtain the first feature vector of the adversarial sample in each iteration, and calculate the labels corresponding to the adversarial sample and the input image in each iteration to obtain the first loss result of the adversarial sample in each iteration;
[0039] Obtain the second feature vector of the target image, and obtain the feature loss according to the first feature vector and the second feature vector; wherein, the target image is used to guide the generation of the target adversarial sample;
[0040] Calculate the adversarial sample and the input image in each iteration to obtain the perturbation loss;
[0041] According to the feature loss and the perturbation loss, obtain the second loss result of the adversarial sample in each iteration.
[0042] In a possible implementation manner, the update module is further used to:
[0043] When the nth iteration is performed, according to the first loss result corresponding to the current iteration, use the backpropagation algorithm to calculate the first loss result to obtain the loss gradient of the adversarial sample relative to the output of the target model in the current iteration;
[0044] Use the improved momentum iterative fast gradient sign method to accumulate the loss gradients corresponding to the first loss result in the previous n iterations;
[0045] According to the loss gradient, predict the gradient of the (n + 1)th iteration to obtain the prediction result;
[0046] If the first gradient direction corresponding to the prediction result is the same as the second gradient direction of the loss gradient in the nth iteration, calculate the prediction result and the loss gradient in the previous n iterations to obtain a calculation result, and update the momentum gradient in the current iteration according to the calculation result to obtain an update result;
[0047] If the first gradient direction corresponding to the prediction result is opposite to the second gradient direction of the loss gradient in the nth iteration, update the momentum gradient in the current iteration according to the prediction result to obtain an update result.
[0048] In a possible implementation manner, the first processing module is further configured to:
[0049] Process the input image to obtain a perturbation region corresponding to the input image, and obtain a target perturbation region of the target pixel position in the perturbation region according to the perturbation region.
[0050] In a possible implementation manner, the first processing module is further configured to:
[0051] Process the input image to obtain a corresponding attention heat map, and select any pixel position in the attention heat map as the center to determine a square region with a preset side length;
[0052] Calculate multiple total weight values for the square regions corresponding to each pixel position in the attention heat map; wherein, the square regions do not exceed the boundary of the attention heat map;
[0053] Compare the multiple total weight values to obtain a perturbation region corresponding to the input image;
[0054] Perform binarization processing on the perturbation region to obtain a target perturbation region of the target pixel position in the perturbation region.
[0055] In a possible implementation manner, the first processing module is further configured to:
[0056] Calculate a corresponding weight mean according to the total weight value corresponding to the perturbation region;
[0057] Obtain the weight value of each pixel position in the perturbation region, compare the weight value of each pixel position with the weight mean to obtain the binarization processing result of each pixel position in the perturbation region;
[0058] Obtain a target perturbation region of the target pixel position in the perturbation region according to the binarization processing result.
[0059] In a third aspect, the present application provides an adversarial sample generation device, including: a memory, a processor;
[0060] The memory stores computer execution instructions;
[0061] The processor executes the computer-executable instructions stored in the memory, such that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.
[0062] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above first aspect and / or various possible implementation manners of the first aspect when being executed by a processor.
[0063] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements the above first aspect and / or various possible implementation manners of the first aspect when being executed by a processor.
[0064] An adversarial sample generation method, apparatus, device, storage medium and program product provided by the present application iteratively processes an input image as an initial adversarial sample input into a trained target model. In each iteration, two loss results are obtained, namely a first loss result and a second loss result. The first loss result mainly plays a role in guiding the generation of the target adversarial sample, providing guidance for the generation of the adversarial sample in the direction of successfully deceiving the black-box model; while the second loss result focuses on optimizing the generation process of the adversarial sample; combining the first loss result and the second loss result can better balance the attack effect and quality of the adversarial sample, improving the attack success rate and transfer ability; at the same time, according to the first loss result, an improved momentum iterative fast gradient sign method is adopted. This method accumulates the loss gradients corresponding to the first loss result in each iteration, predicts the gradients in the next iteration process through these accumulated gradients, and then updates the momentum gradients in the current iteration based on the predicted gradients. The update of the momentum gradients here is an optimization strategy, which makes the adversarial sample generation process more dynamically adaptable on the basis of accurately calculating the required perturbation, further improving the attack success rate and transfer ability of the adversarial sample. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0066] Figure 1 is a flowchart showing an adversarial sample generation method provided by an embodiment of the present application Figure 1 ;
[0067] Figure 2 is a flowchart showing an adversarial sample generation method provided by an embodiment of the present application Figure 2 ;
[0068] Figure 3 is a flowchart showing a method for generating a target perturbation region provided by an embodiment of the present application;
[0069] Figure 4 Flow schematic of an adversarial sample generation method provided by an embodiment of the present application Figure 3 ;
[0070] Figure 5 Schematic diagram of the face recognition model training process provided by an embodiment of the present application;
[0071] Figure 6 Schematic diagram of the structure of an adversarial sample generation device provided by the present application;
[0072] Figure 7 Schematic diagram of the structure of an adversarial sample generation device provided by the present application.
[0073] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be given later. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0074] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0075] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, storage, use, processing, transmission, provision, disclosure, and application, etc. of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0076] In addition, the present application involves big data analysis of user information (including but not limited to personal biometric features, identity data, consumption data, asset data, electronic terminal operation data, etc.), and uses artificial intelligence technology for automated decision-making. For technical solutions that make decisions having a significant impact on personal rights and interests based on the results of automated decision-making, corresponding operation entrances are provided for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.
[0077] With the development of artificial intelligence technology, deep neural network models have achieved extremely remarkable results in many fields such as image recognition and image classification. For example, the face recognition system can use the deep neural network model to automatically detect face images that meet the requirements, and vehicles can perform intelligent recognition of road signs during autonomous driving. However, while artificial intelligence technology brings new technologies, it also poses security risks. Attackers can use various attack methods, such as privacy data leakage, backdoor attacks, and adversarial sample attacks, to tamper with data, resulting in security problems. Among them, adversarial sample attacks have become one of the key research contents of deep neural network models. Adversarial samples refer to adding tiny perturbations that are almost invisible to the naked eye to the input data, which can cause the deep neural network model to produce incorrect results during classification or recognition.
[0078] For example, in recent years, due to the high efficiency and convenience of deep neural network models, online payment methods have penetrated into people's daily lives. Among many online payment methods, face recognition, as a new technology widely used in the transaction systems of major banks, its security is particularly crucial. However, recently, security incidents caused by incorrect face recognition have occurred repeatedly. For example, attackers can mislead the face recognition system by using carefully designed adversarial samples, such as wearing masks or glasses with special patterns, causing it to incorrectly identify the attacker as a legitimate user, and then achieving face unlocking of the mobile phone or illegal access to the bank account.
[0079] In practical applications, deep neural network models generally appear as black-box models with unknown model structures and parameter information. Therefore, some researchers have explored using surrogate models for black-box attacks. This method first selects a surrogate model with a known structure that performs the same task as the black-box model, then attacks the surrogate model to generate adversarial samples, and then uses these samples to perform black-box attacks on the black-box model. This strategy performs well in white-box model attacks, but may have overfitting problems when generating adversarial samples, resulting in poor performance when migrating to other models.
[0080] To solve the above problems, many studies have been dedicated to reducing overfitting to surrogate models to improve the transferability of adversarial examples. Among them, the Momentum Iterative Fast Gradient Sign Method (MI-FGSM), as a method for generating adversarial examples, integrates a momentum term into the iterative attack to prevent the model from overfitting to specific samples, thereby enhancing the transferability and threat of adversarial examples. At the same time, it can also improve the stability of the attack, making the generated adversarial examples more robust under different models and input conditions. However, this method still has technical problems such as low success rate of adversarial example attacks and weak transferability.
[0081] An adversarial example generation method, device, equipment, storage medium and program product provided by this application iteratively processes the input image as the initial adversarial example into the trained target model. In each iteration, two loss results are obtained, namely the first loss result and the second loss result. The first loss result mainly plays a role in guiding the generation of the target adversarial example, providing guidance for the generation of the adversarial example in the direction of successfully deceiving the model; while the second loss result focuses on optimizing the generation process of the adversarial example; combining the first loss result and the second loss result can better balance the attack effect and quality of the adversarial example, improving the attack success rate and transferability; at the same time, according to the first loss result, an improved Momentum Iterative Fast Gradient Sign Method is adopted. This method accumulates the loss gradients corresponding to the first loss result in each iteration, predicts the gradients in the next iteration process through these accumulated gradients, and then updates the momentum gradient in the current iteration based on the predicted gradients. This momentum gradient update is an optimization strategy, which makes the adversarial example generation process more dynamically adaptable on the basis of accurately calculating the required perturbation, further improving the attack success rate and transferability of the adversarial example.
[0082] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0083] Figure 1 Schematic flow of an adversarial example generation method provided by an embodiment of this application Figure 1 As Figure 1 shown, the method includes:
[0084] S101. Obtain an input image, process the input image to obtain a target perturbation region corresponding to the input image.
[0085] In this embodiment, the source of the input image can be obtained from a camera or some datasets (such as face image datasets (CASIA-WebFace dataset, LFW dataset, etc.), animal image datasets (image datasets of birds (Caltech-UCSD Birds 200, CUB-200), etc.), traffic image datasets (KITTI dataset, etc.), etc.). The input image can be any type of image, such as a face input image, an animal input image, a traffic input image, an art work image, etc.
[0086] Optionally, the input image is processed, for example, using a predefined image processing algorithm or a method such as feature extraction based on a deep learning model, to obtain an attention heat map related to the input image; the attention heat map is analyzed to determine the part that has a greater impact on certain attributes of the input image (such as classification results, recognition accuracy, etc.) in the attention heat map, and this is used as the perturbation region corresponding to the input image; the perturbation region is processed to obtain the target perturbation region at the target pixel position in the perturbation region.
[0087] S102. Obtain the trained target model, use the input image as the initial adversarial sample, and input it into the target model for iterative processing to obtain the first loss result and the second loss result of the adversarial sample in each iteration.
[0088] In this embodiment, the trained target model is obtained. The target model can be, for example, a deep neural model (such as a Residual Networks (ResNet) model, a Squeeze-and-Excitation Residual Networks (SEResNet), an Attention56 model, a Mobile Network (MobileNet), etc.); among them, the ResNet model can be, for example, ResNet50, ResNet101, ResNet152; the SEResNet model can be, for example, SEResNet50, SEResNet101, etc.). The purpose of using this target model is: in the black-box attack scenario, this target model is used as a substitute model for the black-box model, and this substitute model can be a white-box model, and the attacker can fully access its structure and parameters. By generating adversarial samples on the substitute model and assuming that the substitute model and the target black-box model have similar characteristics to some extent, the attacker hopes that these adversarial samples can be equally effective on the target black-box model, so that these adversarial samples can be used to attack the actual black-box model subsequently.
[0089] Optionally, use the input image as the initial adversarial sample and input it into the target model for iterative processing. In each iteration, input the current adversarial sample into the target model, calculate and record the first loss result and the second loss result of the adversarial sample in each iteration. Among them, the first loss result is used to guide the generation of the target adversarial sample, and the second loss result is used to optimize the generation process of the target adversarial sample. By using the first loss result and the second loss result to guide the generation of the adversarial sample, it helps the adversarial sample to better balance the attack effect and the sample quality during the generation process, improving the attack success rate and the transfer ability.
[0090] S103. According to the first loss result, use the improved momentum iterative fast gradient sign method to accumulate the loss gradient corresponding to the first loss result in each iteration, and according to the accumulated loss gradient, predict the gradient of the next iteration process. Update the momentum gradient in the current iteration according to the gradient to obtain the update result.
[0091] In this embodiment, after obtaining the first loss result, in each iteration, calculate the corresponding loss gradient for this first loss result. This loss gradient reflects the gradient of the first loss result with respect to the adversarial sample. Use the improved momentum iterative fast gradient sign method to accumulate the loss gradients obtained in each iteration. This accumulation can synthesize the magnitude and direction information of the loss gradients in multiple iteration processes. Predict the gradient of the next iteration process according to the accumulated loss gradient. Since the previously accumulated gradients contain the trend information in the previous iteration process, therefore, based on this information, predict the gradient of the next iteration. After predicting the gradient of the next iteration, use this predicted gradient to update the momentum gradient in the current iteration. The momentum gradient plays a role similar to "inertia" in the optimization algorithm. It can accelerate the convergence of the model and help the model skip local optimal solutions. On the basis of accurately calculating the required perturbation value, the adversarial sample generation process is more dynamically adaptable and can better cope with the subsequent changes in different black-box model architectures. This adaptability improves the transfer ability of the adversarial sample.
[0092] S104. According to the update result, add perturbations to the target perturbation region, traverse the iteration process until the second loss result converges, and obtain the target adversarial sample.
[0093] In this embodiment, according to the information such as the direction and magnitude indicated in the update result, perturbations are added within the target perturbation region. For example, if the update result indicates that the brightness should be increased in the target perturbation region to achieve the optimization purpose, then the corresponding increase operation will be performed on the pixel values in this region. After completing one perturbation addition, the iterative process is entered. In each iteration, according to the current first loss result, the improved momentum iterative fast gradient sign method is used to accumulate the corresponding loss gradients in each iteration, predict the gradient of the next iteration process, update the momentum gradient in the current iteration according to the gradient; and then add perturbations to the target perturbation region again according to the new update result. This process is continuously repeated, and the perturbations in the target perturbation region are adjusted in each iteration. During the continuous iteration process, the second loss result is continuously monitored. When the second loss result converges, that is, when the value of the second loss result no longer changes significantly (or reaches a certain preset convergence criterion, such as the change in the loss value is less than a certain threshold after a certain number of iterations), the finally obtained perturbation magnitude is added to the original input image or the directly obtained sample with perturbations is used as the target adversarial sample. Subsequently, the target adversarial sample can be used as the input of the black box model to test the security or vulnerability of the black box model.
[0094] An adversarial sample generation method provided by an embodiment of the present application determines the region in the input image that needs to be perturbed as the target perturbation region, uses the original input image as the initial adversarial sample, and inputs it into the trained target model for iterative processing to obtain the first loss result and the second loss result of the adversarial sample in each iteration to guide the generation of the adversarial sample; uses the improved momentum iterative fast gradient sign method to accumulate the gradients related to the first loss result in each iteration, and these accumulated gradients help predict the next gradient; then updates the momentum gradient in the current iteration according to the gradient, and adds perturbations in the target perturbation region according to the update result, and traverses the iterative process until the second loss result converges to obtain the target adversarial sample. This method enables the target adversarial sample to better balance the attack effect and sample quality during the generation process by using the first loss result and the second loss result, improving the attack success rate and transferability. At the same time, the improved momentum iterative fast gradient sign method is combined with the gradient of the next iteration process to update the momentum gradient in the current iteration. On the basis of accurately calculating the required perturbations, the adversarial sample generation process is more dynamically adaptable, which helps the generated target adversarial sample to adapt to black box models with different structures during black box attacks, further improving the transferability of the adversarial sample.
[0095] Figure 2 It is a flow schematic diagram of an adversarial sample generation method provided by an embodiment of the present application Figure 2 , such as Figure 2 shown, this embodiment is in Figure 1Based on this, a possible implementation of the adversarial sample generation method will be described in detail. The method includes:
[0096] S201. Process the input image to obtain the corresponding attention heatmap. Select any pixel position in the attention heatmap as the center to determine a square region with a preset side length. Calculate the square regions corresponding to each pixel position in the attention heatmap to obtain multiple total weight values.
[0097] In this embodiment, the input image can be one or multiple, and the number and type of the input image are not limited herein. Taking a face input image as an example, process the input image. For example, first, the input image can be denoised using median filtering to obtain the denoised input image. Then, perform operations such as scaling and normalization on the denoised input image to obtain an image pyramid. By generating the image pyramid, faces can be detected at different scales, which improves the robustness of the subsequent network model in different scenarios and conditions, ensuring that faces can be accurately detected even in complex backgrounds. Then, input the obtained image pyramid into the network model to obtain the position of the face in the image, which is usually represented in the form of a bounding box. Among them, the network model can be, for example, an improved multi-task cascaded convolutional neural network (Multi-Task Cascaded Convolutional Networks, MTCNN) network model. The bounding box defines the rectangular region where the face is located in the image. Then, perform an alignment operation on the face selected by the bounding box. Through alignment, the features of the face (such as eyes, nose, mouth) can be in standardized positions and directions, which is beneficial for subsequent processing (such as feature extraction or face recognition) because it reduces the influence caused by pose changes. Crop the aligned input image to a fixed size. Extracting features on the input image with a fixed size can make the extracted features consistent between different images, simplify the calculation process, and reduce the consumption of computing resources.
[0098] Next, use the Gradient-weighted Class Activation Mapping (Grad-CAM) algorithm to obtain the attention heatmap corresponding to the input image in the trained deep convolutional neural network model. Specifically: Input the input image cropped to a fixed size into the trained deep convolutional neural network model (such as the VGG16 model) to obtain the feature map of the input image, and use gradient feedback to calculate the weight value of each channel in the feature map. Among them, z represents the product of the length and width of the feature map, t represents the category of the input image, and y tRepresents the probability distribution predicted for the category corresponding to the input image in the deep convolutional neural network model. A represents the feature layer, k represents the k-th channel of the feature map, and (i, j) represents the position coordinates of the feature map. Represents the weight value for A k ; Then, the feature map is multiplied by the weight value channel by channel to obtain the weighted feature map for each channel; The weighted feature maps are summed channel by channel, and then the attention heatmap is processed using the Rectified Linear Unit (ReLU) to obtain the corresponding attention heatmap For each pixel position (i, j) in the attention heatmap, a square region with a preset side length (e.g., side length L) is determined with any pixel position in the attention heatmap as the center, and calculations are performed on this square region to obtain the corresponding total weight value The total weight value of the next square region is obtained by sliding this square region over the entire attention heatmap, that is, there are multiple total weight values for one attention heatmap.
[0099] It should be noted that the square region does not exceed the boundary of the attention heatmap, and the area of this square region is smaller than the area of the attention heatmap.
[0100] S202. Compare multiple total weight values to obtain the perturbation region corresponding to the input image, and calculate the corresponding weight mean according to the total weight value corresponding to the perturbation region.
[0101] In this embodiment, since different regions have different importance for the understanding and analysis of the input image, therefore, by comparing the total weight values, the region with the highest weight is selected as the perturbation region, and this perturbation region is used to indicate the region that has a greater impact on the overall features of the input image; Calculate the corresponding weight mean for the total weight value corresponding to this perturbation region
[0102] S203. Obtain the weight value of each pixel position in the perturbation region, compare the weight value of each pixel position with the weight mean, and obtain the binarization processing result of each pixel position in the perturbation region.
[0103] In this embodiment, according to step S201, the input image cropped to a fixed size is input into the trained deep convolutional neural network model (e.g., VGG16 model) to obtain the weight value of each pixel position, and the weight value of each pixel position is compared with the weight mean. The formula is as follows:
[0104]
[0105] Among them, BS(·) is the operation for binarization. If the weight value L of the perturbation region in the attention heatmap at this pixel position Grad-CAM(i, j) is greater than M ij If it is, set it to 1; otherwise, set it to 0. Analyze each pixel position in the perturbation area to obtain the binarization result of each pixel position in the perturbation area. Binarization simplifies the complex weight values into binary forms of 0 and 1 to narrow the range of perturbation addition, so as to effectively improve the attack effect on the premise of maintaining good adversarial sample masking performance.
[0106] S204. Obtain the target perturbation area of the target pixel position in the perturbation area according to the binarization result.
[0107] In this embodiment, according to the binarization result, all pixel positions set to 1 in the perturbation area are used as the target pixel positions. The target perturbation area is composed of all target pixel positions.
[0108] Optionally, Figure 3 is a schematic flowchart of a method for generating a target perturbation area provided by an embodiment of the present application. As Figure 3 shown, input a cropped face input image of a certain size into the VGG16 model. The VGG16 model processes the face input image and outputs an attention area, which can be represented as an attention heat map. Determine a square area with a side length of L centered on any pixel position in the attention area. Then, slide this square area over the entire attention area to obtain the total weight value of the next square area, and continuously compare it with the previous area to determine the area with the largest total weight value, that is, the brown square area in the figure is the perturbation area, and this perturbation area will be used as a mask. Finally, use the idea of square binarization (BS) to generate a target perturbation area (MASK) for this area, and this target perturbation area can be directly used in the input image. This method avoids the global addition of adversarial sample perturbations and improves the attack success rate of adversarial samples.
[0109] S205. Obtain the trained target model, use the input image as the initial adversarial sample, input it into the target model for iterative processing, obtain the first feature vector of the adversarial sample in each iteration, and calculate the labels corresponding to the adversarial sample and the input image in each iteration to obtain the first loss result of the adversarial sample in each iteration.
[0110] In this embodiment, a trained target model is obtained, and the input image is used as an initial adversarial sample and input into the target model for iterative processing, so that the generated target adversarial sample outputs a corresponding category close to the category corresponding to the target image after being input into the target model. In a possible implementation, in this iterative process, taking the nth iteration as an example, the current iterative adversarial sample and the label corresponding to the input image are calculated to obtain the first loss result corresponding to the adversarial sample in the current iteration, where the label corresponding to the input image refers to the category corresponding to the input image. The adversarial sample of the current iteration is input into the target model, and the first feature vector f of the adversarial sample of the current iteration is extracted from the target model adv , and the first feature vector is used to represent the features of the adversarial sample of the current iteration.
[0111] S206. Obtain the second feature vector of the target image, and obtain the feature loss according to the first feature vector and the second feature vector; wherein, the target image is used to guide the generation of the target adversarial sample; calculate the adversarial sample and the input image in each iteration to obtain the perturbation loss.
[0112] Optionally, obtain the second feature vector f of the target image ori , calculate the cosine similarity between the first feature vector and the second feature vector to obtain the corresponding feature loss, and the calculation formula is as follows:
[0113]
[0114] where n represents the number of input images, cos(·) represents the cosine similarity between the first feature vector and the second feature vector, L fea (f, y) represents the feature loss, y represents the label corresponding to the input image (taking values of 1 or -1), and margin represents the margin parameter, which is generally defaulted to 0.
[0115] Calculate the adversarial sample and the input image in the current iteration to obtain the perturbation loss, and the calculation formula is as follows:
[0116] L dis =||x - x adv ||2
[0117] where x represents the input image, x adv represents the adversarial sample in the current iteration, and L dis is the perturbation loss.
[0118] It should be noted that during the iterative process, the feature loss is mainly used to improve the similarity between the adversarial sample of the current iteration and the target image. By calculating the cosine similarity between the first feature vector of the adversarial sample and the second feature vector of the target image, an increase in the cosine similarity means an increase in the similarity between the adversarial sample of the current iteration and the target image, which increases the probability of the subsequent black-box model making an incorrect identification and improves the success rate of the attack. The perturbation loss is mainly used to reduce the difference between the adversarial sample and the input image in each iteration and enhance the concealment of the perturbation. It is necessary to constrain the adversarial sample to be as close as possible to the input image in each iteration and minimize the addition of the perturbation, which will increase the probability of the subsequent black-box model making an incorrect identification, thereby improving the success rate of the attack. By combining the perturbation loss and the feature loss to optimize the loss function to constrain the generation of the perturbation, the transferability of the final target adversarial sample is enhanced, the image quality of the target adversarial sample is ensured, and the concealment of the perturbation is improved.
[0119] S207. Obtain the second loss result corresponding to the adversarial sample in each iteration according to the feature loss and the perturbation loss.
[0120] Optionally, calculate the feature loss and the perturbation loss to obtain the second loss result corresponding to the adversarial sample in each iteration. The calculation formula is as follows:
[0121] L sum = min δ α·L fea +β·L dis
[0122] s.t. x + δ = [0, 1]
[0123] where α and β are the weights of the feature loss and the perturbation loss respectively, used to balance each loss, and δ represents the perturbation size obtained during the current iteration process.
[0124] Optionally, during the process of calculating the second loss result, the determination method of the weight α of the feature loss and the weight β of the perturbation loss can be, for example: conduct experiments through a series of possible weight combinations to observe the influence of different combinations on the attack success rate and the quality of the adversarial sample, so as to select the corresponding target weights.
[0125] S208. When the nth iteration is performed, calculate the first loss result according to the first loss result corresponding to the current iteration by using the backpropagation algorithm to obtain the loss gradient of the adversarial sample in the current iteration with respect to the output of the target model.
[0126] In this embodiment, when the nth iteration is performed, according to the first loss result corresponding to the current iteration Starting from the last layer of the target model, calculate the loss gradient according to the chain rule. The chain rule is the core principle of backpropagation. For each layer, calculate the partial derivative of the first loss result with respect to the input of this layer. This partial derivative reflects the degree of influence of a small change in the input of this layer on the final loss result. During the calculation process, factors such as the weights of this layer, the activation function, and the output of the previous layer need to be considered. As the backpropagation progresses, gradually calculate the loss gradients of all layers from the last layer to the first layer, and finally obtain the loss gradient of the adversarial sample relative to the output of the target model in the current iteration. This gradient will be used to adjust the perturbation of the adversarial sample in the subsequent steps to generate the target adversarial sample.
[0127] S209. Use the improved momentum iterative fast gradient sign method to accumulate the loss gradients corresponding to the first loss results in the previous n iterations, and predict the gradient of the (n + 1)-th iteration according to the loss gradients to obtain the prediction result.
[0128] In this embodiment, during the iteration process, the improved momentum iterative fast gradient sign method (Time Momentum Iterative Fast Gradient Sign Method, T-MI-FGSM), that is, the temporal momentum iterative fast gradient sign method, will accumulate the loss gradients of each iteration and let this loss gradient information affect the future gradient updates. For example, use the improved momentum iterative fast gradient sign method to accumulate the loss gradients corresponding to the first loss results in the previous n iterations, and according to the loss gradients, use the following formula to obtain the gradient of the (n + 1)-th iteration, that is, the prediction result:
[0129]
[0130] where μ represents the momentum decay factor, α represents the step size, g * represents the gradient of the (n + 1)-th iteration (prediction result), and g n represents the loss gradients accumulated in the previous n iterations (i.e., the momentum gradient of the n-th iteration), represents the adversarial sample input in the n-th iteration. It should be noted that the improved momentum iterative fast gradient sign method retains the momentum decay factor and the accumulated loss gradients in the original momentum iterative fast gradient sign method to ensure the stability of the gradient update direction. On this basis, predict the gradient of the (n + 1)-th iteration through the loss gradients accumulated in the previous n iterations, and generate the adversarial sample in the current iteration process in combination with the prediction result.
[0131] S210. If the first gradient direction corresponding to the prediction result is the same as the second gradient direction of the loss gradient in the n-th iteration, calculate the prediction result and the loss gradients in the previous n iterations to obtain a calculation result, and update the momentum gradient in the current iteration according to the calculation result to obtain an update result.
[0132] Optionally, when the first gradient direction in the prediction result is the same as the second gradient direction of the loss gradient in the n-th iteration, according to the formula:
[0133]
[0134] Calculate the prediction result g * and the loss gradients g n in the previous n iterations to obtain a calculation result g n+1 . This calculation result accelerates the update of the gradient and helps to move faster towards the goal (such as generating a target adversarial sample). Update the momentum gradient in the current iteration according to the calculation result to obtain an update result, which prompts the adversarial sample generated in the n-th iteration to be used as the input sample for the (n + 1)-th iteration.
[0135] S211. If the first gradient direction corresponding to the prediction result is opposite to the second gradient direction of the loss gradient in the n-th iteration, update the momentum gradient in the current iteration according to the prediction result to obtain an update result.
[0136] In this embodiment, when the predicted first gradient direction is opposite to the second gradient direction of the loss gradient in the n-th iteration, this first gradient direction is used to slow down the update speed, that is, the momentum gradient in the current iteration is updated according to the prediction result to obtain an update result. Since the first gradient direction is opposite to the current trend, if updated directly in the original way, it may lead to excessive deviation from the correct direction or unstable updates. Therefore, by using this first gradient direction to slow down the update speed, it is possible to search more stably in the parameter space to help improve the quality of the adversarial sample. For example, in adjusting the update step size, the gradient in the opposite direction will reduce the originally large update step size, avoiding crossing the optimal update direction too quickly or falling into the trap of local optimal solutions.
[0137] S212. Add a perturbation to the target perturbation region according to the update result, and traverse the iteration process until the second loss result converges to obtain the target adversarial sample.
[0138] In this embodiment, according to the update result, add a perturbation to the target perturbation region to obtain the perturbation value in each iteration process:
[0139] δ = α·sign(g n+1 )
[0140] and the adversarial examples obtained after adding perturbations:
[0141]
[0142] Traverse the above iterative process until the second loss result converges to obtain the final perturbation value. Calculate the final target adversarial example x by the following formula for the target perturbation region of the input image determined in step S204 and the perturbation value: adv :
[0143] x adv = x + BS(L Grad-CAM , M ij )·δ
[0144] An adversarial example generation method provided by an embodiment of the present application processes an input image to obtain an attention heat map, determines a square region centered on any pixel in the heat map, and calculates the total weight value. The perturbation region is determined by comparing multiple total weight values, and the average weight is calculated. Then, the weight values of the pixels in the perturbation region are compared with the average value to obtain a binarization result, thereby determining the target perturbation region, reducing the range of perturbation addition, and effectively improving the attack effect on the premise of maintaining good masking performance of the adversarial example. The input image is used as the initial adversarial example and input into a trained target model for iteration. Each iteration obtains the first feature vector of the adversarial example, acquires the second feature vector of the target image, calculates the feature loss with the first feature vector, and calculates the perturbation loss through the adversarial example and the input image, thereby obtaining the second loss result. Combining the feature loss and the perturbation loss effectively improves the attack success rate of the adversarial example, enhances the transfer ability, ensures the image quality of the sample, and improves the concealment of the perturbation; during the iteration process, the gradient is processed based on an improved momentum iterative fast gradient sign method, and the momentum gradient is updated according to different situations of the gradient direction, which can make the iterative process more stable and efficient. This method makes better use of the information of previous iterations and the prediction results of the next step, improves the quality of the generated adversarial example, and further improves the attack success rate and transfer ability of the adversarial example.
[0145] Optionally, Figure 4 is a flowchart of an adversarial example generation method provided by an embodiment of the present application Figure 3 , as Figure 4 shown, this embodiment is carried out in Figure 2Based on the embodiments, a possible implementation manner of the adversarial sample generation method is described in detail. The method includes: taking multiple face input images as an example, first processing the multiple face input images to obtain multiple face input images cropped to a fixed size, inputting the multiple face input images cropped to a fixed size into the trained VGG16 model, using the Grad-CAM algorithm to obtain the attention heat map of the picture, and determining the key area (mask) of the input picture; then using the strategy for determining the face attack area to obtain the target perturbation area (MASK) corresponding to each face input image; taking the multiple face input images cropped to a fixed size as multiple initial adversarial samples, inputting them into the face recognition model, and analyzing them separately with the target face image to output the feature loss (L fea ) of each initial adversarial sample. At the same time, analyzing the multiple face input images cropped to a fixed size corresponding to each initial adversarial sample among the multiple initial adversarial samples to obtain the perturbation loss (L dis ) of each initial adversarial sample; according to the feature loss (L fea ) and the perturbation loss (L dis ), obtaining the total loss (L sum ) of each initial adversarial sample. Using the T-MI-FGSM algorithm to add perturbations within the target perturbation area (MASK) to obtain multiple first adversarial samples; then inputting the multiple first adversarial samples into the face recognition model, and repeating the process of obtaining the first adversarial samples from the initial adversarial samples to obtain multiple second adversarial samples. Traverse the iterative process. In this iterative process, the total loss (L sum ) is used to constrain the T-MI-FGSM algorithm to perform multiple iterations within the target perturbation area, and each iteration makes a small adjustment until the total loss reaches a convergence state. At this time, the final perturbation is obtained. Adding this perturbation to the target perturbation area of the multiple face input images cropped to a fixed size to generate the target adversarial sample.
[0146] Optionally, the CASIA-WebFace face dataset is selected as the training set. This dataset covers approximately 490,000 face photos from 10,575 different individuals. These images capture faces in various environments, poses, and expressions, and the sizes of the images are not uniform. The test set consists of 6,000 pairs of images randomly selected from the LFW dataset, with half being positive sample pairs and the other half being negative sample pairs.
[0147] First, seven deep neural models (ResNet50, ResNet101, ResNet152, SEResNet50, SEResNet101, Attention56, MobileNet) were selected as the face recognition models, that is, the surrogate models of the black-box model, and pre-training operations were performed on these seven deep neural models respectively to improve their performance in face recognition tasks. In addition, in order to evaluate the performance of the adversarial sample generation method proposed in the above embodiments in the black-box attack scenario, six models with defense capabilities (Inc-3ens3, Inc-3ens4, IncRes-2ens, HGD, R&P, and NIPS-3) were selected as the black-box models, that is, the attack models.
[0148] At the same time, in order to more accurately measure the performance of the T-MI-FGSM algorithm, it was compared with three baseline algorithms: the Fast Gradient Sign Method (FGSM), the Iterative Fast Gradient Sign Method (I-FGSM), and the Momentum Iterative Fast Gradient Sign Method (MI-FGSM). The attack success rate (ASR), peak signal-to-noise ratio (PSNR), and mean square error (MSE) were selected as evaluation indicators to evaluate the attack effect. Their calculation formulas are shown as follows:
[0149]
[0150] In the formula, m and n represent the size of the image, I represents the original clean image, and MAX i refers to the maximum value of the color of the sampling points in the image, usually set to 255.
[0151] (1) Training of the face recognition model
[0152] To ensure the smooth progress of subsequent adversarial attack experiments, optionally, it is first necessary to fully train the seven face recognition models and verify the recognition effect of the face recognition models. The dataset used for training is the CASIA-WebFace dataset, and the LFW dataset is used for testing. The specific training process is as Figure 5As shown in the figure, in the training stage, first preprocess the CASIA-WebFace dataset to obtain the processed dataset, and then divide it into a training set, a validation set, and a test set according to the ratio of 4:3:3. Train 7 network models on the training set and the validation set, then use the test set to test the models and save the trained models. Finally, use the trained models in the LFW dataset to evaluate their accuracy in the face classification task. When conducting the experiment, select 512 as the batch size and set the number of training rounds to 50. The test results of the 7 face recognition models on the LFW test set are shown in Table 1:
[0153] Table 1
[0154]
[0155] As can be seen from Table 1, the 7 face recognition models participating in the test all achieved a high recognition accuracy of over 99% on the LFW test set. This result indicates that these models have excellent face recognition capabilities after sufficient training.
[0156] (2) Attack experiment on the face recognition model
[0157] Next, in order to verify the effectiveness of the adversarial sample generation method proposed in the embodiment of the present application, further adversarial attack experiments are carried out on the 7 face recognition models that have achieved good recognition effects.
[0158] Experiment for determining the feature loss weight and the perturbation loss weight:
[0159] To ensure the effectiveness of the target adversarial samples in the real scenario, when selecting the feature loss weight and the perturbation loss weight, it is necessary to simultaneously consider maintaining the high attack effect of the target adversarial samples and controlling the difference between the samples and the original images to maintain the image quality. Therefore, take any one of the 7 face recognition models as the experimental model for generating target adversarial samples, and obtain the corresponding target adversarial samples for 6 different parameter configurations. Input the target adversarial samples into the experimental model, and calculate the PSNR and MSE of the target adversarial samples to measure the image quality. The specific experimental results are shown in Table 2:
[0160] Table 2
[0161]
[0162] As can be seen from Table 2, when the parameter α is set in the range of [0.0001, 0.0002], β is 0.005, and when α is in the range of [0.01, 0.02], β is 0.05, the obtained PSNR values are 7.2 and 8.5 respectively, indicating that the image quality is poor. The reason for this is that when the α value is too high or the β value is too low, the feature loss dominates in the total loss, and the influence of the perturbation loss weakens. This makes the experimental model insufficiently constrained by the perturbation, the perturbation introduced in the image becomes more obvious, and the concealment of the perturbation decreases, thus leading to a decline in image quality. To deeply explore the influence of parameter selection on the attack success rate, three parameter combinations with a PSNR value exceeding 20 were selected according to the results of Table 2, and further exploration was carried out using 7 face recognition models. The results of the attack success rate of the target adversarial samples generated by these parameter combinations are shown in Table 3:
[0163] Table 3
[0164]
[0165] As can be seen from Table 3, when the α value is fixed, appropriately increasing the β value helps to improve the attack success rate of the adversarial samples; while when the β value is fixed, appropriately reducing the α value can make the attack effect better. This indicates that in the total loss, appropriately reducing the weight of the feature loss and increasing the weight of the perturbation loss can improve the image quality of the adversarial samples while ensuring the attack success rate.
[0166] Combining the data in Table 2 and Table 3, it can be seen that among the 6 parameter configurations tested, when α is [0.0001, 0.0002] and β is 0.5, the PSNR value of the adversarial samples reaches 27.3. This is a relatively high value, meaning that under this parameter setting, the image quality is well maintained. At the same time, the attack success rate of the adversarial samples in the 7 face recognition models has increased compared with other parameter combinations, and can reach more than 90%.
[0167] Attack comparison experiment of face recognition models
[0168] Optionally, select 7 trained face recognition models as alternative models, and use the processed LFW dataset as the original input image, and input them into these 7 face recognition models respectively. After being processed by these models, the corresponding target adversarial samples of each model are obtained, where the feature loss weight α = [0.0001, 0.0002], the perturbation loss weight β = 0.5, the attenuation coefficient μ = 0.5, and the perturbation iteration times T = 40.
[0169] The adversarial sample generation method proposed in the present invention and three baseline methods are respectively applied to seven face recognition models, and their attack success rates are recorded respectively. The specific experimental results are shown in Table 4. The rows in the table represent the models performing the attack, the columns represent the models being tested. The ones marked with * are white-box attacks, and the rest of the columns represent black-box attacks. The bold numbers indicate the optimal data.
[0170] As can be seen from Table 4, the adversarial sample generation method proposed in the present invention has certain superiority. Under the white-box condition, the attack success rates achieved by the adversarial sample generation method proposed in the present invention on all surrogate models exceed 98.4%. The generated adversarial samples have good transfer attack effects, with the highest attack success rate and the best performance. In contrast, the adversarial samples generated by the FGSM algorithm have the lowest attack success rate and relatively poor effects. Under the black-box condition, on the premise of maintaining an attack success rate close to 100% in the white-box attack, the black-box attack effect of the adversarial sample generation method proposed in the present invention is also significantly better than the three baseline algorithms.
[0171] Table 4
[0172]
[0173]
[0174] To further verify the effectiveness of the adversarial sample generation method proposed in the present invention, three baseline attack algorithms and the adversarial sample generation method proposed in the present invention are used to perform adversarial attacks on six defense models. The attack experimental results are shown in Table 5:
[0175] Table 5
[0176]
[0177]
[0178] As can be seen from Table 5, when the surrogate model is ResNet50, the average attack success rates achieved by the adversarial sample generation method proposed in the present invention increase by 11.1%, 9.0%, and 4.0% respectively compared with the three baseline algorithms. When the surrogate model is ResNet101, the average attack success rates achieved by the adversarial sample generation method proposed in the present invention increase by 16.8%, 15.1%, and 6.6% respectively compared with the three baseline algorithms. From the above results, it can be seen that the attack success rates achieved by the adversarial sample generation method proposed in the present invention for attacking six defense models are all higher than those of the three baseline attack algorithms.
[0179] Figure 6 The structural schematic diagram of an adversarial sample generation device provided for this application is as Figure 6 shown. The adversarial sample generation device 600 provided in this embodiment includes:
[0180] The first processing module 601 is configured to obtain an input image, process the input image, and obtain a target perturbation region corresponding to the input image.
[0181] The second processing module 602 is configured to obtain a trained target model, use the input image as an initial adversarial sample, input it into the target model for iterative processing, and obtain a first loss result and a second loss result of the adversarial sample in each iteration; wherein, the first loss result is used to guide the generation of the target adversarial sample; the second loss result is used to optimize the generation process of the target adversarial sample.
[0182] The update module 603 is configured to, according to the first loss result, use an improved momentum iterative fast gradient sign method to accumulate the loss gradient corresponding to the first loss result in each iteration, predict the gradient of the next iteration process according to the accumulated loss gradient, and update the momentum gradient in the current iteration according to the gradient to obtain an update result.
[0183] The addition module 604 is configured to, according to the update result, add a perturbation to the target perturbation region, traverse the iterative process until the second loss result converges, and obtain the target adversarial sample.
[0184] In a possible implementation manner, the second processing module 602 is further configured to:
[0185] Use the input image as an initial adversarial sample, input it into the target model for iterative processing, obtain a first feature vector of the adversarial sample in each iteration, and calculate the label corresponding to the adversarial sample and the input image in each iteration to obtain a first loss result of the adversarial sample in each iteration.
[0186] Obtain a second feature vector of the target image, and obtain a feature loss according to the first feature vector and the second feature vector; wherein, the target image is used to guide the generation of the target adversarial sample.
[0187] Calculate the perturbation loss between the adversarial sample and the input image in each iteration.
[0188] Obtain a second loss result of the adversarial sample in each iteration according to the feature loss and the perturbation loss.
[0189] In a possible implementation manner, the update module 603 is further configured to:
[0190] When the nth iteration is performed, calculate the first loss result according to the first loss result corresponding to the current iteration by using the backpropagation algorithm, and obtain the loss gradient of the adversarial sample relative to the output of the target model in the current iteration.
[0191] Cumulate the loss gradients corresponding to the first loss results in the previous n iterations using an improved momentum iterative fast gradient sign method;
[0192] Predict the gradient of the (n + 1)-th iteration based on the loss gradient to obtain a prediction result;
[0193] If the first gradient direction corresponding to the prediction result is the same as the second gradient direction of the loss gradient in the n-th iteration, calculate the prediction result and the loss gradients in the previous n iterations to obtain a calculation result, and update the momentum gradient in the current iteration according to the calculation result to obtain an update result;
[0194] If the first gradient direction corresponding to the prediction result is opposite to the second gradient direction of the loss gradient in the n-th iteration, update the momentum gradient in the current iteration according to the prediction result to obtain an update result.
[0195] In a possible implementation manner, the first processing module 601 is further configured to:
[0196] Process the input image to obtain a perturbation region corresponding to the input image, and obtain a target perturbation region at the target pixel position in the perturbation region according to the perturbation region.
[0197] In a possible implementation manner, the first processing module 601 is further configured to:
[0198] Process the input image to obtain a corresponding attention heat map, select any pixel position in the attention heat map as the center to determine a square region with a preset side length;
[0199] Calculate a plurality of total weight values for the square regions corresponding to each pixel position in the attention heat map; wherein, the square regions do not exceed the boundary of the attention heat map;
[0200] Compare the plurality of total weight values to obtain a perturbation region corresponding to the input image;
[0201] Perform binarization processing on the perturbation region to obtain a target perturbation region at the target pixel position in the perturbation region.
[0202] In a possible implementation manner, the first processing module 601 is further configured to:
[0203] Calculate a corresponding weight mean according to the total weight value corresponding to the perturbation region;
[0204] Obtain the weight value of each pixel position in the perturbation region, compare the weight value of each pixel position with the weight mean to obtain the binarization processing result of each pixel position in the perturbation region;
[0205] According to the binarization processing result, a target perturbation region of the target pixel positions in the perturbation region is obtained.
[0206] The adversarial sample generation device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0207] Figure 7 It is a schematic structural diagram of an adversarial sample generation device provided by this application. As Figure 7 shown, the adversarial sample generation method device 700 provided in this embodiment includes: at least one processor 701 and a memory 702. Optionally, the device 700 further includes a communication component 703. Among them, the processor 701, the memory 702, and the communication component 703 are connected through a bus 704.
[0208] In a specific implementation process, at least one processor 701 executes the computer-executable instructions stored in the memory 702, so that at least one processor 701 executes the above method.
[0209] For the specific implementation process of the processor 701, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0210] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0211] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0212] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0213] This application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.
[0214] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.
[0215] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0216] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.
[0217] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0218] In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0219] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program code.
[0220] It should be noted that the terms "first", "second", etc. in the claims, the specification, and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, products, or devices.
[0221] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A method for generating adversarial samples, characterized in that: include: Acquire an input image, process the input image, and obtain a target disturbance region corresponding to the input image; Acquire a trained target model, use the input image as an initial adversarial sample, input it into the target model for iterative processing, and obtain a first loss result and a second loss result of the adversarial sample in each iteration; wherein the first loss result is used to guide the generation of the target adversarial sample; and the second loss result is used to optimize the generation process of the target adversarial sample; According to the first loss result, using an improved momentum iterative fast gradient sign method to accumulate the loss gradient corresponding to the first loss result in each iteration, and according to the accumulated loss gradient, predicting the gradient of the next iteration process, and updating the momentum gradient in the current iteration according to the gradient to obtain an updated result; According to the update result, disturbance is added to the target disturbance region, and an iterative process is repeated until the second loss result converges to obtain a target adversarial sample.
2. The method according to claim 1, characterized in that The input image is used as an initial adversarial sample and input into the target model for iterative processing to obtain a first loss result and a second loss result of the adversarial sample in each iteration, including: The input image is used as an initial adversarial sample and input into the target model for iterative processing to obtain a first feature vector of the adversarial sample in each iteration, and a label corresponding to the adversarial sample and the input image in each iteration is calculated to obtain a first loss result corresponding to the adversarial sample in each iteration; Acquire a second eigenvector of the target image, and obtain a feature loss according to the first eigenvector and the second eigenvector; wherein the target image is used to guide the generation of the target adversarial sample; Calculating the adversarial sample and the input image in each iteration to obtain a perturbation loss; According to the feature loss and the perturbation loss, a second loss result corresponding to the adversarial sample in each iteration is obtained.
3. The method according to claim 1, characterized in that According to the first loss result, the loss gradient corresponding to the first loss result in each iteration is accumulated using an improved momentum iterative fast gradient sign method, and the gradient of the next iteration process is predicted according to the accumulated loss gradient, and the momentum gradient in the current iteration is updated according to the gradient to obtain an updated result, including: When performing the nth iteration, according to the first loss result corresponding to the current iteration, the first loss result is calculated using the back propagation algorithm to obtain the loss gradient of the adversarial sample in the current iteration relative to the target model output; Accumulating the loss gradients corresponding to the first loss results in the first n iterations using an improved momentum iterative fast gradient sign method; According to the loss gradient, predict the gradient of the n+1th iteration to obtain a prediction result; If the first gradient direction corresponding to the prediction result is consistent with the second gradient direction of the loss gradient in the nth iteration, the prediction result and the loss gradient in the previous n iterations are calculated to obtain a calculation result, and according to the calculation result, the momentum gradient in the current iteration is updated to obtain an updated result; If the first gradient direction corresponding to the prediction result is opposite to the second gradient direction of the loss gradient in the nth iteration, the momentum gradient in the current iteration is updated according to the prediction result to obtain an updated result.
4. The method according to claim 1, characterized in that: Processing the input image to obtain a target disturbance region corresponding to the input image includes: The input image is processed to obtain a disturbance region corresponding to the input image, and a target disturbance region at a target pixel position in the disturbance region is obtained according to the disturbance region.
5. The method according to claim 4, characterized in that Processing the input image to obtain a disturbance region corresponding to the input image, and obtaining a target disturbance region at a target pixel position in the disturbance region according to the disturbance region, comprising: Processing the input image to obtain a corresponding attention heat map, selecting any pixel position in the attention heat map as the center to determine a square area with a preset side length; Calculating the square area corresponding to each pixel position in the attention heat map to obtain multiple weight total values; wherein the square area does not exceed the boundary of the attention heat map; Comparing the plurality of weight total values to obtain a disturbance region corresponding to the input image; The disturbance region is binarized to obtain a target disturbance region at a target pixel position in the disturbance region.
6. The method according to claim 5, characterized in that Binarization is performed on the disturbance region to obtain a target disturbance region at a target pixel position in the disturbance region, including: Calculate the corresponding weight mean according to the total weight value corresponding to the disturbance area; Obtaining a weight value of each pixel position in the disturbance area, comparing the weight value of each pixel position with the weight mean, and obtaining a binarization processing result of each pixel position in the disturbance area; According to the binarization result, a target disturbance region at a target pixel position in the disturbance region is obtained.
7. An adversarial sample generation device, characterized in that: include: A first processing module is used to acquire an input image, process the input image, and obtain a target disturbance region corresponding to the input image; A second processing module is used to obtain a trained target model, take the input image as an initial adversarial sample, input it into the target model for iterative processing, and obtain a first loss result and a second loss result of the adversarial sample in each iteration; wherein the first loss result is used to guide the generation of the target adversarial sample; and the second loss result is used to optimize the generation process of the target adversarial sample; An updating module, configured to accumulate the loss gradient corresponding to the first loss result in each iteration by using an improved momentum iterative fast gradient sign method according to the first loss result, and predict the gradient of the next iteration process according to the accumulated loss gradient, and update the momentum gradient in the current iteration according to the gradient to obtain an updated result; An adding module is used to add disturbance to the target disturbance area according to the update result, and traverse the iterative process until the second loss result converges to obtain a target adversarial sample.
8. An adversarial sample generation device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.