Decision black box scene-based adversarial model training method, equipment and medium
By training the improved VAE-Flow model, combining the R-CNN and VAE-Flow parts, the white box proxy model is used for reverse calculation and parameter correction, which solves the problem of adversarial sample generation efficiency and accuracy in black box scenarios, and achieves efficient and accurate adversarial sample generation and target model defect discovery, which improves the robustness of the neural network.
Patent Information
- Application Number
- CN202411995335.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-30
AI Technical Summary
In the black box scenario, it is difficult for the prior art to effectively generate adversarial samples, resulting in a low attack success rate and low query-based methods, especially when only output categories can be obtained.
By training an improved VAE-Flow model, which includes R-CNN and VAE-Flow parts, uses the white box proxy model for inverse calculation and parameter correction, uses the CMA-ES algorithm to generate adversarial samples, and updates the model parameters through batch loss functions.
It realizes rapid and accurate generation of adversarial samples, discovers defects of the target model, discovers security problems in advance, improves the robustness of the neural network, and has better migration on different data sets.
Smart Images

Figure CN120071075A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network adversarial attacks, and more particularly to an adversarial model training method, device, and medium based on a decision black-box scenario. Background Art
[0002] Although deep learning has achieved great development and many deep learning models have been successfully applied in real-world fields such as image recognition and human-computer games, the misjudgment of neural networks for adversarial samples remains a significant security issue that cannot be ignored. Since Szegedy et al. discovered adversarial samples, adversarial attacks and defenses have been a research hotspot in the field of machine learning in recent years. The game between the two has also promoted the continuous development of deep neural networks and provided a basis for the interpretability of neural networks.
[0003] Adversarial attacks can be classified in various ways. For example, according to the different adversarial attack scenarios, attacks can be divided into white-box attacks and black-box attacks. The former requires the attacker to perform an adversarial attack on the premise of having fully obtained all information of the target model. In contrast to white-box attacks, black-box attacks require the attacker to complete an adversarial attack without being able to obtain effective information such as the architecture and gradient of the deep neural network model.
[0004] Currently, one method in the black-box scenario is to utilize the transferability of adversarial samples and generate adversarial samples on a surrogate model for attacking the target. For example, the Nicolas method proposes first training a surrogate white-box model with a dataset labeled by querying the target model, and then using the gradient of the trained surrogate model to generate adversarial perturbations to attack the target model. Although the transfer-based attack method is very efficient, the attack success rate is relatively low. We attribute this to the fact that there are certain deviations in the boundary division between the surrogate model and the target model in terms of modeling, and this deviation is mainly due to the differences in network structure and training datasets.
[0005] Query-based methods solve the black-box optimization problem by iteratively querying the target model to update the adversarial samples. For example, the SimBA method randomly samples perturbations from a predefined orthogonal basis matrix and then adds or subtracts this perturbation to the attacked image. The natural evolution strategy (NES) is adopted to search for the distribution. Although query-based methods generally achieve better attack performance than transfer-based methods, they require frequent access to the target model and are not efficient.
[0006] In addition, previous research has focused more on the case where the probability scores of the output of the target model can be obtained, that is, the score-based black-box attack scenario. For the situation in real-world scenarios where only the output category can be obtained, that is, the decision-based black-box attack scenario, it will greatly reduce the success rate of transfer-based methods and also reduce the efficiency of query-based methods. Summary of the Invention
[0007] The purpose of the present invention is to overcome the defects existing in the above-mentioned prior art, and provide an adversarial model training method, device and medium based on the decision black box scenario, which trains an improved VAE-Flow model that is fully trained for the target model, can quickly and accurately generate adversarial samples, discover the defects of the target model, is beneficial to discovering security problems in advance, and improves the robustness of the neural network.
[0008] The purpose of the present invention can be achieved by the following technical solutions:
[0009] An adversarial model training method based on the decision black box scenario, the method includes:
[0010] Obtain a target model, an adversarial model, a white box proxy model and a training image dataset, and the target model is a black box model;
[0011] Input the original training image into the adversarial model to obtain the position feature, style variable and local feature of the image, and perform reverse calculation through the adversarial model to map the local feature of the image into an adversarial sample;
[0012] Input the adversarial sample into the white box proxy model, and perform reverse gradient derivation of the adversarial sample through the weight information of the white box proxy model to obtain a correction value and calculate the loss, and correct the parameters in the adversarial model;
[0013] According to the training image dataset, use the corrected adversarial model to generate adversarial samples, sample the adversarial samples through the CMA-ES algorithm to obtain a batch of samples, input the batch of samples into the target model, and obtain the classification information of the training images;
[0014] Calculate the probability score of the batch of samples based on the classification information, and calculate and update the parameters of the adversarial model through the batch loss function;
[0015] The updated adversarial model can be used as the final adversarial model to perform adversarial training with the target model.
[0016] Furthermore, the white box proxy model is a substitute model that replaces the adversarial model to study the internal structure and implementation details.
[0017] Furthermore, the adversarial model includes an R-CNN part and a VAE-Flow part; the R-CNN part processes the original training image to obtain the position feature of the image; the VAE-Flow part processes the original training image to obtain the style variable of the image; the VAE-Flow part processes the original training image to obtain the local feature of the image.
[0018] Furthermore, the position feature is represented in the form of an anchor box and does not contain classification labels, and the style variable is the style feature information of the whole picture.
[0019] Furthermore, the VAE-Flow part is divided into a VAE encoder and a Flow decoder. The last layer of the Flow decoder is a conversion layer, making the Flow decoder have a reversible characteristic.
[0020] Furthermore, the Flow decoder uses the position feature and the style variable, and through reverse calculation, maps the local feature of the picture into an adversarial sample.
[0021] Furthermore, through the calculation of the batch loss function, the updated parameter of the adversarial model is the parameter of the last layer of the Flow decoder.
[0022] Further, the calculation of the batch loss function includes:
[0023]
[0024]
[0025]
[0026] where n is the number of batch samples, x i ′ is the i-th batch sample, μ is the probability score of the batch sample, s n is the set of batch samples, Ν(0, C n ) is a normal distribution, P(x i ′ ) is the probability that the i-th batch sample is classified into a certain category, f(x i ) is the ideal probability that the i-th batch sample is classified into a certain category, and k is the number of categories of the batch samples.
[0027] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the adversarial model training method based on the decision black box scenario as described above are implemented.
[0028] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the adversarial model training method based on the decision black box scenario as described above are implemented.
[0029] Compared with the prior art, the beneficial effects of the present invention include:
[0030] 1. The present invention trains an improved adversarial model for a target model, uses the improved adversarial model to quickly and accurately generate adversarial samples, discovers the defects of the target model, is beneficial to discovering security problems in advance, improving the robustness of the neural network, and improving both accuracy and efficiency.
[0031] 2. In the present invention, the adversarial model includes an R-CNN part, which can obtain the main body position features of the picture without using class labels. Since it does not need to use class labels, it is label-independent and has good transferability on different data sets; the VAE-Flow part of the adversarial model maps the picture into a feature variable containing global information, which can greatly improve the classification ability of the classifier and extract the overall style features of the image semantically, with little influence from the data set labels and better transferability for the data set; the Flow decoder in the VAE-Flow part of the adversarial model is reversible, enabling the adversarial samples to have the ability to correct by getting rid of the proxy model error. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flowchart of the method of the present invention;
[0033] Figure 2 is an example diagram of the essential features to skeletal features of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] Embodiment 1
[0036] Previous adversarial samples first emphasized the adversarial success rate and then cropped and constrained the samples to a certain range. From the perspective of human recognition, we believe that adversarial samples and original samples are essentially the same, and for a neural network with a classification intention, they should all be correctly classified. Therefore, we attempt to model this essential information, use it as the starting point for optimization, and then explore adding perturbations to the edge of the classifier, which can also make the constraint of the perturbations more natural and avoid the influence of cropping constraints.
[0037] Inspired by the landmark features of face recognition, such as Figure 2As shown in the figure, we use the similar skeleton information in the picture as the representation of the essential features. First, we use an R-CNN network to locate the position of the main information in the picture, and then extract the global and other relevant information of the picture. We only complete the position localization of the main body, and do not fully obtain the anchor boxes and label information of object detection. Therefore, this feature information is label-independent, which can solve the problem of dataset bias, and the R-CNN itself also has good transferability.
[0038] After simplifying the feature information, we need to extract information from this position feature additionally to ensure the effectiveness of the skeleton feature. When the VAE extracts the picture features, it will pay more attention to the global information of the picture. This global information is semantically related to the style features of the picture, and this global information is also label-independent semantic information. Therefore, acting together with the position feature, it can be used as the skeleton feature for the subsequent generation of adversarial examples.
[0039] This embodiment aims to disclose an adversarial model training method based on the decision black box scenario. The method is as Figure 1 shown, and the specific steps include:
[0040] Step S1, obtain the target model, adversarial model, white box proxy model, and training picture dataset.
[0041] The target model is a black box model, and the white box proxy model is a substitute model for studying the internal structure and implementation details of the black box model, which is used for training in the first stage of the attack process. The attacker can use the internal information (such as weights, gradients, etc.) that can be obtained by the white box model to quickly calculate the loss and update the parameters of the adversarial model.
[0042] Step S2, input the original training picture into the adversarial model to obtain the position feature, style variable, and local feature of the picture, and perform reverse calculation through the adversarial model to map the local feature of the picture into an adversarial example;
[0043] Input the adversarial example into the white box proxy model, and perform reverse gradient derivation of the adversarial example through the weight information of the white box proxy model to obtain the correction value and calculate the loss, and correct the parameters in the adversarial model.
[0044] Step S2 is the first stage of adversarial model training. Training on the white box proxy model can use the internal information such as the weights and gradients of the white box model to help us quickly complete the derivation and loss calculation.
[0045] In this embodiment, the adversarial model is VAE-Flow, which is a reversible generative model. A Flow model is embedded in the VAE, and R-CNN is also added. Thus, this model includes an R-CNN part and a VAE-Flow part. The R-CNN part processes the original training images to obtain the location features of the images; the VAE-Flow part processes the original training images to obtain the style variables of the images, and the VAE-Flow part also processes the original training images to obtain the local features of the images. The VAE-Flow part includes a VAE encoder and a Flow decoder. The encoder is the VAE model, and the decoder is the Flow model.
[0046] In order to enable the adversarial samples to have the ability to correct the errors of the proxy model during the query attack phase, an additional transformation is added to the last layer of the Flow model to ensure the reversible characteristic of the Flow. Thus, the Flow decoder can use the location features and style variables to map the local features of the image into adversarial samples through reverse calculation.
[0047] Step S3: According to the training image dataset, use the corrected adversarial model to generate adversarial samples, sample the adversarial samples through the CMA-ES algorithm to obtain a batch of samples, input the batch of samples into the target model, and obtain the classification information of the training images;
[0048] Calculate the probability scores of the batch of samples based on the classification information, and update the parameters of the adversarial model through calculation using the batch loss function.
[0049] Step S3 is the second stage of the adversarial model training. In this stage, direct queries are made against the target model.
[0050] The calculation of the batch loss function includes:
[0051]
[0052]
[0053]
[0054] where n is the number of batch samples, x′ i is the i-th batch sample, μ is the probability score of the batch samples, s n is the set of batch samples, N(0, C n ) is the normal distribution, P(x′ i ) is the probability that the i-th batch sample is classified into a certain category, f(x i ) is the ideal probability that the i-th batch sample is classified into a certain category, and k is the number of categories of the batch samples.
[0055] Step S4, use the updated adversarial model as the final adversarial model to conduct adversarial training with the target model.
[0056] The trained improved VAE-Flow model can find more adversarial samples for the target model, and both the accuracy and efficiency are improved.
[0057] The following is an example for the actual application scenario:
[0058] For the AlexNet, VGG-16, and ResNet-18 target models, in order to further mitigate the negative impacts that may be brought by the proxy bias of the model architecture, when attacking a target model, we regard the other two models except the ResNet-18 target model as white-box proxy models.
[0059] Adopt two datasets, CIFAR-10 and office-31, as implementation examples. The office-31 dataset is a publicly available benchmark dataset widely used in the field of domain adaptation. This dataset includes 4,110 images of 31 categories collected from three different domains. The images from each source are unbalanced, that is, the number of pictures in the 31 categories is somewhat different from each other, and the categories included in each source are pictures of daily necessities such as backpacks, bicycles, and helmets. This dataset is introduced by us to evaluate the ability of the method of this embodiment to handle dataset bias.
[0060] Conduct two-stage adversarial training on the VAE-Flow model.
[0061] First, input the original image x of the dataset into the R-CNN of the VAE-Flow model, and the main position feature p of the image can be obtained. Since this feature does not need to utilize the class label, it is label-independent and has good transferability on different datasets.
[0062] Utilize the Flow part of the VAE-Flow model to map the image into a feature variable z containing global information, and then, as a condition, input it together with the image x into the Flow part again to map and obtain the local information representation v. Finally, through the reversibility of the Flow, an adversarial sample can be generated reversely. The process can be expressed as:
[0063] z = encoder(x),
[0064] V = g(x),
[0065] x = g -1 (x),
[0066] where z is the feature variable of global information, x is the image, and v is the local information representation.
[0067] The z obtained through the encoder is a better representation of image information. When used as a feature variable, it can significantly improve the classification ability of the classifier. Moreover, this z is semantically highly related to the image style variable, which means that this feature extracts the overall style variable of the image semantically and is less affected by the dataset labels, showing better transferability for datasets.
[0068] More specifically, the process of mapping to obtain the local information representation v and reversely generating adversarial samples through the reversibility of Flow is described as follows: We introduce the body position feature p extracted by R-CNN into the VAE-Flow output p + z as the representation of the backbone feature, and then input it together with x into Flow. In the forward process of Flow, Flow will map x into the local feature representation, and finally complete the generation of the adversarial sample x' through the reverse process of Flow. The final adversarial sample generation can be expressed as:
[0069] x′ = g -1 (x; z, p),
[0070] z = encoder(x; p),
[0071] p = R-CNN(x).
[0072] Input the generated adversarial sample x' into the AlexNet and VGG-16 proxy models, and the corresponding probability scores y' can be obtained. Since information such as the gradients of the white-box proxy model can be directly obtained, Δx can be directly calculated for x' through this gradient, etc., so as to update the parameters of the improved VAE-Flow model. The initial parameters are all initialized with kaiming. The training batch size batch and the number of epochs epoch are set to 64 and 100 respectively. The final loss will converge to 0.021. We can obtain an improved VAE-Flow model that can generate adversarial samples. Due to the transferability, the generated adversarial samples can have a certain attack success rate on the target model, and the training in the first stage ends.
[0073] The improved VAE-Flow trained in the first stage has a certain generation ability. However, due to the architectural differences between the proxy model and the target model, it will lead to proxy errors, thus affecting the attack success rate of the adversarial samples. Therefore, in the second stage, we made corrections for the target model, that is, we adopted a query method to modify the parameters of the last layer added to the Flow model in the improved VAE-Flow.
[0074] In the second stage, the dataset image x is input into the trained improved VAE-Flow model to obtain the adversarial sample. CMA-ES algorithm is used to sample around the generated adversarial sample x' to obtain a batch of adversarial samples x 1 ′, x 2′,..., x 3 ′, input this batch of adversarial samples into the target model, and the corresponding classification labels y can be obtained 1 ′, y 2 ′,..., y n ′, thus calculating the probability score for x′ through this batch of classification label values, and then we can calculate the batch sampling loss through the MDD loss function to complete the adjustment of the parameters of the improved VAE-Flow. Rely on this loss to update the parameters of the last time, and iterate continuously until convergence. In addition, the number of queries is limited to 10,000 times to avoid the reduction in efficiency caused by excessive queries
[0075] The formula for calculating the batch loss is as follows:
[0076] x′ i n+1 ~μ n +s n N(0, C n )
[0077]
[0078]
[0079] where n is the number of batch samples, x′ i is the i-th batch sample, μ is the probability score of the batch samples, s n is the set of batch samples, N(0, C n ) is the normal distribution, P(x′ i ) is the probability that the i-th batch sample is classified into a certain category, f(x i ) is the ideal probability that the i-th batch sample is classified into a certain category, and k is the number of categories of the batch samples
[0080] Finally, complete the two-stage training algorithm to obtain an improved VAE-Flow model that is fully trained for the target model, which can generate adversarial samples quickly and accurately, discover the defects of the target model, is conducive to discovering security problems in advance, and improves the robustness of the neural network
[0081] Embodiment 2
[0082] Based on Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory, and one or more programs are stored in the memory, and the one or more programs include instructions for executing the adversarial model training method based on the decision black box scenario as described above
[0083] At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned adversarial model training method based on the decision black box scenario. Of course, in addition to the software implementation, the present invention does not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and may also be hardware or logic devices.
[0084] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0085] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0086] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for training an adversarial model based on a decision black box scenario, characterized in that: The method comprises: Obtain a target model, an adversarial model, a white-box proxy model, and a training image dataset, wherein the target model is a black-box model; Input the original training image into the adversarial model to obtain the image's position features, style variables, and local features. Perform reverse calculations through the adversarial model to map the image's local features into adversarial samples. Input the adversarial sample into the white-box proxy model, and perform reverse gradient derivation of the adversarial sample through the weight information of the white-box proxy model to obtain the correction value and calculate the loss, and correct the parameters in the adversarial model; According to the training image dataset, the modified adversarial model is used to generate adversarial samples, and the adversarial samples are sampled through the CMA-ES algorithm to obtain batch samples, which are then input into the target model to obtain classification information of the training images. The probability scores of batch samples are calculated using the classification information, and the parameters of the adversarial model are updated through batch loss function calculation; The updated adversarial model can be used as the final adversarial model for adversarial training with the target model.
2. According to claim 1, the adversarial model training method based on the decision black box scenario is characterized in that: The white-box proxy model is an alternative model that replaces the adversarial model and studies the internal structure and implementation details.
3. According to the adversarial model training method based on the decision black box scenario of claim 1, it is characterized in that: The adversarial model includes an R-CNN part and a VAE-Flow part; the R-CNN part processes the original training picture to obtain the position features of the picture; the VAE-Flow part processes the original training picture to obtain the style variables of the picture; the VAE-Flow part processes the original training picture to obtain the local features of the picture.
4. According to claim 3, the adversarial model training method based on the decision black box scenario is characterized in that: The position feature is represented in the form of an anchor frame and does not include a classification label. The style variable is the style feature information of the entire image.
5. According to claim 3, the adversarial model training method based on the decision black box scenario is characterized in that: The VAE-Flow part is divided into a VAE encoder and a Flow decoder. The last layer of the Flow decoder is a conversion layer, which makes the Flow decoder reversible.
6. The adversarial model training method based on a decision black box scenario according to claim 5, characterized in that: The Flow decoder uses the position features and the style variables to map the local features of the image into adversarial samples through reverse calculation.
7. The adversarial model training method based on a decision black box scenario according to claim 5, characterized in that: The parameters of the updated adversarial model are calculated through the batch loss function and are the parameters of the last layer of the Flow decoder.
8. The adversarial model training method based on a decision black box scenario according to claim 1, characterized in that: The batch loss function calculation includes: Where n is the number of batch samples, x i ′ is the i-th batch sample, μ is the probability score of the batch sample, s n is the set of batch samples, N(0, C n ) is a normal distribution, P(x i ′ ) is the probability that the i-th batch sample is classified into a certain category, f(x i ) is the ideal probability that the i-th batch sample is classified into a certain category, and k is the number of categories of the batch samples.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the adversarial model training method based on the decision black box scenario as described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the adversarial model training method based on the decision black box scenario as described in any one of claims 1 to 8 are implemented.