Method for generating adversarial samples of neural network model and related device

The GAN-based training method generates adversarial samples to enhance the robustness of neural networks against perturbations, improving their security in critical applications.

CN114677556BActive Publication Date: 2025-07-15BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210204381.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-03
Publication Date
2025-07-15
Estimated Expiration
2042-03-03

AI Technical Summary

Technical Problem

Deep neural networks are susceptible to tiny perturbations that lead to false outputs, and the prior art is difficult to effectively generate adversarial samples to improve robustness.

Method used

Based on the generative adversarial network, the generator, discriminator and pretrained model are iteratively trained by obtaining the original data set to generate adversarial samples. The discriminator and generator are optimized using multi-dimensional Gaussian distribution and cross-entropy functions. The generator trains through distillation model to improve the black box attack effect.

Benefits of technology

The neural network model's generation efficiency and attack success rate of anti-interference capability are improved, and the model's anti-interference ability is enhanced, and it is suitable for different data sets and models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677556B_ABST
    Figure CN114677556B_ABST
Patent Text Reader

Abstract

The present application provides a method for generating adversarial samples of a neural network model and related devices. The method includes: based on a generative adversarial network, first obtaining an original data set corresponding to the attack requirements of the neural network model; then pre-training the neural network model to obtain a pre-trained model; iteratively training the generator, discriminator and pre-trained model of the generative adversarial network according to the original data set, and finally obtaining a target generator; and generating adversarial samples through the target generator. This method is not limited by the situation of the data set and the specific model. According to the situation of different data sets, the generator of the specified model can be trained, which conveniently improves the generation efficiency of adversarial samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of deep learning, and particularly relates to a method for generating adversarial samples of a neural network model and related devices. Background Art

[0002] In recent years, as an important branch of artificial intelligence, deep neural networks have achieved amazing results in fields such as image recognition, speech recognition, intelligent driving, and medical health. A neural network can transform raw data into a regular pattern by simulating the establishment of connections and information transmission processes of brain neurons. Deep neural networks play a key driving role in today's big data-driven innovations.

[0003] Research has found that deep neural networks are easily disturbed by tiny input perturbations, which are imperceptible to humans but can cause errors in machines. The data that causes errors is called an adversarial sample. An adversarial sample is to add subtle perturbations to the data, which will cause the model to give an incorrect output with a high confidence level. This is also a blind spot in the research of machine learning algorithms. The existence of the adversarial attack phenomenon severely restricts the application scope of neural networks. In scenarios with high security requirements, it is necessary to ensure that the network has sufficient robustness. Therefore, in order to ensure the security of neural networks, the problem of generating adversarial samples is particularly important. Summary of the Invention

[0004] In view of this, the purpose of the present application is to propose a method for generating adversarial samples of a neural network model and related devices.

[0005] Based on the above purpose, the present application provides a method for generating adversarial samples of a neural network model, including:

[0006] Obtaining an original data set corresponding to the attack requirement of the neural network model;

[0007] Pre-training the neural network model according to the original data set to obtain a pre-trained model;

[0008] Iteratively training a generator, a discriminator, and the pre-trained model according to the original data set. In response to determining that the loss of the discriminator after iteration reaches a preset threshold, taking the generator after iteration as the target generator;

[0009] Generating the adversarial sample through the target generator.

[0010] Further, the iteratively training the generator, the discriminator, and the pre-trained model according to the original data set includes:

[0011] Performing the following operations for each round of iterative training:

[0012] Sample from a multi-dimensional Gaussian distribution to obtain multiple candidate solutions for the parameters of the intermediate layer of the pre-trained model;

[0013] Replace the parameters of the intermediate layer of the pre-trained model with the candidate solutions to obtain multiple candidate neural network models corresponding to the multiple candidate solutions;

[0014] Generate a training dataset from the original dataset according to the parameters of the multiple candidate neural network models;

[0015] Select a target neural network model from the multiple candidate neural network models according to the training dataset;

[0016] Input the training dataset into the generator to obtain the perturbation of the training dataset;

[0017] Superimpose the training dataset and the perturbation to obtain a superimposed dataset, input the superimposed dataset into the discriminator to obtain the loss of the discriminator; train the target neural network model according to the superimposed dataset to obtain a new target neural network model;

[0018] Update the multi-dimensional Gaussian distribution according to the parameters of the intermediate layer of the new target neural network model to obtain the multi-dimensional Gaussian distribution in the next iteration; update the discriminator according to the loss of the discriminator to obtain the discriminator in the next iteration; update the generator according to the discriminator in the next iteration to obtain the generator in the next iteration.

[0019] Further, the generator is trained through a distillation model.

[0020] Further, the parameters of the candidate neural network model include: the number of hidden layers of the candidate neural network model, the number of neurons in each layer, the structure of the input layer, and the structure of the output layer.

[0021] Further, the loss of the discriminator is calculated through a cross-entropy function.

[0022] Further, the step of updating the discriminator according to the loss of the discriminator to obtain the discriminator in the next iteration includes:

[0023] Optimize the loss function of the discriminator through the Wasserstein distance according to the loss of the discriminator to obtain the discriminator in the next iteration.

[0024] Further, iteratively training the generator, the discriminator, and the pre-trained model according to the original data set, and in response to determining that the loss of the discriminator after iteration reaches a preset threshold, using the generator after iteration as the target generator, further includes:

[0025] Pre-set a threshold for the number of iteration rounds. When the number of iteration rounds reaches the threshold for the number of iteration rounds, stop the iterative training and output the generator trained in this round of iterative training as the target generator.

[0026] Based on the same concept, the present application further provides an adversarial sample generation device for a neural network model, including:

[0027] An acquisition module configured to acquire an original data set corresponding to the attack requirement of the neural network model;

[0028] A pre-training module configured to pre-train the neural network model according to the original data set to obtain a pre-trained model;

[0029] An iteration module configured to iteratively train the generator, the discriminator, and the pre-trained model according to the original data set, and in response to determining that the loss of the discriminator after iteration reaches a preset threshold, using the generator after iteration as the target generator;

[0030] A generation module configured to generate the adversarial sample through the target generator.

[0031] Based on the same concept, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of the above is implemented.

[0032] Based on the same concept, the present application further provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to implement the method described in any one of the above.

[0033] As can be seen from the above, the adversarial sample generation method for a neural network model provided by the present application is based on a generative adversarial network. First, an original data set corresponding to the attack requirement of the neural network model is acquired; then the neural network model is pre-trained to obtain a pre-trained model; the generator, the discriminator, and the pre-trained model of the generative adversarial network are iteratively trained according to the original data set, and finally a target generator is obtained; and the adversarial sample is generated through the target generator. This method is not limited by the situation of the data set and the specific model, and according to the situation of different data sets, the generator of the specified model can be trained, which conveniently improves the generation efficiency of the adversarial sample. Description of the Drawings

[0034] To more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments or the description of related technologies. Obviously, the drawings in the following description are only embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Figure 1 Flowchart of the method for generating adversarial samples of the neural network model according to the embodiment of the present application;

[0036] Figure 2 Flowchart of the iterative training method according to the embodiment of the present application;

[0037] Figure 3 Schematic structural diagram of the device for generating adversarial samples of the neural network model according to the embodiment of the present application;

[0038] Figure 4 Schematic structural diagram of the electronic device according to the embodiment of the present application. Detailed implementation manners

[0039] To make the objectives, technical solutions, and advantages of the present application more clear and understandable, the following further elaborates on the present application in detail in conjunction with specific embodiments and with reference to the accompanying drawings.

[0040] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meaning understood by those of ordinary skill in the art to which the present application belongs. The "first", "second", and similar terms used in the embodiments of the present application do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0041] As described in the background art section, the problem of generating adversarial samples occupies an important position in the application of neural networks.

[0042] In the process of implementing the present application, the applicant found that adversarial examples refer to input samples formed by deliberately adding subtle disturbances to a dataset, causing the model to give a wrong output with high confidence. In the context of regularization, adversarial training is used to reduce the error rate of the original independent and identically distributed test set - training the network on the training set samples with adversarial perturbations. Deep learning adversarial training is to train the model on adversarial examples. Since the adversarial examples in deep learning are caused by the linear characteristics of the model, a fast method can be designed to generate adversarial examples for adversarial training. By using adversarial examples for training, the misclassification rate on adversarial examples is greatly reduced. At the same time, selecting the adversarial examples generated by the original model as training data can train a model with higher resistance. For misclassified adversarial examples, the confidence of the model obtained by adversarial training is still very high. Therefore, adversarial training can improve the anti-interference ability of deep learning against adversarial examples.

[0043] Adversarial training helps to demonstrate the power of combining positive regularization with large function families. Pure linear models, such as logistic regression, cannot resist adversarial examples because they are restricted to be linear. Neural networks can transform the function from near-linear to locally approximately constant, so that they can flexibly capture the linear trends in the training data while learning to resist local perturbations.

[0044] In view of this, one or more embodiments of the present application provide a scheme for generating adversarial examples of a neural network model. Based on the generative adversarial network, first obtain the original dataset corresponding to the attack requirements of the neural network model; then pre-train the neural network model to obtain a pre-trained model; iteratively train the generator, discriminator, and pre-trained model of the generative adversarial network according to the original dataset, and finally obtain the target generator; and generate adversarial examples through the target generator. The technical solutions of the specific embodiments of the present application will be described below.

[0045] Reference Figure 1 , the method for generating adversarial examples of a neural network model according to an embodiment of this specification includes the following steps:

[0046] Step S101, obtain the original dataset corresponding to the attack requirements of the neural network model;

[0047] In this step, in the process of obtaining the original dataset, it is necessary to select the original dataset according to the usage scenario or test scenario of the neural network model to be optimized.

[0048] In this embodiment, the original data set can be an existing publicly available data set, such as CIFAR-10, MNIST, etc., or a custom data set uploaded by the user. The original data is in the form of "pictures stored as pixel-level matrix data". The original data is used for training the neural network model, such as training neural network models for classification systems, target recognition, etc. In this application, the original data has two uses. One is to train the neural network model to be optimized; the other is to generate the training data set required for iterative training in subsequent steps.

[0049] Step S102: Perform pre-training on the neural network model according to the original data set to obtain a pre-trained model.

[0050] In this step, the original data set is used to perform pre-training on the neural network model to obtain a pre-trained model. Before the optimization starts, the pre-trained model to be optimized is a neural network model trained by the original data set. The pre-trained model can use a classic network structure or a custom network structure. The original data set is input into the neural network model to complete the pre-training in cooperation with the neural network model.

[0051] In some embodiments, different pre-trained models are selected according to the user's device and the requirements of the model application scenario. As an example, neural network models such as ResNet34 and Inception v3 are used. The difference between ResNet34 and Inception v3 lies in the different numbers of structural layers of the two neural network models. In the case where a quickly optimized model is needed, a neural network model with fewer structural layers, such as ResNet34, can be used. In the case where a relatively secure model is needed, a neural network model with more structural layers, such as Inception v3, can be used.

[0052] Step S103: Perform iterative training on the generator, discriminator, and the pre-trained model according to the original data set. In response to determining that the loss of the discriminator after iteration reaches a preset threshold, use the generator after iteration as the target generator.

[0053] In this step, the randomness of the pre-trained model being attacked is modeled through a generative adversarial network. The generative adversarial network includes two networks, namely the generator network and the discriminator network. During training, the role of the discriminator network is to distinguish between the samples generated by the generator and the real samples, while the role of the generator network is to generate generated samples as close as possible to the real samples so as to effectively capture the distribution characteristics of the real data. After training, the generator can be used to generate adversarial samples.

[0054] In this step, iteration is repeated until a pre-set termination condition is met. It is determined whether the pre-set termination condition is satisfied. When the discriminator loss reaches the specified threshold, the algorithm stops running, and the obtained generator is the required target generator.

[0055] In some embodiments, referring to Figure 2 , for the iterative training of the generator, discriminator, and the pre-trained model according to the original data set in the step, it may specifically include:

[0056] The following operations are performed for each round of iterative training:

[0057] Step S201: Sample from a multi-dimensional Gaussian distribution to obtain multiple candidate solutions for the parameters of the intermediate layer of the pre-trained model;

[0058] In this step, first, the change of the intermediate layer of the pre-trained model is modeled as a multi-dimensional Gaussian distribution. Specifically, the parameters of the intermediate layer of the pre-trained model to be optimized are extracted. The intermediate layer of the pre-trained model extracted includes all network layers except the first layer and the last layer. The solution space of the parameters of the intermediate layer of the pre-trained model is modeled as a multi-dimensional Gaussian distribution N(μ, σ 2 C). Where μ is the mean of the distribution, σ is the learning step size, and C is the covariance matrix. The value of the parameters of the intermediate layer of the pre-trained model to be optimized is used as the initial mean μ0 of the Gaussian distribution; then the learning step size σ0 is initialized within a preset interval; as a specific example, the learning step size σ0 is initialized within the interval of 0.0001 - 0.1; in some embodiments, initializing the learning step size σ0 with 0.1 will make the training process obtain results faster.

[0059] Step S202: Replace the parameters of the intermediate layer of the pre-trained model with the candidate solutions to obtain multiple candidate neural network models corresponding to the multiple candidate solutions;

[0060] In this step, all candidate solution sets are sampled and collected within the current multi-dimensional Gaussian distribution. Each candidate solution corresponds to a candidate neural network model; the sampled intermediate layer parameters are used to replace the parameters of the intermediate layer of the pre-trained model to be optimized, and multiple candidate neural network models are obtained.

[0061] Step S203: Generate a training data set from the original data set according to the parameters of the multiple candidate neural network models;

[0062] In this step, the generation of the training dataset is based on the structure and model parameters of the pre-trained model to be optimized. When generating the training dataset, first obtain the specific structure and model parameters of the pre-trained model to be optimized, including the number of hidden layers of the deep neural network, the number of neurons in each layer, the input layer, the output layer, etc. Then process the original dataset according to the above parameters to generate the training dataset.

[0063] Step S204: Select a target neural network model from multiple candidate neural network models according to the training dataset;

[0064] In this step, the model parameters of each candidate neural network model among the multiple candidate neural network models have some differences. Therefore, the candidate neural network model that best suits the training dataset can be selected from the multiple candidate neural network models as the target neural network model.

[0065] Step S205: Input the training dataset into the generator to obtain the perturbation of the training dataset;

[0066] In this step, the generator will generate corresponding perturbations according to the training dataset to complete the iterative training of the adversarial neural network.

[0067] Step S206: Superimpose the training dataset and the perturbation to obtain a superimposed dataset, input the superimposed dataset into the discriminator to obtain the loss of the discriminator; train the target neural network model according to the superimposed dataset to obtain a new target neural network model;

[0068] In this step, after the perturbation and the training dataset are superimposed, adversarial samples are obtained. The adversarial samples and normal samples (i.e., the training dataset) are simultaneously input into the discriminator. The discriminator will perform binary classification on the normal samples and the adversarial samples to determine whether the perturbation added by the generator can cause the discriminator to misclassify, that is, whether it can deceive the discriminator. At the same time, the adversarial samples are input into the target neural network model to train the target neural network model to obtain a new target neural network model.

[0069] In this step, the discriminator discriminates the adversarial samples and the normal samples, and the discrimination result can indicate the authenticity of the adversarial samples relative to the normal samples.

[0070] Step S207: Update the multi-dimensional Gaussian distribution according to the parameters of the intermediate layer of the new target neural network model to obtain the multi-dimensional Gaussian distribution in the next round of iteration; update the discriminator according to the loss of the discriminator to obtain the discriminator in the next round of iteration; update the generator according to the discriminator in the next round of iteration to obtain the generator in the next round of iteration.

[0071] In this step, the intermediate layer parameters are obtained from the new target neural network model, the parameters such as the mean and covariance of the multi-dimensional Gaussian distribution are updated, and a new multi-dimensional Gaussian distribution is calculated.

[0072] Step S104: Generate the adversarial sample through the target generator.

[0073] In this embodiment, the generative adversarial network is obtained through training, which involves the training of two neural networks, namely the discriminator and the generator network. The training of the generative adversarial network is carried out in an alternating training manner for the generator network and the discriminator network, and the two alternately optimize the following objective function in the form of a game:

[0074]

[0075] where p data represents the distribution of the data set, which is a common representation in probability theory; ▽ is the operator for calculating the gradient; E(a) is the mean value of a; λ is the penalty coefficient, x k is the sample (i.e., the synthetic data), is the interpolated sample, ∈ is the interpolation coefficient sampled from the uniform distribution [0, 1]. The latent variable z (usually a random noise obeying the Gaussian distribution) generates the generated sample through the generator network G. For the discriminator D, this is a binary classification problem, and V(D, G) is the common cross-entropy loss in the binary classification problem. In order to ensure that V(D, G) reaches the maximum value, usually the discriminator is trained iteratively k times, and then the generator is iterated 1 time (k usually takes 1). The training steps of the generative adversarial network can be expressed as:

[0076] Initialize the parameters of the two networks, the generator network G and the discriminator network D.

[0077] Extract n samples from the training set, and the generator network generates n samples using the defined noise distribution. Fix the generator network G and train the discriminator network D to distinguish the true from the false as much as possible.

[0078] After the discriminator network D is updated cyclically k times, the generator network G is updated 1 time, so that the discriminator network can hardly distinguish the true from the false.

[0079] After multiple update iterations, in the ideal state, finally the discriminator network D cannot distinguish whether a sample comes from the real training sample set or from the sample generated by the generator network G. At this time, the discrimination probability is 0.5, and the training is completed.

[0080] In this embodiment, the number of samples n and the number of cycles k are selected according to the actual situation.

[0081] Specifically, the discriminator network is used to distinguish normal samples from adversarial samples generated by the generator network, while the generator network endeavors to generate synthetic records that are regarded as "real" by the discriminator network.

[0082] As can be seen from the above, the method for generating adversarial samples of the neural network model according to the embodiments of the present application is based on a generative adversarial network. First, an original data set corresponding to the attack requirements of the neural network model is obtained; then the neural network model is pre-trained to obtain a pre-trained model; the generator, discriminator, and pre-trained model of the generative adversarial network are iteratively trained according to the original data set, and finally a target generator is obtained; and adversarial samples are generated through the target generator. This method is not limited by the situation of the data set and the specific model. According to the situation of different data sets, the generator of the specified model can be trained, which conveniently improves the generation efficiency of adversarial samples.

[0083] In some other embodiments, the generator described in the foregoing embodiments is trained through a distillation model.

[0084] In this embodiment, it is mainly to generate perturbations in the case of a black-box attack. A black-box attack means that, assuming that the adversary does not know the prior knowledge of the training data set or the model, the black-box model is extracted by using data that does not intersect with the training set. In this case, a distillation model can be made according to the output of the black-box model, and the generator is trained by using the distillation model to improve the black-box attack effect of the generator model.

[0085] In some other embodiments, the parameters of the candidate neural network model described in the foregoing embodiments include: the number of layers of the hidden layer of the candidate neural network model, the number of neurons in each layer, the structure of the input layer, and the structure of the output layer.

[0086] In some other embodiments, the loss of the discriminator described in the foregoing embodiments is calculated by a cross-entropy function.

[0087] In this embodiment, the cross-entropy is used as a loss function in the discriminator of the adversarial neural network. p represents the distribution of true labels, and q is the predicted label distribution of the trained model. The cross-entropy function can measure the similarity between p and q. Another advantage of using the cross-entropy as a loss function is that when using the sigmoid function in gradient descent, it can avoid the problem of the learning rate reduction of the mean squared error loss function, because the learning rate can be controlled by the output error.

[0088] In some other embodiments, the updating of the discriminator according to the loss of the discriminator to obtain the discriminator in the next round of iteration includes:

[0089] According to the loss of the discriminator, optimize the loss function of the discriminator through the Wasserstein distance to obtain the discriminator in the next round of iteration.

[0090] In this embodiment, the Wasserstein distance measures the distance between two probability distributions. The advantage of the Wasserstein distance compared to the related KL divergence and JS divergence is that even if the support sets of the two distributions do not overlap or overlap very little, it can still reflect the distance between the two distributions. While the JS divergence is a constant in this situation, and the KL divergence may be meaningless.

[0091] In some other embodiments, for the iterative training of the generator, discriminator, and the pre-trained model according to the original data set described in the foregoing embodiments, and in response to determining that the loss of the discriminator after iteration reaches a preset threshold, using the generator after iteration as the target generator, further includes:

[0092] Pre-set an iteration round threshold, and when the iteration round reaches the iteration round threshold, stop the iterative training and output the generator trained in this round of iteration as the target generator.

[0093] In this embodiment, the iterative operation can end according to the satisfaction of a preset iteration round.

[0094] As can be seen from the above, for the adversarial sample generation method of the neural network model in the embodiments of the present application, the generative adversarial network is used to model the randomness of the target model being attacked. Through the combination of the Wasserstein distance and generative adversarial network technology, while ensuring the effect of the network against known attacks, adversarial samples with a high attack success rate are generated, and at the same time, the attack success rates of semi-white box attacks and black box attacks are improved. This method is not limited by the situation of the data set and the specific model. According to the situation of different data sets, the generator is trained for the specified target model, which conveniently improves the generation efficiency and attack success rate of adversarial samples.

[0095] It should be noted that the method in the embodiments of the present application can be executed by a single device, such as a computer or a server, etc. The method in this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method in the embodiments of the present application, and these multiple devices will interact with each other to complete the described method.

[0096] Note that some embodiments of the present application are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0097] Based on the same inventive concept, corresponding to any of the above method embodiments, the present application further provides a method for generating adversarial samples of a neural network model.

[0098] Refer to Figure 3 , the apparatus for generating adversarial samples of the neural network model includes:

[0099] An acquisition module 301, configured to acquire an original data set corresponding to the attack requirements of the neural network model;

[0100] A pre-training module 302, configured to pre-train the neural network model according to the original data set to obtain a pre-trained model;

[0101] An iteration module 303, configured to iteratively train a generator, a discriminator, and the pre-trained model according to the original data set, and in response to determining that the loss of the discriminator after iteration reaches a preset threshold, use the generator after iteration as the target generator;

[0102] A generation module 304, configured to generate the adversarial samples through the target generator.

[0103] In some other embodiments, the iteration module 303 is further configured to:

[0104] Perform the following operations for each round of iterative training:

[0105] Sample from a multi-dimensional Gaussian distribution to obtain multiple candidate solutions for the parameters of the intermediate layer of the pre-trained model;

[0106] Replace the parameters of the intermediate layer of the pre-trained model with the candidate solutions to obtain multiple candidate neural network models corresponding to the multiple candidate solutions;

[0107] Generate a training data set from the original data set according to the parameters of the multiple candidate neural network models;

[0108] Select a target neural network model from the multiple candidate neural network models according to the training data set;

[0109] Input the training data set into the generator to obtain the perturbation of the training data set;

[0110] Superimpose the training data set and the perturbation to obtain a superimposed data set, input the superimposed data set into the discriminator to obtain the loss of the discriminator; train the target neural network model according to the superimposed data set to obtain a new target neural network model;

[0111] Update the multi-dimensional Gaussian distribution according to the parameters of the intermediate layer of the new target neural network model to obtain the multi-dimensional Gaussian distribution in the next iteration; update the discriminator according to the loss of the discriminator to obtain the discriminator in the next iteration; update the generator according to the discriminator in the next iteration to obtain the generator in the next iteration.

[0112] In some other embodiments, the generator in the iteration module 303 is trained by a distillation model.

[0113] In some other embodiments, the parameters of the candidate neural network model in the iteration module 303 include: the number of hidden layers of the candidate neural network model, the number of neurons in each layer, the structure of the input layer, and the structure of the output layer.

[0114] In some other embodiments, the loss of the discriminator in the iteration module 303 is calculated by a cross-entropy function.

[0115] In some other embodiments, the iteration module 303 is further configured to:

[0116] Optimize the loss function of the discriminator according to the loss of the discriminator by the Wasserstein distance to obtain the discriminator in the next iteration.

[0117] For the convenience of description, when describing the above device, it is divided into various modules according to functions for description. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0118] The device in the above embodiment is used to implement the adversarial sample generation method of the corresponding neural network model in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0119] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the adversarial sample generation method of the neural network model in any of the above embodiments.

[0120] Figure 4 FIG. 1 shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0121] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0122] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.

[0123] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.

[0124] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0125] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0126] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0127] The electronic device in the above embodiment is used to implement the adversarial sample generation method of the corresponding neural network model in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0128] Based on the same inventive concept, corresponding to any of the above-mentioned method embodiments, the present application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the adversarial sample generation method of the neural network model as described in any of the foregoing embodiments.

[0129] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0130] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the adversarial sample generation method of the neural network model as described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0131] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of brevity.

[0132] Additionally, for simplicity of explanation and discussion, and so as not to render the embodiments of the present application difficult to understand, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid rendering the embodiments of the present application difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application may be implemented without these specific details or with variations of these specific details. Accordingly, these descriptions should be regarded as illustrative rather than restrictive.

[0133] Although the present application has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0134] Embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.

Claims

1. A method for generating adversarial samples of a neural network model, characterized in that Comprising: Obtain an original data set corresponding to the attack requirements of the neural network model; Pre-train the neural network model according to the original data set to obtain a pre-trained model; Iteratively train a generator, a discriminator, and the pre-trained model according to the original data set. In response to determining that the loss of the discriminator after iteration reaches a preset threshold, use the generator after iteration as the target generator; Generate the adversarial sample through the target generator; Wherein, the iteratively training the generator, the discriminator, and the pre-trained model according to the original data set includes: Perform the following operations for each round of iterative training: Sample from a multi-dimensional Gaussian distribution to obtain multiple candidate solutions for the parameters of the intermediate layer of the pre-trained model; Replace the parameters of the intermediate layer of the pre-trained model with the candidate solutions to obtain multiple candidate neural network models respectively corresponding to the multiple candidate solutions; Generate a training data set from the original data set according to the parameters of the multiple candidate neural network models; Select a target neural network model from the multiple candidate neural network models according to the training data set; Input the training data set into the generator to obtain a perturbation of the training data set; Superimpose the training data set and the perturbation to obtain a superimposed data set, input the superimposed data set into the discriminator to obtain the loss of the discriminator; train the target neural network model according to the superimposed data set to obtain a new target neural network model; Update the multi-dimensional Gaussian distribution according to the parameters of the intermediate layer of the new target neural network model to obtain the multi-dimensional Gaussian distribution in the next round of iteration; update the discriminator according to the loss of the discriminator to obtain the discriminator in the next round of iteration; update the generator according to the discriminator in the next round of iteration to obtain the generator in the next round of iteration.

2. The method according to claim 1, characterized in that The generator is trained through a distillation model.

3. The method according to claim 1, wherein The parameters of the candidate neural network model include: the number of hidden layers of the candidate neural network model, the number of neurons in each layer, the structure of the input layer, and the structure of the output layer.

4. The method according to claim 1, wherein The loss of the discriminator is calculated through a cross-entropy function.

5. The method according to claim 1, wherein The updating the discriminator according to the loss of the discriminator to obtain the discriminator in the next round of iteration includes: Optimizing the loss function of the discriminator through the Wasserstein distance according to the loss of the discriminator to obtain the discriminator in the next round of iteration.

6. The method according to claim 1, characterized in that, The iteratively training the generator, the discriminator, and the pre-trained model according to the original data set. In response to determining that the loss of the discriminator after iteration reaches a preset threshold, using the generator after iteration as the target generator further includes: Preset an iteration round threshold. When the iteration round reaches the iteration round threshold, stop the iterative training and output the generator trained in this round of iteration as the target generator.

7. An adversarial sample generation device for a neural network model, characterized in that, Comprising: An acquisition module configured to acquire an original data set corresponding to the attack requirements of the neural network model; A pre-training module, configured to pre-train the neural network model according to the original dataset to obtain a pre-trained model; An iteration module, configured to iteratively train the generator, the discriminator, and the pre-trained model according to the original dataset, and in response to determining that the loss of the discriminator after iteration reaches a preset threshold, use the generator after iteration as the target generator; A generation module, configured to generate the adversarial sample through the target generator; Wherein, the iteration module is further configured to: Perform the following operations for each round of iterative training: Sample from a multi-dimensional Gaussian distribution to obtain multiple candidate solutions for the parameters of the intermediate layer of the pre-trained model; Replace the parameters of the intermediate layer of the pre-trained model with the candidate solutions to obtain multiple candidate neural network models corresponding to the multiple candidate solutions respectively; Generate a training dataset from the original dataset according to the parameters of the multiple candidate neural network models; Select a target neural network model from the multiple candidate neural network models according to the training dataset; Input the training dataset into the generator to obtain a perturbation of the training dataset; Superimpose the training dataset and the perturbation to obtain a superimposed dataset, input the superimposed dataset into the discriminator to obtain the loss of the discriminator; train the target neural network model according to the superimposed dataset to obtain a new target neural network model; Update the multi-dimensional Gaussian distribution according to the parameters of the intermediate layer of the new target neural network model to obtain the multi-dimensional Gaussian distribution in the next round of iteration; update the discriminator according to the loss of the discriminator to obtain the discriminator in the next round of iteration; update the generator according to the discriminator in the next round of iteration to obtain the generator in the next round of iteration.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable by the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Generative adversarial network-based adversarial attack sample generation method

    CN111275115A