Adversarial sample generation method based on style generative adversarial network and electronic equipment
By constructing a style mapping network, generating network and discriminative network in the adversarial sample generation method, and being suitable for white box and black box attacks, the existing methods have high computational cost, long generation process and narrow application scenarios are solved, and efficient and low-cost adversarial sample generation is achieved.
Patent Information
- Application Number
- CN202510257711.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-13
AI Technical Summary
The existing adversarial sample generation method based on style generation adversarial networks has high calculation cost, a long generation process, and can only be applied to white box attacks, and has a narrow application scenario.
A method of adversarial sample generation based on style generation adversarial network is proposed. By obtaining the training data set of the current scene, a style mapping network, a generation network and a discriminative network are constructed, and network updates are performed according to the current scene to generate target adversarial samples. This method is suitable for white and black box attacks, reducing calculation costs and generation time.
It realizes efficient generation of adversarial samples that are visually close to real samples on white box attacks, and is suitable for black box attacks, greatly expanding the application scenarios and reducing the time and calculation cost of adversarial samples.
Smart Images

Figure CN120145151A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology. Specifically, this application relates to an adversarial sample generation method and an electronic device based on a style generative adversarial network. Background Art
[0002] Currently, with the rapid development of artificial intelligence technology, neural network models have emerged. Among them, by iteratively training the model parameters in the neural network model based on a large amount of historical sample data, the neural network model can learn the rules from the large amount of historical sample data, so as to use the trained neural network model to classify and identify target images. However, deep learning models are faced with the threat of adversarial samples. An attacker can use network attacks or other attack means. For example, in a face recognition scenario, the attacker bypasses the camera and directly implants the face image after adversarial attack into the digital link to mislead the neural network model to identify the user's identity.
[0003] In order to reduce the impact of adversarial samples on deep learning models, it is necessary to generate adversarial samples and use these adversarial samples to train deep learning models to improve the anti-interference ability. The commonly used method is the optimization-based method. The optimization-based method generates adversarial samples by minimizing the combination of the perturbation size and the loss function. Specifically, such methods use optimization algorithms to find the minimum perturbation so that the input after adding the perturbation can effectively bypass the defense mechanism of the target model and mislead the target model. However, such methods also have obvious disadvantages. First, the computational cost is high. Each time, only the perturbation for a single sample can be optimized, and the generation process is very time-consuming, making it difficult to meet the requirements of generating a large number of adversarial samples in practical applications. Second, such methods also require full access to the architecture, parameters, and gradient information of the target model and can only be applied to white-box attacks, which greatly limits the application scenarios. Summary of the Invention
[0004] In view of the disadvantages of the existing methods, this application proposes an adversarial sample generation method and an electronic device based on a style generative adversarial network, which can solve the problems of high computational cost, long generation process, and narrow application scenarios of the existing adversarial sample generation methods based on a style generative adversarial network, which can only be applied to white-box attacks.
[0005] According to one aspect of the embodiments of this application, the embodiments of this application provide an adversarial sample generation method based on a style generative adversarial network, and the method includes:
[0006] Obtain a training data set according to the current scenario, where the current scenario includes any one of white-box attacks and black-box attacks;
[0007] Construct a network according to the current scenario. The constructed network includes at least a style mapping network, a generation network, and a discriminant network. The style vector value of the style mapping network is incorporated into the generation network.
[0008] Determine the training method corresponding to the current scenario, and iteratively update the network according to the training data set and the training method to obtain a target generation network, and use the target generation network to generate a target adversarial sample.
[0009] In a possible implementation, the constructing the network according to the current scenario includes:
[0010] If it is determined that the current scenario is a white-box attack, then construct a style mapping network, a generation network, and a discriminant network. The network structures of the generation network and the discriminant network are the same.
[0011] If it is determined that the current scenario is a black-box attack, then construct a style mapping network, a generation network, a discriminant network, and a distillation model. The distillation model corresponds to the target model of the black-box attack.
[0012] In a possible implementation, the constructing the generation network includes:
[0013] Input a first vector into the style mapping network to obtain a style vector value. The first vector is sampled based on a standard normal distribution.
[0014] Construct the generation network based on a preset neural network structure, and incorporate the style vector value into the last convolutional layer of the generation network. The preset neural network structure includes a convolutional neural network structure.
[0015] In a possible implementation, the obtaining the training data set according to the current scenario includes:
[0016] If it is determined that the current scenario is a black-box attack, then obtain the target model corresponding to the black-box attack, and generate a training data set according to the training data and output data of the target model. The output data is the output of the target model in response to the input training data.
[0017] In a possible implementation, the training method corresponding to the white-box attack includes a joint training method. The iteratively updating the network according to the training data set and the training method includes:
[0018] Determine the target model corresponding to the current scenario, use the training data set and the generation network to generate adversarial samples, input the adversarial samples into the discriminant network and the target model respectively, and obtain the predicted categories of the outputs.
[0019] Calculate a first loss function according to the predicted category, and iteratively train the generation network and the discriminant network based on the first loss function.
[0020] In a possible implementation, the calculating the first loss function according to the predicted category includes:
[0021] The calculation formula of the first loss function is:
[0022]
[0023] where α and β are parameters, is the first adversarial loss, is the adversarial classification loss of the target model, is the hinge loss;
[0024] x is the input sample, represents the mean of x, D(x) is the predicted category output by the discriminant network for x, and G(x) is the adversarial perturbation generated by the generation network for x; c is the regularization boundary parameter.
[0025] In a possible implementation, the training method corresponding to the black-box attack includes dynamic distillation. The iteratively updating the network according to the training dataset and the training method includes:
[0026] Train the distillation model using the training dataset and a preset loss to obtain an initial distillation model. The preset loss includes cross-entropy loss;
[0027] Alternately update the generation network and the initial distillation model based on the dynamic distillation strategy.
[0028] In a possible implementation, the training the distillation model using the training dataset and a preset loss to obtain an initial distillation model includes:
[0029] Through obtain the initial distillation model. In the formula, f 0 represents the initial distillation model, represents the mean of the sample x in the training dataset, is the cross-entropy loss function, f(x) is the classification result output by the distillation model for the sample x, and b(x) represents the classification result output by the target model for the sample x, represents determining the model obtained when minimizing the loss function as the initial distillation model.
[0030] In a possible implementation, the alternately updating the generation network and the initial distillation model based on the dynamic distillation strategy includes:
[0031] Obtain the distillation model to be updated in the current round, and update the generation network and the discriminant network according to the classification adversarial loss of the distillation model to be updated;
[0032] Generate new adversarial samples using the updated generation network, and update the distillation model to be updated based on the new adversarial samples and the original samples.
[0033] According to one aspect of the embodiments of the present application, the embodiments of the present application provide an electronic device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method described above.
[0034] The beneficial technical effects brought by the technical solutions provided by the embodiments of the present application include:
[0035] An adversarial sample generation method based on a style generative adversarial network provided by the present application has the beneficial effect that the current scenario is obtained, a network is constructed according to the current scenario, and the style vector value of the style mapping network is incorporated into the constructed generation network; a training data set is obtained according to the current scenario; the training method corresponding to the current scenario is determined, and the network is iteratively updated according to the training data set and the training method to obtain a target generation network, and the target adversarial samples are generated using the target generation network. The present application incorporates the style vector value into the generation network, so that adversarial samples can be generated based on the style, which can not only efficiently generate adversarial samples visually close to real samples in white-box attacks, but also be well applicable to black-box attacks, greatly expanding the application scenarios, and can quickly generate multiple adversarial samples, with a short adversarial sample generation time and low computational cost, effectively meeting the need for generating a large number of adversarial samples.
[0036] The additional aspects and advantages of the present application will be partially given in the following description, which will become obvious from the following description, or will be understood through the practice of the present application. Description of the Drawings
[0037] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0038] Figure 1 is a flowchart of the adversarial sample generation method based on the style generative adversarial network provided by the embodiments of the present application;
[0039] Figure 2 is a framework diagram of the adversarial sample generation provided by the embodiments of the present application;
[0040] Figure 3 is a structural diagram of the electronic device provided by the embodiments of the present application. Detailed Embodiments
[0041] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0042] Those skilled in the art of the present technology can understand that unless specifically stated, the "the" and "this" used here may also include plural forms. It should be further understood that the term "including" used in the specification of the present application means that there are the described features, integers, steps, operations, elements and / or components, but does not exclude other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology, etc. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to another element, or it can mean that this element and another element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here means at least one of the items defined by this term. For example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0043] To make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the drawings.
[0044] The embodiments of the present application provide an adversarial sample generation method based on a style generative adversarial network, and this method can be used for mobile phones, computers, servers, the cloud, and other terminals capable of generating adversarial samples.
[0045] As Figure 1 、 Figure 2 shown, the adversarial sample generation method based on the style generative adversarial network of the present application includes:
[0046] S101: Obtain a training data set according to the current scenario.
[0047] Optionally, the current scenario includes any one of white-box attacks and black-box attacks.
[0048] Optionally, when the current scenario is a white-box attack, the training samples of the target model (the target neural network used to attack or train with adversarial samples) corresponding to the current scenario can be obtained, and the training data set can be generated using these training samples. It is also possible to obtain the data processed by the target model and use this data as the data in the training data set.
[0049] Optionally, when the current scenario is a black-box attack, obtain a training data set according to the current scenario, including: if it is determined that the current scenario is a black-box attack, obtain the target model corresponding to the black-box attack, and generate a training data set based on the training data and output data of the target model, where the output data is the output of the target model in response to the input training data.
[0050] Optionally, a part of the data can be randomly selected from multiple data sets (the data in different data sets is different) used to train the target model, and this data is used as the original sample. The original sample is input into the target model to obtain the output data of the target model, which is the prediction result of the target model for the original sample. A training data set is generated using the original sample and the output data.
[0051] S102: Construct a network according to the current scenario. The constructed network includes at least a style mapping network, a generation network, and a discriminant network.
[0052] Optionally, the constructed network can be a style generation adversarial network, and the style vector value of the style mapping network is incorporated into the generation network of this network.
[0053] Optionally, constructing a network according to the current scenario includes: if it is determined that the current scenario is a white-box attack, construct a style mapping network, a generation network, and a discriminant network, and the network structures of the generation network and the discriminant network are the same; if it is determined that the current scenario is a black-box attack, construct a style mapping network, a generation network, a discriminant network, and a distillation model, and the distillation model corresponds to the target model of the black-box attack.
[0054] Optionally, the construction of the generation network includes: inputting the first vector into the style mapping network to obtain the style vector value. The first vector is sampled based on the standard normal distribution; construct the generation network based on a preset neural network, and incorporate the style vector value into the last convolutional layer of the generation network. The preset neural network includes a convolutional neural network. By incorporating the style vector value into the last convolutional layer of the generation network, the weights of the convolutional layer are changed, thereby changing the style of the perturbation.
[0055] Optionally, the style mapping network includes multiple fully connected layers, and the standard normal distribution can be randomly generated by a function with a mean of 0 and a variance of 1.
[0056] In one embodiment, the number of fully connected layers can be 4. Sample the first vector z from the standard normal distribution, and the expression can be represents the standard normal distribution. When processing the input first vector through the style mapping network, first normalize the first vector, and then obtain the style vector value w through four fully connected layers.
[0057] Optionally, the preset neural network can be a CNN (Convolutional Neural Network) structure, through which image data can be efficiently processed, and high-quality adversarial perturbations can be generated after adversarial training.
[0058] Optionally, the discriminative network and the generative network can adopt the same convolutional neural network structure. With the same convolutional neural network, it is convenient to process image data and implement the joint training of the discriminative network and the generative network.
[0059] Optionally, the discriminative network can be a classification neural network. The discriminative network identifies the input object and inputs the classification label of the object to distinguish adversarial samples from real samples using the classification label, thereby ensuring that the generated adversarial samples are visually close to real samples.
[0060] Optionally, during a black-box attack, when the internal information (such as parameters) of the target model cannot be obtained, a distilled model corresponding to the target model can be obtained. The target model is simulated through the trained distilled model.
[0061] S103: Determine the training method corresponding to the current scenario, iteratively update the network according to the training dataset and the training method to obtain the target generative network, and use the target generative network to generate target adversarial samples.
[0062] Optionally, the training methods corresponding to different scenarios are different. Among them, the training method corresponding to the white-box attack can include joint training, and the training method corresponding to the black-box attack can include a dynamic distillation strategy.
[0063] Optionally, the training method corresponding to the white-box attack includes a joint training method. Iteratively updating the network according to the training dataset and the training method includes: determining the target model corresponding to the current scenario, generating adversarial samples using the training dataset and the generative network, inputting the adversarial samples into the discriminative network and the target model respectively, and obtaining the predicted categories of the outputs; calculating the first loss function according to the predicted categories, and iteratively training the generative network and the discriminative network based on the first loss function.
[0064] Optionally, when the scenario is a white-box attack, the sample x in the training dataset can be input into the generative network. The generative network inputs the adversarial perturbation into the sample x to obtain an adversarial sample. The specific expression can be:
[0065] x A = x + G(x), where x A represents the adversarial sample, and G(x) represents the adversarial perturbation.
[0066] Optionally, in the white-box attack scenario, the adversarial examples are respectively input into the discriminative network and the target model, and the discriminative network and the target model are used to distinguish the adversarial examples and output the predicted classes of the adversarial examples. Based on the predicted classes and the adversarial examples, a first loss function is generated.
[0067] Optionally, calculating the first loss function according to the predicted classes includes:
[0068] The calculation formula of the first loss function is:
[0069]
[0070] where α and β are parameters, and the importance can be adjusted by the numerical values of α and β of, is the first adversarial loss, is the adversarial classification loss of the target model, is the hinge loss; x is the input sample, represents the mean of the sample x, D(x) is the predicted class output by the discriminative network for the sample x, and G(x) is the adversarial perturbation generated by the generative network for the sample x; c is the regularization boundary parameter. By limiting the magnitude of the adversarial perturbation.
[0071] Optionally, can be the adversarial classification loss of the classifier in the target model, and the classifier can be the last layer network structure in the target model for completing sample classification. Among them, the expression of the adversarial classification loss can be:
[0072]
[0073] In the formula, t represents the target class (the target class is the class that the target model is expected to misclassify, and it is a label class included in the label classes that the target model can recognize), represents the loss function for training the classifier f. This adversarial classification loss encourages the perturbed image to be misclassified as the target class t.
[0074] Optionally, after obtaining the first loss function, the generative network and the discriminative network are jointly trained based on the first loss function. During joint training, the minimax game method can be used for training. The expression of the minimax game can be: In the formula, different optimization methods are adopted for the minimax game pair generation network G and the discriminant network D: the loss function of the generation network G aims to maximize the probability that the adversarial samples output by the generation network are misjudged as real data, while the loss function of the discriminant network D aims to maximize the recognition ability of the discriminant network for real data and minimize the misjudgment of adversarial samples at the same time. Through this training method, the ability of the generation network to generate adversarial samples can be further improved. After training, the obtained target generation network can independently generate adversarial samples without relying on the target model, realizing an efficient white-box attack.
[0075] Optionally, when the scenario is a black-box attack, the training method corresponding to the black-box attack includes dynamic distillation, and the network is updated iteratively according to the training data set and the training method, including: training the distillation model using the training data set and a preset loss to obtain an initial distillation model, and the preset loss includes cross-entropy loss; alternately updating the generation network and the initial distillation model based on the dynamic distillation strategy.
[0076] Optionally, the training data set may include samples corresponding to the target model and the predicted categories (classification results) output by the target model for the samples. The cross-entropy loss is sampled according to the samples and the predicted categories to train the distillation model to obtain an initial distillation model.
[0077] Optionally, training the distillation model using the training data set and a preset loss to obtain an initial distillation model includes: by obtaining the initial distillation model, where f 0 represents the initial distillation model, represents the mean of the samples x in the training data set, is the cross-entropy loss function, f(x) is the classification result output by the distillation model for the sample x, b(x) represents the classification result output by the target model for the sample x, indicating that the model obtained when minimizing the loss function is determined as the initial distillation model.
[0078] Optionally, when alternately updating based on the dynamic distillation strategy, the updates of the generation network and the discriminant network can be divided into multiple rounds. In each round of update, the generation network and the discriminant network are updated first, and then the model to be updated in the current round is updated based on the generation network and the discriminant network. Specifically, alternately updating the generation network and the initial distillation model based on the dynamic distillation strategy includes: obtaining the distillation model to be updated in the current round, updating the generation network and the discriminant network according to the classification adversarial loss of the distillation model to be updated, the initial distillation model is the distillation model to be updated in the first round of alternating update, and the model to be updated is the distillation model obtained in the previous round of update; using the updated generation network to generate new adversarial samples, and updating the distillation model to be updated based on the new adversarial samples and the original samples.
[0079] Optionally, it can be through Update the generation network and the adversarial network, where G i , D i respectively represent the generation network and the discriminative network obtained from the i-th round of training, represents the distilled model f obtained from the (i - 1)-th round of update i-1 of the adversarial classification loss.
[0080] Optionally, after updating the generation network and the discriminative network, the adversarial perturbation G i generated by the generation network G i (x) and the sample x can be used to update the distilled model to be updated, obtaining the distilled model f i
[0081]
[0082] Perform multiple rounds of alternating updates according to the above method until the number of updates reaches a preset number or meets a preset update condition, and then stop the update. The preset update condition can be that the prediction classes (i.e., classification classes) output by the target model for different adversarial samples for multiple consecutive times (e.g., 20 times) are all the classes that the target model is expected to misclassify.
[0083] The following uses specific embodiments to illustrate the adversarial sample generation method based on the style generation adversarial network of the present application.
[0084] In one embodiment, the performance of the adversarial sample generation method based on the style generation adversarial network was evaluated on the CIFAR-10 dataset under white-box attack and black-box attack settings. The perturbation upper limit was set to 8, the batch size was 128, and the learning rate was 0.001. Three adversarial training methods were used to train different models under each model architecture. The adversarial training methods included: standard FGSM adversarial training (Adv.), ensemble adversarial training (Ens.), and iterative training (Iter.Adv.).
[0085] In the white-box attack, the method of the present application was compared with the existing gradient-based method FGSM and the optimization-based method Opt. The attacked models were Resnet and Wide Resnet respectively. The comparison results are shown in Table 1.
[0086] Table 1
[0087]
[0088] It can be seen from Table 1 that the adversarial sample generation method of the present application achieved the highest attack success rate under three adversarial training methods and two target models, proving the superiority of this method in white-box attacks.
[0089] In black-box attacks, existing FGSM and Opt. need to complete the attack through model migration. To conduct a comparison of black-box attacks, adversarial examples generated by using FGSM and Opt. to attack Wide ResNet are used to test ResNet. A distilled model is trained using the method of this application to perform a black-box attack on ResNet. The experimental results are shown in Table 2.
[0090] Table 2
[0091]
[0092] As can be seen from Table 2, the adversarial example generation method of this application can also achieve a good attack success rate in black-box attacks, and the method of this application is applicable to both white-box attacks and black-box attacks, with wide applicability.
[0093] In an alternative embodiment, an electronic device is provided, as Figure 3 shown. Figure 3 The electronic device 4000 shown in the figure includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of this electronic device 4000 does not constitute a limitation to the embodiments of this application.
[0094] The processor 4001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0095] The bus 4002 may include a path for transmitting information among the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0096] The memory 4003 can be a ROM (ReadOnlyMemory), or other types of static storage devices that can store static information and instructions, a RAM (RandomAccessMemory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable ProgrammableReadOnlyMemory), a CD-ROM (CompactDiscReadOnlyMemory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.
[0097] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0098] Among them, the electronic device can be any kind of electronic product that can perform human-computer interaction with an object. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0099] The electronic device may further include a network device and / or an object device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing (CloudComputing).
[0100] The network where the electronic device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0101] Those skilled in the art of this technology can understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, changed, combined, or deleted. Further, the other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the relevant technologies that are the same as those disclosed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0102] In the description of this application, the directions or positional relationships indicated by the words "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are the exemplary directions or positional relationships based on the drawings, which are for the convenience of describing or simplifying the embodiments of this application, rather than indicating or implying that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application.
[0103] The terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0104] In the description of this application, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0105] In the description of this specification, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0106] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the present application, other similar implementation means based on the technical idea of the present application also fall within the scope of protection of the embodiments of the present application.
Claims
1. A method for generating adversarial samples based on style generative adversarial networks, characterized in that: The method comprises: Acquire a training data set according to a current scenario, wherein the current scenario includes any one of a white box attack and a black box attack; Constructing a network according to the current scene, the constructed network at least includes a style mapping network, a generation network, and a discrimination network, wherein the generation network incorporates the style vector value of the style mapping network; Determine a training method corresponding to the current scenario, iteratively update the network according to the training data set and the training method, obtain a target generation network, and use the target generation network to generate a target adversarial sample.
2. The adversarial sample generation method based on style generative adversarial network according to claim 1 is characterized in that: The constructing a network according to the current scenario includes: If it is determined that the current scene is a white-box attack, a style mapping network, a generation network, and a discriminant network are constructed, and the generation network and the discriminant network have the same network structure; If it is determined that the current scene is a black box attack, a style mapping network, a generation network, a discriminant network and a distillation model are constructed, and the distillation model corresponds to the target model of the black box attack.
3. The method for generating adversarial samples based on style generative adversarial networks according to claim 2, characterized in that: The construction of the generating network includes: Inputting a first vector into the style mapping network to obtain a style vector value, wherein the first vector is obtained by sampling based on a standard normal distribution; The generation network is constructed based on a preset neural network structure, and the style vector value is integrated into the last convolutional layer of the generation network, wherein the preset neural network structure includes a convolutional neural network structure.
4. The method for generating adversarial samples based on style generative adversarial networks according to claim 1, characterized in that: The step of acquiring a training data set according to the current scenario includes: If it is determined that the current scenario is a black box attack, a target model corresponding to the black box attack is obtained, and a training data set is generated according to the training data and output data of the target model, where the output data is output by the target model in response to the input training data.
5. The method for generating adversarial samples based on style generative adversarial networks according to claim 4, characterized in that: The training method corresponding to the white-box attack includes a joint training method, and the iterative network update according to the training data set and the training method includes: Determine a target model corresponding to the current scene, generate adversarial samples using the training data set and the generative network, input the adversarial samples into the discriminant network and the target model respectively, and obtain the output prediction category; A first loss function is calculated according to the predicted category, and the generation network and the discrimination network are iteratively trained based on the first loss function.
6. The method for generating adversarial samples based on style generative adversarial networks according to claim 5, characterized in that: The calculating a first loss function according to the predicted category comprises: The calculation formula of the first loss function is: Among them, α and β are parameters, For the first fight against loss, is the adversarial classification loss of the target model, Hinge loss; x is the input sample, represents the mean of x, D(x) is the predicted category of the discriminant network for the output of x, and G(x) is the adversarial perturbation generated by the generative network for x; c is the regularization boundary parameter.
7. The method for generating adversarial samples based on style generative adversarial networks according to claim 4, characterized in that: The training method corresponding to the black box attack includes dynamic distillation, and the iterative network update according to the training data set and the training method includes: Training the distillation model using the training data set and a preset loss to obtain an initial distillation model, wherein the preset loss includes a cross entropy loss; The generation network and the initial distillation model are updated alternately based on a dynamic distillation strategy.
8. The method for generating adversarial samples based on style generative adversarial networks according to claim 7, characterized in that: The step of training the distillation model using the training data set and the preset loss to obtain an initial distillation model includes: pass Get the initial distillation model, where f0 represents the initial distillation model, represents the mean of the samples x in the training data set, is the cross entropy loss function, f(x) is the classification result output by the distillation model for sample x, b(x) is the classification result output by the target model for sample x, It means that the model obtained by minimizing the loss function is determined as the initial distillation model.
9. The method for generating adversarial samples based on style generative adversarial networks according to claim 7, characterized in that: The step of alternately updating the generation network and the initial distillation model based on a dynamic distillation strategy includes: Obtain the distillation model to be updated in the current round, and update the generation network and the discrimination network according to the classification adversarial loss of the distillation model to be updated; Generate new adversarial samples using the updated generative network, and update the distillation model to be updated based on the new adversarial samples and the original samples.
10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 9.