Robust Content-Style Decoupling Model Training Method and System Based on Adversarial Training
The robustness of the style-content decoupling model is enhanced through adversarial training methods, solving the shortcomings of existing models in terms of security and adversarial attacks, and achieving higher model stability and applicability.
Patent Information
- Application Number
- CN202111222355.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-20
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-10-20
AI Technical Summary
The existing style-content decoupling model has shortcomings in terms of security, cannot effectively resist adversarial attacks, and the traditional methods are not robust enough.
Adversarial training is adopted to enhance the robustness of the model through the design of style encoder and content encoder.
Improves the robustness of the model, making it more resistant to adversarial attacks, while keeping the decoupling ability unaffected, and has applicability and generalization capabilities.
Smart Images

Figure CN113988289B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of style-content decoupling of pictures. Specifically, it relates to a method and system for training a content-style decoupling model robust based on adversarial training, and in particular, a method for enhancing the robustness of a style-content decoupling model by using adversarial learning commonly used in the field of deep learning security. Background Art
[0002] In the field of style-content decoupling, a picture is orthogonally decoupled into a style feature vector and a content feature vector. Among them, the style feature vector only contains the features that can identify the category of the main object in the picture, such as the coat color and breed of a cat in a picture of a cat. The content feature vector contains the remaining features in the picture, such as the posture, orientation, size and background of a cat in a picture of a cat. Style-content decoupling methods have many real-world applications, such as speech synthesis, speech recognition, face recognition, etc. In these applications, security is a very important consideration. Mistaking one person for another has a great security risk.
[0003] Traditional style-content decoupling fields usually do not consider the security of the model, but only consider the quality of decoupling. That is, the degree of independence of the style feature vector and the content feature vector, and whether these two features contain enough features to restore the original picture. Existing picture decoupling tasks usually fall into two categories. One is to use the category features extracted from the picture classification task as the style features, and at the same time use the picture reconstruction task to obtain the content features. The insecurity of this scheme comes from the non-robustness of its classification task. Now there have been sufficient studies finding that simple classification tasks cannot obtain a sufficiently robust model. The second scheme is the GAN model based on AdaIN, which uses the generative adversarial nature in GAN to decouple content-style. This method uses the generative adversarial nature of GAN itself to decouple the model. Although GAN itself has a certain defensive ability against adversarial attacks, it cannot resist direct adversarial attacks. Summary of the Invention
[0004] Aiming at the defects in the prior art, the purpose of the present invention is to provide a method and system for training a content-style decoupling model robust based on adversarial training.
[0005] According to a method for training a content-style decoupling model robust based on adversarial training provided by the present invention, it includes: adopting a style-content decoupling method to train the decoupling ability of the model, and adopting an adversarial training method to enhance the robustness of the model;
[0006] The content-style decoupling model includes a style encoder Es, a content encoder Ec and a decoder;
[0007] The style-content decoupling method: Let the style encoder only extract the style of the image, and let the content encoder only extract the content of the image;
[0008] The adversarial training method: Use adversarial samples to train the model to increase the robustness of the model.
[0009] Preferably, the training method specifically includes the following steps:
[0010] Step S1: Pre-train the style Embedding values for each category and the content Embedding values for each image;
[0011] Step S2: Randomly sample some images from the training dataset;
[0012] Step S3: Generate adversarial samples for each sampled image;
[0013] Step S4: Calculate the decoupling loss and the reconstruction loss function value L of the model using the adversarial samples and the original image samples d ;
[0014] Step S5: Use the total loss function value L d to perform gradient descent to update the model parameter values Θ, where the model parameter values Θ are the neuron weight values in the convolutional layer and the fully connected layer of the neural network model;
[0015] Step S6: Repeat Step S2 - Step S5 until the total loss function value L d converges to obtain a robust content-style decoupling model.
[0016] Preferably, Step S2 includes:
[0017] Step S1.1: Randomly sample a Batch of original images;
[0018] Step S1.2: For a Batch of images, calculate the following Loss function value:
[0019]
[0020] where x i is the original image, i represents the i-th image in a Batch, s i is the style Embedding of the category where x i is located, c i is the content Embedding of x i , G θ is the decoder in the decoupling model, θ is the network parameter value of the decoder, and β is a hyperparameter;
[0021] Step S1.3: According to the Loss function value in Step S1.2, perform gradient descent update on s i , c i , G θ :
[0022]
[0023]
[0024]
[0025] where η is the learning rate;
[0026] Step S1.4: Repeat Step S1.1 - Step S1.3 until the loss value L p converges.
[0027] Preferably, the said Step S3 includes:
[0028] Step S3.1: Add random perturbation to the original picture, x′ i = x i + ∈·ξ, where x′ i represents the picture after adding perturbation, x i represents the original picture, ∈ represents the range of perturbation, and ξ is a uniformly distributed random variable in the interval [-1, 1];
[0029] Step S3.2: Calculate the loss function value of the model, and the loss function is designed as: l adv = ||E S (x′ i ) - E S (x i )|| + ||E c (x′ i ) - E C (x i )||, where E S and E C represent the style encoder and the content encoder respectively;
[0030] Step S3.3: According to the loss function value in Step S3.2, perform gradient ascent on the picture to obtain a new adversarial sample x′ i , and the formula is:
[0031]
[0032] Step S3.4: Repeat Step S3.2 - Step S3.3 multiple times to generate better adversarial samples.
[0033] Preferably, the said Step S4 includes:
[0034] Calculate the clean sample \(x\) through the following formula i and its corresponding adversarial sample \(x'\) i 's reconstruction decoupling loss value:
[0035]
[0036] where \(G\) θ is the decoder, \(E_s\) s is the style encoder, \(E_c\) c is the content encoder, \(s\) i is the style Embedding value obtained in step S1, \(c\) i is the content Embedding value obtained in step S1.
[0037] According to a content-style decoupling model training system based on adversarial training robustness provided by the present invention, it includes: using a style-content decoupling method to train the decoupling ability of the model, and using an adversarial training method to enhance the robustness of the model;
[0038] The content-style decoupling model includes a style encoder \(E_s\), a content encoder \(E_c\) and an encoder;
[0039] The style-content decoupling method: let the style encoder only extract the picture style, and let the content encoder only extract the picture content;
[0040] The adversarial training method: use adversarial samples to train the model to increase the robustness of the model.
[0041] Preferably, the training system specifically includes the following modules:
[0042] Module M1: Pre-train the style Embedding value of each category and the content Embedding value of each picture;
[0043] Module M2: Randomly sample some pictures from the training dataset;
[0044] Module M3: Generate adversarial samples for each sampled picture;
[0045] Module M4: Calculate the decoupling loss and the reconstruction loss function value \(L\) of the model using the adversarial samples and the original picture samples d ;
[0046] Module M5: Use the total loss function value \(L\) d to perform gradient descent to update the model parameter value \(\Theta\), and the model parameter value \(\Theta\) is the neuron weight value in the convolutional layer and the fully connected layer of the neural network model;
[0047] Module M6: Repeat to execute Module M2 - Module M5 until the total loss function value \(L\) dConverge to obtain a robust content-style decoupling model.
[0048] Preferably, the module M1 includes:
[0049] Module M1.1: Randomly sample a batch of original images;
[0050] Module M1.2: For a batch of images, calculate the following Loss function value:
[0051]
[0052] where x i is the original image, i represents the i-th image in a batch, s i is the style Embedding of the category where x i is located, c i is the content Embedding of x i , G θ is the decoder in the decoupling model, θ is the network parameter value of the decoder, and β is a hyperparameter;
[0053] Module 1.3: According to the Loss function value in Module M1.2, perform gradient descent updates on s i , c i , G θ :
[0054]
[0055]
[0056]
[0057] In the formula, η is the learning rate;
[0058] Module M1.4: Repeat Modules M1.1 - M1.3 until the loss value L p converges.
[0059] Preferably, the module M3 includes:
[0060] Module M3.1: Add random perturbations to the original image, x' i = x i + ∈·ξ, where x' i represents the image after adding perturbations, x i represents the original image, ∈ represents the range of perturbations, and ξ is a uniformly distributed random variable in the interval [-1, 1];
[0061] Module M3.2: Calculate the loss function value of the model, and the loss function is designed as: l adv = ||ES (x′ i ) - E S (x i ) || + || E C (x′ i ) - E C (x i ) ||, where E S and E C represent the style encoder and the content encoder respectively;
[0062] Module M3.3: According to the loss function value in Module M3.2, perform gradient ascent on the image to obtain a new adversarial sample x′ i , and the formula is:
[0063]
[0064] Module M3.4: Repeatedly execute Module M3.2 - Module M3.3 to generate better adversarial samples.
[0065] Preferably, the said Module M4 includes:
[0066] Calculate the reconstruction decoupling loss value of the clean sample x i and its corresponding adversarial sample x′ i through the following formula:
[0067]
[0068] where G θ is the decoder, E s is the style encoder, E c is the content encoder, s i is the style Embedding value obtained in Module M1, and c i is the content Embedding value obtained in Module M1.
[0069] Compared with the prior art, the present invention has the following beneficial effects:
[0070] 1. The method provided by the present invention effectively enhances the robustness of the model and is more resistant to adversarial attacks;
[0071] 2. The method provided by the present invention does not affect the decoupling ability of the model while improving the robustness of the model;
[0072] 3. By adjusting the hyperparameters during training, this method can be applied to other models and other datasets, improving the applicability of this method. Description of the Drawings
[0073] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments read in conjunction with the accompanying drawings:
[0074] Figure 1 It is a schematic diagram of the content-style decoupling model in the embodiment of the present invention;
[0075] Figure 2 It is a comparison chart of the robustness of ordinary training and robust training on the SmallNORB dataset in the embodiment of the present invention;
[0076] Figure 3 It is a comparison chart of the decoupling degree of ordinary training and robust training on the SmallNORB dataset in the embodiment of the present invention;
[0077] Figure 4 It is a comparison chart of the robustness of ordinary training and robust training on the MNIST dataset in the embodiment of the present invention;
[0078] Figure 5 It is a comparison chart of the robustness of ordinary training and robust training on the Celeba dataset in the embodiment of the present invention. Detailed Embodiments
[0079] The present invention will be described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0080] The present invention introduces a training method for a content-style decoupling model based on adversarial training robustness, including: using a style-content decoupling method to train the decoupling ability of the model, and using an adversarial training method to enhance the robustness of the model;
[0081] Referring to Figure 1 , the content-style decoupling model includes a style encoder Es, a content encoder Ec, and a decoder;
[0082] Style encoder: A convolutional neural network that can encode a two-dimensional image into a feature vector containing only the style (category) information of the image. The style encoder usually includes multiple convolutional layers, pooling layers, and a final linear connection layer to convert two-dimensional features into one-dimensional image features. The converted image features only contain style features and do not contain content-related features.
[0083] Content Encoder: A convolutional neural network that can extract feature vectors related to object poses, backgrounds, etc., which are independent of style, from images. The style encoder usually consists of multiple convolutional layers, pooling layers, and a final linear connection layer to convert two-dimensional features into one-dimensional image features. The converted image features only contain content features and do not contain style-related features.
[0084] Decoder: A convolutional neural network that can reconstruct corresponding images from style vectors and content vectors. To generate higher-quality images, we usually use the AdaIN structure to remove the style (normalize) the content features of the original image, and then use the mean and variance of the style of the target image to add style (opposite to normalization), so that higher-quality images can be generated.
[0085] The style-content decoupling method: Let the style encoder only extract the image style, and let the content encoder only extract the image content, and the two encoders are not coupled with each other.
[0086] The adversarial training method: Use adversarial samples to train the model to increase the robustness of the model. Adversarial samples are generated by increasing the loss function of the model through the method of gradient ascent.
[0087] The specific training method is as follows:
[0088] Step S1: Pre-train the style Embedding values for each category and the content Embedding values for each image.
[0089] Furthermore, step S1 includes the following sub-steps:
[0090] Step S1.1: Randomly sample a Batch of original images;
[0091] Step S1.2: For a Batch of images, calculate the following Loss function value:
[0092]
[0093] where x i is the original clean image, i represents the i-th image in a Batch, s i is the style Embedding of the category where x i is located, c i is the content Embedding of x i , G θ is the decoder in the decoupling model, θ is the network parameter value of the decoder, and β is a hyperparameter;
[0094] Step S1.3: According to the Loss function value in step S1.2, for s i, c i , G θ Perform gradient descent update:
[0095]
[0096]
[0097]
[0098] Where η is the learning rate;
[0099] Step S1.4: Repeat steps S1.1 - S1.3 until the loss value L p converges.
[0100] Step S2: Randomly sample some pictures from the training dataset;
[0101] Step S3: Generate adversarial samples for each sampled picture;
[0102] Furthermore, step S3 includes the following sub - steps:
[0103] Step S3.1: Add random perturbation to the original picture, x′ i = x i + ∈·ξ, where x′ i represents the picture after adding perturbation, x i represents the original picture, ∈ represents the range of perturbation, and ξ is a uniformly distributed random variable in the interval [-1, 1];
[0104] Step S3.2: Calculate the value of the loss function of the model. The loss function is designed as: l adv = ||E S (x′ i ) - E S (x i )|| + ||E C (x′ i ) - E C (x i )||, where E S and E C represent the style encoder and the content encoder respectively;
[0105] Step S3.3: According to the value of the loss function in step S3.2, perform gradient ascent on the picture to obtain a new adversarial sample x′ i , and the formula is:
[0106]
[0107] Step S3.4: Repeat steps S3.2 - S3.3 multiple times to generate better adversarial samples.
[0108] Step S4: Calculate the decoupling loss and the reconstruction loss function value L of the model using the adversarial samples and the original image samples d ; Calculate the clean sample x i and its corresponding adversarial sample x′ i using the following formula for the reconstruction decoupling loss value:
[0109]
[0110] where G θ is the decoder, E s is the style encoder, E c is the content encoder, s i is the style Embedding value obtained in Step S1, and c i is the content Embedding value obtained in Step S1.
[0111] Step S5: Use the total loss function value L d to perform gradient descent to update the model parameter value Θ, where the model parameter value Θ is the neuron weight value in the convolutional layer and the fully connected layer of the neural network model;
[0112]
[0113] Step S6: Repeat Step S2 - Step S5 until the total loss function value L d converges to obtain a robust content-style decoupling model.
[0114] The present invention also introduces a training system for a robust content-style decoupling model based on adversarial training, including: using a style-content decoupling system to train the decoupling ability of the model, and using an adversarial training system to enhance the robustness of the model; the content-style decoupling model includes a style encoder Es, a content encoder Ec, and a decoder; the style-content decoupling system: allowing the style encoder to only extract the picture style and the content encoder to only extract the picture content; the adversarial training system: using adversarial samples to train the model to increase the robustness of the model.
[0115] Specifically, the training system specifically includes the following modules:
[0116] Module M1: Pretrain the style Embedding value for each category and the content Embedding value for each picture
[0117] Further, Module M1 includes the following sub-modules:
[0118] Module M1.1: Randomly sample a Batch of original pictures
[0119] Module M1.2: For a batch of images, calculate the value of the following Loss function:
[0120]
[0121] where x i is the original clean image, s i is the style Embedding of the category where x i is located, c i is the content Embedding of x i , G θ is the decoder in the decoupling model, θ is the network parameter value of the decoder, and β is a hyperparameter;
[0122] Module M1.3: According to the value of the Loss function in Module M1.2, perform gradient descent updates on s i , c i , G θ :
[0123]
[0124]
[0125]
[0126] where η is the learning rate;
[0127] Module M1.4: Repeat Modules M1.1 - M1.3 until the loss value L converges.
[0128] Module M2: Randomly sample some images from the training dataset;
[0129] Module M3: Generate adversarial samples for each sampled image;
[0130] Furthermore, Module M3 includes the following sub - modules:
[0131] Module M3.1: Add random perturbations to the original image, x′ i = x i + ∈·ξ, where x′ i represents the image after adding perturbations, x i represents the original image, ∈ represents the range of perturbations, and ξ is a uniformly distributed random variable in the interval [-1, 1];
[0132] Module M3.2: Calculate the value of the loss function of the model. The loss function is designed as: l adv = ||E S (x′ i ) - E S (x i)||+||E C (x′ i )-E C (x i )||, where E S and E C represent the style encoder and the content encoder respectively;
[0133] Module M3.3: According to the loss function value in Module M3.2, perform gradient ascent on the picture to obtain a new adversarial sample x′ i , and the formula is:
[0134]
[0135] Module M3.4: Repeat Module M3.2 - Module M3.3 multiple times to generate better adversarial samples.
[0136] Module M4: Calculate the decoupling loss and the reconstruction loss function value L of the model using the adversarial sample and the original picture sample d ; Through the following formula, calculate the reconstruction decoupling loss value of the clean sample x i and its corresponding adversarial sample x′ i :
[0137]
[0138] where G θ is the decoder, E s is the style encoder, E c is the content encoder, s i is the style Embedding value obtained in Module M1, and c i is the content Embedding value obtained in Module M1.
[0139] Module M5: Use the total loss function value L d to perform gradient descent to update the model parameter value Θ, and the model parameter value Θ is the neuron weight value in the convolutional layer and the fully connected layer of the neural network model;
[0140]
[0141] Module M6: Repeat Modules M2 - M5 until the total loss function value L d converges to obtain a robust content-style decoupling model.
[0142] This method improves the robustness of the model. As Figure 2As shown, the dataset used is Small NORB. On the left are the results of ordinary training, and on the right are the training results of this method. When the model trained ordinarily encounters adversarial samples, it cannot accurately reconstruct the corresponding images, such as the reconstruction errors marked in the red boxes in the images. However, the model trained by this method can well defend against the attacks of adversarial samples.
[0143] Compared with the prior art, this method does not affect the decoupling ability of the model while improving the robustness of the model, as Figure 3 shown, both the ordinary training method and this method can exhibit considerable decoupling ability.
[0144] This method can be applied to other models and other datasets by simply slightly adjusting the hyperparameters during training, as Figure 4 shown, this method can also improve the robustness of the model on the MNIST dataset, as Figure 5 shown, this method can also improve the robustness of the model on Celeba. The effects on these two other datasets illustrate that this method has strong enough generalization ability to handle tasks in different fields.
[0145] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as either software modules for implementing the method or structures within the hardware component.
[0146] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.
Claims
1. A training method for a content-style decoupling model robust based on adversarial training, characterized in that, Including: Adopt a style-content decoupling method to train the decoupling ability of the model, and adopt an adversarial training method to enhance the robustness of the model; The content-style decoupling model includes a style encoder Es, a content encoder Ec, and a decoder; The style-content decoupling method: Let the style encoder only extract the image style, and let the content encoder only extract the image content; The adversarial training method: Use adversarial samples to train the model to increase the robustness of the model; The training method specifically includes the following steps: Step S1: Pre-train the style Embedding value of each category and the content Embedding value of each image; Step S2: Randomly sample some images from the training dataset; Step S3: Generate adversarial samples for each sampled image; Step S4: Calculate the decoupling loss and the reconstruction loss function value L of the model using the adversarial samples and the original image samples d ; Step S5: Use the total loss function value L d to perform gradient descent to update the model parameter value Θ, where the model parameter value Θ is the neuron weight value in the convolutional layer and the fully connected layer of the neural network model; Step S6: Repeat Step S2 - Step S5 until the total loss function value L d converges to obtain a robust content-style decoupling model.
2. The training method of the content-style decoupling model based on adversarial training robustness according to claim 1, characterized in that: The step S2 includes: Step S1.1: Randomly sample a Batch of original images; Step S1.2: For a Batch of images, calculate the following Loss function value: where x i is the original image, i represents the i-th image in a batch, and s i is the style Embedding of the category where x i is located, c i is the content Embedding of x i , G θ is the decoder in the decoupling model, θ is the network parameter value of the decoder, and β is the hyperparameter; Step S1.3: According to the Loss function value in Step S1.2, perform gradient descent updates on s i , c i , G θ : In the formula, η is the learning rate; Step S1.4: Repeat steps S1.1 - S1.3 until the loss value L p converges.
3. The training method of the content-style decoupling model based on adversarial training robustness according to claim 1, wherein: The step S3 includes: Step S3.1: Add random perturbations to the original image, x′ i = x i + ∈·ξ, where x′ i represents the image after adding perturbations, x i represents the original image, ∈ represents the range of perturbations, and ξ is a uniformly distributed random variable in the interval [-1, 1]; Step S3.2: Calculate the loss function value of the model. The loss function is designed as: l adv =‖E S (x i ′ ) - E S (x i )‖ + ‖E C (x i ′ ) - E C (x i )‖, where E S and E C represent the style encoder and the content encoder respectively; Step S3.3: According to the loss function value in step S3.2, perform gradient ascent on the image to obtain a new adversarial sample x i ′ , and the formula is: Step S3.4: Repeat step S3.2 - step S3.3 multiple times to generate better adversarial samples.
4. The method for training a content-style decoupling model robust based on adversarial training according to claim 1, wherein: The step S4 includes: The clean sample x is calculated through the following formula i and its corresponding adversarial sample x i ′ for the reconstruction decoupling loss value: Among them, G θ is the decoder, E s is the style encoder, E c is the content encoder, s i is the style Embedding value obtained in step S1, c i is the content Embedding value obtained in step S1.
5. A training system for a content-style decoupling model robust based on adversarial training, characterized in that Including: Adopt a style-content decoupling method to train the decoupling ability of the model, and adopt an adversarial training method to enhance the robustness of the model; The content-style decoupling model includes a style encoder Es, a content encoder Ec, and a decoder; The style-content decoupling method: Let the style encoder only extract the image style, and let the content encoder only extract the image content; The adversarial training method: Use adversarial samples to train the model to increase the robustness of the model; The training system specifically includes the following modules: Module M1: Pre-train the style Embedding value of each category and the content Embedding value of each image; Module M2: Randomly sample some images from the training dataset; Module M3: Generate adversarial samples for each sampled image; Module M4: Calculate the decoupling loss and the reconstruction loss function value L of the model using adversarial samples and original image samples d ; Module M5: Use the total loss function value L d to perform gradient descent to update the model parameter values Θ, where the model parameter values Θ are the neuron weight values in the convolutional layer and fully connected layer of the neural network model; Module M6: Repeatedly execute Module M2 - Module M5 until the total loss function value L d converges to obtain a robust content-style decoupling model.
6. The content-style decoupling model training system based on adversarial training robustness according to claim 5, characterized in that: The module M1 includes: Module M1.1: Randomly sample a Batch of original images; Module M1.2: For a Batch of images, calculate the following Loss function value: where x i is the original image, i represents the i-th image in a batch, and s i is the style Embedding of the category where x i is located, and c i is the content Embedding of x i , G θ is the decoder in the decoupling model, θ is the network parameter value of the decoder, and β is the hyperparameter; Module 1.3: According to the Loss function value in Module M1.2, perform gradient descent updates on s i , c i , G θ : In the formula, η is the learning rate; Module M1.4: Repeat Modules M1.1 - M1.3 until the loss value L p converges.
7. The content-style decoupling model training system based on adversarial training robustness according to claim 5, characterized in that: The module M3 includes: Module M3.1: Add random perturbations to the original image, x′ i = x i + ∈·ξ, where x′ i represents the image after adding perturbations, x i represents the original image, ∈ represents the range of perturbations, and ξ is a uniformly distributed random variable in the interval [-1, 1]; Module M3.2: Calculate the loss function value of the computational model. The loss function is designed as: l adv = ‖E S (x′ i ) - E S (x i )‖ + ‖E C (x′ i ) - E C (x i )‖, where E S and E C represent the style encoder and the content encoder respectively; Module M3.3: Ascend the gradient of the image based on the loss function value in Module M3.2 to obtain a new adversarial sample x i ′ , and the formula is: Module M3.4: Repeat the execution of module M3.2 - module M3.3 to generate better adversarial samples.
8. The training system for the content-style decoupling model based on adversarial training robustness according to claim 5, characterized in that: The module M4 includes: The clean sample x is calculated by the following formula i and its corresponding adversarial sample x i ′ for the reconstruction decoupling loss value: Among them, G θ is the decoder, E s is the style encoder, E c is the content encoder, s i is the style Embedding value obtained in module M1, c i is the content Embedding value obtained in module M1.
Citation Information
Patent Citations
Semi-supervised multi-modal multi-class image translation method
CN110263865A
Adversarial training method and device and application method and device of neural network model
CN112035834A