Training method and device of generative adversarial network, equipment and storage medium

By introducing fine-grained latent features and reversible learning strategies into the generative adversarial network, the problem of cGANs relying on coarse-grained pseudo-labels is solved, and high-quality and accurate conditions are achieved image generation.

CN120579583APending Publication Date: 2025-09-02BEIHANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510693440.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing conditional generation adversarial network (cGANs) training methods rely on coarse-grained pseudo-labels, resulting in poor generation results and difficult to achieve high-quality and accurate image generation.

Method used

By acquiring real images and enhancing images, using the combination of generators and discriminators, the loss function is determined based on real hidden space codewords, enhanced hidden space codewords and false hidden space codewords, and the network parameters are updated when the loss function does not meet the convergence conditions until the convergence conditions are met, and the generation process is controlled using fine-grained potential features and reversible learning strategies.

Benefits of technology

It realizes the generation of high-quality and accurate images under unsupervised conditions, overcomes the limitations of coarse-grained pseudo-labels, and improves the generation effect of the generative adversarial network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579583A_ABST
    Figure CN120579583A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a training method and device for a generative adversarial network, equipment and a storage medium. The method comprises the following steps: acquiring a real image and an enhanced image; obtaining generator input data; inputting the generator input data into a generator for image generation to obtain a false image output by the generator; inputting the real image, the enhanced image and the false image into a discriminator for coding to obtain a real hidden space codeword corresponding to the real image, an enhanced hidden space codeword corresponding to the enhanced image and a false hidden space codeword corresponding to the false image output by the discriminator; determining a loss function of the generative adversarial network based on the real hidden space codeword, the enhanced hidden space codeword and the false hidden space codeword; and if the loss function does not meet the convergence condition, updating network parameters of the generative adversarial network, and returning to the step of inputting the generator input data into the generator for image generation to obtain a false image output by the generator until the updated loss function meets the convergence condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of neural network technology, and in particular to a training method, apparatus, device, and storage medium for generating adversarial networks. Background Art

[0002] Generative Adversarial Nets (GANs) allow for the generation of high-quality and realistic images from low-dimensional latent spaces and have become a mainstay of deep generative models. Conditional Generative Adversarial Nets (cGANs) further provide a viable approach to controlling generation by constructing a continuous latent space conditioned on predefined auxiliary information. Training cGANs is closely tied to the quality of the conditioning dataset, and manual labeling is both expensive and time-consuming. Unsupervised conditional generation, which enables conditional generation without the need for expensive manual labeling, has attracted increasing research attention.

[0003] Currently, most methods predict pseudo-labels before training cGANs, and then use the pseudo-labels to train cGANs. However, due to the coarse granularity or inaccuracy of pseudo-labels, the training effect of cGANs is poor, making it difficult to achieve high-quality generation with accurate conditions. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a training method, apparatus, device and storage medium for generating an adversarial network.

[0005] A first aspect of an embodiment of the present disclosure provides a training method for a generative adversarial network, wherein the generative adversarial network includes a discriminator and a generator, wherein the method includes:

[0006] Acquire a real image and an enhanced image corresponding to the real image;

[0007] Get generator input data;

[0008] Inputting the generator input data into the generator to generate an image, thereby obtaining a false image output by the generator;

[0009] Inputting the real image, the enhanced image, and the false image into the discriminator for encoding, and obtaining a real latent space codeword corresponding to the real image, an enhanced latent space codeword corresponding to the enhanced image, and a false latent space codeword corresponding to the false image output by the discriminator;

[0010] Determining a loss function of the generative adversarial network based on the true latent space codeword, the enhanced latent space codeword, and the false latent space codeword;

[0011] If the loss function does not meet the convergence conditions, update the network parameters of the generative adversarial network, and return to the step of inputting the generator input data into the generator for image generation to obtain a false image output by the generator, until the updated loss function meets the convergence conditions.

[0012] A second aspect of an embodiment of the present disclosure provides a training apparatus for a generative adversarial network, wherein the generative adversarial network includes a discriminator and a generator, wherein the apparatus includes:

[0013] A first acquisition module, configured to acquire a real image and an enhanced image corresponding to the real image;

[0014] The second acquisition module is used to obtain generator input data;

[0015] A first input module, configured to input the generator input data into the generator to generate an image, and obtain a false image output by the generator;

[0016] A second input module is configured to input the real image, the enhanced image, and the false image into the discriminator for encoding, and obtain, output by the discriminator, a real latent space codeword corresponding to the real image, an enhanced latent space codeword corresponding to the enhanced image, and a false latent space codeword corresponding to the false image;

[0017] A first determination module is configured to determine a loss function of the generative adversarial network based on the true latent space codeword, the enhanced latent space codeword, and the false latent space codeword;

[0018] The first update module is used to update the network parameters of the generative adversarial network if the loss function does not meet the convergence conditions, and return to the step of inputting the generator input data into the generator for image generation to obtain a false image output by the generator, until the updated loss function meets the convergence conditions.

[0019] A third aspect of an embodiment of the present disclosure provides an electronic device, which includes: a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the method of the first aspect above.

[0020] A fourth aspect of an embodiment of the present disclosure provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of the first aspect described above can be implemented.

[0021] The technical solution provided by the embodiments of the present disclosure has the following advantages over the prior art:

[0022] The disclosed embodiment can obtain a real image and an enhanced image corresponding to the real image; obtain generator input data; input the generator input data into the generator for image generation, and obtain a false image output by the generator; input the real image, the enhanced image, and the false image into the discriminator for encoding, and obtain the real latent space codeword corresponding to the real image output by the discriminator, the enhanced latent space codeword corresponding to the enhanced image, and the false latent space codeword corresponding to the false image; determine the loss function of the generative adversarial network based on the real latent space codeword, the enhanced latent space codeword, and the false latent space codeword; if the loss function does not meet the convergence condition, update the network parameters of the generative adversarial network, and return to the step of inputting the generator input data into the generator for image generation to obtain the false image output by the generator, until the updated loss function meets the convergence condition. The above technical solution takes into account that the latent space of the generative adversarial network is inherently rich in semantics and can cover the different semantics reflected by pseudo-labels. Therefore, the generation can be controlled by hints from the latent space rather than pseudo-labels. Specifically, in order to overcome coarse-grained pseudo-labels, the image is represented by latent features with fine-grained semantics (i.e., real latent space codewords, enhanced latent space codewords, and false latent space codewords) to achieve accurate conditional generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0024] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0025] Figure 1 is a flowchart of a training method for a generative adversarial network provided by an embodiment of the present disclosure;

[0026] Figure 2 Schematic diagram of a generative adversarial network provided by an embodiment of the present disclosure;

[0027] Figure 3 1 is a schematic diagram of the structure of a training device for a generative adversarial network provided by an embodiment of the present disclosure;

[0028] Figure 4 It is a structural diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0030] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0031] Generative Adversarial Networks (GANs), which allow for the generation of high-quality and realistic images from a low-dimensional latent space, have become a mainstay of deep generative models. Original GANs can only generate images from pure random noise, and methods for controlling generation rely on post-processing within the latent space. cGANs offer a viable approach to controlling generation by constructing a continuous latent space conditioned on predefined auxiliary information. Most cGANs focus on category generation, where the auxiliary information considers categories and attributes, and have recently been extended to other modalities, including text, styles, and sketches. Because cGANs control generation based on auxiliary information, they essentially learn the joint distribution of images and auxiliary information. The applicants' research has found that current cGAN training still heavily relies on manually annotated image labels, which are expensive and often difficult to obtain. To address the bottleneck of requiring large-scale annotated conditional datasets, unsupervised conditional generation has emerged. It attempts to generate images based on implicit representations or underlying data structures, rather than relying on manual labels. One approach uses a feature extractor to generate guided features from the data, then clusters these features and treats the generated clusters as pseudo-labels to guide the conditional generation process. While this approach reduces reliance on manual annotation, its performance is still limited by the quality of feature extraction and clustering accuracy. Another approach employs an end-to-end learning strategy, directly optimizing the latent representation and conditional generation within a single framework. In these models, an additional encoder is used to output pseudo-labels during training, which are then concatenated with the underlying noise to generate conditional images. However, the pseudo-labels are coarse-grained and may not capture the fine-grained semantics required for high-quality, discriminative image generation. Yet another approach employs multiple generator frameworks to handle different pseudo-labels and improve the conditional generation process. Although the pseudo-labels reflect different semantics, the randomness in the conditional generation process is caused by the underlying noise, which is unrelated to the pseudo-labels during learning. This irrelevance hinders effective control of generation, resulting in insufficient discriminability and high-quality generation across labels. This hinders precise control of high-quality unsupervised conditional generation. In summary, the pseudo-labels used in related art are relatively coarse-grained or inaccurate, resulting in poor training results for cGANs and difficulty in achieving high-quality generation with accurate conditions. To address this issue, the present disclosure proposes a training method, apparatus, device, and storage medium for a generative adversarial network. Below, we first explain the training method of the generative adversarial network in detail.

[0032] Figure 1 This is a flow chart of a training method for a generative adversarial network provided by an embodiment of the present disclosure. The method can be executed by an electronic device. The electronic device is deployed with a generative adversarial network, which includes a discriminator and a generator. The electronic device can be exemplarily understood as a device such as a mobile phone, tablet computer, laptop computer, desktop computer, smart TV, etc. Figure 1As shown, the method provided in this embodiment includes the following steps:

[0033] S110: Acquire a real image and an enhanced image corresponding to the real image.

[0034] Specifically, real images are real-world data samples collected from the target distribution.

[0035] For example, the target distribution is P r , from P r Sampling is performed to obtain b d A real image, denoted as That is, use x to uniformly represent the real image, and use x to distinguish different real images. i Represents the i-th real image.

[0036] Specifically, the enhanced image is an image obtained by enhancing a real image.

[0037] Optionally, the enhancement operation may include at least one of the following: resizing, horizontal flipping, color dithering, and cropping.

[0038] For example, Perform an enhancement operation to obtain a corresponding enhanced image, recorded as And Perform another enhancement operation to obtain the corresponding enhanced image, recorded as

[0039] S120. Obtain generator input data.

[0040] Specifically, the generator input data is the input data used by the generator to generate fake images.

[0041] Specifically, there are many ways to obtain generator input data, which is not limited in this disclosure.

[0042] Optionally, S120 includes: S121, generating random semantics through a binary selector.

[0043] Specifically, a binary selector determines the selection of an option by generating a random binary value (0 or 1).

[0044] Specifically, a set of grammatical rules or semantic frameworks can be predefined, which determine the structure of the generated random semantics. A binary selector is used to select options, and the selected main options are combined to form a complete random semantics.

[0045] S122. Generate noise by random Gaussian sampling.

[0046] Specifically, there are many ways to obtain generator input data, which is not limited in this disclosure.

[0047] For example, the binary Gaussian noise is P N , from P N Sampling is performed to obtain b a random variables, denoted as From P N Sampling is performed to obtain b w random variables, denoted as Through affine transformation Get b w noise, among which It is an independent network used to obtain the affine transformation. It can be understood that by introducing the affine transformation, the original random variable can be enhanced or adjusted to make it more suitable for image generation.

[0048] S123. Binary encode the random semantics and noise to obtain generator input data.

[0049] For example, random semantics and Encode and get b g Generator input data, denoted as Ready-to-use c Unified representation of generator input data, in order to distinguish different generator input data, use zcm Represents the mth generator input data.

[0050] It can be understood that based on the fine-grained representation, the discriminator in the embodiment of the present disclosure encodes the input image into latent features (i.e., real latent space codewords, enhanced latent space codewords, and false latent space codewords) to explicitly cover different semantics in the unlabeled image. Therefore, a new binary structure latent space is proposed, which explicitly encodes fine-grained latent features (i.e., generator input data) through a hybrid combination of deterministic binary selectors and random Gaussian sampling. This architecture supports unsupervised discovery of interpretable features while maintaining the flexibility of continuous latent representation. More specifically, we transform the latent variable z c ∈R d Divided into n independent segments:

[0051] z c =[z1,z2,…,z k ,…,z k ] T ,z k ∈R d / n Formula (1)

[0052] Where n is divisible by d, and d is z c Then, a composite random variable is proposed instead of a simple Gaussian distribution or a mixed Gaussian distribution, where each segment zk By a binary selector b k ∈{0,1} control, which determines z k Sampling distribution:

[0053]

[0054] Where c is the number of labels, Represents rounding down, separation between μ and different semantics, and σ 2 The intra-class diversity of control is related. Formula (2) is a formula for converting decimal to binary. For example, if c is 1, the corresponding binary is 0001, and the corresponding b1=1, b2=0, b3=0, b4=0. Formula (3) represents z k is from the mean The variance is σ 2 Normal distribution sampling of I, σ 2 represents the variance of the normal distribution, and I is an n-dimensional identity matrix. Thus, b k = 0 corresponds to a positive mean, b k =1 corresponds to a negative mean. The larger μ is, the more distinct the semantics are, but the generation quality will deteriorate; σ 2 The greater the diversity, the better, but the generated quality will also be worse.

[0055] This design introduces a structured prior, where each binary dimension corresponds to a semantic attribute. In order to ensure that the composite random variables still satisfy the independent and identically distributed property, the disentanglement property between semantics can be maximized. Therefore, the probability density function p m (z c ) can be expressed as:

[0056]

[0057] If the number of labels c is a power of 2, the latent features are independent and identically distributed and satisfy Because c is a power of 2, according to the property of taking the remainder 2, so b k The value of 0 or 1 in all dimensions is independent; that is, b in each dimension k are all standard Bernoulli random variables β, satisfying independent and identical distribution; given another independent and identically distributed standard Gaussian distribution N, then z k It can be expressed as (β+N-0.5)×2, that is, z k are also independent and identically distributed, and we can get

[0058] It is understandable that Gaussian mixture models (GMMs) have been widely used in unsupervised conditional generation tasks. All dimensions in a GMM are used to locate clusters, which leads to the interweaving of semantics across dimensions when controlling generation. Especially for unsupervised settings, pseudo-labels may change during training. This results in an unstructured latent space for the GMM, where each cluster has no clear semantics. Therefore, the binary structured latent space of the present disclosure has advantages over the GMM.

[0059] S130. Input the generator input data into the generator to generate an image, and obtain a false image output by the generator.

[0060] Specifically, the fake image is data generated by the generator based on the generator input data.

[0061] For example, Input the generator and get multiple fake images, recorded as Among them, θ g are the network parameters of the generator.

[0062] S140. Input the real image, enhanced image and false image into the discriminator for encoding, and obtain the real latent space codeword corresponding to the real image, the enhanced latent space codeword corresponding to the enhanced image and the false latent space codeword corresponding to the false image output by the discriminator.

[0063] Specifically, the dimensions of the output of the discriminator and the input of the generator are the same, that is, the dimensions of the real latent space codeword, the enhanced latent space codeword, and the fake latent space codeword are the same.

[0064] For example, the real image Input discriminator, the discriminator is trained on real images Encode and get the real latent space codeword, recorded as The false image Input discriminator, the discriminator is good at distinguishing fake images Encode and get the false latent space codeword, recorded as The first enhanced image Input discriminator, the discriminator is for the first enhanced image Encode and get the first enhanced latent space codeword, denoted as The second enhanced image Input discriminator, the discriminator is for the second enhanced image Encode and get the second enhanced latent space codeword, denoted as Among them, θ d are the network parameters of the discriminator.

[0065] It is understandable that images basically reside on a low-dimensional semantic manifold, where the architecture of GAN enables the latent space to generally have rich semantics in very low dimensions. Therefore, the latent space provides a well-performing option to represent images through fine-grained semantics. Given that GANs usually generate images unidirectionally from the latent space, this application first introduces reversible learning to reversibly transform real images, enhanced images, and false images back to the latent space. In order to achieve conditional generation of discriminative latent features (i.e., real latent space codewords, enhanced latent space codewords, and false latent space codewords), a reversible encoder is incorporated into the discriminator, thereby establishing an adversarial reversible learning strategy. More specifically, this application provides a discriminator and generators A dual role is established, in other words, as the basic role, Used to distinguish real images from fake images, Used to generate fake images, in addition, The input images (i.e., real images, enhanced images, and fake images) are encoded into a latent space with Together we achieve the reversibility of bijective mapping.

[0066] S150. Determine a loss function of a generative adversarial network based on the true latent space codeword, the enhanced latent space codeword, and the false latent space codeword.

[0067] Optionally, S150 includes: S151, determining a reversible loss function based on generator input data and false latent space codewords.

[0068] Optionally, S151 includes: calculating the L2 norm of the generator input data and the false latent space codeword; calculating the square of the L2 norm to obtain a reversible loss function.

[0069] For example, the reversible loss function is calculated by the following formula (5):

[0070]

[0071] It can be understood that, considering the semantics in the latent space, the difference in the latent space has a high-level meaning in revealing the reversible difference compared to the difference between images. Therefore, a reversible function (i.e., reversible difference) is established in the latent space, as shown in Formula (5). cm is obtained by random sampling, so even if the real image is not required, L recip Can also be optimized. By optimizing L recip , discriminator It can accurately map the input image back to the latent space. More importantly, due to the bijective property, if the generator and the discriminator are on the same support, and Construct a bijective mapping to recover fine-grained semantics, then is also bijective so that the input image can be reconstructed. In addition, when When it is possible to generate realistic fake images, that is, when they fall into the same support as real images from the real world, Able to reconstruct real-world images. This feature makes the generator in this application It can generate realistic images and accurately reconstruct real-world images. Optimizing the reversible difference is essentially maximizing the mutual information between the latent features and the generated image, thereby promoting the controllability of unsupervised conditional generation. It should also be pointed out that compared with InfoGAN, this application implicitly maximizes the reversible difference by optimizing the reversible difference. Without the need for an additional variational network. Moreover, the motivation of this application to obtain fine-grained representation is fundamentally different from that of InfoGAN, which relies on coarse-grained pseudo-labels. The reversible difference proposed in this application is consistent with its maximization of mutual information. Specifically, assuming that the discriminator The noise estimate for its coding is consistent with Represents the variational probability distribution, which is a conditional distribution. Indicates that this variational probability distribution conforms to the normal distribution we defined, then z c and the corresponding false image Mutual information The following inequality about reversible differences is satisfied:

[0072]

[0073] Among them, const represents a fixed value.

[0074] In information theory, mutual information Equivalent to the joint probability distribution and the product of its marginal distribution The Kullback-Leibler divergence is:

[0075]

[0076] in, express and The Kullback-Leibler divergence between two distributions further expresses An expectation (e.g., mean) of .

[0077] Since the joint probability distribution is usually unknowable, the above formula (7) can be rewritten as:

[0078]

[0079] Furthermore, since the discriminator The generated picture Re-encoding back to the latent space can be considered as a process of z c The noisy estimate of , so the variational probability distribution can be introduced Then we can get the following formula (9):

[0080]

[0081] Among them, H(z c ) is z c Information entropy.

[0082] S152. Determine an enhanced loss function based on the enhanced latent space codeword.

[0083] Optionally, the generative adversarial network further includes a projection network, wherein determining the enhancement loss function based on the enhanced latent space codeword includes: inputting the enhanced latent space codeword into the projection network for projection to obtain a projected enhanced latent space codeword;

[0084] Calculate the first similarity of a positive sample pair, where the positive sample pair is the projected enhanced latent space codeword corresponding to two different enhanced images of the same real image;

[0085] Calculate the second similarity of the negative sample pair, where the negative sample pair is the projected enhanced latent space codeword corresponding to the enhanced image of two different real images;

[0086] An enhancement loss function is determined based on the first similarity and the second similarity.

[0087] Specifically, when calculating the enhancement loss function, this application relies on the principle that different enhanced images of the same real image have similar representations, and uses the discriminator as a contrast feature extractor. In addition, the contrast loss only operates on the real image and its enhancement, without putting the generated samples into the enhancement pipeline. This design ensures that the generator is purely focused on generating realistic images, while the discriminator is trained to improve the ability to represent discriminative semantics. More specifically, for each real image, two enhanced images are generated through random transformations (such as random cropping and color distortion). The enhanced images are first mapped by the discriminator and then by a simple projection network. Mapping is performed to calculate the contrast loss (i.e., the enhancement loss function). In addition, the binary potential prior is used as a fixed semantic anchor, and the enhanced images corresponding to the same image are encoded as the same binary pseudo codeword. Therefore, the enhancement loss function can be calculated by the following formula (10):

[0088]

[0089] in, is the projected enhanced latent space codeword corresponding to the first enhanced image corresponding to the i1-th real image, is the projected enhanced latent space codeword corresponding to the second enhanced image corresponding to the i2th real image, 2bg represents twice the batch size, τ is the temperature parameter, is an indicator function, which takes the value 1 when i1≠u, otherwise it takes the value 0. The function is to exclude the sample itself (i.e. and similarity).

[0090] The numerator of formula (10) represents the similarity score of the positive sample pairs, which are usually generated by different enhanced versions of the same real image (such as color dithering, cropping, etc.). The denominator represents the sum of the similarity scores of all negative sample pairs, which are similarities between different real images. The goal of the enhancement loss function is to maximize the similarity of positive sample pairs (the larger the numerator, the better). Minimize the similarity of negative sample pairs (the smaller the denominator, the better).

[0091] S153. Determine a discriminator loss function of the discriminator based on the real latent space codeword, the false latent space codeword, and the generator input data.

[0092] Optionally, S153 includes: calculating a first distance between the real image and the generator input data in the latent space based on the real latent space codeword and the generator input data;

[0093] Based on the fake latent space codeword and the generator input data, calculating the second distance between the fake image and the generator input data in the latent space;

[0094] A discriminator loss function is determined based on the first distance and the second distance.

[0095] Specifically, the discriminator loss function can be calculated by the following formula (11) and formula (12):

[0096]

[0097] in, represents the characteristic function of A, represents the conjugate of Φ1, represents the characteristic function of B, represents the conjugate of Φ2, F W express Cumulative distribution function. For calculating the first distance, Used to calculate the second distance, formula (12) represents the difference between A and B. Discriminator The goal is to maximize That is, make the distance between the real image and the generator input data as large as possible (indicating that they are different), and at the same time make the distance between the fake image and the generator input data as small as possible (indicating that they are similar, and actually hope that they are close to the real image).

[0098] S154. Determine a generator loss function of the generator based on the real latent space codeword and the false latent space codeword.

[0099] Optionally, S154 includes: calculating a third distance between the real image and the false image in the latent space based on the real latent space codeword and the false latent space codeword, and using the third distance as the generator loss function.

[0100] Specifically, the generator loss function can be calculated by the following formula (13) and the formula (12) above:

[0101]

[0102] Among them, formula (13) is used to calculate the third distance, and the goal of the generator is to minimize L G , that is, making the fake image as close as possible to the real image.

[0103] S155. Determine the loss function of the generative adversarial network based on the reversible loss function, the enhancement loss function, the discriminator loss function, and the generator loss function.

[0104] Specifically, the loss function of the generative adversarial network is calculated by the following formula (14):

[0105] L=L D +λL recip +θL aug +L g Formula (14)

[0106] Among them, λ is the reversible regularization parameter and γ is the enhancement factor.

[0107] It is understandable that this application follows the design of incorporating anchor points in the discriminator and utilizes z in the dynamic training process. c This design not only stabilizes training and improves generation quality, but also enables the discriminator to effectively map real data to the associated domain, which is crucial for ensuring that the discriminator effectively maintains the consistency between the semantic information in the image data and the structural semantics of the latent space.

[0108] S160. If the loss function does not meet the convergence conditions, update the network parameters of the generative adversarial network and return to the step of inputting the generator input data into the generator for image generation to obtain a false image output by the generator until the updated loss function meets the convergence conditions.

[0109] Specifically, updating the network parameters of the generative adversarial network includes: updating the network parameters θ of the discriminator d , the network parameters θ in the affine transformation w and the network parameters θ of the generator g .

[0110] For example, the Adam optimizer is used to update the network parameters of the generated adversarial network. Among them, l r is the learning rate.

[0111] In summary, this application proposes a new architecture to study fine-grained latent features in the latent space so that different semantics between possible classes can be accurately described. Fine-grained latent features are first extracted through a reversible learning strategy that seamlessly combines the encoder with the discriminator. Then, the fine-grained features are encoded using binary codes, which ensures different semantics during the generation process. An enhanced contrastive learning method is also developed to improve the accuracy and quality of unsupervised conditional generation, providing a new paradigm for unsupervised conditional generation. More specifically, as Figure 2 As shown, the discriminator essentially performs adversarial encoding, which helps align the distribution (to achieve realistic and high-quality generation), as well as represent fine-grained semantic features (to capture accurate and discriminative conditions inherited in unlabeled images). In other words, the discriminator aims to maximize the false latent space codeword and z c The distance between the real latent space codeword and z c The distance between the fake image and the real image is used to distinguish the fake image from the real image by minimizing the fake latent space codeword and z c to satisfy the reversible learning strategy. In addition, the latent features obtained from the enhanced image are optimized for consistency through the projection network. The generator aims to achieve high-quality generation by minimizing the difference between the false latent space codewords and the real latent space codewords. It can be seen that the present application represents the image by latent features with fine-grained semantics to achieve accurate conditional generation. This is supported by reversible learning, which constructs the discriminator as an adversarial encoder to output discriminative latent features. And further proposed a new binarization strategy to encode the latent features, which retains the independent and identically distributed properties for noise sampling. More importantly, the encoded latent features can effectively highlight the possible different conditions in fine-grained semantics. And a contrastive learning strategy is also developed to further enhance the discriminator's ability to automatically represent real-world images through pseudocodes.

[0112] Figure 3This is a schematic diagram of the structure of a training device for a generative adversarial network provided by an embodiment of the present disclosure. The training device for a generative adversarial network can be understood as the above-mentioned electronic device or a part of the functional modules in the above-mentioned electronic device. Figure 3 As shown, the training device for generating an adversarial network includes:

[0113] The generative adversarial network includes a discriminator and a generator, wherein the device includes:

[0114] A first acquisition module 310 is configured to acquire a real image and an enhanced image corresponding to the real image;

[0115] The second acquisition module 320 is used to obtain generator input data;

[0116] A first input module 330 is configured to input the generator input data into the generator for image generation, thereby obtaining a false image output by the generator;

[0117] A second input module 340 is configured to input the real image, the enhanced image, and the false image into the discriminator for encoding, and obtain, as output by the discriminator, a real latent space codeword corresponding to the real image, an enhanced latent space codeword corresponding to the enhanced image, and a false latent space codeword corresponding to the false image;

[0118] A first determination module 350 is configured to determine a loss function of the generative adversarial network based on the true latent space codeword, the enhanced latent space codeword, and the false latent space codeword;

[0119] The first update module 360 ​​is used to update the network parameters of the generative adversarial network if the loss function does not meet the convergence conditions, and return to the step of inputting the generator input data into the generator for image generation to obtain a false image output by the generator, until the updated loss function meets the convergence conditions.

[0120] Optionally, the second acquisition module 320 is specifically configured to generate random semantics through a binary selector;

[0121] Generate noise by random Gaussian sampling;

[0122] The random semantics and the noise are binary-encoded to obtain the generator input data.

[0123] Optionally, the first determination module 350 includes: a first determination submodule, configured to determine a reversible loss function based on the generator input data and the false latent space codeword;

[0124] A second determination submodule, configured to determine an enhanced loss function based on the enhanced latent space codeword;

[0125] A third determination submodule is configured to determine a discriminator loss function of the discriminator based on the true latent space codeword, the false latent space codeword, and the generator input data;

[0126] a fourth determination submodule, configured to determine a generator loss function of the generator based on the true latent space codeword and the false latent space codeword;

[0127] A fifth determination submodule is used to determine the loss function of the generative adversarial network based on the reversible loss function, the enhancement loss function, the discriminator loss function and the generator loss function.

[0128] Optionally, the first determination submodule is specifically configured to calculate the L2 norm of the generator input data and the false latent space codeword;

[0129] The square of the L2 norm is calculated to obtain the reversible loss function.

[0130] Optionally, the generative adversarial network further includes a projection network, wherein the second determination submodule is specifically configured to input the enhanced latent space codeword into the projection network for projection to obtain the projected enhanced latent space codeword;

[0131] Calculating a first similarity of a positive sample pair, wherein the positive sample pair is the projected enhanced latent space codeword corresponding to two different enhanced images of the same real image;

[0132] Calculating a second similarity of a negative sample pair, wherein the negative sample pair is the projected enhanced latent space codeword corresponding to the enhanced images of two different real images;

[0133] The enhancement loss function is determined based on the first similarity and the second similarity.

[0134] Optionally, a third determination submodule is specifically configured to calculate a first distance between the real image and the generator input data in the latent space based on the real latent space codeword and the generator input data;

[0135] Calculating a second distance between the false image and the generator input data in the latent space based on the false latent space codeword and the generator input data;

[0136] The discriminator loss function is determined based on the first distance and the second distance.

[0137] Optionally, the fourth determination submodule is specifically used to calculate a third distance between the real image and the false image in the latent space based on the real latent space codeword and the false latent space codeword, and use the third distance as the generator loss function.

[0138] The device provided in this embodiment can execute the method of any of the above embodiments, and its execution method and beneficial effects are similar, which will not be repeated here.

[0139] An embodiment of the present disclosure further provides an electronic device, comprising: a memory storing a computer program; and a processor for executing the computer program. When the computer program is executed by the processor, the method of any of the above embodiments can be implemented.

[0140] For example, Figure 4 This is a schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 4 , which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present disclosure. The electronic device 400 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0141] like Figure 4 As shown, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0142] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0143] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0144] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0145] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0146] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0147] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method described in any one of the above embodiments.

[0148] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0150] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0151] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0152] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0153] The embodiments of the present disclosure further provide a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of any of the above embodiments can be implemented. The execution method and beneficial effects are similar and will not be repeated here.

[0154] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0155] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. A training method for generating an adversarial network, characterized in that: The generative adversarial network includes a discriminator and a generator, wherein the method includes: Acquire a real image and an enhanced image corresponding to the real image; Get generator input data; Inputting the generator input data into the generator to generate an image, thereby obtaining a false image output by the generator; Inputting the real image, the enhanced image, and the false image into the discriminator for encoding, and obtaining a real latent space codeword corresponding to the real image, an enhanced latent space codeword corresponding to the enhanced image, and a false latent space codeword corresponding to the false image output by the discriminator; Determining a loss function of the generative adversarial network based on the true latent space codeword, the enhanced latent space codeword, and the false latent space codeword; If the loss function does not meet the convergence conditions, update the network parameters of the generative adversarial network, and return to the step of inputting the generator input data into the generator for image generation to obtain a false image output by the generator, until the updated loss function meets the convergence conditions.

2. The method according to claim 1, characterized in that The obtaining of generator input data comprises: Generate random semantics through binary selectors; Generate noise by random Gaussian sampling; The random semantics and the noise are binary-encoded to obtain the generator input data.

3. The method according to claim 1, characterized in that The determining of the loss function of the generative adversarial network based on the true latent space codeword, the enhanced latent space codeword, and the false latent space codeword includes: determining a reversible loss function based on the generator input data and the false latent space codeword; Determining an enhanced loss function based on the enhanced latent space codeword; determining a discriminator loss function of the discriminator based on the true latent space codeword, the false latent space codeword, and the generator input data; Determining a generator loss function of the generator based on the true latent space codeword and the false latent space codeword; A loss function of the generative adversarial network is determined based on the reversible loss function, the enhancement loss function, the discriminator loss function, and the generator loss function.

4. The method according to claim 3, characterized in that The determining of a reversible loss function based on the generator input data and the false latent space codeword comprises: Calculating the L2 norm of the generator input data and the false latent space codeword; The square of the L2 norm is calculated to obtain the reversible loss function.

5. The method according to claim 3, characterized in that The generative adversarial network further includes a projection network, wherein determining the enhanced loss function based on the enhanced latent space codeword includes: Inputting the enhanced latent space codeword into the projection network for projection to obtain a projected enhanced latent space codeword; Calculating a first similarity of a positive sample pair, wherein the positive sample pair is the projected enhanced latent space codeword corresponding to two different enhanced images of the same real image; Calculating a second similarity of a negative sample pair, wherein the negative sample pair is the projected enhanced latent space codeword corresponding to the enhanced images of two different real images; The enhancement loss function is determined based on the first similarity and the second similarity.

6. The method according to claim 3, characterized in that The determining of the discriminator loss function of the discriminator based on the true latent space codeword, the false latent space codeword and the generator input data comprises: Calculating a first distance between the real image and the generator input data in the latent space based on the real latent space codeword and the generator input data; Calculating a second distance between the false image and the generator input data in the latent space based on the false latent space codeword and the generator input data; The discriminator loss function is determined based on the first distance and the second distance.

7. The method according to claim 1, characterized in that The determining of a generator loss function of the generator based on the true latent space codeword and the false latent space codeword comprises: Based on the real latent space codeword and the false latent space codeword, a third distance between the real image and the false image in the latent space is calculated, and the third distance is used as the generator loss function.

8. A training device for generating an adversarial network, characterized in that The generative adversarial network includes a discriminator and a generator, wherein the device includes: A first acquisition module, configured to acquire a real image and an enhanced image corresponding to the real image; The second acquisition module is used to obtain generator input data; A first input module, configured to input the generator input data into the generator to generate an image, and obtain a false image output by the generator; A second input module is configured to input the real image, the enhanced image, and the false image into the discriminator for encoding, and obtain, output by the discriminator, a real latent space codeword corresponding to the real image, an enhanced latent space codeword corresponding to the enhanced image, and a false latent space codeword corresponding to the false image; A first determination module is configured to determine a loss function of the generative adversarial network based on the true latent space codeword, the enhanced latent space codeword, and the false latent space codeword; The first update module is used to update the network parameters of the generative adversarial network if the loss function does not meet the convergence conditions, and return to the step of inputting the generator input data into the generator for image generation to obtain a false image output by the generator, until the updated loss function meets the convergence conditions.

9. An electronic device, characterized in that: include: A processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor performs the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Image semantic communication security protection method and system based on semantic noise

    CN121056237A