A method and device for training a conditional generative adversarial network based on a characteristic function

CN117408313BActive Publication Date: 2026-09-15BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311236238.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2026-09-15
Estimated Expiration
2043-09-22

AI Technical Summary

Benefits of technology

[0017]本公开实施例提供的技术方案与现有技术相比具有如下优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117408313B_ABST
    Figure CN117408313B_ABST
Patent Text Reader

Abstract

The method comprises: obtaining training image data samples and auxiliary information samples; inputting a one-dimensional random sequence sampled from Gaussian random noise and the auxiliary information samples into a generation network to obtain generated image data samples; inputting the training image data samples and the generated image data samples into a discrimination network to obtain a first feature function corresponding to a first joint distribution and a second feature function corresponding to a second joint distribution; calculating a difference measure of the first feature function and the second feature function; calculating a loss function according to the difference measure, and adjusting network parameters of the generation network and the discrimination network to generate a conditional generation adversarial network, so that the conditional generation adversarial network processes real image data to generate target image data. By using the above technical solution, the efficiency and effectiveness of the conditional generation adversarial network training can be improved, thereby improving the subsequent image processing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a training method and apparatus for conditional generative adversarial networks based on feature functions. Background Technology

[0002] Generative Adversarial Networks (GANs) have traditionally been the mainstream model for image generation, thanks to their ability to generate sharp and realistic images from low-dimensional sources. However, the original GAN ​​architecture only allows random image generation from Gaussian noise. A significant variant of GAN aims to control generation through predefined auxiliary information (such as category labels or text), forming Conditional GANs (cGANs). Utilizing this auxiliary information, cGANs have been shown to enhance the generation of realistic images conditional on additional semantic cues. Therefore, cGANs have seen extremely wide applications in recent years, including category-conditional generation, style transfer, and text-to-image translation.

[0003] In related technologies, cGAN will... and auxiliary information Establish a joint distribution among them, that is Most cGANs follow the design of the generator network in GANs, embedding auxiliary information into the intermediate layer between the input noise and the generator, thereby enabling the generator to obtain information from the joint distribution. Mid-sampling. To design a discriminator, existing cGANs are distinguished by different representations of the conditional distribution, because the probability distribution... It can be done or To represent. The former requires auxiliary information. Transform into a discriminator for prediction It can be done by As input or A discriminator is implemented by embedding hidden layers. However, the latter requires a discriminator to predict auxiliary information. For example, through additional explicit classifiers or implicit projections. Although generation can be controlled by predefined auxiliary cues, practical applications of cGANs suffer from training mode collapse and training instability, thus hindering continuous improvement in realistic image generation. Summary of the Invention

[0004] To address the aforementioned technical problems, or at least partially address them, this disclosure provides a training method and apparatus for conditional generative adversarial networks based on feature functions.

[0005] This disclosure provides a training method for a conditional generative adversarial network based on feature functions, the method comprising: Acquire training image data samples and auxiliary information samples; A one-dimensional random sequence is sampled from Gaussian random noise and the auxiliary information sample is input into the generation network to obtain generated image data samples; The training image data samples and the generated image data samples are input into the discriminant network to obtain the first feature function corresponding to the first joint distribution and the second feature function corresponding to the second joint distribution; Calculate the difference measure between the first feature function and the second feature function; The loss function is calculated based on the difference metric, and the network parameters of the generator network and the discriminator network are adjusted to generate a conditional generative adversarial network (GAN), which then processes the real image data to generate the target image data.

[0006] Optionally, the characteristic function of the joint distribution can be determined using the following formula:

[0007] Where x represents a training image data sample or a generated image data sample, and y represents an auxiliary information sample. Indicates the joint distribution; in, express Fourier transform, ,in, and They represent and The corresponding Fourier frequency vector.

[0008] Optionally, if the auxiliary information samples are discretely distributed, then the characteristic function of the joint distribution... .

[0009] Optionally, the difference measure can be determined using the following formula; ; in, The characteristic function is represented by the first... One output.

[0010] Optionally, the loss function is: ; in, .

[0011] Optionally, for any two joint distributions and Given a set There exists a function For any and satisfy and ,in, Indicates the output of the first dimension.

[0012] Optionally, the first feature function and the second feature function are random variables that are real numbers, and the maximum value of the difference measure is the target distance measure.

[0013] Optionally, the method further includes: Acquire the image data to be processed and auxiliary information; A one-dimensional random sequence is sampled from Gaussian random noise and the auxiliary information is input into the generator network in the conditional generative adversarial network to obtain generated image data; The image data to be processed and the generated image data are input into the discriminant network in the conditional generative adversarial network to obtain the discrimination result of the image data to be processed.

[0014] This disclosure also provides a training apparatus for a conditional generative adversarial network based on feature functions, the apparatus comprising: The first acquisition module is used to acquire training image data samples and auxiliary information samples; The first generation module is used to sample a one-dimensional random sequence from Gaussian random noise and input the auxiliary information sample into the generation network to obtain generated image data samples; The processing module is used to input the training image data samples and the generated image data samples into the discriminant network to obtain the first feature function corresponding to the first joint distribution and the second feature function corresponding to the second joint distribution; The calculation module is used to calculate the difference measure between the first feature function and the second feature function; The second generation module is used to calculate a loss function based on the difference metric, and adjust the network parameters of the generation network and the discriminator network to generate a conditional generative adversarial network, so that the conditional generative adversarial network processes real image data to generate target image data.

[0015] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the training method of a feature function-based conditional generative adversarial network as provided in this disclosure.

[0016] This disclosure also provides a computer-readable storage medium storing a computer program for executing a training method for a feature function-based conditional generative adversarial network as provided in this disclosure.

[0017] The technical solution provided in this disclosure has the following advantages compared with the prior art: The training scheme for a conditional generative adversarial network (GAN) based on feature functions provided in this disclosure involves acquiring training image data samples and auxiliary information samples; sampling a one-dimensional random sequence from Gaussian random noise and inputting the auxiliary information samples into a generator network to obtain generated image data samples; inputting the training image data samples and generated image data samples into a discriminator network to obtain a first feature function corresponding to a first joint distribution and a second feature function corresponding to a second joint distribution; calculating a difference metric between the first and second feature functions; calculating a loss function based on the difference metric; and adjusting the network parameters of the generator and discriminator networks to generate a conditional GAN, enabling the GAN to process real image data and generate target image data. By employing this technical solution, the efficiency and effectiveness of GAN training can be improved, thereby enhancing the subsequent image processing results. Attached Figure Description

[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0019] Figure 1 A flowchart illustrating a training method for a conditional generative adversarial network based on feature functions, provided in an embodiment of this disclosure; Figure 2a A schematic diagram illustrating the generation result of a single-label conditional generative adversarial network provided in an embodiment of this disclosure; Figure 2b A schematic diagram illustrating another multi-label conditional generative adversarial network generation result provided in an embodiment of this disclosure; Figure 3a A schematic diagram illustrating the training of a conditional generative adversarial network provided in an embodiment of this disclosure; Figure 3b A schematic diagram illustrating the training of another conditional generative adversarial network provided in this embodiment of the disclosure; Figure 3c A schematic diagram illustrating the training of yet another conditional generative adversarial network provided in this embodiment of the disclosure; Figure 4 This is a schematic diagram of the structure of a training device for a conditional generative adversarial network based on feature functions, provided in an embodiment of this disclosure. Detailed Implementation

[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0022] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0023] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0024] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0025] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0026] Based on the aforementioned background, and considering that most cGAN discriminators are built on cross-entropy adversarial loss, equivalent to the Jensen-Shannon divergence between the generated and real data distributions, it has been theoretically and practically proven that when two probability distributions are misaligned or supported by low dimensionality, comparing the Jensen-Shannon divergence of two distributions in a "bin-to-bin" manner easily reaches its maximum. Therefore, the discriminator suffers from the vanishing gradient problem, which can mislead the generator into simply learning a fixed pattern or completely collapsing during training. For unconditional generation, this problem has been well addressed by introducing the Integral Probability Metric (IPM), a broad class of distance metrics. Building on the theoretical completeness of IPM, the discriminator can act as some bounded function, comparing distributions in a "cross-bin" manner, thus providing smoothness and sufficient gradients for unconditional generation.

[0027] Therefore, it is intuitive to apply IPM to conditional generation, given its theoretical completeness and ability to steadily and continuously improve performance. However, designing an IPM-cGAN is not easy due to the nonlinear coupling between data and auxiliary information. In other words, the bounded function of the discriminator is difficult to explicitly model for conditional generation. or For example, and Connections to enhance random variables Furthermore, cGANs are trained using unconditional IPM-GAN. However, directly combining two random variables at different semantic levels is unreasonable, and its shortcomings are problematic in many cGANs. Although some cGANs use certain IPMs in their implementations, such as Wasserstain distance (optimal transmission distance), their basic theory is based on cross-entropy, which still suffers from mode collapse and training instability caused by "bin-to-bin" comparisons. More importantly, the aforementioned cGANs are based on the existence of probability density functions of random variables, but this premise may not hold in practice, especially when real-world data such as images and videos are inherently located on low-dimensional manifolds.

[0028] Specifically, through To learn the joint distribution and add auxiliary information With data These are connected and used as inputs to both the generator and discriminator, thus providing auxiliary information for both the generation and discrimination processes. Similarly, the Laplace pyramid and temporal GANs will also... Connected to As input to the discriminator, to solve conditional distribution However, due to data and auxiliary information Because they exist at different semantic levels, directly concatenating them may lead to mismatched information aggregation, resulting in unstable training and inefficiency. Furthermore, a proposal is made to... By embedding a discriminator into a hidden layer, deeper information about the data can be extracted and then processed by the embedded... Aggregation. However, the above methods are intended to apply GANs to accomplish specific tasks, such as text-to-image translation and image editing.

[0029] Specifically, cGANs can Decomposed into ,in Predicted by either an implicit or explicit classifier. As a representative of classifier-free methods, projection-cGAN proposes calculating the likelihood ratio, where... Represented by projection, optimization is achieved under cross-entropy loss, possessing theoretical completeness. Due to its simplicity and theoretical completeness, projection-cGAN has been widely applied to many advanced models, including spectral normalization GAN, BigGAN, and self-attention GAN. On the other hand, adding a classifier has been shown to improve generative performance. GANs with auxiliary classifiers are among the most widely used cGANs models, featuring explicit classifiers that can be trained from marginal distributions and prediction accuracy. However, ACGANs are incomplete due to the bias in their learned distributions, a problem that can lead to mode collapse, especially when trained with a large amount of auxiliary information. Therefore, later improvements include using dual auxiliary classifiers, training with contrastive loss, adding auxiliary discriminative classifiers, and implementing multiple regularizations to stabilize training. However, all of the above cGANs are based on the cross-entropy loss function, which may lead to the problem that two well-separated distributions cannot be perfectly compared, potentially resulting in mode collapse and training instability.

[0030] Specifically, IPM has been widely used for unconditional generation, successfully reformulating the cross-entropy loss (predicting both real and generated samples) into a theoretically complete distance metric. IPM-GANs include Wasserstein GAN, Fisher GAN, MMDGAN, and CF GANs. Despite its great potential in addressing training instability in cGANs, research on applying IPM to cGANs remains lacking. Some IPM-GANs have the potential to extend to conditional generation, such as Wasserstein GAN, Fisher GAN, and CF GAN. However, since their IPMs are based on unconditional generation, extensions to conditional generation tend to concatenate data and auxiliary information to apply the unconditional setting. This significantly limits the capabilities of cGANs, as decomposing the joint distribution into marginal and conditional distributions can provide significant improvements. Some cGANs attempt to combine cross-entropy prediction and IPM in a particular way, but this still encounters training instability issues.

[0031] This disclosure proposes a novel cGAN based on the Characteristic Function (CF) of random variables, namely CCF-GAN. This disclosure establishes an empirical CF for both the generated and real joint distributions. By verifying that the empirical CF always exists and uniquely corresponds to the distribution, the difference between the calculated empirical CFs is used as a tool to reflect the differences in the joint distribution. However, CF computation requires extensive sampling in complex domains, especially in high-dimensional domains. Therefore, this disclosure uses a neural network as a tool for computing CF differences, called the Neural Network Characteristic Function (NCF). Based on the NCF, CCF-GAN can be established by explicitly modeling the conditional distribution of the joint distribution, thus laying a solid theoretical foundation for minimizing the differences in the joint distribution. Furthermore, the superior generative performance of CCF-GAN has been verified on both synthetic and real datasets.

[0032] Specifically, to address the issues of pattern collapse and unstable training that arise from the unsuitable method of measuring differences between distributions in Conditional Generative Adversarial Networks (GANs), this disclosure proposes a novel feature function-based Conditional Characteristic Function GAN (CCF-GAN) to reduce the differences between Characteristic Functions (CFs), which can serve as an accurate distance measure for learning the joint distribution. Specifically, the difference between the CFs of two distributions used to measure the difference between them is proven to be theoretically complete and suitable for optimization for the first time. To alleviate the curse of dimensionality when computing CFs, this disclosure uses neural networks, specifically Neural Characteristic Functions (NCFs), to efficiently and theoretically compute the differences. Based on NCFs, the CCF-GAN framework is further constructed, decomposing the joint distribution through conditional distributions to learn data distributions and auxiliary information of varying importance. Furthermore, experimental results on synthetic and real datasets validate the superior performance of CCF-GAN. Specific embodiments are described in detail below.

[0033] Figure 1 This is a flowchart illustrating a training method for a conditional generative adversarial network based on feature functions, provided in an embodiment of this disclosure. This method can be executed by a training device for a conditional generative adversarial network based on feature functions, wherein the device can be implemented in software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method includes: Step 101: Obtain training image data samples and auxiliary information samples.

[0034] Among them, training image data samples refer to real image distribution data, which can be sampled as input for subsequent network generation; auxiliary information samples can be category samples, label samples, etc., which can be selected and set according to the application scenario.

[0035] Step 102: Sample a one-dimensional random sequence from Gaussian random noise and input auxiliary information samples into the generation network to obtain generated image data samples.

[0036] Gaussian noise refers to a sequence that follows a Gaussian random distribution, such as a one-dimensional random sequence with a mean of 0 and a given variance.

[0037] Specifically, a one-dimensional random sequence can be sampled from Gaussian noise and combined with corresponding auxiliary information samples, then input into the generator network for processing to obtain generated image data samples.

[0038] Step 103: Input the training image data samples and the generated image data samples into the discriminant network to obtain the first feature function corresponding to the first joint distribution and the second feature function corresponding to the second joint distribution.

[0039] Step 104: Calculate the difference measure between the first characteristic function and the second characteristic function.

[0040] Step 105: Calculate the loss function based on the difference metric, and adjust the network parameters of the generator network and the discriminator network to generate a conditional generative adversarial network (GAN) so that the GAN can process the real image data and generate the target image data.

[0041] In this embodiment of the disclosure, a first joint distribution corresponding to the training image data samples is obtained, and a first feature function corresponding to the first joint distribution is constructed. A second joint distribution corresponding to the generated image data samples is obtained, and a second feature function corresponding to the second joint distribution is constructed. The network parameters of the generator network and the discriminator network are adjusted by calculating the difference measure between the first feature function and the second feature function and determining the loss function based on the difference measure until the loss value meets the requirements. Then, the training of the generator network and the discriminator network is stopped to obtain a conditional generative adversarial network.

[0042] In this context, it is understandable that during the training process, the loss function is continuously calculated based on the difference metric each time, thereby continuously adjusting the network parameters of the generator network and the discriminator network until the training converges and a well-trained conditional generative adversarial network is obtained.

[0043] Therefore, the efficiency and stability of the generated conditional generative adversarial network are improved, thus enhancing the image generation effect.

[0044] Specifically, the CF uniquely defines a random variable based on the cumulative density function (cdf), given by the following equation: (1) in, express The expectation.

[0045] Specifically, for any random variable, the probability density function (CF) always exists, even if the probability density function (pdf) is not explicitly defined (e.g., the Cantor distribution). When the pdf of a random variable exists, the CF can be expressed as... The inverse Fourier transform, i.e. In problems such as density estimation and generative modeling, random variables... The distribution is usually unknown, and there is only one set from Independent and identically distributed (iid) samples Available; this makes it impossible to perform calculations in CF. Perform continuous integration. Alternatively, it can be calculated as... The empirical CF (ECF) is the population in (1). An unbiased, consistent estimate, thus ensuring that it can be used as a well-defined unknown distribution. The replacement.

[0046] Specifically, the boundary of a cross-cutting network (CF) is another important property. and in The maximum value is reached at this time. In other words, the two distributions... and Automatic alignment within their cross-correlation (CF). In fact, comparing two distributions via their pdfs can lead to biases in optimization, potentially causing gradient vanishing and training instability. This problem has prompted the use of Wasserstein distance, etc., but at the cost of increased computational complexity or the need for additional constraints. However, comparing two CFs naturally avoids the misalignment problem while providing more gradual computation. Therefore, based on the uniqueness between random variables and their CFs, the following difference metric is proposed to compare two distributions (i.e., the first joint distribution) via their CFs. Second joint distribution ).

[0047] (2) in, This indicates a point-by-point distance metric used to compare two CFs. (2) express The distribution of can represent and The difference between them. It has been proven that when Support lies in real numbers China Times, It is an effective distance metric for comparing two distributions. Furthermore, only [data / data] can be obtained. and Discrete random samples, for example, For real images, Used for image generation tasks in image generation. Therefore, their CF can only be replaced by ECF, denoted as . and This basically falls into the category of two-sample testing problems, and under generalized conditions, the equality of their ECFs almost certainly ensures that the two distributions with statistical significance are equal.

[0048] Specifically, and more importantly, for conditional generation involving two joint distributions, such as real images and labels... And the generation of images , represented as and Therefore, the aforementioned ideal properties, including universality and uniqueness, still apply to their corresponding ECFs, because and Then, by utilizing automatic alignment between CFs, a simple and effective method is employed. Pointwise measurement, as shown below: (3) in, It comes from a certain distribution The sample. In formula (3), and Let represent the ECF of the real and generated joint distributions, respectively. They are expressed by the following equation: (4) in, and This is for the convenience of representing column samples. It should be noted that in (3), the number of samples... Distinguishing and This plays a crucial role in demonstrating sufficient differences in probability estimates. Figures 2a-2b The text explains that, without any additional discriminator module, simply by... and By optimizing the generator using the standard Gaussian distribution in (3), it is possible to generate roughly realistic MNIST digital images, where the input image size is 28×28.

[0049] Specifically, compared to real-world images, the 28×28 grayscale digital images from the MNIST dataset represent a simplified scenario. When optimizing images with high dimensionality and diverse content, It must grow exponentially, especially for high-dimensional data, where the curse of dimensionality (COD) problem arises. To solve this problem... A skillful selection is needed, rather than simply using a Gaussian distribution. More importantly, in Figures 2a-2b In this preliminary experiment, directly using and Connecting them together has proven invalid because of the pixel-level images. and semantic tags They are fundamentally different. A method is proposed that decomposes the conditional distribution. and A new method, in this way and and and They can be effectively treated as different levels of importance.

[0050] To address the cod problem when calculating the difference between ECFs, several unconditional generation methods propose reducing image dimensionality by learning embedding functions; for example, allowing the comparison of two embedding distributions... and Explicit enumeration with relatively low dimensions However, this requires an understanding of the function. Additional requirements are imposed, including injective and bijective embeddings, leading to additional hyperparameters and instabilities during GAN training. More importantly, the embedding function... It is essentially implemented using a highly nonlinear discriminator neural network (also known as a critic). Therefore, its extensions to conditional generation are very limited, such as embedding cascaded joint distributions. This is how most IPM-cGANs operate. However, this approach has proven to be inefficient in cGANs.

[0051] Specifically, the most basic operation of CF in formula (1) is... It is from higher dimensions Projection to scalar Therefore, an implicit optimization method is proposed when comparing two complex distributions. Instead of explicit enumeration In (3), the Cramer-Wold theorem states that if and only if... and The distribution for all When both are the same, the two joint distributions They have the same distribution. They can be accessed through them. The NCF network compares two high-dimensional complex distributions using an infinite number of projections in space. Therefore, it does not explicitly sample from a few predefined distributions. Instead, it implicitly searches all possible... And directly output the corresponding projection and In order to compare the projected distribution in the lower dimension.

[0052] More specifically, the generated samples and real samples The input is fed into the NCF network, and the output size of the NCF network is... , representing each sample in Projection in each direction, i.e. and Specifically, Lemma 1 proves the NCF network. Able to generally satisfy the needs of and The projection operation is performed. In this way, NCF implicitly samples... Then output and The difference between the two ECFs can be expressed as: (5) in, Indicates the NCF's One output. In (5), the NCF network The implemented ECF is represented as and By changing the network weights, NCF is able to search... The entire space, thus precisely indicating and The differences between them. Furthermore, existing IPM-GANs are limited by the network's theoretical completeness, such as being constrained by Lipchitz continuity and injectivity requirements, unlike NCF networks which can... The NCF can freely search within the entire space. Therefore, the NCF fully utilizes the universal approximation capability of neural networks when calculating the difference between two joint distributions.

[0053] Lemma 1. For any two joint distributions and Given a set There exists a function For any and satisfy and ,in, Indicates the output of the first dimension.

[0054] Furthermore, with changing Large-scale sampling Compared to other methods, the method that best distinguishes the two ECFs in (5) is the "pre- A "representative" sample is more effective, that is... make Maximize, as follows: (6) Therefore, if the maximum difference in the ECFs of two distributions disappears, then they are equal. Further verification in Lemma 2 reveals that the maximum difference... It is an effective distance metric that can reflect the differences between two distributions.

[0055] Lemma 2. If If there are two joint distributions, then in (6) It is an effective distance metric.

[0056] Furthermore, conditional generation can be achieved by setting... This can be achieved by stacking at different semantic levels. and This can lead to learning problems in cGANs. In this embodiment, the proposed CCF-GAN processes them separately to generate... It can accommodate auxiliary information well. This improves generation performance. More specifically, the CF of the joint distribution can be decomposed as follows: (7) in, This represents training image data samples or generated image data samples. This represents a sample of auxiliary information. Indicates the joint distribution; in, express Fourier transform, ; and They represent and The corresponding Fourier frequency vector.

[0057] Specifically, formula (7) plays a key role in CCF-GAN, effectively [from...] Decomposition Furthermore, in many tasks, auxiliary information is discretely distributed, such as category labels. Therefore, we are able to obtain (7) CF as: (8) in, yes The number of discrete values. Correspondingly, The ECF is: (9) Specifically, the proposed NCF network Calculate in (9) This is due to the data distribution. Typically located in high dimensions, therefore selection is required. The intelligent optimization strategy, and auxiliary information It has a relatively low dimension. Therefore, (9) can be transformed into: (10) Use superscript This indicates that the ECF is computed by the proposed NCF network. Therefore, by substituting (10) into (6), the final loss function for training CCF-GAN is obtained: (11) Specifically, IPMs have been incorporated into existing cGANs through nonlinear transformation functions, which makes... and The decomposition between them becomes tricky. In contrast, CCF-GAN benefits from the direct output of the NCF network. It can be clearly extracted from the joint distribution. The remaining portion is explicitly represented by the conditional distribution. In this way, data distribution and auxiliary information can be explicitly treated with different levels of importance, thus allowing for [further processing] in real-world scenarios. and generated Optimize the difference measure between joint distributions.

[0058] Therefore, the generator network of CCF-GAN aims to minimize equation (11) to reduce the difference between the generated and real joint distributions, while the discriminator network is an NCF network, which maximizes... This constitutes the effective distance metric proposed in Lemma 2.

[0059] Therefore, the steps for training CCF-GAN are as follows: the input includes the real image distribution. Gaussian noise Batch size ; Dimensions Learning rate Number of categories The number of training iterations for the generator network and the discriminator network, respectively. and The output includes the network parameters of the generator network and the discriminator network, respectively. and ;when and When convergence fails, the discriminant network is trained by sampling from the distribution for each training iteration: , and ; Calculate the loss function: ;renew: Train the generative network by sampling from the distribution for each training iteration: , and ; Calculate the loss function: ;renew: .

[0060] in, Represents sampling of a real image. Images are generated by sampling Gaussian noise and inputting it into the generator. This indicates the category label corresponding to the generated image.

[0061] Specifically, the training dynamic process is as follows: Figures 3a-3c As shown, light colors represent the generated image distribution, and dark colors represent the real image distribution. (a) Sampling from a Gaussian distribution (a) OCF-GAN cannot generate images similar to real images. (b) RCF-GAN, which uses t-net to approximate the real distribution, continuously generates images with the real distribution, showing the ineffectiveness of the discriminator. (c) CCF-GAN (the training method of conditional generative adversarial network based on feature function proposed in the embodiments of this disclosure) separates real and fake images in the intermediate stage and finally generates realistic images.

[0062] In some embodiments, image data to be processed and auxiliary information are acquired, and a one-dimensional random sequence is sampled from Gaussian random noise. The auxiliary information is then input into the generator network of the conditional generative adversarial network to obtain generated image data. Alternatively, the image data to be processed and the generated image data can be input into the discriminator network of the conditional generative adversarial network to obtain a discrimination result for the image data to be processed.

[0063] Therefore, the generated conditional generative adversarial network can process the image data to be processed in real time to generate the image data, thereby improving the efficiency and effect of image generation.

[0064] The training scheme for a conditional generative adversarial network (GAN) based on feature functions provided in this disclosure involves acquiring training image data samples and auxiliary information samples; sampling a one-dimensional random sequence from Gaussian random noise and inputting the auxiliary information samples into a generator network to obtain generated image data samples; inputting the training image data samples and generated image data samples into a discriminator network to obtain a first feature function corresponding to a first joint distribution and a second feature function corresponding to a second joint distribution; calculating a difference metric between the first and second feature functions; calculating a loss function based on the difference metric; and adjusting the network parameters of the generator and discriminator networks to generate a conditional GAN, enabling the GAN to process real image data and generate target image data. By employing this technical solution, the efficiency and effectiveness of GAN training can be improved, thereby enhancing the subsequent image processing results.

[0065] Figure 4 This is a schematic diagram of a training device for a conditional generative adversarial network based on feature functions, provided as an embodiment of this disclosure. This device can be implemented in software and / or hardware, and is generally integrated into an electronic device. Figure 4 As shown, the device includes: The first acquisition module 201 is used to acquire training image data samples and auxiliary information samples; The first generation module 202 is used to sample a one-dimensional random sequence from Gaussian random noise and input the auxiliary information sample into the generation network to obtain generated image data samples; Processing module 203 is used to input the training image data samples and the generated image data samples into the discriminant network to obtain the first feature function corresponding to the first joint distribution and the second feature function corresponding to the second joint distribution; Calculation module 204 is used to calculate the difference measure between the first feature function and the second feature function; The second generation module 205 is used to calculate a loss function based on the difference metric and adjust the network parameters of the generation network and the discriminator network to generate a conditional generative adversarial network, so that the conditional generative adversarial network processes real image data to generate target image data.

[0066] Optionally, the characteristic function of the joint distribution can be determined using the following formula: ; Where x represents a training image data sample or a generated image data sample, y represents an auxiliary information sample, and p(x, y) represents a joint distribution; in, Describe the Fourier transform of p(x,y). .

[0067] Optionally, if the auxiliary information samples are discretely distributed, then the characteristic function of the joint distribution... .

[0068] Optionally, the difference measure can be determined using the following formula; ; in, The characteristic function is represented by the first... One output.

[0069] Optionally, the loss function is: ; in, .

[0070] Optionally, for any two joint distributions and Given a set There exists a function For any and satisfy and ,in, Indicates the output of the first dimension.

[0071] Optionally, the first feature function and the second feature function are random variables that are real numbers, and the maximum value of the difference measure is the target distance measure.

[0072] Optionally, the device further includes: The second acquisition module is used to acquire the image data to be processed and auxiliary information; The third generation module is used to sample a one-dimensional random sequence from Gaussian random noise and input the auxiliary information into the generator network in the conditional generative adversarial network to obtain generated image data. The discrimination module is used to input the image data to be processed and the generated image data into the discrimination network in the conditional generative adversarial network to obtain the discrimination result of the image data to be processed.

[0073] The training apparatus for conditional generative adversarial networks based on feature functions provided in this disclosure can execute the training method for conditional generative adversarial networks based on feature functions provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0074] This disclosure also provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the training method for a feature function-based conditional generative adversarial network provided in any embodiment of this disclosure.

[0075] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0076] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0077] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0078] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire training image data samples and auxiliary information samples; sample a one-dimensional random sequence from Gaussian random noise and input the auxiliary information samples into a generator network to obtain generated image data samples; input the training image data samples and generated image data samples into a discriminator network to obtain a first feature function corresponding to a first joint distribution and a second feature function corresponding to a second joint distribution; calculate a difference measure between the first feature function and the second feature function; calculate a loss function based on the difference measure, and adjust the network parameters of the generator network and the discriminator network to generate a conditional generative adversarial network, so that the conditional generative adversarial network processes real image data to generate target image data.

[0079] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0081] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0082] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0083] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0084] According to one or more embodiments of this disclosure, this disclosure provides an electronic device, including: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the training method of the feature function-based conditional generative adversarial network as provided in any of the present disclosure.

[0085] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium storing a computer program for performing a training method for a feature function-based conditional generative adversarial network as described in any of the present disclosure.

[0086] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0087] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0088] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for training a conditional generative adversarial network based on a characteristic function, characterized in that, include: Acquire training image data samples and auxiliary information samples; A one-dimensional random sequence is sampled from Gaussian random noise and the auxiliary information sample is input into the generation network to obtain generated image data samples; The training image data samples and the generated image data samples are input into the discriminant network to obtain the first feature function corresponding to the first joint distribution and the second feature function corresponding to the second joint distribution; Calculate the difference measure between the first feature function and the second feature function; The loss function is calculated based on the difference metric, and the network parameters of the generator network and the discriminator network are adjusted to generate a conditional generative adversarial network, so that the conditional generative adversarial network can process real image data and generate target image data. The characteristic function of the joint distribution is determined by the following formula: ; in, This represents training image data samples or generated image data samples. This represents a sample of auxiliary information. Indicates the joint distribution; in, express Fourier transform, , and They represent and The corresponding Fourier frequency vector; The difference measure is determined by the following formula: ; in, The characteristic function is represented by the first... One output; First joint distribution Second joint distribution ; The loss function is: ; in, .

2. The method according to claim 1, characterized in that, If the auxiliary information samples are discretely distributed, then in the characteristic function of the joint distribution... 。 3. The method according to claim 1, characterized in that, For any two joint distributions and Given a set There exists a function For any and satisfy and ,in, Indicates the output of the first dimension.

4. The method according to claim 1, characterized in that, The first feature function and the second feature function are random variables that are real numbers, and the maximum value of the difference measure is the target distance measure.

5. The method according to any one of claims 1-4, characterized in that, Also includes: Acquire the image data to be processed and auxiliary information; A one-dimensional random sequence is sampled from Gaussian random noise, and the auxiliary information is input into the generator network in the conditional generative adversarial network to obtain generated image data.

6. A training device for a conditional generative adversarial network based on feature functions, characterized in that, include: The first acquisition module is used to acquire training image data samples and auxiliary information samples; The first generation module is used to sample a one-dimensional random sequence from Gaussian random noise and input the auxiliary information sample into the generation network to obtain generated image data samples; The processing module is used to input the training image data samples and the generated image data samples into the discriminant network to obtain the first feature function corresponding to the first joint distribution and the second feature function corresponding to the second joint distribution; The calculation module is used to calculate the difference measure between the first feature function and the second feature function; The second generation module is used to calculate a loss function based on the difference metric, and adjust the network parameters of the generation network and the discriminator network to generate a conditional generative adversarial network, so that the conditional generative adversarial network processes real image data to generate target image data. The characteristic function of the joint distribution is determined by the following formula: ; in, This represents training image data samples or generated image data samples. This represents a sample of auxiliary information. Indicates the joint distribution; in, express Fourier transform, , and They represent and The corresponding Fourier frequency vector; The difference measure is determined by the following formula: ; in, The characteristic function is represented by the first... One output; First joint distribution Second joint distribution ; The loss function is: ; in, .

7. The apparatus according to claim 6, characterized in that, Also includes The second acquisition module is used to acquire the image data to be processed and auxiliary information; The third generation module is used to sample a one-dimensional random sequence from Gaussian random noise and input the auxiliary information into the generator network in the conditional generative adversarial network to obtain generated image data. The discrimination module is used to input the image data to be processed and the generated image data into the discrimination network in the conditional generative adversarial network to obtain the discrimination result of the image data to be processed.