A method for generating images from text using conditional generative adversarial networks based on distribution estimation
By introducing a new loss form based on distribution estimation into the text-generated image model, the problems of overfitting and performance instability of existing models when data volume is limited are solved, and the effect of generating high-quality images and improving training stability is achieved.
Patent Information
- Application Number
- CN202111670694.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-12-31
AI Technical Summary
Existing text-generated image models are prone to overfitting when trained on datasets with limited data volume, and their performance is unstable, the generated image quality is poor, and difficult to reproduce.
A new form of loss for a conditionally generated adversarial network based on distribution estimation is proposed. By introducing a probability distribution, the generator and discriminator are optimized, and the quality and training stability of the model-generated image are improved.
The upper bound of easy-to-calculate loss is obtained through mathematical derivation, implicitly reflecting the impact of generating a large number of images in a single text description, improving the level of high-quality images generated by the model and the stability of overall training.
Smart Images

Figure CN114332565B_ABST
Abstract
Description
Technical Field
[0001] This paper proposes a new loss form of conditional generative adversarial neural network (cGAN) based on distribution estimation for cross-modal text generation image tasks. Background Art
[0002] Humans’ ability to visualize written text plays an important role in many cognitive processes, such as memory, spatial reasoning, etc. Inspired by humans’ ability to visualize, building a cross-modal system that converts between language and vision has become a new pursuit in the field of artificial intelligence.
[0003] Compared with written text, images are a more accurate, efficient and convenient way to share and transmit information. In recent years, the development of deep learning has taken computer vision and image generation technology a step further. The emergence of generative adversarial neural networks (GANs) has enabled image generation tasks to be trained in an unsupervised manner. At the same time, with the further development of generative adversarial networks (GANs), conditional variables such as text descriptions have also been integrated into the framework of image generation tasks. Through conditional generative adversarial neural networks (cGANs), images corresponding to text descriptions can be generated with text descriptions as conditions. Text descriptions can carry dense semantic information about the attributes, spatial positions, relationships, etc. of the current object, and can represent different scenes, thereby realizing the conversion process from language to vision.
[0004] Generating images from textual descriptions (T2I) is a complex computer vision and machine learning task with important applications in multiple fields such as image editing, computer-aided design, and video games.
[0005] The use of conditional generative adversarial neural networks (cGAN) is a mainstream method for achieving text-to-image (T2I). In the past few years, the architecture and performance of the model have been improved. These include more detailed text feature extraction, divided into sentence features and word features; the use of new architectures (such as stacked structures to gradually improve image resolution, the introduction of attention mechanisms and dynamic memory mechanisms in the network, etc.); and the introduction of new multimodal losses for text-to-image generation (T2I). Some excellent algorithms that have emerged in recent years have introduced the above improvements, such as StackGAN++, AttnGAN, DM-GAN, etc., which have greatly improved the quality and resolution of generated images. At the same time, in terms of evaluation indicators, new indicators (R-value, semantic object accuracy, etc.) are developed and defined to evaluate the performance of text-to-image models.
[0006] However, existing models still have some limitations and defects. First, they are trained on some data sets with limited data volume (for example, Oxford-102Flowers and CUB-200Birds), with a total number of images of about 10k, and the total number of images in the data set is too small. The training of the discriminator is often prone to overfitting, which makes it difficult to improve the overall performance of the model after a period of training.
[0007] Another problem is that the performance of the model is unstable. Through observation and statistics of the images generated by the model, it is found that there are still many images of poor quality, and the quantitative results of many methods are difficult to reproduce (even if the code and model are provided). The evaluation indicators of the text-to-image generation (T2I) task are basically based on the distribution of data, and a few low-quality images are difficult to reflect in the evaluation indicators. Consideration should be given to improving the level of the model to generate high-quality images and the stability of the overall training. Summary of the invention
[0008] The purpose of the present invention is to address the deficiencies of the prior art and propose a method for generating images from text using a conditional generative adversarial network based on distribution estimation. A new loss form of a conditional generative adversarial network based on distribution estimation is used to improve the performance of the text-to-image model and the stability of training. The new loss function proposed in the present invention is based on the motivation of generating a large number of images from a single text description, and improving the quality of the overall generated image by penalizing a large number of text-image pairs at the same time, thereby improving the performance of the model.
[0009] However, the actual situation is that the computational cost of the loss involved in generating a large number of images is unbearable. By mathematically deriving the new loss function, using Jason's inequality and the moment generating function formula, we can obtain an upper bound that is easy to calculate. This loss implicitly reflects the impact of infinite image generation from a single text description in terms of the probability distribution of features, and constrains the loss from a distribution perspective, so that the generator and discriminator can be better optimized. Improve the quality of images generated by the model.
[0010] A method for generating images from text using a conditional generative adversarial network based on distribution estimation comprises the following steps:
[0011] Step (1), data preprocessing, extracting features of text data;
[0012] For the training and test datasets of the text generation image task. First, the corresponding natural language text descriptions are added to the CUB and MSCOCO image datasets. The CUB-200 dataset is a bird dataset with 200 categories of bird data. The CUB-200 dataset is divided according to the specified division rules. The training set contains 150 categories and the test set has 50 categories of bird data. The COCO dataset has a total of 91 categories of images, and the training set and test set are also divided according to the specified ratio.
[0013] Feature extraction is performed on the natural language text description to obtain a text feature set. The extracted text feature set includes global sentence-level features and fine-grained word-level features. Specifically, a pre-trained bidirectional long short-term memory network BiLSTM is used to extract semantic features from the natural language text description to form the features of each word, and the features of the sentence are obtained through the hidden state of the final connection.
[0014] Step (2), establishing a multi-stage unconditional and conditional joint generative adversarial neural network (cGAN) and loss function;
[0015] The traditional conditional generative adversarial neural network has only one set of generators and discriminators, which makes it difficult to generate high-resolution images. The present invention adopts a multi-stage conditional generative adversarial neural network model as a benchmark model, and uses its ability to stack generators to gradually improve the resolution of the generated images.
[0016] At the same time, the unconditional generative adversarial neural network and the conditional generative adversarial neural network are jointly trained. For the unconditional generative adversarial neural network, the generator is trained to generate fake images that can deceive the discriminator, and the discriminator can distinguish between real images and fake images. In order to control the image to generate images that meet the description, the conditional generative adversarial neural network is also trained, using the text feature set extracted in step (1) as a conditional variable input to the generator and the discriminator to guide the generator to generate an image distribution that approximates the text condition, while the discriminator can better judge whether the image and text conditions match. The text feature set includes word features and sentence features.
[0017] Step (3), introducing a loss function based on distribution estimation;
[0018] Replace the loss function in step (2) with a new loss function based on distribution estimation. Use the new loss function based on distribution estimation on the loss of the discriminator and the generator respectively. The new loss function assumes that the features of the image generated by a single text description belong to a Gaussian distribution, that is:
[0019]
[0020]
[0021] in, is the characteristic of the image generated by the unconditional generative adversarial neural network, are the characteristics of images generated by conditional generative adversarial neural networks, and are the means of two Gaussian distributions, and is the covariance of the Gaussian distribution, and i represents the i-th text description. The model training is constrained by probability distribution.
[0022] Step (4), model training, optimize the discriminator and generator, and obtain images corresponding to the text description.
[0023] Furthermore, the data preprocessing and text feature extraction in step (1) are specifically as follows:
[0024] Citation datasets (CUB-200, COCO-2014), CUB-200 is a relatively small dataset, containing a total of 200 categories of bird images. According to the specified division into training and test sets, the training set contains 8,855 images and 2,933 images as the test set. Each image describes a single object (bird), and each image has 10 related text descriptions. COCO consists of approximately 123k images, each with 5 descriptions. Among them, 80k images are divided into training sets and 40k images are used as test sets. The COCO dataset is a dataset with more object categories, and the amount of data is several times that of the CUB-200 bird dataset, which can better test the performance of the algorithm in actual scenarios.
[0025] For the natural language text description in the dataset, feature extraction is performed and a pre-trained bidirectional long short-term memory network (BiLSTM) is used to extract a text feature set from the text description. In the bidirectional long short-term memory network, its two hidden states are connected as the features of a word. A feature matrix e∈R of all words in the text description is obtained D×T , where the i-th column vector e of the feature matrix i represents the feature of the i-th word, D represents the dimension of the word feature, and T is the number of words. Connect the last layer of hidden states as the global sentence feature
[0026] Furthermore, the specific method of step (2) is as follows:
[0027] 2-1 uses DM-GAN as the benchmark model. The multi-stage stacked network improves the image resolution by stacking the generator and the discriminator to generate images with richer details. For the generator of the model, given the random noise z~N(0,1) and the conditional variable c, through F0 and F i Get the input of the generator h0=F0(c,z) and h i =F i (h i-1 ,z),h i-1 Input the next stage generator network Fi Get h i , where F i is the neural network of the generator. For the generator G i , generate multi-stage resolution images x i =G i (h i ).
[0028] 2-2 Joint training of conditional and unconditional generative adversarial neural networks. The objective function of the model contains two contents, namely unconditional loss and conditional loss. The unconditional loss determines the visual authenticity of the image, and the conditional loss determines whether the image and text description can match. The i-th stage discriminator D i The loss is defined as follows:
[0029]
[0030] Correspondingly, the generator G of the i-th stage i The loss is also composed of two parts.
[0031]
[0032] where x i is the real image distribution p from the i-th stage datai The image, s i is the generator G i The false image of the generated stage i, c is the conditional variable, and E represents the mathematical expectation.
[0033] Furthermore, the specific method of step (3) is as follows:
[0034] 3-1 To achieve overall optimization of images generated from a single text description, a large number of images with the same text description are generated to optimize the network and improve model performance. The generator loss for generating an image from a single text description is defined as follows:
[0035]
[0036] Therefore, the loss of generating M images is expressed as:
[0037]
[0038] 3-2 However, in the actual calculation process, it is impossible to bear the computational cost of generating too many images. In order to solve this problem, an infinite M is used in the formula. Through mathematical derivation, the loss can be converted into an easily calculated upper bound, which implicitly reflects the constraint of generating a large number of images in the form of probability distribution.
[0039] Let M→∞, the loss of the generator The definition is as follows:
[0040]
[0041] where w u , b u and w c , b c are the weights and biases of the last layer of the discriminator network of the unconditional and conditional generative adversarial neural networks, respectively. It is an image generated by an unconditional generative adversarial neural network, after the discriminator D i Features before the last layer of the network; It is an image generated by a conditional generative adversarial neural network, after the discriminator D i Features before the last layer of the network; i represents the i-th stage, E represents the corresponding mathematical expectation, and N represents the number of samples.
[0042] Assume that the features of the image generated by a single text description belong to a Gaussian distribution, that is:
[0043]
[0044]
[0045] The unconditional loss of formula (5) can be derived as an easily computable upper bound:
[0046]
[0047] The same generator G i The unconditional loss can also be derived as follows:
[0048]
[0049] In the above derivation, formula (8) uses Jensen inequality E[logX]≤logE[X], and formula (9) is obtained by transformation using the moment generating function, which is defined as follows:
[0050]
[0051] For the discriminator D i The conditional and unconditional losses can also be deduced through the same mathematical derivation to obtain the corresponding upper bound of the loss, namely:
[0052]
[0053] where α i and β i is the feature of the real image obtained by the discriminator network, w u , b u and wc , b c are the weights and biases of the last layer of the discriminator network of the unconditional and conditional generative adversarial neural networks, respectively. and Characteristics and The mean of the Gaussian distribution to which , and Characteristics and The covariance of the Gaussian distribution. N represents the number of samples.
[0054] Finally, the loss function is constructed based on the introduction of probability distribution. i and the generator G i (i=0,1,2) all use a new loss function based on distribution estimation.
[0055] Furthermore, the specific method of step (4) is as follows:
[0056] According to the new loss function obtained, the discriminator D is trained i and the generator G i Perform alternating training. When training the discriminator, the generator model is fixed, and the gradient information is only transmitted on the discriminator; when training the generator, the gradient information is always transmitted from the discriminator to the generator, but the discriminator model does not perform gradient updates, and only optimizes the parameters of the generator network. Finally, the model parameters are updated through the back-propagation algorithm (BP) until the model converges.
[0057] The generator model saved after training can generate corresponding high-resolution images based on the specified text description.
[0058] The beneficial effects of the present invention are as follows:
[0059] In order to improve the overall performance of conditional generative adversarial neural networks in the task of generating images from text and generate high-quality images, the present invention proposes a new loss function suitable for conditional generative adversarial neural networks, which is a mechanism for optimizing the network through the probability distribution of image features. This loss implicitly reflects the impact of a single text generating an infinite number of images, and an easily computable upper bound of the loss is obtained through mathematical derivation. By estimating the distribution of image features generated by a single text description, loss calculation and gradient information feedback are achieved. Experiments on multiple models and data sets show that the new loss function based on distribution estimation can effectively improve the performance of conditional generative adversarial neural networks in generating images from text, while reducing the appearance of low-quality images, and improving the overall image effect.
[0060] The present invention adopts a completely end-to-end approach to optimize network performance. The new loss is applied to multiple text-to-image models, and the performance is improved to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 This is the conditional generative adversarial network model structure based on distribution estimation of the present invention.
[0062] Figure 2 The complete flow chart of the text-to-image task implemented by the present invention DETAILED DESCRIPTION
[0063] The method of the present invention and its detailed parameters are further described in detail below.
[0064] A method for generating images from text using a conditional generative adversarial neural network based on distribution estimation, the specific steps are as follows:
[0065] Step (1), data preprocessing, extracting features of text data;
[0066] CUB-200 is a dataset of bird images in 200 categories, totaling 11,788 images. According to the specified division into training and validation sets, the training set contains 8,855 images and 2,933 images as the test set. Each image describes a single object (bird) and each image has 10 associated text descriptions. Since 80% of the birds in this dataset have an object-to-image size ratio less than 0.5, the data is preprocessed and all images are cropped to ensure that the object-to-image size ratio of the bird's bounding box is greater than 0.75. The size of the real image used is 299×299.
[0067] COCO consists of about 123k images, each with 5 descriptions. 80k of them are divided into training sets and 40k of them are used as test sets. After setting up the experiment, we directly use the training and validation sets divided by COCO.
[0068] For the natural language text description in the dataset, a text feature set is extracted. A pre-trained bidirectional long short-term memory network (BiLSTM) is used to extract the text feature set from the text. The text feature set contains the features of words and sentences. In the bidirectional long short-term memory network, each word corresponds to two hidden states, one state for each direction. Therefore, its two hidden states are connected as the features of a word, and finally a word feature matrix e∈R is obtained. D×T , where the i-th column vector e of the matrix irepresents the feature of the i-th word, D = 256 represents the dimension of the word feature, and T = 25 is the number of words. At the same time, the last hidden state of the bidirectional long short-term memory network is connected as the global sentence feature
[0069] Step (2), establishing a multi-stage unconditional and conditional joint generative adversarial neural network and loss function;
[0070] 2-1 uses DM-GAN as the baseline model. The multi-stage stacked network improves the resolution of the image by stacking the generator and the discriminator to generate images with richer details. For the generator of the model, given the random noise z~N(0,1) and the conditional variable c, the dimensions are 100 and 256 respectively.
[0071] Through F0 and F i Get the input h0=F0(c,z) and h0 of the next stage generator i =F i (h i-1 ,z),h i-1 Input the next stage generator network F i Get h i , where F i is the neural network in the generator. F0 consists of a fully connected layer and a four-layer convolutional network. i (i=1,2) consists of a dynamic memory write mechanism, two residual modules and a convolutional layer. i , generating images with multiple resolution stages The resolution sizes are 64×64, 128×128 and 256×256 respectively.
[0072] 2-2 Joint training of conditional and unconditional generative adversarial neural networks. The objective function of the model contains two contents, namely unconditional loss and conditional loss. The i-th stage discriminator D i The loss is defined as follows:
[0073]
[0074] The corresponding generator G of the i-th stage i The loss is also composed of two parts.
[0075]
[0076] where x i is the real image distribution from the i-th stage The image, s i is the generator G iThe false image of the generated stage i, c is the conditional variable, and E represents the mathematical expectation.
[0077] Step (3), introducing a loss function based on distribution estimation;
[0078] In order to achieve overall optimization of the images generated from a single text description, the new loss function derived previously is used. This loss is an easy-to-calculate upper bound that implicitly reflects the impact of a single text generating a large number of images in the form of a probability distribution. The definition is as follows:
[0079]
[0080] where w u , b u and w c , b c are the weights and biases of the last layer of the discriminator network of the unconditional and conditional generative adversarial neural networks, respectively. It is an image generated by an unconditional generative adversarial neural network, after the discriminator D i Features before the last layer of the network; It is an image generated by a conditional generative adversarial neural network, after the discriminator D i Features before the last layer of the network; i represents the i-th stage, E represents the corresponding mathematical expectation, and N represents the number of samples.
[0081] Assume that the features of the image generated by a single text description belong to a Gaussian distribution, that is and Here, we estimate the mean and covariance matrix of the two distributions by generating M′ images with a single text description, where M′=4.
[0082] Generator loss After M tends to infinity, an easily computable form can be derived, and the unconditional loss and conditional loss of the generator are finally defined as follows:
[0083]
[0084]
[0085]
[0086] For the discriminator D i The conditional and unconditional losses can also be deduced through the same mathematical derivation to obtain the corresponding upper bound of the loss, namely:
[0087]
[0088] where αi and β i is the feature obtained by the real image through the discriminator network. u , b u and w c , b c are the weights and biases of the last layer of the discriminator network of the unconditional and conditional generative adversarial neural networks, respectively. and Characteristics and The mean of the Gaussian distribution to which , and Characteristics and The covariance of the Gaussian distribution. N represents the number of samples.
[0089] like Figure 1 As shown in Figure 1, it is a single-stage conditional generative adversarial network based on distribution estimation, and the training process of the text generation image task. Finally, the loss function is constructed by introducing the probability distribution, and the discriminator D of each stage is i and the generator G i (i=0,1,2) all use a new loss function based on distribution estimation.
[0090] Step (4), model training;
[0091] According to the new loss function obtained, the discriminator D is trained i and the generator G i Alternating training is performed. The relevant training parameters are set as follows: the training epoch is 800, the batch size is 20, the Adam optimizer is used, and the initial learning rates of the discriminator and the generator are both 2e-4.
[0092] When the discriminator is trained, the generator model is fixed, and the gradient information is only transmitted on the discriminator; when the generator is trained, the gradient information is always transmitted from the discriminator to the generator, but the discriminator model does not perform gradient updates, and only optimizes the parameters of the generator network. Finally, the model parameters are updated through the back-propagation algorithm (BP) until the model converges.
[0093] The generator model saved after training can generate corresponding high-resolution images according to the specified text description. Figure 2 As shown in the figure, it is the complete process of the model to realize the task of generating images from text.
[0094] The mean and covariance of the generated images are used to calculate the values of the evaluation indicators FID and IS to quantify the performance of the model.
[0095] Table 1 shows the quantitative evaluation results of the distribution estimation-based conditional generative adversarial network (DM-GAN+DE) and its comparison algorithm on the CUB-200 dataset. The image generation quality evaluation uses two indicators, FID (the larger the better) and IS (the smaller the better). The results show that the new loss form of the conditional generative adversarial neural network based on distribution estimation in this paper can effectively improve the performance of text generation image models such as DM-GAN: in terms of the FID indicator, it is reduced from 16.09 to 14.71, and the IS is increased from 4.71 to 4.84.
[0096] This result shows that the new loss form based on distribution estimation proposed in this paper can enable the text-to-image generation model based on the adversarial generative network to generate images with better quality.
[0097] Table 1
[0098]
Claims
1. A method for generating images from text using conditional generative adversarial networks based on distribution estimation, characterized in that The steps include: Step (1), data preprocessing, extracting features of text data; Step (2), establishing a multi-stage unconditional and conditional joint generative adversarial neural network and loss function; Step (3), introducing a loss function based on distribution estimation; Step (4), model training: According to the new loss function obtained, the discriminator D is trained during the training process. i and the generator G i Perform alternating training; Step (2) is specifically implemented as follows: 2-1 uses DM-GAN as the baseline model. The multi-stage stacked network improves the image resolution by stacking generators and discriminators. For the generator of the model, given random noise z~N(0,1) and conditional variable c, the dimensions are 100 and 256 respectively. Through F0 and F i Get the input h0=F0(c,z) and h0 of the next stage generator i =F i (h i-1 ,z),h i-1 Input the next stage generator network F i Get h i , where F i is the neural network in the generator; F0 consists of a fully connected layer and a four-layer convolutional network, F i It consists of a dynamic memory writing mechanism, two residual modules and a convolutional layer; for the generator G i , generating images with multiple resolution stages The resolution sizes are 64×64, 128×128 and 256×256 respectively; 2-2 Joint training of conditional and unconditional generative adversarial neural networks. The objective function of the model contains two contents, namely unconditional loss and conditional loss; the i-th stage discriminator D i The loss is defined as follows: The corresponding generator G of the i-th stage i The loss is also composed of two parts. where x i is the real image distribution from the i-th stage The image, s i is the generator G i The generated false image of the i-th stage, c is the conditional variable, and E represents the mathematical expectation; Step (3) is specifically implemented as follows: In order to achieve overall optimization of the images generated from a single text description, the new loss function derived previously is used. This loss is an easy-to-calculate upper bound that implicitly reflects the impact of a single text generating a large number of images in the form of a probability distribution; the loss of the generator The definition is as follows: where w u , b u and w c , b c They are the weights and biases of the last layer of the discriminator network of the unconditional and conditional generative adversarial neural networks, respectively; It is an image generated by an unconditional generative adversarial neural network, after the discriminator D i Features before the last layer of the network; It is an image generated by a conditional generative adversarial neural network, after the discriminator D i Features before the last layer of the network; Where i represents the i-th stage, E represents the corresponding mathematical expectation, and N represents the number of samples; Assume that the features of the image generated by a single text description belong to a Gaussian distribution, that is and Here, we estimate the mean and covariance matrix of the two distributions by generating M′ images with a single text description, where M′=4; Generator loss After M tends to infinity, an easily computable form is derived, and the unconditional loss and conditional loss of the generator are finally defined as follows: For the discriminator D i The conditional and unconditional losses of , and the corresponding upper bounds of the losses are obtained through the same mathematical derivation, namely: where α i and β i It is the feature obtained by the real image through the discriminator network; w u , b u and w c , b c They are the weights and biases of the last layer of the discriminator network of the unconditional and conditional generative adversarial neural networks, respectively; and Characteristics and The mean of the Gaussian distribution to which , and Characteristics and The covariance of the Gaussian distribution to which it belongs; N represents the number of samples; Finally, the loss function is constructed by introducing the probability distribution, and the discriminator D of each stage is i and the generator G i All use a new loss function based on distribution estimation, where i = 0, 1, 2.
2. The method for generating images from conditional generative adversarial network text based on distribution estimation according to claim 1 is characterized in that Step (1) is specifically implemented as follows: The citation dataset CUB-200 contains 200 categories of bird images, totaling 11,788 images; according to the specified division of training set and validation set, the training set contains 8,855 images and 2,933 images as the test set; each image describes a single object, and each image has 10 related text descriptions; since 80% of the birds in the dataset have an object and image size ratio less than 0.5, the data is preprocessed and all images are cropped to ensure that the object and image size ratio of the bird's bounding box is greater than 0.75; the size of the real image used is 299×299; COCO consists of about 123k images, each with 5 descriptions; 80k of them are divided into training sets and 40k of them are used as test sets; Extract the text feature set from the natural language text description in the data set, and use a pre-trained bidirectional long short-term memory network to extract the text feature set from the text description. The text feature set contains the features of words and sentences. In the bidirectional long short-term memory network, each word corresponds to two hidden states, one state for each direction. Therefore, connect its two hidden states as the features of a word, and finally get a word feature matrix e∈R D×T , where the i-th column vector e of the matrix i represents the feature of the i-th word, D = 256 represents the dimension of the word feature, and T = 25 is the number of words; at the same time, the last hidden state of the bidirectional long short-term memory network is connected as the global sentence feature e∈R D .
Citation Information
Patent Citations
Automatic text generation method and device
CN108334497A
Text image generation method based on StackGAN network
CN111968193A