A cross-modal generation method from text to image based on factorization

The factor-decomposed generative adversarial network (FDGAN) addresses the issue of blurred semantic control in text-to-image models by separating text conditions and noise inputs, enhancing generation control and synthesis performance.

CN116468896BActive Publication Date: 2025-07-15FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310415768.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2025-07-15
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

In the existing cross-modal generation model, text conditions and random noise are coupled together, resulting in unclear semantic control of generated images, making it difficult to achieve accurate generation control and synthesis performance.

Method used

Using a factor decomposition-based generative adversarial network, the output of the generative model is optimized by decoupling text conditions and random noise, using an addition-based instance regularization layer and a deep attention multimodal similarity model.

Benefits of technology

It realizes better generation control and synthesis performance, better semantic consistency of generated results, and improves model complexity and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468896B_ABST
    Figure CN116468896B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of AI-based generated content, and specifically relates to a cross-modal generation method from text to image based on factor decomposition. The present invention uses a generative adversarial network based on factor decomposition; decouples text conditional control and random noise for separate processing, that is, inputs the two into the generative adversarial network based on factor decomposition in different ways: directly inputs the random noise into the generative adversarial network, and embeds the text conditional control into the generative network through an addition-based instance regularization layer to achieve decoupling of text conditional control and random noise; the generative adversarial network includes a basic generator based on factor decomposition and a super-resolution module enhanced by attention, as well as a joint discriminator based on factor decomposition. The joint discriminator is used to discriminate the output of the generative model, thereby optimizing the generative model. The present invention can achieve better conditional control generation and synthesis performance on the basis of the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of AI-based generated content, and particularly relates to a cross-modal generation method from text to image. Background Art

[0002] AI based Generative Content (AIGC) uses generative models, especially cross-modal generative models, to generate realistic images from text descriptions or texture maps. Commonly used cross-modal generative models include Generative Adversarial Nets (GAN) and Diffusion Model. AIGC-related technologies can be applied in multiple fields such as media and publishing, design and mapping, online education, game development, etc., reducing the creation threshold and improving the creation efficiency.

[0003] For cross-modal generation from text to image based on GAN, the text is usually encoded into a feature vector, and then it is concatenated with a random noise to generate an image that is semantically consistent with the input text while also having a certain degree of diversity in the generated image. However, this method couples the input condition and the random noise together, resulting in a blurred corresponding relationship for semantic control of the generated image. For example, in the CUB dataset, the semantic condition usually corresponds to the appearance of the generated bird images, such as the bird species, the colors of different parts, etc. The random noise often corresponds to the pose of the generated birds. But simply concatenating the text condition input and the noise vector will couple these two control inputs together, which is not conducive to the decoupling and control of the generation results. Summary of the Invention

[0004] The purpose of the present invention is to provide a cross-modal generation method from text to image based on factor decomposition that can better achieve generation control and synthesis performance.

[0005] The cross-modal generation method from text to image based on factor decomposition proposed by the present invention uses a factor decomposition-based generative adversarial network; decouples and separately processes the text condition control and the random noise, that is, inputs the two into the factor decomposition-based generative adversarial network in different ways: directly inputs the random noise into the generative adversarial network, and embeds the text condition control into the generative network through the additive instance normalization method in the additive instance regularization layer, so as to achieve the effect of decoupling the text condition control and the random noise. Experimental results on public datasets also show that this decoupled input method can achieve better generation control and synthesis performance.

[0006] In the present invention, the factor decomposition-based generative adversarial network used has a structure as shown in Figure 1As shown in the figure. The Factor Decomposed Generative Adversarial Nets (FDGAN) includes a factor-decomposed basic generator and an attention-enhanced super-resolution module. It also includes a factor-decomposed joint discriminator, which is used to discriminate the output of the generative model, so as to better optimize the generative model.

[0007] The factor-decomposed basic generator is a multi-layer transposed convolutional network; its input is a randomly sampled Gaussian noise, and through multi-layer transposed convolution, a feature h0 with a spatial resolution of 64x64 is obtained. During the transposed convolution process, the sentence-level features of the text visual description are embedded into the transposed convolutional layer; the attention-enhanced super-resolution module takes h0 as the input and obtains an output image with a resolution of 64x64 through a generative module G0; h0 passes through two convolutional blocks and an upsampling layer to obtain h1, and an output image with a resolution of 128x128 is obtained from h1; h1 passes through two convolutional blocks and an upsampling layer to obtain h2, and an output image with a resolution of 256x256 is obtained from h2. During the process of increasing the resolution from low to high, through the attention mechanism F i Attn Embed the word-level features of the text visual description into the super-resolution generation process; the specific process is as follows:

[0008] For the word feature e of the text input and the hidden layer feature h of the previous layer, first pass through a new perceptual layer U to obtain the representation e′ = Ue of the word feature e in the semantic space. The hidden layer feature h consists of a series of representations of sub-regions h = {h1, h2... h N}}, for the context vector c of the cross-modal attention mechanism of the j-th sub-region j is obtained by the following formula:

[0009]

[0010] where:

[0011]

[0012] β ji represents the attention weight between the i-th word in the text and the j-th sub-region of the image hidden layer feature. The cross-modal context attention mechanism can be expressed as:

[0013] F attn (e, h) = (c0, c1,... c N-1 ),

[0014] where N represents the number of sub-regions of the image hidden layer feature h.

[0015] The generated images of three resolutions respectively pass through an image discriminator, and the generation model is optimized through gradient backpropagation; at the same time, for the generated images with a resolution of 256x256, the Deep Attentional Multimodal Similarity Model (DAMSM) is used [1] Optimize the entire generation model. Among them, the discriminator uses a joint discriminator based on factorization, and optimizes the conditional distribution and unconditional distribution of the generated data at the same time. The loss of the i-th discriminator for the generation model is:

[0016]

[0017] In the present invention, in the instance regularization layer based on addition, the instance regularization method based on addition is a new instance normalization method adopted by analyzing the noise characteristics of cross-modal generation on the basis of the existing adaptive instance regularization method; the specific description is as follows:

[0018] The most basic instance normalization method is expressed as:

[0019]

[0020] Among them, γ and β are learnable affine transformation parameters, and u(x) and σ(x) respectively represent calculating the mean and variance for each channel of the input feature x:

[0021]

[0022]

[0023] Among them, ε is a positive constant to avoid division by zero, H and W respectively represent the height and width of the image, and the subscripts nchw respectively represent the meaning of each dimension of the four-dimensional feature: the number of samples, the number of channels, the spatial height, and the spatial width.

[0024] On the basis of the instance normalization method, an adaptive instance normalization method is adopted:

[0025]

[0026] Among them, y is the style image used to guide style transfer.

[0027] In the present invention, for the conditional random perturbation in cross-modal generation, a new instance normalization method is adopted: the instance normalization method based on addition, and its expression:

[0028]

[0029] Among them, c is the conditional feature, T cIt is a transformation that aligns the conditional feature c and the data x, u nc (x) and σ nc (x) represent the mean and variance of x in the sample dimension and the channel dimension respectively. This normalization method can reduce the impact of perturbations on conditional control under the Gaussian perturbation assumption, achieving more accurate conditional control generation.

[0030] Features and advantages of the present invention compared with the prior art:

[0031] Based on the existing technology, the present invention realizes better conditional control generation, can better decouple the control factors corresponding to text conditions and random noise, and achieves more accurate text conditional control generation. At the same time, through the generation network structure disclosed in the present invention, the text-to-image cross-modal generation model based on the generative adversarial network can also be made more efficient.

[0032] References:

[0033] [1]Xu T, Zhang P, Huang Q, et al. Attngan: Fine-grained text to image generation with attentional generative adversarial networks[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 1316 - 1324.

[0034] [2]Wah C, Branson S, Welinder P, et al. The caltech-ucsd birds-200-2011 dataset[J]. 2011.

[0035] [3]Lin T Y, Maire M, Belongie S, et al. Microsoft coco: Common objects in context[C] / / Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6 - 12, 2014, Proceedings, Part V 13. Springer International Publishing, 2014: 740 - 755。 Description of the Drawings

[0036] Figure 1 It is a structural diagram of a factorization-based generative adversarial network.

[0037] Figure 2 This is a comparison of the subjective effects of the present invention and the benchmark model AttnGAN on the CUB dataset. Detailed implementation manners

[0038] Refer to AttnGAN [1] Based on the implementation of AttnGAN, the factorization-based generative model disclosed in the present invention is implemented based on Pytorch, and experiments are carried out on two public datasets, CUB 200 [2] and MS COCO [3] The preprocessing of the input images refers to AttnGAN [1] . For the CUB 200 dataset, data augmentation is achieved by cropping the input images according to the detection boxes in the dataset with a certain perturbation.

[0039] During the training process, the Adam optimizer is used, and the learning rate is set to 1e-3. There are 22 samples in a batch on the CUB 200 dataset, and 600 epochs of training are performed on the dataset. There are 14 samples in a batch on the MS COCO dataset, and 120 epochs of training are performed on the dataset. In terms of evaluation metrics, Inception Score (IS), Fréchet inception distance (FID), and retrieval recall precision (R-precision) are used to evaluate the quality of the generated results, and the number of model parameters, the time T of one forward process in the training mode tr and the time T of one forward process in the test mode te are used to evaluate the efficiency of the model.

[0040] The comparison effects on the two public datasets are shown in Table 1 and Table 2.

[0041] Table 1. Comparative experiments on the CUB200-2011 dataset

[0042]

[0043] Table 2. Comparative experiments on the MS-COCO dataset

[0044]

[0045] From the comparison results of the comparative experiments, it can be seen that adding the factorization module disclosed in the present invention to the commonly used benchmark models AttnGAN, ControlGAN, and DM-GAN can improve the relevant performance indicators while reducing the complexity of the model, including the number of parameters, the training and test times, etc.

[0046] Meanwhile, the subjective effects on the CUB dataset are as Figure 2 shown.

[0047] It can be seen that, compared with the baseline model AttnGAN, the FDGAN disclosed in the present invention has a better decoupling effect in the generation results, and the semantic consistency of the generation results under the same text semantics is better.

Claims

1. A cross-modal generation method from text to image based on factorization, characterized in that, Use a factorization-based generative adversarial network; decouple and separately process text conditional control and random noise, that is, input them into the factorization-based generative adversarial network in different ways: directly input the random noise into the generative adversarial network, and embed the text conditional control into the generative network through the additive instance normalization method in the additive instance regularization layer to achieve the decoupling of text conditional control and random noise; Among them, the factorization-based generative adversarial network FDGAN includes a factorization-based basic generator, an attention-enhanced super-resolution module, and a factorization-based joint discriminator. The joint discriminator is used to discriminate the output of the generative model, thereby optimizing the generative model; The basic generator is a multi-layer transposed convolutional network; its input is a randomly sampled Gaussian noise, and through multi-layer transposed convolutions, a feature h0 with a spatial resolution of 64x64 is obtained. During the transposed convolution process, the sentence-level features of the text visual description are embedded into the transposed convolution layer; The attention-enhanced super-resolution module takes h0 as the input and obtains an output image with a resolution of 64x64 through a generation module G0; h0 passes through two convolutional blocks and an upsampling layer to obtain h1, and an output image with a resolution of 128x128 is obtained from h1; h1 passes through two convolutional blocks and an upsampling layer to obtain h2, and an output image with a resolution of 256x256 is obtained from h2; during the process of increasing the resolution from low to high, through the attention mechanism Embed the word-level features of the text visual description into the super-resolution generation process; the specific process is as follows: For the word feature e of the text input and the hidden layer feature h of the previous layer, first pass through a new perceptron layer U to obtain the representation e' = Ue of the word feature e in the semantic space. The hidden layer feature h consists of a series of representations of sub-regions h = {h1, h2... h N}, and for the context vector c of the cross-modal attention mechanism of the j-th sub-region j is obtained by the following formula: Where: β ji denotes the attention weight between the i-th word in the text and the j-th sub-region of the image hidden layer feature. The cross-modal context attention mechanism is expressed as: F attn (e, h) = (c0, c1, … c N-1 ) Among them, N represents the number of sub-regions of the image hidden layer feature h; The expression of the additive instance normalization method is: Among them, c is the conditional feature, T c is the transformation that aligns the conditional feature c and the data x, u nc (x) and σ nc (x) represent the mean and variance of x in the sample dimension and the channel dimension respectively.

2. The cross-modal generation method from text to image based on factorization according to claim 1, wherein, The attention mechanism embeds the word-level features of the visual description of the text into the super-resolution generation process. The specific process is as follows: For the word feature e of the text input and the hidden layer feature h of the previous layer, first pass through a new perceptron layer U to obtain the representation e' = Ue of the word feature e in the semantic space; the hidden layer feature h consists of representations of a series of sub-regions h = {h1, h2... h N}, for the context vector c of the cross-modal attention mechanism of the j-th sub-region j is obtained by the following formula: Where: β ji denotes the attention weight between the $i$-th word in the text and the $j$-th sub-region of the image hidden layer features; the cross-modal context attention mechanism is expressed as: F attn (e, h) = (c0, c1, … c N-1 ) Among them, N represents the number of sub-regions of the image hidden layer feature h.

Citation Information

Patent Citations

  • Text-to-image generation method based on generative adversarial network

    CN111260740A

  • Text image generation method and system based on multi-stage generative adversarial network

    CN113361251A