A method for generating images from text
By introducing the Transformer module and a multi-stage generation network, combining self-attention and global attention mechanisms, the problems of unreasonable images and unclear details in the existing text image generation method are solved, and the generated image details are clearer and the resolution is higher.
Patent Information
- Application Number
- CN202111109265.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-09-22
AI Technical Summary
The existing text image generation methods have problems such as unreasonable images and unclear details, especially the AttnGAN network-based methods still have shortcomings in the generated image quality and resolution.
A multi-stage generation method based on the Transformer module and AttnGAN network is adopted. By extracting the word features and sentence features of the text, combining random noise for conditional enhancement, and iterative training between multiple generators and discriminators, the self-attention mechanism and global attention module are used to improve image details and resolution.
The resulting image details are clearer, improving image quality and resolution.
Smart Images

Figure CN114022582B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of general image data processing or generation, and particularly to a method for generating images from text based on a Transformer module and an AttnGAN network in the technical fields of computer vision and natural language processing. Background Art
[0002] The rapid development of modern technology has promoted the progress of computer vision and natural language processing in theory and technology. Generating images based on text descriptions is a comprehensive task spanning the two fields of computer vision and natural language processing, with great application potential. It is expected to play an important role in criminal investigation, data enhancement, design, etc. in the future.
[0003] In the early stage, generating images from text mainly combined retrieval and supervised learning. However, this method could only change specific image features. It was not until Reed et al. used a generative adversarial network to first achieve generating images from text, which not only changed the features but also laid a foundation for subsequent development according to the text content.
[0004] StackGAN improves the image resolution by gradually synthesizing images, generating low-resolution images in the first stage and then optimizing the details in the second stage; StackGAN++ reduces the instability of network training by adding regularization of color consistency to the loss. However, the existing GAN-INT-CLS, StackGAN, and StackGAN++ only use the overall text information as text features, which will lead to the loss of important details in the synthesized images. Based on this, AttnGAN was proposed. This network introduces a global attention mechanism and extracts features from the text both globally and locally to obtain sentence features and word features, and inputs both features as text features, improving the relevance between the text and the image. Moreover, due to the global attention mechanism introduced by the AttnGAN network, the quality of the generated images has also been improved.
[0005] Although the above methods have continuously improved in the quality and resolution of the generated images, there are still often problems such as unreasonable images and unclear details. Summary of the Invention
[0006] The present invention solves the problems existing in the prior art and provides a method for generating images from text based on a Transformer module and an AttnGAN network, which solves problems such as unreasonable images and blurred details.
[0007] The technical solution adopted by the present invention is a method for generating images from text, and the method includes the following steps:
[0008] Step 1: Obtain a dataset composed of texts and corresponding images, and perform preprocessing;
[0009] Step 2: Construct a text-to-image network model based on AttnGAN. The network model includes a pre-trained network and a multi-stage generation network. The pre-trained network introduces a Transformer module;
[0010] Step 3: Extract the text features of the text description corresponding to the image. The text features include word features and sentence features. After conditionally enhancing the sentence features, merge them with random noise and input them into the Transformer module to learn spatial and position information;
[0011] Step 4: Input the feature information learned in Step 3 into the first generator to output a 64*64 low-resolution initial generated image. Input the low-resolution initial generated image and the sentence features into the initial discriminator for discrimination;
[0012] Step 5: Downsample the low-resolution initial generated image generated in Step 4 to obtain features. Input the word features into a global attention module to obtain new word features. Input the two features together into a convolutional neural network for learning to obtain fused features. Then input the fused features into the second generator to output a 128*128 image. Input the 128*128 image and the sentence features into the secondary discriminator for discrimination;
[0013] Step 6: Downsample the 128*128 image generated in Step 5 to obtain features. Input the word features into a global attention module to obtain new word features. Input the two features together into a convolutional neural network for learning to obtain new fused features. Input the new fused features into the third generator to output a 256*256 image. Input the 256*256 image and the sentence features into the tertiary discriminator for discrimination;
[0014] Step 7: Take the image generated in Step 6 as the final generated image and output it.
[0015] Preferably, in Step 2, the Transformer module includes an Encoder module and a Decoder module;
[0016] The Encoder module includes three first sub-modules connected in sequence. Any one of the first sub-modules includes a self-attention layer, a normalization layer, and a fully connected layer connected in sequence;
[0017] The Decoder module includes three second sub-modules connected in sequence. Any one of the second sub-modules includes a first self-attention layer, a first normalization layer, a second self-attention layer, a second normalization layer, and a fully connected layer connected in sequence;
[0018] The outputs of the Encoder module are respectively input into the second self-attention layers of three second sub-modules;
[0019] The first first sub-module of the Encoder module and the first second sub-module of the Decoder module respectively correspond to the input ends, and the output end of the third second sub-module of the Decoder module corresponds to the output end.
[0020] Preferably, in step 3, merging the sentence features with random noise and inputting them into the Transformer module to learn the spatial and position information includes the following steps:
[0021] Step 3.1: Obtain a feature vector from the sentence features through the conditional enhancement module
[0022]
[0023] where, is the sentence feature of the text, is the mean vector of the sentence feature vectors of the text, is the covariance matrix of the sentence feature vectors of the text, and ε is the distribution sampling of the unit Gaussian distribution N(0,1);
[0024] Step 3.2: Merge the obtained feature vector with the random noise vector z to obtain e, and use e as the input of the Transformer module;
[0025] Step 3.3: In the Transformer module, convert e in the attention space to obtain three representation vectors: Calculate the weight information
[0026]
[0027] where, α j,i represents the weight information at the i-th position when synthesizing the j-th region of the composite image; finally, obtain the image feature matrix m with the attention mechanism j ,
[0028]
[0029] Step 3.4: Integrate the feature matrix to obtain the Transformer output feature vector h,
[0030] Preferably, in the step 4, the loss function of the first generator is
[0031]
[0032] where denotes the mathematical expectation, G1 and D1 are the first generator and the initial discriminator respectively, λ is a hyperparameter determining the influence of the DAMSM module on the loss function of the generator, h is the output feature vector of the Transformer, e is the text feature vector after combining the sentence feature and the random noise vector, and L DAMSM denotes the loss function obtained by training the DAMSM module of the network.
[0033] Preferably, in the step 4, the loss function of the initial discriminator is L D1 = L1 + L2, where L1 denotes discriminating whether the input image is real, and L2 denotes discriminating whether the input image and the text are semantically consistent.
[0034]
[0035]
[0036] where denotes the mathematical expectation, x is the real image corresponding to the text description, denotes the image generated by the first generator corresponding to the text description, and e is the text feature vector after combining the sentence feature and the random noise vector.
[0037] Preferably, in the step 5, the loss function of the second generator is
[0038]
[0039] where denotes the mathematical expectation, K = S(F(h, e1)), G2 and D2 are the second generator and the secondary discriminator respectively, λ1 is a hyperparameter determining the influence of the DAMSM module on the loss function of the second generator, h is the output feature vector of the Transformer, e1 is the word feature, and L DAMSM denotes the loss function obtained by training the DAMSM module of the network, F is the global attention generation module, and S is the neural network downsampling module.
[0040] Preferably, in the step 5, the loss function of the secondary discriminator is L D2 = L1 + L2, where L1 denotes discriminating whether the input image is real, and L2 denotes discriminating whether the input image and the text are semantically consistent.
[0041]
[0042]
[0043] Among them, represents the mathematical expectation, x represents the real image corresponding to the text description, represents the image generated by the second generator corresponding to the text description, and e represents the text feature vector after the joint of the sentence feature and the random noise vector.
[0044] Preferably, in the step 6, the loss function of the third generator is
[0045]
[0046] where K = S(F(F(h, e1))), G3 and D3 are the third generator and the third discriminator respectively, λ2 represents the hyperparameter that determines the influence of the DAMSM module on the loss function of the third generator, h represents the output feature vector of the Transformer, e1 represents the word feature, e represents the text feature vector after the joint of the sentence feature and the random noise vector, and L DAMSM represents the loss function obtained by training the DAMSM module of the network, and F(h, e1) represents the feature vector learned through the global attention module.
[0047] Preferably, in the step 6, the loss function of the third discriminator is L D3 = L1 + L2, where L1 represents discriminating whether the input image is real, and L2 represents discriminating whether the input image and the text are semantically consistent.
[0048]
[0049]
[0050] where represents the mathematical expectation, x represents the real image corresponding to the text description, represents the image generated by the third generator corresponding to the text description, and e represents the text feature vector after the joint of the sentence feature and the random noise vector.
[0051] Preferably, in the training stage of the text generation image network model based on AtmGAN, the result of passing the word feature through the DAMSM module is compared with the result of passing the 256*256 image through the image decoder, and the text generation image network model is adjusted based on the comparison result.
[0052] The present invention provides a method for generating images from text. Based on the Transformer module and the AttnGAN network, after encoding the text through a text encoder, sentence features and word features are obtained. The sentence features are passed through a conditional enhancement module to obtain feature vectors, which are then fused with random noise vectors and input into the Transformer module for learning to output improved feature vectors. These feature vectors are input into a generator to generate an initial image of 64*64 pixels roughly. The initial synthesized image and the improved feature vectors are input into a discriminator for discrimination, and the generator is trained according to a loss function. Sequentially, the improved feature vectors from the previous step and the word features are input into a neural network for upsampling to obtain fused vectors, which are then input into the generator to obtain images of 128*128 pixels and 256*256 pixels.
[0053] The images generated by the method of the present invention have clearer detail contours compared to the images generated by the previous traditional AttnGAN method. Brief Description of the Drawings
[0054] Figure 1 is the network structure diagram of the present invention;
[0055] Figure 2 is the structure diagram of the Transformer module in the present invention. Detailed Embodiments
[0056] The following further describes the present invention in detail with reference to embodiments, but the protection scope of the present invention is not limited thereto.
[0057] The present invention relates to a method for generating images from text based on the Transformer module and AttnGAN, and the method includes the following steps.
[0058] Step 1: Obtain a dataset composed of text and corresponding images, and perform preprocessing.
[0059] In the present invention, the data in the dataset includes text and corresponding images. The preprocessing includes manually screening the images and text in the dataset, removing text and image data with unclear representations, and modifying text with inaccurate descriptions of the pictures.
[0060] Step 2: Construct a text-to-image network model based on AttnGAN. The network model includes a pre-training network and a multi-stage generation network, and the pre-training network introduces the Transformer module;
[0061] In step 2, the Transformer module includes an Encoder module and a Decoder module;
[0062] The Encoder module includes three first sub-modules connected in sequence. Any one of the first sub-modules includes a self-attention layer, a normalization layer, and a fully-connected layer connected in sequence;
[0063] The Decoder module includes three second sub-modules connected in sequence. Any one of the second sub-modules includes a first self-attention layer, a first normalization layer, a second self-attention layer, a second normalization layer, and a fully-connected layer connected in sequence;
[0064] The outputs of the Encoder module are respectively input into the second self-attention layers of the three second sub-modules;
[0065] The first first sub-module of the Encoder module and the first second sub-module of the Decoder module respectively correspond to the input end, and the output end of the third second sub-module of the Decoder module corresponds to the output end.
[0066] In the training stage of the text-to-image network model based on AttnGAN, the result of passing the word features through the DAMSM module is compared with the result of passing the 256*256 image through the image decoder, and the text-to-image network model is adjusted based on the comparison result.
[0067] In the present invention, the data in the dataset is divided into a training set and a test set. The training set is used to train the text-to-image network model, and the test set is used to test and experience the network performance.
[0068] In the present invention, in order to better extract features and generate images, a pre-trained network introduces a Transformer module, a DAMSM module text encoder, and a picture encoder. The multi-stage generation network part includes three generators and a discriminator;
[0069] Among them, the Transformer module is a neural network based on self-attention. The text encoder is for extracting text features, the picture encoder is for extracting image features, and the DAMSM module is to input the finally synthesized image into the picture encoder to perform a correlation comparison between the local image features and the text features, so as to improve the correlation between the image and the text;
[0070] A standard GAN network is composed of a generation network (also called a generator) and a discriminant network (also called a discriminator). The text-to-image network model of the present invention uses three generators and three discriminators, forming three groups. The generators and discriminators are both convolutional neural networks (CNNs).
[0071] In the present invention, as Figure 2The structure of the Transformer module is shown;
[0072] In the subsequent step 3, the feature vectors (collectively referred to as A) obtained by conditionally enhancing the sentence features and merging them with random noise are input into the Transformer module;
[0073] In the Encoder module, A is flattened into a one-dimensional vector, and after embedding the position information, the corresponding Q, K, and V matrices enter the self-attention layer. In the self-attention layer, weights are calculated for the Q and K matrices, and the corresponding scores are added to the V matrix; after the self-attention layer, a row normalization operation is performed, and then it is sent to the fully connected layer, and the fully connected layer outputs a one-dimensional vector; this is repeated three times, and the vector output by the Encoder module becomes B.
[0074] In the Decoder module, A is flattened into a one-dimensional vector, and after embedding the position information, the corresponding Q, K, and V matrices enter the self-attention layer. After the self-attention layer, a normalization operation is performed and added to the vector B, then it enters another self-attention layer to obtain the vector C, and then the vector C is sent to the fully connected layer;
[0075] The vector output by the Decoder is the vector finally output by the Transformer module.
[0076] Step 3: Extract the text features of the text description corresponding to the image. The text features include word features and sentence features. The sentence features are conditionally enhanced and then merged with random noise and input into the Transformer module to learn spatial and position information; specifically, more spatial and position information.
[0077] In the said step 3, merging the sentence features with random noise and inputting them into the Transformer module to learn spatial and position information includes the following steps:
[0078] Step 3.1: Obtain the feature vector through the conditional enhancement module with the sentence features
[0079]
[0080] Among them, is the sentence feature of the text, is the mean vector of the sentence feature vectors of the text, is the covariance matrix of the sentence feature vectors of the text, and ε is the distribution sampling of the unit Gaussian distribution N(0,1);
[0081] Step 3.2: Combine the obtained feature vectors with the random noise vector z to obtain e, and use e as the input of the Transformer module;
[0082] Step 3.3: In the Transformer module, convert e into the attention space to obtain three representation vectors: Calculate the weight information
[0083]
[0084] where α j,i represents the weight information at the i-th position when synthesizing the j-th region of the synthetic image; finally, obtain the image feature matrix m with the attention mechanism j ,
[0085]
[0086] Step 3.4: Integrate the feature matrix to obtain the Transformer output feature vector h,
[0087] In the present invention, in step 3.3, Query refers to each value in the feature vector (meaning the word to be searched), Key is the other values of the feature vector, and Value can be understood as the correlation between Query and Key; Query, Key, and Value can all be directly obtained, and then α j,i and m j are obtained.
[0088] In the present invention, in step 3.4, the Transformer output feature vector h refers to the spatial and position information.
[0089] In the present invention, both i and j refer to a value (positive integer) from 1 to n, but the values of i and j are different sub-samples in the same large sample.
[0090] Step 4: Input the feature information learned in step 3 into the first generator to output a 64*64 low-resolution initial generated image, and input the low-resolution initial generated image and the sentence feature into the initial discriminator for discrimination;
[0091] In the said step 4, the loss function of the first generator is
[0092]
[0093] where denotes the mathematical expectation, G1 and D1 are the first generator and the initial discriminator respectively, λ is the hyperparameter that determines the influence of the DAMSM module on the loss function of the generator, h is the output feature vector of the Transformer, e is the text feature vector after combining the sentence feature and the random noise vector, L DAMSM denotes the loss function obtained by training the DAMSM module of the network.
[0094] In the said step 4, the loss function of the initial discriminator is L D1 = L1 + L2, where L1 denotes discriminating whether the input image is real, and L2 denotes discriminating whether the input image and the text are semantically consistent.
[0095]
[0096]
[0097] where denotes the mathematical expectation, x denotes the real image corresponding to the text description, denotes the image generated by the first generator corresponding to the text description, and e denotes the text feature vector after combining the sentence feature and the random noise vector.
[0098] Step 5: Downsample the low-resolution initial generated image obtained in step 4 to get features, input the word features into a global attention module to get new word features, input the two features together into a convolutional neural network for learning to get fused features, then input the fused features into the second generator to output a 128*128 image, and input the 128*128 image and the sentence feature into a secondary discriminator for discrimination;
[0099] In the said step 5, the loss function of the second generator is
[0100]
[0101] where denotes the mathematical expectation, K = S(F(h, e1)), G2 and D2 are the second generator and the secondary discriminator respectively, λ1 is the hyperparameter that determines the influence of the DAMSM module on the loss function of the second generator, h is the output feature vector of the Transformer, e1 is the word feature, L DAMSM denotes the loss function obtained by training the DAMSM module of the network, F denotes the global attention generation module, and S denotes the neural network downsampling module.
[0102] In the said step 5, the loss function of the secondary discriminator is L D2 = L1 + L2, where L1 denotes discriminating whether the input image is real, and L2 denotes discriminating whether the input image and the text are semantically consistent.
[0103]
[0104]
[0105] Among them, represents the mathematical expectation, x represents the true image corresponding to the text description, represents the image generated by the second generator corresponding to the text description, and e represents the text feature vector after the joint of the sentence feature and the random noise vector.
[0106] Step 6: Downsample the 128*128 image generated in Step 5 to obtain features. The word features are input into a global attention module to obtain new word features. The two features are input into a convolutional neural network for learning to obtain new fused features. The new fused features are input into the third generator to output a 256*256 image. The 256*256 image and the sentence feature are input into a three-level discriminator for discrimination;
[0107] In the said Step 6, the loss function of the third generator is
[0108]
[0109] where K = S(F(F(h, e1))), G3 and D3 are the third generator and the three-level discriminator respectively, λ2 represents the hyperparameter that determines the influence of the DAMSM module on the loss function of the third generator, h represents the output feature vector of the Transformer, e1 represents the word feature, e represents the text feature vector after the joint of the sentence feature and the random noise vector, L DAMSM represents the loss function obtained by training the DAMSM module of the network, and F(h, e1) represents the feature vector learned through the global attention module.
[0110] In the said Step 6, the loss function of the three-level discriminator is L D3 = L1 + L2, where L1 represents discriminating whether the input image is real, and L2 represents discriminating whether the input image and the text are semantically consistent,
[0111]
[0112]
[0113] where represents the mathematical expectation, x represents the true image corresponding to the text description, represents the image generated by the third generator corresponding to the text description, and e represents the text feature vector after the joint of the sentence feature and the random noise vector.
[0114] In the present invention, for the judgment results of the above three discriminators, L1 represents whether the input image is real, which means calculating a number between 0 and 1. When it is 0, it is not real, and when it is 1, it is real. Similarly, L2 represents whether the input image and the text are semantically consistent, which also means calculating a number between 0 and 1. When it is 0, they are inconsistent, and when it is 1, they are consistent.
[0115] Step 7: Use the image generated in Step 6 as the final generated image and output it.
Claims
1. A method for generating an image from text, characterized in that: The method includes the following steps: Step 1: Obtain a dataset composed of texts and corresponding images, and perform preprocessing; Step 2: Construct a text-to-image network model based on AttnGAN. The network model includes a pre-trained network and a multi-stage generation network. The pre-trained network introduces a Transformer module; Step 3: Extract the text features of the text description corresponding to the image. The text features include word features and sentence features. After conditionally enhancing the sentence features, merge them with random noise and input them into the Transformer module to learn spatial and position information, including the following steps: Step 3.1: Obtain a feature vector ẽ from the sentence features through a conditional enhancement module; , Among them, is the sentence feature of the text, is the mean vector of the sentence feature vector of the text, is the covariance matrix of the sentence feature vector of the text, is the distribution sampling of the unit Gaussian distribution N(0,1); Step 3.2: Combine the obtained feature vector ẽ with the random noise vector to obtain , and use as the input of the Transformer module; ; Step 3.3: In the Transformer module, is transformed in the attention space to obtain three representation vectors: , , , and the weight information is calculated , Among them, , represents the weight information of the i-th position when synthesizing the j-th region of the composite image; finally, an image feature matrix with an attention mechanism is obtained , ; Step 3.4: Integrate the feature matrix to obtain the Transformer output feature vector , ; Step 4: Input the feature information learned in Step 3 into the first generator to output a 64*64 low-resolution initial generated image. Input the low-resolution initial generated image and the sentence features into the initial discriminator for discrimination; Step 5: Downsample the low-resolution initial generated image generated in Step 4 to obtain features. Input the word features into a global attention module to obtain new word features. Input the two features together into a convolutional neural network for learning to obtain fused features. Then input the fused features into the second generator to output a 128*128 image. Input the 128*128 image and the sentence features into the secondary discriminator for discrimination; Step 6: Downsample the 128*128 image generated in Step 5 to obtain features. Input the word features into a global attention module to obtain new word features. Input the two features together into a convolutional neural network for learning to obtain new fused features. Input the new fused features into the third generator to output a 256*256 image. Input the 256*256 image and the sentence features into the tertiary discriminator for discrimination; Step 7: Use the image generated in Step 6 as the final generated image and output it.
2. The method for generating an image from a text according to claim 1, characterized in that: In Step 2, the Transformer module includes an Encoder module and a Decoder module; The Encoder module includes three first sub-modules connected in sequence. Any one of the first sub-modules includes a self-attention layer, a normalization layer, and a fully connected layer connected in sequence; The Decoder module includes three second sub-modules connected in sequence. Any one of the second sub-modules includes a first self-attention layer, a first normalization layer, a second self-attention layer, a second normalization layer, and a fully connected layer connected in sequence; The output of the Encoder module is respectively input into the second self-attention layers of the three second sub-modules; The first first sub-module of the Encoder module and the first second sub-module of the Decoder module respectively correspond to the input ends, and the output end of the third second sub-module of the Decoder module corresponds to the output end.
3. A method for generating an image from a text according to claim 1, characterized in that: In Step 4, the loss function of the first generator is , Among them, represents the mathematical expectation, and are the first generator and the initial discriminator respectively, represents the hyperparameter that determines the influence of the DAMSM module on the generator loss function, represents the output feature vector of the Transformer, represents the text feature vector after combining the sentence feature and the random noise vector, represents the loss function obtained by training the DAMSM module of the network.
4. A method for generating an image from a text according to claim 1, characterized in that: In step 4, the loss function of the initial discriminator is , where represents determining whether the input image is real, represents determining whether the input image and the text are semantically consistent. , , Among them, represents the mathematical expectation, represents that the text description corresponds to the real image, represents that the text description corresponds to the image generated by the first generator, represents the text feature vector after the combination of the sentence feature and the random noise vector.
5. A method for generating an image from a text according to claim 1, characterized in that: In Step 5, the loss function of the second generator is , Among them, represents the mathematical expectation, , and are the second generator and the secondary discriminator respectively, represents the hyperparameter that determines the influence of the DAMSM module on the loss function of the second generator, represents the output feature vector of the Transformer, represents the word feature, represents the loss function obtained by training the DAMSM module of the network, represents the global attention generation module, represents the neural network downsampling module.
6. The method for generating an image from a text according to claim 1, wherein: In the said step 5, the loss function of the secondary discriminator is , where represents determining whether the input image is real, represents determining whether the input image and the text are semantically consistent, , , Among them, represents the mathematical expectation, represents that the text description corresponds to the real image, represents that the text description corresponds to the image generated by the second generator, represents the text feature vector after combining the sentence feature and the random noise vector.
7. A method for generating an image from text according to claim 1, characterized in that: In Step 6, the loss function of the third generator is , Among them, , and are the third generator and the three-level discriminator respectively, represents the hyperparameter that determines the influence of the DAMSM module on the loss function of the third generator, represents the output feature vector of the Transformer, represents the word feature, represents the text feature vector after the combination of the sentence feature and the random noise vector, represents the loss function obtained by training the DAMSM module of the network, represents the feature vector learned through the global attention module.
8. A method for generating an image from a text according to claim 1, wherein: In the said step 6, the loss function of the three-level discriminator is , where represents determining whether the input image is real, represents determining whether the input image and the text are semantically consistent , , Among them, represents the mathematical expectation, represents that the text description corresponds to the real image, represents that the text description corresponds to the image generated by the third generator, represents the text feature vector after combining the sentence feature and the random noise vector.
9. A method for generating an image from a text according to claim 1, characterized in that: In the training stage of the text-to-image network model based on AttnGAN, the result of passing the word features through the DAMSM module is compared with the result of passing the 256*256 image through the image decoder, and the text-to-image network model is adjusted based on the comparison result.
Citation Information
Patent Citations
Text-guided image restoration method and system
CN111861945A
Text image generation method based on StackGAN network
CN111968193A
Cited By
Method for generating image by text based on thinking chain and visual priori guidance
CN121708152A
A Text-to-Image Generation Method Based on Thought Chain and Visual Prior Guidance
CN121708152B