Text image generation method based on encoder joint optimization
Through the joint optimization of encoder and text to generate image method, the network is built using generator and discriminator, and the attention mechanism and vector quantization are added to solve the problem of low image generation accuracy in the existing technology and achieve more efficient image generation effect.
Patent Information
- Application Number
- CN202510827194.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-23
AI Technical Summary
The images generated by existing text-to-image methods have low accuracy, especially in complex layout generation tasks where there is a lack of natural connections and smooth transitions between elements, and they are limited by computing resources and model complexity.
A text-to-image method based on encoder joint optimization is adopted. A text-to-image network with encoder joint optimization is constructed through the generator and discriminator. The attention mechanism and vector quantization are added to learn discrete latent space representation to capture the main information of the image and carefully control the generation process.
The accuracy and detail richness of image generation are improved, and the generation effect of the model is enhanced.
Smart Images

Figure CN120689452A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of text generation, and in particular relates to a method for generating images from text based on encoder joint optimization. Background Art
[0002] With the rapid development of multimedia and network technologies, the amount of cross-media data, including images, videos, and text, is rapidly increasing. Cross-media intelligence technologies have emerged as a response to this trend. They enable the retrieval, recognition, analysis, reasoning, design, and prediction of multimodal information to enhance AI's performance in human perception. Text and images have always been a key focus of cross-media intelligence research, and this research has been widely applied in areas such as online community management, robotics, and intelligent driving.
[0003] Using text-to-image technology, we can convert text information into image data, thereby increasing the diversity and richness of datasets. This can provide more training samples for many tasks, such as computer vision tasks, image classification, and object detection. By using text-to-image technology to expand datasets, we can improve the accuracy and generalization ability of models and promote better performance.
[0004] Since the proposal of generative adversarial networks, text-generated images, image editing and conversion, video generation and prediction, natural language processing, medical image processing, and artistic creation have developed rapidly. However, text-generated images still face some challenges. For example, when generating complex objects, especially in complex image synthesis tasks, due to the complexity of the layout structure, the synthesized objects are often easily placed in unreasonable positions in the image, resulting in chaotic layout. In recent years, some models aim to improve the rationality of layout generation and have established some spatial models. However, there is a lack of natural connections and smooth transitions between elements in current text-generated images, and when dealing with complex layout generation tasks, some models may be limited by computing resources and model complexity, resulting in poor generation effects.
[0005] Therefore, the existing method of generating images through text has the problem of low accuracy of the generated images. Summary of the Invention
[0006] The purpose of this invention is to solve the problem of low accuracy of text-to-image generation in the existing technology. We propose a text-to-image generation method based on encoder joint optimization.
[0007] S1. Obtain training set;
[0008] S2. Build an encoder-jointly optimized text-to-image network to obtain a trained encoder-jointly optimized text-to-image network.
[0009] S3: Input the guidance text into the trained encoder and combine it with the optimized text-to-image network to output the generated image.
[0010] The specific process of obtaining the training set in S1 is:
[0011] S1.1: Get the dataset.
[0012] S1.2: Preprocess the data set to obtain a preprocessed data set;
[0013] S1.3: Construct a codebook and feed the preprocessed data set into the codebook to obtain a training set
[0014] The text generation image network jointly optimized by the encoder in S2 includes: a generator and a discriminator;
[0015] The generator includes: a joint encoder and a decoder;
[0016] The specific process of obtaining the trained encoder and jointly optimizing the text-to-image network is as follows:
[0017] S2.1: Input the images in the training set into the joint encoder to obtain a continuous feature vector A;
[0018] S2.2: Obtain a discrete latent representation B based on the trained codebook and continuous feature vector A in the training set;
[0019] S2.3: Input the discrete potential representation B into the decoder to obtain the intermediate feature C;
[0020] S2.4: Input the intermediate feature C into the discriminator for training. Stop the iteration after N rounds of training to obtain the trained encoder-jointly optimized text-to-image network, where N is a positive integer.
[0021] The beneficial effects of the present invention are:
[0022] 1. Generative adversarial networks learn a discrete latent space representation through vector quantization, which helps the model capture the main information of the image more efficiently.
[0023] 2. Adding the attention mechanism can more finely control the image generation process and generate images with more details. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of the overall structure of the present invention;
[0025] Figure 2 Schematic diagram of the image encoder module structure of the present invention;
[0026] Figure 3It is a schematic diagram of the structure of the image decoder module of the present invention. DETAILED DESCRIPTION
[0027] Specific implementation method 1: Combination Figure 1 The present invention is described, comprising:
[0028] S1. Obtain training set;
[0029] S2. Build an encoder-jointly optimized text-to-image network to obtain a trained encoder-jointly optimized text-to-image network.
[0030] S3: Input the guidance text into the trained encoder and combine it with the optimized text-to-image network to output the generated image.
[0031] The specific process of obtaining the training set in S1 is:
[0032] S1.1: Get the dataset.
[0033] S1.2: Preprocess the data set to obtain a preprocessed data set;
[0034] S1.3: Construct a codebook and feed the preprocessed data set into the codebook to obtain a training set
[0035] The text generation image network jointly optimized by the encoder in S2 includes: a generator and a discriminator;
[0036] The generator includes: a joint encoder and a decoder;
[0037] The specific process of obtaining the trained encoder and jointly optimizing the text-to-image network is as follows:
[0038] S2.1: Input the images in the training set into the joint encoder to obtain a continuous feature vector A;
[0039] S2.2: Obtain a discrete latent representation B based on the trained codebook and continuous feature vector A in the training set;
[0040] S2.3: Input the discrete potential representation B into the decoder to obtain the intermediate feature C;
[0041] S2.4: Input the intermediate feature C into the discriminator for training. Stop the iteration after N rounds of training to obtain the trained encoder-jointly optimized text-to-image network, where N is a positive integer.
[0042] Specific embodiment 2: The difference between this embodiment and specific embodiment 1 is that:
[0043] The dataset obtained in S1.1 includes the COCO dataset; the COCO dataset is the MS COCO open source dataset, which contains 123,287 images, each of which has image annotations of 80 categories and 5 titles;
[0044] The annotations are provided in JSON format and include image name, image size, object category, bounding box coordinates, key points and locations. The dataset can be obtained from: https: / / cocodataset.org / #download;
[0045] The specific process of dividing the dataset is as follows: the dataset is divided into a training set, a validation set, and a test set. In this example, 113,287 images and their corresponding descriptions are used as the training set, 5,000 images and their corresponding descriptions are used as the validation set, and 5,000 images and their corresponding descriptions are used as the test set.
[0046] Other steps and parameters are the same as those in the first embodiment.
[0047] Specific embodiment three: This embodiment differs from specific embodiment one in that:
[0048] In S1.2, the divided data set is preprocessed to obtain a preprocessed data set. The specific process is as follows:
[0049] S1.2.1: Use the torch.nn.functional.normalize function to scale the pixel RGB values of the images in the dataset from the range [0, 255] to [0, 1].
[0050] S1.2.2: Use the center cropping method to crop the images in the dataset scaled to [0, 1] to obtain a cropped image set; the size of the cropped images is 256×256;
[0051] S1.2.3: Annotate the cropped image set to obtain the preprocessed dataset;
[0052] The other steps and parameters are the same as those in the first and second embodiments.
[0053] Specific embodiment 4: This embodiment differs from specific embodiments 1 to 4 in that:
[0054] In S1.2.3, the cropped image set is annotated to obtain the preprocessed dataset. The specific process is as follows:
[0055] S1.2.3.1: Clean the image annotations of the cropped image set. The specific process is as follows:
[0056] Delete stop words and punctuation marks in the image annotations, and then convert all uppercase letters in the image annotations into lowercase letters to obtain the cleaned image annotations;
[0057] S1.2.3.2: Use the Bert word segmentation model to segment the cleaned image annotations to obtain the preprocessed dataset;
[0058] The Bert word segmentation model is used to perform word segmentation processing on the cleaned image annotations. The specific process is: dividing the text of the cleaned image annotations into different tokens;
[0059] The Bert word segmentation model is an existing model integrated in the transformers library. Install the transformers library through pip install transformers, and then import the Bert word segmentation model through from transformers import BertTokenizer. The specific operation process is as follows:
[0060] Use [CLS] to indicate the beginning of a sequence, [SEP] to separate different sentences or paragraphs, and [PAD] to fill shorter sequences. Use the Bert word segmentation model to segment the text of the image annotation into different tokens. These tokens will find the corresponding embedding vectors in the embedding layer of the model. During the word segmentation process, the uniform length is 512 tokens. Use the padding method to expand short tokens, and the [PAD] mark. Use the truncation method to truncate long input tokens. The truncation method deletes the extra tokens from the end of the sequence and only retains the first half of the sequence.
[0061] The other steps and parameters are the same as those in the first to third embodiments.
[0062] Specific embodiment 5: This embodiment differs from specific embodiments 1 to 4 in that:
[0063] The specific process of constructing the codebook in S1.3 and sending the preprocessed data set into the codebook to obtain the training set is as follows:
[0064] S1.3.1: Set the vector capacity of the codebook and initialize the embedding layer of the codebook using nn.Embedding; obtain the codebook;
[0065] The embedding layer represents a lookup vocabulary for retrieving vectors from the codebook. The vocabulary size is 10,000 and the embedding dimension is 300.
[0066] The vector capacity in the codebook in S1.3.1 is 1024, and the dimension of each vector is 256.
[0067] S1.3.2: Train the codebook based on the preprocessed data set to obtain a trained codebook;
[0068] S1.3.3: Use the trained codebook and preprocessed dataset as the training set
[0069] In S1.3.2, the codebook is trained based on the preprocessed data set to obtain a trained codebook. The specific process is as follows:
[0070] S1.3.2.1: Generate the parameter beta using a random number generator and initialize the codebook weights.
[0071] The parameter beta follows a uniform distribution and is used to adjust the weight in the loss function.
[0072] S1.3.2.2: Use the preprocessed dataset as the input vector to the embedding layer of the codebook for forward propagation and vector quantization. The codebook outputs the quantized vector.
[0073] S1.3.2.3: Calculate the codebook loss based on the codebook input and output. When the loss is minimized, the trained codebook is obtained.
[0074] The codebook loss consists of two parts: the first part is the mean square error between the input vector and the quantized vector; the second part is the mean square error between the input vector and the vector in the codebook multiplied by beta;
[0075] The other steps and parameters are the same as those in the first to fourth embodiments.
[0076] Specific embodiment 6: This embodiment differs from specific embodiments 1 to 5 in that:
[0077] In S2.1, the image in the training set is input into the joint encoder to obtain a continuous feature vector A. The specific process is:
[0078] S2.1.1: Input the images in the training set into the first convolutional layer and the first residual layer; obtain the intermediate feature A1;
[0079] S2.1.2: Input the intermediate feature A1 into the second convolutional layer and the second residual layer in sequence to obtain the intermediate feature A2;
[0080] S2.1.3: Input the intermediate feature A2 into the third convolutional layer and the third residual layer in sequence to obtain the intermediate feature A3;
[0081] S2.1.4: Input the intermediate feature A3 into the fourth convolutional layer and the fourth residual layer in sequence to obtain the intermediate feature A4;
[0082] S2.1.5: Input the intermediate feature A4 into the fifth convolutional layer and the fifth residual layer in sequence to obtain the intermediate feature A5;
[0083] S2.1.6: Input intermediate feature A5 into the first attention unit and the first sampling layer in sequence to obtain intermediate feature A6;
[0084] S2.1.7: Add intermediate feature A6, intermediate feature A5, intermediate feature A4, intermediate feature A3, intermediate feature A2, and intermediate feature A1 to obtain intermediate feature A7;
[0085] S2.1.8: Feed the intermediate feature A7 into the second attention layer, the first normalization layer, the Swish function layer, and the sixth convolutional layer in sequence to obtain the intermediate feature A.
[0086] The convolution kernel sizes of the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, and the fifth convolution layer are all 3×3;
[0087] The first residual layer, the second residual layer, the third residual layer, the fourth residual layer, and the fifth residual layer are each composed of a group normalization layer, a Swish activation function layer, and a convolutional layer;
[0088] The first sampling layer includes a sampling layer divided into a first downsampling layer and a first upsampling layer, and the first downsampling layer and the first upsampling layer are both convolution layers with a convolution kernel of 3×3;
[0089] The first convolutional layer has 3 input channels, 128 output channels, a 3×3 kernel size, a stride size of 1, and a padding size of 1. The convolution result is added to the downsampling block, and the image size is reduced from 256×256 to 128×128.
[0090] The second convolutional layer has 128 input channels and 128 output channels, a 3×3 kernel size, a stride of 1, and a padding of 1. A first residual layer is constructed between the first and second layers. This residual layer has 128 input channels and 128 output channels. The result of the residual layer is added to the downsampling block, reducing the image size from 128×128 to 64×64.
[0091] The third convolutional layer has 128 input channels and 256 output channels, a 3×3 kernel size, a stride of 1, and a padding of 1. A second residual layer is constructed between the second and third layers. This residual layer has 128 input channels and 256 output channels. The above results are added to the downsampling block, reducing the image size from 64×64 to 32×32.
[0092] The fourth convolutional layer has 256 input channels and 256 output channels, a 3×3 kernel size, a stride of 1, and a padding of 1. A third residual layer is constructed between the third and fourth layers. This residual layer has 256 input channels and 256 output channels. The above results are added to the downsampling block, reducing the image size from 32×32 to 16×16.
[0093] The fifth convolutional layer has 256 input channels and 512 output channels, a 3×3 kernel size, a stride size of 1, and a padding size of 1. A fourth residual layer is constructed between the fourth and fifth layers. The residual layer has 256 input channels and 512 output channels.
[0094] The above results are added to the attention layer NonLocalBlock, and then input to the fifth residual layer, NonLocalBlock, residual layer, GroupNorm and Swish function layer. Finally, the results are input to the convolution layer with 512 input channels and 256 output channels. The activation function corresponding to the five convolutional layer sequences is ReLU, and each convolution layer is followed by a batch normalization layer with the same number of channels.
[0095] Residual layer: The residual layer contains two parameters and two components. The parameters are the number of input channels and the number of output channels. The components include self.block and self.channel_up. self.block is a sequential model. The first layer in its internal structure is a group normalization layer, which divides each feature map into multiple groups based on the number of input channels and then normalizes each group. The second layer is the Swish activation function. In the self.channel_up component, when the number of input and output channels differs, a 1×1 convolution is used to adjust the number of channels so that the channels match during the residual connection. The third layer is a convolution layer with a kernel size of 3×3, a stride size of 1, and a padding size of 1. The initialization module in the downsampling block includes a kernel size of 3×3, a stride size of 2, and a padding size of 0. The forward propagation uses the F.pad function to pad the input. The padding mode is constant padding, and the padding value is 0. The number of pixels to be filled on the four edges (left, top, right, and bottom) of the input. The left and top edges are filled with 0 pixels, and the right and bottom edges are filled with 1 pixel.
[0096] 2.1.3) Attention Layer: The attention layer first normalizes the input feature map group to obtain a normalized feature map. Three independent one-dimensional convolutional layers, self.q, self.k, and self.v, are used to convolve the normalized feature map, respectively, to obtain three sets of vectors: Query (q), Key (k), and Value (v). The q, k, and v dimensions are combined with the spatial dimensions to form a shape suitable for matrix operations. Batch matrix multiplication is performed. The batch matrix multiplication torch.bmm is used to calculate the dot product between the Query and Key vectors to obtain a preliminary attention score. The attention score at each position is normalized using the softmax function to convert it into a valid probability distribution. Finally, the normalized attention score is weightedly summed on the Value vector to obtain the weighted aggregated feature. The aggregated feature is element-wise added to the original input feature map to form a residual connection.
[0097] 2.1.4) Sampling Layer: The sampling layer consists of a downsampling layer and an upsampling layer. The initialization module for the downsampling layer includes a convolutional layer with a kernel size of 3×3, a stride size of 2, and a padding size of 0. The forward propagation uses the F.pad function to pad the input. The padding mode is constant padding, and the padding value is 0. The four edges of the input (left, top, right, and bottom) are padded with the number of pixels to be padded: 0 pixels on the left and top edges, and 1 pixel on the right and bottom edges. The initialization module for the upsampling layer includes a convolutional layer with a kernel size of 3×3, a stride size of 1, and a padding size of 1. The input feature map is upsampled using the F.interpolate function, with a scale_factor of 2.0, doubling the size of the feature map.
[0098] 2.1.5) Fully connected layer: The output of the convolutional layer is expanded and connected to the fully connected layer. The number of input features of the fully connected layer is 65536, and the number of output features is 256. In the forward propagation, nn.Sequential is used to combine all added layers into an ordered container. The residual layer inherits from nn.Module in PyTorch. It contains two parameters and two components. The parameters are the number of input channels and the number of output channels; the components include self.block and self.channel_up. self.block is a sequential model. The first layer in the internal structure is the group normalization layer. The group normalization layer divides each feature map into multiple groups according to the number of input channels, and then normalizes them within each group; the second layer is the Swish activation function; in the self.channel_up component, when the number of input and output channels is different, a 1×1 convolution is used to adjust the number of channels so that the number of channels can match during the residual connection. The third layer is the convolution layer. The convolution kernel size is 3×3, the stride size is 1, and the padding size is 1.
[0099] The process of attention layer processing in the present invention is as follows:
[0100] Assume that the attention layer inputs a batch of sequence data x, and the sequence data x format is represented as a tensor of shape [B, T, C].
[0101] Where B is the batch size, T is the sequence length, and C is the number of channels (feature dimension).
[0102] The input sequence x is mapped through three different linear layers to obtain the query vector Q, key vector K and value vector V respectively.
[0103] Then the attention score x1 is obtained by performing a dot product calculation using the query vector Q and the key vector K.
[0104] Apply an upper triangular mask to ensure that each position can only see the position before it when calculating attention, and get the masked attention score x2
[0105] The masked attention score x2 is then normalized using the softmax function to obtain the attention weight x3.
[0106] The value vector V is weighted and summed using the normalized attention weights to obtain the weighted value vector representation V1.
[0107] The weighted value vector representation V1 is projected through another linear layer to obtain the final output sequence x4 as the output of the attention layer; the other steps and parameters are the same as those of the specific implementation methods one to five.
[0108] Specific embodiment 7: This embodiment differs from specific embodiments 1 to 6 in that:
[0109] In S2.2, a discrete potential representation B is obtained based on the codebook and the continuous feature vector A. The specific process is as follows:
[0110] The codebook calculates the Euclidean distance and selects a codeword closest to the feature code to represent the input data.
[0111] The other steps and parameters are the same as those in the first to sixth embodiments.
[0112] Specific embodiment eight: This embodiment differs from specific embodiments one to seven in that:
[0113] In S2.3, the discrete potential representation B is input into the decoder to obtain the intermediate feature C. The specific process is as follows:
[0114] S2.3.1: The discrete potential representation B is sequentially input into the seventh convolutional layer and the sixth residual layer to obtain the intermediate feature C1;
[0115] S2.3.2: Input the intermediate feature C1 into the eighth convolutional layer and the seventh residual layer in sequence to obtain the intermediate feature C2;
[0116] S2.3.4: Input the intermediate feature C2 into the ninth convolutional layer and the eighth residual layer in sequence to obtain the intermediate feature C3;
[0117] S2.3.5: Input the intermediate feature C3 into the tenth convolutional layer and the ninth residual layer in sequence to obtain the intermediate feature C4;
[0118] S2.3.6: Input the intermediate feature C4 into the eleventh convolutional layer and the ninth residual layer in sequence to obtain the intermediate feature C5;
[0119] S2.3.7: Add intermediate feature C5, intermediate feature C4, intermediate feature C3, intermediate feature C2, and intermediate feature C1 to obtain intermediate feature C6;
[0120] S2.3.8: Feed the intermediate feature C6 into the third attention layer, the twelfth convolutional layer, the fourth attention layer, the second upsampling layer, the second normalization layer, the second Swish activation layer, and the thirteenth convolutional layer in sequence.
[0121] The encoder contains five layers:
[0122] The first convolutional layer has 512 input channels, 512 output channels, a 3×3 kernel size, a stride size of 1, and a padding size of 1. The convolution result is added to the residual layer and then connected to the attention layer. The input parameter of the attention layer is 512.
[0123] The second convolutional layer has 512 input channels and 256 output channels, a 3×3 kernel size, a stride of 1, and a padding of 1. A residual layer is constructed between the first and second layers, with 128 input channels and 128 output channels. The above results are added to the upsampling block, increasing the image size from 32×32 to 64×64.
[0124] The third convolutional layer has 256 input channels and 256 output channels, a 3×3 kernel size, a stride of 1, and a padding of 1. A residual layer is constructed between the second and third layers. This residual layer has 256 input channels and 256 output channels. The above results are added to the upsampling block, reducing the image size from 64×64 to 128×128.
[0125] The fourth convolutional layer has 256 input channels and 128 output channels, a 3×3 kernel size, a stride of 1, and a padding of 1. A residual layer is constructed between the third and fourth layers, with 256 input channels and 128 output channels. The above results are added to the upsampling block, and the image size is enlarged from 128×128 to 256×256.
[0126] The fifth convolution layer has 128 input channels, 128 output channels, a 3×3 convolution kernel size, a stride size of 1, and a padding size of 1. A residual layer is constructed between the fourth and fifth layers. The residual layer has 128 input channels and 128 output channels. The above results are added to the upsampling block, and the image size is enlarged from 256×256 to 512×512. The results are then input to the GroupNorm and Swish function layers. Finally, the results are input to the convolution layer with 128 input channels, a 3×3 convolution kernel size, a stride size of 1, and a padding size of 1. The activation function corresponding to the five convolution layer sequences is ReLU, and each convolution operation is followed by a batch normalization layer with the same number of channels.
[0127] The other steps and parameters are the same as those in the first to seventh embodiments.
[0128] Specific embodiment 9: This embodiment differs from specific embodiments 1 to 8 in that:
[0129] The convolution kernel sizes of the seventh convolution layer, the eighth convolution layer, the ninth convolution layer, the tenth convolution layer, the eleventh convolution layer, the twelfth convolution layer and the thirteenth convolution layer are all 3×3.
[0130] The other steps and parameters are the same as those in the first to eighth embodiments.
[0131] Specific embodiment 10: This embodiment differs from specific embodiments 1 to 9 in that:
[0132] The discriminator includes: a first circulation layer, a second circulation layer, and a third circulation layer;
[0133] The first recurrent layer includes a fourteenth convolutional layer, a third normalization layer, and a first LeakyReLU activation layer;
[0134] The second recurrent layer includes a fifteenth convolutional layer, a fourth normalization layer, and a second LeakyReLU activation layer;
[0135] The third circulation layer includes a sixteenth convolutional layer, a fifth normalization layer, a third LeakyReLU activation layer, and a seventeenth convolutional layer;
[0136] The convolution kernel sizes of the fourteenth convolution layer, the fifteenth convolution layer, the sixteenth convolution layer, and the seventeenth convolution layer are all 4x4;
[0137] The discriminator is a separate network. The first layer in the initial layer is a convolutional layer. The number of filters in the convolutional layer is initialized to 2, the number of channels of the input image is initialized to 3, the number of output channels is 64, the convolution kernel size is 4, the stride is 2, and the padding is set to 1; followed by an activation function LeakyReLU with a negative slope set to 0.2.
[0138] The discriminator consists of three recurrent layers. The first recurrent layer adds a convolutional layer to the initial layer. The input image has 64 channels, the output has 128 channels, the convolution kernel size is 4, the stride is 2, and the padding is set to 1. The convolutional layer is followed by a normalization layer to 128 channels. The output of the first recurrent layer is then activated with a LeakyReLU with a negative slope of 0.2.
[0139] In the second loop, we add a layer based on the first loop. The convolutional layer is initialized with 4 filters, 256 channels for the input image, and 256 channels for the output. The kernel size is 4, the stride is 2, the padding is 1, and the activation function is LeakyReLU with a negative slope of 0.2. The convolutional layer is followed by a normalization layer to 256 channels. The output of the second loop layer is then activated with a LeakyReLU with a negative slope of 0.2.
[0140] In the third loop, a layer is added based on the second loop. The number of filters is initialized to 8, the number of channels of the input image is initialized to 512, the number of output channels is 512, the convolution kernel size is 4, the stride is 1, the padding is set to 1, the activation function is LeakyReLU, and the negative slope is set to 0.2. The convolution layer is followed by a normalization layer, the number of normalized channels is 512, and then the output of the third loop layer is obtained after activation with LeakyReLU with a negative slope of 0.2. Finally, a convolution layer is added, the number of channels of the input image is initialized to 512, the number of output channels is 1, the convolution kernel size is 4, the stride is 1, and the padding is set to 1. Finally, training is performed in the layer
[0141] The other steps and parameters are the same as those in the first to ninth embodiments.
[0142] Combined with the simulation analysis of specific implementation methods one to ten
[0143] The above only describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above-mentioned specific implementation methods. Although the present invention has been disclosed as above with preferred embodiments, it is not intended to limit the present invention. Any technician familiar with this profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent replacements and improvements made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A text-to-image generation method based on encoder joint optimization, characterized in that: The following steps are involved: S1. Obtain training set; S2. Build an encoder-jointly optimized text-to-image network to obtain a trained encoder-jointly optimized text-to-image network. S3: Input the guidance text into the trained encoder and combine it with the optimized text-to-image network to output the generated image. The specific process of obtaining the training set in S1 is: S1.1: Get the dataset. S1.2: Preprocess the data set to obtain a preprocessed data set; S1.3: Construct a codebook; obtain a trained codebook, and then use the trained codebook and the processed dataset as the training set; The text generation image network jointly optimized by the encoder in S2 includes: a generator and a discriminator; The generator includes: a joint encoder and a decoder; The specific process of obtaining the trained encoder-jointly optimized text-to-image network is as follows: S2.1: Input the images in the training set into the joint encoder to obtain a continuous feature vector A; S2.2: Obtain a discrete latent representation B based on the trained codebook and continuous feature vector A in the training set; S2.3: Input the discrete potential representation B into the decoder to obtain the intermediate feature C; S2.4: Input the intermediate feature C into the discriminator for training. Stop the iteration after N rounds of training to obtain the trained encoder-jointly optimized text-to-image network, where N is a positive integer.
2. The text-to-image generation method based on encoder joint optimization according to claim 1, characterized in that: The dataset obtained in S1.1 is the COCO dataset.
3. The text-to-image generation method based on encoder joint optimization according to claim 2, characterized in that: In S1.2, the divided data set is preprocessed to obtain a preprocessed data set. The specific process is as follows: S1.2.1: Scale the pixel RGB values of the images in the dataset from the range [0, 255] to [0, 1]; S1.2.2: Crop the images in the dataset scaled to [0, 1] to obtain a cropped image set; the size of the cropped images is 256×256; S1.2.3: Annotate the cropped image set to obtain the preprocessed dataset.
4. The text-to-image generation method based on encoder joint optimization according to claim 3 is characterized in that: In S1.2.3, the cropped image set is annotated to obtain the preprocessed dataset. The specific process is as follows: S1.2.3.1: Clean the image annotations of the cropped image set. The specific process is as follows: Delete stop words and punctuation marks in the image annotations, and then convert all uppercase letters in the image annotations into lowercase letters to obtain the cleaned image annotations; S1.2.3.2: Perform word segmentation on the cleaned image annotations to obtain the preprocessed dataset.
5. The text-to-image generation method based on encoder joint optimization according to claim 4 is characterized in that: The specific process of constructing a codebook in S1.3, obtaining a trained codebook, and then using the trained codebook and the processed data set as a training set is as follows: S1.3.1: Set the vector capacity of the codebook and initialize the embedding layer of the codebook; obtain the codebook; The vector capacity in the codebook in S1.3.1 is 1024, and the dimension of each vector is 256. S1.3.2: Train the codebook based on the preprocessed data set to obtain a trained codebook; S1.3.3: Use the trained codebook and preprocessed dataset as the training set; In S1.3.2, the codebook is trained based on the preprocessed data set to obtain a trained codebook. The specific process is as follows: S1.3.2.1: Generate the parameter beta using a random number generator and initialize the codebook weights. The parameter beta follows a uniform distribution; S1.3.2.2: Use the preprocessed dataset as the input vector to the embedding layer of the codebook for forward propagation and vector quantization. The codebook outputs the quantized vector. S1.3.2.3: Calculate the codebook loss based on the codebook input and output. When the loss is minimized, the trained codebook is obtained. The codebook loss consists of two parts: the first part is the mean square error between the input vector and the quantized vector; the second part is the mean square error between the input vector and the vector in the codebook multiplied by beta.
6. The method for generating images from text based on encoder joint optimization according to claim 5, characterized in that: In S2.1, the image in the training set is input into the joint encoder to obtain a continuous feature vector A. The specific process is: S2.1.1: Input the images in the training set into the first convolutional layer and the first residual layer; obtain the intermediate feature A1; S2.1.2: Input the intermediate feature A1 into the second convolutional layer and the second residual layer in sequence to obtain the intermediate feature A2; S2.1.3: Input the intermediate feature A2 into the third convolutional layer and the third residual layer in sequence to obtain the intermediate feature A3; S2.1.4: Input the intermediate feature A3 into the fourth convolutional layer and the fourth residual layer in sequence to obtain the intermediate feature A4; S2.1.5: Input the intermediate feature A4 into the fifth convolutional layer and the fifth residual layer in sequence to obtain the intermediate feature A5; S2.1.6: Input intermediate feature A5 into the first attention unit and the first sampling layer in sequence to obtain intermediate feature A6; S2.1.7: Add intermediate feature A6, intermediate feature A5, intermediate feature A4, intermediate feature A3, intermediate feature A2, and intermediate feature A1 to obtain intermediate feature A7; S2.1.8: Feed the intermediate feature A7 into the second attention layer, the first normalization layer, the Swish function layer, and the sixth convolutional layer, yielding a continuous feature vector A. The convolution kernel sizes of the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, and the fifth convolution layer are all 3×3; The first residual layer, the second residual layer, the third residual layer, the fourth residual layer, and the fifth residual layer are each composed of a group normalization layer, a Swish activation function layer, and a convolutional layer; The first sampling layer includes a sampling layer divided into a first downsampling layer and a first upsampling layer, and the first downsampling layer and the first upsampling layer are both convolution layers with a convolution kernel of 3×3.
7. The method for generating images from text based on encoder joint optimization according to claim 6, characterized in that: In S2.2, a discrete potential representation B is obtained based on the trained codebook and the continuous feature vector A in the training set. The specific process is as follows: S2.2.1: Calculate the Euclidean distance between the continuous feature vector A and the discrete variables in the trained codebook; S2.2.2: Select the discrete variable in the codebook with the smallest Euclidean distance as the discrete latent representation B.
8. The method for generating images from text based on encoder joint optimization according to claim 7, characterized in that: In S2.3, the discrete potential representation B is input into the decoder to obtain the intermediate feature C. The specific process is as follows: S2.3.1: The discrete potential representation B is sequentially input into the seventh convolutional layer and the sixth residual layer to obtain the intermediate feature C1; S2.3.2: Input the intermediate feature C1 into the eighth convolutional layer and the seventh residual layer in sequence to obtain the intermediate feature C2; S2.3.4: Input the intermediate feature C2 into the ninth convolutional layer and the eighth residual layer in sequence to obtain the intermediate feature C3; S2.3.5: Input the intermediate feature C3 into the tenth convolutional layer and the ninth residual layer in sequence to obtain the intermediate feature C4; S2.3.6: Input the intermediate feature C4 into the eleventh convolutional layer and the ninth residual layer in sequence to obtain the intermediate feature C5; S2.3.7: Add intermediate feature C5, intermediate feature C4, intermediate feature C3, intermediate feature C2, and intermediate feature C1 to obtain intermediate feature C6; S2.3.8: Input the intermediate feature C6 into the third attention layer, the twelfth convolutional layer, the fourth attention layer, the second sampling layer, the second normalization layer, the second Swish activation layer, and the thirteenth convolutional layer in sequence;.
9. The method for generating images from text based on encoder joint optimization according to claim 8, characterized in that: The convolution kernel sizes of the seventh convolution layer, the eighth convolution layer, the ninth convolution layer, the tenth convolution layer, the eleventh convolution layer, the twelfth convolution layer and the thirteenth convolution layer are all 3×3.
10. The method for generating images from text based on encoder joint optimization according to claim 9, characterized in that: The discriminator includes in sequence: a first cycle layer, a second cycle layer, and a third cycle layer; The first recurrent layer includes a fourteenth convolutional layer, a third normalization layer, and a first LeakyReLU activation layer; The second recurrent layer includes a fifteenth convolutional layer, a fourth normalization layer, and a second LeakyReLU activation layer; The third circulation layer includes a sixteenth convolutional layer, a fifth normalization layer, a third LeakyReLU activation layer, and a seventeenth convolutional layer; The convolution kernel sizes of the fourteenth convolution layer, the fifteenth convolution layer, the sixteenth convolution layer and the seventeenth convolution layer are all 4x4.