Text generation image method, device and computer equipment
The image features are enhanced through dynamic memory attention model and channel attention residual block model, and the problems of missing details and insufficient utilization of channel characteristics in text-generated images are solved, and high-quality, semantic consistent images are generated.
Patent Information
- Application Number
- CN202210620986.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-01
AI Technical Summary
The text image generation method in the prior art has problems such as missing details of generated images and insufficient utilization of channel feature information.
The dynamic memory attention model and channel attention residual block model are used to enhance the visual and channel characteristics of the initial image features through multiple memory, and combined with batch and instance normalization processing to enhance the image feature representation ability.
Generate high-quality and semantic consistent images, reduce details loss, make full use of channel feature information, and improve the quality and consistency of image generation.
Smart Images

Figure CN114937191B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of word processing, and particularly to a method, apparatus, and computer device for generating images from text. Background Art
[0002] Generative Adversarial Networks (GANs) are widely used in conditional image generation, that is, generating corresponding realistic images according to given conditions (text, sketches, semantic segmentation maps, etc.), and ensuring the quality and diversity of the generated images. Text-to-image synthesis (T2I) is one of the challenging tasks among them, which aims to generate corresponding realistic images through given text descriptions. Moreover, with the rise of cross-modal tasks, this task has attracted some researchers in the fields of natural language processing and computer vision. In addition, this task also has important application values in some application fields, such as art creation, image editing, advertising design, etc.
[0003] In the prior art, in order to solve the problem that the model depends on the initial stage to generate images, a dynamic memory network is introduced into the model. The dynamic memory network mainly refines the details in the image through the semantic correlation between the word sequence and the initial generated image. Although it reduces the dependence on the quality of the initial image during the generation process, it fails to fully utilize the channel semantic information contained in the image; using the Deep Attentional Multimodal Similarity Model (DAMSM) can improve the authenticity of the generated image and promote the generator to better learn the key information in the text description to optimize the overall image quality, but it ignores the representation ability of the image visual features and channel features, resulting in the problem of missing details in the generated image.
[0004] Therefore, in the process of generating images from text, there are problems of missing details in the generated images and insufficient utilization of channel feature information. Summary of the Invention
[0005] The main purpose of this application is to provide a method for generating images from text, aiming to solve the technical problems of missing details in the generated images and insufficient utilization of channel feature information in the prior art.
[0006] This application proposes a method for generating images from text, including:
[0007] Obtain multiple text description sentences, and input the multiple text description sentences into a text encoder for encoding to obtain multiple sentence features and multiple word features;
[0008] Obtain multiple random sampling noises, and input the multiple random sampling noises and the multiple sentence features into an initial generator for fusion to obtain multiple initial image features and multiple initial images;
[0009] Input the multiple initial image features and the multiple word features into a dynamic memory attention model and output multiple first image features, where the dynamic memory attention module is used to enhance the visual features of the multiple initial image features;
[0010] Input the multiple first image features into a channel attention residual block model, and output multiple first refined image features. Convolve the multiple first refined image features to obtain multiple first refined images, where the channel attention residual block model is used to enhance the channel features of the first image features;
[0011] Use the multiple first refined image features as initial image features, and input them and the multiple word features into a dynamic memory attention model to output multiple second refined image features, where the dynamic memory attention module is used to enhance the visual features of the multiple first refined image features;
[0012] Input the multiple second refined image features into a channel attention residual block model, and output multiple third refined image features. Convolve the multiple third refined image features to obtain multiple third refined images, where the channel attention residual block model is used to enhance the channel features of the second refined image features.
[0013] Preferably, the step of obtaining multiple random sampling noises, and inputting the multiple random sampling noises and the multiple sentence features into an initial generator for fusion to obtain multiple initial image features and multiple initial images includes:
[0014] Input the multiple random sampling noises and the multiple sentence features into a fully connected layer respectively for preliminary feature fusion, and output multiple first fusion image features;
[0015] Input the multiple preliminary fusion images into a first upsampling block respectively to perform batch normalization processing on the multiple preliminary fusion image features, and output multiple second fusion image features, where the first upsampling block includes at least three consecutive blocks;
[0016] Input the multiple second fusion image features into a second upsampling block to perform instance normalization processing on the multiple second fusion image features, and obtain multiple third fusion image features;
[0017] Output the multiple third fusion image features as initial image features, and perform a convolution operation on the multiple initial image features to obtain multiple initial images.
[0018] Preferably, the step of inputting the multiple preliminary fusion image features into a first upsampling block to perform batch normalization on the multiple preliminary fusion image features and outputting multiple second fusion image features includes:
[0019] Obtaining the batch value for batch normalization;
[0020] Obtaining the first height value H and the first width value of each of the preliminary fusion image features;
[0021] Obtaining the scaling factor and the translation factor self-learned by the first upsampling block during training;
[0022] Obtaining the feature value of the preliminary fusion image feature currently undergoing normalization;
[0023] Calculating the mean of all the preliminary fusion image features according to the batch value, the first height value, and the first width value, where the calculation formula is:
[0024]
[0025] where, μ c represents the mean of all the preliminary fusion image features, N represents the batch value, H represents the first height value, W represents the first width value, and x nchw represents the feature value of the preliminary fusion image feature currently undergoing normalization;
[0026] Calculating the variance of all the preliminary fusion image features according to the batch value, the first height value, and the first width value, where the calculation formula is:
[0027]
[0028] where, represents the variance of all the preliminary fusion image features,
[0029] Calculating the sample distribution of all the preliminary fusion image features after batch normalization according to the variance and the mean, where the calculation formula is:
[0030]
[0031] where, x' represents the sample distribution of the x-th preliminary fusion image feature after batch normalization, x i represents the i-th preliminary fusion image feature, and ε represents a non-zero constant;
[0032] Generating each of the second fusion image features according to the sample distribution, where the generation function is:
[0033] BN(x) = γ × x' + β;
[0034] Among them, BN(x) represents the x-th second fusion image feature, γ represents the scaling factor, and β represents the translation factor.
[0035] Preferably, the step of inputting the plurality of initial image features and the plurality of word features into the dynamic memory attention model and outputting a plurality of first image features includes:
[0036] Calculate a plurality of weight matrices according to the plurality of initial image features and the plurality of word features;
[0037] Take the plurality of weight matrices as a plurality of dynamic memories and store them in the dynamic memory slots;
[0038] Put the plurality of dynamic memories in the dynamic memory slots into the secondary memory feature enhancement unit to refine the image features in the plurality of dynamic memories and obtain a plurality of memory image features;
[0039] Input the plurality of memory image features into the memory response gate to enhance the insignificant regions in the plurality of memory image features and obtain a plurality of first image features.
[0040] Preferably, the step of putting the plurality of dynamic memories in the dynamic memory slots into the secondary memory feature enhancement unit to refine the image features in the plurality of dynamic memories and obtain a plurality of memory image features includes:
[0041] Take the dynamic memory and the initial image feature as the input of the secondary memory feature enhancement unit and perform the first memory feature enhancement to obtain the first memory image feature;
[0042] Perform secondary memory enhancement on the first memory image feature to obtain the memory image feature.
[0043] Preferably, the step of taking the dynamic memory and the initial image feature as the input of the secondary memory feature enhancement unit and performing the first memory feature enhancement to obtain the first memory image feature includes:
[0044] Perform convolution processing on the dynamic memory to obtain a key vector and a value vector;
[0045] Change the dimension of the initial image feature according to the key vector and the value vector so that the dimension of the initial image feature is the same as the dimensions of the key vector and the value vector;
[0046] Perform dot product processing on the initial image feature and the key vector to obtain a spatial weight matrix;
[0047] Calculate the weighted sum with the value vector according to the spatial weight matrix to obtain the first memory image feature.
[0048] Preferably, the step of inputting the plurality of first image features into the channel attention residual block model and outputting a plurality of first refined image features includes:
[0049] Performing convolution and batch normalization on the plurality of first image features, and inputting them into the gated linear activation function GLU to obtain first intermediate features;
[0050] Performing convolution and batch normalization on the first intermediate features to obtain second intermediate features;
[0051] Concatenating the first intermediate features and the second intermediate features in the channel direction to obtain third intermediate features;
[0052] Taking the third intermediate features as the input of the entire attention unit, and performing global average pooling on the third intermediate features to obtain fourth intermediate features;
[0053] Performing a convolution operation on the fourth intermediate features and inputting them into the activation function ReLU to obtain fifth intermediate features;
[0054] Performing a convolution operation on the fifth intermediate features and inputting them into the second activation function Sigmoid to obtain channel semantic weights;
[0055] Performing element-wise multiplication on the channel semantic weights and the second intermediate features to obtain sixth intermediate features;
[0056] Fusing the sixth intermediate features and the first intermediate features to obtain seventh intermediate features;
[0057] Performing a residual connection between the seventh intermediate features and the first image features to obtain first refined image features.
[0058] The present application also proposes a text-to-image generation device, including:
[0059] A first acquisition module, configured to acquire a plurality of text description statements, and input the plurality of text description statements into a text encoder for encoding to obtain a plurality of sentence features and a plurality of word features;
[0060] A second acquisition module, configured to acquire a plurality of randomly sampled noises, and input the plurality of randomly sampled noises and the plurality of sentence features into an initial generator for fusion to obtain a plurality of initial image features and a plurality of initial images;
[0061] A dynamic memory attention model, configured to input the plurality of initial image features and the plurality of word features into the dynamic memory attention model and output a plurality of first image features, wherein the dynamic memory attention module is used to enhance the visual features of the plurality of initial image features;
[0062] A channel attention residual block model is used to input multiple of the first image features into the channel attention residual block model and output multiple first refined image features, and perform convolution on the multiple first refined image features to obtain multiple first refined images, where the channel attention residual block model is used to enhance the channel features of the first image features;
[0063] The dynamic memory attention model is further used to use multiple first refined image features as initial image features and input them into the dynamic memory attention model together with multiple word features, and output multiple second refined image features, where the dynamic memory attention module is used to enhance the visual features of the multiple first refined image features;
[0064] The channel attention residual block model is further used to input multiple of the second refined image features into the channel attention residual block model and output multiple third refined image features, and perform convolution on the multiple third refined image features to obtain multiple third refined images, where the channel attention residual block model is used to enhance the channel features of the second refined image features.
[0065] This application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above text generation image method are implemented.
[0066] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above text generation image method are implemented.
[0067] The beneficial effects of this application are as follows: First, input multiple of the initial image features and multiple word features into the dynamic memory attention model to enhance the visual features of the initial image features. Then, input the first image features obtained from the initial image features into the channel attention residual model, and fuse the output first refined image features with the initial image features generated in the previous stage to achieve the enhancement of the secondary image visual features, that is, using the secondary memory method can enhance the visual feature representation ability of the initial image features in the spatial dimension, further enhance the semantic consistency between the word level and the feature map, and at the same time use the channel attention in the channel attention residual block to enhance the channel feature representation ability of the feature map in the channel dimension, so as to better guide the picture generation, making this application not only able to generate high-quality images, but also able to generate images with better semantic consistency. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is a schematic flowchart of the text generation image method according to an embodiment of this application.
[0069] Figure 2 Schematic structural diagram of a text generation image device according to an embodiment of the present application.
[0070] Figure 3 Schematic internal structure diagram of a computer device according to an embodiment of the present application.
[0071] The realization, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0072] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0073] As Figures 1 - 3 shown, the present application proposes a text generation image method, including:
[0074] S1. Obtain a plurality of text description statements, and input the plurality of text description statements into a text encoder for encoding to obtain a plurality of sentence features and a plurality of word features;
[0075] S2. Obtain a plurality of randomly sampled noises, and input the plurality of randomly sampled noises and the plurality of sentence features into an initial generator for fusion to obtain a plurality of initial image features and a plurality of initial images;
[0076] S3. Input the plurality of initial image features and the plurality of word features into a dynamic memory attention model and output a plurality of first image features, wherein the dynamic memory attention module is used to enhance the visual features of the plurality of initial image features;
[0077] S4. Input the plurality of first image features into a channel attention residual block model, and output a plurality of first refined image features, and perform convolution on the plurality of first refined image features to obtain a plurality of first refined images, wherein the channel attention residual block model is used to enhance the channel features of the first image features;
[0078] S5. Take the plurality of first refined image features as initial image features, and input them into the dynamic memory attention model together with the plurality of word features, and output a plurality of second refined image features, wherein the dynamic memory attention module is used to enhance the visual features of the plurality of first refined image features;
[0079] S6. Input the plurality of second refined image features into a channel attention residual block model, and output a plurality of third refined image features, and perform convolution on the plurality of third refined image features to obtain a plurality of third refined images, wherein the channel attention residual block model is used to enhance the channel features of the second refined image features.
[0080] As described in the above steps S1 - S6, the CUB-200-2011 and Oxford-102 datasets are used to train and test the feature enhancement generative adversarial network model. Both CUB-200-2011 and Oxford-102 have 10 text description sentences corresponding to each image, which contain the visual attributes of the target and a small amount of background information. The CUB-200-2011 dataset consists of 8,855 training images containing 150 bird categories and 2,933 test images containing another 50 bird categories. The Oxford-102 dataset consists of 7,034 training images containing 82 flower categories and 1,155 test images containing another 20 flower categories. By inputting each text description sentence into the text encoder for encoding, the sentence features and word features of each text description sentence can be obtained. Then, random sampling noise is acquired and fused with the sentence features to obtain the initial image features and the initial image. Since the quality of the initial image features is poor, the details of the generated initial image are severely missing, and the channel feature information cannot be fully utilized. Therefore, for the initial image features, a secondary memory enhancement scheme is proposed. That is, first, multiple of the initial image features and multiple of the word features are input into the dynamic memory attention model to enhance the visual features of the initial image features. Then, the first image features obtained from the initial image features are input into the channel attention residual model, and the output first refined image features are fused with the initial image features generated in the previous stage to achieve the enhancement of the secondary image visual features. This not only improves the visual feature expression ability of the image but also further reduces the dependence on the quality of the initial image generated in the initial stage. By introducing the channel attention residual block model, not only can the image semantics at different levels be obtained (where the image semantics are used for information interaction between feature channels), but also the correlation between similar channel semantics can be enhanced, thereby improving the overall representation ability of the channel features. This can reduce the detail loss of the generated second refined image features and fully utilize the channel feature information.
[0081] In one embodiment, the step S2 of acquiring multiple random sampling noises and inputting the multiple random sampling noises and the multiple sentence features into the initial generator for fusion to obtain multiple initial image features and multiple initial images includes:
[0082] S21. Input the multiple random sampling noises and the multiple sentence features into the fully connected layer respectively for preliminary feature fusion, and output multiple first fusion image features;
[0083] S22. Input the multiple preliminarily fused images into the first upsampling block respectively to perform batch normalization on the multiple preliminarily fused image features, and output multiple second fusion image features, where the first upsampling block includes at least three consecutive blocks;
[0084] S23. Input the multiple second fused image features into a second upsampling block to perform instance normalization on the multiple second fused image features, thereby obtaining multiple third fused image features;
[0085] S24. Output the multiple third fused image features as initial image features, and perform a convolution operation on the multiple initial image features to obtain multiple initial images.
[0086] As described in the above steps S21 - S24, adaptive semantic instance normalization can be used to establish a semantic connection between the generated second fused image features and the given text description, and ensure the individual independence of each generated second fused image feature; the sentence features can be used as the affine transformation parameters after instance normalization of the image features, so as to promote the semantic consistency between the text - image pairs; in the prior art, batch normalization (BatchNormalization, BN) is adopted in most classical models for text - to - image generation to maintain the stability of training, but batch normalization ignores the differences between individual samples. Based on this, in the feature - enhanced generative adversarial network model, by performing instance normalization on the second upsampling block, the style diversity of the third fused image features can be increased, and the impact of a small training batch on the generation result can be reduced. Secondly, combining batch normalization processing with instance normalization processing can also improve the performance of the entire generation network.
[0087] In one embodiment, the step S22 of inputting the multiple preliminary fused image features into a first upsampling block to perform batch normalization on the multiple preliminary fused image features and outputting multiple second fused image features includes:
[0088] S221. Obtain the batch value of batch normalization;
[0089] S222. Obtain the first height value H and the first width value of each of the preliminary fused image features;
[0090] S223. Obtain the scaling factor and translation factor self - learned by the first upsampling block during training;
[0091] S224. Obtain the feature value of the preliminary fused image feature currently undergoing normalization processing;
[0092] S225. Calculate the mean of all the preliminary fused image features according to the batch value, the first height value, and the first width value, where the calculation formula is:
[0093]
[0094] where μ crepresents the mean of all the preliminary fusion image features, N represents the batch value, H represents the first height value, W represents the first width value, and x nchw represents the feature value of the preliminary fusion image feature currently undergoing normalization processing;
[0095] S226. Calculate the variance of all the preliminary fusion image features according to the batch value, the first height value, and the first width value, where the calculation formula is:
[0096]
[0097] where, represents the variance of all the preliminary fusion image features,
[0098] S227. Calculate the sample distribution of all the preliminary fusion image features after batch normalization according to the variance and the mean, where the calculation formula is:
[0099]
[0100] where, x' represents the sample distribution of the x-th preliminary fusion image feature after batch normalization, and x i represents the i-th preliminary fusion image feature, and ε represents a non-zero constant;
[0101] S228. Generate each of the second fusion image features according to the sample distribution, where the generation function is:
[0102] BN(x) = γ × x' + β;
[0103] where, BN(x) represents the x-th second fusion image feature, γ represents the scaling factor, and β represents the translation factor.
[0104] In one embodiment, the step S3 of inputting the plurality of initial image features and the plurality of word features into the dynamic memory attention model and outputting a plurality of first image features includes:
[0105] S31. Calculate a plurality of weight matrices according to the plurality of initial image features and the plurality of word features;
[0106] S32. Store the plurality of weight matrices as a plurality of dynamic memories in the dynamic memory slots;
[0107] S33. Put the plurality of dynamic memories in the dynamic memory slots into the secondary memory feature enhancement unit to refine the image features in the plurality of dynamic memories, and obtain a plurality of memory image features;
[0108] S34. Input multiple memory image features into the memory response gate to enhance the insignificant regions in the multiple memory image features, obtaining multiple first image features.
[0109] As described in the above steps S31 - S34, in order to reduce the impact of the unsatisfactory expression of the initial image features on the generation result, first calculate the weight matrix through the initial image features and word features, then store the weight matrix in the dynamic memory slot, and store the dynamic memory in the secondary memory feature enhancement unit to refine the image features in the multiple dynamic memories, obtaining multiple memory image features; input the memory image features into the memory response gate to enhance the insignificant regions in the multiple memory image features while retaining the useful parts, obtaining the first image features.
[0110] Specifically, the mathematical expression of the dynamic memory attention model is as follows:
[0111] I R = RG(MoM(WG(I i-1 )))
[0112] Among them, I R represents the memory image feature, RG represents the memory response gate, MoM represents the secondary memory feature enhancement unit, WG represents the memory write gate, and I i-1 represents the initial image feature generated in the previous stage;
[0113]
[0114] represents the memory write gate control, σ represents the activation function Sigmoid, w i represents the i-th word feature,
[0115] A represents a 1x256 matrix, B represents a 1x64 matrix, represents the initial image feature obtained after global average pooling;
[0116] (The importance of each word feature to the initial image feature can be calculated through the above formula, and then I i-1 is updated to obtain the weight matrix.)
[0117]
[0118] Among them, W w represents the 1x1 convolution of the first weight value, W m represents the 1x1 convolution of the second weight value, where the values of the first weight value and the second weight value are different.
[0119] I M2 = MoM(I M0 )
[0120] I M2 represents the memory image feature, I M0 represents the dynamic memory;
[0121]
[0122] represents the memory response gating, and b represents the offset in the linear function, i.e., the intercept;
[0123]
[0124] In one embodiment, the step S33 of putting multiple dynamic memories in the dynamic memory slots into the secondary memory feature enhancement unit to refine the image features in the multiple dynamic memories to obtain multiple memory image features includes:
[0125] S331. Take the dynamic memory and the initial image feature as the input of the secondary memory feature enhancement unit, and perform the first memory feature enhancement to obtain the first memory image feature;
[0126] S332. Perform secondary memory enhancement on the first memory image feature to obtain the memory image feature.
[0127] As described in the above steps S331 - S332, by putting the dynamic memory into the secondary memory feature enhancement unit, the secondary visual feature enhancement of the memory can be read, further reducing the image visual deviation and strengthening the semantic connection between words and images. Specifically, perform two visual feature enhancement operations on the dynamic memory. Among them, take the dynamic memory and the initial image feature generated in the previous stage as the input of the secondary memory feature enhancement unit, and perform the first memory feature enhancement on the initial image feature through the attention operation to obtain the first memory image feature, and the first memory image feature is used to supplement the missing local semantic information of the initial image feature. Then perform secondary visual feature enhancement on the first memory image feature to obtain the memory image feature.
[0128] In one embodiment, the step S331 of taking the dynamic memory and the initial image feature as the input of the secondary memory feature enhancement unit and performing the first memory feature enhancement to obtain the first memory image feature includes:
[0129] S3311. Perform convolution processing on the dynamic memory to obtain the key vector and the value vector;
[0130] S3312. Change the dimension of the initial image feature according to the key vector and the value vector so that the dimension of the initial image feature is the same as the dimensions of the key vector and the value vector;
[0131] S3313. Perform a dot product operation on the initial image features and the key vectors to obtain a spatial weight matrix;
[0132] S3314. Calculate the weighted sum with the value vectors according to the spatial weight matrix to obtain the first memory image features.
[0133] As described in the above steps S3311 - S3314, the first visual feature enhancement is to find important memories related to the initial image features to improve the visual semantic expression of the overall image. Convolution operations are used to replace the linear operations in traditional attention to process the dynamic memories, obtaining key vectors and value vectors. And the dimension of the initial image features is transformed from c×h×w to (h*w)×c to make it the same as the dimensions of the key vectors and value vectors. To calculate the importance of each memory slot in the dynamic memory for the initial image features, first perform a dot product operation on the initial image features and the key vectors to obtain a spatial weight matrix. Then, calculate the weighted sum with the value vectors according to the spatial weight matrix to obtain the first memory image features. The mathematical expressions are as follows:
[0134]
[0135] w = σ(Key T ⊙ I i-1 )
[0136]
[0137] I M1 = w ⊙ Value,
[0138] where Key represents the key vectors, Value represents the value vectors, ⊙ represents the dot product operation, σ represents the activation function Sigmoid, and are 1×1 convolution operations used to obtain the key vectors and value vectors of the first memory image features.
[0139] The second visual feature enhancement is to further enhance the correlation between the first memory image features and the initial image features, fuse the important semantic information between the two, ignore the irrelevant semantic information, and make it focus on the most important features. Then, perform a re - fusion process on the first memory image features and the initial image features in the channel direction. Then transform the dimension of the fused feature map from (h*w)×c to the original size c×h×w to obtain the memory image features. The entire fusion process is expressed mathematically as follows:
[0140] I M2 = φ([I M1 ; I i-1 )
[0141] Among them, [;] represents the splicing operation, and φ(*) represents the 1×1 convolution operation.
[0142] In one embodiment, the step of inputting the multiple first image features into the channel attention residual block model and outputting multiple first refined image features includes:
[0143] S41. Perform convolution and batch normalization on the multiple first image features, and input them into the gated linear activation function GLU to obtain a first intermediate feature;
[0144] S42. Perform convolution and batch normalization on the first intermediate feature to obtain a second intermediate feature;
[0145] S43. Splice the first intermediate feature and the second intermediate feature in the channel direction to obtain a third intermediate feature;
[0146] S44. Use the third intermediate feature as the input of the entire attention unit, and perform global average pooling on the third intermediate feature to obtain a fourth intermediate feature;
[0147] S45. Perform a convolution operation on the fourth intermediate feature and input it into the activation function ReLU to obtain a fifth intermediate feature;
[0148] S46. Perform a convolution operation on the fifth intermediate feature and input it into the second activation function Sigmoid to obtain a channel semantic weight;
[0149] S47. Perform element-wise multiplication on the channel semantic weight and the second intermediate feature to obtain a sixth intermediate feature;
[0150] S48. Fuse the sixth intermediate feature and the first intermediate feature to obtain a seventh intermediate feature;
[0151] S49. Perform a residual connection on the seventh intermediate feature and the first image feature to obtain a first refined image feature.
[0152] As described in the above steps S41 - S49, in the entire channel attention residual block model, the first intermediate feature and the second intermediate feature contain different levels of semantic information, and the channel map of each feature can be regarded as a response of a specific attribute. The channel semantic information of the first intermediate feature and the second intermediate feature is used to guide the feature expression in the similar semantic regions of the second intermediate feature to enhance the feature representation ability of the entire image. In the entire channel attention unit, first, the first intermediate feature and the second intermediate feature are spliced in the channel direction to obtain a third intermediate feature, and the third intermediate feature is used as the input of the entire attention unit. Then, a global average pooling operation is performed on the third intermediate feature to focus on the channel information of the third intermediate feature. The mathematical definition is as follows:
[0153] h avg = GAP([O1; O2]);
[0154] Among them, h avg represents the third intermediate feature, O1 represents the first intermediate feature, O2 represents the second intermediate feature, [;] represents the splicing operation; GAP represents global average pooling.
[0155] Then, two consecutive convolutional layers are used to perform information interaction on the third intermediate feature in the channel direction, clarify the correlation between channels, and enhance important channel features. Finally, an activation function is used to judge the importance of each feature channel for the entire feature map to obtain the channel semantic weight. The mathematical expression is as follows:
[0156] w = σ(W1(W2(h avg ))))
[0157] Among them, w is the channel semantic weight, σ represents the Sigmoid function; w represents the channel weight value; W1 and W2 both represent 1×1 convolutional operations. The first convolution W1 is to fuse the channel information between the third intermediate feature and the fourth intermediate feature and perform interaction. The second convolution is to perform channel dimension transformation to match the number of channels of the second intermediate feature, facilitating subsequent channel feature enhancement. Finally, the obtained channel semantic weight is multiplied element-wise with the second intermediate feature. This operation not only enhances the channel semantic expression within the similar semantic region of the second intermediate feature but also enhances the channel features of the second intermediate feature. To further enrich the semantic information in the second intermediate feature, the first intermediate feature is fused with the enhanced sixth intermediate feature, and a residual connection is made with the initially input first image feature to maintain the stability of the entire model training. The mathematical expression is as follows:
[0158] O3 = w × O2;
[0159] h' = O3 + O1 + h;
[0160] Among them, O3 represents the sixth intermediate feature; h is the first image feature, and h' represents the output first refined image feature.
[0161] In addition, the loss of the entire feature enhancement generative adversarial network model is the linear sum of the generator loss, conditional enhancement loss, and DAMSM loss, as follows:
[0162]
[0163] Among them, represents the generator loss, L CA represents the conditional enhancement loss, L DAMSMRepresents the loss of the deep multi-modal attention similarity model. CA means mapping the input sentence vector to an independent Gaussian distribution, and resampling the sentence vector in this distribution for training data augmentation and avoiding overfitting. During this process, the similarity between the two distributions is calculated, and this similarity degree serves as its loss. The DAMSM loss calculates the matching loss between text and image at the word level to determine the semantic matching degree between the generated image and the text description. The entire network loss function is defined as follows:
[0164]
[0165]
[0166] Among them, λ1 and λ2 are hyperparameters, both set to 1. The first half of Equation (1) represents only focusing on the quality of the generated image, and the second half represents determining the quality of the generated image based on the correlation between the generated image and the text description. μ(s) and σ 2 (s) in Equation (2) represent the mean and variance of the sentence vector respectively. At the same time, all discriminator loss functions are defined as follows:
[0167]
[0168] Among them, the unconditional loss is to distinguish between real images and generated images, and the conditional loss is to further determine the matching degree between the generated image and the text description. The purpose of this discriminant is to ensure the consistency between the generated image and the text description while ensuring the quality of the generated image.
[0169] As Figure 3 shown, this application also provides a computer device, which can be a server, and its internal structure can be as Figure 3 shown. The computer device includes a processor, a memory, a display screen, an input device, a network interface, and a database connected through a system bus. Among them, the processor of this computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all the data required for the process of the text generation image method. The network interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes the text generation image method.
[0170] This application also proposes a text generation image device, including:
[0171] The first acquisition module 1 is used to acquire a plurality of text description statements, and input the plurality of text description statements into a text encoder for encoding to obtain a plurality of sentence features and a plurality of word features;
[0172] The second acquisition module 2 is used to acquire a plurality of randomly sampled noises, and input the plurality of randomly sampled noises and the plurality of sentence features into an initial generator for fusion to obtain a plurality of initial image features and a plurality of initial images;
[0173] The dynamic memory attention model 3 is used to input the plurality of initial image features and the plurality of word features into the dynamic memory attention model and output a plurality of first image features, wherein the dynamic memory attention module is used to enhance the visual features of the plurality of initial image features;
[0174] The channel attention residual block model 4 is used to input the plurality of first image features into the channel attention residual block model, and output a plurality of first refined image features, and perform convolution on the plurality of first refined image features to obtain a plurality of first refined images, wherein the channel attention residual block model is used to enhance the channel features of the first image features;
[0175] The dynamic memory attention model is further used to use the plurality of first refined image features as initial image features, and input them into the dynamic memory attention model together with the plurality of word features, and output a plurality of second refined image features, wherein the dynamic memory attention module is used to enhance the visual features of the plurality of first refined image features;
[0176] The channel attention residual block model is further used to input the plurality of second refined image features into the channel attention residual block model, and output a plurality of third refined image features, and perform convolution on the plurality of third refined image features to obtain a plurality of third refined images, wherein the channel attention residual block model is used to enhance the channel features of the second refined image features.
[0177] In one embodiment, the second acquisition module 2 includes:
[0178] The preliminary fusion unit is used to input the plurality of randomly sampled noises and the plurality of sentence features into a fully connected layer respectively for preliminary feature fusion, and output a plurality of first fusion image features;
[0179] The first upsampling block is used to input the plurality of preliminary fusion images into the first upsampling block respectively to perform batch normalization processing on the plurality of preliminary fusion image features, and output a plurality of second fusion image features, wherein the first upsampling block includes at least three consecutive blocks;
[0180] A second upsampling block, configured to input multiple second fused image features into the second upsampling block to perform instance normalization processing on the multiple second fused image features, thereby obtaining multiple third fused image features;
[0181] A convolution unit, configured to output multiple third fused image features as initial image features, and perform convolution operations on the multiple initial image features to obtain multiple initial images.
[0182] In one embodiment, the first upsampling block includes:
[0183] A first obtaining unit, configured to obtain the batch value of batch normalization;
[0184] A second obtaining unit, configured to obtain the first height value H and the first width value of each of the preliminary fused image features;
[0185] A third obtaining unit, configured to obtain the scaling factor and translation factor autonomously learned by the first upsampling block during training;
[0186] A fourth obtaining unit, configured to obtain the feature value of the preliminary fused image feature currently undergoing normalization processing;
[0187] A first calculation unit, configured to calculate the mean of all the preliminary fused image features according to the batch value, the first height value, and the first width value, where the calculation formula is:
[0188]
[0189] where, μ c represents the mean of all the preliminary fused image features, N represents the batch value, H represents the first height value, W represents the first width value, and x nchw represents the feature value of the preliminary fused image feature currently undergoing normalization processing;
[0190] A second calculation unit, configured to calculate the variance of all the preliminary fused image features according to the batch value, the first height value, and the first width value, where the calculation formula is:
[0191]
[0192] where, represents the variance of all the preliminary fused image features,
[0193] A third calculation unit, configured to calculate the sample distribution of all the preliminary fused image features after batch normalization according to the variance and the mean, where the calculation formula is:
[0194]
[0195] Among them, x' represents the sample distribution of the x-th preliminary fusion image feature after batch normalization, and x i represents the i-th preliminary fusion image feature, and ε represents a non-zero constant;
[0196] A generation unit, configured to generate each of the second fusion image features according to the sample distribution, where the generation function is:
[0197] BN(x) = γ × x' + β;
[0198] Among them, BN(x) represents the x-th second fusion image feature, γ represents a scaling factor, and β represents a translation factor.
[0199] In one embodiment, the dynamic memory attention model 3 includes:
[0200] A weight matrix unit, configured to calculate a plurality of weight matrices according to the plurality of initial image features and the plurality of word features;
[0201] A storage unit, configured to store the plurality of weight matrices as a plurality of dynamic memories in a dynamic memory slot;
[0202] A secondary memory feature enhancement unit, configured to put the plurality of dynamic memories in the dynamic memory slot into the secondary memory feature enhancement unit to refine the image features in the plurality of dynamic memories to obtain a plurality of memory image features;
[0203] A memory response gating, configured to input the plurality of memory image features into the memory response gating to enhance the insignificant regions in the plurality of memory image features to obtain a plurality of first image features.
[0204] In one embodiment, the secondary memory feature enhancement unit includes:
[0205] A first memory feature enhancement unit, configured to use the dynamic memory and the initial image feature as the input of the secondary memory feature enhancement unit and perform first memory feature enhancement to obtain a first memory image feature;
[0206] A secondary memory enhancement unit, configured to perform secondary memory enhancement on the first memory image feature to obtain a memory image feature.
[0207] In one embodiment, the first memory feature enhancement unit includes:
[0208] A convolution processing unit, configured to perform convolution processing on the dynamic memory to obtain a key vector and a value vector;
[0209] A dimension change unit for changing the dimension of the initial image feature according to the key vector and the value vector so that the dimension of the initial image feature is the same as that of the key vector and the value vector;
[0210] A dot product processing unit for performing a dot product operation on the initial image feature and the key vector to obtain a spatial weight matrix;
[0211] A weighted sum unit for calculating a weighted sum with the value vector according to the spatial weight matrix to obtain a first memory image feature.
[0212] In one embodiment, the channel attention residual block model includes:
[0213] A first intermediate feature unit for performing convolution and batch normalization on a plurality of first image features and inputting them into a gated linear activation function GLU to obtain a first intermediate feature;
[0214] A second intermediate feature unit for performing convolution and batch normalization on the first intermediate feature to obtain a second intermediate feature;
[0215] A splicing unit for splicing the first intermediate feature and the second intermediate feature in the channel direction to obtain a third intermediate feature;
[0216] A global average pooling unit that takes the third intermediate feature as the input of the entire attention unit and performs global average pooling on the third intermediate feature to obtain a fourth intermediate feature;
[0217] A fifth intermediate feature unit for performing a convolution operation on the fourth intermediate feature and inputting it into an activation function ReLU to obtain a fifth intermediate feature;
[0218] A channel semantic weight unit for performing a convolution operation on the fifth intermediate feature and inputting it into a second activation function Sigmoid to obtain a channel semantic weight;
[0219] An element-wise multiplication processing unit for performing element-wise multiplication on the channel semantic weight and the second intermediate feature to obtain a sixth intermediate feature;
[0220] A seventh intermediate feature unit for fusing the sixth intermediate feature and the first intermediate feature to obtain a seventh intermediate feature;
[0221] A residual connection unit for performing a residual connection on the seventh intermediate feature and the first image feature to obtain a first refined image feature.
[0222] Those skilled in the art can understand, Figure 3The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied.
[0223] An embodiment of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned arbitrary text generation image method is implemented.
[0224] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0225] It should be noted that in this article, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article, or method including that element.
[0226] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of this application by the same token.
Claims
1. A method for generating an image from text, characterized in that, Including: S1. Obtain multiple text description statements, and input the multiple text description statements into a text encoder for encoding to obtain multiple sentence features and multiple word features; S2. Obtain multiple randomly sampled noises, and input the multiple randomly sampled noises and the multiple sentence features into an initial generator for fusion to obtain multiple initial image features and multiple initial images; S3. Input the multiple initial image features and the multiple word features into a dynamic memory attention model and output multiple first image features. In this step, the dynamic memory attention model is used to enhance the visual features of the multiple initial image features. Step S3 specifically includes the following sub-steps: S3.
1. Calculate multiple weight matrices according to the multiple initial image features and the multiple word features; S3.
2. Use the multiple weight matrices as multiple dynamic memories and store them in a dynamic memory slot; S3.
3. Put the multiple dynamic memories in the dynamic memory slot into a secondary memory feature enhancement unit to refine the image features in the multiple dynamic memories to obtain multiple memory image features. Step S3.3 specifically includes the following sub-steps: S3.3.
1. Use the dynamic memory and the initial image feature as the input of the secondary memory feature enhancement unit and perform primary memory feature enhancement to obtain a first memory image feature. Step S3.3.1 specifically includes the following sub-steps: S3.3.1.
1. Perform convolution processing on the dynamic memory to obtain a key vector and a value vector; S3.3.1.
2. Change the dimension of the initial image feature according to the key vector and the value vector so that the dimension of the initial image feature is the same as that of the key vector and the value vector; S3.3.1.
3. Perform dot product processing on the initial image feature and the key vector to obtain a spatial weight matrix; S3.3.1.
4. Calculate the weighted sum with the value vector according to the spatial weight matrix to obtain a first memory image feature; S3.3.
2. Perform secondary memory enhancement on the first memory image feature to obtain a memory image feature; S3.
4. Input the multiple memory image features into a memory response gating to enhance the insignificant regions in the multiple memory image features to obtain multiple first image features; S4. Input the multiple first image features into a channel attention residual block model and output multiple first refined image features. Perform convolution on the multiple first refined image features to obtain multiple first refined images, where the channel attention residual block model is used to enhance the channel features of the first image features; S5. Use the multiple first refined image features as initial image features and input them into the dynamic memory attention model together with the multiple word features to output multiple second refined image features. In this step, the dynamic memory attention model is used to enhance the visual features of the multiple first refined image features; S6. Input the multiple second refined image features into the channel attention residual block model, and output multiple third refined image features. Convolve the multiple third refined image features to obtain multiple third refined images, where the channel attention residual block model is used to enhance the channel features of the second refined image features.
2. The method for generating an image according to the text described in claim 1, characterized in that In step S2, obtain multiple randomly sampled noises, and input the multiple randomly sampled noises and the multiple sentence features into an initial generator for fusion to obtain multiple initial image features and multiple initial images, which specifically includes the following sub-steps: S2.
1. Input the multiple randomly sampled noises and the multiple sentence features into a fully connected layer respectively for preliminary feature fusion, and output multiple first fused image features; S2.
2. Input the multiple first fused image features into a first upsampling block respectively to perform batch normalization on the multiple first fused image features, and output multiple second fused image features, where the first upsampling block includes at least three consecutive blocks; S2.
3. Input the multiple second fused image features into a second upsampling block to perform instance normalization on the multiple second fused image features to obtain multiple third fused image features; S2.
4. Output the multiple third fused image features as initial image features, and perform a convolution operation on the multiple initial image features to obtain multiple initial images.
3. The method for generating an image according to claim 2, wherein In step S2.2, input the multiple first fused image features into a first upsampling block to perform batch normalization on the multiple first fused image features, and output multiple second fused image features, which specifically includes the following sub-steps: S2.2.
1. Obtain the batch value for batch normalization; S2.2.
2. Obtain the first height value H and the first width value of each of the first fused image features; S2.2.
3. Obtain the scaling factor and translation factor learned autonomously by the first upsampling block during training; S2.2.
4. Obtain the feature value of the first fused image feature currently undergoing normalization; S2.2.
5. Calculate the mean of all the first fused image features according to the batch value, the first height value, and the first width value, where the calculation formula is: Among them, μ c represents the mean value of all the first fusion image features, N represents the batch value, H represents the first height value, W represents the first width value, and x nchw represents the feature value of the first fusion image feature currently undergoing normalization processing; S2.2.
6. Calculate the variance of all the first fused image features according to the batch value, the first height value, and the first width value, where the calculation formula is: Among them, represents the variance of all the first fusion image features, S2.2.
7. Calculate the sample distribution of all the first fused image features after batch normalization according to the variance and the mean, where the calculation formula is: Among them, x′ represents the sample distribution of the x-th first fused image feature after batch normalization, and x i represents the i-th first fused image feature, and ε represents a non-zero constant; S2.2.
8. Generate each of the second fused image features according to the sample distribution, where the generation function is: BN(x) = γ × x' + β; where BN(x) represents the x-th second fused image feature, γ represents the scaling factor, and β represents the translation factor.
4. The method for generating an image from a text according to claim 1, wherein In step S4, input the multiple first image features into the channel attention residual block model, and output multiple first refined image features, which specifically includes the following sub-steps: S4.
1. Convolve and perform batch normalization on multiple first image features, and input them into the gated linear unit (GLU) to obtain first intermediate features; S4.
2. Convolve and perform batch normalization on the first intermediate features to obtain second intermediate features; S4.
3. Concatenate the first intermediate features and the second intermediate features in the channel direction to obtain third intermediate features; S4.
4. Use the third intermediate features as the input of the entire attention unit, and perform global average pooling on the third intermediate features to obtain fourth intermediate features; S4.
5. Perform a convolution operation on the fourth intermediate features and input them into the ReLU activation function to obtain fifth intermediate features; S4.
6. Perform a convolution operation on the fifth intermediate features and input them into the second activation function Sigmoid to obtain channel semantic weights; S4.
7. Perform element-wise multiplication on the channel semantic weights and the second intermediate features to obtain sixth intermediate features; S4.
8. Fuse the sixth intermediate features and the first intermediate features to obtain seventh intermediate features; S4.
9. Perform residual connection on the seventh intermediate features and the first image features to obtain first refined image features.
5. A text generation image device, characterized in that, It includes: A first acquisition module, configured to acquire multiple text description statements, and input the multiple text description statements into a text encoder for encoding to obtain multiple sentence features and multiple word features; A second acquisition module, configured to acquire multiple randomly sampled noises, and input the multiple randomly sampled noises and the multiple sentence features into an initial generator for fusion to obtain multiple initial image features and multiple initial images; A dynamic memory attention model, configured to input the multiple initial image features and the multiple word features into the dynamic memory attention model and output multiple first image features, where the dynamic memory attention model is used to enhance the visual features of the multiple initial image features; and is further configured to use the multiple first refined image features as initial image features, and input the multiple first refined image features and the multiple word features into the dynamic memory attention model to output multiple second refined image features, where the dynamic memory attention model is used to enhance the visual features of the multiple first refined image features; A channel attention residual block model, configured to input the multiple first image features into the channel attention residual block model and output multiple first refined image features, and perform convolution on the multiple first refined image features to obtain multiple first refined images, where the channel attention residual block model is used to enhance the channel features of the first image features; The channel attention residual block model is further configured to input the multiple second refined image features into the channel attention residual block model and output multiple third refined image features, and perform convolution on the multiple third refined image features to obtain multiple third refined images, where the channel attention residual block model is used to enhance the channel features of the second refined image features.
6. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method for generating an image from text according to any one of claims 1 to 4 are implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the text generation image method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method for generating image by sensing combined space attention text
CN114387366A
Apparatus and method for detecting scene text in an image
US20190130204A1