Text Generation Image Method Driven by Gated Cross-Word-Visual Attention

By using the gated crossword-visual attention unit in the process of text generation, the problem of inaccurate word importance estimation in the prior art is solved, and richer fine-grained information generation and higher image-text matching are achieved.

CN115438211BActive Publication Date: 2025-06-10SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210947726.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-05
Publication Date
2025-06-10
Estimated Expiration
2042-08-05

AI Technical Summary

Technical Problem

The prior art is difficult to accurately estimate the importance of each word in different image sub-regions during text image generation process, resulting in the neglect of important words and the loss of fine-grained information.

Method used

The importance of each word is estimated and fine-grained information is generated on the image subregion by using a gated crossword-visual attention unit connected by word to visual attention block, selection gate, visual to word attention block.

Benefits of technology

Effectively select important words, enrich the fine-grained information of the image, enhance the matching between the image and text description, and the generated image details are richer, close to the real image, and more in line with the text description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438211B_ABST
    Figure CN115438211B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating images based on gated cross-word-visual attention drive, comprising the following steps: extracting sentence feature vectors and word feature matrices from text descriptions, and obtaining conditional feature vectors by performing conditional enhancement processing on the sentence feature vectors, then inputting the conditional feature vectors and random noise vectors into a visual feature transformer and a generator to obtain a low-resolution image; inputting the word feature matrix and the visual feature matrix into a gated cross-word-visual attention unit to obtain a refined word feature matrix and a refined visual feature matrix, then inputting the refined visual feature matrix into the visual feature transformer and the generator to obtain a high-resolution image; repeating the above steps to obtain an image with a higher resolution; introducing an improved objective function to enhance the authenticity of the generated image and the semantic consistency with the text description, and taking the image with the highest resolution as the final generated image. Through the method of the present invention, higher-quality images can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and particularly relates to a method for generating images from text driven by gated cross-word-visual attention. Background Art

[0002] On the premise of explosive data growth, people urgently need a more efficient information reception method. Text visualization is one of them, which makes it easier for people to obtain and understand complex text information. Therefore, converting text into corresponding images has become an important research hotspot in recent years.

[0003] To strengthen the fusion of text information and image information and generate images with rich fine-grained information, current research methods mainly adopt the attention mechanism. By focusing on relevant words in the text description, fine-grained information is generated in different image sub-regions. However, if the attention mechanism cannot accurately estimate the importance of each word in different image sub-regions at one time, important words will be ignored, easily resulting in the loss of fine-grained information. Summary of the Invention

[0004] The purpose of the present invention is to solve the above-mentioned defects in the prior art, and provide a method for generating images from text driven by gated cross-word-visual attention. The basic unit for realizing the method of generating images from text is a gated cross-word-visual attention unit, which is composed of a word-to-visual attention block, a selection gate, and a visual-to-word attention block connected in series. Among them, first, the semantic information amount of relevant words included in the generated image is obtained through the word-to-visual attention block, providing a basis for selecting important words; then, the importance of each word is determined by comparing the semantic information amount contained in each word with the semantic information amount of relevant words included in the generated image through the selection gate, and important words in the image generation process are selected; finally, fine-grained information of relevant words is generated on the image sub-region through the visual-to-word attention block. By using the gated cross-word-visual attention unit in multiple stages and introducing an improved objective function, it is ensured that the selected important words will not be lost, and the generated images have richer fine-grained information and are more in line with the text description.

[0005] The purpose of the present invention can be achieved by adopting the following technical solutions:

[0006] A method for generating images from text driven by gated cross-word-visual attention, the method for generating images from text includes the following steps:

[0007] S1. Extract the sentence feature vector and the word feature matrix of the first stage from the text description, and obtain the conditional feature vector by processing the sentence feature vector through conditional enhancement. Then, input the conditional feature vector and the random noise vector into the visual feature transformer of the first stage to obtain the visual feature matrix of the first stage. Next, input the visual feature matrix of the first stage into the generator of the first stage to obtain the first-resolution image, and the resolution of the first-resolution image is 64×64;

[0008] S2. Input the word feature matrix and the visual feature matrix of the first stage into the gated cross word-visual attention unit of the first stage to obtain the refined word feature matrix and the refined visual feature matrix of the first stage, and use the refined word feature matrix of the first stage as the word feature matrix of the second stage. Then, input the refined visual feature matrix of the first stage into the visual feature transformer of the second stage to obtain the visual feature matrix of the second stage. Next, input the visual feature matrix of the second stage into the generator of the second stage to obtain the second-resolution image, and the resolution of the second-resolution image is 128×128;

[0009] S3. Input the word feature matrix and the visual feature matrix of the second stage into the gated cross word-visual attention unit of the second stage to obtain the refined word feature matrix and the refined visual feature matrix of the second stage, and use the refined word feature matrix of the second stage as the word feature matrix of the third stage. Then, input the refined visual feature matrix of the second stage into the visual feature transformer of the third stage to obtain the visual feature matrix of the third stage. Next, input the visual feature matrix of the third stage into the generator of the third stage to obtain the third-resolution image, and the resolution of the third-resolution image is 256×256;

[0010] S4. Introduce an improved objective function, enhance the authenticity of the generated image in each stage and the semantic consistency between the generated image and the text description by minimizing the objective function, and use the third-resolution image generated in the third stage as the finally generated high-quality image.

[0011] Furthermore, the word feature matrices of the first, second, and third stages are each composed of multiple word feature vectors. Use N w to represent the number of word feature vectors in the word feature matrices of the first, second, and third stages, and D w to represent the dimension of the word feature vectors of the first, second, and third stages; the visual feature matrices of the first, second, and third stages are each composed of multiple visual feature vectors. Use to represent the number of visual feature vectors in the visual feature matrices of the first, second, and third stages respectively, and D v to represent the dimension of the visual feature vectors of the first, second, and third stages.

[0012] Furthermore, the gated cross-word-visual attention units in the first and second stages are each composed of a word-to-visual attention block, a selection gate, and a visual-to-word attention block connected in series; the visual feature transformer in the first stage is composed of 1 fully connected layer and 4 upsampling blocks connected in series, and the visual feature transformers in the second and third stages are each composed of 2 residual blocks and 1 upsampling block connected in series; the generators in the first, second, and third stages are each composed of 1 3×3 convolutional layer.

[0013] Furthermore, the word-to-visual attention block in the gated cross-word-visual attention unit of the first stage takes the visual feature matrix and word feature matrix of the first stage as inputs, and the output is the local visual feature matrix of the first stage; the word-to-visual attention block in the gated cross-word-visual attention unit of the second stage takes the visual feature matrix and word feature matrix of the second stage as inputs, and the output is the local visual feature matrix of the second stage; the calculation process of the word-to-visual attention block is as follows: First, the input visual feature matrix is subjected to feature mapping through a 1×1 convolutional layer to obtain a visual feature matrix in the word feature semantic space; then, the input word feature matrix and the visual feature matrix in the word feature semantic space are multiplied through matrix multiplication to obtain a similarity matrix; then, the similarity matrix is normalized along the last dimension to obtain an attention weight coefficient matrix; then, the visual feature matrix in the word feature semantic space and the attention weight coefficient matrix are multiplied through matrix multiplication to obtain a visual context feature matrix; finally, the visual context feature matrix and the input word feature matrix are subjected to feature concatenation and passed through two linear transformation layers and a sigmoid activation function to obtain a local visual feature matrix; the expression is as follows:

[0014] V i ′=M v (V i ),i=1,2; (1)

[0015] ɑ i =softmax(W i T V i ′),i=1,2; (2)

[0016]

[0017] Among them, V i represents the visual feature matrix of the i-th stage of the input, with a dimension of W i represents the word feature matrix of the i-th stage of the input, with a dimension of D w ×N w ;V i′ represents the visual feature matrix in the word feature semantic space at the i-th stage, with dimensions of W i T V i ′ represents the similarity matrix at the i-th stage, with dimensions of α i represents the attention weight coefficient matrix at the i-th stage, with dimensions of V i ′ɑ i T represents the visual context feature matrix at the i-th stage, with dimensions of D w ×N w ; represents the local visual feature matrix at the i-th stage of the output, with dimensions of D w ×N w ; M v () represents a 1×1 convolutional layer, and the subscript v at the lower right indicates that the input feature is in the visual feature semantic space; and represent the first and second linear transformation layers, and the subscript w at the lower right indicates that the input feature is in the word feature semantic space, has dimensions of D w ×D w , has dimensions of D w ; σ() represents the sigmoid activation function, represents element-wise multiplication, and the superscript T at the upper right indicates matrix inversion.

[0018] Furthermore, the word-to-visual attention block in the first and second stage gated cross word-visual attention units first distinguishes the image sub-regions containing relevant word semantic information through the attention weight coefficient matrix. For the image sub-regions lacking or containing less relevant word semantic information, their attention weight coefficients are smaller, while for the image sub-regions containing more relevant word semantic information, their attention weight coefficients are larger. Then, it fuses the relevant word semantic information contained in each image sub-region through the visual context feature matrix, and further suppresses the attention weight coefficients of the image sub-regions lacking or containing less relevant word semantic information through the local visual feature matrix, more accurately determining the semantic information volume of the relevant words contained in the generated image, thus providing a basis for selecting important words.

[0019] Furthermore, the selection gate in the first-stage gated cross word-visual attention unit takes the local visual feature matrix and word feature matrix in the first stage as inputs, and the output is the refined word feature matrix in the first stage; the selection gate in the second-stage gated cross word-visual attention unit takes the local visual feature matrix and word feature matrix in the second stage as inputs, and the output is the refined word feature matrix in the second stage; the calculation process of the selection gate is: passing the input local visual feature matrix and word feature matrix through two linear transformation layers and a sigmoid activation function to obtain the refined word feature matrix; the expression is as follows:

[0020]

[0021] Among them, represents the local visual feature matrix of the i-th stage of the input, with a dimension of D w ×N w ; W i represents the word feature matrix of the i-th stage of the input, with a dimension of D w ×N w ; represents the refined word feature matrix of the i-th stage of the output, with a dimension of D w ×N w ; and represent the first and second linear transformation layers, and the lower right subscript w indicates that the input feature is in the word feature semantic space, has a dimension of 1×D w ; σ() represents the sigmoid activation function.

[0022] Furthermore, the selection gates in the gated cross word-visual attention units of the first and second stages compare the semantic information content contained in each word and the semantic information content of the relevant words contained in the generated image through two linear transformation layers and a sigmoid activation function. When the semantic information content of the relevant words contained in the generated image is much less than the semantic information content of the word, the importance of this word is significantly improved, thus realizing the selection of important words.

[0023] Furthermore, the visual-to-word attention block in the first-stage gated cross word-visual attention unit takes the first-stage refined word feature matrix and visual feature matrix as inputs, and the output is the first-stage refined visual feature matrix; the visual-to-word attention block in the second-stage gated cross word-visual attention unit takes the second-stage refined word feature matrix and visual feature matrix as inputs, and the output is the second-stage refined visual feature matrix; the calculation process of the visual-to-word attention block is as follows: First, the input refined word feature matrix is subjected to feature mapping through a 1×1 convolutional layer to obtain a word feature matrix in the visual feature semantic space; then, the word feature matrix in the visual feature semantic space and the input visual feature matrix are multiplied through matrix multiplication to obtain a similarity matrix; then, the similarity matrix is normalized along the last dimension to obtain an attention weight coefficient matrix; then, the word feature matrix in the visual feature semantic space and the attention weight coefficient matrix are multiplied through matrix multiplication to obtain a word context feature matrix; finally, the word context feature matrix and the input visual feature matrix are subjected to feature concatenation and passed through two linear transformation layers and a sigmoid activation function to obtain a refined visual feature matrix; the expression is as follows:

[0024]

[0025] Wherein, represents the input i-th stage refined word feature matrix, with a dimension of D w ×N w ; V i represents the input i-th stage visual feature matrix, with a dimension of represents the i-th stage word feature matrix in the visual feature semantic space, with a dimension of D v ×N w ; represents the i-th stage similarity matrix, with a dimension of β i represents the i-th stage attention weight coefficient matrix, with a dimension of represents the i-th stage word context feature matrix, with a dimension of represents the output i-th stage refined visual feature matrix, with a dimension of M w () represents a 1×1 convolutional layer, and the lower right subscript w indicates that the input feature is in the word feature semantic space; and represent the first and second linear transformation layers, and the lower right subscript v indicates that the input feature is in the visual feature semantic space, has a dimension of D v ×D v , has a dimension of D v; σ() represents the sigmoid activation function, represents element-wise multiplication, and the superscript T represents matrix transpose.

[0026] Furthermore, in the visual-to-word attention blocks of the first and second stage gated cross word-visual attention units, the words containing relevant visual semantic information are first distinguished by the attention weight coefficient matrix. For words lacking or containing less relevant visual semantic information, their attention weight coefficients are smaller, while for words containing more relevant visual semantic information, their attention weight coefficients are larger. Then, the relevant visual semantic information contained in each word is fused through the word context feature matrix, and through the refined visual feature matrix, the attention weight coefficients of the words lacking or containing less relevant visual semantic information are further suppressed, thereby generating more important fine-grained information on each image sub-region.

[0027] Furthermore, the process of step S1 is as follows:

[0028] Define the process of extracting the sentence feature vector and the word feature matrix in the first stage as ENC(), and define the process of conditional enhancement processing to obtain the conditional feature vector as CA(); define the process of the visual feature transformer in the first stage to output the visual feature matrix in the first stage as F 1 (), and define the process of the generator in the first stage to output the first resolution image as G 1 (), and the expressions are as follows:

[0029] (s, W 1 ) = ENC(Text); (8)

[0030] s ca = CA(s); (9)

[0031] V 1 = F 1 (concat(s ca , z)); (10)

[0032] I 1 = G 1 (V 1 ) (11)

[0033] where Text represents the text description; s represents the sentence feature vector with dimension D s ; W 1 represents the word feature matrix in the first stage with dimension D w ×N w ; s ca represents the conditional feature vector with dimension D s ; z represents a standard normal distribution The random noise vector, with a dimension of D z ; V 1 represents the visual feature matrix in the first stage, with a dimension of I 1 represents the first-resolution image output by the generator in the first stage; concat() represents the feature concatenation function that concatenates multiple input feature matrices into one output feature matrix.

[0034] Furthermore, the process of step S2 is as follows:

[0035] Define the process of the gated cross-word-visual attention unit in the first stage using formulas (1)-(7) to output the refined word feature matrix and refined visual feature matrix as GCAU 1 (), define the process of the visual feature transformer in the second stage to output the visual feature matrix in the second stage as F 2 (), define the process of the generator in the second stage to output the second-resolution image as G 2 (), and the expression is as follows:

[0036]

[0037] I 2 = G 2 (V 2 ) (15)

[0038] where represents the refined word feature matrix in the first stage, with a dimension of D w × N w ; represents the refined visual feature matrix in the first stage, with a dimension of W 2 represents the word feature matrix in the second stage, with a dimension of D w × N w ; V 2 represents the visual feature matrix in the second stage, with a dimension of I 2 represents the second-resolution image output by the generator in the second stage; concat() represents the feature concatenation function that concatenates multiple input feature matrices into one output feature matrix.

[0039] Furthermore, the process of step S3 is as follows:

[0040] Define the process of the gated cross-word-visual attention unit in the second stage using formulas (1)-(7) to output the refined word feature matrix and refined visual feature matrix as GCAU 2 (), define the process of the visual feature transformer in the third stage to output the visual feature matrix in the third stage as F3 (), define the process of the generator in the third stage to output the third-resolution image as G 3 (), the expression is as follows:

[0041]

[0042] I 3 = G 3 (V 3 )(19)

[0043] Among them, represents the refined word feature matrix in the second stage, with a dimension of D w ×N w ; represents the refined visual feature matrix in the second stage, with a dimension of W 3 represents the word feature matrix in the third stage, with a dimension of D w ×N w ; V 3 represents the visual feature matrix in the third stage, with a dimension of I 3 represents the third-resolution image output by the generator in the third stage, and also represents the finally generated high-quality image; concat() represents the feature concatenation function, which concatenates multiple input feature matrices into one output feature matrix.

[0044] Furthermore, the word feature matrix in the second stage in step S3 is the same as the refined word feature matrix in the first stage, which ensures that the important words selected in the first stage will not be lost during the image generation process in the second stage, so as to continue to enrich the fine-grained information of the image.

[0045] Furthermore, the process of step S4 is as follows:

[0046] Define the improved objective function as L, and the expression is as follows:

[0047]

[0048] Among them, represents the adversarial loss function of the generator in the i-th stage, L CA represents the conditional enhancement loss function; represents the improved image-text matching loss function in the i-th stage; λ 1 , λ 2 represents the weight coefficient.

[0049] Furthermore, define the expression of the adversarial loss function of the generator in the i-th stage as follows:

[0050]

[0051] Among them, D i () represents a probability value between 0 and 1, and I i represents the image output by the generator in the i-th stage, and I i comes from the data distribution generated by the generator in the i-th stage s represents the sentence feature vector.

[0052] Furthermore, the conditional enhancement loss function is defined as the KL divergence between the standard Gaussian distribution and the Gaussian distribution of the training data, and the expression is as follows:

[0053]

[0054] Among them, ∑(s)) represents the Gaussian distribution of the training data, and μ(s) and ∑(s) represent the mean and diagonal covariance matrix of the sentence feature vector s respectively; represents the standard Gaussian distribution.

[0055] Furthermore, the expression of the improved image-text matching loss function in the i-th stage is defined as follows:

[0056]

[0057] Among them, represents the word-level image-text matching loss in the i-th stage, represents the sentence-level image-text matching loss in the i-th stage.

[0058] Furthermore, the calculation process of the word-level image-text matching loss function in the i-th stage is as follows: First, input the image output by the generator in the i-th stage into the image encoder to obtain the regional feature matrix in the i-th stage; then multiply the word feature matrix and the regional feature matrix in the i-th stage through matrix multiplication to obtain the similarity matrix in the i-th stage; then normalize the similarity matrix along each dimension to obtain the attention weight coefficient matrix in the i-th stage; then multiply the regional feature matrix and the attention weight coefficient matrix in the i-th stage through matrix multiplication to obtain the regional context feature matrix in the i-th stage; finally, input the regional context feature matrix and the word feature matrix in the i-th stage into the region-word matching score function to obtain the word-level image-text matching loss value in the i-th stage; the expression is as follows:

[0059] γ i = softmax(θ 1 softmax(W i T R i ), i = 1, 2, 3; (24)

[0060] C i = Ri γ i T , i = 1, 2, 3; (25)

[0061]

[0062] Among them, W i represents the word feature matrix of the i-th stage, with dimension D w ×N w ; R i represents the region feature matrix of the i-th stage, with dimension D w ×289; W i T R i represents the similarity matrix of the i-th stage, with dimension N w ×289; γ i represents the attention weight coefficient matrix of the i-th stage, with dimension N w ×289; C i represents the region context feature matrix of the i-th stage, with dimension D w ×N w ; represents the n-th region context feature vector in the region context feature matrix of the i-th stage, with dimension D w ; represents the n-th word feature vector in the word feature matrix of the i-th stage, with dimension D w ; represents the region-word matching score function of the i-th stage; represents M pairs of region-word pairs, only matched and the other M - 1 pairs are regarded as unmatched region-word pairs; θ 1 , θ 2 , θ 3 represent the first, second, and third hyperparameters respectively.

[0063] Furthermore, the calculation process of the sentence-level image-text matching loss function in the i-th stage is as follows: First, input the image output by the generator in the i-th stage into the image encoder to obtain the global image feature vector of the i-th stage. Then, input the sentence feature vector and the global image feature vector of the i-th stage into the image-sentence matching score function to obtain the sentence-level image-text matching loss value of the i-th stage. The expression is as follows:

[0064]

[0065] Among them, s i represents the sentence feature vector of the i-th stage, with dimension D s ; g iDenote the global image feature vector of the i-th stage, with dimension D s ; Denote the image-sentence matching score function of the i-th stage; Denote M pairs of image-sentence pairs, where only matched The other M - 1 pairs are regarded as unmatched image-sentence pairs; θ 4 Denote the fourth hyperparameter.

[0066] Furthermore, the adversarial loss function of the generator encourages the generated image to be as realistic as possible, the conditional enhancement loss function avoids overfitting during training, and the improvement of the image-text matching loss function lies in that the word-level image-text matching loss function calculates the word-level image-text matching loss value by using the word feature matrices of the first, second, and third stages, encouraging each word to match the concerned image sub-region as much as possible, and the sentence-level image-text matching loss function encourages the generated image to match the text description as much as possible. By minimizing the improved objective function, it is ensured that the generated image has richer details, is more realistic, and is more in line with the text description.

[0067] The present invention has the following advantages and effects compared with the prior art:

[0068] (1) By using the gated cross-word-visual attention unit in the process of generating an image from text, the present invention first obtains the semantic information amount of relevant words contained in the generated image, then selects important words, and further generates important fine-grained information on the image sub-region.

[0069] (2) By using the gated cross-word-visual attention unit in multiple stages and introducing an improved objective function, it is ensured that the selected important words are not lost, the fine-grained information of the generated image is enriched, and the matching degree between the generated image and the text description is enhanced.

[0070] (3) Experiments were carried out on the benchmark datasets CUB and MS-COCO for generating images from text. The results show that compared with other advanced methods, the method of the present invention has achieved leading performance in the image quality and diversity evaluation indexes IS, FID, and the image-text semantic consistency evaluation index R-precision, and is better than both SEGAN and DM-GAN. At the same time, the images generated by the method of the present invention have richer details, are closer to real images, and are more in line with the text description. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:

[0072] Figure 1 It is a flow chart of the method for generating images from text based on gated cross-word-visual attention disclosed in the present invention;

[0073] Figure 2 It is a flow chart of the method for generating images from text based on gated word-visual attention in the second embodiment;

[0074] Figure 3 It is a framework diagram of the gated cross-word-visual attention unit in the present invention;

[0075] Figure 4 It is a framework diagram of word-to-visual attention in the present invention;

[0076] Figure 5 It is a framework diagram of visual-to-word attention in the present invention;

[0077] Figure 6 It is a framework diagram of the gated word-visual attention unit in the second embodiment. Detailed implementation manners

[0078] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] Embodiment 1

[0080] As Figure 1 shown, this embodiment provides a method for generating images from text based on gated cross-word-visual attention. First, the semantic information amount of relevant words included in the generated image is obtained, then important words are selected, and then important fine-grained information is generated on the image sub-region. By using the gated cross-word-visual attention unit in multiple stages and introducing an improved objective function, it is ensured that the selected important words will not be lost, the fine-grained information of the generated image is enriched, and the matching between the generated image and the text description is enhanced. The specific steps are as follows:

[0081] S1. Extract the sentence feature vector and the word feature matrix of the first stage from the text description, and obtain the conditional feature vector by performing conditional enhancement processing on the sentence feature vector. Then, input the conditional feature vector and the random noise vector into the visual feature transformer of the first stage to obtain the visual feature matrix of the first stage. Then, input the visual feature matrix of the first stage into the generator of the first stage to obtain an image with a resolution of 64×64, specifically:

[0082] The visual feature transformer and generator in the first stage are as follows Figure 1 As shown, the visual feature transformer consists of 1 fully connected layer and 4 upsampling blocks connected in series, and the generator consists of 1 3×3 convolutional layer.

[0083] Define the process of extracting the sentence feature vector and the word feature matrix in the first stage as ENC(), and define the process of conditional enhancement processing to obtain the conditional feature vector as CA(); define the process of the visual feature transformer in the first stage outputting the visual feature matrix in the first stage as F 1 (), and define the process of the generator in the first stage outputting an image with a resolution of 64×64 as G 1 (), and the expressions are as follows:

[0084] (s, W 1 ) = ENC(Test); (1)

[0085] s ca = CA(s); (2)

[0086] V 1 = F 1 (concat(s ca , z)); (3)

[0087] I 1 = G 1 (V 1 ) (4)

[0088] Among them, Test represents the text description; s represents the sentence feature vector with a dimension of D s ; W 1 represents the word feature matrix in the first stage with a dimension of D w ×N w ; s ca represents the conditional feature vector with a dimension of D s ; z represents a random noise vector that follows a standard normal distribution with a dimension of D z ; V 1 represents the visual feature matrix in the first stage with a dimension of I 1 represents the image with a resolution of 64×64 output by the generator in the first stage; concat() represents the feature concatenation function that concatenates multiple input feature matrices into one output feature matrix.

[0089] S2. Input the word feature matrix and visual feature matrix of the first stage into the gated cross-modal word-visual attention unit of the first stage to obtain the refined word feature matrix and refined visual feature matrix of the first stage. Then, use the refined word feature matrix of the first stage as the word feature matrix of the second stage. Next, input the refined visual feature matrix of the first stage into the visual feature transformer of the second stage to obtain the visual feature matrix of the second stage. Finally, input the visual feature matrix of the second stage into the generator of the second stage to obtain an image with a resolution of 128×128. Specifically:

[0090] The gated cross-modal word-visual attention unit of the first stage is as Figure 3 shown, which is composed of a word-to-visual attention block, a selection gate, and a visual-to-word attention block connected in series. The visual feature transformer and generator of the second stage are as Figure 1 shown. The visual feature transformer is composed of 2 residual blocks and 1 upsampling block connected in series, and the generator is composed of 1 3×3 convolutional layer.

[0091] The word-to-visual attention block in the gated cross-modal word-visual attention unit of the first stage is as Figure 4 shown. Taking the visual feature matrix and word feature matrix of the first stage as inputs, the output is the local visual feature matrix of the first stage. The calculation process is as follows: First, perform feature mapping on the input visual feature matrix through a 1×1 convolutional layer to obtain a visual feature matrix in the word feature semantic space. Then, multiply the input word feature matrix and the visual feature matrix in the word feature semantic space through matrix multiplication to obtain a similarity matrix. Next, normalize the similarity matrix along the last dimension to obtain an attention weight coefficient matrix. Then, multiply the visual feature matrix in the word feature semantic space and the attention weight coefficient matrix through matrix multiplication to obtain a visual context feature matrix. Finally, perform feature concatenation on the visual context feature matrix and the input word feature matrix, and pass through two linear transformation layers and a sigmoid activation function to obtain the local visual feature matrix. The expression is as follows:

[0092] V′ 1 =M v (V 1 ); (5)

[0093] α 1 =softmax(W 1 T V′ 1 ); (6)

[0094]

[0095] where V 1 represents the input visual feature matrix of the first stage, and the dimension is W 1 represents the word feature matrix of the first stage of the input, with dimension D w ×N w ; V′ 1 represents the visual feature matrix in the word feature semantic space at the first stage, with dimension W 1 T V′ 1 represents the similarity matrix at the first stage, with dimension α 1 represents the attention weight coefficient matrix at the first stage, with dimension V′ 1 α 1 T represents the visual context feature matrix at the first stage, with dimension D w ×N w ; represents the local visual feature matrix of the first stage of the output, with dimension D w ×N w ; M v () represents a 1×1 convolutional layer, and the subscript v at the lower right indicates that the input feature is in the visual feature semantic space; and represent the first and second linear transformation layers, and the subscript w at the lower right indicates that the input feature is in the word feature semantic space, with dimension D w ×D w , with dimension D w ; σ() represents the sigmoid activation function, represents element-wise multiplication, and T represents matrix transpose.

[0096] The selection gate in the first-stage gated cross word-visual attention unit is as Figure 3 shown, taking the local visual feature matrix and word feature matrix of the first stage as inputs, and the output is the refined word feature matrix of the first stage. The calculation process is as follows: Pass the input local visual feature matrix and word feature matrix through two linear transformation layers and the sigmoid activation function to obtain the refined word feature matrix; the expression is as follows:

[0097]

[0098] where, represents the local visual feature matrix of the first stage of the input, with dimension D w ×N w ; W 1 represents the word feature matrix of the first stage of the input, with dimension D w ×Nw ; represents the refined word feature matrix at the first stage of the output, with a dimension of D w ×N w ; and represent the first and second linear transformation layers. The subscript w in the lower right indicates that the input feature is in the word feature semantic space, with a dimension of 1×D w ; σ() represents the sigmoid activation function.

[0099] The visual-to-word attention block in the first-stage gated cross word-visual attention unit is as Figure 5 shown. Taking the refined word feature matrix and visual feature matrix at the first stage as inputs, the output is the refined visual feature matrix at the first stage. The calculation process is as follows: First, map the input refined word feature matrix through a 1×1 convolutional layer to obtain the word feature matrix in the visual feature semantic space; then multiply the word feature matrix in the visual feature semantic space and the input visual feature matrix through matrix multiplication to obtain the similarity matrix; then normalize the similarity matrix along the last dimension to obtain the attention weight coefficient matrix; then multiply the word feature matrix in the visual feature semantic space and the attention weight coefficient matrix through matrix multiplication to obtain the word context feature matrix; finally, concatenate the word context feature matrix and the input visual feature matrix and pass through two linear transformation layers and the sigmoid activation function to obtain the refined visual feature matrix; the expression is as follows:

[0100]

[0101] where, represents the input refined word feature matrix at the first stage, with a dimension of D w ×N w ; V 1 represents the input visual feature matrix at the first stage, with a dimension of represents the word feature matrix in the visual feature semantic space at the first stage, with a dimension of D v ×N w ; represents the similarity matrix at the first stage, with a dimension of β 1 represents the attention weight coefficient matrix at the first stage, with a dimension of represents the word context feature matrix at the first stage, with a dimension of represents the output refined visual feature matrix at the first stage, with a dimension of M w () represents the 1×1 convolutional layer. The subscript w in the lower right indicates that the input feature is in the word feature semantic space; and represent the first and second linear transformation layers, and the subscript v indicates that the input features are in the visual feature semantic space. has a dimension of D v ×D v , has a dimension of D v ; σ() represents the sigmoid activation function, represents element-wise multiplication, and T represents matrix transpose.

[0102] Define the process of the gated cross-word-visual attention unit in the first stage to output the refined word feature matrix and refined visual feature matrix in the first stage using formulas (5)-(11) as GCAU 1 (), define the process of the visual feature transformer in the second stage to output the visual feature matrix in the second stage as F 2 (), define the process of the generator in the second stage to output an image with a resolution of 128×128 as G 2 (), and the expressions are as follows:

[0103]

[0104] I 2 = G 2 (V 2 ) (15)

[0105] where, represents the refined word feature matrix in the first stage, with a dimension of D w ×N w ; represents the refined visual feature matrix in the first stage, with a dimension of W 2 represents the word feature matrix in the second stage, with a dimension of D w ×N w ; V 2 represents the visual feature matrix in the second stage, with a dimension of I 2 represents the image with a resolution of 128×128 output by the generator in the second stage; concat() represents the feature concatenation function, which concatenates multiple input feature matrices into one output feature matrix.

[0106] S3. Input the word feature matrix and visual feature matrix of the second stage into the gated cross-modal word-visual attention unit of the second stage to obtain the refined word feature matrix and refined visual feature matrix of the second stage. Take the refined word feature matrix of the second stage as the word feature matrix of the third stage. Then, input the refined visual feature matrix of the second stage into the visual feature transformer of the third stage to obtain the visual feature matrix of the third stage. Next, input the visual feature matrix of the third stage into the generator of the third stage to obtain an image with a resolution of 256×256. Specifically:

[0107] The gated cross-modal word-visual attention unit of the second stage is as Figure 3 shown, which is composed of a word-to-visual attention block, a selection gate, and a visual-to-word attention block connected in series. The visual feature transformer and generator of the third stage are as Figure 1 shown. The visual feature transformer is composed of 2 residual blocks and 1 upsampling block connected in series, and the generator is composed of 1 3×3 convolutional layer.

[0108] The word-to-visual attention block in the gated cross-modal word-visual attention unit of the second stage is as Figure 4 shown. Taking the visual feature matrix and word feature matrix of the second stage as inputs, the output is the local visual feature matrix of the second stage. Its calculation process refers to the calculation process of the word-to-visual attention block in the gated cross-modal word-visual attention unit of the first stage in step S2. The expression is as follows:

[0109] V′ 2 =M v (V 2 ); (16)

[0110] α 2 =softmax(W 2 TV′ 2 ); (17)

[0111]

[0112] where V 2 represents the input visual feature matrix of the second stage, with a dimension of W 2 represents the input word feature matrix of the second stage, with a dimension of D w ×N w ; V′ 2 represents the visual feature matrix in the word feature semantic space of the second stage, with a dimension of W 2 T V′ 2 represents the similarity matrix of the second stage, with a dimension of α 2Denote the attention weight coefficient matrix in the second stage, with dimension V′ 2 α 2 T Denote the visual context feature matrix in the second stage, with dimension D w ×N w ; Denote the local visual feature matrix of the output second stage, with dimension D w ×N w ; M v () represents a 1×1 convolutional layer, and the subscript v at the lower right indicates that the input feature is in the visual feature semantic space; and Denote the first and second linear transformation layers, and the subscript w at the lower right indicates that the input feature is in the word feature semantic space, with dimension D w ×D w , with dimension D w ; σ() represents the sigmoid activation function, Denote element-wise multiplication, and T represents matrix transpose.

[0113] The selection gate in the gated cross-word-visual attention unit of the second stage is as Figure 3 shown, taking the local visual feature matrix and word feature matrix of the second stage as inputs, and the output is the refined word feature matrix of the second stage. Its calculation process refers to the calculation process of the selection gate in the gated cross-word-visual attention unit of the first stage in step S2, and the expression is as follows:

[0114]

[0115] where, Denote the local visual feature matrix of the input second stage, with dimension D w ×N w ; W 2 Denote the word feature matrix of the input second stage, with dimension D w ×N w ; Denote the refined word feature matrix of the output second stage, with dimension D w ×N w ; and Denote the first and second linear transformation layers, and the subscript w at the lower right indicates that the input feature is in the word feature semantic space, with dimension 1×D w ; σ() represents the sigmoid activation function.

[0116] The visual-to-word attention block in the gated cross-word-visual attention unit of the second stage is as follows Figure 5 shown. Taking the refined word feature matrix and visual feature matrix of the second stage as inputs, the output is the refined visual feature matrix of the second stage. The calculation process refers to the calculation process of the visual-to-word attention block in the gated cross-word-visual attention unit of the first stage in step S2. The expression is as follows:

[0117]

[0118] Where represents the input refined word feature matrix of the second stage, with a dimension of D w ×N w ; V 2 represents the input visual feature matrix of the second stage, with a dimension of represents the word feature matrix in the visual feature semantic space of the second stage, with a dimension of D v ×N w ; represents the similarity matrix of the second stage, with a dimension of β 2 represents the attention weight coefficient matrix of the second stage, with a dimension of represents the word context feature matrix of the second stage, with a dimension of represents the output refined visual feature matrix of the second stage, with a dimension of M w () represents a 1×1 convolutional layer, and the lower right subscript w indicates that the input feature is in the word feature semantic space; and represent the first and second linear transformation layers, and the lower right subscript v indicates that the input feature is in the visual feature semantic space. The dimension of is D v ×D v , has a dimension of D v ; σ() represents the sigmoid activation function, represents element-wise multiplication, and T represents matrix transpose.

[0119] Define the process of the gated cross-word-visual attention unit of the second stage using formulas (16)-(22) to output the refined word feature matrix and refined visual feature matrix as GCAU 2 (), define the process of the visual feature transformer of the third stage to output the visual feature matrix of the third stage as F 3 (), define the process of the generator of the third stage to output an image with a resolution of 256×256 as G 3 (), and the expression is as follows:

[0120]

[0121] I 3 = G 3 (V 3 ) (26)

[0122] Among them, represents the word feature matrix refined in the second stage, with a dimension of D w × N w ; represents the visual feature matrix refined in the second stage, with a dimension of W 3 represents the word feature matrix in the third stage, with a dimension of D w × N w ; V 3 represents the visual feature matrix in the third stage, with a dimension of I 3 represents the image with a resolution of 256 × 256 output by the generator in the third stage, and also represents the finally generated high-quality image; concat() represents the feature concatenation function, which concatenates multiple input feature matrices into one output feature matrix.

[0123] S4. Introduce an improved objective function, and enhance the authenticity of the images generated in each stage and the semantic consistency between the generated images and the text descriptions by minimizing the objective function, and use the image with a resolution of 256 × 256 generated in the third stage as the finally generated high-quality image. Specifically:

[0124] Define the improved objective function as L, and the expression is as follows:

[0125]

[0126] Among them, represents the adversarial loss function of the generator in the i-th stage, L CA represents the conditional enhancement loss function; represents the improved image-text matching loss function in the i-th stage; λ 1 , λ 2 represents the weight coefficient.

[0127] Define the expression of the adversarial loss function of the generator in the i-th stage as follows:

[0128]

[0129] Among them, D i () represents a probability value between 0 and 1, I i represents the image output by the generator in the i-th stage, I i comes from the data distribution generated by the generator in the i-th stage s represents the sentence feature vector.

[0130] Define the conditional enhancement loss function as the KL divergence between the standard Gaussian distribution and the Gaussian distribution of the training data, and the expression is as follows:

[0131]

[0132] Among them, represents the Gaussian distribution of the training data, and μ(s) and ∑(s) represent the mean and diagonal covariance matrix of the sentence feature vector s respectively; represents the standard Gaussian distribution.

[0133] Define the improved image-text matching loss function expression in the i-th stage as follows:

[0134]

[0135] Among them, represents the word-level image-text matching loss in the i-th stage, represents the sentence-level image-text matching loss in the i-th stage.

[0136] The calculation process of the word-level image-text matching loss function in the i-th stage is as follows: First, input the image output by the generator in the i-th stage into the image encoder to obtain the regional feature matrix in the i-th stage; then multiply the word feature matrix and the regional feature matrix in the i-th stage through matrix multiplication to obtain the similarity matrix in the i-th stage; then normalize the similarity matrix along each dimension to obtain the attention weight coefficient matrix in the i-th stage; then multiply the regional feature matrix and the attention weight coefficient matrix in the i-th stage through matrix multiplication to obtain the regional context feature matrix in the i-th stage; finally, input the regional context feature matrix and the word feature matrix in the i-th stage into the region-word matching score function to obtain the word-level image-text matching loss value in the i-th stage; the expression is as follows:

[0137] γ i = softmax(θ 1 softmax(W i T R i ), i = 1, 2, 3; (31)

[0138] C i = R i γ i T , i = 1, 2, 3; (32)

[0139]

[0140] Among them, W iDenote the word feature matrix of the $i$-th stage, with dimension $D$ w ×N w ; R i Denote the region feature matrix of the $i$-th stage, with dimension $D$ w ×289; W i T R i Denote the similarity matrix of the $i$-th stage, with dimension $N$ w ×289; γ i Denote the attention weight coefficient matrix of the $i$-th stage, with dimension $N$ w ×289; C i Denote the region context feature matrix of the $i$-th stage, with dimension $D$ w ×N w ; Denote the $n$-th region context feature vector in the region context feature matrix of the $i$-th stage, with dimension $D$ w ; Denote the $n$-th word feature vector in the word feature matrix of the $i$-th stage, with dimension $D$ w ; Denote the region-word matching score function of the $i$-th stage; Denote $M$ pairs of region-word pairs, only matching The other $M - 1$ pairs are regarded as unmatched region-word pairs; θ 1 , θ 2 , θ 3 Denote the first, second, and third hyperparameters respectively.

[0141] The calculation process of the sentence-level image-text matching loss function in the $i$-th stage is as follows: First, input the image output by the generator in the $i$-th stage into the image encoder to obtain the global image feature vector of the $i$-th stage. Then, input the sentence feature vector and the global image feature vector of the $i$-th stage into the image-sentence matching score function to obtain the sentence-level image-text matching loss value of the $i$-th stage. The expression is as follows:

[0142]

[0143] where, s i Denote the sentence feature vector of the $i$-th stage, with dimension $D$ s ; g i Denote the global image feature vector of the $i$-th stage, with dimension $D$ s ; Denote the image-sentence matching score function of the $i$-th stage; Denote $M$ pairs of image-sentence pairs, only matching The other $M - 1$ pairs are regarded as unmatched image-sentence pairs; θ4 Represents the fourth hyperparameter.

[0144] In this embodiment, a pre-trained bidirectional long short-term memory network is used as a text encoder to extract sentence feature vectors and a word feature matrix in the first stage from the text description, and the conditional enhancement processing technology of Stack GAN is used to obtain conditional feature vectors; take D z = 100, D s = D w = 256, D v = 64, N w = 18, θ 1 = θ 2 = 5, θ 3 = θ 4 = θ 1 = 10, M = 50, for λ 2 , λ 1 = 1, λ 2 = 5, on the benchmark dataset CUB, take λ 1 = 1, λ 2 = 50, however, the values of all the above parameters do not constitute a limitation to the technical solution of the invention.

[0145] This embodiment is used to generate text-to-image for the benchmark datasets CUB and MS-COCO respectively, and the results are compared with other advanced methods. This embodiment has achieved superior performance on all datasets. The IS index on the benchmark dataset CUB reaches 4.80, the FID index reaches 14.85, and the R-precision index reaches 78.48; on the benchmark dataset MS-COCO, the IS index reaches 31.19, the FID index reaches 24.39, and the R-precision index reaches 90.92. Further, a comparison is made with the methods SEGAN and DM-GAN that are closest to this embodiment, and this embodiment has achieved better results. First, the IS index (high-priority index) of this embodiment on the benchmark dataset CUB is 2.78% higher than that of SEGAN and 1.05% higher than that of DM-GAN. On the benchmark dataset MS-COCO, the IS index is 11.95% higher than that of SEGAN and 2.30% higher than that of DM-GAN. Second, the FID index (low-priority index) of this embodiment on the benchmark dataset CUB is 18.27% lower than that of SEGAN and 7.70% lower than that of DM-GAN. On the MS-COCO dataset, the FID index is 24.44% lower than that of SEGAN and 25.27% lower than that of DM-GAN.

[0146] Embodiment 2

[0147] As Figure 2As shown in the figure, this embodiment provides a method for generating images based on gated word-visual attention drive. Referring to Embodiment 1, there are three differences in the technical means of this embodiment. The first is that the word-to-visual attention block in the gated word-visual attention units of the first and second stages in this embodiment is removed; the second is that the gated word-visual attention unit in the second stage in this embodiment takes the word feature matrix of the first stage and the visual feature matrix of the second stage as inputs; the third is that the calculation process of the word-level image-text matching loss function in the image-text matching loss function is different; specifically, it includes the following steps:

[0148] S1. Refer to step S1 in Embodiment 1;

[0149] S2. Input the word feature matrix and visual feature matrix of the first stage into the gated word-visual attention unit of the first stage to obtain the refined visual feature matrix of the first stage. Then input the refined visual feature matrix of the first stage into the visual feature transformer of the second stage to obtain the visual feature matrix of the second stage. Then input the visual feature matrix of the second stage into the generator of the second stage to obtain an image with a resolution of 128×128. Specifically:

[0150] The gated word-visual attention unit of the first stage is as Figure 6 shown, which is composed of a selection gate and a visual-to-word attention block connected in series; the visual feature transformer and generator of the second stage are as Figure 2 shown. The visual feature transformer is composed of 2 residual blocks and 1 upsampling block connected in series, and the generator is composed of 1 3×3 convolutional layer.

[0151] The selection gate in the gated word-visual attention unit of the first stage is as Figure 6 shown, taking the word feature matrix and visual feature matrix of the first stage as inputs, and the output is the refined word feature matrix of the first stage. Its calculation process is as follows: First, average-pool the input visual feature matrix to obtain the global visual feature matrix. Then, pass the global visual feature matrix and the input word feature matrix through two linear transformation layers and a sigmoid activation function to obtain the refined word feature matrix; the expression is as follows:

[0152]

[0153] Among them, represents the nth visual feature vector in the input visual feature matrix of the first stage, with a dimension of D v ; represents the global visual feature matrix of the first stage, with a dimension of D v ×N w ; W 1 represents the input word feature matrix of the first stage, with a dimension of Dw ×N w ; represents the first-stage refined word feature matrix with a dimension of D w ×N w ; M v () represents a 1×1 convolutional layer, and the subscript v at the lower right indicates that the input feature is in the visual feature semantic space; and represent the first and second linear transformation layers. The subscript w at the lower right indicates that the input feature is in the word feature semantic space, and the subscript v at the lower right indicates that the input feature is in the visual feature semantic space. has a dimension of 1×D w , has a dimension of 1×D v ; repeat() represents matrix column replication, and σ() represents the sigmoid activation function.

[0154] The visual-to-word attention block in the first-stage gated word-visual attention unit refers to the visual-to-word attention block in the visual-to-word attention unit of the first-stage gated cross-word-visual attention unit in Example 1.

[0155] Define the process of the first-stage gated word-visual attention unit using formulas (37), (38), (9)-(11) to output the first-stage refined visual feature matrix as GAU 1 (), define the process of the second-stage visual feature transformer to output the second-stage visual feature matrix as F 2 (), define the process of the second-stage generator to output an image with a resolution of 128×128 as G 2 (), and the expression is as follows:

[0156]

[0157] I 2 = G 2 (V 2 ) (41)

[0158] where represents the first-stage refined visual feature matrix with a dimension of V 2 represents the second-stage visual feature matrix with a dimension of I 2 represents the image with a resolution of 128×128 output by the second-stage generator; concat() represents the feature splicing function that splices multiple input feature matrices into one output feature matrix.

[0159] S3. Input the word feature matrix in the first stage and the visual feature matrix in the second stage into the gated word-visual attention unit in the second stage to obtain the refined visual feature matrix in the second stage. Then, input the refined visual feature matrix in the second stage into the visual feature transformer in the third stage to obtain the visual feature matrix in the third stage. Next, input the visual feature matrix in the third stage into the generator in the third stage to obtain an image with a resolution of 256×256. Specifically:

[0160] The gated word-visual attention unit in the second stage is as shown in Figure 6 and is composed of a selection gate and a visual-to-word attention block connected in series; the visual feature transformer and the generator in the third stage are as shown in Figure 2 . The visual feature transformer is composed of 2 residual blocks and 1 upsampling block connected in series, and the generator is composed of 1 3×3 convolutional layer.

[0161] The selection gate in the gated word-visual attention unit in the second stage is as shown in Figure 6 . Taking the word feature matrix in the first stage and the visual feature matrix in the second stage as inputs, the output is the refined word feature matrix in the second stage. Its calculation process refers to the calculation process of the selection gate in the gated word-visual attention unit in the first stage of step S2, and the expression is as follows:

[0162]

[0163] where represents the nth visual feature vector in the input visual feature matrix in the second stage, with a dimension of D v ; represents the global visual feature matrix in the second stage, with a dimension of D v ×N w ; W 1 represents the input word feature matrix in the first stage, with a dimension of D w ×N w ; represents the output refined word feature matrix in the second stage, with a dimension of D w ×N w ; M v () represents a 1×1 convolutional layer, and the lower right subscript v indicates that the input feature is in the visual feature semantic space; and represent the first and second linear transformation layers. The lower right subscript w indicates that the input feature is in the word feature semantic space, and the lower right subscript v indicates that the input feature is in the visual feature semantic space. The dimension of is 1×D w , and the dimension of is 1×D v; repeat() represents matrix column replication, and σ() represents the sigmoid activation function.

[0164] The visual-to-word attention block in the second-stage gated word-visual attention unit refers to the visual-to-word attention block in the visual-to-word attention block of the second-stage gated cross-word-visual attention unit in step S3 of the first embodiment.

[0165] Define the process of the second-stage gated word-visual attention unit using formulas (42), (43), (20)-(22) to output the second-stage refined visual feature matrix as GAU 2 (), define the process of the third-stage visual feature transformer to output the third-stage visual feature matrix as F 3 (), define the process of the third-stage generator to output an image with a resolution of 256×256 as G 3 (), and the expression is as follows:

[0166]

[0167] I 3 = G 3 (V 3 ) (46)

[0168] Among them, represents the second-stage refined visual feature matrix, with a dimension of V 3 represents the third-stage visual feature matrix, with a dimension of I 3 represents the image with a resolution of 256×256 output by the third-stage generator; concat() represents the feature concatenation function, which concatenates multiple input feature matrices into one output feature matrix.

[0169] S4. Refer to step S4 in the first embodiment. The only difference is the calculation process of the word-level image-text matching loss function in the i-th stage: First, input the image output by the i-th stage generator into the image encoder to obtain the i-th stage regional feature matrix; then multiply the first-stage word feature matrix and the i-th stage regional feature matrix through matrix multiplication to obtain the i-th stage similarity matrix; then normalize the similarity matrix along each dimension to obtain the i-th stage attention weight coefficient matrix; then multiply the i-th stage regional feature matrix and the attention weight coefficient matrix through matrix multiplication to obtain the i-th stage regional context feature matrix; finally, input the i-th stage regional context feature matrix and the first-stage word feature matrix into the region-word matching score function to obtain the i-th stage word-level image-text matching loss value; the expression is as follows:

[0170] γ i= softmax(θ 1 softmax(W 1 T R i ), i = 1, 2, 3; (47)

[0171] C i = R i γ i T , i = 1, 2, 3; (48)

[0172]

[0173] Among them, W 1 represents the word feature matrix in the first stage, with dimension D w × N w ; R i represents the region feature matrix in the i-th stage, with dimension D w × 289; W 1 T R i represents the similarity matrix in the i-th stage, with dimension N w × 289; γ i represents the attention weight coefficient matrix in the i-th stage, with dimension N w × 289; C i represents the region context feature matrix in the i-th stage, with dimension D w × N w ; represents the n-th region context feature vector in the region context feature matrix in the i-th stage, with dimension D w ; represents the n-th word feature vector in the word feature matrix in the first stage, with dimension D w ; represents the region-word matching score function in the i-th stage; represents M pairs of region-word pairs, only matched The other M - 1 pairs are regarded as unmatched region-word pairs; θ 1 , θ 2 , θ 3 represent the first, second, and third hyperparameters respectively.

[0174] The text-to-image generation was performed on the benchmark datasets CUB and MS-COCO using this embodiment, and the results were compared with those of Embodiment 1. The IS metric, FID metric, and R-precision metric of this embodiment on the benchmark dataset CUB were all worse than those of Embodiment 1. Among them, the IS metric (a high-preference metric) reached 4.72, a decrease of 1.67%, the FID metric (a low-preference metric) reached 15.95, an increase of 6.9%, and the R-precision metric (a high-preference metric) reached 77.21, a decrease of 1.62%. The results demonstrate the effectiveness of the method of the present invention.

[0175] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A method for generating images from text driven by gated cross-word-visual attention, characterized in that, the method for generating images from text includes the following steps: S1. Extract the sentence feature vector and the word feature matrix in the first stage from the text description, and obtain the conditional feature vector by conditional enhancement processing of the sentence feature vector. Then, input the conditional feature vector and the random noise vector into the visual feature transformer in the first stage to obtain the visual feature matrix in the first stage. Next, input the visual feature matrix in the first stage into the generator in the first stage to obtain the first-resolution image; S2. Input the word feature matrix and the visual feature matrix in the first stage into the gated cross-word-visual attention unit in the first stage to obtain the refined word feature matrix and the refined visual feature matrix in the first stage. And use the refined word feature matrix in the first stage as the word feature matrix in the second stage. Then, input the refined visual feature matrix in the first stage into the visual feature transformer in the second stage to obtain the visual feature matrix in the second stage. Next, input the visual feature matrix in the second stage into the generator in the second stage to obtain the second-resolution image; S3. Input the word feature matrix and the visual feature matrix in the second stage into the gated cross-word-visual attention unit in the second stage to obtain the refined word feature matrix and the refined visual feature matrix in the second stage. And use the refined word feature matrix in the second stage as the word feature matrix in the third stage. Then, input the refined visual feature matrix in the second stage into the visual feature transformer in the third stage to obtain the visual feature matrix in the third stage. Next, input the visual feature matrix in the third stage into the generator in the third stage to obtain the third-resolution image; S4. Introduce an improved objective function, enhance the authenticity of the generated image in each stage and the semantic consistency between the generated image and the text description by minimizing the objective function, and use the third-resolution image generated in the third stage as the finally generated high-quality image.

2. The method for generating images from text driven by gated cross-word-visual attention according to claim 1, characterized in that, The word feature matrices of the first, second, and third stages are each composed of multiple word feature vectors. Use N w to represent the number of word feature vectors in the word feature matrices of the first, second, and third stages, and D w to represent the dimension of the word feature vectors in the first, second, and third stages; the visual feature matrices of the first, second, and third stages are each composed of multiple visual feature vectors. Use to represent the number of visual feature vectors in the visual feature matrices of the first, second, and third stages respectively, and D v to represent the dimension of the visual feature vectors in the first, second, and third stages.

3. The method for generating images from text driven by gated cross-word-visual attention according to claim 2, characterized in that, the gated cross-word-visual attention units in the first and second stages are each composed of a word-to-visual attention block, a selection gate, and a visual-to-word attention block connected in series; the visual feature transformer in the first stage is composed of 1 fully connected layer and 4 upsampling blocks connected in series, and the visual feature transformers in the second and third stages are each composed of 2 residual blocks and 1 upsampling block connected in series; the generators in the first, second, and third stages are each composed of 1 3×3 convolutional layer.

4. The method for generating images from text driven by gated cross-word-visual attention according to claim 3, characterized in that, the word-to-visual attention block in the gated cross-word-visual attention unit in the first stage takes the visual feature matrix and the word feature matrix in the first stage as inputs, and the output is the local visual feature matrix in the first stage; The word-to-visual attention block in the second-stage gated cross-word-visual attention unit takes the visual feature matrix and word feature matrix of the second stage as inputs, and the output is the local visual feature matrix of the second stage; The calculation process of the word-to-visual attention block is as follows: First, the input visual feature matrix is feature-mapped through a 1×1 convolutional layer to obtain a visual feature matrix in the word feature semantic space; then the input word feature matrix and the visual feature matrix in the word feature semantic space are multiplied through matrix multiplication to obtain a similarity matrix; then the similarity matrix is normalized along the last dimension to obtain an attention weight coefficient matrix; then the visual feature matrix in the word feature semantic space and the attention weight coefficient matrix are multiplied through matrix multiplication to obtain a visual context feature matrix; finally, the visual context feature matrix and the input word feature matrix are feature-concatenated, and through two linear transformation layers and a sigmoid activation function, a local visual feature matrix is obtained; Expression is as follows: V i ′ = M v (V i ), i = 1, 2; (1) Among them, V i represents the visual feature matrix of the i-th stage of the input, with the dimension of W i represents the word feature matrix of the i-th stage of the input, with the dimension of D w ×N w ; V i ' represents the visual feature matrix in the word feature semantic space of the i-th stage, with the dimension of W i T V i ' represents the similarity matrix of the i-th stage, with the dimension of ɑ i represents the attention weight coefficient matrix of the i-th stage, with the dimension of V i 'ɑ i T represents the visual context feature matrix of the i-th stage, with the dimension of D w ×N w ; represents the local visual feature matrix of the i-th stage of the output, with the dimension of D w ×N w ; M v () represents a 1×1 convolutional layer, and the subscript v at the lower right indicates that the input feature is in the visual feature semantic space; and represent the first and second linear transformation layers, and the subscript w at the lower right indicates that the input feature is in the word feature semantic space, whose dimension is D w ×D w , whose dimension is D w ; σ() represents the sigmoid activation function, represents element-wise multiplication, and the superscript T at the upper right indicates matrix transpose.

5. The method for generating an image from text driven by gated cross-word-visual attention according to claim 4, characterized in that, The selection gate in the first-stage gated cross-word-visual attention unit takes the local visual feature matrix and word feature matrix of the first stage as inputs, and the output is the refined word feature matrix of the first stage; the selection gate in the second-stage gated cross-word-visual attention unit takes the local visual feature matrix and word feature matrix of the second stage as inputs, and the output is the refined word feature matrix of the second stage; The calculation process of the selection gate is: passing the input local visual feature matrix and word feature matrix through two linear transformation layers and a sigmoid activation function to obtain a refined word feature matrix; the expression is as follows: Among them, represents the local visual feature matrix of the i-th stage of the input, with a dimension of D w ×N w ; W i represents the word feature matrix of the i-th stage of the input, with a dimension of D w ×N w ; W i r represents the refined word feature matrix of the i-th stage of the output, with a dimension of D w ×N w ; and represent the first and second linear transformation layers. The subscript w in the lower right indicates that the input feature is in the word feature semantic space, with a dimension of 1×D w ; σ() represents the sigmoid activation function.

6. The method for generating an image from text driven by gated cross-word-visual attention according to claim 5, characterized in that, The visual-to-word attention block in the first-stage gated cross-word-visual attention unit takes the refined word feature matrix and visual feature matrix of the first stage as inputs, and the output is the refined visual feature matrix of the first stage; The visual-to-word attention block in the second-stage gated cross-word-visual attention unit takes the refined word feature matrix and visual feature matrix of the second stage as inputs, and the output is the refined visual feature matrix of the second stage; The calculation process of the visual-to-word attention block is as follows: First, the input refined word feature matrix is subjected to feature mapping through a 1×1 convolutional layer to obtain a word feature matrix in the visual feature semantic space; then, the word feature matrix in the visual feature semantic space and the input visual feature matrix are multiplied through matrix multiplication to obtain a similarity matrix; then, the similarity matrix is normalized along the last dimension to obtain an attention weight coefficient matrix; then, the word feature matrix in the visual feature semantic space and the attention weight coefficient matrix are multiplied through matrix multiplication to obtain a word context feature matrix; finally, the word context feature matrix and the input visual feature matrix are subjected to feature concatenation and passed through two linear transformation layers and a sigmoid activation function to obtain a refined visual feature matrix; the expression is as follows: Among them, represents the word feature matrix refined in the i-th stage of the input, with a dimension of D w ×N w ; V i represents the visual feature matrix in the i-th stage of the input, with a dimension of represents the word feature matrix in the visual feature semantic space in the i-th stage, with a dimension of D v ×N w ; represents the similarity matrix in the i-th stage, with a dimension of β i represents the attention weight coefficient matrix in the i-th stage, with a dimension of represents the word context feature matrix in the i-th stage, with a dimension of represents the refined visual feature matrix in the i-th stage of the output, with a dimension of M w () represents a 1×1 convolutional layer, and the subscript w in the lower right indicates that the input feature is in the word feature semantic space; and represent the first and second linear transformation layers, and the subscript v in the lower right indicates that the input feature is in the visual feature semantic space, with a dimension of D v ×D v , with a dimension of D v ; σ() represents the sigmoid activation function, represents element-wise multiplication, and the superscript T indicates matrix inversion.

7. The method for generating an image based on gated cross-word-visual attention driving according to claim 1, wherein, the process of step S1 is as follows: Define the process of extracting the sentence feature vector and the word feature matrix in the first stage as ENC(), and define the process of obtaining the conditional feature vector through conditional enhancement processing as CA(); define the process in which the visual feature transformer in the first stage outputs the visual feature matrix in the first stage as F 1 (), and define the process in which the generator in the first stage outputs the image at the first resolution as G 1 (), and the expression is as follows: (s,W 1 ) = ENC(Text); (8) s ca = CA(s); (9) V 1 = F 1 (concat(s ca , z)); (10) I 1 = G 1 (V 1 ) (11) Among them, Text represents the text description; s represents the sentence feature vector with a dimension of D s ; W 1 represents the word feature matrix in the first stage with a dimension of D w ×N w ; s ca represents the conditional feature vector with a dimension of D s ; z represents a random noise vector that follows a standard normal distribution with a dimension of D z ; V 1 represents the visual feature matrix in the first stage with a dimension of I 1 represents the first-resolution image output by the generator in the first stage; concat() represents the feature concatenation function that concatenates multiple input feature matrices into one output feature matrix.

8. The method for generating an image based on gated cross-word-visual attention driving according to claim 6, wherein, the process of step S2 is as follows: The process of using formulas (1)-(7) to define the gated cross-word-visual attention unit in the first stage to output the refined word feature matrix and the refined visual feature matrix in the first stage is GCAI 1 (), the process of defining the visual feature transformer in the second stage to output the visual feature matrix in the second stage is F 2 (), the process of defining the generator in the second stage to output the second-resolution image is G 2 (), and the expression is as follows: I 2 = G 2 (V 2 ) (15) Among them, represents the word feature matrix refined in the first stage, with a dimension of D w ×N w ; represents the visual feature matrix refined in the first stage, with a dimension of W 2 represents the word feature matrix in the second stage, with a dimension of D w ×N w ; V 2 represents the visual feature matrix in the second stage, with a dimension of I 2 represents the second-resolution image output by the generator in the second stage; concat() represents the feature concatenation function that concatenates multiple input feature matrices into one output feature matrix.

9. The method for generating an image based on gated cross-word-visual attention driving according to claim 6, wherein, the process of step S3 is as follows: The process of using formulas (1)-(7) to output the refined word feature matrix and the refined visual feature matrix of the second stage by the gated cross-word-visual attention unit in the second stage is defined as GCAU 2 (), the process of outputting the visual feature matrix of the third stage by the visual feature transformer in the third stage is defined as F 3 (), the process of outputting the image of the third resolution by the generator in the third stage is defined as G 3 (), the expression is as follows: I 3 = G 3 (V 3 ) (19) Among them, represents the word feature matrix refined in the second stage, with a dimension of D w ×N w ; represents the visual feature matrix refined in the second stage, with a dimension of W 3 represents the word feature matrix in the third stage, with a dimension of D w ×N w ; V 3 represents the visual feature matrix in the third stage, with a dimension of I 3 represents the third-resolution image output by the generator in the third stage, and also represents the finally generated high-quality image; concat() represents the feature concatenation function that concatenates multiple input feature matrices into one output feature matrix.

10. The method for generating an image based on gated cross-word-visual attention driving according to claim 1, wherein, the process of step S4 is as follows: Define the improved objective function as L, and the expression is as follows: Among them, represents the adversarial loss function of the i-th stage generator, L CA represents the conditional enhancement loss function; represents the improved image-text matching loss function of the i-th stage; λ 1 , λ 2 represents the weight coefficient; Define the adversarial loss function expression of the generator in the i-th stage as follows: Among them, D i () represents a probability value between 0 and 1, I i represents the image output by the generator in the i-th stage, I i from the data distribution generated by the generator in the i-th stage s represents the sentence feature vector; Define the conditional enhancement loss function as the KL divergence between the standard Gaussian distribution and the Gaussian distribution of the training data, and the expression is as follows: Among them, represents the Gaussian distribution of the training data, where μ(s) and ∑(s) represent the mean and diagonal covariance matrix of the sentence feature vector s respectively; represents the standard Gaussian distribution; Define the improved image-text matching loss function expression in the i-th stage as follows: Among them, represents the word-level image-text matching loss in the i-th stage, represents the sentence-level image-text matching loss in the i-th stage; The calculation process of the word-level image-text matching loss function in the i-th stage is as follows: First, the image output by the generator in the i-th stage is input into the image encoder to obtain the regional feature matrix in the i-th stage; then, the word feature matrix in the i-th stage and the regional feature matrix are multiplied through matrix multiplication to obtain the similarity matrix in the i-th stage; then, the similarity matrix is normalized along each dimension to obtain the attention weight coefficient matrix in the i-th stage; then, the regional feature matrix in the i-th stage and the attention weight coefficient matrix are multiplied through matrix multiplication to obtain the regional context feature matrix in the i-th stage; finally, the regional context feature matrix in the i-th stage and the word feature matrix are input into the region-word matching score function to obtain the word-level image-text matching loss value in the i-th stage; the expression is as follows: γ i = softmax(θ 1 softmax(W i T R i ), i = 1, 2, 3; (24) C i = R i γ i T , i = 1, 2, 3; (25) Among them, W i represents the word feature matrix in the i-th stage, with a dimension of D w ×N w ; R i represents the region feature matrix in the i-th stage, with a dimension of D w ×289; W i T R i represents the similarity matrix in the i-th stage, with a dimension of N w ×289; γ i represents the attention weight coefficient matrix in the i-th stage, with a dimension of N w ×289; C i represents the region context feature matrix in the i-th stage, with a dimension of D w ×N w ; represents the n-th region context feature vector in the region context feature matrix in the i-th stage, with a dimension of D w ; represents the n-th word feature vector in the word feature matrix in the i-th stage, with a dimension of D w ; represents the region-word matching score function in the i-th stage; represents M pairs of region-word pairs, and only matched The other M - 1 pairs are regarded as unmatched region-word pairs; θ 1 , θ 2 , θ 3 respectively represent the first, second, and third hyperparameters; The calculation process of the sentence-level image-text matching loss function in the i-th stage is as follows: First, the image output by the generator in the i-th stage is input into the image encoder to obtain the global image feature vector in the i-th stage, and then the sentence feature vector in the i-th stage and the global image feature vector are input into the image-sentence matching score function to obtain the sentence-level image-text matching loss value in the i-th stage, and the expression is as follows: where s i represents the sentence feature vector of the i-th stage, with dimension D s ; g i represents the global image feature vector of the i-th stage, with dimension D s ; represents the image-sentence matching score function for the i-th stage; represents M pairs of image-sentence pairs, and only matched while the other M - 1 pairs are regarded as unmatched image-sentence pairs; θ 4 represents the fourth hyperparameter.

Citation Information

Patent Citations

  • Image description generation method fusing visual common sense and enhancing multilayer global features

    CN113378919A

  • Method for generating image semantic description

    WO2020244287A1