Story visualization method based on graph neural network and adaptive attention
Through graph neural networks and adaptive attention models, we construct a character attribute embedding matrix and filter redundant information, solving the problem of insufficient character consistency in existing technologies and generating high-quality images.
Patent Information
- Application Number
- CN202510811102.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-16
AI Technical Summary
Existing story visualization methods have shortcomings in maintaining character consistency and the generated images are of low quality.
A method based on graph neural networks and adaptive attention is adopted to construct a character attribute embedding matrix through a graph convolutional neural network. A gating mechanism is introduced to filter redundant information, and combined with the character attribute classification loss function to improve the image generation quality.
Effectively maintain the character consistency in the generated image sequence and generate high-quality images.
Smart Images

Figure CN120653796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image generation technology, and in particular to a story visualization method based on graph neural network and adaptive attention. Background Art
[0002] With the acceleration of digitalization, China's digital publishing industry is growing rapidly. Comics, a narrative art that combines images and text, requires a creative process that includes storyboarding, scriptwriting, planning key scenes and text, creating storyboards based on the script to depict key scenes, and finally completing the storyboard sketches, coloring, and polishing. This process is not only time-consuming, especially during the storyboarding stage, but also requires the creator to invest a great deal of imagination to draw images that match the story, making it a time-consuming and laborious task.
[0003] Story visualization is a variant of text-to-image generation. It aims to visualize a multi-sentence story text by generating a series of images, with one image for each sentence. Compared with traditional text-to-image generation models that only focus on the generation of a single image, story visualization can generate multiple images with semantic continuity that match the description of the story content, which is more suitable for real-world application scenarios and more challenging. It can well meet the application needs of comic creators, helping them to quickly and automatically generate initial drawings and speed up the workflow of comic creation. At present, some related research works have proposed GAN-based story visualization models to promote the further development of story visualization tasks. As the key elements of the story, characters drive the development of the plot and usually occupy a large area in the generated image. Therefore, one of the challenges of this task is to ensure the consistency of the characters in the generated image.
[0004] To address these challenges, Li et al. first proposed StoryGAN, a framework based on sequential conditional GANs. This framework ensures that generated image sequences are both locally and globally consistent through image and story discriminators. Song et al. built on StoryGAN by introducing a foreground segmentation generation module, which generates low-level foreground shallow features to assist the image generator in generating character-consistent image sequences. Maharana et al. employed an additional auxiliary captioning network and a copy-transformation module to promote semantic alignment and appearance consistency between the story and generated images. Li et al. proposed a novel sentence representation, a discriminator with fused features, and extended spatial attention to selectively integrate different words in the story and focus on generating detailed image regions. However, these methods still have shortcomings in maintaining character consistency within the story, failing to consider how inherent character attributes can enhance and supplement the limited character information within the story text. Furthermore, when using attention mechanisms to capture the relationship between words and image regions, they fail to account for the potential redundant information between the two modalities, which can interfere with the generation of image details and reduce image quality. Consequently, existing story visualization methods lack character consistency and fail to generate high-quality images. Summary of the Invention
[0005] The purpose of this invention is to propose a story visualization method based on graph neural networks and adaptive attention to address the problems that existing story visualization methods are insufficient in maintaining character consistency and cannot generate high-quality images.
[0006] The technical solution adopted by the present invention to solve the above technical problems is:
[0007] The story visualization method based on graph neural network and adaptive attention includes the following steps:
[0008] Input the story text to be processed into the trained story visualization model based on graph neural network and adaptive attention to generate visualization images;
[0009] The trained story visualization model based on graph neural network and adaptive attention is obtained through the following steps:
[0010] Obtain a story text containing multiple sentences and input it into the generator to obtain a generated image. Then, input the generated image into the discriminator for adversarial training to obtain a trained story visualization model based on graph neural network and adaptive attention.
[0011] The generator is a pre-trained generator, which is obtained by the following steps:
[0012] Step 1: Obtain a story text containing multiple sentences, where each sentence in the story text corresponds to a cartoon character image. Then, describe the appearance of the cartoon character in each cartoon character image in text form, i.e., attribute description. Then, use a text encoder to encode the attribute description into attribute embedding, and use the attribute embedding as the initialization node feature of the graph convolutional neural network. ;
[0013] Step 2: For a cartoon character image, any two attribute descriptions in the image are combined into attribute pairs. Then, a correlation matrix is constructed based on the number of occurrences of all attribute pairs in all cartoon characters corresponding to the story text.
[0014] Step 3: Get the weight matrix , and combined with the correlation matrix to update the node features in the graph convolutional neural network to obtain the updated node features ;
[0015] Step 4: Use one-hot encoding to encode the attribute descriptions and obtain the encoding vector, i.e. the attribute label ;
[0016] Step 5: Update the node features With attribute tags The product of is used as the final updated attribute embedding, that is, the attribute feature;
[0017] Step 6: Obtain image features , the specific steps are:
[0018] Step 61: Input each sentence in the story text into the text encoder to obtain the sentence embedding s and word embedding corresponding to the sentence ;
[0019] Step 62: After the sentence embedding s passes through the fully connected layer, batch normalization layer, Tanh function and convolution processing, the feature vector is obtained ;
[0020] Step 63: Embed the sentence into s and perform Gaussian sampling to obtain the feature vector ;
[0021] Step 64: Obtain noise that follows a Gaussian distribution, concatenate the noise with the sentence embedding s, and input it into the gated recurrent unit to obtain the feature vector ;
[0022] Step 65: Eigenvector , eigenvector and eigenvectors After splicing, the image features are obtained by passing through the fully connected layer, batch normalization layer and ReLU function in sequence. ;
[0023] Step 7: Image features Fusion with attribute features to obtain image features ;
[0024] Step 8: Image features After passing through the upsampling layer, convolution layer, batch normalization layer and ReLU function in sequence, the image features are obtained. , ;
[0025] Step 9: Image features Mapping to joint semantic space , perform tensor dimension transformation to obtain image features , , where D represents the feature dimension, H represents the image feature height, W represents the image feature width, and N represents the size of the story. Indicates the number of channels of the image;
[0026] Step 10: Word Embedding Reshape into ;
[0027] Step 11: Utilize image features and Get the similarity matrix between image and text , Expressed as:
[0028] ;
[0029] Step 12: Using the similarity matrix of image and text and , get text-guided attention , Expressed as:
[0030] ;
[0031] Step 13: Utilize image features and text-guided attention , get the door , Expressed as:
[0032]
[0033] Step 14: Close the door Send it to the fully connected layer and get the gated mask through the activation function , Expressed as:
[0034]
[0035] in, 、 and represents the learnable mapping matrix, Represents the element-wise multiplication operation between matrices. Represents the sigmoid activation function;
[0036] Step 15: Based on gated mask , get the cross attention weight matrix , Expressed as:
[0037]
[0038] in, , Indicates the first visual space location and The attention weight of each word; , express The element in row i and column j of ;
[0039] Step 16: Use the cross attention weight matrix obtained in step 15 and , get the weighted visual features containing word information , Expressed as:
[0040]
[0041] Step 17: Weighted visual features and image features Splicing to obtain image features ;
[0042] Step 18: Image features The image features are obtained by sequentially passing through the upsampling layer, convolution layer, batch normalization layer and ReLU function. ;
[0043] Step 19: Image features Replace image features , repeat steps 11 to 18 to obtain the final image features, and then obtain the final generated image.
[0044] Furthermore, the correlation matrix is expressed as:
[0045]
[0046] in, Indicates the Attributes and The probability that an attribute appears in the entire training set at the same time, Indicates the threshold used to filter noise.
[0047] Furthermore, in step 3, updating the node features in the graph convolutional neural network is performed by a propagation rule, which is expressed as:
[0048]
[0049] in, Represents the graph convolutional neural network The learnable weight matrix of the layer, and Represent the node features of input and output respectively.
[0050] Furthermore, the fusion in step seven is expressed as:
[0051]
[0052] in, Represents DFBlock.
[0053] Furthermore, the loss function of the discriminator is:
[0054]
[0055] in, and denote story adversarial loss and image adversarial loss respectively, represents the role attribute classification loss, and Represents the weight coefficient.
[0056] Furthermore, the role attribute classification loss Expressed as:
[0057]
[0058] in, Indicates the number of character attributes, represents the i-th attribute of the real role attribute label, Represents the i-th attribute of the predicted role attribute label.
[0059] Furthermore, the story of the struggle against loss Expressed as:
[0060]
[0061] in, represents the fusion features of word features and story features sampled from the true distribution, represents the fused features of word features and story features sampled from the model distribution, D() represents the discriminator, G() represents the generator, and E represents the expectation.
[0062] Furthermore, the image adversarial loss Expressed as:
[0063]
[0064] in, represents the fusion features of word features and image features sampled from the true distribution, Represents the fusion features of word features and image features sampled from the model distribution.
[0065] Furthermore, the image features Fusion with attribute features is performed through DFBlock, which includes two affine layers, two ReLU activation layers and one convolutional layer.
[0066] Furthermore, the text encoder is a DAMSM model.
[0067] The beneficial effects of the present invention are:
[0068] This application introduces a graph convolutional neural network to construct a character attribute embedding matrix and a character attribute correlation matrix. The character attribute node representation is continuously updated through the information transmission mechanism of the graph convolutional neural network. The character attributes finally obtained based on the graph convolutional neural network are used as supplementary information for the characters in the story text, so as to better maintain the consistency of the characters in the generated image sequence. In addition, this application introduces a gating mechanism when calculating the attention between modalities to dynamically filter out redundant information. Therefore, this application can generate high-quality images. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 Schematic diagram of a graph convolutional neural network encoding character attributes;
[0070] Figure 2 Schematic diagram of the adaptive attention module;
[0071] Figure 3 Schematic diagram of a story visualization model based on graph neural network and adaptive attention. DETAILED DESCRIPTION
[0072] It should be noted that, unless there is any conflict, the various embodiments disclosed in this application can be combined with each other.
[0073] Specific implementation method 1: refer to Figure 1 Specifically describing this embodiment, the story visualization method based on graph neural network and adaptive attention described in this embodiment includes the following steps:
[0074] Input the story text to be processed into the trained story visualization model based on graph neural network and adaptive attention to generate visualization images;
[0075] The trained story visualization model based on graph neural network and adaptive attention is obtained through the following steps:
[0076] Obtain a story text containing multiple sentences and input it into the generator to obtain a generated image. Then, input the generated image into the discriminator for adversarial training to obtain a trained story visualization model based on graph neural network and adaptive attention.
[0077] The generator is a pre-trained generator, which is obtained by the following steps:
[0078] Step 1: Obtain a story text containing multiple sentences, where each sentence in the story text corresponds to a cartoon character image. Then, describe the appearance of the cartoon character in each cartoon character image in text form, i.e., attribute description. Then, use a text encoder to encode the attribute description into attribute embedding, and use the attribute embedding as the initialization node feature of the graph convolutional neural network. ;
[0079] Step 2: For a cartoon character image, any two attribute descriptions in the image are combined into attribute pairs. Then, a correlation matrix is constructed based on the number of occurrences of all attribute pairs in all cartoon characters corresponding to the story text.
[0080] Step 3: Get the weight matrix , and combined with the correlation matrix to update the node features in the graph convolutional neural network to obtain the updated node features ;
[0081] Step 4: Use one-hot encoding to encode the attribute descriptions and obtain the encoding vector, i.e. the attribute label ;
[0082] Step 5: Update the node features With attribute tags The product of is used as the final updated attribute embedding, that is, the attribute feature;
[0083] Step 6: Obtain image features , the specific steps are:
[0084] Step 61: Input each sentence in the story text into the text encoder to obtain the sentence embedding s and word embedding corresponding to the sentence ;
[0085] Step 62: After the sentence embedding s passes through the fully connected layer, batch normalization layer, Tanh function and convolution processing, the feature vector is obtained ;
[0086] Step 63: Embed the sentence into s and perform Gaussian sampling to obtain the feature vector ;
[0087] Step 64: Obtain noise that follows a Gaussian distribution, concatenate the noise with the sentence embedding s, and input it into the gated recurrent unit to obtain the feature vector ;
[0088] Step 65: Eigenvector , eigenvector and eigenvectors After splicing, the image features are obtained by passing through the fully connected layer, batch normalization layer and ReLU function in sequence. ;
[0089] Step 7: Image features Fusion with attribute features to obtain image features ;
[0090] Step 8: Image features After passing through the upsampling layer, convolution layer, batch normalization layer and ReLU function in sequence, the image features are obtained. , ;
[0091] Step 9: Image features Mapping to joint semantic space , perform tensor dimension transformation to obtain image features , , where D represents the feature dimension, H represents the image feature height, W represents the image feature width, and N represents the size of the story. Indicates the number of channels of the image;
[0092] Step 10: Word Embedding Reshape into ;
[0093] Step 11: Utilize image features and Get the similarity matrix between image and text , Expressed as:
[0094] ;
[0095] Step 12: Using the similarity matrix of image and text and , get text-guided attention , Expressed as:
[0096] ;
[0097] Step 13: Utilize image features and text-guided attention , get the door , Expressed as:
[0098]
[0099] Step 14: Close the door Send it to the fully connected layer and get the gated mask through the activation function , Expressed as:
[0100]
[0101] in, 、 and represents a learnable mapping matrix, ⨀ represents the element-wise multiplication operation between matrices, Represents the sigmoid activation function;
[0102] Step 15: Based on gated mask , get the cross attention weight matrix , Expressed as:
[0103]
[0104] in, , Indicates the first visual space location and The attention weight of each word; , express The element in row i and column j of ;
[0105] Step 16: Use the cross attention weight matrix obtained in step 15 and , get the weighted visual features containing word information , Expressed as:
[0106]
[0107] Step 17: Weighted visual features and image features Splicing to obtain image features ;
[0108] Step 18: Image features The image features are obtained by sequentially passing through the upsampling layer, convolution layer, batch normalization layer and ReLU function. ;
[0109] Step 19: Image features Replace image features , repeat steps 11 to 18 to obtain the final image features, and then obtain the final generated image.
[0110] This application uses graph convolutional neural networks (such as Figure 1 As shown in the figure, character attributes are encoded to enhance the representation of story text characters, character attribute classification loss is introduced to promote semantic alignment of text and image, and adaptive attention modules (such as Figure 2 This application improves the dynamic interaction between text and images, enhancing the quality of image generation. This application is applicable to areas such as web comic creation and animation initial plot creation, and has broad application potential, significantly improving creators' work efficiency and the quality of their work.
[0111] A GAN-based story visualization model is used to generate image sequences based on stories. The model consists of three parts: a graph convolutional neural network that encodes character attributes, an adaptive attention module, and a discriminator that incorporates character attribute classification loss.
[0112] Graph Convolutional Neural Networks for Encoding Character Attributes
[0113] In previous story visualization tasks, many works have only used sentence embeddings of text as input, thus neglecting the attribute information of the characters in the story. This information has a significant impact on generating images with story characters that match the description. Since there are connections between attributes, to better model the global correlation between them, this application introduces a graph convolutional neural network to construct a character attribute embedding matrix and a character attribute correlation matrix. The character attribute node representation is continuously updated through the information transfer mechanism of the graph convolutional neural network. The resulting graph convolutional neural network-based character attributes are used as supplementary information for the characters in the story text, better maintaining the consistency of the characters in the generated image sequence.
[0114] Adaptive Attention Module
[0115] Previous work has used semantically-based cross-attention to obtain visually guided textual attention, which is then fused with initial features to enhance the visual representation, resulting in higher-quality story images with finer details and better consistency. However, this approach ignores the potential redundant features between the two modalities, which causes the two modal features to focus more on meaningless information when interacting with each other. Therefore, this application proposes an adaptive attention module that introduces a gating mechanism when calculating attention between modalities to dynamically filter out redundant information.
[0116] Discriminator integrating role attribute classification loss
[0117] This application uses a discriminator with unidirectional output. First, the fine-grained text information and image information at the word level are fused, and then the fused features are input into the discriminator, so that the two goals of image quality and semantic consistency between image and text can be combined into one goal. The goal of the discriminator is to promote the generator to produce better fusion features. The constructed fusion features should have good image generation quality, and there should be an effective connection between the image area and the corresponding words. In order to further enhance the relationship between the generated image area and the corresponding attribute, this application introduces a role classification loss function to perform alignment between attributes and image features.
[0118] This section will introduce the specific implementation methods of the components of the story visualization model based on graph neural networks and adaptive attention.
[0119] Implementation of Graph Convolutional Neural Networks for Encoding Character Attributes
[0120] Get a story text containing multiple sentences, each sentence in the text corresponds to a cartoon character image, and describe the appearance of the cartoon character in each cartoon character image in text, that is, attribute description. Then use the text encoder to encode the attribute description into attribute embedding, and use the attribute embedding as the initial node feature of the graph convolutional neural network (GCN) ,
[0121] After that, for a cartoon character image, any two attribute descriptions in the image are combined into attribute pairs, and then a correlation matrix C is constructed based on the number of occurrences of all attribute pairs in the cartoon character image set.
[0122] First, this application collects attribute descriptions of all characters, such as black eyes and round noses, and then uses a pre-trained text encoder (DAMSM model) to encode these attributes into attribute embeddings and use them as the initial node features of the graph. Then, we count the number of times attribute pairs appear in the training set to construct the correlation matrix C, which is expressed as follows:
[0123] in It represents the probability that the i-th attribute and the j-th attribute appear in the entire training set at the same time, and t represents the threshold used to filter noise.
[0124] Initialize node features given and the correlation matrix After that, GCN passes the correlation matrix and the weight matrix Perform linear transformation and nonlinear transformation of activation function to update the node features. For node i, the calculation process of propagation rule is as follows:
[0125]
[0126] in is the ReLU activation function, is the learnable weight matrix of the GCN layer l, and Corresponding to the node features of input and output respectively.
[0127] Get the weight matrix , and combined with the correlation matrix Update node features in GCN,
[0128] Use one-hot to encode the attribute descriptions and obtain the encoding vector, that is, the attribute label y.
[0129] Finally, we get the updated node features , and multiply it with the attribute label y as the final updated attribute embedding (attribute feature).
[0130] The steps for obtaining image features are:
[0131] Input each sentence in the story text into the text encoder to obtain the sentence embedding (s) and word embedding (w) corresponding to the sentence
[0132] 1. After the sentence embedding (s) passes through the fully connected layer, batch normalization layer, Tanh function and convolution processing, the feature vector S1 is obtained;
[0133] 2. Perform Gaussian sampling on the sentence embedding (s) to obtain the feature vector S2;
[0134] 3. Obtain noise that follows a Gaussian distribution, concatenate the noise with the sentence embedding (s), and input it into the GRU to obtain the feature vector S3.
[0135] After concatenating the feature vectors S1, S2, and S3, they are passed through the fully connected layer, batch normalization layer, and ReLU function in sequence to obtain the image features. .
[0136] In order to better integrate attribute embedding, this application uses DFBlock in the work of Tao et al. to fully integrate image features with attribute features. Among them, DFBlock consists of two affine layers, two ReLU activation layers and a convolutional layer. This fusion process is expressed as follows:
[0137]
[0138] in represents DFBlock,
[0139] Image features After passing through the upsampling layer, convolution layer, batch normalization layer and ReLU function in sequence, the image features are obtained. .
[0140] Implementation of the Adaptive Attention Module
[0141] First, this application uses a convolutional layer to transform the intermediate state image features generated by the generator Mapping to joint semantic space , and then perform tensor dimension transformation to obtain new image features , where D is the feature dimension, H is the visual feature height, W is the visual feature width, and N is the size of the story. Similarly, this application embeds the word Reshape into , where D is the feature dimension and L is the number of words in a sentence. Then, we get the similarity matrix between the image and the text :
[0142]
[0143] Then calculate the text-guided attention :
[0144]
[0145] in, A is normalized according to the word dimension, which represents the attention weight of each word in each visual area.
[0146] In order to select effective information for fusion between modalities, this application adopts a gating mechanism, which first calculates the gate through the attention calculation of visual features and text guidance. , and then send the calculated gate into the fully connected layer and obtain the gate mask through the activation function , the specific calculation process is as follows:
[0147]
[0148]
[0149] in, 、 and is a learnable mapping matrix, Represents the element-wise multiplication operation between matrices. Represents the sigmoid activation function.
[0150] Then, this application uses the gated mask M to calculate the filtered cross attention :
[0151]
[0152] in, , Represents the correlation between the i-th visual space position and the j-th word in the whole story after the gating mechanism. Then, the weighted visual features containing word information are calculated by itself. :
[0153]
[0154] Finally, this application combines the obtained weighted visual features with the image features After splicing together, we get the image feature q,
[0155] The image feature q is passed through the upsampling layer, convolution layer, batch normalization layer and ReLU function in sequence to obtain the image feature p. Replace image features , repeat the above steps (image features To the splicing step), the final image feature p is obtained. Finally, the image feature p and the word embedding w are input into the discriminator for identification.
[0156] Implementation of Discriminator with Fusion Character Attribute Classification Loss
[0157] To further enhance the relationship between the generated image regions and the corresponding attributes, this application introduces a role attribute classification loss function to align role attributes with image features. The discriminator's classifier maps the fused features into binary attribute labels. Given the true attribute labels, this application calculates the discriminator's role attribute classification loss:
[0158]
[0159] in, and Denote the true label and the attribute label predicted by the discriminator, respectively, and n denotes the number of attributes. Therefore, the total loss of the discriminator is:
[0160]
[0161]
[0162]
[0163] in, and denote story adversarial loss and image adversarial loss respectively, and is the weight coefficient. (or ) represents the fusion features of word features and story features (or image features) sampled from the true distribution. (or ) represents the fusion features of word features and story features (or image features) sampled from the model distribution.
[0164] It should be noted that the specific embodiments are merely explanations and illustrations of the technical solutions of the present invention and cannot be used to limit the scope of protection. Any minor changes made based on the claims and description of the present invention shall still fall within the scope of protection of the present invention.
Claims
1. Story visualization method based on graph neural network and adaptive attention, characterized by The following steps are involved: Input the story text to be processed into the trained story visualization model based on graph neural network and adaptive attention to generate visualization images; The trained story visualization model based on graph neural network and adaptive attention is obtained through the following steps: Obtain a story text containing multiple sentences and input it into the generator to obtain a generated image. Then, input the generated image into the discriminator for adversarial training to obtain a trained story visualization model based on graph neural network and adaptive attention. The generator is a pre-trained generator, which is obtained by the following steps: Step 1: Obtain a story text containing multiple sentences, where each sentence in the story text corresponds to a cartoon character image. Then, describe the appearance of the cartoon character in each cartoon character image in text form, i.e., attribute description. Then, use a text encoder to encode the attribute description into attribute embedding, and use the attribute embedding as the initialization node feature of the graph convolutional neural network. ; Step 2: For a cartoon character image, any two attribute descriptions in the image are combined into attribute pairs. Then, a correlation matrix is constructed based on the number of occurrences of all attribute pairs in all cartoon characters corresponding to the story text. Step 3: Get the weight matrix , and combined with the correlation matrix to update the node features in the graph convolutional neural network to obtain the updated node features ; Step 4: Use one-hot encoding to encode the attribute descriptions and obtain the encoding vector, i.e. the attribute label ; Step 5: Update the node features With attribute tags The product of is used as the final updated attribute embedding, that is, the attribute feature; Step 6: Obtain image features , the specific steps are: Step 61: Input each sentence in the story text into the text encoder to obtain the sentence embedding s and word embedding corresponding to the sentence ; Step 62: After the sentence embedding s passes through the fully connected layer, batch normalization layer, Tanh function and convolution processing, the feature vector is obtained ; Step 63: Embed the sentence into s and perform Gaussian sampling to obtain the feature vector ; Step 64: Obtain noise that follows a Gaussian distribution, concatenate the noise with the sentence embedding s, and input it into the gated recurrent unit to obtain the feature vector ; Step 65: Eigenvector , eigenvector and eigenvectors After splicing, the image features are obtained by passing through the fully connected layer, batch normalization layer and ReLU function in sequence. ; Step 7: Image features Fusion with attribute features to obtain image features ; Step 8: Image features After passing through the upsampling layer, convolution layer, batch normalization layer and ReLU function in sequence, the image features are obtained. , ; Step 9: Image features Mapping to joint semantic space , perform tensor dimension transformation to obtain image features , , where D represents the feature dimension, H represents the image feature height, W represents the image feature width, and N represents the size of the story. Indicates the number of channels of the image; Step 10: Word Embedding Reshape into ; Step 11: Utilize image features and Get the similarity matrix between image and text , Expressed as: ; Step 12: Using the similarity matrix of image and text and , get text-guided attention , Expressed as: ; Step 13: Utilize image features and text-guided attention , get the door , Expressed as: Step 14: Close the door Send it to the fully connected layer and get the gated mask through the activation function , Expressed as: in, 、 and represents the learnable mapping matrix, Represents the element-wise multiplication operation between matrices. Represents the sigmoid activation function; Step 15: Based on gated mask , get the cross attention weight matrix , Expressed as: in, , Indicates the first visual space location and The attention weight of each word; , express The element in row i and column j of ; Step 16: Use the cross attention weight matrix obtained in step 15 and , get the weighted visual features containing word information , Expressed as: Step 17: Weighted visual features and image features Splicing to obtain image features ; Step 18: Image features The image features are obtained by sequentially passing through the upsampling layer, convolution layer, batch normalization layer and ReLU function. ; Step 19: Image features Replace image features , repeat steps 11 to 18 to obtain the final image features, and then obtain the final generated image.
2. The story visualization method based on graph neural network and adaptive attention according to claim 1 is characterized in that The correlation matrix is expressed as: in, Indicates the Attributes and The probability that an attribute appears in the entire training set at the same time, Indicates the threshold used to filter noise.
3. The story visualization method based on graph neural network and adaptive attention according to claim 2 is characterized in that In step 3, updating the node features in the graph convolutional neural network is performed through a propagation rule, which is expressed as: in, Represents the graph convolutional neural network The learnable weight matrix of the layer, and Represent the node features of input and output respectively.
4. The story visualization method based on graph neural network and adaptive attention according to claim 3 is characterized in that The fusion in step seven is expressed as: in, Represents DFBlock.
5. The story visualization method based on graph neural network and adaptive attention according to claim 4 is characterized in that The loss function of the discriminator is: in, and denote story adversarial loss and image adversarial loss respectively, represents the role attribute classification loss, and Represents the weight coefficient.
6. The story visualization method based on graph neural network and adaptive attention according to claim 5 is characterized in that The role attribute classification loss Expressed as: in, Indicates the number of character attributes, represents the i-th attribute of the real role attribute label, Represents the i-th attribute of the predicted role attribute label.
7. The story visualization method based on graph neural network and adaptive attention according to claim 6 is characterized in that The story of confronting loss Expressed as: in, represents the fusion features of word features and story features sampled from the true distribution, represents the fused features of word features and story features sampled from the model distribution, D() represents the discriminator, G() represents the generator, and E represents the expectation.
8. The story visualization method based on graph neural network and adaptive attention according to claim 7 is characterized in that The image adversarial loss Expressed as: in, represents the fusion features of word features and image features sampled from the true distribution, Represents the fusion features of word features and image features sampled from the model distribution.
9. The story visualization method based on graph neural network and adaptive attention according to claim 8 is characterized in that The image features Fusion with attribute features is performed through DFBlock, which includes two affine layers, two ReLU activation layers and one convolutional layer.
10. The story visualization method based on graph neural network and adaptive attention according to claim 9 is characterized in that The text encoder is a DAMSM model.
Citation Information
Cited By
Methods and systems for generating one or more emoticons for one or more users
US20240233206A1