Text image generation method and system based on cross-modal information guide fusion

By introducing attention-driven deep text image fusion module and global semantic refinement module into the text generation image method, the problems of insufficient fusion of text visual features and insufficient mining of dependencies are solved, and higher semantic consistency and image quality are achieved.

CN120147479APending Publication Date: 2025-06-13CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510275915.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing text image generation method does not fusion of text visual features during the initialization stage, and insufficient mining of dependencies between image sub-regions during the generation process, resulting in low text semantic consistency and visual authenticity of the generated image.

Method used

Using a method based on cross-modal information guidance fusion, the deep text image fusion module driven by attention-driven and the global semantic refinement module combined with the Mamba module are effectively fused, and the dependencies between image sub-regions are mined.

Benefits of technology

It significantly improves the semantic consistency between the generated image and the text description and greatly improves the quality of the generated image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147479A_ABST
    Figure CN120147479A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, in particular to a text image generation method and system based on cross-modal information guide fusion, and the method comprises the steps: a text encoder processes the text description of a target image to obtain sentence features and word features; enhancing the sentence features to obtain image features; the attention-driven deep text image fusion module fuses the sentence features, the word features and the image features, and further generates an initial image; processing the word features and the features of the initial image by combining a global semantic refinement module of a Mama module to obtain cross-modal fusion features; generating a target image according to the cross-modal fusion features; and updating the cross-modal fusion features and generating an optimized image. According to the method, the technical problem of insufficient text visual feature fusion is solved, the semantic consistency is improved, and the technical defect that a multi-stage generative adversarial network is insufficient in mining of dependency relationships among image sub-regions in the image generation process is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and in particular to a method and system for text-to-image generation based on cross-modal information-guided fusion. Background Art

[0002] Text-to-image generation technology is one of the most challenging tasks in cross-modal, and its goal is to generate semantically consistent images according to the given text description. This technology not only promotes more natural interaction between humans and computers, but also realizes the automatic generation of images through simple text instructions, significantly reducing the threshold of the creation process.

[0003] Currently, most text-to-image methods adopt a multi-stage generation strategy to obtain high-quality images. Under the framework of multi-stage generation, in the initial stage, the resolution of the generated image is mainly increased through upsampling, and in the subsequent stages of the generation process, attention mechanisms are applied to obtain context information between modalities (text-image).

[0004] In summary, the current research urgently needs to solve the following two key problems: one is how to fully integrate text information into visual features in the initialization stage; the other is how to more effectively mine the dependencies between image sub-regions during the generation process to improve the text semantic consistency and visual authenticity of the generated images. Summary of the Invention

[0005] In view of this, the present invention discloses a method and system for text-to-image generation based on cross-modal information-guided fusion to solve the above problems; including:

[0006] A method for text-to-image generation based on cross-modal information-guided fusion, including:

[0007] S1. Obtain the text description of the target image, and process the text description of the target image using a pre-trained text encoder to obtain sentence features and word features; the text encoder uses a bidirectional long short-term memory network text encoder;

[0008] S2. Perform conditional enhancement and noise splicing on the sentence features to obtain image features;

[0009] S3. Process the sentence features, word features, and image features using an attention-driven deep text-image fusion module to obtain initial fusion features;

[0010] S4. Generate an image according to the initial fusion features to obtain an initial image;

[0011] S5. Process the word features and the features of the initial image using a global semantic refinement module combined with a Mamba module to obtain cross-modal fusion features;

[0012] S6. Generate an image based on the cross-modal fusion features to obtain the target image;

[0013] S7. Use the global semantic refinement module combined with the Mamba module to process the word features and the features of the target image to obtain the updated cross-modal fusion features;

[0014] S8. Generate an image based on the updated cross-modal fusion features to obtain the optimized image;

[0015] A text-to-image generation system based on cross-modal information-guided fusion, comprising:

[0016] A text encoder for obtaining the sentence features and word features in the text description of the target image;

[0017] Enhance the sentence features using a conditional enhancement function, and splice the enhanced sentence features with a noise vector to obtain image features;

[0018] An attention-driven deep text-image fusion module for processing the sentence features, word features, and image features to obtain the initial fusion features;

[0019] A first generator for generating an initial image based on the initial fusion features;

[0020] A global semantic refinement module combined with Mamba for processing the word features and the features of the image generated by the generator to generate cross-modal fusion features;

[0021] A second generator for generating a target image based on the cross-modal fusion features;

[0022] A third generator for generating an optimized image based on the updated cross-modal fusion features;

[0023] A discriminator for updating the text-to-image generation system based on the initial image, target image, and optimized image;

[0024] A DAMSM module for updating the text-to-image generation system based on the optimized image and the text description.

[0025] The beneficial effects of the present invention include: By introducing an attention-driven deep text-image fusion module, the technical problem of insufficient fusion of text visual features in the initialization stage of the multi-stage generative adversarial network is effectively solved, and the semantic consistency between the generated image and the text description is significantly improved; By combining the global semantic refinement module of Mamba, the technical defect of insufficient mining of the dependency relationship between image sub-regions in the process of text-to-image generation by the multi-stage generative adversarial network is overcome, thereby greatly improving the quality of the generated image. Description of the Drawings

[0026] Figure 1 Schematic diagram of the step process of the method for generating images from text based on cross-modal information-guided fusion according to the present invention;

[0027] Figure 2 Schematic diagram of the network structure of the system for generating images from text based on cross-modal information-guided fusion according to the present invention;

[0028] Figure 3 Schematic diagram of the structure of the attention-driven deep text-image fusion module according to the present invention;

[0029] Figure 4 Schematic diagram of the structure of the conditional batch normalization layer according to the present invention;

[0030] Figure 5 Schematic diagram of the structure of the global semantic refinement module combined with the Mamba module according to the present invention. Detailed implementation manners

[0031] In order to make the purpose, technical solutions, features and advantages of the present invention clearer and more understandable, the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0032] Embodiment 1

[0033] This embodiment provides a method for generating images from text based on cross-modal information-guided fusion, as Figure 1 shown, including:

[0034] S1. Obtain the text description of the target image, and process the text description of the target image by using a pre-trained text encoder to obtain sentence features and word features; the text encoder uses a bidirectional long short-term memory network text encoder.

[0035] Specifically, extract the sentence feature e sent and the word feature e word :

[0036] e word , e sent = F LSTM (T)

[0037] S2. Perform conditional enhancement and noise splicing on the sentence features to obtain image features.

[0038] Specifically, after extracting the sentence feature e word , further enhance it through the conditional enhancement function F ca to obtain the augmented sentence feature

[0039]

[0040] Select augmented sentence features Perform vector concatenation with a random noise vector z that follows a Gaussian distribution, and through a fully connected layer and a deformation operation, obtain the image feature x:

[0041]

[0042] S3. Use an attention-driven deep text-image fusion module to process the sentence features, word features, and image features to obtain initial fusion features.

[0043] Specifically, the processing of data by the attention-driven deep text-image fusion module includes: using an attention mechanism to calculate the cross-modal context weight information between the word features and the image features, and predicting the channel scaling parameter and the offset parameter according to the sentence features; according to the cross-modal context weight information, the channel scaling parameter, and the offset parameter, perform feature channel scaling and offset changes on the image features to obtain an enhanced vector; perform two non-linear operations, one convolutional operation, and one upsampling operation on the enhanced vector in sequence to obtain the initial fusion features; the non-linear operation uses the LeakyReLU activation function.

[0044] Furthermore, use an attention mechanism to calculate the word feature e word and the cross-modal context weight information m attn between the image feature x, and the formula is:

[0045]

[0046] where N represents the dimension of the image feature, T represents the dimension of the word feature, represents the feature representation of the i-th word, represents the j-th column of the image feature, and U represents a perceptron, which is used to map the word feature to a common semantic space shared with the image feature.

[0047] Furthermore, predicting the channel scaling parameter and the offset parameter according to the sentence features includes: using a multi-layer perceptron (MLPs) to predict the language-conditioned channel scaling parameter γ and the offset parameter β from the sentence feature e sent :

[0048] γ = MLPs(e sent )

[0049] β = MLPs(e sent )

[0050] Furthermore, performing feature channel scaling and offset changes on the image features includes: multiplying the cross-modal context weight information m attnPerform element-wise multiplication operations with the predicted channel scaling parameter γ and offset parameter β respectively, perform feature channel scaling and offset changes on the image features x input in the previous layer, and obtain the enhanced vector after injecting text information.

[0051]

[0052] Furthermore, the initial stage of the upsampling operation includes four fusion modules and four upsampling layers, which are used to increase the resolution of the enhanced vector from 4×4 to 64×64. Each upsampling layer doubles the spatial resolution of the input feature map to ensure that the final high-resolution initial fusion features are generated.

[0053] S4. Generate an image based on the initial fusion features to obtain the initial image.

[0054] S5. Use the global semantic refinement module combined with the Mamba module to process the word features and the features of the initial image to obtain cross-modal fusion features.

[0055] Specifically, the processing of data by the global semantic refinement module combined with the Mamba module includes: mapping the word features to a common semantic space; calculating the similarity between the mapped word features and the features of the image generated by the generator, convolving the similarity with the word features to obtain inter-modal context information; processing the features of the generated image by the Mamba sequence model to obtain intra-modal context information; concatenating the image features generated by the generator, inter-modal context information, and intra-modal context information, and performing multi-scale feature fusion and upsampling processing to obtain cross-modal fusion features.

[0056] Furthermore, map the word feature e word to the common semantic space shared with the image features through the perceptron layer U to obtain the mapped word feature Ue word .

[0057] Furthermore, calculate the similarity between the initial image feature h 0 and the mapped word feature Ue word , perform an inner product operation on the word feature and the similarity to obtain the inter-modal context information h attn :

[0058]

[0059] Furthermore, flatten the image feature to obtain where L = H×W, and input it into the Mamba module after passing through the layer normalization layer. The Mamba module contains two parallel branches: The first branch: expand the input feature to Subsequently, h passes through a one-dimensional convolutional layer, a SiLU activation function, and an SSM layer in sequence to obtain the output of the first branch; Second branch: The input features are extended to Subsequently, h passes through the SiLU activation function to obtain the output of the second branch.

[0060] Further, the outputs of the two branches are combined, and the Hadamard method is used to perform element-wise multiplication and combination of the features of the two branches, and the original shape is restored through linear projection to obtain the image feature h of the context information mamba = F mamba (h).

[0061] Further, the initial image feature h 0 , the inter-modal context information h attn and the image feature h of the context information mamba are concatenated to form the fused feature h concat , and h concat is subjected to residual and upsampling processing to obtain the cross-modal fusion feature h that fuses multi-modal information 1 .

[0062] S6. Generate an image based on the cross-modal fusion feature to obtain the target image.

[0063] S7. Use the global semantic refinement module combined with the Mamba module to process the word features and the features of the target image to obtain the updated cross-modal fusion feature.

[0064] S8. Generate an image based on the updated cross-modal fusion feature to obtain the optimized image.

[0065] Embodiment 2

[0066] This embodiment provides a text-to-image system based on cross-modal information-guided fusion, as Figure 2 shown, including:

[0067] A text encoder for obtaining the sentence features and word features in the text description of the target image.

[0068] A sentence condition enhancement module that enhances the sentence features using a condition enhancement function and concatenates the enhanced sentence features with a noise vector to obtain image features.

[0069] An attention-driven deep text-image fusion module, as Figure 3 shown, for processing the sentence features, word features, and image features to obtain the initial fusion feature.

[0070] A first generator for generating an initial image based on the initial fusion feature.

[0071] The global semantic refinement module combined with Mamba, such as Figure 5 shown, is used to process the word features and the features of the image generated by the generator to generate cross-modal fusion features.

[0072] The second generator is used to generate the target image based on the cross-modal fusion features.

[0073] The third generator is used to generate the optimized image based on the updated cross-modal fusion features.

[0074] The discriminator is used to update the text-to-image generation system based on the initial image, the target image, and the optimized image.

[0075] The DAMSM module is used to update the text-to-image generation system based on the optimized image and the text description.

[0076] Furthermore, the attention-driven deep text-image fusion module includes 2 attention-driven conditional batch normalization layers (ACBN), 2 LeakyReLU activation function layers, and 1 convolutional layer with a convolutional kernel of 3×3 and a stride of 1.

[0077] Furthermore, as Figure 4 shown is the structural schematic diagram of the conditional batch normalization layer.

[0078] Furthermore, the loss function adopted by the discriminator to update the text-to-image generation system consists of the generator loss L G and the discriminator loss The expansion forms of L G and are as follows:

[0079]

[0080] where G i represents the i-th generator, L ca is the conditional enhancement loss, L DAMSM represents the DAMSM loss, which is used to improve semantic consistency. The parameters λ 1 and λ 2 are the conditional enhancement loss weight and the DAMSM loss weight respectively. D i represents the i-th discriminator, represents the unconditional loss function, which is used to judge the authenticity of the input image, represents the conditional loss function, which is used to evaluate the semantic consistency between the input image and the text description.

[0081] Furthermore, the expansion form of is as follows:

[0082]

[0083] Among them, y represents the image generated by the generator, denotes the data distribution of the generated image. The first term inside the parentheses on the right side of the equation is the unconditional loss, which is used to ensure that the generated image is as realistic as possible; the second term is the conditional loss, which is used to maintain the semantic consistency between the forced-generated image and the text description.

[0084] Furthermore, L ca has the following expansion form:

[0085] L ca = D KL (N(μ(e sent ), Σ(e sent )) || N(0, 1))

[0086] Among them, D KL (·) represents the Kullback-Leibler (KL) divergence, and μ(·) and Σ(·) represent the mean and variance functions respectively.

[0087] Furthermore, DAMSM is a text-image semantic similarity model. During training, the loss is constructed by calculating the matching score between the generated image and the text description to continuously optimize the text-to-image generation system, making the images generated by this system more in line with the text description.

[0088] Furthermore, the similarity matrix s is calculated and normalized to obtain

[0089]

[0090] Among them, e is the word feature, v is the image feature, and T represents the dimension of the word feature.

[0091] Furthermore, the regional context vector c related to the i-th word is calculated i :

[0092]

[0093] Among them, N is the image feature dimension, and γ 1 is the weight parameter.

[0094] Furthermore, the cosine similarity is used to calculate the correlation R(c i and the image e i : i , e i ):

[0095]

[0096] Furthermore, according to the correlation R(ci , e i ) Calculate the matching score between the generated image (Q) and the text description (D), which is defined as follows:

[0097]

[0098] where γ 2 is a factor that determines the importance of the relevance of image-text matching.

[0099] Furthermore, calculate the posterior probability of the text D i matching the image Q i as follows:

[0100]

[0101] where M represents the number of image-text pairs, and γ 3 is the smoothing factor. In this batch of text, only D i matches the image Q i , and the other M - 1 texts are regarded as unmatched descriptions. Therefore, the final loss function is defined as the negative log posterior probability of the image (Q) matching its corresponding text description (D) That is:

[0102]

[0103] Obtain the negative log posterior probability of the text description (D) matching its corresponding image (Q) That is:

[0104]

[0105] and The calculation method is the same as , where s represents the sentence feature. The expanded form of L DAMSM is as follows:

[0106]

[0107] and The expanded form is as follows:

[0108]

[0109] where y ∼ p data represents the data distribution of the real image, represents the data distribution of the generated image; the conditional augmentation loss enhances the training data and prevents overfitting by resampling the input sentence vectors from an independent Gaussian distribution.

[0110] Embodiment 3. In this embodiment, a computer device is proposed, which includes a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute the steps of the method for generating an image based on cross-modal information-guided fusion as described in Embodiment 1.

[0111] Embodiment 4. In this embodiment, a storage medium is proposed, which stores computer-readable instructions. It is characterized in that when the computer-readable instructions are executed by a processor, the steps of the method for generating an image based on cross-modal information-guided fusion as described in Embodiment 1 are implemented.

[0112] Finally, it should be noted that the above only describes some embodiments of the present invention. For those skilled in the art, various changes, modifications, substitutions, and deformations can be conceived without departing from the principles and spirits of the present invention. The protection scope of the present invention is defined by the appended claims and their equivalents, and the above actions should all be covered within the protection scope of the present invention.

Claims

1. A method for generating images from text based on cross-modal information guided fusion, characterized in that: include: S1, obtaining a text description of a target image, and processing the text description of the target image using a pre-trained text encoder to obtain sentence features and word features; the text encoder uses a bidirectional long short-term memory network text encoder; S2, conditionally enhance and concatenate the sentence features with noise to obtain image features; S3, using the attention-driven deep text-image fusion module to process sentence features, word features and image features to obtain initial fusion features; S4, generating an image according to the initial fusion features to obtain an initial image; S5, using the global semantic refinement module combined with the Mamba module to process the word features and the features of the initial image to obtain cross-modal fusion features; S6, generating an image according to the cross-modal fusion features to obtain a target image; S7, using the global semantic refinement module combined with the Mamba module to process the word features and the features of the target image to obtain updated cross-modal fusion features; S8. Generate an image according to the updated cross-modal fusion features to obtain an optimized image.

2. The text-to-image method based on cross-modal information guided fusion according to claim 1 is characterized in that: The attention-driven deep text-image fusion module processes data by: using the attention mechanism to calculate the contextual weight information between the modalities of word features and image features, and predicting the channel scaling parameters and offset parameters based on the sentence features; performing channel scaling and offset changes on the image features based on the contextual weight information, channel scaling parameters, and offset parameters to obtain an enhanced vector; performing two nonlinear operations, one convolution operation, and one upsampling operation on the enhanced vector in sequence to obtain the initial fused features; the nonlinear operation uses the LeakyReLU activation function.

3. The text-to-image method based on cross-modal information guided fusion according to claim 2 is characterized in that: The formula for calculating the contextual weight information between the modalities of word features and image features is: Among them, m attn represents context weight information, U represents the perceptron, N represents the dimension of image features, T represents the dimension of word features, represents the feature representation of the i-th word, x represents the image feature, represents the j-th column image feature.

4. The text-to-image method based on cross-modal information guided fusion according to claim 1, characterized in that: The data processing of the global semantic refinement module combined with the Mamba module includes: mapping word features to a common semantic space; calculating the similarity between the mapped word features and the features of the image generated by the generator, convolving the similarity with the word features to obtain inter-modal context information; the Mamba sequence model processes the features of the generated image to obtain intra-modal context information; the image features generated by the generator, the inter-modal context information, and the intra-modal context information are spliced, and multi-scale feature fusion and upsampling are performed to obtain cross-modal fusion features.

5. A text-to-image system based on cross-modal information guided fusion, applying the text-to-image method based on cross-modal information guided fusion according to any one of claims 1 to 4, characterized in that: include: A text encoder is used to obtain sentence features and word features in the text description of the target image; Sentence condition enhancement module, used to enhance sentence features to obtain image features; The attention-driven deep text-image fusion module is used to process sentence features, word features, and image features to obtain initial fusion features; A first generator, used to generate an initial image based on the initial fusion features; Combined with Mamba's global semantic refinement module, it is used to process word features and features of images generated by the generator to generate cross-modal fusion features; A second generator, used for generating a target image based on the cross-modal fusion features; A third generator, for generating an optimized image based on the updated cross-modal fusion features; The discriminator is used to update the text-generated image system based on the initial image, the target image, and the optimized image; DAMSM module, used to update the text-to-image system based on optimized images and text descriptions.

6. The text-to-image system based on cross-modal information guided fusion according to claim 5, characterized in that: The attention-driven deep text-image fusion module consists of 2 attention-driven conditional batch normalization layers, 2 LeakyReLU activation function layers, and 1 convolution layer with a convolution kernel of 3×3 and a stride of 1.

7. The text-to-image system based on cross-modal information guided fusion according to claim 5, characterized in that: The loss function used by the discriminator to update the text-to-image system is composed of the generator loss L G and the discriminator loss Composition, L G and The expanded form is as follows: L ca =D KL (N(µ(e sent ),Σ(e sent ))||N(0,1)) Among them, G i represents the i-th generator, L ca represents the conditional enhancement loss, L DAMSM represents DAMSM loss, which is used to improve semantic consistency. Parameters λ1 and λ2 represent the conditional enhancement loss weight and DAMSM loss weight, respectively. i represents the i-th discriminator, represents the unconditional loss function, represents the conditional loss function, y represents the image generated by the generator, e sent Indicates sentence features. represents the data distribution of generated images, It means unconditional loss. represents the conditional loss, D KL (·) represents the Kullback-Leibler (KL) divergence, μ(·) and Σ(·) represent the mean and variance functions, respectively, y~p data represents the data distribution of real images, Represents the data distribution for generating images.

8. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, characterized in that: When the computer-readable instructions are executed by the processor, the processor performs the steps of the method for generating images from text based on cross-modal information guided fusion as described in any one of claims 1 to 4.

9. A storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by a processor, the steps of the method for generating images from text based on cross-modal information guided fusion as described in any one of claims 1 to 4 are implemented.

Citation Information

Cited By

  • Picture generation and optimization method based on AI

    CN120599074A

  • Cross-modal image fusion method based on deep learning

    CN120852190A

  • A deep learning-based cross-modal image fusion method

    CN120852190B