Text image generation method based on attention modulation and text redescription
By introducing attention modulation and text recapitulation mechanisms into the text-generating image method and combining the large language model to generate layout information, the problems of small objects lost and inaccurate generation positions in the existing methods are solved, high-quality, semantic consistent image generation is achieved, and computing resource requirements are reduced.
Patent Information
- Application Number
- CN202510160113.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-13
AI Technical Summary
Existing text image generation methods are difficult to achieve controllable generation based on layout, and there are often problems of small objects loss and inaccurate generation locations, and require a large number of data sets and computing resources.
Using the method based on attention modulation and text re-explanation, the image layout is automatically generated through a large language model, and the attention modulation module is used to adaptively adjust the attention map during the generation process to correct the incorrect generation area.
It realizes accurate generation of high-quality images with semantic consistency based on text and image layout, which reduces user usage threshold and greatly reduces computing consumption.
Smart Images

Figure CN119991855A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image generation and computer vision technology, and in particular relates to a text-generated image method based on attention modulation and text restatement, aiming to improve the semantic alignment capability and diversified structural expression of text-generated images. Background Art
[0002] In recent years, large-scale diffusion models have been used as the main tool for image generation to synthesize diverse and high-quality images because they outperform generative adversarial networks and autoregressive models and are stable to train. After 2020, major breakthroughs in large-scale diffusion models, especially the emergence of stable diffusion models, have greatly promoted the development of high-quality and semantically consistent text-to-image synthesis. However, text descriptions can only provide coarse-grained color and style information, and cannot give the specific layout of the image. Therefore, research on realizing a layout-based controllable text-to-image generation algorithm is a promising direction.
[0003] Some methods focus on improving the controllability of layout-based text generation images, hoping that objects in the text are generated within a specific layout. The main practice is to inject layout information into the pre-trained stable diffusion model as an additional condition to guide the image generation process. GLIGEN proposes a learnable gated self-attention layer to fuse layout information into the stable diffusion model to achieve controllable text generation images in the open domain. BoxDiff proposes a box-constrained diffusion model and designs box-in constraints, box-out constraints and corner constraints to ensure that objects are generated in the specified area. These two classic methods provide innovative ideas and correct guidance for subsequent research. However, these methods are indirect guidance methods, without directly adjusting the attention content, and often result in the loss of small objects and inaccurate generation positions. Based on this, the effective attention modulation mechanism we proposed can directly and adaptively adjust the attention map in the generation process, effectively alleviating the problem of small object loss.
[0004] Recently, large language models (LLMs) have shown excellent performance in cross-modal tasks due to their powerful context learning and layout reasoning capabilities. Therefore, we use large language models as layout generators, and guide the large language models to return specific object labels and bounding box information through designed text prompts. However, large language models can only provide layout content, but cannot effectively combine high-level spatial relationships of multiple object layout information. Based on this, we propose a text restatement method, which uses the difference between the restated text and the input text to update the latent space vector in the diffusion process, and corrects the incorrect generated content through multiple iterations to alleviate the inaccurate generation problem.
[0005] In summary, current methods for generating images from text can achieve basic generation tasks, but these methods require large datasets containing millions of images and powerful computing resources, which are very expensive to implement. In addition, under the condition of a given layout, they cannot accurately generate images consistent with the text prompts, and have problems such as attribute coupling, unreasonable spatial relationships, and missing objects. Although some of the latest methods can make some progress with additional training, these indirect methods cannot solve the above problems from the root. In short, current generative models cannot plan and generate image layouts based on text content, and accurately generate high-quality and diverse images based on layout conditions. Summary of the invention
[0006] In view of the above problems, the present invention proposes a method for generating images from text based on attention modulation and text restatement, and comprises the following steps:
[0007] 1. A method for generating images from text based on attention modulation and text restatement, characterized in that the attention map in the diffusion process can be adaptively adjusted, and the image layout can be automatically generated using a large language model to reduce the user threshold; and the incorrect generated area is corrected by a text restatement method, comprising the following steps:
[0008] S1, build a large language model prompt framework, and guide the large language model to automatically generate object labels and layout position information through text prompts. Encode the text content into the latent space through the pre-trained text encoder as a guide for the diffusion process in the generation process;
[0009] S2, build a stable diffusion model of latent space, which can receive text encoding and layout information and compress them into a low-dimensional space to achieve a high-efficiency, low-computation diffusion generation process;
[0010] S3, constructs an attention modulation module to modulate the cross-attention map in the diffusion process in S2 to enhance the content generation of the region of interest and alleviate the problem of small object loss;
[0011] S4, build a semantic text regeneration supervision network, use the image translation model to restate the image generated in S2 after adjustment by S3, and perform semantic supervision with the input text to update the hidden features of each iteration step in S2;
[0012] S5, restores the latent features of the last step in the iterative S2 into an image using the pre-trained image decoder.
[0013] 2. A method for generating images from text based on attention modulation and text restatement as claimed in claim 1, characterized in that the step of constructing a large language model prompt framework in step S1 comprises the following steps:
[0014] S11, designing a prompt word framework, which mainly includes: instructions that describe the task and specify the response format, scenario examples and user prompts that help improve the model's task understanding ability;
[0015] S12, input the text content input by the user into the prompt word framework in S11, and use the large language model to generate object labels and layout information. This part can be expressed as;
[0016] O,C=LLMs(P);O={o i |=1,2,...n},C={c i |=1,2,...n} (1)
[0017] 3. The method for generating images from text based on attention modulation and text restatement according to claim 1, wherein step S2 specifically comprises:
[0018] S21, first uses the pre-trained text encoder to encode the text input into a text vector to guide the generation diffusion process. This part can be expressed as;
[0019]
[0020] S22, initializes low-dimensional random noise with a dimension of 64×64, which satisfies the standard Gaussian distribution, and uses it as the starting point of the diffusion process.
[0021] S23, maps the label information obtained in S12 to an image with the same noise dimension and stores it in a database. This part can be expressed as:
[0022]
[0023] S24, builds a pre-trained stable diffusion model and receives the text vector encoded by S21. Taking the random pure Gaussian noise initialized in S22 as the starting point, the diffusion process begins and denoises gradually. This part can be expressed as:
[0024]
[0025] Where t=1...T-1, T, T is the total number of iterations. t represents the noise feature (hidden feature) of the tth step, e is the text vector. ε is the noise prediction network, α t It depends on the choice of mean coefficient.
[0026] 4. The method for generating images from text based on attention modulation and text restatement according to claim 1, wherein step S3 is specifically:
[0027] S31, capture the input feature F of the cross-attention layer at the t-th iteration step in the diffusion process, and construct three linear layers to obtain the cross-attention input amount of the t-th iteration step in the i-th cycle: query key value This part can be expressed as:
[0028]
[0029] Where W q ,W k ,W v They represent the weights of three different linear mapping layers. t is the number of iterations, and i represents the round of the loop, which is related to the convergence of the text restatement.
[0030] S32, get the cross attention map This part can be expressed as:
[0031]
[0032] Where softmax(·) represents the activation function, (·) T represents the transpose operation, and d represents and Dimension, used to normalize softmax(·).
[0033] S33, construct an attention modulation module, and adaptively modulate the cross attention map obtained in S32 according to the layout information obtained in S23. Through the direct modulation method of attention modulation, the content of interest can be generated in a fixed bounding box, while suppressing the area of no interest. And according to the proportion of the bounding box to the entire image area, large objects and small objects are divided, and the attention score of small objects is greatly enhanced to alleviate the problem of object disappearance due to the small attention score of small objects in the diffusion generation process. This part can be expressed as:
[0034]
[0035] M j =Mask(B j )⊙QK T ⊙(1-S j )-(1-Mask(B j ))⊙QK T ⊙(1-S j ) (10)
[0036] Where Mask represents the mask of the bounding box, (·) T represents the transpose operation, S j Indicates the ratio of the jth bounding box area to the entire image.
[0037] S34, use the image decoder to decode the hidden features with iteration step 0, restore from the low-dimensional hidden space to the pixel space, and obtain a high-definition image of the retelling round i. This part can be expressed as:
[0038]
[0039] 5. The method for generating images from text based on attention modulation and text restatement according to claim 1, wherein step S4 is specifically:
[0040] S41, builds a large image translation model and performs a text restatement of the generated image in the i-th round. This part can be expressed as:
[0041]
[0042] S42, using the same text encoder as S21, encodes the i-th round text restatement content, which can be expressed as:
[0043]
[0044] S43 performs high-level semantic supervision on the restated text vector and the initial text vector to indirectly and implicitly correct incorrect semantically generated content. This part can be expressed as:
[0045]
[0046] In the formula, cos(·) represents the cosine similarity calculation, and log represents the logarithmic operation, which can avoid the loss function being too large and causing too much fluctuation in the latent features.
[0047] S44, using high-level semantic supervision loss, updates the latent features of the tth iteration step in the i-th round of diffusion process, and executes a subsequent new round of diffusion process. This part can be expressed as:
[0048]
[0049] S45, using the updated latent features, performing subsequent denoising generation to finally generate an image.
[0050] Beneficial effects: Compared with the prior art, the present invention provides a method for generating images from text based on attention modulation and text restatement, which can accurately generate high-quality images with semantic consistency and close alignment with the image layout according to the text and image layout, and produces the following beneficial effects:
[0051] 1. Lowering the threshold for users: When the existing technology generates images based on layout information, it often requires users to provide fine-grained layout information, which is very difficult for users. The present invention avoids excessively high thresholds for users by designing prompt templates and using a large language model to automatically generate layout information based on text content.
[0052] 2. No need to train the model: Traditional training methods require a large amount of paired text-image data and a large amount of computing resources to train a diffusion model with huge parameters. The difficulty in obtaining paired text-image data, expensive computing resources and long training time have brought difficulties to the specific implementation of the model. This method designs an architecture that does not require training, uses the generation capability of the basic large model, and constructs an external supervision function to update the latent features in the stable diffusion process. The training method of this method greatly reduces memory requirements and resource consumption, and can generate an image within 7 seconds, with excellent real-time generation characteristics.
[0053] 3. User adaptive adjustment: Although existing methods can capture the defects of generated images, they cannot propose effective methods to correct the incorrect generated content. On the contrary, the text restatement stage proposed in this method can effectively alleviate this defect. Text restatement is particularly user-friendly. By designing the number of iterations, the error areas are constantly corrected, and the unreasonable generated content is adaptively adjusted. Finally, the method generates the generated results that users want. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 Generate a graphic flow chart for the inventive text.
[0055] Figure 2 The invention is an effect of generating text containing multiple objects with multiple attributes.
[0056] Figure 3 This is a comparison chart of the visualization results of the present invention and other methods.
[0057] Figure 4 It is a visualization diagram of the ablation experiment results of the main parts of the present invention.
[0058] Figure 5 This is a visualization result diagram of the attention map of the present invention before and after repetition rounds and attention modulation.
[0059] Figure 6 and Figure 7 This is a quantitative comparison result diagram of the controllable text generation image method based on attention modulation and text restatement of the present invention and other methods on four datasets: HRS, CC-500, NSR-1K and DrawBench.
[0060] Figure 8This is a comparison chart of the generation effects of using different large language models in the present invention.
[0061] Fig. 9 and Fig.10 They are respectively the prompt template of the present invention and several layout generation result diagrams. DETAILED DESCRIPTION
[0062] The controllable text-to-image generation method based on attention modulation and text restatement of the present invention uses a stable diffusion model and a large language model to enable the model to automatically generate layout information based on the text, and effectively combine the text content and layout information to generate high-quality images. At the same time, the training-free method can greatly reduce the computational consumption and significantly improve the practical application value of the method.
[0063] The invention will be further described below in conjunction with specific embodiments.
[0064] Embodiment 1:
[0065] like Figure 1 As shown in the figure, the image generation process is as follows: the user inputs text: a yellow dog and a blue chair. First, the user's input text is used to generate the image layout and main object labels through the large language model. Subsequently, the layout and label information are mapped to a layout representation of the same dimension as the random noise. After that, the pre-trained stable diffusion model is loaded and a low-dimensional random noise is initialized. The input text is mapped to a text vector using the pre-trained text encoder, and the text vector and random noise are simultaneously input into the pre-trained stable diffusion model for the diffusion generation process. An attention modulation module is constructed to extract the cross-attention map in the diffusion process, and the attention map and the layout representation of the same dimension are input into the attention modulation module for layout-based modulation. After completing all the iteration steps, the stable diffusion model will output an image. Finally, a semantic text regeneration module is constructed to translate the generated image into a restated text through the image translation model, and encode it into a restated text vector through the text encoder, supervise it with the input text vector, and propagate the loss back to the hidden features through the gradient, and perform gradient penalty to update the hidden features, and continue to complete the subsequent process until the loop stops.
[0066] The specific process of the attention modulation mechanism is as follows: decouple the image layout according to the object label, so that a single label corresponds to its corresponding object layout. Calculate the ratio of the object layout area corresponding to each label to the entire image area. Then, adaptively adjust the cross attention map according to the area ratio. Specifically, in order to emphasize the main objects in the text, we need to adjust the attention map to make it focus more on the aspects of interest. In addition, in order to avoid the problem of small object loss, we
[0067] The specific process of text restatement is as follows: after setting the number of loops, the image of the iteration of each round is generated by attention modulation in the diffusion process, and the iteration step tends to 0. The image is translated into the restatement text through the pre-trained image translation model, and it is encoded into a text vector using the text encoder. The loss of the restatement text vector and the input text vector in the high-dimensional semantic space is calculated, and the loss is updated through the gradient backpropagation method to update the hidden features in the diffusion process, and then continue to enter the next round of loops. When the loop ends or the loss is less than 0.01, the iteration stops and the generated image is input.
[0068] The present invention provides a controllable text-to-image generation method based on attention modulation and text restatement, which can enhance the alignment of the synthesized image with the given text and timely correct the erroneous area by restatement. This method enables the generative model to avoid problems such as attribute coupling, unreasonable spatial relationship expressions and missing objects when facing complex text descriptions. In addition, the image layout is automatically generated according to the text through a large language model, without the need for the user to provide complex layout information, which greatly reduces the user's usage threshold.
[0069] Figure 2 The controllable text-to-image generation method based on attention modulation and text restatement of the present invention is used to generate text containing multiple objects with multiple attributes. It can be seen that this method can accurately generate images that match the text and layout when faced with multiple attributes, multiple objects and complex spatial relationships.
[0070] Figure 3 This is a comparison chart of the visualization results of the controllable text-to-image generation method based on attention modulation and text restatement of the present invention and other methods. It can be found that the controllable text-to-image generation method based on attention modulation and text restatement has higher generation quality and stronger semantic alignment than other methods.
[0071] Figure 4 This is a visualization result diagram of the ablation experiment of the main parts of the controllable text generation image method based on attention modulation and text restatement of the present invention. It can be seen that the attention modulation module and text restatement module proposed in this method are very important.
[0072] Figure 5 This is a visualization result diagram of the attention map of the controllable text generation method based on attention modulation and text restatement before and after restatement iterations and attention modulation. It can be seen that the more restatement rounds, the better the correction effect on the wrong generation area, and the higher the alignment between the image and the text. At the same time, before and after attention modulation, the area of interest becomes larger, and the effect is better.
[0073] Figure 6 and Figure 7This is a quantitative comparison result diagram of the controllable text-to-image generation method based on attention modulation and text restatement of the present invention and other methods on four data sets: HRS, CC-500, NSR-1K and DrawBench. It can be found that the controllable text-to-image generation method based on attention modulation and text restatement has better results than other methods.
[0074] Figure 8 This is a comparison chart of the generation effects of different large language models in the controllable text generation image method based on attention modulation and text restatement of the present invention. It can be found that the continuously updated large language model has a stronger visual planning ability, which improves the performance of the present method.
[0075] Fig. 9 and Fig.10 They are respectively the prompt template and several layout generation result diagrams of the controllable text generation method based on attention modulation and text restatement of the present invention. It can be seen that through the carefully designed prompt template, the visual understanding and planning capabilities of the large language model can be fully mobilized, and the image layout can be reasonably generated.
[0076] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0077] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A text-to-image method based on attention modulation and text restatement, characterized in that The attention map in the diffusion process can be adaptively adjusted, and the image layout can be automatically generated using a large language model to reduce the user threshold; and the incorrect generated area can be corrected through a text restatement method, including the following steps: S1, build a large language model prompt framework, and guide the large language model to automatically generate object labels and layout position information through text prompts. Encode the text content into the latent space through the pre-trained text encoder as a guide for the diffusion process in the generation process; S2, build a stable diffusion model of latent space, which can receive text encoding and layout information and compress them into a low-dimensional space to achieve a high-efficiency, low-computation diffusion generation process; S3, constructs an attention modulation module to modulate the cross-attention map in the diffusion process in S2 to enhance the content generation of the region of interest and alleviate the problem of small object loss; S4, build a semantic text regeneration supervision network, use the image translation model to restate the image generated in S2 after adjustment by S3, and perform semantic supervision with the input text to update the hidden features of each iteration step in S2; S5, restores the latent features of the last step in the iterative S2 into an image using the pre-trained image decoder.
2. The method for generating images from text based on attention modulation and text restatement as claimed in claim 1, characterized in that: The construction of the large language model prompt framework in step S1 includes the following steps: S11, designing a prompt word framework, which mainly includes: instructions that describe the task and specify the response format, scenario examples and user prompts that help improve the model's task understanding ability; S12, input the text content input by the user into the prompt word framework in S11, and use the large language model to generate object labels and layout information. This part can be expressed as; O,C=LLMs(P);O={o i |=1,2,...n},C={C i |=1,2,...n} (1) 3. The method for generating images from text based on attention modulation and text restatement as claimed in claim 1, characterized in that: The step S2 is specifically as follows: S21, first uses the pre-trained text encoder to encode the text input into a text vector to guide the generation diffusion process. This part can be expressed as; S22, initializes low-dimensional random noise with a dimension of 64×64, which satisfies the standard Gaussian distribution, and uses it as the starting point of the diffusion process. S23, maps the label information obtained in S12 to an image with the same noise dimension and stores it in a database. This part can be expressed as: S24, builds a pre-trained stable diffusion model and receives the text vector encoded by S21. Taking the random pure Gaussian noise initialized in S22 as the starting point, the diffusion process begins and denoises gradually. This part can be expressed as: Where t=1...T-1, T, T is the total number of iterations. t represents the noise feature (hidden feature) of the tth step, e is the text vector. ε is the noise prediction network, α t It depends on the choice of mean coefficient.
4. The method for generating images from text based on attention modulation and text restatement as claimed in claim 1, characterized in that: The step S3 is specifically as follows: S31, capture the input feature F of the cross-attention layer at the t-th iteration step in the diffusion process, and construct three linear layers to obtain the cross-attention input amount of the t-th iteration step in the i-th cycle: query key Value V i t , this part can be expressed as: V i t =W v (F) (7) Where W q ,W k ,W v They represent the weights of three different linear mapping layers. t is the number of iterations, and i represents the round of the loop, which is related to the convergence of the text restatement. S32, get the cross attention map This part can be expressed as: Where softmax(·) represents the activation function, (·) T represents the transpose operation, and d represents and Dimension, used to normalize softmax(·). S33, construct an attention modulation module, and adaptively modulate the cross attention map obtained in S32 according to the layout information obtained in S23. Through the direct modulation method of attention modulation, the content of interest can be generated in a fixed bounding box, while suppressing the area of no interest. And according to the proportion of the bounding box to the entire image area, large objects and small objects are divided, and the attention score of small objects is greatly enhanced to alleviate the problem of object disappearance due to the small attention score of small objects in the diffusion generation process. This part can be expressed as: M j =Mask(B j )⊙QK T ⊙(1-S j )-(1-Mask(B j ))⊙QK T ⊙(1-S j ) (10) Where Mask represents the mask of the bounding box, (·) T represents the transpose operation, S j Indicates the ratio of the jth bounding box area to the entire image. S34, use the image decoder to decode the hidden features with iteration step 0, restore from the low-dimensional hidden space to the pixel space, and obtain a high-definition image of the retelling round i. This part can be expressed as:
5. The method for generating images from text based on attention modulation and text restatement as claimed in claim 1, characterized in that: The step S4 is specifically as follows: S41, builds a large image translation model and performs a text restatement of the generated image in the i-th round. This part can be expressed as: S42, using the same text encoder as S21, encodes the i-th round text restatement content, which can be expressed as: S43 performs high-level semantic supervision on the restated text vector and the initial text vector to indirectly and implicitly correct incorrect semantically generated content. This part can be expressed as: In the formula, cos(·) represents the cosine similarity calculation, and log represents the logarithmic operation, which can avoid the loss function being too large and causing too much fluctuation in the latent features. S44, using high-level semantic supervision loss, updates the latent features of the tth iteration step in the i-th round of diffusion process, and executes a subsequent new round of diffusion process. This part can be expressed as:
Citation Information
Patent Citations
Bidirectional text image generation method and system based on semantic consistency
CN113361250A
Method for generating image from text based on cross attention coding
CN115482302A
Non-training layout-to-image generation method based on diffusion model
CN119169146A
Cited By
Knitted product image generation method and device based on text adjustment and visual feedback
CN120219553A
Knitted product image generation method and device based on text adjustment and visual feedback
CN120219553B