Text guide image generation method and device, electronic equipment and storage medium

The method uses a large language model to decompose complex text prompts and employs a block cross-attention mechanism to improve the accuracy and coherence of image generation by addressing attribute confusion and layout inconsistencies in complex scenes.

CN120318348APending Publication Date: 2025-07-15BEIJING QDING INTERCONNECTION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510239408.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing text-guided image generation technology based on the depth diffusion model is prone to attribute confusion and layout inconsistency when dealing with complex scenarios with multiple concepts and multiple attributes, and it is difficult to generate images that meet users' expectations.

Method used

Through a large language model, the text prompt words as the basis of prompt words and sub-prompt words are decomposed, combined with the image generation network of block cross attention mechanism, identify and process the relationship between multiple subjects, and generate a reasonably laid out target image.

Benefits of technology

It effectively avoids attribute confusion and layout inconsistency, improves the accuracy and rationality of multi-concept image generation, and meets the image generation needs in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318348A_ABST
    Figure CN120318348A_ABST
Patent Text Reader

Abstract

The invention provides a text guide image generation method and device, electronic equipment and a storage medium. The method comprises the following steps: receiving a text cue word input by a user, wherein the text cue word comprises picture information of a target image; basic cue words are extracted from the text cue words, a plurality of sub cue words are generated according to the text cue words and the basic cue words, and the sub cue words are used for refining a main body and a background in the target image; determining a picture layout mode of the target image according to the spatial position information in the text cue word and the sub cue word; and inputting the text cue word, the sub cue word and the picture layout mode into a predetermined image generation model, identifying and processing the relationship among the plurality of main bodies by using the image generation model, and generating a target image according with the description of the text cue word. According to the method, confusion among multiple attributes and multiple subjects can be avoided, the layout reasonability among multiple concepts is improved, and the image generation requirement in a complex scene is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image generation technology, and in particular to a text-guided image generation method, device, electronic device and storage medium. Background Art

[0002] Text-guided image generation technology based on deep diffusion models has made significant progress in the field of image generation. By inputting text prompts, the model is able to generate images that match the description. However, traditional generation methods face many challenges when generating complex scenes, especially tasks involving multiple concepts, different attributes, and complex semantics.

[0003] The existing technology mainly has the following problems: First, the attribute confusion problem. When the text prompt contains multiple subjects with different attributes (for example, "a boy in a blue shirt" and "a girl in a red skirt"), the existing generation model often has difficulty distinguishing and accurately identifying the independent features of each subject, resulting in attribute confusion in the image generation process. Secondly, the multi-concept layout problem. In complex scenes, text prompts involve multiple concepts (such as multiple subjects and background elements). How to reasonably layout these elements so that they present a natural and logical spatial relationship in the image is a difficult problem in the design of the generation algorithm. The existing single end-to-end generation model is difficult to meet the generation requirements of complex scenes, especially when the relative positions and proportions between multiple elements are not fully considered, the layout of the generated image may not meet the user's expectations. Summary of the invention

[0004] In view of this, the embodiments of the present application provide a text-guided image generation method, device, electronic device and storage medium to solve the problems existing in the prior art, such as easy attribute confusion, inconsistent layout between multiple concepts, and inability to meet user expectations.

[0005] In a first aspect of an embodiment of the present application, a text-guided image generation method is provided, comprising: receiving a text prompt word input by a user, wherein the text prompt word includes picture information of a target image; extracting a basic prompt word from the text prompt word, and generating a plurality of sub-prompt words based on the text prompt word and the basic prompt word, wherein the sub-prompt words are used to refine the subject and background in the target image; determining a picture layout mode of the target image based on spatial position information in the text prompt word and the sub-prompt word; inputting the text prompt word, the sub-prompt word and the picture layout mode into a predetermined image generation model, using the image generation model to identify and process the relationship between a plurality of subjects, and generating a target image that conforms to the description of the text prompt word.

[0006] In a second aspect of the embodiments of the present application, there is provided a text-guided image generation device, including: a receiving module, configured to receive a text prompt input by a user, where the text prompt includes the picture information of the target image; an extraction module, configured to extract a basic prompt from the text prompt, and generate a plurality of sub-prompts according to the text prompt and the basic prompt, where the sub-prompts are used to refine the main body and background in the target image; a determination module, configured to determine the picture layout mode of the target image according to the spatial position information in the text prompt and the sub-prompts; a generation module, configured to input the text prompt, the sub-prompts and the picture layout mode into a predetermined image generation model, and use the image generation model to identify and process the relationships between multiple main bodies, and generate a target image that conforms to the description of the text prompt.

[0007] In a third aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0008] In a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0009] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects:

[0010] By receiving the text prompt input by the user, where the text prompt includes the picture information of the target image; extracting the basic prompt from the text prompt, and generating a plurality of sub-prompts according to the text prompt and the basic prompt, where the sub-prompts are used to refine the main body and background in the target image; determining the picture layout mode of the target image according to the spatial position information in the text prompt and the sub-prompts; inputting the text prompt, the sub-prompts and the picture layout mode into a predetermined image generation model, and using the image generation model to identify and process the relationships between multiple main bodies, and generating a target image that conforms to the description of the text prompt. The present application can avoid confusion between multiple attributes and multiple main bodies, improve the layout rationality between multiple concepts, and meet the image generation requirements in complex scenarios. Description of the Drawings

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0012] Figure 1It is a schematic flowchart of the text-guided image generation method provided by an embodiment of the present application;

[0013] Figure 2 It is a schematic flowchart of generating a target image by using an image generation model through the denoising process of global latent variables and block latent variables provided by an embodiment of the present application;

[0014] Figure 3 It is a schematic structural diagram of the text-guided image generation device provided by an embodiment of the present application;

[0015] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0016] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0017] The deep diffusion model (such as the diffusion model) is a generative model that has achieved remarkable breakthroughs in the field of image generation in recent years and can generate high-quality images based on text prompts. In this task, given a text description, the model will generate an image that matches the content of the text.

[0018] Existing text-guided image generation methods based on the deep diffusion model usually adopt an end-to-end generation method. That is to say, the model directly inputs the text prompt into the generation network, and generates an image after a series of processes. However, these traditional end-to-end models have certain limitations when dealing with complex text descriptions with multiple concepts, multiple attributes, and multiple subjects, which are specifically manifested in the following aspects:

[0019] Attribute confusion problem: When the text prompt contains multiple subjects with different attributes (such as "a boy wearing a blue top" and "a girl wearing a red skirt"), there may be confusion between these different attributes. The deep learning model needs to be able to accurately identify the independent features of each subject and avoid mixing the features of different subjects together.

[0020] Layout rationality between multiple concepts: The image may contain multiple concepts (such as multiple subjects or background elements), and the relative positions, sizes, angles, etc. of these elements all require reasonable layout. Existing methods may have an inconsistent layout when dealing with such complex scenarios, resulting in the generated image not meeting the user's expectations.

[0021] Therefore, the main technical problems faced by the text-guided image generation task based on the deep diffusion model include:

[0022] The multi-attribute and multi-subject problems in the text: How to identify and separate multiple different attributes in the text (such as "boy" and "girl") and their descriptive features. How to ensure that the model can understand the independence of these subjects and accurately generate images that conform to their descriptive features for each subject.

[0023] The problem of layout rationality: In the generated image, how to rationally layout multiple concepts so that the image content can not only accurately reflect the text prompt but also be visually coordinated. Especially in complex scenes, how to design the relative positions, sizes, angles, etc. between the subjects to meet the complex requirements of users for image generation.

[0024] In view of the problems existing in the prior art, the purpose of this application is to design an image generation algorithm guided by complex text, which has more accurate image generation ability in the target task compared with traditional algorithms. The input of this application is a text that describes the target image information to be generated, and the output is an image that conforms to the description information. The technical solution of this application mainly includes two parts: the generation of sub-prompts and the layout of the picture based on the large language model, and the image generation based on the block cross-attention model. The technical solution of this application mainly includes the following contents:

[0025] First, given a complex text prompt, use the large language model to restate it. Through restatement, the model decomposes the original complex prompt into a basic prompt and multiple sub-prompts, simplifying the understanding of the task. In this way, multiple subjects in the text and their features can be effectively identified, and independent descriptive information can be assigned to each subject.

[0026] Then, use the large language model to generate the picture layout according to the number and content of the sub-prompts. Specifically, the model will design elements such as their relative positions and proportions in the image according to the descriptions of different subjects in the text to ensure the rationality of the layout.

[0027] Finally, input the basic prompt, sub-prompts, and the picture layout into the image generation network based on the block cross-attention mechanism. This network can effectively identify and process the relationships between multiple subjects through the block cross-attention mechanism, so as to generate the target image that conforms to the original description and has a reasonable layout.

[0028] The innovation of the technical solution of this application lies in restating the text and generating the layout through the large language model, so as to consider the relationships between multiple attributes and subjects when generating images. At the same time, the image generation process is optimized through the block cross-attention mechanism to ensure that the multi-concept layout in the image is more reasonable and avoid the problems of attribute confusion and layout incoordination in traditional methods.

[0029] The content of the technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] Figure 1 is a schematic flowchart of the text-guided image generation method provided by an embodiment of this application. As Figure 1 shown, the text-guided image generation method may specifically include:

[0031] S101, receiving a text prompt input by a user, where the text prompt includes the picture information of the target image;

[0032] S102, extracting a basic prompt from the text prompt, and generating a plurality of sub-prompts according to the text prompt and the basic prompt, where the sub-prompts are used to refine the main body and background in the target image;

[0033] S103, determining the picture layout mode of the target image according to the spatial position information in the text prompt and the sub-prompts;

[0034] S104, inputting the text prompt, the sub-prompts, and the picture layout mode into a predetermined image generation model, and using the image generation model to identify and process the relationship between multiple main bodies, so as to generate a target image that conforms to the description of the text prompt.

[0035] In some embodiments, extracting a basic prompt from the text prompt includes:

[0036] processing the text prompt by using a predetermined large language model to analyze the semantic information in the text prompt;

[0037] identifying and extracting key element information for generating the target image from the text prompt, and using the key element information as the basic prompt, where the key element information includes the main body, background, and characteristic attributes of the target image.

[0038] Specifically, this embodiment details how to extract a basic prompt from the text prompt, use a large language model (such as LLaMa) to perform semantic analysis on the text, identify and extract key element information for generating the target image, and use this key information as the basic prompt to guide the subsequent image generation process.

[0039] In some examples, the user provides a text prompt to the system, describing the target image they hope to generate. For example, the text prompt input by the user is: "On the grass in the park, there is an orange cat standing on the left and a black dog squatting on the right." This text describes multiple elements in the target image, including the image background ("the grass in the park") and the positions, colors, and actions of two main bodies ("the orange cat" and "the black dog").

[0040] Process and analyze the input text prompt using a pre - determined large - language model (such as LLaMa). The large - language model will first perform semantic analysis to understand the main information and context relationships in the text. Specifically, LLaMa will identify the following key information from the text:

[0041] "An orange cat standing"

[0042] "A black dog squatting"

[0043] "The grass in the park"

[0044] Through semantic analysis, the model identifies these key elements and understands their importance in the image and the relationships between them. The key elements include the subjects (the cat and the dog), the background (the grass in the park), and the characteristic attributes such as the actions, colors, and positions of each subject.

[0045] After extracting the above - mentioned key information, the large - language model further processes it into basic prompts, which provide clear guidance for subsequent image generation. Based on the information identified above, the basic prompts can be decomposed into:

[0046] Subjects: Orange cat (standing), black dog (squatting)

[0047] Background: The grass in the park

[0048] Attributes: The colors, actions of the cat and the dog, and their relative positions (the cat is on the left, the dog is on the right)

[0049] Finally, the basic prompts extracted from the text (such as "An orange cat stands on the grass in the park, with blue sky and white clouds, and bright sunshine" and "A black dog squats on the grass in the park, with blue sky and white clouds, and bright sunshine") are used as the core information for generating images. These basic prompts will be passed as input to the subsequent image generation model to provide accurate and detailed descriptions for image generation.

[0050] Through the analysis and processing of the large - language model in this embodiment, the text prompt is accurately refined into basic prompts and effectively guides the key elements in the image generation process. This method effectively avoids attribute confusion and layout problems through in - depth understanding and decomposition of the text prompt, and can ensure the accuracy and rationality of the generated image, meeting the user's needs for complex image generation.

[0051] In some embodiments, multiple sub - prompts are generated based on the text prompt and the basic prompt, including:

[0052] Based on the text prompt and the basic prompt, use the large - language model to perform semantic analysis on the basic prompt;

[0053] According to the descriptions of each image element in the basic prompt, multiple sub-prompts are generated, where the sub-prompts are used to refine the main body, background, and characteristic attributes in the target image, so as to guide the image generation model to generate the target image containing the details of the image elements.

[0054] Specifically, this embodiment further details how to generate multiple sub-prompts based on the text prompt and the basic prompt, and perform semantic analysis on the basic prompt through a large language model (such as LLaMa) to generate refined sub-prompts to guide the image generation model to generate the target image with more details.

[0055] In some examples, assume that the user enters the following text prompt: "On the grass in the park, there is an orange cat standing on the left and a black dog squatting on the right." This prompt describes the background of the target image (the grass in the park) and the basic information of two main bodies (the orange cat and the black dog), including their positions, colors, and actions, etc. Based on this text prompt, through processing and analysis by a large language model (LLaMa), the basic prompt is extracted:

[0056] Main body 1: Orange cat (standing)

[0057] Main body 2: Black dog (squatting)

[0058] Background: Grass in the park

[0059] Spatial information: Orange cat on the left, black dog on the right

[0060] Furthermore, use the large language model to perform further semantic analysis on the basic prompt in order to refine the description of each image element. The large language model will generate multiple sub-prompts, which describe the main body and background in more detail, enabling the image generation model to better understand the specific performance of each element and ensuring that the generated image accurately reflects the details in the text description. Specific sub-prompts may include:

[0061] Sub-prompt 1: "An orange cat stands on the grass in the park, with blue sky and white clouds, and the sun is shining brightly."

[0062] Sub-prompt 2: "A black dog squats on the grass in the park, with blue sky and white clouds, and the sun is shining brightly."

[0063] Each sub-prompt not only refines the description of the main bodies (the cat and the dog), but also provides further information about the background elements for generating the image. For example, the descriptions of "blue sky and white clouds" and "the sun is shining brightly" help the model understand the environmental factors in the image, making the finally generated image more natural, with depth and a sense of hierarchy.

[0064] With these refined sub-prompts, the image generation model can render the features of each element more precisely when generating an image (such as the action, color, background environment, etc. of the subject), so as to generate the target image that meets the user's requirements.

[0065] In the image generation stage, the multiple generated sub-prompts will be passed as input to the image generation model. Through these sub-prompts, the model can accurately understand the details and spatial relationships of each element, so as to generate an image that conforms to the description of the text prompt. For example, in the generated image, the orange cat will accurately stand on the left side of the grass in the park, while the black dog will squat on the right side, and the background will show blue sky, white clouds and sunny weather.

[0066] By refining the basic prompt to generate multiple sub-prompts, this embodiment improves the accuracy and detail richness of image generation. Each sub-prompt provides a more specific and clear description for the image generation model, helping the generation model better understand the attributes and spatial relationships of image elements, and then generating the target image that meets the user's expectations, avoiding the inaccurate generation problems caused by unclear descriptions or insufficient details in traditional methods.

[0067] In some embodiments, determining the layout method of the target image according to the spatial position information in the text prompt and sub-prompts includes:

[0068] Analyze the spatial position information in the text prompt and sub-prompts, and extract the relative position relationships between various image elements;

[0069] Determine the layout method of the image elements in the target image according to the relative position relationships, where the layout methods include left-right layout, up-down layout, and four-grid layout;

[0070] According to the layout method, guide the image generation model to generate the target image that conforms to the spatial relationship corresponding to the spatial position information according to the position, proportion and arrangement relationships of the image elements.

[0071] Specifically, this embodiment details how to determine the layout method of the target image according to the spatial position information in the text prompt and sub-prompts to ensure that the generated image conforms to the description input by the user, especially in the case where multiple image elements (such as the subject and the background) need to be reasonably laid out.

[0072] In some examples, assume that the text prompt input by the user is: "On the grass in the park, there is an orange cat standing on the left and a black dog squatting on the right." The system will use a large language model (such as LLaMa) to process it and extract the following sub-prompts:

[0073] Sub - prompt 1: "An orange cat stands on the grass in the park, with blue sky, white clouds and bright sunshine."

[0074] Sub - prompt 2: "A black dog squats on the grass in the park, with blue sky, white clouds and bright sunshine."

[0075] The actions, colors, positions and background information of each subject have been clearly described in these sub - prompts.

[0076] Furthermore, the system then analyzes the spatial position information in the text prompt and sub - prompts. For example, in the text prompt, "left" and "right" are mentioned. The system will recognize that the orange cat should be located on the left side of the picture, while the black dog should be located on the right side of the picture. These spatial position information helps to determine the relative position relationship of the elements in the image.

[0077] In some examples, the analysis results are as follows:

[0078] Subject 1 (orange cat): The position is on the left side of the image, described as "standing on the grass in the park".

[0079] Subject 2 (black dog): The position is on the right side of the image, described as "squatting on the grass in the park".

[0080] Background: On the grass in the park, with blue sky, white clouds and bright sunshine.

[0081] Furthermore, based on the analysis of the spatial position information, the system will determine the final image layout according to the preset picture layout method. In this embodiment, since the text input by the user clearly mentions the two positions of "left" and "right", the system automatically determines to adopt a left - right layout.

[0082] In some examples, the preset layout methods can include:

[0083] Left - right layout: Suitable for situations where the description contains clear left - right relationships.

[0084] Up - down layout: Suitable for situations where the description contains clear up - down relationships.

[0085] Four - grid layout: Suitable for situations where multiple regions need to be included in the image and evenly distributed.

[0086] According to the user's prompt, the system determines the left - right layout and provides clear guidance for subsequent image generation based on this.

[0087] After determining the layout method, the system inputs this spatial position information into the image generation model. The image generation model guides the position, proportion and arrangement relationship of the elements in the image according to the layout method, ensuring that the generated image conforms to the spatial relationship corresponding to the spatial position information.

[0088] For example, in some examples, according to the left - right layout, the orange cat is on the left side of the image, and the black dog is on the right side of the image. The system adjusts the elements in the image so that the positional relationship between the cat and the dog meets the expectation, while ensuring that the spacing, size, and proportion between them are appropriate, forming a natural layout. In this way, the generated image will accurately reflect the user's description, with the orange cat on the left and the black dog on the right, and the background presenting a scene of park grassland, blue sky, white clouds, and bright sunshine.

[0089] Through the process of analyzing and determining the screen layout method based on spatial position information in this embodiment, the system can effectively convert the text description input by the user into a spatial layout that can be understood and accurately presented by the image generation model. This method can not only handle complex image layout requirements but also ensure that the relative positions, proportions, and arrangements of the generated image elements are reasonable, meeting the user's expectations.

[0090] In some embodiments, an image generation model is used to identify and process the relationships between multiple subjects to generate a target image that conforms to the description of the text prompt, including:

[0091] In the image generation model, through the denoising process of global latent variables and block - level latent variables, a target image that conforms to the description of the text prompt is gradually generated.

[0092] Specifically, the extracted basic prompt and sub - prompts are input into a predetermined image generation model (such as SDXL). The image generation model will use these prompts to guide the image generation process. In particular, the model will gradually generate the image by processing global latent variables and block - level latent variables.

[0093] In the image generation model, first, a rough model of the image is built through global latent variables. Global latent variables contain the overall information of the image, such as the main body of the image, the basic features of the background, and their relative positions. The system performs denoising processing on the global latent variables according to the prompts provided by the user, thereby generating a preliminary image that conforms to the user's description. The specific steps are as follows:

[0094] Starting from the global latent variables, combined with the information in the text prompt (such as "orange cat" and "black dog"), denoising processing begins.

[0095] The model gradually improves the global latent variables through a series of denoising steps (such as a 20 - step denoising process) to obtain a preliminary image structure.

[0096] To generate the details of each subject and its surroundings more precisely, the image generation model then divides the global latent variables according to the spatial layout of the image to obtain block - level latent variables. Each block - level latent variable is responsible for generating the details of a certain part of the image, such as the areas where the "orange cat" and "black dog" are located.

[0097] Further, the processed chunk latent variables are reassembled together to generate a complete latent variable for subsequent image generation.

[0098] After the above steps, the image generation model will continue with denoising processing, gradually optimizing the latent variable until the final image that matches the user's description is generated. The final image will accurately display all the elements in the text prompt: the orange cat on the left, the black dog on the right, and the park grass, blue sky, white clouds, and sunny weather in the background.

[0099] In this embodiment, through the denoising process of the global latent variable and chunk latent variables, the relationships between multiple subjects are effectively identified and processed. During the image generation process, the model can not only correctly identify each element but also reasonably layout their positions and refine the specific manifestations of each element. The finally generated image can accurately reflect the user's text description, avoiding problems such as attribute confusion and unreasonable layout, and meeting the user's requirements for complex images.

[0100] In some embodiments, through the denoising process of the global latent variable and chunk latent variables in the image generation model, a target image that conforms to the description of the text prompt is gradually generated, including:

[0101] According to the predetermined global prompt and latent variable, use the image generation model to perform denoising processing on the latent variable to obtain the global latent variable;

[0102] Cut the global latent variable into multiple latent variable chunks according to the screen layout method, and input the latent variable chunks and the corresponding sub-prompts into the image generation model to generate new latent variable chunks;

[0103] Stitch the new latent variable chunks according to the original layout positions to obtain the stitched latent variable chunks;

[0104] Perform weighted summation on the stitched latent variable chunks and the global latent variable according to a preset weight to obtain a new latent variable;

[0105] Use the image generation model to perform denoising processing on the new latent variable. After the denoising processing reaches the preset number of rounds, generate the target image that conforms to the text prompt.

[0106] Specifically, this embodiment details how to gradually generate a target image that conforms to the description of the text prompt through the denoising process of the global latent variable and chunk latent variables. In the embodiment, the system uses an image generation model (such as SDXL) to perform denoising on the latent variable, and cuts and stitches the latent variable according to the screen layout, and finally generates the target image that matches the user's description.

[0107] Figure 21 is a flow chart of generating a target image by using an image generation model through a denoising process of global latent variables and block latent variables provided by an embodiment of the present application. Figure 2 As shown, the method may specifically include:

[0108] Suppose the text prompt word entered by the user is: "On the grass in the park, there is an orange cat standing on the left and a black dog squatting on the right." This text describes multiple elements in the image (two subjects and the background) and their relative positions and features. The system first extracts the basic prompt words and sub-prompt words through a large language model (such as LLaMa):

[0109] Basic prompt words: orange cat (standing), black dog (crouching), park grass (background).

[0110] Sub-prompt 1: "An orange cat is standing on the grass in the park. The sky is blue and the sun is shining."

[0111] Sub-prompt 2: "A black dog is squatting on the grass in the park. The sky is blue and the sun is shining."

[0112] Furthermore, in the image generation model, the global cue words and latent variables are first input into the model for denoising. The global latent variable contains the high-level structure and overall information of the entire image (such as the basic features of the subject, the background environment, etc.). The model gradually optimizes the global latent variable through denoising to obtain the global latent variable _ti, which provides a preliminary image framework for subsequent processing. This process is carried out through multiple rounds of iterations, for example, 20 steps of denoising are performed to gradually refine the image content.

[0113] Next, the system divides the global latent variable into multiple latent variable blocks according to the predetermined layout of the image (such as left-right layout, top-bottom layout, etc.). For example, in some examples, assuming a left-right layout is used, the global latent variable will be divided into two blocks:

[0114] Latent variable _t_ block 1: corresponds to the area where the orange cat is located.

[0115] Latent variable _t_ block 2: corresponds to the area where the black dog is located.

[0116] Then, the system inputs latent variable_t_block1 and sub-prompt word 1 (such as "an orange cat is standing on the grass in the park, with blue sky and white clouds and bright sunshine") into the image generation model, and after denoising, generates latent variable_t-1_block1. Similarly, latent variable_t_block2 and sub-prompt word 2 (such as "a black dog is squatting on the grass in the park, with blue sky and white clouds and bright sunshine") are input into the image generation model to generate latent variable_t-1_block2.

[0117] Further, after obtaining the two denoised latent variable sub - latent variables latent_variable_t - 1_chunk1 and latent_variable_t - 1_chunk2, the system stitches them according to a preset layout (such as a left - right layout) to form a new latent variable latent_variable_t - 1_chunk. This step ensures that the relative positions and proportional relationships of the various elements in the image conform to the user's description.

[0118] Next, the system performs a weighted sum of the stitched latent variable latent_variable_t - 1_chunk and the global latent variable (the denoised latent variable t - 1). Specifically, the weight of the latent variable latent_variable_t - 1_chunk is 0.7, and the weight of the global latent variable latent_variable_t - 1 is 0.3. After the weighted sum, a new latent variable t - 1 is obtained, which contains refined image information and the overall structure.

[0119] Further, after completing the weighted sum of the latent variables, the image generation model continues to perform further denoising on the new latent variable t - 1 until a preset number of denoising rounds (such as 20 rounds) is reached. After these denoising steps, the target image that conforms to the description of the text prompt is finally generated. This image will contain an orange cat standing on the left side of the park grass, a black dog squatting on the right side, and a sunny environment with blue sky and white clouds in the background.

[0120] In this embodiment, through the denoising process of the global latent variable and the sub - latent variables, the description in the text prompt is effectively transformed into a detailed image generation process. The denoising and weighted sum of each sub - latent variable ensure that the relationships and positions between the elements in the image are accurately reproduced. The finally generated image not only meets the user's requirements for each element but also reasonably presents their spatial layout, avoiding the problems of attribute confusion or inconsistent layout that may occur in traditional methods.

[0121] In some embodiments, the image generation model is a deep - learning model trained using an image generation network based on a chunk - based cross - attention mechanism.

[0122] The following is an embodiment of the apparatus of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the method embodiment of the present application.

[0123] Figure 3 It is a schematic structural diagram of a text - guided image generation apparatus provided by an embodiment of the present application. As Figure 3 shown, the text - guided image generation apparatus includes:

[0124] A receiving module 301, configured to receive a text prompt input by a user, where the text prompt includes the scene information of the target image;

[0125] An extraction module 302 is configured to extract a basic prompt from a text prompt, and generate a plurality of sub-prompts based on the text prompt and the basic prompt, where the sub-prompts are used to refine the main body and background in the target image;

[0126] A determination module 303 is configured to determine the layout mode of the target image according to the spatial position information in the text prompt and the sub-prompts;

[0127] A generation module 304 is configured to input the text prompt, the sub-prompts, and the layout mode into a predetermined image generation model, and use the image generation model to identify and process the relationships between multiple main bodies, and generate a target image that conforms to the description of the text prompt.

[0128] In some embodiments, Figure 3 the extraction module 302 processes the text prompt using a predetermined large language model, analyzes the semantic information in the text prompt; identifies and extracts key element information for generating the target image from the text prompt, and uses the key element information as the basic prompt, where the key element information includes the main body, background, and feature attributes of the target image.

[0129] In some embodiments, Figure 3 the extraction module 302 performs semantic analysis on the basic prompt using the large language model according to the text prompt and the basic prompt; generates a plurality of sub-prompts according to the descriptions of each image element in the basic prompt, where the sub-prompts are used to refine the main body, background, and feature attributes in the target image, so as to guide the image generation model to generate a target image including the details of the image elements.

[0130] In some embodiments, Figure 3 the determination module 303 analyzes the spatial position information in the text prompt and the sub-prompts, extracts the relative position relationships between the respective image elements; determines the layout mode of the image elements in the target image according to the relative position relationships, where the layout mode includes a left-right layout, an up-down layout, and a four-grid layout; according to the layout mode, guides the image generation model to generate a target image that conforms to the spatial relationship corresponding to the spatial position information according to the positions, proportions, and arrangement relationships of the image elements.

[0131] In some embodiments, Figure 3 the generation module 304 gradually generates a target image that conforms to the description of the text prompt through the denoising process of the global latent variable and the block latent variable in the image generation model.

[0132] In some embodiments, Figure 3The generation module 304 denoises the latent variables using an image generation model based on a predetermined global prompt and latent variables to obtain global latent variables; divides the global latent variables into multiple latent variable blocks according to the screen layout method, and inputs the latent variable blocks and corresponding sub-prompts into the image generation model to generate new latent variable blocks; splices the new latent variable blocks according to the original layout positions to obtain spliced latent variable blocks; performs weighted summation on the spliced latent variable blocks and the global latent variables according to a preset weight to obtain new latent variables; denoises the new latent variables using the image generation model, and after the denoising process reaches a preset number of rounds, generates a target image that conforms to the text prompt.

[0133] In some embodiments, the image generation model is a deep learning model trained using an image generation network based on a block cross-attention mechanism.

[0134] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0135] Figure 4 is a schematic structural diagram of the electronic device 4 provided by the embodiments of the present application. As Figure 4 shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above various method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above various device embodiments are implemented.

[0136] Exemplarily, the computer program 403 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 402 and executed by the processor 401 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 403 in the electronic device 4.

[0137] The electronic device 4 can be a desktop computer, a notebook, a palm computer, a cloud server, and other electronic devices. The electronic device 4 can include but is not limited to the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely an example of the electronic device 4, and does not constitute a limitation to the electronic device 4. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.

[0138] The processor 401 can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0139] The memory 402 can be an internal storage unit of the electronic device 4. For example, the hard disk or memory of the electronic device 4. The memory 402 can also be an external storage device of the electronic device 4, such as a plug-in hard disk equipped on the electronic device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 402 can also include both the internal storage unit and the external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device. The memory 402 can also be used to temporarily store data that has been output or will be output.

[0140] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0141] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0142] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0143] In the embodiments provided in this application, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0144] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0145] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0146] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0147] The above embodiments are only used to illustrate the technical solutions of this application, rather than limiting them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A text-guided image generation method, characterized in that, Including: Receiving a text prompt input by a user, where the text prompt includes the picture information of the target image; Extracting a basic prompt from the text prompt, and generating multiple sub-prompts according to the text prompt and the basic prompt, where the sub-prompts are used to refine the main body and background in the target image; Determining the picture layout mode of the target image according to the spatial position information in the text prompt and the sub-prompts; Inputting the text prompt, the sub-prompts, and the picture layout mode into a predetermined image generation model, and using the image generation model to identify and process the relationships between multiple main bodies, and generating a target image that conforms to the description of the text prompt.

2. The method according to claim 1, characterized in that, The extracting a basic prompt from the text prompt includes: Processing the text prompt by using a predetermined large language model to analyze the semantic information in the text prompt; Identifying and extracting key element information for generating the target image from the text prompt, and using the key element information as the basic prompt, where the key element information includes the main body, background, and characteristic attributes of the target image.

3. The method according to claim 2, wherein The generating multiple sub-prompts according to the text prompt and the basic prompt includes: Performing semantic analysis on the basic prompt by using a large language model according to the text prompt and the basic prompt; Generating multiple sub-prompts according to the descriptions of each image element in the basic prompt, where the sub-prompts are used to refine the main body, background, and characteristic attributes in the target image, so as to guide the image generation model to generate a target image containing the details of the image elements.

4. The method according to claim 1, wherein The determining the picture layout mode of the target image according to the spatial position information in the text prompt and the sub-prompts includes: Analyzing the spatial position information in the text prompt and the sub-prompts, and extracting the relative position relationship between each image element; Determining the picture layout mode of the image elements in the target image according to the relative position relationship, where the picture layout mode includes left-right layout, up-down layout, and four-grid layout; According to the picture layout mode, guiding the image generation model to generate a target image that conforms to the spatial relationship corresponding to the spatial position information according to the position, proportion, and arrangement relationship of the image elements.

5. The method according to claim 1, characterized in that The using the image generation model to identify and process the relationships between multiple main bodies and generating a target image that conforms to the description of the text prompt includes: Gradually generating a target image that conforms to the description of the text prompt through the denoising process of global latent variables and block latent variables in the image generation model.

6. The method according to claim 5, characterized in that, The gradually generating a target image that conforms to the description of the text prompt through the denoising process of global latent variables and block latent variables in the image generation model includes: According to a predetermined global prompt and latent variables, using the image generation model to perform denoising processing on the latent variables to obtain global latent variables; Split the global latent variable into multiple latent variable blocks according to the described screen layout method, and input the latent variable blocks and corresponding sub-prompts into the image generation model to generate new latent variable blocks; Stitch the new latent variable blocks according to the original layout positions to obtain the stitched latent variable blocks; Perform weighted summation on the stitched latent variable blocks and the global latent variable according to a preset weight to obtain a new latent variable; Use the image generation model to perform denoising processing on the new latent variable. After the denoising processing reaches the preset number of rounds, generate a target image that conforms to the text prompt.

7. The method according to any one of claims 1 to 6, characterized in that, The image generation model is a deep learning model trained by an image generation network based on a block cross-attention mechanism.

8. A text-guided image generation device, characterized in that, It includes: A receiving module for receiving a text prompt input by a user, where the text prompt includes screen information of a target image; An extraction module for extracting a basic prompt from the text prompt and generating multiple sub-prompts according to the text prompt and the basic prompt, where the sub-prompts are used to refine the main body and background in the target image; A determination module for determining the screen layout method of the target image according to the spatial position information in the text prompt and the sub-prompts; A generation module for inputting the text prompt, the sub-prompts, and the screen layout method into a predetermined image generation model, and using the image generation model to identify and process the relationships between multiple main bodies, and generating a target image that conforms to the description of the text prompt.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Image generation method and device, equipment, storage medium and product

    CN120997628A

  • Image generation method and device, electronic equipment and storage medium

    CN122176079A