Multi-condition guide text image generation method based on decoupling and multi-domain guide strategy
By decoupling the generation of appearance and structure, and using multi-domain guidance strategies, the problem that image generation in the prior art is difficult to ensure structure alignment and semantic consistency at the same time, and image generation that is flexible to adapt to various structural controls is achieved, with strong generalization and low operation overhead.
Patent Information
- Application Number
- CN202510160176.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art is difficult to generate multi-condition-guided image generation that ensures both structural alignment and semantic consistency, and the training-based model is inflexible and difficult to adapt to new conditions.
The text image generation method based on decoupling and multi-domain guidance is adopted to generate the image by decoupling appearance and structure formation, and to use multi-domain guidance in joint airspace and frequency domains to achieve better appearance consistency and structural alignment.
This method can adapt to arbitrary structural control without training, generate images aligned with text descriptions and structures, has powerful generalization, can complete a variety of downstream tasks, and reduce computing overhead.
Smart Images

Figure CN120047565A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image generation and computer vision, and in particular relates to a multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance, aiming to improve the semantic alignment and layout accuracy of combinatorial text-to-image generation. Background Art
[0002] In recent years, the emergence of diffusion models has promoted significant progress in image generation tasks. Subsequently, after training on a large amount of paired text-image data using a large amount of computing resources, large-scale diffusion models have been derived, which have achieved satisfactory results in the fields of text-to-image generation and image editing. However, text can only achieve the generation of basic content because it only provides coarse-grained information such as style and color, which is not sufficient to accurately describe the fine-grained spatial layout to achieve controllable text-to-image generation. Therefore, only using the previous large-scale diffusion models results in poor controllability.
[0003] Some works have tried to solve the above controllability defects. For example, ControlNet introduces a lightweight adapter outside the pre-trained text-to-image diffusion model to accept a series of input structure controls from users (such as edge maps, depth maps, and segmentation maps). However, ControlNet needs to train different adapters for different control conditions, which is very time-consuming and laborious. Based on this, some fine-tuning-based methods have tried to design a unified architecture to achieve multi-condition guided image generation. However, these methods overemphasize structure control and weaken the expression of text content, which will inevitably lead to semantic misalignment with text prompts. Because these methods simply use the hybrid features of text conditions and structure controls as the guiding conditions for image generation, resulting in pixel-level structure control dominating the generation process while ignoring the appearance expression in text prompts. In addition, these methods usually only focus on spatial domain features and ignore the crucial frequency domain representation, which leads to poor detail texture and structural consistency.
[0004] Recently, some methods have attempted to address the resource consumption issues of the above-mentioned training-based methods without training. Among them, Plug-and-Play (PnP) proposed a training-free controllable text-to-image generation method that adjusts attention features by injecting intermediate features. However, PnP only designed a single structural guidance branch to achieve structural alignment with the control signal, which inevitably leads to appearance leakage, that is, the appearance of the control signal penetrates into the generated image. Another recently proposed training-free method, called FreeControl, is a two-stage network that first models the linear subspace of intermediate features and then generates images through appearance and structural supervision in this subspace. Due to its non-end-to-end architecture, FreeControl has the complexity of modeling the subspace in the first stage, resulting in inflexibility and poor practical applications. This is mainly because a specific theme requires a corresponding subspace. However, this subspace needs to generate dozens of seed images for this object, which is very laborious and costly.
[0005] In summary, existing methods cannot generate multi-condition-guided image generation that simultaneously ensures structural alignment and semantic consistency. In addition, these training-based models are unable to adapt to new conditions due to inflexibility, resulting in difficulties in implementation. Based on this, the present method proposes a multi-condition-guided text-to-image generation method based on decoupling and multi-domain guidance, which can well adapt to various structural conditions without training and can simultaneously ensure high alignment with structural conditions and text, achieving excellent results compared with existing methods. In addition, the present method can be plugged into various generative large models to better complete common downstream tasks. Summary of the Invention
[0006] In view of the above problems, the present invention proposes a multi-condition-guided text-to-image generation method based on decoupling and multi-domain guidance, and includes the following steps:
[0007] 1. A multi-condition-guided text-to-image generation method based on a decoupling and multi-domain guidance strategy, characterized in that it can decouple the formation of appearance and structure and use multi-domain guidance that combines spatial and frequency domains to achieve better appearance consistency. It includes the following steps:
[0008] S1, construct a pre-trained StableDiffusion model as the generation branch and initialize a random noise. Input the user text and the random noise into the generation model for diffusion generation. Extract three components of the self-attention layer in the diffusion generation process: query, key, and value, and obtain the self-attention map in the generation process according to the query and the key;
[0009] S2. Construct the appearance guidance branch, load the pre-trained StableDiffusion model, and copy the randomly initialized noise in the generation model as the initial noise of the appearance guidance branch. Input it together with the user text into the appearance guidance branch for the diffusion process. Extract the values of the self-attention layers during the appearance guidance process;
[0010] S3. Construct the structure guidance branch, load the pre-trained StableDiffusion model. Obtain the noise related to the condition through the denoising diffusion implicit model (DDIM) inversion method for the input text, and input it into the StableDiffusion model. Extract the queries and keys of the self-attention layers during the structure guidance process to obtain the self-attention map of the structure guidance branch;
[0011] S4. Use the principal component extraction technique to extract the principal components of the self-attention maps and values during the generation process as the structural representation and appearance representation of the generation branch respectively. At the same time, use the principal component extraction technique to extract the principal component of the values of the appearance guidance branch as the appearance guidance representation, and use the principal component extraction technique to extract the principal component of the self-attention map of the structure guidance branch as the structure guidance representation. Extract the four-band features of the structural guidance representations of the generation branch and the structure guidance branch through wavelet transform, and add the features related to the high frequency in the four bands to obtain the frequency-domain structural representation;
[0012] S5. Calculate and sum the appearance supervision loss, the spatial-domain structural supervision loss, and the frequency-domain structural supervision loss. The total loss punishes and updates the latent features of the generation branch through gradient propagation for subsequent denoising until the iteration step is 0, and use the image decoder to restore it to the pixel space to obtain the finally generated image.
[0013] 2. A multi-condition guided text-to-image generation method based on a decoupled and multi-domain guidance strategy as claimed in claim 1, wherein the construction of the generation branch in step S1 includes the following steps:
[0014] S11. Initialize a low-dimensional random noise that follows a standard Gaussian distribution, which can be expressed as;
[0015]
[0016] S12. Encode the input text using the text encoder to obtain the mapped text embedding vector, which can be expressed as;
[0017]
[0018] where T is the input text, E(·) represents the encoding of the text encoder, is the text embedding vector.
[0019] S13. Construct a generation branch using the pre-trained StableDiffusion model, and extract the queries, keys, and values of the self-attention layer of the generation branch, which can be expressed as;
[0020]
[0021] where t is the iteration time step, Q t , V t , K t represent the query, key, and value respectively, ∈ θ is the generation branch, and z t is the hidden feature at time step t.
[0022] S13. Calculate the attention map of the generation branch based on the queries and keys obtained in step S13, which can be expressed as:
[0023]
[0024] where A t is the attention map, and d represents the dimension of Q and K, which is used to normalize the softmax value.
[0025] 3. A multi-condition guided text-to-image generation method based on a decoupled and multi-domain guidance strategy according to claim 1, wherein the step S2 is specifically as follows:
[0026] S21. Copy the initial random noise of the generation branch as the initial noise of the appearance guidance branch. This part can be expressed as;
[0027]
[0028] S22. Construct an appearance guidance branch and extract the values in the self-attention layer. This part can be expressed as;
[0029]
[0030] where t is the iteration time step, V t a represent the values of the appearance guidance branch respectively, ∈ a is the appearance guidance branch, is the hidden feature of the appearance guidance branch at time step t
[0031] 4. A multi-condition guided text-to-image generation method based on a decoupled and multi-domain guidance strategy according to claim 1, wherein the step S3 is specifically as follows:
[0032] S31. Use the DDIM inversion method for the structural control condition to obtain the initial noise of the structure guidance branch. This part can be expressed as;
[0033]
[0034] where DDIM(·) is the inversion process, and I c represents the input structural control, is the initial noise of the structure-guided branch.
[0035] S32. Extract the queries and keys of the self-attention blocks in the structure-guided branch, which can be expressed as:
[0036]
[0037] S33. Calculate the attention map of the structure-guided branch based on the queries and keys obtained in step S32, which can be expressed as;
[0038]
[0039] 5. A multi-condition-guided text-to-image generation method based on a decoupling and multi-domain guidance strategy according to claim 1, wherein the step S4 is specifically as follows:
[0040] S41. According to the values obtained in steps S13 and S22, use the principal component extraction technique to extract the appearance representations of the generation branch and the appearance-guided branch, which can be expressed as:
[0041]
[0042] where PCA(·) is the principal component analysis representation, A t and A t a respectively represent the appearance representations of the generation branch and the appearance-guided branch.
[0043] S42. According to the values obtained in steps S13 and S32, use the principal component extraction technique to extract the structure representations of the generation branch and the structure-guided branch, which can be expressed as:
[0044]
[0045] where PCA(·) is the principal component analysis representation, S t and S t s respectively represent the structure representations of the generation branch and the structure-guided branch.
[0046] S43. Use wavelet transform to extract the high-frequency structure representations from the structure representations of the generation branch and the structure-guided branch respectively, which can be expressed as:
[0047]
[0048] where DWT(·) is the wavelet transform, h t and h t s respectively represent the high-frequency appearance representations of the generation branch and the structure guidance branch.
[0049] 6. A multi-condition guided text-to-image generation method based on a decoupling and multi-domain guidance strategy according to claim 1, wherein the step S5 is specifically as follows:
[0050] S51, Calculate the appearance supervision loss, which can be expressed as:
[0051]
[0052] where N represents the number of principal component extractions.
[0053] S52, Calculate the spatial domain structure supervision loss, which can be expressed as:
[0054]
[0055] S53, Calculate the frequency domain structure supervision loss, which can be expressed as:
[0056]
[0057] S54, Calculate the sum of the losses of the three parts and update the hidden features of the generation branch, which can be expressed as:
[0058]
[0059] where α a , and are hyperparameters used to balance the three guidance branches. λ t The hyperparameter represents the guidance strength.
[0060] S55, The hidden features are used for subsequent generation. When the iteration step is 0, use the image decoder to map it from the low-dimensional hidden space to the pixel space to obtain an image. The generation process is guided by the decoupled appearance guidance branch and the structure guidance branch. The former uses principal component supervision for appearance representation, cleverly preserves the appearance in the input text, and effectively avoids weakening the expression of text content. The latter guides the generation process through multi-domain structure supervision to achieve better structural alignment with the control signal.
[0061] Beneficial effects: Compared with the prior art, the present invention provides a multi-condition guided text-to-image generation method based on a decoupling and multi-domain guidance strategy. This method can generate images that conform to the text description and are structurally aligned with the text according to any structural signal and text content, and produce the following beneficial effects:
[0062] 1. Do not require training to adapt to arbitrary structural control: Benefiting from the powerful multi-domain structural guidance strategy proposed by this method, which combines frequency-domain and spatial-domain supervision, it realizes strong guidance for the generation of the generative branch structure. By means of implicit structural representation, it enhances the structural consistency and high alignment between the generated image and the input structural control. This implicit guidance method enables this method to flexibly adapt to various types of structural control and generate images with diverse structures.
[0063] 2. Plug-and-play: This method does not require training and can be plugged into different types of generative models to achieve different generative effects. The cost of training large models is expensive. This method can greatly reduce the computational overhead, keep up with the replacement of the era's large models, and ensure excellent image generation effects.
[0064] 3. Strong generalization ability: This method can be generalized to complete many common downstream tasks, such as image editing, image deblurring, image coloring, and image restoration. For single-object image editing tasks, we can achieve accurate editing. For multi-object image editing tasks, we can achieve partial editing. For image coloring and image deblurring, we support the edge map as the structural control input and increase the penalty factors of the two guidance branches. In the image restoration task, we choose GLIGEN as our basic generative model and provide bounding box information to obtain good generative effects. Relying on the decoupled appearance guidance branch, structural guidance branch, and multi-domain structural supervision, this method has strong generalization ability and can complete many downstream tasks. Description of the Drawings
[0065] Figure 1 It is the flowchart of text-to-image generation guided by the structural control of the present invention.
[0066] Figure 2 It is the analysis of appearance representation and structural representation and the analysis of frequency-domain features of the present invention.
[0067] Figure 3 It is the generated result diagram of single-control single-object generation, single-control multi-object generation, and multi-control multi-object completed by the present invention.
[0068] Figure 4 It is the generated result diagram of plugging the present invention into three different versions of the StableDiffusion model with eight different structural control inputs.
[0069] Figure 5 It is the quantitative result diagram of generating images with four different structural control inputs on the ImageNet-R-TI2I dataset by plugging the present invention into three different versions of the StableDiffusion model.
[0070] Figure 6 andFigure 7 This is a graph showing the qualitative and quantitative comparison results between the present invention and other methods on the ImageNet-R-TI2I dataset.
[0071] Figure 8 This is a graph showing the quantitative comparison results between the present invention and other methods on the ImageNet-R-TI2I dataset under single condition and single evaluation metric.
[0072] Figure 9 and Figure 10 are respectively the ablation experiment result graph of the present invention for the number of principal components and the ablation experiment result graphs of three supervision functions.
[0073] Figure 11 and Figure 12 are respectively the qualitative and quantitative ablation experiment result graphs of the present invention for the selection of appearance representation.
[0074] Figure 13 The result graph of the present invention on downstream tasks. Detailed implementation manners
[0075] The multi-condition guided text-to-image generation method based on the decoupling and multi-domain guidance strategy of the present invention can accurately express the appearance information in the text while being highly aligned with the control condition in terms of structure by using the decoupled appearance and structure guidance branches and the multi-domain guidance strategy, thus avoiding the problem of overemphasizing the spatial structure while ignoring the text expression.
[0076] The invention will be further described below in conjunction with specific embodiments.
[0077] Embodiment 1:
[0078] As Figure 1 shown, the overall architecture of this method includes three parts: the generation process, the appearance guidance branch, and the structure guidance branch. The user inputs the text: a sculpture of a castle and a spatial guidance condition: an edge map of a castle. Three exactly the same parallel Stable Diffusion models are used to initialize the generation process, the appearance guidance branch, and the structure guidance branch. The text information is input into the generation process and the appearance guidance branch, and at the same time, a random noise that satisfies the Gaussian distribution is initialized as the initial noise for the generation process and the appearance guidance branch. Subsequently, after restoring the spatial control condition into the latent feature by means of DDIM proliferation, it is input into the structure guidance branch for the diffusion process.
[0079] During the denoising process of three Stable Diffusion models, we separately extract the query Q, key K, and value V of the first self-attention block of the decoder during the generation process, the value V of the appearance guidance branch, and the Q and key K of the structure guidance branch. Subsequently, we use principal component analysis to extract the first three principal components of these components as appearance representations and structure representations. Then, we use appearance supervision, spatial domain structure supervision, and frequency domain structure supervision to supervise the appearance representations and structure representations of the generation process and the guidance branches respectively. After calculating the loss, we update the latent features of the generation process by backpropagation of gradients. When the denoising iteration step approaches 0, the denoising is completed, and the latent features are restored to the pixel space using an image decoder to obtain an image.
[0080] Figure 2 This is the analysis of appearance representation and structure representation, as well as the analysis of frequency domain features, for the multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention. As Figure 2 shown, on the left side, we extract Q, K, V, and the attention map S from the self-attention layer of the U-Nets decoder every 200 iteration time steps, and use principal component analysis (PCA) to obtain the first three principal components. Subsequently, we map these principal components to pseudo-color images using pseudo-color technology. The same color regions in the pseudo-color images represent similar semantic information, while the brighter regions correspond to the regions with higher values in the original feature map. We find that the principal components of V express richer appearance information (the same color regions are roughly distributed on the surface of the castle), while the self-attention map S represents stronger structure information (the edge part of the castle in the self-attention map is the brightest). Inspired by the above findings, we select the principal components of V as the appearance representation for appearance supervision, and the principal components of S as the structure representation for structure representation to guide the generation process.
[0081] In addition, to study the relationship between frequency domain features and structure control, we use discrete wavelet transform (DWT) to extract the high-frequency features of the pseudo-color images. The visualization results are as Figure 2 shown on the right side. We find that the high-frequency features are roughly consistent with the control signals in structure, which helps to achieve more comprehensive structure supervision in the structure guidance branch. This key finding inspires this method to use basic frequency features as additional structure guidance to obtain better structural consistency.
[0082] Figure 3 These are the generation result graphs of single-control single-object generation, single-control multi-object generation, and multi-control multi-object generation completed by the multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention.
[0083] Figure 4The multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention is plug-and-play on three different versions of the StableDiffusion model, and eight different structures control the generated result images.
[0084] Figure 5 The quantitative result images of the multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention, which is plug-and-play on three different versions of the StableDiffusion model, for generating images with four different structures controlled on the ImageNet-R-TI2I dataset.
[0085] Figure 6 and Figure 7 The qualitative and quantitative comparison result images of the multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention and other methods on the ImageNet-R-TI2I dataset.
[0086] Figure 8 The quantitative comparison result images of the multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention and other methods on the ImageNet-R-TI2I dataset with single condition and single evaluation index.
[0087] Figure 9 and Figure 10 They are respectively the ablation experiment result images of the multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention for the number of principal components and the ablation experiment result images of three supervision functions.
[0088] Figure 11 and Figure 12 They are respectively the qualitative and quantitative ablation experiment result images of the multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention for the selection of appearance representation.
[0089] Figure 13 The result images of the multi-condition guided text-to-image generation method based on decoupling and multi-domain guidance strategies of the present invention on downstream tasks.
[0090] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0091] Although the specific implementation manners of the present invention are described above, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative labor on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A multi-condition guided text-to-image method based on decoupling and multi-domain guided strategies, characterized in that It is able to decouple the generation of appearance and structure, and achieve better appearance consistency by combining multi-domain guidance of spatial and frequency domains. It includes the following steps: S1, build a pre-trained stable diffusion model as the generation branch and initialize a random noise. Input the user text and random noise into the generation model for diffusion generation. Extract the three components of the self-attention layer in the diffusion generation process: query Q, key K and value V, and obtain the self-attention map in the generation process based on query Q and key K; S2, build the appearance guidance branch, load the pre-trained stable diffusion model, and copy the random noise initialized in the generative model as the initial noise of the appearance guidance branch, and input it into the appearance guidance branch together with the user text for diffusion process. Extract the value of the self-attention layer in the appearance guidance process; S3, build the structure guidance branch and load the pre-trained stable diffusion model. The input text is inverted by the denoising diffusion implicit model (DDIM) method to obtain the condition-related noise, which is input into the stable diffusion model. The query and key of the self-attention layer in the structure guidance process are extracted to obtain the self-attention map of the structure guidance branch; S4, using principal component extraction technology to extract the principal components of the self-attention map and value in the generation process, as the structure representation and appearance representation of the generation branch respectively. At the same time, using principal component extraction technology to extract the principal components of the appearance-guided branch value as the appearance-guided representation, using principal component extraction technology to extract the principal components of the self-attention map of the structure-guided branch as the structure-guided representation. The structure-guided representation of the generation branch and the structure-guided branch is extracted through wavelet transform to extract four frequency band features, and the high-frequency-related features in the four frequency bands are added to obtain the frequency domain structure representation; S5, calculate and sum the appearance supervision loss, spatial domain structure supervision loss and frequency domain structure supervision loss. The total loss penalizes and updates the hidden features of the generated branch through gradient propagation, and performs subsequent denoising until the iteration step is 0, and uses the image decoder to restore it to the pixel space to obtain the final generated image.
2. The method for generating images from texts with multi-condition guidance based on decoupling and multi-domain guidance strategy as claimed in claim 1, characterized in that: The construction generation branch in step S1 includes the following steps: S11, initialize a low-dimensional random noise that satisfies the standard Gaussian distribution. This part can be expressed as; S12, uses the text encoder to encode the input text to obtain the mapped text embedding vector, which can be expressed as; Where T is the input text, E(·) represents the text encoder encoding, is the embedding vector. S13, uses the pre-trained stable diffusion model to build the generation branch and extracts the query, key and value of the self-attention layer of the generation branch. This part can be expressed as; Where t is the iteration time step, Q t ,V t ,K t Represent query, key and value respectively, ∈ θ To generate a branch, z t is the latent feature at time step t. S13, based on the query and key obtained in step S13, the attention graph of the generated branch is calculated. This part can be expressed as: Among them A t is the attention map, d represents the dimension of Q and K, which is used to normalize the softmax value.
3. The method for generating images from texts with multi-condition guidance based on decoupling and multi-domain guidance strategy as claimed in claim 1, characterized in that: The step S2 is specifically as follows: S21, the initial random noise of the copy generation branch is used as the initial noise of the appearance guidance branch. This part can be expressed as; S22, builds the appearance guidance branch and extracts the value from the attention layer. This part can be expressed as; Where t is the iteration time step, V t a They represent the values of the appearance-guided branches, ∈ a Bootstrap branches for appearance, is the hidden feature of the appearance guidance branch at time step t.
4. The method for generating images from texts with multi-condition guidance based on decoupling and multi-domain guidance strategy as claimed in claim 1, characterized in that: The step S3 is specifically as follows: S31, the structure control condition is used to obtain the initial noise of the structure guidance branch using the DDIM inversion method, which can be expressed as; Where DDIM(·) is the inversion process, I c Represents the structure control of the input, Initial noise for structure-guided branching. S32, extracts the query and key of the self-attention block in the structure guidance branch, which can be expressed as: S33, based on the query and key obtained in step S32, the attention graph of the structure-guided branch is calculated, which can be expressed as; 5. The method for generating images from texts with multiple conditions based on decoupling and multi-domain guidance strategy as claimed in claim 1, characterized in that: The step S4 is specifically as follows: S41, according to the values obtained in steps S13 and S22, the appearance representation of the generative branch and the appearance guided branch is extracted using the principal component extraction technique, which can be expressed as: Where PCA(·) is the principal component analysis representation, A t and denote the appearance representations of the generative branch and the appearance-guided branch, respectively. S42, according to the values obtained in steps S13 and S32, the principal component extraction technique is used to extract the structural representation of the generating branch and the structural guiding branch, which can be expressed as: Where PCA(·) is the principal component analysis representation, S t and They represent the structural representations of the generative branch and the structure-guided branch respectively. S43, using wavelet transform to extract high-frequency structure representation from the structure representation of the generating branch and the structure guiding branch respectively, this part can be expressed as: Where DWT(·) is the wavelet transform, h t and The high-frequency appearance representations of the generative branch and the structure-guided branch, respectively.
6. The method for generating images from texts with multiple conditions based on decoupling and multi-domain guidance strategy as claimed in claim 1, characterized in that: The step S5 is specifically as follows: S51, calculate the appearance supervision loss, which can be expressed as: Where N represents the number of principal components extracted. S52, calculates the spatial domain structure supervision loss, which can be expressed as: S53, calculates the frequency domain structure supervision loss, which can be expressed as: S54, calculate the sum of the losses of the three parts and update the hidden features of the generated branch. This part can be expressed as: Where α a , and is a hyperparameter used to balance the three guided branches. t The hyperparameter represents the strength of guidance. S55, latent features are used for subsequent generation. When the iteration step is 0, the image decoder is used to map it from the low-dimensional latent space to the pixel space to obtain an image. The generation process is guided by the decoupled appearance guidance branch and structure guidance branch. The former uses the principal component to supervise the appearance representation, cleverly retains the appearance in the input text, and effectively avoids weakening the text content expression. The latter guides the generation process and the control signal to achieve better structural alignment through multi-domain structural supervision.
Citation Information
Patent Citations
Method for generating image containing expected identifier based on pre-trained text graph model, computer equipment, readable storage medium and program product
CN118520134A
Text-driven image translation method and system based on frequency spectrum reconstruction
CN118537209A
Organ and medical image segmentation result correction system based on DS-ASPP and CBEM
CN118840382A
Image stylization method based on diffusion model
CN119006310A
Hyperspectral image super-resolution method based on global guide condition diffusion model
CN119027317A
Cited By
Image generation system and method for multi-object space condition
CN120655767A
Diffusion generation security optimization method and device based on multi-modal region semantic alignment
CN121458831A