Multi-subject semantic attribute binding image generation method based on diffusion model

By optimizing the input and cross-attention mechanism of the diffusion model and combining layout and reference images, the semantic alignment problem of the diffusion model in multi-object, multi-attribute scenarios is solved, achieving efficient semantic attribute binding and spatial control, and improving the accuracy and consistency of generated images.

CN121685718APending Publication Date: 2026-03-17FOSHAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511669797.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-03-17

Smart Images

  • Figure CN121685718A_ABST
    Figure CN121685718A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-subject semantic attribute binding image generation method based on a diffusion model. The method comprises the steps of improving independent coding of a single sentence into segmentation of the sentence into a plurality of sub-sentences and independent coding of the sub-sentences; improving a denoising process and improving a cross attention mechanism; clear semantic representation is obtained by independently coding clauses, a denoising process is divided into a reconstruction branch and a generation branch, intermediate potential representation is spliced and fused to improve consistency and visual quality, and meanwhile, independent attention guidance is realized by combining an improved cross attention mechanism and mask division, so that feature aliasing is avoided, and the accuracy and the reliability of the system are improved. And the local precision and the space decoupling capability are enhanced. According to the multi-entity semantic modeling and image-text alignment method, a finer and more stable solution is provided while high efficiency and universality are ensured, and the potential in multi-entity semantic modeling and image-text alignment is shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of image generation, and in particular to a multi-subject semantic attribute binding image generation method based on a diffusion model. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence generative technology, deep generative models have played an increasingly important role in visual tasks such as image generation, editing, restoration, and video generation, driving innovation in digital creation methods. Among them, diffusion models, due to their stepwise denoising mechanism inspired by non-equilibrium thermodynamics, have demonstrated excellent generation quality and stability, and have become one of the most promising generative technologies today.

[0003] However, diffusion models still have shortcomings when dealing with complex text-to-image generation tasks involving multiple objects and attributes. These shortcomings mainly manifest as semantic interleaving, attribute mismatch, and layout chaos, making it difficult for the model to accurately reproduce text semantics and resulting in semantic alignment failures. This limits its application effectiveness in complex scenarios. Typical failure cases include attribute chaos, semantic omission, and object mixing.

[0004] To enhance the semantic representation and attribute binding capabilities of diffusion models in complex text-to-image generation tasks, researchers have proposed semantic attribute binding methods, which can be mainly divided into three categories: 1) Attention mechanism-based optimization methods: By adjusting the latent representation in the denoising process, we can ensure that the text prompts correspond one-to-one with the image activation regions, thereby reducing semantic omissions; or we can combine syntactic analysis and positive and negative loss to optimize the cross-attention map to improve the accuracy of semantic separation and attribute binding.

[0005] 2) Semantic structure parsing-based methods: Some scholars use language parsers to extract the hierarchical structure of text and combine cross-attention layers to manipulate key-value pairs to enhance the expression of attribute combinations; Alternatively, attribute binding ability can be improved without additional training through vector decoupling and embedding optimization.

[0006] 3) Spatial constraint-based methods: transform spatial information into various layout constraints embedded in the diffusion process to control the position and scale of objects; or combine bounding box prediction and attention masking to generate a reasonable layout that conforms to the text description.

[0007] Although the above methods have made some progress in semantic alignment and image quality, they still generally suffer from the following problems: 1) Reliance on additional input or retraining leads to high costs and deployment difficulties; 2) Insufficient semantic control granularity makes it difficult to accurately achieve binding of multiple subjects and attributes; 3) In complex scenarios, semantic interleaving and attribute conflicts are prone to occur, which limits its real-time application value.

[0008] In light of the above issues, there is an urgent need for a method that requires no additional training or fine-tuning, while ensuring the quality of generation and achieving efficient semantic attribute binding and spatial control. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings and deficiencies of the prior art and provide a multi-subject semantic attribute binding image generation method based on a diffusion model. This method introduces layout images and reference images, clarifies the position and attribute division of semantic entities, and combines spatial layout, image features and text description to guide the process, thereby improving semantic consistency and attribute accuracy and alleviating the semantic alignment problem under complex text instructions.

[0010] To achieve the above objectives, the technical solution provided by this invention is as follows: a multi-subject semantic attribute binding image generation method based on a diffusion model. This method achieves multi-subject semantic image generation based on an optimized diffusion model. The optimization of the diffusion model improves the input, denoising process, and cross-attention mechanism of the original diffusion model. Specifically, the improvement to the input is: instead of independently encoding a single sentence, the sentence is segmented into multiple sub-sentences and encoded independently, achieving multi-subject semantic decoupling. Combined with the latent representation guidance of the layout image and the reference image, the model's ability to understand and generate multiple subjects and attributes in complex scenes is optimized. The improvement to the denoising process is: the original denoising process is divided into a generation branch and a reconstruction branch, used for latent representation fusion and denoising in the latent space, mitigating global semantic leakage and improving image consistency and visual quality. The improvement to the cross-attention mechanism is: based on the division of semantic regions, the cross-attention calculation, which was originally a single sentence uniformly applied to the global latent representation, is expanded to multiple sentences applying to different latent representation regions, enabling different regions to independently obtain corresponding semantic guidance, thereby avoiding feature aliasing and interference and improving semantic consistency in multi-subject scenes. The specific implementation of the multi-subject semantic attribute binding image generation method includes the following steps: 1) By using part-of-speech segmentation and semantic unit independent encoding, the input text is divided into structured independent clauses. Then, each independent clause is input into the optimized diffusion model to generate corresponding reference images and layout images. Next, the latent representation is obtained through the DDIM inversion method. Finally, the latent representations are fused by mask concatenation to obtain a latent representation that can be used as model input. 2) Input the latent representation obtained after fusion in step 1) into the optimized diffusion model to obtain the final target image. The denoising process of the optimized diffusion model is divided into a reconstruction branch and a generation branch. The reconstruction branch uses only the latent representation obtained by the DDIM inversion method to restore the original image, while the generation branch uses the latent representation obtained after fusion in step 1) to generate the target content through an improved cross-attention mechanism combined with the semantic guidance of multiple text prompt words. In each denoising iteration, the two branches output intermediate latent representations respectively, and fuse the latent space dimensions through a mask to obtain a new latent representation input for continued iteration, and finally output the target image.

[0011] Furthermore, step 1) includes the following steps: 1.1) Text part-of-speech segmentation and independent encoding: The input text is parsed and reconstructed, and it is broken down into multiple semantically clear clauses. Each clause retains only a single subject and its related attributes, as well as a summary sentence consisting only of all subject words. These clauses are then input into the text encoder for independent encoding to obtain a clear and distinct semantic representation. 1.2) Obtaining the layout image: The independent encoding of the sentence containing only all the subject words is input into the optimized diffusion model to generate a layout image without attributes. The DINO and SAM models are used to extract the subject position and mask information to obtain the spatial layout of the multi-subject scene. 1.3) Obtaining the Subject Reference Image: Independent encodings of descriptive sentences containing only their respective attributes are input into the optimized diffusion model. The subject reference image is generated through a mask-controlled noise initialization and region constraint generation process. That is, the initial latent representation required to generate the reference image is divided into target regions and non-target regions. The non-target regions are filled with noise obtained by inverting a fixed black image, which does not contain structural information or randomness, thus suppressing the generation activation of the non-target regions. The target regions retain the original random noise, guiding the model to gradually generate image content in the corresponding regions based on the descriptive sentences containing only their respective attributes. During the denoising process, the model only gradually generates reference image content in the target regions, while the non-target regions remain stable due to the lack of effective noise perturbation and semantic response. 1.4) Latent Representation Fusion: The generated layout image and reference image are inverted into the latent space using the DDIM inversion method to obtain their respective latent representations, and then stitched together according to the mask. The layout image only provides the latent representation of the background region, and the reference image only provides the latent representation of the reference image region. The two latent representations are stitched together. Random noise is further introduced into the contact area of ​​their respective latent representations to enhance smoothness and obtain the initial latent representation required for the generated image.

[0012] Furthermore, in step 2), during the process of gradually generating the image through diffusion, the denoising process is divided into a reconstruction branch and a generation branch. The two branches will extract the corresponding latent representations in each denoising process, fuse and stitch them together, obtain a new latent representation input, continue iterating, and finally output the target image. For the reconstruction branch, the layout image is first inverted to the noise space using the DDIM inversion method to obtain the initial latent representation. Then, without introducing any text encoding, the initial latent representation is gradually denoised to restore the layout image. The intermediate latent representation is saved in each denoising process and is used to be concatenated with the latent representation of the generated branch in subsequent steps, thereby achieving the preservation of structural information and the fusion of multi-branch features. For the generation branch, an improved cross-attention mechanism is adopted to extend the original cross-attention calculation, which was uniformly applied to the latent representation by a single sentence, to cross-attention calculation where multiple sentences apply to different latent representation regions. In each denoising process, the latent representation information passes through various parts of the optimized diffusion model. When passing through the improved cross-attention mechanism, the latent representation information is projected into a global query vector and multiplied by the region mask of each subject to obtain multiple local query vectors to locate the spatial region corresponding to each subject. Then, each local query vector is cross-attention calculated with the key vector and value vector obtained by mapping the corresponding subject's text description, thereby obtaining a series of sub-hidden states: ; ; In the formula, Represents the global query vector. For the first The key vector obtained by projecting the complete sentence of each subject. Scaling factor For the first The position mask of the layout image where each main sentence is located. For the first The attention score matrix calculated by each subject. For the first The value vector obtained by projecting the complete sentence of each subject. Indicates the first The sub-hidden states corresponding to each subject; during the attention calculation process, the position mask... It is used to explicitly limit the spatial response range of attention weights; Finally, based on the mask information, the formula is used to... The individual hidden states are merged into the final hidden state. : ; Hidden state It will then be fed into the subsequent functional modules of the diffusion model for processing and information transmission until a denoising process is completed; In each denoising process, the two branches each perform a denoising operation to obtain their respective intermediate latent representations. Then, the latent representations are concatenated using the masks of each subject to obtain a new latent representation, which is used as the input for the next denoising iteration of the generation branch. After obtaining the latent representation in each denoising process, both branches perform the above denoising operation to gradually complete the image generation.

[0013] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. This invention addresses both text structure and image semantics to construct a stable input representation, effectively mitigating the problems of attribute mixing and semantic interweaving. By inputting each clause into a text encoder for independent encoding, it generates a separate and clearly structured semantic representation, eliminating semantic interference between different entities and improving the descriptive parsing capabilities in complex scenarios with multiple subjects and attributes.

[0014] 2. This invention divides the diffusion model denoising process into a reconstruction branch and a generation branch. The intermediate latent representations of the two branches are spliced ​​and fused in the latent space, effectively mitigating the global semantic leakage problem and enhancing the collaborative expression of the background and target regions, thereby significantly improving the consistency and visual quality of the generated image.

[0015] 3. This invention improves the cross-attention mechanism in the original diffusion model by combining the generated mask to divide the semantic region, realizing independent attention-guided computation, avoiding feature aliasing and interference, thereby enhancing the local accuracy and spatial decoupling capability of semantic expression.

[0016] 4. In the task of generating complex objects, the scores of the three indicators of color, texture and shape reached 69.43%, 58.67% and 53.60% respectively, which is an improvement of 1.04%, 1.33% and 2.68% over the existing methods. The scores of the three categories in the large model evaluation reached 81.17%, which is an improvement of 2.47% over the existing methods.

[0017] 5. The method of the present invention requires no additional training, has low computational overhead and low resource consumption, and has strong applicability and promotion value. Attached Figure Description

[0018] Figure 1 This is a diagram illustrating the overall architecture of the method of the present invention.

[0019] Figure 2 Example diagram for text part-of-speech segmentation and independent encoding.

[0020] Figure 3The diagram shows the layout; DINO is an object detection model, and SAM is an image segmentation model.

[0021] Figure 4 This is a schematic diagram for obtaining a reference image.

[0022] Figure 5 This is a schematic diagram of potential representation fusion; in the diagram, DDIM inversion is a method of inverting an image into a noisy space.

[0023] Figure 6 A flowchart for generating the target image; in the diagram, To reconstruct the potential representation of the initial input to the branch. To generate the preprocessed fused latent representation of the initial input for the branch, Let T be the potential representation of the next time step obtained after the reconstruction branch. Let T be the potential representation of the next time step obtained through the generating branch. For time T-1 Potential representation and The latent representation is the latent noise state fused from the subject mask, and * represents the dot product calculation. Detailed Implementation

[0024] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0025] This embodiment discloses a multi-agent semantic attribute binding image generation method based on a diffusion model. This method achieves multi-agent semantic image generation based on an optimized diffusion model. The optimization of the diffusion model improves the input, denoising process, and cross-attention mechanism of the original diffusion model. Specifically, the input is improved by dividing the sentence into multiple sub-sentences and encoding them independently, thus decoupling the multi-agent semantics. Combined with the latent representation guidance of the layout image and the reference image, the model's ability to understand and generate multiple agents and attributes in complex scenes is optimized. The denoising process is improved by dividing the original denoising process into a generation branch and a reconstruction branch, which are used to perform latent representation fusion and denoising in the latent space, mitigating global semantic leakage and improving image consistency and visual quality. The cross-attention mechanism is improved by expanding the cross-attention calculation, which was originally a single sentence acting uniformly on the global latent representation, to multiple sentences acting on different latent representation regions, based on the division of semantic regions. This allows different regions to independently obtain corresponding semantic guidance, thereby avoiding feature aliasing and interference and improving semantic consistency in multi-agent scenes.

[0026] like Figure 1As shown, the specific implementation of the multi-subject semantic attribute binding image generation method includes the following steps: 1) By segmenting text into part-of-speech tags and encoding semantic units independently, the input text is divided into structured independent clauses. Each independent clause is then input into an optimized diffusion model to generate corresponding reference and layout images. The latent representation is then obtained through the DDIM inversion method. Finally, the latent representations are fused using mask concatenation to obtain a latent representation that can be used as model input. This process includes the following steps: 1.1) Text part-of-speech segmentation and independent encoding: such as Figure 2 As shown, the input text is parsed and reconstructed, breaking it down into multiple semantically clear clauses. Each clause retains only a single subject and its related attributes, as well as a summary sentence consisting only of all subject words. These clauses are then input into a text encoder for independent encoding, resulting in a clearly separated semantic representation. 1.2) Obtain the layout image: such as Figure 3 As shown, the independent encoding of a sentence containing only all the subject words is input into the optimized diffusion model to generate a layout image without attributes. The DINO and SAM models are then used to extract the subject position and mask information to obtain the spatial layout of a multi-subject scene. 1.3) Obtain the subject reference image: such as Figure 4 As shown, the independent encodings of descriptive sentences containing only their respective attributes are input into the optimized diffusion model. A reference image of the subject is generated using a noise initialization and region constraint generation process controlled by a mask. That is, the initial latent representation required to generate the reference image is divided into target regions and non-target regions. The non-target regions are filled with noise obtained by inverting a fixed black image, which does not contain structural information or randomness, thus suppressing the generation activation of the non-target regions. The target regions retain the original random noise, guiding the model to gradually generate image content in the corresponding regions based on the descriptive sentences containing only their respective attributes. During the denoising process, the model only gradually generates reference image content in the target regions, while the non-target regions remain stable due to the lack of effective noise perturbation and semantic response. 1.4) Latent representation fusion: such as Figure 5 As shown, the DDIM inversion method is used to invert the generated layout image and reference image into the latent space to obtain their respective latent representations, which are then stitched together according to a mask. The layout image provides only the latent representation of the background region, and the reference image provides only the latent representation of the reference image region. The latent representations are then stitched together. Random noise is further introduced into the contact area between their latent representations to enhance smoothness, thereby improving the fusion smoothness and overall generation quality. The fused latent representations are shown below. Figure 5 As shown on the right, it ultimately forms a potential representation that integrates multiple subject attributes and a reasonable spatial layout, providing accurate initial conditions for the subsequent diffusion generation stage.

[0027] 2) Input the latent representation obtained after fusion in step 1) into the optimized diffusion model to obtain the final target image. The denoising process of the optimized diffusion model is divided into a reconstruction branch and a generation branch. The reconstruction branch uses only the latent representation obtained by the DDIM inversion method to restore the original image, while the generation branch uses the latent representation obtained after fusion in step 1) to generate the target content through an improved cross-attention mechanism combined with the semantic guidance of multiple text prompt words. In each denoising iteration, the two branches output intermediate latent representations respectively, and fuse the latent space dimensions through a mask to obtain a new latent representation input for continued iteration, and finally output the target image.

[0028] Specifically, in the process of gradually generating an image through diffusion, the denoising process is divided into a reconstruction branch and a generation branch. In each denoising process, the two branches extract the corresponding latent representations, fuse and stitch them together to obtain a new latent representation input for further iteration, ultimately outputting the target image. The detailed process of the generation stage is as follows: Figure 6 As shown; For the reconstruction branch, the layout image is first inverted to the noise space using the DDIM inversion method to obtain the initial latent representation. Subsequently, without introducing any text encoding, the initial latent representation is gradually denoised to restore the layout image, and the intermediate latent representation is saved in each denoising process for concatenation with the latent representation of the generated branch in subsequent steps, thereby achieving the preservation of structural information and the fusion of multi-branch features. For the generation branch, an improved cross-attention mechanism is adopted to extend the original cross-attention calculation, which was uniformly applied to the latent representation by a single sentence, to cross-attention calculation where multiple sentences apply to different latent representation regions respectively; firstly, the preprocessed and fused latent representations are... As the initial input to the generation branch, the latent representation information passes through various parts of the optimized diffusion model in each denoising process. During the improved cross-attention mechanism, the latent representation information is projected into a global query vector and multiplied by the region masks of each subject to obtain multiple local query vectors, which are used to locate the spatial region corresponding to each subject. Then, each local query vector undergoes cross-attention calculation with the key vector and value vector mapped from the corresponding subject's text description, resulting in a series of sub-hidden states. ; ; In the formula, Represents the global query vector. For the first The key vector obtained by projecting the complete sentence of each subject. Scaling factor For the first The position mask of the layout image where each main sentence is located. For the first The attention score matrix calculated by each subject. For the first The value vector obtained by projecting the complete sentence of each subject. Indicates the first The sub-hidden states corresponding to each subject; during the attention calculation process, the position mask... It is used to explicitly limit the spatial response range of attention weights; Finally, based on the mask information, the formula is used to... The individual hidden states are merged into the final hidden state. : ; Hidden state It will then be fed into the subsequent functional modules of the diffusion model for processing and information transmission until a denoising process is completed; In each denoising process, the two branches each perform a denoising operation once, obtaining their respective intermediate latent representations, such as Figure 6 The reconstructed branch and the generated branch respectively obtained the time at time T-1 from time T. and The latent representation is then multiplied by the masks of each subject to obtain a new latent representation. The latent representation is used to generate the input for the next denoising iteration of the branch; after obtaining the latent representation in each denoising process, both branches perform the above denoising operation to gradually complete the image generation.

[0029] To verify the superiority of the multi-subject semantic attribute binding image generation method described in this embodiment, three subsets of the T2I-CompBench dataset—color, texture, and shape—were selected as the test set for the experiment. The experimental results are shown in Table 1.

[0030] Table 1. Experimental results from the T2I-CompBench dataset.

[0031] Table 1 presents the quantitative comparison results of each method on three subsets, where the underlined scores represent the scores of the second-best method in different metrics, and "+" indicates the performance improvement of the method of this invention compared to the second-best method. It can be observed that the method of this invention achieves significant performance improvements in all three tasks: color, texture, and shape. The BLIP-VQA metrics are improved by 1.04%, 1.33%, and 2.68%, respectively, and the VQAScore is improved by 2.47%, demonstrating superior image-text consistency compared to the baseline method.

[0032] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating a multi-agent semantic attribute binding image based on a diffusion model, characterized by, The method is based on an optimized diffusion model to realize multi-subject semantic image generation, and the optimization of the diffusion model is to improve the input, denoising process and cross attention mechanism of the original diffusion model; wherein, the improvement of the input is: the original independent coding of only a single sentence is improved to split the sentence into multiple sub-sentences and independently code each sub-sentence, realize multi-subject semantic decoupling, and combine the latent representation guidance of the layout image and the reference image to optimize the model's understanding and generation ability of multiple subjects and multiple attributes in complex scenes; the improvement of the denoising process is: the original denoising process is divided into a generation branch and a reconstruction branch, which are used to fuse and denoise the latent representation in the latent space, alleviate the global semantic leakage, and improve the image consistency and visual quality; the improvement of the cross attention mechanism is: based on the division of semantic regions, the original cross attention calculation of a single sentence uniformly acting on the global latent representation is extended to multiple sentences respectively acting on different latent representation regions, so that different regions can independently obtain corresponding semantic guidance, thereby avoiding feature aliasing and interference and improving the semantic consistency in multi-subject scenes; The specific implementation of the multi-subject semantic attribute binding image generation method includes the following steps: 1) By text part-of-speech segmentation and semantic unit independent coding, the input text is divided into structured independent clauses, then each independent clause is input into the optimized diffusion model to generate the corresponding reference image and layout image, then the latent representation is obtained by the DDIM inversion method, and finally the latent representation is fused by mask splicing to obtain a latent representation that can be used as input to the model; 2) The latent representation obtained after fusion in step 1) is input into the optimized diffusion model to obtain the final target image; wherein, the denoising process of the optimized diffusion model is divided into a reconstruction branch and a generation branch, the reconstruction branch only uses the latent representation obtained by the DDIM inversion method to restore the original image, and the generation branch uses the latent representation obtained after fusion in step 1) to generate target content by an improved cross attention mechanism combined with multi-text prompt semantics guidance, in each denoising iteration, the two branches output intermediate latent representations, and the latent space dimensions are fused by masks to obtain new latent representations for further iteration, and finally output the target image. 2.The diffusion model based multi-agent semantic attribute binding image generation method of claim 1, wherein, Step 1) includes the following steps: 1.1) Text part-of-speech segmentation and independent coding: parse and reconstruct the input text, and split it into multiple semantically clear sub-clauses, each sub-clause only retains a single subject and its related attributes, and a summary sentence composed of all subject words, and input them into the text encoder for independent coding to obtain clear and separated semantic representations; 1.2) Obtain the layout image: input the independent coding of the sentence with only all subject words into the optimized diffusion model to generate a layout image without attributes, and use the DINO and SAM models to extract subject position and mask information to obtain the spatial layout of the multi-subject scene; 1.3) Obtain the subject reference image: input the independent encoding of the description sentence containing only the respective attribute into the optimized diffusion model, generate the reference image of the subject through the noise initialization and region constraint generation process controlled by the mask, that is, divide the initial latent representation required to generate the reference image into target regions and non-target regions, where the non-target region is filled with noise obtained by inverting the fixed black image, which does not contain structural information and randomness, and the generation activation of the non-target region is suppressed; while the target region retains the original random noise, guiding the model to gradually generate image content in the corresponding region according to the description sentence containing only the respective attribute; in the denoising process, the model gradually generates reference image content only in the target region, while the non-target region remains stable due to the lack of effective noise disturbance and semantic response; 1.4) Latent representation fusion: use the DDIM inversion method to invert the generated layout image and the reference image into the latent space to obtain their respective latent representations, and splice them according to the mask, where the layout image only provides the latent representation of the background region, and the reference image only provides the latent representation of the reference image region, and the two are spliced into latent representations; further introduce random noise in the contact area of the respective latent representations to enhance smoothness, obtaining the initial latent representation required for generating the image. 3.The diffusion model based multi-agent semantic attribute binding image generation method of claim 1, wherein, In step 2), in the process of diffusion step-by-step image generation, the denoising process is divided into reconstruction branch and generation branch, and the two branches will extract the corresponding latent representation for fusion and splicing in each denoising process, to obtain a new latent representation for further iteration, and finally output the target image; For the reconstruction branch, first invert the layout image into the noise space by the DDIM inversion method to obtain the initial latent representation; Then, without introducing any text encoding, gradually denoise the initial latent representation to restore the layout image, and save the intermediate latent representation at each denoising process for splicing with the latent representation of the generation branch in the subsequent steps, thereby realizing the preservation of structural information and multi-branch feature fusion; For the generation branch, the improved cross-attention mechanism is used to extend the original cross-attention calculation of a single sentence acting on the latent representation to the cross-attention calculation of multiple sentences acting on different latent representation regions, in each denoising process, the latent representation information will pass through each part in the optimized diffusion model, and when passing through the improved cross-attention mechanism, the latent representation information will be projected into a global query vector, and then multiplied by the region mask of each subject to obtain multiple local query vectors to locate the corresponding spatial region of each subject, then each local query vector will be cross-attention calculated with the key vector and value vector mapped from the corresponding subject text description, to obtain a series of sub-hidden states: ; ; wherein, represents the global query vector, is the key vector obtained by projecting the th subject complete sentence, is a scaling factor, is the position mask of the layout image where the th subject sentence is located, is the attention score matrix computed for the th subject, is the value vector obtained by projecting the th subject complete sentence, represents the sub-hidden state corresponding to the th subject; during the attention computation, the position mask is used to explicitly limit the spatial response range of the attention weight. Finally , according to the mask information, the formula is used to combine the sub-hidden states into the final hidden state :​ ; Hidden state The data will continue to be sent to the subsequent functional modules of the diffusion model for processing and information transmission until the denoising process is completed. In each denoising process, the two branches perform a denoising operation respectively to obtain respective intermediate latent representations, and then the latent representations are spliced through the respective subject masks to obtain a new latent representation for generating the next denoising iteration input of the branch; after obtaining the latent representations in each denoising process, the two branches perform the above denoising operation to gradually complete image generation.