Bimodal cooperative control layout controllable main body consistency advertisement generation method
By adopting a dual-modal collaborative control method based on diffusion model in advertising generation, combined with VAE decoder, the problem of difficulty in realizing layout controllability and subject consistency in the prior art is solved, and high-precision and high-controllability advertising image generation is achieved.
Patent Information
- Application Number
- CN202510211618.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to achieve layout controllability and subject consistency simultaneously in advertising generation, and it has failed to fully utilize the potential of semantic space to improve the quality of advertising image generation.
Using a dual-mode collaborative control method based on diffusion model, by introducing semantic spatial modal control of subject consistency forward sampling and inverse update of layout conditions in the denoising generation process, a VAE decoder is used to generate advertising images with controllable layout and consistent subjects.
It realizes high-precision and high-controllability layout controllable subject consistent advertising generation in advertising generation, meeting the actual needs of diversified and high-quality advertising content generation.
Smart Images

Figure CN120219002A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image generation, and particularly relates to a method for generating layout-controllable and subject-consistent advertisements with dual-modal collaborative control based on a diffusion model. Background Art
[0002] In today's digital age, advertisements have become a key means of commercial promotion. Among them, high-quality advertisement images with subject-consistent layouts are crucial for attracting consumers' attention and conveying brand information. Existing image generation technologies, especially text-to-image diffusion models, although have made remarkable progress in generating high-fidelity images, still have many deficiencies when dealing with advertisement scenarios. In typical advertisement creation, such as the production of a series of product promotion posters, it is necessary to ensure that each product subject maintains a consistent appearance in different images and is accurately placed at the expected position to conform to the pre-designed layout. However, traditional models are difficult to meet these requirements.
[0003] At the same time, there is a disconnection in the existing technology in dealing with layout controllability and subject consistency, and the two are not organically combined to meet the actual needs of advertisement generation. Moreover, regarding the role of the semantic space in text-to-image models, existing research has limited attention and fails to fully explore its potential to improve the quality of advertisement image generation.
[0004] According to the applicant's retrieval and novelty search, the following several patents related to the present invention in the technical field of image generation are retrieved, and they are respectively:
[0005] 1. CN113361659B, an image controllable generation method and system based on latent space principal component analysis.
[0006] 2. CN117218489A, an image sample generation method based on key frame point detection.
[0007] The above-mentioned Patent 1 provides an image controllable generation method and system based on latent space principal component analysis. This method first randomly samples latent vectors in the latent space and inputs them into a generative adversarial network to generate images, then performs image transformation operations on the generated images for target attributes such as brightness change, size scaling, horizontal movement, vertical movement, etc. to obtain an image set with target attribute changes, and then constructs a latent vector set corresponding to the image set by minimizing the reconstruction loss and using gradient backpropagation. The reconstruction loss function is also optimized in the frequency domain. After that, the singular value decomposition method is used to perform principal component analysis on the latent vector set to find the direction with the largest variance change in the latent vector set. Finally, the latent vectors are moved in different degrees along the attribute change direction and then input into the generative adversarial network to output the images with controlled target attributes.
[0008] The above-mentioned Patent 2 provides an image sample generation method and system based on key point detection. First, image preprocessing is performed on the target detection image sample data set to enhance the data. Then, using an independent noise vector, the key point generation network generates information such as the position of key points and the width of the frame, and determines the part scale and appearance embedding. Subsequently, based on this, the initial mask of the part is calculated, etc., and the mask, foreground, and background are generated. A discriminant network is built to train the generation network. Finally, the trained model is used to generate image samples, supplement the few-sample database, alleviate the overfitting problem of the target detection network, and improve the detection accuracy. Each step has its unique network structure and processing method.
[0009] Although the above-mentioned Related Patent 1 can control the target attributes of an image by moving the latent vector along the direction of attribute change, the controllable attributes are mainly concentrated on relatively conventional attributes such as brightness, size, horizontal and vertical positions, etc. For some more complex and abstract image attributes, such as the emotional atmosphere of the image and the delicate changes in style, it is difficult to effectively control and generate them through the existing method steps. In addition, the singular value decomposition method is used for principal component analysis to find the direction with the largest variance change in the latent vector set. The computational complexity of singular value decomposition itself is not low. As the scale of the latent vector set increases, its computational cost will increase significantly, which may also affect the overall image generation speed. Summary of the Invention
[0010] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a layout-controllable and consistent image generation method based on bimodal collaborative control, and define a layout-to-consistent image (L2CI) generation task for the problems existing in text-image diffusion models in scenarios such as advertising production.
[0011] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0012] A layout-controllable and subject-consistent advertisement generation method based on bimodal collaborative control, comprising the following steps:
[0013] Step 1: Input information such as prompts required to generate the target advertisement image into the diffusion model. The diffusion model encodes the prompts and initializes a random latent space feature matrix as the starting point of the generation process. The specific steps can be described as follows:
[0014] The information input into the diffusion model in Step 1.1 includes the border information of the position of each object in the target advertisement image, the description information of the overall advertisement image, and the hyperparameter information required for the model to run. Among them, the border information of the position of each object in the target advertisement image is given in the form of the upper left coordinate and the lower right coordinate of the border. For example, the input form of the border information of object i is The descriptive information of the entire advertising image is given in text form, including the nouns that need to be generated with controllable layout conditions, and number the nouns according to their positions in the text. For example, if a noun is in the $i$-th position in the text, the noun is numbered as object $i$. The image information of the main body for which consistent generation is to be performed is given in image form, which is used to specify the main body of the finally generated advertising image. The hyperparameter information required for model operation is normally used to determine the random seed of the model, the scheduler selection and the number of time steps in the denoising generation process, and the step size $\alpha$ for denoising generation at each time step. t , such as the intensity of bimodal collaborative control.
[0015] Step 1.2: The model uses the CLIP text encoder and the T5-XXL text encoder to encode the prompt words input by the user, and concatenates the encoding outputs of the two encoders in the feature dimension to obtain the text features in the semantic space of the user input. And the model will initialize a random latent space Gaussian noise matrix as the latent space features of the user input according to the random seed in the hyperparameters, and at the same time select the scheduler for denoising generation according to the hyperparameters to determine the step size $\alpha$ for denoising at each time step $t$. t . The model processes various information input into the model from the perspective of the feature space.
[0016] Step 2, input an image containing a specific main body into the diffusion model, and introduce a CLIP image encoder to encode the image to obtain the image features of the main body in the image in the semantic space. The specific steps can be described as follows:
[0017] In this step, the image information of the main body for which consistent generation is to be performed input by the model will be encoded by a pre-trained CLIP image encoder to extract the image features of the picture, and the extracted features will be transformed into image features with the same dimension as the text features in the semantic space through a linear projection layer and a Layer normalization layer.
[0018] Step 3, use the diffusion model to process the latent space feature matrix, perform N-step denoising generation, and apply a bimodal collaborative control strategy for semantic space features and latent space features to control the model to denoise the latent space feature matrix at each time step of the denoising process. The specific steps can be described as follows:
[0019] Step 3.1: The diffusion model will use a neural network with a U-Net architecture as the denoising network to predict the noise residuals in the latent space features at each time step. In the prediction process, a bimodal collaborative control strategy is used to guide the prediction results to generate in the direction of layout controllable main body consistency, and the noise residuals in the latent space features are removed according to the prediction results. The bimodal collaborative control function is expressed as follows:
[0020]
[0021] Among them, p(z t-1 |z t ) represents the layout controllable subject consistency advertisement generation result at time step t, p f and p b respectively represent the semantic space modal control of subject consistency forward sampling and the latent space modal control of layout conditional reverse update. z t is the latent space feature at time step t, represents the latent space feature after the modal control of subject consistency, and z t-1 represents the latent space feature after the denoising generation of bimodal collaborative control.
[0022] Step 3.2: Use the semantic space modal control p f of subject consistency forward sampling mentioned in Step 3.1 to perform forward sampling of subject consistency. The model will adopt a decoupled cross-attention mechanism for text features and image features in the semantic space in this step, and an additional cross-attention layer is added outside each cross-attention layer of the denoising network to receive image features in the semantic space, which is used to integrate subject consistency information into the latent space features. At the same time, the original cross-attention layer continues to receive text features in the semantic space, which is used to integrate user prompt information into the latent space features. The formula for the semantic space modal control of subject consistency forward sampling is as follows:
[0023]
[0024] Q = z t W q , K = cW k , V = cW v , K ′ = c i W k ′ , V ′ = c i W v ′
[0025] Among them, W q , W k , W v respectively represent the query weight matrix, key weight matrix, and value weight matrix of the original diffusion model cross-attention layer. c represents the text features in the semantic space. The corresponding Q, K, and V respectively represent the query vector, key vector, and value vector of the original diffusion model cross-attention layer. W k ′ and W v ′They respectively represent the key weight matrix and value weight matrix of the additional cross-attention layer in the IP-Adapter, obtained from the pre-trained IP-Adapter model, c i represents the image features in the semantic space, corresponding to K ′ and V ′ respectively represent the key vector and value vector in the additional cross-attention layer. d represents the dimension of the feature vector obtained after calculating the query vector and the key vector during the cross-attention calculation process. λ is a user-specified hyperparameter used to determine the intensity of semantic space modality control for subject consistency.
[0026] This step decouples the text features and image features in the semantic space. In traditional diffusion models, the text features and image features are generally directly concatenated and then a cross-attention layer is used to integrate the semantic space information into the latent space information. However, the effectiveness of this method is insufficient and it will damage the image generation quality. Therefore, the present invention adopts a semantic space decoupling method similar to that in IP-Adapter, and uses different cross-attention layers to respectively integrate the text features and image features in the semantic space into the latent space, greatly improving the image generation quality and the subject consistency generation effect.
[0027] Step 3.3: Use the latent space modality control p of the reverse update of the layout condition mentioned in Step 3.1 b to perform the reverse update of the layout condition. The model will construct three losses in this step according to the border information of the position of each object in the target advertisement image input by the user: in-frame loss, border loss, out-of-frame loss, denoted as Update the features of the forward sampling semantic space modality control with subject consistency obtained in Step 3.2 Complete the reverse update latent space modality control of the layout condition at a specific time step. The formulas for updating the latent space feature denoising generation direction using three different types of loss functions are as follows:
[0028]
[0029] Among them, represents the loss function of the in-frame loss, represents the loss function of the out-of-frame loss, represents the loss function of the border loss, represents the total loss function after comprehensively considering the three spatial constraints.
[0030] The three losses constructed by the model according to the border information of the position of each object in the target advertisement image input by the user: in-frame loss border loss out-of-frame loss The calculation process is as follows:
[0031] To ensure that the generated object is included in the bounding box of each object in the target advertisement image specified by the user, an in-box loss is introduced to constrain the P points with the highest attention values in the bounding box where the object is located, and a loss function is constructed. The calculation formula of the in-box loss is as follows:
[0032]
[0033] where i represents the position number of the object that needs to be controlled by the layout condition in the user prompt, s i represents the object i that needs to be controlled by the layout condition, S represents the set of objects that need to be controlled by the layout condition, represents the attention map corresponding to the object i at time step t, M i represents the mask matrix of the bounding box where the object i is located, P is a hyperparameter provided by the user, representing the strength of the spatial constraint, topk(·, P) represents the P points with the highest attention values in the specified area, represents the loss value of the in-box loss of the object i, represents the loss value of the in-box losses of all objects.
[0034] Loss function will limit the bounding box where the object i is located to contain as many as possible the P points with the highest attention values in the object generation area, making the generated object close to the position provided by the user. At the same time, only constraining a few P high-attention points is sufficient to affect the generation of the position where the object i is located and reduce the impact on the image generation quality.
[0035] To ensure that the generated object does not exceed the bounding box specified by the user, an out-of-box loss is introduced to constrain the P points with the highest attention values outside the object bounding box, and a loss function is constructed. The calculation formula of the out-of-box loss is as follows:
[0036]
[0037] where, represents the loss value of the out-of-box loss of the object i, represents the loss value of the out-of-box losses of all objects. The definitions of other symbols are the same as the symbol definitions in the calculation process.
[0038] The loss can only ensure that the generated object is included in the bounding box provided by the user, but cannot ensure that the generated object does not exceed the bounding box specified by the user. The loss function will limit the P points with the highest attention values outside the bounding box where the object i is located, and minimize the attention values of the P highest attention value points outside the bounding box as much as possible to prevent the object generation range from exceeding the bounding box.
[0039] To avoid the object being generated too small within the border specified by the user, a border loss is introduced to constrain the projections of the generated object on the x-axis and y-axis of the image plane. The calculation formulas for the spatial constraint and the loss function are as follows:
[0040] m x (k) = max j=1,…,H {M i (j, k)}
[0041]
[0042] where m x represents the projection of the mask of the border of object i on the x-axis, M i (j, k) represents the mask corresponding to the border of object i, and max j=1,…,H {M i (j, k)} means traversing the mask matrix from bottom to top in the y-axis direction and selecting the element with the largest value in the mask matrix, and projecting it to the element m x (k) at the corresponding position k on the x-axis. represents the projection of the attention map of object i on the x-axis, means traversing the attention map matrix from bottom to top in the y-axis direction and selecting the element with the largest value in the attention map matrix, and projecting it to the element and respectively represent the left and right boundaries of the projection of the border corresponding to object i on the x-axis. L is a hyperparameter specified by the user for the intensity of the border loss. represents the error term between the projection m k (k) of the mask matrix and the projection of the attention map matrix under the condition that the ordinate k = 1 in the x-axis dimension W. means sampling L error terms and uniformly between the given x-axis coordinates represents the border loss of object i under the x-axis projection.
[0043] Similarly, by projecting the mask matrix and the attention map matrix of object i onto the y-axis, the border loss of object i under the x-axis projection can be obtained
[0044] m y (j) = max k=1,…,W {M i (j, k)}
[0045]
[0046] where the symbol represents the same as The symbols in the calculation process are the same.
[0047]
[0048] Among them, and respectively represent the bounding box loss of object i under the projections on the x-axis and y-axis. s i represents object i, and S represents the set of objects that require layout condition control. represents the bounding box loss of all objects.
[0049] The loss sum and the loss can only ensure that the generated objects are within the bounding boxes specified by the user. However, the objects may also be generated too small within the bounding boxes and do not fill the entire bounding box area. Therefore, the
[0050] loss is introduced to constrain the attention of the objects to be generated on the projections on the x-axis and y-axis respectively, ensuring that the size of the generated objects can fill the bounding box area specified by the user.
[0051] In step 4, the VAE decoder is used to decode the result generated by the model denoising from the latent space to the pixel space, obtaining an advertisement image with consistent layout and controllable main body.
[0052]
[0053] Among them, z0 represents the latent space feature at time step 0 obtained after all denoising generation steps. represents the VAE decoder. represents the image in the pixel space finally generated by the diffusion model.
[0054] Compared with the prior art, the present invention has at least the following beneficial effects:
[0055] 1. The present invention combines the layout controllable generation of the diffusion model with the main body consistency generation for the first time, and innovatively proposes a layout controllable and consistent advertisement generation model. It successfully defines and solves the layout-to-consistent image (L2CI) generation task.
[0056] 2. During the denoising process of the diffusion model, a dual-modal collaborative control function is adopted, and the denoising process at each time step is divided into two main stages. In the first stage, semantic space mode control for subject consistency forward sampling is performed. In the semantic space, text features and image features are decoupled, and different cross-attention layers are used to integrate the text features and image features in the semantic space into the latent space features of the model. In the second stage, latent space mode control for layout controllable backward update is performed. According to the bounding box information of the position of each object in the target advertisement image input by the user, three losses are constructed: in-box loss Bounding box loss Out-of-box loss Three different types of loss functions are formed to control the object generation in the bounding box area specified by the user.
[0057] 3. The layout controllable subject consistency advertisement generation model proposed by the present invention starts from a unified perspective, comprehensively analyzes and processes the semantic space and the latent space, and realizes the synchronous optimization of the semantic space and the latent space through a dual-modal collaborative control function during the denoising process. This design effectively coordinates the dual requirements of layout control and generation consistency. Through this method, the present invention provides a user with a high-precision and highly controllable advertisement customization generation solution, meeting the actual needs of diverse and high-quality advertisement content generation. Brief Description of the Drawings
[0058] Figure 1 It is a schematic structural diagram of the present invention.
[0059] Figure 2 It is a flowchart of the algorithm of the present invention.
[0060] Figure 3 It is a diagram showing the effect of the present invention. Detailed Embodiment
[0061] The following describes the embodiments of the present invention in detail with reference to the drawings and embodiments.
[0062] Refer to Figure 1 and Figure 2 As shown, the main steps of the present invention are as follows:
[0063] Step1. The user inputs information such as prompts required to generate a target advertisement image into the diffusion model. The text encoder in the diffusion model encodes the user input prompts, and initializes a random latent space feature matrix as the starting point of the image generation process.
[0064] Step2. Introduce a CLIP image encoder to encode the image to obtain the subject image features in the semantic space.
[0065] Step3. The diffusion model incorporates semantic space features into the latent space features and uses the semantic space modality control p of the subject consistency forward sampling f Perform forward sampling of subject consistency.
[0066] Step4. The diffusion model further processes the latent space features and uses the latent space modality control p of the layout condition reverse update b Perform reverse update of the layout condition.
[0067] Step5. Repeat Step3 and Step4. At each time step of denoising generation by the diffusion model, perform semantic space modality control of subject consistency and latent space modality control of layout condition to form a bimodal collaborative control to denoise the latent space features of the diffusion model.
[0068] Step6. Use the VAE decoder to decode the denoising generation result of the latent space features after N-step bimodal collaborative control from the latent space to the pixel space to generate an advertisement image with controllable layout and consistent subject.
[0069] The implementation effect of the present invention is as Figure 3 shown. Given the prompt "a small fan is [*]" and the position bounding box of the subject to be generated (small fan), when "[*]" is replaced with "located in the corner", the subject to be generated will be generated in the corner of the specified position; when replaced with "blowing on two people", it will be generated between the two people at the specified position; when replaced with "located on the table", it will be generated on the tabletop at the specified position. In addition, no matter how the prompt changes, the small fan in the generated picture is the same fan, ensuring the consistency of the generation effect.
[0070] The present invention adopts the above technical solutions and proposes a layout controllable and consistent advertisement generation method based on bimodal collaborative control, which takes into account both the layout controllable generation task and the consistent generation task in the image generation process, and solves the technical problem that traditional advertisement generation models cannot accurately control the image layout and ensure multi- Figure 1 consistency at the same time. By introducing a bimodal collaborative control function in the denoising generation stage, the present invention effectively regulates the relative positions of various objects in the image, making the advertisement image more consistent in the overall visual effect. The model is evaluated and analyzed using a real image dataset. The results show that compared with the existing traditional methods, the method proposed by the present invention shows significant advantages in layout controllable consistency, generated subject consistency, compliance with user prompts, and diversity of generated backgrounds.
Claims
1. A method for generating layout-controllable subject-consistent advertisements with dual-modal collaborative control, characterized in that: The steps include: Step 1: input the prompt words required to generate the target advertising image into the diffusion model, the diffusion model encodes the prompt words, and initializes a random latent space feature matrix as the starting point of the generation process; Step 2, input an image containing a specific subject into the diffusion model, introduce a CLIP image encoder to encode the image, and obtain the image features of the subject in the image in the semantic space; Step 3, using a diffusion model to process the latent space feature matrix, performing N-step denoising generation, and at each time step of the denoising process, applying a dual-modal collaborative control strategy control model for semantic space features and latent space features to denoise the latent space feature matrix; Step 4: Use the VAE decoder to decode the denoising result of the model from the latent space to the pixel space to obtain an advertising image with a consistent layout and controllable subject.
2. The method for generating layout controllable subject consistency advertisements of dual-modal collaborative control according to claim 1 is characterized in that: In step 1 and step 2, the input information of the model includes the bounding box information of the location of each object in the target advertising image, the overall description information of the advertising image, the image information of the consistent subject to be generated, and the hyperparameter information required for the model operation.
3. The method for generating layout controllable subject consistency advertisements of dual-modal collaborative control according to claim 1, characterized in that: In step 1, the model uses CLIP text encoder and T5-XXL text encoder to encode the prompt word input by the user, and splices the encoding outputs of the two encoders in the feature dimension to obtain the text features in the semantic space of the user input; and the model initializes a random latent space Gaussian noise matrix as the latent space features of the user input, and at the same time selects the scheduler generated by denoising according to the hyperparameters to determine the step length α of denoising for each time step t t .
4. The method for generating layout controllable subject consistency advertisements of dual-modal collaborative control according to claim 1 is characterized in that: In step 3, the diffusion model uses a neural network with a U-Net architecture as a denoising network to predict the noise residual in the latent space features at each time step. During the prediction process, a bimodal collaborative control strategy is used to control the prediction results to be generated in the direction of consistency of the layout controllable body, and the noise residual in the latent space features is removed according to the prediction results.
5. The method for generating layout controllable subject consistency advertisements of dual-modal collaborative control according to claim 1 or 4, characterized in that: The denoising process is implemented as follows: Step 3.1: The diffusion model uses a bimodal collaborative control strategy to control the prediction results to be generated in the direction of consistency of the layout controllable subjects; the bimodal collaborative control function is expressed as follows: Among them, p(z t-1 |z t ) represents the layout controllable subject consistency advertisement generation result at time step t, p f and p b They represent the semantic space modal control of subject consistency forward sampling and the latent space modal control of layout conditional reverse update, respectively. t is the latent space feature at time step t, represents the latent space features after the subject-consistent semantic space modal control processing, z t-1 represents the latent space features after being processed by bimodal collaborative control; Step 3.2: Use the subject-consistent forward sampling semantic space modality control p f Perform subject-consistent forward sampling and use a decoupled cross-attention mechanism for text features and image features in the semantic space; Step 3.3: Use the layout condition to reversely update the latent space modal control p b Performs the reverse update of layout conditions.
6. The method for generating layout controllable subject consistency advertisements of dual-modal collaborative control according to claim 5 is characterized in that: In step 3.2, the forward sampling formula for subject consistency is as follows: Q=z t W q ,K=cW k ,V=cW v ,K ′ =c i W k ′ ,V ′ =c i W v ′ Where W q , W k , W v represents the query weight matrix, key weight matrix, and value weight matrix of the cross-attention layer of the original diffusion model, respectively; c represents the text feature in the semantic space, and the corresponding Q, K, and V represent the query vector, key vector, and value vector of the cross-attention layer of the original diffusion model, respectively; W k ′ and W v ′ They represent the key weight matrix and value weight matrix of the additional cross-attention layer in IP-Adapter, obtained from the pre-trained IP-Adapter model, respectively. i Represents the image features in the semantic space, and the corresponding K ′ and V ′ They represent the key vector and value vector in the additional cross-attention layer respectively; d represents the dimension of the feature vector obtained after the query vector and the key vector are calculated during the cross-attention calculation process, and λ is a user-specified hyperparameter used to determine the strength of the semantic space modal control of subject consistency.
7. The method for generating layout controllable subject consistency advertisements of dual-modal collaborative control according to claim 5 is characterized in that: In step 3.3, the model constructs three different losses: inner-box loss, outer-box loss, and outer-box loss according to the box information of each object in the target advertisement image input by the user. b The formula for updating the latent space feature denoising generation direction is as follows: in, represents the loss function for the in-box loss, represents the loss function for out-of-box loss, represents the loss function of bounding box loss, Represents the total loss function after comprehensively considering the three spatial constraints.
8. The method for generating layout controllable subject consistency advertisements of dual-modal collaborative control according to claim 7, characterized in that: In step 3.3, the reverse update method of the layout conditions is as follows: The calculation formula of the loss in the box is as follows: Among them, i represents the position number of the object in the user prompt word that needs to be controlled by the layout condition, and s i represents the object i that needs layout condition control, S represents the set of objects that need layout condition control, represents the attention map corresponding to object i at time step t, M i represents the mask matrix of the border of the location of object i. P is a hyperparameter provided by the user, indicating the strength of the spatial constraint. topk(·,P) represents the P points with the highest attention value in the specified area. Indicates the loss value of the object i in the frame, Indicates the loss value of all object boxes; The calculation formula for out-of-frame loss is as follows: in, Indicates the loss value of object i outside the frame, Indicates the loss value of all objects outside the frame. Other symbols are defined in the same way. Symbolic definition of computational procedures; The calculation formula for the border loss is as follows: the mask matrix and attention map matrix of object i are projected onto the x-axis to obtain the border loss of object i under the y-axis projection m x (k)=max j=1,…,H {M i (j,k)} Where m x Represents the projection of the mask of the bounding box of object i on the x-axis, M i (j,k) represents the mask corresponding to the bounding box of object i, max j=1,…,H {M i (j,k)} means traversing the mask matrix from bottom to top along the y-axis direction, and selecting the largest element in the mask matrix to project to the element m at the corresponding position k on the x-axis x (k); represents the projection of the attention map of object i on the x-axis, It means traversing the attention map matrix from bottom to top along the y-axis direction, and selecting the largest element in the attention map matrix to project to the element at the corresponding position k on the x-axis and They represent the left and right boundaries of the projection of the bounding box corresponding to object i on the x-axis. L is a user-specified hyperparameter used to determine the strength of the bounding box loss. Represents the projection m of the mask matrix on the x-axis dimension W under the condition of ordinate k = 1 k (k) and attention map matrix projection The error term between Indicates uniform distribution at a given x-axis coordinate and Sample L error terms between Represents the border loss of object i under x-axis projection; Project the mask matrix and attention map matrix of object i onto the y-axis to obtain the border loss of object i under the x-axis projection m y (j)=max k=1,…,W {M i (j,k)} in, and Respectively represent the border loss of object i under x-axis and y-axis projection, s i represents object i, S represents the set of objects that need to be controlled by layout conditions, The loss value representing the bounding box loss of all objects.
9. The method for generating layout controllable subject consistency advertisements of dual-modal collaborative control according to claim 1, characterized in that: In step 4, after completing the denoising generation of the bimodal collaborative control of all time steps of the diffusion model, the generated result exists in the form of latent space features. The model decodes the image from the latent space to the pixel space through the VAE decoder to generate an advertising image with controllable layout and consistent subject.
10. The method for generating layout controllable subject consistent advertisements of dual-modal collaborative control according to claim 1, characterized in that: The VAE decoder calculation formula is as follows: Where z0 represents the latent space feature at time step 0 after all denoising generation steps. represents the VAE decoder, Represents the image in pixel space that is ultimately generated by the diffusion model.