Layout-controllable image personalized generation method based on diffusion model

Through the personalized generation method of layout controllable images based on diffusion model, combined with dynamic adaptive visual encoder, static detail refining module and position perception module, the problems of detail preservation and position controllability in image personalized generation are solved, and customized image generation with high precision and controllable layout are achieved.

CN120014117APending Publication Date: 2025-05-16UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510118789.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to achieve fine-grained subject details and position controllability in image personalization generation, especially in image generation with controllable layout, which limits the potential and freedom of practical application scenarios.

Method used

A personalized generation method of layout controllable images based on diffusion model is adopted. A network composed of dynamic adaptive visual encoder, static detail refining module, position perception module and pre-trained variational autoencoder is achieved, combined with the lightweight fine-tuning of the adapter, high-precision control of the details and position of the reference subject is achieved.

Benefits of technology

It improves the detail fidelity and layout controllability of the reference subject, enhances the consistency and authenticity of the generated images, adapts to complex layout requirements, reduces calculation overhead, and improves generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014117A_ABST
    Figure CN120014117A_ABST
Patent Text Reader

Abstract

The invention discloses a layout-controllable image personalized generation method based on a diffusion model. The layout-controllable image personalized generation method comprises the following steps: 1, acquiring video and image data and corresponding text description, masks and bounding box annotations; 2, constructing a diffusion model adapter, and embedding reference main body features, bounding boxes and text description; 3, performing off-line training on the constructed diffusion model adapter; and 4, performing generation by using the trained model to realize the target of performing main body driven customized generation on a given image main body. According to the method and the device, the capability of generating any reference object at any position is realized in a manner of introducing the position information and the reference main body characteristics by using the lightweight adapter, and the main body characteristic keeping capability and the position controllability are improved, so that a user is allowed to autonomously generate a highly customized image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image generation, in particular to a layout-controllable image personalized generation method based on a diffusion model. Background Art

[0002] In recent years, image generation technology has developed rapidly thanks to the rise of diffusion models, and has been widely used in many fields such as product design, art creation, and e-commerce. Compared with traditional text-based image tasks, subject-driven personalized image generation is more challenging. In addition to generating high-quality images consistent with the text, it is also necessary to maintain the detailed features of the reference subject in the reference image.

[0003] Classic image personalization generation is mainly achieved through reconstruction-based methods that learn text embeddings bound to reference concepts or encoder-based methods that design an encoder to extract and fuse visual representations with text. However, the former often requires lengthy fine-tuning time and cannot be generalized to unseen examples to achieve fine-tuning-free testing, while the latter often ignores fine-grained subject details and produces suboptimal generation performance in zero-shot generation scenarios.

[0004] Although image personalized generation technology has made significant progress, there are still a series of unresolved problems and challenges in finer-grained control, including the preservation of subject details and position controllability. In particular, layout-controllable image personalized generation has not been fully explored, which greatly limits the potential and freedom in practical application scenarios. These problems and challenges provide directions for further research. Summary of the invention

[0005] In view of the shortcomings of the above-mentioned prior art, the present invention proposes a personalized generation method of layout-controllable images based on a diffusion model, which aims to improve the detail fidelity of the reference subject and increase the layout controllability, so as to generate highly customized images with controllable layout and detailed fidelity.

[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:

[0007] The invention discloses a method for generating personalized layout-controllable images based on a diffusion model, which comprises the following steps:

[0008] Step 1: Get the video dataset and image datasets ,in, Indicates video frames, Indicates the total number of video frames; represents the jth image, and N represents the total number of images;

[0009] Respectively and After preprocessing, the corresponding The corresponding reference subject bounding box , Reference Subject Tag , text description , instance segmentation map and mask as well as The corresponding reference subject bounding box , Reference Subject Tag , text description , instance segmentation map and mask ;

[0010] Step 2: Construct a position-controllable personalized generation network based on a diffusion model, including: a dynamic adaptive visual encoder, a static detail extraction module, a position perception module, a pre-trained variational autoencoder, and an adapter; according to the sampling probability, determine the input data of the network as or ;

[0011] Step 2.1: The dynamic adaptive visual encoder or Processing, the corresponding Reference body tag Corresponding dynamic features or Reference body tag Corresponding dynamic features ;

[0012] Step 2.2: The static detail extraction module or Processing, the corresponding Reference body tag The corresponding static characteristics or Reference body tag The corresponding static characteristics ;

[0013] Step 2.3: The location sensing module or Processing, the corresponding Reference body tag Corresponding position marker or Reference body tag Corresponding position marker ;

[0014] Step 2.4: The pre-trained variational autoencoder or Processing, the corresponding Hidden variable characteristics or Hidden variable characteristics ;

[0015] Step 2.5, the adapter or Processing, the corresponding Dynamic latent variable characteristics of or Dynamic latent variable characteristics of ;

[0016] Step 3: Training phase:

[0017] Using formula (1) to construct The loss function at time step t is Or use formula (2) to construct The loss function at time step t is , and is used to train the dynamic adaptive visual encoder, the linear projection layer in the static detail extraction module, the multilayer perceptron in the position perception module, and the adapter until the loss function or Until convergence, the trained dynamic adaptive visual encoder, static detail extraction module, position perception module, and adapter are obtained and combined with the pre-trained variational autoencoder to form a trained position controllable personalized generation model;

[0018] (1)

[0019] (2)

[0020] In formula (1) and formula (2), is the noise predicted by the network, It follows a standard Gaussian distribution Gaussian noise, represents the time step; is the identity matrix;

[0021] Step 4: Reasoning phase:

[0022] Step 4.1: The pre-trained variational autoencoder is subjected to Gaussian noise Processing is performed to obtain the latent variable features of time step t And with the given reference image , Bounding Box , and text description P are input into the trained position-controllable personalized generation model for processing to obtain dynamic latent variable features containing position information and reference image features After that, the cross attention calculation is performed together with the given text description P containing S reference main text tags to obtain the cross attention graph at time step t ; Set the bounding box Normalized to After the same size, the normalized bounding box is obtained , thereby utilizing and Construct the position loss function for time step t together ;

[0023] Step 4.2: Convert to bounding box mask ,Will The four corner points of are converted into corresponding corner masks and form a corner mask set , thereby utilizing , and Together we construct the scale loss function for time step t ;

[0024] Step 4.3: Use formula (3) to get the constraint loss function of the bounding box at time step t :

[0025] (3)

[0026] Using formula (4) Minimize to obtain the updated latent variable features at time step t :

[0027] (4)

[0028] In formula (4), represents the step size of time step t; is a scale factor; Express The gradient of

[0029] Step 4.4 The input is decoded into the pre-trained variational autoencoder to obtain a customized image with controllable layout.

[0030] The method for generating personalized layout-controllable images based on a diffusion model according to the present invention is also characterized in that the preprocessing in step 1 includes the following steps:

[0031] Step 1.1: Use the image description model to and Process and obtain Text description of and Text description of ;

[0032] Step 1.2: Use the named entity recognition model to Extract The labels of all reference subjects in ,in, Represents the i-th video frame The label of the kth reference subject in , n represents the i-th video frame The total number of reference subjects in ;

[0033] Using named entity recognition model Extract All reference subject labels in ,in, Represents the jth image Middle Reference subject tags, Represents the jth image The total number of reference subjects in ;

[0034] Step 1.3: Use the target detection model to and Process and get the i-th video frame The bounding boxes of all reference bodies in ,in, Represents the i-th video frame The bounding box of the kth reference subject in , and correspond;

[0035] Using target detection models and Processed in The bounding boxes of all reference bodies in ,in, Indicates the jth image The bounding box of the reference subject, and correspond;

[0036] Step 1.4: and the bounding box of the kth reference subject Input the segmentation model for processing to obtain the i-th video frame The instance segmentation map corresponding to the kth reference subject in and mask , and with correspond;

[0037] Will and The bounding box of the reference body Input the segmentation model for processing to obtain the jth image Middle The instance segmentation map corresponding to the reference subject and mask , and with correspond.

[0038] Furthermore, the dynamic adaptive visual encoder in step 2.1 comprises: a pre-trained visual encoder, a perceptual resampler and a first pre-trained text encoder;

[0039] When the input data is When , execute steps 2.1.1 to 2.1.4;

[0040] When the input data is When , execute steps 2.1.5 to 2.1.8;

[0041] Step 2.1.1: Middle Pair Perform similar frame matching and obtain Similar frames matched in And input into the segmentation model for processing, we get Instance segmentation map of ;in, express middle The corresponding segmentation map;

[0042] Step 2.1.2: Input the pre-trained visual encoder for feature extraction and obtain k reference subject labels Corresponding fine-grained visual features ;in, express middle corresponding fine-grained visual features;

[0043] Step 2.1.3: Set the query vector to be learned to q, and As the key vector and value vector respectively, they are input into the perceptual resampler for attention calculation, and the Corresponding reference subject dynamic features ;in, express middle Corresponding dynamic features;

[0044] Step 2.1.4: Send it to the first pre-trained text encoder for encoding, and get Text description of Text features ;

[0045] Step 2.1.5: After data enhancement, the jth enhanced image is obtained And input into the segmentation model for processing, we get Instance segmentation map of ;in, express middle The corresponding segmentation map;

[0046] Step 2.1.6: Input the pre-trained visual encoder for feature extraction and obtain of Reference body tags Corresponding fine-grained visual features ;in, express middle corresponding fine-grained visual features;

[0047] Step 2.1.7: Set the query vector q to be learned. After being used as the key vector and value vector respectively, they are input into the perceptual resampler for attention calculation, and we get Corresponding reference subject dynamic features ;in, express middle Corresponding dynamic features;

[0048] Step 2.1.8: Send it to the first pre-trained text encoder for encoding, and get Text description of Text features .

[0049] Furthermore, the static detail extraction module in step 2.2 includes: a pre-trained UNet network and a linear projection layer;

[0050] When the input data is When , execute step 2.2.1-step 2.2.2;

[0051] When the input data is When and Process and obtain middle The corresponding static characteristics ;in, express middle The corresponding static features;

[0052] Step 2.2.1. and Input the self-attention layer in the pre-trained UNet network for calculation, and get Corresponding self-attention features ;in, express middle The corresponding self-attention features;

[0053] Step 2.2.2: and Multiply them and input them into the linear projection layer for processing to get middle The corresponding static characteristics ;in, express middle The corresponding static features;

[0054] Furthermore, the location perception module in step 2.3 comprises: a second pre-trained text encoder, a Fourier encoding layer and a multi-layer perceptron;

[0055] When the input data is When , execute steps 2.3.1 to 2.3.5;

[0056] When the input data is When Corresponding , , Process and obtain middle Location marker ;

[0057] Step 2.3.1. of Input into the second pre-trained text encoder for processing, and get Text embedding features ,in, express The corresponding text embedding features;

[0058] Step 2.3.2: The corresponding bounding box Input into the Fourier coding layer for encoding to obtain the encoded position information ;in, express Corresponding location information;

[0059] Step 2.3.3: and After concatenation in the feature dimension, it is input into the multi-layer perceptron for encoding, and the obtained Corresponding text position marker , thus obtaining middle Text position marker ;

[0060] Step 2.3.4: and After concatenation in the feature dimension, it is input into the multi-layer perceptron for encoding, and the obtained Corresponding image position marker , thus obtaining middle Image location markers ;

[0061] Step 2.3.5: Use formula (5) to get middle Location marker :

[0062] (5)

[0063] In formula (5), Represents a splicing operation, Represents a multilayer perceptron.

[0064] Furthermore, the adapter in step 2.5 includes: a static cross-attention layer, a gated self-attention layer, and a dynamic cross-attention layer;

[0065] When the input data is When , execute steps 2.5.1 to 2.5.3;

[0066] When the input data is When and the corresponding , , , Calculate and get Dynamic latent variable characteristics of ;

[0067] Step 2.5.1: As the query vector, As key and value vectors, they are input into the static cross attention layer for cross attention calculation, and we get Static latent variable characteristics ;

[0068] Step 2.5.2: and Input to the gated self-attention layer, and use formula (6) to get The corresponding latent variable features containing position information :

[0069] (6)

[0070] In formula (6), represents the self-attention calculation, is a coefficient that adjusts the strength of layout control. is the activation function, is a scalar to be learned that is initialized to 0;

[0071] Will and Input to the gated self-attention layer, and use formula (2) to get The corresponding latent variable features containing position information ;

[0072] Step 2.5.3: Use formula (7) to get Dynamic latent variable characteristics of :

[0073] (7)

[0074] In formula (7), represents the cross attention calculation, is the coefficient for adjusting the weight of the reference image;

[0075] Furthermore, in step 4.1, formula (8) is used to construct ;

[0076] (8)

[0077] In formula (8), Represents the bounding box specified by the s-th reference body text tag Normalized to The bounding box obtained after the same size, express The position coordinates within; express The position coordinates in Indicates that the sth reference body text tag in P is The attention value at the (u, v) position in ; Indicates that the sth reference body text tag in P is In The attention value at the position.

[0078] Further, the step 4.2 includes:

[0079] Step 4.2.1: Bounding box Normalized to After the same size, the normalized bounding box is obtained ,in, represents the normalized s-th bounding box;

[0080] Applying a binary all-one mask will Convert to bounding box mask ;in, express Central position coordinates The bounding box mask at ;

[0081] Step 4.2.2, apply a binary all-1 mask to The four corner points of ;in, represents the sth corner point mask set, , Respectively The mask of the four corner points of ;

[0082] Step 4.2.3: Cross attention map Project it onto the x-axis and y-axis of the two-dimensional rectangular plane where it is located to obtain the x-axis attention vector and the y-axis attention vector ;

[0083] The bounding box mask Project to On the x-axis and y-axis of the two-dimensional rectangular plane, get the x-axis bounding box mask vector and the y-axis bounding box mask vector ;

[0084] Set the corner mask Project to On the x-axis and y-axis of the two-dimensional rectangular plane, get the x-axis corner point mask vector and the y-axis corner mask vector ;

[0085] Step 4.2.4: Use formula (9) to get the x-axis scale loss of time step t :

[0086] (9)

[0087] In formula (9), express Width;

[0088] Step 4.2.5: According to formula (9), we can get the y-axis scale loss of time step t. ;

[0089] Step 4.2.6: Use formula (10) to construct :

[0090] (10).

[0091] The electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the layout controllable image personalized generation method, and the processor is configured to execute the program stored in the memory.

[0092] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the method for generating personalized layout-controllable images when the computer program is executed by a processor.

[0093] Compared with the prior art, the beneficial effects of the present invention are:

[0094] 1. The present invention designs a layout-controllable personalized generation framework based on a diffusion model, which is an exploratory work to add position controllable capabilities to the personalized generation model. While maintaining the fidelity of the reference subject, it can give the model the ability to generate any object at any position, thereby improving the accuracy of positioning the generated content, thereby adapting to more complex layout requirements and adapting to the generation tasks of specific spatial configurations.

[0095] 2. The present invention innovatively designs a dynamic and static complementary visual feature refinement mechanism. One branch uses a dynamic adaptive visual encoder to extract dynamic detail information from video and data-enhanced image data, allowing the model to learn multi-perspective subject features; the other branch uses a static detail extraction module to further refine the static detail features, greatly improving the detail restoration of the reference subject and improving the consistency and authenticity of the generated image.

[0096] 3. This invention proactively proposes the goal of adding layout controllability to the personalized generation framework to improve generation controllability, and defines layout-controllable personalized generation tasks. Through fine-tuning of the gated sub-attention layer in the training phase and cross-attention regulation of the bounding box constraints in the reasoning phase, a dual position control signal is applied to the model, giving the model a robust position controllable generation capability, thereby enhancing the stability and adaptability of generating complex scenes and always maintaining a high visual quality of the generated image.

[0097] 4. The present invention adopts a lightweight adapter fine-tuning scheme, which has the ability to generalize to unseen reference subjects and different reference image distributions. It can be directly generated without fine-tuning in the test phase, which greatly reduces the computational overhead of generating images and improves generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] Figure 1 It is a framework diagram of the method of the present invention;

[0099] Figure 2 It is a structural diagram of the static detail extraction module of the present invention;

[0100] Figure 3 This is a structural diagram of the internal attention block of the adapter of the present invention. DETAILED DESCRIPTION

[0101] In this embodiment, a layout-controllable image personalized generation method based on a diffusion model is a more fine-grained controllable image generation. Its main features are that the ability to retain details of the reference subject is improved through alternating training of video and image data and the introduction of a static detail extraction module. The gated self-attention of the position-aware module and the cross-attention regulation of the bounding box constraints in the reasoning stage give the model position-controllable generation capabilities. In the framework of lightweight fine-tuning of the adapter, the fine-tuning-free generation of a specified reference subject at a specified position is achieved. Specifically, Figure 1 As shown, it includes the following steps:

[0102] Step 1: Get the video dataset and image datasets ,in, Indicates video frames, Indicates the total number of video frames; represents the jth image, N represents the total number of images; and After preprocessing, the corresponding and Corresponding reference subject bounding boxes, reference subject labels, text descriptions, instance segmentation maps, and masks.

[0103] Step 1.1: Use the image description model to and Process and obtain Text description of and Text description of ;

[0104] Step 1.2: Use the named entity recognition model to Extract The labels of all reference subjects in ,in, Represents the i-th video frame The label of the kth reference subject in , n represents the i-th video frame The total number of reference subjects in ;

[0105] Using named entity recognition model Extract All reference subject labels in ,in, Represents the jth image Middle Reference subject tags, Represents the jth image The total number of reference subjects in .

[0106] Step 1.3: Use the target detection model to and Process and get the i-th video frame The bounding boxes of all reference bodies in ,in, Represents the i-th video frame The bounding box of the kth reference subject in , and correspond;

[0107] Using target detection models and Processed in The bounding boxes of all reference bodies in ,in, Indicates the jth image The bounding box of the reference subject, and correspond;

[0108] Step 1.4: and the bounding box of the kth reference subject Input the segmentation model for processing to obtain the i-th video frame The instance segmentation map corresponding to the kth reference subject in and mask , and with correspond;

[0109] Will and The bounding box of the reference body Input the segmentation model for processing to obtain the jth image Middle The instance segmentation map corresponding to the reference subject and mask , and with correspond.

[0110] Step 2: Construct a position-controllable personalized generation network based on a diffusion model, including: a dynamic adaptive visual encoder, a static detail extraction module, a position perception module, a pre-trained variational autoencoder, and an adapter; according to the sampling probability, determine the input data of the network as or ,The purpose of sampling and using video and image datasets is to enable the ,model to learn the dynamic features of different perspectives or postures of the ,reference subject from the video data, and learn the changes of object texture, color, ,geometry from the data-enhanced image data, and the text description adapted to ,text changes is used for generation, improving the stability, generalization and ,diversity of the generation.

[0111] Step 2.1, the dynamic adaptive visual encoder comprises: a pre-trained visual encoder, a perceptual resampler and a first pre-trained text encoder;

[0112] When the input data is When , execute steps 2.1.1 to 2.1.4;

[0113] When the input data is When , execute steps 2.1.5 to 2.1.8;

[0114] Step 2.1.1: Middle Pair Perform similar frame matching and obtain Similar frames matched in And input into the segmentation model for processing, we get Instance segmentation map of ;in, express middle The corresponding segmentation map;

[0115] Step 2.1.2: The pre-trained visual encoder is used for feature extraction. The CLIP visual encoder is used to extract unpooled visual features from the penultimate layer of the encoder, which contains 256 visual blocks and 1 category label. k reference subject labels Corresponding fine-grained visual features ;in, express middle Corresponding fine-grained visual features.

[0116] Step 2.1.3, set the query vector to be learned to q. In this example, the vector length is 16. As the key vector and value vector respectively, they are input into the perceptual resampler for attention calculation, and we get Corresponding reference subject dynamic features ;in, express middle Corresponding dynamic features;

[0117] Step 2.1.4: Send it to the first pre-trained text encoder for encoding, and get Text description of Text features .

[0118] Step 2.1.5: After data enhancement, the jth enhanced image is obtained And input into the segmentation model for processing, we get Instance segmentation map of ;in, express middle Corresponding segmentation map. In this example, the data enhancement method uses rotation, flipping, color and geometric transformation to simulate the deformation and texture transformation of the object, so that the model can better adapt to some text descriptions with texture changes for generation;

[0119] Step 2.1.6: Input the pre-trained visual encoder for feature extraction and obtain of Reference body tags Corresponding fine-grained visual features ;in, express middle Corresponding fine-grained visual features.

[0120] Step 2.1.7: Set the query vector q to be learned. After being used as the key vector and value vector respectively, they are input into the perceptual resampler for attention calculation, and we get Corresponding reference subject dynamic features ;in, express middle Corresponding dynamic features;

[0121] Step 2.1.8: Send it to the first pre-trained text encoder for encoding, and get Text description of Text features .

[0122] Step 2.2: Figure 2 As shown, the static detail extraction module includes: a pre-trained UNet network and a linear projection layer;

[0123] When the input data is When , execute step 2.2.1-step 2.2.2;

[0124] When the input data is When and Process and obtain middle The corresponding static characteristics ;in, express middle The corresponding static features;

[0125] Step 2.2.1. and Input the self-attention layer in the pre-trained UNet network for calculation, and get Corresponding self-attention features ;in, express middle Corresponding self-attention features; In this example, the pre-trained UNet network is used for feature extraction. First, the feature space of the UNet network and the UNet network in the diffusion model backbone is naturally aligned, which can be better adapted. Second, the UNet network can extract multi-scale features for refining the main features. Compared with other feature extractors, it has better ability to retain main details. Other general feature extractors such as CLIP and DINO are also feasible.

[0126] Step 2.2.2: and Multiply them and input them into the linear projection layer for processing to get middle The corresponding static characteristics ;in, express middle The corresponding static features;

[0127] The static image features extracted using the static detail extraction module can further refine the subject details and have better detail preservation capabilities.

[0128] Step 2.3, the location perception module includes: a second pre-trained text encoder, a Fourier encoding layer and a multi-layer perceptron;

[0129] When the input data is When , execute steps 2.3.1 to 2.3.5;

[0130] When the input data is When Corresponding , , Process and obtain middle Location marker ;

[0131] Step 2.3.1. of Input into the second pre-trained text encoder for processing, and get Text embedding features ,in, express The corresponding text embedding features;

[0132] Step 2.3.2: The corresponding bounding box Input into the Fourier coding layer for encoding to obtain the encoded position information ;in, express Corresponding location information;

[0133] Step 2.3.3: and After concatenation in the feature dimension, it is input into the multi-layer perceptron for encoding, and we get Corresponding text position marker , thus obtaining middle Text position marker .

[0134] Step 2.3.4: and After concatenation in the feature dimension, it is input into the multi-layer perceptron for encoding, and we get Corresponding image position marker , thus obtaining middle Image location markers ;

[0135] Step 2.3.5: Use formula (5) to get middle Location marker :

[0136] (5)

[0137] In formula (5), Represents a splicing operation, Represents a multilayer perceptron.

[0138] Step 2.4: Pre-trained variational autoencoder pair or Processing, corresponding Hidden variable characteristics or Hidden variable characteristics ;

[0139] Step 2.5: Figure 3 As shown, the adapter contains: static criss-cross attention layer, gated self-attention layer, dynamic criss-cross attention layer;

[0140] When the input data is When , execute steps 2.5.1 to 2.5.3;

[0141] When the input data is When and the corresponding , , , Calculated .

[0142] Step 2.5.1: As the query vector, As key and value vectors, they are input into the static cross attention layer for cross attention calculation, and we get Static latent variable characteristics ;

[0143] Step 2.5.2: and Input to the gated self-attention layer, and use formula (6) to get The corresponding latent variable features containing position information :

[0144] (6)

[0145] In formula (6), represents the self-attention calculation, is a coefficient that adjusts the strength of layout control. is the activation function, is a scalar to be learned that is initialized to 0;

[0146] Will and Input to the gated self-attention layer, and use formula (2) to get The corresponding latent variable features containing position information .

[0147] Step 2.5.3: Use formula (7) to obtain dynamic latent variable features :

[0148] (7)

[0149] In formula (7), represents the cross attention calculation, is the coefficient that adjusts the weight of the reference image and is set to 0.6 in this example.

[0150] Step 3: Training phase:

[0151] like Figure 1 As shown in the trainable parameters, we use formula (1) to construct The loss function at time step t is Or use formula (2) to construct The loss function at time step t is , and is used to train the dynamic adaptive visual encoder, the linear projection layer in the static detail extraction module, the multilayer perceptron in the position perception module, and the adapter until the loss function or Until convergence, the trained dynamic adaptive visual encoder, static detail extraction module, position perception module, and adapter are obtained and combined with the pre-trained variational autoencoder to form a trained position controllable personalized generation model;

[0152] (1)

[0153] (2)

[0154] In formula (1) and formula (2), is the noise predicted by the network, It follows a standard Gaussian distribution Gaussian noise, represents the time step; is the identity matrix.

[0155] Step 4: Reasoning phase:

[0156] Step 4.1: Pre-trained variational autoencoder for Gaussian noise Processing is performed to obtain the latent variable features of time step t And given a reference image , Bounding Box , and text description P are input into the trained position-controllable personalized generation model for processing to obtain dynamic latent variable features containing position information and reference image features After that, the cross attention calculation is performed together with the given text description P containing S reference body text tags to obtain the cross attention graph at time step t . Given a bounding box Normalized to Size , thus according to and , use formula (8) to construct the position loss function at time step t ;

[0157] (8)

[0158] In formula (8), Represents the bounding box specified by the s-th reference body text tag Normalized to The bounding box of size is obtained, express The position coordinates within; express The position coordinates in Indicates that the sth reference body text tag in P is The attention value at the (u, v) position in ; Indicates that the sth reference body text tag in P is In The attention value at the position.

[0159] Step 4.2: Convert to bounding box mask ,Will The four corner points of are converted into corresponding corner masks and form a corner mask set , thus according to , and , use equations (9) and (10) to construct the scale loss function of time step t ;

[0160] Step 4.2.1. Given a bounding box Normalized to After the size is obtained , applying a binary all-1 mask will Convert to bounding box mask ;in, express Central position coordinates The bounding box mask at ;

[0161] Step 4.2.2, apply a binary all-1 mask to The four corner points of ;in, , Respectively Masks of the four corner points;

[0162] Step 4.2.3: Cross attention map Project it onto the x-axis and y-axis of the two-dimensional rectangular plane where it is located to obtain the x-axis attention vector and the y-axis attention vector ;

[0163] The bounding box mask Project to On the x-axis and y-axis of the two-dimensional rectangular plane, get the x-axis bounding box mask vector and the y-axis bounding box mask vector ;

[0164] Set the corner mask Project to On the x-axis and y-axis of the two-dimensional rectangular plane, get the x-axis corner point mask vector and the y-axis corner mask vector ;

[0165] Step 4.2.4: Use formula (9) to get the x-axis scale loss of time step t :

[0166] (9)

[0167] In formula (9), express Width;

[0168] Step 4.2.5: According to formula (9), we can get the y-axis scale loss of time step t. ;

[0169] Step 4.2.6: Use formula (10) to construct :

[0170] (10)

[0171] Step 4.3: Use formula (3) to get the constraint loss function of the bounding box at time step t :

[0172] (3)

[0173] Using formula (4) Minimize to obtain the updated latent variable features at time step t :

[0174] (4)

[0175] In formula (4), represents the step size of time step t; is a scale factor; Express gradient.

[0176] During the inference stage, the cross-attention regulation of the training-free bounding box constraints can further improve the position controllability of the generated images. At the same time, it can effectively alleviate the problems of subject missing, multi-subject conflict, and identity mixing in the multi-subject generation framework, thereby improving the consistency and quality of the generated images.

[0177] Step 4.4 After inputting the decoding part of the pre-trained variational autoencoder for decoding, a customized image with controllable layout is generated.

[0178] In summary, the present invention complementarily extracts the dynamic change features and static detail features of the reference subject through the dynamic adaptive visual encoder and the static detail extraction module, and then introduces the position-controllable generation capability through the two-stage joint operation of the position perception module in the training phase and the cross-attention regulation of the bounding box constraints in the reasoning phase, which greatly improves the ability of the generated image to preserve the details of the reference subject and the accuracy of the position-controlled generation. At the same time, the training method using adapter fine-tuning efficiently realizes the position-controllable personalized generation at a very low training cost while maintaining the original capabilities of the pre-trained model, reducing the computational overhead and improving the authenticity and consistency of the generated image.

[0179] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0180] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the above method when executed by a processor.

Claims

1. A layout controllable image personalized generation method based on a diffusion model, characterized in that: The following steps are involved: Step 1: Get the video dataset and image datasets ,in, Indicates video frames, Indicates the total number of video frames; represents the jth image, and N represents the total number of images; Respectively and After preprocessing, the corresponding The corresponding reference subject bounding box , Reference Subject Tag , text description , instance segmentation map and mask as well as The corresponding reference subject bounding box , Reference Subject Tag , text description , instance segmentation map and mask ; Step 2: Construct a position-controllable personalized generation network based on a diffusion model, including: a dynamic adaptive visual encoder, a static detail extraction module, a position perception module, a pre-trained variational autoencoder, and an adapter; according to the sampling probability, determine the input data of the network as or ; Step 2.1: The dynamic adaptive visual encoder or Processing, corresponding Reference body tag Corresponding dynamic features or Reference body tag Corresponding dynamic features ; Step 2.2: The static detail extraction module or Processing, corresponding Reference body tag The corresponding static characteristics or Reference body tag The corresponding static characteristics ; Step 2.3: The location sensing module or Processing, corresponding Reference body tag Corresponding position marker or Reference body tag Corresponding position marker ; Step 2.4: The pre-trained variational autoencoder or Processing, corresponding Hidden variable characteristics or Hidden variable characteristics ; Step 2.5, the adapter or Processing, corresponding Dynamic latent variable characteristics of or Dynamic latent variable characteristics of ; Step 3: Training phase: Using formula (1) to construct The loss function at time step t is Or use formula (2) to construct The loss function at time step t is , and is used to train the dynamic adaptive visual encoder, the linear projection layer in the static detail extraction module, the multilayer perceptron in the position perception module, and the adapter until the loss function or Until convergence, the trained dynamic adaptive visual encoder, static detail extraction module, position perception module, and adapter are obtained and combined with the pre-trained variational autoencoder to form a trained position controllable personalized generation model; (1) (2) In formula (1) and formula (2), is the noise predicted by the network, It follows a standard Gaussian distribution Gaussian noise, represents the time step; is the identity matrix; Step 4: Reasoning phase: Step 4.1: The pre-trained variational autoencoder is subjected to Gaussian noise Processing is performed to obtain the latent variable features of time step t And with the given reference image , Bounding Box , and text description P are input into the trained position-controllable personalized generation model for processing to obtain dynamic latent variable features containing position information and reference image features After that, the cross attention calculation is performed together with the given text description P containing S reference body text tags to obtain the cross attention graph at time step t ; Set the bounding box Normalized to After the same size, the normalized bounding box is obtained , thereby utilizing and Construct the position loss function for time step t together ; Step 4.2: Convert to bounding box mask ,Will The four corner points of are converted into corresponding corner masks and form a corner mask set , thereby utilizing , and Together we construct the scale loss function for time step t ; Step 4.3: Use formula (3) to get the constraint loss function of the bounding box at time step t : (3) Using formula (4) Minimize to obtain the updated latent variable features at time step t : (4) In formula (4), represents the step size of time step t; is a scale factor; Express The gradient of Step 4.4 The input is decoded into the pre-trained variational autoencoder to obtain a customized image with controllable layout.

2. The method for generating personalized layout-controllable images based on a diffusion model according to claim 1, characterized in that: The pre-treatment in step 1 comprises the following steps: Step 1.1: Use the image description model to and Process and obtain Text description of and Text description of ; Step 1.2: Use the named entity recognition model to Extract The labels of all reference subjects in ,in, Represents the i-th video frame The label of the kth reference subject in , n represents the i-th video frame The total number of reference subjects in ; Using named entity recognition model Extract All reference subject labels in ,in, Represents the jth image Middle Reference subject tags, Represents the jth image The total number of reference subjects in ; Step 1.3: Use the target detection model to and Process and get the i-th video frame The bounding boxes of all reference bodies in ,in, Represents the i-th video frame The bounding box of the kth reference subject in , and correspond; Using target detection models and Processed in The bounding boxes of all reference bodies in ,in, Indicates the jth image The bounding box of the reference subject, and correspond; Step 1.4: and the bounding box of the kth reference subject Input the segmentation model for processing to obtain the i-th video frame The instance segmentation map corresponding to the kth reference subject in and mask , and with correspond; Will and The bounding box of the reference body Input the segmentation model for processing to obtain the jth image Middle The instance segmentation map corresponding to the reference subject and mask , and with correspond.

3. The method for generating personalized layout-controllable images based on a diffusion model according to claim 2, characterized in that: The dynamic adaptive visual encoder in step 2.1 comprises: a pre-trained visual encoder, a perceptual resampler and a first pre-trained text encoder; When the input data is When , execute steps 2.1.1 to 2.1.4; When the input data is When , execute steps 2.1.5 to 2.1.8; Step 2.1.1: Middle Pair Perform similar frame matching and obtain Similar frames matched in And input into the segmentation model for processing, we get Instance segmentation map of ;in, express middle The corresponding segmentation map; Step 2.1.2: Input the pre-trained visual encoder for feature extraction and obtain k reference subject labels Corresponding fine-grained visual features ;in, express middle corresponding fine-grained visual features; Step 2.1.3: Set the query vector to be learned to q, and As the key vector and value vector respectively, they are input into the perceptual resampler for attention calculation, and the Corresponding reference subject dynamic features ;in, express middle Corresponding dynamic features; Step 2.1.4: Send it to the first pre-trained text encoder for encoding, and get Text description of Text features ; Step 2.1.5: After data enhancement, the jth enhanced image is obtained And input into the segmentation model for processing, we get Instance segmentation map of ;in, express middle The corresponding segmentation map; Step 2.1.6: Input the pre-trained visual encoder for feature extraction and obtain of Reference body tags Corresponding fine-grained visual features ;in, express middle corresponding fine-grained visual features; Step 2.1.7: Set the query vector q to be learned. After being used as the key vector and value vector respectively, they are input into the perceptual resampler for attention calculation, and we get Corresponding reference subject dynamic features ;in, express middle Corresponding dynamic features; Step 2.1.8: Send it to the first pre-trained text encoder for encoding, and get Text description of Text features .

4. The method for generating personalized layout-controllable images based on a diffusion model according to claim 3, characterized in that: The static detail extraction module in step 2.2 includes: a pre-trained UNet network and a linear projection layer; When the input data is When , execute step 2.2.1-step 2.2.2; When the input data is When and Process and obtain middle The corresponding static characteristics ;in, express middle The corresponding static features; Step 2.2.

1. and Input the self-attention layer in the pre-trained UNet network for calculation, and get Corresponding self-attention features ;in, express middle The corresponding self-attention features; Step 2.2.2: and Multiply them and input them into the linear projection layer for processing to get middle The corresponding static characteristics ;in, express middle The corresponding static features.

5. The method for generating personalized layout-controllable images based on a diffusion model according to claim 4, characterized in that: The location perception module in step 2.3 comprises: a second pre-trained text encoder, a Fourier encoding layer and a multi-layer perceptron; When the input data is When , execute steps 2.3.1 to 2.3.5; When the input data is When Corresponding , , Process and obtain middle Location marker ; Step 2.3.

1. of Input into the second pre-trained text encoder for processing, and get Text embedding features ,in, express The corresponding text embedding features; Step 2.3.2: The corresponding bounding box Input into the Fourier coding layer for encoding to obtain the encoded position information ;in, express Corresponding location information; Step 2.3.3: and After concatenation in the feature dimension, it is input into the multi-layer perceptron for encoding, and the obtained Corresponding text position marker , thus obtaining middle Text position marker ; Step 2.3.4: and After concatenation in the feature dimension, it is input into the multi-layer perceptron for encoding, and the obtained Corresponding image position marker , thus obtaining middle Image location markers ; Step 2.3.5: Use formula (5) to get middle Location marker : (5) In formula (5), Represents a splicing operation, Represents a multilayer perceptron.

6. The method for generating personalized layout-controllable images based on a diffusion model according to claim 5, characterized in that: The adapter in step 2.5 includes: a static cross-attention layer, a gated self-attention layer, and a dynamic cross-attention layer; When the input data is When , execute steps 2.5.1 to 2.5.3; When the input data is When and the corresponding , , , Calculate and get Dynamic latent variable characteristics of ; Step 2.5.1: As the query vector, As key and value vectors, they are input into the static cross attention layer for cross attention calculation, and we get Static latent variable characteristics ; Step 2.5.2: and Input to the gated self-attention layer, and use formula (6) to get The corresponding latent variable features containing position information : (6) In formula (6), represents the self-attention calculation, is a coefficient that adjusts the strength of layout control. is the activation function, is a scalar to be learned that is initialized to 0; Will and Input to the gated self-attention layer, and use formula (2) to get The corresponding latent variable features containing position information ; Step 2.5.3: Use formula (7) to get Dynamic latent variable characteristics of : (7) In formula (7), represents the cross attention calculation, is the coefficient that adjusts the weight of the reference image.

7. The method for generating personalized layout-controllable images based on a diffusion model according to claim 6, characterized in that: In step 4.1, formula (8) is used to construct ; (8) In formula (8), Represents the bounding box specified by the s-th reference body text tag Normalized to The bounding box obtained after the same size, express The position coordinates within; express The position coordinates in Indicates that the sth reference body text tag in P is The attention value at the (u, v) position in ; Indicates that the sth reference body text tag in P is In The attention value at the position.

8. The method for generating personalized layout-controllable images based on a diffusion model according to claim 7, characterized in that: The step 4.2 comprises: Step 4.2.1: Bounding box Normalized to After the same size, the normalized bounding box is obtained ,in, represents the normalized s-th bounding box; Applying a binary all-one mask will Convert to bounding box mask ;in, express Central position coordinates The bounding box mask at ; Step 4.2.2, apply a binary all-1 mask to The four corner points of ;in, represents the sth corner point mask set, , Respectively The mask of the four corner points of ; Step 4.2.3: Cross attention map Project it onto the x-axis and y-axis of the two-dimensional rectangular plane where it is located to obtain the x-axis attention vector and the y-axis attention vector ; The bounding box mask Project to On the x-axis and y-axis of the two-dimensional rectangular plane, get the x-axis bounding box mask vector and the y-axis bounding box mask vector ; Set the corner mask Project to On the x-axis and y-axis of the two-dimensional rectangular plane, get the x-axis corner point mask vector and the y-axis corner mask vector ; Step 4.2.4: Use formula (9) to get the x-axis scale loss of time step t : (9) In formula (9), express Width; Step 4.2.5: According to formula (9), we can get the y-axis scale loss of time step t. ; Step 4.2.6: Use formula (10) to construct : (10)。 9. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the method for generating a personalized layout-controllable image according to any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating personalized layout-controllable images according to any one of claims 1 to 8 are executed.

Citation Information

Cited By

  • Desktop layer object generation method based on conditional diffusion model

    CN120655856A