Multi-modal controllable portrait generation method and system
By decoupling and preprocessing the original portrait image by multimodal input conditions, extracting features using visual language model and reference image segmentation network, and generating controllable portraits in combination with U-Net network, the problem of poor flexibility and accuracy of image generation under multimodal input conditions in the prior art is solved, and more efficient multi-task generation and personalized image creation are achieved.
Patent Information
- Application Number
- CN202510236206.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-01
AI Technical Summary
The existing portrait generation methods have poor flexibility and accuracy in generating images under multimodal input conditions, especially under decoupling of multimodal input conditions, which is difficult to achieve efficient multi-source information processing.
By decoupling the original portrait image by multimodal input conditions, including text prompts, reference images and hand-drawn layout decoupling, features are extracted using visual language models and reference image segmentation network, preprocessing is combined with VAE and CLIP encoder, and stitching and loss function training is performed through U-Net network to generate controllable portraits.
It realizes flexibility and accuracy improvement under multimodal input conditions, and the generated images are more consistent with the given layout, text description and reference images, supporting multi-task generation and personalized image creation.
Smart Images

Figure CN120235973A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and particularly to a multi-modal controllable portrait generation method and system. Background Art
[0002] In recent years, the progress of diffusion models has driven major breakthroughs in the field of generative artificial intelligence, especially showing unprecedented potential in visual tasks that combine multiple controls during the image generation process. Through the technology of generating images from text precisely guided by semantic layout and personalized image customization methods driven by theme features, researchers have successfully achieved fine-grained control over the generated content. This technological breakthrough has promoted the rapid development of portrait generation technology, making portrait generation gradually become a research direction that has received much attention.
[0003] Currently, mainstream portrait generation methods have achieved good results in single-modal input scenarios, but there are still obvious bottlenecks in dealing with multi-source heterogeneous information, making them show obvious limitations in complex scenarios involving multi-source information, especially in tasks that require advanced multi-modal collaboration. For example, most theme-driven methods mainly face single-image input, and alternative solutions such as λ-ECLIPSE that support sequential multi-image input are still cumbersome for users; in addition, existing portrait generation technologies generally rely on a strict text-image pairing input mechanism, which not only increases the system redundancy but also easily leads to the neglect of certain control conditions during the generation process, resulting in poor flexibility and accuracy in generating images under decoupled multi-modal input conditions. Summary of the Invention
[0004] To solve the problem of poor flexibility and accuracy in generating images under decoupled multi-modal input conditions in the above-mentioned existing technologies, the present invention proposes a multi-modal controllable portrait generation method and system, which can effectively improve the flexibility and accuracy of generating images under decoupled multi-modal input conditions.
[0005] To achieve the above technical effects, the technical solution of the present invention is as follows:
[0006] A multi-modal controllable portrait generation method, comprising the following steps:
[0007] S1. Decouple the multi-modal input conditions of the original portrait picture to obtain a multi-modal decoupling result;
[0008] S2. Preprocess the multi-modal decoupling result to obtain an image embedding result and a text embedding result;
[0009] S3. Concatenate the image embedding results to obtain a concatenated embedding;
[0010] S4. Input the spliced embedding and the text embedding result into a preset portrait generation network to output a controllable portrait generation result.
[0011] Preferably, the multi-modal input condition decoupling of the original portrait picture includes text prompt condition decoupling, image condition decoupling, and hand-drawn layout condition decoupling, and the multi-modal decoupling result includes a text prompt, a reference image set, and a hand-drawn layout set.
[0012] Preferably, the text prompt condition decoupling includes:
[0013] S111. Set a standard formatted query condition;
[0014] S112. Input the query condition and the original portrait picture into a preset vision-language model, and use the vision-language model to output the text prompt.
[0015] Preferably, the image condition decoupling includes:
[0016] S121. Input the original portrait picture into a preset reference image segmentation network, and the reference image segmentation network outputs reference images of different parts in the original portrait picture;
[0017] S122. Perform data augmentation on the reference image to obtain the reference image set.
[0018] Preferably, the hand-drawn layout condition decoupling includes:
[0019] S131. Obtain the hand-drawn layout of the original portrait picture;
[0020] S132. Extract the contour coordinates of the color blocks in the hand-drawn layout, and fit the contour coordinates into an ellipse and a rectangle. The calculation expressions for the center (x c , y c ), the major semi-axis a, and the minor semi-axis b of the ellipse are as follows:
[0021]
[0022] where, (x min , y min ) and (x max , y max ) are the coordinates of the minimum bounding rectangle of the enclosed area respectively;
[0023] The calculation expressions for the coordinates (x i , y i ) of the four corners of the rectangle are as follows:
[0024] x i = x r + cos(θr +α i )·d i
[0025] y i =y r +sin(θ r +α i )·d i
[0026] where (x r , y r ) represents the center coordinates of the matrix, θ r represents the rotation angle, α i represents the angle from the center to each corner of the rectangle, and d i represents the distance from the center to each corner.
[0027] S133. Combine the ellipse and the rectangle to form the hand-drawn layout set.
[0028] Preferably, the preprocessing of the multi-modal decoupling result includes:
[0029] S211. Define the set of the original portrait pictures as H i , the reference image set as R i , the hand-drawn layout set as L i , and the text prompt as T i ;
[0030] S212. Input the set of the original portrait pictures H i , the reference image set R i , and the hand-drawn layout set L i into the VAE encoder, and the VAE encoder maps H i , R i , and L i to the latent space, and the obtained image embedding results are as follows:
[0031] z human =E(H i ), z ref =E(R i ), z layout =E(L i )
[0032] where z human represents the latent vector of the set of the original portrait pictures, z ref represents the latent vector of the reference image set, z layout represents the latent vector of the hand-drawn layout set, E(.) represents the VAE encoder, and denote z human , z ref and zlayout is the image embedding result;
[0033] and input the text prompt T i into the CLIP encoder, and the calculation expression for outputting the text embedding result Y is as follows:
[0034] Y = CLIP(T i )
[0035] where CLIP(.) represents the CLIP encoder.
[0036] Preferably, the splicing of the image embedding result includes:
[0037] S311. Concatenate the latent vector z human of the set of original portrait pictures and the latent vector z ref of the reference image set in the selected spatial dimension to obtain the target latent vector Z gt The calculation expression is as follows:
[0038] Z gt = Concat(z human , z ref )
[0039] S312. Set part of the latent vector z ref of the reference image set to zero based on the Drop mask to obtain the reference latent vector Concatenate the reference latent vector and the latent vector z layout of the hand-drawn layout set in the selected spatial dimension to obtain the source latent vector Z src The calculation expression is as follows:
[0040]
[0041] S313. Concatenate the target latent vector Z gt and the source latent vector Z src at the channel level to obtain the final concatenated embedding Z′. The calculation expression is as follows:
[0042]
[0043] Preferably, the portrait generation network is a U-Net network, and the U-Net network is trained using the total loss function L total The calculation expression of the total loss function L total is as follows:
[0044] L total = L MSE ·wi
[0045] Among them, L MSE represents the mean squared error loss function, and w i represents the weight calculated based on the signal-to-noise ratio at time step t i ;
[0046] The calculation expression of the mean squared error loss function L MSE is as follows:
[0047]
[0048] Among them, N represents the batch size, represents the predicted noise, and ∈ i represents the actual noise;
[0049] The weight w i calculated based on the signal-to-noise ratio at time step t i has the following calculation expression:
[0050]
[0051] The total loss function L total has the following calculation expression for backpropagation through the U-Net network:
[0052]
[0053] Among them, θ represents the network parameters, η represents the learning rate, represents the total loss gradient with respect to the network parameters;
[0054] The calculation expression of the total loss gradient is as follows:
[0055]
[0056] Among them, max_grad_norm represents the threshold for gradient clipping.
[0057] Preferably, the output of the controllable portrait generation result includes:
[0058] S411. Define C = {C1, C2,..., C N} to represent N color blocks in the hand-drawn layout set L i , each color block corresponds to a unique key area, and calculate the binary mask M i of each color block as follows:
[0059]
[0060] Among them, C iDenote the \(i\)-th color patch in set \(C\), and \((x, y)\) represents the coordinates of the color patch;
[0061] S412. Calculate the cross-attention map \(A\) for each text description in the text description set (i) as follows:
[0062] \(A\) (i) = Attention(\(Q\) i , \(K\) i , \(V\) i )
[0063] where Attention(.) represents the cross-attention extraction function, \(Q\) i represents the query vector extracted from the input image features, \(K\) i represents the key corresponding to the text description, and \(V\) i represents the value vector corresponding to the text description;
[0064] S413. Modulate the cross-attention map \(A\) i using the binary mask \(M\) (i) to obtain the cross-attention modulation map The calculation expression is as follows:
[0065]
[0066] where \(\odot\) represents element-wise multiplication;
[0067] S414. Inject the cross-attention modulation map into the portrait generation network to guide the portrait generation network to output a controllable portrait generation result affected by the cross-attention modulation map .
[0068] The present invention also proposes a multi-modal controllable portrait generation system, including:
[0069] A multi-modal input condition decoupling module for decoupling the multi-modal input conditions of the original portrait picture to obtain a multi-modal decoupling result;
[0070] A preprocessing module for preprocessing the multi-modal decoupling result to obtain an image embedding result and a text embedding result;
[0071] A splicing module for splicing the image embedding results to obtain a spliced embedding;
[0072] An output module for inputting the spliced embedding and the text embedding result into a preset portrait generation network and outputting a controllable portrait generation result.
[0073] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0074] The present invention proposes a multi-modal controllable portrait generation method and system. First, the multi-modal input conditions of the original portrait image are decoupled, realizing the independent feature extraction of the multi-modal input in the original portrait image, avoiding the limitation of the input conditions, and enabling the unified regulation of the portrait generation process for multi-modal input conditions. Secondly, the image embedding result and the text embedding result are obtained through the preprocessing of the multi-modal decoupling result, which not only retains the detailed information of the visual features but also integrates the semantic constraints of the text description to form a multi-dimensional feature representation. Furthermore, by splicing the image embedding results, the spatial alignment and weight allocation of cross-modal features are realized, avoiding the semantic conflicts between multi-modal inputs. Finally, by inputting the spliced embedding and text embedding results into a preset portrait generation network, the controllable recombination of the decoupled features of multi-modal input conditions is realized under the framework of the diffusion model, effectively improving the flexibility and accuracy of the generated images. Brief Description of the Drawings
[0075] Figure 1 It represents a flowchart of a multi-modal controllable portrait generation method proposed in an embodiment of the present invention;
[0076] Figure 2 It represents a schematic diagram of the dataset preparation process proposed in an embodiment of the present invention;
[0077] Figure 3 It represents an example diagram of text description proposed in an embodiment of the present invention;
[0078] Figure 4 It represents a schematic diagram of the data training process proposed in an embodiment of the present invention;
[0079] Figure 5 It represents a schematic diagram of generating a human image under decoupled multi-modal conditions proposed in an embodiment of the present invention;
[0080] Figure 6 It represents a qualitative comparison diagram of a multi-modal controllable portrait generation method proposed in an embodiment of the present invention and a theme-driven benchmark;
[0081] Figure 7 It represents a qualitative comparison diagram of a multi-modal controllable portrait generation method proposed in an embodiment of the present invention and a layout-guided text-to-image benchmark;
[0082] Figure 8 It represents a schematic diagram of the visual result proposed in an embodiment of the present invention;
[0083] Figure 9 It represents a schematic diagram of an additional visual result proposed in an embodiment of the present invention;
[0084] Figure 10The figure shows a structural block diagram of a multi-modal controllable portrait generation system proposed in an embodiment of the present invention. DETAILED DESCRIPTION
[0085] The drawings are for illustrative purposes only and are not to be construed as limiting the present patent;
[0086] It is understandable to those skilled in the art that some well-known contents may be omitted in the drawings;
[0087] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0088] Example 1
[0089] like Figure 1 As shown, this embodiment proposes a multi-modal controllable portrait generation method, comprising the following steps:
[0090] S1. Decoupling the original portrait image with multimodal input conditions to obtain a multimodal decoupling result;
[0091] In S1, existing portrait generation methods support joint image and text input and rely on paired formats, with two main purposes: (1) to enhance generation performance through the interaction of text descriptions or specified categories; (2) to enable text-based editing or migration of specific subject images. In contrast, this step proposes an innovative method to independently align each modality by decoupling multimodal input conditions. In addition, this step introduces hand-drawn layout to further control the spatial arrangement. Figure 2 As shown, the multimodal input condition decoupling of the original portrait image includes text prompt condition decoupling, image condition decoupling and hand-drawn layout condition decoupling. The method of the present invention supports three decoupled input conditions, and the multimodal decoupling results include text prompts, reference image sets and hand-drawn layout sets.
[0092] The text prompt condition decoupling includes:
[0093] S111. Set standard formatted query conditions;
[0094] S112. Input the query condition and the original portrait image into a preset visual language model, and use the visual language model to output the text prompt.
[0095] Currently, visual language models have shown excellent capabilities in image annotation, providing high accuracy and exceptionally clear description details. The present invention uses the visual language model CogVLM2 to expand the fine-grained text description of each person image, and processes attributes such as face, tops, bottoms, and shoes through standardized and formatted component-level query conditions. These attributes are combined and converted into the final descriptive sentence output, such as Figure 3As shown, an example of a fine-grained text description of the top attributes. In the present invention, CogVLM2 is used to extract various attributes corresponding to each component. Finally, these attributes are combined and transformed into a complete descriptive sentence.
[0096] The image condition decoupling includes:
[0097] S121. Input the original portrait picture into a preset reference image segmentation network, and the reference image segmentation network outputs the reference images of different components in the original portrait picture;
[0098] S122. Perform data augmentation on the reference image to obtain the reference image set.
[0099] Among them, the different components in the original portrait picture include hair, face, top, bottom, and shoes. To achieve the segmentation of different components in the original portrait picture, in this step, two image segmentation models, SAM and SCHP, are combined and compared through cross-validation. First, the present invention calculates the structural similarity index (SSIM) to evaluate the similarity between the component masks extracted by the two models. If the similarity exceeds 0.75, it is considered that the extraction is successful, and the smaller mask is preferentially selected; if it is lower than 0.75, it is marked as a failure, and the data will enter the manual review or data cleaning process. Subsequently, the present invention performs data augmentation on the extracted reference image, including rotation, flipping, and scaling, to obtain the reference image set.
[0100] Based on the SCHP image segmentation model, the present invention calculates the contour coordinates of the color blocks through S3 and fits them into basic shapes such as ellipses and rectangles to construct a coarse-grained hand-drawn layout set. Mathematically, the basic shapes are classified into ellipses and rectangles, and this distinction enhances user-friendliness and accessibility. More specifically, the hand-drawn layout condition decoupling includes:
[0101] S131. Obtain the hand-drawn layout of the original portrait picture;
[0102] S132. Extract the contour coordinates of the color blocks in the hand-drawn layout, and fit the contour coordinates into ellipses and rectangles, where the center (x c , y c ) of the ellipse, the major semi-axis a, and the minor semi-axis b are calculated as follows:
[0103]
[0104] Among them, (x min , y min ) and (x max , y max ) are the coordinates of the minimum bounding rectangle of the enclosed area;
[0105] The calculation expressions for the coordinates (x i , y i ) of the four corners of the rectangle are as follows:
[0106] x i = x r + cos(θ r + α i )·d i
[0107] y i = y r + sin(θ r + α i )·d i
[0108] Among them, (x r , y r ) represents the central coordinates of the matrix, θ r represents the rotation angle, α i represents the angle from the center to each corner of the rectangle, and d i represents the distance from the center to each corner.
[0109] S133. Combine the ellipse and the rectangle to form the hand-drawn layout set.
[0110] Through the above strategy, the present invention has completed the preparation work of three modalities: text, image, and layout, thus constructing the basis of the ComposeHuman dataset; in order to achieve decoupling, the present invention introduces a random "text or image" input modality. In this process, the present invention randomly selects several reference images from the corresponding human body part set and checks the relevant labels to extract and delete their associated text descriptions. This process establishes the comprehensive input conditions of the present invention.
[0111] S2. Preprocess the multi-modal decoupling result to obtain an image embedding result and a text embedding result;
[0112] S3. Concatenate the image embedding results to obtain a concatenated embedding;
[0113] S4. Input the concatenated embedding and the text embedding result into a preset portrait generation network to output a controllable portrait generation result.
[0114] In this embodiment, the present invention allows the use of text or reference images to decouple and control any component in the hand-drawn portrait layout, and seamlessly integrates these control factors during the generation process. The hand-drawn layout uses color-block geometric shapes such as ellipses and rectangles, which can be easily drawn by users, thus providing a more flexible and accessible way to define the spatial layout. In addition, the present invention also proposes the ComposeHuman dataset, which provides decoupled text and reference image annotations for different human body components, making it have a wider application in the portrait generation task. Through large-scale experiments on multiple datasets, the results show that the present invention can generate portrait images that better conform to the given layout, text description, and reference image, demonstrating its multi-tasking ability and controllability.
[0115] Embodiment 2
[0116] This embodiment further elaborates on the steps of a multi-modal controllable portrait generation method proposed in the above embodiment. Although many layout-to-image methods have demonstrated excellent performance, they mainly rely on text as the sole auxiliary condition. In addition, most existing theme-driven methods are limited to a single image reference. In contrast, the method of the present invention supports multi-image references and can generate multi-modal and controllable portraits. Refer to Figure 2 and Figure 4 , the present invention uses CogVLM2, SAM, and SCHP to enrich the image-based try-on dataset through fine-grained text descriptions, hand-drawn layouts, and human body component sets. By decoupling multi-modal conditions, the present invention realizes portrait generation in text-only, image-only, and hybrid modalities by seamlessly stitching input information in the spatial and channel dimensions. More specifically, the preprocessing of the multi-modal decoupling result in S2 includes:
[0117] S211. Define the set of original portrait pictures as H i ={h1, h2,..., h n}, the reference image set as R i ={r1, r2,..., r n}, the hand-drawn layout set as l i ={l1, l2,..., l n}, and the text prompt as T i ={t1, t2,..., t n}; where h n is an element of H i , r n is an element of R i , l n is an element of L i , and t n is an element of T i ;
[0118] For the input containing multiple reference images, the present invention performs non-overlapping stitching at the pixel level and then executes S212;
[0119] S212. Input the set H i of the original portrait pictures, i the reference image set R i and the hand-drawn layout set L i into the VAE encoder, and the VAE encoder maps H i , R i and L
[0120] to the latent space, and the image embedding results are respectively as follows: human z i = E(H ref ), z i = E(R layout ), z i = E(L
[0121] where z human represents the latent vector of the set of the original portrait pictures, z ref represents the latent vector of the reference image set, z layout represents the latent vector of the hand-drawn layout set, E(.) represents the VAE encoder, and denote z human , z ref and z layout as the image embedding results;
[0122] And input the text prompt T i into the CLIP encoder, and the calculation expression for outputting the text embedding result Y is as follows:
[0123] Y = CLIP(T i )
[0124] where CLIP(.) represents the CLIP encoder.
[0125] S3 The splicing of the image embedding results includes:
[0126] S311. Splice the latent vector z human of the set of the original portrait pictures and the latent vector z ref of the reference image set in the selected -1 spatial dimension to obtain the target latent vector Z gt The calculation expression is as follows:
[0127] Z gt = Concat(z human , z ref )
[0128] S312. Set the latent vector z of the partial reference image set to zero based on the Drop mask to obtain the reference latent vector ref On the selected spatial dimension, splice the reference latent vector and the latent vector z of the hand-drawn layout set layout to obtain the source latent vector Z src The calculation expression is as follows:
[0129]
[0130] S313. Splice the target latent vector Z gt and the source latent vector Z src at the channel level to obtain the calculation expression of the final spliced embedding Z′ as follows:
[0131] Z′ = Concat(Z gt , Z src ).
[0132] S4 The portrait generation network is a U-Net network, and the total loss function L total is used to train the U-Net network. The calculation expression of the total loss function L total is as follows:
[0133] L total = L MSE · w i
[0134] where L MSE represents the mean squared error loss function, and w i represents the weight calculated based on the signal-to-noise ratio at time step t i ;
[0135] The calculation expression of the mean squared error loss function L MSE is as follows:
[0136]
[0137] where N represents the batch size, represents the predicted noise, and ∈ i represents the actual noise;
[0138] The calculation expression of the weight w i calculated based on the signal-to-noise ratio at time step t i is as follows:
[0139]
[0140] The total loss function L total The computational expression for backpropagation through the U-Net network is as follows:
[0141]
[0142] where θ represents the network parameters, η represents the learning rate, represents the total loss gradient with respect to the network parameters;
[0143] The computational expression for the total loss gradient is as follows:
[0144]
[0145] where max_grad_norm represents the threshold for gradient clipping.
[0146] In the inference stage, the present invention combines the cross-attention control obtained from the hand-drawn layout to improve the quality and accuracy of the generated text description. More specifically, the output of the controllable portrait generation result includes:
[0147] S411. Define C = {C1, C2,..., C N} to represent the N color blocks in the hand-drawn layout set L i Each color block corresponds to a unique key area, and calculate the binary mask M i of each color block as follows:
[0148]
[0149] where C i represents the i-th color block in the set C, and (x, y) represents the coordinates of the color block;
[0150] S412. Calculate the cross-attention map A (i) of each text description in the text description set as follows:
[0151] A (i) = Attention(Q i , K i , V i )
[0152] where Attention(.) represents the cross-attention extraction function, Q i represents the query vector extracted from the input image features, K i represents the key corresponding to the text description, and V i represents the value vector corresponding to the text description;
[0153] S413. Modulate the cross-attention map A i using the binary mask M (i) to obtain the cross-attention modulation map The calculation expression is as follows:
[0154]
[0155] Where ⊙ represents element-wise multiplication, which is used to increase or decrease attention in the white area of the mask; this modulation adjusts attention by enhancing or suppressing the attention distribution in specific regions;
[0156] S414. Inject the cross-attention modulation mapping into the portrait generation network, guiding the portrait generation network to output a controllable portrait generation result affected by the cross-attention modulation mapping so as to improve the quality and accuracy of the generated text description.
[0157] In summary, to address the limitations of existing methods, the present invention proposes an innovative multi-modal portrait generation method, aiming to enhance flexibility and accuracy under decoupled multi-modal input conditions. The present invention introduces the concept of "hand-draw layouts", allowing users to draw with colored blocks composed of simple geometric shapes (such as ellipses and rectangles) to specify the spatial positions of various human body parts. This method uses unique color information to divide regions, effectively avoiding feature confusion between different human body parts, thereby achieving more precise spatial layout control.
[0158] The core innovation of the present invention lies in a flexible multi-modal input mechanism that allows users to describe human body parts through text or images, and at the same time supports unpaired input methods, such as Figure 5 shown. In addition, the present invention supports pixel-level fusion of multiple reference images, eliminating the cumbersome process of sequential input processing. On this basis, the present invention designs a data decoupling pipeline that integrates text, reference images, and hand-drawn layouts. By spatially aligning latent features, the generated images are more consistent with the expected multi-modal conditions in terms of content and structure. At the same time, the present invention applies attention modulation during the inference process to further enhance spatial consistency and text semantic consistency. In addition, the present invention constructs a multi-modal dataset ComposeHuman, covering portrait images, hand-drawn layouts, fine-grained text descriptions, and human body part assembly information. Experimental results show that the present invention is superior to existing methods in terms of flexibility and quality, especially performing well in diverse portrait generation tasks. Its high degree of customizability also enables users to add or remove accessories (such as hats, bags) in real time by adjusting the hand-drawn layout, thus realizing personalized image creation.
[0159] Example 3
[0160] To verify the effectiveness of the multi-modal controllable portrait generation method proposed in the present invention, the method of the present invention and the prior art method are respectively subjected to qualitative comparison and quantitative comparison;
[0161] For the qualitative comparison, it includes:
[0162] The present invention comparatively analyzes the generation methods centered on theme-driven and layout guidance, and the results are as Figure 6 and Figure 7 shown. Figure 6 For the qualitative comparison with the theme-driven benchmark, the present invention shows a high degree of accuracy in matching specific features of a given reference clothing image; Figure 7 For the qualitative comparison with the layout-guided text-to-image benchmark, the present invention shows a high degree of consistency with the text description and spatial layout arrangement in the generated output. Although the GLIGEN method, InstanceDiffusion method, and MIGC method can approximate the spatial layout to a certain extent, they are insufficient in precisely controlling human poses. Among them, the InstanceDiffusion method often generates unrealistic human figures, while the DenseDiffusion method often ignores key semantic details. In contrast, the present invention can better conform to the layout specifications and text input. In addition, in the case of providing a clothing reference image, the present invention can always capture delicate visual details, and its performance is significantly better than the AnyDoor method and the IP-Adapte method. See Figure 8 , which further demonstrates the ability of the present invention in generating realistic human images, and its generated results can achieve seamless alignment with multi-modal inputs.
[0163] For the quantitative analysis, it includes: text-to-human generation based on layout - the present invention conducts a quantitative comparison with existing layout-guided text-to-image generation methods on the VITON-HD dataset. As shown in Table 1, the present invention proposed in the present invention is superior to the baseline model in all metrics and can generate human images that are consistent with the hand-drawn layout space and match the text description. This shows that the method has significant advantages in terms of robustness and diversity. Quantitative comparison with other layout-guided methods. The present invention compares various metrics on the VITON-HD dataset. The best and second-best results are shown in bold and underlined respectively.
[0164] Table 1 Quantitative comparison table of the method of the present invention and existing layout-guided text-to-image generation methods
[0165]
[0166] Subject-driven Human Generation - The present invention has been quantitatively compared with advanced subject-driven image generation methods on the DressCode dataset. As shown in Table 2, the method of the present invention performs excellently in all metrics. This highlights the powerful ability of the present invention to accurately locate each component in the target image while precisely retaining the subject features. Quantitative comparison with other subject-driven methods. The present invention compares various metrics on the DressCode dataset. The best and second-best results are shown in bold and underlined respectively.
[0167] Table 2 Quantitative Comparison Table of the Method of the Present Invention and Advanced Subject-driven Image Generation Methods
[0168]
[0169] In summary, the present invention has the following advantages:
[0170] 1. An innovative controllable layout-to-human generation method (Layout-to-Human) is proposed, which uses hand-drawn geometric shapes as layout information and combines text and image references to describe different human body parts, thereby generating highly consistent and realistic human portrait images in text-only, image-only, or mixed-modal tasks.
[0171] 2. A multi-modal decoupled layout-to-human dataset ComposeHuman is constructed. By annotating text labels for each body part and adopting a semi-supervised reference image extraction method to obtain decoupled multi-modal references. In addition, the present invention converts traditional segmentation into hand-drawn layouts to achieve more flexible and detailed spatial organization.
[0172] 3. A large number of layout-guided generation and subject-driven generation task experiments on the VITON-HD, DressCode, and DeepFashion datasets show that the present invention can generate human portrait images that better conform to the given layout, text description, and reference image, demonstrating its multi-task generation ability and controllability.
[0173] See Figure 9 , which shows visual examples of human image generation guided by hand-drawn layouts in a multi-modal input framework. Through the given hand-drawn layout as a guide, a decoupled multi-modal input of choosing one from two of image and text is performed for multiple attributes of the human portrait (face, upper garment, lower garment, and shoes), and the final output human portrait result is obtained. These results show the effectiveness of the present invention in processing decoupled multi-modal inputs, making it possible to generate real and high-quality human images. The generated output is precisely aligned with the spatial configuration of the hand-drawn layout, faithfully follows the details specified in the text description, and accurately captures the unique features of the reference image.
[0174] Example 4
[0175] See Figure 10 , this embodiment proposes a multi-modal controllable portrait generation system, including:
[0176] A multi-modal input condition decoupling module, which is used to decouple the multi-modal input conditions of the original portrait picture to obtain a multi-modal decoupling result;
[0177] A preprocessing module, which is used to preprocess the multi-modal decoupling result to obtain an image embedding result and a text embedding result;
[0178] A splicing module, which is used to splice the image embedding results to obtain a spliced embedding;
[0179] An output module, which is used to input the spliced embedding and the text embedding result into a preset portrait generation network and output a controllable portrait generation result.
[0180] In this embodiment, first, the multi-modal input conditions of the original portrait picture are decoupled, realizing the independent feature extraction of the multi-modal input in the original portrait picture, avoiding the limitation of the input conditions, and enabling the portrait generation process to uniformly regulate the multi-modal input conditions; second, the image embedding result and the text embedding result are obtained through the preprocessing of the multi-modal decoupling result, which not only retains the detailed information of the visual features but also integrates the semantic constraints of the text description to form a multi-dimensional feature representation; furthermore, by splicing the image embedding results, the spatial alignment and weight allocation of cross-modal features are realized, avoiding the semantic conflict between multi-modal inputs; finally, by inputting the spliced embedding and the text embedding result into a preset portrait generation network, the controllable recombination of the decoupled features of the multi-modal input conditions is realized under the diffusion model framework, effectively improving the flexibility and accuracy of the generated images.
[0181] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A multi-modal controllable portrait generation method, characterized in that: The following steps are involved: S1. Decoupling the original portrait image with multimodal input conditions to obtain a multimodal decoupling result; S2. preprocessing the multimodal decoupling result to obtain an image embedding result and a text embedding result; S3. splicing the image embedding results to obtain a spliced embedding; S4. Input the concatenated embedding and the text embedding results into a preset portrait generation network, and output a controllable portrait generation result.
2. The multi-modal controllable portrait generation method according to claim 1, characterized in that: The multimodal input condition decoupling of the original portrait image includes text prompt condition decoupling, image condition decoupling and hand-drawn layout condition decoupling, and the multimodal decoupling result includes text prompt, reference image set and hand-drawn layout set.
3. The multi-modal controllable portrait generation method according to claim 2, characterized in that: The text prompt condition decoupling includes: S111. Set standard formatted query conditions; S112. Input the query condition and the original portrait image into a preset visual language model, and use the visual language model to output the text prompt.
4. The multi-modal controllable portrait generation method according to claim 2, characterized in that: The image condition decoupling includes: S121. Inputting the original portrait image into a preset reference image segmentation network, and outputting reference images of different parts in the original portrait image by the reference image segmentation network; S122. Perform data enhancement on the reference image to obtain the reference image set.
5. The multi-modal controllable portrait generation method according to claim 2, characterized in that: The hand-drawn layout condition decoupling includes: S131. Obtaining the hand-drawn layout of the original portrait image; S132. Extract the outline coordinates of the color block in the hand-painted layout, and fit the outline coordinates into an ellipse and a rectangle, where the center (x c ,y c ), the calculation expressions of the major semi-axis a and the minor semi-axis b are as follows: Among them, (x min ,y min ) and (x max ,y max ) are the coordinates of the minimum circumscribed rectangle of the enclosing area; The coordinates of the four corners of the rectangle (x i ,y i ) is calculated as follows: x i =x r +cos(θ r +a i )·d i y i =y r +sin(θ r +a i )·d i Among them, (x r ,y r ) represents the center coordinate of the matrix, θ r represents the rotation angle, α i represents the angle from the center to each corner of the rectangle, d i Represents the distance from the center to each corner. S133. Combining the ellipse and rectangle into the hand-drawn layout set.
6. The multi-modal controllable portrait generation method according to claim 2, characterized in that: The preprocessing of the multimodal decoupling result includes: S211. Define the set of original portrait images as H i , the reference image set is R i , the hand-drawn layout set is L i , the text prompt is T i ; S212. The set H of the original portrait pictures i , the reference image set R i And the hand-drawn layout set L i Input to the VAE encoder, the VAE encoder converts H i , R i and L i Mapped to the latent space, the image embedding results are as follows: z human =E(H i ),z ref =E(R i ),z layout =E(L i ) Among them, z human The latent vector representing the set of original portrait images, z ref represents the latent vector of the reference image set, z layout represents the latent vector of the hand-drawn layout set, E(.) represents the VAE encoder, and z human 、z ref and z layout embedding a result for the image; And the text prompt T i Input to the CLIP encoder, the calculation expression of outputting the text embedding result Y is as follows: Y=CLIP(T i ) Among them, CLIP(.) represents the CLIP encoder.
7. The multi-modal controllable portrait generation method according to claim 6, characterized in that: The step of splicing the image embedding results comprises: S311. The latent vector z of the set of the original portrait pictures is transformed into human and the latent vector z of the reference image set ref Splice and get the target potential vector Z gt The calculation expression is as follows: From ge =Concat(from human ,from ref ) S312. Based on the Drop mask, the potential vector z of the partial reference image set is ref Set to zero to get the reference potential vector The reference latent vector is transformed into and the latent vector z of the hand-drawn layout set layout Splice and get the source potential vector Z src The calculation expression is as follows: S313. The target potential vector Z gt and the source latent vector Z src The concatenation is performed at the channel level, and the calculation expression of the final concatenated embedding Z′ is as follows: Z′=Concat(Z gt ,Z src )。 8. The multi-modal controllable portrait generation method according to claim 7, characterized in that: The portrait generation network is a U-Net network, using the total loss function L total The U-Net network is trained, and the total loss function L total The calculation expression is as follows: L total =L MSE ·w i Among them, L MSE represents the mean square error loss function, w i Represents the time step t i The weight of the signal-to-noise ratio calculation; The mean square error loss function L MSE The calculation expression is as follows: Where N is the batch size, represents the prediction noise, ∈ i represents the actual noise; The time step t i The weight w of the signal-to-noise ratio calculation i The calculation expression is as follows: The total loss function L total The computational expression for back propagation through the U-Net network is as follows: Among them, θ represents the network parameters, η represents the learning rate, represents the total loss gradient with respect to the network parameters; The calculation expression of the total loss gradient is as follows: Among them, max_grad_norm represents the threshold of gradient clipping.
9. The multi-modal controllable portrait generation method according to claim 7, characterized in that: The output controllable portrait generation result includes: S411. Define C = {C1, C2, ..., C N } represents the hand-drawn layout set L i There are N color blocks in the image, each of which corresponds to a unique key area, and the binary mask M of each color block is calculated. i as follows: Among them, C i represents the i-th color block in the set C, (x, y) represents the coordinates of the color block; S412. Calculate the cross attention map A of each text description in the text description set (i) as follows: A (i) =Attention(Q i ,K i ,V i ) Among them, Attention(.) represents the cross attention extraction function, Q i represents the query vector extracted from the input image features, K i Indicates the key corresponding to the text description, V i A vector of values representing the corresponding text description; S413. Using the binary mask M i For the cross attention map A (i) Modulate and get the cross attention modulation map The calculation expression is as follows: Among them, ⊙ represents element-wise multiplication; S414. Mapping the cross attention modulation Injected into the portrait generation network, guiding the portrait generation network output to be modulated by the cross attention mapping The controllable portrait generation results affected by this.
10. A multi-modal controllable portrait generation system, characterized in that: include: A multimodal input condition decoupling module is used to decouple the multimodal input conditions of the original portrait image to obtain a multimodal decoupling result; A preprocessing module, used for preprocessing the multimodal decoupling result to obtain an image embedding result and a text embedding result; A splicing module, used for splicing the image embedding results to obtain a spliced embedding; An output module is used to input the spliced embedding and the text embedding results into a preset portrait generation network and output a controllable portrait generation result.