Text-based Multimodal Face Generation Method, Device, Equipment, and Storage Medium

By constructing a multimodal-guided identity constraint framework, combining the decoupling of global identity embedding features and multimodal local identity embedding features, the problems of identity consistency and detail learning in text-to-face generation are solved, and personalized face generation with stable identity consistency is achieved.

CN119722837BActive Publication Date: 2025-07-22BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411715791.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-07-22
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

The prior art is difficult to balance text-image and identity consistency in text-to-face generation, resulting in poor identity consistency and difficult for the generation model to learn facial details.

Method used

By decoupling the combination of global identity embedding features and multimodal local identity embedding features, a multimodal-guided identity constraint framework is built, and the mask of residual cross attention balances loss is used to separate identity information and background, and fine-grained semantic editing is achieved.

Benefits of technology

Improve the identity consistency and detail retention ability of text-to-face generation. Users can adjust local details through text to broaden the application scenarios, and the generated images maintain stable identity consistency under the prompts of diversified text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722837B_ABST
    Figure CN119722837B_ABST
Patent Text Reader

Abstract

The present disclosure provides a text-based multi-modal face generation method, apparatus, device, and storage medium, belonging to the technical field of face image generation. The method includes: determining a subject image based on a reference image and a subject mask corresponding to the reference image, and determining a decoupled global identity embedding feature based on the subject image. The reference image is an initial face image. Determining a multi-modal local identity embedding feature based on the reference image and a mask image corresponding to the reference image. The multi-modal local identity embedding feature is a text embedding-like feature. Determining a target generated face image based on the decoupled global identity embedding feature and the multi-modal local identity embedding feature. The text-based multi-modal face generation method, apparatus, device, and storage medium provided by the present disclosure can improve the accuracy of text-to-face generation and meet actual requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of face image generation, and more specifically, relates to a text-based multi-modal face generation method, device, equipment, and storage medium. Background Art

[0002] With the continuous development of artificial intelligence technology, especially the breakthroughs in the field of deep learning, the technology of generating images from text has made further progress. These technologies can generate corresponding images based on given text descriptions, including face generation. The personalized text-to-face generation technology has further developed on this basis, aiming to generate face images with specific identity features and personalized elements. The personalized text-to-face generation technology has important value in fields such as artistic creation, 3D human body generation, and image animation.

[0003] However, the existing technologies still face some problems. For example, the method based on fine-tuning can transform visual concepts, but it is time-consuming and prone to overfitting, reducing the editability of semantic prompts. Due to the characteristics of the text embedding space, it is difficult for the generation model to balance text-image and identity consistency, and it cannot meet the actual needs. The second is the method based on training a general encoder, which relies on a large-scale dataset to train the mapping network. Although it improves the inference efficiency, it is difficult to learn facial details, resulting in poor identity consistency in text-to-face generation. Summary of the Invention

[0004] The purpose of the present disclosure is to provide a text-based multi-modal face generation method, device, equipment, and storage medium to improve the identity consistency in text-to-face generation and meet the actual needs.

[0005] In the first aspect of the embodiments of the present disclosure, a text-based multi-modal face generation method is provided, including:

[0006] Determine a subject image based on a reference image and a subject mask corresponding to the reference image, and determine a decoupled global identity embedding feature based on the subject image. The reference image is an initial face image.

[0007] Determine a multi-modal local identity embedding feature based on the reference image and a mask image corresponding to the reference image. The multi-modal local identity embedding feature is a text embedding type feature.

[0008] Determine a target generated face image based on the decoupled global identity embedding feature and the multi-modal local identity embedding feature.

[0009] In the second aspect of the embodiments of the present disclosure, a text-based multi-modal face generation device is provided, including:

[0010] A first computing module, configured to determine a subject image based on a reference image and a subject mask corresponding to the reference image, and determine a decoupled global identity embedding feature based on the subject image. The reference image is an initial face image.

[0011] A second computing module, configured to determine a multi-modal local identity embedding feature based on the reference image and a mask image corresponding to the reference image. The multi-modal local identity embedding feature is a text embedding-like feature.

[0012] A face generation module, configured to determine a target generated face image based on the decoupled global identity embedding feature and the multi-modal local identity embedding feature.

[0013] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned text-based multi-modal face generation method are implemented.

[0014] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned text-based multi-modal face generation method are implemented.

[0015] The beneficial effects of the text-based multi-modal face generation method, device, equipment, and storage medium provided by the embodiments of the present disclosure are as follows:

[0016] On the one hand, the embodiments of the present disclosure construct a multi-modal guided identity constraint framework by decoupling the global identity embedding feature and the multi-modal local identity embedding feature, which can more accurately capture identity information at different levels in the face image, ensure the stability of identity features and the retention of facial details, and have semantic editing capabilities. The combination of the decoupled global identity embedding feature and the multi-modal local identity embedding feature makes the generated face image more accurate and rich in restoring identity features. In addition, by using text information to determine the multi-modal local identity embedding feature, users can precisely adjust the local details of the generated face by inputting specific text, broadening the application scenarios of face generation.

[0017] On the other hand, the embodiments of the present disclosure utilize a mask balance loss based on residual cross-attention (corresponding to relevant parts of the subject mask and the mask image), which can effectively separate identity information from non-identity-related backgrounds and balance the influence of text prompts and identity constraints.

[0018] On yet another hand, the embodiments of the present disclosure also achieve smooth control of fine-grained semantic editing.

[0019] In summary, the embodiments of the present disclosure implement a method for personalized text-to-face generation that ensures stable identity consistency based on multimodal learning, enabling the generated face results to ensure stable subject identity consistency on the basis of aligning with diverse text prompts. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of a text-based multimodal face generation method provided by an embodiment of the present disclosure;

[0022] Figure 2 It is a schematic flowchart of a text-based multimodal face generation method provided by another embodiment of the present disclosure;

[0023] Figure 3 It is a schematic diagram for comparing face image generation provided by an embodiment of the present disclosure;

[0024] Figure 4 It is a structural block diagram of a text-based multimodal face generation device provided by an embodiment of the present disclosure;

[0025] Figure 5 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0027] To make the objectives, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the accompanying drawings.

[0028] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a text-based multimodal face generation method provided by an embodiment of the present disclosure, and the method may include S101 to S103.

[0029] S101: Determine the subject image based on the reference image and the subject mask corresponding to the reference image, and determine the decoupled global identity embedding feature based on the subject image. The reference image is the initial face image.

[0030] In this embodiment, determining the subject image based on the reference image and the subject mask corresponding to the reference image includes:

[0031] Determine the subject mask based on the reference image.

[0032] Calculate the subject image by performing calculations on the corresponding pixel positions of the subject mask and the reference image.

[0033] In this embodiment, determining the subject mask based on the reference image may include: extracting the subject mask corresponding to the reference image based on a face parsing model.

[0034] In this embodiment, calculating the subject image by performing calculations on the corresponding pixel positions of the subject mask and the reference image may include: performing pixel-level multiplication calculations on the corresponding pixel positions of the subject mask and the reference image to obtain the subject image.

[0035] In this embodiment, the reference image is a face image. The subject mask corresponds to the reference image. The subject mask divides the reference image into a face region and a background region, with the background region assigned a value of 0 and the face region assigned a value of 1. The subject image is a face image with background irrelevant information removed.

[0036] As Figure 2 shown, exemplarily, use the pre-trained face parsing model ParseNet to extract the subject mask of the reference image , and perform pixel-level multiplication calculations with the corresponding pixel positions of the reference image to obtain the subject image after that.

[0037] In this embodiment, determining the decoupled global identity embedding feature based on the subject image includes:

[0038] Determine the first identity feature vector based on the subject image and the first encoder.

[0039] Perform vector editing on the first identity feature vector to generate the second identity feature vector.

[0040] Determine the decoupled global identity embedding feature based on the second identity feature vector and the mapping network.

[0041] In this embodiment, performing vector editing on the first identity feature vector may include:

[0042] Perform vector editing on the first identity feature vector based on the attribute direction and attribute intensity.

[0043] In this embodiment, the first identity feature vector is a decoupled identity feature vector extracted based on the subject image. The second identity feature vector is the identity feature vector after vector editing of the first identity feature vector. The attribute direction is the facial attribute category, and the attribute intensity is the facial attribute parameter. Since a vector is a vector with magnitude and direction, the vector can be controlled by adjusting parameters such as the attribute direction and the attribute intensity. The mapping network is a multi-dimensional mapping network. Through the mapping network, data processing is performed on the second identity feature vector to change the data dimension of the second identity feature vector, and a decoupled global identity embedding feature that conforms to the input format is obtained.

[0044] As Figure 2 shown, exemplarily, the subject image is input into a pre-trained StyleGAN encoder to extract a decoupled identity feature vector , and by adjusting the attribute direction and the attribute intensity (set to 0 during training), the translated feature vector is calculated . Project it onto a mapping network composed of four linear layers to generate a decoupled global identity embedding .

[0045] S102: Determine the multi-modal local identity embedding feature based on the reference image and the mask image corresponding to the reference image. The multi-modal local identity embedding feature is a text embedding type feature.

[0046] In this embodiment, determining the multi-modal local identity embedding feature based on the reference image and the mask image corresponding to the reference image may include:

[0047] Obtain the mask image based on the reference image.

[0048] Determine the multi-modal local identity embedding feature based on the mask image.

[0049] As Figure 2 shown, in this embodiment, the mask image is a mask image obtained by dividing the subject attribute regions of the reference image. For example, the reference image is divided into local regions such as background, eyes, hair, nose, mouth, chin, cheeks, etc. The segmentation model BiSeNet can be used to segment the reference image to generate a mask image with clearly divided subject attribute regions.

[0050] Exemplarily, an animation production studio uses a hand-drawn character sketch as a reference image, which depicts the image of a mysterious magician. Obtain the corresponding mask image of this reference image, divide it into different regions, and extract the features of each region, such as the deep eyes, flowing long hair, straight nose, smiling mouth, firm chin, and round cheeks of the magician. Perform data preprocessing on these features to obtain multi-modal local identity embedding features (text embedding features). The multi-modal local identity embedding features can guide the subsequent production software to accurately generate a facial image that conforms to the image of this magician, including the details and styles of each local region, thus providing strong support for the creation of animation characters.

[0051] S103: Determine the target generated face image based on the decoupled global identity embedding features and multi-modal local identity embedding features.

[0052] In this embodiment, determining the target generated face image based on the decoupled global identity embedding features and multi-modal local identity embedding features may include:

[0053] Input the decoupled global identity embedding features and multi-modal local identity embedding features into the target diffusion model to generate the target generated face image.

[0054] As Figure 2 shown, exemplarily, in a scene of post-production of a film and television work, we have a photo of a lady as a reference image, that is, the initial face image. Now, it is required to put a police cap on the lady during the face generation process. Through the reference image and the corresponding body mask of the reference image, the body image can be determined, that is, the key parts such as the face and hair of the beautiful woman in the photo are highlighted, and the background irrelevant information is removed. Feature extraction and analysis are performed on the body image to obtain decoupled global identity embedding features, which contain the overall appearance characteristics of the lady, such as the facial contour, hairstyle, etc. The decoupled global identity embedding features are used as identity constraints to ensure identity consistency.

[0055] At the same time, use the reference image and its corresponding mask image (the mask image divides the face of the reference image into local regions such as hair, eyes, mouth, nose, etc.) to determine the multi-modal local identity embedding features, which are a type of text embedding features, such as information about the shape and color of the eyes. Input the decoupled global identity embedding features and multi-modal local identity embedding features into a specific target diffusion model. The target diffusion model generates the target face image according to these features, adding a handsome police cap on the head while maintaining the original facial features of the lady. The style and wearing angle of the cap conform to the real visual effect, meeting the creative requirements for the picture in film and television production.

[0056] As can be seen above, on the one hand, in this embodiment, a multi-modal guided identity constraint framework is constructed by decoupling the global identity embedding feature and the multi-modal local identity embedding feature to ensure the stability of the identity feature and the retention of facial details, and has the semantic editing ability. The combination of decoupling the global identity embedding feature and the multi-modal local identity embedding feature makes the generated face image more accurate and rich in identity feature restoration. In addition, by using text information to determine the multi-modal local identity embedding feature, the user can precisely adjust the local details of the generated face by inputting specific text, which broadens the application scenarios of face generation.

[0057] On the other hand, this embodiment uses a mask balance loss based on residual cross-attention (corresponding to the subject mask and the relevant part of the mask image), which can effectively separate the identity information from the non-identity related background and balance the influence of text prompts and identity constraints.

[0058] On yet another hand, this embodiment also realizes smooth control of fine-grained semantic editing.

[0059] In summary, this embodiment realizes a method for personalized text-to-face generation based on multi-modal learning to ensure stable identity consistency, making the generation result of the face ensure stable identity consistency of the subject on the basis of aligning with diverse text prompts.

[0060] As Figure 2 shown, in an embodiment of the present disclosure, determining the multi-modal local identity embedding feature based on the reference image and the mask image corresponding to the reference image includes:

[0061] Determining a facial prompt and a scene prompt based on the reference image, and generating a mask image based on the reference image.

[0062] Performing data processing on the facial prompt and the scene prompt to obtain a text token sequence, and performing feature extraction on the mask image to obtain a face attribute feature.

[0063] Concatenating the face attribute feature and the text token sequence to obtain a multi-modal token sequence, and determining the multi-modal local identity embedding feature based on the multi-modal token sequence.

[0064] In this embodiment, performing data processing on the facial prompt and the scene prompt to obtain a text token sequence, and performing feature extraction on the mask image to obtain a face attribute feature includes:

[0065] Concatenating the facial prompt and the scene prompt to obtain a text prompt, and inputting the text prompt into a second encoder to obtain a text token sequence;

[0066] Determining the face attribute feature corresponding to the mask image based on the mask image and a third encoder;

[0067] Determine an attribute feature sequence based on the face attribute features and the first query feature; the first query feature is the query feature vector corresponding to the reference image.

[0068] In this embodiment, determining the facial prompt and the scene prompt based on the reference image may include:

[0069] Determine a problem description based on the reference image, input the problem description into the vision-language model, and obtain the facial prompt and the scene prompt.

[0070] In this embodiment, determining the facial prompt and the scene prompt based on the reference image may also include:

[0071] Determine the scene prompt based on the reference image and the generation constraint. The generation constraint is a personalized generation constraint condition for the reference image.

[0072] In this embodiment, the second encoder may be a text encoder. The third encoder may be an image encoder. The problem description may include the question text about the reference image. The facial prompt is multimodal data, and the facial prompt may include the text feature description and the image feature data for the reference image.

[0073] Exemplarily, the problem description may include: "Describe the facial identity features of this person" and "Describe the pose, action, accessories, clothing, and background environment of this person", etc. Input the reference image and the problem description into the vision-language model LLaVA, and the facial prompt corresponding to the reference image can be generated and the scene prompt . The generation constraint is a personalized generation condition, which can be set by the user according to specific requirements.

[0074] Exemplarily, use the segmentation model BiSeNet to segment the reference image to generate a mask image with clearly divided subject attribute regions. Input the mask image into the image encoder to obtain the face attribute features , and input the into an 8-layer learnable attention network, input the learnable query vector (i.e., the first query feature), continuously strengthen it, obtain the detailed face attribute features of each part, and perform data processing on the strengthened face attribute features to obtain the attribute feature sequence.

[0075] After splicing the attribute feature sequence and the text token sequence, input them into a Multilayer Perceptron (MLP) network. Embed the facial attribute features of each part in the attribute feature sequence (such as the mouth, nose, eyebrows, ears) into the corresponding prompt positions of the text token sequence through the attention network. Process the multimodal token sequence (i.e., the text token sequence) after embedding the attribute feature sequence (i.e., the image detail features) through an 8-layer MLP network with residual connections to generate the multimodal local identity embedding features 。

[0076]

[0077] Among them, is the text encoder, is the attention network, is the face prompt, is the scene prompt, is the face attribute feature, is the MLP network.

[0078] In this embodiment, by combining the face prompt and the scene prompt related to the reference image and integrating detailed face attribute features, image information can be captured more accurately. This embodiment considers generation constraints and can meet the personalized needs of users. Different constraint conditions can generate unique face images in different scenarios. This embodiment uses an encoder, an attention network, and an MLP network to effectively fuse text and image features, making the multi-modal local identity embedding features more comprehensive and rich, and improving the subsequent face generation effect.

[0079] In one embodiment of the present disclosure, determining a target generated face image based on the decoupled global identity embedding feature and the multi-modal local identity embedding feature includes:

[0080] Determining a first hidden state based on the second query feature and the multi-modal local identity embedding feature.

[0081] Performing a residual connection on the first hidden state and the decoupled global identity embedding feature to obtain a second hidden state.

[0082] Determining the target generated face image based on the second hidden state.

[0083] Exemplarily, given the second query feature , the hidden state after text embedding cross-attention is obtained (i.e., the first hidden state);

[0084]

[0085]

[0086]

[0087]

[0088] Among them, z is the second query feature, is the first hidden state, , and are the query matrix, key matrix, and value matrix of the text cross-attention module respectively, is the weight matrix corresponding to Q, is the weight matrix corresponding to K, is the weight matrix corresponding to V, is the attention network, is the multi-modal local identity embedding feature, is the attention score calculation formula, and d is the matrix vector dimension.

[0089] Take and as the key-value vectors for residual connection after cross-attention to obtain the second hidden state of the reference image after double identity conditional constraints ;

[0090]

[0091]

[0092]

[0093]

[0094] Among them, , and are the query, key, and value matrices in the identity features respectively. is corresponding weight matrix, is corresponding weight matrix, is corresponding weight matrix, is a hyperparameter.

[0095] This embodiment uses different queries for cross-attention between images and text prompts.

[0096] This embodiment determines the first hidden state and the second hidden state through a specific calculation method, and can accurately fuse the multi-modal local identity embedding feature and the decoupled global identity embedding feature. This helps to comprehensively consider various identity-related information when generating the target face image, from local details to overall features, making the identity features of the generated face image more accurate and rich, and avoiding information loss or confusion.

[0097] This embodiment uses double identity conditional constraints, adds the decoupled global identity embedding feature to the first hidden state to obtain the second hidden state, and effectively enhances the grasp of the identity information of the reference image. This double constraint mechanism can better guide the process of generating the target face image, make the generated image more in line with expectations, improve the image quality, and more vividly present the identity features of the target subject.

[0098] In one embodiment of the present disclosure, the text-based multimodal face generation method further includes:

[0099] Training an initial diffusion model based on a training data set to obtain a target diffusion model.

[0100] The training data set includes decoupled global identity embedding features, multimodal local identity embedding features, and noise images.

[0101] The training process of the initial diffusion model uses a first loss function and a second loss function.

[0102] In this embodiment, training the initial diffusion model based on the training data set includes:

[0103] Generating first prediction data based on the decoupled global identity embedding features, multimodal local identity embedding features, and noise images.

[0104] Determining a target loss based on the first prediction data, and updating the initial diffusion model based on the target loss to obtain a target diffusion model.

[0105] In this embodiment, generating first prediction data based on the decoupled global identity embedding features, multimodal local identity embedding features, and noise images includes:

[0106] Determining a noise image based on a reference image and true noise.

[0107] Generating predicted noise based on the decoupled global identity embedding features, multimodal local identity embedding features, and noise images, generating a noise reconstruction image based on the decoupled global identity embedding features, multimodal local identity embedding features, and noise images, and using the predicted noise and the noise reconstruction image as the first prediction data.

[0108] In this embodiment, determining the target loss based on the first prediction data includes:

[0109] Within a first loss step length, calculating a first loss between true noise and predicted noise based on the first loss function. The first loss function is:

[0110]

[0111] Wherein, represents the first loss, 0.4S, S) represents the first loss step length, 0.4S < s < S represents within the first loss step length range, S is the total number of denoising steps, s is the denoising index, and the value range of s is (0, S], is the true noise, is the predicted noise, is the noise decoded image after adding noise to the reference image at time t.

[0112] Within the second loss step size, calculate the second loss between the main image and the noise reconstruction image based on the second loss function. The second loss function is:

[0113]

[0114] wherein, denotes the second loss, (0, denotes the second loss step size, 0 < s ≤ 0.4S indicates within the range of the second loss step size, is the decoded image obtained by decoding the predicted noise image, is the initial state before denoising.

[0115] Determine the target loss based on the first loss and the second loss.

[0116] In this embodiment, the denoising time step size S can be set to 1000 steps (i.e., the total number of denoising steps). During the denoising process, the initial value of s is 1000, and the loss functions within two different time step sizes of 0.6S (600 steps) and 0.4S (400 steps) are used to guide the diffusion generation process.

[0117] Exemplarily, within the first 600 steps of denoising (i.e., s decreases from 1000 to 400 one by one), the mean square error between the predicted noise and the real noise is used as the first loss , and let the first loss guide the initial diffusion model to gradually reduce the noise in the image, so as to determine the basic structure and element distribution for the generated image;

[0118] Within the last 400 steps of denoising, within the main mask range, use the mean square error between the prediction and the actual to replace and learn the pixel-level details of the reference image within the mask range.

[0119] In this embodiment, based on the first loss function and the second loss function, a third loss function is obtained. The third loss function is:

[0120]

[0121] wherein, is the target loss.

[0122] In this embodiment, it further includes: using a fourth loss function and a fifth loss function in the training process of the initial diffusion model.

[0123] In this embodiment, the fourth loss function is a mask editing loss function, and the mask editing loss function includes the mean square error between the predicted noise and the real noise in the main mask​ The mean squared error within a range.

[0124] The fifth loss function is the diversity loss, which includes the noise loss when the multi-modal local identity embedding features and the multi-modal local identity embedding features are used as conditional constraints simultaneously and when only the multi-modal local identity embedding features are used as conditional constraints.

[0125] In this embodiment, the text-based multi-modal face generation method further includes constructing a comprehensive portrait dataset.

[0126] Exemplarily, in order to enhance the robustness of the initial diffusion model during training, the construction process of the comprehensive portrait dataset covers data collection, data cleaning, and prompt generation. The comprehensive portrait dataset consists of real images and generated images, and the data sources include 181,431 training images such as FFHQ, CelebA-HQ, and stylega2.

[0127] Process 181K face images using the dlib alignment method, and filter out face images with small faces, extreme angles, and incomplete face regions. Ask the vision-language model "Describe the facial identity features of this person" and "Describe the pose, action, accessories, clothing, and background environment of this person" to obtain the facial prompt and the scene prompt , to ensure the richness of the image prompts. The comprehensive portrait dataset has decoupled editable attribute vectors, providing fine-grained visual manipulation capabilities for model training.

[0128] In this embodiment, by using different loss functions in stages, determining the basic structure using the mean squared error between the predicted noise and the real noise in the early stage, and learning the pixel-level details within the mask in the later stage, it can accurately guide the diffusion generation process, making the generated image have a reasonable structure and rich details. The target loss function comprehensively considers various factors, effectively balancing the influence of identity information and text prompts, avoiding the generated image being biased towards one aspect, and ensuring that the generated face conforms to the identity characteristics and meets the text description.

[0129] The comprehensive portrait dataset in this embodiment contains a large number of images from various sources, which provides rich materials for model training after processing and enhances the generalization ability of the model. The constructed comprehensive portrait dataset has both face detail prompts and editability, facilitating fine-grained visual manipulation of the generated image, meeting diverse generation requirements, and overcoming the limitations of existing public datasets.

[0130] Corresponding to the text-based multi-modal face generation method in the above embodiment, Figure 2 This is the structural block diagram of the text-based multi-modal face generation device provided by an embodiment of the present disclosure. For the sake of illustration, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 2, the text-based multi-modal face generation device 20 includes: a first calculation module 21, a second calculation module 22, and a face generation module 23.

[0131] Among them, the first calculation module 21 is used to determine a subject image based on a reference image and a subject mask corresponding to the reference image, and determine a decoupled global identity embedding feature based on the subject image. The reference image is an initial face image.

[0132] The second calculation module 22 is used to determine a multi-modal local identity embedding feature based on the reference image and a mask image corresponding to the reference image. The multi-modal local identity embedding feature is a text embedding type feature.

[0133] The face generation module 23 is used to determine a target generated face image based on the decoupled global identity embedding feature and the multi-modal local identity embedding feature.

[0134] In an embodiment of the present disclosure, the first calculation module 21 is specifically used to determine a subject mask based on the reference image.

[0135] Calculate the subject mask with the corresponding pixel positions of the reference image to obtain the subject image.

[0136] In an embodiment of the present disclosure, the first calculation module 21 is specifically further used to determine a first identity feature vector based on the subject image and a first encoder.

[0137] Perform vector editing on the first identity feature vector to generate a second identity feature vector.

[0138] Determine the decoupled global identity embedding feature based on the second identity feature vector and a mapping network.

[0139] In an embodiment of the present disclosure, the second calculation module 21 is specifically used to determine a face prompt and a scene prompt based on the reference image, and generate a mask image based on the reference image.

[0140] Perform data processing on the face prompt and the scene prompt to obtain a text token sequence, and perform feature extraction on the mask image to obtain face attribute features.

[0141] Concatenate the face attribute features and the text token sequence to obtain a multi-modal token sequence, and determine the multi-modal local identity embedding feature based on the multi-modal token sequence.

[0142] In an embodiment of the present disclosure, the second calculation module 21 is specifically further used to concatenate the face prompt and the scene prompt to obtain a text prompt, input the text prompt into a second encoder to obtain a text token sequence;

[0143] Determine the face attribute features corresponding to the mask image based on the mask image and a third encoder;

[0144] Determine an attribute feature sequence based on face attribute features and a first query feature; the first query feature is a query feature vector corresponding to a reference image.

[0145] In an embodiment of the present disclosure, the face generation module 23 is specifically configured to determine a first hidden state based on a second query feature and a multi-modal local identity embedding feature.

[0146] Perform a residual connection on the first hidden state and the decoupled global identity embedding feature to obtain a second hidden state.

[0147] Determine a target generated face image based on the second hidden state.

[0148] In an embodiment of the present disclosure, the text-based multi-modal face generation device 20 further includes:

[0149] A model training module for training an initial diffusion model based on a training data set to obtain a target diffusion model.

[0150] The training data set includes a decoupled global identity embedding feature, a multi-modal local identity embedding feature, and a noise image.

[0151] The training process of the initial diffusion model uses a first loss function and a second loss function.

[0152] See Figure 5 , Figure 5 is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. As Figure 3 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 complete mutual communication through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is configured to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of the respective modules in the above-mentioned device embodiments, such as Figure 4 the functions of the modules 21 to 23 shown.

[0153] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0154] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0155] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0156] In a specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first embodiment and the second embodiment of the text-based multi-modal face generation method provided by the embodiments of the present disclosure, and may also execute the implementation manner of the electronic device 300 described in the embodiments of the present disclosure, which will not be elaborated herein.

[0157] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing related hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0158] The computer-readable storage medium can be the internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk equipped on the electronic device, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0159] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0160] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0161] In several embodiments provided in the present application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, and can also be an electrical, mechanical or other form of connection.

[0162] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.

[0163] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0164] The above are only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A text-based multimodal face generation method, characterized in that, Including: Determine a subject image based on a reference image and a subject mask corresponding to the reference image; Determine a first identity feature vector based on the subject image and a first encoder; Perform vector editing on the first identity feature vector based on an attribute direction and an attribute intensity to generate a second identity feature vector; determine a decoupled global identity embedding feature based on the second identity feature vector and a mapping network; The attribute direction is a facial attribute category, and the attribute intensity is a facial attribute parameter; the reference image is an initial face image; Determine a multimodal local identity embedding feature based on the reference image and a mask image corresponding to the reference image; the multimodal local identity embedding feature is a text embedding-like feature; Determine a target generated face image based on the decoupled global identity embedding feature and the multimodal local identity embedding feature; wherein, determining the target generated face image based on the decoupled global identity embedding feature and the multimodal local identity embedding feature includes: inputting the decoupled global identity embedding feature and the multimodal local identity embedding feature into a target diffusion model to generate the target generated face image; The method further includes: Determine a noise image based on the reference image and real noise; Generate predicted noise based on the decoupled global identity embedding feature, the multimodal local identity embedding feature, and the noise image, generate a noise reconstruction image based on the decoupled global identity embedding feature, the multimodal local identity embedding feature, and the noise image, and use the predicted noise and the noise reconstruction image as first prediction data; Within a first loss step, calculate a first loss between the real noise and the predicted noise based on a first loss function; Within a second loss step, calculate a second loss between the subject image and the noise reconstruction image based on a second loss function; Determine a target loss based on the first loss and the second loss; update an initial diffusion model based on the target loss to obtain the target diffusion model.

2. The text-based multi-modal face generation method according to claim 1, wherein, The determining the subject image based on the reference image and the subject mask corresponding to the reference image includes: Determine the subject mask based on the reference image; Calculate the subject image by performing calculations on corresponding pixel positions of the subject mask and the reference image.

3. The text-based multi-modal face generation method according to claim 1, wherein The determining the multimodal local identity embedding feature based on the reference image and the mask image corresponding to the reference image includes: Determine a facial prompt and a scene prompt based on the reference image, and generate the mask image based on the reference image; Perform data processing on the facial prompt and the scene prompt to obtain a text token sequence, and perform feature extraction on the mask image to obtain a face attribute feature; Concatenate the face attribute feature and the text token sequence to obtain a multimodal token sequence, and determine the multimodal local identity embedding feature based on the multimodal token sequence.

4. The text-based multi-modal face generation method according to claim 3, wherein The performing data processing on the facial prompt and the scene prompt to obtain a text token sequence, and performing feature extraction on the mask image to obtain a face attribute feature includes: Concatenate the facial prompt and the scene prompt to obtain a text prompt, and input the text prompt into a second encoder to obtain the text token sequence; Determine the face attribute features corresponding to the masked image based on the masked image and the third encoder; Determine an attribute feature sequence based on the face attribute features and the first query feature; the first query feature is the query feature vector corresponding to the reference image.

5. The text-based multi-modal face generation method according to claim 4, wherein The determining the target generated face image based on the decoupled global identity embedding feature and the multimodal local identity embedding feature includes: Determine a first hidden state based on the second query feature and the multimodal local identity embedding feature; Perform a residual connection on the first hidden state and the decoupled global identity embedding feature to obtain a second hidden state; Determine the target generated face image based on the second hidden state.

6. A text-based multi-modal face generation device, characterized in that, Including: A first calculation module, configured to determine a subject image based on a reference image and the subject mask corresponding to the reference image; Determine a first identity feature vector based on the subject image and a first encoder; perform vector editing on the first identity feature vector based on an attribute direction and an attribute intensity to generate a second identity feature vector; determine a decoupled global identity embedding feature based on the second identity feature vector and a mapping network; The attribute direction is a facial attribute category, and the attribute intensity is a facial attribute parameter; the reference image is an initial face image; A second calculation module, configured to determine a multimodal local identity embedding feature based on the reference image and the masked image corresponding to the reference image; the multimodal local identity embedding feature is a text embedding type feature; A face generation module, configured to determine a target generated face image based on the decoupled global identity embedding feature and the multimodal local identity embedding feature; wherein, the determining the target generated face image based on the decoupled global identity embedding feature and the multimodal local identity embedding feature includes: inputting the decoupled global identity embedding feature and the multimodal local identity embedding feature into a target diffusion model to generate a target generated face image; A model training module, configured to determine a noise image based on a reference image and real noise; generate predicted noise based on the decoupled global identity embedding feature, the multimodal local identity embedding feature, and the noise image, generate a noise reconstruction image based on the decoupled global identity embedding feature, the multimodal local identity embedding feature, and the noise image, and use the predicted noise and the noise reconstruction image as first prediction data; Within a first loss step, calculate a first loss between the real noise and the predicted noise based on a first loss function; within a second loss step, calculate a second loss between the subject image and the noise reconstruction image based on a second loss function; determine a target loss based on the first loss and the second loss; update an initial diffusion model based on the target loss to obtain a target diffusion model.

7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Speaking face video generation method and device based on multi-modal information control

    CN117456587A

  • Portrait generation system and method with multi-modal fine-grained identity reservation

    CN118298065A