Personalized face image generation method and system
By introducing a two-layer text encoding module and a two-layer face image encoding module in the stable diffusion model, combined with the two-stage training method, the problems of low similarity and low accuracy in the generation of personalized face images are solved, and higher image generation accuracy is achieved.
Patent Information
- Application Number
- CN202510130142.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-05
AI Technical Summary
In the prior art, in the generation of personalized face images, the generated face images have low similarity to targets and low image accuracy.
A two-layer text encoding module and a two-layer face image encoding module are introduced. Through a two-stage training method, the cross-attention mechanism is decoupled to improve the expression ability of text features and image features.
The problem of inconsistent faces in the subject words in the text prompt and the image prompt is solved, which significantly improves the accuracy of face image generation.
Smart Images

Figure CN120070634A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of personalized image, and more specifically, to a method and system for generating personalized face images. Background Art
[0002] In recent years, large text-to-image diffusion models have powerful generation capabilities. In particular, the Stable Diffusion generation model can create high-fidelity images. However, it is very tricky to generate the required images only using text prompts because it usually involves complex prompt engineering. A natural choice is to use image prompts because images can express more content and details compared with text. Therefore, personalized generation based on diffusion models has emerged. Personalized generation means providing a small number of specific concept pictures to the diffusion model so that it can learn and master this new concept. After such training, the model will be able to generate new pictures related to this concept according to text prompts, including images in different scenarios and styles. Face personalized image generation is a special case of personalized generation, which focuses on face attributes with strong semantics and has been widely applied in real-world scenarios.
[0003] Currently, significant progress has been made in personalized image synthesis using methods such as Textual Inversion, DreamBooth, and LoRA. However, their applicability in the real world is hindered by high storage requirements and long fine-tuning processes. IP-Adapter realizes theme-driven text-to-image generation without additional fine-tuning, which is flexible, effective, and lightweight. For each cross-attention layer in the U-Net of the diffusion model, it only adds an additional cross-attention layer to the image features, thus separating the cross-attention layers of text features and image features to decouple the cross-attention mechanism, and then realizing the image prompt function of the pre-trained text-to-image diffusion model. However, since IP-Adapter does not process text features and uses simple text descriptions such as "a photo of a girl", "a photo of a man" during training, when used for face personalized image generation, simple theme words such as "girl", "man" in the text prompt may be inconsistent with the faces in the image prompt, thus affecting the similarity between the generated face and the provided target person, that is, the low image accuracy. At the same time, the image features of IP-Adapter are only extracted by the CLIP model and cannot fully express the face image features, which will also cause the problem of low image accuracy. Summary of the Invention
[0004] The present invention aims to overcome the defects of low similarity between the generated face image and the target and low image accuracy in face personalized image generation, and provides a method and system for generating personalized face images.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] A personalized face image generation method, comprising the following steps:
[0007] Input the reference image and a preset text prompt set into a double-layer text encoding module to obtain a first text prompt containing the topic word vector s * ;
[0008] Input the text prompt into a pre-trained stable diffusion model, and use the LDM loss to train the double-layer text encoding module;
[0009] Input the reference image into a double-layer face image encoding module to obtain an image prompt, and input the reference image into the trained double-layer text encoding module to obtain a second text prompt;
[0010] Input the image prompt and the second text prompt into the stable diffusion model through a decoupled cross-attention mechanism, and use the LDM loss to train the double-layer face image encoding module;
[0011] Input the reference image and a personalized prompt text prompt into the stable diffusion model including the trained double-layer text encoding module and double-layer face image encoding module to generate a personalized face image.
[0012] Furthermore, the present invention also proposes a personalized face image generation system that applies the personalized face image generation method proposed by the present invention. Among them, the system includes:
[0013] A text prompt module, including a double-layer text encoding module, for generating a text prompt containing the topic word vector s * according to the input reference image and a preset text prompt set;
[0014] An image prompt module, including a double-layer face image encoding module, for generating an image prompt according to the input reference image;
[0015] A personalized generation module, including a stable diffusion model, for generating a personalized face image according to the text prompt and the image prompt.
[0016] Furthermore, the present invention also proposes a device, including a memory and a processor, and computer-readable instructions are stored in the memory. Among them, when the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the personalized face image generation method as described in the present invention.
[0017] Furthermore, the present invention also proposes a storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, all or part of the steps of the personalized face image generation method as described in the present invention are implemented.
[0018] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0019] The present invention introduces IP-Adapter into the StableDiffusion model, and additionally adds a text encoder and an image encoder on the basis of IP-Adapter to form a double-layer text encoding module and a double-layer face image encoding module, and uses a two-stage training method to solve the problem of inconsistency between the subject words in the text prompt and the face in the image prompt, and further improves the accuracy of face image generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart of the personalized face image generation method shown according to an embodiment of the present invention.
[0021] Figure 2 It is an architecture diagram of the StableDiffusion model shown according to an embodiment of the present invention.
[0022] Figure 3 It is an example and effect comparison diagram of personalized face image generation.
[0023] Figure 4 It is another example and effect comparison diagram of personalized face image generation.
[0024] Figure 5 It is another example and effect comparison diagram of personalized face image generation.
[0025] Figure 6 It is another example and effect comparison diagram of personalized face image generation.
[0026] Figure 7 It is another example and effect comparison diagram of personalized face image generation.
[0027] Figure 8 It is another example and effect comparison diagram of personalized face image generation.
[0028] Figure 9 It is another example and effect comparison diagram of personalized face image generation.
[0029] Figure 10 It is an architecture diagram of the personalized face image generation system shown according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. It should be emphasized that all the faces appearing in the drawings are virtual character cartoon images and do not involve real faces.
[0031] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0032] The terms used in the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0033] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0034] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0035] Embodiment 1
[0036] This embodiment proposes a personalized face image generation method, as Figure 1 shown, is a flowchart of the personalized face image generation method of this embodiment.
[0037] In the personalized face image generation method proposed in this embodiment, the following steps are included:
[0038] S1. Input a reference image and a preset text prompt set into a double-layer text encoding module to obtain a first text prompt containing the topic word vector s * ;
[0039] S2. Input the text prompt into a pre-trained stable diffusion model, and use the LDM loss to train the double-layer text encoding module;
[0040] S3. Input the reference image into the double-layer face image encoding module to obtain an image prompt, and input the reference image into the trained double-layer text encoding module to obtain a second text prompt;
[0041] S4. Input the image prompt and the second text prompt into the StableDiffusion model through the decoupled cross-attention mechanism, and use the LDM loss to train the double-layer face image encoding module;
[0042] S5. Input the reference image and the personalized prompt text prompt into the StableDiffusion model including the trained double-layer text encoding module and double-layer face image encoding module to generate a personalized face image.
[0043] In this embodiment, IP-Adapter is introduced into the StableDiffusion model, and a text encoder and an image encoder are added on the basis of IP-Adapter to form a double-layer text encoding module and a double-layer face image encoding module, and a two-stage training method is used to solve the problem of inconsistency between the subject words in the text prompt and the face in the image prompt, and improve the accuracy of face image generation.
[0044] Compared with IP-Adapter, in this embodiment, the text features are processed through the double-layer text encoding module, avoiding the problem of inconsistency between simple text descriptions such as "a photo of a girl" and "a photo of a man" and the face in the image prompt, and ensuring the accuracy of the generated face and the provided target person image; and, in cooperation with the double-layer face image encoding module to extract multi-dimensional image features, the generated image can express the face image information, further ensuring the problem of image accuracy.
[0045] Exemplarily, as Figure 2 shown, is the architecture diagram of the StableDiffusion model of this embodiment.
[0046] In an optional embodiment, the double-layer text encoding module includes:
[0047] The first text encoder E 1 , which is used to generate a subject word vector s * representing the face features according to the input reference image;
[0048] The second text encoder Text encoder, which takes the text prompt set except the subject word vector s * as the input, converts the self-words in the text prompt set into text feature vectors containing text semantic information, and then splices them with the subject word vector s * to generate a text prompt.
[0049] Among them, the second text encoder Text encoder is a text encoder based on the CLIP model within the StableDiffusion model.
[0050] In this embodiment, a first text encoder is added on the basis of IP-Adapter, which is used to generate a topic word vector s about the face features of the reference image * , so as to enrich the topic words and further assist in the generation of personalized images.
[0051] As an exemplary illustration, a fully connected layer is connected to the output end of the first text encoder in this embodiment, which is used to convert the text features output by the first text encoder into word vectors.
[0052] In an optional embodiment, the training of the double-layer text encoding module using the LDM loss includes:
[0053] Generating a training image prompt for the reference image through the double-layer face image encoding module;
[0054] Inputting the training image prompt and the first text prompt into the U-Net model in the StableDiffusion model through the decoupled cross-attention mechanism, and outputting the first generated image;
[0055] Calculating its loss value L based on the LDM loss DM ; Its expression is:
[0056]
[0057] Among them, x represents the reference image; ∈ represents the noise during the noise addition process, which follows a standard normal distribution; t is the time step; ∈ θ (·) represents the predicted noise during the denoising process of the U-Net model, x t is a completely noisy image, and y is the text prompt input into the decoupled cross-attention mechanism; only the first text encoder is updated during the training process.
[0058] As an exemplary illustration, the StableDiffusion model includes two processes: noise addition and denoising. During the noise addition process, the noise ∈ is gradually iteratively applied to the reference image x until a completely noisy image x t is generated; during the denoising process, the completely noisy image x t , the time step t, and the text prompt y are input, and the noise ∈ added at each step is gradually iteratively predicted θ , and the L2 loss between ∈ and ∈ θ is used for training.
[0059] During the training process of the first stage, only the first text encoder E is trained 1and the fully connected layer FC1 connected thereto, while the parameters of the original U-Net model and the second text encoder Text encoder in the Stable Diffusion model remain frozen.
[0060] In an optional embodiment, the first text encoder is initialized using the structure and parameters of Arcface.
[0061] Among them, ArcFace is a mature deep metric learning method for face recognition. Based on the convolutional neural network (CNN) architecture, it includes multiple convolutional layers, pooling layers and residual blocks.
[0062] In this embodiment, initializing with the structure and parameters of Arcface means that the architecture of the first text encoder E 1 uses the ArcFace model, and before training the network model of the first text encoder E 1 the weights and biases of each node are initially assigned using the parameters of the pre-trained ArcFace model.
[0063] In an optional embodiment, the double-layer face image encoding module includes:
[0064] The first face image encoder E 2 , which is used to extract multi-dimensional face image feature vectors from the input reference image, and perform mapping through a fully connected layer to obtain multi-dimensional first image feature vectors;
[0065] The second face image encoder Image Encoder set in the Stable Diffusion model, the second face image encoder includes an image encoder based on the CLIP model, which is used to extract features from the input reference image and output a second image feature vector through a linear layer;
[0066] And a fully connected layer FC2, which is used to map the first image feature vector and the second image feature vector to obtain an image prompt.
[0067] In this embodiment, on the basis of IP-Adapter, the first face image encoder E 2 is added, which takes the given reference image x 0 as input to obtain a multi-dimensional face image feature vector n 1 = E 2 (x). Further, in cooperation with the face image feature vector n 2 generated by the second face image encoder Image Encoder, an image prompt (ImageFeatures) is obtained through the fully connected layer FC2 to assist in the generation of personalized images.
[0068] The second face image encoder, Image Encoder, is an image encoder built into the StableDiffusion model. Optionally, it adopts the image encoder of the CLIP model and projects the global image embedding into a feature sequence through a small Linar layer.
[0069] In an optional embodiment, training the double-layer face image encoding module using the LDM loss includes:
[0070] Input the image prompt and the second text prompt into the decoupled cross-attention mechanism. After aligning the image prompt and the second text prompt with the semantics of the generated image respectively through the decoupled cross-attention mechanism, input them into the U-Net model in the StableDiffusion model to output a second generated image;
[0071] Calculate its loss value L DM ′ according to the LDM loss; its expression is:
[0072]
[0073] where z is the image prompt input to the decoupled cross-attention mechanism; during the training process, freeze the parameters of the first text encoder, the second text encoder, the text cross-attention layer in the decoupled cross-attention mechanism, and the U-Net model.
[0074] During the training process of the second stage, keep the text prompt branch, the text feature cross-attention layer, and the U-Net model that have completed training in the first stage frozen, and only train the double-layer face image encoding module.
[0075] In an optional embodiment, the first face image encoder E 2 is initialized using the structure and parameters of Arcface.
[0076] In the inference stage, that is, in step S5, input the reference image and the personalized prompt text except the subject word vector s * and, under the action of the trained first text encoder E 1 , the first face image encoder E 2 and Image Encoder, any generated face image with high accuracy and conforming to the text prompt can be output.
[0077] As an example, compare the personalized face image generation method proposed in this embodiment with IP-Adapter. As Figures 3 to 9 shown, the personalized face images obtained by inputting different reference images and personalized prompt texts are shown.
[0078] As can be seen from the figure, compared with the IP-Adapter, the personalized face image generation method proposed in this embodiment can more effectively express the face image features, and its image accuracy is higher.
[0079] Embodiment 2
[0080] This embodiment proposes a personalized face image generation system, which is applied to the personalized face image generation method proposed in Embodiment 1. As Figure 10 shown, it is the architecture diagram of the personalized face image generation system of this embodiment.
[0081] The personalized face image generation system proposed in this embodiment includes:
[0082] A text prompt module, which includes a double-layer text encoding module, and is used to generate a text prompt containing the topic word vector s * according to the input reference image and the preset text prompt set;
[0083] An image prompt module, which includes a double-layer face image encoding module, and is used to generate an image prompt according to the input reference image;
[0084] A personalized generation module, which includes a StableDiffusion model, and is used to generate a personalized face image according to the text prompt and the image prompt.
[0085] Among them, the reference image and the preset text prompt set are input into the double-layer text encoding module to obtain a first text prompt containing the topic word vector s * ; the text prompt is input into the pre-trained StableDiffusion model, and the double-layer text encoding module is trained in the first stage by using the LDM loss;
[0086] The reference image is input into the double-layer face image encoding module to obtain an image prompt, and the reference image is input into the trained double-layer text encoding module to obtain a second text prompt; the image prompt and the second text prompt are input into the StableDiffusion model through the decoupled cross-attention mechanism, and the double-layer face image encoding module is trained in the second stage by using the LDM loss to obtain a trained personalized face image generation system.
[0087] It can be understood that the system in this embodiment corresponds to the method in Embodiment 1 above, and the optional items in Embodiment 1 above also apply to this embodiment, so they will not be repeated here.
[0088] Embodiment 3
[0089] This embodiment provides a computer device, including a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute all or part of the steps of the personalized face image generation method proposed in Embodiment 1.
[0090] Embodiment 4
[0091] This embodiment provides a storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, all or part of the steps of the personalized face image generation method proposed in Embodiment 1 are implemented.
[0092] Exemplarily, the storage medium includes but is not limited to various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0093] Exemplarily, the instructions, programs, code sets, or instruction sets can be implemented using conventional programming languages.
[0094] Exemplarily, the processor includes but is not limited to smartphones, personal computers, servers, network devices, etc., and is used to execute all or part of the steps of the personalized face image generation method described in Embodiment 1.
[0095] Each embodiment in the present invention is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, refer to the partial description of the method embodiments. The device embodiments described above are merely exemplary. The modules described as separate components may or may not be physically separated. When implementing the solution of the present invention, the functions of the various modules can be implemented in the same or multiple software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solution of this embodiment.
[0096] Obviously, the above embodiments of the present invention are merely examples given for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A personalized face image generation method, characterized in that: The following steps are involved: The reference image and the preset text prompt set are input into the double-layer text encoding module to obtain the theme word vector s * The first text prompt; Inputting the text prompt into a pre-trained stable diffusion model, and training the two-layer text encoding module using LDM loss; Inputting the reference image into a two-layer face image encoding module to obtain an image prompt, and inputting the reference image into a trained two-layer text encoding module to obtain a second text prompt; Input the image prompt and the second text prompt into the stable diffusion model via a decoupled cross attention mechanism, and train the two-layer face image encoding module using LDM loss; The reference image and the personalized prompt text prompt are input into the stable diffusion model including the trained double-layer text encoding module and the double-layer face image encoding module to generate a personalized face image.
2. The method for generating a personalized face image according to claim 1, characterized in that: The double-layer text encoding module comprises: The first text encoder is used to generate a topic word vector s representing facial features based on the input reference image * ; The second text encoder is used to divide the topic word vector s * The text prompt set other than s is taken as input, and the self-words in the text prompt set are converted into text feature vectors containing text semantic information, and then compared with the subject word vector s * Splice to generate text prompts.
3. The method for generating a personalized face image according to claim 2, characterized in that: The method of training the double-layer text encoding module by using the LDM loss includes: Passing the reference image through the double-layer face image encoding module to generate a training image prompt; Input the training image prompt and the first text prompt into a U-Net model in a stable diffusion model through a decoupled cross attention mechanism, and output a first generated image; Calculate the loss value L based on LDM loss DM ; Its expression is: Where x represents the reference image; ∈ represents the noise in the noise adding process, which obeys the standard normal distribution; t is the time step; ∈ θ (·) represents the predicted noise of the U-Net model during the denoising process, x t is the completely noisy image generated by the denoising process, and y is the text prompt input to the decoupled cross-attention mechanism; only the first text encoder is updated during training.
4. The method for generating a personalized face image according to claim 2, characterized in that: The first text encoder is initialized using the Arcface structure and parameters.
5. The method for generating a personalized face image according to any one of claims 1 to 4, characterized in that: The double-layer face image encoding module comprises: A first face image encoder is used to extract a multi-dimensional face image feature vector from an input reference image, and map it through a fully connected layer to obtain a multi-dimensional first image feature vector; A second face image encoder is provided in the stable diffusion model, wherein the second face image encoder includes an image encoder based on a CLIP model, and is used for extracting features from an input reference image and outputting a second image feature vector through a linear layer; And, a fully connected layer is used to map the first image feature vector and the second image feature vector to obtain an image prompt.
6. The method for generating a personalized face image according to claim 5, characterized in that: The method of training the double-layer face image encoding module by using the LDM loss includes: Inputting the image prompt and the second text prompt into a decoupled cross attention mechanism, aligning the image prompt and the second text prompt with the semantics of the generated image through the decoupled cross attention mechanism, and then inputting them into a U-Net model in a stable diffusion model, and outputting a second generated image; Calculate the loss value L based on LDM loss DM ′; its expression is: Among them, z is the image prompt input to the decoupled criss-cross attention mechanism; the first text encoder, the second text encoder, the text criss-cross attention layer in the decoupled criss-cross attention mechanism, and the parameters of the U-Net model are frozen during training.
7. The method for generating a personalized face image according to claim 5, characterized in that: The first face image encoder is initialized using the structure and parameters of Arcface.
8. A personalized face image generation system, using the personalized face image generation method according to any one of claims 1 to 7, characterized in that: include: The text prompt module includes a double-layer text encoding module, which is used to generate a theme word vector s based on the input reference image and the preset text prompt set. * Text prompts; An image prompt module, including a double-layer face image encoding module, is used to generate image prompts based on an input reference image; The personalized generation module includes a stable diffusion model and is used to generate a personalized face image according to the text prompt and the image prompt.
9. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, characterized in that: When the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the personalized facial image generation method according to any one of claims 1 to 7.
10. A storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, all or part of the steps of the personalized facial image generation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Image description text generation method based on generative adversarial network
CN112818159A
Image scene feature-based image description text generation method and system
CN117036545A
Method for generating face image under guidance of text description based on multi-level features
CN119295586A
Personalized text-to-image generation
US20240355022A1
System and method for variable encoding based on image content
US6028962A
Cited By
Model graph generation method and device, equipment, storage medium and product
CN121352935A