A personalized face image generation method and system
By introducing a two-layer text encoding and image encoding module into the stable diffusion model, the problem of inconsistency between text prompts and image prompts is solved, and high-accuracy face image generation is achieved.
Patent Information
- Application Number
- CN202510130142.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-02-05
AI Technical Summary
In existing technologies for generating personalized facial images, the keywords in the text prompts are inconsistent with the faces in the image prompts, resulting in low similarity between the generated facial images and the target, and low image accuracy.
An IP-Adapter is introduced into the stable diffusion model, and a two-layer text encoding module and a two-layer face image encoding module are added. Through a two-stage training method, text features are processed and multi-dimensional image features are extracted to ensure the consistency between text prompts and image prompts.
It improves the accuracy of face image generation, resulting in more similar faces to the target person and higher image accuracy.
Smart Images

Figure CN120070634B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of personalized image technology, and more specifically, to a method and system for generating personalized face images. Background Technology
[0002] In recent years, large-scale text-to-image diffusion models have demonstrated powerful generative capabilities, especially stable diffusion models, which can create high-fidelity images. However, generating the desired images using only text prompts is extremely challenging, as it typically involves complex prompting engineering. A natural choice is to use image prompts, as images can convey more content and detail compared to text, leading to the development of personalized generation based on diffusion models. Personalized generation involves providing a diffusion model with a small number of images related to a specific concept, enabling it to learn and master this new concept. After such training, the model will be able to generate new images related to that concept based on text prompts, including images from different scenes and styles. Personalized face image generation is a special case of personalized generation, focusing on facial attributes with strong semantics and has been widely applied in real-world scenarios.
[0003] Currently, significant progress has been made in personalized image synthesis using methods such as Textual Inversion, DreamBooth, and LoRA. However, their applicability in the real world is hindered by high storage requirements and lengthy fine-tuning processes. IP-Adapter, on the other hand, implements topic-driven text-to-image generation without requiring additional fine-tuning, offering flexibility, efficiency, and lightweight characteristics. For each cross-attention layer in the U-Net diffusion model, it adds only an additional cross-attention layer for image features, thus separating the cross-attention layers of text and image features to decouple the cross-attention mechanism and achieve image cues for the pre-trained text-to-image diffusion model. However, because IP-Adapter does not process text features, its training uses simple text descriptions like "a photo of a girl" or "a photo of a man." When used for personalized face image generation, simple topic words like "girl" and "man" in the text cues may not match the face in the image cues, affecting the resemblance between the generated face and the provided target person, resulting in low image accuracy. At the same time, the image features of the IP-Adapter are extracted only through the CLIP model, which cannot fully represent the features of the face image, and will also cause the problem of low image accuracy. Summary of the Invention
[0004] To overcome the shortcomings of low similarity and low accuracy of generated face images in personalized face image generation, this invention provides a method and system for generating personalized face images.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] A method for generating personalized facial images includes the following steps:
[0007] The reference image and a preset set of text prompts are input into a two-layer text encoding module to obtain a vector s containing topic words. * The first text prompt;
[0008] The text prompt is input into a pre-trained stable diffusion model, and the two-layer text encoding module is trained using LDM loss.
[0009] The reference image is input into the two-layer face image encoding module to obtain an image prompt, and the reference image is input into the trained two-layer text encoding module to obtain a second text prompt;
[0010] The image prompt and the second text prompt are input into the stable diffusion model through a decoupled cross-attention mechanism, and the two-layer face image encoding module is trained using LDM loss;
[0011] The reference image and personalized prompt text are input into the stable diffusion model, which includes a trained two-layer text encoding module and a two-layer face image encoding module, to generate a personalized face image.
[0012] Furthermore, this invention also proposes a personalized face image generation system, applying the personalized face image generation method proposed in this invention. The system includes:
[0013] The text prompt module includes a two-layer text encoding module, used to generate topic word vectors s based on the input reference image and a preset set of text prompts. * The text prompt;
[0014] The image prompting module includes a two-layer face image encoding module, which is used to generate image prompts based on the input reference image;
[0015] The personalized generation module includes a stable diffusion model for generating personalized face images based on the text prompts and image prompts.
[0016] Furthermore, the present invention also proposes an apparatus comprising a memory and a processor, wherein the memory stores computer-readable instructions, wherein when executed by the processor, the computer-readable instructions cause the processor to perform all or part of the steps of the personalized face image generation method as described in the present invention.
[0017] Furthermore, the present invention also proposes a storage medium storing computer-readable instructions thereon, wherein the computer-readable instructions, when executed by a processor, implement all or part of the steps of the personalized face image generation method as described in the present invention.
[0018] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0019] This invention introduces an IP-Adapter into a stable diffusion model, and adds a text encoder and an image encoder on top of the IP-Adapter to form a two-layer text encoding module and a two-layer face image encoding module. It uses a two-stage training method to solve the problem of inconsistency between the topic words in the text prompts and the faces in the image prompts, and further improves the accuracy of face image generation. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating a personalized face image generation method according to an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of a stable diffusion model according to an embodiment of the present invention.
[0022] Figure 3 An example of personalized face image generation and a comparison of the results.
[0023] Figure 4 Another example of personalized face image generation and a comparison of the results.
[0024] Figure 5 Another example of personalized face image generation and a comparison of the results.
[0025] Figure 6 Another example of personalized face image generation and a comparison of the results.
[0026] Figure 7 Another example of personalized face image generation and a comparison of the results.
[0027] Figure 8 Another example of personalized face image generation and a comparison of the results.
[0028] Figure 9 Another example of personalized face image generation and a comparison of the results.
[0029] Figure 10 This is an architectural diagram of a personalized face image generation system according to an embodiment of the present invention. Detailed Implementation
[0030] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. It should be emphasized that the faces appearing in the drawings are all virtual cartoon characters and do not involve real human faces.
[0031] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0032] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0033] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0034] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0035] Example 1
[0036] This embodiment proposes a personalized face image generation method, such as... Figure 1 The diagram shown is a flowchart of the personalized face image generation method in this embodiment.
[0037] The personalized face image generation method proposed in this embodiment includes the following steps:
[0038] S1. Input the reference image and the preset set of text prompts into the two-layer text encoding module to obtain the topic word vector s. * The first text prompt;
[0039] S2. Input the text prompt into the pre-trained stable diffusion model and train the two-layer text encoding module using LDM loss;
[0040] S3. Input the reference image into the two-layer face image encoding module to obtain image prompts, and input the reference image into the trained two-layer text encoding module to obtain second text prompts;
[0041] S4. The image prompt and the second text prompt are input into the stable diffusion model through a decoupled cross-attention mechanism, and the LDM loss is used to train the two-layer face image encoding module.
[0042] S5. Input the reference image and personalized prompt text into the stable diffusion model containing the trained two-layer text encoding module and two-layer face image encoding module to generate a personalized face image.
[0043] This embodiment introduces an IP-Adapter into the stable diffusion model, and adds a text encoder and an image encoder on the basis of the IP-Adapter to form a two-layer text encoding module and a two-layer face image encoding module. It uses a two-stage training method to solve the problem of inconsistency between the topic words in the text prompts and the faces in the image prompts, thereby improving the accuracy of face image generation.
[0044] Compared to the IP-Adapter, this embodiment processes text features through a dual-layer text encoding module, avoiding inconsistencies between simple text descriptions like "a photo of a girl" or "a photo of a man" and image prompts about faces, thus ensuring the accuracy of the generated face and the provided target person image. Furthermore, by combining this with a dual-layer face image encoding module to extract multi-dimensional image features, the generated image can express face image information, further ensuring image accuracy.
[0045] For example, such as Figure 2 The diagram shown is an architecture diagram of the stable diffusion model in this embodiment.
[0046] In an optional embodiment, the two-layer text encoding module includes:
[0047] The first text encoder E1 is used to generate topic word vectors s representing facial features based on the input reference image. * ;
[0048] A second text encoder is used to extract the topic word vector s. * The text prompt set other than the one mentioned above is used as input. After converting the words in the text prompt set into text feature vectors containing text semantic information, it is compared with the topic word vector s. * Generate a text prompt by concatenating the text.
[0049] The second text encoder is a CLIP-based text encoder within the stable diffusion model.
[0050] This embodiment adds a first text encoder to the IP-Adapter to generate topic word vectors s about the facial features of the reference image. * This enriches the keywords and further assists in the generation of personalized images.
[0051] As an example, the output of the first text encoder in this embodiment is connected to a fully connected layer, which is used to convert the text features output by the first text encoder into word vectors.
[0052] In an optional embodiment, training the two-layer text encoding module using LDM loss includes:
[0053] The reference image is used to generate training image prompts through the dual-layer face image encoding module;
[0054] The training image prompts and the first text prompts are decoupled from the cross-attention mechanism and then input into the U-Net model in the stable diffusion model to output the first generated image.
[0055] The loss value L is calculated based on LDM loss. DM Its expression is:
[0056]
[0057] Where x represents the reference image; ∈ represents the noise added during the noise-addition process, which follows a standard normal distribution; t is the time step; ∈ θ (·) represents the predicted noise of the U-Net model during the denoising process, x t Given a completely noisy image, y is the text cue input to the decoupled cross-attention mechanism; only the first text encoder is updated during training.
[0058] As an example, the stable diffusion model includes two processes: adding noise and denoising. In the noise addition process, noise ∈ is applied to the reference image x iteratively until a completely noisy image x is generated. t During the denoising process, input a completely noisy image x. t Given time step t and text prompt y, the noise ∈ added at each step is predicted iteratively. θ Using ∈ and ∈ θ The L2 loss between them is used for training.
[0059] During the first phase of training, only the first text encoder E1 and its connected fully connected layer FC1 are trained, while the parameters of the original U-Net model and the second text encoder in the stable diffusion model remain frozen.
[0060] In an alternative embodiment, the first text encoder is initialized using Arcface's structure and parameters.
[0061] ArcFace is a mature deep metric learning method for face recognition. It is based on a convolutional neural network (CNN) architecture and includes multiple convolutional layers, pooling layers, and residual blocks.
[0062] In this embodiment, the initialization of ArcFace structure and parameters refers to the process of using the ArcFace model to initialize the structure and parameters of each node before training the network model of the first text encoder E1.
[0063] In an optional embodiment, the dual-layer face image encoding module includes:
[0064] The first face image encoder E2 is used to extract multi-dimensional face image feature vectors from the input reference image and map them through a fully connected layer to obtain a multi-dimensional first image feature vector.
[0065] The second face image encoder, which is set within the stable diffusion model, includes a CLIP-based image encoder for extracting features from the input reference image and outputting a second image feature vector through a linear layer.
[0066] Additionally, a fully connected layer FC2 is used to map the first image feature vector and the second image feature vector to obtain image cues.
[0067] This embodiment adds a first face image encoder E2 to the IP-Adapter. This encoder takes a given reference image x0 as input and obtains a multi-dimensional face image feature vector n1 = E2(x). Furthermore, the face image feature vector n2 generated by the second face image encoder (Image Encoder) is passed through a fully connected layer FC2 to obtain image cues (ImageFeatures) to assist in the generation of personalized images.
[0068] The second face image encoder, Image Encoder, is an image encoder built into the stable diffusion model. Optionally, it can employ the CLIP model image encoder and embeds the global image into the feature sequence through a small Linar layer.
[0069] In an optional embodiment, training the two-layer face image coding module using LDM loss includes:
[0070] The image prompt and the second text prompt are input into the decoupled cross-attention mechanism. The image prompt and the second text prompt are semantically aligned with the generated image by the decoupled cross-attention mechanism and then input into the U-Net model in the stable diffusion model to output the second generated image.
[0071] The loss value L is calculated based on LDM loss. DM Its expression is:
[0072]
[0073] Where z is the image cue input to the decoupled cross-attention mechanism; during training, the parameters of the first text encoder, the second text encoder, the text cross-attention layer in the decoupled cross-attention mechanism, and the U-Net model are frozen.
[0074] During the second phase of training, the text prompt branch, text feature cross-attention layer and U-Net model that were trained in the first phase are kept frozen, and only the two-layer face image encoding module is trained.
[0075] In an alternative embodiment, the first face image encoder E2 is initialized using the structure and parameters of Arcface.
[0076] In the inference phase, specifically step S5, the input reference image and the topic word vector s are... * The personalized prompt text, under the training of the first text encoder E1, the first face image encoder E2, and the image encoder, can output any generated face image that is highly accurate with the reference image and conforms to the text prompt.
[0077] As an example, the personalized face image generation method proposed in this embodiment is compared with the IP-Adapter. Figures 3-9 The image shown is a personalized face image obtained by inputting different reference images and personalized prompt text.
[0078] As shown in the figure, compared with the IP-Adapter, the personalized face image generation method proposed in this embodiment can more effectively express the features of face images and has higher image accuracy.
[0079] Example 2
[0080] This embodiment proposes a personalized face image generation system, applied to the personalized face image generation method proposed in Embodiment 1. For example... Figure 10 The diagram shown is an architecture diagram of the personalized face image generation system in this embodiment.
[0081] The personalized face image generation system proposed in this embodiment includes:
[0082] The text prompt module includes a two-layer text encoding module, used to generate topic word vectors s based on the input reference image and a preset set of text prompts. * The text prompt;
[0083] The image prompting module includes a two-layer face image encoding module, which is used to generate image prompts based on the input reference image;
[0084] The personalized generation module includes a stable diffusion model for generating personalized face images based on the text prompts and image prompts.
[0085] Specifically, the reference image and a preset set of text prompts are input into a two-layer text encoding module to obtain a vector s containing topic words. * The first text prompt; the text prompt is input into the pre-trained stable diffusion model, and the two-layer text encoding module is trained in the first stage using LDM loss;
[0086] The reference image is input into the two-layer face image encoding module to obtain an image prompt, and the reference image is input into the trained two-layer text encoding module to obtain a second text prompt. The image prompt and the second text prompt are input into the stable diffusion model through a decoupled cross-attention mechanism, and the two-layer face image encoding module is trained in the second stage using LDM loss to obtain a trained personalized face image generation system.
[0087] It is understood that the system in this embodiment corresponds to the method in Embodiment 1 above, and the options in Embodiment 1 above are also applicable to this embodiment, so they will not be described again here.
[0088] Example 3
[0089] This embodiment proposes a computer device, including a memory and a processor. The memory stores computer-readable instructions, which, when executed by the processor, cause the processor to perform all or part of the steps of the personalized face image generation method proposed in Embodiment 1.
[0090] Example 4
[0091] This embodiment proposes a storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, they implement all or part of the steps of the personalized face image generation method proposed in Embodiment 1.
[0092] By way of example, the storage medium includes, but is not limited to, USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks or optical disks, and other media capable of storing program code.
[0093] By way of example, the instructions, programs, code sets, or instruction sets may be implemented using conventional programming languages.
[0094] By way of example, the processor includes, but is not limited to, smartphones, personal computers, servers, network devices, etc., for performing all or part of the steps of the personalized face image generation method described in Example 1.
[0095] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely exemplary. The modules described as separate components may or may not be physically separate. When implementing the present invention, the functions of each module can be implemented in one or more software and / or hardware. Alternatively, some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0096] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for generating personalized facial images, characterized in that, Includes the following steps: The reference image and a set of preset text prompts are input into a two-layer text encoding module to obtain a result containing topic word vectors. The first text prompt; The text prompt is input into a pre-trained stable diffusion model, and the two-layer text encoding module is trained using LDM loss. The reference image is input into the two-layer face image encoding module to obtain an image prompt, and the reference image is input into the trained two-layer text encoding module to obtain a second text prompt; The image prompt and the second text prompt are input into the stable diffusion model through a decoupled cross-attention mechanism, and the two-layer face image encoding module is trained using LDM loss; The reference image and personalized prompt text are input into the stable diffusion model, which includes a trained two-layer text encoding module and a two-layer face image encoding module, to generate a personalized face image. The dual-layer text encoding module includes: The first text encoder is used to generate topic word vectors representing facial features based on the input reference image. ; A second text encoder is used to extract text from the topic word vectors. Taking the text prompt set other than the one mentioned above as input, the self-words in the text prompt set are converted into text feature vectors containing text semantic information, and then compared with the topic word vector. Concatenate to generate text prompts; The dual-layer face image encoding module includes: The first face image encoder is used to extract multi-dimensional face image feature vectors from the input reference image and map them through a fully connected layer to obtain a multi-dimensional first image feature vector. The second face image encoder is set within the stable diffusion model. The second face image encoder includes a CLIP-based image encoder for extracting features from the input reference image and outputting a second image feature vector through a linear layer. In addition, a fully connected layer is used to map the feature vectors of the first image and the second image to obtain image cues.
2. The personalized face image generation method according to claim 1, characterized in that, The training of the two-layer text encoding module using LDM loss includes: The reference image is used to generate training image prompts through the dual-layer face image encoding module; The training image prompts and the first text prompts are decoupled from the cross-attention mechanism and then input into the U-Net model in the stable diffusion model to output the first generated image. Calculate its loss value based on LDM loss. Its expression is: in, Indicates a reference image; The noise generated during the noise generation process follows a standard normal distribution. For time steps; This represents the predicted noise in the U-Net model during the denoising process. A completely noisy image generated during the noise-adding process. The text prompts are input to the decoupled cross-attention mechanism; only the first text encoder is updated during training.
3. The personalized face image generation method according to claim 1, characterized in that, The first text encoder is initialized using Arcface's structure and parameters.
4. The personalized face image generation method according to claim 1, characterized in that, The training of the two-layer face image coding module using LDM loss includes: The image prompt and the second text prompt are input into the decoupled cross-attention mechanism. The image prompt and the second text prompt are semantically aligned with the generated image by the decoupled cross-attention mechanism and then input into the U-Net model in the stable diffusion model to output the second generated image. Calculate its loss value based on LDM loss. Its expression is: in, Image cues are input to the decoupled cross-attention mechanism; during training, the parameters of the first text encoder, the second text encoder, the text cross-attention layer in the decoupled cross-attention mechanism, and the U-Net model are frozen.
5. The personalized face image generation method according to claim 1, characterized in that, The first face image encoder is initialized using the Arcface structure and parameters.
6. A personalized face image generation system, employing the personalized face image generation method according to any one of claims 1 to 5, characterized in that, include: The text prompt module includes a two-layer text encoding module, used to generate topic word vectors based on the input reference image and a preset set of text prompts. The text prompt; The image prompting module includes a two-layer face image encoding module, which is used to generate image prompts based on the input reference image; A personalized generation module, including a stable diffusion model, is used to generate personalized face images based on the text prompts and image prompts; The dual-layer text encoding module includes: The first text encoder is used to generate topic word vectors representing facial features based on the input reference image. ; A second text encoder is used to extract text from the topic word vectors. Taking the text prompt set other than the one mentioned above as input, the self-words in the text prompt set are converted into text feature vectors containing text semantic information, and then compared with the topic word vector. Concatenate to generate text prompts; The dual-layer face image encoding module includes: The first face image encoder is used to extract multi-dimensional face image feature vectors from the input reference image and map them through a fully connected layer to obtain a multi-dimensional first image feature vector. The second face image encoder is set within the stable diffusion model. The second face image encoder includes a CLIP-based image encoder for extracting features from the input reference image and outputting a second image feature vector through a linear layer. In addition, a fully connected layer is used to map the feature vectors of the first image and the second image to obtain image cues.
7. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, characterized in that, When the computer-readable instructions are executed by the processor, the processor performs all or part of the steps of the personalized face image generation method as described in any one of claims 1 to 6.
8. A storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by a processor, they implement all or part of the steps of the personalized face image generation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image description text generation method based on generative adversarial network
CN112818159A
Method for generating face image under guidance of text description based on multi-level features
CN119295586A