Face image generation method and device, and electronic equipment

By obtaining the makeup features of the beauty makeup picture and adding it to the initial face image, and using the diffusion model to generate the target face image, the problems of inaccurate makeup applications and poor identity characteristics in the prior art are solved, and the face image generation with high authenticity is achieved.

CN119963427APending Publication Date: 2025-05-09WONDERSHARE TECH (HUNAN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411958123.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to accurately apply the makeup effect on the beauty photos to the user's photos while maintaining the user's identity characteristics, and the generated images often have significant differences from the target map and reference map in terms of makeup and identity characteristics.

Method used

By obtaining the makeup features of the beauty picture and adding it to the initial face image, the target face image is generated using a pre-trained diffusion model. The method includes obtaining the makeup features of the beauty picture, encoding makeup features, and identity features of the initial face image, and inputting them into the diffusion model to generate the target face image.

Benefits of technology

It realizes that the beauty effects are applied accurately while maintaining the user's identity characteristics, and the authenticity and accuracy of the generated target face images are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963427A_ABST
    Figure CN119963427A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a face image generation method. The face image generation method comprises the following steps: acquiring makeup features of a beauty makeup picture; inputting an initial face image; and adding makeup features of the beauty makeup picture to the initial face image to generate a target face image. According to the face image generation method provided by the embodiment of the invention, the makeup feature of the beauty makeup picture is added to the face image, so that the face image required by the user is accurately generated. The embodiment of the invention further provides a face image generation device and electronic equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology, and more specifically, to a method, device, and electronic device for generating a facial image. Background Art

[0002] With the rapid development of computer vision technology, image generation technology has become a research hotspot. In particular, given certain conditions (such as text descriptions, reference images), the technology of generating images or video content that meet these conditions has received widespread attention.

[0003] However, in practical applications, users usually want to apply the makeup effects in beauty photos to their own photos while keeping their identity features unchanged. This requires the generation system to accurately capture and apply makeup features while maintaining the identity features of the user's face. Current technologies still face challenges in fine control and identity preservation. For example, when a user wants to modify only a part of the makeup, it often causes the entire image to change, or the generated image has significant differences in makeup and identity features from the target image and reference image. Summary of the invention

[0004] In response to the problems existing in the above-mentioned prior art, the embodiments of the present application provide a facial image generation method, device, and electronic device, which accurately generate the facial image required by the user by adding the makeup features of a beauty picture to the facial image.

[0005] In a first aspect, an embodiment of the present application provides a method for generating a face image, comprising the following steps:

[0006] Obtain makeup features of beauty makeup pictures;

[0007] Input an initial face image; and

[0008] The makeup features of the beauty makeup image are added to the initial face image to generate a target face image.

[0009] Furthermore, obtaining makeup features of the beauty makeup image includes:

[0010] Obtaining the beauty makeup picture;

[0011] Performing makeup removal processing on the makeup image to obtain a makeup removal image; and

[0012] The makeup features of the beauty image are obtained by subtracting the features of the makeup removal image from the features of the beauty image.

[0013] Furthermore, the step of adding the makeup features of the beauty makeup image to the initial face image to obtain a target face image includes:

[0014] The makeup features of the beauty makeup image and the initial face image are input into a pre-trained diffusion model to generate the target face image.

[0015] Furthermore, after the inputting of the initial face image, the method further comprises:

[0016] Encoding the makeup features of the beauty makeup image to obtain the makeup code of the beauty makeup image;

[0017] Encoding the initial face image to obtain a corresponding identity code; and

[0018] Generate matching prompt words based on the initial face image.

[0019] Furthermore, the makeup features of the beauty makeup image and the initial face image are input into a pre-trained diffusion model to generate the target face image, including:

[0020] The makeup code of the beauty makeup image, the identity code corresponding to the initial face image and the prompt word are input into a pre-trained diffusion model to generate the target face image.

[0021] Furthermore, the step of inputting the makeup features of the beauty makeup image and the initial face image into a pre-trained diffusion model to generate the target face image includes:

[0022] According to the makeup features of the beauty makeup image and the initial face image, a target face image is generated through inverse diffusion reasoning of the pre-trained diffusion model.

[0023] Furthermore, the training process of the diffusion model includes:

[0024] Constructing a training data set, including obtaining a preset number of face images and beauty makeup images to form sample pairs as the training data set;

[0025] The training data set is input into the diffusion model for training, and the diffusion model is optimized by a preset loss function to obtain the trained diffusion model.

[0026] In a second aspect, the embodiment of the present application further provides a facial image generation device, comprising:

[0027] Makeup acquisition module, used to obtain makeup features of beauty pictures;

[0028] A face image input module, used for inputting an initial face image; and

[0029] The image generation module is used to add the makeup features of the beauty makeup image to the initial face image to generate a target face image.

[0030] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the facial image generation method according to the first aspect when executing the program.

[0031] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, wherein the computer program is used to implement the facial image generation method according to the first aspect above.

[0032] The embodiments of the present application bring the following beneficial effects:

[0033] In the facial image generation method provided in the embodiment of the present application, the makeup features of the beauty image are first obtained, and an initial facial image is input. Finally, the makeup features of the beauty image are added to the initial facial image to generate a target facial image. The facial image generation method provided in the embodiment of the present application accurately generates the facial image required by the user by adding the makeup features of the beauty image to the facial image, thereby ensuring the authenticity of the generated facial image. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0035] Figure 1 A schematic diagram of a process for generating a facial image according to an embodiment of the present invention;

[0036] Figure 2 A schematic diagram of the framework structure of a face image generation model used in the face image generation method provided in an embodiment of the present application;

[0037] Figure 3 A schematic diagram of the framework structure of a makeup removal image generation model used in the facial image generation method provided in an embodiment of the present application;

[0038] Figure 4 A schematic diagram of the structure of a makeup encoder used in the facial image generation method provided in an embodiment of the present application;

[0039] Figure 5 A schematic diagram of the structure of an image identity encoder used in the face image generation method provided in an embodiment of the present application;

[0040] Figure 6A structural block diagram of a facial image generation device provided in an embodiment of the present application;

[0041] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0042] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments described in the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.

[0044] In the specification and claims of this application and the above-mentioned drawings, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "multiple" means two or more. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to the specific circumstances.

[0045] Figure 1 FIG. 1 is a flowchart of a method for generating a face image according to an embodiment of the present application. Figure 1 As shown, the face image generation method of the embodiment of the present application is used for gesture recognition training, including the following steps:

[0046] S101: Obtain makeup features of the beauty makeup image;

[0047] The beauty pictures here are reference beauty pictures, that is, reference pictures of the beauty effects that users expect to achieve.

[0048] S102: Input an initial face image;

[0049] The initial face image is usually a selfie provided by the user, but can also be a face image of another person.

[0050] S103: Adding makeup features of the beauty makeup image to the initial face image to generate a target face image.

[0051] That is, for example, when a user provides a selfie, a virtual beauty picture is generated by adding makeup features of a beauty picture to the selfie to generate a beauty picture of the user's desired effect.

[0052] Therefore, in the facial image generation method provided in the embodiment of the present application, the makeup features of the beauty picture are first obtained, and an initial facial image is input, and finally the makeup features of the beauty picture are added to the initial facial image to generate a target facial image. The facial image generation method provided in the embodiment of the present application accurately generates the facial image required by the user by adding the makeup features of the beauty picture to the facial image, thereby ensuring the authenticity of the generated facial image.

[0053] Furthermore, in some embodiments of the present application, the step of obtaining makeup features of the beauty makeup image includes:

[0054] Obtaining the beauty makeup picture;

[0055] Performing makeup removal processing on the makeup image to obtain a makeup removal image; and

[0056] The makeup features of the beauty image are obtained by subtracting the features of the makeup removal image from the features of the beauty image.

[0057] Specifically, when it is necessary to obtain the makeup features of a beauty image, the beauty image is first obtained, and then the beauty image is subjected to makeup removal processing to obtain a makeup removal image, and the features of the makeup removal image are subtracted from the features of the beauty image to obtain the makeup features of the beauty image, thereby facilitating the accurate generation of a facial image desired by the user.

[0058] Furthermore, in some embodiments of the present application, the step of adding the makeup features of the beauty makeup image to the initial face image to obtain a target face image includes:

[0059] The makeup features of the beauty makeup image and the initial face image are input into a pre-trained diffusion model to generate the target face image.

[0060] Reference Figure 2The facial image generation method provided in the embodiment of the present application is based on, for example, a selfie given by a user, applies the makeup on a beauty picture given by the user to the selfie, and generates several photos of the user himself with similar makeup. The process and modules used are introduced as follows: the first step is to generate the corresponding makeup removal image based on the beauty picture given by the user using the makeup removal image generation method; the second step is to send the beauty picture to the CLIP visual feature extractor to obtain the global and local features of the beauty picture, and similarly obtain the features of the makeup removal image, and subtract the features of the makeup removal image from the features of the beauty picture to obtain the makeup features that are relatively independent of the facial identity features, and send them to the makeup encoder to obtain the makeup code; the third step is to send the user picture to the face recognizer to obtain the facial identity latent space vector, and send this vector to the face identity encoder to obtain the identity code of the user's face. This code is connected to the makeup code in the previous step using the concat method and sent to the cross-attention part of the UNet of the diffusion model; the fourth step is to send the user picture to the prompt word generator to generate appropriate prompt words according to the gender and age of the user's face, such as "a makeup picture of a girl", "a makeup picture of a girl", etc. The rest of the prompt word can be added with relevant content according to the specific application; the fifth step is to send the user graph to Openpose face ControlNet, which is used to add constraints on the spatial position of the facial features; the sixth step is to use the DDIM inversion algorithm to convert the user graph into the latent space of the diffusion model, input it into the UNet of the diffusion model, and finally obtain the user makeup image after adding makeup effect by combining the above-mentioned features.

[0061] Reference Figure 3 , Figure 3 The schematic diagram of the framework structure of the generation model of the makeup removal image adopted by the face image generation method provided in the embodiment of the present application is a schematic diagram of the framework structure of the generation model of the makeup removal image. The main purpose of the generation model of the makeup removal image is to remove the makeup on the face of the person in the beauty image and restore the effect of the photo without makeup. The specific modules are introduced as follows: 1. The prompt word generation module includes face age and gender judgment. According to the judgment result, prompt words such as "a boy", "an old woman", "a man" are generated. This prompt word is passed into the diffusion model as a positive prompt word, and its corresponding encoded latent space vector is Figure 11. It is represented by p'; 2. p* is the latent space vector of the negative prompt word passed into the diffusion model. This is a special prompt word used to express the command to remove makeup; 3. The beauty picture is converted into a latent space tensor through the VAE encoder, and the tensor is passed into the diffusion model after the DDIM inversion algorithm (DDIMInversion); 4. The output of the diffusion model is converted to the image space through the VAE decoder to obtain the picture after makeup removal; 5. LoRA is a large model fine-tuning technology, which is used here to adjust the cross-attention and self-attention modules in UNet.

[0062] The training process of the makeup removal image generation model is as follows:

[0063] 1. First, prepare training samples. The samples are paired. Each single person picture corresponds to a photo of the same person with makeup (that is, without makeup, the two photos are exactly the same). You can use a large language model to help generate several prompt words, and then use some image editing algorithms based on diffusion models, such as InfEdit, EDICT, etc., to generate the corresponding makeup pictures;

[0064] 2. Based on Figure 2 The model framework is to remove the LoRA module first, and randomly initialize p*. The word sequence length of p* is generally recommended to be set to 5.

[0065] 3. Find the optimal p* based on the following loss function,

[0066] L id It measures the similarity of faces. When t>0.2T, this loss function is defined as 0, because the quality of the generated image is not high at this time. When t≤0.2T, it is defined as Here f t2i Refers to the entire Wensheng graph model, x represents the input image, f t2i (x,p) is the makeup removal map generated based on the input x under the control of condition p; f id refers to the face recognition module, cosine(u,v) is the cosine similarity; ∈θ refers to the UNet model in the Wensheng graph model; ω is generally a real number greater than 1, and the recommended setting is 7.5; λ id The recommended setting is 0.8.

[0067] 4. Lock p*, add the LoRA module back to the framework, and train LoRA according to the same process and loss function as in step 3, except that this step is to find the optimal LoRA weights, and the loss function is

[0068] Furthermore, in some embodiments of the present application, after the inputting of the initial face image, the method further includes:

[0069] Encoding the makeup features of the beauty makeup image to obtain the makeup code of the beauty makeup image;

[0070] Encoding the initial face image to obtain a corresponding identity code; and

[0071] Generate matching prompt words based on the initial face image.

[0072] As mentioned above, it is necessary to obtain the makeup code, identity code and prompt words of the beauty image as input for subsequent processing to ensure the accuracy and authenticity of the generated target face image.

[0073] Reference Figure 4 , Figure 4 This is a schematic diagram of the structure of the makeup encoder used in the face image generation method provided in the embodiment of the present application. In this module, the features of the last several layers of the CLIP visual model are used as the input features of this module, such as Figure 3 As shown in the figure, makeup features are mapped to the latent space of the diffusion model through two multi-layer perceptrons (MLP) and a self-attention module (SelfAttention) to obtain makeup encoding (makeup embedding). For the features extracted from the CLIP visual model, the features of the last 12 layers can be preferably taken, and the number of words mapped to the latent space is recommended to be 6.

[0074] Reference Figure 5 , Figure 5 A schematic diagram of the structure of the image identity encoder used in the face image generation method provided in the embodiment of the present application. Identity encoding (id embedding) is the face recognition module mapping the face to the latent space vector expressing the face. The vector can be used to calculate the similarity of the face, so it contains the information for identifying the face. The latent space vector is mapped to the latent space of the text encoder through a self-attention module (SelfAttention) and a multi-layer perceptron (MLP). The point in the latent space of the text encoder can be regarded as a word, and preferably 4 words are used to express the face identity information.

[0075] Furthermore, in some embodiments of the present application, the makeup features of the beauty makeup image and the initial face image are input into a pre-trained diffusion model to generate the target face image, including:

[0076] The makeup code of the beauty makeup image, the identity code corresponding to the initial face image and the prompt word are input into a pre-trained diffusion model to generate the target face image.

[0077] As described above, when the makeup code, identity code and prompt word of the beauty image are obtained, they are input into the diffusion model for inference calculation to generate the target face image, so as to further ensure the accuracy and authenticity of the generated target face image.

[0078] Furthermore, in some embodiments of the present application, the step of inputting the makeup features of the beauty makeup image and the initial face image into a pre-trained diffusion model to generate the target face image includes:

[0079] According to the makeup features of the beauty makeup image and the initial face image, a target face image is generated through inverse diffusion reasoning of the pre-trained diffusion model.

[0080] SD is based on the diffusion model framework and image-text latent space representation. It plans the sampling steps according to a predefined scheduler and gradually generates the target image from the noise. Specifically, the sampling steps of the diffusion model are set to T, z t represents the latent space variable at step t. The inference stage is a reverse diffusion process, starting from z T Start with z t Calculate z t-1 , step by step, we gradually get z 0 The process of 0 After the transformation from latent space to image space, the generated image is obtained, z T Initialized as a sample from a standard Gaussian distribution, the planner signal-to-noise ratio parameters are {αt,σt} T t=1 , different planners have different settings.

[0081] Furthermore, in some embodiments of the present application, the training process of the diffusion model includes:

[0082] Constructing a training data set, including obtaining a preset number of face images and beauty makeup images to form sample pairs as the training data set;

[0083] The training data set is input into the diffusion model for training, and the diffusion model is optimized by a preset loss function to obtain the trained diffusion model.

[0084] Specifically, prepare training samples. Here, we use the training samples generated by the makeup removal image in the training process, but add the corresponding makeup removal image to each pair of samples. Then, what needs to be trained in the training process is as follows Figure 3The makeup encoder, face identity encoder, and cross-attention module of UNet are shown in the figure. The above modules are added to each cross-attention module in UNet. The newly added modules are used to associate the encoded outputs of the makeup encoder and face identity encoder with the image-related data inside UNet for calculation; the loss function is defined as Here c represents the control condition, L id The meanings of and other symbols have been described above and will not be repeated here.

[0085] Figure 6 FIG. 2 is a structural block diagram of a facial image generating device 200 provided in an embodiment of the present application. Figure 6 As shown, the face image generation device 200 of the embodiment of the present application includes: a makeup acquisition module 210, a face image input module 220 and an image generation module 230, wherein:

[0086] A makeup acquisition module 210 is used to acquire makeup features of a beauty makeup image;

[0087] A face image input module 220, used to input an initial face image; and

[0088] The image generation module 230 is used to add the makeup features of the beauty makeup image to the initial face image to generate a target face image.

[0089] In the facial image generation device provided in the embodiment of the present application, the makeup features of the beauty makeup picture are first obtained, and an initial facial image is input. Finally, the makeup features of the beauty makeup picture are added to the initial facial image to generate a target facial image. The facial image generation method provided in the embodiment of the present application accurately generates the facial image required by the user by adding the makeup features of the beauty makeup picture to the facial image, thereby ensuring the authenticity of the generated facial image.

[0090] It should be noted that the specific implementation method of the face image generation device of the embodiment of the present application is similar to the specific implementation method of the face image generation method of the embodiment of the present application. Please refer to the description of the method part for details, and no further details will be given here.

[0091] Figure 7 Schematic diagram of the structure of an electronic device 300 according to an embodiment of the present application.

[0092] like Figure 7As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from the storage part 302 to a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0093] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as needed, so that a computer program read therefrom is installed into the storage section 308 as needed.

[0094] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 309, and / or installed from a removable medium 311. When the computer program is executed by a central processing unit (CPU) 301, the above-mentioned functions defined in the electronic device of the present application are executed.

[0095] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electronic device, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0096] In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction-executing electronic device, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction-executing electronic device, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0097] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the processing receiving device, method and computer program product according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the aforementioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based electronic device that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0098] The units or modules involved in the embodiments of the present application may be implemented by software or hardware. The units or modules described may also be provided in a processor, and the processor is used to implement the face image generation method when executing the program:

[0099] Obtain makeup features of beauty makeup pictures;

[0100] Input an initial face image; and

[0101] The makeup features of the beauty makeup image are added to the initial face image to generate a target face image.

[0102] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently and not be assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the face image generation method described in the present application:

[0103] Obtain makeup features of beauty makeup pictures;

[0104] Input an initial face image; and

[0105] The makeup features of the beauty makeup image are added to the initial face image to generate a target face image.

[0106] As another aspect, the present application further provides a computer program product, which may be included in the electronic device described in the above embodiment; or may exist independently without being installed in the electronic device. The above computer program product stores one or more programs, and when the above programs are used by one or more processors to execute the facial image generation method described in the present application:

[0107] Obtain makeup features of beauty makeup pictures;

[0108] Input an initial face image; and

[0109] The makeup features of the beauty makeup image are added to the initial face image to generate a target face image.

[0110] The above description is only a preferred embodiment of the present application, and does not limit the patent scope of the present application. All equivalent structural changes made by using the contents of the present application specification and drawings under the application concept of the present application, or directly / indirectly used in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for generating a face image, characterized in that: The following steps are involved: Obtain makeup features of beauty makeup pictures; Input the initial face image; and The makeup features of the beauty makeup image are added to the initial face image to generate a target face image.

2. The method for generating a facial image according to claim 1, characterized in that: The step of obtaining makeup features of the beauty makeup image includes: Obtaining the beauty makeup picture; Performing makeup removal processing on the makeup image to obtain a makeup removal image; and The makeup features of the beauty image are obtained by subtracting the features of the makeup removal image from the features of the beauty image.

3. The method for generating a facial image according to claim 2, characterized in that The step of adding the makeup features of the beauty makeup image to the initial face image to obtain a target face image includes: The makeup features of the beauty makeup image and the initial face image are input into a pre-trained diffusion model to generate the target face image.

4. The method for generating a facial image according to claim 3, characterized in that: After the initial face image is input, the method further includes: Encoding the makeup features of the beauty makeup image to obtain the makeup code of the beauty makeup image; Encoding the initial face image to obtain a corresponding identity code; and Generate matching prompt words based on the initial face image.

5. The method for generating a facial image according to claim 4, characterized in that: Inputting the makeup features of the beauty makeup image and the initial face image into a pre-trained diffusion model to generate the target face image, including: The makeup code of the beauty makeup image, the identity code corresponding to the initial face image and the prompt word are input into a pre-trained diffusion model to generate the target face image.

6. The method for generating a facial image according to claim 3, characterized in that: The step of inputting the makeup features of the beauty makeup image and the initial face image into a pre-trained diffusion model to generate the target face image includes: According to the makeup features of the beauty makeup image and the initial face image, a target face image is generated through inverse diffusion reasoning of the pre-trained diffusion model.

7. The method for generating a facial image according to claim 6, characterized in that: The training process of the diffusion model includes: Constructing a training data set, including obtaining a preset number of face images and beauty makeup images to form sample pairs as the training data set; The training data set is input into the diffusion model for training, and the diffusion model is optimized by a preset loss function to obtain the trained diffusion model.

8. A facial image generating device, characterized in that: include: Makeup acquisition module, used to obtain makeup features of beauty pictures; A face image input module, used for inputting an initial face image; and The image generation module is used to add the makeup features of the beauty makeup image to the initial face image to generate a target face image.

9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the method for generating a facial image according to any one of claims 1 to 7 when executing the program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is used to implement the facial image generation method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Face privacy protection method, system and device and storage medium

    CN117912120A