Face driving model training and face image synthesis method
Through unsupervised training across identity images and multi-path driving strategies, combined with feature vector group weighted summing, the problem of information loss in face image synthesis is solved, and the effect of effectively maintaining complete information during face image synthesis is achieved.
Patent Information
- Application Number
- CN202410223320.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-08-29
AI Technical Summary
In the prior art, face image synthesis cannot effectively maintain the complete information of face images, including face identity information, background information and attribute information of non-face areas.
The unsupervised training strategy and multipath driving strategy across identity images are adopted, and the motion vector synthesis of face images is achieved by designing identity consistency loss, semantic alignment of face key point loss, pseudo-label loss and multipath training regularization loss function, and combining two sets of feature vector groups for weighted summing, the motion vector synthesis of face images is realized.
Effectively maintaining complete information of face images, improving the robustness of face-driven models in different appearances and image features, and achieving better face-driven performance and source face image feature retention capabilities.
Smart Images

Figure CN120563333A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of digital human synthesis, and in particular to a method for training a face-driven model and synthesizing a face image. Background Art
[0002] Face-driven generation involves estimating the motion state of a given driving face image, as depicted by its pose and expression, and transferring this motion information to a source face image, thereby synthesizing a new image with the same expression and pose as the driving face image. A key requirement is to preserve the background features and facial attributes of the source face image, including details such as hair color, texture, facial identity information, face shape, and skin color. These attributes remain unchanged despite changes in facial pose and expression, and are referred to as source motion invariance features. Face-driven generation, due to its rich entertainment features, has a wide range of applications in film and television production, digital entertainment, social media, video conferencing, and other fields.
[0003] With the rapid development of generative models, face-driven generation has made significant progress in recent years. However, due to the limited training methods and the uncertainty in synthesizing the driving face images, current face image synthesis based on face-driven generation still cannot effectively preserve the complete facial image information, such as facial identity information, background information, and attributes of non-face areas. Summary of the Invention
[0004] The embodiments of the present invention provide a face-driven model training and face image synthesis method to at least solve the problem in the related art that face image synthesis cannot effectively maintain the complete information of the face image.
[0005] According to one embodiment of the present invention, a face driving model training method is provided, comprising: inputting a first identity face image as a driving face image together with a source face image into a face driving model to obtain a first synthetic face image; inputting a second identity face image as a driving face image together with the source face image into the face driving model to obtain a second synthetic face image, so as to train the face driving model.
[0006] According to another embodiment of the present invention, a facial image synthesis method is provided, which is applied to a face driving model, including: constructing a first feature vector group of a source facial image and a second feature vector group of a driving facial image; performing weighted summation on the first feature vector group and the second feature vector group to obtain a motion vector of the face in a latent space; and obtaining a motion feature stream of the source facial image based on the motion vector of the face in the latent space to obtain a synthesized facial image.
[0007] According to yet another embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when run.
[0008] According to another embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any one of the above method embodiments.
[0009] The present invention provides a face-driven model training method. This method uses a first-identity face image as a driving face image and inputs it into the face-driven model along with a source face image to obtain a first synthesized face image. A second-identity face image is used as a driving face image and inputted into the face-driven model along with the source face image to obtain a second synthesized face image for training the face-driven model. This method solves the problem in related art where face image synthesis cannot effectively preserve the complete information of the face image, achieving the technical effect of effectively preserving the complete information of the face image during the face image synthesis process. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 1 is a hardware structure block diagram of a computer terminal for a face-driven model training method according to an embodiment of the present invention;
[0011] Figure 2 is a flowchart of a face-driven model training method according to an embodiment of the present invention;
[0012] Figure 3 The process of the face image synthesis method of the embodiment of the present invention is as follows Figure 1 ;
[0013] Figure 4 The process of the face image synthesis method of the embodiment of the present invention is as follows Figure 2 ;
[0014] Figure 5 1 is a system principle diagram of the face-driven model training method according to an embodiment of the present invention;
[0015] Figure 6 2 is a schematic diagram of the principle of a face image encoder according to an embodiment of the present invention;
[0016] Figure 7 2 is a schematic diagram of the principle of image motion synthesis according to an embodiment of the present invention;
[0017] Figure 8 Schematic diagram of the principle of the cross-identity training strategy of an embodiment of the present invention;
[0018] Figure 9This is a flowchart of face image synthesis according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in combination with embodiments.
[0020] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0021] In related technologies, there are generally two technical approaches to achieving face-driven generation: one is motion capture and a graphics rendering pipeline. For example, sensor-based motion capture technology accurately captures the details of an actor's expression and posture, and feeds this information into a rendering model to synthesize the movements and image of a specific actor (such as the movie Avatar). This technology has been widely used in the film and television production industry. The other is based on a generative deep learning model, which learns the facial motion transfer and synthesis process by pre-training on a large number of facial videos. The former method has high accuracy and good results, but requires a lot of face and computer rendering costs, and is usually used in the film and television industry with high quality requirements and high profit output. The latter method has slightly lower generation results, but has the advantages of generalization and low deployment costs. It can achieve driving effects on any face and generally has considerable application potential for consumer-level face-driven generation applications.
[0022] With the rapid development of generative models, face-driven generation has made significant progress in recent years. However, current face image synthesis based on face-driven generation still fails to effectively preserve the complete facial image information, which generally includes facial identity information, background information, and attribute information of non-facial areas. This non-facial area attribute information generally includes details such as hair color and texture, facial identity information, facial shape, and skin color. This is primarily due to two technical characteristics of existing methods that limit their generalization performance.
[0023] First, the training methods of existing methods are insufficient. Existing methods typically only use self-supervised training methods, using facial videos as training data to generate image pairs. Each time, two frames of images are sampled from the same video as training data. The two images usually have different postures and expressions but the same background, skin color, hair, and facial identity information. In this scenario, the difference between the two images seen during model training does not include detailed differences such as background, appearance, and hair, but only differences in expression and posture. However, when testing the model, the driving image is usually a different facial image, and has different detailed features from the source facial image, such as face shape, skin color, background, hat, headwear, etc. These features are sometimes coupled with the intermediate facial representation, resulting in the model's inability to decouple motion information and non-motion representations. To address this shortcoming, the embodiment of the present invention proposes using two different sets of representations to respectively describe the motion information of the face in the driving facial image and the motion of the image area outside the facial part in the source facial image (including hair attached to the head, shoulders, etc.). The motion-related feature weight coefficients of the driving face image and the target motion-related feature vector group are weighted and summed, and the motion-invariant feature weight coefficients of the source face image and the source motion-invariant feature vector group are weighted and summed, and the motion vector of the face in the latent space is obtained by adding them together;
[0024] Secondly, the model has uncertainty when synthesizing driving face images. For example, when driven by two different people, the result image synthesized from the source face image may have a generated face shape or image features that are inconsistent with the attribute information of the source image due to the differences in the driving face images, or artifacts may be caused by occlusion. The embodiments of the present invention propose a cross-identity image unsupervised training strategy and a multi-path driving strategy. During the training process, the model is constrained to keep the generated image with image features similar to the source face image, and corresponding loss functions are designed, including identity consistency loss, semantically aligned face key point loss, pseudo-label loss, and multi-path training regularization loss, thereby improving the robustness of the model for driving face images with different appearances and image features, and achieving better face driving performance and source face image feature retention capabilities.
[0025] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 FIG is a hardware structure block diagram of a computer terminal for the face driving model training method according to an embodiment of the present invention. Figure 1 As shown, the computer terminal may include one or more ( Figure 1Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0026] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the face-driven model training method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0027] The transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a communications provider of a computer terminal. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0028] In an embodiment of the present invention, a face-driven model training method is provided. Figure 2 FIG. 1 is a flow chart of a face-driven model training method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0029] Step S202: The first identity face image is used as a driving face image and inputted into a face driving model together with the source face image to obtain a first synthesized face image.
[0030] Step S204 : The second identity face image is used as a driving face image and inputted into the face driving model together with the source face image to obtain a second synthesized face image for training the face driving model.
[0031] In the embodiment of the present invention, unlike the related art which only uses two frames from the same video as input for network training, the embodiment of the present invention uses three image pairs as input for network training, wherein two of the three image pairs are different frames from the same face video, respectively denoted as the source face image and the driving face image (corresponding to the first identity face image), which have the same appearance and identity information, and the third image is a face image with different appearance and identity information, denoted as the cross-identity face image (corresponding to the second identity face image).
[0032] The embodiment of the present invention adopts a cross-identity image unsupervised training strategy. In the face-driven model training, cross-identity training is introduced by randomly selecting a face image with different identity information from the source face image as a driving face image.
[0033] In an exemplary embodiment, after obtaining the first synthetic facial image and the second synthetic facial image, it also includes: using the second synthetic facial image as a new source facial image, using the first synthetic facial image as a new driving facial image, inputting the new source facial image and the new driving facial image into the facial driving model to obtain a third synthetic facial image.
[0034] In an embodiment of the present invention, a multi-path driving strategy is adopted, the first synthetic face image is selected as a new driving image, the second synthetic face image is used as the source face image, training input is performed, the face driving result is synthesized, and a third synthetic face image is obtained.
[0035] In an embodiment of the invention, a cross-identity image unsupervised training strategy and a multi-path driving strategy are proposed. During the training process, the face driving model is constrained to keep the generated image having image features similar to the source face image, and corresponding loss functions are designed, including identity consistency loss, semantically aligned face key point loss, pseudo-label loss, and multi-path training regularization loss, thereby improving the robustness of the model for driving face images with different appearances and image features, and achieving better face driving performance and source face image feature retention capabilities.
[0036] In an exemplary embodiment, during the process of training the face-driven model, it also includes: calculating the identity information loss between the third synthetic face image and the source face image based on the two-dimensional image face recognition network; and adjusting the third synthetic face image output by the face-driven model according to the identity information loss.
[0037] In an exemplary embodiment, during the process of training the face-driven model, it also includes: respectively obtaining the facial geometric feature parameters of the source face image, the second identity face image, and the third synthetic face image; calculating the facial key point loss of the third synthetic face image and the source face image based on the facial geometric feature parameters; and adjusting the third synthetic face image output by the face-driven model based on the facial key point loss.
[0038] In an exemplary embodiment, the facial geometric feature parameters include at least one of the following: facial shape parameters; facial expression parameters; and facial posture parameters.
[0039] In an exemplary embodiment, during the process of training the face driving model, it also includes: inputting the second identity face image as the driving face image and the source face image into the untrained face driving model to obtain a fourth synthetic face image; calculating the L1 distance loss between the third synthetic face image and the fourth synthetic face image; and adjusting the third synthetic face image output by the face driving model according to the L1 distance loss.
[0040] In this embodiment, a pseudo-label loss is introduced, using a face-driven model that has not undergone cross-identity training—that is, a model retrained without the source motion-invariant feature vector set—as the pseudo-label generation module. A fourth synthesized face image is obtained, and this driven result serves as the pseudo-label. The L1 distance loss is calculated between the third and fourth synthesized face images to improve training stability.
[0041] In an exemplary embodiment, during the process of training the face-driven model, it also includes: calculating a norm regularization loss of the third synthetic face image and the first identity face image; and adjusting the training parameters of the face-driven model according to the norm regularization loss.
[0042] In an embodiment of the invention, methods for calculating a norm regularization loss include but are not limited to using absolute distance loss (L1), mean squared error loss (MSE), and structural similarity index (SSIM).
[0043] Through the above steps, a face-driven model training method is provided. A first-identity face image is input into the face-driven model along with a source face image as a driving face image to obtain a first synthesized face image. A second-identity face image is input into the face-driven model along with the source face image as a driving face image to obtain a second synthesized face image for training the face-driven model. This method solves the problem in related art where face image synthesis cannot effectively preserve the complete information of the face image, achieving the technical effect of effectively preserving the complete information of the face image during the face image synthesis process.
[0044] The embodiment of the present invention also provides a face image synthesis method, which is applied to a face driving model. Figure 3 The process of the face image synthesis method of the embodiment of the present invention is as follows Figure 1 ,like Figure 3 As shown, the process includes the following steps:
[0045] Step S302 : constructing a first feature vector group of the source face image and a second feature vector group of the driving face image.
[0046] In an exemplary embodiment, the feature parameters of the first feature vector group include at least one of the following: skin color feature parameters of the source facial image; hair feature parameters of the source facial image; and background feature parameters of the source facial image.
[0047] In an exemplary embodiment, the feature parameters of the second feature vector group include at least one of the following: a facial expression parameter driving the facial image; and a facial posture parameter driving the facial image.
[0048] Step S304: performing weighted summation on the first eigenvector group and the second eigenvector group to obtain a motion vector of the face in the latent space.
[0049] In an exemplary embodiment, a first eigenvector group and a second eigenvector group are weightedly summed to obtain a motion vector of a face in a latent space, including: encoding a source face image and a driving face image respectively to obtain a first feature weight coefficient and a second feature weight coefficient; and weighted summing the first feature weight coefficient and the first eigenvector group, the second feature weight coefficient, and the second eigenvector group respectively to obtain a motion vector of the face in a latent space.
[0050] Step S306 , obtaining a motion feature stream of the source face image according to the motion vector of the face in the latent space, to obtain a synthesized face image.
[0051] Figure 4 The process of the face image synthesis method of the embodiment of the present invention is as follows Figure 2 ,like Figure 4As shown in FIG, after obtaining the motion feature stream of the source face image according to the motion vector of the face in the latent space, the method further includes: performing feature distortion and image synthesis on the face coding features according to the motion feature stream to obtain a synthesized face image. Figure 4 As shown, the process of the face image synthesis method includes the following steps:
[0052] Step S402: constructing a first feature vector group of the source face image and a second feature vector group of the driving face image.
[0053] Step S404: performing weighted summation on the first eigenvector group and the second eigenvector group to obtain a motion vector of the face in the latent space.
[0054] Step S406: Obtain a motion feature stream of the source face image based on the motion vector of the face in the latent space.
[0055] Step S408 , performing feature distortion and image synthesis on the facial coding features according to the motion feature stream to obtain a synthesized facial image.
[0056] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method described in each embodiment of the present invention.
[0057] In this embodiment, a face-driven model training device or a face image synthesis device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. The details that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation by hardware, or a combination of software and hardware, is also possible and conceivable.
[0058] In one embodiment, the face driving model training device provided by the present invention includes a first training module and a second training module. The first training module is configured to input a first identity face image as a driving face image, along with a source face image, into the face driving model to obtain a first synthesized face image. The second training module is configured to input a second identity face image as a driving face image, along with a source face image, into the face driving model to obtain a second synthesized face image for training the face driving model.
[0059] In one embodiment, the face image synthesis device provided by the present invention includes: a vector acquisition module, a coefficient acquisition module, and a feature stream acquisition module, wherein the vector acquisition module is used to construct a first feature vector group of the source face image and a second feature vector group of the driving face image, the coefficient acquisition module is used to perform weighted summation on the first feature vector group and the second feature vector group to obtain the motion vector of the face in the latent space, and the feature stream acquisition module is used to obtain the motion feature stream of the source face image based on the motion vector of the face in the latent space to obtain a synthesized face image.
[0060] In actual implementation, the above-mentioned face-driven model training device or face image synthesis device may further include other modules, wherein the functional division and naming of different modules may also be different, as long as the steps of the face-driven model training method or face image synthesis method in the above-mentioned embodiment can be implemented. It should be noted that each of the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but is not limited to this: the above-mentioned modules are all located in the same processor; or the above-mentioned modules are located in different processors in any combination.
[0061] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when running.
[0062] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0063] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0064] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0065] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0066] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0067] In order to enable those skilled in the art to better understand the technical solution of the present invention, it is described below in conjunction with scenario embodiments.
[0068] Example 1
[0069] In the embodiment of the present invention, the face image x is driven d That is, the first identity face image in the above embodiment, the cross-identity face image x o That is, the second identity face image in the above embodiment, the source face image x s That is, the source face image in the above embodiment, the synthesized face image x s→d That is, the first synthetic face image in the above embodiment, the cross-identity synthetic face image x s→o That is, the second synthetic face image in the above embodiment, the multi-path driven synthetic face image x mul This is the third synthetic face image in the above embodiment, the cross-identity pseudo-label image is the fourth synthetic face image in the above embodiment, the target motion-related feature vector group is the second feature vector group in the above embodiment, and the source motion-invariant feature vector group is the first feature vector group in the above embodiment.
[0070] Figure 5 : is a system principle diagram of the face driven model training method according to an embodiment of the present invention. Figure 5 As shown, the driving face image x d : A face image used to provide expression and posture. Source face image x s : The driven face image is usually used to provide face attributes and identity details. It is necessary to ensure that the identity information and attributes of the synthesized face are consistent with the source face image. o: A driving face image, which is not the same person as the source face image, and is used to provide a face image with expression and posture. Driving face generation model: A model used to synthesize face images, which contains the synthesis network involved in the implementation of the present invention, the target motion-related feature vector group and the source motion-invariant feature vector group. Training strategy: includes a multi-path inference regularization loss function, a multi-path driving strategy, and a cross-identity training strategy.
[0071] In the embodiment of the present invention, the face image x is driven d and source face image x s These are usually two randomly sampled face images from the same face video. Cross-identity face images provide input for the subsequent cross-identity image unsupervised training strategy.
[0072] The driven face generation model consists of two main parts, namely the image coding mapping part and the image motion synthesis part. Figure 6 FIG. 1 is a schematic diagram of the principle of a face image encoder according to an embodiment of the present invention. Figure 6 As shown in the figure, in the image coding mapping part, any face image will be encoded by the convolutional neural network to obtain four outputs, namely the motion correlation feature weight coefficient a, the motion invariance feature weight coefficient b, the face latent space z and the encoder feature. Figure 7 FIG. 1 is a schematic diagram of the principle of image motion synthesis according to an embodiment of the present invention. Figure 7 As shown in the figure, in the image motion synthesis part, the four sets of encoder outputs are used as input, and a, b, z and two learnable orthogonal decoupling vector groups are used to obtain the combined latent motion representation d. d and the encoder features are fed into the feature flow motion prediction network, the feature warping module, and the image synthesis network to synthesize the driven face image.
[0073] The training strategy adopted by the face-driven model training method of the embodiment of the present invention is as follows: In the embodiment of the present invention, an optimizer is used to train the network so that the network has synthesis capabilities, and an accurate loss function needs to be designed. Figure 5 As shown, there are three loss paths, namely similarity reconstruction loss, cross-identity training strategy, and multi-path inference regularization loss function.
[0074] Among them, the similarity reconstruction loss function is also called the similarity loss function, which is used for driving faces of the same identity. The cross-identity loss function of the cross-identity training strategy, including identity consistency loss, semantically aligned face key point loss, and pseudo-label loss, are all used for cross-identity face driving training. The cross-identity loss function proposed in the embodiment of the present invention is applicable to the multi-path driving strategy and cross-identity training strategy in the embodiment of this application. The multi-path inference regularization loss function is used for multi-path inference loop reconstruction.
[0075] Example 2
[0076] In the second embodiment, the training strategy of the face-driven model of the embodiment of the present invention is introduced in detail.
[0077] Related technologies use only self-supervised training, using facial videos as training data to generate image pairs. Each time, two frames are sampled from the same video as training data. These two images typically have different poses and expressions but share the same background, skin color, hair, and facial identity information. These features are sometimes coupled with the intermediate representation of the face, resulting in the model's inability to decouple motion information from non-motion representations. Therefore, when the trained model is put into use (including during testing), the person driving the facial image and the person being driven by the source facial image are often different people, leading to inconsistencies between training and testing, and consequently, performance loss.
[0078] Therefore, in the embodiment of the present invention, cross-identity training is introduced during training and these losses are proposed to stabilize the training and improve the performance of the model during testing. The embodiment of the present invention proposes a cross-identity image unsupervised training strategy, such as Figure 5 As shown, the present invention uses three image pairs as input, where two of the three image pairs are different frames from the same face video, respectively denoted as the source face image and the driving face image, which have the same appearance and identity information, and the third image is a face image with different appearance and identity information, denoted as a cross-identity face image.
[0079] The embodiment of the present invention adopts a cross-identity image unsupervised training strategy, such as Figure 5 As shown in the figure, in the face-driven model training, cross-identity training is introduced by randomly selecting face images with different identity information from the source face images as driving face images. During training, identity consistency loss, semantically aligned face key point loss, and pseudo-label loss are added to the results of cross-identity synthesis to improve the stability of training and the model's ability to maintain motion invariance features.
[0080] Figure 8 This is a schematic diagram of the principle of the cross-identity training strategy of an embodiment of the present invention. Figure 8 As shown, by constructing face image pairs of the same appearance identity as same-identity training, and constructing face image pairs of different appearance identities as cross-identity training image pairs, identity consistency loss, semantically aligned face key point loss, and pseudo label loss are added to the results of cross-identity synthesis during training.
[0081] In one embodiment, if Figure 8 As shown, identity consistency loss is introduced and the synthetic face image x is evaluated by the two-dimensional image face recognition network. s→o With the source face image xs loss of identity information.
[0082] In one embodiment, if Figure 8 As shown, the semantic alignment of facial key points loss is introduced, and x is extracted through the 3D face model s The face shape parameters, combined with the o The extracted expression and posture parameters, combined with the shape parameters, expression parameters and posture parameters, can be used to obtain a face model with the shape of the source face and the expression and posture of the driving face image through a three-dimensional face model. The face key point coordinates are obtained using this model. The coordinates are consistent with x s→o The root mean square error loss is calculated based on the facial key point coordinates to obtain the semantically aligned facial key point loss.
[0083] In one embodiment, if Figure 8 As shown in , the pseudo label loss is introduced, and the face-driven model that has not been trained across identities is used as the pseudo label generation module, that is, the model that is retrained without the source motion invariance feature vector group is removed. o Driver x s The result of the drive is used as a pseudo label to calculate x s→o The L1 distance loss with the driving result improves the stability of training.
[0084] In related technologies, the face-driven model has uncertainty when synthesizing the driving face image. For example, when driven by two different people, the result image synthesized from the source face image may have a generated face shape or image features that are inconsistent with the attribute information of the source image due to the differences in the driving face images, or artifacts may be caused by occlusion.
[0085] The embodiments of the present invention propose a cross-identity image unsupervised training strategy and a multi-path driving strategy. During the training process, the face driving model is constrained to maintain the generated image with image features similar to the source face image, and corresponding loss functions are designed, including identity consistency loss, semantically aligned face key point loss, pseudo-label loss, and multi-path training regularization loss, thereby improving the robustness of the model for driving face images with different appearances and image features, and achieving better face driving performance and source face image feature preservation capabilities.
[0086] like Figure 5As shown, an embodiment of the present invention proposes a regularization method, namely a multi-path driving strategy and a multi-path training regularization loss function. Using a pair of face images of different appearances and identities as input, outputting a cross-identity synthesized face image as a new source face image, introducing a face image with the same identity information as the source face image as a driving image, and obtaining the path-driven synthesis result. The path-driven synthesis result and the driving image are calculated with loss to construct a multi-path training regularization loss. The multi-path training regularization loss function calculates the multi-path driven synthesized face image and the original driving face image x d The norm regularization loss is used to help train the model. The calculation methods include but are not limited to absolute distance loss L1, root mean square error loss MSE, and structural similarity loss SSIM.
[0087] The embodiment of the present invention adopts a multi-path driving strategy and a multi-path training regularization loss function. In the face driving model training, the cross-identity synthetic face image obtained through training is used as the new source face image, and a face image with the same facial identity information is introduced as the driving image. The multi-path training regularization loss is added to the path driving result to improve the model's recovery and reconstruction of background and facial motion-irrelevant details.
[0088] In one embodiment, if Figure 5 As shown, a multi-path driving strategy is adopted to select the driving face image x d As a new driving image, training input is performed to obtain a cross-identity face driving image, which is used as the source face image to synthesize the face driving result, denoted as x mul .
[0089] like Figure 5 As shown, the multi-path training regularization loss function calculates x mul With the original driver image x d The norm regularization loss is used to help train the model. The calculation methods include but are not limited to absolute distance loss L1, root mean square error loss MSE, and structural similarity loss SSIM.
[0090] The above embodiment uses two groups of feature vectors to represent the modeling of target motion correlation features and source motion invariance features in facial images, and adds vector mutual orthogonal constraints therein to help learn the decoupling of facial motion and facial attributes. The motion representation of the latent space is obtained by merging the target motion correlation feature weight coefficients of the driving image and retaining the source motion invariant feature weight coefficients of the source facial image. The present invention combines two efficient training strategies and a loss function, namely, through the cross-identity image unsupervised training strategy and the multi-path driving strategy, as well as the multi-path training regularization loss function, to effectively improve the model's ability to synthesize accurate expressions and postures in the face of different identity images during testing, while maintaining the consistency and coherence of the motion invariant features in the source facial image, thereby achieving high-fidelity face-driven generation of details.
[0091] Example 3
[0092] In the third embodiment, a face driving model trained by the method in the above embodiment is used to perform face image synthesis.
[0093] In the embodiment of the present invention, the face image x is driven d That is, the first identity face image in the above embodiment, the cross-identity face image x o That is, the second identity face image in the above embodiment, the source face image x s That is, the source face image in the above embodiment, the synthesized face image x s→d That is, the first synthetic face image in the above embodiment, the cross-identity synthetic face image x s→o That is, the second synthetic face image in the above embodiment, the multi-path driven synthetic face image x mul This is the third synthetic face image in the above embodiment, the cross-identity pseudo-label image is the fourth synthetic face image in the above embodiment, the target motion-related feature vector group is the second feature vector group in the above embodiment, and the source motion-invariant feature vector group is the first feature vector group in the above embodiment.
[0094] Face driving is a common task in image generation. Consider two face images, a and b. Face driving involves grafting the expression and head posture of face a onto face b. This means making face b make the same expression and head movements as face a. The resulting new face image is still of person b, but with a different head posture and expression from the original image b. Therefore, face driving generally involves three images: a source face image, a driving face image (also commonly referred to as a target image), and a synthesized face image. The synthesized face image has the same expression and head posture as the driving face image, as well as the same facial identity, wrinkles, skin color, earrings, hair, shoulders, and other image details as the source face image.
[0095] In this embodiment of the present invention, unlike related art methods that only use two frames from the same video as input for network training, three image pairs are used as input for network training. Two of the three image pairs are different frames from the same face video, respectively denoted as the source face image and the driving face image, which have the same appearance and identity information. The third image is a face image with different appearance and identity information, denoted as the cross-identity face image.
[0096] The embodiment of the present invention utilizes the idea of decoupled learning to decompose the above-mentioned face driving problem into target motion-related features and source motion-invariant features, and converts the operation of the pixel space into encoding changes in the potential decoupled space. Combined with the cross-identity image unsupervised training strategy and multi-path training regularization loss, high-fidelity face driving can be achieved with facial details and background details.
[0097] Figure 9 FIG. 1 is a flow chart of face image synthesis according to an embodiment of the present invention. Figure 9 As shown, the following steps are included:
[0098] Step S902 : constructing a source motion invariant feature vector group of the source face image and a target motion related feature vector group of the driving face image.
[0099] Two sets of feature vectors are used to represent the target motion-related features and the source motion-invariant features in facial images. Orthogonal constraints are added to these sets to learn the decoupling of facial motion and facial attributes. For the face-driven generation task, two sets of mutually orthogonal vectors are defined, one for the target motion-related feature vectors and one for the source motion-invariant feature vectors.
[0100] In actual implementation, the process of defining two mutually orthogonal vector groups is as follows:
[0101] S9021 uses matrix orthogonal triangular decomposition to perform orthogonal triangular decomposition on a learnable predefined matrix to obtain an upper triangular matrix and a standard orthogonal vector group matrix, where each column vector of the standard orthogonal vector group matrix is orthogonal to each other.
[0102] S9022, separate the first m (m is a positive integer) column vectors {m i |1≤i≤m,m i ∈R N} is defined as the target motion related feature vector group M, and the remaining n column vectors {i i |1≤i≤n,i i ∈R N} is defined as the source motion invariant feature vector group Q, and the relative size of m and n is set to m>=n.
[0103] Step S904 , performing information encoding on the source face image and the driving face image to obtain two sets of weight coefficients, which are respectively allocated as weight coefficients of motion-related features and source motion-invariant features.
[0104] use Figure 6 The face encoder shown encodes information of the face image (including the source face image and the driving face image) and maps the face image to two one-dimensional weight vectors, which respectively represent the motion feature weight coefficients of the known face image. and motion invariance feature weight coefficient where a d is a one-dimensional vector containing m parameters, b d is a one-dimensional vector containing n parameters.
[0105] Step S906: The motion-related feature weight coefficients are weighted and summed with the target motion-related feature vector group, and the source motion-invariant feature weight coefficients are weighted and summed with the source motion-invariant feature vector group, and the motion vector of the face in the latent space is obtained by adding them together.
[0106] Among them, the face latent space z s It is a hypothetical intermediate state of facial movement, corresponding to the frontal face and neutral expression. The intermediate state plus the motion vector represents the movement process from the source face to the driving face.
[0107] In actual implementation, the process of calculating the motion vector of the face in the latent space is as follows:
[0108] S9061, select source face image x s and driving face image x d , and obtain the motion feature weight coefficient a of the source face image s and motion invariance feature weight coefficient b s And the motion feature weight coefficient a that drives the face image d and motion invariance feature weight coefficient b d ;
[0109] S9062, multiply the target motion related feature vector group M and the source motion invariant feature vector group Q by a respectively. d with b s , get the motion vector of the face in the latent space:
[0110]
[0111] Among them, M and Q are obtained during model training and have nothing to do with specific face images.
[0112] In an embodiment of the present invention, the motion vector of the face in the latent space is obtained by merging the target motion correlation feature weight coefficients of the driving image and retaining the source motion invariant feature weight coefficients of the source face image. The motion vector can be mapped from the latent space motion representation to the image feature motion flow through a dual-branch feature motion flow prediction network, and the final face image is obtained by combining multi-resolution image fusion.
[0113] Step S908: Obtain a motion feature stream of the source face image based on the motion vector of the face in the latent space.
[0114] The latent motion flow network is used to estimate the motion feature flow adapted to the source face image from the motion vector of the face in the latent space, and the face coding feature is obtained according to the information coding in step S904.
[0115] S9081, using the potential motion flow network to estimate the motion feature flow φ at different resolutions using the potential motion representation calculated from S3 s , where s is the resolution width and height of the motion stream output by different network layers, s = 4, 8, 16, 32, 64, 128, 256;
[0116] S9082, using motion flows at different resolutions to perform feature backward warping operations on encoder features of the source face image with a resolution of s×s, to obtain deformation features F at different resolutions s , as the face encoding feature.
[0117] Step S910 , performing feature distortion and image synthesis on the facial coding features according to the motion feature stream to obtain a synthesized facial image.
[0118] The image synthesis network is used to transform the deformation features F of different resolutions s The face images are projected into different resolutions, and the final driven face image x is obtained by upsampling and adding (layer by layer progressive addition). s→d , the calculation formula is as follows:
[0119]
[0120] Among them, toRGB is a two-dimensional convolution layer with an output channel of 3, which maps the deformation features from multiple channels to an image with 3 channels represented by RGB. 8-k For the upsampling module, the image resolution is upsampled by a factor of 2 8-k times.
[0121] In summary, the embodiments of the present invention provide a face-driven model training and face image synthesis method, which is mainly aimed at the defect that the current face-driven generation technology is still unable to better maintain the face identity information, background information, and attribute information of non-face areas. The present invention defines two groups of decoupled motion-related feature vectors and invariant feature vectors, encodes and predicts the face image to obtain the weight coefficients of motion-related features and motion-invariant features, and calculates the potential motion vector of the face by weighted summation of the weight coefficients and the feature vector group. Two networks (latent motion flow network and image synthesis network) are used to learn motion feature flow synthesis and face image synthesis from the latent motion vector and source face image features respectively. The embodiment of the present invention implicitly learns the motion of each region in the source face image, constructs a more reasonable global face feature motion expression, and effectively improves the model's ability to synthesize accurate facial expressions and postures when facing cross-identity face driving through a novel cross-identity image unsupervised training strategy and a multi-path driven training strategy, while maintaining the consistency of the face's identity features and attribute features, and realizing face-driven generation with high fidelity details.
[0122] In this embodiment, a cross-identity unsupervised training strategy, a multi-path driving strategy, and a multi-path training regularization loss function are used to enhance motion decoupling and facial attribute feature preservation. This effectively improves the facial driving model's ability to synthesize accurate expressions and postures across different identity images during testing, while maintaining the consistency and coherence of the motion-invariant features in the source facial images, achieving high-fidelity facial driving generation.
[0123] In an embodiment of the present invention, two sets of feature vectors are used to represent the target motion correlation features and source motion invariance features in facial images, and the addition of a mutual orthogonality constraint between the vectors facilitates learning the decoupling of facial motion from facial attributes. A motion representation of the latent space is obtained by combining the motion correlation feature weight coefficients of the driving image while retaining the motion invariance feature weight coefficients of the source facial image. This motion representation is mapped from the latent space motion representation to an image feature motion stream through two parallel networks, and the final facial image is obtained by combining multi-resolution image fusion.
[0124] The face-driven model training and face image synthesis methods provided by the embodiments of the present invention are applicable to scenarios such as digital virtual humans, conference virtual humans, digital humans, and mobile digital human-driven applications. These methods address the problem of related face-driven technologies failing to adequately preserve facial identity information, background information, and attribute information of non-face areas, enabling high-fidelity face-driven processing with both facial and background details. These methods can be run on computers, including but not limited to Windows, Linux, and Mac OS operating systems.
[0125] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A face-driven model training method, characterized in that: include: The first identity face image is used as a driving face image and inputted into the face driving model together with the source face image to obtain a first synthesized face image; The second identity face image is used as a driving face image and is input into the face driving model together with the source face image to obtain a second synthetic face image to train the face driving model.
2. The method according to claim 1, characterized in that After obtaining the first synthesized face image and the second synthesized face image, the method further includes: The second synthesized face image is used as a new source face image, the first synthesized face image is used as a new driving face image, the new source face image and the new driving face image are input into the face driving model to obtain a third synthesized face image.
3. The method according to claim 2, characterized in that Also includes: Calculating identity information loss between the third synthesized face image and the source face image based on a two-dimensional image face recognition network; The third synthesized face image output by the face driving model is adjusted according to the identity information loss.
4. The method according to claim 2, characterized in that Also includes: Respectively obtaining facial geometric feature parameters of the source facial image, the second identity facial image, and the third synthesized facial image; Calculating the facial key point loss of the third synthesized facial image and the source facial image based on the facial geometric feature parameters; The third synthesized face image output by the face-driven model is adjusted according to the face key point loss.
5. The method according to claim 4, characterized in that in, The facial geometric feature parameters include at least one of the following: Face shape parameters; Facial expression parameters; Face pose parameters.
6. The method according to claim 2, characterized in that Also includes: inputting the second identity face image as the driving face image and the source face image into an untrained face driving model to obtain a fourth synthesized face image; Calculating an L1 distance loss between the third synthesized face image and the fourth synthesized face image; The third synthesized face image output by the face driven model is adjusted according to the L1 distance loss.
7. The method according to claim 2, characterized in that Also includes: Calculating a norm regularization loss of the third synthesized face image and the first identity face image; Adjusting training parameters of the face-driven model according to the one-norm regularization loss.
8. A facial image synthesis method, applied to a face-driven model, characterized in that: include: Constructing a first feature vector group of the source face image and a second feature vector group of the driving face image; Performing a weighted summation on the first eigenvector group and the second eigenvector group to obtain a motion vector of the face in a latent space; According to the motion vector of the face in the latent space, a motion feature stream of the source face image is obtained to obtain a synthesized face image.
9. The method according to claim 8, characterized in that in, The characteristic parameters of the first characteristic vector group include at least one of the following: Skin color feature parameters of the source facial image; Hair feature parameters of the source facial image; Background feature parameters of the source facial image.
10. The method according to claim 8, characterized in that in, The characteristic parameters of the second characteristic vector group include at least one of the following: Facial expression parameters of the driving facial image; The facial posture parameters of the driving facial image.
11. The method according to claim 8, characterized in that The step of performing weighted summation on the first eigenvector group and the second eigenvector group to obtain a motion vector of the face in the latent space includes: Encoding the source face image and the driving face image respectively to obtain a first feature weight coefficient and a second feature weight coefficient; A weighted sum is performed on the first feature weight coefficient and the first feature vector group, and the second feature weight coefficient and the second feature vector group, respectively, to obtain a motion vector of the face in the latent space.
12. The method according to claim 8, characterized in that After obtaining the motion feature stream of the source face image according to the motion vector of the face in the latent space, the method further includes: Feature distortion and image synthesis are performed on the facial coding features according to the motion feature stream to obtain the synthesized facial image.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 12 is implemented.