A face driving and live broadcast method, device, computer equipment and storage medium

By synthesizing facial images in face driving and combining the background features of the source image, the problems of high training difficulty and low quality are solved, and higher quality face driving effects are achieved.

CN113486787BActive Publication Date: 2025-09-19GUANGZHOU HUYA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110756772.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-05
Publication Date
2025-09-19
Estimated Expiration
2041-07-05

AI Technical Summary

Technical Problem

In existing technologies, face-driven training is difficult and of low quality, mainly due to the many interference factors caused by background changes in video data.

Method used

By acquiring the source image and the driving image, the facial appearance features and posture expression features are extracted, the synthetic facial image is synthesized, and the target driving image is generated by combining the background features of the source image to avoid interference from background changes.

Benefits of technology

The difficulty of face-driven training is reduced, the quality of face-driven is improved, and the generated target-driven images are made more vivid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113486787B_ABST
    Figure CN113486787B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a facial driving and live broadcasting method, apparatus, computer equipment and storage medium. The method comprises: obtaining a source image and a driving image, wherein the source image and the driving image include facial data of different objects; synthesizing at least one synthetic facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the driving image; synthesizing a target driving image based on the facial features of each synthetic facial image and the background features of the source image. The technical solution of the embodiment of the present invention realizes the recombination of the source image and the driving image, thereby driving the presentation of facial expressions under different facial models. It can be applied to application scenarios such as live broadcasting, and solves the problem of the existing difficulty and low quality of facial driving training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a face driving and live broadcast method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the development of artificial intelligence technology, smart terminals are widely used in various aspects such as learning, entertainment, and work. For example, live broadcasts can be carried out through smart terminals to achieve information exchange in various aspects such as learning, entertainment, or work.

[0003] In applications such as live broadcasting, virtual reality, and designing character expressions, the current character's face can be used to drive another character to present the same expression.

[0004] However, in the existing technology, when performing face driving, it is necessary to train through videos, and the background in the video data is constantly changing, causing many interference factors, making face driving training difficult and of low quality. Summary of the Invention

[0005] Embodiments of the present invention provide a face driving and live broadcasting method, apparatus, computer equipment, and storage medium, which can reduce the difficulty of face driving training and improve the quality of face driving.

[0006] In a first aspect, an embodiment of the present invention provides a face driving method, comprising:

[0007] Acquire a source image and a driving image, wherein the source image and the driving image include facial data of different objects;

[0008] synthesizing at least one synthesized facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the driving image;

[0009] The target driving image is synthesized based on the facial features of each synthesized facial image and the background features of the source image.

[0010] In a second aspect, an embodiment of the present invention further provides a live broadcast method, comprising:

[0011] Receive live video data uploaded by the anchor client, and extract multiple live video frames from the live video data, wherein the live video frames include facial data of the live user;

[0012] Obtaining a reference image selected by the host client, the reference image including facial data of the set subject;

[0013] synthesizing at least one synthesized facial image corresponding to each live video frame based on the facial appearance features extracted from the reference image and the facial posture and expression features extracted from each live video frame;

[0014] Based on the facial features of each synthesized facial image corresponding to each live video frame and the background features of the reference image, a target driving image corresponding to each live video frame is synthesized;

[0015] After each target driving image is used to replace each live video frame in the live video data, the live video data is published in the live broadcast room of the anchor client.

[0016] In a third aspect, an embodiment of the present invention further provides a face driving device, comprising:

[0017] An image acquisition module, configured to acquire a source image and a driving image, wherein the source image and the driving image include facial data of different objects;

[0018] A synthetic facial image synthesis module, configured to synthesize at least one synthetic facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the driving image;

[0019] The target driving image synthesis module is used to synthesize the target driving image according to the facial features of each synthesized facial image and the background features of the source image.

[0020] In a fourth aspect, an embodiment of the present invention further provides a live broadcast device, comprising:

[0021] A live video frame extraction module is used to receive live video data uploaded by the anchor client and extract multiple live video frames from the live video data, wherein the live video frames include facial data of the live user;

[0022] A reference image acquisition module is used to obtain a reference image selected by the anchor client, wherein the reference image includes facial data of a set object;

[0023] A synthetic facial image synthesis module is used to synthesize at least one synthetic facial image corresponding to each live video frame based on the facial appearance features extracted from the reference image and the facial posture and expression features extracted from each live video frame;

[0024] A target driven image synthesis module is used to synthesize target driven images corresponding to each live video frame based on the facial features of each synthesized facial image corresponding to each live video frame and the background features of the reference image;

[0025] The live video data publishing module is used to replace each live video frame in the live video data with each target driving image, and then publish the live video data in the live broadcast room of the anchor client.

[0026] In a fifth aspect, an embodiment of the present invention further provides a computer device, the computer device comprising:

[0027] one or more processors;

[0028] a storage device for storing one or more programs,

[0029] When one or more programs are executed by one or more processors, the one or more processors implement the face driving method or live broadcast method provided by any embodiment of the present invention.

[0030] In a sixth aspect, an embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the face driving method or live broadcast method provided by any embodiment of the present invention is implemented.

[0031] The technical solution of the embodiment of the present invention obtains a source image and a driving image, wherein the source image and the driving image include facial data of different objects; synthesizes at least one synthetic facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the driving image; synthesizes a target driving image based on the facial features of each synthetic facial image and the background features of the source image, thereby solving the problem of facial driving and reducing the difficulty of facial driving training by recombining the source image and the driving image; and synthesizes the target driving image by adding the background features of the source image to the synthetic facial image, thereby improving the quality of facial driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flowchart of a face driving method provided in Example 1 of the present invention;

[0033] Figure 2a is a flowchart of a face driving method in embodiment 2 of the present invention;

[0034] Figure 3 This is a flow chart of a live broadcast method in Embodiment 3 of the present invention;

[0035] Figure 4 1 is a structural diagram of a face driving device in a fourth embodiment of the present invention;

[0036] Figure 5 This is a structural diagram of a live broadcast device in Embodiment 5 of the present invention;

[0037] Figure 6 A structural diagram of a computer device provided in Example 6 of the present invention. DETAILED DESCRIPTION

[0038] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0039] It should also be noted that, for ease of description, only the part relevant to the present invention, rather than all of the content, is shown in the accompanying drawings. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processing or methods depicted as flow charts. Although flow charts describe various operations (or steps) as sequential processing, many operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of various operations can be rearranged. When its operation is completed, processing can be terminated, but can also have additional steps not included in the accompanying drawings. Processing can correspond to methods, functions, procedures, subroutines, subprograms, etc.

[0040] Example 1

[0041] Figure 1 This is a flow chart of a facial driving method provided by the first embodiment of the present invention. This embodiment is applicable to situations where the facial gestures and expressions of the host's face are used to drive the host's selected characters to present the same expression during live broadcasts. This method can be executed by a facial driving device, which can be implemented by software and / or hardware and can generally be integrated into a computer device used for facial driving (for example, various smart terminals or servers, etc.). Figure 1 As shown, the method includes:

[0042] Step 110: Acquire a source image and a driving image.

[0043] The source image and the driving image include facial data of different subjects. Specifically, the source image can include facial shape data of the intended person. This means that by obtaining facial appearance features from the source image, the target driving image can have the same facial appearance as the source image. The driving image can include facial expression data of the intended person. This means that by obtaining facial posture and expression features from the driving image, the target driving image can have the same expression as the driving image. The source image and the driving image can be high-definition images of different people.

[0044] For example, the source image can be an image of a desired character selected by the host in a live broadcast application scenario. The driving image can be images of the live broadcast with different facial expressions. By combining the source image and the driving image, the host's expression can be specifically rendered for different facial shapes.

[0045] Step 120: Synthesize at least one synthesized facial image based on the facial appearance features extracted from the source image and the facial posture and expression features extracted from the driving image.

[0046] Facial appearance features may be data related to a person's facial features. Specifically, facial appearance features may include a facial shape vector and facial texture vector of the person in the source image. Facial posture and expression features may be data related to a person's facial posture and expression. Specifically, facial posture and expression features may include a facial expression vector and facial angle vector of the person in the driving image. Facial appearance features and facial posture and expression features may be vector data obtained through, for example, a deep learning network.

[0047] In an embodiment of the present invention, a synthesized facial image may be an image synthesized by synthesizing facial appearance features in a source image and facial posture and expression features in a driving image. Specifically, the synthesized facial image may be a specific embodiment of facial appearance features in a source image and facial posture and expression features in a driving image in different ways. For example, the synthesized facial image may be a rendered image generated by facial appearance features in a source image and facial posture and expression features in a driving image; or the synthesized facial image may be a depth image generated by facial appearance features in a source image and facial posture and expression features in a driving image. The rendered image may be an image generated with the expression in the driving image and the facial shape of the person in the source image. The depth image may be a rendered image supplemented with facial depth information, such as depth information of the person's facial features.

[0048] Step 130: synthesize the target driving image based on the facial features of each synthesized facial image and the background features of the source image.

[0049] Facial features may include facial appearance features and facial posture and expression features in the synthesized facial image. Furthermore, facial features may also include depth features of facial features. Background features of the source image may include features other than facial appearance features. For example, background features may include features related to a person's hair, neck, clothing, and the scene in which the person is located. The target driving image may be the image corresponding to the expression in the driving image when specifically represented in the source image.

[0050] In an embodiment of the present invention, in addition to features related to the human face, the target drive image also has background features, which can improve the quality of face driving and make the generated target drive image more vivid and realistic. The target drive image is obtained based on the source image and the drive image. The background does not change, and there is no need for complex interference data cleaning operations, which can reduce the difficulty of face driving. Moreover, because image acquisition technology is far superior to video acquisition technology and also superior to the acquisition technology of captured images in videos, it is easier to obtain high-definition source images and drive images, thereby improving the image quality of the target drive image and further improving the quality of face driving.

[0051] The technical solution of the embodiment of the present invention obtains a source image and a driving image, wherein the source image and the driving image include facial data of different objects; synthesizes at least one synthetic facial image based on the facial appearance features extracted from the source image and the facial posture and expression features extracted from the driving image; and synthesizes a target driving image based on the facial features of each synthetic facial image and the background features of the source image, thereby solving the problem of the existing difficulty and low quality of facial driving training, achieving the effect of improving the quality of facial driving, making the generated target driving image more vivid and reducing the difficulty of facial driving.

[0052] Example 2

[0053] Figure 2a This is a flow chart of a face driving method in the second embodiment of the present invention. This embodiment is based on the above embodiment and is refined. Specifically:

[0054] In an optional implementation of the embodiment of the present invention, synthesizing at least one synthesized facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the drive image includes:

[0055] Inputting the source image and the driving image into the face reconstruction network respectively, obtaining the face shape vector set and the face texture vector set in the source image, and obtaining the face expression vector set and the face angle vector set in the driving image;

[0056] The facial shape vector set, facial texture vector set, facial expression vector set and facial angle vector set are input into the facial rendering model to obtain a synthetic facial rendering image and a synthetic facial depth image synthesized by the facial rendering model.

[0057] In an optional implementation of the embodiment of the present invention, synthesizing a target driving image based on facial features of each synthesized facial image and background features of the source image includes:

[0058] Input the source image into a pre-trained feature encoder to obtain the background feature encoding of the source image;

[0059] Inputting each synthesized facial image and background feature code into a pre-trained neural network model to obtain a target driving image synthesized by the neural network model;

[0060] The feature encoder and the neural network model are jointly trained in an unsupervised manner using the same training sample set. The training samples include: source sample images and driving sample images. The objects to which the facial data in the source sample images and the driving sample images belong are the same or different.

[0061] Correspondingly, such as Figure 2a As shown, the method of this embodiment may include:

[0062] Step 210: Acquire a source image and a driving image.

[0063] The source image and the driving image include facial data of different objects.

[0064] Step 220: Input the source image and the driving image into the face reconstruction network respectively, obtain the facial shape vector set and the facial texture vector set in the source image, and obtain the facial expression vector set and the facial angle vector set in the driving image.

[0065] The facial reconstruction network can be used to obtain data on all key points of the face in the image. For example, the facial reconstruction network can be a deep learning network, and the process of obtaining facial data can be to convert the input facial image into the parameters of a three-dimensional facial deformation statistical model by training a convolutional neural network to perform three-dimensional face reconstruction. After reconstruction, the generated three-dimensional deformation statistical model parameters can be used as features to perform face recognition tests, and the parameters can be adjusted based on the test results to further improve the accuracy of the three-dimensional deformation statistical model parameters. The three-dimensional deformation statistical model is a classic statistical model of three-dimensional facial shape and albedo. For a given facial image, the three-dimensional deformation statistical model can estimate parameters such as projection, illumination, shape, and albedo through a network encoder, and then map these parameters to three-dimensional shape, texture, expression, and posture through a decoder.

[0066] In an embodiment of the present invention, a facial shape vector set and a facial texture vector set can be obtained from a source image through a facial reconstruction network. The facial shape vector set and the facial texture vector set can be used to represent facial appearance features in the source image, and specifically, can be data related to facial features. A facial expression vector set and a facial angle vector set can be obtained from a driving image through a facial reconstruction network. The facial expression vector set and the facial angle vector set can be used to represent facial pose features in the driving image, and specifically, can be data related to facial expression and head posture.

[0067] Step 230: Input the facial shape vector set, facial texture vector set, facial expression vector set, and facial angle vector set into a facial rendering model to obtain a synthesized facial rendering image and a synthesized facial depth image synthesized by the facial rendering model.

[0068] The facial rendering model can be implemented through an open source facial database, such as the Basel Face Model (BFM). Specifically, the functions glBegin and glEnd in the OpenGL module in the BFM model can be used to draw a face based on a facial shape vector set, a facial texture vector set, a facial expression vector set, and a facial angle vector set to synthesize a synthetic facial rendering image. The direction of the model can also be controlled based on theta and scale parameters in the OpenGL module; different viewing angles can also be set according to the size of the model. The synthetic facial rendering image is used to specifically reflect information such as facial shape, facial texture, facial expression, and head posture.

[0069] In an embodiment of the present invention, the synthesized facial depth image can be an image formed by further supplementing the depth information of the human face with the facial rendering model based on the synthesized facial rendering image. Specifically, the synthesized facial rendering image can be first generated into a grayscale image, and then the grayscale image is adjusted for facial depth based on the facial depth information contained in the facial shape vector set and the facial texture vector set, such as the depth information of the facial features and the height of the nose, to generate a synthesized facial depth image. This is beneficial for adjusting the lighting of the face when generating the target drive image and improving the generation quality of the target drive image. By synthesizing the facial rendering image and the synthesized facial depth image to generate the target drive image, the human face in the target drive image can be made more three-dimensional and more realistic.

[0070] Step 240: Input the source image into a pre-trained feature encoder to obtain background feature encoding of the source image.

[0071] In an embodiment of the present invention, to enhance the comprehensiveness and authenticity of the generated target driving image, background features from the source image can be incorporated into the target driving image generation process. Background features can include information beyond the synthesized facial rendering image and synthesized facial depth image, such as the character's hair and neck information, as well as information about the scene in which the character is located.

[0072] The feature encoder may be an encoder that extracts features from the source image. Specifically, the feature encoder may not focus on information about facial shape, facial texture, facial expression, and facial angle in the source image, but may only focus on background features in the source image. For example, the feature encoder may extract features of pixel positions related to background features in the source image, while ignoring features of pixel positions related to facial shape, facial texture, facial expression, and facial angle in the source image. Alternatively, the feature encoder may extract all features from the source image, and when using the feature information obtained by the feature encoder, the features of facial shape, facial texture, facial expression, and facial angle may be suppressed, with background features being primarily used.

[0073] Step 250: Input each synthesized facial image and background feature code into a pre-trained neural network model to obtain a target driving image synthesized by the neural network model.

[0074] The feature encoder and the neural network model are jointly trained in an unsupervised manner using the same training sample set. The training samples include: source sample images and driving sample images. The objects to which the facial data in the source sample images and the driving sample images belong are the same or different.

[0075] In an embodiment of the present invention, a neural network model can be used to synthesize a target drive image based on each synthesized facial image and background feature coding. Specifically, the neural network model can be used to synthesize the target drive image based on a synthesized facial rendering image, a synthesized facial depth image, and background feature coding. To avoid the influence of the background feature coding on the face of the target drive image, facial information can be first synthesized using the synthesized facial rendering image and the synthesized facial depth image, and then the background feature coding can be added to generate the target drive image.

[0076] Exemplarily, a process for synthesizing a target driving image based on a synthesized facial rendering image, a synthesized facial depth image, and background feature encoding can be as follows: a feature encoder obtains all features in a source image and inputs them into a neural network model; since the facial features contained in all features are in the source image and are not required for synthesizing the target driving image, the facial features in all features can be suppressed; specifically, facial information can be synthesized by first synthesizing a facial rendering image and a synthesized facial depth image, and then all features can be added to enhance the synthesized facial information, suppress the facial features in all features, enhance the background feature encoding in all features, and synthesize the target driving image. For example, facial features in all features can be suppressed based on the pixels corresponding to the synthesized facial information in the source image; and background feature encoding in all features can be enhanced based on the remaining pixels in the source image.

[0077] In embodiments of the present invention, an unsupervised approach is used to train feature encoders and neural network models, improving the accuracy of training results even when background feature encoding and driver results are unavailable. Specifically, the unsupervised approach can calculate the difference between the original situation and the final result for the target object and use this difference to adjust parameters during training; smaller differences indicate more accurate training results.

[0078] Exemplarily, the unsupervised training of the feature encoder may be to compare the source background feature encoding and the generated background feature encoding extracted from the source sample image and the target driving image generated based on the source sample image and the driving sample image, respectively, to determine the difference value. The unsupervised training of the neural network model may be to compare the facial shape vector set and facial texture vector set in the target driving image generated based on the source sample image and the driving sample image with the facial shape vector set and facial texture vector set in the source sample image, to determine the difference value; and to compare the facial expression vector set and facial angle vector set in the target driving image generated based on the source sample image and the driving sample image with the facial expression vector set and facial angle vector set in the driving sample image, to determine the difference value.

[0079] In embodiments of the present invention, the feature encoder and neural network model are trained using training samples consisting of source sample images and driving sample images, thereby improving the method for generating target driving images from different frames in a video. Specifically, when generating target driving images using a video as a training sample, the previous frame in the video is typically used as the pre-driving image, and the next frame in the video is used as the post-driving image. This means that the same person's expression is used to drive the same person, and there are no paired source sample images and driving sample images. Furthermore, when generating target driving images using a video as a training sample, because the video is dynamic, the backgrounds in the pre-driving and post-driving images differ. Background changes can cause interference when generating the target driving image, resulting in poor quality of the target driving image. However, to prevent background changes from affecting the generation of the target driving image, video data cleaning is required, which increases the difficulty of generating the target driving image and prevents the target driving image from incorporating background information. The technical solution in the embodiments of the present invention can avoid a series of problems brought about by training through videos. For example, the human objects in the source sample image and the driving sample image can be different, and the target driving image can contain background information. The generation of the target driving image through training with pictures will not be affected by the dynamic background, which can improve the quality of the target driving image. In addition, training in an unsupervised manner can ensure the accuracy of the target driving image synthesis when there are no labeled results.

[0080] Based on the above implementation, optionally, the neural network model is a Unet neural network model; the Unet neural network model includes: a connected neural network encoder and a neural network decoder, the output end of the feature encoder is connected to the input end of the neural network decoder; each synthesized facial image is input to the input end of the neural network encoder; the neural network encoder is used to generate facial features of each synthesized facial image and transmit them to the neural network decoder; the neural network decoder is used to encode the facial features of each synthesized facial image and the background features of the source image to synthesize the target driving image.

[0081] Among them, the Unet neural network model (a neural network model with a U-shaped structure) includes two parts: a neural network encoder and a neural network decoder. The neural network encoder can receive each synthesized facial image, that is, receive a synthesized facial rendering image and a synthesized facial depth image; the neural network encoder can be composed of four downsampling layers, which can perform feature extraction and generate facial features. The neural network decoder can receive the output of the neural network encoder and the background feature encoding of the source image; the neural network decoder can be composed of four upsampling layers, which can perform feature fusion, and can ensure that the final synthesized target drive image incorporates more low-level features, making the edge information of the image more refined. A jump link can be used between the neural network encoder and the neural network decoder, and the input position of the background feature encoding is at the layer where the downsampling layer ends and the upsampling begins. For example, the background feature encoding can be R 512×16×16 The feature map has 64 channels, with a width and height of 16, where R 512×16×16 Represents a 512×16×16 three-dimensional real vector; the output of the four downsampling layers is R 512×16×16 The features can be concatenated with the background feature code to generate R 1024×16×16 The feature map is input to the first upsampling layer.

[0082] In an optional implementation manner of an embodiment of the present invention, before synthesizing the target driving image based on the facial features of each synthesized facial image and the background features of the source image, the method further includes: sequentially obtaining current training samples from the training sample set, and obtaining the current source sample image and the current driving sample image in the current training sample; using a facial reconstruction network and a facial rendering model to obtain a current synthesized facial rendering image and a current synthesized facial depth image corresponding to the current source sample image and the current driving sample image; using the currently trained feature encoder and neural network model to synthesize the current target driving image based on the current synthesized facial rendering image and the current synthesized facial depth image; calculating a target loss function based on the result obtained by re-inputting the current target driving image into the facial reconstruction network and the feature encoder; after adjusting the parameters of the feature encoder and the neural network model using the target loss function, returning to executing the operation of sequentially obtaining the current training samples in the training sample set until the training end condition is met.

[0083] Embodiment 2 of the present invention provides a training process for a face-driven model. The current training sample is sequentially obtained from the training sample set, and the current source sample image and the current driven sample image in the current training sample are obtained. The current source sample image and the current driven sample image can be input into a face reconstruction network, and the face appearance features of the current source sample image and the face posture and expression features of the current driven sample image can be extracted through the face reconstruction network. Specifically, the face shape vector set and the face texture vector set of the current source sample image, and the face expression vector set and the face angle vector set of the current driven sample image can be extracted through the face reconstruction network.

[0084] The facial rendering model can synthesize the current synthesized facial rendering image and the current synthesized facial depth image based on facial appearance features and facial posture and expression features. Specifically, the facial rendering model can synthesize the current synthesized facial rendering image and the current synthesized facial depth image based on a facial shape vector set, a facial texture vector set, a facial expression vector set, and a facial angle vector set.

[0085] The currently trained feature encoder can be used to extract the original background feature code of the current source sample image; the currently trained neural network model can be used to synthesize the current target driving image based on the current synthesized face rendering image, the current synthesized face depth image and the original background feature code.

[0086] According to the result obtained by re-inputting the current target driving image into the face reconstruction network and the feature encoder, the target loss function is calculated. Specifically, it can be: re-inputting the current target driving image into the face reconstruction network to obtain the currently generated facial appearance features and facial posture and expression features corresponding to the current source sample image and the current driving sample image, respectively, and respectively comparing the facial appearance features and facial posture and expression features with the features corresponding to the current source sample image and the current driving sample image, and calculating the target loss function; re-inputting the current target driving image into the feature encoder to obtain the currently generated background features corresponding to the current source sample image, and comparing the background features with the features corresponding to the current source sample image, and calculating the target loss function.

[0087] Exemplarily, the training end condition may be that the result of the target loss function is less than a preset loss value, which can make the target driving image obtained under the current parameters more accurate.

[0088] On the basis of the above embodiment, optionally, according to the result obtained by re-inputting the current target driving image into the face reconstruction network and the feature encoder, a target loss function is calculated, including: re-inputting the current target driving image into the face reconstruction network to obtain a comparison face shape vector set, a comparison face texture vector set, a comparison face expression vector set and a comparison face angle vector set corresponding to the current target driving image; re-inputting the current target driving image into the feature encoder to obtain a comparison background feature code corresponding to the current target driving image; calculating a first loss function based on the comparison face shape vector set and the original face shape vector set obtained by inputting the current source sample image into the face reconstruction network; calculating a first loss function based on the comparison face texture vector set and the comparison face angle vector set; calculating a first loss function based on the comparison face shape vector set and the comparison face texture vector set. The second loss function is calculated based on the original facial texture vector set obtained by inputting the previous source sample image into the face reconstruction network; the third loss function is calculated based on the comparison of the facial expression vector set and the original facial expression vector set obtained by inputting the current driving sample image into the face reconstruction network; the fourth loss function is calculated based on the comparison of the facial angle vector set and the original facial angle vector set obtained by inputting the current driving sample image into the face reconstruction network; the fifth loss function is calculated based on the comparison of the background feature code and the original background feature code obtained by inputting the current source sample image into the feature encoder; the target loss function is calculated based on the first loss function, the second loss function, the third loss function, the fourth loss function and the fifth loss function.

[0089] The comparison facial shape vector set, comparison facial texture vector set, comparison facial expression vector set, and comparison facial angle vector set corresponding to the current target driving image can be compared with the original facial shape vector set, original facial texture vector set, and original facial expression vector set and original facial angle vector set corresponding to the current source sample image, respectively, to calculate a loss function to implement neural network model training. The comparison background feature code corresponding to the previous target driving image can be compared with the original background feature code corresponding to the current source sample image to calculate a loss function to implement feature encoder training.

[0090] For example, I A Represents the current source sample image, I B Represents the current driving sample image, G(I A , I B ) represents the current target driving image. The calculation of each loss function can be composed of one or more types of loss functions. For example, the target loss function can be calculated by using the first type of loss function and the second type of loss function.

[0091] Specifically, the first loss function can be obtained by Indicates; where L S is the first loss function, A vector set of primitive face shapes. To compare the face shape vector set, is the first type of loss function between the original face shape vector set and the comparison face shape vector set, It is the second type of loss function between the original face shape vector set and the comparison face shape vector set.

[0092] The second loss function can be obtained by Indicates that L T is the second loss function, For the original face texture vector set, To compare the facial texture vector set, is the first type of loss function between the original face texture vector set and the comparison face texture vector set, It is the second type of loss function between the original facial texture vector set and the comparison facial texture vector set.

[0093] The third loss function can be obtained by Indicates that L E is the third loss function, Original facial expression vector set, To compare facial expression vector sets, is the first type of loss function between the original facial expression vector set and the comparison facial expression vector set, is the second type of loss function between the original facial expression vector set and the compared facial expression vector set.

[0094] The fourth loss function can be obtained by Indicates that L P is the fourth loss function, is the original face angle vector set, To compare the facial angle vector set, is the first type of loss function between the original face angle vector set and the comparison face angle vector set, It is the second type of loss function between the original face angle vector set and the compared face angle vector set.

[0095] The fifth loss function can be obtained by Indicates that L feat is the fifth loss function, Encode the original background features, To compare the background feature code, is the first type of loss function between the original background feature encoding and the compared background feature encoding, It is the second type of loss function between the original background feature encoding and the compared background feature encoding.

[0096] The target loss function can be expressed as L = a × L feat +b×L S +c×L E +d×L P +e×L T In practice, the best model training results are achieved when a, b, c, d, and e are set according to the relationship a<e<b≤d<c.

[0097] Among them, the L1 loss function is also called the minimum absolute deviation or absolute value loss function, which minimizes the sum of the absolute differences between each target value and the corresponding estimated value. L2 is also called the norm loss function or the minimum square error, which minimizes the sum of the squares of the differences between each target value and the corresponding estimated value. L2 squares the error, and the error of the model will be much larger than L1, so it is more sensitive, and the model needs to be adjusted to minimize the error; however, it is very likely that there are outliers in the sample, and when the model is sensitive, it may cause the model training to deviate from the target. In order to circumvent the above problems, the embodiment of the present invention adopts a combination of L1 and L2 to calculate the loss function, so that the model has a certain sensitivity while avoiding the deviation of the model training from the target.

[0098] The technical solution of the embodiment of the present invention is as follows: obtaining a source image and a driving image; inputting the source image and the driving image into a face reconstruction network respectively, obtaining a face shape vector set and a face texture vector set in the source image, and obtaining a face expression vector set and a face angle vector set in the driving image; inputting the face shape vector set, the face texture vector set, the face expression vector set and the face angle vector set into a face rendering model together, obtaining a synthetic face rendering image and a synthetic face depth image synthesized by the face rendering model; inputting the source image into a pre-trained feature encoder, obtaining a background feature code of the source image; inputting each synthetic face image and the background feature code into a pre-trained neural network model together, obtaining a target driving image synthesized by the neural network model, thereby solving the problem of the existing face driving training being difficult and of low quality, and achieving the effect of improving the quality of the target driving image, ensuring the accuracy of the target driving image synthesis, and reducing the difficulty of model training.

[0099] Example 3

[0100] Figure 3 This is a flow chart of a live broadcast method in the third embodiment of the present invention. This embodiment is applicable to the situation where the facial expression of the host's face drives the host's selected character to show the same expression during the live broadcast. This method can be executed by a live broadcast device, which can be implemented by software and / or hardware and can generally be integrated into a computer device used for live broadcast (for example, various smart terminals or servers, etc.). Figure 3 As shown, the method of this embodiment may include:

[0101] Step 310: Receive live video data uploaded by the anchor client, and extract multiple live video frames from the live video data, where the live video frames include facial data of the live user.

[0102] The live video data can be real-time live video data from the host client or pre-recorded live video data. The time interval for extracting live video frames can be short, for example, 10 frames or 100 frames can be extracted at an average rate of 1 second. For live video frames with the face of the live user (host), face driving can be performed; for live video frames without the face of the live user (host), face driving is not required.

[0103] Step 320: Obtain a reference image selected by the anchor client, where the reference image includes facial data of the set object.

[0104] Among them, the reference image can be understood as the source image in the above embodiment, and the set object in the reference image can be the expected facial appearance of the anchor.

[0105] Step 330: Based on the facial appearance features extracted from the reference image and the facial posture and expression features extracted from each live video frame, synthesize at least one synthesized facial image corresponding to each live video frame.

[0106] The facial appearance of the person in the reference image can be combined with the facial expression of the person in the live video frame to obtain an image of the person in the reference image having the expression of the live video frame. The method for extracting facial appearance features, facial posture and expression features, and the method for obtaining the synthesized facial image can be the same as any of the corresponding methods in the above embodiments and will not be repeated here.

[0107] Step 340: synthesize the target driving images corresponding to the live video frames according to the facial features of the synthesized facial images corresponding to the live video frames and the background features of the reference image.

[0108] The method for extracting background features and the method for synthesizing the target driving image may be the same as any corresponding method in the above embodiments, and will not be described in detail here.

[0109] Step 350: After replacing each live video frame in the live video data with each target driving image, the live video data is published in the live broadcast room of the anchor client.

[0110] Among them, using the replaced live video data to publish in the live broadcast room can achieve the effect of hiding the real anchor's appearance and protect the anchor's privacy; the real anchor can also be replaced with the expected character to express the anchor's facial expressions, which can make the live broadcast character more popular with users watching the live broadcast and enhance the entertainment and popularity of the live broadcast.

[0111] The technical solution of the embodiment of the present invention receives live video data uploaded by an anchor client, and extracts multiple live video frames from the live video data, wherein the live video frames include facial data of the live user; obtains a reference image selected by the anchor client, wherein the reference image includes facial data of a set object; synthesizes at least one synthetic facial image corresponding to each live video frame based on facial appearance features extracted from the reference image and facial posture and expression features extracted from each live video frame; synthesizes target driven images corresponding to each live video frame based on facial features of each synthetic facial image corresponding to each live video frame and background features of the reference image; after replacing each live video frame in the live video data with each target driven image, the live video data is published in the live broadcast room of the anchor client, thereby solving the problem of live broadcasting through facial drive technology, protecting the privacy of the anchor, and improving the entertainment and popularity of the live broadcast.

[0112] Example 4

[0113] Figure 4 FIG. 1 is a structural diagram of a face driving device in the fourth embodiment of the present invention. Figure 4 As shown, the face driving device includes: an image acquisition module 410, a synthetic face image synthesis module 420 and a target driving image synthesis module 430.

[0114] in:

[0115] An image acquisition module 410 is configured to acquire a source image and a driving image, wherein the source image and the driving image include facial data of different objects;

[0116] A synthetic facial image synthesis module 420 is configured to synthesize at least one synthetic facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the drive image;

[0117] The target driving image synthesis module 430 is configured to synthesize a target driving image based on the facial features of each synthesized facial image and the background features of the source image.

[0118] Optionally, the facial image synthesis module 420 includes:

[0119] A vector set acquisition unit, configured to input the source image and the driving image into a face reconstruction network, respectively, to acquire a face shape vector set and a face texture vector set in the source image, and to acquire a face expression vector set and a face angle vector set in the driving image;

[0120] The image acquisition unit is used to input the facial shape vector set, facial texture vector set, facial expression vector set and facial angle vector set into the facial rendering model, and obtain the synthetic facial rendering image and synthetic facial depth image synthesized by the facial rendering model.

[0121] Optionally, the target driven image synthesis module 430 includes:

[0122] A background feature code acquisition unit is used to input the source image into a pre-trained feature encoder to obtain the background feature code of the source image;

[0123] A target driving image acquisition unit is used to input each synthesized facial image and background feature code into a pre-trained neural network model to obtain a target driving image synthesized by the neural network model;

[0124] The feature encoder and the neural network model are jointly trained in an unsupervised manner using the same training sample set. The training samples include: source sample images and driving sample images. The objects to which the facial data in the source sample images and the driving sample images belong are the same or different.

[0125] Optionally, the neural network model is a Unet neural network model;

[0126] The Unet neural network model includes: a connected neural network encoder and a neural network decoder, wherein the output of the feature encoder is connected to the input of the neural network decoder; each synthesized facial image is input to the input of the neural network encoder;

[0127] A neural network encoder, configured to generate facial features of each synthesized facial image and transmit the features to a neural network decoder;

[0128] The neural network decoder is used to encode the facial features of each synthesized facial image and the background features of the source image to synthesize the target driving image.

[0129] Optionally, the device further includes:

[0130] A sample image acquisition module is used to sequentially acquire current training samples from the training sample set and obtain the current source sample image and the current driving sample image in the current training sample before synthesizing the target driving image based on the facial features of each synthesized facial image and the background features of the source image;

[0131] A synthetic image acquisition module is used to obtain a current synthetic face rendering image and a current synthetic face depth image corresponding to the current source sample image and the current driving sample image using a face reconstruction network and a face rendering model;

[0132] A current target driving image synthesis module is used to synthesize a current target driving image based on a current synthesized face rendering image and a current synthesized face depth image using a currently trained feature encoder and a neural network model;

[0133] a target loss function determination module, configured to calculate a target loss function based on a result obtained by re-inputting the current target driving image into the face reconstruction network and the feature encoder;

[0134] The parameter adjustment module is used to adjust the parameters of the feature encoder and the neural network model using the target loss function, and then return to execute the operation of sequentially obtaining the current training samples in the training sample set until the training end condition is met.

[0135] Optional target loss function determination module, specifically used for:

[0136] Re-inputting the current target driving image into the face reconstruction network to obtain a comparison face shape vector set, a comparison face texture vector set, a comparison face expression vector set, and a comparison face angle vector set corresponding to the current target driving image;

[0137] Re-inputting the current target driving image into the feature encoder to obtain a comparison background feature code corresponding to the current target driving image;

[0138] Calculating a first loss function based on the compared facial shape vector set and the original facial shape vector set obtained by inputting the current source sample image into the facial reconstruction network;

[0139] Calculating a second loss function based on the compared facial texture vector set and the original facial texture vector set obtained by inputting the current source sample image into the facial reconstruction network;

[0140] Calculating a third loss function based on the compared facial expression vector set and the original facial expression vector set obtained by inputting the current driving sample image into the face reconstruction network;

[0141] Calculating a fourth loss function based on the compared facial angle vector set and the original facial angle vector set obtained by inputting the current driving sample image into the facial reconstruction network;

[0142] A fifth loss function is calculated based on the compared background feature code and the original background feature code obtained by inputting the current source sample image into the feature encoder;

[0143] The target loss function is calculated based on the first loss function, the second loss function, the third loss function, the fourth loss function and the fifth loss function.

[0144] The face driving device provided in the embodiment of the present invention can execute the face driving method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0145] Example 5

[0146] Figure 5 This is a structural diagram of a live broadcast device in the fifth embodiment of the present invention. Figure 5 As shown, the live broadcast device includes: a live video frame extraction module 510, a reference image acquisition module 520, a synthetic face image synthesis module 530, a target drive image synthesis module 540 and a live video data publishing module 550. Among them:

[0147] The live video frame extraction module 510 is used to receive live video data uploaded by the anchor client and extract multiple live video frames from the live video data, wherein the live video frames include facial data of the live user;

[0148] The reference image acquisition module 520 is used to acquire the reference image selected by the anchor client, wherein the reference image includes facial data of a set object;

[0149] A synthetic facial image synthesis module 530 is configured to synthesize at least one synthetic facial image corresponding to each live video frame based on facial appearance features extracted from the reference image and facial posture and expression features extracted from each live video frame;

[0150] The target driven image synthesis module 540 is used to synthesize the target driven images corresponding to the live video frames according to the facial features of the synthesized facial images corresponding to the live video frames and the background features of the reference image;

[0151] The live video data publishing module 550 is used to replace each live video frame in the live video data with each target driving image, and then publish the live video data in the live broadcast room of the anchor client.

[0152] The live broadcast device provided in the embodiment of the present invention can execute the live broadcast method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0153] Example 6

[0154] Figure 6 A schematic diagram of the structure of a computer device provided in Example 6 of the present invention is shown in FIG. Figure 6 As shown, the computer device includes a processor 60, a memory 61, an input device 62 and an output device 63; the number of processors 60 in the computer device can be one or more. Figure 6 In the figure, a processor 60 is used as an example; the processor 60, memory 61, input device 62 and output device 63 in the computer device can be connected by a bus or other means. Figure 6 The bus connection is taken as an example.

[0155] The memory 61 is a computer-readable storage medium that can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the face driving method or live broadcast method in the embodiment of the present invention (for example, Figure 4 The image acquisition module 410, the synthetic face image synthesis module 420 and the target drive image synthesis module 430 shown in FIG; or Figure 5 The processor 60 executes the software programs, instructions, and modules stored in the memory 61 to execute various functional applications and data processing of the computer device, that is, to implement the above-mentioned face driving method or live broadcast method, namely:

[0156] Acquire a source image and a driving image, wherein the source image and the driving image include facial data of different objects;

[0157] synthesizing at least one synthesized facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the driving image;

[0158] Based on the facial features of each synthesized facial image and the background features of the source image, the target driving image is synthesized. Or,

[0159] Receive live video data uploaded by the anchor client, and extract multiple live video frames from the live video data, wherein the live video frames include facial data of the live user;

[0160] Obtaining a reference image selected by the host client, the reference image including facial data of the set subject;

[0161] synthesizing at least one synthesized facial image corresponding to each live video frame based on the facial appearance features extracted from the reference image and the facial posture and expression features extracted from each live video frame;

[0162] Based on the facial features of each synthesized facial image corresponding to each live video frame and the background features of the reference image, a target driving image corresponding to each live video frame is synthesized;

[0163] After each target driving image is used to replace each live video frame in the live video data, the live video data is published in the live broadcast room of the anchor client.

[0164] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, and such remote memory may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0165] The input device 62 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the computer device. The output device 63 may include a display device such as a display screen.

[0166] Example 7

[0167] The seventh embodiment of the present invention further discloses a computer storage medium having a computer program stored thereon. When the program is executed by a processor, a face driving method or a live broadcast room method is implemented, namely:

[0168] Acquire a source image and a driving image, wherein the source image and the driving image include facial data of different objects;

[0169] synthesizing at least one synthesized facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the driving image;

[0170] Based on the facial features of each synthesized facial image and the background features of the source image, the target driving image is synthesized. Or,

[0171] Receive live video data uploaded by the anchor client, and extract multiple live video frames from the live video data, wherein the live video frames include facial data of the live user;

[0172] Obtaining a reference image selected by the host client, the reference image including facial data of the set subject;

[0173] synthesizing at least one synthesized facial image corresponding to each live video frame based on the facial appearance features extracted from the reference image and the facial posture and expression features extracted from each live video frame;

[0174] Based on the facial features of each synthesized facial image corresponding to each live video frame and the background features of the reference image, a target driving image corresponding to each live video frame is synthesized;

[0175] After each target driving image is used to replace each live video frame in the live video data, the live video data is published in the live broadcast room of the anchor client.

[0176] The computer storage medium of the embodiment of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.

[0177] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0178] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0179] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0180] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will appreciate that the present invention is not limited to the specific embodiments herein, and that various obvious changes, readjustments, and substitutions are possible for those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present invention. The scope of the present invention is determined by the scope of the appended claims.

Claims

1. A face driving method, characterized in that: include: Acquire a source image and a driving image, wherein the source image and the driving image include facial data of different objects, wherein the source image and the driving image are high-definition images of different people; At least one synthetic facial image is synthesized based on facial appearance features extracted from a source image and facial posture and expression features extracted from a driving image, including: inputting the source image and the driving image into a facial reconstruction network respectively, obtaining a facial shape vector set and a facial texture vector set in the source image, and obtaining a facial expression vector set and a facial angle vector set in the driving image; inputting the facial shape vector set, the facial texture vector set, the facial expression vector set, and the facial angle vector set into a facial rendering model, and obtaining a synthetic facial rendering image and a synthetic facial depth image synthesized by the facial rendering model, wherein the synthetic facial depth image is an image formed by the facial rendering model by further supplementing the depth information of the human face on the basis of the synthetic facial rendering image; A target driving image is synthesized based on the facial features of each synthesized facial image and the background features of the source image, including: inputting the source image into a pre-trained feature encoder to obtain the background feature code of the source image, wherein the background feature is information other than the synthesized facial rendering image and the synthesized facial depth image; inputting each synthesized facial image and the background feature code into a pre-trained neural network model to obtain the target driving image synthesized by the neural network model; wherein the feature encoder and the neural network model are jointly trained in an unsupervised manner using the same training sample set, and the training samples include: source sample images and driving sample images, and the objects to which the facial data in the source sample images and the driving sample images belong are the same or different.

2. The method according to claim 1, characterized in that The neural network model is a Unet neural network model; The Unet neural network model includes: a connected neural network encoder and a neural network decoder, the output end of the feature encoder is connected to the input end of the neural network decoder; each synthesized facial image is input to the input end of the neural network encoder; The neural network encoder is used to generate facial features of each synthesized facial image and transmit them to the neural network decoder; The neural network decoder is used to encode the facial features of each synthesized facial image and the background features of the source image to synthesize the target driving image.

3. The method according to claim 1, characterized in that Before synthesizing the target driving image based on the facial features of each synthesized facial image and the background features of the source image, the method further includes: Obtaining the current training samples in the training sample set in sequence, and obtaining the current source sample image and the current driving sample image in the current training sample; Using the face reconstruction network and the face rendering model, obtaining a current synthesized face rendering image and a current synthesized face depth image corresponding to the current source sample image and the current driving sample image; Using the currently trained feature encoder and the neural network model, synthesizing a current target driving image based on a current synthesized facial rendering image and a current synthesized facial depth image; Calculating a target loss function based on a result obtained by re-inputting the current target driving image into the face reconstruction network and the feature encoder; After adjusting the parameters of the feature encoder and the neural network model using the target loss function, the operation of sequentially acquiring the current training samples in the training sample set is returned to be executed until the training end condition is met.

4. The method according to claim 3, characterized in that Based on the result obtained by re-inputting the current target driving image into the face reconstruction network and the feature encoder, a target loss function is calculated, including: Re-inputting the current target driving image into the face reconstruction network, obtaining a comparison face shape vector set, a comparison face texture vector set, a comparison face expression vector set, and a comparison face angle vector set corresponding to the current target driving image; Re-inputting the current target driving image into the feature encoder to obtain a comparison background feature code corresponding to the current target driving image; Calculating a first loss function based on the compared facial shape vector set and an original facial shape vector set obtained by inputting the current source sample image into the facial reconstruction network; Calculating a second loss function based on the compared facial texture vector set and an original facial texture vector set obtained by inputting the current source sample image into the facial reconstruction network; Calculating a third loss function based on the compared facial expression vector set and the original facial expression vector set obtained by inputting the current driving sample image into the facial reconstruction network; Calculating a fourth loss function based on the compared facial angle vector set and an original facial angle vector set obtained by inputting the current driving sample image into the facial reconstruction network; Calculating a fifth loss function based on the compared background feature code and the original background feature code obtained by inputting the current source sample image into the feature encoder; A target loss function is calculated based on the first loss function, the second loss function, the third loss function, the fourth loss function and the fifth loss function.

5. A live broadcast method, characterized in that: include: Receive live video data uploaded by the anchor client, and extract multiple live video frames from the live video data, wherein the live video frames include facial data of the live user; Obtaining a reference image selected by the anchor client, wherein the reference image includes facial data of a set subject; wherein the live video frame and the reference image are high-definition images of different people; At least one synthetic facial image corresponding to each live video frame is synthesized based on facial appearance features extracted from a reference image and facial posture and expression features extracted from each live video frame, including: inputting the reference image and each live video frame into a facial reconstruction network, obtaining a facial shape vector set and a facial texture vector set in the reference image, and obtaining a facial expression vector set and a facial angle vector set in each live video frame; inputting the facial shape vector set, the facial texture vector set, the facial expression vector set, and the facial angle vector set into a facial rendering model, and obtaining a synthetic facial rendering image and a synthetic facial depth image synthesized by the facial rendering model, wherein the synthetic facial depth image is an image formed by the facial rendering model by further supplementing the depth information of the human face on the basis of the synthetic facial rendering image; Based on the facial features of each synthesized facial image corresponding to each live video frame and the background features of the reference image, a target driving image corresponding to each live video frame is synthesized, including: inputting the reference image into a pre-trained feature encoder to obtain a background feature code of the reference image, wherein the background feature is information other than the synthesized facial rendering image and the synthesized facial depth image; inputting each synthesized facial image and the background feature code into a pre-trained neural network model to obtain a target driving image synthesized by the neural network model; wherein the feature encoder and the neural network model are jointly trained in an unsupervised manner using the same training sample set, and the training samples include: a source sample image and a driving sample image, and the objects to which the facial data in the source sample image and the driving sample image belong are the same or different; After each target driving image is used to replace each live video frame in the live video data, the live video data is published in the live broadcast room of the anchor client.

6. A facial driving device, characterized in that: include: An image acquisition module, configured to acquire a source image and a drive image, wherein the source image and the drive image include facial data of different objects, wherein the source image and the drive image are high-definition images of different people; A synthetic facial image synthesis module, configured to synthesize at least one synthetic facial image based on facial appearance features extracted from the source image and facial posture and expression features extracted from the driving image; A target driving image synthesis module is used to synthesize a target driving image based on the facial features of each synthesized facial image and the background features of the source image; Wherein, the synthetic facial image synthesis module includes: A vector set acquisition unit, configured to input the source image and the driving image into a face reconstruction network, respectively, to acquire a face shape vector set and a face texture vector set in the source image, and to acquire a face expression vector set and a face angle vector set in the driving image; An image acquisition unit, configured to input the facial shape vector set, facial texture vector set, facial expression vector set, and facial angle vector set into a facial rendering model, and acquire a synthesized facial rendering image and a synthesized facial depth image synthesized by the facial rendering model, wherein the synthesized facial depth image is an image formed by the facial rendering model by further supplementing the depth information of the human face on the basis of the synthesized facial rendering image; The target-driven image synthesis module includes: a background feature code acquisition unit, configured to input a source image into a pre-trained feature encoder to acquire background feature codes of the source image, wherein the background features are information other than the synthesized facial rendering image and the synthesized facial depth image; A target driving image acquisition unit, configured to input the synthesized facial images and the background feature codes into a pre-trained neural network model to obtain a target driving image synthesized by the neural network model; The feature encoder and the neural network model are jointly trained in an unsupervised manner using the same training sample set. The training samples include: source sample images and driving sample images. The objects to which the facial data in the source sample images and the driving sample images belong are the same or different.

7. A live broadcast device, characterized in that: include: A live video frame extraction module is used to receive live video data uploaded by the anchor client and extract multiple live video frames from the live video data, wherein the live video frames include facial data of the live user; A reference image acquisition module is used to acquire a reference image selected by the anchor client, wherein the reference image includes facial data of a set object; wherein the live video frame and the reference image are high-definition images of different people; A synthetic facial image synthesis module is used to synthesize at least one synthetic facial image corresponding to each live video frame based on the facial appearance features extracted from the reference image and the facial posture and expression features extracted from each live video frame; A target driven image synthesis module is used to synthesize target driven images corresponding to each live video frame based on the facial features of each synthesized facial image corresponding to each live video frame and the background features of the reference image; A live video data publishing module, configured to replace each live video frame in the live video data with each target driving image, and then publish the live video data in the live broadcast room of the anchor client; Among them, the synthetic facial image synthesis module is specifically used to input the reference image and each live video frame into the facial reconstruction network respectively, obtain the facial shape vector set and facial texture vector set in the reference image, and obtain the facial expression vector set and facial angle vector set in each live video frame; input the facial shape vector set, facial texture vector set, facial expression vector set and facial angle vector set into the facial rendering model, and obtain the synthetic facial rendering image and synthetic facial depth image synthesized by the facial rendering model, wherein the synthetic facial depth image is an image formed by the facial rendering model by further supplementing the depth information of the human face on the basis of the synthetic facial rendering image; The target-driven image synthesis module includes: a background feature code acquisition unit, configured to input a reference image into a pre-trained feature encoder to acquire background feature codes of the reference image, wherein the background features are information other than the synthesized facial rendering image and the synthesized facial depth image; A target driving image acquisition unit, configured to input the synthesized facial images and the background feature codes into a pre-trained neural network model to obtain a target driving image synthesized by the neural network model; The feature encoder and the neural network model are jointly trained in an unsupervised manner using the same training sample set. The training samples include: source sample images and driving sample images. The objects to which the facial data in the source sample images and the driving sample images belong are the same or different.

8. A computer device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a face driving method according to any one of claims 1 to 4; or a live broadcast method according to claim 5.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute a face driving method as described in any one of claims 1 to 4; or a live broadcast method as described in claim 5.

Citation Information

Patent Citations

  • Image processing method and device, processor, electronic equipment and storage medium

    CN110399849A

  • Face driving and live streaming method and device, electronic equipment and storage medium

    CN111402399A