A method and device for generating a realistic virtual human based on voice drive

By extracting the voice characteristics of the driving audio and facial information of the source video, a 3DMM model is constructed and the consistency of lip movement is strengthened. The virtual human video is generated using the conditional generation and adversarial network, which solves the problems of lack of dynamic expression details of virtual humans, inconsistent audio and lip movements, and poor video quality, and achieves high-quality virtual human video generation.

CN116206607BActive Publication Date: 2025-06-10BEIHANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310081778.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-08
Publication Date
2025-06-10
Estimated Expiration
2043-02-08

AI Technical Summary

Technical Problem

In the virtual human generation task based on audio-driven, there are problems such as lack of dynamic expression details of virtual humans, inconsistent audio and lip movements, and poor video quality.

Method used

A real virtual human generation method based on voice-driven is proposed. By obtaining source video and driver audio, head posture, facial shape information and texture information are extracted, driver voice features are extracted using the speech recognition model, facial expression parameters and blinking action information are synthesized, 3DMM model renderings are constructed, and lip movement consistency is strengthened through the Wav2Lip module, and the virtual human video is finally generated using the conditional generation adversarial network.

Benefits of technology

It realizes the generation of virtual human videos with natural and delicate dynamic expressions, solves the problem of inconsistent audio and lip movements, and improves the quality of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206607B_ABST
    Figure CN116206607B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for generating a realistic virtual human based on voice driving: inputting a source video and driving audio; using the person in the source video as the prototype of the virtual human, and extracting head pose, facial shape information, and texture information from the source video; using the driving audio as the content for the virtual human to speak, inputting the driving audio, and synthesizing facial expression parameters and blink action information synchronized with the driving audio; using the facial expression parameters, blink action information, head pose, facial shape information, and texture information to construct a rendered image of the virtual human 3DMM model; introducing the Wav2Lip module to enhance the speech-lip consistency of the lip information in the rendered image of the 3DMM model to obtain a virtual human lip enhancement result image; inputting the Mel spectrogram features of the driving audio, the virtual human lip enhancement result image, and a reference background, and using a conditional generative adversarial network to generate a virtual human video. The present invention helps to improve the quality of virtual human video generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video generation, and in particular, to a method and device for generating a realistic virtual human based on voice driving. Background Art

[0002] The task of generating a virtual human based on voice driving refers to generating a video sequence that maintains the identity information of the source image, lip movement, facial expression, and is consistent with the voice content given a piece of voice and an input image. Driven by the technological waves of artificial intelligence, virtual reality, etc., virtual human generation has received increasing attention and is widely used in fields such as human-computer interaction, film and television production, virtual anchors, and intelligent employees.

[0003] Although virtual human generation technology has been improved to a certain extent, the realistic virtual human based on audio driving is still in the research foundation stage. A realistic virtual human needs to have the following characteristics: 1. Having natural and delicate dynamic expressions, and realistic expression display is one of the key factors for generating a realistic virtual human video; 2. The voice is consistent with the lip movement; 3. The virtual human image quality requirements are relatively high. The human visual system is very sensitive to the temporal discontinuity in the video, which requires the generated video output to be stable and the video frames to be continuous. Limited by the complexity of the human face structure and the diversity of expressions and lip movements, the generation of a realistic virtual human has become one of the key points and difficulties in the field of computer vision perception research.

[0004] At the present stage, the virtual human generation algorithms based on voice driving mainly conduct research from several aspects: First, directly find the synchronization relationship between audio and lip movement; Second, with the assistance of intermediate information such as two-dimensional key points or three-dimensional information, learn the corresponding relationship from audio to intermediate information and from intermediate information to the virtual human. Compared with the first type of algorithm, this type of method can more conveniently handle problems such as expression synchronization. Summary of the Invention

[0005] In order to solve the problems of lack of virtual human dynamic expression details, inconsistency between audio and lip movement, and poor video quality in the virtual human generation task based on audio driving, the present invention proposes a method and device for generating a realistic virtual human based on voice driving.

[0006] According to the first aspect of the present invention, there is provided a method for generating a realistic virtual human based on voice driving, including:

[0007] Obtaining a source video and a driving audio, where the source video is a video of a person speaking, and the driving audio is a voice audio unrelated to the source video;

[0008] Extracting head pose, facial shape information, and texture information from the source video;

[0009] Using the driving audio as the speech content for the virtual human to speak, based on the existing speech recognition model, extracting generalizable driving speech features from the driving audio, where the driving speech features are the outputs of the hidden layer of the model when the driving audio is used as the input;

[0010] Using the driving speech features, synthesizing facial expression parameters and blink action information synchronized with the driving audio, and constructing a virtual human 3DMM model rendering diagram using the facial expression parameters, blink action information, head pose, facial shape information, and texture information;

[0011] Introducing the Wav2Lip module to enhance the speech-lip consistency of the lip information in the virtual human 3DMM model rendering diagram to obtain a virtual human lip enhancement result diagram;

[0012] Inputting the Mel spectrogram features of the driving audio, the virtual human lip enhancement result diagram, and the reference background, and generating a virtual human video using a conditional generative adversarial network.

[0013] According to the second aspect of the present invention, there is provided a generation device for a realistic virtual human based on speech driving, including:

[0014] An acquisition module for acquiring a source video and a driving audio;

[0015] A human face model parameter extraction module for extracting head pose, facial shape information, and texture information from the source video, extracting driving speech features from the driving audio, synthesizing facial expression parameters and blink action information synchronized with the audio according to the driving speech feature time series, and constructing a virtual human 3DMM model rendering diagram using the facial expression parameters, blink action information, head pose, facial shape information, and texture information;

[0016] A lip shape enhancement module for enhancing the speech-lip consistency of the lip information in the virtual human 3DMM model rendering diagram to obtain a virtual human lip enhancement result diagram;

[0017] A video generation module for inputting the Mel spectrogram of the driving audio, the virtual human lip enhancement result diagram, and the reference background, and generating a virtual human video using a conditional generative adversarial network renderer.

[0018] In the third aspect of the present invention, a non-transitory computer-readable storage medium is proposed, where the non-transitory computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the method shown in the first aspect above is implemented.

[0019] Compared with the prior art, the present invention has the following advantages:

[0020] 1. Construct rich 3D face information from multiple dimensions. Extract facial expression information from audio and actively set blinking actions, and obtain head pose, identity features, and texture information from the input video. The 3D model constructed by this method can explicitly control blinking actions and facial expressions, enriching the 3D face model, which helps to generate virtual humans with natural expressions and rich expression details.

[0021] 2. Coarse-to-fine two-stage lip synchronization synthesis. In the coarse-grained stage, lip actions with lower accuracy are represented by facial expression parameters of the 3D face model; in the fine-grained stage, the consistency of lip movement information is enhanced for the 3D model rendering based on the driving audio, avoiding interference from the original lip actions of the person in the source video when directly modifying image frames, which helps to strengthen the synthesis of lip movement and can thus well solve the problem of inconsistent audio-lip movement.

[0022] 3. Solve the problem of video generation quality from two perspectives. First, modify the input sequence length of cGAN, and perform synthesis by referring to the previous and next n frames to improve the stability between video frames; second, add audio information to cGAN to strengthen the constraint of the driving speech features on the lip graphics, making the lip actions more natural, which in turn helps to improve the video generation quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic flowchart of a method for generating a realistic virtual human based on speech driving provided by an embodiment of the present invention;

[0024] Figure 2 is a schematic flowchart of the image rendering process of a cGAN renderer provided by an embodiment of the present invention;

[0025] Figure 3 is a schematic diagram of the representation of blinking action information provided by an embodiment of the present invention;

[0026] Figure 4 is a schematic flowchart of the process for obtaining a reference background provided by an embodiment of the present invention;

[0027] Figure 5 is a schematic structural diagram of a device for generating a realistic virtual human based on speech driving provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as a limitation of the present application.

[0029] The virtual human generation method and device according to the embodiments of the present application will be described below with reference to the accompanying drawings.

[0030] Embodiment 1

[0031] Figure 1 The flowchart of the realistic virtual human generation method based on voice driving provided by an embodiment of the present invention is shown in the figure, and includes:

[0032] Obtain a source video and a driving audio. The source video is a video of a person speaking, and the driving audio is a voice audio unrelated to the source video;

[0033] Extract the head pose, facial shape information, and texture information from the source video;

[0034] Using the driving audio as the voice content of the virtual human to speak, based on the existing speech recognition model DeepSpeech, with a frame length of 80 milliseconds, extract generalizable driving speech features from the driving audio, where the driving speech features are the outputs of the hidden layer of the DeepSpeech model when the driving audio is used as the input;

[0035] Use the driving speech features to synthesize facial expression parameters and blink action information synchronized with the driving audio, and use the facial expression parameters, blink action information, head pose, facial shape information, and texture information to construct a virtual human 3DMM model rendering;

[0036] Input the virtual human 3DMM model rendering into the Wav2Lip audio-lip synchronization model to strengthen the speech-lip consistency of the lip information and obtain a virtual human lip enhancement result map;

[0037] Input the Mel spectrogram features of the driving audio, the virtual human lip enhancement result map, and the reference background, and use the conditional generative adversarial network to generate a virtual human video.

[0038] In this embodiment, the construction of the virtual human 3DMM model rendering includes:

[0039] Use OpenFace to detect the source video to obtain the head pose and the AU value of the blink action of the person in the source video, where the head pose includes the Euler angle R and the translation amount t;

[0040] Use Deep3DFaceReconstruction to extract the shape parameters, texture parameters, and expression parameters of the BFM model of the face of the person in the source video from the RGB image of the source video, where the shape parameters and texture parameters are the facial shape information and texture information;

[0041] Train the speech-facial action extraction network in a self-supervised manner.

[0042] Use the trained speech-facial action extraction network to predict the facial expression parameters and AU45 value of the 3DMM model synchronized with the driving speech according to the driving speech features, that is, the AU value of the blinking action, and obtain the facial expression parameters and blinking action information synchronized with the driving audio.

[0043] Use the Euler angles R and translation amount t of the extracted head pose to transform the vertices of the 3DMM model, move the 3DMM model to the same position as the face of the person in the RGB frame, input the above facial expression parameters synchronized with the driving audio, as well as the shape parameters and texture parameters extracted from the source video into Face3D for rendering, and modify the RGB value of the eyes in the rendering result according to the AU value of the blinking action synchronized with the driving audio to represent the blinking action. As Figure 3 shown, the eye area will be black when closing the eyes and white when opening the eyes, and obtain the rendering diagram of the virtual human 3DMM model.

[0044] In this embodiment, the training of the speech-facial action extraction network in a self-supervised manner includes:

[0045] Extract the source video speech features from the audio content of the source video in the same way as extracting the driving speech features.

[0046] Train the speech-facial action extraction network. Taking 8 frames as a time period, input the time series of the source video speech features in the current time period, the start time of the current time period, and the expression parameters and AU values of the blinking action at the end of the previous time period. Using the AU value of the blinking action of the person in the source video in the current time period and the expression parameters in the 3DMM model parameters extracted from the source video as supervision, perform regression on the 3DMM expression parameters of the face and the AU values of the blinking action synchronized with the driving audio.

[0047] In this embodiment, the obtaining of the enhanced result diagram of the virtual human lips includes:

[0048] Use the audio processing library librosa to process the driving audio to obtain the Mel spectrogram features of the driving audio with a frame length of 40 ms and 80 channels.

[0049] Input the Mel spectrogram features of the driving audio and the rendering diagram of the virtual human 3DMM model into the existing trained Wav2Lip model to enhance the lip movement of the rendering diagram and obtain the enhanced result diagram of the virtual human lips.

[0050] In this embodiment, generating a virtual human video by using the Mel spectrum features of the driving audio, the enhanced result map of the virtual human's lips, and the reference background by using a conditional generative adversarial network includes:

[0051] Input the expression parameters, facial shape information, texture information, head pose of the 3DMM model extracted from the source video, and the AU value of the person's blinking action into Face3D for rendering, and modify the RGB values of the eyes in the rendering result according to the same rule as obtaining the rendering map of the virtual human 3DMM model to obtain the rendering map of the source video 3DMM model;

[0052] As Figure 4 shown, extract all RGB image frames of the source video, cover the facial areas covered by the rendering map of the source video 3DMM model corresponding to each frame in all RGB image frames, and perform a certain expansion and covering around the facial areas to obtain the reference background;

[0053] Train a conditional generative adversarial network cGAN renderer in a self-supervised manner;

[0054] Refer to Figure 2 , use the trained cGAN renderer, input an image sequence composed of 1 frame corresponding to the moment to be synthesized and the enhanced result maps of the lips corresponding to 5 frames before and after it, the reference background at this moment, and the Mel spectrum features of the driving audio at this moment to obtain the synthesized RGB frame at this moment;

[0055] Encode the synthesized RGB frames at all moments into a video to obtain the virtual human video.

[0056] In this embodiment, training a conditional generative adversarial network cGAN renderer in a self-supervised manner includes:

[0057] Extract Mel spectrum feature training data from the source video. In the same way as extracting the Mel spectrum features of the driving audio, use the audio content of the source video as the input to obtain the Mel spectrum features of the source video;

[0058] Extract lip enhancement result map training data from the source video. Input the Mel spectrum features of the source video and the rendering map of the source video 3DMM model into the used Wav2Lip model to obtain the lip enhancement result map of the source video;

[0059] Train a conditional generative adversarial network cGAN renderer. Input an image sequence composed of 1 frame corresponding to a certain moment in the video and the enhanced result maps of the lips of the source video corresponding to 5 frames before and after it, and the reference background at this moment, and use the Mel spectrum features of the source video corresponding to this moment as the conditional input, with the goal of reconstructing the original RGB image of the video frame corresponding to the current moment.

[0060] Furthermore, based on the method for generating a realistic virtual human driven by voice provided in the above embodiments, an embodiment of the present invention further provides a device for generating a realistic virtual human driven by voice. Figure 5 As shown in the structural schematic diagram of a device for generating a realistic virtual human driven by voice according to an embodiment of the present application, Figure 5 it includes:

[0061] An acquisition module, configured to acquire a source video and a driving audio;

[0062] A face model parameter extraction module, configured to extract head pose, facial shape information, and texture information from the source video, extract driving speech features from the driving audio, synthesize facial expression parameters and blink action information synchronized with the audio according to the driving speech feature time series, and use the facial expression parameters, blink action information, head pose, facial shape information, and texture information to construct a rendered image of a virtual human 3DMM model;

[0063] A lip enhancement module, configured to enhance the speech lip consistency of the lip information in the rendered image of the virtual human 3DMM model to obtain a rendered image of the enhanced lips of the virtual human;

[0064] A video generation module, configured to input the Mel spectrogram of the driving audio, the rendered image of the enhanced lips of the virtual human, and a reference background, and use a conditional generative adversarial network renderer to generate a virtual human video.

[0065] To implement the above embodiments, the present invention also proposes a non-transitory computer-readable storage medium.

[0066] The non-transitory computer-readable storage medium provided by the embodiment of the present invention stores a computer program; when the computer program is executed by a processor, it can implement the method for generating a realistic virtual human driven by voice as Figure 1 shown in any one.

[0067] In the description of this specification, the descriptions with reference to terms such as "an embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0068] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many variations without departing from the gist of the present invention, and all of these fall within the scope of protection of the present invention.

Claims

1. A method for generating a realistic virtual human based on voice driving, characterized in that, the steps are as follows: Obtain a source video and a driving audio. The source video is a video of a person speaking, and the driving audio is a voice audio unrelated to the source video; Extract the head pose, facial shape information, and texture information from the source video; Use the driving audio as the voice content of the virtual human to speak. Based on the existing speech recognition model, extract the generalized driving speech features from the driving audio, where the driving speech features are the output of the hidden layer of the model when the driving audio is used as the input; Use the driving speech features to synthesize facial expression parameters and blink action information synchronized with the driving audio, and use the facial expression parameters, blink action information, head pose, facial shape information, and texture information to construct a virtual human 3DMM model rendering; Introduce the Wav2Lip module to enhance the speech lip consistency of the lip information of the virtual human 3DMM model rendering to obtain a virtual human lip enhancement result map; Input the expression parameters, facial shape information, texture information, head pose of the 3DMM model extracted from the source video, and the AU value of the blink action of the person into Face3D for rendering, and modify the RGB value of the eyes in the rendering result according to the AU value of the blink action extracted from the source video to obtain a source video 3DMM model rendering; Extract all RGB image frames of the source video, cover the face area covered by the source video 3DMM model rendering corresponding to each frame in all RGB image frames, and perform a certain expansion and covering around the face area to obtain a reference background; Extract Mel spectrogram feature training data from the source video. In the same way as extracting the Mel spectrogram features of the driving audio, use the audio content of the source video as the input to obtain the Mel spectrogram features of the source video; Extract lip enhancement result map training data from the source video, input the Mel spectrogram features of the source video and the source video 3DMM model rendering into the Wav2Lip module to obtain a source video lip enhancement result map; Train a conditional generative adversarial network cGAN renderer. Input an image sequence composed of 1 frame corresponding to a certain moment in the source video and the source video lip enhancement result maps corresponding to the previous and next n frames, and the reference background at this moment. Use the Mel spectrogram features of the source video corresponding to this moment as the conditional input, and the goal is to reconstruct the original RGB image of the video frame corresponding to the current moment; Use the virtual human lip enhancement result map, the reference background, and the Mel spectrogram features of the driving audio as the input, and use the trained cGAN renderer to generate a virtual human video.

2. The virtual human generation method according to claim 1, characterized in that, constructing a virtual human 3DMM model rendering includes: Use OpenFace to detect the source video to obtain the head pose and the AU value of the blink action of the person in the source video; Extract the shape parameters, texture parameters, and expression parameters of the 3DMM model of the person's face in the source video from the RGB images of the source video by using Deep3DFaceReconstruction, where the shape parameters and texture parameters are the facial shape information and texture information; Train the speech-facial action extraction network in a self-supervised manner; Use the trained speech-facial action extraction network to predict the facial expression parameters of the 3DMM model and the AU values of the blinking actions synchronized with the driving audio according to the driving speech features, and obtain the facial expression parameters and blinking action information synchronized with the driving audio; Input the facial expression parameters synchronized with the driving audio, as well as the head pose, facial shape information, and texture information extracted from the source video, into Face3D for rendering, and modify the RGB values of the eyes in the rendering result according to the AU values of the blinking actions synchronized with the driving audio to represent the blinking actions, so as to obtain the rendering diagram of the virtual human 3DMM model.

3. The virtual human generation method according to claim 1, characterized in that, obtain the enhanced result diagram of the virtual human lips, including: Process the driving audio by using the audio processing library librosa, convert the original audio signal from the time domain to the frequency domain, and obtain the audio features with stronger representation ability after compression, that is, the Mel spectrogram features of the driving audio; Input the Mel spectrogram features of the driving audio and the rendering diagram of the virtual human 3DMM model into the Wav2Lip module to enhance the lip actions of the rendering diagram, and obtain the enhanced result diagram of the virtual human lips.

4. The virtual human generation method according to claim 2, characterized in that, the training of the speech-facial action extraction network in a self-supervised manner includes: Extract the source video speech features from the audio content of the source video in the same way as extracting the driving speech features; Train the speech-facial action extraction network, input the time series of the source video speech features, and regress the 3DMM expression parameters of the face and the AU values of the blinking actions synchronized with the driving audio with the AU values of the blinking actions of the person in the source video and the expression parameters in the 3DMM model parameters extracted from the source video as the supervision.

5. The virtual human generation method according to claim 1, characterized in that, generate a virtual human video by using the trained cGAN renderer, including: Use the trained cGAN renderer, input the image sequence composed of 1 frame corresponding to the moment to be synthesized and the enhanced result diagrams of the lips corresponding to the previous and next n frames, the reference background at this moment, and the Mel spectrogram features of the driving audio at this moment, and obtain the synthesized RGB frame at this moment; Encode the synthesized RGB frames at all moments into a video to obtain the virtual human video.

6. A realistic virtual human generation device based on speech driving, characterized in that, implement the realistic virtual human generation method based on speech driving as described in any one of claims 1-5, including: An acquisition module for acquiring a source video and a driving audio; A face model parameter extraction module, which is used to extract head pose, facial shape information and texture information from the source video, extract driving speech features from the driving audio, synthesize facial expression parameters and blink action information synchronized with the audio according to the driving speech feature time series, and use the facial expression parameters, blink action information, head pose, facial shape information and texture information to construct a rendering of the virtual human 3DMM model; A lip enhancement module, which is used to enhance the speech lip consistency of the lip information in the rendering of the virtual human 3DMM model to obtain a rendering of the enhanced virtual human lips; A video generation module, which is used to input the Mel spectrogram of the driving audio, the rendering of the enhanced virtual human lips and the reference background, and use a conditional generative adversarial network renderer to generate a virtual human video.

7. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, it implements the speech-driven realistic virtual human generation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Virtual character speaking video synthesis method and device, equipment and storage medium

    CN115116109A

  • Image generation method and apparatus, device, and storage medium

    WO2022242381A1