Face multi-frame dynamic fusion implementation method based on audio driving

By extracting facial expressions and head movement information combined with speech characteristics, using multi-frame dynamic fusion and diffusion models to generate realistic and natural face videos, the problem of insufficient naturalness of facial expressions and head movement in the prior art is solved, and the authenticity and nature of virtual speaking faces are improved.

CN120451343APending Publication Date: 2025-08-08ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358265.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the voice-driven face generation, the prior art mainly focuses on the accuracy of lip movement, ignores the naturalness of facial expressions and head movement, resulting in insufficient authenticity and naturalness of virtual speaking face generation.

Method used

By inputting reference audio, video and driver audio, facial expressions and head movement information are extracted, combined with speech semantics and acoustic features, a multi-frame dynamic fusion method and diffusion model are used to generate face motion sequences that conform to the speaking style, and a combination of lip motion optimization and neural rendering modules to generate realistic and natural face videos.

Benefits of technology

Realistic and natural facial expressions and head movements are generated, improving the authenticity and nature of virtual speaking faces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451343A_ABST
    Figure CN120451343A_ABST
Patent Text Reader

Abstract

The invention discloses a face multi-frame dynamic fusion implementation method based on audio driving, and relates to the technical field of face multi-frame dynamic fusion, the face multi-frame dynamic fusion implementation method comprises a face motion sequence generation method and a face video generation method, and the face motion sequence generation method comprises the following steps: inputting a reference audio, a reference video and a driving audio; extracting facial expression and head motion information of a person from the reference video; extracting semantic and acoustic features of voice from the reference audio and the driving audio; and combining the extracted facial expression and head motion information with the semantic and acoustic features of the voice by using a multi-frame dynamic fusion method. According to the face multi-frame dynamic fusion implementation method based on audio driving, the learned speaking style is combined with the input driving audio, and vivid and natural facial expressions and head motions can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of facial multi-frame dynamic fusion, and in particular to a method for realizing facial multi-frame dynamic fusion based on audio drive. Background Art

[0002] In recent years, speech-driven face generation technology, also known as talking face generation technology, has become a research focus in academia and has achieved significant progress. However, while current research has achieved significant breakthroughs in improving the synchronization between speech and lip movements, most work still focuses primarily on the accuracy of lip movements, relatively neglecting the generation of motions in other parts of the face, such as facial expressions and head movements. In the future virtual world, the realism and naturalness of virtual digital humans will become important performance indicators. Therefore, the naturalness of facial expressions and head movements has a profound impact on the practical application of virtual talking face generation technology. Intuitively, the correlation between speech and non-lip movements such as facial expressions and head movements is not straightforward, making it difficult to directly establish a universal correlation model. Speech primarily conveys acoustic information, while facial expressions and head movements reflect a person's emotions and attitudes.

[0003] To this end, we propose an audio-driven multi-frame dynamic fusion implementation method for faces. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for realizing dynamic fusion of multiple facial frames based on audio drive to solve the problems raised in the above background technology.

[0005] To achieve the above-mentioned purpose, the present invention provides the following technical solution: a method for realizing dynamic fusion of multiple facial frames based on audio drive, comprising a method for generating a facial motion sequence and a method for generating a facial video. The method for generating a facial motion sequence comprises the following steps: Input a reference audio, a reference video and a driving audio; extracting facial expression and head movement information of a person from the reference video; extracting semantic and acoustic features of speech from the reference audio and the driving audio; Using a multi-frame dynamic fusion method, the extracted facial expression and head movement information are combined with the semantic and acoustic features of the speech to learn the person's unique speaking style; Through the diffusion model technology, a facial motion sequence that conforms to the speaking style is generated.

[0006] Furthermore, the facial expressions and head movement information of the characters are extracted from the reference video, including using image recognition algorithms to identify the positions and changes of key points on the characters' faces in the video, as well as the rotation and translation movement parameters of the head.

[0007] Furthermore, the semantic and acoustic features of speech are extracted from the audio, including using natural language processing technology to analyze the semantic content of speech, and using signal processing technology to extract the frequency, amplitude, and duration acoustic features of the audio.

[0008] Furthermore, a multi-frame dynamic fusion method includes weighted fusion of facial expressions and head movement information and speech features of different frames to reflect dynamic changes during the speaking process.

[0009] Furthermore, the face video generation method includes the following steps: Input driving audio, face motion sequence and source image; The lip movement optimization module generates more accurate lip movement data based on audio features and expression coefficients, and combines this with head posture coefficients to form complete facial movement information. Through the neural rendering module, the complete facial motion information is mapped into a two-dimensional video, specifically through the mapping of 3DMM coefficients to key points and the image-driven network, to generate the final facial video with personalized speaking style.

[0010] Furthermore, the lip movement optimization module uses a deep learning model to analyze and process audio features and expression coefficients to generate accurate data that is more consistent with actual lip movements during speaking.

[0011] Furthermore, the mapping of 3DMM coefficients to key points in the neural rendering module converts the coefficients of the three-dimensional deformable model into the coordinates of facial key points on the two-dimensional image to achieve precise control of facial movement.

[0012] Furthermore, the image-driven network uses the feature information of the source image and combines it with facial motion information to generate the final face video, so that the generated video matches the appearance characteristics of the source image while maintaining the personalized speaking style.

[0013] Compared with the existing technology, the beneficial effects of the present invention are: by learning and modeling the audio and video frame sequences in the reference video, it aims to capture and learn the personalized speaking style (including expressions and head movements) of the characters in the reference video; by combining the learned speaking style with the input driving audio, it can generate realistic and natural facial expressions and head movements. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a flow chart of the method for generating a facial motion sequence according to the present invention; Figure 2 This is a flow chart for generating face videos with personalized speaking styles according to the present invention. DETAILED DESCRIPTION

[0015] like Figure 1-Figure 2 As shown, a facial multi-frame dynamic fusion implementation method based on audio drive includes a facial motion sequence generation method and a facial video generation method. Example

[0016] The method for generating a face motion sequence comprises the following steps: Select a clear audio clip with rich speech content as a reference. For example, you can obtain it from a professional voice database, or record a clear speech of a specific person. Ensure that the audio sampling rate is appropriate, generally 44.1kHz or higher is recommended to ensure audio quality. Select a video that clearly demonstrates the subject's facial expressions and head movements, consistent with the reference audio. The video resolution should be as high as possible, such as 1080p or higher, to more accurately extract facial landmarks and motion information. This can be captured with a high-definition camera or selected from a high-quality video library. Obtain the target audio for generating the corresponding facial motion sequence. The content and style of the audio should be consistent with the desired speaking style. To ensure audio quality, perform pre-processing operations such as noise reduction. Audio processing software such as Audacity can be used. Using deep learning models from the OpenCV library, such as convolutional neural network-based facial landmark detection models (e.g., the 68-point facial landmark detector from the dlib library), the model feeds a reference video frame by frame. The model outputs the coordinates of facial landmarks, such as the eyes, eyebrows, nose, and mouth, for each frame. By analyzing the positional changes of these landmarks between frames, the model can determine changes in facial expression, such as a raised corner of the mouth indicating a smile, or a furrowed brow indicating thought. Calculating Head Motion Parameters: To obtain the rotational and translational parameters of the head, 3D reconstruction techniques can be used. Reference videos from multiple viewpoints are first used, or markers with known locations are added during video capture. Based on this information, computer vision algorithms, such as the Structure from Motion (SfM) algorithm based on feature point matching, are used to calculate the head's rotation angle (around the x, y, and z axes) and translation distance in 3D space. Semantic content analysis: Using natural language processing toolkits such as AllenNLP, the reference and driving audio files are converted into text. Through lexical analysis, syntactic analysis, and semantic understanding, information such as keywords, grammatical structures, and semantic relationships is extracted. For example, named entity recognition identifies key entities such as names of people and places, while dependency parsing identifies grammatical dependencies between words. Acoustic feature extraction: Using the Librosa library, we extract acoustic features such as frequency, amplitude, and duration. For frequency features, we can calculate the fundamental frequency (F0) of the audio using the autocorrelation method or the YIN algorithm. Amplitude features can be obtained by calculating the root mean square (RMS) value of the audio signal. Duration features can be determined by detecting silence in the audio to determine the duration of each speech segment. Multi-frame dynamic fusion: Build a multi-frame dynamic fusion model based on a recurrent neural network (RNN), specifically using a long short-term memory network (LSTM). Model Input: The extracted facial expression and head movement information for each frame, along with the corresponding speech semantics and acoustic features, are organized into a suitable format and fed into the LSTM model. For example, the coordinates of facial key points, head movement parameters, the semantic vector representation of the audio, and the acoustic feature vector are concatenated to form an input vector sequence. Weight Adjustment and Training: During model training, a large number of reference audio-video pairs are used as training data. Using a backpropagation algorithm, the weights of different frames in the model are continuously adjusted, enabling the model to accurately capture the dynamic changes during speech and learn the person's unique speaking style. For example, for important semantic turning points in speech, corresponding facial expressions and head movements are given higher weights to highlight the role of these key frames. Model training: The DDPM model is trained using the speaking style features learned by the multi-frame dynamic fusion model as conditions. During training, noise is gradually added to clean facial motion sequence samples, while the model learns how to recover the original facial motion sequence that matches the specific speaking style from the noise. Sequence Generation: The DDPM model takes the processed features of the driving audio and the learned speaking style parameters as input and generates a facial motion sequence that matches the speaking style. During the generation process, model hyperparameters, such as noise intensity and diffusion steps, can be adjusted to optimize the quality of the generated sequence. For example, appropriately reducing the noise intensity can make the generated sequence smoother and more natural. Example

[0017] The face video generation method includes the following steps: Ensure that the audio format is consistent with the driving audio used when generating the facial motion sequence. Its format should be a common audio format, such as WAV, MP3, etc. If necessary, perform audio format conversion and quality check again to ensure that it can be correctly processed by subsequent modules. Face motion sequence: The sequence data obtained using the above-mentioned face motion sequence generation method should have a data structure that contains information such as the coordinate changes of facial key points and head motion parameters corresponding to each frame, to facilitate subsequent integration with other data. Source Image: Select a clear image of a person's face as the source image. The image should include all facial features and have a simple background. PNG or JPEG is recommended. The resolution should match the desired resolution of the final video, or be resized accordingly during subsequent processing.

[0018] Model Construction: A convolutional neural network (CNN) is used to construct the lip movement optimization module. The network structure can include multiple convolutional layers, pooling layers, and fully connected layers. For example, the convolutional layers first extract audio features and expression coefficients. Then, the pooling layers reduce the dimensionality of the feature maps, reducing computational effort. Finally, the fully connected layers integrate and map the extracted features to output more accurate lip movement data. Model training: We collect a large number of samples containing audio, facial expressions (focusing on lip movements), and real lip movement data as training data. During training, we use audio features and expression coefficients as input, and real lip movement data as labels. We use a backpropagation algorithm to continuously adjust the model parameters, enabling the model to accurately generate data that matches actual lip movement during speech based on the audio and expression information. Data Integration: The generated lip motion data is integrated with the head posture coefficients. Matrix operations can be used to fuse lip motion parameters such as displacement and rotation with head posture parameters to form complete facial motion information. For example, assuming the lip motion has a displacement dx in the x-direction and the head has a rotation angle θ in the x-direction, these can be combined using a specific conversion formula to obtain the final comprehensive motion parameters in the x-direction. Neural Rendering: Mapping 3DMM coefficients to keypoints: Based on the principle of 3D Deformable Model (3DMM), code is written to map 3DMM coefficients to facial keypoint coordinates on a 2D image. First, obtain the 3DM.

[0019] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A method for realizing dynamic fusion of multiple facial frames based on audio drive, characterized in that: The method includes a face motion sequence generation method and a face video generation method. The face motion sequence generation method includes the following steps: Input a reference audio, a reference video and a driving audio; extracting facial expression and head movement information of a person from the reference video; extracting semantic and acoustic features of speech from the reference audio and the driving audio; Using a multi-frame dynamic fusion method, the extracted facial expression and head movement information are combined with the semantic and acoustic features of the speech to learn the person's unique speaking style; Through the diffusion model technology, a facial motion sequence that conforms to the speaking style is generated.

2. The method for realizing multi-frame dynamic fusion of face based on audio drive according to claim 1, characterized in that: Extracting facial expressions and head movement information of characters from reference videos, including using image recognition algorithms to identify the position and changes of key points on the characters' faces in the video, as well as the rotation and translation movement parameters of the head.

3. The method for realizing multi-frame dynamic fusion of face based on audio drive according to claim 2, characterized in that: Extracting semantic and acoustic features of speech from audio, including using natural language processing technology to analyze the semantic content of speech, and using signal processing technology to extract the frequency, amplitude, and duration acoustic features of audio.

4. The method for realizing multi-frame dynamic fusion of face based on audio drive according to claim 3, characterized in that: The multi-frame dynamic fusion method includes weighted fusion of facial expressions and head movement information and speech features of different frames to reflect dynamic changes during speaking.

5. The method for realizing multi-frame dynamic fusion of face based on audio drive according to claim 4, characterized in that: The face video generation method includes the following steps: Input driving audio, face motion sequence and source image; The lip movement optimization module generates more accurate lip movement data based on audio features and expression coefficients, and combines this with head posture coefficients to form complete facial movement information. Through the neural rendering module, the complete facial motion information is mapped into a two-dimensional video, specifically including the mapping of 3DMM coefficients to key points and the image-driven network to generate the final facial video with personalized speaking style.

6. The method for realizing multi-frame dynamic fusion of face based on audio drive according to claim 5, characterized in that: The lip movement optimization module uses a deep learning model to analyze and process audio features and expression coefficients to generate accurate data that is more consistent with actual lip movements during speaking.

7. The method for realizing multi-frame dynamic fusion of faces based on audio drive according to claim 5, characterized in that: The mapping of 3DMM coefficients to key points in the neural rendering module converts the coefficients of the three-dimensional deformable model into the coordinates of facial key points on the two-dimensional image to achieve precise control of facial movement.

8. The method for realizing multi-frame dynamic fusion of faces based on audio drive according to claim 5, characterized in that: The image-driven network uses the feature information of the source image and combines it with facial motion information to generate the final face video, so that the generated video matches the appearance characteristics of the source image while maintaining the personalized speaking style.