Emotional Animation Generation Method Based on U-net Attention-enhanced Decoder
By integrating facial key point prediction and the use of CBAM modules in the U-net attention enhancement decoder, the problems of low video quality and inaccurate facial expressions in the prior art are solved, and high-fidelity head video generation is achieved.
Patent Information
- Application Number
- CN202410448894.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-04-15
AI Technical Summary
In the prior art, when generating high-quality videos with head movements and facial expressions, the details of the eye and skin texture are not processed enough, resulting in low video quality and insufficient facial expressions.
The emotional animation generation method based on the U-net attention enhancement decoder is adopted to generate high-fidelity speaking head videos by integrating facial emotions through facial key point prediction, and the CBAM module enhances the decoder's attention mechanism.
Improve the quality of the generated video, so that the output image can maintain more details, such as the complex skin texture and facial shadows of the target person, enhancing the naturalness of facial emotional expression.
Smart Images

Figure CN118505860B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of virtual digital animation, and particularly relates to an emotional animation generation method based on a U-net attention-enhanced decoder. Background Art
[0002] Face animation generation aims to generate a series of highly natural and lip-synchronized character animations from given speech or text, images or videos through certain transformation methods. In recent years, with the rapid development of face animation technology based on deep learning, virtual digital human technology has been applied in many industries, such as news media, commercial customer service, etc. In news media, using virtual real news anchors to visualize information news can quickly and accurately convey the latest domestic and international news to users in a timely and easy-to-understand manner, achieving the rapid production of a large amount of news. In virtual customer service, using virtual characters that conform to the corporate image to answer user questions and explain related industries can reduce labor consumption and quickly solve user doubts, improving the economic benefits of the enterprise. However, face animation driven by audio and video is limited by bandwidth and cost, and is not suitable for specific application fields such as video conferencing with limited bandwidth and high-cost video production, having limitations. Currently, face animation technology driven by audio and images has become a research hotspot. This driving method is relatively easy to obtain. With the rapid development of artificial intelligence technology, the capture quality of audio and images by photographic devices such as mobile phones and cameras is very high, which can save resources and can be used in many application fields, having generalization. Currently, it is possible to achieve lip movement synchronization with audio in the generation of talking face videos driven by audio and images, but the coordination between head movement and facial expressions and audio, as well as the video quality, still need to be improved. However, there are few methods for generating both facial expressions and head movements. Regarding the quality of animated videos, there have been certain research reports, but the generated video results still need to be improved in terms of eye and skin texture details. Facial rendering of eye and skin texture details plays an important role in emotional expression. In order to better reflect the facial expressions of the target person, it is also necessary to improve the video quality by rendering the skin texture of the target person with high fidelity to generate more realistic video frames. Summary of the Invention
[0003] The purpose of the present invention is to provide an emotional animation generation method based on a U-net attention-enhanced decoder, and finally output high-fidelity video frames of the target person with both lip, head, and facial emotional animations, which can improve the quality of the generated video and enable the output image to retain more details.
[0004] The technical solution adopted by the present invention is an emotional animation generation method based on a U-net attention-enhanced decoder, specifically as follows:
[0005] Step 1: Prediction of facial key points integrating facial emotions;
[0006] Step 2: Decoding of facial key points based on the U-net attention-enhanced decoder. In the stage of decoding facial key points to generate a talking head video, the predicted facial key point video frames obtained from Step 1 and the target person's image are input into the U-net-based attention-enhanced decoder to generate a realistic talking head video.
[0007] The features of the present invention also lie in:
[0008] Specifically, Step 1 is as follows:
[0009] Step 1.1: Construct and train a facial emotion animation network;
[0010] Step 1.2: Select an audio segment and obtain the predicted facial key point video frames corresponding to the audio.
[0011] Specifically, Step 1.1 is as follows:
[0012] Step 1.1.1: Download the latest publicly available multi-modal emotion and sentiment detection dataset, abbreviated as the MEAD dataset, which contains high-quality head videos and audio with different emotions for the same text.
[0013] Step 1.1.2: Convert the talking head videos in the MEAD dataset of Step 1.1.1 to 62.5fps and set the audio sampling rate to 16KHz;
[0014] Step 1.1.3: Extract the 3D facial key point information of the head videos in the MEAD dataset processed in Step 1.1.2 through the face_alignment library, then perform normalization processing with standard facial key points, and finally store it in a document;
[0015] Step 1.1.4: On the basis of the MEAD dataset processed in Step 1.1.2, train an emotion encoder for extracting emotion-assisted features through cross-reconstruction emotion disentanglement technology;
[0016] Step 1.1.5: Use the emotion encoder in Step 1.1.4 to extract the emotion feature information in the audio of the MEAD dataset processed in Step 1.1.2 and store it in a document;
[0017] Step 1.1.6: Construct a facial emotion animation network through a recurrent network LSTM and a multi-layer perceptron;
[0018] Step 1.1.7: Input the emotional feature information in the audio of the MEAD dataset obtained in Step 1.1.5 into the facial emotion animation network constructed in Step 1.1.6. Through mapping, obtain the corresponding facial landmarks, and then compare the loss with the facial key points of the MEAD dataset extracted in Step 1.1.3 to train the facial emotion animation network constructed in Step 1.1.6.
[0019] Specifically, Step 1.2 is as follows:
[0020] Step 1.2.1: Select any segment of audio from the MEAD dataset in Step 1.1.1.
[0021] Step 1.2.2: Preprocess the audio selected in Step 1.2.1, modify the frame rate of the audio to 62.5 Hz, and the sampling rate of the speech waveform to 16 KHz.
[0022] Step 1.2.3: Input the audio processed in Step 1.2.2 into the pre-trained text, speaker style, and emotion encoders respectively, and obtain the content, speaker style, and emotion features in the audio, namely vectors of 80, 256, and 128 dimensions;
[0023] Step 1.2.4: Input the content features extracted in Step 1.2.3 into the lip motion network in the MakeItTalk method to obtain the relative displacement of the 3D static landmarks of the lip animation, namely a vector of 204 dimensions;
[0024] Step 1.2.5: Input the speaker style features and content features extracted in Step 1.2.3 into the head motion network in the MakeItTalk method to obtain the relative displacement of the 3D static landmarks of the head animation, namely a vector of 204 dimensions;
[0025] Step 1.2.6: Input the emotion features extracted in Step 1.2.3 into the finally trained facial emotion animation network in 1.1.7 to obtain the relative displacement of the 3D static landmarks of the facial emotion animation, namely a vector of 204 dimensions;
[0026] Step 1.2.7: Add and fuse the relative displacements of the 3D static landmarks obtained from Steps 1.2.4, 1.2.5, and 1.2.6 together with the standard facial key points to obtain the predicted facial key point video frames corresponding to the audio.
[0027] Specifically, Step 2 is as follows:
[0028] Step 2.1: Construct and train an attention-enhanced decoder based on U-net;
[0029] Step 2.2 Input the randomly selected target person image and the predicted facial key-point video frames into the attention-enhanced decoder based on U-net to generate a realistic talking head video.
[0030] Specifically, Step 2.1 is as follows:
[0031] Step 2.1.1: Download the publicly available VoxCeleb2 dataset and MEAD dataset, both of which have high-quality head videos of different speakers and corresponding audio.
[0032] Step 2.1.2: Convert the talking head videos in the dataset of Step 2.1.1 to 25fps.
[0033] Step 2.1.3: Extract the 2D facial key-point information of the head videos in the dataset processed in Step 2.1.2 through the face_alignment library, then perform normalization with standard facial key points, and finally store it in a document.
[0034] Step 2.1.4: Build an attention-enhanced decoder based on U-net by adding a CBAM module to the U-net model. A CBAM module is added before each upsampling layer of U-net except for the last upsampling layer of U-net.
[0035] Step 2.1.5: Select any video frame from the dataset processed in Step 2.1.2, randomly grab two different images, one as the input target image and the other as the output target image. Adopt the facial alignment method to obtain key-point information by predicting the heat map of facial key points of the output target image; extract the lip key-point information of the output target image from the document storing the 2D facial key-point information in Step 2.1.3 to replace the lip information in the key-point information obtained from the heat map, and finally obtain 74-bit facial key-point information.
[0036] Step 2.1.6: Connect the facial key-point information obtained in Step 2.1.5 and the input target image randomly grabbed in Step 2.1.5 through channels, input them into the attention-enhanced decoder based on U-net constructed in Step 2.1.4, and finally generate a highly imitated target image. Compare the generated image with the output target image randomly grabbed in Step 2.1.5 for loss to train the attention-enhanced decoder based on U-net, so that the decoder can finally generate high-fidelity images.
[0037] Step 2.2 is specifically as follows: the randomly selected target person image and the predicted facial key point video frame obtained from step 1.2.7 are input together with the channel connection into the U-net-based attention enhanced decoder trained in step 2.1.6, and finally a high-fidelity talking head video of the target person with both lip and head and facial emotion animation is generated.
[0038] The beneficial effects of the present invention are:
[0039] (1) The invention discloses a method for generating emotional animation based on U-net attention-enhanced decoder, which is a method for constructing and training a facial emotional animation network, and the method can be applied to any character image. The network solves the problem of inconsistency and unnaturalness between head movement and facial expression and audio.
[0040] (2) The emotional animation generation method based on the U-net attention enhancement decoder of the present invention uses a U-net attention enhancement decoder. The decoder adds multiple convolutional attention modules on the basis of the U-net model, namely the CBAM module. The module consists of two sub-modules: spatial attention and channel attention. The spatial attention enables the neural network to pay more attention to the pixel areas in the image that play a decisive role in facial expressions and lip shapes, while ignoring unimportant areas. Channel attention is used to process the distribution relationship of feature map channels. This module is added before sampling on the U-net model, aiming to enhance the representation ability of features by focusing on important information in decoding and suppressing unnecessary information from encoded features, reduce the loss of useful information, improve the performance of the model, and enable the output image to retain more details, such as: complex skin texture and facial shadows of the target person. The decoder can further improve the quality of the generated head video and enhance the facial emotional expression. The method finally outputs a high-fidelity video frame of the target person with both lips and head and facial emotional animation, which can improve the quality of the generated video and enable the output image to retain more details. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flow chart of the emotional animation generation method based on U-net attention enhancement decoder of the present invention;
[0042] Figure 2 It is a comparison diagram of randomly captured character images after using the method of the present invention and high-fidelity video frames generated with the same audio emotions of different characters. DETAILED DESCRIPTION
[0043] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] The present invention provides an emotional animation generation method based on U-net attention enhancement decoder, specifically:
[0045] Step 1: Prediction of facial keypoints integrating facial emotions. In the stage of facial keypoint prediction, regarding the facial emotion network, there are few publicly reported research results at present, and most methods lack generalization ability, that is, they are only applicable to the character images in the dataset. For this reason, the present invention discloses a method for constructing and training a facial emotion animation network, which can be applicable to any character image and enables the finally predicted video frames of facial keypoints to have emotional expressions. This stage solves the problem of incoordination and unnaturalness between head movement, facial expression and audio.
[0046] Step 1 is specifically implemented according to the following steps:
[0047] Step 1.1: Construct and train a facial emotion animation network
[0048] Step 1.1 is specifically implemented according to the following steps:
[0049] Step 1.1.1: Download the latest publicly available multi-modal emotion and emotion detection dataset, abbreviated as the MEAD dataset, which contains high-quality head videos and audio with the same text for different emotions.
[0050] Step 1.1.2: Convert the talking head videos in the MEAD dataset of Step 1.1.1 to 62.5 fps and set the audio sampling rate to 16KHz;
[0051] Step 1.1.3: Extract the 3D facial keypoint information of the head videos in the MEAD dataset processed in Step 1.1.2 through the face_alignment library, then perform normalization processing with standard facial keypoints, and finally store it in a document;
[0052] Step 1.1.4: On the basis of the MEAD dataset processed in Step 1.1.2, train an emotion encoder for extracting emotion auxiliary features through cross-reconstruction emotion disentanglement technology;
[0053] Step 1.1.5: Use the emotion encoder in Step 1.1.4 to extract the emotion feature information in the audio of the MEAD dataset processed in Step 1.1.2 and store it in a document;
[0054] Step 1.1.6: Construct a facial emotion animation network through a recurrent network LSTM and a multi-layer perceptron;
[0055] Step 1.1.7: Input the emotional feature information in the audio of the MEAD dataset obtained in Step 1.1.5 into the facial emotional animation network constructed in Step 1.1.6. Through mapping, obtain the corresponding facial landmarks, and then compare the loss with the facial key points of the MEAD dataset extracted in Step 1.1.3 to train the facial emotional animation network constructed in Step 1.1.6.
[0056] Step 1.2: Select any audio segment from the MEAD dataset in Step 1.1.1 to obtain the predicted facial key point video frames corresponding to the audio.
[0057] Step 1.2 is specifically implemented according to the following steps:
[0058] Step 1.2.1: Select any audio segment from the MEAD dataset in Step 1.1.1.
[0059] Step 1.2.2: Preprocess the audio selected in Step 1.2.1, modify the frame rate of the audio to 62.5 Hz, and the sampling rate of the speech waveform to 16 KHz.
[0060] Step 1.2.3: Input the audio processed in Step 1.2.2 into the pre-trained text (AutoVC), speaker style, and emotion encoders respectively to obtain the content, speaker style, and emotion features in the audio, that is, vectors of 80, 256, and 128 dimensions.
[0061] Step 1.2.4: Input the content features extracted in Step 1.2.3 into the lip movement network (content animation network) in the MakeItTalk method to obtain the relative displacement of the 3D static landmarks of the lip animation, that is, a vector of 204 (68×3) dimensions.
[0062] Step 1.2.5: Input the speaker style features and content features extracted in Step 1.2.3 into the head movement network (speaker style animation network) in the MakeItTalk method to obtain the relative displacement of the 3D static landmarks of the head animation, that is, a vector of 204 (68×3) dimensions.
[0063] Step 1.2.6: Input the emotion features extracted in Step 1.2.3 into the finally trained facial emotion animation network in 1.1.7 to obtain the relative displacement of the 3D static landmarks of the facial emotion animation, that is, a vector of 204 (68×3) dimensions.
[0064] Step 1.2.7: Add and fuse the relative displacements of the 3D static landmarks obtained from Steps 1.2.4, 1.2.5, and 1.2.6 together with the standard facial key points to obtain the predicted facial key point video frames corresponding to the audio.
[0065] Step 2: Facial keypoint decoding based on the U-net attention enhanced decoder. In the stage of decoding facial keypoints to generate a talking head video, the predicted facial keypoint video frames obtained from Step 1 and the target person's image are input into the U-net based attention enhanced decoder to generate a realistic talking head video.
[0066] Step 2 is specifically implemented according to the following steps:
[0067] Step 2.1 Construct and train a U-net based attention enhanced decoder.
[0068] Step 2.1 is specifically implemented according to the following steps:
[0069] Step 2.1.1 Download the publicly available VoxCeleb2 dataset and MEAD dataset. These datasets all have high-quality head videos of different speakers and corresponding audio.
[0070] Step 2.1.2 Convert the talking head videos in the dataset from Step 2.1.1 to 25fps.
[0071] Step 2.1.3 Extract the 2D facial keypoint information of the head videos in the dataset processed in Step 2.1.2 through the face_alignment library, then normalize it with standard facial keypoints, and finally store it in a document.
[0072] Step 2.1.4 Construct a U-net based attention enhanced decoder by adding a CBAM module to the U-net model. A CBAM module is added before each upsampling layer of the U-net except for the last upsampling layer of the U-net.
[0073] Step 2.1.5 Randomly select two different images from an arbitrary video frame in the dataset processed in Step 2.1.2. One is used as the input target image and the other is used as the output target image. In the way of facial alignment, obtain the keypoint information by predicting the heatmap of facial keypoints from the output target image; extract the lip keypoint information of the output target image from the document storing the 2D facial keypoint information in Step 2.1.3 to replace the lip information in the keypoint information obtained from the heatmap, and finally obtain 74-bit facial keypoint information.
[0074] Step 2.1.6: Connect the facial key point information obtained in Step 2.1.5 and the randomly captured input target image through channels, and input them into the U-net based attention enhanced decoder constructed in Step 2.1.4. Finally, generate a highly imitated target image, and use the contrast loss between the generated image and the output target image randomly captured in Step 2.1.5 to train the U-net based attention enhanced decoder, so that the decoder can finally generate high-fidelity images.
[0075] Step 2.2: Input the randomly selected target person image (256×256) and the predicted face key point video frame into the U-net based attention enhanced decoder to generate a realistic talking head video.
[0076] Specifically, Step 2.2: Connect the randomly selected target person image (256×256) and the predicted face key point video frame obtained from Step 1.2.7 through channels and input them together into the U-net based attention enhanced decoder trained in Step 2.1.6, and finally generate a high-fidelity talking head video of the target person with both lip, head and facial emotion animations.
[0077] Effect display: Randomly capture person images on the website and process them to the specified size, i.e., 256×256. Randomly select a happy audio from the MEAD dataset, and input the audio and the image into the emotion animation generation method based on the U-net attention enhanced decoder, and finally generate high-fidelity video frames of the same audio emotion for different persons, as Figure 2 shown.
[0078] Parameter settings when using the method of the present invention: The facial emotion animation network and the U-net based attention enhanced decoder are trained using the Adam optimizer based on pytorch, the learning rate is set to 10 -4 and the weight decay is set to 10 -6 , and the batch size is 16.
[0079] Embodiment 1
[0080] The emotion animation generation method based on the U-net attention enhanced decoder is specifically as follows:
[0081] Step 1: Predict the face key points integrating facial emotions;
[0082] Step 2: Decode the face key points based on the U-net attention enhanced decoder. In the stage of decoding the face key points to generate a talking head video, input the predicted face key point video frame obtained in Step 1 and the target person image into the U-net based attention enhanced decoder to generate a realistic talking head video.
[0083] Example 2
[0084] A method for generating emotional animations based on a U-net attention enhanced decoder, specifically:
[0085] Step 1: Prediction of facial key points integrating facial emotions;
[0086] Step 1 is specifically as follows:
[0087] Step 1.1: Construct and train a facial emotion animation network;
[0088] Step 1.2: Select an audio segment and obtain the predicted facial key point video frames corresponding to this audio.
[0089] Step 2: Decoding of facial key points based on a U-net attention enhanced decoder. In the stage of decoding facial key points to generate a talking head video, input the predicted facial key point video frames obtained from Step 1 and the target person's image into the attention enhanced decoder based on U-net to generate a realistic talking head video.
[0090] Example 3
[0091] A method for generating emotional animations based on a U-net attention enhanced decoder, specifically:
[0092] Step 1: Prediction of facial key points integrating facial emotions;
[0093] Step 1 is specifically as follows:
[0094] Step 1.1: Construct and train a facial emotion animation network;
[0095] Step 1.1 is specifically as follows:
[0096] Step 1.1.1: Download the latest publicly available multi-modal emotion and emotion detection dataset, abbreviated as the MEAD dataset. This dataset contains high-quality head videos and audio with different emotions for the same text.
[0097] Step 1.1.2: Convert the talking head videos in the MEAD dataset of Step 1.1.1 to 62.5fps and set the audio sampling rate to 16KHz;
[0098] Step 1.1.3: Extract the 3D facial key point information of the head videos in the MEAD dataset processed in Step 1.1.2 through the face_alignment library, then normalize it with standard facial key points, and finally store it in a document;
[0099] Step 1.1.4: Based on the MEAD dataset processed in Step 1.1.2, train an emotion encoder that extracts emotion-assisted features through cross-reconstruction emotion disentanglement technology;
[0100] Step 1.1.5: Use the emotion encoder in Step 1.1.4 to extract the emotion feature information in the audio of the MEAD dataset processed in Step 1.1.2 and store it in a document;
[0101] Step 1.1.6: Construct a facial emotion animation network through a recurrent network LSTM and a multi-layer perceptron;
[0102] Step 1.1.7: Input the emotion feature information in the audio of the MEAD dataset obtained in Step 1.1.5 into the facial emotion animation network constructed in Step 1.1.6, obtain the corresponding facial landmarks through mapping, and then train the facial emotion animation network constructed in Step 1.1.6 through contrast loss with the facial key points of the MEAD dataset extracted in Step 1.1.3.
[0103] Step 1.2: Select an audio and obtain the predicted facial key point video frames corresponding to the audio.
[0104] Step 2: Facial key point decoding based on the U-net attention-enhanced decoder. In the stage of decoding facial key points to generate a talking head video, input the predicted facial key point video frames obtained in Step 1 and the target person's image into the U-net-based attention-enhanced decoder to generate a realistic talking head video.
Claims
1. Emotional animation generation method based on U-net attention enhanced decoder, characterized in that: Specifically: Step 1: Prediction of facial key points with facial emotions; Step 2: Facial key point decoding based on U-net attention enhanced decoder. In the stage of decoding facial key points to generate talking head video, the predicted facial key point video frame obtained from step 1 and the target person image are input into the U-net based attention enhanced decoder to generate a realistic talking head video. Step 2 is as follows: Step 2.1 Build and train a U-net-based attention-enhanced decoder; Step 2.1 is as follows: Step 2.1.1, download the public VoxCeleb2 dataset and MEAD dataset, which have high-quality head videos and corresponding audio of different speakers; Step 2.1.2, convert the talking head video in the dataset of step 2.1.1 to 25fps; Step 2.1.3, extract the 2D face key point information of the head video in the data set processed in step 2.1.2 through the face_alignment library, then normalize it with standard facial key points, and finally store it in the document; Step 2.1.4: Build a U-net-based attention-enhanced decoder by adding a CBAM module to the U-net model. Except for the last upsampling layer of U-net, a CBAM module is added before each upsampling layer of U-net. Step 2.1.5, select any video frame from the data set processed in step 2.1.2, randomly grab two different images, one as the input target image and the other as the output target image; use the face alignment method to predict the heat map of the face key points of the output target image to obtain the key point information; extract the lip key point information of the output target image from the document storing the 2D face key point information in step 2.1.3, and use it to replace the lip information in the key point information obtained from the heat map, and finally obtain 74-bit facial key point information; Step 2.1.6, connect the facial key point information obtained in step 2.1.5 and the input target image randomly captured in step 2.1.5 through the channel, input them into the U-net-based attention enhancement decoder constructed in step 2.1.4, and finally generate a high-fidelity target image. The generated image is compared with the output target image randomly captured in step 2.1.5 to train the U-net-based attention enhancement decoder, so that the decoder finally generates a high-fidelity image; Step 2.2 Input the randomly selected target person image and the predicted facial keypoint video frame into the U-net based attention enhanced decoder to generate a realistic talking head video.
2. The emotional animation generation method based on U-net attention enhanced decoder according to claim 1 is characterized in that: Step 1 is as follows: Step 1.1, build and train the facial emotion animation network; Step 1.2: Select a segment of audio and obtain a video frame of predicted facial key points corresponding to the audio.
3. The emotional animation generation method based on U-net attention enhanced decoder according to claim 2 is characterized in that: Step 1.1 is as follows: Step 1.1.1, download the latest publicly available multimodal emotion and emotion detection dataset, referred to as the MEAD dataset, which contains high-quality head videos and audios with the same text of different emotions; Step 1.1.2, convert the talking head video in the MEAD dataset of step 1.1.1 to 62.5fps and set the audio sampling rate to 16KHz; Step 1.1.3, extract the 3D face key point information of the head video in the MEAD dataset processed in step 1.1.2 through the face_alignment library, then normalize it with standard facial key points, and finally store it in the document; Step 1.1.4: Based on the MEAD dataset processed in step 1.1.2, the emotion encoder for extracting emotion auxiliary features is trained by cross-reconstruction emotion disentanglement technology; Step 1.1.5, use the emotion encoder of step 1.1.4 to extract the emotion feature information in the audio of the MEAD dataset processed in step 1.1.2, and store it in the document; Step 1.1.6, construct the facial emotion animation network through the recursive network LSTM and multi-layer perceptron; Step 1.1.7: Input the emotional feature information in the MEAD dataset audio obtained in step 1.1.5 into the facial emotion animation network constructed in step 1.1.6, obtain the corresponding facial landmarks through mapping, and then perform contrast loss with the MEAD dataset facial key points extracted in step 1.1.3 to train the facial emotion animation network constructed in step 1.1.
6.
4. The emotional animation generation method based on U-net attention enhanced decoder according to claim 3 is characterized in that: Step 1.2 is as follows: Step 1.2.1, select any audio segment from the MEAD dataset in step 1.1.1; Step 1.2.2, pre-process the audio selected in step 1.2.1, modify the frame rate of the audio to 62.5 Hz, and the sampling rate of the speech waveform to 16 KHz; Step 1.2.3, input the audio processed in step 1.2.2 into the pre-trained text, speaker style, and emotion encoders respectively, to obtain the content, speaker style, and emotion features in the audio, i.e., vectors of 80, 256, and 128 dimensions respectively; Step 1.2.4, input the content features extracted in step 1.2.3 into the lip motion network in the MakeItTalk method to obtain the relative displacement of the 3D static landmarks of the lip animation, that is, a 204-dimensional vector; Step 1.2.5, input the speaker style features and content features extracted in step 1.2.3 into the head movement network in the MakeItTalk method to obtain the relative displacement of the 3D static landmarks of the head animation, that is, a 204-dimensional vector; Step 1.2.6, input the emotion features extracted in step 1.2.3 into the facial emotion animation network finally trained in step 1.1.7, and obtain the relative displacement of the 3D static landmarks of the facial emotion animation, that is, a 204-dimensional vector; Step 1.2.7, add and fuse the relative displacement of the 3D static landmarks obtained from steps 1.2.4, 1.2.5, and 1.2.6 with the standard facial key points to obtain a video frame of predicted facial key points corresponding to the audio.
5. The emotional animation generation method based on U-net attention enhanced decoder according to claim 1 is characterized in that: Step 2.2 is specifically as follows: the randomly selected target person image and the predicted facial key point video frame obtained from step 1.2.7 are input together with the channel connection into the U-net-based attention enhanced decoder trained in step 2.1.6, and finally a high-fidelity talking head video of the target person with both lip and head and facial emotion animation is generated.
Citation Information
Patent Citations
Face animation generation method based on memory sharing and attention enhancement
CN117152310A
Fine-grained emotion control speaking face video generation method, system and equipment based on audio and single image driving and medium
CN117409121A