Emotional speaking head video generation model training method and system

By preprocessing and feature fusion of the video training set, combining Mamba decoder and LLaVA model to optimize the generation model, the problem of insufficient accuracy in the generation of emotional speaking head videos is solved, and higher quality facial expressions and emotional expressions are achieved.

CN120374813APending Publication Date: 2025-07-25XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510508689.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, when generating emotional speaking videos, it is difficult to accurately capture subtle changes and dynamic transitions of emotions, and the coupled information of facial expressions and pronunciation is interfered with, resulting in inaccurate generation.

Method used

By obtaining the video training set for preprocessing, combining geometric features, emotional features and facial 3D expression sequences, the feature fusion is performed using the Mamba decoder and Audio-to-ExpressionTransformer model, and the loss function calculation and optimization generation model is performed through the LLaVA model.

Benefits of technology

It significantly improves the accuracy and consistency of facial expression generation, improves the accuracy of emotional expression, and the generated video is more in line with the target emotional state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374813A_ABST
    Figure CN120374813A_ABST
Patent Text Reader

Abstract

The invention discloses an emotional speaking head video generation model training method and system, and the method comprises the steps: S1, obtaining a video training set, and carrying out the preprocessing of the video training set, and obtaining a video source image, a source audio, and the head posture of the video source image; s2, inputting the video source image, the source audio and the head posture of the video source image into an emotional speaking head video generation model to obtain an emotional speaking head video; s3, carrying out loss function calculation based on the emotional speaking head video and the video training set, and reversely optimizing an emotional speaking head video generation model; and S4, executing the steps S1 to S3 until the loss function is minimum, and outputting an optimal emotion speaking head video generation model. According to the invention, accurate emotion expression of the emotional speaking head video can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of training of emotional talking head video generation models, and particularly to a method and system for training an emotional talking head video generation model. Background Art

[0002] In recent years, due to the increasingly wide application of digital humans in industries such as education and entertainment, the generation of realistic head talking videos with emotions has attracted much attention. With the development of generation models, especially generative adversarial networks (GANs) and diffusion models, the head talking video generation technology has also developed rapidly and has good research prospects. Audio-driven emotional head talking video generation aims to generate realistic head talking videos with corresponding emotions synchronized with speech through external emotional cues. In interpersonal communication, human speech expressions often contain rich emotional factors. Therefore, in the process of simulating real human conversations, the integration of emotional elements is crucial and can endow language with vividness and authenticity.

[0003] Currently, existing work still faces certain challenges in terms of emotional sources and generation quality. Specifically, in terms of emotional sources, currently, many methods rely on emotional labels (such as happy, angry, sad, etc.) to drive emotional expressions. However, the information provided by emotional labels is relatively limited and cannot capture the subtle changes and dynamic transitions of emotions. For example, the same emotional label (such as "happy") may correspond to multiple different facial expressions and speech intonations, and the label itself cannot fully express these differences. Another common method is to extract emotional information from the driving video. Although this method can capture richer emotional details, the extraction process is relatively complicated and is easily interfered by the coupled information of facial movements. For example, the change of facial expressions may be affected by both emotions and speech pronunciation at the same time, making it difficult to completely separate the two, resulting in inaccurate generated emotional expressions. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for training an emotional talking head video generation model, aiming to solve the accuracy of emotional expression in the generation of emotional talking head videos.

[0005] The present invention provides a method for training an emotional talking head video generation model, including: S1. Obtain a video training set, and preprocess the video training set to obtain video source images, source audio, and the head poses of the video source images; S2. Input the video source images, source audio, and the head poses of the video source images into the emotional talking head video generation model to obtain an emotional talking head video; S3. Calculate a loss function based on the emotional talking head video and the video training set, and reversely optimize the emotional talking head video generation model; S4. Execute steps S1 to S3 until the loss function is minimized and the optimal emotional talking head video generation model is output.

[0006] The present invention also provides a training system for an emotional talking head video generation model, including: A preprocessing module, configured to obtain a video training set, and preprocess the video training set to obtain a video source image, source audio, and the head pose of the video source image; An input module, configured to input the video source image, source audio, and the head pose of the video source image into the emotional talking head video generation model to obtain an emotional talking head video; A loss function calculation module, configured to calculate a loss function based on the emotional talking head video and the video training set, and inversely optimize the emotional talking head video generation model; An output module, configured to execute the preprocessing module, the input module, and the loss function calculation module until the loss function is minimized and the optimal emotional talking head video generation model is output.

[0007] By adopting the embodiments of the present invention, it is possible to accurately express emotions in the generated emotional talking head video.

[0008] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it is implemented in accordance with the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Description of the Drawings

[0009] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 is a flowchart of a method for training an emotional talking head video generation model according to an embodiment of the present invention; Figure 2 is a specific flow schematic diagram of a method for training an emotional talking head video generation model according to an embodiment of the present invention; Figure 3 is a schematic diagram of a training system for an emotional talking head video generation model according to an embodiment of the present invention.

[0011] Description of the reference numerals: 310: Preprocessing module, 320: Input module, 330: Loss function calculation module, 340: Output module. Detailed Embodiments

[0012] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the protection scope of the present invention.

[0013] Embodiment 1 As Figure 1 shown, according to an embodiment of the present invention, a method for training an emotional talking head video generation model is provided, specifically including: S1. Obtain a video training set, and preprocess the video training set to obtain a video source map, source audio, and the head pose of the video source map; S2. Input the video source map, source audio, and the head pose of the video source map into the emotional talking head video generation model to obtain an emotional talking head video; S3. Calculate a loss function based on the emotional talking head video and the video training set, and reversely optimize the emotional talking head video generation model; S4. Execute steps S1 to S3 until the loss function is minimized to output an optimal emotional talking head video generation model.

[0014] In this embodiment, the emotional talking head video generation model in S2 specifically includes: A geometric feature module for obtaining depth map geometric features after inputting the video source map; An emotional feature module for obtaining emotional deformation features after inputting the video source map; An attention fusion module for weighted fusion of the depth map geometric features and the emotional deformation features to obtain fusion features; A facial 3D expression sequence module for obtaining a facial 3D expression sequence after inputting the source audio and the head pose of the video source map; A first decoder for decoding the fusion features and the facial 3D expression sequence to obtain an emotional talking head video.

[0015] Further, the obtaining of the depth map geometric features after inputting the video source map specifically includes: as Figure 2 shown, input the video source map into a source map encoder to obtain parameters, input the parameters into FLAME to decode the geometric shape of the video source map and input it into a second decoder to obtain the depth map geometric features of the video source map; the parameters include: shape parameters, expression parameters, and head pose parameters.

[0016] In this embodiment, FLAME is a conventional 3D face parametric model, designed specifically for high-precision face modeling, animation, and expression synthesis. It combines shape parameters (β), expression parameters (ψ), and head pose parameters (θ) for joint modeling. The face is described by three types of parameters: β (Shape), which controls the static face shape (such as face shape, nose height, etc.); ψ (Expression), which controls dynamic expressions (such as smiling, frowning, etc.); and θ (Pose), which controls the rigid movements (rotation and translation) of the head and neck. Given these three parameters, FLAME generates a three-dimensional face mesh from the given shape parameters, expression parameters, and head pose parameters. The three-dimensional face mesh models the face through vertices and faces, providing geometric information in the two-dimensional image space.

[0017] Furthermore, the specific steps for obtaining the emotional deformation features after inputting the video source image are as follows: Figure 2 As shown, the video source image is input into the key point detector to obtain face key points. After adding the face key points to the emotion vector, the result is input into the emotion deformation network to obtain the emotional deformation features.

[0018] In this embodiment, the emotion vector is obtained by acquiring the emotion labels in the emotion dataset, that is, the emotion vector. Then, inputting it into the emotion deformation network enhances the expression ability of facial emotions.

[0019] Furthermore, the specific steps for obtaining the facial 3D expression sequence after inputting the source audio and the head pose of the video source image are as follows: Figure 2 As shown, the source audio is respectively input into the acoustic feature encoder and the speech encoder to obtain acoustic features and speech features. The acoustic features and the face key points are concatenated to obtain the concatenated features. The concatenated features and the emotion vector are input into the Emotion Adaptation Module (EAM) for emotion fusion, and then input into the third decoder for dimensionality elevation to obtain the first dimensionality elevation feature. The speech features and the head pose are concatenated and input into the fourth decoder for decoding to obtain the second dimensionality elevation feature. The first dimensionality elevation feature and the second dimensionality elevation feature are input into the stacked Transformer layers for processing to obtain the facial 3D expression sequence.

[0020] In this embodiment, the EAM incorporates emotion information into the features through the learned parameters, making the generated expressions more in line with the specified emotional state.

[0021] In this embodiment, the Facial 3D Expression Sequence Module adopts the Audio-to-Expression Transformer (A2ET) model. The purpose of designing A2ET is to learn audiovisual synchronous expression features from two modalities, video and audio, that is, to map the emotion-related vectors in speech and vision to E iAbove, the facial 3D expression sequence obtained by the A2ET model is presented in the form of key points, with a shape of (batch size, number of key points, 3), where 3 represents the 3 components in the three-dimensional coordinates xyz.

[0022] The batch size is a set of samples input into the model at one time during training. Training in batches can increase efficiency.

[0023] The number of key points is the face key points obtained by the source image through the key point detector.

[0024] In this embodiment, the first decoder is a Mamba decoder.

[0025] The addition of the Mamba decoder can effectively promote the mutual learning between facial actions, thereby enhancing the consistency of facial expressions. The Mamba decoder makes the generated facial expressions more natural and coherent by capturing the dependencies between actions.

[0026] In this embodiment, in S3, calculating the loss function based on the emotional talking head video and the video training set, and reversely optimizing the emotional talking head video generation model specifically includes: Inputting the video training set into the pre-trained LLaVA (Large Language Vision Assistant) model to obtain emotion-related texts, inputting the emotional talking head video into the image encoder for encoding to obtain image features, inputting the emotion text into the text encoder to obtain text features, calculating the loss function for the text features and the image features, and reversely optimizing the emotional talking head video generation model.

[0027] S3 is further specifically implemented through the following steps: S301: Input 8 videos in the video training set into the pre-trained LLaVA model to obtain eight emotion-related texts, and input the eight emotion-related texts into the text encoder in longCLIP-L to obtain eight text features; The present invention uses the LLaVA model to make up for the lack of emotion sources. LLaVA is a multi-modal large model, and its core architecture includes a large language model (LLM) and a visual encoder. The training method of the pre-trained LLaVA model is to extract the video into frames one by one. Each image passes through the visual encoder to obtain visual features, and the visual features pass through a projection layer to obtain visual tokens, which are input into the large language model.

[0028] S302: Input the emotional talking head video into the image encoder in longCLIP-L for encoding to obtain image features, and calculate the Hadamard product between the image feature f v and the eight text features f t to obtain the maximum probability distribution. The formula for the maximum probability distribution is as follows: ; Among them, denotes transpose. Take the dot product of the transpose of the image feature and eight text features , and then normalize it to the range of 0 - 1 using softmax to obtain these 8 values. Select the maximum value as the value most similar to the emotional talking head video and the video training set. The maximum value indicates which emotional text the current frame best matches. Input the most similar value into the cross-entropy loss calculation formula.

[0029] The cross-entropy loss calculation formula is as follows: ; The calculation method is calculated using the method for calculating p. n represents the total number of samples in a batch, and m represents the number of emotion categories. denotes the output score of the i-th sample in the j-th emotion category, that is, the similarity. denotes the actual label of the i-th sample in the j-th emotion category, encoded using the one-hot encoding method.

[0030] Calculate the cross-entropy loss between the maximum value and the true label of the dataset to minimize the difference between the generated image and the true image, thereby improving the training performance.

[0031] S303. By calculating the cross-entropy loss, reverse-optimize the emotional talking head generation model. The emotion-related text can guide the model during the training process, thereby enhancing the emotional expression ability of the generated video. This design not only enhances the model's ability to understand emotions but also makes the generated facial expressions more in line with the target emotional state.

[0032] The quantitative evaluation of the above method specifically includes: The method proposed in the present invention is verified on the MEAD dataset and the LRW dataset respectively. The generation effect of the emotional talking head is verified on the MEAD dataset, and the generation effect of audio-visual synchronization is verified on the LRW dataset. The improvement in the generation quality of the proposed method can be fully demonstrated on both datasets, reaching 22.11 and 21.77 respectively in terms of the PSNR metric, which is nearly 1.5 points and 0.05 points higher than the current MakeIttalk, EAMM, EAT, and Hallo. On the MEAD dataset, the emotion accuracy is reflected by ACCemo. The method proposed in the present invention enables the emotion accuracy to reach 80.54, while the emotion accuracy of the ground truth is 84.37, which is very close to the result of the ground truth.

[0033] The beneficial effects are as follows: EAM incorporates emotional information into features through the learned parameters, making the generated expressions more consistent with the specified emotional states. The input of the emotional vector into the emotional deformation network enhances the facial emotional expression ability. By calculating the cross-entropy loss between the maximum value and the true labels of the dataset, the difference between the generated image and the true image is minimized, thereby improving the training performance.

[0034] Combining depth map geometric features, introducing the Mamba decoder, and leveraging the emotional guidance of the LLaVA model significantly improve the generation of facial details, expression consistency, and the accuracy of emotional expression. Compared with existing methods, the method of the present invention demonstrates stronger robustness and higher generation quality in the facial expression generation task.

[0035] Embodiment 2 According to an embodiment of the present invention, a training system for an emotional talking head video generation model is provided. Figure 3 It is a schematic diagram of the training system for the emotional talking head video generation model according to an embodiment of the present invention, as Figure 3 shown, and specifically includes: A preprocessing module 310, configured to obtain a video training set, and preprocess the video training set to obtain video source images, source audio, and the head poses of the video source images. An input module 320, configured to input the video source images, source audio, and the head poses of the video source images into the emotional talking head video generation model to obtain an emotional talking head video. A loss function calculation module 330, configured to calculate a loss function based on the emotional talking head video and the video training set, and reversely optimize the emotional talking head video generation model. An output module 340, configured to execute the preprocessing module, the input module, and the loss function calculation module until the loss function is minimized to output an optimal emotional talking head video generation model.

[0036] The embodiment of the present invention is a system embodiment corresponding to the above method embodiment. The specific operations of each module can be understood with reference to the description of the method embodiment, and will not be elaborated herein.

[0037] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the technical solutions of the embodiments of the present invention to deviate from the scope of the present solution.

Claims

1. A method for training an emotional talking head video generation model, characterized in that, including, S1. Obtain a video training set, preprocess the video training set to obtain a video source image, source audio, and the head pose of the video source image; S2. Input the video source image, source audio, and the head pose of the video source image into an emotion talking head video generation model to obtain an emotion talking head video; S3. Calculate a loss function based on the emotion talking head video and the video training set, and reversely optimize the emotion talking head video generation model; S4. Execute steps S1 to S3 until the loss function is minimized and the optimal emotion talking head video generation model is output.

2. The method according to claim 1, wherein The emotion talking head video generation model specifically includes: A geometric feature module for obtaining depth map geometric features after inputting the video source image; An emotion feature module for obtaining emotion deformation features after inputting the video source image; An attention fusion module for weighted fusion of the depth map geometric features and the emotion deformation features to obtain fusion features; A facial 3D expression sequence module for obtaining a facial 3D expression sequence after inputting the source audio and the head pose of the video source image; A first decoder for decoding the fusion features and the facial 3D expression sequence to obtain an emotion talking head video.

3. The method according to claim 2, characterized in that, The obtaining of depth map geometric features after inputting the video source image specifically includes: Inputting the video source image into a source image encoder to obtain parameters, inputting the parameters into FLAME to decode the geometry of the video source image, and inputting it into a second decoder to obtain the depth map geometric features of the video source image.

4. The method according to claim 2, wherein The obtaining of emotion deformation features after inputting the video source image specifically includes: Inputting the video source image into a key point detector to obtain face key points, adding the face key points to an emotion vector, and inputting it into an emotion deformation network to obtain emotion deformation features.

5. The method according to claim 4, characterized in that The obtaining of a facial 3D expression sequence after inputting the source audio and the head pose of the video source image specifically includes: Inputting the source audio into an acoustic feature encoder and a speech encoder respectively to obtain acoustic features and speech features, concatenating the acoustic features and the face key points to obtain concatenated features, inputting the concatenated features and the emotion vector into an emotion adaptation module for emotion fusion, then inputting it into a third decoder for dimensionality elevation to obtain a first dimensionality elevation feature, concatenating the speech features and the head pose and inputting it into a fourth decoder for decoding to obtain a second dimensionality elevation feature, and inputting the first dimensionality elevation feature and the second dimensionality elevation feature into a stacked Transformer layer for processing to obtain a facial 3D expression sequence.

6. The method according to claim 1, characterized in that, The calculating of a loss function based on the emotion talking head video and the video training set and reversely optimizing the emotion talking head video generation model specifically includes: Inputting the video training set into a pre-trained LLaVA model to obtain emotion-related text, inputting the emotion talking head video into an image encoder for encoding to obtain image features, inputting the emotion-related text into a text encoder to obtain text features, calculating a loss function for the text features and the image features, and reversely optimizing the emotion talking head video generation model.

7. The method according to claim 3, characterized in that The source image encoder includes: ResNet50 and a fully connected layer connected in sequence.

8. The method according to claim 3, wherein The inputting of the video source image into the source image encoder to obtain parameters specifically includes: Inputting the video source image into the source image encoder to obtain shape parameters, expression parameters, and pose parameters.

9. An emotional talking head video generation model training system, characterized in that including, A preprocessing module, configured to obtain a video training set, preprocess the video training set to obtain a video source image, source audio, and the head pose of the video source image; An input module, configured to input the video source image, source audio, and the head pose of the video source image into an emotional talking head video generation model to obtain an emotional talking head video; A loss function calculation module, configured to calculate a loss function based on the emotional talking head video and the video training set, and reversely optimize the emotional talking head video generation model; An output module, configured to execute the preprocessing module, the input module, and the loss function calculation module until the loss function is minimized and output an optimal emotional talking head video generation model.