Facial expression generation method and system based on diffusion model and continuous controllable emotion

By using a diffusion-based facial generation system, expression vectors are extracted and integrated with a reference feature network. The generator then generates target emotional expression vectors, solving the problems of uneditable expression intensity and insufficient generalization ability in existing technologies, and achieving high-quality generation of speaking face videos.

CN122115655APending Publication Date: 2026-05-29SUPER ROBOT RESEARCH INSTITUTE (HUANGPU) +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUPER ROBOT RESEARCH INSTITUTE (HUANGPU)
Filing Date
2026-03-13
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies cannot flexibly edit the intensity and emotion of facial expressions in speaking face generation, and it is difficult to efficiently learn emotional information from data without emotion annotations to improve the generalization performance of the model.

Method used

By constructing a face generation system based on a diffusion model, extracting expression vectors and establishing an emotion label mapping relationship, introducing a reference feature network for feature integration, using an emotion editing condition generator to generate neutral and target emotion expression vectors, and combining time, audio, and expression embedding to regulate video generation, fine-grained emotion control is achieved.

Benefits of technology

It enables explicit and continuous facial expression editing, supports single-sample generation, can generalize to unseen identity images, improves emotional accuracy, linearity of emotional intensity editing, and video quality, and ensures lip synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122115655A_ABST
    Figure CN122115655A_ABST
Patent Text Reader

Abstract

The application discloses a face video generation method and system based on a diffusion model and capable of continuously controlling emotions, and the method comprises the following steps: obtaining a speaker face video, performing 3D face reconstruction on video frames, and extracting expression vectors; a corresponding mapping relationship between the extracted expression vectors and emotion labels is established, an emotion editing condition generator is constructed and trained, and corresponding expression vectors are generated according to the emotion labels; a diffusion model is constructed and trained, noise frames, identity frames and motion frames are spliced in the channel dimension to serve as the input of the diffusion model, and time embedding, audio embedding and expression embedding are fused for regulation and control; a reference feature network is constructed, and features output by each layer of the reference feature network are integrated into the diffusion model based on a cross attention and a spatial attention mechanism; and the face video frames output by the diffusion model are subjected to super-resolution processing to obtain a speaker face video. The application can explicitly and continuously edit the emotions of a speaker face video and realize fine-grained control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video generation technology, specifically to a method and system for generating facial motion with continuous and controllable emotion based on a diffusion model. Background Technology

[0002] Audio-driven face animation tasks involve generating speaking face videos using the speaker's identity image and audio, and have wide applications in digital assistants, film and television production, and virtual video conferencing. With the development of AI-generated content, especially the emergence of Generative Adversarial Networks (GANs) and diffusion models, speaking face generation has become a research hotspot in recent years.

[0003] Most existing methods primarily focus on lip synchronization and video quality, neglecting the control of emotion. However, emotion-driven speaking face generation is a crucial aspect of producing realistic and expressive animated faces. While some research on speaking face generation has begun to focus on incorporating facial emotions, it lacks the flexibility to edit these expressions, such as adjusting the intensity of the expressions.

[0004] Early research on emotion-driven speaking face generation only used one-hot encoded emotion labels as the emotion source, lacking the ability to edit intensity. Emotionally-Driven Video Portraits (EVP) methods provide a way to edit emotion categories and intensities through interpolation in the emotion space, but cannot generalize to unseen faces. Emotion-Aware Motion Models (EAMM) methods not only support emotion editing but also achieve single-sample generation, but their emotion sources depend on other videos, making them less convenient to use. Furthermore, large-scale unannotated audio-visual datasets are easier to obtain than emotionally-annotated data. These datasets are typically collected under non-experimental recording conditions and contain diverse facial and background information; using them as training data can enhance the model's generalization performance. Therefore, how to effectively and flexibly control the emotion of synthetic videos, and how to efficiently learn emotional information from data lacking emotion annotation and obtained under non-experimental recording conditions to improve the model's generalization performance, have become critical issues that urgently need to be addressed. Summary of the Invention

[0005] To overcome the shortcomings and deficiencies of existing technologies, this invention provides a method and system for generating facial motion with continuous and controllable emotion based on a diffusion model. This invention reconstructs facial geometry from videos, extracts expression vectors, and creates training pairs consisting of video emotion tags and corresponding expression vectors. Subsequently, given input audio and expression vectors, a time-based denoising network is trained to generate videos that match the input. To maintain the consistency of the target identity, a reference feature network is introduced to extract feature maps from the target identity, which are then integrated into the denoising network through cross-attention. During inference, an emotion editing conditional generator is trained using emotion tags and expression vector pairs. The neutral expression vector generated by the emotion editing conditional generator and the target emotion expression vector are used to calculate the final expression vector, which is then used by the conditional denoising network to synthesize videos corresponding to the desired emotion and its intensity. This invention enables explicit and continuous editing of the emotion in speaking face videos, achieving fine-grained control.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] This invention provides a method for generating facial movements with continuous and controllable emotion based on a diffusion model, comprising the following steps:

[0008] Acquire a video of the speaking person's face, perform 3D face reconstruction on the video frames, and extract expression vectors;

[0009] The extracted facial expression vectors are mapped to corresponding emotion tags. An emotion editing condition generator is constructed and trained. The emotion editing condition generator generates corresponding facial expression vectors based on the emotion tags.

[0010] A diffusion model is constructed and trained. Noise frames, identity frames, and motion frames are concatenated in the channel dimension as input to the diffusion model. Temporal embedding, audio embedding, and facial expression embedding are fused for regulation.

[0011] A reference feature network is constructed to extract facial feature maps. The feature maps output by each layer of the reference feature network are integrated into the diffusion model based on cross attention and spatial attention mechanisms.

[0012] The facial video frames output by the diffusion model are obtained, and the speaking facial video is obtained through super-resolution processing.

[0013] As a preferred technical solution, the emotion editing condition generator includes an expression generator, a global discriminator, and a local discriminator;

[0014] The facial expression generator generates a sequence of facial expression vectors based on the emotion label noise vector;

[0015] The global discriminator uses a temporal convolutional network to extract global features, divides the expression vector sequence into real sequences and generated sequences, and predicts sentiment labels;

[0016] The local discriminator is used to ensure that the facial expression vector generated at each time step is similar to the real sample.

[0017] As a preferred technical solution, a global discriminator is trained based on global adversarial loss, a local discriminator is trained based on local adversarial loss, and the distance between generated samples and real samples is minimized based on mean squared error loss.

[0018] As a preferred technical solution, the emotion editing condition generator generates corresponding expression vectors based on emotion tags, specifically including:

[0019] By setting the driving audio, identity image, and emotional tags with intensity, the emotional editing condition generator produces neutral emotional expression vectors. and target non-neutral emotional expression vector The calculated emotional intensity is The target non-neutral emotional expression vector is represented as:

[0020] ;

[0021] ;

[0022] in, This represents the editing direction vector from neutral sentiment to the target non-neutral sentiment.

[0023] As a preferred technical solution, the adjustment is achieved by integrating time embedding, audio embedding, and facial expression embedding, specifically including:

[0024] Obtain audio features, input the audio features and time information into the diffusion model, represented as:

[0025] ;

[0026] in, and For the continuous hidden states of the diffusion model, Indicates group normalization, The scaling and offset parameters represent time. Scaling and offset parameters for audio features;

[0027] The facial expression information is input into the diffusion model and represented as:

[0028] ;

[0029] in, These are the scaling and offset parameters for facial expression information.

[0030] As a preferred technical solution, the feature maps output by each layer of the reference feature network are integrated into the diffusion model based on cross-attention and spatial attention mechanisms, specifically including:

[0031] The face feature map extracted by the reference feature network is combined with the continuous hidden states of the diffusion model through cross-attention and spatial attention mechanisms. By mixing, we obtain a continuous hidden state that is more relevant to facial features.

[0032] As a preferred technical solution, the diffusion model employs a time-based UNet denoising backbone network, performing cross-interaction and spatial attention calculations between the output of each residual block and the feature map of the corresponding layer of the reference feature network.

[0033] As a preferred technical solution, the loss function in the training step of the diffusion model is:

[0034] ;

[0035] in, Indicates the total loss. This represents the main loss, used for predicting noise. This represents the variational lower bound loss. This represents the loss in the lip region, used to minimize the noise prediction error in the lip region. This represents the loss in the eye region, used to minimize the noise prediction error in the eye region. , and To weigh the parameters.

[0036] This invention also provides a facial motion generation system based on a diffusion model with continuous and controllable emotion, used to implement the above-mentioned facial motion generation method based on a diffusion model with continuous and controllable emotion, including: a speaking face video acquisition module, an expression vector extraction module, an emotion editing condition generator training module, a diffusion model training module, a reference feature network construction module, a feature map integration module, and a super-resolution processing module;

[0037] The speaking face video acquisition module is used to acquire speaking face videos;

[0038] The expression vector extraction module is used to perform 3D face reconstruction on video frames and extract expression vectors;

[0039] The emotion editing condition generator training module is used to train the emotion editing condition generator, establish a corresponding mapping relationship between the extracted expression vectors and emotion tags, and generate the corresponding expression vectors according to the emotion tags.

[0040] The diffusion model training module is used to train the diffusion model. It concatenates noise frames, identity frames, and motion frames in the channel dimension as input to the diffusion model, and integrates temporal embedding, audio embedding, and facial expression embedding for regulation.

[0041] The reference feature network construction module is used to construct a reference feature network, which extracts facial feature maps.

[0042] The feature map integration module is used to integrate the feature maps output by each layer of the reference feature network into the diffusion model based on cross-attention and spatial attention mechanisms;

[0043] The super-resolution processing module is used to acquire the face video frames output by the diffusion model, and then to obtain the speaking face video through super-resolution processing.

[0044] The present invention also provides a computer device, including a processor and a memory for storing processor-executable programs, wherein when the processor executes the program stored in the memory, it implements the above-described method for generating facial movements based on a diffusion model of continuous and controllable emotion.

[0045] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0046] This invention reconstructs facial geometry from videos, extracts expression vectors, and creates training pairs consisting of video emotion labels and corresponding expression vectors. Subsequently, given input audio and expression vectors, a time-based denoising network is trained to generate videos matching the input. To maintain consistency in the target identity, a reference feature network is introduced to extract feature maps from the target identity, which are then integrated into the denoising network through cross-attention. During inference, an emotion editing conditional generator is trained using emotion labels and expression vector pairs. Neutral expression vectors generated by the emotion editing conditional generator and target emotion expression vectors are used to calculate the final expression vector, which is then used by the conditional denoising network to synthesize videos corresponding to the desired emotion and its intensity. This invention can effectively learn emotional information from unlabeled data, using the extracted expression vectors as conditions for generating emotional faces. It can learn rich facial information from unlabeled data, enabling explicit and continuous editing of the emotion in speaking face videos, achieving fine-grained control of emotion category and intensity while ensuring accurate lip synchronization. Furthermore, it supports single-sample generation and can generalize to identity images not seen during training. It outperforms existing emotion-driven face video generation methods in terms of emotion accuracy, linearity of emotion intensity editing, video quality, and lip synchronization. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating the facial motion generation method based on a diffusion model for continuous and controllable emotion generation according to the present invention.

[0048] Figure 2 This is a schematic diagram of the overall implementation architecture of the facial motion generation method based on the diffusion model for continuous and controllable emotion generation according to the present invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0050] Example 1

[0051] like Figure 1 , Figure 2 As shown, this embodiment provides a method for generating facial movements with continuous and controllable emotion based on a diffusion model, including the following steps:

[0052] S1: Capture speaking face videos using image acquisition equipment, and use the DECA method based on FLAME face 3D parametric modeling to reconstruct 3D faces from the video frames, extracting features including identity shape parameters. Expression parameters and head pose parameters For facial parameters, this embodiment only uses expression parameters for emotion control, while identity shape parameters and head posture parameters are discarded.

[0053] In this embodiment, the FLAME model is a 3D facial statistical model that integrates linear identity, shape, and expression modeling. The FLAME 3D model allows for linear editing during expression modeling. This embodiment uses the DECA method based on FLAME face 3D parametric modeling to reconstruct facial geometry from the video, extract expression vectors, and create training pairs consisting of video emotion tags and corresponding expression vectors. For each frame in the speaking face video, the DECA method can regress the FLAME parameters of the face. Subsequently, expression vectors extracted from the training video are used... It is used as an emotional condition and embedded in the process of synthesizing the speaker's face;

[0054] S2: Establish a one-to-one mapping relationship between the extracted facial expression vectors and the corresponding emotion tags to form training data pairs of facial expression vectors and emotion tags, so as to train the emotion editing condition generator. The generator is a conditional generative adversarial network based on the LSTM architecture, which can generate the corresponding facial expression vector sequence according to the given emotion category and intensity, so as to use the emotion condition to regulate the emotion type and intensity of the output video during the inference of the denoising backbone network.

[0055] In this embodiment, sentiment labels are established for each training video. With expression vector sequence Correspondence Train the emotion editing condition generator;

[0056] In this embodiment, control of seven basic emotion categories is supported, including: anger, contempt, disgust, fear, happiness, sadness, and surprise. The intensity of each emotion can be set to... Continuous adjustment within the range;

[0057] In this embodiment, the emotion editing condition generator includes:

[0058] (1) Facial Expression Generator Based on LSTM architecture, given sentiment labels and noise vector generator The expression vector for the current frame is synthesized using the output of the previous frame. Generate a sequence of facial expression vectors The introduced noise vector is randomly generated using uniform sampling to enhance the randomness of the generated samples;

[0059] (2) Global discriminator A Temporal Convolutional Network (TCN) is used to capture long-term dependencies and global features of the sequence, and the entire expression vector sequence is evaluated to distinguish it from the real sequence. With the generated sequence And predict sentiment tags;

[0060] (3) Local discriminator Employing an MLP architecture ensures that the facial expression vectors generated at each time step are accurate and reliable. Compared with real samples resemblance.

[0061] In this embodiment, the training loss of the emotion editing condition generator includes:

[0062] (1) Global combat losses Used to train the global discriminator Specifically, it is expressed as:

[0063] ;

[0064] (2) Localized combat losses Used to train local discriminators Specifically, it is expressed as:

[0065] ;

[0066] (3) Mean square error loss Minimize the distance between generated samples and real samples.

[0067] Training loss of the sentiment editing condition generator Represented as:

[0068] ;

[0069] in, Indicates the trade-off parameters;

[0070] During the inference phase, given the driving audio, identity image, and sentiment labels with intensity, the sentiment editing conditional generator generates neutral sentiment expression vectors. And the target non-neutral emotional expression vector with an emotional intensity of 1 by default. According to the formula , The calculated emotional intensity is The target non-neutral emotional expression vector, where, Indicates the intensity of emotion. This represents the editing direction vector from neutral sentiment to the target non-neutral sentiment.

[0071] In this embodiment, the generator and discriminator losses are used alternately during training to train the generator and discriminator respectively. In order to generate facial expression vector samples that the discriminator cannot distinguish between real and fake, the generator will eventually generate facial expression vectors that are close to real samples. Then, during the inference process, the discriminator is discarded and only the generator is used to generate facial expression vectors. During the inference process, it is used to modulate the input of the diffusion model so that the generated video conforms to the given final facial expression vector.

[0072] S3. Construct a diffusion model to realize the diffusion process. The noise frame, identity frame and motion frame are concatenated in the channel dimension as input. Temporal embedding, audio embedding and facial expression embedding are injected through the conditional residual module.

[0073] In this embodiment, the input signal after splicing noise frames, identity frames and motion frames is modulated based on the time signal. The noise frame is modulated so that time, audio and facial expression information are gradually incorporated into it during the denoising process, so that the output result matches the given conditions. The diffusion model can finally output face video frames corresponding to audio and face identity.

[0074] In this embodiment, the given length is video frame sequence The UNet denoising backbone network model receives three types of images as input: noisy frames. (at time step) (Identified frame obtained by adding noise to the target frame) (from (random sampling) and motion frames (Right now and (used for smooth video generation), final input ,in, Indicates a channel splicing operation;

[0075] In this embodiment, time and audio condition injection specifically includes:

[0076] For a given sequence of video frames, the original audio source is preprocessed to correspond to the number of video frames, and the original audio is encoded using a pre-trained audio encoder to obtain audio features. Audio and temporal information are injected into the backbone network through the conditional residual module, specifically as follows:

[0077] ;

[0078] in, and For the continuous hidden states of UNet, Indicates group normalization, The scaling and offset parameters represent time. For the scaling and offset parameters of the audio, to ensure the accuracy of lip movements, past and future audio segments are spliced ​​with the current frame audio for calculation;

[0079] Facial expression conditional injection uses the same conditional injection method as audio, specifically as follows:

[0080] ;

[0081] in, Scaling and offset parameters for the emoji embedding.

[0082] The diffusion model implements both a diffusion process and a denoising process. The diffusion process progressively adds noise to the data to disrupt its structure, while the denoising process learns to reverse this process to recover the data. In the diffusion process, given data from a distribution... samples In a series of time steps Noise is gradually added, with a small amount of Gaussian noise introduced at each step t, creating an increasingly noisy sample sequence. This process is controlled by the following Gaussian transition: ,in It is a variance scheduling that typically increases over time. Indicates a Gaussian distribution;

[0083] At the same time, it allows at any step Directly from raw data points Calculate samples ,in, , , It is Gaussian noise.

[0084] The reverse process is the generation stage, which starts from pure Gaussian noise samples. Initially, noise is removed iteratively using learned parameters to recover data that follows the data distribution. The sample.

[0085] Each step of the reverse process is defined as: .

[0086] S4. Construct a reference feature network to extract facial feature maps from identity images. During training and inference, the feature maps output by each layer are integrated into the backbone network through cross attention and spatial attention mechanisms. The purpose is to enable the backbone network to refer to identity features during denoising, ensuring that the denoising result is highly matched with the identity features.

[0087] In this embodiment, the diffusion model adopts a time-based UNet denoising backbone network, and the reference feature network adopts a UNet structure similar to the diffusion model network but does not contain temporal information. The network input of the UNet denoising backbone network is a 128×128 image, which is downsampled three times to obtain feature maps of 64×64, 32×32 and 16×16 sizes. Cross and spatial attention calculations are performed between the output of each residual block and the feature map of the corresponding layer of the reference feature network. To reduce computational cost, cross attention operation is only performed when the feature map size is 64, 32 or 16.

[0088] Specifically, the face feature map extracted by the reference feature network is connected to the continuous hidden states of UNet through cross-attention and spatial attention mechanisms. By mixing, a continuous hidden state that is more relevant to facial features is obtained;

[0089] In this embodiment, the backbone network is trained using training data, and the loss function includes the main loss. Variational lower bound loss lip loss and eye loss ;

[0090] Among them, the main loss Used for noise prediction ;

[0091] Variational lower bound loss Includes the KL divergence term;

[0092] Lip area loss Minimize the noise prediction error in the lip region;

[0093] Eye area loss Minimize noise prediction error in the eye region;

[0094] The final loss function is: ,in, , and To weigh the parameters;

[0095] In this embodiment, the feature fusion of the reference feature network and the backbone network adopts a cross-attention mechanism. The cross-attention operation is performed when the feature map size is 64, 32 or 16, and spatial attention is introduced to make the model focus on important regions in the feature map.

[0096] S5. Obtain the face video frames output by the time-series-based UNet denoising backbone network, process them through the super-resolution module built on generative adversarial network, obtain high-resolution speaking face video, and display it on the monitor.

[0097] In this embodiment, rich facial information can be learned from data without emotion annotation. The expression vectors extracted by DECA are used as conditions to enable the backbone network to synthesize videos with different emotions, overcoming the limitation of insufficient diversity in existing emotion annotation datasets. Since the FLAME model has linear editing characteristics in the expression space, when given a series of expression vectors that change continuously along the direction from neutral emotion to non-neutral emotion, it can synthesize speaking face videos with continuously changing expressions, realizing fine-grained continuous control of emotion category and intensity.

[0098] Specifically, this embodiment utilizes an image acquisition device to collect multiple sets of speaking face videos as a training dataset. The training data comes from the MEAD and CREMA speaking face video datasets. The training videos are sampled at 25 frames per second, and the audio preprocessing is 16kHz. To improve the quality of the synthesized video, the same face alignment is used for all training videos. Specifically, the videos are aligned with the tip of the nose as the center and adjusted to a uniform 128×128 pixel resolution. The FAN method is used to obtain 68 facial landmarks per frame for calculation. and ;

[0099] An emotion editing conditional generator and a denoising network were built in a Python integrated development environment. For the emotion editing conditional generator, videos with the highest emotion intensity were selected from the MEAD dataset for face reconstruction. Then, expression vectors were extracted from the reconstruction results to train the emotion editing conditional generator. The emotion editing conditional generator was trained for 200 epochs with a batch size of 8 and a generator learning rate of [missing information]. The discriminator learning rate is The reconstruction loss weight is set to ;

[0100] For the denoising network, training is performed for 1000 epochs with a batch size of 10 and a learning rate of... The loss weights are set as follows: , , Number of steps during training The reasoning process takes 50 steps. Cosine scheduling is used to ensure The value increases smoothly from a small value in the early steps to a larger value in the later steps;

[0101] On the hardware side, the sentiment editing conditional generator is trained and inferred on a single RTX 3090 GPU, while the denoising network is trained in parallel on four A800 GPUs and inferred on a single RTX 3090 GPU. During inference, both the sentiment editing conditional generator and the denoising network use no more than 12GB of VRAM, allowing for efficient operation on consumer-grade GPUs such as the RTX 3060 (12GB) or RTX 3070 (12GB).

[0102] To enable the model to learn identity and background diversity and perform single-sample generation, processed HDTF videos (containing diverse identities and backgrounds but no emotional information) are incorporated into the training process along with sentiment datasets such as MEAD during the training of the denoising network. This allows the model to control the sentiment category and intensity on unseen images.

[0103] To measure the generation quality of the method of this invention on the test set, the following evaluation metrics are used for quantitative comparison:

[0104] (1) Emotion editing capability: The emotional accuracy (EmoAcc) of the generated video was evaluated using an emotion classifier network; the linearity of intensity editing was evaluated using the LIE metric based on LPIPS and the linearity of emotion intensity editing (FLIE) metric based on FLAME.

[0105] (2) Video quality: The generated results were analyzed using Fraser video distance (FVD), peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and cumulative probabilistic fuzzy detection (CPBD);

[0106] (3) Audio and video synchronization: The audio and video synchronization of the synthesis results is estimated using SyncNet confidence.

[0107] Table 1 below shows the comparison results of the sentiment editing capabilities of each method. The numbers in parentheses represent the number of sentiment intensity levels used in the calculation of LIE and FLIE, such as method (15) and method (6) in the table:

[0108] Table 1. Comparison of Emotion Editing Capabilities of Various Methods

[0109]

[0110] Table 2 below shows the quantitative comparison results between the method of the present invention and existing methods in terms of video quality and audio-visual synchronization:

[0111] Table 2. Quantitative comparison of this method with existing methods in terms of video quality and audio-visual synchronization.

[0112]

[0113] Comparison results with existing methods MEAD, EVP, EAMM, MakeItTalk, EMOspeaker, Diffused Heads, and NeRFFaceSpeech show that the method of the present invention performs well in terms of emotional accuracy, linearity of emotional intensity editing, overall video quality, and lip synchronization. In particular, the method of the present invention outperforms Diffused Heads, which also uses a diffusion model architecture, indicating that the introduction of the reference feature network significantly improves the quality of the generated video.

[0114] Example 2

[0115] This embodiment provides a facial motion generation system based on a diffusion model with continuous and controllable emotion, used to implement the above-mentioned facial motion generation method based on a diffusion model with continuous and controllable emotion, including: a speaking face video acquisition module, an expression vector extraction module, an emotion editing condition generator training module, a diffusion model training module, a reference feature network construction module, a feature map integration module, and a super-resolution processing module.

[0116] In this embodiment, the speaking face video acquisition module is used to acquire speaking face video;

[0117] In this embodiment, the expression vector extraction module is used to perform 3D face reconstruction on video frames and extract expression vectors;

[0118] In this embodiment, the emotion editing condition generator training module is used to train the emotion editing condition generator, establish a corresponding mapping relationship between the extracted expression vectors and emotion tags, and generate the corresponding expression vectors according to the emotion tags.

[0119] In this embodiment, the diffusion model training module is used to train the diffusion model. The noise frame, identity frame and motion frame are concatenated in the channel dimension as the input of the diffusion model, and time embedding, audio embedding and facial expression embedding are fused for regulation.

[0120] In this embodiment, the reference feature network construction module is used to construct a reference feature network, which extracts facial feature maps.

[0121] In this embodiment, the feature map integration module is used to integrate the feature maps output by each layer of the reference feature network into the diffusion model based on cross attention and spatial attention mechanisms;

[0122] In this embodiment, the super-resolution processing module is used to acquire the face video frames output by the diffusion model, and then to obtain the speaking face video through super-resolution processing.

[0123] Example 3

[0124] This embodiment provides a computing device, which may be a desktop computer, laptop computer, smartphone, PDA handheld terminal, tablet computer or other terminal device with display function. The computing device includes a processor and a memory. The memory stores one or more programs. When the processor executes the program stored in the memory, it implements the facial motion generation method based on diffusion model emotion continuous and controllable according to Embodiment 1.

[0125] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A method for generating facial movements with continuous and controllable emotion based on a diffusion model, characterized in that, Includes the following steps: Acquire a video of the speaking person's face, perform 3D face reconstruction on the video frames, and extract expression vectors; The extracted facial expression vectors are mapped to corresponding emotion tags. An emotion editing condition generator is constructed and trained. The emotion editing condition generator generates corresponding facial expression vectors based on the emotion tags. A diffusion model is constructed and trained. Noise frames, identity frames, and motion frames are concatenated in the channel dimension as input to the diffusion model. Temporal embedding, audio embedding, and facial expression embedding are fused for regulation. A reference feature network is constructed to extract facial feature maps. The feature maps output by each layer of the reference feature network are integrated into the diffusion model based on cross attention and spatial attention mechanisms. The facial video frames output by the diffusion model are obtained, and the speaking facial video is obtained through super-resolution processing.

2. The facial motion generation method based on diffusion model with continuous and controllable emotion generation according to claim 1, characterized in that, The emotion editing condition generator includes an expression generator, a global discriminator, and a local discriminator; The facial expression generator generates a sequence of facial expression vectors based on the emotion label noise vector; The global discriminator uses a temporal convolutional network to extract global features, divides the expression vector sequence into real sequences and generated sequences, and predicts sentiment labels; The local discriminator is used to ensure that the facial expression vector generated at each time step is similar to the real sample.

3. The facial motion generation method based on diffusion model with continuous and controllable emotion generation according to claim 2, characterized in that, A global discriminator is trained based on global adversarial loss, a local discriminator is trained based on local adversarial loss, and the distance between generated samples and real samples is minimized based on mean squared error loss.

4. The facial motion generation method based on diffusion model with continuous and controllable emotion generation according to claim 1, characterized in that, The emotion editing condition generator generates corresponding expression vectors based on emotion tags, specifically including: By setting the driving audio, identity image, and emotional tags with intensity, the emotional editing condition generator produces neutral emotional expression vectors. and target non-neutral emotional expression vector The calculated emotional intensity is The target non-neutral emotional expression vector is represented as: ; ; in, This represents the editing direction vector from neutral sentiment to the target non-neutral sentiment.

5. The facial motion generation method based on diffusion model with continuous and controllable emotion generation according to claim 1, characterized in that, The modulation is achieved by integrating time embedding, audio embedding, and facial expression embedding, specifically including: Obtain audio features, input the audio features and time information into the diffusion model, represented as: ; in, and For the continuous hidden states of the diffusion model, Indicates group normalization, The scaling and offset parameters represent time. Scaling and offset parameters for audio features; The facial expression information is input into the diffusion model and represented as: ; in, These are the scaling and offset parameters for facial expression information.

6. The facial motion generation method based on diffusion model with continuous and controllable emotion generation according to claim 5, characterized in that, The feature maps output from each layer of the reference feature network are integrated into the diffusion model based on cross-attention and spatial attention mechanisms, specifically including: The face feature map extracted by the reference feature network is combined with the continuous hidden states of the diffusion model through cross-attention and spatial attention mechanisms. By mixing, we obtain a continuous hidden state that is more relevant to facial features.

7. The facial motion generation method based on diffusion model with continuous and controllable emotion generation according to claim 6, characterized in that, The diffusion model employs a time-based UNet denoising backbone network, performing cross-interaction and spatial attention computation between the output of each residual block and the feature map of the corresponding layer of the reference feature network.

8. The facial motion generation method based on diffusion model with continuous and controllable emotion generation according to claim 1, characterized in that, In the training steps of the diffusion model, the loss function is: ; in, Indicates the total loss. This represents the main loss, used for predicting noise. This represents the variational lower bound loss. This represents the loss in the lip region, used to minimize the noise prediction error in the lip region. This represents the loss in the eye region, used to minimize the noise prediction error in the eye region. , and To weigh the parameters.

9. A facial motion generation system based on a diffusion model for continuous and controllable emotion generation, characterized in that, The method for generating facial motion based on diffusion model with continuous and controllable emotion as described in any one of claims 1-8 includes: a speaking face video acquisition module, an expression vector extraction module, an emotion editing condition generator training module, a diffusion model training module, a reference feature network construction module, a feature map integration module, and a super-resolution processing module. The speaking face video acquisition module is used to acquire speaking face videos; The expression vector extraction module is used to perform 3D face reconstruction on video frames and extract expression vectors; The emotion editing condition generator training module is used to train the emotion editing condition generator, establish a corresponding mapping relationship between the extracted expression vectors and emotion tags, and generate the corresponding expression vectors according to the emotion tags. The diffusion model training module is used to train the diffusion model. It concatenates noise frames, identity frames, and motion frames in the channel dimension as input to the diffusion model, and integrates temporal embedding, audio embedding, and facial expression embedding for regulation. The reference feature network construction module is used to construct a reference feature network, which extracts facial feature maps. The feature map integration module is used to integrate the feature maps output by each layer of the reference feature network into the diffusion model based on cross-attention and spatial attention mechanisms; The super-resolution processing module is used to acquire the face video frames output by the diffusion model, and then to obtain the speaking face video through super-resolution processing.

10. A computer device comprising a processor and a memory for storing a processor-executable program, characterized in that, When the processor executes the program stored in the memory, it implements the facial motion generation method based on the diffusion model with continuous and controllable emotion as described in any one of claims 1-8.