A method for simulating speech-driven facial animation based on facial muscle linkage

By constructing a PPMF encoder and an FDCP-based decoder, combined with Mouth Mapping, Mouth2Face and Refine Decoder modules, it simulates the complex dynamics of facial muscle activity, solving the problem that facial animation generation task is simplified to infinitely thin surface skin deformation in the prior art, and achieving high-quality facial animation generation.

CN119579742BActive Publication Date: 2025-07-01SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411722697.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-07-01
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

The prior art ignores the intricate and personalized dynamics of facial muscle activity, resulting in an out-of-balance between lip movement and driving audio, and the facial animation generation task is reduced to infinitely thin surface skin deformation without underlying structure.

Method used

Using a speech-driven facial animation simulation method based on facial muscle linkage, a PPMF encoder and an FDCP-based decoder are constructed, combined with Mouth Mapping, Mouth2Face and Refine Decoder modules, the complex dynamics of facial muscle activities are simulated and personalized facial animation is generated.

Benefits of technology

It realizes the complex dynamics of simulated facial muscle activities, generates realistic and coordinated facial animations, improves the synchronization between lip movement and speech signals, and enhances the details and nature of facial animations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119579742B_ABST
    Figure CN119579742B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for speech-driven facial animation simulation based on facial muscle linkage, belonging to the field of artificial intelligence, including the steps of: S1, constructing a PPMF encoder; S2, constructing a decoder based on FDCP to decode the features F provided by the PPMF to obtain facial animations; S3, training the speech-driven 3D face animation framework DCPTalk; S4, model optimization; S5, model quantitative evaluation. The present invention proposes the DCPTalk framework and, based on the linkage characteristics of facial muscle groups, proposes Mouth2Face. The mouth movement has a strong correlation with the speech signal and is easily synthesized with vocal tract dynamics. In order to further enhance the details of facial movements, a Refine Decoder is used to simulate the skin deformation on the surface to refine the facial animations. The inherent body characteristics and the body characteristics related to the facial muscle group movements are embedded into Mouth2Face to construct a personalized facial muscle control system, and at the same time, the external drive signal is modulated by the speaking style. Qualitative and quantitative experiments and user studies show that DCPTalk is superior to the existing state-of-the-art methods. P The mouth movement has a strong correlation with the speech signal and is easily synthesized with vocal tract dynamics. In order to further enhance the details of facial movements, a Refine Decoder is used to simulate the skin deformation on the surface to refine the facial animations. The inherent body characteristics and the body characteristics related to the facial muscle group movements are embedded into Mouth2Face to construct a personalized facial muscle control system, and at the same time, the external drive signal is modulated by the speaking style. Qualitative and quantitative experiments and user studies show that DCPTalk is superior to the existing state-of-the-art methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology. Specifically, it relates to a method for simulating speech-driven facial animation based on facial muscle linkage. Background Art

[0002] Speech-driven 3D facial animation has become an important focus of scientific interest in both academic and industrial fields. Its applications include multiple fields such as film production, computer games, education and training, and telemedicine. Speech-driven three-dimensional human face animation technology, including methods based on 3DMM and methods based on meshes, has a high degree of diversity. The 3DMM-based methods usually predict intermediate 3DMM coefficients under speech conditions and transform them into 3D space. However, these 3DMM coefficients, as intermediate variables, lack clear semantics and may cause a mismatch between lip movement and the driving audio.

[0003] Facial expressions are generated by 268 muscles exerting traction on soft tissues. The muscles in the mouth area are closely connected to the surrounding muscles. Muscles are the basic driving force for facial expression changes and can ensure coordinated changes in the movements of different areas of the face. In addition, the description of expression details is closely related to the deformation of the surface layer of the skin. However, the synthesis of facial expressions is not just a process of combining facial muscle activities with surface skin deformation; it is also affected by personalized factors such as inherent physical characteristics and acquired speaking styles. For people with a longer or wider facial shape, the traction force of the muscles may be reduced because longer muscle fibers have a weaker traction force than shorter muscle fibers. The increase in collagen under the facial skin further limits the stretchability of these fibers.

[0004] Existing methods usually simplify the facial animation generation task to the deformation of an infinitely thin surface skin without an underlying structure, thus ignoring the intricate and personalized dynamics of facial muscle activities. The structure and activities of facial muscles have been widely studied and applied in various fields. However, the interactions and forces between these muscles have not been thoroughly quantified. Summary of the Invention

[0005] To solve the problem that the above-mentioned existing technology ignores the intricate and personalized dynamics of facial muscle activities, simplifies the facial animation generation task to the deformation of an infinitely thin surface skin without an underlying structure, the present invention provides a method for simulating speech-driven facial animation based on facial muscle linkage.

[0006] To achieve the above technical solution, the present invention provides a method for simulating speech-driven facial animation based on facial muscle linkage, including the steps:

[0007] S1: Construct a PPMF encoder, which consists of an audio feature extractor and a pseudo facial key point extractor, and use a module similar to the Transformer decoder to fuse and align the personalized pseudo facial key point feature F L and the audio feature F A ;

[0008] S2: Construct a decoder based on FDCP to decode the F P feature provided by PPMF to obtain facial animation. This decoder consists of Mouth Mapping, Mouth2Face and Refine Decoder; among them, synthesizing mouth movements from the driving signal, evoking facial animation using the mouth movements, and refining the facial animation are respectively implemented by Mouth Mapping, Mouth2Face and Refine Decoder;

[0009] S3: Train the speech-driven 3D face animation framework DCPTalk. First, use the loss function to train the Mouth2Face module to establish the mapping rule between mouth movement and facial animation; then fix the parameters of the trained Mouth

[0010] 2Face, and start training other components of DCPTalk, and the loss functions for training Mouth2Face and other components are given respectively; among them, the loss function includes reconstruction loss velocity loss and facial key point loss

[0011] S4: Model optimization. Introduce BIWI, Multiface and VOCASET to provide comprehensive analysis and optimization for DCPTalk, and compare VOCA, MeshTalk, FaceFormer, CodeTalker, FaceDiffuse, DiffSpeaker, TalkingStyle and SelfTalk with the method of the present invention; then train DCPTalk on a single NVIDIA A100 GPU;

[0012] S5: Model quantitative evaluation. According to the methods of FaceFormer, CodeTalker and SelfTalk, evaluate the synchronization between speech content and lip movement by calculating the lip vertex error (LVE).

[0013] Furthermore, in step S1, the fusion step of the pseudo facial key point feature F L and the audio feature F A includes:

[0014] S1a: Process F using a multi - head self - attention layer with linear biases (ALiBi). L Perform the processing;

[0015] S1b: Use a multi - head cross - attention layer to align the F obtained from the self - attention layer A and with F L Align;

[0016] S1c: After passing through the feed - forward layer, obtain the personalized pseudo - multi - modal feature F P = [f1 P , …, f t P , …, f T P ∈ R T×d .

[0017] Furthermore, the audio feature extractor consists of an audio encoder A and an audio feature projection layer F A and is expressed by the formula as follows:

[0018] A = Audio Encoder(χ; θ A )

[0019] F A = Audio Projection(Α; ψ A ).

[0020] Furthermore, the pseudo - facial key - point extractor includes an audio encoder and an Audio2lmk decoder, and finally outputs 3D pseudo - facial key - points L = [l1, …, l t , …, l T ∈ R T×68×3 ; Use a personalized factor P ∈ R N for personalized facial key - point modulation; Using the personalized factor P, modulate the pseudo - facial key - points L in the representation space through element - wise addition, and the formula is expressed as follows:

[0021] L = Audio2lmk Decoder(Α; θ L )

[0022] F L = Personlized Modulation(L, P; ψ L ).

[0023] Furthermore, in step S2, the process of Mouth Mapping for synthesizing accurate mouth movements is expressed by the formula as:

[0024]

[0025] Mouth2Face converts the mouth movement through a mouth encoder into a signal Q that controls the facial muscle activity. By using an element-wise quantization function to discretize the facial muscle control signal Q, a facial muscle control command is obtained to activate the relevant facial muscles, thereby obtaining facial animation The process is expressed by the formula as:

[0026]

[0027]

[0028]

[0029] Refine Decoder predicts the displacement of each vertex based on F P Then, the vertex displacement is added to the vertex position of the human face caused by the mouth movement in a basic way. The refinement stage is expressed by the formula as:

[0030]

[0031] Furthermore, Mouth2Face uses the mouth actions and the facial animation Y = [y1,..., y t ,..., y T ∈ R T×V×3 obtained from the real scene for training.

[0032] Furthermore, in step S3, the process of training DCPTalk through the loss function includes the training of the Mouth2Face module, which is expressed by the formula as:

[0033]

[0034]

[0035]

[0036]

[0037]

[0038]

[0039] Furthermore, in step S4, for the pre-trained models lacking BIWI, VOCASET, or Multiface, their official source codes are used to retrain their models.

[0040] Further, in step S4, the model parameters are iteratively updated using the Adam optimizer, with beta1 = 0.9, beta2 = 0.999, and the learning rate is 1×10 -4 ; the feature dimension d is 256; the facial muscle control commands H and P are set to 32 and 128 respectively; for the loss function, λ0 = 1.0, λ lmk = 1×10 -5 , λ1 = 0.1, λ2 = 2.0, λ3 = 1.0, β1 = 0.3, β2 = 4.0, and β3 = 10.0.

[0041] In summary, the present invention has the following beneficial effects compared with the prior art:

[0042] The present invention proposes a new framework called DCPTalk to simulate the complex dynamics of facial muscle activities and depict personalized facial animations. Based on the linkage relationship of facial muscles, the present invention proposes Mouth2Face to simulate the facial muscle control system and generate realistic and coordinated facial animations induced by mouth movements accordingly. Mouth movements have a strong correlation with speech signals and are easily synthesized with vocal tract dynamics. To further enhance the details of facial movements, the present invention uses surface skin deformation to refine the facial animations generated by Mouth2Face, and uses the Refine Decoder to simulate the skin deformation of the surface layer to refine the facial animations. In addition, personalized factors, including inherent physical characteristics and acquired speaking styles, directly determine the uniqueness and realism of facial animations. The inherent physical characteristics related to the movement of facial muscle groups are embedded into Mouth2Face to construct a personalized facial muscle control system, and at the same time, the external driving signals are modulated using the speaking style. Qualitative and quantitative experiments and user studies show that DCPTalk is superior to existing state-of-the-art methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0044] Figure 1 is a flowchart of a method for simulating speech-driven facial animations based on facial muscle linkage provided by the present invention;

[0045] Figure 2 is a schematic diagram of the DCPTalk framework. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0047] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form can also include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0048] See Figure 1 As shown, the present invention provides a method for simulating voice-driven facial animation based on facial muscle linkage, where DCPTalk is the personalized face dynamic coupling characteristic of voice-driven 3D face animation, including the steps:

[0049] S1: Construct a PPMF encoder, which is composed of an audio feature extractor and a pseudo-facial key point extractor;

[0050] As a preference, the audio feature extractor is composed of an audio encoder A and an audio feature projection layer F A Specifically, the audio encoder is an improved large-scale pre-trained self-supervised speech feature extraction model WavLM, which includes a temporal convolutional module (TCN) encoder and a Transformer encoder. A linear interpolation layer is embedded behind the CNN encoder of WavLM to adapt to the frame rate of the target face animation. During training, the parameters of the CNN encoder are fixed, and the parameters of the Transform encoder are optimized. The linear mapping layer is responsible for converting the high-dimensional features extracted by WavLM into lower-dimensional features and adapting them to the face animation generation task. The formula expression of this process is as follows:

[0051] A = Audio Encoder(χ; θ A ) (1)

[0052] F A = Audio Projection(Α; ψ A ) (2)

[0053] where χ represents the original audio input, A = [a1,…,a t ,…,a T ∈ R T×1024 , F A = [f1 A ,…,f t A ,…,f TA ∈ R T×d , θ A and ψ A represent the parameters of the audio encoder and the audio feature projection layer, T is the frame length, and d is the feature dimension.

[0054] As a preference, the pseudo facial key point extractor includes an audio encoder and an Audio2lmk decoder, and finally outputs 3D pseudo facial key points L = [l1, …, l t , …, l T ∈ R T×68×3 . The audio encoder of the pseudo facial key point extractor shares parameters with the audio encoder in the audio feature extractor. Facial landmarks represent facial expression movements and their dynamics. To maintain the unique speaker style, the present invention uses a personalized factor P ∈ R N , that is, a hot vector of the identity label, for personalized facial key point modulation, where N represents the number of subjects used to train DCPTalk. The personalized modulation module includes a linear pseudo facial key point projection layer and a linear personalized factor projection layer. Using the personalized factor P, the pseudo facial key points L in the representation space are modulated by element-wise addition. The relevant formula is expressed as follows:

[0055] L = Audio2lmk Decoder(Α; θ L )(3)

[0056] F L = Personlized Modulation(L, P; ψ L )(4)

[0057] where, F L = [f1 L , …, f t L , …, f T L ∈ R T×d is the personalized pseudo facial key point feature, θ L and ψ L represent the parameters of the Audio2lmk decoder and the personalized modulation module respectively.

[0058] To fuse and align the personalized pseudo facial key point feature F L and the audio feature F A , the present invention uses a module similar to the Transformer decoder. It includes the steps: S1a, using the multi-head self-attention layer with linear bias (ALiBi) for F LProcess it, which assigns higher weights to historical actions closer to the current action; S1b, use a multi-head cross-attention layer to align F obtained by the self-attention layer A and with F L to facilitate information interaction and integration between different modalities; S1c, after passing through a feed-forward layer, obtain personalized pseudo-multi-modal feature F P =[f1 P ,…,f t P ,…,f T P ∈R T×d .

[0059] S2: Construct a decoder based on FDCP to decode the F provided by PPMF P features to obtain facial animation. This decoder is composed of Mouth Mapping, Mouth2Face, and Refine Decoder;

[0060] As Figure 2 shown on the right side of P , the present invention proposes a decoder based on FDCP for decoding facial animation from F M . The speech-driven facial animation process is divided into three stages: synthesizing mouth movements from the driving signal, using the mouth movements to evoke facial animation, and refining the facial animation. These three stages are respectively implemented by Mouth Mapping, Mouth2Face, and Refine Decoder. Compared with the movement of the upper face, mouth movements are easier to generate, which benefits from their direct association with speech pronunciation and vocal tract dynamics. Therefore, the Mouth Mapping module is constructed; Mouth2Face is the core of the FDCP decoder, which can use mouth movements to evoke facial animation; facial animation not only involves the activation of facial muscles but also involves the dynamic changes of the surface skin structure. Therefore, Refine Decoders are used to further refine the facial animation at the vertex level (Y

[0061] As a preference, Mouth Mapping is composed of linear layers for synthesizing precise mouth movements where V M is the number of vertices in the mouth area. The definition of the above process is as follows:

[0062]

[0063] where, θ M is the parameter of the mouth mapping module.

[0064] As a preference, asFigure 2 As shown, Mouth2Face includes a Transformer-based mouth encoder, a Transformer-based face decoder, and a set of facial muscle control commands. Based on the prior knowledge of face animation, the present invention utilizes the mapping rules between pre-trained mouth movements and face animations. Specifically, the mouth encoder converts the mouth movement into a signal Q = [q1, …, q t , …, q T ∈ R T×H×P that controls the facial muscle activities, where H and P respectively represent the number of commands for controlling facial animations and the dimension of each command. In a neural control system, the interaction between neurons usually relies on electrical impulses to transmit information. By using an element-wise quantization function to discretize the facial muscle control signal Q, the facial muscle control commands Facial muscle control commands are indistinguishable during the optimization process. Activate the relevant facial muscles to obtain the facial animation where V is the number of vertices of the entire face. The formula for this process is as follows:

[0065]

[0066]

[0067]

[0068] where φ M and φ F are the parameters of the mouth encoder and the face encoder respectively.

[0069] As a preference, Mouth2Face uses the mouth actions and the facial animation Y = [y1, …, y t , …, y T ∈ R T×V×3 obtained from the real scene for training.

[0070] As a preference, the Refine Decoder consists of a linear layer that predicts the displacement of each vertex based on F P and then adds the vertex displacement to the vertex positions of the face caused by the mouth movement in a basic way. The refinement stage can be expressed as follows:

[0071]

[0072] where is the facial animation in the refinement stage, and φ RThey are the parameters for fine-grained decoding.

[0073] S3: Train the speech-driven 3D face animation framework DCPTalk. First, use the loss function to train the Mouth2Face module to establish the mapping rules between mouth movements and facial animations. Then, fix the parameters of the trained Mouth2Face and start training other components of DCPTalk, and the loss functions for training Mouth2Face and other components are given respectively.

[0074] As an optimization, use the loss function to train the Mouth2Face module, and this loss function consists of a self-reconstruction loss and an intermediate command-level loss as follows:

[0075]

[0076] Among them, the self-reconstruction loss can be expressed as:

[0077]

[0078] where Y = [y1,…,y t ,…,y T ∈ R T×V×3 represents the ground truth facial movement.

[0079] Use the command-level loss to learn the set of facial expression control commands Q. The relevant formula is defined as follows:

[0080]

[0081] where λ0 is the weighting factor and sg(·) represents the stop gradient operation.

[0082] As an optimization, training other components of DCPTalk includes the steps of fixing the parameters of Mouth2Face and the CNN encoder. The loss function includes a reconstruction loss a velocity loss and a facial key point loss Among them, the reconstruction loss is used to measure the difference between the synthesized facial animation and the ground truth, including mouth movements mouth movement-induced facial animation and fine facial animation The velocity loss is used to ensure consistency on the time axis, which includes mouth movements induced facial animation and delicate facial animation Facial key point loss Obtained from the input audio. DCPTalk The loss function expression of

[0083]

[0084] where λ lmk is the weighting factor.

[0085] Reconstruction loss can be expressed by the following formula:

[0086]

[0087] where λ1, λ2, λ3 are weighting factors.

[0088] Velocity loss can be expressed by the following formula:

[0089]

[0090] where β1, β2, β3 are weighting factors.

[0091] Facial key point loss can be expressed by the following formula:

[0092]

[0093] where L = [l1,…,l t ,…,l T ∈ R T×68×3 represents the real 3D facial key points extracted by the Openface 2.0 tool.

[0094] S4: Model optimization, introducing BIWI, Multiface, and VOCASET to provide comprehensive analysis and optimization for DCPTalk; they contain paired audio-3D mesh sequences of English speech. And VOCA, MeshTalk, FaceFormer, CodeTalker, FaceDiffuse, DiffSpeaker, TalkingStyle, and SelfTalk are all compared with the method of the present invention. For some pre-trained models lacking BIWI, VOCASET, or Multiface, their official source codes are used to retrain their models. For example, MeshTalk needs to be retrained on BIWI and VOCASET.

[0095] As a preference, DCPTalk is trained on a single NVIDIA A100 GPU. The model parameters are iteratively updated using the Adam optimizer with beta1 = 0.9, beta2 = 0.999, and a learning rate of 1×10 -4 . The feature dimension d is set to 256. For facial muscle control commands, H and P are set to 32 and 128 respectively. For the loss function in all experiments of the present invention, λ0 = 1.0, λ lmk = 1×10 -5 , λ1 = 0.1, λ2 = 2.0, λ3 = 1.0, β1 = 0.3, β2 = 4.0, and β3 = 10.0.

[0096] Among them, BIWI includes audio-visual recordings and their corresponding detailed dynamic 3D facial geometries, which are collected from 14 individuals, including 8 females and 6 males, who are required to verbally express 40 English sentences. Each sentence is recorded in two scenarios: one in a neutral context and the other in an emotional context. The 3D facial geometry sequences are recorded at a frame rate of 25 frames per second, and each geometry sequence consists of 23,370 vertices. The present invention focuses on FaceFormer, CodeTalker, and SelfTalk, and only uses the subset with emotions. Specifically, the dataset is divided into different subsets: BIWI-Train, BIWI-Val, and two test subsets - BIWI-Test-A and BIWI-Test-B. These subsets are composed of 192 sequences (6 subjects × 32 sentences), 24 sequences (6 subjects × 4 sentences), 24 sequences (6 subjects × 4 sentences), and 32 sequences (8 unseen subjects × 4 sentences) respectively.

[0097] VOCASET consists of 480 pairs of audio-3D geometries from 12 subjects. Each subject contributes 40 sequences, each sequence lasting between 3 and 5 seconds and recorded at a speed of 60 frames per second. The 3D head geometry is described by 5023 vertices and 9976 faces. For a fair comparison, following FaceFormer, CodeTalker, and SelfTalk, the dataset is divided into a training subset (VOCASET-Train) containing 320 sequences (8 subjects × 40 sentences), a validation subset (VOCASET-Val) containing 40 sentences (2 subjects × 20 sentences), and a test subset (VOCASET-Test) containing 40 sentences (2 subjects × 20 sentences).

[0098] Multiface is a new multi-view high-resolution face dataset collected from 13 individuals. Each subject records 50 speech-balanced sentences at 30fps. Under uniform illumination, approximately 150 different camera views are captured per frame and used to reconstruct a 3D head mesh. The 3D head mesh consists of 6,172 vertices and 12,294 surfaces. For fair comparison, following the method of Face Diffuser, the dataset is divided into different subsets: Multiface-Train, Multiface-Val, Multiface-Test-A, and Multiface-Test-B, which are composed of 360 sequences (9 subjects × 40 sentences), 45 sequences (9 subjects × 5 sentences), 45 sequences (9 subjects × 5 sentences), and 20 sequences (4 subjects × 5 sentences), respectively.

[0099] S5: Model quantitative evaluation. According to the methods of FaceFormer, CodeTalker, and SelfTalk, the synchronization between speech content and lip movement is evaluated by calculating the Lip Vertex Error (LVE). Among them, the Lip Vertex Error (LVE) is the average of the maximum L2 error induced by the vertices in the lip region over all frames.

[0100] In addition, facial expressions directly reflect the naturalness and vividness of facial animation and play a crucial role in enhancing the interactive experience and effectiveness of the proposed method. The Upper Face Dynamic Deviation (FDD) metric can be used to measure the upper face dynamic deviations that occur in a series of animations and compare them with the ground truth.

[0101] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A voice-driven facial animation simulation method based on facial muscle linkage, characterized in that: Includes steps: S1: Construct a PPMF encoder, which consists of an audio feature extractor and a pseudo facial key point extractor, and uses a module similar to the Transformer decoder to fuse and align personalized pseudo facial key point features F L and audio feature F A ; S2: Construct a decoder based on FDCP to decode the F P Features are used to obtain facial animation. The decoder consists of Mouth Mapping, Mouth2Face, and Refine Decoder. The three stages of synthesizing mouth movements from driving signals, using mouth movements to evoke facial animation, and refining facial animation are implemented by Mouth Mapping, Mouth2Face, and Refine Decoder respectively. S3: Training the voice-driven 3D face animation framework DCPTalk, first using the loss function The Mouth2Face module is trained to establish the mapping rules between mouth movement and facial animation. Then the parameters of the trained Mouth2Face are fixed and the other components of DCPTalk are trained. The loss functions for training Mouth2Face and other components are given respectively. The loss function includes reconstruction loss Speed ​​loss and facial keypoint loss S4: Model optimization, introducing BIWI, Multiface and VOCASET to provide comprehensive analysis and optimization for DCPTalk, and integrating VOCA, MeshTalk, FaceFormer, CodeTalker, FaceDiffuse, DiffSpeaker, Both TalkingStyle and SelfTalk are compared with the method of the present invention; then DCPTalk is trained on a single NVIDIA A100 GPU; S5: Quantitative evaluation of the model, based on the methods of FaceFormer, CodeTalker and SelfTalk, evaluates the synchronization between speech content and lip movement by calculating the lip vertex error LVE.

2. The voice-driven facial animation simulation method based on facial muscle linkage according to claim 1, characterized in that: In step S1, the pseudo facial key point feature F L and audio feature F A The fusion steps include: S1a: Using a multi-head self-attention layer ALiBi with a linear bias to L to process; S1b: Use a multi-head cross attention layer to convert the F obtained by the self-attention layer A and With F L Alignment; S1c: After the feed-forward layer, the personalized pseudo multimodal feature F is obtained P =[f1 P ,…,f t P ,…,f T P ]∈R T×d .

3. The voice-driven facial animation simulation method based on facial muscle linkage according to claim 2, characterized in that: The audio feature extractor consists of an audio encoder A and an audio feature projection layer F A The composition is expressed as follows: A=Audio Encoder(x;θ) A ) F A =Audio Projection(A;ψ A )。 4. The voice-driven facial animation simulation method based on facial muscle linkage according to claim 2, characterized in that: The pseudo facial key point extractor consists of an audio encoder and an Audio2lmk decoder, which finally outputs 3D pseudo facial key points L = [l1,…,l t ,…,l T ]∈R T×68×3 ; Use the personalization factor P∈R N Personalized facial keypoint modulation; Using the personalization factor P, the pseudo facial key point L in the representation space is modulated by element-wise additive modulation, and the formula is expressed as follows: L=Audio2lmk Decoder(Α;θ L ) F L =Personalized Modulation(L,P;ψ L )。 5. The voice-driven facial animation simulation method based on facial muscle linkage according to claim 1, characterized in that: In step S2, the process of Mouth Mapping for synthesizing accurate mouth movements is expressed by the formula: Mouth2Face uses the mouth encoder to convert mouth movements Converted to a signal Q that controls facial muscle activity by using an element-wise quantization function Discrete facial muscle control signal Q to obtain facial muscle control command Activate relevant facial muscles to achieve facial animation The process is expressed by the formula: Refine Decoder is based on F P The refinement stage predicts the displacement of each vertex and then adds the vertex displacement to the vertex position of the face caused by the mouth movement in a basic way by formulating:

6. The voice-driven facial animation simulation method based on facial muscle linkage according to claim 5, characterized in that: Mouth2Face uses mouth movements separately and the facial animation Y=[y1,…,y t ,…,y T ]∈R T×V×3 Conduct training.

7. The voice-driven facial animation simulation method based on facial muscle linkage according to claim 1, characterized in that: In step S3, the process of DCPTalk training through the loss function includes Mouth2Face module training, which is expressed by the formula:

8. The voice-driven facial animation simulation method based on facial muscle linkage according to claim 1, characterized in that: In step S4, for the lack of pre-trained models of BIWI, VOCASET or Multiface, their official source codes are used to retrain their models.

9. The voice-driven facial animation simulation method based on facial muscle linkage according to claim 1, characterized in that: In step S4, the model parameters are iteratively updated using the Adam optimizer with beta1 = 0.9, beta2 = 0.999 and a learning rate of 1×10 -4 ; The feature dimension d is 256; The facial muscle control commands H and P are set to 32 and 128 respectively; The loss function, λ0=1.0, λ lmk =1×10 -5 , λ1=0.1, λ2=2.0, λ3=1.0, β1=0.3, β2=4.0 and β3=10.0.

Citation Information

Patent Citations

  • Audio-driven face animation generation method and system fused with emotion coding

    CN113378806A

  • Voice-driven face animation generation method and system

    CN115457169A