A speech-driven 3D face animation method based on multi-modal synchronous alignment

By using a multimodal synchronous alignment method, the semantic, emotional, and geometric information in voice-driven 3D facial animation is synchronized and coordinated, solving the problems of insufficient cross-modal consistency and expressive controllability in existing technologies, and improving the realism and expressiveness of the generated facial animation.

CN121074207BActive Publication Date: 2026-02-03GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511622790.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-03
Estimated Expiration
2045-11-07

AI Technical Summary

Technical Problem

Existing voice-driven 3D facial animation methods are insufficient in maintaining cross-modal consistency and expressive controllability, and have difficulty synchronizing semantic, emotional and geometric information, resulting in facial animations that lack naturalness and realism.

Method used

Employing a multimodal synchronous alignment-based approach, this method utilizes a priori-guided emotion-content decoupling strategy to extract semantic and emotional features from audio signals using pre-trained content encoders and emotion encoders. By combining the mesh refinement module and lip-reading module of the Transformer architecture, cross-modal consistency constraints and hierarchical perception reconstruction constraints are constructed to generate highly expressive 3D facial animations.

Benefits of technology

It achieves synchronous coordination of semantic content, emotional tone, and realistic geometry, improving the accuracy of lip-sync, the realism of emotional expression, and the fidelity of 3D meshes, thereby enhancing the naturalness and expressiveness of generated facial animations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074207B_ABST
    Figure CN121074207B_ABST
Patent Text Reader

Abstract

The application discloses a speech-driven 3D face animation method based on multi-modal synchronous alignment, relates to the technical field of computer graphics and deep learning, and comprises a priori guide-based emotion-content decoupling strategy, content features and emotion features of an audio signal are acquired, a grid refinement module based on a Transformer architecture is used to generate a time-sequentially coherent fusion deformation coefficient sequence by using multi-modal features, the grid refinement module based on the Transformer architecture is used to directly predict vertex positions of a face grid by using the fusion deformation coefficient, a grid sequence is obtained, and a lip reading module is used to reconstruct content features from mouth region motion, and a 3D face animation is generated by using cross-modal consistency constraints and hierarchical perception reconstruction constraints. Therefore, the speech-driven 3D face animation method based on multi-modal synchronous alignment can realize high-quality face animation which is expressive, semantically consistent and visually coherent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer graphics and deep learning, and in particular to a voice-driven 3D facial animation method based on multimodal synchronous alignment. Background Technology

[0002] Speech-generated 3D facial animation has significant applications in virtual reality, virtual digital humans, and human-computer interaction. Traditional methods are mainly based on pre-defined phoneme-viewpoint mapping rules, but they have limitations in presenting diverse speech features and the realism of emotional expression. In recent years, the development of machine learning has made it possible to more efficiently model complex temporal and emotional speech motion dynamics using data-driven neural networks, significantly advancing the field.

[0003] Currently, existing methods for voice-driven 3D facial animation can be divided into procedural and learning-based methods. Procedural methods typically rely on rule-based systems or manual mapping from phonemes to facial movements, resulting in limited flexibility and expressiveness, especially in different speech scenarios and emotional contexts. Learning-based methods, on the other hand, utilize data-driven neural architectures to directly model the complex relationship between audio and motion, thereby improving temporal dynamics and expressiveness. However, they still exhibit significant limitations, ignoring the inherent emotional variability in natural speech, leading to generated facial movements that fail to accurately match actual facial expressions in new emotional states.

[0004] To address the aforementioned issues, the TalentTalk method uses a specialized encoder to separate emotional tones from speech content, thereby achieving facial expression synthesis. However, its lip synchronization remains suboptimal due to the lack of clear consistency between visual clarity and linguistic semantics. In contrast, the SelfTalk method emphasizes semantic consistency between lip movements and speech, but still fails to account for the influence of emotional expression and intensity on lip movements. Furthermore, many methods generate 3D facial meshes by directly decoding and fusing deformation coefficients using parametric facial models (such as FLAME), but this approach has limited representativeness and restricts geometric accuracy.

[0005] In summary, existing methods often struggle to maintain cross-modal consistency and controllability of expression due to entanglement in audio representations and insufficient geometric precision. Therefore, there is an urgent need for a unified framework that can clearly define the semantic, emotional, and geometric information of synchronized speech for expressing speech-driven 3D facial animation, thereby improving the realism, naturalness, and accuracy of emotional expression in the animation. Summary of the Invention

[0006] The purpose of this invention is to provide a voice-driven 3D facial animation method based on multimodal synchronous alignment. By introducing a unified framework for voice-driven 3D facial animation, semantic content, audio emotion, and geometric realism are synchronously coordinated to effectively achieve multi-factor alignment, thereby solving the problem of lack of naturalness in facial animations generated by existing methods and improving the realism and expressiveness of generated facial animations.

[0007] To achieve the above objectives, this invention provides a voice-driven 3D facial animation method based on multimodal synchronous alignment, comprising the following steps:

[0008] S1. Based on a priori-guided sentiment-content decoupling strategy, a pre-trained content encoder and sentiment encoder are used as priors to extract content features and sentiment features of audio signals through independent fine-tuning.

[0009] S2. Based on content features, emotional features, emotional level and character features, a temporally coherent sequence of fusion deformation coefficients is generated through a mesh refinement module based on the Transformer architecture, and the fusion deformation coefficients at each time step exist in the FLAME parameter space.

[0010] S3. The mesh refinement module based on the Transformer architecture directly predicts the vertex position of the facial mesh by using the fusion deformation coefficient to obtain the mesh sequence, and then uses the lip reading module to reconstruct the content features from the movement of the lip region.

[0011] S4. Construct cross-modal consistency constraints and hierarchical perception reconstruction constraints to generate 3D facial animation;

[0012] Cross-modal consistency constraints include content-lip consistency loss, sentiment-content unpacking loss, and sentiment-content classification loss; hierarchical perception reconstruction constraints include coefficient loss and vertex loss.

[0013] Furthermore, S1 includes resampling the audio signal to extract audio features and normalizing it using linear interpolation.

[0014] Furthermore, in S2, the mesh refinement module based on the Transformer architecture includes a face mesh generation module, which applies masked multi-head self-attention and a 1024-dimensional feedforward layer, as well as a single-layer Transformer decoder to generate vertex features for each face mesh; and a weight prediction module, which directly uses a fully connected layer to linearly map the vertex features of the face network to vertex numbers.

[0015] Furthermore, in S3, the lip reading module consists of six stacked Transformer encoders used to capture lip temporal dynamics and semantic cues.

[0016] Furthermore, in S4, the content-lip consistency loss ensures spatial alignment by minimizing the Euclidean distance between the reconstructed content features and the original features, as follows:

[0017] ;

[0018] In the formula, For content-lip consistency loss, For the content features to be reconstructed, Original features;

[0019] The emotion-content decoupling loss is achieved by using a cross-reconstruction-based emotion-content decoupling loss function, as follows:

[0020] ;

[0021] In the formula, Unraveling the loss of emotion-content , Each corresponds to a different audio signal marker. Indicates semantic content and emotional characteristics The input audio signal constituted , To provide a realistic animation representation of the corresponding audio signal, For content encoders, For emotion encoder, For Transformer decoders;

[0022] The sentiment-content classification loss, used to guide the outputs of the sentiment encoder and content encoder, is as follows:

[0023] ;

[0024] In the formula, For sentiment-content classification loss, For connection time series classification loss term, The cross-loss term for sentiment features, The content features neutral emotional characteristics. The emotional characteristics of audio;

[0025] Coefficient loss ensures that the generated coefficient sequence matches the actual ground conditions, as follows:

[0026] ;

[0027] In the formula, For coefficient loss, for The predicted fusion deformation coefficient at any given time; For the first The weights of each basic deformation; It is the first The fusion deformation coefficient of the basic deformation; The number of basic deformations, with a value of 52;

[0028] Vertex loss is used for the mesh vertices generated based on ground reality supervision, as follows:

[0029] ;

[0030] In the formula, For vertex loss, For the first The weight of each facial vertex, For the predicted first facial vertex coordinates For the first The true coordinates of each facial vertex This represents the number of vertices in the facial mesh.

[0031] Therefore, the present invention employs the above-mentioned voice-driven 3D facial animation method based on multimodal synchronous alignment, which has the following technical effects:

[0032] (1) This invention proposes a new framework for voice-driven 3D facial animation based on multimodal synchronous alignment, which can synchronously coordinate semantic content, emotional tone and realistic geometric human mesh, and effectively achieve multi-factor alignment in a unified process. Compared with the prior art, this invention shows significant advantages in the accuracy of lip-sync, the authenticity of emotional expression and the fidelity of 3D mesh. This verifies the effectiveness and universality of the method of this invention, and makes it easier to promote and apply and meet actual needs.

[0033] (2) This invention introduces a priori guided emotion-content decoupling strategy and develops a lip-reading module based on neutral speech. Through a cross-modal semantic alignment mechanism, it strengthens the semantic consistency between audio and facial movements, achieves robust lip-sync, significantly improves semantic-temporal consistency, and ensures accurate lip-reading matching even under emotional interference.

[0034] (3) This invention ensures the geometric accuracy and motion coherence of the generated results through the consistency constraints of coefficients and vertices, providing a flexible architectural foundation for tasks such as emotion-controlled avatar generation and voice-aware facial editing.

[0035] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0036] Figure 1This is a flowchart of a voice-driven 3D facial animation method based on multimodal synchronous alignment. Detailed Implementation

[0037] The present invention will be explained in more detail through the following embodiments. The purpose of disclosing the present invention is to protect all changes and modifications within the scope of the present invention. The present invention is not limited to the following embodiments.

[0038] Example 1

[0039] like Figure 1 As shown, this invention provides a voice-driven 3D facial animation method based on multimodal synchronous alignment, including audio feature decoupling, lip synchronization constraint and mesh optimization, which can synchronize semantic content, emotional expression and geometric details to generate highly expressive and realistic facial animations.

[0040] This invention proposes a priori-guided sentiment-content decoupling strategy, which utilizes structured audio prior knowledge to decompose speech into separate semantic and sentiment features, as detailed below:

[0041] Using a pre-trained Wav2Vec 2.0[3] encoder and As an audio prior, it is independently fine-tuned to extract semantic and emotional representations from the audio signal:

[0042] ;

[0043] In the formula, For semantic representation, For the expression of emotion, , These are the content encoder and the emotion encoder, respectively. The audio signal is used. Specifically, the audio signal is resampled to a 16 kHz waveform, the corresponding audio features are extracted, and then normalized using linear interpolation.

[0044] Given two different audio and By exchanging their semantic and emotional features, new hybrid deformation coefficients can be flexibly synthesized. This enables feature recombination for building 3D facial animations.

[0045] in,

[0046] ;

[0047] In the formula, superscript , Tags representing different audio frequencies Indicates emotional level, Indicates the speaker's style; This represents a Transformer decoder with 832 hidden units and a 4-head Transformer layer, which merges unwrapped features to regress 52-dimensional hybrid shape coefficients; and, the generated... It resides in the structured parameter space of FLAME. This untangled structure enables flexible recombination across modes.

[0048] To enhance consistency and verify separation, this embodiment employs a cross-reconstruction loss similar to selftalk. This constraint mechanism enables the decoder to accurately reconstruct facial expression parameters from the mixed representations.

[0049] For the extracted content and sentiment representation, the fused deformation coefficient sequence is first predicted by decoder D. Although each fused deformation coefficient vector In the FLAME parameter space, and using the standard FLAME generation formula Direct mapping to mesh vertices is a method that is inherently limited by the fixed expressive power of the FLAME model and cannot capture high-frequency geometric changes. Therefore, this embodiment introduces a mesh refinement module based on the Transformer architecture, which uses mesh vertex features as input to directly predict high-fidelity vertex positions.

[0050] .

[0051] This indirect decoding strategy allows the network to transcend the geometric constraints of FLAME, generating temporally coherent and more accurate 3D facial mesh sequences that better reflect subtle speech dynamics and emotional changes. Specifically, the facial mesh generation module first applies masked multi-head self-attention and a 1024-dimensional feedforward layer, along with a single-layer Transformer decoder (512 hidden size, 4 heads), to generate features for each facial vertex. Then, the weight prediction module linearly maps the per-vertex features to vertex numbers using another fully connected layer, thereby achieving the goal of directly predicting facial vertex positions from facial vertex features.

[0052] In the predicted grid sequence Based on this, this embodiment uses a lip-reading module. LipReader Reconstructing semantic content specifically from lip region movements:

[0053] ;

[0054] In the formula, For the content characteristics of audio, To extract a predefined binary mask of the lip region from a complete facial mesh, LipReader uses a six-stacked Transformer encoder to capture the temporal dynamics of lip movements to decode latent content features, while taking into account the significant modulation of lip articulation by emotional expression. This embodiment also restricts LipReader training to samples labeled with neutral emotions to ensure the extraction of neutral content features. It reflects clear vocal movements, unaffected by emotional interference.

[0055] To facilitate cross-modal alignment, this embodiment uses a shared text decoder, implemented by fully connected layers, which can map 1024-dimensional latent features to a 32-dimensional vocabulary space, thus mapping audio content features. Neutral content features predicted by lips They are jointly decoded into text, and consistency is explicitly enforced between the predictions of both, as follows:

[0056] ;

[0057] This emotion control training strategy improves semantic consistency between modalities and makes visual-voice cues more reliable supervision.

[0058] The total loss function constructed in this invention for:

[0059] ;

[0060] In the formula, For content-lip consistency loss, ensure semantic consistency between lip movements and speech content; Unlocking the loss of emotion-content, clearly separating emotion and content in the potential space; , These are coefficient loss and vertex loss, respectively, which together ensure accurate and physically reasonable mesh reconstruction at the parameter and vertex levels; To provide auxiliary supervision for sentiment-content classification loss, thereby improving the accuracy of semantic and sentiment differentiation; , which are weighting coefficients used to control the relative importance of each loss in the total loss.

[0061] Content-lip consistency loss, sentiment-content unpacking loss, and sentiment-content classification loss are cross-modal consistency constraints, while coefficient loss and vertex loss are hierarchical perception reconstruction constraints, as detailed below:

[0062] Content-lip consistency loss is minimized by reconstructing content features. and original features The Euclidean distance between them ensures alignment in space:

[0063] .

[0064] The sentiment-content decoupling loss adopts a sentiment-content decoupling loss function based on cross-reconstruction:

[0065] );

[0066] In the formula, Indicates semantic content and emotional characteristics The input audio signal constituted and These correspond to real animation representations, and decoder D fuses content features. With emotional characteristics This is used to reconstruct the corresponding animation. The loss function ensures both the expressiveness and semantic accuracy of the animation while maintaining independent control over emotion and content.

[0067] Considering the challenge of clearly identifying the separability of the sentiment latent space and content latent space during the deentanglement process, this embodiment introduces a sentiment-content classification loss to guide the sentiment encoder. and content encoder Output:

[0068] ;

[0069] The first constraint The Connected Temporal Classification (CTC) method is used, and the predicted neutral sentiment content features are processed through a shared TextDecoder. Align with the corresponding text sequence; second constraint Then, the cross-entropy loss function is applied to analyze sentiment features. Extract accurate sentiment category predictions from it.

[0070] The coefficient loss ensures that the generated coefficient sequence matches the actual ground conditions.

[0071] ;

[0072] In the formula, It is the first The fusion deformation coefficient of the basic deformation, for The predicted fusion deformation coefficient at any given time; For the first The weights of each basic deformation are determined by the weight prediction module. It was predicted; The number of basic deformations is set to 52 in this embodiment. Among them, the weight prediction module... The fusion deformation coefficient related features of the connections are projected from 832 dimensions to 52 dimensions using fully connected layers.

[0073] For missing vertices, the generated mesh vertices will be based on ground-based monitoring:

[0074] ;

[0075] In the formula, Vertex-specific weight prediction module The predicted first The weight of each facial vertex, For the predicted first facial vertex coordinates For the first The true coordinates of each facial vertex The number of grid vertices. Vertex-specific weight prediction module. The per-vertex features generated by the Transformer decoder are linearly mapped to vertex numbers using a fully connected layer.

[0076] By employing the aforementioned cross-modal monitoring and hierarchical perception reconstruction strategies, we can enhance temporal coherence and expressive realism, thereby improving the effectiveness of the method in lip-syncing, emotional expression, and grid fidelity.

[0077] Example 2

[0078] This invention utilizes a voice-driven 3D facial animation method based on multimodal synchronous alignment to extract facial expression and head posture information from video frames and predict character-driven frame coefficients, generating dynamic character animations with realism and emotional expressiveness. Specifically, it includes:

[0079] (1) Input data and initialization:

[0080] By initiating multi-source heterogeneous data input, the system first receives an RGB video stream (1920×1080 resolution @30fps) and a synchronized 16-bit PCM audio stream (sampling rate 44.1kHz). The video frames are processed in real-time by MediaPipe Face Mesh to detect 468 3D facial feature points, while the audio signal is converted into 80-dimensional Mel-frequency spectral features through short-time Fourier transform. All input data undergoes normalization before entering the processing pipeline, mapping pixel values ​​to the [-1,1] range, and scaling the audio amplitude to a -3dBFS reference.

[0081] (2) Facial animation generation framework predicts blending shape coefficients (fusion deformation coefficients):

[0082] To simulate facial animation deformation, the system employs a transformer architecture to predict updated blending shape coefficients based on the input audio features. Specifically, multimodal feature fusion is performed on speech content features, emotional features, emotional level, and character style. The fused multimodal features are then transformed into... This enables the generation of temporal continuity coefficients.

[0083] (3) Mesh mapping:

[0084] The mesh refinement module based on the Transformer architecture updates the blending shape coefficients. As input, the vertex positions of the facial mesh are directly predicted. Meanwhile, the training of the lip reader is restricted to samples labeled with neutral emotions to ensure that the extracted features reflect clear articulation movements without being affected by emotions.

[0085] (4) Cross-modal consistency constraints:

[0086] To maintain consistency across multimodal features, this embodiment introduces cross-modal consistency constraints. Specifically, this includes aligning the content features predicted from the lip reader. With original features Furthermore, it utilizes a cross-reconstruction method to decouple the emotion and content of audio while introducing a classification loss to guide the emotion encoder. and content encoder The output of .

[0087] (5) Reconstruction constraints of hierarchical perception:

[0088] To generate accurate geometric animations, the system employs two types of constraints: coefficient constraints and vertex constraints. Coefficient constraints ensure that the generated coefficient sequence matches the actual ground conditions by aligning the generated hybrid shape coefficient sequence with real motion capture data; vertex constraints ensure natural transitions in facial expressions by minimizing the vertices of the supervised mesh generated.

[0089] (6) Training and optimization:

[0090] In this embodiment, the total loss function consists of content-lip consistency loss, sentiment-content unpacking loss, sentiment-content classification loss, coefficient loss, and vertex loss. Each loss term is weighted by parameters. , , , and Weighted algorithms are applied and optimized using the Adam optimizer, with a learning rate set to [value missing]. During training, the parameters are set to... , , , , , .

[0091] (7) Implementation details and hardware environment:

[0092] This implementation utilizes a pre-trained Wav2Vec2.0 model to resample to 16 kHz and extracts audio features (content features and emotional features) from the waveform through linear interpolation normalization. All model training is performed on a server equipped with an NVIDIA RTX A6000 graphics card to ensure computational efficiency and processing power.

[0093] (8) Dataset and Validation:

[0094] To verify the effectiveness of the method of this invention, this embodiment was tested on the 3D-ETF dataset, which was constructed from two widely used 2D audiovisual corpora: RAVDESS and HDTF. The 3D-ETF dataset provides synchronized speech audio, facial fusion deformation coefficients, and grid sequences covering a wide range of speakers and emotional expressions. The RAVDESS subset consists of high-quality recordings from 24 actors (12 males and 12 females) performing scripted speeches in eight emotional categories: neutral, calm, happy, sad, angry, fearful, disgusted, and surprised. The HDTF subset contains approximately 16 hours of high-resolution spoken facial video footage collected from YouTube, including over 300 audio clips and 10,000 unique sentences.

[0095] Through the above specific implementation methods, the present invention realizes the generation of dynamic character animations with realism and emotional expressiveness from videos, with good spatial and temporal consistency, and can be widely used in a variety of application scenarios.

[0096] Therefore, the present invention employs the above-mentioned voice-driven 3D facial animation method based on multimodal synchronous alignment, which can effectively solve the problem that current voice-driven 3D facial animation technology is difficult to maintain cross-modal consistency and expressive controllability, significantly improve the realism and expressiveness of the generated facial animation, and meet the needs of application fields such as virtual reality, games and virtual digital humans.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A voice-driven 3D facial animation method based on multimodal synchronous alignment, characterized in that, Includes the following steps: S1. Based on a priori-guided sentiment-content decoupling strategy, a pre-trained content encoder and sentiment encoder are used as priors to extract content features and sentiment features of audio signals through independent fine-tuning. S2. Based on content features, emotional features, emotional level and character features, a temporally coherent sequence of fusion deformation coefficients is generated through a mesh refinement module based on the Transformer architecture, and the fusion deformation coefficients at each time step exist in the FLAME parameter space. S3. The mesh refinement module based on the Transformer architecture directly predicts the vertex position of the facial mesh by using the fusion deformation coefficient to obtain the mesh sequence, and then uses the lip reading module to reconstruct the content features from the movement of the lip region. S4. Construct cross-modal consistency constraints and hierarchical perception reconstruction constraints to generate 3D facial animation; Cross-modal consistency constraints include content-lip consistency loss, sentiment-content unpacking loss, and sentiment-content classification loss; hierarchical perception reconstruction constraints include coefficient loss and vertex loss. In S4, the content-lip consistency loss ensures spatial alignment by minimizing the Euclidean distance between the reconstructed content features and the original features, as follows: ; In the formula, For content-lip consistency loss, For the content features to be reconstructed, Original features; The emotion-content decoupling loss is achieved by using a cross-reconstruction-based emotion-content decoupling loss function, as follows: ; In the formula, Unraveling the loss of emotion-content , Each corresponds to a different audio signal marker. Indicates semantic content and emotional characteristics The input audio signal constituted , To provide a realistic animation representation of the corresponding audio signal, For content encoders, For emotion encoder, For Transformer decoders; The sentiment-content classification loss, used to guide the outputs of the sentiment encoder and content encoder, is as follows: ; In the formula, For sentiment-content classification loss, For connection time series classification loss term, The cross-loss term for sentiment features, The content features neutral emotional characteristics. The emotional characteristics of audio; Coefficient loss ensures that the generated coefficient sequence matches the actual ground conditions, as follows: ; In the formula, For coefficient loss, for The predicted fusion deformation coefficient at any given time; For the first The weights of each basic deformation; It is the first The fusion deformation coefficient of the basic deformation; The number of basic deformations, with a value of 52; Vertex loss is used for the mesh vertices generated based on ground reality supervision, as follows: ; In the formula, For vertex loss, For the first The weight of each facial vertex, For the predicted first facial vertex coordinates For the first The true coordinates of each facial vertex This represents the number of vertices in the facial mesh.

2. The voice-driven 3D facial animation method based on multimodal synchronous alignment according to claim 1, characterized in that, S1 involves resampling the audio signal to extract audio features and normalizing it using linear interpolation.

3. The voice-driven 3D facial animation method based on multimodal synchronous alignment according to claim 1, characterized in that, In S2, the mesh refinement module based on the Transformer architecture includes a face mesh generation module, which applies masking multi-head self-attention and a 1024-dimensional feedforward layer, as well as a single-layer Transformer decoder to generate vertex features for each face mesh. The weight prediction module directly uses a fully connected layer to linearly map the vertex features of the face network to vertex numbers.

4. The voice-driven 3D facial animation method based on multimodal synchronous alignment according to claim 1, characterized in that, In S3, the lip reading module consists of six stacked Transformer encoders used to capture lip temporal dynamics and semantic cues.

Citation Information

Patent Citations

  • Voice-driven facial animation simulation method based on facial muscle linkage

    CN119579742A

  • Multi-modal driven virtual digital human face animation generation method and system

    CN120298559A