Voice-driven three-dimensional face animation generation method and system based on feature enhancement
By combining global-local encoders and regional cross-attention mechanisms with speech and emotion features, and utilizing global temporal context information for animation generation, this technology solves the problems of slow generation speed and insufficient animation quality in existing technologies. It achieves efficient and natural 3D face animation generation, which is suitable for applications such as virtual reality and game animation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-31
AI Technical Summary
In existing voice-driven 3D face animation generation methods, autoregressive models are slow to generate and have large cumulative errors, while non-autoregressive models cannot effectively utilize global contextual information, resulting in inaccurate lip movements and unnatural expressions in the generated animations.
By employing a global-local encoder and a region cross-attention mechanism, combined with semantic and sentiment feature extraction, and utilizing global temporal context information for animation generation through spiral convolution and adaptive instance normalization modulation, and introducing SENet for feature recalibration, high-quality 3D face animation sequences are generated.
It enables the rapid generation of accurate lip-syncing, natural facial expressions, and realistic details in 3D facial animations, improving generation efficiency and the realism of the animations. It adapts to different accents and background noise, and is suitable for fields such as virtual reality and game animation.
Smart Images

Figure CN121767516A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of animation generation technology, specifically a method and system for generating 3D face animations based on feature enhancement and voice-driven methods. Background Technology
[0002] The statements in this section merely refer to the background art related to this invention and do not necessarily constitute prior art.
[0003] Voice-driven 3D facial animation generation refers to the process of using any input speech to drive a given 3D character model, generating realistic facial animations synchronized with the input speech, including lip movements, facial expressions, and head movements. This technology has wide applications in game animation, film and television production, virtual reality, and human-computer interaction.
[0004] Early voice-driven 3D facial animation primarily focused on establishing mapping rules between speech signals and visual responses. These methods typically required animators to have extensive modeling experience, and the mapping rules needed to be established multiple times for different animated characters. In recent years, deep learning-based methods have attracted widespread attention from researchers. These methods are broadly classified into two categories: autoregressive and non-autoregressive.
[0005] Among them, the autoregressive method utilizes the prior Animation sequence of frames to generate the first The animation is generated frame by frame, and the entire facial animation sequence is generated recursively. In the training and inference process, the generation of each frame of the action depends on the previous action sequence, so the generation speed is slow and there will be accumulated error as the action is generated frame by frame.
[0006] Non-autoregressive methods generate the entire animation sequence at once using the input speech and face templates, without recursive operations. These methods have a fast generation speed, but they cannot establish a global connection with the context. Summary of the Invention
[0007] This invention provides a method and system for generating 3D face animations based on feature enhancement and driven by speech. It solves the problems of slow speed and large cumulative error of autoregressive models and lack of global context utilization in existing speech-driven 3D face animation methods, and enables the rapid generation of high-quality animation sequences with accurate lip movements, natural expressions and realistic details.
[0008] The first aspect of this invention discloses a voice-driven 3D face animation generation method based on feature enhancement, comprising the following steps: Obtain the 3D face template S of the target character and the driving voice sequence A; After normalization, the face template S is used to obtain a structured template facial feature representation containing multi-scale features using a global-local encoder. The speech sequence A is driven by semantic feature extraction and sentiment feature extraction branches to obtain semantic and phonetic features. Emotional features in speech ; Structured facial feature representation Enhancement is achieved through a region cross-attention mechanism, utilizing both semantic and phonetic features. Emotional features in speech The enhanced facial features are subjected to cross-modal fusion and dynamic modulation, and the modulation results are fused to obtain fused motion features with temporal consistency and semantic integrity. ; Fusion motion characteristics The squeeze-excitation network (SENet) performs global context modeling and recalibration along the temporal dimension to generate facial motion features enhanced with global temporal context information. ; Facial movement characteristics The motion offsets of the vertices of the 3D face template are decoded by a motion decoder and then denormalized to generate the final 3D face animation sequence.
[0009] Furthermore, the global-local encoder comprises a global encoder and a local encoder configured in parallel, used to extract global large-scale features and local fine-grained detail features of the face, respectively. Both the global encoder and the local encoder are constructed based on helical convolution operators, and a structured facial feature representation is obtained by aggregating their output features. .
[0010] Furthermore, the spiral convolution operator in the global encoder and local encoder is specifically as follows: for each vertex on the face template, a length of [length missing] is pre-generated based on its topological neighborhood relationship. The spiral sequence is used to concatenate the features of each vertex on the spiral sequence and then perform feature aggregation through convolution operations.
[0011] Furthermore, structured facial feature representation Enhancement is achieved through a region cross-attention mechanism, specifically: right Perform standard convolution processing to generate a product with... M Guiding feature map of each channel Further processing via convolution generates M Dedicated convolutional kernels for each region; Using region-specific convolutional kernels respectively Convolution is performed to divide the structured facial feature representation into... M From the semantic regions, we obtain the region feature matrix; Based on regional feature matrix and The attention interaction between the components is calculated to update the enhanced facial feature representation.
[0012] Furthermore, utilizing semantic and phonetic features The enhanced facial features are then subjected to cross-modal fusion and dynamic modulation, specifically as follows: An adaptive instance normalization method is used, through The global statistics are used to modulate the distribution of enhanced facial features; as shown in the following formula: ; in, Facial features enhanced by regional attention mechanisms and These are operations for calculating the mean and standard deviation, respectively.
[0013] Furthermore, utilizing emotional features in speech The enhanced facial features are subjected to cross-modal fusion and dynamic modulation, and the modulation results are then fused. Specifically, this involves using... The enhanced facial features are scaled and offset modulated as shown in the following formula: ; in, To incorporate facial motion features that incorporate emotional information. Facial features enhanced by regional attention mechanisms and These are operations for calculating the mean and standard deviation, respectively. will utilize The results and applications of modulation The modulation results are added or weighted and fused to obtain the fused motion characteristics.
[0014] Furthermore, the fused motion features are processed by the Squeeze-Excitement Network (SENet) for global context modeling and recalibration along the temporal dimension, generating facial motion features enhanced with global temporal context information. Specifically: Global average pooling is performed on the fused motion features along the time dimension to squeeze out the time channel descriptors; The time channel descriptor is passed sequentially through a dimensionality-reduced fully connected layer, an activation function layer, and a dimensionality-restored fully connected layer to learn the adaptive weights for each time channel; The adaptive weights are multiplied channel-by-channel with the fused motion features to generate facial motion features enhanced with global temporal context information. .
[0015] A second aspect of the present invention discloses a voice-driven 3D face animation generation system based on feature enhancement, comprising: The task objective module is configured to: obtain the 3D face template of the target character. With driving speech sequences ; The preprocessing module is configured as: face template After normalization, a structured template facial feature representation containing multi-scale features is obtained using a global-local encoder. ; drive speech sequences Semantic and phonetic features are obtained through semantic feature extraction and sentiment feature extraction branches, respectively. Emotional features in speech ; The feature enhancement module is configured as: structured facial feature representation. Enhancement is achieved through a region cross-attention mechanism, utilizing both semantic and phonetic features. Emotional features in speech The enhanced facial features are subjected to cross-modal fusion and dynamic modulation, and the modulation results are fused to obtain fused motion features with temporal consistency and semantic integrity. The feature enhancement module is also configured to: pass the fused motion features through a squeeze-excitation network (SENet) for global context modeling and recalibration along the temporal dimension, generating facial motion features enhanced based on global temporal context information. ; The decoding output module is configured to: facial motion features The motion offsets of the vertices of the 3D face template are decoded by a motion decoder and then denormalized to generate the final 3D face animation sequence.
[0016] A third aspect of the present invention discloses a computer program product including computer-readable instructions, which, when executed on an electronic device, enable the electronic device to implement the above-described feature-enhanced voice-driven 3D face animation generation method.
[0017] A fourth aspect of the present invention discloses an electronic device, including at least one processor and a memory connected to the processor, the memory being used to store a computer program; the processor being used to execute the computer program, enabling the electronic device to implement the above-described feature-enhanced voice-driven three-dimensional face animation generation method.
[0018] Compared with existing technologies, one or more of the above technical solutions have the following beneficial effects: 1. Through the collaborative design of a global-local encoder and a region cross-attention mechanism, the large-scale structure and local details of the face are captured simultaneously at the feature level. Furthermore, the motion relationships between various semantic regions of the face are explicitly modeled, resulting in accurately synchronized lip movements, vivid and rich expressions, and natural coordination among facial features in the generated animation. Simultaneously, by utilizing emotional feature branches and independent modulation, the animation faithfully reflects the emotional nuances of the speech, significantly enhancing its appeal and realism.
[0019] 2. A non-autoregressive paradigm is adopted to generate the complete sequence in one step, avoiding the error accumulation problem of traditional autoregressive methods and significantly improving generation efficiency. An adaptive instance normalization (AdaIN) modulation mechanism based on global context is used to achieve accurate alignment of speech and facial features at the global scale. Subsequently, a temporal SENet module is used to recalibrate the fused features, further enhancing the discriminative power and temporal consistency of the features.
[0020] 3. Using neutral face templates as the driving force, and through normalization and general motion pattern learning, the trained model can be effectively generalized to previously unseen 3D face characters. Robust feature extraction based on a large-scale pre-trained speech model enables the system to adapt to input speech with different accents, speaking speeds, and background noise, demonstrating broad practical application value. Attached Figure Description
[0021] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0022] Figure 1 A schematic diagram illustrating a feature-enhanced voice-driven 3D face animation generation process provided for one or more embodiments of the present invention; Figure 2 A schematic diagram of the model architecture during animation generation provided for one or more embodiments of the present invention; Figure 3 A schematic diagram of a global-local encoder architecture provided in one or more embodiments of the present invention; Figure 4 A schematic diagram of spiral convolution path selection provided for one or more embodiments of the present invention; Figure 5 A schematic diagram of SENet provided for one or more embodiments of the present invention; Figure 6 Qualitative comparison results of different methods provided in one or more embodiments of the present invention on the vocaset dataset; Figure 7 Qualitative comparison of different methods provided in one or more embodiments of the present invention on the BIWI dataset; Figure 8 Qualitative results of ablation experiments performed on the vocaset dataset provided by one or more embodiments of the present invention; Figure 9 Qualitative results of ablation experiments performed on the BIWI dataset, provided by one or more embodiments of the present invention. Detailed Implementation
[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0024] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0025] As introduced in the background section, early voice-driven 3D face animation primarily focused on establishing mapping rules between voice signals and visual responses. These methods typically required animators to have ample modeling experience, and the mapping rules needed to be established multiple times for different animated characters. In recent years, deep learning-based methods have received widespread attention from researchers. These methods are broadly classified into autoregressive and non-autoregressive categories. Among them, autoregressive methods utilize... Animation sequence of frames to generate the first The animation is generated frame by frame, recursively to produce the entire facial animation sequence. In this type of method, the generation of each frame's action depends on the preceding action sequence during training and inference, resulting in slow generation speed and accumulated errors as actions are generated frame by frame. Non-autoregressive methods, on the other hand, generate the entire animation sequence at once using the input speech and facial template, without recursion. These methods have a faster generation speed but cannot establish a global connection with the context.
[0026] This solution provides a speech-driven 3D face animation generation method and system based on feature enhancement, addressing the problems of slow generation speed and accumulated errors in autoregressive models, and the inability of non-autoregressive models to utilize contextual information. Using speech sequences and 3D face templates as input, it fully leverages multi-scale features of the face, interactive information between facial feature regions, and global contextual information of the speech feature sequence to mine and fuse the inherent information of the face and speech sequences, generating an animation sequence with accurate lip movements, emotion matching, natural expressions, and realistic details in a single operation.
[0027] This solution constructs a Global-Local Encoder (GLE) based on spiral convolution. This encoder can simultaneously capture and fuse large-scale basic features such as identity information of 3D faces as well as fine-grained features such as local details, significantly improving the ability to acquire 3D face features.
[0028] This scheme designs a Feature Region Cross-based Modal Interaction Module (FRCMIM). The module adopts a Transformer architecture. A Feature Region Cross-Attention mechanism is proposed, which divides facial features into several regions based on semantic consistency and calculates cross-attention at the region level. This fully captures the correlation and dependency between different facial regions, enhances information interaction between different regions, and significantly improves the quality of animation generation. In cross-modal fusion, the Adaptive Instance Normalization (AdaIN) concept is adopted, using the features of the entire speech sequence as the modulation signal to dynamically adjust facial features in one step. This allows the modulation process to fully utilize the contextual information of speech features at the global scale, significantly improving modality alignment accuracy. The Squeeze-and-Excitation Network (SENet) is introduced, which squeezes the fused features along the time dimension and learns adaptive weights for the time channels. This allows for further optimization and feature updates using global contextual information, effectively improving the accuracy of animation generation.
[0029] This approach introduces soft-DTW loss into speech-driven 3D face animation generation, which can guide the model to capture the motion trends of real animation sequences, thereby effectively improving the visual naturalness and realism of facial movements.
[0030] For the task of voice-driven 3D face animation generation, the goal of this solution is: given any neutral face template It can utilize any segment containing Frame speech sequence to drive Generate corresponding 3D facial motion sequences .
[0031] The process of this solution is as follows: Figure 1 As shown, the speech-driven 3D face animation generation method based on feature enhancement includes the following steps: Obtain the 3D face template of the target character With driving speech sequences Face template After normalization, a structured template facial feature representation containing multi-scale features is obtained using a global-local encoder. ; drive speech sequences Semantic and phonetic features are obtained through semantic feature extraction and sentiment feature extraction branches, respectively. Emotional features in speech ; Structured facial feature representation Enhancement is achieved through a region cross-attention mechanism, utilizing both semantic and phonetic features. Emotional features in speech The enhanced facial features are subjected to cross-modal fusion and dynamic modulation, and the modulation results are fused to obtain fused motion features with temporal consistency and semantic integrity. ; Fusion motion characteristics The squeeze-excitation network (SENet) performs global context modeling and recalibration along the temporal dimension to generate facial motion features enhanced with global temporal context information. ; Facial movement characteristics The motion offsets of the vertices of the 3D face template are decoded by a motion decoder and then denormalized to generate the final 3D face animation sequence.
[0032] The model architecture of this solution is as follows: Figure 2 As shown. It simultaneously receives two inputs: a 3D face template of the target character. (Template 3D Face) S ) and driving speech sequence (Driven audio) A ).
[0033] First, face template After normalization, the data is input into the Global-Local Encoder to obtain a structured template facial feature representation.
[0034] Meanwhile, speech sequence The input is fed into two independent speech feature extraction branches: HuBERT and Emotion2vec. HuBERT is used to extract the semantic and phonetic features of the speech. Emotion2Vec is responsible for extracting emotional features from speech. .
[0035] Secondly, facial features and voice features Emotional characteristics The input is fed into the Feature Region Cross-based Modal Interaction Module (FRCMIM).
[0036] In this module, a region attention mechanism is first used to enhance facial features, resulting in an enhanced facial feature representation. ; then calculate With speech features The statistical distribution parameters, mean and variance, were determined; then, the AdaIN method was used to calculate the values of these parameters. and right Cross-modal fusion and dynamic modulation are performed, and the two modulation results are fused to generate motion feature representations with temporal consistency and semantic integrity. The fused motion features are then passed through the SENet module to model global temporal context information. By learning channel-level attention weights, the fused motion features are adaptively recalibrated to generate enhanced motion feature representations with strong discriminative power.
[0037] Finally, the enhanced motion features are decoded into motion offsets of template vertices by a motion decoder, and the final 3D face motion sequence is generated through denormalization. .
[0038] Accurate extraction of latent speech features is crucial in speech-driven 3D face animation generation. This scheme constructs a two-branch speech feature extraction structure, used to extract semantic and phonetic features and emotional features from the original input speech sequence, respectively. This extraction method effectively decouples the linguistic and emotional information contained in the speech, generating a comprehensive feature representation that is both rich in content and discriminative in emotion.
[0039] This scheme uses a pre-trained Hubert model to extract semantic and phonetic information from speech sequences. The Hubert model extracts latent representations from unlabeled audio data using a self-supervised approach. A linear projection layer is then added after Hubert. The weights of the linear projection layer are randomly initialized, mapping the Hubert output to a feature space of the same dimension as the template face features. This process is represented as follows: ; in, For the input speech sequence, For semantic and phonetic feature extraction models, For linear projection layers, The semantic and phonetic features obtained.
[0040] This scheme introduces the Emotion2vec method to extract emotional features from speech sequences. A linear alignment layer is added after Emotion2vec, with its weights also randomly initialized, to align the output dimension of Emotion2vec with the dimension of the template face features. This process is represented as follows: ; in, For the input speech sequence, For emotion embedding extractor, For linear alignment layers, The extracted emotional features.
[0041] About global-local encoders.
[0042] The accuracy of facial feature extraction from face templates directly affects the realism and richness of detail in subsequent voice-driven facial animations. Existing methods often have limited effectiveness in lip-syncing and facial expression vividness, mainly due to poor feature space representation or modeling capabilities of the models.
[0043] To address this issue, this solution designs a global-local encoder based on spiral convolution. This allows for complementary learning of large-scale global facial features by the global encoder and fine-grained local facial details by the local encoder, effectively capturing facial identity information, personality details, and other feature information at different scales. Spiral convolution fully considers the topological relationships between facial vertices. By normalizing and sorting the vertex neighborhood into a specific spiral sequence, unordered graph data can be transformed into a structured representation. This enables the network to learn unique feature representations at different locations within the receptive field, significantly enriching the model's ability to express facial expressions and lip movements.
[0044] The global-local encoder employs a dual-branch spiral convolutional encoder structure, designed to simultaneously capture both large-scale and fine-grained detail features of the face. For example... Figure 3 As shown, the global-local encoder consists of a global encoder. Local Encoder And the Global-LocalInformationAggregation Module. constitute.
[0045] The global encoder consists of several cascaded spiral downsampling modules and a fusion layer (MLP layer). Each downsampling module encapsulates a spiral convolutional layer for feature extraction and a pooling layer for reducing spatial resolution. After multi-level feature extraction and dimensionality reduction, the resulting feature maps are flattened and aggregated through the fusion layer to output the final large-scale facial feature representation.
[0046] The local encoder employs the same architecture as the global encoder, but with a smaller receptive field designed to capture high-frequency, fine-grained geometric details. The feature vectors output from the two encoders are concatenated and then aggregated via a global-local information aggregation module composed of fully connected layers to obtain the facial features. The process is represented as follows: ; in, For human face templates, and These are the global encoder and the local encoder, respectively. This is a global-local information aggregation module. This solution will include a global encoder. The output channel dimension of each layer is set to [64, 64, 64, 64]; to maintain the sensitivity of local features, the local encoder is... The number of channels in each layer is set to [4, 8, 16, 32].
[0047] Since the global encoder and the local encoder have the same network architecture, this embodiment will introduce their core calculation process in a unified manner.
[0048] Before inputting a given face template into the global-local encoder, a regularization operation is performed on it to extract facial motion information and remove identity features. Specifically, the mean of the training set sequence is subtracted from the input face template and divided by the standard deviation.
[0049] To accommodate the efficient spiral convolution operations in the global-local encoder, the originally unordered local neighborhood structure of each vertex on the regularized face template needs to be mapped to a standard ordered 1D sequence. Therefore, a spiral queue needs to be predefined. Figure 4 As shown, the length of the spiral sequence is fixed at... For each vertex on the face template, it is considered the center vertex, and the search proceeds outward from this point. Vertices in the outer neighborhoods (ring 1, ring 2, etc.) are sequentially concatenated with the center vertex, ultimately generating a string of length [length missing]. A spiral sequence. This process generates a length of [length missing] for each vertex in the mesh. A spiral queue. To visualize this process, we introduce... and The definition is as follows: ; ; ; in, Represents the central vertex. Initialize it to the center vertex itself, and then, based on the adjacency relationship and step size with the center vertex, set the center vertex's... Ring neighborhood is represented as ,For example Represents the relationship with the central vertex The neighborhood with a distance step size of 1, and so on. Represents the set of vertices The neighborhood selection operator is used to select a set. The neighboring vertices of all vertices in the equation. This represents the historical path of the spiral, that is, from... arrive The union of all vertices. This represents the set difference operation.
[0050] From the above formula, we can conclude that Only includes with Adjacent and not included The new vertex in the spiral. This constraint ensures that the spiral expands strictly outward, avoiding the repeated selection of vertices.
[0051] Based on the pre-computed spiral sequence, a spiral convolution operator is defined to achieve feature aggregation. Specifically, the first... Vertices in each downsampling module The feature update process is defined as follows: ; in, Represents vertices In the The output of each downsampling module. Represents vertices The length is spiral sequence, This indicates a cascade operation, which follows a spiral sequence, combining the features of all vertices in the neighborhood of the previous layer. Concatenate them into a long vector. Indicates the first The convolutional functions in each downsampling module are used to fuse and map the concatenated features, thereby effectively encoding short-range correlation features within the neighborhood structure. This represents a mesh simplification algorithm that downsamples the convolutional mesh. After multiple downsampling calculations, the resulting small-scale mesh is flattened and then fused through a fully connected fusion layer.
[0052] Regarding cross-modal modulation modules based on feature region intersections.
[0053] To address the shortcomings of existing voice-driven 3D face animation methods, which often neglect the connections between different facial feature regions, struggle to fully capture the global contextual information of the speech signal, and suffer from insufficient alignment accuracy between speech and face data, leading to inadequate mining, fusion, and utilization of multimodal feature information, this solution designs a cross-modal modulation module (FRCMIM) based on feature region intersection in the proposed model. This module aims to enhance feature interaction between different facial regions and fully utilize contextual information to calibrate and update facial features, thereby achieving deep cross-modal fusion and enhancement of facial and speech features and effectively improving the quality of animation generation.
[0054] First, to model the motion correlations between different facial regions, a feature region cross-attention mechanism is proposed in the FRCMIM module. By dividing the facial features obtained from the global-local encoder into several independent semantic regions and calculating cross-attention at the region level, the network can fully capture and model the correlations and dependencies between different facial regions. This mechanism significantly enhances the information interaction between facial regions, effectively improving the coordination and vividness of animation generation.
[0055] Secondly, in the design of the cross-modal fusion strategy, the AdaIN algorithm is introduced to modulate facial features of the speech sequence in one step. This design can fully utilize the contextual information of speech features at the global scale, improving the alignment accuracy between speech modalities and facial modalities.
[0056] Finally, to further enhance the model's temporal modeling capabilities, a SENet module is introduced. This module compresses the fused multimodal features along the time dimension and adaptively learns the weights of the temporal channels using activation operations. This process enables the network to calibrate and update features using global contextual information, thereby further optimizing the accuracy of animation generation.
[0057] To dynamically model the semantic relationships between different regions of a face, this scheme introduces a dynamic region partitioning convolution method and calculates the cross-attention between regions based on the partitioning results, so as to achieve fine-grained regional feature interaction.
[0058] Specifically, firstly, the facial feature map from the global-local encoder... Apply a standard Convolution to generate with Guiding feature map of each channel This guiding feature is used to instruct the filter generation module, which processes the input features. Perform convolution operation and output a set of data. Region-specific convolution kernels Each convolutional kernel operates only on its corresponding local region.
[0059] Subsequently, these non-shared convolutional kernels are used to process the input features. Perform region convolution operations to obtain the feature matrices of the divided regions: ; in Indicates the first The feature vectors of each region.
[0060] To further capture the semantic dependencies between different regions, this scheme proposes Cross-Region Attention to compute the interaction relationships between regional features, thereby achieving comprehensive facial feature recognition. The updating and enhancement process can be represented as: ; in, This indicates facial features enhanced through the regional attention mechanism. This represents the scaling factor for the feature dimension, used to stabilize the gradient.
[0061] To effectively fuse facial and speech features, this scheme introduces the AdaIN concept to achieve cross-modal feature modulation and global information alignment. Specifically, it integrates facial features enhanced by the region attention mechanism. and speech feature sequences extracted from the Hubert branch. First, calculate the global statistics for both in terms of feature dimensions: mean. with standard deviation To obtain the global representation of facial features in each frame: the mean of facial features. with standard deviation and the mean of speech features and standard deviation Then, similar to AdaIN, feature modulation and fusion are performed between speech features and facial features. In the first stage, speech feature sequences are utilized. Modulate facial features The distribution yields the modulation results of the first stage: ; In the second stage, to incorporate emotional information, the emotional features extracted by the Emotion2vec branch are further integrated. Similarly, utilizing emotional characteristics Facial features Modulation is performed to obtain the modulation result of the second stage: ; in, Facial movement features that incorporate emotional information.
[0062] Finally, the two modulation results are fused to obtain a cross-modal facial motion feature representation, as follows: ; To further enhance the model's temporal modeling capabilities and fully utilize global temporal context information to further calibrate and enhance cross-modal facial motion features, this scheme introduces the SENet module and proposes a strategy of squeezing along the temporal dimension and adaptively learning temporal channel weights using activation operations.
[0063] like Figure 5 As shown, the SENet module analyzes facial motion features that incorporate speech information. The operation is divided into two parts: one is the squeezing operation of global time information, and the other is the incentive operation of learning channel adaptive weights.
[0064] The compression operation captures global temporal information by compressing the temporal dimension of facial motion features. Specifically, it compresses the input facial motion features along the temporal dimension. Perform adaptive average pooling, reducing its dimension from Reduce to ,in For batch size, The length of this sequence of facial motion features, for The feature dimension. This operation, by performing average pooling on the features along the time dimension and calculating the mean of each feature channel over all time steps, yields a feature representation that captures the global temporal context, mathematically expressed as follows: ; in, b This represents the sample currently being calculated. t For the current time step, c This is the current feature channel.
[0065] The purpose of the activation operation is to dynamically adjust the importance of each feature channel by learning a set of adaptive weights. This process is achieved through a two-layer fully connected network with a bottleneck structure.
[0066] First, the compressed feature representation Dimensionality reduction is achieved through a reduction layer, as follows: ; in, To reduce the weights of the layers, for The feature dimension, i.e., the initial dimension. For the reduced target dimension (here, , (Dimensional reduction rate) This is a bias term.
[0067] Then, a nonlinear transformation is introduced through a nonlinear function to obtain the activated features. : ; Next, the activated features After an expansion layer, the dimensions are restored to their initial dimensions: ; in, These are the weights for the extended layer.
[0068] Finally, the application The activation function generates the channel weight matrix. The weights range from [0, 1], as shown below: ; Use the channel weights described above to adjust the original input. This allows us to obtain facial motion features enhanced with global temporal context information. , means as follows: ; in, This indicates element-wise multiplication.
[0069] The enhanced facial motion features obtained from the cross-modal modulation module based on feature region intersection are input into a linear motion decoder and mapped to motion offsets of vertices of a 3D face template. Then, by performing an inverse normalization operation on the motion offsets, the final 3D face animation sequence is generated.
[0070] The process is represented as follows: ; in, It is a linear motion decoder. and These are the standard deviation and mean of the samples in the face training set, respectively. For human face templates.
[0071] The loss function of this model consists of four parts: soft-DTW loss, reconstruction loss, velocity loss, and KL divergence loss. Among them, the soft-DTW loss is introduced into the speech-driven 3D face animation generation. This loss calculates the soft alignment path between the generated action sequence and the real action sequence by the model, guiding the model to capture the motion trend of the real action sequence, thereby effectively improving the visual naturalness of facial movements.
[0072] This approach introduces soft-DTW loss to constrain the real action sequence and the model-generated predicted action sequence. By aligning them in time, it aims to make the predicted action trajectory as consistent as possible with the real overall action trajectory.
[0073] Assume the model generates the predicted action sequence as follows The actual action sequence is First, calculate the Euclidean distance between all frame pairs of the two sequences and construct the distance matrix. ,in, .
[0074] Then, the cost matrix is constructed using the distance matrix. and with , The cost matrix is initialized using this as an initialization condition.
[0075] Subsequently, recursive calculation As shown in the following formula: ; in, The soft minimum is defined as: ; This represents the smoothness parameter; the larger the value, the smoother the alignment path.
[0076] Based on the above definition, the final form of the soft-DTW loss is: ; The smaller the loss value, the higher the temporal dynamic fit between the predicted action trajectory generated by the model and the real action trajectory, indicating that the two have achieved a better global alignment effect.
[0077] This approach uses reconstruction loss to constrain the predicted and actual values, namely: ; in, Indicates the first The vertex coordinates of a real human face in a frame. Indicates the generated first The vertex coordinates of a face in a frame. Indicates the length of the action sequence.
[0078] This approach introduces velocity loss to constrain the motion transition between two frames in the predicted motion sequence, making it as similar as possible to the transition between corresponding two frames in the actual motion sequence, thereby achieving a smoother animation effect. ; By calculating the KL divergence loss, the distribution of speech features can be made closer to that of facial features, reducing the variability in the feature space and enabling smoother integration of cross-modal information during feature fusion.
[0079] ; in, The distribution of facial features, This represents the distribution of speech features.
[0080] In summary, the total loss function of this model is: ; in, The weights of each loss term.
[0081] experiment.
[0082] Datasets: This approach uses three datasets widely used in voice-driven 3D face animation tasks: VOCASET, BIWI, and Meshtalk datasets.
[0083] The VOCASET dataset contains 480 3D facial sequences from 12 subjects. These sequences were recorded at 60fps by a scanning device while the subjects spoke according to a script. Each sequence lasted 3-4 seconds, and each frame of the face mesh contained 5023 vertices. This approach uses 8 subjects as the training set, 2 subjects as the validation set, and 2 subjects as the test set.
[0084] Unlike VOCASET, a subset of samples in the BIWI dataset contains emotional information in their expressions. The BIWI dataset contains 14 subjects, each reading 40 sentences. Each sentence is read twice: once without any emotion, and once with the emotion matching the sentence. Each frame of the recorded face mesh contains 23,370 vertices. For this dataset, our approach removes samples with neutral expressions and uses only those containing emotional information. The training set contains 192 sentences from 6 subjects; the validation set contains 24 sentences from the same 6 subjects as the training set; the test set is divided into two subsets: BIWI-Test-A contains 24 sentences from 6 subjects, and BIWI-Test-B contains 32 sentences from 8 subjects. The subjects in BIWI-Test-A are the same as those in the training set, but the sentences they read are different. The subjects in BIWI-Test-B are not present in the training set. BIWI-Test-A is used for quantitative evaluation and analysis, while BIWI-Test-B is used for qualitative analysis.
[0085] This approach uses a publicly available subset of the Meshtalk dataset for experiments, comprising 13 subjects, each recording 50 sentences at 30fps. Each frame's facial geometry is represented using a 3D mesh with 6172 vertices. The training-validation-test set partitioning is as follows: 40 sentences (360 sequences) from 9 subjects are used for training; another 5 sentences (45 sequences) from these 9 subjects are used for validation; the last five sentences from these 9 subjects are assigned to Meshtalk-Test-A for quantitative evaluation; and the remaining 4 subjects and their 5 sentences not present in the training, validation, or test set A are assigned to Meshtalk-Test-B for qualitative analysis.
[0086] For the HuBERT model, this scheme adopts the HuBERT-large model with 24 transformer layers. For the global-local encoder, the spiral length is... Set to 9, mesh simplification algorithm The QEM algorithm was used. During training, the Adam optimizer was employed with a learning rate set to 1e-4. The channel reduction rate was set to 0.7. The weights in the loss function were... Set to 1e-5, Set to 10, Set to 1e-2. A total of 120 rounds were trained, with the batch size set to 1. In the first ten rounds, the HuberT model was frozen so that it did not participate in parameter updates.
[0087] This scheme uses five indicators to quantitatively evaluate the method, namely: 1) : Represents the average vertex error of mouth movements; 2) : Represents the average maximum Euclidean distance between the lip region vertices of all corresponding frames between the 3D face motion sequence generated by the model and the real motion sequence; 3) : Indicates the average error of the upper half of the face vertices; 4) : Indicates the average generation error of the entire face; 5) The algorithm calculates an optimal matching normalization path between a generated face sequence and a real face sequence in the time dimension, and measures the similarity between the two sequences by calculating the error between corresponding frames.
[0088] Comparative experiment.
[0089] This approach was compared with three other methods: FaceFormer, CodeTalker, and TalkingStyle. To ensure fairness in the evaluation, this approach employed a testing method tailored to the input requirements of each method. Specifically, FaceFormer, CodeTalker, and TalkingStyle use speech and an identity label as input, thus requiring speech from the test set and identity labels from the training set for animation generation. For this approach, 3D facial motion sequences are synthesized using speech from the test set and neutral face templates from the subjects.
[0090] Table 1: Performance comparison of different methods on the BIWI-TEST-A test set
[0091] Table 1 shows the performance comparison of our proposed solution with the three advanced methods mentioned above on the BIWI-TEST-A test set. It is evident that our proposed model exhibits the best overall performance (optimal in the first four metrics, and second best in the last metric). Furthermore, our proposed model also demonstrates a significant advantage in generation speed. The autoregressive methods Faceformer, CodeTalker, and TalkingStyle require 0.1 seconds, 0.004 seconds, and 0.002 seconds to generate one frame of animation, respectively, while our proposed solution generates one frame in just 0.0003 seconds. This is because our proposed solution is a non-autoregressive method, capable of generating multiple animation frames simultaneously. Therefore, this proposed solution is highly suitable for applications requiring rapid generation, such as virtual anchors, live streaming, and conferences.
[0092] This solution was qualitatively compared with Faceformer, CodeTalker, and TalkingStyle methods. For fairness, this solution used the same voice, character template, and character style as input when rendering the results generated by Faceformer, CodeTalker, and TalkingStyle. However, when generating animations using this solution, since the model does not require character style as input, it only used the same character template and voice as the aforementioned models.
[0093] Figure 6 and Figure 7 The images show face animation sequences generated by different methods on the vocaset and BIWI datasets. Figure 6 The first and second rows are FaceTalk_170731_00024_TA subjects, and the third and fourth rows are FaceTalk_170809_00138_TA subjects. Figure 7 The first row represents subjects F1, the second row represents subjects F6, the third row represents subjects F7, and the fourth row represents subjects M2. Figure 6 and Figure 7 In the image, the first column contains the words spoken by the subject in the animated sequence, with the highlighted parts indicating the pronunciation during the capture of the facial animation frame. The second column contains the actual animation frames, and columns three through six contain animation frames generated using different methods. Figure 6 and Figure 7 It can be seen that this scheme produces more vivid mouth movements when pronouncing the voiceless consonants / k / , / s / , and / t / , as well as the voiced consonants / d / and / v / . When pronouncing / k / , / s / , and / t / , the animation generated by this scheme shows more obvious mouth opening movements, while when pronouncing the voiced consonant / v / , the pursing and closing of the lips are more obvious.
[0094] Ablation experiments were conducted on this scheme to verify the effect of the design of each module and the loss function on the gain of the model. The experimental results are shown in Table 2.
[0095] Table 2: Comparison of quantitative results of ablation experiments performed on the BIWI-TEST-A test set
[0096] Table 2 presents the quantitative analysis results of the ablation experiments conducted on the BIWI-TEST-A test set. The first row shows the removal of the region attention module; it can be seen that removing this module reduces the accuracy of all metrics because it prevents the construction of contextual relationships between facial regions. The second row shows the removal of the Senet module; it can be seen that removing the Senet module reduces the accuracy of all metrics because the model cannot dynamically adjust the temporal weights of facial features during training, i.e., it cannot utilize contextual information to enhance facial features. The third to fifth rows show the removal of velocity loss, KL divergence loss, and soft-DTW loss, respectively. It can be seen that removing these three losses separately reduces the accuracy of all metrics to varying degrees. This is because removing different loss functions leads to abrupt changes in the movement of facial vertices, resulting in a lack of smooth transitions between adjacent frames in the generated sequence, or poor fusion of facial features with speech features.
[0097] Figure 8 and Figure 9 Qualitative results of ablation experiments on the vocaset and BIWI datasets are presented respectively. It is evident that, compared to the complete model of this scheme, the quality of the animation sequences generated by the model after removing different modules or loss functions is significantly reduced. This demonstrates the necessity and superiority of the modular component and loss function design of this scheme.
[0098] This solution addresses the issues of slow generation speed and accumulated errors in autoregressive models, as well as the inability of non-autoregressive models to fully utilize contextual information. A newly constructed global-local encoder based on helical convolution significantly improves the model's ability to capture facial features at different scales, greatly enriching its ability to express facial expressions and lip-sync details. A newly designed cross-modal modulation module based on feature region intersection fully utilizes the interaction information between facial feature regions and the global contextual information of speech feature sequences, achieving enhanced and deep fusion of speech and facial features, effectively improving the accuracy of animation generation. The newly introduced soft-DTW loss further enhances the realism and naturalness of the generated facial movements. Experiments show that the facial animations generated by this solution outperform existing methods in terms of lip-sync synchronization, accurate expression of emotion, vividness of expression, realism of detail, and generation speed.
[0099] Correspondingly, a feature-enhanced speech-driven 3D face animation generation system includes: The task objective module is configured to: obtain the 3D face template of the target character. With driving speech sequences ; The preprocessing module is configured as: face template After normalization, a structured template facial feature representation containing multi-scale features is obtained using a global-local encoder. ; drive speech sequences Semantic and phonetic features are obtained through semantic feature extraction and sentiment feature extraction branches, respectively. Emotional features in speech ; The feature enhancement module is configured as: structured facial feature representation. Enhancement is achieved through a region cross-attention mechanism, utilizing both semantic and phonetic features. Emotional features in speech The enhanced facial features are subjected to cross-modal fusion and dynamic modulation, and the modulation results are fused to obtain fused motion features with temporal consistency and semantic integrity. The feature enhancement module is also configured to: perform global context modeling and recalibration along the time dimension using a squeeze-excitation network on the fused motion features to generate facial motion features enhanced based on global temporal context information. ; The decoding output module is configured to: facial motion features The motion offsets of the vertices of the 3D face template are decoded by a motion decoder and then denormalized to generate the final 3D face animation sequence.
[0100] Correspondingly, a computer program product includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the aforementioned feature-enhanced voice-driven 3D face animation generation method.
[0101] Accordingly, an electronic device includes at least one processor and a memory connected to the processor, the memory being used to store computer programs; the processor is used to execute the computer programs, enabling the electronic device to implement the above-described feature-enhanced voice-driven 3D face animation generation method.
[0102] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for speech-driven 3D facial animation generation based on feature enhancement, characterized in that, Comprising the following steps: Acquiring a three-dimensional face template of a target role with driving speech sequences ; Face template After normalization, a structured template face feature representation containing multi-scale features is obtained using a global-local encoder ; driving speech sequence Semantic and prosodic features are obtained through a semantic feature extraction branch and a prosodic feature extraction branch, respectively and emotional features in speech ; Structured facial feature representation Enhanced by regional cross-attention mechanism, respectively using semantic and phonetic features and emotional features in speech , cross-modal fusion and dynamic modulation are performed on the enhanced facial features, and the modulation results are fused to obtain fusion motion features with temporal consistency and semantic integrity ; Fused motion features The extruded-excitation network performs global context modeling and re-scaling along the time dimension, and generates facial motion features enhanced based on global temporal context information ; Facial motion features The motion offsets of the three-dimensional face template vertices are decoded by the motion decoder, and after the inverse normalization operation, the final three-dimensional face animation sequence is generated.
2. The feature enhancement based speech-driven 3D face animation generation method of claim 1, wherein, The global-local encoder comprises a parallel global encoder and a local encoder, used to extract large-scale global features and fine-grained local details of the face, respectively. Both the global and local encoders are constructed based on helical convolution operators, and a structured facial feature representation is obtained by aggregating their output features. .
3. The feature enhancement based speech-driven 3D face animation generation method of claim 1, wherein, The spiral convolution operator in the global encoder and the local encoder, specifically: for each vertex on the face template, a spiral sequence with a length of is generated in advance based on the topological neighborhood relationship thereof, features of each vertex on the spiral sequence are spliced, and feature aggregation is achieved through convolution operation.
4. The feature enhancement based speech-driven 3D face animation generation method of claim 1, wherein, Structured facial feature representation Enhanced by a region cross-attention mechanism, specifically: right Perform standard convolution processing to generate a product with... M Guiding feature map of each channel Further processing via convolution generates M Dedicated convolutional kernels for each region; Using region-specific convolution kernels performing convolution processing, dividing the structured facial feature representation into M semantic regions to obtain a region feature matrix; Based on the attention interaction calculation between the regional feature matrix and The enhanced facial feature representation is obtained by updating.
5. The feature enhancement based speech-driven 3D face animation generation method of claim 1, wherein, Utilizing semantic and phonetic features The enhanced facial features are cross-modally fused and dynamically modulated, specifically: Adaptive instance normalization method is adopted to modulate the distribution of enhanced facial features by global statistics of ; as shown in the following formula: ; wherein, is the facial feature enhanced by the region attention mechanism, and are operations for calculating the mean and standard deviation, respectively.
6. The feature enhancement based speech-driven 3D face animation generation method of claim 1, wherein, Utilizing affective features in speech The enhanced facial features are cross-modally fused and dynamically modulated, and the modulation results are fused, specifically: The enhanced facial features are scaled and offset modulated, as shown in the following formula: ; wherein, is a facial motion feature fused with emotion information, is a facial feature enhanced by a region attention mechanism, and are operations for calculating mean and standard deviation, respectively. The results obtained by using The results obtained by using The results obtained by using 7. The feature enhancement based speech-driven 3D face animation generation method of claim 1, wherein, The fused motion features are subjected to global context modeling and re-scaling along the time dimension by a squeeze-and-excitation network to generate facial motion features enhanced based on global time context information , specifically: Global average pooling is performed on the fused motion features along the time dimension, and a time channel descriptor is obtained by squeezing; The time channel descriptor sequentially passes through a dimension-reduced fully connected layer, an activation function layer, and a dimension-restored fully connected layer, and adaptive weights of each time channel are learned; The adaptive weight is multiplied with the fused motion feature channel by channel to generate a face motion feature enhanced based on global temporal context information .
8. A speech-driven three-dimensional face animation generation system based on feature enhancement, characterized by, Comprising; The task target module is configured to: acquire a three-dimensional face template of a target role with driving speech sequences ; The pre-processing module is configured to obtain a face template After normalization, a structured template face feature representation containing multi-scale features is obtained by using a global-local encoder ; driving speech sequence The semantic and prosodic features are obtained through the semantic feature extraction branch and the emotional feature extraction branch respectively and the emotional features in the speech ; a feature enhancement module configured to structure facial feature representation enhanced by a region cross-attention mechanism, respectively using semantic and phonetic features and emotional features in speech cross-modal fusion and dynamic modulation of the enhanced facial features, and fusion of the modulation results to obtain fusion motion features with temporal consistency and semantic integrity; The feature enhancement module is further configured to perform global context modeling and re-parameterization along a time dimension on the fused motion features through a squeeze-and-excitation network, to generate facial motion features enhanced based on global time context information ; The decoding output module is configured to: face motion features The motion offset of the three-dimensional face template vertex is decoded by the motion decoder, and the final three-dimensional face animation sequence is generated through the inverse normalization operation.
9. A computer program product, characterised in that, The computer readable instructions, when executed on an electronic device, cause the electronic device to implement the steps of the feature enhancement-based speech-driven three-dimensional face animation generation method of any one of claims 1-7.
10. An electronic device, comprising: The electronic device comprises at least one processor and a memory connected to the processor, and the memory is used to store a computer program; the processor is used to execute the computer program, so that the electronic device can implement the steps of the feature enhancement-based speech-driven three-dimensional face animation generation method of any one of claims 1-7.