A voice-driven three-dimensional face vertex sequence generation method
Patent Information
- Application Number
- CN202611307944.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-27
- Publication Date
- 2026-09-25
AI Technical Summary
然而,欧氏距离的最小化并不等同于感知质量的最优化,约束几何上的逼近无法保证生成的动作序列在语言学上的可读性,且在面对噪声干扰或复杂语音时,容易产生抖动或不自然的伪影
[0008]本发明提供的方法的有益效果在于:设计采用双流架构分别提取基础声学特征与深层语义表征,并通过跨模态注意力与自适应门控机制实现声学节奏与语义内容的动态融合,从而增强特征表示的语义清晰度与抗噪性。利用预训练的教师网络,设计多维语义感知优化目标从语言学可读性的角度约束学生网络模型的学习过程以最大限度地减少从语音到三维动作映射过程中的语义歧义。
Smart Images

Figure CN122820928A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of computer vision and computer graphics, and in particular to a voice-driven method for generating 3D face vertex sequences. Background Technology
[0002] Speech-driven 3D facial animation aims to automatically generate realistic 3D facial movement sequences synchronized with the input speech signal. With the explosive growth of applications such as the metaverse, virtual reality (VR), live streaming of digital humans, film and television production, and low-bandwidth video calls, the market has an urgent need for high-fidelity, highly automated lip-syncing technology. This technology can significantly reduce the high cost of traditional manual keyframe animation, giving virtual characters the ability to "speak," and is a key component in achieving natural human-computer interaction.
[0003] Currently, deep learning-based speech-driven face animation methods have become mainstream. Early studies mostly used convolutional neural networks (CNNs) or recurrent neural networks (RNNs) to establish a mapping from speech acoustic features (such as MFCC and Mel-frequency spectra) to facial vertex displacements. In recent years, with the rise of the Transformer architecture, leveraging its powerful long-sequence modeling capabilities to capture the temporal dependencies between speech and facial movements has become a research hotspot. For example, a pre-trained Wav2Vec model is used to extract audio features, and facial movement sequences are generated through a self-attention mechanism.
[0004] However, existing acoustic feature-driven methods still face significant challenges in generating high-quality, highly intelligible lip-sync animations, limiting their application in sophisticated scenarios. While audio features can characterize prosodic information such as rhythm, loudness, and pitch, different phonemes may have similar audio features in actual pronunciation, or the same phoneme may correspond to completely different lip shapes in different contexts. Facial action sequences generated based on audio features often tend to produce "averaged" lip shapes, resulting in loose closure of closed-lip sounds (such as / b / , / p / , / m / ) and unclear distinction in vowel opening. Although the generated animations are rhythmically synchronized, they are difficult to recognize by lip reading. In terms of training supervision, existing methods mainly rely on geometric reconstruction loss (such as the mean square error of vertex positions, MSE) to constrain the generation results. However, minimizing the Euclidean distance is not equivalent to optimizing perceptual quality. Constraining geometric approximation cannot guarantee the linguistic readability of the generated action sequences, and they are prone to jitter or unnatural artifacts when faced with noise interference or complex speech.
[0005] Therefore, it is necessary to provide a speech-driven 3D face animation generation method that can improve the accuracy, readability, and naturalness of the generated face animation lip movements. Summary of the Invention
[0006] The purpose of this invention is to provide a voice-driven method for generating 3D face vertex sequences, which can improve the accuracy of lip movements, semantic readability, and naturalness of voice-driven 3D face animation.
[0007] The method for generating a speech-driven 3D face vertex sequence provided by this invention includes: pre-training a teacher network to construct a student network model including a feature processing module and a prediction module; the feature processing module receives a speech-driven signal, extracts hidden states from the speech-driven signal and resamples them in the time dimension to align with the target video, projects the aligned hidden states onto a feature space to obtain acoustic features, extracts semantic features from the aligned hidden states, and calculates an attention weight matrix and context features based on the acoustic and semantic features; the prediction module concatenates the acoustic and context features and generates gating coefficients through a gating network, fuses the acoustic and context features according to the gating coefficients to generate fused features, integrates the fused features with temporal information and maps them to the displacement space of 3D vertices to obtain predicted vertex offsets, and generates a predicted 3D face vertex sequence based on the predicted vertex offsets and the static template of the target face; the teacher network is used to predict the character sequence probability distribution of the speech-driven signal, and the text encoder of the teacher network is used to extract the speech... The contextual hidden state of the voice-driven signal is used as a reference semantic feature; the probability distribution of the character sequence is greedily decoded to obtain the text transcription result, and the connection temporal classification loss is measured by combining the text transcription result and the predicted 3D face vertex sequence; lip motion features are extracted based on the predicted vertex offset generated by the student network model, and the lip-semantic consistency loss between the lip motion features and the reference semantic features is measured; the attention loss entropy is measured based on the attention weight matrix; the real vertex sequence of the target face speaking is obtained, and the vertex position loss and velocity loss between the real vertex sequence and the predicted 3D face vertex sequence are measured; a total loss function constraining the training of the student network model is constructed by combining the connection temporal classification loss, lip-semantic consistency loss, attention loss entropy, vertex position loss, and velocity loss; the student network model is trained using the total loss function until the total loss function converges, and the trained student network model is used as a face animation sequence generation model to generate a 3D face vertex sequence based on the voice-driven signal.
[0008] The beneficial effects of the method provided by this invention are as follows: It employs a dual-stream architecture to extract basic acoustic features and deep semantic representations separately, and dynamically fuses acoustic rhythm and semantic content through cross-modal attention and adaptive gating mechanisms, thereby enhancing the semantic clarity and noise resistance of the feature representations. Utilizing a pre-trained teacher network, a multi-dimensional semantic perception optimization objective is designed to constrain the learning process of the student network model from a linguistic readability perspective, minimizing semantic ambiguity in the mapping process from speech to three-dimensional actions.
[0009] Extracting semantic features from aligned hidden states includes: performing multi-layer nonlinear transformation on aligned hidden states using a speech semantic processor and then performing layer normalization to obtain semantic features; the speech semantic processor consists of a multi-layer perceptron, each layer of which contains a linear layer, a ReLU activation function, and Dropout.
[0010] The attention weight matrix and context features are calculated based on acoustic and semantic features, including: projecting acoustic and semantic features onto the query and semantic features as keys and values using a pre-defined cross-modal attention projection matrix. , , , Represents the query matrix. Indicates acoustic characteristics, The query projection matrix represents cross-modal attention. Represents the key matrix, Represents semantic features, The key projection matrix represents cross-modal attention. Represents a value matrix, The value projection matrix represents cross-modal attention, where, , , The output dimension is consistent; the attention weight matrix is calculated based on the projected output: , Represents the attention weight matrix. Represents the normalized exponential function, Key matrix The transpose of the matrix, Represents the scaling factor; calculates contextual features: , Indicates contextual features.
[0011] The predicted vertex offset is obtained by integrating the fused features with temporal information and mapping them to the displacement space of the three-dimensional vertices. This process includes: inputting the fused features into a single-layer Transformer decoder and using the fused features as the target sequence and memory sequence of the decoder, and performing full-sequence temporal integration of the fused features using a self-attention mechanism; mapping the decoder output to the displacement space of the three-dimensional vertices through a fully connected projection layer to obtain the frame-by-frame predicted vertex offset, which represents the dynamic deformation of the fused features relative to the static face template.
[0012] The connection-time classification loss is measured by summing the likelihood functions of all valid CTC alignment paths that can be collapsed into text transcription results and taking the negative logarithm, specifically satisfying the following formula: ,in, This represents the connection-time classification loss function. This represents the set of valid CTC alignment paths that can be mapped to text transcription results after folding. Indicates the number of frames. This indicates that given fusion features and parameters Under the conditions, at the first Frame prediction of the first CTC legal alignment path The probability of a character in a frame. Indicates the CTC legal alignment path in the th order. The corresponding character on the frame, This indicates the fusion feature.
[0013] Extracting lip motion features based on the predicted vertex offsets generated by the student network model includes: determining the predicted vertex results based on the predicted vertex offsets generated by the student network model; flattening the predicted vertex results and projecting them onto a set dimension; aligning the projected results by interpolation according to the frame rate of the predicted 3D face vertex sequence; and performing lip reading recognition on the aligned results to obtain lip motion features.
[0014] Measuring attention loss entropy based on the attention weight matrix includes: calculating the Shannon entropy for the attention distribution of each query frame based on the attention weight matrix and taking the mean.
[0015] Obtain the true vertex sequence of the target face speaking, including: reading the true 3D facial scan of the target face speaking from a public dataset, using the target face and speaking content as indices. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a voice-driven method for generating a three-dimensional face vertex sequence, as provided in an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed following the word and its equivalents, but do not exclude other elements or objects.
[0018] To address the problems existing in the prior art, embodiments of the present invention provide a voice-driven method for generating three-dimensional face vertex sequences.
[0019] See the instruction manual appendix Figure 1 The method includes: S1: Pre-train the teacher network to build a student network model that includes a feature processing module and a prediction module.
[0020] The feature processing module receives the speech-driven signal, extracts the hidden state from the speech-driven signal and resamples it in the time dimension to align it with the target video, projects the aligned hidden state into the feature space to obtain acoustic features, extracts semantic features from the aligned hidden state, and calculates the attention weight matrix and context features based on the acoustic features and semantic features.
[0021] After the prediction module concatenates acoustic features and contextual features, it generates gating coefficients through a gating network. Based on the gating coefficients, it fuses the acoustic features and contextual features to generate fused features. After integrating the temporal information of the fused features, it maps them to the displacement space of the three-dimensional vertices to obtain the predicted vertex offsets. Based on the predicted vertex offsets and the static template of the target face, it generates the predicted three-dimensional face vertex sequence.
[0022] In one possible embodiment, the teacher network refers to the text encoder Wav2Vec2ForCTC. The teacher network freezes its parameters after pre-training and processes the received input data with the frozen parameters.
[0023] After receiving the voice-driven signal input by the operator, the student network model performs the following processing: firstly, it extracts the hidden state through a pre-trained Wav2Vec 2.0 encoder.
[0024] In this embodiment, the pre-trained Wav2Vec 2.0 encoder is pre-trained on a large-scale unlabeled speech corpus covering 53 languages using a self-supervised contrastive learning approach. By performing a contrastive prediction task on masked latent speech representations, the encoder learns general and robust speech representations without manual annotation. The encoder's low-level convolutional feature extractor is frozen (not participating in backpropagation), with only the upper-level Transformer context network participating in fine-tuning. After receiving the speech-driven signal, the encoder outputs a contextual hidden state with a frame rate of approximately 50Hz and a dimension of D=1024 through multiple layers of one-dimensional convolutions and multiple Transformers. The hidden state simultaneously contains low-level acoustic prosodic information and high-level phonemic semantic information, which are highly entangled.
[0025] After extracting the hidden states, complementary resampling and alignment are performed along the time dimension to ensure that the feature sequence length of the hidden states corresponds one-to-one with the number of frames in the target video. The specific implementation is as follows: VOCASET: Within the encoder, one-dimensional linear interpolation (mode=linear, align_corners=True) is used to resample the feature frame rate of the hidden state (approximately 50Hz) to the target video, resulting in a video frame rate of 30FPS. BIWI: Adjacent frame pairs are reassembled, downsampling the 50Hz features to 25FPS in the time dimension by factor, correspondingly concatenating the feature channel dimension from 1024 to 2048 (i.e., each target frame is composed of two adjacent frames). According to the formula... After alignment processing, the aligned hidden state is obtained. , Indicates the number of frames. This represents a one-dimensional feature interpolation function used to resample the feature tensor along the time dimension. This represents the original context hidden state without temporal resampling. This represents the target output length parameter for the interpolation operation. Indicates the number of frames sent in the target video. The interpolation algorithm uses one-dimensional linear interpolation mode. align_corners=True means that the first and last corner points (i.e., boundary frames) of the input and output tensors are aligned during the interpolation calculation to ensure that the first and last boundary information of the temporal features is completely preserved.
[0026] It should be noted that if there is a discrepancy between the sequence length of the semantic features and the acoustic features during the stage of fusing semantic features and acoustic features, the temporal resolution of the acoustic features will be used as the benchmark to perform linear interpolation alignment on the semantic features to ensure strict frame-by-frame correspondence.
[0027] A two-stream architecture is used to extract acoustic and semantic features in parallel from aligned hidden states: Acoustic features are obtained by projecting the hidden state onto the feature space through a linear mapping. These acoustic features preserve the prosody, pitch, and intensity information of the speech, serving as the query source for cross-modal attention. , Indicates acoustic characteristics, This represents the hidden layer dimension of the feature mapping within the student network model. , This represents the learnable parameters.
[0028] A speech semantic processor is designed to mine implicit factors and lexical information in hidden states. The speech semantic processor is composed of a multilayer perceptron (MLP), with each layer containing a linear layer, a ReLU activation function, and Dropout. After performing multilayer nonlinear transformations on the aligned hidden states using the speech semantic processor and then performing layer normalization, semantic features can be extracted. , Represents semantic features, Representation layer normalization, This represents a multi-layer nonlinear transformation.
[0029] A multi-head attention mechanism is employed, using acoustic features as queries and semantic features as keys and values, and projecting the acoustic and semantic features using a pre-defined cross-modal attention projection matrix: , , , Represents the query matrix. Indicates acoustic characteristics, The query projection matrix represents cross-modal attention. Represents the key matrix, Represents semantic features, The key projection matrix represents cross-modal attention. Represents a value matrix, The value projection matrix represents cross-modal attention, where, , , The output dimensions are consistent.
[0030] Calculate the attention weight matrix based on the projected output: , Represents the attention weight matrix. Represents the normalized exponential function, The key matrix represents the cross-modal attention computation. The transpose of the matrix is used to query the matrix. Perform matrix dot product calculation. Represents the scaling factor. Calculate contextual features: , Indicates contextual features.
[0031] In one possible implementation, the number of heads in the multi-head attention mechanism is 4, and each head has a dimension of... .
[0032] A multi-head attention mechanism can automatically retrieve the most relevant semantic cues from the context based on the acoustic state. In this embodiment, (Acoustic features) represent the rhythmic skeleton of the speech-driven signal. and (Semantic features) represent the content of the speech-driven signal. This is achieved through calculation... and The dot product similarity can find the best-matching contextual information in the semantic sequence for each frame of acoustic signal. For example, when acoustic features are difficult to identify due to unclear pronunciation, the attention mechanism will give higher weights to relevant phonemes based on the contextual semantics, thereby utilizing... The semantic information in the current frame is used to complete and enhance the features.
[0033] After obtaining acoustic and linguistic features, based on the different degrees of dependence of different speech segments (such as silence, vowels, and consonants) on acoustic and semantic information, an adaptive gating mechanism is designed to dynamically adjust the fusion ratio of acoustic and linguistic features: First, acoustic features are concatenated with contextual features, and then gating coefficients are generated through a gating network. : , This represents the Sigmoid activation function. Indicates feature splicing, Represents the weight matrix. This represents the bias vector. The gated network consists of a linear layer and a sigmoid activation function. The weight matrix and bias vector in the gated network layer are learnable parameters, initialized with default settings and updated during training. For example, That is, 2048×1024; That is, 1024.
[0034] Through training, the adaptive gating mechanism can learn the following strategy: when the input speech is a closed-lip consonant (such as / b / , / p / ) or a plosive, due to the extremely short transient nature of its acoustic characteristics, it will automatically amplify... The value relies more on the contextual semantics within the contextual features to guide mouth closure; however, when the input speech consists of long vowels or silent segments, the energy distribution of acoustic features is more stable and decisive, which reduces... The residual gating design preserves more of the original acoustic characteristics to maintain the natural jitter and rhythm of the animation. This residual gating design effectively avoids information bias caused by a single feature dominating the overall design.
[0035] Finally, the calculation of the fused features satisfies the following formula: , , This indicates the result of fusing acoustic features and contextual features based on gating to determine fusion weights. This refers to a feedforward network, which consists of linear mappings and activation functions, used to perform deep nonlinear spatial transformations on fused features. This indicates the fusion feature.
[0036] The design employs a Transformer-based decoding mechanism to decode fused features and generate a 3D face vertex sequence. Specifically, the fused features are input into a single-layer Transformer decoder, serving as both the target and memory sequences. A self-attention mechanism is used to perform temporal integration of the fused features across the entire sequence. The decoder output is mapped to the 3D vertex displacement space via a fully connected projection layer to obtain frame-by-frame predicted vertex offsets. These predicted vertex offsets represent the dynamic deformation of the fused features relative to the static face template. The predicted vertex offsets are then added to the static template of the target face to obtain the predicted 3D face vertex sequence. The static template of the target face refers to the static 3D face template (vertex coordinates under neutral expression) of a specific subject with a defined identity from the dataset used by the student network model; each subject's static 3D face template is unique.
[0037] S2: The teacher network is used to predict the character sequence probability distribution of the speech-driven signal, and the text encoder of the teacher network is used to extract the contextual hidden state of the speech-driven signal as a reference semantic feature.
[0038] The application of a teacher network to predict the character sequence probability distribution of a speech-driven signal refers to the frame-by-frame character vocabulary probability given by the frozen Wav2Vec2ForCTC for the speech-driven signal. The application of a teacher network to extract the contextual hidden state of a speech-driven signal using a text encoder refers to feeding the same speech-driven signal into a frozen text encoder to extract its contextual hidden state.
[0039] S3: Greedy decoding is performed on the probability distribution of the character sequence to obtain the text transcription result. The connection temporal classification loss is measured by combining the text transcription result and the predicted 3D face vertex sequence. Lip motion features are extracted based on the predicted vertex offsets generated by the student network model, and the lip-semantic consistency loss between the lip motion features and the reference semantic features is measured. The attention loss entropy is measured based on the attention weight matrix. The real vertex sequence of the target face speaking is obtained, and the vertex position loss and velocity loss between the real vertex sequence and the predicted 3D face vertex sequence are measured. The total loss function constraining the training of the student network model is constructed by combining the connection temporal classification loss, lip-semantic consistency loss, attention loss entropy, vertex position loss, and velocity loss.
[0040] The design aims to optimize multi-dimensional semantic perception by forcing the student network model to generate 3D face vertex sequences that contain machine-interpretable textual information. Specifically, this includes: To ensure the generated 3D face vertex sequences have clear linguistic meaning, a design is implemented to map the generated lip features back to the text space and calculate the Connectionist Temporal Classification (CTC) loss. An accurate supervision signal is constructed using a decode-reencode label generation strategy. Specifically, since the teacher network (text encoder) outputs unaligned text probabilities (Logits), direct use may introduce errors. Therefore, the probability distribution of the character sequence output by the teacher network is greedily decoded to obtain the most probable text string. Then, a pre-trained token segmenter is used to re-encode this text string into a standard token ID sequence, which serves as the true label for the CTC loss. This strategy eliminates the alignment ambiguity between the teacher network output and the student network model's predicted output, ensuring the strict linguistic correctness of the supervision signal. Greedy decoding refers to first obtaining the frame-by-frame token by taking the maximum probability (argmax) for each frame in the vocabulary dimension, and then folding it according to the CTC rules (merging consecutive repeated tokens and removing blanks) to obtain the most probable text string. This process is a deterministic maximum probability path decoding. The pre-trained tokenizer refers to the Wav2Vec2CTCTokenizer that is paired with the text encoder Wav2Vec2ForCTC of the teacher network. The pre-trained tokenizer uses a character-level vocabulary (including special symbols such as blank and word separators) consistent with the teacher network / student network model. Its function is to re-encode the text string obtained by greedy decoding into a standard token ID sequence as the true label of CTC loss, thereby completing the "decoding-recoding" closed loop and ensuring that the supervision label and the prediction result of the student network model are in the same symbol space.
[0041] set up For a given real text transcription, the probability distribution of the symbol sequence is... The CTC loss aims to maximize the likelihood probability of the target sequence, specifically satisfying the following formula: ,in, This represents the connection-time classification loss function. This represents the set of valid CTC alignment paths that can be mapped to text transcription results after folding. Indicates the number of frames. This indicates that given fusion features and parameters Under the conditions, at the first Frame prediction of the first CTC legal alignment path The probability of a character in a frame. Indicates the CTC legal alignment path in the th order. The corresponding character on the frame, This represents the fusion feature. The pseudo-labels obtained from the real text transcription are processed by the frozen teacher network through decoding and recoding of the speech-driven signal. They share the same character vocabulary with the student network, thereby eliminating alignment ambiguity between the teacher's output probability and the student's prediction, and ensuring the linguistic consistency of the supervision signal. This represents all learnable parameters (weights of the non-frozen part) of the student network participating in the CTC loss calculation, and minimizes them. That is to Optimization is performed so that the latent features of the generated 3D face vertex sequence contain linguistic information that can be used to reconstruct the target text.
[0042] The lip-semantic consistency loss is designed to constrain the lip motion features of the 3D face vertex sequence generated by the student network model to remain consistent with the standard text features in the latent space. The lip-semantic consistency loss is calculated according to the following formula: , This represents the lip-semantic consistency loss function. Indicates the first Lip motion features extracted from the generating vertex on the frame. Indicates the first Reference semantic features of the teacher network output on the frame. This represents calculating the square of the L2 norm between feature vectors, i.e., using the mean squared error to measure the semantic deviation between hidden layer features. Here, "generated vertices" refers to the predicted vertex offsets generated by the student network model, whose topology and vertex number are consistent with the ground truth sequence. Lip vertices are indexed from the generated vertices using a lip mask. After flattening the lip vertices, they are projected onto a set dimension (e.g., 1024-dimensional) via a linear mapping layer. The projection result is then aligned using interpolation according to the frame rate of the predicted 3D face vertex sequence. The alignment result is then recognized by a lip reading recognition Transformer to obtain the lip motion features. The speech-driven signal is input into the teacher network, and its contextual hidden state is extracted by the text encoder as a reference semantic feature. When calculating the lip-semantic consistency loss, the parameters of the teacher network are frozen throughout, which makes... It can provide a documented, objective semantic reference benchmark. The lip-semantic consistency loss function is measured by mean squared error (MSE). and The bias in the hidden space is minimized. This forces the generated lip movement features to converge towards standard textual semantic features in a high-dimensional manifold space, thereby endowing visual lip shapes with semantic attributes that can be understood by machines.
[0043] To prevent excessive dispersion of attention weights across modalities, an entropy regularization term is designed. Specifically, for the attention weight matrix, the entropy that measures the attention loss is: , The attention loss entropy function is represented by the following: Represents the cross-modal attention weight matrix The Line 1 Column element, representing the first Frame to the first The attention weights assigned to each semantic keyframe satisfy the following conditions: , This indicates the prevention of logarithmic underflow (i.e., avoiding...) The minimal normal number added. The attention loss entropy is calculated by taking the Shannon entropy of the attention distribution for each query frame based on the attention weight matrix and then minimizing it. This can sharpen the attention distribution, allowing the model to focus more on key semantic features. (Design) Intended to penalize cross-modal attention weight matrix When the distribution is too flat, the model cannot determine the uncertainty of the semantic factors corresponding to the acoustic frame. The attention distribution becomes sharp, that is, it generates high response values at key semantic correspondence points. Through training, it helps the student network model establish accurate acoustic-phoneme alignment relationship and reduce the ambiguity of generated lip movements.
[0044] The accuracy of the generated face mesh in spatial geometry is determined by setting vertex position loss constraints. Specifically, let... For the true vertex sequence, For the predicted 3D face vertex sequence, where, Indicates the number of frames. This represents the number of vertices. The vertex position loss is designed using mean squared error (MSE) to measure the Euclidean distance between the true vertex sequence and the predicted 3D face vertex sequence: , This represents the vertex position loss function. Indicates the first Frame number The predicted coordinates of each vertex. Indicates the first Frame number The true coordinates of each vertex. The true vertex sequence is derived from ground truth 3D face scans in a publicly available dataset, indexed by the target face and speech content. The true vertex sequence and the predicted 3D face vertex sequence use the same face topology, have the same number of vertices, and have the same number of frames after temporal alignment. The true vertex sequence and the predicted 3D face vertex sequence share the same static template base of the same speaker; the only difference between them is in the dynamic displacement. The purpose of designing vertex position loss constraints is to force the lip movements of the face animation obtained from the generated predicted 3D face vertex sequence to approximate the true data.
[0045] Design velocity loss constraint on the rate of change of vertex displacement between adjacent frames: , Represents the velocity loss function. This represents the predicted instantaneous velocity vector at the vertex. This represents the real instantaneous velocity vector of the vertex. The velocity loss design forces the student network model to learn realistic motion dynamics during training, smooths the temporal generation results, eliminates the jitter in the time dimension of the final generated face animation, and ensures the continuity of the action.
[0046] The total loss function for training the constrained student network model is constructed by combining connection-time classification loss, lip-semantic consistency loss, attention loss entropy, vertex position loss, and velocity loss. , Represents the total loss function. , , , , This represents the balancing weight hyperparameter for each loss term.
[0047] The principle for setting the balancing weight hyperparameters is as follows: the numerical dimensions of each loss term differ greatly—the geometric loss and velocity loss are the mean square error of the vertex coordinate space (numerically very small), the CTC loss is on the order of log probability, and the attention entropy is 0 to... The magnitude is significant. Therefore, geometry / velocity is given high weights, making them dominant in the total loss after weighting, ensuring that spatial geometric accuracy and temporal smoothness are the primary optimization objectives; semantic loss ( , , Assign it a small weight so that it can be used as a linguistic regularization / guide term to provide semantic supervision without overwhelming the geometric objective or causing training instability.
[0048] Specifically, the loss terms originate from different mathematical spaces, with significant differences in their numerical dimensions and orders of magnitude. Direct summation would lead to the model training being dominated by the loss term with the largest value, or even crash due to gradient imbalance. Therefore, the setting of the balancing weight hyperparameters must adhere to geometric dominance and semantic regularization; that is, while ensuring accurate 3D face vertex positions and smooth motion, semantic supervision is introduced as an auxiliary constraint. For example, the values of the balancing weight hyperparameters for each loss term can be set as follows: , =0.001、 =0.0001、 The weights were determined through experimental tuning and can achieve a balance between geometric accuracy, temporal coherence and semantic intelligibility.
[0049] The reason for the large differences in the settings of the balancing weight hyperparameters for each loss term is as follows: The numerical magnitudes of vertex position loss and velocity loss are too small to generate effective gradients. For example, the mean square error (MSE) of vertex coordinates is typically within a certain range. At this level, such tiny gradients are insufficient to drive parameter updates in deep networks during backpropagation. Therefore, assigning a weight of 1000 times to the corresponding loss term essentially increases the numerical order of magnitude of the geometric error. The interval is set to ensure that the generated data from the model dominates the total loss and that the generated data is geometrically strictly aligned with the real scanned data.
[0050] The CTC loss measures the negative logarithm of the classification probability, typically in the tens. Without suppression, the model might sacrifice geometric accuracy (the geometric loss is only 0.00001) to reduce this small loss, resulting in distorted facial features. Assigning a weight of 0.0001 compresses the contribution of the CTC loss to... The level transforms the speech recognition task into a weakly supervised signal, used only to guide the model to understand the pronunciation content, rather than replacing the main objective of geometric reconstruction.
[0051] The lip-semantic consistency loss calculates the distance between vertex features and text features in the latent space. Since acoustic and semantic features have different distributions, excessively large weights can lead to training oscillations. Assigning a weight of 0.001 as a regularization term gently narrows the distance between visual lip shapes and textual semantics, improving the semantic intelligibility of lip shapes without forcing perfect alignment, thus preserving generative diversity.
[0052] Attention loss entropy is the degree of concentration of attention distribution. Its value range is limited (about 0~4). Assigning a weight of 0.01 is enough to penalize "attention dispersion", prompting the model to focus on key phoneme frames, thereby improving the interpretability of the model and preventing the attention mechanism from failing (i.e., paying attention to all frames equally).
[0053] By minimizing the total loss function, the trained student network model can maximize the semantic intelligibility and naturalness of lip movements in speech-driven 3D face animation while ensuring geometric accuracy.
[0054] S4: Use the total loss function to constrain the student network model for training until the total loss function converges. Use the trained student network model as a face animation sequence generation model to generate a 3D face vertex sequence based on the speech-driven signal.
[0055] The trained student network model for generating face animation sequences can be directly used to generate 3D face vertex sequences in practical applications. That is, the speech-driven signal to be synthesized and the static face template of the target speaker are input into the face animation sequence generation model. After processing by the model, a 3D face vertex sequence can be obtained. Applying the obtained 3D face vertex sequence to drive 3D face animation will complete a speech-driven 3D face animation with accurate lip movements, readability and naturalness.
[0056] It should be noted that when applying the speech-driven 3D face vertex sequence generation method proposed in this invention, the model can be trained using publicly available datasets to acquire the ability to generate speech-driven 3D face animations with accurate lip movements and semantically readable, natural results. In practical applications, when the trained model generates a 3D face animation of the target speaker based on the speech-driven signal, the target speaker's face data used must be authorized data.
[0057] This invention proposes a speech-driven 3D face vertex sequence generation method that addresses the issues of high-level semantic loss and ambiguous lip movements in existing speech-driven 3D face animation methods. The method employs a dual-stream architecture to extract basic acoustic features and deep semantic representations separately. Furthermore, it utilizes cross-modal attention and adaptive gating mechanisms to dynamically fuse acoustic rhythm and semantic content, thereby enhancing the semantic clarity and noise resistance of the feature representations. A pre-trained text encoder is used as the teacher network, and a multi-dimensional semantic perception optimization objective is designed to constrain the learning process of the student network model from a linguistic readability perspective, minimizing semantic ambiguity in the speech-to-3D motion mapping process.
[0058] The speech-driven 3D face vertex sequence generation method of this invention solves the problem that semantic features are difficult to explicitly and decoupledly used as controllable driving signals. Through a specially designed semantic processing method, semantic components are explicitly separated from coupled features to obtain a pure high-level semantic representation. Then, through cross-modal attention and adaptive gating mechanisms, semantic features are injected directionally and on demand into acoustic features. This transforms the passive mapping of semantics, which is poorly utilized in existing technologies, into active guidance with controllable semantic injection. This results in the generation of mouth shapes that better conform to the physical laws of pronunciation on key phoneme frames such as closed-lip consonants and complex vowels, significantly improving semantic intelligibility and noise resistance.
[0059] Extensive experiments were conducted on the VOCASET and BIWI public benchmark datasets. Experimental results show that the speech-driven 3D face vertex sequence generation method of this invention can significantly improve the lip-sync accuracy, semantic readability, and naturalness of speech-driven 3D face animation. Specifically, the evaluation metrics are lip vertex error (LVE) and facial dynamic deviation (FDD), both of which are preferred when lower. The experimental results of the method of this invention compared with several existing methods are shown in the table below: The data in the table above shows that the method of the present invention outperforms methods such as SelfTalk in all three evaluation indicators, especially the FDD method, which reflects the overall dynamic distribution, with the most significant improvement (approximately 14.08%). This confirms the effectiveness of the method of the present invention in improving semantic intelligibility and temporal consistency in the generation of voice-driven 3D face animation.
[0060] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0061] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0062] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.
[0063] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. A voice-driven method for generating a three-dimensional face vertex sequence, characterized in that, The method comprises the following steps: a pre-trained teacher network is used to construct a student network model comprising a feature processing module and a prediction module; the feature processing module receives a speech-driven signal, extracts a hidden state from the speech-driven signal, and resamples the hidden state in a time dimension to align with a target video, projects the aligned hidden state to a feature space to obtain acoustic features, extracts semantic features from the aligned hidden state, and calculates an attention weight matrix and context features based on the acoustic features and the semantic features; the prediction module generates gating coefficients through a gating network after splicing the acoustic features and the context features, fuses the acoustic features and the context features according to the gating coefficients to generate fused features, integrates the timing information of the fused features, and maps the fused features to a displacement space of three-dimensional vertices to obtain predicted vertex offsets, and generates a predicted three-dimensional face vertex sequence according to the predicted vertex offsets and a static template of a target face; the teacher network is applied to predict a character sequence probability distribution of the speech-driven signal, and a text encoder of the teacher network is applied to extract a context hidden state of the speech-driven signal as a reference semantic feature; a text transcription result is obtained by performing greedy decoding on the character sequence probability distribution, and a connection time sequence classification loss is measured based on the text transcription result and the predicted three-dimensional face vertex sequence; a lip-semantic consistency loss between a lip motion feature extracted from the predicted vertex offset generated by the student network model and the reference semantic feature is measured, and an attention loss entropy is measured based on the attention weight matrix; a real vertex sequence of a target face speaking is obtained, and a vertex position loss and a speed loss between the real vertex sequence and the predicted three-dimensional face vertex sequence are measured; and a total loss function for constraining the training of the student network model is constructed based on the connection time sequence classification loss, the lip-semantic consistency loss, the attention loss entropy, the vertex position loss, and the speed loss. The student network model is trained using the total loss function until the total loss function converges, and the trained student network model is used as a face animation sequence generation model for generating a three-dimensional face vertex sequence based on a speech-driven signal.
2. The method according to claim 1, characterized in that, The semantic features are extracted from the aligned hidden state, comprising: performing multi-layer nonlinear transformation on the aligned hidden state using a speech semantic processor, and then performing layer normalization to obtain semantic features; the speech semantic processor is composed of multiple perception machines, and each perception machine comprises a linear layer, a ReLU activation function, and a Dropout.
3. The method according to claim 1, characterized in that, The attention weight matrix and the context features are calculated based on the acoustic features and the semantic features, comprising: Using the acoustic features as the query and the semantic features as the key and value, a preset cross-modal attention projection matrix is used to project the acoustic features and the semantic features: , , , Represents the query matrix. Indicates acoustic characteristics, The query projection matrix represents cross-modal attention. Represents the key matrix. Represents semantic features, The key projection matrix represents cross-modal attention. Represents a value matrix, The value projection matrix represents cross-modal attention, where, , , The output dimensions are consistent; Calculate the attention weight matrix based on the projected output: , Represents the attention weight matrix. Represents the normalized exponential function, Key matrix The transpose of the matrix, Indicates the scaling factor; Calculate contextual features: , Indicates contextual features.
4. The method according to claim 1, characterized in that, the fused features are input into a single-layer Transformer decoder, and the fused features are used as target sequences and memory sequences of the decoder, and full-sequence timing integration is performed on the fused features using a self-attention mechanism; The decoder output is mapped to the displacement space of the three-dimensional vertices through a fully connected projection layer to obtain the frame-by-frame predicted vertex offsets, which represent the dynamic deformation of the fused features relative to the static face template.
5. The method of claim 1, wherein, The connection-time classification loss is measured by summing the likelihood functions of all valid CTC alignment paths that can be folded into the text transcription result and taking the negative logarithm, specifically satisfying the following formula: ,in, This represents the connection-time classification loss function. This represents the set of valid CTC alignment paths that can be mapped to text transcription results after folding. Indicates the number of frames. This indicates that given fusion features and parameters Under the conditions, at the first Frame prediction of the first CTC legal alignment path The probability of a character in a frame. Indicates the CTC legal alignment path in the th order. The corresponding character on the frame, This indicates the fusion feature.
6. The method of claim 1, wherein, Lip motion features are extracted based on the predicted vertex offsets generated by the student network model, including: The predicted vertex result is determined based on the predicted vertex offset generated by the student network model; The predicted vertex results are flattened and projected onto a set dimension, and the projection results are interpolated and aligned according to the frame rate of the predicted 3D face vertex sequence. The alignment results are used to perform lip reading recognition to obtain lip movement features.
7. The method of claim 1, wherein, Measuring the attention loss entropy based on the attention weight matrix includes: calculating the Shannon entropy for the attention distribution of each query frame based on the attention weight matrix and taking the mean value.
8. The method of claim 1, wherein, Obtain the true vertex sequence of the target face for speech, including: Using the target face and speech content as indexes, the ground truth of 3D facial scans of the target face speaking is read from a public dataset.