Voice-driven three-dimensional virtual image expression animation generation method
By using the bidirectional Mamba module to process audio features and static 3D head grid templates in the voice-driven three-dimensional virtual image expression animation generation technology, the problems of insufficient lip synchronization and high computational complexity in the existing technology are solved, and higher quality animation generation and lower computing power consumption are achieved.
Patent Information
- Application Number
- CN202510225498.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing speech-driven three-dimensional virtual image expression animation generation technology uses insufficient lip sound synchronization and high computational complexity when processing long speech sequences, making it difficult to meet the needs of high-quality animation generation and low computing power consumption.
Using a method based on Mamba technology, the audio features and static three-dimensional head grid templates are processed through the bidirectional Mamba module to generate a three-dimensional head grid sequence with lip sound synchronization, reducing the computational complexity and making it grow linearly as the audio length grows.
It achieves better lip synchronism in long audio inputs, reduces computational complexity, and improves the efficiency and quality of animation generation.
Smart Images

Figure CN120163908A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision, graphics, and deep learning, and specifically relates to a method for generating speech-driven three-dimensional virtual character expression animations based on Mamba technology. Background Art
[0002] Audio-driven three-dimensional face animation generation can be applied in fields such as movies, animation production, and games. Lip-syncing face animations can enhance the user experience in virtual reality and augmented reality, thereby making users more immersed. It can also be used to create virtual digital humans, such as virtual voice assistants, virtual e-commerce anchors, etc. Most existing methods are based on the CNN or Transformer architectures. CNN needs to gradually increase the receptive field layer by layer, and its effect is not good when generalizing to long speech sequences. Although Transformer can achieve good results, its computational complexity grows quadratically with the sequence length, and for long audio, a large amount of computing power resources are required. Summary of the Invention
[0003] In view of the deficiencies in the existing speech-driven three-dimensional virtual character expression animation generation technology, the present invention provides a method for generating speech-driven three-dimensional virtual character expression animations based on Mamba technology. This method can accept speech segments, static three-dimensional head mesh templates, and speaking styles as inputs, and output a lip-syncing three-dimensional head mesh sequence corresponding to the speaking style and the speech signal. When this method accepts long audio as input, the output result has better lip-syncing performance, and the computational complexity grows linearly with the audio length.
[0004] To achieve the above objectives, the present invention is implemented using the following technical solutions.
[0005] A method for generating speech-driven three-dimensional virtual character expression animations includes the following steps:
[0006] Step 1: Input a speech segment, which is converted into audio features through an audio processing module inside the network framework;
[0007] Step 2: The encoder module encodes the audio features extracted by the audio processing module, the static three-dimensional head mesh template, and the input one-hot encoded speaking style, encodes them into the latent space, and concatenates them. The concatenated feature vector sequence is denoted as where L represents the sequence length and D represents the feature vector dimension;
[0008] Step 3: Process the concatenated feature vector sequence through a bidirectional Mamba module For the feature vectors at each position, capture the context information and output the sequence of feature vectors after information exchange with the same dimension. The sequence of feature vectors after information exchange is denoted as β 1:T ; The bidirectional Mamba module exchanges the information between different position sequences in α 1:T through three independent paths: The first path is responsible for capturing information forward; the second path captures information backward by reversing the input sequence along the time dimension, and then reverses the sequence of feature vectors back along the time dimension after capture; the third path is a skip connection that directly passes the input information to the result; perform a weighted sum of the results of the first path and the second path, and then element-wise multiply the result with the result of the third path to obtain the final result. The output sequence of feature vectors of the bidirectional Mamba module structure is
[0009] Step 4: The decoder decodes the sequence of feature vectors β 1:T output by the bidirectional Mamba module to obtain the offset of each vertex position of each frame of the face facial mesh, and add it to the static three-dimensional head mesh template to obtain a three-dimensional head mesh sequence;
[0010] Step 5: Complete the mapping from audio to the three-dimensional head mesh sequence through the above steps, and use the audio-head mesh sequence dataset to perform end-to-end training on the entire network model.
[0011] Further, in step 1, the audio processing module uses the deep learning method wav2vec 2.0 to extract speech features, and adds a layer of linear interpolation after the temporal convolutional layer in the wav2vec 2.0 model to interpolate the audio features to the fps value of the corresponding mesh sequence to achieve the alignment of audio and mesh sequences for easy training.
[0012] Further, the method of step 2 is as follows: The encoder module includes three parts: an audio encoder, a mesh encoder, and a style encoder. The audio encoder is used to encode the audio features into the latent space while compressing the feature dimension; the mesh encoder encodes the input static three-dimensional head mesh template into the latent space based on the speaking style to capture facial features and is used to assist in generating a mesh sequence with a certain speaking style.
[0013] Further, in step 2, the FLAME model is used as the representation method of the static three-dimensional head mesh template.
[0014] Further, in step 2, a linear layer is used as the implementation method of each encoder to encode each part into the latent space. And replicate and expand the encoded mesh feature vectors and speaking style feature vectors along the time dimension, and then concatenate the three part feature vector sequences along the time dimension.
[0015] Furthermore, in step 4, a multi-layer perceptron is used to implement the decoder module, which outputs the offsets of each vertex in each frame relative to the static three-dimensional head mesh. This offset is added to the static three-dimensional head mesh template to obtain the complete three-dimensional head mesh sequence generated by the model.
[0016] Furthermore, in step 5, the audio-head mesh sequence dataset used for training contains paired audio and three-dimensional head mesh sequences, and includes various speaking styles.
[0017] Furthermore, in step 5, the loss function used to train the model consists of three terms: the reconstruction loss L rec , the velocity loss L vel and the lip loss L lip . The loss function for training the model is L = 1×L rec +1×L vel +0.3×L lip . BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of a method for generating speech-driven three-dimensional virtual character facial animation based on the Mamba model in an example of the present invention.
[0020] Figure 2 It is a schematic diagram of the overall network structure of the method of the present invention.
[0021] Figure 3 It is a schematic diagram of the bidirectional Mamba module structure proposed by the method of the present invention.
[0022] Figure 4 It is a diagram showing the generation results of the same word by different individuals and speaking styles of the present invention.
[0023] Figure 5 It is a comparison result of the method of the present invention and other methods in the field using two evaluation indicators, LVE (Lip Vertex Error) and FDD (Upper-Face Dynamics Deviation), on the test sets of two public datasets, VOCASET and BIWI. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To more clearly illustrate the purpose, technical solutions, and advantages of the embodiments of the present invention, the following will clearly and completely describe the technologies in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope included in the present invention.
[0025] The present invention proposes a method for generating voice-driven three-dimensional virtual character expression animations based on the Mamba model. The method includes the following steps:
[0026] Step 101: Input a voice segment. The voice segment passes through the audio processing module inside the network framework and is converted into audio features, which are presented in the form of a sequence of feature vectors.
[0027] It should be noted that the original input audio is converted into audio features through the audio processing module. In specific implementations, the audio processing module can adopt various implementation methods, such as manual extraction and using machine learning models, etc. The present invention does not limit this. Among them, various audio features can be selected for the module output, such as the time-domain features and frequency-domain features of the audio, but they should meet the actual requirements of subsequent model training and analysis and retain the key information in the voice to the greatest extent.
[0028] As a preferred method: Use the deep learning method wav2vec 2.0 as the audio processing module to extract voice features. Because the audio sampling rate in wav2vec 2.0 is 50Hz. For the convenience of training, a linear interpolation layer is added after the temporal convolutional layer in the wav2vec 2.0 model to interpolate the 50Hz audio features into the fps value (30fps or 60fps) of the corresponding sequence, realizing the alignment of audio and mesh sequences for easy training.
[0029] Step 102: The encoder module encodes the audio features, static three-dimensional head mesh template, and the input speaking style one-hot encoding extracted by the audio processing module, encodes them into the latent space, and splices them. The spliced sequence of feature vectors is represented as where L represents the sequence length and D represents the feature vector dimension.
[0030] It should be noted that the encoder consists of three parts: an audio encoder, a mesh encoder, and a style encoder. The audio encoder encodes audio features into the latent space while compressing the feature dimension; the mesh encoder encodes the input static 3D head mesh template into the latent space based on the speaking style which can be reflected to a certain extent by the facial vertex positions, capturing facial features for assisting in generating a mesh sequence with a certain speaking style. The representation of the above-mentioned input static 3D head mesh template depends on the dataset and can be represented in various ways (such as blend shape models, parametric 3D head models, etc.), and the present invention does not limit this; the style encoder encodes the one-hot encoding representing the speaking style into the latent space. The present invention does not limit the specific implementation manners of each encoder.
[0031] As a preferred manner: The representation of the static 3D head mesh template adopts the FLAME model (Faces Learned with an Articulated Model and Expressions, FLAME).
[0032] As a preferred manner: A linear layer is used as the implementation manner of each encoder to encode each part into the latent space. And the encoded mesh feature vector and speaking style feature vector are replicated and extended along the time dimension, and then the three part feature vector sequences are concatenated along the time dimension.
[0033] Step 103: Process the concatenated α through a bidirectional Mamba module 1:T , for the feature vector at each position, capture rich context information and output a feature vector sequence after information exchange with the same dimension. The feature vector sequence after information exchange is denoted as β 1:T .
[0034] It should be noted that the present invention does not limit the specific implementation manner of the bidirectional Mamba module. It should be designed based on the Mamba structure and can enable the feature vector at a certain position to perceive the context in both the front and back directions.
[0035] As a preferred manner: The present invention example gives a design manner for implementing the bidirectional Mamba module. The feature vector sequence after the encoder can represent all the information input by each part. As Figure 3 shown, the bidirectional Mamba module exchanges α through three independent paths 1:TThe information between sequences at different positions. The first path is responsible for capturing information in the forward direction; the second path captures information in the reverse direction by reversing the input sequence along the time dimension, and then reverses the sequence of feature vectors back along the time dimension after capture; the third path is a skip connection that directly passes the input information to the result. The example of the present invention designs a bidirectional Mamba module based on the original Mamba block, which adds a second path relative to the original Mamba block. After obtaining the results of the three paths, the results of the first path and the second path are weighted and summed, and then multiplied element-wise with the result of the third path to obtain the final result. The output sequence of feature vectors of the bidirectional Mamba module structure is
[0036] Step 104: The decoder decodes the β output via the bidirectional Mamba module 1:T to obtain the offset of each vertex of the face facial mesh for each frame, and adds it to the static three-dimensional head mesh template to obtain a three-dimensional head mesh sequence.
[0037] It should be noted that: the role of the decoder is to map the sequence of latent space feature vectors to a sequence of head vertex displacements. The present invention does not limit the implementation manner of the decoder module, and the output sequence of vertex displacements should be able to reflect the input audio pronunciation features and speaking styles.
[0038] As a preferred method: the present invention uses a multi-layer perceptron to implement the decoder module, and outputs the offset of each vertex of each frame relative to the static three-dimensional head mesh. Adding this offset to the static template to obtain the complete three-dimensional head mesh sequence generated by the model.
[0039] Step 105: After completing the mapping from audio to three-dimensional head mesh sequence through the above steps, use the audio-head mesh sequence dataset to perform end-to-end training on the entire network model.
[0040] It should be noted that: the present invention does not limit the dataset used, but the dataset used should contain paired audio and three-dimensional head mesh sequences, and contain multiple speakers, that is, there are multiple speaking styles.
[0041] As a preferred method: the training set of the VOCASET dataset and the BIWI dataset is used to train the model. The loss function used to train the model contains three terms: the reconstruction loss L rec , the velocity loss L vel and the lip loss L lip . The calculation methods of each loss are as follows:
[0042]
[0043] where T represents the length of the sequence, V represents the number of vertices of the three-dimensional head mesh, Vlip represents the number of vertices in the lip region of the 3D head mesh, y t,v represents the true value of the position of the v-th vertex at position t in the sequence, and is the predicted value of its corresponding model. The loss function of the final trained model is L = 1×L rec +1×L vel +0.3×L lip .
[0044] The result obtained by the above method is as Figure 4 shown.
[0045] Step 106: Use the corresponding evaluation metrics to evaluate the method proposed in the present invention on the test set of the dataset. Compare it with other methods in the same field.
[0046] It should be noted that: the present invention does not limit the evaluation metrics used.
[0047] As a preferred method: Use LVE (Lip Vertex Error) and FDD (Upper-Face Dynamics Deviation) to compare the method of the present invention with other methods on the test sets of VOCASET and BIWI. Among them, the methods of VOCA, FaceFormer, CodeTalker, and SelfTalk are all methods for generating sequences of speech-driven 3D head meshes. The calculation methods of LVE and FDD are as follows:[[]]
[0048]
[0049] where V lip represents the set of vertices in the lip region, represents the predicted lip vertex position of the model, and v represents the true position of the corresponding vertex in the dataset.
[0050]
[0051] where V up represents the set of vertices in the upper half face region, represents the sequence of true positions of the vertices, represents the sequence of predicted positions of the same vertex, and dyn represents the standard deviation of the sequence.
[0052] The result of comparing with other methods using the above metrics and dataset is as Figure 5 shown, and the lower the value, the better the effect.
[0053] The above working mode is only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of the present invention.
Claims
1. A method for generating a three-dimensional virtual image expression animation driven by voice, comprising the following steps: Step 1: Input a speech clip, which is converted into audio features through the audio processing module inside the network framework; Step 2: The encoder module encodes the audio features extracted by the audio processing module, the static 3D head mesh template, and the input speech style one-hot encoding, encodes them into the latent space, and concatenates them. The concatenated feature vector sequence is represented as Where L represents the sequence length and D represents the feature vector dimension; Step 3: Process the concatenated feature vector sequence through the bidirectional Mamba module For the feature vector of each position, the context information is captured and the feature vector sequence after information exchange of the same dimension is output. The feature vector sequence after information exchange is expressed as β 1:T ; The bidirectional Mamba module exchanges α through three independent paths 1:T The first path is responsible for capturing information in the forward direction; the second path captures information in the reverse direction by reversing the input sequence along the time dimension, and then reverses the feature vector sequence back along the time dimension after capturing; the third path is a skip connection, which directly transmits the input information to the result; the results of the first and second paths are weighted and then multiplied element by element with the result of the third path to obtain the final result. The bidirectional Mamba module structure outputs the feature vector sequence as Step 4: The decoder processes the feature vector sequence β output by the bidirectional Mamba module 1:T Decode to obtain the position offset of each vertex of the face mesh of each frame, and add it to the static 3D head mesh template to obtain a 3D head mesh sequence; Step 5: Complete the mapping of audio to three-dimensional head mesh sequence through the above steps, and use the audio-head mesh sequence dataset to perform end-to-end training on the entire network model.
2. The method for generating a three-dimensional virtual image expression animation driven by voice according to claim 1, characterized in that: In step 1, the audio processing module uses the deep learning method wav2vec 2.0 to extract speech features, adds a layer of linear interpolation after the temporal convolution layer in the wav2vec2.0 model, and interpolates the audio features into the fps value of the corresponding grid sequence to achieve alignment of the audio and grid sequence for easy training.
3. The method for generating a three-dimensional virtual image expression animation driven by voice according to claim 1, characterized in that: The method of step 2 is as follows: the encoder module consists of three parts: an audio encoder, a grid encoder and a style encoder. The audio encoder is used to encode audio features into a latent space while compressing feature dimensions. The grid encoder encodes the input static three-dimensional head grid template into a latent space based on the speaking style, captures facial features, and is used to assist in generating a grid sequence with a certain speaking style.
4. The method for generating a voice-driven three-dimensional virtual image expression animation according to claim 1, characterized in that: In step 2, the static 3D head mesh template is represented using the FLAME model.
5. The method for generating a three-dimensional virtual image expression animation driven by voice according to claim 1, characterized in that: In step 2, a linear layer is used as the implementation method of each encoder to encode each part into the latent space and copy and expand the encoded grid feature vector and speech style feature vector along the time dimension, and then concatenate the three-part feature vector sequence along the time dimension.
6. The method for generating a voice-driven three-dimensional virtual image expression animation according to claim 1, characterized in that: In step 4, a multi-layer perceptron is used to implement a decoder module, which outputs the offset of each vertex of each frame relative to the static three-dimensional head mesh. The offset is added to the static three-dimensional head mesh template to obtain a complete three-dimensional head mesh sequence generated by the model.
7. The method for generating a voice-driven three-dimensional virtual image expression animation according to claim 1, characterized in that: In step 5, the audio-head mesh sequence dataset used for training contains paired audio and three-dimensional head mesh sequences and contains a variety of speaking styles.
8. The method for generating a voice-driven three-dimensional virtual image expression animation according to claim 1, characterized in that: In step 5, the loss function used to train the model includes three items: reconstruction loss L rec , speed loss L vel and lip loss L lip , the loss function of the training model is L = 1 × L rec +1×L vel +0.3×L lip . .
Citation Information
Patent Citations
Audio-driven face animation generation method and system fused with emotion coding
CN113378806A
Voice-driven three-dimensional face dynamic simulation method
CN117197297A
Emotion-controlled three-dimensional virtual image expression animation generation method
CN117765137A
Voice-driven three-dimensional face animation generation method and device based on linear attention mechanism
CN118015154A
Robot action prediction method based on MAMBA and selective memory three-dimensional space
CN118769250A