A training method and device for a 3D virtual digital human lip animation generation model
By obtaining audio and text data sets and using the Transformer model to process features, real-life digital human lip animations are solved, and virtual digital human facial animations are not realistic enough, costly and technically difficult in the existing technology, and efficient and low-cost animation generation is achieved.
Patent Information
- Application Number
- CN202311559668.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-11-21
AI Technical Summary
In the prior art, virtual digital human facial animation is not realistic enough, is costly and technically difficult, especially in terms of understanding and personalized expression based on text speech.
By obtaining the audio data set, the corresponding text data set and the BlendShape parameter set, the audio features and text features are extracted, and then fused them and processed through the Transformer model, finally generating a realistic digital human lip animation.
It realizes the generation of more realistic digital human lip animations, reducing cost and technical difficulty, while improving processing speed and low latency, and is suitable for real-time application scenarios.
Smart Images

Figure CN117292031B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to a method and device for training a 3D virtual digital human lip animation generation model. Background Art
[0002] With the rapid development of computer vision technology, digital twins have been widely used in various fields. Digital twin is a technology that models and simulates objects, systems or processes in the real world in a digital way to achieve their monitoring, analysis and optimization. As a large reliance, the metaverse provides a broader space for the development of digital twins. The metaverse is a virtual digital world, which is constructed by technologies such as digital twins, virtual reality, and augmented reality, and can simulate and reproduce various scenes and objects in the real world. In the metaverse, digital twins can exist as a mirror image of the real world, interact and simulate with the real world, and play an important role in various fields.
[0003] In modern life such as social, shopping and gaming, attractive and animatable 3D characters are important entrances to the digital world. Creating a digital human requires a lot of time and experience. The creation of virtual characters includes: character prototype, modeling, generation and rendering, etc.; the driving methods of virtual characters include: manually making animations, motion capture technology and AI intelligent driving technology. Expressive 3D character facial animations are an important part of modern computer-generated movies and digital games. Currently, vision-based performance capture, that is, driving the facial expressions of digital humans by observing the actions of human actors, is an indispensable part of most production processes.
[0004] In practical applications, enabling digital humans to simulate real human emotions and behavioral details is still an industry research topic that is constantly being optimized. How to enable digital humans to have more realistic expressions and lip performances based on the understanding of text and speech is a problem that needs to be solved currently. Although the quality obtained from the capture system is steadily improving, the cost of producing high-quality facial animations is still very high.
[0005] First, the data demand is large: Training an effective voice-driven model requires a large amount of data as input, including paired voice signals and corresponding lip animations. Collecting and preparing this data is a time-consuming and laborious task.
[0006] Second, difficult in personalized expression: The existing technology mainly generates lip-sync animations of digital humans by learning the speech and mouth shape patterns of the average population. However, the mouth shapes and pronunciation habits of different people are different, which makes personalized expression very difficult. Moreover, in many modern games, the dozens of hours of dialogue spoken by game characters are too expensive for vision-based systems. Most importantly, the rendering quality is not realistic enough. Although certain progress has been made in the existing technology, the generated lip-sync animations are still not accurate or natural enough, lacking the subtle expressions and oral movements of real people. Therefore, the common practice is to only produce key animations, such as animations, using vision-based systems and relying on audio- and text-record-based systems to produce a large amount of game-related materials. However, the quality of the animations produced by such systems still needs to be improved.
[0007] Finally, emerging real-time applications in telepresence and virtual reality avatars pose additional challenges due to the lack of recordings, wide variations in user voices and physical settings, and strict latency requirements. Summary of the Invention
[0008] For this reason, the present application provides a method for training a 3D virtual digital human lip-sync animation generation model to solve the problems of insufficient realism, high cost, and great technical difficulty of virtual digital human facial animations existing in the prior art.
[0009] To achieve the above object, the present application provides the following technical solutions:
[0010] In a first aspect, a method for training a 3D virtual digital human lip-sync animation generation model includes:
[0011] Step 1: Obtain an audio data set, as well as a text data set and a BlendShape parameter set corresponding to the audio data set;
[0012] Step 2: Process the audio data set and the text data set so that the audio data set and the text data set correspond to the BlendShape parameter set;
[0013] Step 3: Extract the audio features of the processed audio data set and the text features of the text data set;
[0014] Step 4: Concatenate or merge the audio features and the text features to obtain a fused feature, and adjust and map the fused feature through a first linear layer;
[0015] Step 5: Input the adjusted and mapped fused feature into a Transformer model to obtain an enhanced semantic vector;
[0016] Step 6: Adjust and map the enhanced semantic vector through a second linear layer to obtain a final feature;
[0017] Step 7: Input the final feature into an activation function to obtain Blendshape parameters;
[0018] Step 8: Calculate the loss value between the Blendshape parameters and the original BlendShape GT parameters, and update the parameters through backpropagation. When the loss value reaches the optimum, stop training to obtain a 3D virtual digital human lip animation generation model.
[0019] Preferably, in step 2, when processing the audio data set to make the audio data set correspond to the BlendShape parameter set, it specifically includes:
[0020] Step 201: Determine whether the audio data is forced to align with the BlendShape parameters;
[0021] Step 202: If the audio data set cannot be forced to align with the BlendShape parameter set, discard the audio data;
[0022] Step 203: Determine whether the audio data and the BlendShape parameters meet the requirements of the number of frames and the number of characters;
[0023] Step 204: If the requirements of the number of frames and the number of characters are not met, discard the audio data.
[0024] Preferably, in step 3, the Bert model is used to extract the text features of the text data set.
[0025] Preferably, between step 6 and step 7, it further includes:
[0026] Pass the final feature through a Dropout layer and then input it into the activation function.
[0027] Preferably, in step 7, the activation function is tanh.
[0028] In a second aspect, a 3D virtual digital human lip animation generation model training device includes:
[0029] A data acquisition module, configured to acquire an audio data set, as well as a text data set and a BlendShape parameter set corresponding to the audio data set;
[0030] A data processing module, configured to process the audio data set and the text data set to make the audio data set and the text data set correspond to the BlendShape parameter set;
[0031] A feature extraction module, configured to extract the audio features of the processed audio data set and the text features of the text data set;
[0032] A feature fusion module, configured to splice or merge the audio feature and the text feature to obtain a fusion feature, and adjust and map the fusion feature through a first linear layer;
[0033] An enhanced semantic vector acquisition module, configured to input the adjusted and mapped fusion feature into a Transformer model to obtain an enhanced semantic vector;
[0034] A feature adjustment module, configured to adjust and map the enhanced semantic vector through a second linear layer to obtain a final feature;
[0035] A parameter acquisition module, configured to input the final feature into an activation function to obtain Blendshape parameters;
[0036] A backpropagation module, configured to calculate a loss value between the Blendshape parameters and the original BlendShape GT parameters, and update the parameters through backpropagation. When the loss value reaches the optimum, stop training to obtain a 3D virtual digital human lip animation generation model.
[0037] In a third aspect, a computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a method for training a 3D virtual digital human lip animation generation model are implemented.
[0038] In a fourth aspect, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of a method for training a 3D virtual digital human lip animation generation model are implemented.
[0039] Compared with the prior art, the present application has at least the following beneficial effects:
[0040] The present application provides a method and device for training a 3D virtual digital human lip animation generation model. By obtaining an audio data set, a corresponding text data set, and a BlendShape parameter set, and extracting the audio features of the audio data set and the text features of the text data set; fusing the audio features and the text features, and adjusting and mapping them through a first linear layer and then inputting them into a Transformer model to obtain an enhanced semantic vector; adjusting and mapping the enhanced semantic vector through a second linear layer to obtain a final feature; inputting the final feature into an activation function to obtain Blendshape parameters; calculating a loss value between the Blendshape parameters and the original BlendShape GT parameters, and updating the parameters through backpropagation. When the loss value reaches the optimal value, stop training to obtain a 3D virtual digital human lip animation generation model. The 3D virtual digital human lip animation generation model provided by the present application can generate more realistic digital human lip animations, with high efficiency and low cost. Description of the Drawings
[0041] To more intuitively illustrate the prior art and the present application, several exemplary drawings are given below. It should be understood that the specific shapes and structures shown in the drawings generally should not be regarded as limiting conditions when implementing the present application; for example, those skilled in the art are capable of making routine adjustments or further optimizations to the addition / deletion / attribution division of certain units (components), specific shapes, positional relationships, connection methods, dimensional proportional relationships, etc. based on the technical concept disclosed in the present application and the exemplary drawings.
[0042] Figure 1 It is a flowchart of a method for training a 3D virtual digital human lip animation generation model provided in Embodiment 1 of the present application;
[0043] Figure 2 It is a schematic diagram of the deep neural network structure provided in Embodiment 1 of the present application. Detailed Embodiments
[0044] The following further details the present application through specific embodiments in conjunction with the drawings.
[0045] In the description of the present application: Unless otherwise specified, "a plurality of" means two or more. The terms "first", "second", "third", etc. in the present application are intended to distinguish the objects being referred to, and do not have special significance in terms of technical connotation (for example, it should not be understood as emphasizing the importance or order, etc.). Expressions such as "including", "comprising", "having", etc. also mean "not limited to" (certain units, components, materials, steps, etc.).
[0046] Terms such as "upper", "lower", "left", "right", "middle", etc. cited in this application are usually for the convenience of intuitive understanding with reference to the attached drawings, rather than absolute limitations on the positional relationship in the actual product. Without departing from the technical concept disclosed in this application, changes in these relative positional relationships shall also be regarded as falling within the scope of the expression of this application.
[0047] Embodiment 1
[0048] Please refer to Figure 1 , this embodiment provides a method for training a 3D virtual digital human lip animation generation model, including:
[0049] Step 1: Obtain an audio data set, as well as the corresponding text data set and BlendShape parameter set for the audio data set;
[0050] Specifically, the sources of the data set are divided into two aspects: one is to screen from public data sets; the other is to independently shoot the data set through a face capture software. In this embodiment, the main source of the data set is the second one, which solves the problem of poor generalization ability caused by insufficient or single data sets.
[0051] The data set includes a training set, a validation set, and a test set. Among them, the data scale of the training set is about 120 minutes of voice data, as well as the corresponding text data and BlendShape parameters (output at 25 FPS, one second of audio corresponds to 25 groups of BlendShape parameters); the data scale of the validation set is about 20 minutes of voice data, as well as the corresponding text and BlendShape parameters; the data scale of the test set is about 5 minutes of voice data, as well as the corresponding text and BlendShape parameters.
[0052] Since the fbx animation file is output through the face capture software instead of the direct BlendShape parameter, it is necessary to convert the fbx animation file into a BlendShape parameter to facilitate subsequent model training.
[0053] Step 2: Process the audio data set and the text data set so that the audio data set and the text data set correspond to the BlendShape parameter set;
[0054] Specifically, for all the data sets after collection and screening, it is necessary to unify their formats so that the audio data and the corresponding text data correspond to the BlendShape parameters, thereby ensuring the accuracy of subsequent model training. Among them, the requirements for audio data are as follows:
[0055] Step 201: Determine whether the audio data set is forcibly aligned with the BlendShape parameter set;
[0056] Step 202: If the audio data cannot be forced to align with the BlendShape parameters, discard the audio data;
[0057] Step 203: Determine whether the audio data and the BlendShape parameters meet the requirements of the number of frames and the number of characters;
[0058] Step 204: If the requirements of the number of frames and the number of characters are not met, discard the audio data.
[0059] Step 3: Extract the audio features of the processed audio data set and the text features of the text data set;
[0060] Specifically, in this step, feature extraction technology is used to convert audio data into distinguishable and representative audio features. The audio features can capture the speech information and pronunciation details in the audio, providing accurate input for the subsequent model. The text data is parameterized by the Bert model, aiming to decouple and represent the text data through a self-designed encoder. Decoupled representation is a method of decomposing the input data during the encoding stage and removing redundant information to better learn and utilize the data features.
[0061] The Bert model is a Transformer-based language model that can perform encoding and representation learning on text sequences. It captures the context relationships and semantic information in the text sequence through multiple layers of self-attention mechanisms and feed-forward neural networks, and extracts text features. The Bert model can convert the input text encoding into a high-dimensional text feature representation. After the Bert model is built, as long as the corresponding text data is fed into the Bert model, each Transformer layer of the Bert model will output a corresponding number of hidden vectors, which are passed down layer by layer until the text features are finally output.
[0062] Step 4: Concatenate or merge the audio features and the text features to obtain a fused feature, and adjust and map the fused feature through a first linear layer;
[0063] Specifically, the audio features and the text features encoded by Bert are concatenated or merged to obtain a fused feature, and then the fused feature passes through a linear layer, which is used to adjust and map the feature dimensions to meet the input requirements of the subsequent Transformer. This step realizes the fusion of audio and text, more comprehensively expressing the association between speech input and text.
[0064] Step 5: Input the adjusted and mapped fused feature into the Transformer model to obtain an enhanced semantic vector;
[0065] Specifically, for the input fusion features, it is necessary to enhance the semantic vector representation of each word therein. Therefore, the improved Transformer takes each word as a Query respectively, weights the semantic information of all words in the fusion features, and obtains the enhanced semantic vectors of each word. To enhance the diversity of Attention, different Self-Attention modules are further used to obtain the enhanced semantic vectors of each word in the fusion features under different semantic spaces, and multiple enhanced semantic vectors of each word are linearly combined to obtain a final enhanced semantic vector with the same length as the original word vector (the enhanced semantic vector can be regarded as feature extraction for each word in the input text).
[0066] In this step, the Transformer model can effectively capture the context information and semantic associations of the input fusion features through the self-attention mechanism and the feed-forward neural network, improving the modeling ability for audio and text features.
[0067] Step 6: Adjust and map the enhanced semantic vectors through the second linear layer to obtain the final features;
[0068] Specifically, the features processed by the Transformer model go through a linear layer. This linear layer is used to adjust and map the output of the Transformer to obtain the final feature representation.
[0069] To prevent overfitting, a Dropout layer can be added after the second linear layer. Dropout is a regularization technique that can reduce the dependency relationship between neurons during the training process by randomly discarding the outputs of some neurons, improving the generalization ability of the network.
[0070] Step 7: Input the final features into the activation function to obtain the Blendshape parameters;
[0071] Specifically, the activation function is preferably tanh. After the Dropout layer, tanh is used as the activation function. The tanh function can map the predicted values to the range of [-1, 1], which is suitable for representing the Blendshape parameters and gives them a symmetric range of variation.
[0072] The output of the tanh function is the Blendshape parameter, which is used to represent facial expressions. The Blendshape parameter can control the facial deformation of the 3D virtual digital human and is used to generate realistic mouth animations.
[0073] Step 8: Calculate the loss value between the Blendshape parameters and the original BlendShape GT parameters, and update the parameters through backpropagation. When the loss value reaches the optimum, stop the training to obtain the 3D virtual digital human lip animation generation model.
[0074] In this step, calculate the loss value between the BlendShape parameters output by the tanh function and the original BlendShape GT parameters, and perform backpropagation to update the model parameters during multiple iterative trainings. When the loss value reaches the optimum, the model stops training and saves the optimum model weights, that is, the 3D virtual digital human lip animation generation model is obtained.
[0075] When performing facial parameter inference, input the target text audio to be inferred into the 3D virtual digital human lip animation generation model for automatic inference to generate the corresponding BlendShape parameters, and perform 3D rendering through Unity to obtain the lip movements of the corresponding virtual digital human. Finally, optimize the inference results. By analyzing the weighted average of the results of multiple models and running data processing combined with the dataset distribution, the final inference results are presented more smoothly and realistically.
[0076] The 3D virtual digital human lip animation generation model trained through this embodiment can generate credible and expressive 3D facial animations based on human audio. To make the results look natural, this embodiment takes into account complex and interdependent phenomena, including phoneme coarticulation, lexical stress, and the interaction between facial muscles and skin tissues, and adopts a data-driven method to train the independently developed deep neural network in an end-to-end manner (see Figure 2 ) to replicate the relevant effects observed in the training data and achieve the correspondence between text audio and BlendShape parameters.
[0077] The 3D virtual digital human lip animation generation model trained in this embodiment has a high processing speed and low latency, can achieve the voice-driven effect in real-time application scenarios, makes the generated digital human lip animation more realistic, not only reduces the technical difficulty of controlling the 3D character facial animation, but also reduces the cost.
[0078] In summary, the 3D virtual digital human lip animation generation model training method provided in this embodiment has the following advantages:
[0079] First, the independently developed deep neural network can identify the input of different audio files, solve the gap in input audio timbre caused by different users, and solve the problem of poor generalization ability caused by insufficient or single dataset.
[0080] Second, it solves the cumbersome process that requires a large amount of computer vision system resources and human resources to generate 3D virtual digital human facial animations;
[0081] Third, through the deep deep learning network structure to process the audio-text vector and the processing method of high-dimensional mapping and BlendShape parameter matching, it not only retains the significant features of the audio text, but also highly matches the BlendShape parameters;
[0082] Fourth, it has built its own audio, text and corresponding facial BlendShape parameter datasets, which are more adaptable to the overall network model, making the final inference effect more realistic.
[0083] Embodiment 2
[0084] This embodiment provides a training device for a 3D virtual digital human lip animation generation model, including:
[0085] A data acquisition module for acquiring an audio dataset and a text dataset and a BlendShape parameter set corresponding to the audio dataset;
[0086] A data processing module for processing the audio dataset and the text dataset so that the audio dataset and the text dataset correspond to the BlendShape parameter set;
[0087] A feature extraction module for extracting the audio features of the processed audio dataset and the text features of the text dataset;
[0088] A feature fusion module for splicing or combining the audio features and the text features to obtain a fusion feature, and adjusting and mapping the fusion feature through a first linear layer;
[0089] An enhanced semantic vector acquisition module for inputting the adjusted and mapped fusion feature into a Transformer model to obtain an enhanced semantic vector;
[0090] A feature adjustment module for adjusting and mapping the enhanced semantic vector through a second linear layer to obtain a final feature;
[0091] A parameter acquisition module for inputting the final feature into an activation function to obtain Blendshape parameters;
[0092] A backpropagation module for calculating the loss value between the Blendshape parameters and the original BlendShape GT parameters, and updating the parameters through backpropagation. When the loss value reaches the optimum, stop training to obtain a 3D virtual digital human lip animation generation model.
[0093] For the specific limitations of a training device for a 3D virtual digital human lip animation generation model, reference can be made to the limitations of a method for training a 3D virtual digital human lip animation generation model in the above text, which will not be elaborated here.
[0094] Embodiment III
[0095] This embodiment provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a method for training a 3D virtual digital human lip animation generation model are implemented.
[0096] Embodiment IV
[0097] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a method for training a 3D virtual digital human lip animation generation model are implemented.
[0098] The technical features of the above embodiments can be combined arbitrarily (as long as there is no contradiction in the combination of these technical features). For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described; these embodiments not explicitly written out should also be considered to be within the scope described in this specification.
[0099] In the above text, the present application has been described in a relatively specific and detailed manner through general descriptions and specific embodiments. It should be understood that based on the technical concept of the present application, several conventional adjustments or further innovations can be made to these specific embodiments; but as long as they do not deviate from the technical concept of the present application, the technical solutions obtained by these conventional adjustments or further innovations also fall within the protection scope of the claims of the present application.
Claims
1. A training method for a 3D virtual digital human lip animation generation model, characterized in that Including: Step 1: Obtain an audio dataset, as well as the corresponding text dataset and BlendShape parameter set for the audio dataset; Step 2: Process the audio dataset and the text dataset so that the audio dataset and the text dataset correspond to the BlendShape parameter set; Step 3: Extract the audio features of the processed audio dataset and the text features of the text dataset; Specifically, use feature extraction techniques to transform the audio dataset into distinguishable and representative audio features; Parameterize the text dataset through the Bert model to obtain text features; Step 4: Concatenate or merge the audio features and the text features to obtain a fused feature, and adjust and map the fused feature through a first linear layer; Specifically, concatenate or merge the audio features and the text features to obtain a fused feature, and then pass the fused feature through a linear layer, which is used to adjust and map the feature dimension; Step 5: Input the adjusted and mapped fused feature into the Transformer model to obtain an enhanced semantic vector; Specifically, the Transformer model takes each word as a Query, weights the semantic information of all words in the fused feature, obtains the enhanced semantic vector of each word, then uses different Self-Attention modules to obtain the enhanced semantic vector of each word in different semantic spaces, and linearly combines the multiple enhanced semantic vectors of each word to obtain a final enhanced semantic vector with the same length as the original word vector; Step 6: Adjust and map the enhanced semantic vector through a second linear layer to obtain a final feature; Step 7: Input the final feature into an activation function to obtain Blendshape parameters; Step 8: Calculate the loss value between the Blendshape parameters and the original BlendShape GT parameters, and update the parameters through backpropagation. When the loss value reaches the optimum, stop training to obtain a 3D virtual digital human lip animation generation model.
2. The method for training a 3D virtual digital human lip animation generation model according to claim 1, wherein In step 2, when processing the audio dataset so that the audio dataset corresponds to the BlendShape parameter set, it specifically includes: Step 201: Determine whether the audio data is forced to align with the BlendShape parameters; Step 202: If the audio dataset cannot be forced to align with the BlendShape parameter set, discard the audio data; Step 203: Determine whether the audio data and the BlendShape parameters meet the requirements of the frame number and the number of words; Step 204: If the requirements of the frame number and the number of words are not met, discard the audio data.
3. The training method of the 3D virtual digital human lip animation generation model according to claim 1, wherein Between step 6 and step 7, it further includes: Pass the final feature through a Dropout layer and then input it into the activation function.
4. The training method of the 3D virtual digital human lip animation generation model according to claim 1, wherein In step 7, the activation function is tanh.
5. A training device for a 3D virtual digital human lip animation generation model, characterized in that, Including: A data acquisition module, configured to acquire an audio data set, as well as a text data set and a BlendShape parameter set corresponding to the audio data set; A data processing module, configured to process the audio data set and the text data set so that the audio data set and the text data set correspond to the BlendShape parameter set; A feature extraction module, configured to extract audio features of the processed audio data set and text features of the text data set; specifically, using feature extraction technology to convert the audio data set into audio features with distinctiveness and representativeness; parameterizing the text data set through a Bert model to obtain text features; A feature fusion module, configured to splice or combine the audio features and the text features to obtain a fused feature, and adjust and map the fused feature through a first linear layer; specifically, splicing or combining the audio features and the text features to obtain a fused feature, and then passing the fused feature through a linear layer, where the linear layer is used to adjust and map the feature dimension; An enhanced semantic vector acquisition module, configured to input the adjusted and mapped fused feature into a Transformer model to obtain an enhanced semantic vector; specifically, the Transformer model takes each word as a Query, weights the semantic information of all words in the fused feature to obtain enhanced semantic vectors of each word, then uses different Self-Attention modules to obtain enhanced semantic vectors of each word in different semantic spaces in the fused feature, and linearly combines multiple enhanced semantic vectors of each word to obtain a final enhanced semantic vector with the same length as the original word vector; A feature adjustment module, configured to adjust and map the enhanced semantic vector through a second linear layer to obtain a final feature; A parameter acquisition module, configured to input the final feature into an activation function to obtain Blendshape parameters; A backpropagation module, configured to calculate a loss value between the Blendshape parameters and the original BlendShape GT parameters, and update the parameters through backpropagation. When the loss value reaches the optimum, stop training to obtain a 3D virtual digital human lip animation generation model.
6. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.