Facial animation generation method, device and readable storage medium for emotional expression
By training facial animation generation and facial expression classification models, and combining voice input and loss function optimization, the problem of stiff facial animation in digital humans was solved, achieving natural and emotionally rich facial expressions and improving the user interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG LAB
- Filing Date
- 2023-01-03
- Publication Date
- 2026-05-12
AI Technical Summary
The facial animations of existing digital humans are stiff and unnatural, which reduces the user experience during interaction.
By acquiring user-input speech, a trained facial expression animation generation model is used to predict the PCA coefficients of facial expression animations in 3D faces, which are then projected into facial expression animation data and finally redirected onto the target digital human. The training parameters are optimized by combining an expression classification model and a loss function to improve the naturalness and emotional consistency of facial expression animations.
It improves the naturalness and emotional expression of digital human facial expressions, enhances the user interaction experience, and reduces the difficulty of animation production.
Smart Images

Figure CN115984434B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital human intelligence technology, and in particular to a method, apparatus, and readable storage medium for generating facial animations that express emotions. Background Technology
[0002] With the expanding application of artificial intelligence technologies such as natural language processing, speech recognition, and computer vision, virtual digital human technology is also developing towards greater intelligence and diversification. Early applications of digital humans were primarily in the entertainment industry, such as film, animation, and games. Today, digital humans have been successfully applied to various sectors including banking, healthcare, education, government affairs, and telecommunications. Among these applications, the ability to express emotions and communicate interactively is fundamental to enabling digital humans to interact with the real world. However, the stiff and unnatural facial animations in related technologies degrade the user experience during interaction. Summary of the Invention
[0003] This application provides a method, apparatus, and readable storage medium for generating facial animations that express emotions. The digital facial expressions are natural, improving the user experience during interaction.
[0004] This application provides a method for generating facial animations that express emotions, comprising: acquiring user-inputted speech; inputting the speech into a trained facial animation generation model to output predicted PCA coefficients of facial animations of a 3D face; the trained facial animation generation model is trained using a speech sample set as input; projecting the predicted PCA coefficients of the facial animations into facial animation data of a 3D face; and redirecting the projected facial animation data onto a target digital human.
[0005] Furthermore, the facial animation generation method for expressing emotions also includes obtaining a trained facial animation generation model through the following method:
[0006] A speech sample set is obtained, which includes the vertex coordinates of a real 3D face mesh model, and the vertex coordinates of the real 3D face mesh model include the feature key points of the real 3D face mesh model; the speech sample set is input into an expression animation generation model to predict the PCA coefficients of the 3D face expressions corresponding to the speech sample set; the predicted PCA coefficients are converted into predicted vertex coordinates of the 3D face mesh model; the difference between the vertex coordinates of the real 3D face mesh model and the predicted vertex coordinates of the 3D face mesh model is compared to obtain a vertex distance loss function; the training parameters of the expression animation generation model are adjusted according to the vertex distance loss function until a preset termination condition is met to obtain the trained expression animation generation model.
[0007] Furthermore, the vertex coordinates of the predicted 3D face mesh model include the feature key points of the predicted 3D face mesh model; before adjusting the training parameters of the expression animation generation model according to the vertex distance loss function until a preset termination condition is met to obtain the trained expression animation generation model, the expression animation generation method further includes: comparing the difference between the feature key points of the real 3D face mesh model and the feature key points of the predicted 3D face mesh model to obtain a feature-related loss function; obtaining a total loss function based on the vertex distance loss function and the feature-related loss function; adjusting the training parameters of the expression animation generation model according to the vertex distance loss function until a preset termination condition is met to obtain the trained expression animation generation model includes: adjusting the training parameters of the expression animation generation model according to the total loss function until a preset termination condition is met to obtain the trained expression animation generation model.
[0008] Furthermore, the key feature points of the predicted 3D face mesh model include the upper and lower lip key point pairs of the predicted 3D face mesh model; comparing the difference between the key feature points of the real 3D face mesh model and the key feature points of the predicted 3D face mesh model to obtain a feature-related loss function includes: determining the mean square error or root mean square error between the distance between the upper and lower lip key point pairs of the real 3D face mesh model and the distance between the upper and lower lip key point pairs of the predicted 3D face mesh model as the lip closure loss function; obtaining the total loss function based on the vertex distance loss function and the feature-related loss function includes: obtaining a first total loss function based on the vertex distance loss function and the lip closure loss function; adjusting the training parameters of the expression animation generation model according to the total loss function until a preset termination condition is met to obtain the trained expression animation generation model includes: adjusting the training parameters of the expression animation generation model according to the first total loss function until a preset termination condition is met to obtain the trained expression animation generation model;
[0009] And / or,
[0010] The key features of the predicted 3D face mesh model include vertex displacement key points of adjacent frames of the predicted 3D face mesh model; comparing the differences between the key features of the real 3D face mesh model and the key features of the predicted 3D face mesh model to obtain a feature-related loss function includes: determining the mean square error or root mean square error between the vertex displacement key points of adjacent frames of the real 3D face mesh model and the vertex displacement key points of adjacent frames of the predicted 3D face mesh model as a temporal continuity loss function; obtaining a total loss function based on the vertex distance loss function and the feature-related loss function includes: obtaining a second total loss function based on the vertex distance loss function and the temporal continuity loss function; adjusting the training parameters of the facial animation generation model according to the total loss function until a preset termination condition is met to obtain the trained facial animation generation model includes: adjusting the training parameters of the facial animation generation model according to the second total loss function until a preset termination condition is met to obtain the trained facial animation generation model.
[0011] Furthermore, before adjusting the training parameters of the facial expression animation generation model according to the vertex distance loss function until a preset termination condition is met to obtain the trained facial expression animation generation model, the facial animation generation method for emotional expression further includes:
[0012] Using a trained facial expression classification model and predicted PCA coefficients, the predicted sentiment classification probabilities of real facial expression animations and predicted facial expression animations are respectively predicted. The trained facial expression classification model is obtained by training the model with facial expression animation samples of real 3D faces corresponding to the speech sample set, and the facial expression animation samples are labeled with sentiment. The difference between the predicted sentiment classification probabilities of real facial expression animations and predicted facial expression animations is compared to obtain the sentiment consistency loss function. Based on the vertex distance loss function and the sentiment consistency loss function, a third total loss function is obtained.
[0013] The step of adjusting the training parameters of the facial animation generation model according to the vertex distance loss function until a preset termination condition is met to obtain the trained facial animation generation model includes: adjusting the training parameters of the facial animation generation model according to the third total loss function until a preset termination condition is met to obtain the trained facial animation generation model.
[0014] Furthermore, the step of using the trained expression classification model and the predicted PCA coefficients to predict the sentiment classification probability of the real facial expression animation and the predicted sentiment classification probability of the facial expression animation respectively includes: inputting the predicted PCA coefficients into the trained expression classification model to output the predicted sentiment classification probability of the facial expression animation; obtaining a real 3D face animation sample set corresponding to the speech sample set, wherein the real 3D face animation sample set includes real 3D face expressions; projecting the real 3D face expressions into the PCA coefficients of a real 3D deformable face model; inputting the PCA coefficients of the real 3D deformable face model into the trained expression classification model to output the sentiment classification probability of the real facial expression animation.
[0015] The step of obtaining the third total loss function based on the vertex distance loss function and the sentiment consistency loss function includes: determining the mean square error or root mean square error of the predicted sentiment classification probability of the facial animation and the sentiment classification probability of the actual facial animation as the sentiment consistency loss function.
[0016] Furthermore, the facial animation generation method for emotional expression further includes: obtaining a trained expression classification model in the following manner: acquiring a three-dimensional facial animation sample set containing the speech sample set, wherein the real three-dimensional facial animation sample set includes real three-dimensional facial expressions; projecting the real three-dimensional facial expressions into the PCA coefficients of a real three-dimensional deformable face model; inputting the PCA coefficients of the real three-dimensional deformable face model into the expression classification model to output the emotional classification probability of the real facial animation; using cross-entropy, determining the current loss function based on the emotional label and the predicted emotional classification probability of the facial animation; adjusting the training parameters of the expression classification model according to the current loss function until a preset termination condition is met to obtain a trained expression classification model.
[0017] Furthermore, the step of projecting the predicted PCA coefficients of the 3D facial expression into 3D facial expression animation data includes: converting the predicted PCA coefficients of the 3D facial expression into a mixed deformation coefficient expression basis according to the pre-constructed correspondence between the PCA coefficients of the 3D facial expression and the expression basis of the mixed deformation coefficients; the expression basis is used to reflect facial movement, and each basis of the 3D deformable face model corresponds to a set of mixed deformation coefficients; multiplying the PCA coefficients of the expression basis with the PCA coefficients predicted by the expression animation generation model and summing them to obtain the coefficients of the expression basis; the step of redirecting the projected facial expression animation data onto the target digital human includes: redirecting the coefficients of the expression basis to the target digital human.
[0018] This application provides a facial animation generation device for expressing emotions, comprising:
[0019] The system includes a voice acquisition module for acquiring user-input voice; a prediction module for inputting the voice into a trained facial animation generation model to output predicted PCA coefficients of the facial animation of a 3D face; the trained facial animation generation model is trained using a voice sample set as input; a facial animation data projection module for projecting the predicted PCA coefficients of the facial animation into facial animation data of a 3D face; and a redirection module for redirecting the projected facial animation data onto a target digital human.
[0020] This application provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the method described in any of the preceding claims.
[0021] In some embodiments, the facial animation generation method for emotional expression of this application obtains user-input speech; inputs the speech into a trained facial animation generation model to output predicted PCA (Principal Component Analysis) coefficients of the facial animation of a 3D face; the trained facial animation generation model is trained using a speech sample set as input; the predicted PCA coefficients of the facial animation are projected as facial animation data of a 3D face; and the projected facial animation data is redirected onto a target digital human. Thus, by using speech input into the trained facial animation generation model to output predicted PCA coefficients of the facial animation of a 3D face, projecting the predicted PCA coefficients of the facial animation into facial animation data of a 3D face, and redirecting the projected facial animation data onto a target digital human, the generation of the entire 3D facial animation can be completed. This makes the facial animation of the digital human face natural, improving the user experience during interaction. Attached Figure Description
[0022] Figure 1 The diagram shown is a flowchart illustrating the facial animation generation method for emotional expression provided in an embodiment of this application.
[0023] Figure 2 As shown Figure 1 The diagram illustrates the process of obtaining a trained facial expression animation generation model using the facial animation generation method for emotional expression.
[0024] Figure 3 As shown Figure 2 The diagram shows another flowchart of step 250 above, which yields the trained facial animation generation model.
[0025] Figure 4 As shown Figure 2 The diagram shows another step 250 above in obtaining the trained facial expression animation generation model.
[0026] Figure 5 As shown Figure 4 The diagram shows the structure of the trained facial expression animation generation model.
[0027] Figure 6 The diagram shown is a flowchart illustrating the process of obtaining a trained facial expression classification model according to an embodiment of this application.
[0028] Figure 7 As shown Figure 2 The trained facial animation generation model shown and Figure 6 The diagram shows the application process of the trained facial expression classification model.
[0029] Figure 8 As shown Figure 1 The flowcharts for steps 130 and 140 of the facial animation generation method for expressing emotions are shown below.
[0030] Figure 9 The diagram shown is a schematic representation of the facial animation generation device for emotional expression provided in an embodiment of this application.
[0031] Figure 10 The diagram shown is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0033] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.
[0034] To address the technical problem of stiff and unnatural digital human facial animations that negatively impact user experience during interaction, this application provides a method for generating facial animations that express emotions. The method involves: acquiring user-inputted speech; feeding the speech into a trained facial expression animation generation model to output predicted PCA (Principal Component Analysis) coefficients of the 3D facial expression animation; training the facial expression animation generation model using a speech sample set as input; projecting the predicted PCA coefficients of the facial expression animation into 3D facial expression animation data; and redirecting the projected facial expression animation data onto a target digital human.
[0035] In this embodiment, since the trained facial expression animation generation model is trained using a set of speech samples, the facial expression animations of the digital human face generated by the trained model are richer and more vivid. Furthermore, speech input is used to train the facial expression animation generation model to output the predicted PCA coefficients of the 3D facial expression animation; these predicted PCA coefficients are then projected onto the 3D facial expression animation data; and the projected expression animation data is redirected onto the target digital human to complete the generation of the entire 3D facial expression animation. This makes the facial expression animation of the digital human face natural, improving the user experience during interaction. Simultaneously, using speech input to train the facial expression animation generation model to complete the generation of the entire 3D facial expression animation significantly reduces the difficulty of animation production.
[0036] The facial animation generation method for emotional expression in this application is applied to an electronic device. Specifically, the electronic device can be a desktop computer, a portable computer, a smart mobile terminal, etc. No limitation is made here; any electronic device that can implement the embodiments of this application falls within the protection scope of this application. The aforementioned electronic devices can be used, including but not limited to, applications in film and television creation, game development, human-computer interaction, banking, healthcare, education, government affairs, and communications. Further examples are not provided here.
[0037] Figure 1 The diagram shown is a flowchart illustrating a method for generating facial animations expressing emotions, as provided in an embodiment of this application. Figure 1 As shown, the method for generating facial animations that express emotions includes the following steps 110 to 140:
[0038] Step 110: Obtain the user's voice input.
[0039] Step 120: Input the speech into the trained facial expression animation generation model to output the PCA coefficients of the predicted 3D facial expression animation. The trained facial expression animation generation model is trained by inputting the speech sample set into the facial expression animation generation model.
[0040] During the training process of the facial expression animation generation model, a trained facial expression classification model and the facial expression animation generation model are used to predict the PCA coefficients of the three-dimensional facial expressions corresponding to the speech sample set, thereby obtaining an emotion consistency loss function. This emotion consistency loss function is used to adjust the training parameters of the facial expression animation generation model. See below for detailed explanation.
[0041] During the training process of the facial expression animation generation model, the model is used to predict the PCA coefficients of the three-dimensional facial expressions corresponding to the speech sample set, resulting in a vertex distance loss function. This vertex distance loss function is used to adjust the training parameters of the facial expression animation generation model. See below for detailed explanation.
[0042] During the training process of the facial expression animation generation model, the model is used to predict the PCA coefficients of the 3D facial expressions corresponding to the speech sample set, resulting in a feature-related loss function. This feature-related loss function is used to adjust the training parameters of the facial expression animation generation model. The feature correlation can reflect the key feature points of the 3D face mesh model. Example 1: The feature correlation can be the upper and lower lips of the 3D face mesh model; the corresponding feature-related loss function is the lip closure loss function. Example 2: The feature correlation can be the vertex displacements of adjacent frames in the 3D face mesh model; the corresponding feature-related loss function is the temporal continuity loss function. See below for detailed explanation.
[0043] In the usage stage of step 120 above, it is not necessary to use the trained facial expression classification model. After the speech is input into the trained facial expression animation generation model, the facial expression animation generation model performs feature extraction and other processing on the speech to predict the PCA coefficients of the facial expression animation.
[0044] The facial expression animation generation model can be, but is not limited to, a deep neural network structure. This model is used to extract speech features and obtain the PCA coefficients of the predicted 3D facial expression animation.
[0045] Step 130: Project the predicted PCA coefficients of the facial expression animation into 3D facial expression animation data. The facial expression animation data can consist of multiple animation frames.
[0046] Step 140: Redirect the projected facial animation data onto the target digital human. The target digital human is the digital human to be redirected, which can be set according to user requirements. The projected facial animation data will ultimately be displayed on the target digital human.
[0047] Figure 2 As shown Figure 1 The diagram illustrates the process of obtaining a trained facial expression animation generation model using a facial animation generation method that focuses on emotional expression. Figure 2 In the illustrated embodiment, the trained facial expression animation generation model is obtained through the following steps 210 to 250:
[0048] Step 210: Obtain a speech sample set. The speech sample set includes the vertex coordinates of a real 3D face mesh model, which includes the feature keypoints of the real 3D face mesh model. These vertex coordinates reflect the facial information of the 3D face mesh model. A vertex is a component of the 3D face mesh model, which can be viewed as multiple small triangles (quadrilaterals), each of which can be considered a vertex. The more vertices, the more detailed the 3D face mesh model.
[0049] Step 220: Input the speech sample set into the facial expression animation generation model to predict the PCA coefficients of the 3D facial expression animation corresponding to the speech sample set; the predicted PCA coefficients of the 3D facial expression animation corresponding to the speech sample set are used to reflect the predicted facial expression animation data.
[0050] The aforementioned facial expression animation generation model may, but is not limited to, include an encoder and a decoder. Step 220 may further include first extracting speech features from a speech sample set. Then, the encoder encodes the speech features into a latent expression input to the decoder. Finally, the decoder can obtain the PCA coefficients of the 3D facial expression animation.
[0051] In some embodiments, the decoder described above may use an ASR (Automated Speech Recognition) model. For example, the ASR model may be, but is not limited to, a wav2vec2.0 model. Using a wav2vec2.0 model trained on a large-scale speech corpus as a pre-trained model for the encoder to extract features from the speech sample set can improve the generalization ability of the facial animation generation model.
[0052] For example, librosa is used to extract MFCC (Mel Frequency Cepstrum Coefficient) features, with a sampling rate of 16000Hz, a sliding window size of 0.02s, and a sliding window step size of 0.02s. The frame rate for extracting MFCC features from the speech sample set is 50fps, meaning each frame contains 0.02s of speech signal. Librosa is a Python toolkit for audio and music analysis and processing, offering a wide range of powerful functions including common time-frequency processing, feature extraction, and sound graph rendering.
[0053] The decoder comprises four 1D temporal convolutions and three fully connected layers to learn the mapping from speech features to PCA coefficients of facial expression animation, thereby obtaining the predicted PCA coefficients of the 3D facial expression animation. Furthermore, the predicted PCA coefficients of the 3D facial expression animation can be projected onto a 3D face mesh model using the following formula:
[0054]
[0055] In the formula, For realistic 3D facial expressions, S0 represents a neutral expression, e i (i = 0, 1, ..., m) are the PCA coefficients of the actual 3D facial expression projection. i (i = 0, 1, ..., m are the first m principal components related to expression in the 3DMM (3D Morphable Model). For the PCA coefficients predicted by 3DMM, t i Let be the i-th expression PCA component in 3DMM, and m represent the first m components. In this embodiment, m = 64, where i represents the sequence number and m represents the sequence number. The expression animation generation model of this embodiment is very lightweight, has low computational resource requirements, and runs quickly.
[0056] Step 230: Convert the predicted PCA coefficients into the vertex coordinates of the predicted 3D face mesh model.
[0057] Step 230 above may further include reprojecting the predicted PCA coefficients into the vertex coordinates of the predicted 3D face mesh model based on the 3DMM model used; and projecting the vertex coordinates of the predicted 3D face mesh model into the vertex coordinates of the predicted 3D face mesh model.
[0058] Step 240: Compare the vertex coordinates of the real 3D face mesh model with the vertex coordinates of the predicted 3D face mesh model to obtain the vertex distance loss function. The degree of difference between the vertex coordinates of the real 3D face mesh model and the predicted 3D face mesh model reflects the difference between the vertex coordinates of the real 3D face mesh model and the predicted 3D face mesh model. This degree of difference can be, but is not limited to, the variance, standard deviation, root mean square error, or root mean square error between the two.
[0059] in,
[0060] In the formula, L dist Let L represent the vertex distance loss function, dist represent the loss function, and p represent the distance. i t Let be the 3D coordinates of the i-th vertex of the real 3D face mesh model at time t. Let t be the 3D coordinates of the i-th vertex of the 3D face mesh model predicted at time t, and N be the number of vertices of the 3D face mesh model or the mesh of the 3D face mesh model.
[0061] Step 250: Adjust the training parameters of the facial animation generation model according to the vertex distance loss function until the preset termination condition is met, and obtain the trained facial animation generation model.
[0062] The preset termination condition can be set according to user needs. In some embodiments, the preset termination condition may be that the loss value of the loss function converges to a preset threshold, resulting in a trained facial animation generation model. In other embodiments, the preset termination condition may be that the current loss of the loss function reaches a set accuracy, or that the training of the trained facial animation generation model reaches a set number of iterations, etc. The loss function includes one or more of the following: vertex distance loss function, lip closure loss function, temporal continuity loss function, emotional consistency loss function, first total loss function, second total loss function, third total loss function, and fourth total loss function. See below for detailed explanation.
[0063] In this embodiment, by using the vertex key points of the real 3D face mesh model and the predicted 3D face mesh model, the trained real 3D face mesh model and the predicted 3D face mesh model are made closer, thus achieving synchronization of speech and facial animation data. Simultaneously, using a facial animation generation model trained with a speech sample set ensures that the predicted facial animation data maintains consistent emotional expression with the real facial animation data, significantly improving the prediction accuracy of the trained facial animation generation model and the vividness of the generated facial animation data.
[0064] Step 250 above has at least four specific embodiments, as follows:
[0065] (1) In the first specific embodiment of step 250 above, the vertex distance loss function is used to adjust the training parameters of the facial animation generation model until the preset termination condition is met, thus obtaining the trained facial animation generation model. In this way, by using only the vertex distance loss function to adjust the training parameters of the facial animation generation model, the predicted 3D face mesh model is closer to the real 3D face mesh model, thereby aligning the 3D face mesh model and achieving synchronization of voice and facial animation data.
[0066] Figure 3 As shown Figure 2 The diagram shows another step 250 above, which involves obtaining a trained facial expression animation generation model. (See attached diagram.) Figure 3 As shown, the vertex coordinates of the predicted 3D face mesh model include the feature keypoints of the predicted 3D face mesh model. Feature keypoints are partial features within the vertices of the 3D face mesh model. Feature keypoints may include, but are not limited to, one or more of the upper and lower lip keypoint pairs of the predicted 3D face mesh model and vertex displacement keypoints of adjacent frames of the predicted 3D face mesh model.
[0067] Step 251a involves comparing the differences between the feature key points of the real 3D face mesh model and the feature key points of the predicted 3D face mesh model to obtain the feature-related loss function. The degree of difference between the feature key points of the real 3D face mesh model and the feature key points of the predicted 3D face mesh model reflects the differences between their respective feature key points. This degree of difference can be, but is not limited to, the variance, standard deviation, root mean square error, or root mean square error of the two models.
[0068] Step 252: Obtain the total loss function based on the vertex distance loss function and the feature-related loss function.
[0069] Step 253: Adjust the training parameters of the facial animation generation model according to the total loss function until the preset termination condition is met, and obtain the trained facial animation generation model.
[0070] (2) In the second specific embodiment of step 250 above, step 251a may further include a first step, determining the degree of difference between the distance between the upper and lower lip keypoint pairs of the real 3D face mesh model and the distance between the predicted upper and lower lip keypoint pairs of the 3D face mesh model, such as, but not limited to, mean square error or root mean square error, as the lip closure loss function. Correspondingly, step 252 may further include a second step, obtaining a first total loss function based on the vertex distance loss function and the lip closure loss function. For example, the vertex distance loss function and the lip closure loss function are linearly weighted to obtain the first total loss function. Step 253 may further include a third step, adjusting the training parameters of the facial animation generation model according to the first total loss function until a preset termination condition is met, thereby obtaining a trained facial animation generation model.
[0071] The first step above specifically uses the following formula to obtain the lip closure loss function:
[0072]
[0073] In the formula, L lip Denotes the lip closure loss function, d i t Let be the distance between the i-th pair of upper and lower lip keypoints in the real 3D face mesh model at time t. Let M be the distance between the corresponding key points of the i-th upper and lower lips of face t, and M be the number of M key point pairs of the lips. For example, in this embodiment, 5 key points of the inner circle of the upper lip and 5 key points of the corresponding inner circle of the lower lip can be taken as key point pairs of the upper and lower lips to calculate the lip closure loss.
[0074] In this embodiment, visually, there is a significant difference between when a user pronounces words with their lips closed or open. Therefore, by using a lip closure loss function in addition to the vertex distance loss function, the accuracy of the generated animated lips can be improved.
[0075] (3) In the third specific embodiment of step 250 above, step 251a can further include a first step of determining the degree of difference between the vertex displacement key points of adjacent frames of the real 3D face mesh model and the vertex displacement key points of adjacent frames of the predicted 3D face mesh model, such as, but not limited to, mean square error or root mean square error, as the temporal continuity loss function. Correspondingly, step 252 can further include a second step of obtaining a second total loss function based on the vertex distance loss function and the temporal continuity loss function. For example, the vertex distance loss function and the temporal continuity loss function are linearly weighted to obtain the second total loss function. Step 252 can further include a third step of adjusting the training parameters of the facial animation generation model according to the second total loss function until a preset termination condition is met, thereby obtaining a trained facial animation generation model.
[0076] Specifically, in the first step described above, the time continuity loss function is obtained using the following formula:
[0077]
[0078] In the formula, L time The time continuity loss function is used to calculate the mean square error between the vertex displacements of adjacent frames before and after the actual 3D face mesh model and the vertex displacements of adjacent frames before and after the predicted 3D face mesh model.
[0079] In this embodiment, based on the vertex distance loss function, the animation between adjacent frames is made continuous and smooth through the temporal continuity loss function, so as to improve the smoothness of the animation and the stability of training.
[0080] Figure 4 As shown Figure 2 The diagram shows another step in the process of obtaining the trained facial expression animation generation model, specifically step 250. Figure 4 As shown, before step 250 above, the method for generating facial animations expressing emotions also includes steps 241 to 243.
[0081] Step 241: Using the trained expression classification model and the predicted PCA coefficients, predict the sentiment classification probability of the real expression animation and the predicted sentiment classification probability of the expression animation, respectively. The trained expression classification model is obtained by training the expression classification model with expression animation samples of real three-dimensional human faces corresponding to the speech sample set. The expression animation samples are labeled with sentiment.
[0082] Figure 5 As shown Figure 4 The diagram shows the structure of the trained facial expression animation generation model. Figure 5As shown, the trained facial expression classification model can be, but is not limited to, a deep neural network structure, containing three fully connected layers. The first and second fully connected layers use ReLU as the activation function, and the last fully connected layer is connected to a softmax layer as the activation function. The facial expression classification model obtained by the deep neural network in this embodiment is very lightweight, requires low computational resources, and runs quickly.
[0083] Step 242: Compare the predicted sentiment classification probabilities of the real facial animation with the predicted sentiment classification probabilities of the actual facial animation to obtain the sentiment consistency loss function. The degree of difference between the predicted sentiment classification probabilities of the real facial animation and the predicted sentiment classification probabilities of the actual facial animation reflects the difference between the two. This degree of difference can be, but is not limited to, the variance, standard deviation, root mean square error, or root mean square error of the two.
[0084] (4) In the fourth specific embodiment of step 250 above, step 241 further includes the following four steps: First, input the predicted PCA coefficients into the trained expression classification model to output the predicted emotion classification probability of the expression animation. Second, obtain the real 3D face animation sample set corresponding to the speech sample set, wherein the real 3D face animation sample set includes real 3D face expressions. Third, project the real 3D face expressions into the PCA coefficients of the real 3D deformable face model. Fourth, input the PCA coefficients of the real 3D deformable face model into the trained expression classification model to output the emotion classification probability of the real expression animation.
[0085] Step 242 above may further include a fifth step, which determines the mean square error or root mean square error of the predicted emotion classification probability of the facial animation and the actual emotion classification probability of the facial animation as the emotion consistency loss function.
[0086] Specifically, in the fifth step above, the following formula is used to obtain the emotional consistency loss function:
[0087]
[0088] In the formula, L emo For loss of emotional consistency, ∈ t The expression classification model predicts the emotion classification probability of the actual 3D face mesh model at time t. The emotion classification probability of the true 3D face mesh model at time t, predicted by the expression classification model.
[0089] Step 243: Based on the vertex distance loss function and the sentiment consistency loss function, obtain the third total loss function.
[0090] Example 1: The vertex distance loss function, lip closure loss function, and sentiment consistency loss function are linearly weighted to obtain the third total loss function.
[0091] Example 2: The vertex distance loss function, temporal continuity loss function, and sentiment consistency loss function are linearly weighted to obtain the third total loss function.
[0092] Example 3: By linearly weighting the vertex distance loss function, lip closure loss function, temporal continuity loss function, and sentiment consistency loss function, a fourth total loss function is obtained:
[0093] L=λ1L dist +λ1L lip +λ1L time +λ1L emo
[0094] In the formula, λ1, λ2, λ3, and λ4 represent the proportions of the vertex distance loss function, lip closure loss function, temporal continuity loss function, and sentiment consistency loss function, respectively. For example, λ1, λ2, λ3, and λ4 can be, but are not limited to, 1, 1, 1, and 1, respectively.
[0095] In Example 3, all loss functions are weighted and then backpropagated to guide the training of the facial animation generation model. Training stops and the model is saved when the loss value of the fourth total loss function decreases to a predetermined threshold. See below for details. Figure 7 As shown.
[0096] In this embodiment, the training data includes approximately 10 hours of voice and facial animation pairs (one male and one female) collected using motion capture equipment, evenly covering data from seven emotional states: neutral, happy, sad, surprised, disgusted, fearful, and angry. The 3D facial animation is projected using the PCA coefficients of the FLAME model during preprocessing. Of course, these seven emotional states are merely illustrative; other emotional states such as anger and contempt also fall within the scope of protection of this embodiment, and will not be listed here.
[0097] Accordingly, step 250 may further include step 251b, which adjusts the training parameters of the facial animation generation model according to the third total loss function until the preset termination condition is met, thereby obtaining the trained facial animation generation model.
[0098] In the embodiments of this application, the combination of lip-syncing and emotional expression makes the facial animation generation model tend to be rich in emotion during training, resulting in sufficiently expressive animation and improving the emotional expression of the generated animation.
[0099] exist Figures 2 to 4 In the embodiment shown, the speech sample set contains samples with a uniform distribution of 7 different emotions, and the speech sample set is divided into a training set and a validation set according to a preset ratio. The Adam optimizer is used and the training is performed for 20 epochs with a learning rate of 1e-4.
[0100] The process of testing the trained facial expression animation generation model using the validation set of the above-mentioned speech sample set is as follows:
[0101] Step 1: Divide the speech sample set into a training set and a validation set according to a preset ratio. The training set is used to train the trained facial expression animation generation model, and the validation set is used to test the accuracy of the trained model. Step 2: Train the trained facial expression animation generation model using the training set to obtain the final trained model. The preset ratio reflects the fact that the training set is larger than the validation set. The preset ratio is the ratio of the training set to the validation set. Preset ratios include 6:4, 7:3, or 8:2.
[0102] The method also includes the following steps: Step 1: Test the trained facial expression animation generation model using a validation set, obtain the test results, and determine the warning accuracy of the trained facial expression animation generation model based on the test results. Then, compare the test results with the labels of the validation set to verify the accuracy of the trained facial expression animation generation model and the extracted speech features. Step 2: If the warning accuracy of the trained facial expression animation generation model does not meet a preset threshold, adjust the model parameters of the trained facial expression animation generation model. Each trained facial expression animation generation model has different model parameters. Specific model parameters in the trained facial expression animation generation model include, for example, the number of neurons and hidden layers in the neural network. Step 3: If the warning accuracy of the trained facial expression animation generation model meets the preset threshold, then the trained facial expression animation generation model is considered the final trained facial expression animation generation model. This ensures that the warning accuracy of the trained facial expression animation generation model is relatively good, facilitating subsequent use.
[0103] The preset threshold can be determined based on the precision required by the user. The higher the precision required by the user, the larger the preset threshold. Optionally, the preset threshold can be greater than 70%. For example, the preset threshold can be 85%.
[0104] In this embodiment, data corresponding to preset parameters of the trained facial animation generation model can be obtained without the need for manual input of parameters. This reduces the professional threshold of the trained facial animation generation model, making it less complex to use, simpler for users, and improving the user experience.
[0105] Figure 6The diagram shown is a flowchart illustrating the process of obtaining a trained facial expression classification model according to an embodiment of this application. Figure 7 As shown Figure 2 The trained facial animation generation model shown and Figure 6 The diagram shows the application process of the trained facial expression classification model. Figure 6 and Figure 7 As shown, the method for generating facial animations expressing emotions further includes: obtaining a trained expression classification model by using steps 310 to 350 as follows:
[0106] Step 310: Obtain a real 3D face animation sample set containing the voice sample set. The real 3D face animation sample set includes real 3D face expressions.
[0107] The 3D facial animation sample set used to train the expression classification model includes 3D facial expressions and emotion labels.
[0108] Step 320: Project the real 3D facial expression onto the PCA coefficients of the real 3D deformable face model.
[0109] The projection of a real 3D face onto the PCA of a real 3DMM can be described by the formula:
[0110]
[0111] In the formula, For realistic 3D facial expressions, S0 represents a neutral expression, e i (i = 0, 1, ..., m) are the PCA coefficients of the facial animation, t i (i = 0, 1, ..., m) represents the first m principal components related to facial expressions in the 3DMM model.
[0112] The process of projecting a realistic 3D face mesh model onto the PCA coefficients of a realistic 3DMM can be achieved using the least squares method. The optimization objective formula is as follows: through iterative fitting, the face synthesized from the calculated PCA coefficients is made to match the original PCA coefficients:
[0113]
[0114] Where S is the original 3D face mesh model in the training set.
[0115] Depending on the requirements, the 3DMM models that can be used include FLAME, Basel Face Model (BFM), Surrey Face Model (SFM), Face Warehouse, and Large Scale Facial Model (LSFM).
[0116] For example, the 3DMM model used is FLAME, and 64 PCA coefficients from facial animation are used to represent facial expressions. For example, the first fully connected layer of the expression classification model has an input size of 64 and an output size of 32, the second fully connected layer has an input size of 32 and an output size of 16, and the third fully connected layer has an input size of 16 and an output size of 7. After passing through the final Softmax layer, the predicted probabilities of the emotions being neutral, happy, sad, surprised, disgusted, fearful, and angry are obtained.
[0117] Step 330: Input the PCA coefficients of the real 3D deformable face model into the expression classification model to output the emotion classification probability of the real facial animation.
[0118] Step 340: Using cross-entropy, determine the current loss function based on the sentiment label and the predicted sentiment classification probability of the facial animation.
[0119] Step 350: Adjust the training parameters of the expression classification model according to the current loss function until the preset termination condition is met, and obtain the trained expression classification model.
[0120] In this embodiment, the facial expression classification model needs to simultaneously predict the emotion of the real facial expression and the probability of the emotion of the expression output by the facial expression generation network. Furthermore, the real 3D facial expressions are projected onto the PCA coefficients of the real 3DMM, thereby reducing the dimensionality of each expression to a latent vector representation. The PCA coefficients of the facial animation are used as input to output the facial expression classification probability. During training, cross-entropy is used to obtain the current loss function, guiding the training of the facial expression classification model. Once the loss value of the current loss function meets the preset termination condition, the facial expression classification model is saved, resulting in a trained facial expression classification model. This trained facial expression classification model ensures that the predicted animation has consistent emotional expression with the real animation.
[0121] Figure 6 and Figure 7 The embodiments are similar to Figures 2 to 4 The illustrated embodiment, compared to Figures 2 to 4 The illustrated embodiment, in Figure 6 and Figure 7 In this embodiment, the 3D facial animation sample set includes samples evenly distributed across 7 different emotions. The 3D facial animation sample set is divided into a training set and a validation set according to a preset ratio. The Adam optimizer is used, and the set is trained for 20 epochs with a learning rate of 1e-4. The preset ratio reflects that the training set is larger than the validation set. The preset ratio is the ratio of the training set to the validation set. Preset ratios include 6:4, 7:3, or 8:2.
[0122] The process of testing the trained expression classification model using the validation set of this 3D face animation sample set is similar to the process of testing the trained expression animation generation model using the validation set of the aforementioned speech sample set. The only difference is that the 3D face animation sample set and the trained expression classification model are used as the processing objects, while the aforementioned speech sample set and the trained expression animation generation model are used as the processing objects. Apart from the difference in the processing objects, the rest of the process is the same. Both can refer to the process of testing the trained expression animation generation model using the validation set of the aforementioned speech sample set, and will not be repeated here.
[0123] Figure 8 As shown Figure 1 The diagram illustrates the detailed process of steps 130 and 140 in the method for generating facial animations that express emotions. Figure 8 As shown, in step 131, based on the correspondence between the pre-constructed PCA coefficients of the 3D facial expression animation and the expression basis of the hybrid deformation coefficients, the predicted PCA coefficients of the 3D facial expression animation are converted into the expression basis of the hybrid deformation coefficients; the expression basis is used to reflect facial movement, and each basis of the 3D deformable face model corresponds to a set of hybrid deformation coefficients; in step 132, the PCA coefficients of the expression basis are multiplied by the PCA coefficients predicted by the expression animation generation model and summed to obtain the coefficients of the expression basis.
[0124] The Blendshape blending deformation coefficients contain a set of expression bases with specific semantics, generally including facial motion units defined by FACS (Facial Action Coding System). In this embodiment, the expression bases use Apple's ARKit Blendshapes specification, which contains 52 expression bases, is widely applicable, and is well-compatible with various digital humans. Of course, in other implementations, other specifications can be used, such as NVIDIA's audio2face-defined Blendshapes and Faceware's Blendshapes, or a custom Blendshapes set can be defined.
[0125] The Blendshapes model described above is a linear facial model where each basis vector is not orthogonal but represents an individual's facial expression. Non-orthogonality means the same expression can be derived from different combinations of expression bases, which is detrimental to the stability of network training. In contrast, the bases of 3DMMs are orthogonal rather than semantic. Blendshapes are easier for animators to manually edit, while 3DMMs are better suited for mathematical calculations. Thus, the PCA coefficients of the predicted 3D facial expression animations from the trained facial animation generation model are then converted into hybrid deformation coefficients, making the trained facial animation generation model more stable and applicable to various 3D facial mesh models.
[0126] The aforementioned expression basis is used to reflect semantic features. The expression basis can be a multi-dimensional vector, with each dimension representing an expression. For example, the predicted PCA coefficients, such as 0.5, 0.1, and 0.9 for blinking expressions, are converted into the coefficients of the expression basis on a 3D face mesh model. This allows it to be applied to various 3D face mesh models, improving the applicability.
[0127] The correspondence between the PCA coefficients and the hybrid deformation coefficients of the pre-constructed 3D facial expression animation is obtained in the following way: obtain the preset expression base; project the preset 3D face into the preset 3D facial expression PCA coefficients; establish the correspondence between the preset expression base and the preset 3D facial expression PCA coefficients.
[0128] The aforementioned correspondence may include, but is not limited to, one of the following: a list of correspondences between a preset expression base and preset PCA coefficients of a 3D facial expression; a functional relationship between a preset expression base and preset PCA coefficients of a 3D facial expression; and a mapping relationship between a preset expression base and preset PCA coefficients of a 3D facial expression.
[0129] For example, each PCA component t i Corresponding to a set of blendshape coefficients w:[w0 i w1 i ,…w m i ]
[0130] The trained facial animation generation model outputs PCA coefficients. After that, it is necessary to calculate the blendshape coefficients [w0, w1, ..., w i ,…w m ].
[0131] The calculation method is as follows:
[0132]
[0133]
[0134] …
[0135]
[0136] …
[0137]
[0138] In the formula, Here, n represents the number of PCA coefficients, and m represents the number of blendshapes. In this embodiment, n = 64 and m = 52. The blendshape coefficients [w0, w1, ..., w...] are... i ,…w m The redirection can be completed by applying it to the target digital person.
[0139] The calculation of the PCA coefficients of the preset 3D facial expression and the mapping of the preset Blendshape is the same as the process of projecting the PCA coefficients of the real 3D facial mesh model onto the real 3DMM, and the least squares method can be used.
[0140] Step 141: Redirect the coefficients of the expression base to the target digital man.
[0141] In this embodiment, by calculating the hybrid deformation coefficient for facial expression retargeting, the facial expressions corresponding to the obtained projected facial animation data can be better applied to the face of the target digital human.
[0142] In some other embodiments of step 130 above, step 130 may include, but is not limited to, projecting the predicted PCA coefficients of the 3D facial expression animation into 3D facial expression animation data according to the correspondence between the PCA coefficients of the 3D facial expression animation and the 3D facial expression animation data. Then, step 140 is performed to redirect the projected facial expression animation data onto the target digital human.
[0143] Figure 9 The diagram shown is a schematic representation of the facial animation generation device for emotional expression provided in an embodiment of this application.
[0144] like Figure 9 As shown, the facial animation generation device for emotional expression includes the following modules.
[0145] The speech acquisition module 51 is used to acquire the speech input by the user; the prediction module 52 is used to input the speech into the trained facial expression animation generation model to output the PCA coefficients of the predicted facial expression animation of the 3D face; the trained facial expression animation generation model is trained by inputting the speech sample set into the facial expression animation generation model; the facial expression animation data projection module 53 is used to project the predicted PCA coefficients of the facial expression animation into facial expression animation data of the 3D face; the redirection module 54 is used to redirect the projected facial expression animation data onto the target digital human.
[0146] In some embodiments, the facial animation generation device for emotional expression further includes a first training module for obtaining a trained facial expression animation generation model in the following manner:
[0147] The speech sample set acquisition module is used to acquire a speech sample set, which includes the vertex coordinates of a real 3D face mesh model. The vertex coordinates of the real 3D face mesh model include the feature key points of the real 3D face mesh model.
[0148] The facial animation generation model prediction module is used to input the speech sample set into the facial animation generation model to predict the PCA coefficients of the three-dimensional facial expressions corresponding to the speech sample set.
[0149] The vertex coordinate prediction module is used to convert the predicted PCA coefficients into the vertex coordinates of the predicted 3D face mesh model.
[0150] The vertex difference comparison module is used to compare the degree of difference between the vertex coordinates of the real 3D face mesh model and the vertex coordinates of the predicted 3D face mesh model, and obtain the vertex distance loss function.
[0151] The facial animation generation model adjustment module is used to adjust the training parameters of the facial animation generation model according to the vertex distance loss function until the preset termination condition is met, thus obtaining the trained facial animation generation model.
[0152] In some embodiments, the vertex coordinates of the predicted 3D face mesh model include the feature keypoints of the predicted 3D face mesh model.
[0153] Before adjusting the training parameters of the facial expression animation generation model according to the vertex distance loss function until the preset termination condition is met and the trained facial expression animation generation model is obtained, the facial animation generation device for emotional expression also includes:
[0154] The feature key point difference comparison module is used to compare the degree of difference between the feature key points of the real 3D face mesh model and the feature key points of the predicted 3D face mesh model, and obtain the feature-related loss function.
[0155] The total loss function acquisition module is used to obtain the total loss function based on the vertex distance loss function and the feature-related loss function;
[0156] The facial animation generation model adjustment module is specifically used to adjust the training parameters of the facial animation generation model according to the total loss function until the preset termination condition is met, thus obtaining a trained facial animation generation model.
[0157] In some embodiments, the feature keypoints of the predicted 3D face mesh model include the upper and lower lip keypoint pairs of the predicted 3D face mesh model.
[0158] The feature key point difference comparison module is specifically used to determine the mean square error or root mean square error between the distance between the upper and lower lip key point pairs of the real 3D face mesh model and the distance between the upper and lower lip key point pairs of the predicted 3D face mesh model as the lip closure loss function; the total loss function acquisition module is specifically used to obtain the first total loss function based on the vertex distance loss function and the lip closure loss function; the expression animation generation model adjustment module is specifically used to adjust the training parameters of the expression animation generation model according to the first total loss function until the preset termination condition is met, and obtain the trained expression animation generation model.
[0159] And / or,
[0160] The key features of the predicted 3D face mesh model include vertex displacement key points of adjacent frames of the predicted 3D face mesh model; the feature key point difference comparison module is specifically used to determine the mean square error or root mean square error between the vertex displacement key points of adjacent frames of the real 3D face mesh model and the vertex displacement key points of adjacent frames of the predicted 3D face mesh model as the temporal continuity loss function; the total loss function acquisition module is used to obtain the second total loss function based on the vertex distance loss function and the temporal continuity loss function; the expression animation generation model adjustment module is specifically used to adjust the training parameters of the expression animation generation model according to the second total loss function until the preset termination condition is met, and obtain the trained expression animation generation model.
[0161] In some embodiments, before adjusting the training parameters of the facial expression animation generation model according to the vertex distance loss function until a preset termination condition is met and a trained facial expression animation generation model is obtained, the facial animation generation device for emotional expression further includes:
[0162] The emotion classification probability prediction module is used to predict the emotion classification probability of the real facial animation and the predicted emotion classification probability of the facial animation using a trained expression classification model and the predicted PCA coefficients, respectively. The trained expression classification model uses facial animation samples of 3D faces corresponding to the speech sample set, and the facial animation samples are labeled with emotion.
[0163] The sentiment classification probability difference comparison module is used to compare the degree of difference between the predicted sentiment classification probability of the real facial animation and the predicted sentiment classification probability of the facial animation, and obtain the sentiment consistency loss function.
[0164] The third total loss function acquisition module is used to obtain the third total loss function based on the vertex distance loss function and the sentiment consistency loss function;
[0165] The facial animation generation model adjustment module is specifically used to adjust the training parameters of the facial animation generation model according to the third total loss function until the preset termination condition is met, thus obtaining the trained facial animation generation model.
[0166] In some embodiments, the emotion classification probability prediction module is specifically used to input the predicted PCA coefficients into a trained expression classification model to output the predicted emotion classification probability of the expression animation; obtain a real 3D face animation sample set corresponding to the speech sample set, the real 3D face animation sample set including real 3D face expressions; project the real 3D face expressions into the PCA coefficients of a real 3D deformable face model; input the PCA coefficients of the real 3D deformable face model into the trained expression classification model to output the emotion classification probability of the real expression animation;
[0167] The third module for obtaining the total loss function is specifically used to determine the mean square error or root mean square error between the predicted emotion classification probability of the facial animation and the actual emotion classification probability of the facial animation as the emotion consistency loss function.
[0168] In some embodiments, the facial animation generation method for emotional expression further includes: a second training module, used to obtain a trained expression classification model in the following manner:
[0169] The 3D face animation sample set acquisition module is used to acquire a 3D face animation sample set containing a voice sample set. The real 3D face animation sample set includes real 3D face expressions.
[0170] A module for obtaining PCA coefficients of a realistic 3D deformable face model is used to project realistic 3D facial expressions into PCA coefficients of a realistic 3D deformable face model.
[0171] The module for obtaining the sentiment classification probability of realistic facial animation is used to input the PCA coefficients of a realistic 3D deformable face model into the facial expression classification model to output the sentiment classification probability of realistic facial animation.
[0172] The current loss function determination module is used to determine the current loss function based on the sentiment label and the predicted sentiment classification probability of the facial animation using cross-entropy.
[0173] The trained facial expression classification model module is used to adjust the training parameters of the facial expression classification model according to the current loss function until the preset termination condition is met, thus obtaining the trained facial expression classification model.
[0174] In some embodiments, the PCA coefficient acquisition module of the real 3D deformable face model is specifically used to convert the predicted PCA coefficients of the 3D face expression into a mixed deformation coefficient expression basis according to the pre-constructed correspondence between the PCA coefficients of the 3D face expression and the expression basis of the mixed deformation coefficients; the expression basis is used to reflect facial movement, and each basis of the 3D deformable face model corresponds to a set of mixed deformation coefficients; the PCA coefficients of the expression basis are multiplied and summed with the PCA coefficients predicted by the expression animation generation model to obtain the coefficients of the expression basis; the redirection module is specifically used to redirect the coefficients of the expression basis to the target digital human.
[0175] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0176] Figure 10 The diagram shown is a block diagram of the electronic device 60 provided in an embodiment of this application. Figure 10 As shown, the electronic device 60 includes one or more processors 61 for implementing the above-described method for generating facial animations with emotional expression. In some embodiments, the electronic device 60 may include a computer-readable storage medium 69, which may store a program that can be invoked by the processor 61, and may include a non-volatile storage medium. In some embodiments, the electronic device 60 may include memory 68 and an interface 67. In some embodiments, the electronic device 60 may also include other hardware depending on the specific application.
[0177] The computer-readable storage medium 69 of this application embodiment stores a program that, when executed by the processor 61, is used to implement the facial animation generation method for emotional expression as described above.
[0178] This application may take the form of a computer program product implemented on one or more computer-readable storage media 69 (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing program code. The computer-readable storage media 69 includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented using any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media 69 include, but are not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0179] The above are merely preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification shall be included within the scope of protection of this specification.
[0180] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitations, an element qualified by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
Claims
1. A method for generating facial animations that express emotions, characterized in that, include: Get the user's voice input; The speech is input into a trained facial expression animation generation model to output the PCA coefficients of the predicted 3D facial expression animation. The trained facial expression animation generation model is trained using a speech sample set as input. The facial expression animation generation model includes an encoder and a decoder. The encoder encodes speech features into latent expressions and inputs them into the decoder, which obtains the PCA coefficients of the 3D facial expression animation. The decoder contains four 1D temporal convolutions and three fully connected layers to learn the mapping from speech features to PCA coefficients of facial expression animation, so as to obtain the predicted PCA coefficients of facial expression animation of the 3D face. The predicted PCA coefficients of the facial animation are projected into 3D facial animation data; Redirect the projected facial animation data onto the target digital human; The trained facial expression animation generation model is obtained in the following way: The mean square error or root mean square error between the distance between the upper and lower lip keypoint pairs of the real 3D face mesh model and the distance between the upper and lower lip keypoint pairs of the predicted 3D face mesh model is determined as the lip closure loss function. The vertex distance loss function and the lip closure loss function are linearly weighted to obtain the first total loss function. The training parameters of the expression animation generation model are adjusted according to the first total loss function until the preset termination condition is met, and the trained expression animation generation model is obtained. The upper and lower lip keypoint pairs refer to the 5 keypoints of the inner circle of the upper lip and the corresponding 5 keypoints of the inner circle of the lower lip.
2. The facial animation generation method for expressing emotions as described in claim 1, characterized in that, The method for generating facial animations that express emotions also includes: The trained facial expression animation generation model is obtained in the following way: Obtain a speech sample set, which includes the vertex coordinates of a real 3D face mesh model, and the vertex coordinates of the real 3D face mesh model include the feature key points of the real 3D face mesh model; The speech sample set is input into the facial expression animation generation model to predict the PCA coefficients of the three-dimensional facial expressions corresponding to the speech sample set; The predicted PCA coefficients are converted into vertex coordinates of the predicted 3D face mesh model; By comparing the vertex coordinates of the real 3D face mesh model with the vertex coordinates of the predicted 3D face mesh model, the vertex distance loss function is obtained. The training parameters of the facial animation generation model are adjusted according to the vertex distance loss function until the preset termination condition is met, thus obtaining the trained facial animation generation model.
3. The facial animation generation method for expressing emotions as described in claim 2, characterized in that, The vertex coordinates of the predicted 3D face mesh model include the feature key points of the predicted 3D face mesh model; Before adjusting the training parameters of the facial expression animation generation model according to the vertex distance loss function until a preset termination condition is met to obtain the trained facial expression animation generation model, the facial expression animation generation method further includes: By comparing the difference between the feature key points of the real 3D face mesh model and the feature key points of the predicted 3D face mesh model, a feature-related loss function is obtained. Based on the vertex distance loss function and the feature-related loss function, the total loss function is obtained; The step of adjusting the training parameters of the facial animation generation model according to the vertex distance loss function until a preset termination condition is met to obtain the trained facial animation generation model includes: The training parameters of the facial animation generation model are adjusted according to the total loss function until the preset termination condition is met, thus obtaining the trained facial animation generation model.
4. The facial animation generation method for expressing emotions as described in claim 3, characterized in that, The predicted 3D face mesh model's key features include vertex displacement key points of adjacent frames of the predicted 3D face mesh model; comparing the differences between the key features of the real 3D face mesh model and the predicted 3D face mesh model to obtain a feature-related loss function includes: determining the mean square error or root mean square error between the vertex displacement key points of adjacent frames of the real 3D face mesh model and the vertex displacement key points of adjacent frames of the predicted 3D face mesh model as a temporal continuity loss function; obtaining a total loss function based on the vertex distance loss function and the feature-related loss function includes: obtaining a second total loss function based on the vertex distance loss function and the temporal continuity loss function; adjusting the training parameters of the facial animation generation model according to the total loss function until a preset termination condition is met to obtain the trained facial animation generation model includes: adjusting the training parameters of the facial animation generation model according to the second total loss function until a preset termination condition is met to obtain the trained facial animation generation model.
5. The facial animation generation method for expressing emotions as described in any one of claims 2 to 4, characterized in that, Before adjusting the training parameters of the facial expression animation generation model according to the vertex distance loss function until a preset termination condition is met to obtain the trained facial expression animation generation model, the facial animation generation method for emotional expression further includes: Using the trained facial expression classification model and the predicted PCA coefficients, the predicted sentiment classification probability of the real facial expression animation and the predicted sentiment classification probability of the facial expression animation are respectively predicted; the trained facial expression classification model is obtained by training the facial expression classification model using facial expression animation samples of real three-dimensional human faces corresponding to the speech sample set, and the facial expression animation samples are labeled with sentiment. The sentiment consistency loss function is obtained by comparing the difference between the predicted sentiment classification probability of the real facial animation and the predicted sentiment classification probability of the facial animation. Based on the vertex distance loss function and the sentiment consistency loss function, a third total loss function is obtained; The step of adjusting the training parameters of the facial animation generation model according to the vertex distance loss function until a preset termination condition is met to obtain the trained facial animation generation model includes: The training parameters of the facial animation generation model are adjusted according to the third total loss function until the preset termination condition is met, thus obtaining the trained facial animation generation model.
6. The facial animation generation method for expressing emotions as described in claim 5, characterized in that, The step of using the trained facial expression classification model and the predicted PCA coefficients to predict the sentiment classification probability of the real facial expression animation and the predicted sentiment classification probability of the facial expression animation, respectively, includes: The predicted PCA coefficients are input into the trained facial expression classification model to output the predicted sentiment classification probability of the facial expression animation. Obtain the real three-dimensional face animation sample set corresponding to the voice sample set, wherein the real three-dimensional face animation sample set includes real three-dimensional face expressions; The real 3D facial expressions are projected as PCA coefficients of a real 3D deformable face model; The PCA coefficients of the real 3D deformable face model are input into the trained expression classification model to output the emotion classification probability of the real facial animation. The process of obtaining the third total loss function based on the vertex distance loss function and the sentiment consistency loss function includes: The mean square error or root mean square error of the predicted emotion classification probability of the facial animation and the actual emotion classification probability of the facial animation is determined as the emotion consistency loss function.
7. The facial animation generation method for expressing emotions as described in claim 5, characterized in that, The method for generating facial animations that express emotions also includes: The trained facial expression classification model is obtained using the following method: Obtain a 3D face animation sample set containing the voice sample set, wherein the real 3D face animation sample set includes real 3D face expressions; The real 3D facial expressions are projected as PCA coefficients of a real 3D deformable face model; The PCA coefficients of the real 3D deformable human face model are input into the expression classification model to output the emotion classification probability of the real facial animation. Using cross-entropy, the current loss function is determined based on the sentiment tags and the predicted sentiment classification probabilities of the facial animations; The training parameters of the expression classification model are adjusted according to the current loss function until the preset termination condition is met, thus obtaining a trained expression classification model.
8. The facial animation generation method for expressing emotions as described in any one of claims 1 to 4, characterized in that, The process of projecting the predicted PCA coefficients of 3D facial expressions into 3D facial expression animation data includes: Based on the pre-constructed correspondence between the PCA coefficients of 3D facial expressions and the expression basis of mixed deformation coefficients, the predicted PCA coefficients of 3D facial expressions are converted into expression basis of mixed deformation coefficients; the expression basis is used to reflect facial movement, and each basis of the 3D deformable face model corresponds to a set of mixed deformation coefficients. The coefficients of the expression basis are obtained by multiplying the PCA coefficients of the expression basis with the PCA coefficients predicted by the expression animation generation model and summing them. The step of redirecting the projected facial expression animation data onto the target digital human includes: The coefficients of the expression base are redirected to the target digital human.
9. A facial animation generation device for expressing emotions, characterized in that, include: The voice acquisition module is used to acquire the user's voice input. A prediction module is used to input the speech into a trained facial expression animation generation model to output the predicted PCA coefficients of the facial expression animation of a 3D face; the trained facial expression animation generation model is trained using a speech sample set as input; the facial expression animation generation model includes an encoder and a decoder; the encoder encodes speech features into latent expressions and inputs them into the decoder, and the decoder obtains the PCA coefficients of the facial expression animation of a 3D face; The decoder contains four 1D temporal convolutions and three fully connected layers to learn the mapping from speech features to PCA coefficients of facial expression animation, so as to obtain the predicted PCA coefficients of facial expression animation of the 3D face. The facial animation data projection module is used to project the predicted PCA coefficients of facial animation into 3D facial animation data. The redirection module is used to redirect the projected facial animation data onto the target digital human. The trained facial expression animation generation model is obtained in the following way: The mean square error or root mean square error between the distance between the upper and lower lip keypoint pairs of the real 3D face mesh model and the distance between the upper and lower lip keypoint pairs of the predicted 3D face mesh model is determined as the lip closure loss function. The vertex distance loss function and the lip closure loss function are linearly weighted to obtain the first total loss function. The training parameters of the expression animation generation model are adjusted according to the first total loss function until the preset termination condition is met, and the trained expression animation generation model is obtained. The upper and lower lip keypoint pairs refer to the 5 keypoints of the inner circle of the upper lip and the corresponding 5 keypoints of the inner circle of the lower lip.
10. A computer-readable storage medium, characterized in that, It stores a program that, when executed by a processor, implements the facial animation generation method for emotional expression as described in any one of claims 1-8.