Facial action driving method and device, readable storage medium and electronic equipment

By extracting the emotional characteristics of the target sound and generating three-dimensional facial movements, the problem of poor realism in the prior art is solved, and more natural and emotional facial movements are achieved.

CN120215693APending Publication Date: 2025-06-27UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510229313.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing facial action driving methods have poor sense of reality when driving facial movements and cannot be as emotional as real human facial movements.

Method used

By obtaining the emotional characteristics of the target sound and using a preset three-dimensional facial action model based on these emotional characteristics, the corresponding three-dimensional facial action model is finally controlled to perform these actions.

Benefits of technology

It improves the authenticity and emotional richness of facial movements, making facial movements more natural and real.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215693A_ABST
    Figure CN120215693A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of facial action driving, and particularly relates to a facial action driving method and device, a computer readable storage medium and electronic equipment. The method comprises the following steps: acquiring a target sound for driving a facial action; performing emotion feature extraction processing on the target sound to obtain an emotion feature corresponding to the target sound; based on the emotion features, processing the target sound through a preset three-dimensional facial action model to obtain a three-dimensional facial action corresponding to the target sound; and controlling a preset virtual digital human to execute the three-dimensional face action. Through the method and the device, the emotion features in the sound can be extracted, and the corresponding three-dimensional facial action is generated based on the emotion features in the facial action driving process, so that the facial action can be emotional like a real human facial action, and is more real and natural.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of facial motion driving, and in particular relates to a facial motion driving method, device, computer-readable storage medium, and electronic device. Background Art

[0002] In the prior art, there are already relatively mature methods for driving facial motions through sound. However, the existing facial motion driving methods mainly drive facial motions based on the content in the sound. Although the facial motions are consistent with the content in the sound, compared with real human facial motions, they still appear monotonous and rigid, with poor realism. Summary of the Invention

[0003] In view of this, embodiments of this application provide a facial motion driving method, device, computer-readable storage medium, and electronic device to solve the problem of poor realism existing in the existing facial motion driving methods.

[0004] The first aspect of the embodiments of this application provides a facial motion driving method, which may include:

[0005] Obtain a target sound for driving a facial motion;

[0006] Perform emotional feature extraction processing on the target sound to obtain an emotional feature corresponding to the target sound;

[0007] Based on the emotional feature, process the target sound through a preset three-dimensional facial motion model to obtain a three-dimensional facial motion corresponding to the target sound;

[0008] Control a preset virtual digital human to execute the three-dimensional facial motion.

[0009] In a specific implementation manner of the first aspect, the step of based on the emotional feature, processing the target sound through a preset three-dimensional facial motion model to obtain a three-dimensional facial motion corresponding to the target sound may include:

[0010] Perform sound feature extraction processing on the target sound to obtain a sound feature corresponding to the target sound;

[0011] Perform action encoding processing on the previous model output to obtain first processed data;

[0012] Perform periodic position encoding processing on the first processed data to obtain second processed data;

[0013] Perform biased causal multi-head self-attention processing on the second processed data to obtain third processed data;

[0014] Perform biased cross-modal multi-head self-attention processing on the sound features and the third processed data to obtain fourth processed data;

[0015] Perform fusion processing on the emotion features and the fourth processed data to obtain fifth processed data;

[0016] Perform forward feedback processing on the fifth processed data to obtain sixth processed data;

[0017] Perform action decoding processing on the sixth processed data to obtain the three-dimensional facial action corresponding to the target sound.

[0018] In a specific implementation manner of the first aspect, before performing sound feature extraction processing on the target sound, it may further include:

[0019] Obtain a sound feature extraction model corresponding to the first language;

[0020] Fine-tune the sound feature extraction model using the corpus of the second language to obtain the fine-tuned sound feature extraction model;

[0021] Correspondingly, the performing sound feature extraction processing on the target sound to obtain sound features corresponding to the target sound may include:

[0022] Perform sound feature extraction processing on the target sound through the fine-tuned sound feature extraction model to obtain the sound features corresponding to the target sound.

[0023] In a specific implementation manner of the first aspect, before processing the target sound through a preset three-dimensional facial action model, it may further include:

[0024] Determine the first loss function of the three-dimensional facial action model; wherein, the first loss function is a loss function for truth value constraint of the three-dimensional facial action model;

[0025] Determine the second loss function of the three-dimensional facial action model; wherein, the second loss function is a loss function for inter-frame constraint of the three-dimensional facial action model;

[0026] Fuse the first loss function and the second loss function to obtain the fusion loss function of the three-dimensional facial action model;

[0027] Train the initial three-dimensional facial action model based on the fusion loss function to obtain the trained three-dimensional facial action model.

[0028] In a specific implementation manner of the first aspect, the determining the first loss function of the three-dimensional facial action model may include:

[0029] Determine the output error of the three-dimensional facial motion model; wherein, the output error is the absolute value of the difference between the expected output and the actual output;

[0030] Determine the wing-shaped function of the output error as the first loss function of the three-dimensional facial motion model.

[0031] In a specific implementation manner of the first aspect, determining the second loss function of the three-dimensional facial motion model may include:

[0032] Determine the first inter-frame difference of the three-dimensional facial motion model; wherein, the first inter-frame difference is the difference between two consecutive expected outputs;

[0033] Determine the second inter-frame difference of the three-dimensional facial motion model; wherein, the second inter-frame difference is the difference between two consecutive actual outputs;

[0034] Determine the second loss function of the three-dimensional facial motion model according to the difference between the first inter-frame difference and the second inter-frame difference.

[0035] In a specific implementation manner of the first aspect, before performing the emotional feature extraction process on the target sound, it may further include:

[0036] Obtain an emotional feature extraction model corresponding to the first language;

[0037] Perform low-rank adaptation adjustment on the emotional feature extraction model using the corpus of the second language to obtain the low-rank adaptation adjusted emotional feature extraction model;

[0038] Correspondingly, performing the emotional feature extraction process on the target sound to obtain the emotional feature corresponding to the target sound may include:

[0039] Perform emotional feature extraction process on the target sound through the low-rank adaptation adjusted emotional feature extraction model to obtain the emotional feature corresponding to the target sound.

[0040] A second aspect of the embodiments of the present application provides a facial motion driving device, which may include:

[0041] A sound acquisition module, configured to acquire a target sound for driving facial motion;

[0042] An emotional feature extraction module, configured to perform an emotional feature extraction process on the target sound to obtain an emotional feature corresponding to the target sound;

[0043] A facial action model processing module, configured to process the target sound based on the emotion feature through a preset three-dimensional facial action model to obtain a three-dimensional facial action corresponding to the target sound;

[0044] A facial action execution control module, configured to control a preset virtual digital human to execute the three-dimensional facial action.

[0045] In a specific implementation manner of the second aspect, the facial action model processing module may include:

[0046] A voice feature extraction unit, configured to perform voice feature extraction processing on the target sound to obtain a voice feature corresponding to the target sound;

[0047] An action encoding unit, configured to perform action encoding processing on the previous model output to obtain first processing data;

[0048] A position encoding unit, configured to perform periodic position encoding processing on the first processing data to obtain second processing data;

[0049] A first attention processing unit, configured to perform biased causal multi-head self-attention processing on the second processing data to obtain third processing data;

[0050] A second attention processing unit, configured to perform biased cross-modal multi-head self-attention processing on the voice feature and the third processing data to obtain fourth processing data;

[0051] A data fusion unit, configured to perform fusion processing on the emotion feature and the fourth processing data to obtain fifth processing data;

[0052] A forward feedback unit, configured to perform forward feedback processing on the fifth processing data to obtain sixth processing data;

[0053] An action decoding unit, configured to perform action decoding processing on the sixth processing data to obtain the three-dimensional facial action corresponding to the target sound.

[0054] In a specific implementation manner of the second aspect, the facial action driving device may further include:

[0055] A model fine-tuning module, configured to obtain a voice feature extraction model corresponding to a first language; use the corpus of a second language to fine-tune the voice feature extraction model to obtain the fine-tuned voice feature extraction model;

[0056] Correspondingly, the voice feature extraction unit may be specifically configured to: perform voice feature extraction processing on the target sound through the fine-tuned voice feature extraction model to obtain the voice feature corresponding to the target sound.

[0057] In a specific implementation of the second aspect, the facial action driving device may further include:

[0058] A first loss function determination module, configured to determine a first loss function of the three-dimensional facial action model; wherein, the first loss function is a loss function for truth value constraint of the three-dimensional facial action model;

[0059] A second loss function determination module, configured to determine a second loss function of the three-dimensional facial action model; wherein, the second loss function is a loss function for inter-frame constraint of the three-dimensional facial action model;

[0060] A fused loss function determination module, configured to fuse the first loss function and the second loss function to obtain a fused loss function of the three-dimensional facial action model;

[0061] A model training module, configured to train the initial three-dimensional facial action model based on the fused loss function to obtain the trained three-dimensional facial action model.

[0062] In a specific implementation of the second aspect, the first loss function determination module may specifically be configured to: determine an output error of the three-dimensional facial action model; wherein, the output error is the absolute value of the difference between the expected output and the actual output; determine the wing function of the output error as the first loss function of the three-dimensional facial action model.

[0063] In a specific implementation of the second aspect, the second loss function determination module may specifically be configured to: determine a first inter-frame difference of the three-dimensional facial action model; wherein, the first inter-frame difference is the difference between two consecutive expected outputs; determine a second inter-frame difference of the three-dimensional facial action model; wherein, the second inter-frame difference is the difference between two consecutive actual outputs; determine the second loss function of the three-dimensional facial action model according to the difference between the first inter-frame difference and the second inter-frame difference.

[0064] In a specific implementation of the second aspect, the facial action driving device may further include:

[0065] A low-rank adaptation adjustment module, configured to obtain an emotion feature extraction model corresponding to a first language; perform low-rank adaptation adjustment on the emotion feature extraction model using a corpus of a second language to obtain the low-rank adaptation adjusted emotion feature extraction model;

[0066] Accordingly, the emotion feature extraction module may specifically be configured to: perform emotion feature extraction processing on the target sound through the emotion feature extraction model adjusted by low-rank adaptation, so as to obtain the emotion feature corresponding to the target sound.

[0067] A third aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned facial action driving methods are implemented.

[0068] A fourth aspect of the embodiments of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above-mentioned facial action driving methods are implemented.

[0069] A fifth aspect of the embodiments of the present application provides a computer program product, and when the computer program product runs on an electronic device, the electronic device is enabled to execute the steps of any of the above-mentioned facial action driving methods.

[0070] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The embodiments of the present application obtain a target sound for driving facial actions; perform emotion feature extraction processing on the target sound to obtain an emotion feature corresponding to the target sound; based on the emotion feature, process the target sound through a preset three-dimensional facial action model to obtain a three-dimensional facial action corresponding to the target sound; control a preset virtual digital human to execute the three-dimensional facial action. Through the embodiments of the present application, the emotion feature in the sound can be extracted, and in the process of facial action driving, the corresponding three-dimensional facial action can be generated based on the emotion feature, so that the facial action can be as emotional as a real human facial action and more realistic and natural. Description of the Drawings

[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0072] Figure 1 It is a flowchart of an embodiment of a facial action driving method in the embodiments of the present application;

[0073] Figure 2 It is a schematic diagram of a human three-dimensional facial action based on hybrid deformation expression;

[0074] Figure 3Schematic structural diagram of a three-dimensional facial action model;

[0075] Figure 4 Schematic flowchart for determining a loss function and performing model training;

[0076] Figure 5 Structural diagram of an embodiment of a facial action driving device in an embodiment of the present application;

[0077] Figure 6 Schematic block diagram of an electronic device in an embodiment of the present application. Detailed implementation manners

[0078] To make the objectives, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0079] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0080] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0081] It should be further understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0082] As used in this specification and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0083] In addition, in the description of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0084] In the prior art, there are already relatively mature methods for driving facial movements through sound. However, the existing facial movement driving methods mainly drive facial movements based on the content in the sound. Although the facial movements are consistent with the content in the sound, compared with real human facial movements, they still appear monotonous and rigid, with poor realism.

[0085] In view of this, the embodiments of the present application provide a facial movement driving method, device, computer-readable storage medium, and electronic device to solve the problem of poor realism existing in the existing facial movement driving methods.

[0086] In the embodiments of the present application, the emotional features in the sound can be extracted, and corresponding three-dimensional facial movements can be generated based on the emotional features during the process of driving facial movements, so that the facial movements can be as emotional as real human facial movements and be more realistic and natural.

[0087] The execution subject of the embodiments of the present application can be an electronic device, including but not limited to computing devices such as mobile phones, tablet computers, desktop computers, laptops, handheld computers, robots, and servers.

[0088] Please refer to Figure 1 , an embodiment of a facial movement driving method in the embodiments of the present application may include:

[0089] Step S101, obtain a target sound for driving facial movements.

[0090] The target sound is the sound that the electronic device is about to play. During the process of playing the target sound, a preset virtual digital human needs to synchronously make corresponding facial movements so that its mouth shape and expression are both consistent with the target sound, thereby producing the effect that the virtual digital human is speaking like a human.

[0091] Step S102, perform emotional feature extraction processing on the target sound to obtain emotional features corresponding to the target sound.

[0092] In the embodiments of the present application, any existing emotional feature extraction model in the prior art can be used for emotional feature extraction processing, which may include but is not limited to the BERT model and other models. The embodiments of the present application do not make specific limitations in this regard.

[0093] Considering that although the BERT model can be used as an emotion feature extraction model after training, due to the fact that it requires a large amount of data to achieve good generalization performance, the generalization performance of the model trained on the public dataset is not good. The wav2vec English model can not only extract English content, but also extract certain emotion features. Based on this phenomenon, a more generalizable emotion feature extraction model can be trained by full fine-tuning of the wav2vec English model, which is denoted as wav2vec_emo_english.

[0094] In a specific implementation manner of the embodiment of the present application, the wav2vec_emo_english can be directly used to perform emotion feature extraction processing on the target sound, so as to obtain the emotion features corresponding to the target sound.

[0095] In another specific implementation manner of the embodiment of the present application, considering that although wav2vec_emo_english has good emotion generalization for other non-English languages (such as Chinese), the spatial aggregation is not absolutely ideal. Therefore, feature compensation can be further introduced to repair this deficiency in spatial aggregation. When the language corresponding to the existing emotion feature extraction model (such as wav2vec_emo_english) (denoted as the first language) is inconsistent with the language corresponding to the target sound (denoted as the second language), an emotion feature extraction model corresponding to the first language can be obtained, and the emotion feature extraction model can be adjusted by Low-Rank Adaptation (LoRA) using the corpus of the second language to obtain the LoRA-adjusted emotion feature extraction model.

[0096] LoRA adjustment simulates the change amount of parameters through low-rank decomposition, so as to indirectly train a large model with a very small number of parameters. By multiplying the two matrices A and B before and after, the first matrix A is responsible for dimensionality reduction, the second matrix B is responsible for dimensionality increase, the middle layer dimension is r, the trainable layer dimension is the same as the pre-trained model layer dimension d. First, the dimension d is reduced to r through a fully connected layer, and then mapped back to the d dimension from r through a fully connected layer, where r << d, r is the rank of the matrix. In this way, the matrix calculation changes from d * d to d * r + r * d, and the number of parameters is greatly reduced.

[0097] For an existing emotional feature extraction model, LoRA adjustment can be performed on the query vector Query (denoted as Q), key vector Key (denoted as K), and value vector Value (denoted as V) involved in its self-attention mechanism. Taking Q as an example, the left branch is the original frozen network, and the right branch introduces parameters W_A and W_B for learning. Among them, the fully connected layer parameters of W_A are (d, r), and the fully connected layer parameters of W_B are (r, d). The parameters of the right branch are learnable, denoted as Q 右 , then the updated query vector (denoted as Q ′ ) can be obtained: Q ′ = Q + Q 右 . The adjustment processes of K and V are the same as that of Q. The updated key vector (denoted as K ′ ) and value vector (denoted as V ′ ) can be obtained, which will not be elaborated here. When applying, use Q ′ , K ′ and V ′ to replace the original Q, K, and V.

[0098] When performing emotional feature extraction processing on the target sound, the LoRA-adjusted emotional feature extraction model can be used to perform emotional feature extraction processing on the target sound, so as to obtain the emotional features corresponding to the target sound.

[0099] Step S103: Based on the emotional features, process the target sound through a preset three-dimensional facial motion model to obtain the three-dimensional facial motion corresponding to the target sound.

[0100] Among them, the three-dimensional facial motion can be a human three-dimensional facial motion expressed based on BlendShape (BS). The reason why human facial motions are rich and diverse is that some points on the face have displaced in space. For example, when some points at the corners of the mouth move obliquely upward, visually, a person has a smiling facial motion. Based on this principle, the human facial motion system can be designed to extract and produce some basic facial motions, which can be regarded as facial deformations of people in different emotional states. With these basic facial motions, various composite facial motions can be combined by adjusting their weight coefficients. The magnitude of the weight coefficients determines the contribution degree of each basic facial motion in the composite facial motion.

[0101] Figure 2The figure shows a schematic diagram of human three-dimensional facial actions expressed based on BlendShape. Denote the number of basic facial actions as N, where N is a positive integer, and its specific value can be flexibly set according to the actual situation. For example, it can be set to 52 or other values, and the embodiments of this application do not make specific limitations on this. Under different combinations of weight coefficients, linear weighting of these basic facial actions can obtain various different composite facial actions.

[0102] Since there are already many models for generating corresponding facial actions according to sound in the prior art, but their three-dimensional facial actions are expressed based on vertices or 3DMM. In the embodiments of this application, based on these models, the three-dimensional facial actions can be expressed based on BlendShape, so as to obtain the three-dimensional facial action model (3D Face Animation) in the embodiments of this application.

[0103] The three-dimensional facial action model can adopt the autoregressive method, that is, taking the closed-mouth state as the initialization, and continuously predicting the current state according to the previous prediction and the current speech features. As an example, Figure 3 A possible schematic diagram of the model structure of the three-dimensional facial action model is shown. As shown in the figure, the three-dimensional facial action model can include but is not limited to a sound feature extraction module, an action encoding module (BS Encoder), a periodic positional encoding module (Periodic Positional Encoding), a biased causal multi-head self-attention module (Biased Causal Multi-Head Self-Attention), a biased cross-modal multi-head self-attention module (Biased Cross-Modal Multi-Head Self-Attention), a biased cross-emotion multi-head self-attention module (Biased Cross-Emo Multi-Head Self-Attention), a feed-forward module (Feed Forward), and an action decoding module (BSDecoder), etc.

[0104] Compared with the prior art, the core difference of this model structure is that a biased cross-emotion multi-head self-attention module is added, and the emotional features and the output of the biased cross-modal multi-head self-attention module can be fused through the self-attention mechanism, as shown in the following formula:

[0105]

[0106] Among them, is the Q value corresponding to the output of the biased cross-modal multi-head self-attention module, K E and V EThe K value and V value corresponding to the emotional feature, B A is the alignment bias, Att is the self-attention mechanism processing function, d k is K E The number of dimensions of

[0107] Based on Figure 3 the model structure shown, the process of processing the target sound through the three-dimensional facial action model may include: performing sound feature extraction processing on the target sound to obtain the sound features corresponding to the target sound; performing action encoding processing on the previous model output to obtain the first processed data; performing periodic position encoding processing on the first processed data to obtain the second processed data; performing biased causal multi-head self-attention processing on the second processed data to obtain the third processed data; performing biased cross-modal multi-head self-attention processing on the sound features and the third processed data to obtain the fourth processed data; performing fusion processing on the emotional feature and the fourth processed data to obtain the fifth processed data; performing forward feedback processing on the fifth processed data to obtain the sixth processed data; performing action decoding processing on the sixth processed data to obtain the three-dimensional facial action corresponding to the target sound. The specific processing process may refer to the model for generating the corresponding facial action according to the sound in the prior art, and the embodiments of the present application will not elaborate on this.

[0108] It should be noted that, in the embodiments of the present application, the input of the action encoding module (i.e., the previous facial action) and the output of the action decoding module (i.e., the currently predicted facial action) are both human three-dimensional facial actions expressed based on BlendShape. In the human three-dimensional facial actions expressed based on BlendShape, since the weight coefficients of the basic facial actions are all positive values, therefore, after the action decoding processing, a sigmoid activation function processing process can be added to ensure that the weight coefficients are all positive values.

[0109] In the sound feature extraction module, any one of the sound feature extraction models in the prior art can be used to perform the sound feature extraction processing, which may include but is not limited to the wav2vec model and other models, and the embodiments of the present application do not make specific limitations on this.

[0110] When the language (the first language) corresponding to the existing sound feature extraction model is the same as the language (the second language) corresponding to the target sound, the existing sound feature extraction model can be directly used to perform the sound feature extraction processing on the target sound, so as to obtain the sound features corresponding to the target sound.

[0111] In the case where the first language is inconsistent with the second language, based on the existing voice feature extraction model, the corpus of the second language can be used to fine-tune it, so as to obtain a fine-tuned voice feature extraction model, thereby improving the adaptability to the second language. For example, the language corresponding to the existing voice feature extraction model can be English, and the language corresponding to the target voice can be Chinese. Then, the existing voice feature extraction model can be fine-tuned using the Chinese corpus, so as to obtain a fine-tuned voice feature extraction model, thereby improving the adaptability to Chinese. When performing voice feature extraction processing, the fine-tuned voice feature extraction model can be used to perform voice feature extraction processing on the target voice, so as to obtain the voice features corresponding to the target voice.

[0112] The three-dimensional facial action model can be pre-trained through the corresponding data set, and the loss function used in the training process can be flexibly set according to the actual situation. The embodiments of the present application do not make specific limitations on this.

[0113] In a specific implementation manner of the embodiments of the present application, it can be determined according to the process as Figure 4 shown to determine the loss function and perform model training:

[0114] Step S401: Determine the first loss function of the three-dimensional facial action model.

[0115] Among them, the first loss function is the loss function for performing groundtruth constraint on the three-dimensional facial action model. Specifically, the output error of the three-dimensional facial action model can be determined, and the wing function of the output error can be determined as the first loss function of the three-dimensional facial action model, as shown in the following formula:

[0116]

[0117] Among them, y is the expected output of the three-dimensional facial action model, that is, the groundtruth, is the actual output of the three-dimensional facial action model, is the output error, that is, the absolute value of the difference between the expected output and the actual output, WingLoss is the wing function, WingLoss has a much higher response to subtle differences than other functions, and can more finely optimize the tiny errors of facial action changes, and L1 is the first loss function of the three-dimensional facial action model.

[0118] Step S402: Determine the second loss function of the three-dimensional facial action model.

[0119] Among them, the second loss function is the loss function for performing inter-frame constraints on the 3D facial action model. Specifically, the first inter-frame difference and the second inter-frame difference of the 3D facial action model can be determined respectively, and the second loss function of the 3D facial action model can be determined according to the difference between the first inter-frame difference and the second inter-frame difference, as shown in the following formula:

[0120]

[0121] Among them, (y i -y i-1 ) is the first inter-frame difference, that is, the difference between two consecutive expected outputs, is the second inter-frame difference, that is, the difference between two consecutive actual outputs, is the difference between the first inter-frame difference and the second inter-frame difference, and L2 is the second loss function of the 3D facial action model.

[0122] Step S403: Fuse the first loss function and the second loss function to obtain the fused loss function of the 3D facial action model.

[0123] The specific fusion method can be flexibly set according to the actual situation, and the embodiments of the present application do not make specific limitations on this. As an example, for the sake of simplicity, the sum of the first loss function and the second loss function can be directly used as the fused loss function of the 3D facial action model.

[0124] Step S404: Train the initial 3D facial action model based on the fused loss function to obtain the trained 3D facial action model.

[0125] In the embodiments of the present application, it can be trained with a sufficient number of training samples, and each training sample includes a set of voice samples and corresponding 3D facial action labels. Using the voice samples of the training samples as inputs and the corresponding 3D facial action labels as expected outputs, the 3D facial action model can be trained to obtain the trained 3D facial action model.

[0126] During the training process, for each training sample, the 3D facial action model can be used to process the voice samples of the training sample to obtain the actual output of the training sample, and then the fused loss function can be used to calculate the training loss value according to the expected output and the actual output in the training sample. After calculating the training loss value, the model parameters of the 3D facial action model can be adjusted according to the training loss value.

[0127] In an embodiment of the present application, it is assumed that in the initial state, the model parameters of the three-dimensional facial action model are W1. The training loss value is backpropagated to modify the model parameters W1 of the three-dimensional facial action model, and the modified model parameters W2 are obtained. After modifying the parameters, the next training process is continued. During this training process, the training loss value is recalculated, and the training loss value is backpropagated to modify the model parameters W2 of the three-dimensional facial action model, and the modified model parameters W3 are obtained, and so on. The above process is continuously repeated, and the model parameters can be modified in each training process until the preset training conditions are met. Among them, the training conditions can be that the number of training times reaches a preset number threshold, and the number threshold can be set according to the actual situation. For example, it can be set to several thousand, tens of thousands, hundreds of thousands or even larger values; the training conditions can also be that the three-dimensional facial action model converges; since it is possible that the number of training times has not reached the number threshold, but the three-dimensional facial action model has already converged, which may lead to unnecessary repeated work; or the three-dimensional facial action model may never converge, which may lead to an infinite loop and the training process cannot end. Based on the above two situations, the training conditions can also be that the number of training times reaches the number threshold or the three-dimensional facial action model converges. When the training conditions are met, the trained three-dimensional facial action model can be obtained.

[0128] Through the above process, the voice sample of the training sample and the corresponding three-dimensional facial action label are used as the learning objects of the three-dimensional facial action model. After the training process, the three-dimensional facial action model can establish a mapping relationship between the voice sample and the three-dimensional facial action label, so that when facing a new voice, the corresponding three-dimensional facial action can also be obtained according to this mapping relationship.

[0129] After the training of the three-dimensional facial action model is completed, the three-dimensional facial action model can be used to process the target voice, so as to obtain the three-dimensional facial action corresponding to the target voice.

[0130] Step S104: Control the preset virtual digital human to perform three-dimensional facial actions.

[0131] A virtual digital human (Digital Human / Meta Human) refers to a digital human image created by using digital technology and similar to the human image. After obtaining the three-dimensional facial action corresponding to the target voice, the virtual digital human can be controlled to generate the corresponding BlendShape expression, so that the virtual digital human can show three-dimensional facial actions as close as possible to the real facial actions of humans.

[0132] In summary, the embodiment of the present application obtains a target sound for driving facial movements; performs emotional feature extraction processing on the target sound to obtain emotional features corresponding to the target sound; based on the emotional features, processes the target sound through a preset three-dimensional facial movement model to obtain a three-dimensional facial movement corresponding to the target sound; and controls a preset virtual digital human to execute the three-dimensional facial movement. Through the embodiment of the present application, emotional features in the sound can be extracted, and during the process of facial movement driving, corresponding three-dimensional facial movements can be generated based on the emotional features, making the facial movements as emotional as real human facial movements and more natural and realistic.

[0133] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0134] Corresponding to the facial movement driving method described in the above embodiments, Figure 5 FIG. shows a structural diagram of an embodiment of a facial movement driving device provided by an embodiment of the present application.

[0135] In this embodiment, a facial movement driving device may include:

[0136] A sound acquisition module 501, configured to acquire a target sound for driving facial movements;

[0137] An emotional feature extraction module 502, configured to perform emotional feature extraction processing on the target sound to obtain emotional features corresponding to the target sound;

[0138] A facial movement model processing module 503, configured to process the target sound through a preset three-dimensional facial movement model based on the emotional features to obtain a three-dimensional facial movement corresponding to the target sound;

[0139] A facial movement execution control module 504, configured to control a preset virtual digital human to execute the three-dimensional facial movement.

[0140] In a specific implementation manner of the embodiment of the present application, the facial movement model processing module may include:

[0141] A sound feature extraction unit, configured to perform sound feature extraction processing on the target sound to obtain sound features corresponding to the target sound;

[0142] An action encoding unit, configured to perform action encoding processing on the previous model output to obtain first processed data;

[0143] A position encoding unit, configured to perform periodic position encoding processing on the first processed data to obtain second processed data;

[0144] A first attention processing unit for performing biased causal multi-head self-attention processing on the second processed data to obtain third processed data;

[0145] A second attention processing unit for performing biased cross-modal multi-head self-attention processing on the voice feature and the third processed data to obtain fourth processed data;

[0146] A data fusion unit for fusing the emotion feature and the fourth processed data to obtain fifth processed data;

[0147] A forward feedback unit for performing forward feedback processing on the fifth processed data to obtain sixth processed data;

[0148] An action decoding unit for performing action decoding processing on the sixth processed data to obtain the three-dimensional facial action corresponding to the target voice.

[0149] In a specific implementation manner of the embodiment of the present application, the facial action driving device may further include:

[0150] A model fine-tuning module for obtaining a voice feature extraction model corresponding to a first language; fine-tuning the voice feature extraction model with the corpus of a second language to obtain the fine-tuned voice feature extraction model;

[0151] Correspondingly, the voice feature extraction unit may specifically be used for: performing voice feature extraction processing on the target voice through the fine-tuned voice feature extraction model to obtain the voice feature corresponding to the target voice.

[0152] In a specific implementation manner of the embodiment of the present application, the facial action driving device may further include:

[0153] A first loss function determination module for determining a first loss function of the three-dimensional facial action model; wherein, the first loss function is a loss function for performing true value constraint on the three-dimensional facial action model;

[0154] A second loss function determination module for determining a second loss function of the three-dimensional facial action model; wherein, the second loss function is a loss function for performing inter-frame constraint on the three-dimensional facial action model;

[0155] A fusion loss function determination module for fusing the first loss function and the second loss function to obtain a fusion loss function of the three-dimensional facial action model;

[0156] A model training module, configured to train the initial three-dimensional facial action model based on the fusion loss function to obtain the trained three-dimensional facial action model.

[0157] In a specific implementation manner of the embodiment of the present application, the first loss function determination module may be specifically configured to: determine the output error of the three-dimensional facial action model; wherein, the output error is the absolute value of the difference between the expected output and the actual output; determine the wing function of the output error as the first loss function of the three-dimensional facial action model.

[0158] In a specific implementation manner of the embodiment of the present application, the second loss function determination module may be specifically configured to: determine the first inter-frame difference of the three-dimensional facial action model; wherein, the first inter-frame difference is the difference between two consecutive expected outputs; determine the second inter-frame difference of the three-dimensional facial action model; wherein, the second inter-frame difference is the difference between two consecutive actual outputs; determine the second loss function of the three-dimensional facial action model according to the difference between the first inter-frame difference and the second inter-frame difference.

[0159] In a specific implementation manner of the embodiment of the present application, the facial action driving device may further include:

[0160] A low-rank adaptation adjustment module, configured to obtain an emotion feature extraction model corresponding to a first language; perform low-rank adaptation adjustment on the emotion feature extraction model using a corpus of a second language to obtain the low-rank adaptation adjusted emotion feature extraction model;

[0161] Correspondingly, the emotion feature extraction module may be specifically configured to: perform emotion feature extraction processing on the target sound through the low-rank adaptation adjusted emotion feature extraction model to obtain the emotion feature corresponding to the target sound.

[0162] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices, modules, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein again.

[0163] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0164] Figure 6 A schematic block diagram of an electronic device provided by an embodiment of the present application is shown. For the sake of illustration, only the parts related to the embodiment of the present application are shown.

[0165] Such as Figure 6As shown, the electronic device 6 of this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, it implements the steps in the above-mentioned embodiments of various facial action driving methods, such as Figure 1 the steps S101 to S104 shown. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each module / unit in the above-mentioned device embodiments, such as Figure 5 the functions of the modules 501 to 504 shown.

[0166] Exemplarily, the computer program 62 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 61 and executed by the processor 60 to complete this application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 62 in the electronic device 6.

[0167] The electronic device 6 can include, but is not limited to, computing devices such as mobile phones, tablet computers, desktop computers, laptops, handheld computers, robots, and servers. Those skilled in the art can understand that Figure 6 these are merely examples of the electronic device 6 and do not constitute a limitation on the electronic device 6. It may include more or fewer components than shown, or combine certain components, or have different components. For example, the electronic device 6 may further include input / output devices, network access devices, buses, etc.

[0168] The processor 60 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0169] The memory 61 may be an internal storage unit of the electronic device 6, such as a hard disk or memory of the electronic device 6. The memory 61 may also be an external storage device of the electronic device 6, such as a plug-in hard disk equipped on the electronic device 6, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 61 may also include both an internal storage unit and an external storage device of the electronic device 6. The memory 61 is used to store the computer program and other programs and data required by the electronic device 6. The memory 61 may also be used to temporarily store data that has been output or is to be output.

[0170] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be described in detail here.

[0171] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0172] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0173] In the embodiments provided in the present application, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0174] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0175] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0176] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing the relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0177] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A facial action driving method, characterized in that: include: Obtaining a target sound for driving facial movements; Performing an emotion feature extraction process on the target sound to obtain an emotion feature corresponding to the target sound; Based on the emotional features, the target sound is processed by a preset three-dimensional facial action model to obtain a three-dimensional facial action corresponding to the target sound; Control the preset virtual digital human to perform the three-dimensional facial action.

2. The facial action driving method according to claim 1, characterized in that: The method of processing the target sound based on the emotion feature by using a preset three-dimensional facial action model to obtain a three-dimensional facial action corresponding to the target sound includes: Performing sound feature extraction processing on the target sound to obtain a sound feature corresponding to the target sound; Performing action encoding processing on the previous model output to obtain first processed data; Performing periodic position encoding processing on the first processed data to obtain second processed data; Performing biased causal multi-head self-attention processing on the second processed data to obtain third processed data; performing biased cross-modal multi-head self-attention processing on the sound feature and the third processed data to obtain fourth processed data; Performing fusion processing on the emotion feature and the fourth processed data to obtain fifth processed data; Performing forward feedback processing on the fifth processed data to obtain sixth processed data; The sixth processed data is subjected to motion decoding processing to obtain the three-dimensional facial motion corresponding to the target sound.

3. The facial action driving method according to claim 2, characterized in that: Before performing sound feature extraction processing on the target sound, the method further includes: Acquire a sound feature extraction model corresponding to the first language; Using the corpus of the second language to fine-tune the sound feature extraction model to obtain the fine-tuned sound feature extraction model; Accordingly, the performing sound feature extraction processing on the target sound to obtain a sound feature corresponding to the target sound includes: The target sound is subjected to sound feature extraction processing by using the finely tuned sound feature extraction model to obtain the sound feature corresponding to the target sound.

4. The facial action driving method according to claim 1, characterized in that: Before processing the target sound through the preset three-dimensional facial action model, the method further includes: Determine a first loss function of the three-dimensional facial action model; wherein the first loss function is a loss function for performing truth constraints on the three-dimensional facial action model; Determine a second loss function of the three-dimensional facial action model; wherein the second loss function is a loss function for performing inter-frame constraints on the three-dimensional facial action model; fusing the first loss function and the second loss function to obtain a fusion loss function of the three-dimensional facial action model; The initial three-dimensional facial action model is trained based on the fusion loss function to obtain the trained three-dimensional facial action model.

5. The facial action driving method according to claim 4, characterized in that: The determining of a first loss function of the three-dimensional facial action model comprises: Determining an output error of the three-dimensional facial action model; wherein the output error is an absolute value of a difference between an expected output and an actual output; A wing function of the output error is determined as the first loss function of the three-dimensional facial action model.

6. The facial action driving method according to claim 4, characterized in that: The determining of a second loss function of the three-dimensional facial action model comprises: Determining a first inter-frame difference of the three-dimensional facial action model; wherein the first inter-frame difference is a difference between two consecutive frames of expected output; Determine a second inter-frame difference of the three-dimensional facial action model; wherein the second inter-frame difference is a difference between two consecutive frames of actual output; The second loss function of the 3D facial action model is determined according to a difference between the first inter-frame difference and the second inter-frame difference.

7. The facial action driving method according to any one of claims 1 to 6, characterized in that: Before performing emotional feature extraction processing on the target sound, the method further includes: Obtaining a sentiment feature extraction model corresponding to the first language; Using the corpus of the second language to perform low-rank adaptation adjustment on the sentiment feature extraction model, to obtain the sentiment feature extraction model after low-rank adaptation adjustment; Accordingly, the performing of emotional feature extraction processing on the target sound to obtain the emotional feature corresponding to the target sound includes: The target sound is subjected to emotional feature extraction processing by using the emotional feature extraction model adjusted by low-rank adaptation to obtain the emotional feature corresponding to the target sound.

8. A facial action driving device, characterized in that: include: A sound acquisition module, used to acquire a target sound for driving facial movements; An emotion feature extraction module, used to perform emotion feature extraction processing on the target sound to obtain an emotion feature corresponding to the target sound; A facial action model processing module, used to process the target sound through a preset three-dimensional facial action model based on the emotional features to obtain a three-dimensional facial action corresponding to the target sound; The facial action execution control module is used to control the preset virtual digital human to execute the three-dimensional facial action.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the facial action driving method according to any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the facial action driving method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Face animation driving method and apparatus, and readable storage medium and electronic device

    WO2026179392A1