Facial action driving method and device, computer readable storage medium and robot

By processing the target sound into human three-dimensional facial movements and mapping the robot three-dimensional facial movements, the problem that the existing technology cannot be applied to bionic robots is solved, and the concrete control of facial movements is achieved.

CN120215694APending Publication Date: 2025-06-27UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510233727.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing facial action driving method cannot be applied to physical bionic robots because it is difficult to disassemble facial action to the control of the motor in a concrete way.

Method used

By obtaining the target sound, using the preset three-dimensional facial action model to process the sound, obtain the human three-dimensional facial action, and then perform mapping processing to obtain the robot three-dimensional facial action, and finally control the facial motor to perform the action.

Benefits of technology

It realizes the mapping of human three-dimensional facial movements into robot three-dimensional facial movements, thereby decomposing facial movements into motor control in a concrete way, which is suitable for physical bionic robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215694A_ABST
    Figure CN120215694A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of facial action driving, and particularly relates to a facial action driving method and device, a computer readable storage medium and a robot. The method comprises the following steps: acquiring a target sound for driving a facial action; processing the target sound through a preset three-dimensional facial action model to obtain a first facial action corresponding to the target sound; wherein the first facial action is a human three-dimensional facial action based on hybrid deformation expression; performing mapping processing on the first facial action to obtain a second facial action corresponding to the first facial action; wherein the second facial action is a robot three-dimensional facial action based on hybrid deformation expression; and controlling a face motor of the robot to execute the second face action. According to the method and the device, the three-dimensional facial action based on hybrid deformation expression is adopted, and the human three-dimensional facial action is mapped into the robot three-dimensional facial action, so that the method and the device can be better applied to an entity bionic robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of facial motion driving, and particularly relates to a facial motion driving method, device, computer-readable storage medium, and robot. Background Art

[0002] In the prior art, there are already relatively mature methods for driving facial motions through sound. However, these existing facial motion driving methods mainly target virtual digital humans, and their three-dimensional facial motions are expressed based on vertices or 3D Morphable Models (3DMMs). In this processing method, it is difficult to disassemble facial motions into motor control in a concrete manner, so it cannot be applied to physical bionic robots. Summary of the Invention

[0003] In view of this, embodiments of this application provide a facial motion driving method, device, computer-readable storage medium, and robot to solve the problem that existing facial motion driving methods cannot be applied to physical bionic robots.

[0004] The first aspect of the embodiments of this application provides a facial motion driving method, which may include:

[0005] Obtain a target sound for driving a facial motion;

[0006] Process the target sound through a preset three-dimensional facial motion model to obtain a first facial motion corresponding to the target sound; wherein, the first facial motion is a human three-dimensional facial motion expressed based on blend shapes.

[0007] Perform a mapping process on the first facial motion to obtain a second facial motion corresponding to the first facial motion; wherein, the second facial motion is a robot three-dimensional facial motion expressed based on blend shapes.

[0008] Control the facial motors of the robot to execute the second facial motion.

[0009] In a specific implementation manner of the first aspect, the performing a mapping process on the first facial motion to obtain a second facial motion corresponding to the first facial motion may include:

[0010] Perform a mapping process on the first facial motion through a preset facial motion mapping network to obtain the second facial motion corresponding to the first facial motion;

[0011] wherein, the facial motion mapping network is a neural network pre-trained for performing facial motion mapping.

[0012] In a specific implementation of the first aspect, before processing the target sound through a preset three-dimensional facial action model, it may further include:

[0013] Determine a first loss function of the three-dimensional facial action model; wherein, the first loss function is a loss function for true value constraint of the three-dimensional facial action model;

[0014] Determine a second loss function of the three-dimensional facial action model; wherein, the second loss function is a loss function for inter-frame constraint of the three-dimensional facial action model;

[0015] Fuse the first loss function and the second loss function to obtain a fused loss function of the three-dimensional facial action model;

[0016] Train the initial three-dimensional facial action model based on the fused loss function to obtain the trained three-dimensional facial action model.

[0017] In a specific implementation of the first aspect, the determining the first loss function of the three-dimensional facial action model may include:

[0018] Determine the output error of the three-dimensional facial action model; wherein, the output error is the absolute value of the difference between the expected output and the actual output;

[0019] Determine the wing function of the output error as the first loss function of the three-dimensional facial action model.

[0020] In a specific implementation of the first aspect, the determining the second loss function of the three-dimensional facial action model may include:

[0021] Determine a first inter-frame difference of the three-dimensional facial action model; wherein, the first inter-frame difference is the difference between two consecutive expected outputs;

[0022] Determine a second inter-frame difference of the three-dimensional facial action model; wherein, the second inter-frame difference is the difference between two consecutive actual outputs;

[0023] Determine the second loss function of the three-dimensional facial action model according to the difference between the first inter-frame difference and the second inter-frame difference.

[0024] In a specific implementation of the first aspect, the processing the target sound to obtain a first facial action corresponding to the target sound may include:

[0025] Perform sound feature extraction processing on the target sound to obtain a sound feature corresponding to the target sound;

[0026] Perform action encoding processing on the previous model output to obtain first processed data;

[0027] Perform periodic positional encoding processing on the first processed data to obtain second processed data;

[0028] Perform biased causal multi-head self-attention processing on the second processed data to obtain third processed data;

[0029] Perform biased cross-modal multi-head self-attention processing on the voice feature and the third processed data to obtain fourth processed data;

[0030] Perform forward feedback processing on the fourth processed data to obtain fifth processed data;

[0031] Perform action decoding processing on the fifth processed data to obtain the first facial action corresponding to the target voice.

[0032] In a specific implementation manner of the first aspect, before performing voice feature extraction processing on the target voice, it may further include:

[0033] Obtain a voice feature extraction model corresponding to the first language;

[0034] Fine-tune the voice feature extraction model using the corpus of the second language to obtain the fine-tuned voice feature extraction model;

[0035] Correspondingly, the performing voice feature extraction processing on the target voice to obtain the voice feature corresponding to the target voice may include:

[0036] Perform voice feature extraction processing on the target voice through the fine-tuned voice feature extraction model to obtain the voice feature corresponding to the target voice.

[0037] A second aspect of the embodiments of the present application provides a facial action driving device, which may include:

[0038] A voice acquisition module, configured to acquire a target voice for driving facial actions;

[0039] A facial action model processing module, configured to process the target voice through a preset three-dimensional facial action model to obtain a first facial action corresponding to the target voice; wherein, the first facial action is a human three-dimensional facial action based on blend shape expression;

[0040] A facial action mapping module, configured to perform mapping processing on the first facial action to obtain a second facial action corresponding to the first facial action; wherein, the second facial action is a robot three-dimensional facial action based on blend shape expression;

[0041] A facial motor control module for controlling the facial motors of the robot to perform the second facial action.

[0042] In a specific implementation of the second aspect, the facial action mapping module may specifically be used to: perform mapping processing on the first facial action through a preset facial action mapping network to obtain the second facial action corresponding to the first facial action; wherein, the facial action mapping network is a neural network pre-trained for facial action mapping.

[0043] In a specific implementation of the second aspect, the facial action driving device may further include:

[0044] A first loss function determination module for determining the first loss function of the three-dimensional facial action model; wherein, the first loss function is a loss function for truth value constraint of the three-dimensional facial action model;

[0045] A second loss function determination module for determining the second loss function of the three-dimensional facial action model; wherein, the second loss function is a loss function for inter-frame constraint of the three-dimensional facial action model;

[0046] A fused loss function determination module for fusing the first loss function and the second loss function to obtain the fused loss function of the three-dimensional facial action model;

[0047] A facial action model training module for training the initial three-dimensional facial action model based on the fused loss function to obtain the trained three-dimensional facial action model.

[0048] In a specific implementation of the second aspect, the first loss function determination module may specifically be used to: determine the output error of the three-dimensional facial action model; wherein, the output error is the absolute value of the difference between the expected output and the actual output; determine the wing function of the output error as the first loss function of the three-dimensional facial action model.

[0049] In a specific implementation of the second aspect, the second loss function determination module may specifically be used to: determine the first inter-frame difference of the three-dimensional facial action model; wherein, the first inter-frame difference is the difference between two consecutive expected outputs; determine the second inter-frame difference of the three-dimensional facial action model; wherein, the second inter-frame difference is the difference between two consecutive actual outputs; determine the second loss function of the three-dimensional facial action model according to the difference between the first inter-frame difference and the second inter-frame difference.

[0050] In a specific implementation manner of the second aspect, the facial action model processing module may include:

[0051] A voice feature extraction unit, configured to perform voice feature extraction processing on the target voice to obtain voice features corresponding to the target voice;

[0052] An action encoding unit, configured to perform action encoding processing on the previous model output to obtain first processed data;

[0053] A position encoding unit, configured to perform periodic position encoding processing on the first processed data to obtain second processed data;

[0054] A first attention processing unit, configured to perform biased causal multi-head self-attention processing on the second processed data to obtain third processed data;

[0055] A second attention processing unit, configured to perform biased cross-modal multi-head self-attention processing on the voice features and the third processed data to obtain fourth processed data;

[0056] A forward feedback unit, configured to perform forward feedback processing on the fourth processed data to obtain fifth processed data;

[0057] An action decoding unit, configured to perform action decoding processing on the fifth processed data to obtain the first facial action corresponding to the target voice.

[0058] In a specific implementation manner of the second aspect, the facial action driving device may further include:

[0059] A voice feature extraction model fine-tuning module, configured to obtain a voice feature extraction model corresponding to a first language; use a corpus of a second language to fine-tune the voice feature extraction model to obtain the fine-tuned voice feature extraction model;

[0060] Correspondingly, the voice feature extraction unit may specifically be configured to: perform voice feature extraction processing on the target voice through the fine-tuned voice feature extraction model to obtain the voice features corresponding to the target voice.

[0061] A third aspect of the embodiments of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned facial action driving methods are implemented.

[0062] A fourth aspect of the embodiments of the present application provides a robot, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of any of the above-mentioned facial action driving methods are implemented.

[0063] The fifth aspect of the embodiments of the present application provides a computer program product, which, when running on a robot, enables the robot to execute the steps of any of the above-mentioned facial action driving methods.

[0064] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: The embodiments of the present application obtain a target sound for driving facial actions; process the target sound through a preset three-dimensional facial action model to obtain a first facial action corresponding to the target sound; wherein, the first facial action is a human three-dimensional facial action expressed based on blend shapes; perform a mapping process on the first facial action to obtain a second facial action corresponding to the first facial action; wherein, the second facial action is a robot three-dimensional facial action expressed based on blend shapes; control the facial motors of the robot to execute the second facial action. Through the embodiments of the present application, three-dimensional facial actions expressed based on blend shapes are adopted, and human three-dimensional facial actions can be mapped to robot three-dimensional facial actions, so that facial actions can be disassembled into motor control in a concrete manner, which can be better applied to physical bionic robots. Description of the Drawings

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0066] Figure 1 It is a flowchart of an embodiment of a facial action driving method in the embodiments of the present application;

[0067] Figure 2 It is a schematic diagram of a human three-dimensional facial action expressed based on blend shapes;

[0068] Figure 3 It is a schematic diagram of the model structure of a three-dimensional facial action model;

[0069] Figure 4 It is a schematic flowchart for determining a loss function and performing model training;

[0070] Figure 5 It is a structural diagram of an embodiment of a facial action driving device in the embodiments of the present application;

[0071] Figure 6 It is a schematic block diagram of a robot in the embodiments of the present application. Detailed Embodiments

[0072] To make the objectives, features, and advantages of the present application more apparent and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0073] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0074] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0075] It should be further understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0076] As used in this specification and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0077] In addition, in the description of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0078] In the prior art, there are already relatively mature methods for driving facial movements through sound. However, these existing facial movement driving methods mainly target virtual digital humans, and their three-dimensional facial movements are expressed based on vertices or three-dimensional deformation models (3D Morphable Model, 3DMM). In this processing method, it is difficult to disassemble the facial movements concretely to the control of the motor, so it cannot be applied to physical bionic robots.

[0079] In view of this, embodiments of the present application provide a facial action driving method, device, computer-readable storage medium, and robot, so as to solve the problem that existing facial action driving methods are not applicable to physical bionic robots.

[0080] In the embodiments of the present application, three-dimensional facial actions based on blend shape expressions are adopted, and human three-dimensional facial actions can be mapped to robot three-dimensional facial actions, so that facial actions can be disassembled into motor control in a figurative manner, which is more applicable to physical bionic robots.

[0081] The execution subject of the embodiments of the present application can be a robot, especially a physical bionic robot that can make various human-like facial actions.

[0082] Please refer to Figure 1 , an embodiment of a facial action driving method in the embodiments of the present application may include:

[0083] Step S101, obtain a target sound for driving a facial action.

[0084] The target sound is the sound that the robot is going to play. During the process of playing the target sound, the robot needs to synchronously make corresponding facial actions so that its mouth shape and expression are both consistent with the target sound, thereby producing the effect that the robot is speaking like a human.

[0085] Step S102, process the target sound through a preset three-dimensional facial action model to obtain a first facial action corresponding to the target sound.

[0086] Among them, the first facial action is a human three-dimensional facial action based on blend shape (BS) expression. The reason why human facial actions are rich and diverse is that some points on the face have displaced in space. For example, when some points at the corners of the mouth move obliquely upward, visually, a person has a smiling facial action. Based on this principle, the human facial action system can be designed to extract and produce some basic facial actions, which can be regarded as facial deformations of people in different emotional states. With these basic facial actions, various composite facial actions can be combined by adjusting their weight coefficients, and the size of the weight coefficient determines the contribution degree of each basic facial action in the composite facial action.

[0087] Figure 2The figure shows a schematic diagram of human 3D facial movements expressed based on BlendShape. Denote the number of basic facial movements as N, where N is a positive integer, and its specific value can be flexibly set according to the actual situation. For example, it can be set to 52 or other values, and the embodiments of the present application do not make specific limitations on this. Under different combinations of weight coefficients, linear weighting of these basic facial movements can obtain various different composite facial movements.

[0088] Since there are already many models for generating corresponding facial movements according to sounds in the prior art, but their 3D facial movements are expressed based on vertices or 3DMM. In the embodiments of the present application, based on these models, 3D facial movements can be expressed based on BlendShape, so as to obtain the 3D facial movement model (3D Face Animation) in the embodiments of the present application.

[0089] The 3D facial movement model can adopt an autoregressive method, that is, with the closed-mouth state as the initialization, continuously predict the current state according to the previous prediction and the current speech features. As an example, Figure 3 shows a schematic diagram of a possible model structure of the 3D facial movement model. As shown in the figure, the 3D facial movement model can include, but is not limited to, a sound feature extraction module (wav2vec), an action encoding module (BS Encoder), a periodic positional encoding module (Periodic Positional Encoding), a biased causal multi-head self-attention module (Biased Causal Multi-Head Self-Attention), a biased cross-modal multi-head self-attention module (Biased Cross-Modal Multi-HeadSelf-Attention), a feed-forward module (Feed Forward), and an action decoding module (BSDecoder), etc.

[0090] Based on Figure 3For the model structure shown, the process of processing the target sound through the three-dimensional facial motion model may include: performing sound feature extraction processing on the target sound to obtain sound features corresponding to the target sound; performing motion encoding processing on the previous model output to obtain first processed data; performing periodic position encoding processing on the first processed data to obtain second processed data; performing biased causal multi-head self-attention processing on the second processed data to obtain third processed data; performing biased cross-modal multi-head self-attention processing on the sound features and the third processed data to obtain fourth processed data; performing forward feedback processing on the fourth processed data to obtain fifth processed data; and performing motion decoding processing on the fifth processed data to obtain a first facial motion corresponding to the target sound. The specific processing process may refer to the model for generating corresponding facial motions according to sounds in the prior art, and this is not elaborated in the embodiments of the present application.

[0091] It should be noted that, in the embodiments of the present application, the input of the motion encoding module (i.e., the previous facial motion) and the output of the motion decoding module (i.e., the currently predicted facial motion) are both three-dimensional human facial motions expressed based on BlendShape. In the three-dimensional human facial motions expressed based on BlendShape, since the weight coefficients of the basic facial motions are all positive values, after the motion decoding processing, a processing process of adding a sigmoid activation function may be further performed to ensure that the weight coefficients are all positive values.

[0092] In the sound feature extraction module, any sound feature extraction model in the prior art can be used to perform sound feature extraction processing, which may include but is not limited to the wav2vec model and other models, and the embodiments of the present application do not make specific limitations on this.

[0093] When the language corresponding to the existing sound feature extraction model (denoted as the first language) is the same as the language corresponding to the target sound (denoted as the second language), the existing sound feature extraction model can be directly used to perform sound feature extraction processing on the target sound, so as to obtain sound features corresponding to the target sound.

[0094] In the case where the first language is inconsistent with the second language, on the basis of the existing voice feature extraction model, the second language corpus can be used to fine-tune it, so as to obtain the fine-tuned voice feature extraction model, thereby improving the adaptability to the second language. For example, the language corresponding to the existing voice feature extraction model can be English, and the language corresponding to the target voice can be Chinese. Then, the Chinese corpus can be used to fine-tune the existing voice feature extraction model, so as to obtain the fine-tuned voice feature extraction model, thereby improving the adaptability to Chinese. When performing voice feature extraction processing, the fine-tuned voice feature extraction model can be used to perform voice feature extraction processing on the target voice, so as to obtain the voice features corresponding to the target voice.

[0095] The three-dimensional facial action model can be pre-trained through the corresponding data set, and the loss function used in the training process can be flexibly set according to the actual situation, and the embodiments of the present application do not make specific limitations on this.

[0096] In a specific implementation manner of the embodiments of the present application, it can be determined according to the process as Figure 4 shown to determine the loss function and perform model training:

[0097] Step S401: Determine the first loss function of the three-dimensional facial action model.

[0098] Among them, the first loss function is a loss function for constraining the ground truth of the three-dimensional facial action model. Specifically, the output error of the three-dimensional facial action model can be determined, and the wing function of the output error can be determined as the first loss function of the three-dimensional facial action model, as shown in the following formula:

[0099]

[0100] Among them, y is the expected output of the three-dimensional facial action model, that is, the ground truth, is the actual output of the three-dimensional facial action model, is the output error, that is, the absolute value of the difference between the expected output and the actual output, WingLoss is the wing function, and WingLoss responds much higher to subtle differences than other functions, and can more finely optimize the small errors of facial action changes. L1 is the first loss function of the three-dimensional facial action model.

[0101] Step S402: Determine the second loss function of the three-dimensional facial action model.

[0102] Among them, the second loss function is the loss function for performing inter-frame constraints on the 3D facial action model. Specifically, the first inter-frame difference and the second inter-frame difference of the 3D facial action model can be determined respectively, and the second loss function of the 3D facial action model can be determined according to the difference between the first inter-frame difference and the second inter-frame difference, as shown in the following formula:

[0103]

[0104] Among them, (y i -y i-1 ) is the first inter-frame difference, that is, the difference between two consecutive expected outputs, is the second inter-frame difference, that is, the difference between two consecutive actual outputs, is the difference between the first inter-frame difference and the second inter-frame difference, and L2 is the second loss function of the 3D facial action model.

[0105] Step S403: Fuse the first loss function and the second loss function to obtain the fused loss function of the 3D facial action model.

[0106] The specific fusion method can be flexibly set according to the actual situation, and the embodiments of the present application do not make specific limitations in this regard. As an example, for the sake of simplicity, the sum of the first loss function and the second loss function can be directly used as the fused loss function of the 3D facial action model.

[0107] Step S404: Train the initial 3D facial action model based on the fused loss function to obtain the trained 3D facial action model.

[0108] In the embodiments of the present application, it can be trained with a sufficient number of training samples, and each training sample includes a set of voice samples and corresponding first facial action labels. Using the voice samples of the training samples as inputs and the corresponding first facial action labels as expected outputs to train the 3D facial action model, the trained 3D facial action model can be obtained.

[0109] During the training process, for each training sample, the 3D facial action model can be used to process the voice samples of the training sample to obtain the actual output of the training sample, and then the fused loss function can be used to calculate the training loss value according to the expected output and the actual output in the training sample. After calculating the training loss value, the model parameters of the 3D facial action model can be adjusted according to the training loss value.

[0110] In the embodiments of the present application, it is assumed that in the initial state, the model parameters of the three-dimensional facial action model are W1. The training loss value is backpropagated to modify the model parameters W1 of the three-dimensional facial action model, and the modified model parameters W2 are obtained. After modifying the parameters, the next training process is continued. In this training process, the training loss value is recalculated, and the training loss value is backpropagated to modify the model parameters W2 of the three-dimensional facial action model, and the modified model parameters W3 are obtained, and so on. The above process is continuously repeated, and the model parameters can be modified in each training process until the preset training conditions are met. Among them, the training conditions can be that the number of training times reaches the preset number threshold, and the number threshold can be set according to the actual situation. For example, it can be set to thousands, tens of thousands, hundreds of thousands or even larger values; the training conditions can also be that the three-dimensional facial action model converges; since it is possible that the three-dimensional facial action model has converged before the number of training times reaches the number threshold, which may lead to unnecessary repeated work; or the three-dimensional facial action model may never converge, which may lead to an infinite loop and the training process cannot end. Based on the above two situations, the training conditions can also be that the number of training times reaches the number threshold or the three-dimensional facial action model converges. When the training conditions are met, the trained three-dimensional facial action model can be obtained.

[0111] Through the above process, the voice sample of the training sample and the corresponding first facial action label are used as the learning objects of the three-dimensional facial action model. After the training process, the three-dimensional facial action model can establish a mapping relationship between the voice sample and the first facial action label, so that when facing a new voice, the corresponding first facial action can also be obtained according to the mapping relationship.

[0112] After the training of the three-dimensional facial action model is completed, the three-dimensional facial action model can be used to process the target voice, so as to obtain the first facial action corresponding to the target voice.

[0113] Step S103: Perform mapping processing on the first facial action to obtain a second facial action corresponding to the first facial action.

[0114] Among them, the second facial action is a three-dimensional facial action of the robot expressed based on BlendShape. Generally, the BlendShape of the robot has many fewer dimensions (i.e., the number of basic facial actions) than that of humans. For example, the BlendShape of the robot can be 40-dimensional, and the BlendShape of humans can be 52-dimensional, resulting in a large expression loss when driving the robot.

[0115] In the embodiments of the present application, in order to alleviate this lack of expression, a neural network for facial action mapping can be pre-trained and denoted as the facial action mapping network. The first facial action is mapped through the facial action mapping network to obtain a second facial action corresponding to the first facial action.

[0116] The specific network structure of the facial action mapping network can be flexibly set according to the situation, and the embodiments of the present application do not make specific limitations thereto. As an example, for simplicity, a fully connected neural network structure can be adopted.

[0117] In the training process of the facial action mapping network, taking Chinese as an example, for each Chinese phoneme, a corresponding first facial action and a second facial action can be constructed respectively and used as a training sample. Using the first facial action of the training sample as the input and the corresponding second facial action as the expected output, the facial action mapping network is trained to obtain a trained facial action mapping network. During the training process, for each training sample, the facial action mapping network can be used to process the first facial action of the training sample to obtain the actual output of the training sample. Then, a preset loss function can be used to calculate the training loss value according to the expected output and the actual output in the training sample. After calculating the training loss value, the parameters of the facial action mapping network can be adjusted according to the training loss value until the preset training conditions are met, thereby obtaining a trained facial action mapping network.

[0118] Through the facial action mapping network, the lack of expression of the robot can be compensated to a certain extent, so as to achieve a reasonable facial action driving effect.

[0119] Step S104: Control the facial motor of the robot to execute the second facial action.

[0120] After obtaining the second facial action, the facial motor of the robot can be controlled to generate a corresponding BlendShape expression, so that the robot can show a second facial action that is as close as possible to the real facial action of humans.

[0121] In summary, the embodiment of the present application obtains a target sound for driving facial movements; processes the target sound through a preset three-dimensional facial movement model to obtain a first facial movement corresponding to the target sound; wherein, the first facial movement is a human three-dimensional facial movement based on blend shape expression; performs a mapping process on the first facial movement to obtain a second facial movement corresponding to the first facial movement; wherein, the second facial movement is a robot three-dimensional facial movement based on blend shape expression; controls the facial motor of the robot to execute the second facial movement. Through the embodiment of the present application, a three-dimensional facial movement based on blend shape expression is adopted, and the human three-dimensional facial movement can be mapped to the robot three-dimensional facial movement, so that the facial movement can be concretely disassembled into the control of the motor, which can be better applied to physical bionic robots.

[0122] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0123] Corresponding to the facial movement driving method described in the above embodiments, Figure 5 FIG. shows a structural diagram of an embodiment of a facial movement driving device provided by an embodiment of the present application.

[0124] In this embodiment, a facial movement driving device may include:

[0125] A sound acquisition module 501, configured to acquire a target sound for driving facial movements;

[0126] A facial movement model processing module 502, configured to process the target sound through a preset three-dimensional facial movement model to obtain a first facial movement corresponding to the target sound; wherein, the first facial movement is a human three-dimensional facial movement based on blend shape expression;

[0127] A facial movement mapping module 503, configured to perform a mapping process on the first facial movement to obtain a second facial movement corresponding to the first facial movement; wherein, the second facial movement is a robot three-dimensional facial movement based on blend shape expression;

[0128] A facial motor control module 504, configured to control the facial motor of the robot to execute the second facial movement.

[0129] In a specific implementation manner of the embodiment of the present application, the facial movement mapping module may be specifically configured to: perform a mapping process on the first facial movement through a preset facial movement mapping network to obtain the second facial movement corresponding to the first facial movement; wherein, the facial movement mapping network is a neural network pre-trained for facial movement mapping.

[0130] In a specific implementation manner of the embodiment of the present application, the facial action driving device may further include:

[0131] A first loss function determination module, configured to determine a first loss function of the three-dimensional facial action model; wherein, the first loss function is a loss function for performing ground-truth constraint on the three-dimensional facial action model;

[0132] A second loss function determination module, configured to determine a second loss function of the three-dimensional facial action model; wherein, the second loss function is a loss function for performing inter-frame constraint on the three-dimensional facial action model;

[0133] A fused loss function determination module, configured to fuse the first loss function and the second loss function to obtain a fused loss function of the three-dimensional facial action model;

[0134] A facial action model training module, configured to train the initial three-dimensional facial action model based on the fused loss function to obtain the trained three-dimensional facial action model.

[0135] In a specific implementation manner of the embodiment of the present application, the first loss function determination module may specifically be configured to: determine the output error of the three-dimensional facial action model; wherein, the output error is the absolute value of the difference between the expected output and the actual output; determine the wing function of the output error as the first loss function of the three-dimensional facial action model.

[0136] In a specific implementation manner of the embodiment of the present application, the second loss function determination module may specifically be configured to: determine the first inter-frame difference of the three-dimensional facial action model; wherein, the first inter-frame difference is the difference between two consecutive expected outputs; determine the second inter-frame difference of the three-dimensional facial action model; wherein, the second inter-frame difference is the difference between two consecutive actual outputs; determine the second loss function of the three-dimensional facial action model according to the difference between the first inter-frame difference and the second inter-frame difference.

[0137] In a specific implementation manner of the embodiment of the present application, the facial action model processing module may include:

[0138] A voice feature extraction unit, configured to perform voice feature extraction processing on the target voice to obtain voice features corresponding to the target voice;

[0139] An action encoding unit, configured to perform action encoding processing on the previous model output to obtain first processed data;

[0140] A position encoding unit, configured to perform periodic position encoding processing on the first processed data to obtain second processed data;

[0141] A first attention processing unit, configured to perform biased causal multi-head self-attention processing on the second processed data to obtain third processed data;

[0142] A second attention processing unit, configured to perform biased cross-modal multi-head self-attention processing on the voice feature and the third processed data to obtain fourth processed data;

[0143] A forward feedback unit, configured to perform forward feedback processing on the fourth processed data to obtain fifth processed data;

[0144] An action decoding unit, configured to perform action decoding processing on the fifth processed data to obtain the first facial action corresponding to the target voice.

[0145] In a specific implementation manner of the embodiment of the present application, the facial action driving device may further include:

[0146] A voice feature extraction model fine-tuning module, configured to obtain a voice feature extraction model corresponding to a first language; use a corpus of a second language to fine-tune the voice feature extraction model to obtain the fine-tuned voice feature extraction model;

[0147] Correspondingly, the voice feature extraction unit may specifically be configured to: perform voice feature extraction processing on the target voice through the fine-tuned voice feature extraction model to obtain the voice feature corresponding to the target voice.

[0148] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices, modules, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0149] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0150] Figure 6 FIG. shows a schematic block diagram of a robot provided by an embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the present application are shown.

[0151] As Figure 6 shown, the robot 6 in this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, the steps in the foregoing various embodiments of the facial action driving method are implemented, for exampleFigure 1 Steps S101 to S104 shown. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each module / unit in the above device embodiments, such as Figure 5 the functions of the modules 501 to 504 shown.

[0152] Exemplarily, the computer program 62 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 61 and executed by the processor 60 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 62 in the robot 6.

[0153] Those skilled in the art can understand that Figure 6 this is only an example of the robot 6 and does not constitute a limitation on the robot 6. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the robot 6 may further include input / output devices, network access devices, a bus, etc.

[0154] The processor 60 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0155] The memory 61 can be an internal storage unit of the robot 6, such as the hard disk or memory of the robot 6. The memory 61 can also be an external storage device of the robot 6, such as a plug-in hard disk equipped on the robot 6, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 61 can also include both the internal storage unit and the external storage device of the robot 6. The memory 61 is used to store the computer program and other programs and data required by the robot 6. The memory 61 can also be used to temporarily store the data that has been output or will be output.

[0156] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0157] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0158] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0159] In the embodiments provided in this application, it should be understood that the disclosed device / robot and method can be implemented in other ways. For example, the device / robot embodiments described above are only illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0160] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0161] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0162] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the above-mentioned embodiment methods of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0163] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A facial action driving method, characterized in that: include: Obtaining a target sound for driving facial movements; The target sound is processed by a preset three-dimensional facial action model to obtain a first facial action corresponding to the target sound; wherein the first facial action is a three-dimensional facial action of a human being expressed based on a blended deformation; Mapping the first facial action to obtain a second facial action corresponding to the first facial action; wherein the second facial action is a robot three-dimensional facial action based on mixed deformation expression; The facial motors of the robot are controlled to perform the second facial action.

2. The facial action driving method according to claim 1, characterized in that: The mapping process is performed on the first facial action to obtain a second facial action corresponding to the first facial action, including: Mapping the first facial action through a preset facial action mapping network to obtain the second facial action corresponding to the first facial action; The facial action mapping network is a pre-trained neural network used for facial action mapping.

3. The facial action driving method according to claim 1, characterized in that: Before processing the target sound through the preset three-dimensional facial action model, the method further includes: Determine a first loss function of the three-dimensional facial action model; wherein the first loss function is a loss function for performing truth constraints on the three-dimensional facial action model; Determine a second loss function of the three-dimensional facial action model; wherein the second loss function is a loss function for performing inter-frame constraints on the three-dimensional facial action model; fusing the first loss function and the second loss function to obtain a fusion loss function of the three-dimensional facial action model; The initial three-dimensional facial action model is trained based on the fusion loss function to obtain the trained three-dimensional facial action model.

4. The facial action driving method according to claim 3, characterized in that: The determining of a first loss function of the three-dimensional facial action model comprises: Determining an output error of the three-dimensional facial action model; wherein the output error is an absolute value of a difference between an expected output and an actual output; A wing function of the output error is determined as the first loss function of the three-dimensional facial action model.

5. The facial action driving method according to claim 3, characterized in that: The determining of a second loss function of the three-dimensional facial action model comprises: Determining a first inter-frame difference of the three-dimensional facial action model; wherein the first inter-frame difference is a difference between two consecutive frames of expected output; Determine a second inter-frame difference of the three-dimensional facial action model; wherein the second inter-frame difference is a difference between two consecutive frames of actual output; The second loss function of the 3D facial action model is determined according to a difference between the first inter-frame difference and the second inter-frame difference.

6. The facial action driving method according to any one of claims 1 to 5, characterized in that: The processing of the target sound to obtain a first facial action corresponding to the target sound includes: Performing sound feature extraction processing on the target sound to obtain a sound feature corresponding to the target sound; Performing action encoding processing on the previous model output to obtain first processed data; Performing periodic position encoding processing on the first processed data to obtain second processed data; Performing biased causal multi-head self-attention processing on the second processed data to obtain third processed data; performing biased cross-modal multi-head self-attention processing on the sound feature and the third processed data to obtain fourth processed data; Performing forward feedback processing on the fourth processed data to obtain fifth processed data; Performing action decoding processing on the fifth processed data to obtain the first facial action corresponding to the target sound.

7. The facial action driving method according to claim 6, characterized in that: Before performing sound feature extraction processing on the target sound, the method further includes: Acquire a sound feature extraction model corresponding to the first language; Using the corpus of the second language to fine-tune the sound feature extraction model to obtain the fine-tuned sound feature extraction model; Accordingly, the performing sound feature extraction processing on the target sound to obtain a sound feature corresponding to the target sound includes: The target sound is subjected to sound feature extraction processing by using the finely tuned sound feature extraction model to obtain the sound feature corresponding to the target sound.

8. A facial action driving device, characterized in that: include: A sound acquisition module, used to acquire a target sound for driving facial movements; A facial action model processing module, used for processing the target sound through a preset three-dimensional facial action model to obtain a first facial action corresponding to the target sound; wherein the first facial action is a three-dimensional facial action of a human being expressed based on a blended deformation; A facial action mapping module, used for mapping the first facial action to obtain a second facial action corresponding to the first facial action; wherein the second facial action is a robot three-dimensional facial action based on a blended deformation expression; The facial motor control module is used to control the facial motor of the robot to perform the second facial action.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the facial action driving method according to any one of claims 1 to 7 are implemented.

10. A robot comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the facial action driving method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Expression control method and system of expression robot

    CN121696952A