Facial action intensity prediction method and system of digital human, terminal and medium
By constructing a multi-layer neural network model, using historical audio data and facial expression parameters, the precise prediction of the intensity of facial movements of digital people is achieved, solving the problem of lack of external factors in the existing technology in facial expression generation, and improving the interactivity between digital people and users.
Patent Information
- Application Number
- CN202411255837.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology generates digital facial expressions of people without the ability to respond to external factors and cannot accurately predict the intensity of facial movements, resulting in the inability to generate dynamic facial expressions in real time.
By obtaining the historical audio data of digital people, converting them into facial expression videos, and analyzing facial motion expression parameters. Then, the historical audio data is transformed, the training data set is constructed, and the multi-layer neural network model is trained to obtain the facial motion intensity prediction model. This model can output facial action intensity values based on current audio data.
It realizes accurate prediction of the intensity of facial movements of digital people, improves the interactivity and expressiveness between digital people and users, and supports real-time dynamic facial expression generation.
Smart Images

Figure CN120088374A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, system, terminal and computer-readable storage medium for predicting the facial action intensity of a digital human. Background Art
[0002] With the rapid development of VR (Virtual Reality) and AR (Augmented Reality) technologies, digital human technology is increasingly widely used in fields such as games and online education. Among them, the facial expressions of digital humans play a crucial role in enhancing the user experience and strengthening emotional communication. Currently, the facial expressions of most 3D digital humans mainly rely on the ARKit Blendshape technology, which enables developers to capture and track the facial expressions of users, thereby achieving highly realistic facial animations.
[0003] Currently, the existing method for generating the facial expressions of digital humans based on ARKit Blendshape achieves vivid and realistic facial animations by accurately capturing and tracking the facial expressions of users. However, this method mainly generates the facial expressions of digital humans statically based on the actual facial expressions of users, lacks the ability to respond to external factors such as voice input, and cannot accurately predict the corresponding facial action intensity, resulting in limited support for real-time dynamic facial expression generation.
[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0005] The main objective of the present invention is to provide a method for predicting the facial action intensity of a digital human, aiming to solve the problem that the existing technology for generating the facial expressions of digital humans lacks the ability to respond to external factors, cannot accurately predict the corresponding facial action intensity, and thus cannot generate dynamic facial expressions in real time.
[0006] To achieve the above objective, the present invention provides a method, system, terminal and medium for predicting the facial action intensity of a digital human. The method for predicting the facial action intensity of a digital human includes the following steps:
[0007] Obtain the historical audio data of the digital human, perform conversion processing on the historical audio data to obtain a facial expression video, and perform analysis processing on the facial expression video to obtain facial action expression parameters;
[0008] Perform feature transformation on the historical audio data to obtain a target feature vector, construct a training data set based on the target feature vector and the facial action expression parameters, and perform model training based on the training data set to obtain a facial action intensity prediction model;
[0009] Obtain the current target audio data of the digital human, input the target audio data into the facial action intensity prediction model, output the corresponding facial action intensity value, and obtain the target facial action intensity according to the facial action intensity value.
[0010] Optionally, for the facial action intensity prediction method of the digital human, wherein, obtaining the historical audio data of the digital human and performing conversion processing on the historical audio data to obtain a facial expression video specifically includes:
[0011] Obtain the historical audio data of the digital human, and perform preprocessing on the historical audio data to obtain target historical audio data;
[0012] Extract features from the target historical audio data to obtain audio features, and perform mapping processing on the audio features to obtain corresponding facial expression parameters;
[0013] Construct a three-dimensional face model, drive the three-dimensional face model according to the facial expression parameters to obtain multiple facial animation frames, and synthesize all the facial animation frames to obtain a facial expression video.
[0014] Optionally, for obtaining the historical audio data of the digital human and performing preprocessing on the historical audio data to obtain target historical audio data, specifically includes:
[0015] Obtain the historical audio data of the digital human, and perform noise reduction processing on the historical audio data to obtain noise-reduced audio;
[0016] Obtain a preset audio standard, and perform standardization processing on the noise-reduced audio according to the preset audio standard to obtain target historical audio data.
[0017] Optionally, for the facial action intensity prediction method of the digital human, wherein, analyzing and processing the facial expression video to obtain facial action expression parameters specifically includes:
[0018] Extract facial actions from the facial expression video to obtain facial action information, and perform action recognition according to the facial action information to obtain corresponding facial actions;
[0019] Extract key points of the facial action to obtain facial key points, and perform mapping processing on the facial key points to obtain facial action expression parameters.
[0020] Optionally, for the facial action intensity prediction method of the digital human, wherein, after extracting key points of the facial action to obtain facial key points and performing mapping processing on the facial key points to obtain facial action expression parameters, it further includes:
[0021] Match the facial key points with the facial regions corresponding to the facial action expression parameters to obtain a matching result, and perform numerical calculation on the facial action expression parameters according to the matching result to obtain a facial action parameter value.
[0022] Optionally, in the facial action intensity prediction method of the digital human, wherein, the historical audio data is subjected to feature transformation to obtain a target feature vector, and a training data set is constructed according to the target feature vector and the facial action expression parameters, specifically including:
[0023] Extract features from the historical audio data to obtain a high-dimensional feature vector, and perform optimization processing on the high-dimensional feature vector to obtain a target feature vector, wherein the optimization processing includes downsampling processing, shearing processing, and alignment processing;
[0024] Perform corresponding processing on the target feature vector and the facial action parameter value to obtain a target correspondence, and construct a training data set according to the target correspondence.
[0025] Optionally, in the facial action intensity prediction method of the digital human, wherein, the model is trained according to the training data set to obtain a facial action intensity prediction model, specifically including:
[0026] Create a facial action intensity training model, and input a set of training samples in the training data set into the facial action intensity training model;
[0027] Extract features from the audio data of the training sample to obtain an audio feature, and perform feature mapping on the audio feature to obtain a corresponding feature vector;
[0028] Perform facial action intensity prediction on the audio data according to the feature vector to obtain a predicted intensity value, calculate the loss value between the predicted intensity value and the facial action intensity value corresponding to the audio data according to the cross-entropy loss function, and adjust the parameters of the facial action intensity training model according to the loss value;
[0029] Input the next set of training samples into the facial action intensity training model until the training situation of the facial action intensity training model meets the preset conditions to obtain a trained facial action intensity prediction model.
[0030] Optionally, in the facial action intensity prediction method of the digital human, wherein, the facial action intensity prediction system of the digital human includes:
[0031] A data analysis module, configured to obtain historical audio data of a digital human, perform conversion processing on the historical audio data to obtain a facial expression video, and perform analysis processing on the facial expression video to obtain facial action expression parameters;
[0032] A model training module, configured to perform feature transformation on the historical audio data to obtain a target feature vector, construct a training data set according to the target feature vector and the facial action expression parameters, and perform model training according to the training data set to obtain a facial action intensity prediction model;
[0033] An intensity prediction module, configured to obtain current target audio data of the digital human, input the target audio data into the facial action intensity prediction model, output a corresponding facial action intensity value, and obtain a target facial action intensity according to the facial action intensity value.
[0034] In addition, to achieve the above object, the present invention further provides a terminal, wherein the terminal includes: a memory, a processor, and a facial action intensity prediction program of the digital human stored on the memory and executable on the processor. When the facial action intensity prediction program of the digital human is executed by the processor, the steps of the above-mentioned facial action intensity prediction method of the digital human are implemented.
[0035] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a facial action intensity prediction program of the digital human. When the facial action intensity prediction program of the digital human is executed by a processor, the steps of the above-mentioned facial action intensity prediction method of the digital human are implemented.
[0036] In the present invention, historical audio data of a digital human is obtained, conversion processing is performed on the historical audio data to obtain a facial expression video, and analysis processing is performed on the facial expression video to obtain facial action expression parameters; feature transformation is performed on the historical audio data to obtain a target feature vector, a training data set is constructed according to the target feature vector and the facial action expression parameters, and model training is performed according to the training data set to obtain a facial action intensity prediction model; current target audio data of the digital human is obtained, the target audio data is input into the facial action intensity prediction model, a corresponding facial action intensity value is output, and a target facial action intensity is obtained according to the facial action intensity value. The present invention realizes accurate prediction of the facial action intensity of the digital human and improves the interactivity between the digital human and the user. Description of the Drawings
[0037] Figure 1 is a flowchart of a preferred embodiment of the facial action intensity prediction method of the digital human in the present invention;
[0038] Figure 2 It is a schematic diagram of generating ARKit Blendshapes parameter values of a preferred embodiment of the present invention;
[0039] Figure 3 It is a schematic diagram of extracting feature vectors of audio data of a preferred embodiment of the present invention;
[0040] Figure 4 It is a schematic diagram of the overall process of the facial action intensity prediction method of the digital human in the present invention;
[0041] Figure 5 It is a structural diagram of a preferred embodiment of the facial action intensity prediction system of the digital human in the present invention;
[0042] Figure 6 It is a schematic diagram of the operating environment of a preferred embodiment of the terminal of the present invention. Specific Embodiments
[0043] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention.
[0044] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0045] In addition, if there are descriptions such as "first", "second", etc. involved in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0046] The facial action intensity prediction method of the digital human according to a preferred embodiment of the present invention, as Figure 1 shown, the facial action intensity prediction method of the digital human includes the following steps:
[0047] Step S10: Obtain the historical audio data of the digital human, perform conversion processing on the historical audio data to obtain a facial expression video, and perform analysis processing on the facial expression video to obtain facial action expression parameters.
[0048] The step S10 includes:
[0049] Step S11: Obtain the historical audio data of the digital human, perform preprocessing on the historical audio data to obtain target historical audio data;
[0050] Step S12: Extract features from the target historical audio data to obtain audio features, and perform mapping processing on the audio features to obtain corresponding facial expression parameters;
[0051] Step S13: Construct a three-dimensional face model, drive the three-dimensional face model according to the facial expression parameters to obtain multiple facial animation frames, and synthesize all the facial animation frames to obtain a facial expression video;
[0052] Step S14: Extract facial actions from the facial expression video to obtain facial action information, and perform action recognition according to the facial action information to obtain corresponding facial actions;
[0053] Step S15: Extract key points from the facial actions to obtain facial key points, and perform mapping processing on the facial key points to obtain facial action expression parameters.
[0054] Specifically, the existing digital human facial expression generation technology based on ARKit Blendshapes (ARKit Blendshapes is a function for facial tracking in the ARKit framework. By detecting facial expression changes and converting them into deformation parameters of a 3D model, each Blendshape represents a specific facial expression or action, and the corresponding numerical range is usually between 0 and 1, which is used to control the intensity of the expression or action) lacks the ability to respond to external factors (such as voice input), limiting the interaction experience between the digital human and the user. Moreover, most of the facial expression prediction models based on deep learning are static, with limited support for real-time dynamic facial expression generation, making it difficult to achieve realistic dynamic expression effects. There is also a lack of a 3D digital human facial action intensity prediction model driven by voice, which restricts the application and development of digital human technology in aspects such as voice interaction and emotional expression. Additionally, the generated 3D digital human facial expression data cannot be directly mapped to ARKit Blendshape, requiring additional conversion and processing steps, increasing complexity and cost. To address the problems existing in the prior art, the present invention solves them by constructing a prediction model for the 3D digital human facial action intensity based on a multi-layer neural network, namely the facial action intensity prediction model. The facial action intensity prediction model aims to learn the complex relationship between the audio feature vectors provided in the dataset and the numerical values of the corresponding ARKit Blendshapes, i.e., the facial action parameter values. Moreover, the facial action intensity prediction model has different weights and biases to perform a non-linear transformation on the input data.
[0055] Before obtaining the facial action intensity prediction model, a training dataset for model training needs to be constructed, wherein the training set includes the feature vectors of audio data and the corresponding facial action parameter values; the processing process for the facial action parameter values is as Figure 2As shown, obtain the historical audio data of the digital human, preprocess the historical audio data to obtain the target historical audio data, where the preprocessing includes noise reduction processing and normalization processing, that is, perform noise reduction processing on the historical audio data to obtain the denoised audio; obtain the preset audio standard, and perform normalization processing on the denoised audio according to the preset audio standard to obtain the target historical audio data; afterwards, it is necessary to generate a facial expression video of the digital human according to the target historical audio data. For example, use the sadtalker method to analyze the speech features in the audio, and synthesize facial expressions corresponding to the audio content according to the speech features, so that the facial movements of the digital human match the voice naturally. Specifically, extract features from the target historical audio data to obtain audio features (such as pitch and frequency, etc., which will be used to drive the generation of facial expressions), and perform mapping processing on the audio features to obtain the corresponding facial expression parameters; at the same time, detect facial landmarks in the input picture or video, and construct a corresponding 3D face model, drive the 3D face model according to the facial expression parameters to obtain multiple facial animation frames, and synthesize all the facial animation frames to obtain the facial expression video.
[0056] After obtaining the corresponding facial expression video, it is necessary to obtain the corresponding facial action expression parameters according to the facial expression video. For example, by inputting the generated facial expression video into the Mediapipe system (the MediaPipe system is a cross-platform open-source framework for building multimodal machine learning pipelines, which can accurately identify various subtle facial movements, including the movements of parts such as eyes, mouth, and eyebrows), through the powerful facial action recognition function of the mediapipe system, extract the key facial action information, that is, perform facial action extraction on the facial expression video to obtain the facial action information, and perform action recognition according to the facial action information to obtain the corresponding facial action; afterwards, generate the corresponding ARKit Blendshapes parameters, that is, facial action expression parameters, according to the recognized facial action. Specifically, extract key points from the facial action to obtain facial key points, and perform mapping processing on the facial key points to obtain the facial action expression parameters, and the facial action expression parameters can drive the facial model of the 3D digital human to make it show realistic facial expressions.
[0057] Furthermore, the facial key points can be matched with the facial regions corresponding to the facial action expression parameters to obtain a matching result. For example, jawOpen (jaw opening and closing) can be determined by analyzing the displacement of the key points in the chin region, while eyeBlinkLeft (left eye closing and blinking) is estimated by analyzing the degree of closure of the eyelid key points. Then, numerical calculations are performed on the facial action expression parameters according to the matching result to obtain facial action parameter values. Each facial action parameter value will be between 0 and 1, which is used to represent the intensity of the facial expression.
[0058] Step S20: Perform feature transformation on the historical audio data to obtain a target feature vector. Construct a training data set based on the target feature vector and the facial action expression parameters, and perform model training based on the training data set to obtain a facial action intensity prediction model.
[0059] Specifically, in the embodiments of the present invention, the processing process of the feature vector of the audio data is as Figure 3 shown. That is, feature extraction is performed on the historical audio data to obtain a high-dimensional feature vector. For example, a high-dimensional feature vector of each frame is extracted from the historical audio data through the wav2vec2 model. Then, the high-dimensional feature vector is optimized to obtain a target feature vector. The optimization process includes downsampling, shearing, and alignment processes, aiming to make the number of its frames the same as the number of facial action parameter values. Subsequently, the target feature vector is corresponding processed with the facial action parameter values to obtain a target correspondence, so that the subsequent model can learn the direct relationship between them, thus making better predictions. And a training data set is constructed according to the target correspondence, which provides a basis for further studying the facial action intensity of the language-driven digital human. The target feature vector represents the speech feature at each time point, while the facial action parameter value represents the actual action intensity of the digital human's face at the corresponding time point.
[0060] After that, use this data training set to train a multi-layer neural network model, that is, a facial action intensity training model. Specifically, the training process is as follows: create a facial action intensity training model, and input a set of training samples in the training data set into the facial action intensity training model; extract features from the audio data of the training samples to obtain audio features, and perform feature mapping on the audio features to obtain corresponding feature vectors; predict the facial action intensity of the audio data according to the feature vectors to obtain a predicted intensity value, calculate the loss value between the predicted intensity value and the facial action intensity value corresponding to the audio data according to the cross-entropy loss function, and adjust the parameters of the facial action intensity training model according to the loss value. For example, the weights and biases in the facial action intensity training model can be adjusted through the backpropagation algorithm and the gradient descent optimization algorithm to minimize the error between the predicted value and the actual value; then, input the next set of training samples into the facial action intensity training model, and repeat the above process, which will not be elaborated here. Through continuous iterative training, the generalization ability and prediction accuracy of the model are improved until the training situation of the facial action intensity training model meets the preset conditions. The preset conditions include that the loss value meets the preset requirements or the number of training times reaches the preset number. Among them, the preset requirements can be determined according to the accuracy of the facial action intensity prediction model, which will not be elaborated here. The preset number can be the maximum number of training times of the facial action intensity training model. For example, 2000 times, etc. Finally, a trained facial action intensity prediction model is obtained.
[0061] Step S30: Obtain the current target audio data of the digital human, input the target audio data into the facial action intensity prediction model, output the corresponding facial action intensity value, and obtain the target facial action intensity according to the facial action intensity value.
[0062] Specifically, the present invention realizes the accurate prediction of the facial action intensity of the digital human based on voice input by using the trained facial action intensity prediction model. Specifically, obtain the current target audio data of the digital human, input the target audio data into the facial action intensity prediction model, output the corresponding facial action intensity value, and obtain the target facial action intensity according to the facial action intensity value, thereby improving the interactivity and expressiveness between the digital human and the user. Because during the real-time conversation with the digital human, it is required that the digital human has a natural mouth shape when speaking, which conforms to normal expression habits, rather than a fixed mouth shape. And the content expressed by the digital human comes from a large model, and the output format is generally text or voice. When input into the digital human, it is the voice playback content. The existing methods cannot realize real-time language-driven facial actions, while the present invention can drive the facial actions of the digital human in real time through language, and it can be used in the digital human scenario application of real-time conversation.
[0063] In summary, the present invention proposes a language-driven multi-layer neural network prediction model, namely a facial action intensity prediction model, which can accurately predict the real-time facial expression action intensity of a digital human by learning the non-linear mapping relationship between the input data and the output result, thereby improving the interactivity between the digital human and the user.
[0064] Furthermore, the overall process of the facial action intensity prediction method for the digital human in the present invention is as Figure 4 shown. First, obtain the historical audio data of the digital human, preprocess the historical audio data to obtain target historical audio data. After that, it is necessary to generate a facial expression video of the digital human according to the target historical audio data. For example, use the sadtalker method to analyze the speech features in the audio and synthesize a facial expression corresponding to the audio content according to the speech features, so that the facial actions of the digital human match the voice naturally. After obtaining the corresponding facial expression video, it is necessary to obtain the corresponding facial action expression parameters according to the facial expression video. For example, by inputting the generated facial expression video into the Mediapipe system, through the powerful facial action recognition function of the Mediapipe system, extract the key facial action information, and perform action recognition according to the facial action information to obtain the corresponding facial action. Then generate the corresponding ARKit Blendshapes parameters, that is, facial action expression parameters, according to the recognized facial action, and obtain the numerical value of the facial action expression parameters, that is, the facial action parameter value.
[0065] Secondly, extract features from the historical audio data to obtain a high-dimensional feature vector. For example, extract the high-dimensional feature vector of each frame from the historical audio data through the wav2vec2 model. After that, perform optimization processing on the high-dimensional feature vector to obtain a target feature vector. Among them, the optimization processing includes downsampling processing, clipping processing, and alignment processing, aiming to make the number of its frames the same as the number of facial action parameter values. Subsequently, perform corresponding processing on the target feature vector and the facial action parameter value to obtain a target correspondence relationship, and construct a training data set according to the target correspondence relationship. Then, use this data training set to train a multi-layer neural network model, namely a facial action intensity training model, to obtain a trained facial action intensity prediction model.
[0066] Finally, use the trained facial action intensity prediction model to accurately predict the facial action intensity of the digital human based on voice input, obtain the current target audio data of the digital human, input the target audio data into the facial action intensity prediction model, output the corresponding facial action intensity numerical value, and subsequently obtain the target facial action intensity according to the facial action intensity numerical value, thereby improving the interactivity and expressiveness between the digital human and the user.
[0067] Further, as Figure 5 shown, based on the above-mentioned facial action intensity prediction method of the digital human, the present invention also correspondingly provides a facial action intensity prediction system for the digital human. The facial action intensity prediction system for the digital human includes:
[0068] A data analysis module 51, configured to obtain historical audio data of the digital human, perform conversion processing on the historical audio data to obtain a facial expression video, and perform analysis processing on the facial expression video to obtain facial action expression parameters;
[0069] A model training module 52, configured to perform feature transformation on the historical audio data to obtain a target feature vector, construct a training data set according to the target feature vector and the facial action expression parameters, and perform model training according to the training data set to obtain a facial action intensity prediction model;
[0070] An intensity prediction module 53, configured to obtain current target audio data of the digital human, input the target audio data into the facial action intensity prediction model, output a corresponding facial action intensity value, and obtain a target facial action intensity according to the facial action intensity value.
[0071] Further, as Figure 6 shown, based on the above-mentioned facial action intensity prediction method of the digital human, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0072] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as a hard disk or a memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk equipped on the terminal, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes installed on the terminal, etc. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a facial action intensity prediction program 40 for the digital human is stored on the memory 20, and the facial action intensity prediction program 40 for the digital human can be executed by the processor 10, so as to implement the facial action intensity prediction method of the digital human in this application.
[0073] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips, which are used to run the program code stored in the memory 20 or process data, such as executing the facial action intensity prediction method of the digital human, etc.
[0074] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display the information on the terminal and to display a visual user interface. The components 10-30 of the terminal communicate with each other through a system bus.
[0075] In one embodiment, when the processor 10 executes the program 40 for predicting the facial action intensity of the digital human in the memory 20, the following steps are implemented:
[0076] Obtain the historical audio data of the digital human, perform conversion processing on the historical audio data to obtain a facial expression video, and perform analysis processing on the facial expression video to obtain facial action expression parameters;
[0077] Perform feature transformation on the historical audio data to obtain a target feature vector, construct a training data set according to the target feature vector and the facial action expression parameters, and perform model training according to the training data set to obtain a facial action intensity prediction model;
[0078] Obtain the current target audio data of the digital human, input the target audio data into the facial action intensity prediction model, output the corresponding facial action intensity value, and obtain the target facial action intensity according to the facial action intensity value.
[0079] Among them, the obtaining of the historical audio data of the digital human and the performing of conversion processing on the historical audio data to obtain a facial expression video specifically include:
[0080] Obtain the historical audio data of the digital human, perform preprocessing on the historical audio data to obtain target historical audio data;
[0081] Perform feature extraction on the target historical audio data to obtain audio features, and perform mapping processing on the audio features to obtain corresponding facial action expression parameters;
[0082] Construct a three-dimensional face model, drive the three-dimensional face model according to the facial action expression parameters to obtain a plurality of facial animation frames, and synthesize all the facial animation frames to obtain a facial expression video.
[0083] Among them, obtaining the historical audio data of the digital human, preprocessing the historical audio data to obtain target historical audio data specifically includes:
[0084] Obtain the historical audio data of the digital human, perform noise reduction processing on the historical audio data to obtain noise-reduced audio;
[0085] Obtain a preset audio standard, and perform standardization processing on the noise-reduced audio according to the preset audio standard to obtain target historical audio data.
[0086] Among them, analyzing and processing the facial expression video to obtain facial action expression parameters specifically includes:
[0087] Extract facial actions from the facial expression video to obtain facial action information, and perform action recognition according to the facial action information to obtain corresponding facial actions;
[0088] Extract key points from the facial actions to obtain facial key points, and perform mapping processing on the facial key points to obtain facial action expression parameters.
[0089] Among them, after extracting key points from the facial actions to obtain facial key points, performing mapping processing on the facial key points to obtain facial action expression parameters, it further includes:
[0090] Match the facial key points with the facial regions corresponding to the facial action expression parameters to obtain a matching result, and perform numerical calculation on the facial action expression parameters according to the matching result to obtain facial action parameter values.
[0091] Among them, performing feature transformation on the historical audio data to obtain a target feature vector, and constructing a training data set according to the target feature vector and the facial action expression parameters specifically includes:
[0092] Extract features from the historical audio data to obtain a high-dimensional feature vector, and perform optimization processing on the high-dimensional feature vector to obtain a target feature vector, where the optimization processing includes downsampling processing, shearing processing, and alignment processing;
[0093] Correspond the target feature vector with the facial action parameter values to obtain a target correspondence, and construct a training data set according to the target correspondence.
[0094] Among them, training a model according to the training data set to obtain a facial action intensity prediction model specifically includes:
[0095] Create a facial action intensity training model, and input a set of training samples in the training dataset into the facial action intensity training model;
[0096] Extract features from the audio data of the training samples to obtain audio features, and perform feature mapping on the audio features to obtain corresponding feature vectors;
[0097] Predict the facial action intensity of the audio data according to the feature vectors to obtain predicted intensity values, calculate the loss value between the predicted intensity values and the facial action intensity values corresponding to the audio data according to the cross-entropy loss function, and adjust the parameters of the facial action intensity training model according to the loss value;
[0098] Input the next set of training samples into the facial action intensity training model until the training situation of the facial action intensity training model meets the preset conditions, and obtain a trained facial action intensity prediction model.
[0099] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a facial action intensity prediction program for a digital human, and when the facial action intensity prediction program for the digital human is executed by a processor, the steps of the above-mentioned facial action intensity prediction method for the digital human are implemented.
[0100] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0101] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above-mentioned embodiment methods can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable computer-readable storage medium, and when the program is executed, it can include the processes of the above-mentioned method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disc, etc.
[0102] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all these improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for predicting the intensity of facial movements of digital humans, characterized in that: The digital human facial action intensity prediction method comprises: Acquire historical audio data of the digital human, convert the historical audio data to obtain a facial expression video, and analyze the facial expression video to obtain facial action expression parameters; Performing feature conversion on the historical audio data to obtain a target feature vector, constructing a training data set according to the target feature vector and the facial action expression parameters, and performing model training according to the training data set to obtain a facial action intensity prediction model; The current target audio data of the digital human is obtained, the target audio data is input into the facial action intensity prediction model, the corresponding facial action intensity value is output, and the target facial action intensity is obtained according to the facial action intensity value.
2. The method for predicting the intensity of digital human facial movements according to claim 1, characterized in that: The step of acquiring historical audio data of a digital human and converting the historical audio data to obtain a facial expression video specifically includes: Acquire historical audio data of the digital human, and pre-process the historical audio data to obtain target historical audio data; Extracting features from the target historical audio data to obtain audio features, and mapping the audio features to obtain corresponding facial expression parameters; A three-dimensional face model is constructed, and the three-dimensional face model is driven according to the facial expression parameters to obtain a plurality of facial animation frames, and all the facial animation frames are synthesized to obtain a facial expression video.
3. The method for predicting the intensity of digital human facial movements according to claim 2, characterized in that: The step of obtaining the historical audio data of the digital human and preprocessing the historical audio data to obtain the target historical audio data specifically includes: Acquire historical audio data of the digital human, and perform noise reduction processing on the historical audio data to obtain noise-reduced audio; A preset audio standard is obtained, and the noise reduction audio is standardized according to the preset audio standard to obtain target historical audio data.
4. The method for predicting the intensity of digital human facial movements according to claim 1, characterized in that: The analyzing and processing the facial expression video to obtain facial action expression parameters specifically includes: Extracting facial movements from the facial expression video to obtain facial movement information, and performing movement recognition based on the facial movement information to obtain corresponding facial movements; Key points of the facial action are extracted to obtain facial key points, and the facial key points are mapped to obtain facial action expression parameters.
5. The method for predicting the intensity of digital human facial movements according to claim 4, characterized in that: The facial action is subjected to key point extraction to obtain facial key points, and the facial key points are subjected to mapping processing to obtain facial action expression parameters, and then the following steps are further included: The facial key points are matched with the facial area corresponding to the facial action expression parameters to obtain a matching result, and the facial action expression parameters are numerically calculated according to the matching result to obtain facial action parameter values.
6. The method for predicting the intensity of digital human facial movements according to claim 5, characterized in that: The feature conversion of the historical audio data to obtain a target feature vector and constructing a training data set according to the target feature vector and the facial action expression parameters specifically includes: Extracting features from the historical audio data to obtain a high-dimensional feature vector, and optimizing the high-dimensional feature vector to obtain a target feature vector, wherein the optimization process includes downsampling, shearing, and alignment; The target feature vector and the facial action parameter value are processed correspondingly to obtain a target corresponding relationship, and a training data set is constructed according to the target corresponding relationship.
7. The method for predicting the intensity of digital human facial movements according to claim 1, characterized in that: The performing model training according to the training data set to obtain a facial action intensity prediction model specifically includes: Creating a facial action intensity training model, and inputting a set of training samples in the training data set into the facial action intensity training model; Extracting features from the audio data of the training sample to obtain audio features, and performing feature mapping on the audio features to obtain corresponding feature vectors; Predicting the intensity of facial movements of the audio data according to the feature vector to obtain a predicted intensity value, calculating a loss value between the predicted intensity value and a facial movement intensity value corresponding to the audio data according to a cross entropy loss function, and adjusting parameters of the facial movement intensity training model according to the loss value; The next set of training samples is input into the facial action intensity training model until the training status of the facial action intensity training model meets the preset conditions, thereby obtaining a trained facial action intensity prediction model.
8. A digital human facial action intensity prediction system, characterized in that: The digital human facial action intensity prediction system comprises: A data analysis module is used to obtain the historical audio data of the digital human, convert the historical audio data to obtain a facial expression video, and analyze the facial expression video to obtain facial action expression parameters; A model training module, used to perform feature conversion on the historical audio data to obtain a target feature vector, construct a training data set according to the target feature vector and the facial action expression parameters, and perform model training according to the training data set to obtain a facial action intensity prediction model; The intensity prediction module is used to obtain the current target audio data of the digital human, input the target audio data into the facial action intensity prediction model, output the corresponding facial action intensity value, and obtain the target facial action intensity according to the facial action intensity value.
9. A terminal, characterized in that: The terminal includes a memory, a processor, and a digital human facial movement intensity prediction program stored in the memory and executable on the processor. When the digital human facial movement intensity prediction program is executed by the processor, the steps of the digital human facial movement intensity prediction method as claimed in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, wherein the computer-readable storage medium stores a digital human facial action intensity prediction program, and when the digital human facial action intensity prediction program is executed by a processor, the steps of the digital human facial action intensity prediction method as described in any one of claims 1-7 are implemented.