Digital population animation generation method, device, electronic device and storage medium

By obtaining audio data for vertex action calculation and key point feature processing, the Blendshapes technology is used to generate lip animation, which solves the fluency and versatility problems in traditional technologies, and achieves natural and smooth lip animation effects and reduces costs.

CN119991892BActive Publication Date: 2025-07-18HANGZHOU FUYUN NETWORK TECHNOLOGY CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510406947.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The prior art is difficult to take into account the smoothness and versatility of lip animations at the same time. The animation effect of traditional phoneme solutions is stiff, and the deep learning solutions are costly and not universal.

Method used

By obtaining audio data, vertex action calculation, selecting key point features and inputting deformation parameter prediction model, and using Blendshapes technology to generate lip animations, which are suitable for different digital characters.

Benefits of technology

Generate smooth and natural lip animations, reducing the high cost of retraining new characters and improving the realism and versatility of the animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991892B_ABST
    Figure CN119991892B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, electronic device, and storage medium for generating digital lip-sync animations. The method for generating digital lip-sync animations includes: obtaining audio data to be recognized, performing vertex motion operations on the audio data to be recognized to obtain vertex motion data; selecting key-point features from the vertex motion data, and inputting the key-point features into a trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key-point features; and generating a digital lip-sync animation based on the deformation parameter prediction values. Through the present application, the problem of how to balance the playback smoothness and generality of lip-sync animations is solved, the high cost of retraining new characters is avoided, and a smooth and natural lip-sync animation effect is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and particularly to a method, apparatus, electronic device, and storage medium for generating digital human mouth animations. Background Art

[0002] In the current digital age, 3D digital humans are increasingly widely used in fields such as virtual anchors, online education, and virtual customer service. To achieve natural interaction of 3D digital humans, especially oral expression, accurately driving their mouth movements is crucial.

[0003] Traditional voice-driven mouth animation technologies mainly rely on phoneme-based schemes and deep learning-based schemes. Although the phoneme-based scheme is simple to implement, the animation effect is rigid and it is difficult to show natural and smooth mouth transitions. The deep learning-based scheme, on the other hand, relies on high-quality actor mouth datasets for model training. Although it can generate more natural mouth animations, most of them are closed-source and charged, and the generated vertex-driven animations cannot be directly applied to different digital characters, resulting in high costs and poor versatility.

[0004] Currently, no effective solution has been proposed for the problem of how to balance the playback smoothness and versatility of mouth animations in related technologies. Summary of the Invention

[0005] Embodiments of the present application provide a method, apparatus, electronic device, and storage medium for generating digital human mouth animations to at least solve the problem of how to balance the playback smoothness and versatility of mouth animations in related technologies.

[0006] In a first aspect, embodiments of the present application provide a method for generating digital human mouth animations, including:

[0007] Obtaining audio data to be recognized, performing vertex motion operations on the audio data to be recognized to obtain vertex motion data;

[0008] Selecting key point features from the vertex motion data and inputting the key point features into a trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features;

[0009] Generating the digital human mouth animation based on the deformation parameter prediction values.

[0010] In some of these embodiments, the deformation parameter prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer; the inputting the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features includes:

[0011] Input the key point features into the first hidden layer to map the key point features to first features;

[0012] Input the first features into the second hidden layer to map the first features to second features;

[0013] Input the second features into the regression prediction layer to map the second features to the predicted values of the deformation parameters.

[0014] In some embodiments, the method further includes:

[0015] Obtain vertex deformation parameters based on the acquired training video data;

[0016] Extract training audio data based on the training video data; select training features from the training audio data, and input the training features into the initial prediction model to obtain training parameter results;

[0017] Perform error optimization processing on the initial prediction model based on the training parameter results and the vertex deformation parameters, and generate the deformation parameter prediction model.

[0018] In some embodiments, the initial prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer; the step of inputting the training features into the initial prediction model to obtain training parameter results includes:

[0019] Map the training features to first training features via the first hidden layer based on a preset number of first features, where the number of first features is the number of the first training features;

[0020] Map the first training features to second training features via the second hidden layer based on a preset number of second features, where the number of second features is the number of the second training features; wherein, the number of second features is less than the number of first features;

[0021] Map the second training features to training parameter results via the regression prediction layer based on a preset number of prediction features, where the number of prediction features is the number of the training parameter results; wherein, the number of prediction features is less than the number of second features.

[0022] In some embodiments, the regression prediction layer includes a compression activation function; the step of mapping the second training features to training parameter results based on a preset number of prediction features includes:

[0023] Map the second training features to training initial parameters based on the number of prediction features;

[0024] Using the compression activation function, compress the initial training parameters to a preset weight range to obtain the training parameter result, where the initial training parameters and the training parameter results are in one-to-one correspondence.

[0025] In some embodiments, the error optimization process for the initial prediction model based on the training parameter result and the vertex deformation parameter, and generating the deformation parameter prediction model includes:

[0026] Calculate the loss function result based on the vertex deformation parameter and the parameter prediction value;

[0027] Backpropagate the gradient of the loss function result to the initial prediction model for iterative training to generate the deformation parameter prediction model.

[0028] In some embodiments, the selection of key point features from the vertex motion data includes:

[0029] Based on a preset key point quantity threshold, select the key point features from the vertex motion data, where the quantity of the key point features is less than the key point quantity threshold.

[0030] In a second aspect, an embodiment of the present application provides a digital human mouth animation generation device, including:

[0031] A data acquisition module, configured to acquire the audio data to be recognized, perform vertex action operations on the audio data to be recognized, and obtain vertex motion data;

[0032] A parameter prediction module, configured to select key point features from the vertex motion data, and input the key point features into the trained deformation parameter prediction model to obtain a deformation parameter prediction value corresponding to the key point features;

[0033] An animation generation module, configured to generate the digital human mouth animation based on the deformation parameter prediction value.

[0034] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the digital human mouth animation generation method as described in the first aspect above.

[0035] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the digital human mouth animation generation method as described in the first aspect above.

[0036] Compared with the related art, the digital lip animation generation method provided by the embodiments of the present application analyzes and calculates the vertex output results of the existing deep learning solutions, and converts them into Blendshapes sequences, which are common to each digital human model, solves the problem of how to balance the playback smoothness and universality of lip animations, avoids the high cost of retraining new characters, and generates smooth and natural lip animation effects.

[0037] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0039] Figure 1 is a hardware structure block diagram of a terminal of the digital lip animation generation method according to an embodiment of the present invention;

[0040] Figure 2 is a flowchart of the digital lip animation generation method according to an embodiment of the present application;

[0041] Figure 3 is a schematic diagram of the result of the digital lip animation generation method according to an embodiment of the present application;

[0042] Figure 4 is a structure block diagram of the digital lip animation generation device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] In order to make the purpose, technical solutions, and advantages of the present application clearer, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided by the present application without creative efforts belong to the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and time-consuming, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as insufficient disclosure of the content of the present application.

[0044] References to "embodiments" in this application mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase may not necessarily refer to the same embodiment when it appears in various places in the specification, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.

[0045] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meaning as understood by those with ordinary skills in the technical field to which this application pertains. The words such as "a", "an", "one", "the", and the like involved in this application do not indicate a limitation in quantity and can represent a singular or plural number. The terms "including", "comprising", "having", and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include steps or units not listed, or may further include other steps or units inherent to these processes, methods, products, or devices. The terms "connected", "coupled", and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means greater than or equal to two. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0046] The method embodiments provided in this embodiment can be executed on a terminal, a computer, or a similar computing device. Taking running on a terminal as an example, Figure 1 is the hardware structure block diagram of the terminal of the digital human mouth animation generation method according to the embodiments of the present invention. As Figure 1 shown, the terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than those Figure 1 shown in the figure, or may have a structure different from that Figure 1The different configurations shown.

[0047] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the digital lip-sync animation generation method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0048] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RadioFrequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0049] This embodiment provides a digital lip-sync animation generation method. Figure 2 It is a flowchart of the digital lip-sync animation generation method according to the embodiments of the present application, as Figure 2 shown, and the process includes the following steps:

[0050] Step S201, obtain the audio data to be recognized, perform vertex motion operations on the audio data to be recognized, and obtain vertex motion data;

[0051] The audio data to be recognized can be obtained by recording equipment or from existing video files. Specifically, it can be extracted from a video with a clear front face and a clear reading of a person, or it can be extracted from a video in the VOCASET data set. The audio data to be recognized can also be speech synthesized by TTS (Text To Speech), etc. For the acquired audio data to be recognized, the extracted audio data to be recognized is processed using existing deep learning schemes (such as Faceformer, MeshTalk, etc.). These deep learning schemes have been trained to extract features from speech and generate corresponding 3D face vertex motion data. The vertex motion data describes the motion trajectory of each vertex of the face during the speaking process. The vertex motion data is usually used to drive the lip animation of 3D digital people. This step can process audio data from various sources, including real human voices and TTS synthesized sounds, so it has strong flexibility; through the precise calculation of the deep learning scheme, vertex motion data that is highly matched with the speech content can be obtained, providing a solid foundation for the subsequent lip animation generation; using the existing deep learning algorithm for vertex motion calculation, a large amount of vertex motion data can be generated in a short time, improving the overall processing efficiency.

[0052] Step S202, selecting key point features from vertex motion data, and inputting the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features;

[0053] Among them, a series of key point features are selected from the vertex motion data. The key points are usually located in important parts of the face, such as the philtrum, the lowest vertex of the upper lip, the outer protrusion of the upper lip, the upper vertex of the lower lip, the lower vertex of the lower lip, the left corner of the mouth, the right corner of the mouth, the tip of the nose, the left and right corners of the eyes, the midpoint of the cheek, etc. These key points can represent the main deformation of the face when speaking, and the key point features are the coordinate information of the above key points. The selected key point features are input into the trained deformation parameter prediction model. The deformation parameter prediction model is a multi-layer perceptron (MLP) or other neural network model, which is used to map the key point features to the deformation parameters of Blendshapes. The input layer of the model receives the key point features (the coordinate information of the key points), and then extracts and maps the features through multiple hidden layers, and finally outputs the deformation parameter prediction values, which are the deformation parameter prediction values corresponding to the key point features. These prediction values can be the weight values of one or more Blendshapes, which are used to represent the deformation degree of each part of the face when speaking. It should be explained that Blendshapes is a commonly used technology in 3D modeling and animation, which allows animators to create animations by changing the shape of the model; the specific process of producing complex animation effects through Blendshapes is to define multiple deformation targets in the 3D model, each deformation target represents a specific shape or expression of the model, these deformation targets can include various facial expressions such as smile, frown, blink, etc., and then by adjusting the weight of each deformation target, they can be mixed together to produce complex animation effects, and the adjustment of weights can be achieved by manual setting or automatic calculation using algorithms. This step can greatly reduce the computational complexity and improve the running efficiency of the algorithm by selecting key point features for input instead of using all vertex data; since key point features are usually located in important parts of the face, these features have great similarities between different characters, so the model can be better generalized to different characters; by predicting the Blendshapes weight value output by the deformation parameter model, a more natural and smooth face lip animation can be generated, avoiding the problem of stiff animation effects of traditional phoneme schemes.

[0054] Step S203, generating a digital population animation based on the predicted values of the deformation parameters.

[0055] Among them, according to the predicted values of Blendshapes deformation parameters, the facial shape of the 3D digital human model is adjusted. By blending different deformation targets and performing smooth transitions according to weights, a lip-sync animation synchronized with the speech is generated. The generated lip-sync animation sequence is input into a rendering engine (such as Unreal Engine, abbreviated as UE) for final rendering and output. At this time, the user can see the 3D digital human lip-sync animation that is perfectly synchronized with the speech. Through the Blendshapes technology in this step, more realistic and delicate facial expressions and lip movements can be simulated, making the animation effect of the 3D digital human more vivid and lifelike. And because Blendshapes is universal for each digital human, this application can be applied to different 3D digital human models without retraining the model, making the algorithm highly versatile and scalable.

[0056] Through the above steps, the present application first receives the audio data to be recognized, processes the audio data using existing deep learning solutions (such as Faceformer or MeshTalk, etc.), and generates vertex motion data. These vertex data represent the dynamic changes of the 3D digital human face during speech. Compared with the phoneme scheme in the prior art, the present application avoids the problems of rigid animation effects and unnatural action transitions. And compared with other deep learning solutions, although this step is also based on deep learning, the subsequent processing flow makes the results more general and practical. After obtaining the vertex motion data, this step selects the features of key points from it, and inputs these key point features into a pre-trained deformation parameter prediction model. This model predicts the deformation parameter values corresponding to the key point features through a regression task. Compared with the existing deep learning solutions that directly output vertex animation sequences, the present application greatly reduces the complexity of data processing by selecting key points and predicting deformation parameters, while improving the convergence speed and prediction accuracy of the model. And because the predicted deformation parameters are universal for each digital character, the animation data can be applied to different digital characters, avoiding the high cost of retraining new characters. According to the obtained predicted deformation parameter values, the Blendshapes technology is used to generate the lip-sync animation of the digital human. The predicted deformation parameter values are used to drive the changes of these Blendshapes, thereby generating a realistic lip-sync animation. Compared with the vertex-driven scheme in the prior art, the lip-sync animation generated by the present application is more natural and smooth, and is easy to migrate and use among different digital characters. In addition, due to the universality of the Blendshapes technology, the animation data generated by this method can be seamlessly integrated into various 3D rendering engines (such as UE), further broadening the application scope. Therefore, through vertex motion calculation, key point feature selection and deformation parameter prediction, and the application of the Blendshapes technology, the present application effectively solves the limitations of the prior art in voice-driven 3D digital human lip-sync animation, not only improving the realism and smoothness of the animation, but also greatly reducing the production cost and time, providing strong support for applications in fields such as real-time dialogue digital humans, digital human animation generation, and digital human broadcasting.

[0057] In some of these embodiments, the deformation parameter prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer; inputting the key point features into the trained deformation parameter prediction model to obtain the predicted deformation parameter values corresponding to the key point features includes:

[0058] Input the key point features into the first hidden layer to map the key point features to the first feature;

[0059] Input the first feature into the second hidden layer to map the first feature to the second feature;

[0060] Input the second feature into the regression prediction layer to map the second feature to the predicted deformation parameter value.

[0061] Among them, the deformation parameter prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer. The first hidden layer includes a linear layer and an activation function; the input of the first hidden layer is key point features, such as the coordinates of positions like the philtrum, the corners of the mouth, the tip of the nose, etc.; the first hidden layer maps the input key point features to the first feature through the linear layer, which helps to extract the potential relationships between key points; the first hidden layer uses the ReLU activation function to increase the non-linear expression ability of the model and filter out negative value features. The second hidden layer also includes a linear layer and an activation function; the input of the second hidden layer is the first feature output by the first hidden layer; the second hidden layer maps the first feature to a feature space in another dimension, that is, the second feature, to further extract and integrate feature information; the second hidden layer also uses the ReLU activation function to maintain the non-linear characteristics of the model. The regression prediction layer includes a linear layer and an activation function; the input of the regression prediction layer is the second feature output by the second hidden layer; the second feature is mapped to the final predicted deformation parameter value through the linear layer; by using the Sigmoid activation function, the output value is compressed between 0 and 1 to conform to the range of Blendshapes values. In this embodiment, through hierarchical processing, the model structure is clear, and each layer has its specific function and role. Through the feature extraction and integration of multiple hidden layers, the model can more accurately predict the deformation parameters of the mouth animation, thereby generating a more natural and smooth mouth animation effect. And the use of the ReLU activation function increases the non-linear expression ability of the model, enabling the model to fit more complex mouth animation deformation relationships.

[0062] In some embodiments, the method further includes:

[0063] Based on the obtained training video data, obtain the vertex deformation parameters;

[0064] Based on the training video data, extract the training audio data; select training features from the training audio data and input the training features into the initial prediction model to obtain the training parameter results;

[0065] Based on the training parameter results and the vertex deformation parameters, perform error optimization processing on the initial prediction model and generate the deformation parameter prediction model.

[0066] Among them, the training video data is a video with a duration of about 1 hour, showing a clear frontal face and the person reading clearly (or using the videos in the VOCASET dataset). The training audio data is extracted from the training video data. Both the training video data and the training audio data are used to train the deformation parameter prediction model. The training audio data is extracted from the training video data using an audio extraction tool (such as FFmpeg, etc.); through the Live Link Face software or manual annotation, the Blendshapes data corresponding to the training audio data, that is, the vertex deformation parameters, are obtained from the training video data, and then the corresponding relationship between the vertex coordinates and the Blendshapes is established; effective training features are selected from the extracted training audio data, such as the coordinates of positions like the philtrum, corners of the mouth, and the tip of the nose. The selected training features are input into the initial prediction model (such as a multi-layer perceptron MLP, Multi-layer Perceptron) for training to obtain the training parameter results. The training parameter results are compared with the actual vertex deformation parameters, and the error is calculated (such as the mean squared error MSE, Mean SquaredErrorLoss). According to the error, backpropagation is performed to update the parameters of the initial prediction model to reduce the error. Through multiple iterations of training and optimization, the final deformation parameter prediction model is generated. In this embodiment, obtaining the vertex deformation parameters is the key data for training the deformation parameter prediction model. By comparing the training parameter results with the actual vertex deformation parameters and calculating the error for model optimization, the prediction accuracy of the model can be significantly improved, making the deformation parameter prediction model more reliable and stable. The generated deformation parameter prediction model can be directly applied to scenarios such as real-time dialogue digital humans and digital human-generated animations, improving the realism and smoothness of the digital human lip-sync animation.

[0067] In some of these embodiments, the initial prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer; inputting the training features into the initial prediction model to obtain the training parameter results, including:

[0068] Based on a preset first feature quantity, the training features are mapped to first training features via the first hidden layer, and the first feature quantity is the number of first training features;

[0069] Based on a preset second feature quantity, the first training features are mapped to second training features via the second hidden layer, and the second feature quantity is the number of second training features; wherein, the second feature quantity is less than the first feature quantity;

[0070] Based on a preset prediction feature quantity, the second training features are mapped to the training parameter results via the regression prediction layer, and the prediction feature quantity is the number of training parameter results; wherein, the prediction feature quantity is less than the second feature quantity.

[0071] Among them, the initial prediction model (i.e., the multi-layer perceptron MLP) includes a first hidden layer, a second hidden layer, and a regression prediction layer. The training features are input into the first hidden layer, which maps the input features to a new feature space through a linear transformation (i.e., the product of the weight matrix and the input features) to obtain the first training features. This mapping process can be achieved through matrix multiplication, where the dimension of the weight matrix is the number of first features. Then, the ReLU activation function is used to perform a non-linear transformation on the first training features to increase the expressive power of the model. The role of the first hidden layer is to transform the input training features into a feature space that is more suitable for subsequent processing. The ReLU activation function can introduce non-linearity, enabling the model to learn more complex mapping relationships. For example, the first hidden layer maps the input features to 30 features, and the number of first features is 30 (this number of first features can be adjusted up or down as needed during actual training).

[0072] The first training features are input into the second hidden layer, which also maps the first training features to another feature space through a linear transformation to obtain the second training features. This process can also be achieved through matrix multiplication, where the dimension of the weight matrix is the number of second features. Then, the ReLU activation function is used again to perform a non-linear transformation on the second training features. The second hidden layer further abstracts and transforms the features, extracting higher-level features. By reducing the number of features (i.e., the number of second features is less than the number of first features), the complexity and computational amount of the model can be reduced, and it also helps to prevent overfitting. For example, the second hidden layer maps 30 input features to 10 features, and the number of second features is 10 (this number of second features can be adjusted up or down as needed during actual training).

[0073] The second training features are input into the regression prediction layer, which maps the second training features to the final prediction result (i.e., the predicted values of Blendshapes) through a linear transformation. This process can also be achieved through matrix multiplication, where the dimension of the weight matrix is the number of prediction features. The regression prediction layer maps the high-level features to the final prediction result. By performing step-by-step training (i.e., training only one value of Blendshapes at a time, or modifying the output to multiple values of multiple Blendshapes), the training difficulty of the model can be reduced, and the convergence speed and performance of the model can be improved. For example, 10 input features are mapped to 1 feature, and the number of prediction features is 1. By selecting training key points and performing step-by-step training, the training difficulty of the deformation parameter prediction model is reduced, and the convergence speed and performance of the model are improved.

[0074] In some of these embodiments, the regression prediction layer includes a compression activation function; based on a preset number of prediction features, mapping the second training features to the training parameter result, including:

[0075] Map the second training feature to the initial training parameter based on the number of predicted features;

[0076] Use a compression activation function to compress the initial training parameter to a preset weight range to obtain the training parameter result, and the initial training parameter and the training parameter result are in one-to-one correspondence.

[0077] Among them, according to the preset number of predicted features (i.e., the number of Blendshapes or the feature dimension that the model wants to output), the second training feature output by the second hidden layer is further mapped to the initial training parameter. After obtaining the initial training parameter, a compression activation function (such as the Sigmoid activation function) is used to compress the initial training parameter to a preset weight range between 0 and 1. The Sigmoid function is a commonly used activation function that can map any real value to between 0 and 1, thereby ensuring that the output value is within a reasonable range and contributing to the stability and convergence of the model. In this embodiment, by using the Sigmoid activation function to compress the training parameter to between 0 and 1, it is ensured that the output value of the model is within a reasonable range, thereby avoiding problems such as model instability or difficulty in convergence caused by too large or too small output values.

[0078] In some embodiments, based on the training parameter result and the vertex deformation parameter, perform error optimization processing on the initial prediction model and generate a deformation parameter prediction model, including:

[0079] Calculate the loss function result based on the vertex deformation parameter and the parameter prediction value;

[0080] Backpropagate the gradient of the loss function result to the initial prediction model for iterative training to generate a deformation parameter prediction model.

[0081] Among them, a loss function (such as the mean squared error MSELoss) is used to measure the difference between the predicted value and the actual vertex deformation parameter, that is, the error between the vertex deformation parameter and the parameter prediction value, to obtain the loss function result. Backpropagation is a method used in deep learning training to calculate the gradient of model parameters. By using the chain rule, the impact of the error on each neuron is calculated layer by layer starting from the output layer, and then the gradient of each parameter is obtained. Through multiple iterations of backpropagation and parameter update, the neural network gradually adjusts the weights and biases, enabling the model to better fit the training data and improve its generalization ability for new data, generating a trained deformation parameter prediction model. In this embodiment, by calculating the loss function and backpropagating the gradient to update the model parameters, the difference between the predicted value and the actual value can be gradually reduced, thereby improving the prediction accuracy of the model; during the training process, the model not only learns the specific patterns in the training data but also enhances its generalization ability through optimization methods such as gradient descent, that is, it can also make good predictions for new data.

[0082] In some of these embodiments, key-point features are selected from vertex motion data, including:

[0083] Based on a preset key-point quantity threshold, key-point features are selected from vertex motion data, and the quantity of key-point features is less than the key-point quantity threshold.

[0084] Among them, according to the requirements of the complexity of the model and the training efficiency, a key-point quantity threshold is preset. This threshold determines the upper limit of the quantity of key-point features selected from vertex motion data. From the original vertex motion data, according to the preset key-point quantity threshold, representative or key points sensitive to facial expression changes are selected as features. These key points may include several key positions such as the philtrum, the lowest vertex of the upper lip, the protruding point of the upper lip, the upper vertex of the lower lip, the lower vertex of the lower lip, the left corner of the mouth, the right corner of the mouth, the tip of the nose, the left and right eye corners, the midpoint of the cheek, etc. By selecting key-point features in this embodiment, the input dimension of the model can be greatly reduced, thereby reducing the training difficulty and improving the training efficiency; the selected key-point features are usually sensitive to facial expression changes, so they can more accurately reflect facial motion information, thereby improving the prediction performance of the model.

[0085] The embodiments of the present application will be described and illustrated below through preferred embodiments.

[0086] (I) Prepare a dataset:

[0087] 1.1 Original audio data: Prepare a video with a duration of about 1 hour of a frontal face being clear and the person reading clearly (it can also be a video in the VOCASET dataset), and extract the speech in the video.

[0088] Mouth shape action data corresponding to the speech: Use existing deep learning solutions (algorithms such as Faceformer or MeshTalk for speech-driven human face vertex motion are all acceptable) to input the audio data extracted in 1.1, and use existing deep learning solutions (algorithms such as Faceformer or MeshTalk for speech-driven human face vertex motion are all acceptable) to perform inference to obtain vertex action data. (If the video used in the previous step is from the VOCASET dataset, this step can be omitted because the VOCASET dataset has scanned result data of human face vertex animations).

[0089] 1.2 Blendshapes data corresponding to the speech: Use the Live Link Face software or manual annotation method for the video in 1.1 to obtain the Blendshapes data corresponding to the original audio data.

[0090] (II) Build a model from vertex coordinates to Blendshapes:

[0091] 2.1 Data processing: Since the shape of the vertex data of one frame output by Faceformer is 5023X3 and the number of vertices is too large, it is difficult to converge during training. In actual training, the method of only using some key points can be selected to reduce the training difficulty. (Positions of several key points selected in this experiment: the philtrum, the lowest vertex of the upper lip, the protruding point of the upper lip, the upper vertex of the lower lip, the lower vertex of the lower lip, the left corner of the mouth, the right corner of the mouth, the tip of the nose, the left and right eye corners, the midpoint of the cheek, etc.).

[0092] 2.2 Model selection and training:

[0093] Here, a multi-layer perceptron (MLP) model can be selected for the regression task, which consists of multiple linear layers and non-linear activation functions. The specific structure is as follows:

[0094] Input layer: The input feature is the coordinates of the selected key points.

[0095] The first hidden layer: The linear layer maps the input features to 30 features (the 30 features can be adjusted up and down as needed during actual training); the activation function is the ReLU activation function.

[0096] The second hidden layer: The linear layer maps 30 input features to 10 features (the 10 features can be adjusted up and down as needed during actual training); the activation function is the ReLU activation function.

[0097] Regression prediction layer: The linear layer maps 10 input features to 1 feature; the Sigmoid activation function is applied after the regression prediction layer to compress the output value between 0 and 1.

[0098] The above structure is for training one value of Blendshapes. It can also be modified to output multiple values of multiple Blendshapes. However, if only one value is trained at a time, it is easier to converge and the model effect is better. Train the model for multiple required values step by step.

[0099] 2.3 During training, the input of the above model is the selected point information, and the predicted result output is the predicted value of Blendshapes between 0 and 1. The predicted value of Blendshapes and the corresponding actual Blendshapes value use MSELoss (mean squared error) to calculate the loss and perform backpropagation.

[0100] 2.4 Inferring Blendshapes: Figure 3 It is a schematic diagram of the result of the digital lip animation generation method according to the embodiment of the present application. Please refer to Figure 3, when in use, first input the audio into the Faceformer model to output the vertex motion animation sequence, and then use the vertex motion animation sequence as the input to the model to output the Blendshapes, which can drive digital mouth animation such as UE as Figure 3 the digital mouth animation.

[0101] This preferred embodiment is used for voice-driven 3D digital mouth actions. It can also input text for TTS synthetic voice, input the synthetic voice into this solution to output the Blendshapes sequence, and input the Blendshapes sequence into UE to render the digital human animation sequence.

[0102] This embodiment also provides a digital mouth animation generation device, which is used to implement the above embodiments and preferred embodiments, and those that have been described will not be repeated. As used below, terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0103] Figure 4 is a structural block diagram of the digital mouth animation generation device according to an embodiment of the present application, as Figure 4 shown, the device includes:

[0104] A data acquisition module 10, configured to acquire the audio data to be recognized, perform vertex motion operations on the audio data to be recognized, and obtain vertex motion data;

[0105] A parameter prediction module 20, configured to select key point features from the vertex motion data, and input the key point features into the trained deformation parameter prediction model to obtain the deformation parameter prediction values corresponding to the key point features;

[0106] An animation generation module 30, configured to generate a digital mouth animation based on the deformation parameter prediction values.

[0107] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combination form.

[0108] This embodiment also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0109] Optionally, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0110] Optionally, in this embodiment, the above processor may be configured to perform the following steps by a computer program:

[0111] S1, obtain the audio data to be recognized, perform vertex motion operations on the audio data to be recognized, and obtain vertex motion data;

[0112] S2, select key-point features from the vertex motion data, and input the key-point features into the trained deformation parameter prediction model to obtain the predicted deformation parameter values corresponding to the key-point features;

[0113] S3, generate a digital mouth shape animation based on the predicted deformation parameter values.

[0114] It should be noted that the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.

[0115] In addition, in combination with the digital mouth shape animation generation method in the above embodiments, an embodiment of the present application can be implemented by providing a storage medium. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the digital mouth shape animation generation methods in the above embodiments is implemented.

[0116] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0117] The above embodiments only represent several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for generating a digital population-based animation, characterized in that, Including: Obtain the audio data to be recognized, perform vertex motion operations on the audio data to be recognized, and obtain vertex motion data; Select key-point features from the vertex motion data, including: Based on a preset key-point quantity threshold, select the key-point features from the vertex motion data, and the quantity of the key-point features is less than the key-point quantity threshold; Input the key-point features into the trained deformation parameter prediction model to obtain a deformation parameter prediction value corresponding to the key-point features, including: The deformation parameter prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer; The key-point features are input into the first hidden layer, and the key-point features are mapped to a first feature; The first feature is input into the second hidden layer, and the first feature is mapped to a second feature; The second feature is input into the regression prediction layer, and the second feature is mapped to the deformation parameter prediction value; Generate the digital mouth shape animation based on the deformation parameter prediction value.

2. The digital population type animation generation method according to claim 1, wherein The method further includes: Based on the obtained training video data, obtain vertex deformation parameters; Based on the training video data, extract training audio data; select training features from the training audio data, and input the training features into an initial prediction model to obtain a training parameter result; Based on the training parameter result and the vertex deformation parameters, perform error optimization processing on the initial prediction model and generate the deformation parameter prediction model.

3. The digital population type animation generation method according to claim 2, wherein The initial prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer; the inputting the training features into the initial prediction model to obtain a training parameter result includes: Based on a preset first feature quantity, map the training features to first training features via the first hidden layer, and the first feature quantity is the number of the first training features; Based on a preset second feature quantity, map the first training features to second training features via the second hidden layer, and the second feature quantity is the number of the second training features; wherein, the second feature quantity is less than the first feature quantity; Based on a preset prediction feature quantity, map the second training features to a training parameter result via the regression prediction layer, and the prediction feature quantity is the number of the training parameter result; wherein, the prediction feature quantity is less than the second feature quantity.

4. The digital population animation generation method according to claim 3, wherein The regression prediction layer includes a compression activation function; the mapping the second training features to a training parameter result based on a preset prediction feature quantity includes: Based on the prediction feature quantity, map the second training features to training initial parameters; Use the compression activation function to compress the training initial parameters into a preset weight range to obtain the training parameter result, and the training initial parameters and the training parameter result are in one-to-one correspondence.

5. The digital population animation generation method according to claim 2, characterized in that The performing error optimization processing on the initial prediction model based on the training parameter result and the vertex deformation parameters and generating the deformation parameter prediction model includes: Calculate a loss function result based on the vertex deformation parameters and the parameter prediction value; Backpropagate the gradient of the loss function result to the initial prediction model for iterative training to generate the deformation parameter prediction model.

6. A digital population type animation generation device, characterized in that, It includes: A data acquisition module, configured to acquire audio data to be recognized, perform vertex motion operations on the audio data to be recognized, and obtain vertex motion data; A parameter prediction module, configured to select key point features from the vertex motion data, including: Based on a preset key point quantity threshold, select the key point features from the vertex motion data, and the quantity of the key point features is less than the key point quantity threshold; Input the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features, including: The deformation parameter prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer; Input the key point features into the first hidden layer to map the key point features to a first feature; Input the first feature into the second hidden layer to map the first feature to a second feature; Input the second feature into the regression prediction layer to map the second feature to the deformation parameter prediction values; An animation generation module, configured to generate the digital lip-sync animation based on the deformation parameter prediction values.

7. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the digital lip-sync animation generation method according to any one of claims 1 to 5.

8. A storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is configured to execute the digital lip-sync animation generation method according to any one of claims 1 to 5 when running.

Citation Information

Patent Citations

  • Three-dimensional face animation generation method and device based on audio driving and medium

    CN116309988A