Digital population animation generation method and device, electronic device and storage medium
By extracting key point features and predicting deformation parameters of vertex motion data of digital vernal animations, a general Blendshapes sequence is generated, which solves the problems of fluency and universality of digital vernal animations and realizes cost-effective animation generation.
Patent Information
- Application Number
- CN202510406947.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The prior art is difficult to take into account the smoothness and versatility of digital verb animations at the same time, resulting in high cost and inapplicable to different digital characters.
By obtaining the audio data to be identified, performing vertex action operations to obtain vertex motion data, selecting key point features and inputting them into the trained deformation parameter prediction model, and generating a Blendshapes sequence, which is common to each digital character model.
It realizes the smooth and natural effect of digital verbal animations, and avoids the high cost of retraining new characters, and has high versatility and scalability.
Smart Images

Figure CN119991892A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a method, device, electronic device and storage medium for generating digital population animation. Background Art
[0002] In the current digital age, 3D digital humans are increasingly being used in a wide range of fields, such as virtual anchors, online education, and virtual customer service. In order to achieve natural interaction with 3D digital humans, especially oral expression, it is crucial to accurately drive their lip movements.
[0003] Traditional voice-driven lip animation technology mainly relies on phoneme solutions and deep learning solutions. Although the phoneme solution is simple to implement, the animation effect is stiff and it is difficult to show natural and smooth lip transitions. The deep learning solution relies on high-quality actor lip data sets for model training. Although it can generate more natural lip animations, most of them are closed source and charged, and the generated vertex-driven animations cannot be directly applied to different digital characters. There are problems of high cost and poor versatility.
[0004] Currently, no effective solution has been proposed for the problem of how to balance the playback smoothness and versatility of lip-sync animation in related technologies. Summary of the invention
[0005] The embodiments of the present application provide a method, device, electronic device and storage medium for generating a digital mouth shape animation, so as to at least solve the problem in the related art of how to simultaneously take into account the playback fluency and versatility of the lip shape animation.
[0006] In a first aspect, an embodiment of the present application provides a method for generating a digital population animation, comprising: Acquire audio data to be recognized, perform vertex motion calculation on the audio data to be recognized, and obtain vertex motion data; Selecting key point features from the vertex motion data, and inputting the key point features into a trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features; The digital population animation is generated based on the predicted values of the deformation parameters.
[0007] In some embodiments, the deformation parameter prediction model includes a first hidden layer, a second hidden layer and a regression prediction layer; the step of inputting the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features includes: Inputting the key point feature into the first hidden layer, mapping the key point feature to the first feature; Inputting the first feature into the second hidden layer, mapping the first feature to the second feature; The second feature is input into the regression prediction layer, and the second feature is mapped to the deformation parameter prediction value.
[0008] In some embodiments, the method further comprises: Based on the acquired training video data, vertex deformation parameters are obtained; Based on the training video data, extracting training audio data; selecting training features from the training audio data, inputting the training features into an initial prediction model, and obtaining training parameter results; Based on the training parameter results and the vertex deformation parameters, the initial prediction model is subjected to error optimization processing, and the deformation parameter prediction model is generated.
[0009] In some embodiments, the initial prediction model includes a first hidden layer, a second hidden layer and a regression prediction layer; the inputting the training features into the initial prediction model to obtain the training parameter results includes: Based on a preset first feature quantity, mapping the training features to first training features via the first hidden layer, where the first feature quantity is the number of the first training features; Based on a preset second feature quantity, mapping the first training feature to a second training feature via the second hidden layer, the second feature quantity being the number of the second training features; wherein the second feature quantity is less than the first feature quantity; Based on a preset number of prediction features, the second training feature is mapped to the training parameter result via the regression prediction layer, and the number of prediction features is the number of the training parameter results; wherein the number of prediction features is less than the second number of features.
[0010] In some embodiments, the regression prediction layer includes a compression activation function; and mapping the second training feature to a training parameter result based on a preset number of prediction features includes: Based on the number of predicted features, mapping the second training features to initial training parameters; The compression activation function is used to compress the initial training parameters into a preset weight range to obtain the training parameter results, and the initial training parameters correspond to the training parameter results one by one.
[0011] In some embodiments, the error optimization process is performed on the initial prediction model based on the training parameter result and the vertex deformation parameter, and the deformation parameter prediction model is generated, including: Calculating a loss function result based on the vertex deformation parameters and the parameter prediction values; The gradient of the loss function result is back-transferred to the initial prediction model for iterative training to generate the deformation parameter prediction model.
[0012] In some embodiments, the step of extracting key point features from the vertex motion data includes: Based on a preset key point quantity threshold, the key point features are selected from the vertex motion data, and the number of the key point features is less than the key point quantity threshold.
[0013] In a second aspect, an embodiment of the present application provides a digital population animation generation device, comprising: A data acquisition module, used for acquiring audio data to be recognized, performing vertex motion calculation on the audio data to be recognized, and obtaining vertex motion data; A parameter prediction module, used for selecting key point features from the vertex motion data, and inputting the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features; The animation generation module is used to generate the digital population animation based on the predicted value of the deformation parameter.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for generating digital population animation as described in the first aspect above is implemented.
[0015] In a fourth aspect, an embodiment of the present application provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating a digital population animation as described in the first aspect above.
[0016] Compared with the related art, the digital population animation generation method provided in the embodiment of the present application analyzes and calculates the vertex output results of the existing deep learning solution, and converts them into a Blendshapes sequence, which is universal for all digital character models. It solves the problem of how to simultaneously take into account the smoothness and universality of the playback of lip-sync animation, avoids the high cost of retraining new characters, and generates a smooth and natural lip-sync animation effect.
[0017] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 It is a hardware structure block diagram of a terminal of a method for generating a digital population animation according to an embodiment of the present invention; Figure 2 is a flow chart of a method for generating a digital population animation according to an embodiment of the present application; Figure 3 is a schematic diagram of the result of the method for generating a digital population animation according to an embodiment of the present application; Figure 4 It is a structural block diagram of a digital population animation generation device according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for ordinary technicians in the field related to the contents disclosed in the present application, some changes such as design, manufacturing or production based on the technical contents disclosed in the present application are only conventional technical means, and should not be understood as insufficient contents disclosed in the present application.
[0020] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0021] Unless otherwise defined, the technical terms or scientific terms involved in this application should be understood by people with ordinary skills in the technical field to which this application belongs. The words "one", "a", "a", "the" and the like involved in this application do not indicate a quantity limitation, and may indicate the singular or plural. The terms "include", "comprise", "have" and any of their variations involved in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application means greater than or equal to two. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. The terms "first", "second", "third" and the like involved in the present application are merely used to distinguish similar objects and do not represent a specific ordering of the objects.
[0022] The method embodiment provided in this embodiment can be executed in a terminal, a computer or a similar computing device. Taking running on a terminal as an example, Figure 1 FIG. 1 is a hardware structure diagram of a terminal of a method for generating a digital population animation according to an embodiment of the present invention. Figure 1 As shown, the terminal may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is for illustration only and does not limit the structure of the above terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0023] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the digital population animation generation method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0024] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (Radio Frequency, referred to as RF) module, which is used to communicate with the Internet wirelessly.
[0025] This embodiment provides a method for generating a digital population animation. Figure 2 is a flow chart of a method for generating a digital population animation according to an embodiment of the present application. Figure 2 As shown, the process includes the following steps: Step S201, obtaining audio data to be recognized, performing vertex motion calculation on the audio data to be recognized, and obtaining vertex motion data; The audio data to be recognized can be obtained by recording equipment or from existing video files. Specifically, it can be extracted from a video with a clear front face and a clear reading of a person, or it can be extracted from a video in the VOCASET data set. The audio data to be recognized can also be speech synthesized by TTS (Text To Speech), etc. For the acquired audio data to be recognized, the extracted audio data to be recognized is processed using existing deep learning schemes (such as Faceformer, MeshTalk, etc.). These deep learning schemes have been trained to extract features from speech and generate corresponding 3D face vertex motion data. The vertex motion data describes the motion trajectory of each vertex of the face during the speaking process. The vertex motion data is usually used to drive the lip animation of 3D digital people. This step can process audio data from various sources, including real human voices and TTS synthesized sounds, so it has strong flexibility; through the precise calculation of the deep learning scheme, vertex motion data that is highly matched with the speech content can be obtained, providing a solid foundation for the subsequent lip animation generation; using the existing deep learning algorithm for vertex motion calculation, a large amount of vertex motion data can be generated in a short time, improving the overall processing efficiency.
[0026] Step S202, selecting key point features from vertex motion data, and inputting the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features; Among them, a series of key point features are selected from the vertex motion data. The key points are usually located in important parts of the face, such as the philtrum, the lowest vertex of the upper lip, the outer protrusion of the upper lip, the upper vertex of the lower lip, the lower vertex of the lower lip, the left corner of the mouth, the right corner of the mouth, the tip of the nose, the left and right corners of the eyes, the midpoint of the cheek, etc. These key points can represent the main deformation of the face when speaking, and the key point features are the coordinate information of the above key points. The selected key point features are input into the trained deformation parameter prediction model. The deformation parameter prediction model is a multi-layer perceptron (MLP) or other neural network model, which is used to map the key point features to the deformation parameters of Blendshapes. The input layer of the model receives the key point features (the coordinate information of the key points), and then extracts and maps the features through multiple hidden layers, and finally outputs the deformation parameter prediction values, which are the deformation parameter prediction values corresponding to the key point features. These prediction values can be the weight values of one or more Blendshapes, which are used to represent the deformation degree of each part of the face when speaking. It should be explained that Blendshapes is a commonly used technology in 3D modeling and animation, which allows animators to create animations by changing the shape of the model; the specific process of producing complex animation effects through Blendshapes is to define multiple deformation targets in the 3D model, each deformation target represents a specific shape or expression of the model, these deformation targets can include various facial expressions such as smile, frown, blink, etc., and then by adjusting the weight of each deformation target, they can be mixed together to produce complex animation effects, and the adjustment of weights can be achieved by manual setting or automatic calculation using algorithms. This step can greatly reduce the computational complexity and improve the running efficiency of the algorithm by selecting key point features for input instead of using all vertex data; since key point features are usually located in important parts of the face, these features have great similarities between different characters, so the model can be better generalized to different characters; by predicting the Blendshapes weight value output by the deformation parameter model, a more natural and smooth face lip animation can be generated, avoiding the problem of stiff animation effects of traditional phoneme schemes.
[0027] Step S203, generating a digital population animation based on the predicted values of the deformation parameters.
[0028] Among them, according to the predicted value of the Blendshapes deformation parameter, the facial shape of the 3D digital human model is adjusted, and by mixing different deformation targets and performing smooth transition according to the weights, a lip animation synchronized with the speech is generated; the generated lip animation sequence is input into the rendering engine (such as Unreal Engine, UE for short) for final rendering and output, at which point the user can see the 3D digital human mouth animation perfectly synchronized with the speech. This step uses the Blendshapes technology to simulate more realistic and delicate facial expressions and lip movements, which makes the animation effect of the 3D digital human more vivid and realistic, and because Blendshapes is universal for all digital characters, this application can be applied to different 3D digital human models without retraining the model, which makes the algorithm highly versatile and scalable.
[0029] Through the above steps, the present application first receives the audio data to be recognized, processes the audio data using existing deep learning solutions (such as Faceformer or MeshTalk, etc.), and generates vertex motion data. These vertex data represent the dynamic changes of the 3D digital human face when speaking. Compared with the phoneme solution in the prior art, the present application avoids the problems of stiff animation effects and unnatural motion transitions. Compared with other deep learning solutions, although this step is also based on deep learning, the subsequent processing flow makes the result more universal and practical. After obtaining the vertex motion data, this step selects the features of key points from it, and inputs these key point features into a pre-trained deformation parameter prediction model. The model predicts the deformation parameter values corresponding to the key point features through regression tasks. Compared with the existing deep learning solution that directly outputs the vertex animation sequence, the present application greatly reduces the complexity of data processing by selecting key points and predicting deformation parameters, while improving the convergence speed and prediction accuracy of the model. Moreover, since the predicted deformation parameters are universal for each digital character, the animation data can be applied to different digital characters, avoiding the high cost of retraining new characters. According to the obtained deformation parameter prediction values, the Blendshapes technology is used to generate the lip animation of the digital human. The deformation parameter prediction values are used to drive the changes of these Blendshapes, thereby generating realistic lip animation. Compared with the vertex drive scheme in the prior art, the lip animation generated by this application is more natural and smooth, and easy to migrate and use between different digital characters. In addition, due to the versatility of the Blendshapes technology, the animation data generated by this method can be seamlessly integrated into various 3D rendering engines (such as UE), further broadening the scope of application. Therefore, this application effectively solves the limitations of the prior art in voice-driven 3D digital human mouth shape animation through vertex action calculation, key point feature selection and deformation parameter prediction, and the application of Blendshapes technology, which not only improves the realism and fluency of the animation, but also greatly reduces the production cost and time, and provides strong support for applications in the fields of real-time dialogue with digital humans, digital human animation generation, and digital human broadcasting.
[0030] In some embodiments, the deformation parameter prediction model includes a first hidden layer, a second hidden layer and a regression prediction layer; the key point features are input into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features, including: Inputting the key point feature into the first hidden layer, mapping the key point feature to the first feature; Inputting the first feature into the second hidden layer, mapping the first feature to the second feature; The second feature is input into the regression prediction layer, and the second feature is mapped to the deformation parameter prediction value.
[0031] Among them, the deformation parameter prediction model includes a first hidden layer, a second hidden layer and a regression prediction layer. The first hidden layer includes a linear layer and an activation function; the input of the first hidden layer is a key point feature, such as the coordinates of the philtrum, the corner of the mouth, the tip of the nose and other positions; the first hidden layer maps the input key point feature to the first feature through the linear layer, which helps to extract the potential relationship between the key points; the first hidden layer uses the ReLU activation function to increase the nonlinear expression ability of the model and filter out negative value features. The second hidden layer also includes a linear layer and an activation function; the input of the second hidden layer is the first feature output by the first hidden layer; the second hidden layer maps the first feature to another dimension of feature space, namely the second feature, through the linear layer, to further extract and integrate feature information; the second hidden layer also uses the ReLU activation function to maintain the nonlinear characteristics of the model. The regression prediction layer includes a linear layer and an activation function; the input of the regression prediction layer is the second feature output by the second hidden layer; the second feature is mapped to the final deformation parameter prediction value through the linear layer; the output value is compressed to between 0 and 1 by using the Sigmoid activation function to meet the range of the Blendshapes value. Through layered processing, the model structure of this embodiment is clear and each layer has its specific function and role. Through feature extraction and integration of multiple hidden layers, the model can more accurately predict the deformation parameters of lip animation, thereby generating a more natural and smooth lip animation effect. In addition, the use of the ReLU activation function increases the nonlinear expression ability of the model, so that the model can fit more complex lip animation deformation relationships.
[0032] In some embodiments, the method further comprises: Based on the acquired training video data, vertex deformation parameters are obtained; Based on the training video data, training audio data is extracted; training features are selected from the training audio data, and the training features are input into the initial prediction model to obtain training parameter results; Based on the training parameter results and vertex deformation parameters, the initial prediction model is error optimized and a deformation parameter prediction model is generated.
[0033] The training video data is a video of about 1 hour with a clear frontal face and a clear reading by the person (or a video in the VOCASET dataset). The training audio data is extracted from the training video data. Both the training video data and the training audio data are used to train the deformation parameter prediction model. The training audio data is extracted from the training video data using an audio extraction tool (such as FFmpeg, etc.); the Blendshapes data corresponding to the training audio data is obtained from the training video data through Live Link Face software or manual annotation, so as to establish a corresponding relationship between vertex coordinates and Blendshapes; effective training features are selected from the extracted training audio data, such as the coordinates of the philtrum, corners of the mouth, and tip of the nose, and the selected training features are input into the initial prediction model (such as a multi-layer perceptron MLP, Multi-layer Perceptron) for training and obtaining training parameter results. The training parameter results are compared with the actual vertex deformation parameters, and the error (such as mean square error MSE, Mean SquaredErrorLoss) is calculated. Back propagation is performed based on the error to update the parameters of the initial prediction model to reduce the error. Through multiple iterations of training and optimization, the final deformation parameter prediction model is generated. In this embodiment, the vertex deformation parameters are obtained as key data for training the deformation parameter prediction model. By comparing the training parameter results with the actual vertex deformation parameters and calculating the error to optimize the model, the prediction accuracy of the model can be significantly improved, making the deformation parameter prediction model more reliable and stable. The generated deformation parameter prediction model can be directly applied to scenes such as real-time conversations with digital humans and digital human animation generation, so as to improve the fidelity and fluency of digital human animation.
[0034] In some embodiments, the initial prediction model includes a first hidden layer, a second hidden layer, and a regression prediction layer; the training features are input into the initial prediction model to obtain training parameter results, including: Based on a preset first feature quantity, mapping the training features to first training features via a first hidden layer, where the first feature quantity is the number of first training features; Based on a preset second feature quantity, mapping the first training feature to the second training feature via the second hidden layer, the second feature quantity being the number of the second training features; wherein the second feature quantity is less than the first feature quantity; Based on a preset number of prediction features, the second training feature is mapped to the training parameter result via the regression prediction layer, and the number of prediction features is the number of training parameter results; wherein the number of prediction features is less than the number of second features.
[0035] Among them, the initial prediction model (i.e., multi-layer perceptron MLP) includes a first hidden layer, a second hidden layer, and a regression prediction layer. The training features are input into the first hidden layer, which maps the input features to a new feature space through a linear transformation (i.e., the product of the weight matrix and the input features) to obtain the first training features. This mapping process can be achieved through matrix multiplication, where the dimension of the weight matrix is the number of first features. Afterwards, the first training features are nonlinearly transformed through the ReLU activation function to increase the expressive power of the model. The function of the first hidden layer is to convert the input training features into a feature space that is more suitable for subsequent processing. The ReLU activation function can introduce nonlinearity, so that the model can learn more complex mapping relationships. For example, the first hidden layer maps the input features to 30 features, and the number of first features is 30 (in actual training, the number of first features can be adjusted up and down as needed).
[0036] The first training feature is input to the second hidden layer, which also maps the first training feature to another feature space through linear transformation to obtain the second training feature. This process can also be achieved through matrix multiplication, where the dimension of the weight matrix is the number of second features. After that, the second training feature is again nonlinearly transformed through the ReLU activation function. The second hidden layer further abstracts and transforms the features to extract higher-level features. By reducing the number of features (that is, the number of second features is less than the number of first features), the complexity and computational complexity of the model can be reduced, and it also helps prevent overfitting. For example, the second hidden layer maps 30 input features to 10 features, and the number of second features is 10 (in actual training, the number of second features can be adjusted up and down as needed).
[0037] The second training feature is input to the regression prediction layer, which maps the second training feature to the final prediction result (i.e., the predicted value of Blendshapes) through linear transformation. This process can also be achieved through matrix multiplication, where the dimension of the weight matrix is the number of predicted features. The regression prediction layer maps high-level features to the final prediction result. Through step-by-step training (i.e., only one Blendshapes value is trained at a time, and the output can also be modified to multiple values of multiple Blendshapes), the training difficulty of the model can be reduced and the convergence speed and performance of the model can be improved. For example, 10 input features are mapped to 1 feature, and the number of predicted features is 1. By selecting training key points and step-by-step training, the training difficulty of the deformation parameter prediction model is reduced, and the convergence speed and performance of the model are improved.
[0038] In some embodiments, the regression prediction layer includes a compression activation function; based on a preset number of prediction features, mapping the second training feature to a training parameter result includes: Based on the number of predicted features, mapping the second training features to initial training parameters; The compression activation function is used to compress the initial training parameters into the preset weight range to obtain the training parameter results. The initial training parameters correspond to the training parameter results one by one.
[0039] Among them, according to the preset number of prediction features (i.e., the number of Blendshapes or the feature dimension you want the model to output), the second training feature output by the second hidden layer is further mapped to the initial training parameters. After obtaining the initial training parameters, a compression activation function (such as a Sigmoid activation function) is used to compress the initial training parameters to a preset weight range between 0 and 1. The Sigmoid function is a commonly used activation function that can map any real value to between 0 and 1, thereby ensuring that the output value is within a reasonable range and contributing to the stability and convergence of the model. This embodiment compresses the training parameters to between 0 and 1 using the Sigmoid activation function, thereby ensuring that the output value of the model is within a reasonable range, thereby avoiding the problem of model instability or difficulty in convergence due to excessively large or small output values.
[0040] In some embodiments, based on the training parameter results and the vertex deformation parameters, the initial prediction model is subjected to error optimization processing, and a deformation parameter prediction model is generated, including: Calculate the loss function result based on vertex deformation parameters and parameter prediction values; The gradient of the loss function result is transferred back to the initial prediction model for iterative training to generate a deformation parameter prediction model.
[0041] Among them, a loss function (such as mean square error MSELoss) is used to measure the difference between the predicted value and the actual vertex deformation parameter, that is, the error between the vertex deformation parameter and the parameter prediction value, to obtain the loss function result. Backpropagation is a method used to calculate the gradient of model parameters in deep learning training. The chain rule is used to calculate the impact of the error on each neuron layer by layer starting from the output layer, and then the gradient of each parameter is obtained. Through multiple iterations of backpropagation and parameter updates, the neural network gradually adjusts the weights and biases, so that the model can better fit the training data and improve the generalization ability of new data, and generate a trained deformation parameter prediction model. This embodiment updates the model parameters by calculating the loss function and backpropagating the gradient, which can gradually reduce the difference between the predicted value and the actual value, thereby improving the prediction accuracy of the model; during the training process, the model not only learns the specific patterns in the training data, but also enhances its generalization ability through optimization methods such as gradient descent, that is, it can also make better predictions for new data.
[0042] In some embodiments, key point features are extracted from vertex motion data, including: Based on a preset key point quantity threshold, key point features are selected from the vertex motion data, and the number of key point features is less than the key point quantity threshold.
[0043] Among them, according to the complexity of the model and the requirements of training efficiency, a key point number threshold is preset. This threshold determines the upper limit of the number of key point features selected from the vertex motion data. From the original vertex motion data, according to the preset key point number threshold, key points that are representative or sensitive to changes in facial expressions are selected as features. These key points may include the philtrum, the lowest vertex of the upper lip, the outer protrusion of the upper lip, the upper vertex of the lower lip, the lower vertex of the lower lip, the left corner of the mouth, the right corner of the mouth, the tip of the nose, the left and right corners of the eyes, the midpoint of the cheek and other parts. By selecting key point features, this embodiment can greatly reduce the input dimension of the model, thereby reducing the difficulty of training and improving training efficiency; the selected key point features are usually sensitive to changes in facial expressions, and therefore can more accurately reflect facial motion information, thereby improving the prediction performance of the model.
[0044] The embodiments of the present application are described and illustrated below through preferred embodiments.
[0045] 1. Prepare the dataset: 1.1 Original audio data: Prepare a video of about 1 hour in length with a clear frontal face and a clear reading by the person (it can also be a video in the VOCASET dataset), and extract the speech in the video.
[0046] Lip movement data corresponding to speech: Use existing deep learning solutions (Faceformer or MeshTalk, etc., which can drive facial vertex motion with speech) to extract audio data from 1.1, and use existing deep learning solutions (Faceformer or MeshTalk, etc., which can drive facial vertex motion with speech) to infer vertex motion data. (If the previous step uses a video from the VOCASET dataset, this step can be omitted, because the VOCASET dataset contains scanned data of facial vertex animations).
[0047] 1.2 Blendshapes data corresponding to speech: Use Live Link Face software or manual annotation to obtain the Blendshapes data corresponding to the original audio data of the video in 1.1.
[0048] (ii) Build vertex coordinates into Blendshapes model: 2.1 Data processing: Since the shape of one frame of vertex data output by Faceformer is 5023X3, it will be difficult to converge when the number of vertices is too large. In actual training, you can choose to use only some key points to reduce the difficulty of training. (This experiment selects several key points at the following locations: philtrum, the lowest vertex of the upper lip, the outer protrusion of the upper lip, the upper vertex of the lower lip, the lower vertex of the lower lip, the left corner of the mouth, the right corner of the mouth, the tip of the nose, the left and right corners of the eyes, the midpoint of the cheek, etc.).
[0049] 2.2 Model selection and training: Here you can choose a multi-layer perceptron (MLP) model for regression tasks, which consists of multiple linear layers and nonlinear activation functions. The specific structure is as follows: Input layer: The input features are the coordinates of the selected key points.
[0050] First hidden layer: The linear layer maps the input features to 30 features (these 30 features can be adjusted up or down as needed in actual training); the activation function is the ReLU activation function.
[0051] Second hidden layer: The linear layer maps 30 input features to 10 features (these 10 features can be adjusted up or down as needed in actual training); the activation function is the ReLU activation function.
[0052] Regression prediction layer: The linear layer maps 10 input features to 1 feature; the Sigmoid activation function is applied after the regression prediction layer to compress the output value between 0 and 1.
[0053] The above structure is to train one value of Blendshapes, and the output can also be modified to multiple values of multiple Blendshapes. However, if only one value is trained at a time, it is easier to converge and the model effect is better. The model of multiple required values can be trained step by step.
[0054] 2.3 During training, the input of the above model is the selected point information, and the output prediction result is the Blendshapes prediction value between 0 and 1. The Blendshapes prediction value and the corresponding actual Blendshapes value use MSELoss (mean square error) to calculate the loss and perform back propagation.
[0055] 2.4 Inference to get Blendshapes: Figure 3 This is a schematic diagram of the result of the method for generating a digital population animation according to an embodiment of the present application. Figure 3 When using it, first input the audio to the Faceformer model to output the vertex motion animation sequence, then use the vertex motion animation sequence as input to the model, and output Blendshapes. Blendshapes can drive UE, etc. Figure 3Digital population animation.
[0056] This preferred embodiment is used to drive the 3D digital population action by voice. Text can also be input for TTS synthesized speech. The synthesized speech is input into this solution to output a Blendshapes sequence. The Blendshapes sequence is input into the UE to render a digital human animation sequence.
[0057] This embodiment also provides a digital population animation generation device, which is used to implement the above-mentioned embodiments and preferred implementations, and will not be repeated here. As used below, the terms "module", "unit", "subunit", etc. can implement a combination of software and / or hardware for a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0058] Figure 4 is a structural block diagram of a digital population animation generating device according to an embodiment of the present application, such as Figure 4 As shown, the device comprises: The data acquisition module 10 is used to acquire the audio data to be recognized, perform vertex motion calculation on the audio data to be recognized, and obtain vertex motion data; The parameter prediction module 20 is used to select key point features from the vertex motion data and input the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features; The animation generation module 30 is used to generate a digital population animation based on the predicted values of the deformation parameters.
[0059] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0060] This embodiment further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0061] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0062] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program: S1, obtaining audio data to be recognized, performing vertex motion calculation on the audio data to be recognized, and obtaining vertex motion data; S2, selecting key point features from vertex motion data, and inputting the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features; S3, generating a digital population animation based on the predicted values of the deformation parameters.
[0063] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0064] In addition, in combination with the digital population animation generation method in the above embodiment, the present application embodiment can provide a storage medium for implementation. The storage medium stores a computer program; when the computer program is executed by a processor, any digital population animation generation method in the above embodiment is implemented.
[0065] Those skilled in the art should understand that the technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0066] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for generating a digital population animation, characterized in that: include: Acquire audio data to be recognized, perform vertex motion calculation on the audio data to be recognized, and obtain vertex motion data; Selecting key point features from the vertex motion data, and inputting the key point features into a trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features; The digital population animation is generated based on the predicted values of the deformation parameters.
2. The method for generating digital population animation according to claim 1, characterized in that: The deformation parameter prediction model includes a first hidden layer, a second hidden layer and a regression prediction layer; the step of inputting the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features includes: Inputting the key point feature into the first hidden layer, mapping the key point feature to the first feature; Inputting the first feature into the second hidden layer, mapping the first feature to the second feature; The second feature is input into the regression prediction layer, and the second feature is mapped to the deformation parameter prediction value.
3. The method for generating digital population animation according to claim 1, characterized in that: The method further comprises: Based on the acquired training video data, vertex deformation parameters are obtained; Based on the training video data, extracting training audio data; selecting training features from the training audio data, inputting the training features into an initial prediction model, and obtaining training parameter results; Based on the training parameter results and the vertex deformation parameters, the initial prediction model is subjected to error optimization processing, and the deformation parameter prediction model is generated.
4. The method for generating digital population animation according to claim 3, characterized in that: The initial prediction model includes a first hidden layer, a second hidden layer and a regression prediction layer; the training features are input into the initial prediction model to obtain training parameter results, including: Based on a preset first feature quantity, mapping the training features to first training features via the first hidden layer, where the first feature quantity is the number of the first training features; Based on a preset second feature quantity, mapping the first training feature to a second training feature via the second hidden layer, the second feature quantity being the number of the second training features; wherein the second feature quantity is less than the first feature quantity; Based on a preset number of prediction features, the second training feature is mapped to the training parameter result via the regression prediction layer, and the number of prediction features is the number of the training parameter results; wherein the number of prediction features is less than the second number of features.
5. The method for generating digital population animation according to claim 4, characterized in that: The regression prediction layer includes a compression activation function; the second training feature is mapped to a training parameter result based on a preset number of prediction features, including: Based on the number of predicted features, mapping the second training features to initial training parameters; The compression activation function is used to compress the initial training parameters into a preset weight range to obtain the training parameter results, and the initial training parameters correspond to the training parameter results one by one.
6. The method for generating digital population animation according to claim 3, characterized in that: The step of performing error optimization processing on the initial prediction model based on the training parameter result and the vertex deformation parameter and generating the deformation parameter prediction model comprises: Calculating a loss function result based on the vertex deformation parameters and the parameter prediction values; The gradient of the loss function result is back-transferred to the initial prediction model for iterative training to generate the deformation parameter prediction model.
7. The method for generating digital population animation according to claim 1, characterized in that: The step of selecting key point features from the vertex motion data comprises: Based on a preset key point quantity threshold, the key point features are selected from the vertex motion data, and the number of the key point features is less than the key point quantity threshold.
8. A digital population animation generation device, characterized in that: include: A data acquisition module, used for acquiring audio data to be recognized, performing vertex motion calculation on the audio data to be recognized, and obtaining vertex motion data; A parameter prediction module, used for selecting key point features from the vertex motion data, and inputting the key point features into the trained deformation parameter prediction model to obtain deformation parameter prediction values corresponding to the key point features; The animation generation module is used to generate the digital population animation based on the predicted value of the deformation parameter.
9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the method for generating a digital population animation according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method for generating a digital population animation according to any one of claims 1 to 7 when running.
Citation Information
Patent Citations
Three-dimensional face animation generation method and device based on audio driving and medium
CN116309988A
Voice-driven face model processing method and device and electronic equipment
CN117853620A
Shape reconstruction and editing using anatomically constrained implicit shape models
US20250037375A1