Training method of digital human generation model, and digital human video generation method and system
By using the first model to learn images in the digital life generation model and the second model fits the pictures output by the first model, the problem of excessive demand for computing resources is solved, and the generation of digital human videos on low-computing devices is realized, and the application scenarios and audience are expanded.
Patent Information
- Application Number
- CN202311867891.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-12-29
AI Technical Summary
The existing digital human synthesis model has too much demand for device computing resources, resulting in limited application scenarios, and users need to upload high-quality videos, forming a threshold for use and limited audience groups.
Two preset models are adopted, where the first model is responsible for learning the image reference training data, and the second model only needs to fit the picture output by the first model. By adjusting the parameters to reduce the model parameters and calculation amount, iterative learning to convergence and generate digital human videos.
It reduces the computing needs of digital life-generation models, enables them to run on low-computing devices, enriches application scenarios, lowers the threshold for use, and expands the audience.
Smart Images

Figure CN120279122A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of computers and image processing, and particularly relates to a method for training a digital human generation model, a method for generating a digital human video, and a system therefor. Background Art
[0002] With the increasing demand for digital experiences, the trend of using digital characters (such as virtual assistants) is also accelerating. Digital humans have always been a prominent topic. A digital human can be a virtual character, a humanoid robot, or any digital entity with voice, vision, touch, and locomotion capabilities, and is widely used in multiple fields such as healthcare, financial services, industrial automation, and customer service.
[0003] Digital human synthesis mostly uses a digital human synthesis model, which can synthesize a digital human video according to the voice input by the user. However, during the digital human synthesis process, users need to upload high-quality videos, which has a certain usage threshold and the audience group is too limited. At the same time, due to the large number of parameters and computational amount of the digital human synthesis model, the demand for the computing resources of the device is very large, resulting in the digital human not being widely applicable in various scenarios. Summary of the Invention
[0004] The present application provides a method for training a digital human generation model, a method for generating a digital human video, and a system therefor, which can solve the problem that in the existing digital human synthesis process, the demand for the computing resources of the device is too large and it cannot be widely applicable in various scenarios.
[0005] In a first aspect, the present application provides a method for training a digital human generation model, the method comprising:
[0006] Obtaining reference training data, the reference training data including reference pictures and reference audio;
[0007] Inputting the reference training data into a preset first model and a second model respectively, to instruct the first model to generate a first fitting picture according to the reference training data, and to instruct the second model to generate a second fitting picture according to the reference training data; wherein, the parameter value of the second model is less than the parameter value of the first model;
[0008] Judging a first true value corresponding to the first model according to the similarity between the first fitting picture and the reference picture, and training the first model according to the first true value;
[0009] Judging a second true value corresponding to the second model according to the similarity between the first fitting picture and the second fitting picture, and training the second model according to the second true value;
[0010] Train the first model and the second model until convergence, and output the trained second model as the target model; wherein, the target model is used to generate a digital human video according to the target audio input by the user.
[0011] In this application, the first model and the second model respectively generate a first fitted picture and a second fitted picture according to the reference training data, adjust the parameters of the first model based on the similarity difference between the first fitted picture and the reference picture, and determine the output result difference between the first model and the second model based on the similarity between the first fitted picture and the second fitted picture, so as to adjust the parameters of the second model; the first model and the second model are trained iteratively at the same time, and the first model and the second model are trained until convergence; in this way, the second model does not need to directly learn the real images in the reference training data, but only needs to fit the first fitted picture output by the first model, so as to use fewer model parameters and computational amounts to achieve an output effect close to that of the first model, enabling the second model to be applicable to low-computing-power devices, thus enriching the application scenarios of digital humans.
[0012] In some possible implementation manners, determining a second true value corresponding to the second model according to the similarity between the first fitted picture and the second fitted picture, and training the second model according to the second true value includes:
[0013] Obtain the first fitted picture and the second fitted picture;
[0014] Calculate a second true value corresponding to the second model according to the similarity between the first fitted picture and the second fitted picture; the second true value is used to represent the output result difference between the first model and the second model;
[0015] If the second true value is less than a second preset threshold, adjust the parameter value of the second model according to the similarity between the first fitted picture and the second fitted picture;
[0016] Train the second model with the adjusted parameter value until convergence, and input the next group of the reference training data into the second model.
[0017] In some possible implementation manners, calculating a second true value corresponding to the second model according to the similarity between the first fitted picture and the second fitted picture includes:
[0018] Calculate the structural similarity and the perceptual similarity according to the first fitted picture and the second fitted picture respectively;
[0019] Obtain a first weight value through the structural similarity and a first coefficient, and obtain a second weight value through the perceptual similarity and a second coefficient;
[0020] Obtain a second true value based on the first weight value and the second weight value.
[0021] In some possible implementation manners, the method further includes:
[0022] If the second true value is equal to or greater than a preset threshold, train the second model until convergence, and output the second model as the target model.
[0023] In some possible implementation manners, determining the first true value corresponding to the first model according to the similarity between the first fitted picture and the reference picture includes:
[0024] Calculate a first loss value based on the similarity between the first fitted picture and the reference picture;
[0025] Calculate a second loss value based on the difference between the first fitted picture and the reference picture;
[0026] Obtain the first true value based on the first loss value and the second loss value; the first true value is used to characterize the difference between the first fitted picture output by the first model and the reference picture.
[0027] In some possible implementation manners, after obtaining the first true value based on the first loss value and the second loss value, further includes:
[0028] When training the first model, if the first true value is less than a first preset threshold, adjust the parameter value of the first model according to the similarity between the first fitted picture and the reference picture;
[0029] Train the first model with the adjusted parameter value until convergence, and input the next set of the reference training data into the first model.
[0030] In a second aspect, the present application provides a method for generating a digital human video, including:
[0031] Obtain a picture input by a user and target audio;
[0032] Synthesize a template video based on the picture, and preprocess the template video to obtain a target picture; wherein, the target picture is a face picture with the mouth region removed;
[0033] Import the target picture and the target audio into a target model to generate a predicted face picture;
[0034] Stitch the predicted face picture and the target audio to obtain a digital human video.
[0035] After the target model in this application obtains a picture, it outputs a template video, and then performs some preprocessing operations on the template video. First, the video is converted into pictures, then face detection operations are performed, and then the detected faces are cropped and saved. After that, the voice input of the user is obtained, and then the voice and the face pictures are input into the trained target model. The target model outputs the face corresponding to the voice according to the voice and the image, and the output face image is stitched back into the original video to obtain a digital human video.
[0036] In a third aspect, this application provides a digital human video generation system, and the system includes:
[0037] An acquisition module, configured to acquire reference training data, where the reference training data includes reference pictures and reference audio;
[0038] A training module, which inputs the reference training data into a preset first model and a second model respectively, to instruct the first model to generate a first fitting picture according to the reference training data, and to instruct the second model to generate a second fitting picture according to the reference training data; wherein, the parameter value of the second model is smaller than the parameter value of the first model;
[0039] Judge the first true value corresponding to the first model according to the similarity between the first fitting picture and the reference picture, and train the first model according to the first true value;
[0040] Judge the second true value corresponding to the second model according to the similarity between the first fitting picture and the second fitting picture, and train the second model according to the second true value;
[0041] Train the first model and the second model until convergence, and output the trained second model as the target model; wherein, the target model is used to generate a digital human video according to the target audio input by the user.
[0042] In some possible implementation manners, the system further includes:
[0043] A generation module, configured to synthesize a template video based on the picture input by the user, and perform preprocessing on the template video to obtain a target picture; wherein, the target picture is a face picture with the mouth area removed;
[0044] Import the target picture and the target audio input by the user into the target model to generate a predicted face picture;
[0045] Stitch the predicted face picture and the target audio to obtain a digital human video.
[0046] The present application also provides a digital human video generation system, which includes an acquisition module, a training module, and a generation module. The acquisition module is used to acquire reference training data, as well as pictures and target audio input by the user. The training module is used to train a preset second model to obtain a target model with a relatively small number of parameters and capable of running on low-computing-power devices. The generation module is used to control the target model to synthesize and output a digital human video according to the pictures and target audio input by the user.
[0047] Fourthly, the present application also provides an electronic device, including:
[0048] A processor;
[0049] A memory for storing executable instructions of the processor;
[0050] Wherein, the processor is configured to execute the training method of the digital human generation model described in the first aspect and / or the digital human video generation method described in the second aspect by executing the executable instructions.
[0051] As can be seen from the above, the present application provides a training method for a digital human generation model, a digital human video generation method and a system. The training method includes: acquiring reference training data; respectively inputting the reference training data into a preset first model and a second model to instruct the first model to generate a first fitting picture according to the reference training data, and instructing the second model to generate a second fitting picture according to the reference training data; judging a first true value corresponding to the first model according to the similarity between the first fitting picture and the reference picture, and training the first model according to the first true value; judging a second true value corresponding to the second model according to the similarity between the first fitting picture and the second fitting picture, and training the second model according to the second true value; training the first model and the second model until convergence, and outputting the trained second model as the target model. The second model does not need to directly learn the real images in the reference training data, but only needs to fit the first fitting picture output by the first model, so as to use fewer model parameters and calculations to achieve an output effect close to that of the first model, enabling the second model to be applicable to low-computing-power devices, thus enriching the application scenarios of digital humans. Description of the Drawings
[0052] Figure 1 It is a schematic diagram of a training method for a digital human generation model provided by some embodiments of the present application;
[0053] Figure 2 It is a schematic diagram of the training process of the first model provided by some embodiments of the present application;
[0054] Figure 3 It is a schematic diagram of judging the first true value corresponding to the first model provided by some embodiments of the present application;
[0055] Figure 4 Schematic diagram of the second model training process provided by some embodiments of the present application;
[0056] Figure 5 Schematic diagram of the interaction process between the first model and the second model provided by some embodiments of the present application;
[0057] Figure 6 Schematic diagram of a digital human video generation system provided by some embodiments of the present application Figure 1 ;
[0058] Figure 7 Schematic diagram of a digital human video generation system provided by some embodiments of the present application Figure 2 . Detailed implementation manners
[0059] For the convenience of describing the technical solutions of the application, some concepts related to the present application will be described first below.
[0060] With the increasing demand for digital experiences, the trend of using digital characters (such as virtual assistants) is also accelerating. Digital humans have always been a prominent topic. A digital human can be a virtual character, a humanoid robot, or any digital entity with voice, vision, touch, and mobility capabilities, and is widely used in multiple fields such as healthcare, financial services, industrial automation, and customer service.
[0061] Digital human synthesis mostly uses digital human synthesis models, which can synthesize digital human videos according to the speech input by users. The current digital human synthesis model only supports training through videos. Therefore, when synthesizing digital human videos, it is necessary for the corresponding user to record a high-quality video data; because the training effect of the digital human is directly related to the quality of the captured video, the captured video has requirements such as clear shooting, no background noise, and no occlusion, which forms a relatively high usage threshold and is not conducive to the popularization and use of digital human technology. Therefore, in the process of digital human synthesis, under the requirement of users to upload high-quality videos, there is a certain usage threshold, and the audience group is too limited.
[0062] At the same time, the current digital human synthesis model has a large number of parameters and a large amount of calculations. The training and use of the model both require high-end graphics cards such as A100 and 3090ti. Whether it is the training or use of the model, a large amount of computing power cost is required. The high demand for computing resources greatly restricts the application scenarios of digital humans, resulting in their use only in the computer terminal or in an environment with high-speed network, and causing digital humans not to be widely used in various scenarios.
[0063] Therefore, the present application provides a training method for a digital human generation model, a digital human video generation method and system. In the present application, there are two preset models, namely a first model and a second model. The first model is responsible for learning real images in the reference training data. As for the subsequent second model for generating digital human videos, it does not need to directly learn real images in the reference training data, but only needs to fit the first fitting image output by the first model. By comparing the similarity difference between the first fitting image output by the first model and the real image, and comparing the similarity difference between the second fitting image output by the second model and the first fitting image output by the first model, the parameters of the first model and the second model are adjusted respectively, and the first model and the second model are trained iteratively at the same time. The first model and the second model are trained until convergence, and the trained second model is used as the target model, so as to use fewer model parameters and computational amounts to make the second model achieve an output effect close to that of the first model, so that the second model can be applied to low-computing power devices, thereby enriching the application scenarios of digital humans.
[0064] See Figure 1 , the present application provides a training method for a digital human generation model, and the method includes:
[0065] S100: Obtain reference training data, where the reference training data includes reference pictures and reference audio.
[0066] Among them, in the learning process of the digital human model, there is a clear corresponding relationship between the mouth shape of humans and the semantic information in the voice, and it has nothing to do with factors such as loudness, timbre, and noise in the voice. In addition, the appearance information of people needs to be considered. The skin color, tooth shape, lip shape, etc. of different people have personal characteristics. Therefore, the model to be trained requires two input information, speech and face. The speech information is used to predict the output mouth shape information, and the face picture is used to control the output appearance information. The two features jointly determine the output result.
[0067] In some embodiments of the present application, the reference training data is used to enable the preset first model and second model to learn the corresponding relationship between speech and mouth shape, and keep the appearance information of the input image unchanged to obtain accurate effects and good stability. Therefore, to ensure better training effects of the subsequent first model and second model, the selected reference training data needs to be clear, expose the mouth area, and at the same time, the mouth shape should correspond accurately to the speech, without deviation, without too much noise, without the voice of other people talking, and without covering the mouth.
[0068] The reference training data is obtained by self-recording and network downloading, and data cleaning work is carried out before training to delete problematic segments to ensure the cleanliness of the training data.
[0069] After obtaining the reference training data, data preprocessing work needs to be carried out first. The data preprocessing work of this application mainly includes audio feature extraction and face cropping work.
[0070] Among them, for audio feature extraction, an audio feature extraction model is used. The ProsodyEncoder module can be adopted. The ProsodyEncoder module can extract audio features through a word-level vector quantization bottleneck.
[0071] The data preprocessing work also includes cropping the video. First, convert the video into an image sequence, then use a face detection program to detect the face position, then crop the face according to the position information, and use it as the real picture for subsequent training. Then, mask the mouth area of the face picture and use it as the input picture, that is, the reference picture, for subsequent training. The model repairs the masked area based on the masked picture and voice information to obtain a complete face image.
[0072] S200: Input the reference training data into a preset first model and a second model respectively, to instruct the first model to generate a first fitted picture according to the reference training data, and to instruct the second model to generate a second fitted picture according to the reference training data; wherein, the parameter value of the second model is less than the parameter value of the first model.
[0073] In this application, there are two preset models, namely the first model and the second model. The audio and image encoding networks of the first model adopt a large number of Resnet structures. While this network structure outputs convolutional results, it superimposes the input content on the output, which can greatly alleviate the problem of gradient disappearance and enhance the learning ability of the network. However, it also greatly increases the number of model parameters and the amount of calculation. The second model adopts a common convolutional structure. The number of channels in the middle network layer of the second model is 0.25 - 0.4 times that of the first model. The parameter value of the second model is less than the parameter value of the first model. In some embodiments, the running speed of the second model is 19 times that of the first model, meeting the requirements for running on low-computing-power devices such as mobile phones.
[0074] Input the reference training data into the first model and the second model respectively, and obtain the first fitted picture generated by the first model according to the reference training data and the second fitted picture generated by the second model according to the reference training data respectively.
[0075] S300: Determine the first true value corresponding to the first model according to the similarity between the first fitted picture and the reference picture, and train the first model according to the first true value.
[0076] After obtaining the first fitted image, compare the similarity between the first fitted image and the reference image to determine the difference between the first fitted image output by the first model and the reference image, so as to adjust the parameters of the first model, facilitating the subsequent training and learning of the first model based on the current reference image.
[0077] As Figure 2 shown, in some embodiments, determining the first true value corresponding to the first model according to the similarity between the first fitted image and the reference image includes:
[0078] S301: Calculate the first loss value based on the similarity between the first fitted image and the reference image;
[0079] S302: Calculate the second loss value based on the difference between the first fitted image and the reference image;
[0080] S303: Obtain the first true value based on the first loss value and the second loss value; the first true value is used to characterize the difference between the first fitted image output by the first model and the reference image.
[0081] During the training and learning process of the first model, there are two loss functions, namely the first loss function (concat) and the second loss function. Determine the first true value corresponding to the first model based on the first loss function and the second loss function. The first true value is used to characterize the difference between the first fitted image output by the first model and the reference image.
[0082] As Figure 3 shown, in some embodiments, the first loss function can be the reconstruction loss L Recon , and the second loss function can be the GAN loss L GAN , where the calculation formulas are as follows:
[0083] L GAN (G T ) = E x,y [logD(x,y)] + E x [log(1 - D(x,G T (x)))] (1)
[0084] L Recon (G T ) = E x,y [||y - G T (x)||1] (2)
[0085]
[0086] where G TLet \(G\) represent the generator in the first model, \(D\) represent the discriminator in the first model, \(y\) be the real face image, \(x\) be the input audio features and image features, i.e., the reference training data. T \(G(x)\) is the first fitted picture; the generator is used to learn the correspondence between semantics and lip shapes, as well as the face fitting ability, and the discriminator is used to evaluate the quality of the pictures synthesized by the generator.
[0087] \(L\) GAN (G T ) is the GAN loss, and its calculation requires sending the reference picture and the first fitted picture into the discriminator for encoding, and then obtaining a label to calculate the loss. Recon (G T ) is the reconstruction loss, i.e., the L1 loss. Subtract the reference picture from the first fitted picture to calculate the absolute difference, which is used to evaluate the authenticity of the first fitted picture. is the overall loss, which contains \(L\) GAN and \(L\) Recon two loss functions. The purpose of the generator during the training process is to reduce \(L\) GAN loss, that is, to make the first fitted picture more similar to the reference picture. The purpose of the discriminator is to maximize \(L\) GAN , aiming to distinguish the reference picture and the first fitted picture as much as possible. The generator is also a process of continuously reducing \(L\) Recon during the training process.
[0088] In some embodiments, after obtaining the first true value based on the first loss value and the second loss value, it further includes:
[0089] S304: When training the first model, if the first true value is less than the first preset threshold, adjust the parameter values of the first model according to the similarity between the first fitted picture and the reference picture;
[0090] S305: Train the first model with the adjusted parameter values until convergence, and input the next set of the reference training data into the first model.
[0091] After obtaining the first true value, before performing iterative learning training on the first model, it is necessary to first compare the first true value with the first preset threshold to determine whether to adjust the parameters of the first model. If the first true value is less than the first preset threshold, adjust the parameter values of the first model according to the similarity between the first fitted picture and the reference picture, train the first model with the adjusted parameter values until convergence, and input the next set of the reference training data into the first model. If the first true value is not less than the first preset threshold, use the current first model to train until convergence, and input the next set of the reference training data into the first model.
[0092] S400: Determine the second true value corresponding to the second model according to the similarity between the first fitted image and the second fitted image, and train the second model according to the second true value.
[0093] As Figure 4 and Figure 5 shown, after obtaining the first fitted image and the second fitted image, compare the similarity between the first fitted image and the second fitted image to determine the difference between the second fitted image output by the second model and the first fitted image output by the first model, so as to adjust the parameters of the second model, facilitating subsequent training and learning of the second model based on the current first fitted image.
[0094] In this application, the second model no longer fits with the real image, i.e., the reference image, but fits with the first fitted image output by the first model. It only needs to fit a structure similar to the output of the first model, greatly reducing the learning difficulty, making the operation speed of the second model faster, and meeting the requirements for running on low-computing-power devices such as mobile phones.
[0095] In some embodiments, determining the second true value corresponding to the second model according to the similarity between the first fitted image and the second fitted image, and training the second model according to the second true value includes:
[0096] S401: Obtain the first fitted image and the second fitted image;
[0097] S402: Calculate the second true value corresponding to the second model according to the similarity between the first fitted image and the second fitted image; the second true value is used to characterize the difference in the output results between the first model and the second model.
[0098] During the training and learning process of the second model, it is necessary to calculate the structural similarity and perceptual similarity between the second model and the first model to evaluate the difference between the first fitted image output by the first model and the second fitted image output by the second model. The second true value is used to characterize the difference in the output results between the first model and the second model.
[0099] In some embodiments, calculating the second true value corresponding to the second model according to the similarity between the first fitted image and the second fitted image includes:
[0100] S4021: Calculate the structural similarity and perceptual similarity respectively according to the first fitted image and the second fitted image;
[0101] S4022: Obtain the first weight value through the structural similarity and the first coefficient, and obtain the second weight value through the perceptual similarity and the second coefficient;
[0102] S4023: Obtain a second true value based on the first weight value and the second weight value.
[0103] In some embodiments, the structural similarity, perceptual similarity, and the second true value can be calculated through the following loss function:
[0104]
[0105]
[0106] L KD (p t , p s ) = λ SSIM L SSIM + λ feature L feature + λ e L e (6)
[0107] Where p t , p s are the first fitted image output by the first model and the second fitted image output by the second model. L SSIM The structural similarity loss and the perceptual loss are used to evaluate the difference between p t , p s . The perceptual loss is the feature reconstruction loss L feature and the style reconstruction loss L e . μ t μ s represents the average value of brightness, is the standard deviation of contrast, σ ts is the structural similarity covariance, and C2C1 is a fixed constant. is the pre-trained VGG network, which is used here to evaluate the similarity of the features of p t , p s . L e is also evaluated using the middle layer of the VGG network. λ SSIM , λ feature , λ e are the coefficients of various losses, which are used to define the weights of various losses.
[0108] After obtaining the second true value and before performing iterative learning and training on the second model, it is necessary to first compare the second true value with the first preset threshold to determine whether to adjust the parameters of the second model.
[0109] S403: If the second true value is less than the second preset threshold, adjust the parameter value of the second model according to the similarity between the first fitted image and the second fitted image;
[0110] S404: Train the second model with the adjusted parameter values until convergence, and input the next set of the reference training data into the second model.
[0111] S405: If the second true value is equal to or greater than the preset threshold, train the second model until convergence, and output the second model as the target model.
[0112] If the second true value is less than the second preset threshold, adjust the parameter values of the second model according to the similarity between the first fitted picture and the second fitted picture, train the second model with the adjusted parameter values until convergence, and input the next set of the reference training data into the second model. If the second true value is equal to or greater than the preset threshold, train the second model until convergence, and output the second model as the target model.
[0113] It should be noted that the first model and the second model in this application are trained simultaneously, and the two models are iterated cyclically without a sequential order. When training, the first model and the second model are trained together. The second model learns based on the first model, and the first model is used to assist the second model in learning, enabling the second model to obtain similar fitting capabilities with fewer parameter quantities and computational amounts. After training, it is deployed to low-computing-power devices. If the first model is trained first and then the second model is trained based on the first model, to avoid the problem of model collapse and gradient disappearance due to a too large difference in model scales, the second model needs to maintain a dynamic balance with the first model, and the parameter quantity of the first model is similar to that of the second model, which will result in the second model being unable to be applied to low-computing-power devices; if the first model is a model that has been trained, it cannot guide the learning of the second model in the initial stage, resulting in problems easily occurring in the starting stage. In addition, it is also prone to overfitting.
[0114] In this application, the first model and the second model are trained and learned together. The first model can determine the fitting situation of the second model in the initial stage of training, and they are synchronously iterated and learned. The two models are jointly learned and closely combined to obtain a good fitting effect. The second model can fit the parameters more flexibly. Therefore, there is no problem that the first model and the second model must be balanced and the parameter quantities cannot differ too much. In this way, the second model can adopt a smaller parameter quantity so that the trained second model can be deployed on low-computing-power devices, enriching the application scenarios of digital humans.
[0115] S500: Train the first model and the second model until convergence, and output the trained second model as the target model; wherein, the target model is used to generate a digital human video according to the target audio input by the user.
[0116] In this application, the first model and the second model respectively generate a first fitted picture and a second fitted picture based on reference training data, adjust the parameters of the first model based on the similarity difference between the first fitted picture and the reference picture, and determine the output result difference between the first model and the second model based on the first fitted picture and the second fitted picture, so as to adjust the parameters of the second model; the first model and the second model are trained iteratively at the same time, and the first model and the second model are trained until convergence; in this way, the second model does not need to directly learn the real images in the reference training data, but only needs to fit the first fitted picture output by the first model, so as to use fewer model parameters and calculation amounts to achieve an output effect close to that of the first model, so that the second model can be applied to low-computing-power devices, thus enriching the application scenarios of digital humans.
[0117] In some embodiments, this application provides a digital human video generation method, including:
[0118] S600: Obtain a picture and target audio input by a user;
[0119] S601: Synthesize a template video based on the picture, and preprocess the template video to obtain a target picture; wherein, the target picture is a face picture with the mouth area removed;
[0120] S602: Import the target picture and the target audio into a target model to generate a predicted face picture;
[0121] S603: Stitch the predicted face picture and the target audio to obtain a digital human video.
[0122] After outputting the trained second model as the target model, the user only needs to upload a photo with a clear face showing the front to synthesize a digital human video with accurate lip-sync and realistic effects. The following details the digital human work process. After the user uploads a picture, a template video is synthesized based on the picture uploaded by the user. Thereafter, the template video is preprocessed. The processing process is similar to the training process to obtain a face crop picture and a mask picture. In addition, the user also needs to provide audio data, which can be directly recorded with a mobile phone or existing audio data can be uploaded. Then, the face mask picture and the audio data are sent into the target model to obtain a digital human video.
[0123] After the target model in this application obtains a picture, it outputs a template video, and then performs some preprocessing operations on the template video. First, the video is converted into pictures, and then face detection operations are performed. Then, the detected face is cropped and saved. Thereafter, the voice input of the user is obtained, and then the voice and the face picture are input into the trained target model. The target model outputs the face corresponding to the voice according to the voice and the image, and the output face image is stitched back into the original video to obtain a digital human video.
[0124] This application can synthesize digital human videos with accurate mouth shapes and natural movements. It only requires one picture to synthesize a video corresponding to the audio. Because the target model has a low number of parameters and a fast running speed, it can run offline on low-computing devices such as mobile phones and can be widely applied to scenarios such as news broadcasts, customer service, and AI assistants, greatly improving the interaction effect.
[0125] As Figure 6 shown, in some embodiments, this application provides a digital human video generation system, and the system includes:
[0126] An acquisition module, configured to acquire reference training data, where the reference training data includes reference pictures and reference audio;
[0127] A training module, which inputs the reference training data into a preset first model and a second model respectively, to instruct the first model to generate a first fitting picture according to the reference training data, and to instruct the second model to generate a second fitting picture according to the reference training data; wherein, the parameter value of the second model is less than the parameter value of the first model;
[0128] Determine the first true value corresponding to the first model according to the similarity between the first fitting picture and the reference picture, and train the first model according to the first true value;
[0129] Determine the second true value corresponding to the second model according to the similarity between the first fitting picture and the second fitting picture, and train the second model according to the second true value;
[0130] Train the first model and the second model until convergence, and output the trained second model as the target model; wherein, the target model is used to generate a digital human video according to the target audio input by the user.
[0131] The reference training data in the acquisition module includes high-definition videos self-recorded by the database and high-definition data downloaded from the Internet. Data cleaning work was carried out before training. The video content to be cleaned includes misaligned mouth shapes, multiple people speaking, large environmental noises, no human figures, no human faces, serious occlusions, etc. Clean data can help the model converge faster and obtain better video synthesis effects. It should be noted that the reference training data here refers to general training video data, not the videos of specific models that need to be synthesized into digital humans.
[0132] This application provides a digital human video generation system, which includes an acquisition module and a training module. The acquisition module is used to acquire reference training data. The training module is used to train a preset second model to obtain a target model with a lower number of parameters and capable of running on low-computing-power devices. Users can generate digital human videos based on the target audio input to the target model.
[0133] As Figure 7 shown, in some embodiments, the digital human video generation system provided by this application further includes:
[0134] A generation module, which is used to synthesize a template video based on the user input picture, and preprocess the template video to obtain a target picture. Wherein, the target picture is a face picture with the mouth area removed;
[0135] Import the target picture and the target audio input by the user into the target model to generate a predicted face picture;
[0136] Stitch the predicted face picture and the target audio to obtain a digital human video.
[0137] Among them, an image driving unit and a lip-sync generation unit are configured in the generation module. The generation module controls the image driving unit to drive the picture synthesis template video, and controls the lip-sync generation unit to drive the template video according to the input voice to synthesize a digital human video matching the audio.
[0138] The image driving unit is used to convert the input image into a video with head movements and expressions. This unit is configured with an image driving model, which is trained with a large amount of unlabeled video data and can use a video to drive a picture, transfer the head movements and expressions of the video to the image, and synthesize a video with actions and expressions. This video is used as the input of the subsequent module and as the template video for subsequent digital human driving.
[0139] Among them, the generation module further includes a digital human driving unit, which can use the input audio data to drive the template video synthesized by the image driving unit to synthesize a video matching the voice. By training with the recorded video data, the corresponding relationship between the voice and the lip-sync can be learned.
[0140] This application also provides a digital human video generation system, which includes an acquisition module, a training module, and a generation module. The acquisition module is used to acquire reference training data, as well as the pictures and target audio input by the user. The training module is used to train a preset second model to obtain a target model with a lower number of parameters and capable of running on low-computing-power devices. The generation module is used to control the target model to synthesize and output a digital human video according to the pictures and target audio input by the user.
[0141] The digital human video generation system proposed in this application can train a digital human model and generate digital human videos based on only one picture. Moreover, the digital human model only needs to run on the server during the training process, and can be performed on low-computing-power devices such as mobile phones during the process of generating digital human videos, without the support of high-computing-power graphics cards. Based on the digital human video generation system, it is possible to generate digital human videos with accurate lip movements and natural actions according to audio, and this technology can run on low-computing-power devices and can run offline, greatly reducing the usage threshold and promoting the implementation of digital humans in more application scenarios.
[0142] In some embodiments, the present application further provides an electronic device, including:
[0143] a processor;
[0144] a memory for storing executable instructions of the processor;
[0145] wherein, the processor is configured to execute the above-mentioned training method of the digital human generation model and the digital human video generation method by executing the executable instructions.
[0146] As can be seen from the above embodiments, the present application provides a training method of a digital human generation model, a digital human video generation method and a system. The training method includes: obtaining reference training data; respectively inputting the reference training data into a preset first model and a second model to instruct the first model to generate a first fitting picture according to the reference training data, and instructing the second model to generate a second fitting picture according to the reference training data; judging the first true value corresponding to the first model according to the similarity between the first fitting picture and the reference picture, and training the first model according to the first true value; judging the second true value corresponding to the second model according to the similarity between the first fitting picture and the second fitting picture, and training the second model according to the second true value; training the first model and the second model until convergence, and outputting the trained second model as the target model. The second model does not need to directly learn the real images in the reference training data, and only needs to fit the first fitting picture output by the first model, so as to use fewer model parameters and computational amounts to achieve an output effect close to that of the first model, so that the second model can be applied to low-computing-power devices, thus enriching the application scenarios of digital humans.
Claims
1. A training method for a digital life generation model, characterized in that, The method includes: Obtaining reference training data, where the reference training data includes reference pictures and reference audio; Inputting the reference training data into a preset first model and a second model respectively, to instruct the first model to generate a first fitted picture according to the reference training data, and to instruct the second model to generate a second fitted picture according to the reference training data; wherein, the parameter value of the second model is less than the parameter value of the first model; Judging the first true value corresponding to the first model according to the similarity between the first fitted picture and the reference picture, and training the first model according to the first true value; Judging the second true value corresponding to the second model according to the similarity between the first fitted picture and the second fitted picture, and training the second model according to the second true value; Training the first model and the second model until convergence, and outputting the trained second model as the target model; wherein, the target model is used to generate a digital human video according to the target audio input by the user.
2. The training method of the digital life generation model according to claim 1, characterized in that Judging the second true value corresponding to the second model according to the similarity between the first fitted picture and the second fitted picture, and training the second model according to the second true value includes: Obtaining the first fitted picture and the second fitted picture; Calculating the second true value corresponding to the second model according to the similarity between the first fitted picture and the second fitted picture; the second true value is used to characterize the difference between the output results of the first model and the second model; If the second true value is less than a second preset threshold, adjusting the parameter value of the second model according to the similarity between the first fitted picture and the second fitted picture; Training the second model with the adjusted parameter value until convergence, and inputting the next set of the reference training data into the second model.
3. The training method of the digital life generation model according to claim 2, wherein Calculating the second true value corresponding to the second model according to the similarity between the first fitted picture and the second fitted picture includes: Calculating the structural similarity and the perceptual similarity according to the first fitted picture and the second fitted picture respectively; Obtaining a first weight value through the structural similarity and a first coefficient, and obtaining a second weight value through the perceptual similarity and a second coefficient; Obtaining the second true value based on the first weight value and the second weight value.
4. The training method of the digital life generation model according to claim 2, characterized in that, The method further includes: If the second true value is equal to or greater than the preset threshold, training the second model until convergence, and outputting the second model as the target model.
5. The training method of the digital life generation model according to claim 1, characterized in that, Judging the first true value corresponding to the first model according to the similarity between the first fitted picture and the reference picture includes: Calculating a first loss value based on the similarity between the first fitted picture and the reference picture; Calculating a second loss value based on the difference between the first fitted picture and the reference picture; Obtaining the first true value based on the first loss value and the second loss value; the first true value is used to characterize the difference between the first fitted picture output by the first model and the reference picture.
6. The training method of the digital life generation model according to claim 5, wherein After obtaining the first true value based on the first loss value and the second loss value, it further includes: When training the first model, if the first true value is less than the first preset threshold, adjust the parameter value of the first model according to the similarity between the first fitted picture and the reference picture; Train the first model with the adjusted parameter value until convergence, and input the next set of the reference training data into the first model.
7. A method for generating digital human videos, characterized in that, Including: Obtain a picture and a target audio input by a user; Based on the picture, synthesize a template video, and preprocess the template video to obtain a target picture; wherein, the target picture is a face picture with the mouth region removed; Import the target picture and the target audio into a target model to generate a predicted face picture; Stitch the predicted face picture and the target audio to obtain a digital human video.
8. A digital human video generation system, characterized in that The system includes: An acquisition module, configured to acquire reference training data, where the reference training data includes a reference picture and a reference audio; A training module, configured to input the reference training data into a preset first model and a second model respectively, to instruct the first model to generate a first fitted picture according to the reference training data, and to instruct the second model to generate a second fitted picture according to the reference training data; wherein, the parameter value of the second model is less than that of the first model; Discriminate the first true value corresponding to the first model according to the similarity between the first fitted picture and the reference picture, and train the first model according to the first true value; Discriminate the second true value corresponding to the second model according to the similarity between the first fitted picture and the second fitted picture, and train the second model according to the second true value; Train the first model and the second model until convergence, and output the trained second model as a target model; wherein, the target model is used to generate a digital human video according to a target audio input by a user.
9. The digital human video generation system according to claim 8, wherein The system further includes: A generation module, configured to synthesize a template video based on a picture input by a user, and preprocess the template video to obtain a target picture; wherein, the target picture is a face picture with the mouth region removed; Import the target picture and the target audio input by the user into the target model to generate a predicted face picture; Stitch the predicted face picture and the target audio to obtain a digital human video.
10. An electronic device, characterized in that, Including: A processor; A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the training method of the digital human generation model according to any one of claims 1-6, and / or execute the digital human video generation method according to claim 7 by executing the executable instructions.
Citation Information
Patent Citations
Digital human generation model, model training method and digital human generation method
CN114419702A
Visual dubbing using synthetic models
US11562597B1