Lightweight digital human generation model training method and digital human video generation method

By employing two model training methods—the first model learns from real images, and the second model fits the output of the first model—the problem of high computational resource requirements for digital human synthesis models is solved. This enables the generation of high-quality digital human videos on low-computing-power devices, expanding the application scenarios.

CN122244250APending Publication Date: 2026-06-19NANJING SILICON INTELLIGENCE TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING SILICON INTELLIGENCE TECH CO LTD
Filing Date
2023-12-29
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing digital human synthesis models have high computational resource requirements, which limits their application scenarios, makes it difficult for users to upload high-quality videos, and restricts their audience.

Method used

Two model training methods are used: the first model learns from real images, and the second model only fits the images output by the first model. By adjusting the parameters, the two models converge. The second model does not need to directly learn from real images and uses fewer parameters and less computation.

Benefits of technology

It enables the generation of high-quality digital human videos on low-computing-power devices, lowering the barrier to entry and enriching the application scenarios of digital humans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244250A_ABST
    Figure CN122244250A_ABST
Patent Text Reader

Abstract

This application provides a training method for a lightweight digital human generation model and a digital human video generation method. A first model and a second model generate a first fitted image and a second fitted image, respectively, based on reference training data. The parameters of the first model are adjusted based on the similarity difference between the first fitted image and the reference image. The output differences between the first and second models are determined based on the first and second fitted images, thereby adjusting the parameters of the second model. The first and second models are trained iteratively simultaneously until convergence. In this way, the second model does not need to directly learn from real images in the reference training data; it only needs to fit the first fitted image output by the first model. This achieves a near-first model output effect using fewer model parameters and less computation, making the second model applicable to low-computing devices and thus enriching the application scenarios of digital humans.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The original application has the application number 202311867891.6 and the original application date is December 29, 2023. The entire contents of the original application are incorporated herein by reference. Technical Field

[0002] This application relates to the fields of computer and image processing, and in particular to a training method for a lightweight digital human generation model and a method for digital human video generation. Background Technology

[0003] As people's demand for digital experiences grows, the trend of using digital characters (such as virtual assistants) is also accelerating. Digital humans have always been a fascinating topic. Digital humans can be virtual characters, humanoid robots, or any digital entity with voice, vision, touch, and mobility capabilities, and are widely used in many fields such as healthcare, financial services, industrial automation, and customer service.

[0004] Digital human synthesis often employs digital human synthesis models, which can synthesize digital human videos based on user-input speech. However, the process of digital human synthesis requires users to upload high-quality videos, presenting a certain barrier to entry and limiting the target audience. Furthermore, the large number of parameters and computational demands of digital human synthesis models place significant demands on device computing resources, preventing the widespread application of digital humans in various scenarios. Summary of the Invention

[0005] This application provides a training method for a lightweight digital human generation model and a digital human video generation method, which can solve the problem that the existing digital human synthesis process requires too much computing resources and cannot be widely applied in various scenarios.

[0006] In a first aspect, this application provides a training method for a lightweight digital human generation model, the method comprising: Acquire reference training data and preprocess the reference training data. The preprocessed reference training data includes reference images and reference audio. The preprocessing includes audio feature extraction and image sequence processing. The reference training data is input into a preset first model and a second model respectively, so as to instruct the first model to generate a first fitted image based on the reference training data, and to instruct the second model to generate a second fitted image based on the reference training data; wherein the parameter values ​​of the second model are smaller than the parameter values ​​of the first model; The first true value corresponding to the first model is determined based on the similarity between the first fitted image and the reference image, and the first model is trained based on the first true value. The second true value corresponding to the second model is determined based on the similarity between the first fitted image and the second fitted image, and the second model is trained based on the second true value. The first model and the second model are trained until convergence, and the trained second model is output as the target model; wherein, the target model is used to generate a lightweight digital human video based on the target audio input by the user.

[0007] In this application, the first model and the second model generate a first fitted image and a second fitted image respectively based on the reference training data. The parameters of the first model are adjusted based on the similarity difference between the first fitted image and the reference image. The output difference between the first model and the second model is determined based on the similarity between the first fitted image and the second fitted image, thereby adjusting the parameters of the second model. The first model and the second model are trained iteratively simultaneously until they converge. In this way, the second model does not need to directly learn the real images in the reference training data, but only needs to fit the first fitted image output by the first model. This allows it to achieve an output effect close to that of the first model with fewer model parameters and less computation, making the second model applicable to low-computing devices and thus enriching the application scenarios of digital humans.

[0008] In some possible implementations, the step of determining the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image, and training the second model based on the second true value, includes: The second model is processed based on the comparison between the second true value and the second preset threshold.

[0009] In some possible implementations, determining the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image, and training the second model based on the second true value includes: Obtain the first fitted image and the second fitted image; The second true value corresponding to the second model is calculated based on the similarity between the first fitted image and the second fitted image; the second true value is used to characterize the difference between the output results of the first model and the second model. If the second true value is less than the second preset threshold, the parameter values ​​of the second model are adjusted according to the similarity between the first fitted image and the second fitted image. The second model, with its parameter values ​​adjusted, is trained until convergence, and the next set of reference training data is input into the second model.

[0010] In some possible implementations, calculating the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image includes: Calculate structural similarity and perceptual similarity based on the first fitted image and the second fitted image, respectively. A first weight value is obtained by using the structural similarity and the first coefficient, and a second weight value is obtained by using the perceptual similarity and the second coefficient; The second true value is obtained based on the first weight value and the second weight value.

[0011] In some possible implementations, the method further includes: If the second true value is equal to or greater than a preset threshold, the second model is trained until convergence, and the second model is output as the target model.

[0012] In some possible implementations, determining the first true value corresponding to the first model based on the similarity between the first fitted image and the reference image includes: A first loss value is calculated based on the similarity between the first fitted image and the reference image; The second loss value is calculated based on the difference between the first fitted image and the reference image; The first true value is obtained based on the first loss value and the second loss value; the first true value is used to characterize the difference between the first fitted image output by the first model and the reference image.

[0013] In some possible implementations, after obtaining the first true value based on the first loss value and the second loss value, the method further includes: When training the first model, if the first true value is less than the first preset threshold, the parameter values ​​of the first model are adjusted according to the similarity between the first fitted image and the reference image. The first model, after its parameter values ​​have been adjusted, is trained until convergence, and the next set of reference training data is input into the first model.

[0014] Secondly, this application provides a lightweight digital human video generation method, including: Obtain the image and target audio input by the user; A template video is synthesized based on the image, and a target image is obtained by preprocessing the template video; wherein, the target image is a face image with the mouth area removed; The target image and the target audio are imported into the target model to generate a predicted face image; The predicted face image is spliced ​​with the target audio to obtain a lightweight digital human video.

[0015] In this application, after the target model obtains the image, it outputs a template video. The template video undergoes preprocessing: first, the video is converted into an image; then, face detection is performed, and the detected faces are cropped and saved. Next, the user's voice input is obtained, and the voice and face images are input into the trained target model. The target model outputs the face corresponding to the voice based on the voice and image, and the output face image is stitched back into the original video to obtain a lightweight digital human video.

[0016] Thirdly, this application provides a lightweight digital human video generation system, the system comprising: The acquisition module is used to acquire reference training data and preprocess the reference training data. The preprocessed reference training data includes reference images and reference audio. The preprocessing includes audio feature extraction and image sequence processing. The training module inputs the reference training data into a preset first model and a second model, respectively, to instruct the first model to generate a first fitted image based on the reference training data, and to instruct the second model to generate a second fitted image based on the reference training data; wherein the parameter values ​​of the second model are smaller than the parameter values ​​of the first model; The first true value corresponding to the first model is determined based on the similarity between the first fitted image and the reference image, and the first model is trained based on the first true value. The second true value corresponding to the second model is determined based on the similarity between the first fitted image and the second fitted image, and the second model is trained based on the second true value. The first model and the second model are trained until convergence, and the trained second model is output as the target model; wherein, the target model is used to generate a lightweight digital human video based on the target audio input by the user.

[0017] In some possible implementations, the system further includes: A generation module is used to synthesize a template video based on images input by the user, and to preprocess the template video to obtain a target image; wherein the target image is a face image with the mouth area removed; The target image and the target audio input by the user are imported into the target model to generate a predicted face image; The predicted face image is spliced ​​with the target audio to obtain a lightweight digital human video.

[0018] This application also provides a lightweight digital human video generation system, the system including an acquisition module, a training module, and a generation module; the acquisition module is used to acquire reference training data and user-input images and target audio; the training module is used to train a preset second model to obtain a target model with low parameter count and capable of running on low-computing-power devices; the generation module is used to control the target model to synthesize and output a lightweight digital human video based on user-input images and target audio.

[0019] Fourthly, this application also provides an electronic device, comprising: processor; Memory for storing the executable instructions of the processor; The processor is configured to execute the training method of the lightweight digital human generation model described in the first aspect and / or the lightweight digital human video generation method described in the second aspect by executing the executable instructions.

[0020] As described above, this application provides a training method for a lightweight digital human generation model and a digital human video generation method. The training method includes: acquiring reference training data; inputting the reference training data into a preset first model and a second model respectively, to instruct the first model to generate a first fitted image based on the reference training data, and to instruct the second model to generate a second fitted image based on the reference training data; determining a first real value corresponding to the first model based on the similarity between the first fitted image and the reference image, and training the first model based on the first real value; determining a second real value corresponding to the second model based on the similarity between the first fitted image and the second fitted image, and training the second model based on the second real value; training the first model and the second model until convergence, and outputting the trained second model as the target model. The second model does not need to directly learn the real images in the reference training data, but only needs to fit the first fitted image output by the first model, so as to achieve an output effect close to that of the first model with fewer model parameters and less computation, making the second model applicable to low-computing-power devices, thereby enriching the application scenarios of digital humans. Attached Figure Description

[0021] Figure 1 A schematic diagram illustrating a training method for a lightweight digital human generation model provided in some embodiments of this application; Figure 2 This is a schematic diagram of the first model training process provided in some embodiments of this application; Figure 3 A schematic diagram illustrating the determination of the first true value corresponding to the first model, provided in some embodiments of this application; Figure 4 This is a schematic diagram of the second model training process provided in some embodiments of this application; Figure 5 This is a schematic diagram illustrating the interaction process between the first model and the second model provided in some embodiments of this application; Figure 6 This application provides a schematic diagram of a lightweight digital human video generation system as part of its embodiments. Figure 1 ; Figure 7 This application provides a schematic diagram of a lightweight digital human video generation system as part of its embodiments. Figure 2 . Detailed Implementation

[0022] To facilitate the explanation of the technical solution of this application, some concepts involved in this application will be explained first below.

[0023] As people's demand for digital experiences grows, the trend of using digital characters (such as virtual assistants) is also accelerating. Digital humans have always been a fascinating topic. Digital humans can be virtual characters, humanoid robots, or any digital entity with voice, vision, touch, and mobility capabilities, and are widely used in many fields such as healthcare, financial services, industrial automation, and customer service.

[0024] Digital human synthesis primarily employs digital human synthesis models, which can synthesize digital human videos based on user-input speech. Currently, these models only support training via video, requiring users to record high-quality video data for the synthesis process. Since the training effect is directly related to the quality of the video, the video must be clear, free of background noise, and unobstructed. This creates a high barrier to entry, hindering the widespread adoption of digital human technology. Therefore, the requirement for users to upload high-quality videos in the digital human synthesis process presents a significant barrier to entry, limiting the target audience.

[0025] At the same time, the current digital human synthesis models have a large number of parameters and computational load. The training and use of the models require high-end graphics cards such as A100 and 3090ti. Both the training and use of the models require a lot of computing power. The high demand for computing resources greatly limits the application scenarios of digital humans, causing them to be used only on computers or in environments with high-speed networks, which prevents digital humans from being widely used in various scenarios.

[0026] Therefore, this application provides a lightweight digital human generation model training method and a digital human video generation method. This application includes two preset models: a first model and a second model. The first model is responsible for learning real images from the reference training data. The second model, which generates digital human videos, does not need to directly learn real images from the reference training data; it only needs to fit the first fitted image output by the first model. By comparing the similarity difference between the first fitted image output by the first model and the real image, and by comparing the similarity difference between the second fitted image output by the second model and the first fitted image output by the first model, the parameters of the first and second models are adjusted respectively. The first and second models are trained iteratively simultaneously until convergence. The trained second model is used as the target model, achieving output effects close to the first model with fewer model parameters and less computation. This allows the second model to be applied to low-computing devices, thereby enriching the application scenarios of digital humans.

[0027] See Figure 1 This application provides a training method for a lightweight digital human generation model, the method comprising: S100: Obtain reference training data and preprocess the reference training data.

[0028] In the digital human model learning process, there is a clear correspondence between human lip movements and semantic information in speech, independent of factors such as loudness, timbre, and noise. Furthermore, facial features must be considered, as different people have individual characteristics such as skin color, tooth shape, and lip shape. Therefore, the training model requires both speech and facial input. Speech information is used to predict the output lip movement, while facial images control the output facial features; both features jointly determine the output result. In some embodiments of this application, reference training data is used to enable the preset first and second models to learn the correspondence between speech and lip movements, keeping the facial features of the input image unchanged to achieve accurate results and good stability. Therefore, to ensure better training results for the subsequent first and second models, the selected reference training data needs to be clear, showing the mouth area, and the lip movements must accurately correspond to the speech without any deviation, excessive noise, or the sounds of other people speaking, and must not obscure the mouth.

[0029] The reference training data was obtained through self-recording and network download. Data cleaning was performed before training to delete problematic segments and ensure the cleanliness of the training data.

[0030] After obtaining the reference training data, the reference training data needs to be preprocessed. The preprocessed reference training data includes reference images and reference audio. The preprocessing includes audio feature extraction and image sequence processing.

[0031] After obtaining the reference training data, data preprocessing is required. The data preprocessing work in this application mainly includes audio feature extraction and face cropping.

[0032] The audio feature extraction uses an audio feature extraction model, such as the ProsodyEncoder module, which can extract audio features by overcoming the bottleneck of word-level vector quantization.

[0033] Data preprocessing also includes cropping the video. First, the video is converted into an image sequence. Then, a face detection program is used to detect the face position. The face is then cropped based on the position information and used as the real image for training. After that, the mouth area of ​​the face image is masked and used as the input image, i.e., the reference image, for subsequent training. The model repairs the masked area based on the masked image and voice information to obtain the complete face image.

[0034] S200: The reference training data is input into a preset first model and a second model respectively, so as to instruct the first model to generate a first fitted image based on the reference training data, and to instruct the second model to generate a second fitted image based on the reference training data; wherein the parameter values ​​of the second model are less than the parameter values ​​of the first model.

[0035] This application includes two pre-defined models: a first model and a second model. The first model employs a large amount of ResNet architecture in its audio and image coding networks. This network architecture superimposes the input content onto the output convolutional results, which can significantly alleviate the gradient vanishing problem and enhance the network's learning ability, but it also greatly increases the number of model parameters and computational cost. The second model uses a standard convolutional structure. The number of channels in the intermediate network layers of the second model is 0.25 to 0.4 times that of the first model, and the parameter values ​​of the second model are smaller than those of the first model. In some embodiments, the running speed of the second model is 19 times that of the first model, meeting the requirements for operation on low-computing-power devices such as mobile phones.

[0036] The reference training data is input into the first model and the second model respectively, resulting in a first fitted image generated by the first model based on the reference training data, and a second fitted image generated by the second model based on the reference training data.

[0037] S300: Determine the first true value corresponding to the first model based on the similarity between the first fitted image and the reference image, and train the first model based on the first true value.

[0038] After obtaining the first fitted image, the similarity between the first fitted image and the reference image is compared to determine the difference between the first fitted image output by the first model and the reference image, so as to adjust the parameters of the first model and facilitate the subsequent training and learning of the first model based on the current reference image.

[0039] like Figure 2 As shown, in some embodiments, determining the first true value corresponding to the first model based on the similarity between the first fitted image and the reference image includes: S301: Calculate a first loss value based on the similarity between the first fitted image and the reference image; S302: Calculate a second loss value based on the difference between the first fitted image and the reference image; S303: The first true value is obtained based on the first loss value and the second loss value; the first true value is used to characterize the difference between the first fitted image output by the first model and the reference image.

[0040] During the training and learning process, the first model has two loss functions: the first loss function (concat) and the second loss function. Based on the first loss function and the second loss function, the first true value corresponding to the first model is determined. The first true value is used to characterize the difference between the first fitted image output by the first model and the reference image.

[0041] like Figure 3 As shown, in some embodiments, the first loss function can be... loss The second loss function can be the GAN loss. The calculation formula is as follows: (1) (2) (3) in Let represent the generator in the first model, D represent the discriminator in the first model, y represent the real face image, and x represent the input audio features and image features, i.e., the reference training data. The first image is the fitted image; the generator is used to learn the correspondence between semantics and lip shape, as well as the face fitting ability; the discriminator is used to evaluate the quality of the image synthesized by the generator.

[0042] For GAN loss, its calculation requires feeding the reference image and the first fitted image into a discriminator for encoding, and then obtaining a label to calculate the loss. The reconstruction loss, also known as L1 loss, is calculated by subtracting the reference image from the first fitted image to determine the absolute difference, which is used to evaluate the realism of the first fitted image. For the overall loss, containing and Two loss functions, the purpose of training the generator is to reduce The loss function aims to make the first fitted image resemble the reference image more closely; the discriminator's goal is to maximize... The goal is to distinguish between the reference image and the first fitted image as much as possible. The generator during training is also a continuously decreasing... The process.

[0043] In some embodiments, after obtaining the first true value based on the first loss value and the second loss value, the method further includes: S304: When training the first model, if the first true value is less than the first preset threshold, adjust the parameter values ​​of the first model according to the similarity between the first fitted image and the reference image. S305: Train the first model with adjusted parameter values ​​until convergence, and input the next set of reference training data into the first model.

[0044] After obtaining the first true value, before iteratively training the first model, the first true value needs to be compared with a first preset threshold to determine whether the parameters of the first model need to be adjusted. If the first true value is less than the first preset threshold, the parameter values ​​of the first model are adjusted according to the similarity between the first fitted image and the reference image. The first model with adjusted parameter values ​​is trained until convergence, and the next set of reference training data is input into the first model. If the first true value is not less than the first preset threshold, the current first model is trained until convergence, and the next set of reference training data is input into the first model.

[0045] S400: Determine the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image, and train the second model based on the second true value.

[0046] In some embodiments, the step of determining the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image, and training the second model based on the second true value, includes: The second model is processed based on the comparison between the second true value and the second preset threshold.

[0047] like Figure 4 and Figure 5 As shown, after obtaining the first fitted image and the second fitted image, the similarity between the first fitted image and the second fitted image is compared to determine the difference between the second fitted image output by the second model and the first fitted image output by the first model, so as to adjust the parameters of the second model and facilitate the subsequent training and learning of the second model based on the current first fitted image.

[0048] In this application, the second model no longer fits to the real image, i.e., the reference image, but rather to the first fitted image output by the first model. It only needs to fit a structure similar to the output of the first model, which greatly reduces the learning difficulty. This makes the second model run faster and meets the requirements for running on low-computing-power devices such as mobile phones.

[0049] In some embodiments, determining the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image, and training the second model based on the second true value includes: S401: Obtain the first fitted image and the second fitted image; S402: Calculate the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image; the second true value is used to characterize the difference between the output results of the first model and the second model; During the training and learning process, the second model needs to calculate the structural similarity and perceptual similarity between the second model and the first model to evaluate the difference between the first fitted image output by the first model and the second fitted image output by the second model. The second true value is used to characterize the difference between the output results of the first model and the second model.

[0050] In some embodiments, calculating the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image includes: S4021: Calculate the structural similarity and perceptual similarity based on the first fitted image and the second fitted image, respectively; S4022: Obtain a first weight value through the structural similarity and the first coefficient, and obtain a second weight value through the perceptual similarity and the second coefficient; S4023: Obtain the second true value based on the first weight value and the second weight value.

[0051] In some embodiments, structural similarity, perceptual similarity, and the second true value can be calculated using the following loss function: (4) (5) (6) in The first fitted image output by the first model and the second fitted image output by the second model. Structural similarity loss and perceptual loss are used to evaluate The difference between them, the perceptual loss is the feature reconstruction loss. and style reconstruction loss Two kinds, The average value representing brightness. The standard deviation of contrast. Structural similarity covariance It is a fixed constant. This is a pre-trained VGG network used here for evaluation. Feature similarity, We also use the VGG network intermediate layer for evaluation. , , These are the coefficients for various losses, used to define the weights of each loss.

[0052] After obtaining the second true value, before iteratively training the second model, it is necessary to compare the second true value with the first preset threshold to determine whether the parameters of the second model need to be adjusted.

[0053] S403: If the second true value is less than the second preset threshold, adjust the parameter values ​​of the second model according to the similarity between the first fitted image and the second fitted image; S404: Train the second model with adjusted parameter values ​​until convergence, and input the next set of reference training data into the second model.

[0054] S405: If the second true value is equal to or greater than a preset threshold, train the second model until convergence, and output the second model as the target model.

[0055] If the second true value is less than a second preset threshold, the parameter values ​​of the second model are adjusted according to the similarity between the first fitted image and the second fitted image. The second model with adjusted parameter values ​​is trained until convergence, and the next set of reference training data is input into the second model. If the second true value is equal to or greater than the preset threshold, the second model is trained until convergence, and the second model is output as the target model.

[0056] It is worth noting that the first and second models in this application are trained simultaneously, iterating over both models without any specific order. During training, the first and second models are trained together, with the second model learning based on the first model. The first model is used to assist the second model's learning, allowing the second model to achieve a similar fitting ability with fewer parameters and less computation. After training, the model is deployed to low-computing-power devices. If the first model is trained first, and then the second model is trained based on the first model, to avoid model collapse and gradient vanishing phenomena due to a large difference in model size, the second model needs to maintain a dynamic balance with the first model. If the number of parameters in the first and second models is similar, the second model will not be applicable to low-computing-power devices. If the first model is a pre-trained model, it cannot guide the learning of the second model in the initial stage, leading to problems in the early stages and increasing the risk of overfitting.

[0057] This application trains the first model and the second model together. The first model can determine the fitting of the second model in the initial stage of training and learn iteratively in sync. The two models learn together and are closely combined to obtain a good fitting effect. The second model can fit the parameters more flexibly. Therefore, there is no problem that the first model and the second model must be balanced and the number of parameters cannot be too different. In this way, the second model can use a smaller number of parameters, so that the trained second model can be deployed on low computing power devices, enriching the application scenarios of digital humans.

[0058] S500: Train the first model and the second model until convergence, and output the trained second model as the target model; wherein, the target model is used to generate a lightweight digital human video based on the target audio input by the user.

[0059] In this application, the first model and the second model generate a first fitted image and a second fitted image based on reference training data, respectively. The parameters of the first model are adjusted based on the similarity difference between the first fitted image and the reference image. The output difference between the first model and the second model is determined based on the first fitted image and the second fitted image, thereby adjusting the parameters of the second model. The first model and the second model are trained iteratively simultaneously until they converge. In this way, the second model does not need to directly learn the real images in the reference training data, but only needs to fit the first fitted image output by the first model. This allows it to achieve an output effect close to that of the first model with fewer model parameters and less computation, making the second model applicable to low-computing devices and thus enriching the application scenarios of digital humans.

[0060] In some embodiments, this application provides a lightweight digital human video generation method, including: S600: Acquires the image and target audio input by the user; S601: Based on the image, a template video is synthesized, and the template video is preprocessed to obtain a target image; wherein, the target image is a face image with the mouth area removed; S602: Import the target image and the target audio into the target model to generate a predicted face image; S603: The predicted face image and the target audio are spliced ​​together to obtain a lightweight digital human video.

[0061] After the second model trained according to the above embodiments is output as the target model, the user only needs to upload a clear frontal photo to synthesize a lightweight digital human video with accurate lip movements and realistic effects. The digital human workflow is described in detail below. After the user uploads the image, a template video is synthesized based on the uploaded image. The template video is then preprocessed, similar to the processing and training process, to obtain a face crop image and a mask image. In addition, the user needs to provide audio data, which can be recorded directly with a mobile phone or uploaded existing audio data. Afterwards, the face mask image and audio data are fed into the target model to obtain the lightweight digital human video.

[0062] In this application, after the target model obtains the image, it outputs a template video. The template video undergoes preprocessing: first, the video is converted into an image; then, face detection is performed, and the detected faces are cropped and saved. Next, the user's voice input is obtained, and the voice and face images are input into the trained target model. The target model outputs the face corresponding to the voice based on the voice and image, and the output face image is stitched back into the original video to obtain a lightweight digital human video.

[0063] This application can synthesize lightweight digital human videos with accurate lip movements and natural actions. It can synthesize a video corresponding to the audio with only one image. Because the target model has low parameters and fast running speed, it can run offline on low computing power devices such as mobile phones. It can be widely used in scenarios such as news broadcasting, customer service, and AI assistants, greatly improving the interactive effect.

[0064] like Figure 6 As shown, in some embodiments, this application provides a lightweight digital human video generation system, the system comprising: The acquisition module is used to acquire reference training data and preprocess the reference training data. The preprocessed reference training data includes reference images and reference audio. The preprocessing includes audio feature extraction and image sequence processing. The training module inputs the reference training data into a preset first model and a second model, respectively, to instruct the first model to generate a first fitted image based on the reference training data, and to instruct the second model to generate a second fitted image based on the reference training data; wherein the parameter values ​​of the second model are smaller than the parameter values ​​of the first model; The first true value corresponding to the first model is determined based on the similarity between the first fitted image and the reference image, and the first model is trained based on the first true value. The second true value corresponding to the second model is determined based on the similarity between the first fitted image and the second fitted image, and the second model is trained based on the second true value. The first model and the second model are trained until convergence, and the trained second model is output as the target model; wherein, the target model is used to generate a lightweight digital human video based on the target audio input by the user.

[0065] The reference training data in the acquisition module includes high-definition videos recorded by the database itself and high-definition data downloaded from the internet. Data cleaning was performed before training, removing issues such as misaligned lip movements, multiple voices, significant environmental noise, no human figures or faces, and severe occlusion. Clean data helps the model converge faster and achieve better video synthesis results. It should be noted that the reference training data here refers to general training video data, not videos of specific models for digital human synthesis.

[0066] This application provides a lightweight digital human video generation system, which includes an acquisition module and a training module. The acquisition module is used to acquire reference training data. The training module is used to train a preset second model to obtain a target model with low parameter count and capable of running on low-computing-power devices. Users can generate lightweight digital human videos based on target audio input from the target model.

[0067] like Figure 7 As shown, in some embodiments, the lightweight digital human video generation system provided in this application further includes: A generation module is used to synthesize a template video based on images input by the user, and to preprocess the template video to obtain a target image; wherein the target image is a face image with the mouth area removed; The target image and the target audio input by the user are imported into the target model to generate a predicted face image; The predicted face image is spliced ​​with the target audio to obtain a lightweight digital human video.

[0068] The generation module is equipped with an image driving unit and a lip-sync generation unit. The generation module controls the image driving unit to drive the image to synthesize a template video, and controls the lip-sync generation unit to drive the template video according to the input speech to synthesize a lightweight digital human video that matches the audio.

[0069] The image-driven unit is used to convert input images into videos with head movements and expressions. This unit is equipped with an image-driven model, which is trained using a large amount of unlabeled video data. This model can drive an image with a video, transferring the head movements and expressions from the video to the image, and synthesizing a video with movements and expressions. This video is used as input for subsequent modules and as a template video for subsequent digital human-driven processes.

[0070] The generation module also includes a digital human driving unit, which uses input audio data to drive the template video synthesized by the image driving unit, synthesizing a video that matches the speech. By training with recorded video data, it can learn the correspondence between speech and lip movements.

[0071] This application also provides a lightweight digital human video generation system, the system including an acquisition module, a training module, and a generation module; the acquisition module is used to acquire reference training data and user-input images and target audio; the training module is used to train a preset second model to obtain a target model with low parameter count and capable of running on low-computing-power devices; the generation module is used to control the target model to synthesize and output a lightweight digital human video based on user-input images and target audio.

[0072] The lightweight digital human video generation system proposed in this application can train a digital human model and generate digital human videos based on only a single image. Furthermore, the digital human model only requires server training; the video generation process can be performed on low-computing-power devices such as mobile phones, without the need for high-performance graphics cards. Based on this system, lightweight digital human videos with accurate lip movements and natural gestures can be generated from audio. This technology can run on low-computing-power devices and can operate offline, significantly lowering the barrier to entry and enabling the application of digital humans in more scenarios.

[0073] In some embodiments, this application also provides an electronic device, including: processor; Memory for storing the executable instructions of the processor; The processor is configured to execute the aforementioned lightweight digital human generation model training method and lightweight digital human video generation method by executing the executable instructions.

[0074] As can be seen from the above embodiments, this application provides a training method for a lightweight digital human generation model and a digital human video generation method. The training method includes: acquiring reference training data; inputting the reference training data into a preset first model and a second model respectively, to instruct the first model to generate a first fitted image based on the reference training data, and to instruct the second model to generate a second fitted image based on the reference training data; determining a first real value corresponding to the first model based on the similarity between the first fitted image and the reference image, and training the first model based on the first real value; determining a second real value corresponding to the second model based on the similarity between the first fitted image and the second fitted image, and training the second model based on the second real value; training the first model and the second model until convergence, and outputting the trained second model as the target model. The second model does not need to directly learn the real images in the reference training data, but only needs to fit the first fitted image output by the first model, so as to achieve an output effect close to that of the first model with fewer model parameters and less computation, making the second model applicable to low-computing-power devices, thereby enriching the application scenarios of digital humans.

Claims

1. A method for training a lightweight digital human generation model, characterized in that, The method includes: Acquire reference training data and preprocess the reference training data. The preprocessed reference training data includes reference images and reference audio. The preprocessing includes audio feature extraction and image sequence processing. The reference training data is input into a preset first model and a second model respectively, so as to instruct the first model to generate a first fitted image based on the reference training data, and to instruct the second model to generate a second fitted image based on the reference training data; wherein the parameter values ​​of the second model are smaller than the parameter values ​​of the first model; The first true value corresponding to the first model is determined based on the similarity between the first fitted image and the reference image, and the first model is trained based on the first true value. The second true value corresponding to the second model is determined based on the similarity between the first fitted image and the second fitted image, and the second model is trained based on the second true value. The first model and the second model are trained until convergence, and the trained second model is output as the target model; wherein, the target model is used to generate a lightweight digital human video based on the target audio input by the user.

2. The method of claim 1, wherein, The step of determining the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image, and training the second model based on the second true value, includes: The second model is processed based on the comparison between the second true value and the second preset threshold.

3. The method of claim 2, wherein, The second model is then trained based on the similarity between the first fitted image and the second fitted image to determine the second true value corresponding to the second model, including: Obtain the first fitted image and the second fitted image; The second true value corresponding to the second model is calculated based on the similarity between the first fitted image and the second fitted image; the second true value is used to characterize the difference between the output results of the first model and the second model. If the second true value is less than the second preset threshold, the parameter values ​​of the second model are adjusted according to the similarity between the first fitted image and the second fitted image. The second model, with its parameter values ​​adjusted, is trained until convergence, and the next set of reference training data is input into the second model.

4. The method of claim 3, wherein, The calculation of the second true value corresponding to the second model based on the similarity between the first fitted image and the second fitted image includes: Calculate structural similarity and perceptual similarity based on the first fitted image and the second fitted image, respectively. A first weight value is obtained by using the structural similarity and the first coefficient, and a second weight value is obtained by using the perceptual similarity and the second coefficient; The second true value is obtained based on the first weight value and the second weight value.

5. The method of claim 3, wherein the method further comprises: The method further includes: If the second true value is equal to or greater than a preset threshold, the second model is trained until convergence, and the second model is output as the target model.

6. The method of claim 1, wherein, The first true value corresponding to the first model is determined based on the similarity between the first fitted image and the reference image, including: A first loss value is calculated based on the similarity between the first fitted image and the reference image; The second loss value is calculated based on the difference between the first fitted image and the reference image; The first true value is obtained based on the first loss value and the second loss value; the first true value is used to characterize the difference between the first fitted image output by the first model and the reference image.

7. The method of claim 6, wherein, After obtaining the first true value based on the first loss value and the second loss value, the process further includes: When training the first model, if the first true value is less than the first preset threshold, the parameter values ​​of the first model are adjusted according to the similarity between the first fitted image and the reference image. The first model, after its parameter values ​​have been adjusted, is trained until convergence, and the next set of reference training data is input into the first model.

8. A lightweight digital human video generation method, characterized by, The lightweight digital human video generation method uses the target model in the training method of the lightweight digital human generation model according to any one of claims 1 to 7, and the lightweight digital human video generation method includes: Obtain the image and target audio input by the user; A template video is synthesized based on the image, and a target image is obtained by preprocessing the template video; wherein, the target image is a face image with the mouth area removed; The target image and the target audio are imported into the target model to generate a predicted face image; The predicted face image is spliced ​​with the target audio to obtain a lightweight digital human video.

9. A lightweight digital human video generation system, characterized by, The system includes: The acquisition module is used to acquire reference training data and preprocess the reference training data. The preprocessed reference training data includes reference images and reference audio. The preprocessing includes audio feature extraction and image sequence processing. The training module inputs the reference training data into a preset first model and a second model, respectively, to instruct the first model to generate a first fitted image based on the reference training data, and to instruct the second model to generate a second fitted image based on the reference training data; wherein the parameter values ​​of the second model are smaller than the parameter values ​​of the first model; The first true value corresponding to the first model is determined based on the similarity between the first fitted image and the reference image, and the first model is trained based on the first true value. The second true value corresponding to the second model is determined based on the similarity between the first fitted image and the second fitted image, and the second model is trained based on the second true value. The first model and the second model are trained until convergence, and the trained second model is output as the target model; wherein, the target model is used to generate digital human videos based on target audio input by the user; A generation module is used to synthesize a template video based on images input by the user, and to preprocess the template video to obtain a target image; wherein the target image is a face image with the mouth area removed; The target image and the target audio input by the user are imported into the target model to generate a predicted face image; The predicted face image is spliced ​​with the target audio to obtain a lightweight digital human video.

10. An electronic device, characterized in that, include: processor; Memory for storing the executable instructions of the processor; The processor is configured to execute the training method of the lightweight digital human generation model according to any one of claims 1-7, and / or execute the lightweight digital human video generation method according to claim 8, by executing the executable instructions.