Method and device for optimizing and generating lip-shaped driving video effect and storage medium
By randomly selecting and limiting the angle difference between the reference graph and the target graph in the training of the generative digital human lip-synchronous model, the problem of slow convergence and poor effect in the prior art is solved, and more efficient training and better robustness are achieved.
Patent Information
- Application Number
- CN202510130118.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
AI Technical Summary
When the existing generative digital humans are synchronized, the model converges too slowly and has poor effect due to the random selection of reference pictures, especially when the head of the character is rotated at a large angle.
By randomly selecting reference maps with angle differences from the target map during the training process and limiting the angle differences within the set range, ensure the generalization of the model while avoiding overfitting.
This method inhibits overfitting of the model, improves robustness, and can train the model faster, saving 50% of training time under the same effect.
Smart Images

Figure CN120075548A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital humans, and specifically provides a method, a device and a storage medium for optimizing and generating lip-driven video effects. Background Art
[0002] In today's era of rapid technological development, artificial intelligence (AI) has become an indispensable part of our lives. From smartphones, self-driving cars to smart home systems, AI technology has penetrated into every aspect of our lives. Among these AI technologies, Generative Digital Humans, as an emerging technology, is gradually attracting people's attention.
[0003] Generative digital humans are a character model based on artificial intelligence technology. They can generate virtual characters with specific characteristics and behaviors by learning and analyzing large amounts of data. These digital humans can be applied to various scenarios, such as film and television, games, and advertising, to bring users a more realistic and immersive experience.
[0004] The basic framework of the generative digital human currently on the market is to use the reference image Ref img, the target image targetimg, and the mask Figure 3 The information is used as the model input. The ref img is usually randomly selected and is the same video, but a different frame from the target frame.
[0005] However, we found that in the same video, when a person is talking, the head and face can turn sideways at a large angle. The difference between turning 50 degrees to the left and turning 50 degrees to the right is very large. The introduction of the reference image is to provide the generative model with the person's head posture and facial texture reference. However, purely randomly selecting the reference image Ref img is likely to bring a large deviation. It is inappropriate to randomly select the reference image Ref img without any method.
[0006] The current mainstream generative lip sync drivers, such as Musetalk, wav2lip, faceswap and other lip sync models, all randomly select the reference image Ref img. Although Musetalk adds a constraint of at least 5 frames apart, it actually still randomly selects the reference image Ref img, and cannot solve the technical problem of randomly selecting the reference image Ref img. Excessive differences will cause the model to converge too slowly and affect the effect.
[0007] In view of this, this application is hereby filed. Summary of the invention
[0008] In view of the above technical problems, the present invention provides a method, device and storage medium for optimizing the generation of lip-driven video effects, which can suppress overfitting in model training and has better robustness.
[0009] Specifically, the following technical solutions are adopted:
[0010] In a first aspect, the present invention provides a method for optimizing the generation of lip-driven video effects, including: inputting a voice file and a video file into a lip-sync deep learning model to output a voice video, in which the dynamic lip effect of the digital human in the voice video matches the text content in the voice file;
[0011] The training process of the lip-sync deep learning model includes:
[0012] Performing data preprocessing on the training data sample set, and randomly selecting a reference image Ref img from the training data sample set whose angular difference from the target image target img is limited within a set range;
[0013] Using the reference image Ref img, the target image target img and the mask image as inputs to the lip-sync deep learning model for training.
[0014] As an optional implementation manner of the present invention, in a method for optimizing the generation of lip-driven video effects of the present invention, the performing data preprocessing on the training data sample set and randomly selecting a reference image Ref img from the training data sample set whose angular difference from the target image target img is limited within a set range includes:
[0015] Calculating the facial landmark coordinates in the target image target img;
[0016] Calculating the head pose deflection angle using the facial landmarks;
[0017] Randomly selecting a picture from the same video as the target image target img but in a different frame as a candidate reference image;
[0018] Calculating the angular difference between the candidate reference image and the target image target img, and determining whether the angular difference is within the set range;
[0019] If the judgment result is yes, determining the candidate reference image as the reference image Ref img; if the judgment result is no, looping to select a candidate reference image and determine whether the angular difference is within the set range until a candidate reference image with an angular difference within the set range is selected as the reference image Ref img.
[0020] As an alternative embodiment of the present invention, in a method for optimizing the generation of a lip-driven video effect of the present invention, during the training process of using the reference image Ref img, the target image target img, and the mask image as the inputs of the lip synchronization deep learning model, when the inputs and outputs of the entire time series segment occur, the same reference image Ref img is selected for multiple target images target img of the entire time series segment.
[0021] As an alternative embodiment of the present invention, in a method for optimizing the generation of a lip-driven video effect of the present invention, during the data preprocessing process for the training data sample set, when it is determined that the face is in a deflected state by calculating the head pose deflection angle using the face landmarks, the face is rotated to be upright based on the head pose deflection angle and then input into the lip synchronization deep learning model for training.
[0022] As an alternative embodiment of the present invention, in a method for optimizing the generation of a lip-driven video effect of the present invention, the training process of the lip synchronization deep learning model includes using a multi-scale loss function loss to evaluate the effect of the lip synchronization deep learning model and determining whether the lip synchronization deep learning model has completed training.
[0023] As an alternative embodiment of the present invention, in a method for optimizing the generation of a lip-driven video effect of the present invention, the training process of the lip synchronization deep learning model includes using a multi-scale loss function loss to evaluate the effect of the lip synchronization deep learning model, including:
[0024] The first loss function L1 loss(pred, target) corresponding to the mask region of the lower half of the face;
[0025] The second loss function L1loss(lip_pred, target_lip) corresponding to the lip region;
[0026] The third loss function L1 loss(teeth_pred, target_teeth) corresponding to the teeth region;
[0027] The perceptual loss function perceptual_loss of the mask region, lip region, and teeth region of the lower half of the face calculated by the Vgg19 model.
[0028] As an alternative embodiment of the present invention, in a method for optimizing the generation of a lip-driven video effect of the present invention, the pixel height of the mask region of the lower half of the face is greater than the pixel height of the lip region, the pixel height of the lip region is greater than the pixel height of the teeth region, and the pixel widths of the mask region, lip region, and teeth region of the lower half of the face are scaled proportionally according to their respective pixel heights.
[0029] In a second aspect, the present invention provides an apparatus for optimizing the generation of a lip-driven video effect, including a lip synchronization deep learning model module: inputting a voice file and a video file into the lip synchronization deep learning model module, and outputting a voice video, in which the dynamic lip effect of the digital human in the voice video matches the text content in the voice file;
[0030] A deep learning model training module for training the lip synchronization deep learning model, and the training process of the lip synchronization deep learning model includes:
[0031] Performing data preprocessing on the training data sample set, and randomly selecting a reference image Ref img from the training data sample set, where the angular difference between the reference image Ref img and the target image target img is limited within a set range;
[0032] Using the reference image Ref img, the target image target img, and the mask image as the inputs of the lip synchronization deep learning model for training.
[0033] In a third aspect, the present invention provides an electronic device, including a processor and a memory, where the memory is used to store a computer-executable program, and when the computer program is executed by the processor, the processor executes the method for optimizing the generation of a lip-driven video effect.
[0034] In a fourth aspect, the present invention provides a computer-readable recording medium storing a computer-executable program, and when the computer-executable program is executed, the method for optimizing the generation of a lip-driven video effect is implemented.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] In the training process of the lip synchronization deep learning model of the method for optimizing the generation of a lip-driven video effect of the present invention, a reference image Ref img with an angular difference from the target image target img is selected. The existence of the difference is for the generalization of the model, but too large a difference will cause the model to converge too slowly and affect the effect. The present invention randomly selects the reference image Ref img to ensure that there is an angular difference from the target image target img, and at the same time limits the angular difference between the reference image Ref img and the target image target img, avoiding the situation that too large a difference will cause the model to converge too slowly and affect the effect.
[0037] Therefore, the method for optimizing the generation of lip-driven video effects according to the present invention can suppress overfitting during model training and has higher robustness. At the same time, the method for optimizing the generation of lip-driven video effects according to the present invention can train the model faster. In the case of the same effect, the solution of this embodiment can save 50% of the training time. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Flowchart of the method for optimizing the generation of lip-driven video effects according to Embodiment 1 of the present invention;
[0039] Figure 2 Schematic structural diagram of the electronic device according to Embodiment 2 of the present invention;
[0040] Figure 3 Schematic diagram of the computer-readable recording medium according to Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention.
[0042] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely represents some embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0043] It should be noted that, without conflict, the embodiments in the present invention and the features and technical solutions in the embodiments may be combined with each other.
[0044] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0045] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the present invention is usually placed during use, or the orientation or positional relationship commonly understood by those skilled in the art. Such terms are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.
[0046] Example 1
[0047] See Figure 1 As shown, this embodiment provides a method for optimizing the generation of lip-driven video effects, including: inputting a voice file and a video file into a lip-sync deep learning model to output a voice video, where the dynamic lip effect of the digital human in the voice video matches the text content in the voice file;
[0048] The training process of the lip-sync deep learning model includes:
[0049] Perform data preprocessing on the training data sample set, and randomly select a reference image Ref img from the target images targetimg in the training data sample set whose angular difference from the target image target img is restricted within a set range;
[0050] Use the reference image Ref img, the target image target img, and the mask image as the inputs of the lip-sync deep learning model for training.
[0051] In the method for optimizing the generation of lip-driven video effects in this embodiment, during the training process of the lip-sync deep learning model, a reference image Ref img with an angular difference from the target image target img is selected. The existence of the difference is for the generalization of the model, but too large a difference will cause the model to converge too slowly and affect the effect. This embodiment randomly selects the reference image Ref img to ensure the existence of an angular difference from the target image target img, and at the same time restricts the angular difference between the reference image Ref img and the target image target img to avoid the situation that too large a difference will cause the model to converge too slowly and affect the effect.
[0052] Therefore, the method for optimizing the generation of lip-driven video effects in this embodiment can suppress overfitting during model training and has higher robustness; at the same time, the method for optimizing the generation of lip-driven video effects in this embodiment can train the model faster. In the case of the same effect, the solution of this embodiment can save 50% of the training time.
[0053] The lip-sync deep learning model in this embodiment mainly includes:
[0054] ·MuseTalk
[0055] MuseTalk is a real-time audio-driven lip-sync model. This model can automatically adjust the facial image of the digital human according to the input audio signal to make its lip shape highly synchronized with the audio content.
[0056] Musetalk is a real-time high-quality audio-driven lip-sync model ft-mse-vae trained in the latent space, where:
[0057] Modify unseen faces according to the input audio, and the size of the face area is 256x256.
[0058] Support audio in multiple languages such as Chinese, English, Japanese, etc.
[0059] Support real-time inference at 30fps+ on NVIDIA Tesla V100.
[0060] Support modifying only the center point of the face area, which significantly affects the generation result. MuseTalk is trained in the latent space, where the images are encoded by a frozen VAE. The audio is encoded by a frozen whisper-tiny model. The architecture of the generation network draws on UNet stable-diffusion-v1-4, where the audio embedding is fused with the image embedding through cross-attention. It uses an architecture very similar to Stable Diffusion, but the uniqueness of MuseTalk is that it is not a diffusion model.
[0061] ·Wav2Lip
[0062] Wav2Lip is a deep learning algorithm for generating lip-sync videos. It can automatically add accurate lip movements to a given face video according to the input audio stream.
[0063] The Wav2Lip model is built based on the Generative Adversarial Network (GAN), which consists of two main parts: a generator and a discriminator. The generator is responsible for generating realistic facial animations based on the input audio waveform, while the discriminator is responsible for distinguishing the generated animations from real facial animations. The detailed description of its main structure and working principle is as follows:
[0064] 1. Discriminator (D_{SyncNet}): The first stage is to train a discriminator that can judge whether the sound and lip movements are synchronized. The goal of this discriminator is to improve the ability to judge the synchronization of sound and lip movements.
[0065] 2. Generator (encoding-decoding model structure): The second stage adopts an encoding-decoding model structure, including a generator and two discriminators. The generator tries to generate facial animations synchronized with the audio, while the two discriminators are responsible for judging the synchronization and visual quality of the generated animations and real animations respectively.
[0066] 3. Main modules: The Wav2Lip model includes three main modules:
[0067] (1). Identity Encoder: Responsible for encoding the random reference frame to extract identity features.
[0068] (2). Speech Encoder: Encodes the input speech segment into facial animation features.
[0069] (3). Face Decoder: Upsamples the encoded features and finally generates facial animation.
[0070] ·Faceswap
[0071] The basic principle of the Faceswap model is to extract the facial features from an image or video through deep learning technology and apply them to another image or video to achieve face replacement. Specifically, it first extracts the features of the input face image through a deep neural network, then uses a generative adversarial network (GAN) to generate a replacement face that matches the target face, and finally fuses the replacement face with the original image through image fusion technology to obtain the final face-swapping effect.
[0072] Deep learning models: Faceswap mainly relies on deep learning technology, especially convolutional neural networks (CNNs) and generative adversarial networks (GANs). These models can automatically learn feature representations from a large amount of data, thus enabling efficient extraction and generation of facial features.
[0073] Face detection and alignment: Before extracting facial features, it is necessary to accurately detect and align the input face. This involves the use of computer vision techniques such as Haar features, Dlib, etc.
[0074] Image fusion: After generating the replacement face, it is necessary to seamlessly fuse it with the original image. This requires the use of image processing techniques such as Gaussian mixture models, Poisson fusion, etc. to achieve a natural transition.
[0075] In addition, the above lip-sync deep learning model is only used as an exemplary illustration, and a method for optimizing the generation of lip-driven video effects in this embodiment can be applied to other lip-sync deep learning models.
[0076] As an alternative implementation of this embodiment, in a method for optimizing the generation of lip-driven video effects in this embodiment, the data preprocessing of the training data sample set, and randomly selecting a reference image Ref img from the training data sample set whose angular difference from the target image target img is limited within a set range includes:
[0077] Calculate the facial landmark coordinates in the target image (target img);
[0078] Calculate the head pose deflection angle using the facial landmarks;
[0079] Randomly select an image from the same video as the target image (target img), but from a different frame than the target image, as the candidate reference image;
[0080] Calculate the angular difference between the candidate reference image and the target image (target img), and determine whether the angular difference is within a set range;
[0081] If the judgment result is yes, determine the candidate reference image as the reference image (Ref img). If the judgment result is no, loop to select a candidate reference image and determine whether the angular difference is within the set range until a candidate reference image with an angular difference within the set range is selected as the reference image (Ref img).
[0082] In an optimized method for generating lip-driven video effects in this embodiment, considering the model training speed, the calculation of the angular difference in the above process is placed in the data preprocessing stage.
[0083] Furthermore, in an optimized method for generating lip-driven video effects in this embodiment, during the training process of taking the reference image (Ref img), the target image (target img), and the mask image as the inputs of the lip synchronization deep learning model, when there are inputs and outputs for an entire time series segment, the same reference image (Ref img) is selected for multiple target images (target img) in this entire time series segment.
[0084] Due to restricting the influence degree of the reference image (Ref img) on model prediction, too large an influence is likely to cause overfitting. Experiments have shown that when there are inputs and outputs for an entire time series segment in the model, the same number of reference images (Ref img) and target images (target img) will cause overfitting during training, resulting in the predicted image being highly similar to the reference image (Ref img). Selecting the same reference image (Ref img) for multiple target images (target img) in an entire time series segment is a suitable choice.
[0085] As an alternative implementation of this embodiment, in an optimized method for generating lip-driven video effects in this embodiment, during the data preprocessing of the training data sample set, when it is determined that the face is in a deflected state by calculating the head pose deflection angle using the facial landmarks, the face is rotated to be upright based on the head pose deflection angle and then input into the lip synchronization deep learning model for training.
[0086] Experiments have shown that when the range of face deflection in the training set is large and there is a large amount of such data, straightening the face during the data preprocessing stage can accelerate convergence.
[0087] As an alternative implementation of this embodiment, an optimized method for generating lip-driven video effects in this embodiment, the training process of the lip synchronization deep learning model includes using a multi-scale loss function loss to evaluate the effect of the lip synchronization deep learning model and determining whether the lip synchronization deep learning model is trained.
[0088] Using a multi-scale loss function loss in this embodiment during model training has a better effect than using a single-scale loss function loss and can suppress the problem of lip jitter.
[0089] Specifically, in an optimized method for generating lip-driven video effects in this embodiment, the training process of the lip synchronization deep learning model includes using a multi-scale loss function loss to evaluate the effect of the lip synchronization deep learning model, including:
[0090] The first loss function L1 loss(pred, target) corresponding to the mask region of the lower half of the face;
[0091] The second loss function L1loss(lip_pred, target_lip) corresponding to the lip region;
[0092] The third loss function L1 loss(teeth_pred, target_teeth) corresponding to the tooth region;
[0093] The perceptual loss function perceptual_loss of the mask region, lip region, and tooth region of the lower half of the face calculated by the Vgg19 model.
[0094] Specifically, in an optimized method for generating lip-driven video effects in this embodiment, the pixel height of the mask region of the lower half of the face is greater than the pixel height of the lip region, the pixel height of the lip region is greater than the pixel height of the tooth region, and the pixel widths of the mask region, lip region, and tooth region of the lower half of the face are scaled proportionally according to their respective pixel heights.
[0095] Specific example solutions of an optimized method for generating lip-driven video effects in this embodiment include:
[0096] 1. Calculate the coordinates of the face landmarks (landmark).
[0097] 2. Calculate the head pose deflection angle using the face landmarks (landmark).
[0098] 3. Calculate the angular difference between the reference image Ref img and the target image target img, limit the difference within a certain range, and at the same time, the difference must be maintained.
[0099] Note: The existence of the angular difference is for the generalization of the model. However, too large an angular difference will cause the model to converge too slowly and affect the effect.
[0100] In addition, considering the model training speed, the calculation of the angle in the above process is placed in the data preprocessing stage.
[0101] 4. Constrain the influence degree of the reference image Ref img on the model prediction. Too large an influence is likely to cause overfitting. Experiments have shown that when there are input and output of the entire time series segment in the model, for the same number of reference images Ref img and target images target img, it will cause the model training to overfit, resulting in the predicted image being highly similar to the reference image Ref img, and one reference image Ref img is a suitable choice.
[0102] 5. Straighten the face of the original data during model training. Experiments have shown that when the face deflection range in the training set is large and there is a large amount of such data, straightening the face in the preprocessing stage can accelerate convergence.
[0103] 6. Using multi-scale loss during training has a better effect than using a single-scale loss. Specifically:
[0104] a) The L1 loss (pred, target) corresponding to the mask of the lower half of the face;
[0105] b) The L1 loss (lip_pred, target_lip) corresponding to the lips;
[0106] c) The L1 loss (teeth_pred, target_teeth) corresponding to the teeth;
[0107] d) The perceptual_loss of the above three regions calculated by the Vgg19 model.
[0108] e) In terms of size, they are 256, 100, and 50 pixels in height, and the width is scaled proportionally.
[0109] This embodiment also provides a device for optimizing the generation effect of the lip shape driving video, including a lip shape synchronization deep learning model module: inputting a voice file and a video file into the lip shape synchronization deep learning model module, and outputting a voice video, in which the digital human lip shape dynamic effect in the voice video matches the text content in the voice file;
[0110] A deep learning model training module for training the lip-sync deep learning model. The training process of the lip-sync deep learning model includes:
[0111] Perform data preprocessing on the training data sample set, and randomly select a reference image Ref img from the training data sample set whose angular difference from the target image target img is limited within a set range.
[0112] Use the reference image Ref img, the target image target img, and the mask image as the inputs of the lip-sync deep learning model for training.
[0113] In the device for optimizing the generation of lip-driven video effects in this embodiment, during the training process of the lip-sync deep learning model by the deep learning model training module, a reference image Ref img with an angular difference from the target image target img is selected. The existence of the difference is for the generalization of the model, but too large a difference will cause the model to converge too slowly and affect the effect. Therefore, in this embodiment, the reference image Ref img is randomly selected to ensure an angular difference from the target image target img, and at the same time, the angular difference between the reference image Ref img and the target image target img is restricted to avoid the model from converging too slowly and affecting the effect due to too large a difference.
[0114] Therefore, the device for optimizing the generation of lip-driven video effects in this embodiment can suppress overfitting and has higher robustness during model training; at the same time, the device for optimizing the generation of lip-driven video effects in this embodiment can train the model faster. In the case of the same effect, the solution of this embodiment can save 50% of the training time.
[0115] The lip-sync deep learning model trained by the deep learning model training module in this embodiment includes, but is not limited to, MuseTalk, Wav2Lip, and Faceswap.
[0116] As an alternative implementation of this embodiment, the deep learning model training module in this embodiment performs data preprocessing on the training data sample set, and randomly selects a reference image Ref img from the training data sample set whose angular difference from the target image target img is limited within a set range, including:
[0117] Calculate the facial landmark coordinates in the target image target img;
[0118] Calculate the head pose deflection angle using the facial landmarks;
[0119] Randomly select a picture from the same video as the target img but from a different frame than the target img as the candidate reference picture;
[0120] Calculate the angular difference between the candidate reference picture and the target img, and determine whether the angular difference is within a set range;
[0121] If the judgment result is yes, determine the candidate reference picture as the reference picture Ref img. If the judgment result is no, loop to select a candidate reference picture and determine whether the angular difference is within the set range until a candidate reference picture with an angular difference within the set range is selected as the reference picture Ref img.
[0122] For an apparatus for optimizing the generation of lip-driven video effects in this embodiment, considering the model training speed, the calculation of the angular difference in the above process is placed in the data preprocessing stage.
[0123] Furthermore, for an apparatus for optimizing the generation of lip-driven video effects in this embodiment, during the training process of the deep learning model training module using the reference picture Ref img, the target picture target img, and the mask picture as the inputs of the lip synchronization deep learning model, when there are inputs and outputs for the entire time series segment, the same reference picture Ref img is selected for multiple target pictures target img in the entire time series segment.
[0124] Due to restricting the influence degree of the reference picture Ref img on model prediction, too large an influence is likely to cause overfitting. Experiments have shown that when there are inputs and outputs for the entire time series segment in the model, the same number of reference pictures Ref img and target pictures target img will cause overfitting in training, resulting in the predicted picture being highly similar to the reference picture Ref img. Selecting the same reference picture Ref img for multiple target pictures target img in the entire time series segment is a suitable choice.
[0125] As an alternative implementation of this embodiment, during the data preprocessing process of the deep learning model training module of this embodiment for the training data sample set, when it is determined that the human face is in a deflected state by calculating the head pose deflection angle using facial landmarks, the human face is rotated to be upright based on the head pose deflection angle and then input into the lip synchronization deep learning model for training.
[0126] Experiments have shown that when the deflection range of the human face in the training set is large and there is a large amount of such data, rotating the human face to be upright in the data preprocessing stage can accelerate convergence.
[0127] As an alternative implementation of this embodiment, the training process of the lip-sync deep learning model in the deep learning model training module of this embodiment includes using the multi-scale loss function loss to evaluate the effect of the lip-sync deep learning model and determining whether the lip-sync deep learning model is trained.
[0128] When training the model in this embodiment, using the multi-scale loss function loss has a better effect than using a single-scale loss function loss, and can suppress the problem of lip jitter.
[0129] Specifically, for a device for optimizing the generation of lip-driven video effects in this embodiment, the training process of the lip-sync deep learning model including using the multi-scale loss function loss to evaluate the effect of the lip-sync deep learning model includes:
[0130] The first loss function L1 loss(pred,target) corresponding to the mask region of the lower half of the face;
[0131] The second loss function L1loss(lip_pred,target_lip) corresponding to the lip region;
[0132] The third loss function L1 loss(teeth_pred,target_teeth) corresponding to the tooth region;
[0133] The perceptual loss function perceptual_loss of the mask region, lip region, and tooth region of the lower half of the face calculated by the Vgg19 model.
[0134] Specifically, for a device for optimizing the generation of lip-driven video effects in this embodiment, the pixel height of the mask region of the lower half of the face is greater than the pixel height of the lip region, the pixel height of the lip region is greater than the pixel height of the tooth region, and the pixel widths of the mask region, lip region, and tooth region of the lower half of the face are scaled proportionally according to their respective pixel heights.
[0135] Example 2
[0136] Next, an embodiment of the electronic device of the present invention is described. This electronic device can be regarded as a specific physical implementation manner of the above method and device embodiments of the present invention. For the details described in the embodiment of the electronic device of the present invention, they should be regarded as a supplement to the above method or device embodiments; for the details not disclosed in the embodiment of the electronic device of the present invention, they can be implemented with reference to the above method or device embodiments.
[0137] Figure 2It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory. The memory is used to store computer-executable programs. When the computer program is executed by the processor, the processor executes a method for optimizing and generating a lip-driven video effect in Embodiment 1.
[0138] As Figure 2 shown, the electronic device is presented in the form of a general-purpose computing device. The processor can be one or multiple and work collaboratively. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity and can also be the sum of multiple physical devices.
[0139] The memory stores computer-executable programs, usually machine-readable code. The computer-readable program can be executed by the processor so that the electronic device can execute the method of the present invention or at least some of the steps in the method.
[0140] The memory includes volatile memory, such as a random access storage unit (RAM) and / or a cache storage unit, and can also be non-volatile memory, such as a read-only storage unit (ROM).
[0141] Optionally, in this embodiment, the electronic device further includes an I / O interface, which is used for the electronic device to exchange data with external devices. The I / O interface can represent one or more of several bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in multiple bus structures.
[0142] It should be understood that Figure 2 the shown electronic device is only an example of the present invention. The electronic device of the present invention may further include elements or components not shown in the above example. For example, some electronic devices further include a display unit such as a display screen, and some electronic devices further include human-computer interaction elements, such as buttons, keyboards, etc. As long as the electronic device can execute the computer-readable program in the memory to implement the method of the present invention or at least some of the steps of the method, it can be considered as the electronic device covered by the present invention.
[0143] Figure 3 It is a schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. As Figure 3As shown, a computer-readable recording medium stores a computer-executable program which, when executed, implements a method for optimizing the generation of a lip-driven video effect according to Embodiment 1 of the present invention. The computer-readable recording medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable recording medium may also be any readable medium other than the readable recording medium, and this readable medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable recording medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0144] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., connected through the Internet using an Internet service provider).
[0145] From the above description of the embodiments, those skilled in the art can easily understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. included in the system. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, an electronic control unit, a client, a server, etc. However, it should be noted that the computer software for executing the method of the present invention is not limited to being executed in one or a specific number of hardware entities. It can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product may be stored in a computer-readable recording medium (which may be a CD-ROM, a USB flash drive, a removable disk, etc.), or may be distributed and stored on a network, as long as it enables an electronic device to execute the method according to the present invention.
[0146] The above embodiments are only used to illustrate the present invention and do not limit the technical solutions described in the present invention. Although this specification has described the present invention in detail with reference to the above respective embodiments, the present invention is not limited to the above specific embodiments. Therefore, any modification or equivalent replacement of the present invention; and all technical solutions and their improvements that do not depart from the spirit and scope of the invention are covered by the scope of the claims of the present invention.
Claims
1. A method for optimizing and generating lip-driven video effects, characterized in that: include: Input the voice file and video file into the lip synchronization deep learning model, output the voice video, and the dynamic effect of the digital human's lip shape in the voice video matches the text content in the voice file; The training process of the lip synchronization deep learning model includes: performing data preprocessing on a training data sample set, randomly selecting a reference image Ref img whose angle difference with the target image target img is limited to a set range based on a target image target img in the training data sample set; and training the reference image Ref img, the target image target img and the mask image as inputs of the lip synchronization deep learning model.
2. The method for optimizing and generating lip-driven video effects according to claim 1, characterized in that: The data preprocessing is performed on the training data sample set, and the reference image Ref i mg is randomly selected based on the target image target img in the training data sample set, and the angle difference between the target image target img and the reference image Ref i mg is limited to a set range, including: Calculate the coordinates of the facial landmarks in the target image target i mg; Calculate the head posture deflection angle using facial landmarks; Randomly select images from the same video as the target image target img but in different frames as candidate reference images; Calculate the angle difference between the candidate reference image and the target image target img, and determine whether the angle difference is within a set range; If the judgment result is yes, the candidate reference image is determined to be the reference image Ref img. If the judgment result is no, the candidate reference image is selected in a loop and it is determined whether the angle difference is within the set range until a candidate reference image with an angle difference within the set range is selected as the reference image Ref img.
3. The method for optimizing and generating lip-driven video effects according to claim 2, characterized in that: In the process of training the reference image Ref img, the target image target img and the mask image as inputs of the lip synchronization deep learning model, when the input and output of the entire time sequence segment appear, the same reference image Ref img is selected for multiple target images target img of the entire time sequence segment.
4. The method for optimizing and generating lip-driven video effects according to claim 2, characterized in that: During the data preprocessing process for the training data sample set, when the face landmark is used to calculate the head posture deflection angle and it is determined that the face is in a deflected state, the face is straightened based on the head posture deflection angle and then input into the lip synchronization deep learning model for training.
5. The method for optimizing and generating lip-driven video effects according to claim 2, characterized in that: The training process of the lip synchronization deep learning model includes using a multi-scale loss function loss to evaluate the effect of the lip synchronization deep learning model and determine whether the lip synchronization deep learning model has completed training.
6. The method for optimizing and generating lip-driven video effects according to claim 5, characterized in that: The training process of the lip synchronization deep learning model includes using a multi-scale loss function loss to evaluate the effect of the lip synchronization deep learning model, including: The first loss function corresponding to the mask area of the lower half of the face is L1 loss(pred,target); The second loss function corresponding to the lip area is L1 loss (lip_pred, target_lip); The third loss function corresponding to the tooth area is L1 loss (teeth_pred, target_teeth); The perceptual loss function perceptual_loss of the mask area, lip area and tooth area of the lower half of the face calculated by the Vgg19 model.
7. The method for optimizing and generating lip-driven video effects according to claim 6, characterized in that: The pixel height of the mask area of the lower half of the face is greater than the pixel height of the lip area, and the pixel height of the lip area is greater than the pixel height of the teeth area. The pixel widths of the mask area, lip area and teeth area of the lower half of the face are scaled proportionally according to their respective pixel heights.
8. A device for optimizing and generating lip-driven video effects, characterized in that: Including lip synchronization deep learning model module: input voice files and video files into the lip synchronization deep learning model module, output voice video, and the dynamic effect of the digital human lip shape in the voice video matches the text content in the voice file; A deep learning model training module is used to train the lip synchronization deep learning model. The training process of the lip synchronization deep learning model includes: Perform data preprocessing on the training data sample set, and randomly select a reference image Ref i mg whose angle difference with the target image target img is limited to a set range based on the target image target img in the training data sample set; The reference image Ref img, the target image target img and the mask image are used as inputs of the lip sync deep learning model for training.
9. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program, characterized in that: When the computer program is executed by the processor, the processor performs a method for optimizing generation of lip-driven video effects as described in any one of claims 1-7.
10. A computer-readable recording medium storing a computer-executable program, characterized in that: When the computer executable program is executed, a method for optimizing and generating lip-driven video effects as described in any one of claims 1 to 7 is implemented.