Virtual human dubbing image generation method and device, equipment and storage medium
By performing spatial deformation and lip lip repair processing on the source image, driving audio and reference images, the problem of insufficient dubbing image quality in the prior art is solved, and a virtual human dubbing image with high definition and high synchronization is generated, which is suitable for the fields of virtual reality and augmented reality.
Patent Information
- Application Number
- CN202510471866.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-15
AI Technical Summary
The dubbing images generated by the prior art under zero sample conditions are of poor quality, high noise, and cannot effectively deal with complex conditions such as lighting changes and dynamic backgrounds, resulting in inaccurate and natural lip movements.
By acquiring source images, driving audio and reference images, spatial deformation processing and lip lip repair are performed, and high-quality virtual human dubbing images are generated using an adaptive channel spatial deformation operator and feature decoder, combining perceived loss, adversarial network loss and mouth lip synchronization loss optimization generation process.
It significantly improves the clarity and delicateness of dubbing images, improves the synchronization accuracy of audio and lip styling, and generates realistic virtual face animations, which are suitable for fields such as virtual reality and augmented reality.
Smart Images

Figure CN120495132A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer processing technology, and in particular to a method, device, equipment and storage medium for generating a virtual human dubbing image. Background Art
[0002] Facial visual dubbing technology aims to synchronize the lip movements, expressions, and facial movements of the characters in the video with the audio content, thereby achieving a more natural and realistic visual effect. For zero-shot learning, achieving facial visual dubbing in high-resolution videos remains a key challenge.
[0003] In related technologies, Wav2Lip designed a face decoder consisting of six deconvolution layers. This decoder's architecture aims to gradually restore low-dimensional latent feature maps to face images, gradually increasing the spatial resolution of the image through deconvolution operations. TRVD combines upsampling layers and adaptive instance normalization (AdaIN) in its face decoder. The upsampling layers increase the image resolution through methods such as interpolation or transposed convolution, while AdaIN is used to adaptively adjust the normalization parameters based on the input features, effectively fusing features from different sources.
[0004] There are at least the following shortcomings and deficiencies in the relevant existing technical solutions: Models such as Wav2Lip and TRVD require large amounts of high-quality, well-labeled video and audio data during training to ensure they can accurately learn the mapping between audio and lip sync. However, in zero-shot conditions, the correlation between mouth texture details and the driving audio is low, making it extremely challenging to directly generate high-frequency texture details. As a result, the lip movements generated by the model in low-quality, blurry, or noisy videos may not be accurate or natural. Summary of the Invention
[0005] The present invention provides a method, device, equipment and storage medium for generating a virtual human dubbing image, which is used to solve the defects of poor quality and high noise of the dubbing images generated in the prior art. By performing spatial deformation processing and lip shape repair processing on the source image, driving audio and reference image, the blurring phenomenon in the generation process is effectively reduced, the clarity and fineness of the dubbing image are improved, and the accuracy of the synchronization between audio and lip shape is further improved.
[0006] In a first aspect, the present invention provides a method for generating a virtual human dubbing image, comprising the following steps: Acquire a source image, acquire a driving audio, and acquire a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's voice emitted by the user's mouth, and the reference image is represented by an image describing the user's lip posture and the user's facial expression; Performing spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of a virtual person; wherein the facial image features indicate that a head posture of the virtual person is aligned with a head posture of the user, and that a mouth shape of the virtual person is synchronized with a mouth shape of the user; The facial image features of the virtual person are repaired to generate a dubbing image of the virtual person; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted from the virtual person's mouth when the virtual person's head posture is the same as the user's head posture, the virtual person's mouth shape is the same as the user's mouth shape posture, and / or the virtual person's facial expression is the same as the user's facial expression.
[0007] Preferably, according to the method for generating a virtual human dubbing image provided by the present invention, performing spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of the virtual human includes: Extracting audio features of the driving audio, extracting source features of the source image using a first feature encoder, and extracting reference features of the reference image using a second feature encoder; Aligning the source feature with the reference feature to obtain an aligned feature; A spatial deformation process is performed based on the alignment features, the audio features, and the reference features to generate the facial image features of the virtual person.
[0008] Preferably, according to a method for generating a virtual human dubbing image provided by the present invention, performing spatial deformation processing based on the alignment features, the audio features, and the reference features to generate the facial image features of the virtual human includes: performing splicing processing on the alignment feature and the audio feature to obtain a head posture audio splicing feature; Using an adaptive channel spatial deformation operator and the head posture audio splicing feature, the reference feature is spatially deformed to generate a deformation feature; The deformed features and the source features are spliced together to generate the facial image features of the virtual person.
[0009] Preferably, according to a method for generating a virtual human dubbing image provided by the present invention, the repairing of the facial image features of the virtual human to generate the virtual human dubbing image includes: The facial image features are input into a feature decoder for decoding and repair processing to generate the dubbing image of the virtual person.
[0010] Preferably, according to the method for generating a virtual human dubbing image provided by the present invention, after the step of repairing the facial image features of the virtual human to generate the virtual human dubbing image, the method includes: Calculating a perceptual loss and an adversarial network loss for generating the dubbing image, and calculating a lip synchronization loss; Based on the perceptual loss, the adversarial network loss, and the lip synchronization loss, a total loss for generating the facial image features of the virtual person is determined.
[0011] Preferably, according to the method for generating a virtual human dubbing image provided by the present invention, determining the total loss of generating the facial image features of the virtual human based on the perceptual loss, the adversarial network loss, and the lip synchronization loss includes: Obtaining preset perceptual loss weights and obtaining preset lip-sync loss weights; performing a first calculation process on the perception loss and the perception loss weight to obtain a first perception loss, and performing a second calculation process on the lip synchronization loss and the lip synchronization loss weight to obtain a second lip synchronization loss; The first perceptual loss, the second lip synchronization loss, and the adversarial network loss are summed to obtain a total loss for generating the facial image features of the virtual person.
[0012] In a second aspect, the present invention further provides a device for generating a virtual human dubbing image, comprising the following modules: an acquisition module, configured to acquire a source image, acquire a driving audio, and acquire a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's voice emitted from the user's mouth, and the reference image is represented by an image describing the user's lip posture and facial expression; a deformation module, configured to perform spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of a virtual person; wherein the facial image features indicate that the head posture of the virtual person is aligned with the head posture of the user, and the lip shape of the virtual person is synchronized with the lip shape of the user; A repair module is used to repair the facial image features of the virtual person to generate a dubbing image of the virtual person; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted from the virtual person's mouth when the head posture of the virtual person is the same as the head posture of the user, the mouth shape of the virtual person is the same as the mouth shape posture of the user, and / or the facial expression of the virtual person is the same as the facial expression of the user.
[0013] In a third aspect, the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for generating a virtual human dubbing image as described above is implemented.
[0014] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for generating a virtual human dubbing image.
[0015] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described methods for generating a virtual human dubbing image.
[0016] The present invention provides a method, device, equipment and storage medium for generating a virtual human dubbing image, which comprises obtaining a source image, obtaining driving audio and obtaining a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's voice emitted by the user's mouth, and the reference image is represented by an image describing a user's lip posture and a user's facial expression; spatial deformation processing is performed on the source image, the driving audio and the reference image to generate facial image features of the virtual human; wherein the facial image features represent that the virtual human's head posture is aligned with the user's head posture, and the virtual human's mouth shape is synchronized with the user's lip posture; and the facial image features of the virtual human are repaired to generate a dubbing image of the virtual human; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted by the virtual human's mouth when the virtual human's head posture is the same as the user's head posture, the virtual human's mouth shape is the same as the user's mouth shape, and / or the virtual human's facial expression is the same as the user's facial expression. It is used to solve the defects of poor quality and high noise in the dubbing images generated in the existing technology. By performing spatial deformation processing and lip repair processing on the source image, driving audio and reference image, it effectively reduces the blurring phenomenon in the generation process, improves the clarity and fineness of the dubbing image, and further improves the accuracy of audio and lip synchronization. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1This is one of the flow charts of a method for generating a virtual human dubbing image provided by the present invention.
[0019] Figure 2 This is the second schematic diagram of a method for generating a virtual human dubbing image provided by the present invention.
[0020] Figure 3 It is a structural schematic diagram of a virtual human dubbing image generation device provided by the present invention.
[0021] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0023] In the related art, there are at least the following technical problems: Lipsync3D achieves personalized 3D speaking face generation through pose and lighting normalization. This method utilizes a 3D face model to generate corresponding lip movements based on driving audio, and combines pose information to generate synchronized facial expressions. While Lipsync3D significantly improves data efficiency and the naturalness of generated facial animations, its reliance on high-quality 3D face models limits its generalization across diverse individuals and scenarios.
[0024] 1) Generating high-resolution, detailed facial animations, such as those used in the TRVD model, requires high computational requirements in model design. This makes it difficult to generate high-quality lip movements while maintaining real-time performance, and the model may struggle to find the optimal balance between real-time performance and generation quality. Compared to dubbing for specific audiences, zero-shot dubbing, when relying on direct generation methods, often produces blurrier results, lacking clear textures and details, which impacts visual quality and user experience.
[0025] 2) Existing models generally cannot handle complex conditions such as changing lighting, dynamic backgrounds, dangling earrings, flowing hair, and camera motion. As a result, in these cases, the model may not correctly generate facial movements or even generate artifacts outside the face, affecting the overall visual realism.
[0026] 3) Models such as Lipsync3D rely on high-quality 3D facial models. The construction of these models usually requires specialized equipment and complex processing, which increases the complexity of the system, places higher demands on computing resources, increases costs and difficulty, and limits the popularity and application scope of the technology.
[0027] The following combination Figures 1-4 The present invention describes a method, device, equipment and storage medium for generating a virtual human dubbing image, which is used to address the defects of poor quality and high noise in the dubbing images generated in the prior art. By performing spatial deformation processing and lip shape repair processing on the source image, driving audio and reference image, the blurring phenomenon in the generation process is effectively reduced, the clarity and fineness of the dubbing image are improved, and the accuracy of audio and lip shape synchronization is further improved.
[0028] Figure 1 This is one of the flow charts of a method for generating a virtual human dubbing image provided by the present invention. Figure 1 As shown, the method may include but is not limited to steps S100 to S300: S100, obtaining a source image, obtaining a driving audio, and obtaining a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's voice emitted by the user's mouth, and the reference image is represented by an image describing the user's lip posture and facial expression; S200, performing spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of a virtual person; wherein the facial image features indicate that the head posture of the virtual person is aligned with the head posture of the user, and the mouth shape of the virtual person is synchronized with the mouth shape and posture of the user; S300, repairing the facial image features of the virtual person to generate a dubbing image of the virtual person; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted from the virtual person's mouth when the head posture of the virtual person is the same as the head posture of the user, the mouth shape of the virtual person is the same as the mouth shape posture of the user, and / or the facial expression of the virtual person is the same as the facial expression of the user.
[0029] In step S100 of some embodiments, a source image is obtained, a driving audio is obtained, and a reference image is obtained; wherein the source image is represented as an image describing the user's head posture, the driving audio is represented as audio describing the user's sound emitted by the user's mouth, and the reference image is represented as an image describing the user's lip posture and the user's facial expression.
[0030] The source image is an image describing the user's head posture. Acquiring such an image usually requires the use of a camera device, such as a smart phone, a computer camera, or a professional photographic device, which is not specifically limited in the embodiments of the present invention.
[0031] The driving audio is the audio that describes the sound emitted by the user's mouth. Acquiring this audio requires a recording device, such as a microphone, a mobile phone, or a professional recording device, which is not specifically limited in the embodiment of the present invention.
[0032] The reference image is an image that describes the user's lip posture and facial expression. Acquiring this image also requires the use of a camera device, which is not specifically limited in the embodiment of the present invention.
[0033] In some embodiments of the present invention, Figure 2 As shown, get a source image , a driving audio (384 is the dimension of the audio feature, whisper features are used in this invention) and five reference images .
[0034] In step S200 of some embodiments, spatial deformation processing is performed on the source image, the driving audio, and the reference image to generate facial image features of the virtual person.
[0035] The facial image feature indicates that the head posture of the virtual person is aligned with the head posture of the user, and the mouth shape of the virtual person is synchronized with the mouth shape and posture of the user.
[0036] It can be understood that the specific execution steps may be: extracting the audio features of the driving audio, extracting the source features of the source image using a first feature encoder, and extracting the reference features of the reference image using a second feature encoder; Aligning the source feature with the reference feature to obtain an aligned feature; A spatial deformation process is performed based on the alignment features, the audio features, and the reference features to generate the facial image features of the virtual person.
[0037] In some embodiments, the driving audio is input into an audio encoder, and feature extraction processing is performed using the audio encoder to extract audio features.
[0038] Extract driving audio through the audio encoder Audio Encoder Audio features , Encoded voice content.
[0039] The source image is input into the first feature encoder for feature extraction processing to extract source features, and the reference image is input into the second feature encoder for feature extraction processing to extract reference features.
[0040] Specifically, the source image is extracted through two different feature encoders Feature Encoder and reference images Source feature and reference features .
[0041] An embodiment of aligning the source feature and the reference feature to obtain an alignment feature: Bundle and Concatenate and input into an Alignment Encoder to calculate alignment features . Encoded and Alignment information of head poses between them.
[0042] In some embodiments of the present invention, performing spatial deformation processing based on the alignment features, the audio features, and the reference features to generate the facial image features of the virtual person includes: performing splicing processing on the alignment feature and the audio feature to obtain a head posture audio splicing feature; Using an adaptive channel spatial deformation operator and the head posture audio splicing feature, the reference feature is spatially deformed to generate a deformation feature; The deformed features and the source features are spliced together to generate the facial image features of the virtual person.
[0043] Furthermore, the alignment feature and the audio feature are spliced to obtain a head posture audio splicing feature.
[0044] Aligned features are the coordinates of facial key points (such as eyes, nose, mouth, etc.) extracted from the source and reference images. Ensure that these feature points are spatially aligned, that is, the corresponding points in the source and reference images match.
[0045] Audio features are features extracted from the driving audio, such as Mel-Frequency Cepstral Coefficients (MFCC). These features can capture key information in the audio signal, such as pitch, loudness, and timbre. The alignment features and audio features are concatenated in the channel dimension. Specifically, if the alignment features is a tensor of shape (H, W, Calign), and the audio features is a tensor of shape (T, Caudio), then the feature shape of the head pose audio splicing feature after splicing will be (H, W, Calign+Caudio), where H and W are the height and width of the feature map, Calign is the number of channels of the alignment feature, and Caudio is the number of channels of the audio feature.
[0046] By concatenating the aligned features with the audio features, we can leverage both visual and auditory information to generate more realistic facial animations. This concatenation enables the subsequent spatial deformation process to consider both head pose and audio-driven mouth movements.
[0047] Furthermore, we define the adaptive channel space deformation operator: Design a neural network module that can dynamically adjust deformation parameters based on input features. This module can be based on a deep learning model (such as a convolutional neural network or Transformer) to learn complex relationships between input features.
[0048] The previously concatenated head pose audio splicing feature is input into the adaptive channel space deformation operator. This input contains the position information of facial key points and audio driving information.
[0049] The adaptive channel-space warping operator computes a set of warping parameters based on the input features. These parameters describe how to warp the reference features to match the target expression or mouth shape.
[0050] The calculated deformation parameters are then used to perform spatial deformation on the reference feature to obtain the deformed feature. The deformation process can be an affine transformation, a nonlinear transformation, or other complex geometric transformations, depending on the implementation of the deformation operator.
[0051] In some embodiments, an Adaptive Adaptive Transformation (AdaAT) operator is used. and ,Will Perform spatial deformation The AdaAT operator can deform feature maps with misaligned spatial layouts by performing feature channel-specific deformations, computing different affine coefficients in different feature channels.
[0052] By applying the adaptive channel space deformation operator, the shape and position of the reference features can be dynamically adjusted according to the head pose and audio-driven information, thereby generating more natural and smooth facial animation effects.
[0053] Furthermore, the deformation feature and the source features Perform splicing processing to generate the facial image features of the virtual person so that In
[15] , the feature map of the reference image is spatially warped to synchronize the mouth shape with the driving audio and to align the head pose with the source image.
[0054] By splicing the deformation features with the source features, the target expression or mouth shape of the reference image can be fused with the details and texture information of the source image, thereby generating a highly realistic virtual human facial animation effect.
[0055] In step S300 of some embodiments, the facial image features of the virtual person are repaired to generate a dubbing image of the virtual person; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted from the virtual person's mouth when the virtual person's head posture is the same as the user's head posture, the virtual person's mouth shape is the same as the user's mouth shape posture, and / or the virtual person's facial expression is the same as the user's facial expression.
[0056] It is understandable that, combined with Figure 2 As shown, in the repairing stage, the facial image features of the virtual person are repaired to generate a dubbing image of the virtual person.
[0057] Specifically, the facial image features are input into a feature decoder for decoding and repair processing to generate the dubbing image of the virtual person.
[0058] By repairing part ,from and Generate dubbing images To achieve this goal, we first concatenate and Then, a feature decoder with convolutional layers is used to repair the masked mouth and generate a dubbing image. .
[0059] More specifically, the previously generated virtual human facial image features are input into the feature decoder. These features include facial key points, expression details, and texture information after audio-driven and spatial deformation processing. A mask is used to cover the mouth area in the virtual human facial image. This mask can be a binary matrix with the same size as the feature map of the facial image, where the value of the masked mouth area is 1 and the value of other areas is 0. The purpose of the mask is to protect the mouth area in the source image from the influence of the subsequent restoration process, ensuring that the audio-driven mouth movements can be clearly reflected in the final dubbing image.
[0060] Apply a feature decoder: A feature decoder with convolutional layers is used to process the masked facial image features. The feature decoder typically consists of multiple convolutional layers, activation functions, and possibly residual connections to gradually restore image details and structure.
[0061] During decoding, the feature decoder uses the concatenated audio and image features to repair the masked mouth area and generate a complete dubbed image.
[0062] Generate dubbing image: After being processed by the feature decoder, the restored facial image features are obtained.
[0063] Converting these features into the final dubbed image is typically accomplished through a pixel-wise classification or regression layer that maps each feature point to a specific pixel value or color.
[0064] Through these steps, high-quality virtual human voiceover images can be restored and generated from the encoded and deformed facial image features. This process combines audio-driven information with the characteristics of the original facial image, resulting in a generated voiceover image that not only has realistic expressions and lip syncing effects, but also maintains the details and texture information of the source image.
[0065] In some embodiments, the user's head posture is detected and the virtual person's head posture is matched. The user's lip posture is detected and the virtual person's mouth shape is matched. If necessary, the user's facial expression can also be detected and the virtual person's facial expression is matched. This can be achieved through emotion recognition algorithms, such as deep learning models that analyze the user's facial muscle movements and expression changes and apply this information to the virtual person's facial model. Ensure that the audio produced by the virtual person has the same content as the audio of the user's voice. The user's voice can be converted to text using speech recognition technology, and then text-to-speech (TTS) technology can be used to generate an audio signal with the same content as the original audio. Alternatively, the user's audio signal can be directly transmitted to the virtual person, allowing it to repeat the user's voice content in lip sync.
[0066] The head pose, lip pose, facial expression and audio content obtained in the above steps are combined to generate the final dubbing image representation.
[0067] This dubbing image representation will be a dynamic sequence of video frames in which the virtual person's head pose is the same as the user's, the mouth shape is consistent with the user's, and the audio signal with the same content as the user is emitted.
[0068] This process generates a highly realistic virtual human voice-over video, where the virtual human's movements and expressions are closely synchronized with the real user, while also emitting the same audio signal as the user. This technology has broad application prospects in virtual reality, augmented reality, online conferencing, and other fields, providing a more immersive and interactive experience.
[0069] In some embodiments of the present invention, after the step of repairing the facial image features of the virtual person to generate a dubbing image of the virtual person, the method includes: Calculating a perceptual loss and an adversarial network loss for generating the dubbing image, and calculating a lip synchronization loss; Based on the perceptual loss, the adversarial network loss, and the lip synchronization loss, a total loss for generating the facial image features of the virtual person is determined.
[0070] It can be understood that the perceptual loss is calculated at two image scales, in the deformation part and the restoration part. Specifically, the dubbing image and real images Downsample to and , then, pair the images and The input is fed into a pre-trained VGG-19 network to compute the perceptual loss.
[0071] The pre-trained VGG-19 network is a deep convolutional neural network model pre-trained on large-scale image datasets (such as ImageNet). VGG-19 contains multiple convolutional layers, each of which consists of a 3×3 convolution kernel and appropriate padding. These convolutional layers are divided into 5 stages, each of which consists of multiple convolutional layers, and there are ReLU activation layers between the convolutional layers.
[0072] At the end of each stage, except the last one, there is a 2×2 max pooling layer (MaxPoolingLayer) with a stride of 2, which helps to reduce the spatial size of the feature map, thereby reducing the number of parameters and computation.
[0073] After all convolution and pooling layers, there is a flatten layer (FlattenLayer), which flattens the multi-dimensional feature map into a one-dimensional vector for input to the fully connected layer.
[0074] VGG-19 contains 3 fully connected layers (FullyConnectedLayer), the first and second fully connected layers are followed by ReLU activation layers.
[0075] The output of the last fully connected layer is sent to the Softmax layer (SoftmaxLossLayer) to calculate the probability of each category.
[0076] Since the model has been pre-trained on a large-scale dataset, it has learned rich image feature representations. Therefore, in practical applications, it can reduce training time and computing resources and achieve better performance.
[0077] The perceptual loss is calculated as follows: in, represents the perceptual loss for generating dubbed images, Represents the VGG-19 network layer, yes The characteristic size of the layer, Indicates a dubbing image, represents the real image, represents the image features of the downsampled dubbing image, Represents the image features of the real image downsampled, and N represents the number of channel layers.
[0078] In some embodiments, using the effective LS-GAN loss, the adversarial network loss (GAN loss) is calculated as follows: in, represents the generated adversarial network loss, G represents the generator, D represents the discriminator, Indicates a dubbing image, Represents a real image. GAN loss is used on a single frame and five consecutive frames.
[0079] Generator (G): The generator's main task is to receive a random noise vector as input and attempt to generate samples that are as close as possible to the real data distribution. These samples can be in various forms, such as images, audio, and text. By learning the statistical properties of real data, the generator continuously adjusts its parameters to produce increasingly realistic samples.
[0080] Discriminator (D): The discriminator's primary task is to determine whether the input data is real or generated by the generator. It receives a sample as input and outputs a probability value, indicating the probability that the sample is real. By learning the differences between real and generated data, the discriminator continuously optimizes its parameters to improve the accuracy of its judgment.
[0081] Training process: Initialization: First, initialize the parameters of the generator and discriminator. Usually, these parameters are randomly initialized.
[0082] Fix the generator and train the discriminator: During training, the generator parameters are first fixed to remain constant. Then, the discriminator is trained using both real and generated data. The discriminator's goal is to maximize its ability to distinguish between real and generated data. That is, for real data, the discriminator should output a probability close to 1; for generated data, the discriminator should output a probability close to 0.
[0083] Fix the discriminator and train the generator: Next, fix the parameters of the discriminator so that they remain unchanged. Then, use the data generated by the generator to train the generator. The goal of the generator is to maximize the discriminator's prediction probability for the generated data, that is, to make the discriminator mistakenly classify the generated data as real data as much as possible.
[0084] Repeat training: The above two steps are performed alternately until the preset number of training rounds is reached or the performance of the generator and discriminator no longer improves significantly.
[0085] Discriminator loss: The discriminator loss function usually consists of two parts: one is the negative log-likelihood of the probability that a real sample is judged to be real; the other is the negative log-likelihood of the probability that a generated sample is judged to be generated. By minimizing this loss function, the discriminator can better distinguish between real data and generated data.
[0086] Generator loss: The loss function of the generator is usually the negative log-likelihood of the discriminator’s predicted probability for the generated samples. By minimizing this loss function, the generator can generate more realistic samples.
[0087] The generative adversarial network continuously learns and optimizes its own parameters through the mutual confrontation and cooperation between the generator and the discriminator, thereby achieving the goal of generating increasingly realistic data samples, that is, the generated audio images of virtual people are more realistic.
[0088] In some embodiments of the present invention, lip sync loss is mainly used to improve the synchronization of lip movements in videos composed of dubbed images. Deep speech features are used to replace audio spectrograms and retrain SycNet to improve the accuracy of lip sync. The formula for calculating lip sync loss is as follows: in, Expressed as lip sync loss, Indicates a dubbing image, Indicates driving audio.
[0089] In some embodiments of the present invention, the step of determining the total loss of generating the facial image features of the virtual person based on the perceptual loss, the adversarial network loss, and the lip synchronization loss specifically includes: Obtaining preset perceptual loss weights and obtaining preset lip-sync loss weights; performing a first calculation process on the perception loss and the perception loss weight to obtain a first perception loss, and performing a second calculation process on the lip synchronization loss and the lip synchronization loss weight to obtain a second lip synchronization loss; The first perceptual loss, the second lip synchronization loss, and the adversarial network loss are summed to obtain a total loss for generating the facial image features of the virtual person.
[0090] It is understandable that the preset perceptual loss weight is , the preset lip loss weight is .
[0091] The perception loss and the perception loss weight are multiplied to obtain a first perception loss, and the lip synchronization loss and the lip synchronization loss weight are multiplied to obtain a second lip synchronization loss.
[0092] The first perceptual loss, the second lip synchronization loss, and the adversarial network loss are summed to obtain a calculation formula for the total loss for generating the facial image features of the virtual person as follows: Where, Denotes the total loss for generating the facial image features of the virtual person, Denoted as the first perceptual loss, Expressed as the second lip sync loss, represents the generated adversarial network loss, Expressed as the perceptual loss weight, Expressed as the lip shape loss weight, represents the perceptual loss for generating dubbed images, Represented as lip sync loss.
[0093] In an embodiment of the present invention, the perceptual loss weight is set =10, lip shape loss weight =0.1.
[0094] In the embodiments of the present invention, perceptual loss is a perceptual quality metric used to measure the similarity between two images in image processing and computer vision tasks. It measures image similarity by comparing the differences between the target image and the generated image in high-level feature maps, rather than relying solely on pixel-level differences. Perceptual loss can better reflect the semantic content and style characteristics of the image, making the generated image more consistent with human visual perception. GAN loss is composed of a generator and a discriminator. Through training with a large amount of sample data, the generator's generation ability and the discriminator's discrimination ability are gradually improved in the confrontation. The goal is to make the generator produce fake data as realistic as possible, while allowing the discriminator to better distinguish between real and fake data.
[0095] The embodiments provided by the present invention have at least the following technical effects: 1) The network designed in this paper achieves high-resolution facial dubbing and outperforms existing techniques in preserving texture detail. By spatially warping the feature maps of the reference image, it better preserves and reproduces high-frequency texture details, significantly improving the visual quality and realism of the dubbed image.
[0096] 2) In this embodiment of the present invention, a comprehensive loss function combining perceptual loss, GAN loss, and lip synchronization loss is employed during training. The perceptual loss is calculated at multiple scales using a pre-trained VGG-19 network, improving the perceived quality of the image. The GAN loss employs the LS-GAN strategy, enhancing the realism of the generated image. The lip synchronization loss ensures precise alignment of the audio and lip movements. This method outperforms existing methods in visual metrics such as structural similarity (SSIM), peak signal-to-noise ratio (PSNR), and perceptual image patch similarity (LPIPS). Furthermore, in user studies, the dubbed images generated by the proposed model also performed well in authenticity scores, achieving high user acceptance.
[0097] 3) The network designed in this invention not only performs well in the current high-resolution video face dubbing task, but its modular and adaptive design also opens up possibilities for future applications in different scenarios and multimodal data. By further optimizing the synchronization network and deformation strategy, the model proposed in this invention has the potential to play a role in a wider range of visual generation tasks. The network designed in this invention consists of a deformation part and a patching part, with clear division of labor. The deformation part uses the AdaAT operator to achieve spatial deformation of specific channels, generating a deformation feature map that is highly synchronized with the driving audio and the source image head posture; the patching part uses a feature decoder to fuse the deformed mouth features with the source image to generate a natural and smooth dubbing image. This modular design improves the flexibility and generation effect of the network.
[0098] The present invention provides a method, device, equipment and storage medium for generating a virtual human dubbing image, which comprises obtaining a source image, obtaining driving audio and obtaining a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's voice emitted by the user's mouth, and the reference image is represented by an image describing a user's lip posture and a user's facial expression; spatial deformation processing is performed on the source image, the driving audio and the reference image to generate facial image features of the virtual human; wherein the facial image features represent that the virtual human's head posture is aligned with the user's head posture, and the virtual human's mouth shape is synchronized with the user's lip posture; and the facial image features of the virtual human are repaired to generate a dubbing image of the virtual human; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted by the virtual human's mouth when the virtual human's head posture is the same as the user's head posture, the virtual human's mouth shape is the same as the user's mouth shape, and / or the virtual human's facial expression is the same as the user's facial expression. It is used to solve the defects of poor quality and high noise in the dubbing images generated in the existing technology. By performing spatial deformation processing and lip repair processing on the source image, driving audio and reference image, it effectively reduces the blurring phenomenon in the generation process, improves the clarity and fineness of the dubbing image, and further improves the accuracy of audio and lip synchronization.
[0099] The following describes a virtual human dubbing image generation device provided by the present invention. The virtual human dubbing image generation device described below and the virtual human dubbing image generation method described above can refer to each other.
[0100] like Figure 3 This is a structural diagram of a virtual human dubbing image generation device provided by the present invention, which includes the following modules: An acquisition module 310 is configured to acquire a source image, a driving audio, and a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's mouth producing a user's voice, and the reference image is represented by an image describing the user's lip posture and facial expression; a deformation module 320 configured to perform spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of a virtual person; wherein the facial image features indicate that the head posture of the virtual person is aligned with the head posture of the user, and the lip shape of the virtual person is synchronized with the lip shape of the user; The repair module 330 is used to repair the facial image features of the virtual person and generate a dubbing image of the virtual person; wherein the dubbing image represents an image of the virtual person's mouth emitting audio with the same content as the audio of the user's voice when the virtual person's head posture is the same as the user's head posture, the virtual person's mouth shape is the same as the user's mouth shape posture, and / or the virtual person's facial expression is the same as the user's facial expression.
[0101] Preferably, the virtual human dubbing image generation device provided by the present invention is further configured to extract audio features of the driving audio, extract source features of the source image using a first feature encoder, and extract reference features of the reference image using a second feature encoder; Aligning the source feature with the reference feature to obtain an aligned feature; A spatial deformation process is performed based on the alignment features, the audio features, and the reference features to generate the facial image features of the virtual person.
[0102] Preferably, the virtual human dubbing image generation device provided by the present invention performs splicing processing on the alignment feature and the audio feature to obtain a head posture audio splicing feature; Using an adaptive channel spatial deformation operator and the head posture audio splicing feature, the reference feature is spatially deformed to generate a deformation feature; The deformed features and the source features are spliced together to generate the facial image features of the virtual person.
[0103] Preferably, according to the apparatus for generating a virtual human dubbing image provided by the present invention, the facial image features are input into a feature decoder for decoding and repairing, thereby generating the dubbing image of the virtual human.
[0104] Preferably, the virtual human dubbing image generation device provided by the present invention calculates the perceptual loss and adversarial network loss of generating the dubbing image, and calculates the lip synchronization loss of the mouth; Based on the perceptual loss, the adversarial network loss, and the lip synchronization loss, a total loss for generating the facial image features of the virtual person is determined.
[0105] Preferably, the virtual human dubbing image generation device provided by the present invention obtains a preset perception loss weight and a preset lip-sync loss weight; performing a first calculation process on the perception loss and the perception loss weight to obtain a first perception loss, and performing a second calculation process on the lip synchronization loss and the lip synchronization loss weight to obtain a second lip synchronization loss; The first perceptual loss, the second lip synchronization loss, and the adversarial network loss are summed to obtain a total loss for generating the facial image features of the virtual person.
[0106] The present invention provides a method, device, equipment and storage medium for generating a virtual human dubbing image, which comprises obtaining a source image, obtaining driving audio and obtaining a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's voice emitted by the user's mouth, and the reference image is represented by an image describing a user's lip posture and a user's facial expression; spatial deformation processing is performed on the source image, the driving audio and the reference image to generate facial image features of the virtual human; wherein the facial image features represent that the virtual human's head posture is aligned with the user's head posture, and the virtual human's mouth shape is synchronized with the user's lip posture; and the facial image features of the virtual human are repaired to generate a dubbing image of the virtual human; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted by the virtual human's mouth when the virtual human's head posture is the same as the user's head posture, the virtual human's mouth shape is the same as the user's mouth shape, and / or the virtual human's facial expression is the same as the user's facial expression. It is used to solve the defects of poor quality and high noise in the dubbing images generated in the existing technology. By performing spatial deformation processing and lip repair processing on the source image, driving audio and reference image, it effectively reduces the blurring phenomenon in the generation process, improves the clarity and fineness of the dubbing image, and further improves the accuracy of audio and lip synchronization.
[0107] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4As shown, the electronic device may include: a processor (processor) 410 , a communication interface (Communications Interface) 420 , a memory (memory) 430 and a communication bus 440 , wherein the processor 410 , the communication interface 420 , and the memory 430 communicate with each other via the communication bus 440 . The processor 410 can call the logic instructions in the memory 430 to execute the method for generating a virtual human dubbing image, which includes: obtaining a source image, obtaining a driving audio, and obtaining a reference image; wherein the source image is represented as an image describing the user's head posture, the driving audio is represented as an audio describing the user's voice emitted by the user's mouth, and the reference image is represented as an image describing the user's lip posture and the user's facial expression; performing spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of the virtual human; wherein the facial image features represent that the head posture of the virtual human is aligned with the user's head posture, and the lip shape of the virtual human's mouth is synchronized with the user's lip posture; performing repair processing on the facial image features of the virtual human to generate a dubbing image of the virtual human; wherein the dubbing image represents an image of the virtual human's mouth emitting audio with the same content as the audio of the user's voice when the virtual human's head posture is the same as the user's head posture, the virtual human's mouth shape is the same as the user's lip posture, and / or the virtual human's facial expression is the same as the user's facial expression.
[0108] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0109] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the virtual human dubbing image generation method provided by the above methods, which includes: obtaining a source image, obtaining a driving audio, and obtaining a reference image; wherein, the source image is represented as an image describing the user's head posture, the driving audio is represented as an audio describing the user's voice emitted by the user's mouth, and the reference image is represented as an image describing the user's lip posture and the user's facial expression; the source image, the driving audio, and the reference image are spatially processed. The method further comprises performing a deformation process between the virtual person and the user, generating facial image features of the virtual person, wherein the facial image features represent that the head posture of the virtual person is aligned with the head posture of the user, and the mouth shape of the virtual person is synchronized with the mouth shape posture of the user; and performing a repair process on the facial image features of the virtual person to generate a dubbing image of the virtual person, wherein the dubbing image represents an image in which the virtual person's mouth emits an audio having the same content as the audio of the user's voice when the head posture of the virtual person is the same as the head posture of the user, the mouth shape of the virtual person is the same as the mouth shape posture of the user, and / or the facial expression of the virtual person is the same as the facial expression of the user.
[0110] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the virtual human dubbing image generation method provided by the above methods, the method comprising: obtaining a source image, obtaining a driving audio, and obtaining a reference image; wherein the source image is represented as an image describing the user's head posture, the driving audio is represented as an audio describing the user's voice emitted by the user's mouth, and the reference image is represented as an image describing the user's lip posture and the user's facial expression; spatial deformation processing is performed on the source image, the driving audio, and the reference image to generate a virtual human face image. The facial image feature represents that the head posture of the virtual person is aligned with the head posture of the user, and the mouth shape of the virtual person is synchronized with the mouth shape posture of the user; the facial image feature of the virtual person is repaired to generate a dubbing image of the virtual person; wherein, the dubbing image represents an image of the virtual person's mouth emitting audio with the same content as the audio of the user's voice when the head posture of the virtual person is the same as the head posture of the user, the mouth shape of the virtual person is the same as the mouth shape posture of the user, and / or the facial expression of the virtual person is the same as the facial expression of the user.
[0111] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0112] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for generating a virtual human dubbing image, characterized in that: include: Acquire a source image, acquire a driving audio, and acquire a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's voice emitted by the user's mouth, and the reference image is represented by an image describing the user's lip posture and the user's facial expression; Performing spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of a virtual person; wherein the facial image features indicate that a head posture of the virtual person is aligned with a head posture of the user, and that a mouth shape of the virtual person is synchronized with a mouth shape of the user; The facial image features of the virtual person are repaired to generate a dubbing image of the virtual person; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted from the virtual person's mouth when the virtual person's head posture is the same as the user's head posture, the virtual person's mouth shape is the same as the user's mouth shape posture, and / or the virtual person's facial expression is the same as the user's facial expression.
2. The method for generating a virtual human dubbing image according to claim 1, wherein: The performing spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of a virtual person includes: Extracting audio features of the driving audio, extracting source features of the source image using a first feature encoder, and extracting reference features of the reference image using a second feature encoder; Aligning the source feature with the reference feature to obtain an aligned feature; A spatial deformation process is performed based on the alignment features, the audio features, and the reference features to generate the facial image features of the virtual person.
3. The method for generating a virtual human dubbing image according to claim 2, wherein: The performing of spatial deformation processing based on the alignment feature, the audio feature, and the reference feature to generate the facial image feature of the virtual person includes: performing splicing processing on the alignment feature and the audio feature to obtain a head posture audio splicing feature; Using an adaptive channel spatial deformation operator and the head posture audio splicing feature, the reference feature is spatially deformed to generate a deformation feature; The deformed features and the source features are spliced together to generate the facial image features of the virtual person.
4. The method for generating a virtual human dubbing image according to claim 1, wherein: The repairing process of the facial image features of the virtual person to generate a dubbing image of the virtual person includes: The facial image features are input into a feature decoder for decoding and repair processing to generate the dubbing image of the virtual person.
5. The method for generating a virtual human dubbing image according to claim 1, wherein: After the step of repairing the facial image features of the virtual person to generate a dubbing image of the virtual person, the method includes: Calculating a perceptual loss and an adversarial network loss for generating the dubbing image, and calculating a lip synchronization loss; Based on the perceptual loss, the adversarial network loss, and the lip synchronization loss, a total loss for generating the facial image features of the virtual person is determined.
6. The method for generating a virtual human dubbing image according to claim 5, wherein: The determining, based on the perceptual loss, the adversarial network loss, and the lip synchronization loss, of a total loss for generating the facial image features of the virtual person comprises: Obtaining preset perceptual loss weights and obtaining preset lip-sync loss weights; performing a first calculation process on the perception loss and the perception loss weight to obtain a first perception loss, and performing a second calculation process on the lip synchronization loss and the lip synchronization loss weight to obtain a second lip synchronization loss; The first perceptual loss, the second lip synchronization loss, and the adversarial network loss are summed to obtain a total loss for generating the facial image features of the virtual person.
7. A device for generating a virtual human dubbing image, characterized in that: include: an acquisition module, configured to acquire a source image, acquire a driving audio, and acquire a reference image; wherein the source image is represented by an image describing a user's head posture, the driving audio is represented by an audio describing a user's voice emitted from the user's mouth, and the reference image is represented by an image describing the user's lip posture and facial expression; a deformation module, configured to perform spatial deformation processing on the source image, the driving audio, and the reference image to generate facial image features of a virtual person; wherein the facial image features indicate that the head posture of the virtual person is aligned with the head posture of the user, and the lip shape of the virtual person is synchronized with the lip shape of the user; A repair module is used to repair the facial image features of the virtual person to generate a dubbing image of the virtual person; wherein the dubbing image represents an image of an audio having the same content as the audio of the user's voice emitted from the virtual person's mouth when the head posture of the virtual person is the same as the head posture of the user, the mouth shape of the virtual person is the same as the mouth shape posture of the user, and / or the facial expression of the virtual person is the same as the facial expression of the user.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the method for generating a virtual human dubbing image according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating a virtual human dubbing image according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a virtual human dubbing image according to any one of claims 1 to 6 is implemented.