Wav2lip model training method, image frame generation method, electronic device, and storage medium

CN115713579BActive Publication Date: 2026-09-25KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211326787.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2026-09-25
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

(1)、该模型无法满足高质量的说话人视频生成

Benefits of technology

[0023]一种计算机程序产品,包括计算机指令,所述计算机指令在被处理器执行时实施如上任一项所述的Wav2Lip模型的训练方法或图像帧生成方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115713579B_ABST
    Figure CN115713579B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a Wav2Lip model training method, an image frame generation method, an electronic device and a storage medium. The method comprises the following steps: determining a training sample, the training sample comprising an original image frame, a real image frame and an audio file, the original image frame containing a face of a speaker, and the real image frame containing real lip shapes of the speaker expressing the audio file; performing a training process of a Wav2Lip model based on the training sample, the training process comprising: the Wav2Lip model outputting a generated image frame based on the original image frame and the audio file; inputting the generated image frame and the real image frame into a multi-scale image quality discriminator to determine whether the generated image frame and the real image frame are real images at multiple scales by the image quality discriminator; determining a loss function value of the Wav2Lip model based on the determination result; and configuring model parameters of the Wav2Lip model so that the loss function value is lower than a preset threshold. The embodiment of the application can improve the image quality and the training stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and more specifically, to a Wav2Lip model training method, an image frame generation method, an electronic device, and a storage medium. Background Technology

[0002] Generate speaker image frames (such as videos) based on given audio and a given speaker image, and generate speaker video generation where the speaker's lip movements in the image frames correspond to the audio content. This is called speaker video generation and can be applied to scenarios such as digital virtual humans, game and anime character dubbing, and lip-sync speech translation.

[0003] The Wav2Lip model is a lip-sync algorithm based on Generative Adversarial Networks (GANs) to synchronize lip movements with speech in videos. The Wav2Lip model can not only output lip-synced videos that match the target speech based on static images, but also directly perform lip-sync transformation on dynamic videos to output videos that match the input speech. The Wav2lip model incorporates a pre-trained lip-sync expert model during the training phase, enabling it to achieve highly accurate lip-sync results on any speech.

[0004] In practice, the applicant discovered that the Wav2lip model has at least the following two problems: (1) The model cannot meet the requirements for generating high-quality speaker videos. The model performs reasonably well when the input image resolution is low, but when faced with generating higher resolution images, the overall image is blurry, especially the lower half of the face.

[0005] (2) The model is not stable enough during the training phase, and its performance varies greatly depending on the dataset. Because the model contains multiple loss functions, it is crucial to ensure the quality of image generation while synchronizing audio and lip movements. The weight design of each part is critical. In practical applications, the model's performance is very unstable. Summary of the Invention

[0006] The present invention provides a Wav2Lip model training method, an image frame generation method, an electronic device, and a storage medium.

[0007] The technical solution of the embodiments of the present invention is as follows: A training method for a Wav2Lip model, the method comprising: Determine training samples, which include original image frames, real image frames, and audio files. The original image frames contain the speaker's face, and the real image frames contain the speaker's actual lip shape when describing the audio file. Based on the training samples, the training process of the Wav2Lip model is performed. The training process includes: inputting the original image frames and the audio file into the Wav2Lip model, so that the Wav2Lip model outputs generated image frames based on the original image frames and the audio file; inputting the generated image frames and the real image frames into a multi-scale image quality discriminator, so that the multi-scale image quality discriminator determines whether the generated image frames and the real image frames are real images at multiple scales. Based on the discrimination results of the multi-scale image quality discriminator, the loss function value of the Wav2Lip model is determined; Configure the model parameters of the Wav2Lip model so that the loss function value is lower than a preset threshold.

[0008] In an exemplary embodiment, determining whether the generated image frame and the real image frame are real images at multiple scales includes: Perform (n-1) average pooling operations on the generated image frame to obtain (n-1) generated image frames after average pooling operations, where n is a positive integer of at least 2; Perform (n-1) average pooling operations on the real image frame to obtain a real image frame after (n-1) average pooling operations; The multi-scale image quality discriminator includes n sub-discriminators. The input images of the first sub-discriminator are the real image frame and the generated image frame. The input images of the remaining (n-1) sub-discriminators are respectively the generated image frame obtained after the (n-1) average pooling operations and the real image frame obtained after the (n-1) average pooling operations.

[0009] In the exemplary implementation, k is a positive integer greater than or equal to 1 and less than or equal to n; The k-th sub-discriminator includes: The first convolutional layer is used to extract features from the input image of the k-th sub-discriminator; A downsampling layer is used to perform downsampling on the features; The second convolutional layer is used to output a discrimination result based on the downsampled features.

[0010] In the exemplary implementation, it also includes: Determine the loss function value of the multi-scale image quality discriminator. ; Based on the above Update the model parameters of the multi-scale image quality discriminator; wherein: ; The discrimination result of the real image frame input to the k-th sub-discriminator; The discrimination result of the generated image frame input to the k-th sub-discriminator.

[0011] In an exemplary implementation, determining the loss function value of the Wav2Lip model based on the discrimination result of the multi-scale image quality discriminator includes: Determine the loss function value of the Wav2Lip model. ;in: ; in , , and These are the preset coefficients; For reconstruction loss function; The lip alignment loss function; For adversarial loss function; This is the feature matching loss function.

[0012] In the exemplary implementation, it also includes: Sure ,in: ; The total number of layers in each sub-discriminator is T; i is the layer number. The discrimination result of the i-th layer of the k-th sub-discriminator for the real image frame input to the k-th sub-discriminator; This represents the discrimination result of the i-th layer of the k-th sub-discriminator for the generated image frame input to the k-th sub-discriminator.

[0013] An image frame generation method, comprising: Identify the audio test file and the first image frame containing the speaker's face; The audio test file and the first image frame are input into the Wav2Lip model so that the Wav2Lip model generates a second image frame with lip movements synchronized with the audio test file based on the audio test file and the first image frame, wherein the Wav2Lip model is trained according to the training method of the Wav2Lip model described in any of the above embodiments. The second image frame is received from the Wav2Lip model.

[0014] A training device for a Wav2Lip model, comprising: The first determining module is used to determine training samples, the training samples including original image frames, real image frames and audio files, the original image frames containing the speaker's face, and the real image frames containing the speaker's real lip shape when describing the audio file. The training module is used to execute the training process of the Wav2Lip model based on the training samples. The training process includes: inputting the original image frames and the audio file into the Wav2Lip model, so that the Wav2Lip model outputs generated image frames based on the original image frames and the audio file; inputting the generated image frames and the real image frames into a multi-scale image quality discriminator, so that the multi-scale image quality discriminator determines whether the generated image frames and the real image frames are real images at multiple scales. The second determining module is used to determine the loss function value of the Wav2Lip model based on the discrimination result of the multi-scale image quality discriminator. The configuration module is used to configure the model parameters of the Wav2Lip model so that the loss function value is lower than a preset threshold.

[0015] In an exemplary embodiment, the step of determining whether the generated image frame and the real image frame are real images at multiple scales includes: performing (n-1) average pooling operations on the generated image frame to obtain a generated image frame after (n-1) average pooling operations, where n is a positive integer of at least 2; performing (n-1) average pooling operations on the real image frame to obtain a real image frame after (n-1) average pooling operations; wherein the multi-scale image quality discriminator includes n sub-discriminators, the input image of the first sub-discriminator of the n sub-discriminators is: the real image frame and the generated image frame; the input images of the remaining (n-1) sub-discriminators of the n sub-discriminators are respectively: the generated image frame after (n-1) average pooling operations obtained according to the order of average pooling operations, and the real image frame after (n-1) average pooling operations obtained according to the order of average pooling operations.

[0016] In an exemplary implementation, k is a positive integer greater than or equal to 1 and less than or equal to n; the kth sub-discriminator includes: a first convolutional layer for extracting features from the input image of the kth sub-discriminator; a downsampling layer for performing downsampling on the features; and a second convolutional layer for outputting a discrimination result based on the downsampled features.

[0017] In an exemplary embodiment, the second determining module is further configured to determine the loss function value of the multi-scale image quality discriminator. Based on the above Update the model parameters of the multi-scale image quality discriminator; wherein: ; The discrimination result of the real image frame input to the k-th sub-discriminator; The discrimination result of the generated image frame input to the k-th sub-discriminator.

[0018] In an exemplary implementation, the second determining module is used to determine the loss function value of the Wav2Lip model. ;in: ;in , , and These are the preset coefficients; For reconstruction loss function; The lip alignment loss function; For adversarial loss function; This is the feature matching loss function.

[0019] In an exemplary embodiment, the second determining module is used to determine ,in: The total number of layers in each sub-discriminator is T; i is the layer number. The discrimination result of the i-th layer of the k-th sub-discriminator for the real image frame input to the k-th sub-discriminator; This represents the discrimination result of the i-th layer of the k-th sub-discriminator for the generated image frame input to the k-th sub-discriminator.

[0020] An image frame generation apparatus, comprising: The determination module is used to determine the audio test file and the first image frame containing the speaker's face; An input module is used to input the audio test file and the first image frame into the Wav2Lip model, so that the Wav2Lip model generates a second image frame with lip movements synchronized with the audio test file based on the audio test file and the first image frame, wherein the Wav2Lip model is trained according to the training method of the Wav2Lip model as described above. A receiving module is used to receive the second image frame from the Wav2Lip model.

[0021] An electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the training method of the Wav2Lip model or the image frame generation method as described in any of the preceding claims.

[0022] A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, can implement the training method of the Wav2Lip model or the image frame generation method as described in any of the preceding claims.

[0023] A computer program product includes computer instructions that, when executed by a processor, implement a training method for the Wav2Lip model or an image frame generation method as described in any of the preceding claims.

[0024] As can be seen from the above technical solution, in the embodiments of the present invention, training samples are determined, including original image frames, real image frames, and audio files. The original image frames contain the speaker's face, and the real image frames contain the speaker's actual lip shape as described in the audio file. The training process of the Wav2Lip model is executed based on the training samples. The training process includes: the Wav2Lip model outputs generated image frames based on the original image frames and audio files; the generated image frames and real image frames are input into a multi-scale image quality discriminator, which determines whether the generated image frames and real image frames are real images at multiple scales; based on the discrimination results, the loss function value of the Wav2Lip model is determined; and the model parameters of the Wav2Lip model are configured so that the loss function value is lower than a preset threshold. Therefore, the embodiments of the present invention utilize a multi-scale image quality discriminator to discriminate images at different scales, thereby improving the discriminator's capabilities, providing stronger supervision for generating high-quality images, and improving image quality.

[0025] Furthermore, considering the feature matching loss introduced by the multi-scale discriminator, the embodiments of the present invention further set the feature matching loss, thereby improving training stability. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a demonstrative structural diagram of the Wav2Lip model.

[0028] Figure 2 This is an exemplary flowchart of the Wav2Lip model training method according to an embodiment of the present invention.

[0029] Figure 3 This is an exemplary structural diagram of a multi-scale image quality discriminator according to an embodiment of the present invention.

[0030] Figure 4 This is an exemplary structural diagram of the sub-discriminator according to an embodiment of the present invention.

[0031] Figure 5 This is an exemplary flowchart of the image frame generation method according to an embodiment of the present invention.

[0032] Figure 6 This is an exemplary structural diagram of the Wav2Lip model training device according to an embodiment of the present invention.

[0033] Figure 7 This is an exemplary structural diagram of the image frame generation apparatus according to an embodiment of the present invention.

[0034] Figure 8 This is an exemplary structural diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.

[0036] For the sake of brevity and intuitiveness, the present invention will be described below through several representative embodiments. Numerous details in the embodiments are provided only to aid in understanding the present invention. However, it is obvious that the implementation of the present invention may not be limited to these details. To avoid unnecessarily obscuring the present invention, some embodiments are not described in detail, but only outlines are given. In the following text, "comprising" means "including but not limited to," and "according to..." means "at least according to..., but not limited to only according to...". Due to Chinese language habits, unless the quantity of a component is specifically indicated below, it means that the component can be one or more, or can be understood as at least one. The terms "first," "second," "third," "fourth," etc. (if present), in the specification, claims, and accompanying drawings of the embodiments of the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented, for example, in orders other than those illustrated or described herein.

[0037] Figure 1 This is a demonstrative architecture diagram of the Wav2Lip model. The Wav2Lip model includes a generator and two discriminators: a lip-sync discriminator and a visual quality discriminator.

[0038] To generate lip-sync images that are precisely synchronized with the audio, the model input consists of two parts: (1) the original image frames (generally including the speaker's face), such as a video frame sequence; and (2) the audio file, generally a Melspectrogram segment. These two parts of data are input into the generator according to a specific organization format. The generator includes an image encoder, an audio encoder, a fusion unit, and an image decoder. The image encoder in the generator receives the original image frames and extracts the intermediate features of the original image frames. The audio encoder in the generator receives the audio signal and extracts the intermediate features of the audio. The fusion unit performs feature fusion (concat) processing on the intermediate features of the original image frames and the intermediate features of the audio, and the fused features are sent to the image decoder for decoding. The image decoder outputs lip-sync images synchronized with the audio (GeneratedFrames), which are called generated image frames.

[0039] During the training phase of the Wav2Lip model, the generator generates lip-synced image frames based on the original image frames and audio files in the training data, and sends these generated image frames to two discriminators. These two discriminators include: (1) The pre-trained lip-sync discriminator takes the audio file and the generated image frame as input to determine whether the generated image frame and the audio file are synchronized. The lip-sync discriminator needs to be pre-trained in advance to enhance the ability to distinguish between lip-sync and audio synchronization. (2) Image quality discriminator, which receives the generated image frame output by the generator and the speaker's real lip image (called real image frame) synchronized with the audio file in the training data, to judge its authenticity and drive better lip generation.

[0040] During the training phase, the generator's input consists of two parts: the original image frames and the audio files from the training data. The image encoder and audio encoder respectively obtain their respective feature information, which is then fused in a fusion unit. The image decoder then generates lip-synced image frames (i.e., the generated image frames from the training phase). Real image frames (usually included in the training data) and generated image frames from the training phase are input to an image quality discriminator. The image quality discriminator performs binary classification, determining whether the image (real or generated) is a real image or a generated image, thus improving image quality. Furthermore, the generated image frames from the training phase and the audio files from the training data are input to a pre-trained lip-sync discriminator to determine lip-sync accuracy. During training, the model parameters within the lip-sync discriminator are frozen and do not participate in training or updates.

[0041] During the inference phase, given an audio and video clip (e.g., an image or animation), the generator can output lip-synced generated image frames (e.g., a video).

[0042] The loss function of existing generators mainly consists of three parts: L1 reconstruction loss, lip-sync loss (Lsync), and adversarial loss (Ladv). L1 reconstruction loss stems from the loss in reconstructing the generated image frame based on the original image frame; lip-sync loss originates from the loss of the lip-sync discriminator; and adversarial loss arises from adversarial interaction with the image quality discriminator, aiming to make the generated image deceive the discriminator as much as possible. It is evident that the image quality discriminator plays a crucial supervisory role in image quality. Currently, existing image quality discriminators typically consist of multiple convolutional blocks (NonNormConv2d), each block including a convolutional layer and a LeakyReLU activation layer, used to improve visual quality and synchronization accuracy.

[0043] However, existing image quality discriminators take a single-scale image as input, limiting their capabilities and making it impossible to discriminate images at different scales. This results in the Wav2lip model failing to meet the requirements for generating high-quality speaker videos. In particular, while the Wav2lip model performs reasonably well when the input image resolution is low, it becomes blurry when generating higher-resolution images, especially the lower half of the face.

[0044] Figure 2 This is an exemplary flowchart of the Wav2Lip model training method according to an embodiment of the present invention.

[0045] like Figure 2 As shown, the method includes: Step 201: Determine the training samples. The training samples include original image frames, real image frames, and audio files. The original image frames contain the speaker's face, and the real image frames contain the speaker's actual lip shape in the audio file.

[0046] The following example illustrates the data implementation process of the Wav2Lip model's training samples. In data preparation, a Lip Reading Sentences 2 (RS2) dataset is obtained, where each sentence is no longer than 100 characters. Then, a continuous real-world video sequence of faces is acquired, retaining only the upper half of the face (the lower half is occluded). A random reference frame is also selected (generated by randomly moving the video). The two image sequences are then fused together along the channel dimension to form the original image frame. Additionally, real-world image frames containing the speaker's lip movements from an audio file (e.g., spoken or sung) are obtained. The audio file and the original image frames are then input into the generator. The generator generates lip-shape image frames synchronized with the audio file, i.e., generating the occluded lower half of the face, where the lip movements are already synchronized with the audio.

[0047] Step 202: Based on the training samples, perform the training process of the Wav2Lip model. The training process includes: inputting the original image frames and audio files into the Wav2Lip model so that the Wav2Lip model can output image frames based on the original image frames and audio files; inputting the generated image frames and real image frames into a multi-scale image quality discriminator so that the multi-scale image quality discriminator can determine whether the generated image frames and real image frames are real images at multiple scales.

[0048] In this embodiment of the invention, a multi-scale image quality discriminator is introduced into the Wav2lip model to improve the quality of the generated image frames. Image size is generally represented by the total number of pixels along the width and height of the image. Considering that characteristics that are not easily seen or extracted at one scale may be easily discovered or extracted at another scale, it is preferable to use a multi-scale representation for the image and process them separately at different scales. For example, a single scale may specifically include: 256 (pixels) * 256 (pixels), 128 (pixels) * 128 (pixels), or 64 (pixels) * 64 (pixels), etc.

[0049] In one implementation, determining whether a generated image frame and a real image frame are real images at multiple scales includes: performing (n-1) average pooling operations on the generated image frame to obtain a generated image frame after (n-1) average pooling operations, where n is a positive integer of at least 2; performing (n-1) average pooling operations on the real image frame to obtain a real image frame after (n-1) average pooling operations; wherein the multi-scale image quality discriminator includes n sub-discriminators, the input images of the first sub-discriminator among the n sub-discriminators are: a real image frame and a generated image frame; the input images of the remaining (n-1) sub-discriminators among the n sub-discriminators are respectively: a generated image frame after (n-1) average pooling operations obtained according to the order of average pooling operations, and a real image frame after (n-1) average pooling operations obtained according to the order of average pooling operations.

[0050] As can be seen, the embodiments of the present invention no longer utilize a single-scale image quality discriminator, but introduce a multi-scale image quality discriminator into the Wav2lip model to discriminate images at different scales. The discriminator's capabilities are improved, and it has a stronger supervisory role in generating high-quality images, thus improving image quality.

[0051] Figure 3 This is an exemplary structural diagram of a multi-scale image quality discriminator according to an embodiment of the present invention. Figure 3 In this paper, we will take a multi-scale image quality discriminator containing three sub-discriminators as an example for description.

[0052] like Figure 3 As shown, the multi-scale image quality discriminator includes a first sub-discriminator, a second sub-discriminator, and a third sub-discriminator, each corresponding to its respective image scale. Each sub-discriminator makes a decision on the input image to determine whether it is a real image (e.g., outputting a Boolean value of true) or a generated image (e.g., outputting a Boolean value of false). Specifically, an image with a scale of 256 pixels * 256 pixels is the input image of the first sub-discriminator. A first average pooling process is performed on the 256 pixel * 256 pixel image to reduce the image scale, for example, to obtain a 128 pixel * 128 pixel image, which is the input image of the second sub-discriminator. A second average pooling process is performed on the 128 pixel * 128 pixel image to further reduce the image scale, for example, to obtain a 64 pixel * 64 pixel image, which is the input image of the third sub-discriminator. During training, the input image frames include real image frames and generated image frames from the training process. During the inference phase, the input image frames include generated image frames from the inference phase.

[0053] exist Figure 3In this paper, an image quality discriminator comprising three sub-discriminators is used as an example for illustration. Those skilled in the art will recognize that the number of sub-discriminators in the image quality discriminator can be two or more than three, and the embodiments of the present invention are not limited in this respect.

[0054] Figure 4 This is an exemplary structural diagram of a sub-discriminator according to an embodiment of the present invention. This sub-discriminator is applicable to the k-th sub-discriminator, where k is a positive integer greater than or equal to 1 and less than or equal to n; and n is a positive integer greater than or equal to 2.

[0055] like Figure 4 As shown, the k-th sub-discriminator includes: a first convolutional layer for extracting features from the input image of the k-th sub-discriminator; a downsampling layer for performing downsampling on the features; and a second convolutional layer for outputting a discrimination result based on the downsampled features.

[0056] Figure 4 This example uses a 3-layer sub-discriminator. In practice, the 3-layer sub-discriminator can have more layers. For example, combining 3... Figure 4 By connecting the sub-discriminators shown in the diagram in series, a 9-layer sub-discriminator structure can be formed; by connecting 4 such sub-discriminators... Figure 4 The sub-discriminators shown can be connected in series to form a 12-layer sub-discriminator, and so on.

[0057] As can be seen, the embodiments of the present invention also propose an optimized structure for the sub-discriminator, which facilitates the implementation of a multi-scale image quality discriminator. Furthermore, the number of layers in the sub-discriminator of the embodiments of the present invention is easily expandable.

[0058] Step 203: Determine the loss function value of the Wav2Lip model based on the discrimination results of the multi-scale image quality discriminator.

[0059] Step 204: Configure the model parameters of the Wav2Lip model so that the loss function value is lower than the preset threshold.

[0060] In one implementation, the method further includes: determining the loss function value of the multi-scale image quality discriminator. ;based on Update the model parameters of the multi-scale image quality discriminator; where: ; The discrimination result of the real image frame input to the k-th sub-discriminator; The discrimination result of the generated image frame input to the k-th sub-discriminator.

[0061] Therefore, by setting a loss function for the multi-scale image quality discriminator, the model parameters of the Wav2Lip model can be updated during training, thereby improving image quality.

[0062] In one implementation, determining the loss function value of the Wav2Lip model based on the discrimination results of a multi-scale image quality discriminator includes: determining the loss function value of the Wav2Lip model. ;in: ;in , , and These are the preset coefficients; For reconstruction loss function; The lip alignment loss function; For adversarial loss function; This is the feature matching loss function. , and The calculation method refers to common processing methods in this field, and will not be repeated here.

[0063] It is evident that by taking into account the feature matching loss introduced by the multi-scale image quality discriminator, further adjustments to the feature matching loss can improve training stability.

[0064] In one implementation, it further includes: determining ,in: The total number of layers in each sub-discriminator is T; i is the layer number. The discrimination result of the i-th layer of the k-th sub-discriminator for the real image frame input to the k-th sub-discriminator; This represents the discrimination result of the i-th layer of the k-th sub-discriminator for the generated image frame input to the k-th sub-discriminator.

[0065] For example: when a multi-scale image quality discriminator has, for example... Figure 3 The structure shown is such that the first sub-discriminator (i.e., k=1) has the following characteristics: Figure 4 When the structure shown is: During training, the input image for the first sub-discriminator includes: the real image frame (scale 256*256) input to the first sub-discriminator and the generated image frame (scale 256*256).

[0066] For the first layer (i.e., the first convolutional layer) of the first sub-discriminator. The image features extracted from the generated 256*256 image frames are: These are image features extracted from real image frames of 256*256.

[0067] For the second layer (i.e., the downsampling layer) of the first sub-discriminator. The result is the downsampling of features extracted from the generated 256*256 image frame. The result is the downsampling of features extracted from a real 256*256 image frame.

[0068] For the third layer (i.e. the second convolutional layer) of the first sub-discriminator. The result is: the downsampling result of the generated image frame based on 256*256, and the discrimination result of the generated image frame based on 256*256; The result is: the discrimination result of a 256*256 real image frame based on the downsampling result of the real image frame.

[0069] During training, the input images for the second sub-discriminator include: real image frames (scale 128*128) and generated image frames (scale 128*128) input to the second sub-discriminator.

[0070] For the first layer (i.e., the first convolutional layer) of the second sub-discriminator. The image features extracted from the generated 128*128 image frames are: These are image features extracted from a real 128*128 image frame.

[0071] For the second layer (i.e., the downsampling layer) of the second sub-discriminator. The result is the downsampling of features extracted from the generated 128*128 image frame. The result is the downsampling of features extracted from a 128*128 real image frame. For the third layer (i.e., the second convolutional layer) of the second sub-discriminator. The result is: the downsampling result of the generated image frame based on 128*128, and the discrimination result of the generated image frame based on 128*128; The result is: the discrimination result of the 128*128 real image frame based on the downsampling result of the real image frame.

[0072] Similarly, during training, the input images for the third sub-discriminator include: real image frames (64*64 pixels) and generated image frames (64*64 pixels). The output of each layer of the third sub-discriminator can be determined in a similar manner. Likewise, the output of each layer of the other sub-discriminators can be determined, which will not be elaborated here.

[0073] Therefore, the embodiments of the present invention also realize a fast calculation method for feature matching loss.

[0074] Based on the Wav2Lip model obtained by the above training method, the embodiments of the present invention also propose an image frame generation method.

[0075] Figure 5 This is an exemplary flowchart of an image frame generation method according to an embodiment of the present invention. Figure 5 As shown, the method includes: Step 501: Determine the audio test file and the first image frame containing the speaker's face.

[0076] Step 502: Input the audio test file and the first image frame into the Wav2Lip model, so that the Wav2Lip model generates a second image frame with lip movements synchronized with the audio test file based on the audio test file and the first image frame, wherein the Wav2Lip model according to... Figure 2 The training method shown for the Wav2Lip model was used for training.

[0077] Here, as Figure 2 The input to the Wav2Lip model obtained by the training method shown consists of two parts: (1) a first image frame, such as a video frame sequence; and (2) an audio test file, typically a Mel spectrum segment. These two parts of data are input into the generator according to a specific organization format. The generator includes an image encoder, an audio encoder, a fusion unit, and an image decoder. The image encoder in the generator receives the first image frame and extracts the intermediate features of the first image frame. The audio encoder in the generator receives the audio test file and extracts the intermediate audio features. The fusion unit performs feature fusion processing on the intermediate features of the original image frame and the intermediate audio features, and the fused features are sent to the image decoder for decoding. The image decoder outputs an image frame with lip-syncing and audio synchronization, called the second image frame.

[0078] Step 503: Receive the second image frame from the Wav2Lip model.

[0079] The present invention also proposes a Wav2Lip model training device. Figure 6 This is an exemplary structural diagram of the Wav2Lip model training device according to an embodiment of the present invention. Figure 6 As shown, the Wav2Lip model training device 600 includes: The first determining module 601 is used to determine training samples, which include original image frames, real image frames, and audio files. The original image frames contain the speaker's face, and the real image frames contain the speaker's actual lip shape as described in the audio file. The training module 602 is used to perform the training process of the Wav2Lip model based on the training samples. The training process includes: inputting the original image frames and audio files into the Wav2Lip model so that the Wav2Lip model outputs generated image frames based on the original image frames and audio files; inputting the generated image frames and real image frames into a multi-scale image quality discriminator so that the multi-scale image quality discriminator determines whether the generated image frames and real image frames are real images at multiple scales. The second determining module 603 is used to determine the loss function value of the Wav2Lip model based on the discrimination result of the multi-scale image quality discriminator. The configuration module 604 is used to configure the model parameters of the Wav2Lip model so that the loss function value is lower than a preset threshold.

[0080] In one implementation, determining whether a generated image frame and a real image frame are real images at multiple scales includes: performing (n-1) average pooling operations on the generated image frame to obtain a generated image frame after (n-1) average pooling operations, where n is a positive integer of at least 2; performing (n-1) average pooling operations on the real image frame to obtain a real image frame after (n-1) average pooling operations; wherein the multi-scale image quality discriminator includes n sub-discriminators, the input images of the first sub-discriminator among the n sub-discriminators are: a real image frame and a generated image frame; the input images of the remaining (n-1) sub-discriminators among the n sub-discriminators are respectively: a generated image frame obtained after (n-1) average pooling operations according to the order of average pooling operations, and a real image frame obtained after (n-1) average pooling operations according to the order of average pooling operations.

[0081] In one implementation, k is a positive integer greater than or equal to 1 and less than or equal to n; the kth sub-discriminator includes: a first convolutional layer for extracting features from the input image of the kth sub-discriminator; a downsampling layer for performing downsampling on the features; and a second convolutional layer for outputting a discrimination result based on the downsampled features.

[0082] In one embodiment, the second determining module 603 is further configured to determine the loss function value of the multi-scale image quality discriminator. ;based on Update the model parameters of the multi-scale image quality discriminator; where: ; The discrimination result of the real image frame input to the k-th sub-discriminator; The discrimination result of the generated image frame input to the k-th sub-discriminator.

[0083] In one implementation, the second determining module 603 is used to determine the loss function value of the Wav2Lip model. ;in: ;in , , and These are the preset coefficients; For reconstruction loss function; The lip alignment loss function; For adversarial loss function; This is the feature matching loss function.

[0084] In one embodiment, the second determining module 603 is used to determine ,in: The total number of layers in each sub-discriminator is T; i is the layer number. The discrimination result of the i-th layer of the k-th sub-discriminator for the real image frame input to the k-th sub-discriminator; This represents the discrimination result of the i-th layer of the k-th sub-discriminator for the generated image frame input to the k-th sub-discriminator.

[0085] Figure 7 This is an exemplary structural diagram of an image frame generation apparatus according to an embodiment of the present invention. For example... Figure 7 As shown, the image frame generation apparatus 700 includes: a determining module 701, used to determine an audio test file and a first image frame containing a speaker's face; an input module 702, used to input the audio test file and the first image frame into a Wav2Lip model, so that the Wav2Lip model generates a second image frame with synchronized lip movements with the audio test file based on the audio test file and the first image frame, wherein the Wav2Lip model is trained according to the training method of the Wav2Lip model described above; and a receiving module 703, used to receive the second image frame from the Wav2Lip model.

[0086] This invention also provides a computer-readable medium storing instructions that, when executed by a processor, can perform steps in the above-described Wav2Lip model training method or image frame generation method. In practical applications, the computer-readable medium may be included in the device / apparatus / system described in the above embodiments, or it may exist independently and not assembled into that device / apparatus / system. The aforementioned computer-readable storage medium carries one or more programs, which, when executed, can implement the Wav2Lip model training method or image frame generation method described in the above embodiments. According to the embodiments disclosed in this invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof, but this is not intended to limit the scope of protection of this invention. In the embodiments disclosed in this invention, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0087] like Figure 8 As shown, embodiments of the present invention also provide an electronic device in which a training device or an image frame generation device for the Wav2Lip model of the present invention can be integrated. For example... Figure 8 The diagram illustrates an exemplary structural diagram of an electronic device according to an embodiment of the present invention. Specifically, the electronic device may include a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, and a computer program stored in the memory and executable on the processor. When the program in the memory 802 is executed, the training method of the Wav2Lip model or the image frame generation method described above can be implemented.

[0088] In practical applications, this electronic device may also include components such as a power supply 803, an input unit 804, and an output unit 805. Those skilled in the art will understand that... Figure 8The structure of the electronic device shown does not constitute a limitation on the electronic device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements. The processor 801 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and lines. It performs various server functions and processes data by running or executing software programs and / or modules stored in memory 802, and by calling data stored in memory 802, thereby providing overall monitoring of the electronic device. Memory 802 can be used to store software programs and modules, i.e., the aforementioned computer-readable storage medium. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in memory 802. Memory 802 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function, etc.; the data storage area may store data created according to the use of the server, etc. Furthermore, memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory 802 may also include a memory controller to provide processor 801 with access to memory 802.

[0089] The electronic device also includes a power supply 803 that supplies power to various components. This power supply is logically connected to the processor 801 via a power management system, enabling functions such as charging, discharging, and power consumption management. The power supply 803 may also include one or more DC or AC power supplies, a recharging system, a power fault detection circuit, a power converter or inverter, a power status indicator, or other arbitrary components. The electronic device may also include an input unit 804, which can receive input digital or character information and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. The electronic device may also include an output unit 805, which can display information input by the user or information provided to the user, as well as various graphical user interfaces (GUIs), which can be composed of graphics, text, icons, video, and any combination thereof.

[0090] This invention also provides a computer program product comprising computer instructions that, when executed by a processor, implement the training method for the Wav2Lip model or the image frame generation method as described in any of the above embodiments. The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments disclosed in this invention. In this regard, each block in a flowchart or block diagram may represent a module, program segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in the order specified in the different figures. For example, two blocks shown connected together may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0091] This document describes the principles and implementation methods of the present invention using specific embodiments. These embodiments are merely illustrative of the method and core concepts of the present invention and are not intended to limit the invention. Those skilled in the art can make changes to the specific implementation methods and application scope based on the ideas, spirit, and principles of the present invention. Any modifications, equivalent substitutions, or improvements made should be included within the scope of protection of this invention.

Claims

1. A training method for a Wav2Lip model, characterized in that, The method includes: Determine training samples, which include original image frames, real image frames, and audio files. The original image frames contain the speaker's face, and the real image frames contain the speaker's actual lip shape when describing the audio file. Based on the training samples, the training process of the Wav2Lip model is performed. The training process includes: inputting the original image frames and the audio file into the Wav2Lip model, so that the Wav2Lip model outputs generated image frames based on the original image frames and the audio file; inputting the generated image frames and the real image frames into a multi-scale image quality discriminator, so that the multi-scale image quality discriminator determines whether the generated image frames and the real image frames are real images at multiple scales. Based on the discrimination results of the multi-scale image quality discriminator, the loss function value of the Wav2Lip model is determined; Configure the model parameters of the Wav2Lip model so that the loss function value is lower than a preset threshold; The step of determining whether the generated image frame and the real image frame are real images at multiple scales includes: Perform (n-1) average pooling operations on the generated image frame to obtain (n-1) generated image frames after average pooling operations, where n is a positive integer of at least 2; Perform (n-1) average pooling operations on the real image frame to obtain a real image frame after (n-1) average pooling operations; The multi-scale image quality discriminator includes n sub-discriminators. The input images of the first sub-discriminator are the real image frame and the generated image frame. The input images of the remaining (n-1) sub-discriminators are respectively the generated image frame obtained after (n-1) average pooling operations and the real image frame obtained after (n-1) average pooling operations. The first sub-discriminator takes as input real image frames and generated image frames at the original scale, while the other sub-discriminators take as input real image frames and generated image frames that have been scaled down by average pooling.

2. The method according to claim 1, characterized in that, k is a positive integer greater than or equal to 1 and less than or equal to n; The k-th sub-discriminator includes: The first convolutional layer is used to extract features from the input image of the k-th sub-discriminator; A downsampling layer is used to perform downsampling on the features; The second convolutional layer is used to output a discrimination result based on the downsampled features.

3. The method according to claim 1 or 2, characterized in that, Also includes: Determine the loss function value of the multi-scale image quality discriminator. ; Based on the above Update the model parameters of the multi-scale image quality discriminator; wherein: ; The discrimination result of the real image frame input to the k-th sub-discriminator; The discrimination result of the generated image frame input to the k-th sub-discriminator.

4. The method according to claim 1 or 2, characterized in that, The loss function value of the Wav2Lip model is determined based on the discrimination result of the multi-scale image quality discriminator, including: Determine the loss function value of the Wav2Lip model. ;in: ; in , , and These are the preset coefficients; For reconstruction loss function; The lip alignment loss function; For adversarial loss function; This is the feature matching loss function.

5. The method according to claim 4, characterized in that, Also includes: Sure ,in: ; The total number of layers in each sub-discriminator is T; i is the layer number. The discrimination result of the i-th layer of the k-th sub-discriminator for the real image frame input to the k-th sub-discriminator; This represents the discrimination result of the i-th layer of the k-th sub-discriminator for the generated image frame input to the k-th sub-discriminator.

6. A method for generating image frames, characterized in that, include: Identify the audio test file and the first image frame containing the speaker's face; The audio test file and the first image frame are input into the Wav2Lip model, so that the Wav2Lip model generates a second image frame with lip movements synchronized with the audio test file based on the audio test file and the first image frame, wherein the Wav2Lip model is trained according to the training method of the Wav2Lip model according to any one of claims 1-5. The second image frame is received from the Wav2Lip model.

7. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the training method of the Wav2Lip model according to any one of claims 1-5 or the image frame generation method according to claim 6.

8. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instructions are executed by the processor, they implement the training method of the Wav2Lip model as described in any one of claims 1-5 or the image frame generation method as described in claim 6.

9. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the training method of the Wav2Lip model as described in any one of claims 1-5 or the image frame generation method as described in claim 6.