High-fidelity and high-synchronization speaking face generation model training method and system

Through difficult example mining, selecting the reference image with the lowest matching degree with the mouth of the pose image, and training it in combination with the generation of adversarial network model and the loss function guided by the target resolution face image, the problem of low synchronization and fidelity of the speaking face generation method in the prior art is solved, and high-fidelity and high-synchronization speaking face generation is achieved.

CN119963703APending Publication Date: 2025-05-09INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510078966.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, the speaking face generation method has the problem that the lip shape of the face is low and the sound synchronization in the audio.

Method used

By obtaining the audio to be driven, the pose image and the reference image candidate set, difficult mining is carried out based on the pose image and the image candidate set, the identity reference image with the lowest matching degree of the mouth of the pose image is obtained, and it is input with the pose image and audio to the generation adversarial network model. The loss function guided by the target resolution face image is supervised and trained to improve the synchronization and fidelity of the model.

Benefits of technology

It realizes the generation of speaking faces with high fidelity and high synchronization, which improves the generation quality and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963703A_ABST
    Figure CN119963703A_ABST
Patent Text Reader

Abstract

The invention provides a high-fidelity and high-synchronization speaking face generation model training method and system, which are applied to the technical field of image processing, and the method comprises the steps: obtaining a to-be-driven audio, a pose image and a reference image candidate set; difficult case mining is carried out based on the pose image and the image candidate set, an identity reference image corresponding to the pose image is obtained, and the mouth matching degree between the identity reference image and the pose image is the lowest; the identity reference image, the pose image and the audio to be driven are input to a speaking face generation model, a generated speaking face image output by the speaking face generation model is obtained, and the speaking face generation model is based on a generative adversarial network model; on the basis of a loss function guided by the target resolution ratio face image, supervising the generated speaking face image model so as to train a speaking face generation model; according to the invention, a speaking face image with fidelity and synchronism can be generated at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a high-fidelity and high-synchronization speaking face generation model training method and system. Background Art

[0002] The talking face generation algorithm is a technology that drives and synthesizes facial images in videos through audio signals, so that the mouth shape and expression of the face are synchronized with the sound in the audio.

[0003] The generation of speaking faces based on generative adversarial networks is generally performed by taking an identity reference image and a pose image as input and combining them with given audio. Existing methods generally use a random selection method when selecting identity reference images. This method will result in a large number of samples with the same mouth shape between the reference image and the pose image during training. The model tends to directly copy the mouth shape, affecting the synchronization of the generated speaking faces.

[0004] It can be seen that the speaking face generation method in the related art has a technical problem that the lip shape of the face and the sound in the audio are not synchronized well. Summary of the invention

[0005] The present invention provides a high-fidelity and high-synchronization speaking face generation model training method and system, which are used to solve the defect of the speaking face generation method in the prior art that the lip shape of the face and the sound in the audio are low in synchronization, and realize the generation of speaking faces with both fidelity and synchronization.

[0006] The present invention provides a high-fidelity and high-synchronization speaking face generation model training method, comprising the following steps. The audio to be driven, the pose image and the reference image candidate set are obtained; based on the pose image and the image candidate set, hard example mining is performed to obtain the identity reference image corresponding to the pose image, wherein the mouth matching degree between the identity reference image and the pose image is the lowest; the identity reference image, the pose image and the audio to be driven are input into the speaking face generation model to obtain the generated speaking face image output by the speaking face generation model, wherein the speaking face generation model is based on a generative adversarial network model; based on the loss function guided by the target resolution face image, the generated speaking face image model is supervised to train the speaking face generation model, wherein the resolution of the target resolution face image is higher than the resolution of the generated speaking face image.

[0007] According to a high-fidelity and high-synchronization speaking face generation model training method provided by the present invention, the hard example mining is performed based on the pose image and the image candidate set to obtain the identity reference image corresponding to the pose image, including: performing facial key point detection on the pose image and each candidate image in the reference image candidate set to obtain the first facial key point corresponding to the pose image and the second facial key point corresponding to each candidate image; determining the first affine transformation matrix corresponding to the first facial key point and the second affine transformation matrix corresponding to the second facial key point; aligning the pose image with the first facial key point, and each candidate image with the second facial key point, based on the first affine transformation matrix and the second affine transformation matrix, to obtain the first mouth corresponding key point of the pose image and the second mouth corresponding key point of each candidate image; respectively determining the distance between the first mouth corresponding key point of the pose image and the second mouth corresponding key point of each candidate image; and taking the candidate image corresponding to the second mouth corresponding key point with the largest distance as the identity reference image.

[0008] According to a high-fidelity and high-synchronization speaking face generation model training method provided by the present invention, the candidate image corresponding to the second mouth corresponding key point with the largest distance is used as the identity reference image, comprising: in, represents the identity reference image, is the candidate image, represents the first affine transformation matrix, represents the second affine transformation matrix, represents the first mouth corresponding key point of the pose image, The second mouth corresponding key point of each candidate image is represented.

[0009] According to a high-fidelity and high-synchronization speaking face generation model training method provided by the present invention, the speaking face generation model includes: an image encoder, a pre-trained audio feature extractor and a generator; the identity reference image, the posture image and the audio to be driven are input into the speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, including: blocking a target part of the posture image to obtain an occluded posture image; splicing the occluded posture image with the identity reference image based on RGB channels to obtain a spliced ​​image; inputting the spliced ​​image into the image encoder to obtain an image feature map output by the image encoder; inputting the audio to be driven into the pre-trained audio feature extractor to obtain audio features output by the pre-trained audio feature extractor; inputting the image feature map as a noise condition into the generator, and inputting the audio features as a style condition into the generator; and outputting a generated speaking face image corresponding to the audio to be driven through the generator.

[0010] According to a high-fidelity and high-synchronization speaking face generation model training method provided by the present invention, the generation adversarial network model includes: a discriminator and a generator; before the loss function guided by the target resolution face image supervises the generation speaking face image model, the method further includes: obtaining a real speaking face image and a target resolution face image; the loss function guided by the target resolution face image is specifically: in, represents the loss function guided by the target resolution face image, represents the discriminator, represents the generator, represents the real speaking face image, represents the target resolution face image, represents the generation of a speaking face image, Indicates that the generator is minimized and the discriminator is maximized; represents the expected log-likelihood of the real speaking face image, represents the expected log-likelihood of the target resolution face image, represents the expected log-likelihood of generating the speaking face image.

[0011] According to a high-fidelity and high-synchronization speaking face generation model training method provided by the present invention, after the loss function guided by the target resolution facial image supervises the generated speaking face image model, the method also includes: randomly initializing the speaking face image model to obtain an initialized speaking face image model; using the global loss function as the initial loss function for training to improve the lip-syncing ability of the initialized speaking face image model; when the lip-syncing ability of the initialized speaking face image model no longer increases, replacing the initial loss function with the detail loss function for training to improve the quality of the generated facial image of the initialized speaking face image model; and ending the training when the quality of the generated facial image no longer increases or the lip-syncing ability decreases.

[0012] The present invention also provides a high-fidelity and high-synchronization speaking face generation model training system, comprising the following modules: an acquisition module, used to acquire audio to be driven, a pose image and a reference image candidate set; a mining module, used to perform hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image, wherein the mouth matching degree between the identity reference image and the pose image is the lowest; an input module, used to input the identity reference image, the pose image and the audio to be driven into the speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, wherein the speaking face generation model is based on a generative adversarial network model; a supervision module, used to supervise the generated speaking face image model based on a loss function guided by a target resolution face image to train the speaking face generation model, wherein the resolution of the target resolution face image is higher than the resolution of the generated speaking face image.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the high-fidelity and high-synchronization speaking face generation model training method as described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described high-fidelity and high-synchronization speaking face generation model training methods.

[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned high-fidelity and high-synchronization speaking face generation model training methods.

[0016] The high-fidelity and high-synchronization speaking face generation model training method and system provided by the present invention ensure the diversity and richness of model training data by obtaining audio to be driven, pose images and reference image candidate sets; perform hard case mining based on pose images and image candidate sets to obtain identity reference images with the lowest matching degree with the mouth of the pose image, thereby increasing the training difficulty by selecting the most difficult-to-match image, thereby forcing the model to learn more refined feature representations, and improving the model's adaptability and robustness to different identities and expressions; input the identity reference image, pose image and audio to be driven into a speaking face generation model based on a generative adversarial network, and can generate audio-synchronized speaking face images; supervise the generation speaking face image model based on a loss function guided by a target resolution face image, and ensure that the generated speaking face image is consistent with the target high-resolution face image in terms of resolution, details and overall quality, thereby improving the generation quality and fidelity of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 It is a schematic diagram of the overall process of the high-fidelity and high-synchronization speaking face generation model training method provided by the present invention.

[0019] Figure 2 It is a flow chart of the high-fidelity and high-synchronization speaking face generation model training method provided by the present invention.

[0020] Figure 3 It is a schematic diagram of a process of obtaining an image that is most dissimilar to the mouth of a pose image from a reference image candidate set based on hard example mining provided by the present invention.

[0021] Figure 4 It is a schematic diagram of the flow of generating a speaking face image of the method provided by the present invention.

[0022] Figure 5 It is a schematic diagram of the process of adjusting the loss function provided by the present invention.

[0023] Figure 6 It is a structural schematic diagram of the high-fidelity and high-synchronization speaking face generation model training system provided by the present invention.

[0024] Figure 7 It is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] The current speaking face generation algorithms are divided into two categories. The methods based on neural radiation fields or Gaussian sputtering can generate a specific person's speaking face through a long video of a specific person, which is relatively inefficient. The methods based on generative models generally use a generative model, such as a generative adversarial network and a diffusion model, to generate a general speaking face. Among them, the diffusion model takes a long time to train and infer, and the generation speed is slow. The speaking face generation based on the generative adversarial network generally uses an identity reference image and a pose image as input, combined with the given audio for generation.

[0027] Existing methods generally use random selection when selecting identity reference images. This method results in a large number of samples with the same mouth shape in the reference image and the pose image during training. The model tends to directly copy the mouth shape, affecting the synchronization of the generated speaking face. In addition, due to the scarcity of existing high-resolution speaking face datasets, the judgment ability of the discriminator in the generative adversarial network is limited, which limits the generation ability of the generator.

[0028] In view of the defects of low fidelity and synchronization of speaking faces generated by existing speaking face generation algorithms, the present invention selects an identity reference image with a low matching degree with the pose image to improve the synchronization, uses a loss function guided by a high-resolution face image to improve the fidelity of the model, and uses a two-stage training strategy to enable the model to generate speaking faces with both fidelity and synchronization, so that the quality of the generated speaking faces is improved and more realistic.

[0029] refer to Figure 1 , Figure 1 The figure is a schematic diagram of the overall process of the high-fidelity and high-synchronization speaking face generation model training method provided by the present invention, which specifically includes the following steps: S010, obtaining audio, pose image and reference image candidate set to be driven.

[0030] S020, based on hard example mining, obtains the image that is least similar to the mouth of the pose image from the reference image candidate set to improve the model's lip-syncing ability.

[0031] S030, using the identity reference image, the posture image and the audio as conditions, using a generative network to generate a speaking face image.

[0032] S040, supervising the generated result with a loss function guided by a high-resolution face image to train the generated network.

[0033] S050, based on a global-to-detail generation method, adjusting the loss function during the training process to generate a high-fidelity and highly synchronized speaking face.

[0034] In the embodiments below, the image least similar to the mouth of the pose image is obtained from a reference image candidate set based on hard example mining, the generation method of speaking face image, the loss function construction guided by high-resolution face image, and the adjustment of the loss function during the training process based on a generation method from global to detail are described in detail.

[0035] Optionally, the high-fidelity and high-synchronization speaking face generation model training method of the embodiment of the present application can be executed by a server, or by a terminal device, or jointly by a server and a terminal device, taking the example of the high-fidelity and high-synchronization speaking face generation model training method of the embodiment being executed by a server.

[0036] Figure 2 is a flow chart of the high-fidelity and high-synchronization speaking face generation model training method provided by the present invention, such as Figure 2 As shown, the method includes the following steps.

[0037] Step 201, obtaining audio to be driven, pose image and reference image candidate set.

[0038] The driving audio provides the speech information to be synthesized, the pose image provides the facial posture and expression information to be synthesized, and the reference image candidate set provides a variety of reference information for generating the speaking face image.

[0039] In some embodiments, the audio to be driven refers to the audio input used to drive the speaker face generation model, which contains the speaker's voice information and rhythm.

[0040] For example, methods for obtaining audio to be driven usually include: recording equipment acquisition, that is, using professional recording equipment (such as a microphone, a voice recorder, etc.) to record the speaker's voice; uploading existing audio resources, that is, obtaining from an existing audio resource library, such as a music library, a voice library, etc.

[0041] A pose image refers to an image containing the speaker's facial posture and expression, which is used to provide dynamic information about the speaker's face.

[0042] For example, the method of obtaining the posture image generally includes: camera shooting, using the camera to shoot the speaker's face image in real time; image database uploading, that is, obtaining the image containing the speaker's facial posture and expression from the image database.

[0043] The reference image candidate set refers to a set of candidate images that are used to match the pose image to find the closest or least similar candidate images to the pose image.

[0044] For example, the method of obtaining a candidate set of reference images usually includes: image database search, that is, searching for images similar to the pose image in a large image database. These databases may contain a large number of image samples, and images similar to the pose image can be found through keyword search, image recognition and other technologies.

[0045] Step 202 , performing hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image.

[0046] Among them, the mouth has the lowest match between the identity reference image and the pose image.

[0047] Hard example mining refers to identifying samples from the training data set that are difficult to be correctly classified or predicted by the model, and using these samples for further training of the model to improve the robustness and accuracy of the model. The core idea is to optimize the performance of the model by focusing on samples that the model predicts incorrectly.

[0048] According to a high-fidelity and high-synchronization speaking face generation model training method and system provided by the present invention, hard example mining is performed based on pose images and image candidate sets to obtain identity reference images corresponding to the pose images, including: Performing facial key point detection on the pose image and each candidate image in the reference image candidate set, respectively, to obtain the first facial key point corresponding to the pose image and the second facial key point corresponding to each candidate image; Determine a first affine transformation matrix corresponding to the first facial key points, and a second affine transformation matrix corresponding to the second facial key points; Align the pose image with the first facial key point, and each candidate image with the second facial key point based on the first affine transformation matrix and the second affine transformation matrix, respectively, to obtain the first mouth corresponding key point of the pose image and the second mouth corresponding key point of each candidate image; Determine the distance between the first mouth corresponding key point of the pose image and the second mouth corresponding key point of each candidate image respectively; The candidate image corresponding to the second mouth corresponding key point with the largest distance is used as the identity reference image.

[0049] refer to Figure 3 , Figure 3 It is a schematic diagram of a process of obtaining an image that is most dissimilar to the mouth of a pose image from a reference image candidate set based on hard example mining provided by the present invention.

[0050] In the present invention, if Figure 3 As shown, the reference image is compared with the pose image in terms of mouth shape, and the image with the least similar mouth shape is selected from the reference image candidate set as the identity reference image. The specific steps are as follows: Step 1: Obtain the input pose image and reference image candidate set.

[0051] Step 2: Detect facial key points on the pose image and all candidate images.

[0052] Step 3: Calculate the corresponding affine transformation matrix based on the facial key points.

[0053] Step 4: Align the pose image and all candidate images and their key points to the same pose through the affine transformation matrix.

[0054] Step 5: Take the aligned mouth key points and calculate the distance between the pose image and the corresponding mouth of all candidate images.

[0055] Step 6: Select the candidate reference image with the largest distance as the input identity reference image.

[0056] Thus, the model generation is made more difficult to avoid directly copying the mouth from the reference image.

[0057] The input pose image and the reference image candidate set are selected from a fixed window near the pose image. The selected window is centered on the pose image and has a size of 150 frames.

[0058] Among them, the affine transformation matrix is ​​calculated from the key points of the face. By selecting five key points, namely the left eye, right eye, nose tip, left corner of the mouth, and right corner of the mouth, and the corresponding five template key points, the affine transformation matrix from the identity image pose to the template image pose is solved by the least squares method.

[0059] In some embodiments, a facial key point detection algorithm is used to detect facial key points in the pose image and all candidate images. Facial key points generally include feature points such as the corners of the eyes, the corners of the mouth, and the tip of the nose. The detected facial key points are used to calculate the affine transformation matrix from the candidate image to the pose image. The affine transformation matrix is ​​used to align the candidate image to the pose of the pose image. The calculated affine transformation matrix is ​​applied to align the candidate image and its key points to the pose of the pose image, so that all images will have the same pose and expression (different identities). In the aligned image, the position of the mouth key point is extracted. The distance of the mouth key point between the pose image and each candidate image is calculated (for example, using Euclidean distance, Manhattan distance, etc.). From all candidate images, the image farthest from the mouth key point of the pose image is selected as the final identity reference image.

[0060] This selection process is designed to increase the difficulty for the model to generate talking faces and prevent the model from directly copying the mouth area of ​​the reference image, thereby encouraging the model to learn how to generate more realistic and natural talking faces based on audio signals and posture information.

[0061] Through the embodiments of the present invention, the difficulty of model training can be effectively increased, and the robustness and accuracy of the model when generating speaking faces can be improved.

[0062] According to a high-fidelity and high-synchronization speaking face generation model training method and system provided by the present invention, the candidate image corresponding to the second mouth corresponding key point with the largest distance is used as the identity reference image, including: in, represents the identity reference image, is a candidate image. represents the first affine transformation matrix, represents the second affine transformation matrix, Indicates the key point corresponding to the first mouth of the pose image, Indicates the second mouth corresponding key point of each candidate image.

[0063] In an embodiment of the present invention, Represents each candidate image in the reference image candidate set Perform the subsequent operation and select the candidate image that maximizes the result of the operation , as the identity reference image .

[0064] In some embodiments, facial key point detection is performed on the pose image and each candidate image to obtain the coordinates of the mouth key points. The detected key points are used to calculate the affine transformation matrix, and the candidate image is aligned to the pose of the pose image. For each candidate image, the distance between the new coordinates of its mouth key points after affine transformation and the new coordinates of the mouth key points in the pose image is calculated. The candidate image with the largest distance is selected as the final reference image.

[0065] Through the embodiments of the present invention, the candidate image with the largest difference from the mouth of the pose image can be effectively selected as the reference image, thereby increasing the difficulty and generalization ability of the model training.

[0066] Step 203 , input the identity reference image, the posture image, and the audio to be driven into the speaking face generation model to obtain a generated speaking face image output by the speaking face generation model.

[0067] Among them, the speaking face generation model is based on the generative adversarial network model.

[0068] In an embodiment of the present invention, input data includes: an identity reference image, a posture image, and an audio to be driven, wherein the identity reference image includes posture features of a reference speaker, and the identity reference image is a candidate image with the lowest matching degree with the mouth of the posture image, and is used to increase the difficulty of model generation and avoid directly copying the mouth from the reference image; the posture image captures the speaker's head posture and facial expression, and it is usually close or synchronized in time with the audio to be driven, and the posture image is used to guide the generation of the posture and expression of the speaker's face image; the audio to be driven includes an audio signal of the speaking content.

[0069] The speaking face generation model is based on a generative adversarial network and is used to generate realistic speaking face images. The speaking face generation model includes: a generator and a discriminator. The generator is responsible for synthesizing realistic speaking face images from the input identity, posture and audio features. The discriminator is used to distinguish the generated image from the real image, thereby helping the generator to improve the quality of the generated image.

[0070] During the training process, the generator and the discriminator are continuously optimized through competition and cooperation until the generator is able to produce synthetic images that are difficult to distinguish from real images.

[0071] According to a high-fidelity and high-synchronization speaking face generation model training method and system provided by the present invention, the speaking face generation model includes: an image encoder, a pre-trained audio feature extractor and a generator; The identity reference image, the posture image, and the audio to be driven are input into the speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, including: Occlude the target part of the pose image to obtain an occluded pose image; The occluded pose image and the identity reference image are spliced ​​based on the RGB channel to obtain a spliced ​​image; Input the spliced ​​image into an image encoder to obtain an image feature map output by the image encoder; Inputting the audio to be driven into a pre-trained audio feature extractor to obtain audio features output by the pre-trained audio feature extractor; Input the image feature map as the noise condition to the generator, and input the audio feature as the style condition to the generator; The generator outputs a generated speaking face image corresponding to the audio to be driven.

[0072] refer to Figure 4 , Figure 4 It is a schematic diagram of the flow of generating a speaking face image of the method provided by the present invention.

[0073] In the embodiment of the present invention, Figure 4 As shown in the figure, the network that generates the predicted speaking face image is styleGAN2 (an improved version of the Generative Adversarial Network (GAN)), which takes advantage of its ability to generate high-quality images. The specific steps for generating speaking face images are as follows: Step 1: Mask the lower half of the face of the input pose image and concatenate it with the reference identity image on the RGB channel.

[0074] Step 2: The concatenated input is fed into the image encoder to obtain the output image vector and feature map.

[0075] Step 3: Extract audio features from the input audio features using a pre-trained audio feature extractor.

[0076] Step 4: Input the audio features into the generator as style conditions and the image feature maps into the generator as noise conditions.

[0077] Step 5: The generator outputs the speaking face image corresponding to the audio.

[0078] Among them, the input pose image and the reference identity image are spliced ​​to obtain a six-channel image, which is input into the image encoder to obtain the output image vector and its corresponding feature map, and the size gradually decreases.

[0079] Among them, the audio is extracted through the pre-trained audio feature extractor whisper, which can extract audio information containing many audio corresponding content features.

[0080] The audio features are input into the generator as style conditions, and the image feature map is input into the generator as noise conditions. The audio features are mapped to the latent space of the generator through a mapping network, and the latent space controls the mouth shape of the generated image. The image feature map is input into the image as a noise condition, providing conditions such as spatial information.

[0081] In some embodiments, first, the input pose image is processed to block its lower half (usually including the mouth and chin area). The purpose of this is to make the model rely more on the identity information in the reference identity image rather than the lower face information in the pose image during the subsequent generation process. The blocked upper half of the pose image is spliced ​​with the lower half of the reference identity image on the RGB channel. It should be noted here that when splicing, it is necessary to ensure that the two images are consistent in size, resolution, color space, etc. to avoid obvious seams or color mismatches in the spliced ​​image.

[0082] The spliced ​​image is input into the image encoder. The image encoder is a deep learning model that can convert the input image into a high-dimensional vector representation (i.e., image vector) and feature map. The image vector output by the image encoder is usually used to represent the global information of the image, while the feature map contains the local information and spatial structure of the image. This information will be used as input for the subsequent generation process.

[0083] Use a pre-trained audio feature extractor to extract features of the input audio. The audio feature extractor converts the input audio signal into a series of feature vectors that capture information such as speech content, intonation, and rhythm in the audio.

[0084] The extracted audio features are input into the generator as style conditions. The style conditions are used to guide the generator to generate a speaking face image that matches the audio features. The image feature map is input into the generator as a noise condition. The noise condition here is not random noise in the traditional sense, but a feature map extracted from the image, which contains the spatial structure and local information of the image. This information will be used as an additional input to the generation process to enrich the content of the generated image.

[0085] Based on the input style and noise conditions, the generator will output a speaker face image corresponding to the audio. The speaker face image contains dynamic features such as mouth shape and expression corresponding to the speech content in the input audio.

[0086] Among them, the generator outputs the speaking face image corresponding to the audio, specifically: in, is the generated speaking face image, is the reference image, is the pose image, The characteristics of the input audio, is a generative network.

[0087] Through the embodiments of the present invention, by blocking the lower half of the face of the pose image and splicing it with the reference identity image, the consistency of the identity information of the generated image can be ensured, while allowing the model to rely more on audio features to generate mouth shapes and expressions, thereby improving the realism of the generated image. The audio features are input into the generator as style conditions, which can guide the generator to generate a speaking face image synchronized with the audio content, thereby enhancing the synchronization between the image and the audio.

[0088] Step 204 , based on the loss function guided by the target resolution face image, supervise the generation of the speaking face image model to train the speaking face generation model.

[0089] The resolution of the target resolution face image is higher than the resolution of the generated speaking face image.

[0090] Due to the lack of high-resolution talking face datasets, in order to improve the discriminant ability of the discriminator and thus enhance the generation ability of the generator, an embodiment of the present invention introduces a high-quality face image dataset (i.e., real talking face images) as true value data during the discriminator training process, allowing the discriminator to simultaneously determine the generation of talking faces and high-quality face images, thereby improving the discriminator's judgment ability, thereby enhancing the generator's generation ability in the confrontation.

[0091] According to a high-fidelity and high-synchronization speaking face generation model training method and system provided by the present invention, the generation adversarial network model includes: a discriminator and a generator; Prior to supervising the model for generating a speaking face image based on a loss function guided by the target resolution face image, the method further includes: Obtaining a real speaking face image and a target resolution face image; The loss function guided by the target resolution face image is specifically: in, represents the loss function guided by the target resolution face image, represents the discriminator, represents a generator, represents the real speaking face image, represents the target resolution face image, Indicates generating a speaking face image, Indicates minimization optimization of the generator and maximization optimization of the discriminator; represents the expected log-likelihood of the real speaking face image, represents the expected log-likelihood of the target resolution face image, represents the expected log-likelihood of generating a speaking face image.

[0092] Here, after obtaining a dataset of real speaking face images and a dataset of target resolution face images, the high-resolution face image dataset and the dataset of real speaking face videos are mixed and sent together to the discriminator for training.

[0093] Among them, other loss functions include reconstruction loss and audio lip alignment loss. The reconstruction function includes two parts: pixel loss and perceptual loss. Pixel loss is responsible for reconstructing spatial information, and perceptual loss is responsible for reconstructing detail information. The reconstruction loss function is specifically: in, To reconstruct the loss function, is the activation layer of the pre-trained perception network, and They are respectively generated images (i.e. generated speaking face images) and real images (i.e. real speaking face images). and are the weights of reconstruction loss and perceptual loss, respectively.

[0094] The audio lip alignment loss function uses a pre-trained lip alignment network to determine whether the input audio and the face lip shape match. Specifically: in, is the audio lip alignment loss function, and are the image encoder and audio encoder of the lip alignment expert network, is the real image frame and the following four frames, It is the corresponding audio feature and the following four frames.

[0095] Through the embodiments of the present invention, real speaking face images are introduced as true value data in the discriminator training process, so that the discriminator can simultaneously judge the generated speaking face and the real speaking face image, thereby improving the judgment ability of the discriminator and enhancing the generation ability of the generator in the confrontation.

[0096] According to a high-fidelity and high-synchronization speaking face generation model training method and system provided by the present invention, after supervising the generation of the speaking face image model based on the loss function guided by the target resolution face image, the method further includes: Randomly initializing the speaking face image model to obtain an initialized speaking face image model; The global loss function is used as the initial loss function for training to improve the lip-syncing ability of the initialized speaking face image model; When the lip-syncing ability of the initialized speaking face image model no longer increases, the initial loss function is replaced with the detail loss function for training to improve the quality of the generated face images of the initialized speaking face image model; When the quality of the generated face images no longer increases or the lip-syncing ability decreases, the training ends.

[0097] refer to Figure 5 , Figure 5 It is a schematic diagram of the process of adjusting the loss function provided by the present invention.

[0098] In the embodiment of the present invention, a two-stage model training method is used to first improve the lip-syncing ability of the model and then improve the model's ability to generate high-quality faces, thereby achieving the generation of speaking faces with both high fidelity and high synchronization. The specific steps are: Step 1: Randomly initialize the model.

[0099] Step 2: Use the global loss function as the initial loss function to improve lip-syncing ability and start training.

[0100] Step 3: When the model’s ability to generate lip-syncing images no longer improves, the loss function is replaced with a detail loss function to improve the quality of the generated faces.

[0101] Step 4: When the quality of the images generated by the model no longer increases or the lip-syncing ability decreases, the training ends.

[0102] Among them, the global loss function is specifically: in, is the global loss function, It is against loss, is the reconstruction loss, It's lip sync loss, , , They are the weights corresponding to the three loss functions.

[0103] Among them, the detail loss function is specifically: in, is the detail loss function, is the adversarial loss guided by the high-resolution face image, is the reconstruction loss, , are the weights corresponding to the two loss functions respectively.

[0104] In the second stage, when improving the model's ability to generate high-quality faces, the training time is one-third of that in the first stage, and the model's lip-syncing ability will not be destroyed.

[0105] In some embodiments, the model's parameters such as weights and biases are randomly initialized, and a global loss function is designed to evaluate the difference in global features between the model-generated image and the real image (or target image), especially the evaluation of lip-syncing ability. Global features include the overall structure of the image, color distribution, and the synchronization of the mouth shape with the audio content. By minimizing the global loss function, the model gradually learns how to generate a talking face image that is synchronized with the audio. At this stage, the main task of the model is to improve the lip-syncing ability.

[0106] When the model's lip-syncing ability reaches a certain level, further improvement may become difficult. At this point, the global loss function is replaced by a detail loss function, which focuses more on local details in the image, such as texture, edges, and finer changes in mouth shape. By minimizing the detail loss function, the model can further improve the quality of the generated image, making the generated image more realistic and delicate.

[0107] When the model's performance on the detail loss function stops improving, or the lip-syncing ability decreases, it can be considered that the model has reached its performance ceiling. At this point, the training process ends.

[0108] Through the above-mentioned embodiments of the present invention, the present invention improves the fidelity and synchronization of the speaking person's face.

[0109] The present invention selects pose images from a reference image candidate set based on hard case mining. The affine transformation matrix is ​​calculated by key point detection, and all candidate images (reference images) and pose images and key points are mapped to a pose. The distances between the candidate reference image and the pose image and the lip key points are compared, and the largest one is selected as the reference image. The reference image and the pose image selected by this method have large differences in lip shapes, which increases the difficulty of model training, prevents the model from directly copying the lip part of the reference image, and improves the synchronization of generating the speaking face.

[0110] The present invention uses a loss function guided by high-resolution face images to supervise the results generated by the generative network, introduces a high-resolution face image dataset, and calculates the adversarial loss together with the speaking face dataset to improve the performance of the model discriminator, thereby enabling the generator to have a more powerful ability to generate high-fidelity faces.

[0111] The present invention is based on a generation method from global to detailed. The loss function is adjusted during the training process. The initial loss function focuses on lip syncing. When the lip syncing ability of the model no longer increases, the focus is on improving the image quality, so that the model can generate speaking faces with both high fidelity and high synchronization.

[0112] refer to Figure 6 , Figure 6 It is a structural schematic diagram of the high-fidelity and high-synchronization speaking face generation model training system provided by the present invention.

[0113] An acquisition module 601 is used to acquire audio to be driven, pose images and a candidate set of reference images; A mining module 602 is used to perform hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image, wherein the mouth matching degree between the identity reference image and the pose image is the lowest; An input module 603 is used to input the identity reference image, the posture image and the audio to be driven into the speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, wherein the speaking face generation model is based on a generative adversarial network model; The supervision module 604 is used to supervise the generation of the speaking face image model based on the loss function guided by the target resolution face image, so as to train the speaking face generation model.

[0114] Specifically, the high-fidelity and high-synchronization speaking face generation model training system provided by the present invention can implement all the method steps implemented by the above-mentioned high-fidelity and high-synchronization speaking face generation model training method embodiment, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.

[0115] Figure 7 is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as Figure 7As shown, the electronic device may include: a processor (processor) 710 , a communication interface (Communications Interface) 720 , a memory (memory) 730 and a communication bus 740 , wherein the processor 710 , the communication interface 720 , and the memory 730 communicate with each other through the communication bus 740 . The processor 710 can call the logic instructions in the memory 730 to execute a high-fidelity and high-synchronization speaking face generation model training method, which includes: obtaining audio to be driven, a pose image and a reference image candidate set; performing hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image, wherein the mouth match between the identity reference image and the pose image is the lowest; inputting the identity reference image, the pose image and the audio to be driven into the speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, wherein the speaking face generation model is based on a generative adversarial network model; supervising the generated speaking face image model based on a loss function guided by a target resolution face image to train the speaking face generation model, wherein the resolution of the target resolution face image is higher than the resolution of the generated speaking face image.

[0116] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0117] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the high-fidelity and high-synchronization speaking face generation model training method provided by the above methods, the method including: obtaining audio to be driven, a pose image and a reference image candidate set; performing hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image, wherein the mouth match between the identity reference image and the pose image is the lowest; inputting the identity reference image, the pose image and the audio to be driven into the speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, wherein the speaking face generation model is based on a generative adversarial network model; supervising the generated speaking face image model based on a loss function guided by a target resolution face image to train the speaking face generation model, wherein the resolution of the target resolution face image is higher than the resolution of the generated speaking face image.

[0118] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the high-fidelity and high-synchronization speaking face generation model training method provided by the above-mentioned methods, the method comprising: obtaining audio to be driven, a pose image and a reference image candidate set; performing hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image, wherein the mouth match between the identity reference image and the pose image is the lowest; inputting the identity reference image, the pose image and the audio to be driven into a speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, wherein the speaking face generation model is based on a generative adversarial network model; supervising the generated speaking face image model based on a loss function guided by a target resolution face image to train the speaking face generation model, wherein the resolution of the target resolution face image is higher than the resolution of the generated speaking face image.

[0119] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0120] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A high-fidelity and high-synchronization speaking face generation model training method, characterized in that: include: Obtain the audio to be driven, the pose image and the candidate set of reference images; Performing hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image, wherein the mouth matching degree between the identity reference image and the pose image is the lowest; Inputting the identity reference image, the posture image, and the to-be-driven audio into a speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, wherein the speaking face generation model is based on a generative adversarial network model; Based on a loss function guided by a target resolution face image, the generated speaking face image model is supervised to train the speaking face generation model, wherein the resolution of the target resolution face image is higher than the resolution of the generated speaking face image.

2. The high-fidelity and high-synchronization speaking face generation model training method according to claim 1, characterized in that: The performing hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image includes: Performing facial key point detection on the pose image and each candidate image in the reference image candidate set respectively to obtain a first facial key point corresponding to the pose image and a second facial key point corresponding to each candidate image; Determine a first affine transformation matrix corresponding to the first facial key points, and a second affine transformation matrix corresponding to the second facial key points; Align the pose image with the first facial key point, and each candidate image with the second facial key point based on the first affine transformation matrix and the second affine transformation matrix, respectively, to obtain a first mouth corresponding key point of the pose image and a second mouth corresponding key point of each candidate image; respectively determining the distance between the first mouth corresponding key point of the pose image and the second mouth corresponding key point of each candidate image; The candidate image corresponding to the second mouth corresponding key point with the largest distance is used as the identity reference image.

3. The high-fidelity and high-synchronization speaking face generation model training method according to claim 2, characterized in that: The step of taking the candidate image corresponding to the second mouth corresponding key point with the largest distance as the identity reference image includes: in, represents the identity reference image, is the candidate image, represents the first affine transformation matrix, represents the second affine transformation matrix, represents the first mouth corresponding key point of the pose image, The second mouth corresponding key point of each candidate image is represented.

4. The high-fidelity and high-synchronization speaking face generation model training method according to claim 1, characterized in that: The speaking face generation model includes: an image encoder, a pre-trained audio feature extractor and a generator; The step of inputting the identity reference image, the posture image, and the to-be-driven audio into a speaker face generation model to obtain a generated speaker face image output by the speaker face generation model includes: Blocking a target part of the pose image to obtain a blocked pose image; splicing the occluded pose image and the identity reference image based on RGB channels to obtain a spliced ​​image; Inputting the spliced ​​image into the image encoder to obtain an image feature map output by the image encoder; Inputting the to-be-driven audio into the pre-trained audio feature extractor to obtain audio features output by the pre-trained audio feature extractor; Inputting the image feature map as a noise condition into the generator, and inputting the audio feature as a style condition into the generator; The generator outputs a generated speaking face image corresponding to the audio to be driven.

5. The high-fidelity and high-synchronization speaking face generation model training method according to claim 1, characterized in that: The generative adversarial network model includes: a discriminator and a generator; Before the loss function guided by the target resolution face image supervises the generation of the speaking face image model, the method further includes: Obtaining a real speaking face image and a target resolution face image; The loss function guided by the target resolution face image is specifically: in, represents the loss function guided by the target resolution face image, represents the discriminator, represents the generator, represents the real speaking face image, represents the target resolution face image, represents the generation of a speaking face image, Indicates that the generator is minimized and the discriminator is maximized; represents the expected log-likelihood of the real speaking face image, represents the expected log-likelihood of the target resolution face image, represents the expected log-likelihood of generating the speaking face image.

6. The high-fidelity and high-synchronization speaking face generation model training method according to claim 1, characterized in that: After the loss function guided by the target resolution face image supervises the generation of the speaking face image model, the method further includes: Randomly initializing the speaker face image model to obtain an initialized speaker face image model; Using the global loss function as the initial loss function for training, so as to improve the lip-syncing ability of the initialized speaking face image model; When the lip-syncing ability of the initialized speaking face image model no longer increases, replacing the initial loss function with a detail loss function for training to improve the quality of the generated face image of the initialized speaking face image model; When the quality of the generated facial image no longer increases or the lip-syncing ability decreases, the training ends.

7. A high-fidelity and high-synchronization speaking face generation model training system, characterized in that: include: An acquisition module is used to acquire audio to be driven, pose images and a candidate set of reference images; A mining module, configured to perform hard example mining based on the pose image and the image candidate set to obtain an identity reference image corresponding to the pose image, wherein the mouth matching degree between the identity reference image and the pose image is the lowest; An input module, used for inputting the identity reference image, the posture image and the to-be-driven audio into a speaking face generation model to obtain a generated speaking face image output by the speaking face generation model, wherein the speaking face generation model is based on a generative adversarial network model; A supervision module is used to supervise the generated speaking face image model based on a loss function guided by a target resolution face image to train the speaking face generation model, wherein the resolution of the target resolution face image is higher than the resolution of the generated speaking face image.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the high-fidelity and high-synchronization speaking face generation model training method as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the high-fidelity and high-synchronization speaking face generation model training method as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the high-fidelity and high-synchronization speaking face generation model training method as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Digital human modeling method and system combining Gaussian sputtering image and GAN model

    CN121280577A

  • Digital human modeling method and system combining gaussian sputtering image and gan model

    CN121280577B

  • 3D field speaking face generation defense method, system, device and medium

    CN121640543A