Face video generation method for an audio-driven face dialogue generation model

By using the audio-driven face dialogue generation model and attention mechanism in face video generation, the problems of unnatural lip generation and low pixel values ​​are solved, and face video generation with high quality and good synchronization effects are achieved.

CN119028369BActive Publication Date: 2025-06-17JINHUA INSTITUTE OF ZHEJIANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411029024.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-06-17
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

In the prior art, when generating face videos, lip shape generation is exaggerated and unnatural, and the low pixel value leads to blur and artifacts, making it difficult to generate high-quality face sequences.

Method used

Using an audio-driven face dialogue generation model, the lip-sound synchronization discrimination network and the audio-driven lip-shaped network QAW based on Quality-Attention-Wav2Lip are used to optimize the lip-shaped generation and face video quality by establishing a lip-shaped synchronization discrimination network and the Quality-Attention-Wav2Lip-based audio-driven lip-shaped network QAW, combining spatial and position attention mechanisms.

Benefits of technology

It effectively improves the synchronization effect of lip shape generation and the image quality of the overall face. The generated face video has natural head movement and lip sound synchronization effects, which significantly improves the authenticity and clarity of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119028369B_ABST
    Figure CN119028369B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating a face video of a face dialogue generation model driven by audio. The method includes: establishing a lip-sync discrimination network and an audio-driven lip shape network based on quality attention; training the network using a lip-sync training set, constructing an overall loss function of the audio-driven lip shape network based on quality attention according to the discrimination loss function of the lip-sync discrimination network, and completing the training until the overall loss function converges; obtaining a reply audio according to the text or audio to be replied; inputting the reply audio and the face image of the person to be generated into the trained network, outputting a face video of the current person when reading the current reply audio, and finally displaying it on a display. The method of the present invention effectively improves the synchronization effect of lip shape generation and the image quality of the overall face, and can conduct a dialogue with the customer, aiming to generate a real face video with natural head movement and good lip-sync effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for generating a face video, and more particularly to a method for generating a face video based on an audio-driven face dialogue generation model. Background Art

[0002] The face animation technology aims to generate a series of highly natural face sequences given speech or text. The generated video not only needs to ensure the realism of video frames and textures, but also the temporal continuity between video frames. The audio-driven face animation technology is the focus of most researchers. People are sensitive to subtle changes in the face. If the facial movement is inconsistent with the speech, it will cause a sense of incongruity to the user. How to improve the realism of face animation technology remains an urgent problem to be solved in computer animation.

[0003] To solve the problem of unnatural facial movement of the character during speech in the model-synthesized face animation, Wav2Lip receives a series of image frame sequences as input. While retaining the facial movement information of the character in the original frame sequence, it generates face animation by guiding the change of lip movement through audio features. Compared with previous models, the Wav2Lip model has made significant progress in generating natural lip movement. This model can generate real face videos with natural head movement and good lip-sync effect. However, there are still problems such as exaggerated lip generation and mismatch with the previous frames. In addition, the generated lip pixel values are too low, resulting in blurring and artifacts. The detailed features generated on the lower half of the face still do not match the overall features naturally, and even post-image enhancement cannot generate high-quality face sequences. Summary of the Invention

[0004] To solve the problems existing in the background art, the present invention provides a method for generating a face video based on an audio-driven face dialogue generation model.

[0005] The technical solution adopted by the present invention is as follows:

[0006] The method for generating a face video based on an audio-driven face dialogue generation model of the present invention includes:

[0007] S1) Establish a lip-sync discrimination network and a quality-attention-based audio-driven lip shape network QAW (Quality-Attention-Wav2Lip).

[0008] S2) Collect videos of several people speaking and divide each video into the person's audio and consecutive-frame face images. Construct a lip-sync training set from each group of person's audio and consecutive-frame face images.

[0009] S3) Input the lip-sync training set into the lip-sync discrimination network for training to make it more accurate in discriminating lip-sync faces until the discrimination loss function of the lip-sync discrimination network converges, and obtain the trained lip-sync discrimination network. The convergence of the discrimination loss function of the lip-sync discrimination network is constructed based on the synchronization discrimination weight file and the similarity function output by the lip-sync discrimination network during training.

[0010] S4) Input the lip-sync training set into the quality attention-based audio-driven lip shape network QAW for training, construct the overall loss function of the quality attention-based audio-driven lip shape network QAW based on the discrimination loss function, obtain the overall loss value of the quality attention-based audio-driven lip shape network QAW and use the gradient descent method to update the network parameters until the loss value of the overall loss function converges, and obtain the trained quality attention-based audio-driven lip shape network QAW.

[0011] S5) Input the text or audio to be replied into the dialogue system. For each piece of text or audio to be replied, the dialogue system processes it and outputs the reply text of the text or audio to be replied, and converts the reply text into a reply audio; input the reply audio and the face image of the person to be generated into the trained quality attention-based audio-driven lip shape network QAW. The trained quality attention-based audio-driven lip shape network QAW processes it and outputs the face video of the current person when reading the current reply audio, and finally displays it on the monitor to realize the generation of the face video.

[0012] In the step S1), the lip-sync discrimination network includes an audio encoder and an image encoder. The audio encoder includes several convolutional layers Conv connected in sequence and several two-dimensional convolutions Conv2d. The image encoder includes several two-dimensional convolutions Conv2d connected in sequence and several convolutional layers Conv. Each convolutional layer Conv is composed of three two-dimensional convolutions Conv2d connected in sequence. Each two-dimensional convolution Conv2d includes a convolutional block, a batch normalization layer BN (Batch Normalization), and a ReLU activation function connected in sequence; for each group of person audio and consecutive frame face images in the lip-sync training set, convert the person audio into a mel spectrogram and input it into the audio encoder for processing and then output the first audio feature vector, and input the consecutive frame face images into the image encoder to obtain the first lip shape feature vector; use the cosine similarity method (CosineSimilarity) to obtain the similarity function between the first audio feature vector and the first lip shape feature vector to represent the matching degree of audio and video, specifically as follows:

[0013]

[0014] Among them, Similarity(A,B) represents the similarity function between the first audio feature vector A and the first lip feature vector B; A i and B i respectively represent the i-th dimensional first audio feature vector output by the audio encoder and the i-th dimensional first lip feature vector output by the image encoder.

[0015] In the step S1), the quality attention-based audio-driven lip network QAW includes a face generator and a video quality discriminator connected in sequence. The lip-sync training set is input into the face generator, and the face generator outputs the generated face video after processing. The generated face video and the real face video are jointly input into the video quality discriminator to obtain the adversarial loss function and the quality evaluation loss function of the lower half of the face containing the lips of the video quality discriminator. When the adversarial loss function of the video quality discriminator converges, the video quality discriminator training is completed; the reconstruction loss function is obtained according to the generated face video and the real face video. Finally, the overall loss function of the quality attention-based audio-driven lip network QAW is constructed based on the discriminant loss function of the lip-sync discriminant network, the quality evaluation loss of the video quality discriminator, the FSIM loss function of the face dialogue generation model, and the reconstruction loss function.

[0016] The quality attention-based audio-driven lip network QAW extracts sound waveform information from audio and extracts character lip movement information from video, and uses a generative adversarial network to train the model to obtain the mapping relationship between the two, so that the model can generate a series of lip movement sequences matching the input audio.

[0017] The described face generator includes an audio encoder, a face encoder, and a face decoder. The audio encoder includes a first convolutional layer Conv, a second convolutional layer Conv, a third convolutional layer Conv, a fourth convolutional layer Conv, a fifth convolutional layer Conv, a sixth convolutional layer Conv, and a first two-dimensional convolutional layer Conv2d connected in sequence. Each convolutional layer Conv is composed of three two-dimensional convolutional layers Conv2d connected in sequence. Each two-dimensional convolutional layer Conv2d includes a convolutional block, a batch normalization layer BN, and a ReLU activation function connected in sequence. The face encoder includes a second two-dimensional convolutional layer Conv2d, a seventh convolutional layer Conv, an eighth convolutional layer Conv, a ninth convolutional layer Conv, a first spatial self-attention mechanism SSA (Spatial Self Attention), a tenth convolutional layer Conv, a second spatial self-attention mechanism SSA, an eleventh convolutional layer Conv, a twelfth convolutional layer Conv, a thirteenth convolutional layer Conv, a third two-dimensional convolutional layer Conv2d, and a fourth two-dimensional convolutional layer Conv2d connected in sequence. The face decoder includes a fifth two-dimensional convolutional layer Conv2d, a first spatial attention mechanism SA (Spatial Attention), a first fusion function Concat, a fourteenth convolutional layer Conv, a second spatial attention mechanism SA, a second fusion function Concat, a fifteenth convolutional layer Conv, a third spatial attention mechanism SA, a third fusion function Concat, a sixteenth convolutional layer Conv, a fourth spatial attention mechanism SA, a fourth fusion function Concat, a seventeenth convolutional layer Conv, a fifth spatial attention mechanism SA, a fifth fusion function Concat, an eighteenth convolutional layer Conv, a sixth spatial attention mechanism SA, a sixth fusion function Concat, a nineteenth convolutional layer Conv, a first coordinate attention mechanism CA (Coordinate Attention), a seventh spatial attention mechanism SA, a seventh fusion function Concat, a twentieth convolutional layer Conv, a second coordinate attention mechanism CA, an eighth spatial attention mechanism SA, an eighth fusion function Concat, a twenty-first convolutional layer Conv, a third coordinate attention mechanism CA, a ninth spatial attention mechanism SA, a ninth fusion function Concat, a sixth two-dimensional convolutional layer Conv2d, and a seventh two-dimensional convolutional layer Conv2d connected in sequence.

[0018] For each group of character audio and consecutive-frame face images in the lip synchronization training set, after converting the character audio into a Mel spectrogram and inputting it into the audio encoder of the face generator for processing to output a second audio feature vector, input the consecutive-frame face images into the face encoder for processing to output a second lip feature vector. Input the second audio feature vector into the fifth two-dimensional convolution Conv2d of the face decoder for processing. Input the output of the fifth two-dimensional convolution Conv2d processing and the second lip feature vector into the first spatial attention mechanism SA of the face encoder for processing. Input the output of the first spatial attention mechanism SA and the second lip feature vector into the first fusion function Concat for processing and then input it into the fourteenth convolutional layer Conv for processing. Input the output of the fourteenth convolutional layer Conv and the output of the thirteenth convolutional layer Conv into the second spatial attention mechanism SA for processing. Input the output of the second spatial attention mechanism SA and the output of the thirteenth convolutional layer Conv into the second fusion function Concat for processing and then output it to the fifteenth convolutional layer Conv for processing. Input the output of the fifteenth convolutional layer Conv and the output of the twelfth convolutional layer Conv into the third spatial attention mechanism SA for processing. Input the output of the third spatial attention mechanism SA and the output of the twelfth convolutional layer Conv into the third fusion function Concat for processing and then output it to the sixteenth convolutional layer Conv for processing. Input the output of the sixteenth convolutional layer Conv and the output of the eleventh convolutional layer Conv into the fourth spatial attention mechanism SA for processing. Input the output of the fourth spatial attention mechanism SA and the output of the eleventh convolutional layer Conv into the fourth fusion function Concat for processing and then output it to the seventeenth convolutional layer Conv for processing. Input the output of the seventeenth convolutional layer Conv and the output of the tenth convolutional layer Conv into the fifth spatial attention mechanism SA for processing. Input the output of the fifth spatial attention mechanism SA and the output of the tenth convolutional layer Conv into the fifth fusion function Concat for processing and then output it to the eighteenth convolutional layer Conv for processing. Input the output of the eighteenth convolutional layer Conv and the output of the ninth convolutional layer Conv into the sixth spatial attention mechanism SA for processing. Input the output of the sixth spatial attention mechanism SA and the output of the ninth convolutional layer Conv into the sixth fusion function Concat for processing and then sequentially output it to the nineteenth convolutional layer Conv and the first position attention mechanism CA for processing. Input the output of the first position attention mechanism CA and the output of the eighth convolutional layer Conv into the seventh spatial attention mechanism SA for processing. Input the output of the seventh spatial attention mechanism SA and the output of the eighth convolutional layer Conv into the seventh fusion function Concat for processing and then sequentially output it to the twentieth convolutional layer Conv and the second position attention mechanism CA for processing.The output of the second position attention mechanism CA and the output of the seventh convolutional layer Conv are jointly input into the eighth spatial attention mechanism SA for processing. The output of the eighth spatial attention mechanism SA and the output of the seventh convolutional layer Conv are jointly input into the eighth fusion function Concat for processing and then sequentially output to the twenty-first convolutional layer Conv and the third position attention mechanism CA for processing. The output of the third position attention mechanism CA and the output of the second two-dimensional convolution Conv2d are jointly input into the ninth spatial attention mechanism SA for processing. The output of the ninth spatial attention mechanism SA and the output of the second two-dimensional convolution Conv2d are jointly input into the ninth fusion function Concat for processing and then sequentially output to the sixth two-dimensional convolution Conv2d and the seventh two-dimensional convolution Conv2d for processing, and the final output is used as the output of the face generator.

[0019] a) The face encoder, which is composed of residual convolutional layers, mainly encodes the face reference frame F in the face video. The reference frame F obtained from the face video needs to be concatenated with the pose frame covering the lower half of the face in the channel dimension. In particular, to make the face generation more continuous and reduce the appearance of exaggerated lip shapes and unnatural lower contours, the reference frame generally uses the first T frames of the training frames. If there are less than T frames in front of the training frames, they are randomly selected. b) The audio encoder, which is a standard convolutional neural network, encodes the corresponding speech segment. Its input is a Mel-frequency Cepstral Coefficient (MFCC) heat map with a size of M×T×1, and the output is an audio embedding with a size of h. c) The face decoder connects the feature vector generated by the audio encoder with a size of h, and connects the fused feature vector with the face embedding to generate a joint embedding with a size of 2×h. The joint embedding is used as the input of the face decoder. The face decoder is composed of residual convolutional layers and transposed convolutional layers for upsampling.

[0020] The spatial attention mechanism SA has two input ends, which are the input of the face feature vector and the input of the audio feature vector respectively; in addition, it also includes an average pooling layer AvgPool, a max pooling layer MaxPool, a fusion function Concat, a two-dimensional convolution Conv2d with a convolution kernel size of 7×7, an activation function sigmoid, and an output end. After the vectors are input to the input ends, the face feature vectors are processed through the average pooling layer AvgPool and the max pooling layer MaxPool respectively, and the processed outputs are jointly input to the fusion function Concat for splicing processing in the channel dimension. The spliced vectors are then sequentially input to the two-dimensional convolution Conv2d and the activation function sigmoid for processing. The output of the activation function sigmoid is multiplied element-wise with the audio feature vector and then added together, and finally output to the output end. In particular, the dimension of the finally output vector is the same as that of the audio feature vector.

[0021] The spatial self-attention mechanism SSA includes an input end, an average pooling layer AvgPool, a max pooling layer MaxPool, a fusion function Concat, a two-dimensional convolution Conv2d with a convolution kernel size of 7×7, an activation function sigmoid, and an output end. After the feature vector is input to the input end, it is processed through the average pooling layer AvgPool and the max pooling layer MaxPool respectively, and then jointly input to the fusion function Concat for splicing in the channel dimension. The spliced vector is sequentially input to the two-dimensional convolution Conv2d with a convolution kernel size of 7×7 and the activation function Sigmoid for processing; the output of the Sigmoid layer is multiplied element-wise with the feature vector at the input end and then added together, and finally output to the output end.

[0022] The position attention mechanism CA includes an input end, two average pooling layers AvgPool, a fusion function Concat, a two-dimensional convolution Conv2d, two convolutional layers Conv, two activation functions Sigmoid, and an output end. After the feature vector is input to the input end, it passes through the two average pooling layers AvgPool in the width direction and the height direction respectively, and then the outputs are jointly input to the fusion function Concat for splicing in the channel dimension. The spliced vector is input to the two-dimensional convolution Conv2d. The output of the two-dimensional convolution Conv2d is divided into a feature vector of size H and a feature vector of size W in the width direction, and they are respectively input to the two convolutional layers Conv in the H direction and the W direction. The outputs of the two convolutional layers Conv are then respectively input to the two activation functions Sigmoid. The outputs of the two activation functions Sigmoid are multiplied with the feature vector at the input end and then added to the feature vector at the input end, and finally output to the output end.

[0023] The described video quality discriminator includes a number of convolutional groups connected in sequence, and each convolutional group includes a convolutional block and the activation function LeakyReLU connected in sequence. In specific implementation, the video quality discriminator includes 17 convolutional groups connected in sequence.

[0024] The overall loss function L all is specifically as follows:

[0025] L all = (1 - w s - w g - w f )L recon + w s ·L sync + w g ·L gen + w f ·L FSIM

[0026] Among them, w s , w g and w f respectively represent the synchronization penalty weight, the adversarial loss weight, and the FSIM loss weight; L recon , L sync , L gen and L FSIM respectively represent the reconstruction loss function, the discriminant loss function of the lip synchronization discriminant network after training is completed, the quality evaluation loss function of the lower half face including the lip shape of the video quality discriminator, and the quality evaluation loss index loss FSIM of the face dialogue generation model.

[0027] The discriminant loss function of the lip synchronization discriminant network is obtained according to the synchronization discriminant weight file. The generated lip shape and audio are input into the lip synchronization discriminant network loaded with the synchronization discriminant weight file to obtain the audio feature vector and the lip shape feature vector respectively, and then the discriminant loss function of the lip synchronization discriminant network is calculated by further using the similarity function.

[0028] The loss function generated by the video quality discriminator consists of two parts: a) the adversarial loss of the generator, that is, the original adversarial loss, and the generator hopes that the generated video is considered real by the discriminator; b) the adversarial loss of the discriminator, and the discriminator can correctly distinguish between real videos and generated videos.

[0029] The described reconstruction loss function L recon is a loss function composed of losses obtained according to the Manhattan L1 distance between each frame of face image corresponding in the generated face video and the real face video.

[0030] The quality evaluation loss function L gen of the lower half face including the lip shape of the video quality discriminator is specifically as follows:

[0031]

[0032] The adversarial loss function \(L\) of the lower half face including the lip shape of the described video quality discriminator disc is as follows:

[0033]

[0034] where \(x\) i~Lg represents the \(i\)-th number of the one-dimensional vector generated by inputting the generated face video into the video quality discriminator, \(N\) represents the total number of elements in the one-dimensional vector generated by the video quality discriminator, \(D()\) represents the probability that the video quality discriminator discriminates the input as true; \(x\) i~LG represents the \(i\)-th number of the one-dimensional vector generated by inputting the real face video into the video quality discriminator.

[0035] The quality evaluation loss index loss \(L\) of the described face dialogue generation model FSIM is as follows:

[0036]

[0037] where \(G\) g and \(G\) G respectively represent the gradient magnitudes of the corresponding video frames in the generated face video and the real face video obtained by using the Sobel operator, \(P\) g and \(P\) G respectively represent the corresponding video frame phase consistencies calculated according to the gradient magnitudes of the corresponding video frames in the generated face video and the real face video; \(T1\) and \(T2\) respectively represent the first and second adjustment parameters for controlling the weights of the phase consistency and the gradient magnitude; the feature similarity index (FSIM, feature similarity index measure) is the latter part of the quality evaluation loss index loss \(L\) FSIM . \(L\) represents the video frame in the generated face video or the real face video, Sobel x and Sobel y respectively represent the gradient maps of the Sobel operator in the horizontal \(X\) direction and the vertical \(Y\) direction of the image, and \(\epsilon\) represents a preset minimum value.

[0038] The FSLM loss \(L\) FSIM , for the image feature similarity picture quality evaluation of the video generated by the generator and the real video, aims to make some detailed features of the generated video tend to be similar to the real features.

[0039] In the described step S5), the dialogue system includes a natural language processing module and a ChatGPT module. After the text to be replied is input into the dialogue system, it is directly input into the ChatGPT module to generate the replied text. After the audio to be replied is input into the dialogue system, it is first input into the natural language processing module for processing and then converted into the text in the audio to be replied, and then input into the ChatGPT module to generate the replied text. Each segment of the replied text is input into the text-to-speech conversion model TTS (Text to Speech) for processing and then the replied audio is output. The TTS model is a model based on deep learning and can generate high-quality and natural speech.

[0040] The replied audio and the face image of the person to be generated are input into the face generator of the quality attention-based audio-driven lip network QAW that has been trained. After being processed by the trained face generator, the face video of the current person when reading the current replied audio is output.

[0041] The present invention designs a lip-sync discrimination network and a quality-attention-wav2lip-based audio-driven lip network. The method selects the first T frames of the pose frames as the reference frame F of the face decoder in the quality-attention-based audio-driven lip quality-attention-wav2lip network, making the face generation more continuous and reducing the appearance of exaggerated lip shapes and unnatural lower contours. The face feature dimensions of the input and output of the network are increased to 384×384 to ensure the clarity of lip shape generation. In addition, a spatial self-attention structure is added to the face encoder, and channel self-attention is increased to obtain sufficient and effective reference frame features; position attention is added to the face decoder part to encode the channel relationship and long-term dependence through accurate position information, improving the image quality and connection of the lower half of the face and the overall face features. For artifacts and unnatural lip shapes, the present invention introduces the FSLM loss for optimizing image generation and repair.

[0042] The beneficial effects of the present invention are:

[0043] The method of the present invention effectively improves the synchronization effect of lip shape generation and the image quality of the overall face, and can conduct a dialogue with the customer, aiming to generate a real face video with natural head movement and good lip-sync effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flowchart of the method of the present invention;

[0045] Figure 2 It is a structural diagram of the lip-sync discrimination network of the present invention;

[0046] Figure 3Structural diagram of the audio-driven lip network QAW based on quality attention of the present invention;

[0047] Figure 4 Structural diagrams of the spatial attention mechanism SA, spatial self-attention mechanism SSA, and position attention mechanism CA of the present invention. Detailed implementation manners

[0048] The following uses specific specific embodiments to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0049] As Figure 1 shown, the method for generating a face video of the audio-driven face dialogue generation model of the present invention is specifically as follows:

[0050] S1) Establish a lip-sync discrimination network and an audio-driven lip network QAW based on quality attention.

[0051] As Figure 2 shown, the lip-sync discrimination network includes an audio encoder and an image encoder. The audio encoder includes a plurality of convolutional layers Conv and a plurality of two-dimensional convolutions Conv2d connected in sequence. The image encoder includes a plurality of two-dimensional convolutions Conv2d and a plurality of convolutional layers Conv connected in sequence. Each convolutional layer Conv is composed of three two-dimensional convolutions Conv2d connected in sequence. Each two-dimensional convolution Conv2d includes a convolutional block, a batch normalization layer BN, and a ReLU activation function connected in sequence. For each group of human audio and consecutive-frame face images in the lip-sync training set, after converting the human audio into a mel spectrogram and inputting it into the audio encoder for processing, a first audio feature vector is output. The consecutive-frame face images are input into the image encoder to obtain a first lip feature vector. The cosine similarity method is used to obtain a similarity function between the first audio feature vector and the first lip feature vector to represent the matching degree between the audio and the video, specifically as follows:

[0052]

[0053] where Similarity(A,B) represents the similarity function between the first audio feature vector A and the first lip feature vector B; A i and B irespectively represent the i-th dimensional first audio feature vector output by the audio encoder and the i-th dimensional first lip feature vector output by the image encoder, where n = 1024.

[0054] The lip-sync discriminant network can help the face generator generate a more accurate lip-sync face video. In specific implementation, the audio encoder includes five consecutive convolutional layers Conv and two two-dimensional convolutional layers Conv2d. The dimension of the input audio is sequentially changed from (80, 16, 1) to (80, 16, 32), (27, 16, 64), (9, 6, 128), (3, 3, 256), (3, 3, 512), (1, 1, 1024), (1, 1, 1024); the image encoder includes one two-dimensional convolutional layer Conv2d and seven consecutive convolutional layers Conv. The dimension of the input image is sequentially changed from (384, 384, 3*T) to (384, 384, 16), (192, 192, 32), (96, 96, 64), (48, 48, 128), (24, 24, 256), (12, 12, 512), (6, 6, 1024), (1, 1, 1024), where T represents the duration of the person's audio and the number of frames of the face images in consecutive frames, T = 5, n = 1024.

[0055] As Figure 3 shown, the quality attention-based audio-driven lip network QAW includes a face generator and a video quality discriminator connected in sequence. The lip-sync training set is input into the face generator. After being processed by the face generator, the generated face video is output. The generated face video and the real face video are jointly input into the video quality discriminator to obtain the adversarial loss function and the quality evaluation loss function of the lower half of the face containing the lip shape of the video quality discriminator. When the adversarial loss function of the video quality discriminator converges, the video quality discriminator training is completed; the reconstruction loss function is obtained according to the generated face video and the real face video. Finally, the overall loss function of the quality attention-based audio-driven lip network QAW is constructed based on the discriminant loss function of the lip-sync discriminant network, the quality evaluation loss of the video quality discriminator, the FSIM loss function of the face dialogue generation model, and the reconstruction loss function.

[0056] The quality attention-based audio-driven lip network QAW extracts the sound waveform information from the audio and the lip movement information of the character from the video, and uses the generative adversarial network to train the model to obtain the mapping relationship between the two, so that the model can generate a series of lip movement sequences matching the input audio.

[0057] The face generator includes an audio encoder, a face encoder, and a face decoder. The audio encoder includes a first convolutional layer Conv, a second convolutional layer Conv, a third convolutional layer Conv, a fourth convolutional layer Conv, a fifth convolutional layer Conv, a sixth convolutional layer Conv, and a first two-dimensional convolutional layer Conv2d connected in sequence. Each convolutional layer Conv is composed of three two-dimensional convolutional layers Conv2d connected in sequence. Each two-dimensional convolutional layer Conv2d includes a convolutional block, a batch normalization layer BN, and a ReLU activation function connected in sequence. The face encoder includes a second two-dimensional convolutional layer Conv2d, a seventh convolutional layer Conv, an eighth convolutional layer Conv, a ninth convolutional layer Conv, a first spatial self-attention mechanism SSA, a tenth convolutional layer Conv, a second spatial self-attention mechanism SSA, an eleventh convolutional layer Conv, a twelfth convolutional layer Conv, a thirteenth convolutional layer Conv, a third two-dimensional convolutional layer Conv2d, and a fourth two-dimensional convolutional layer Conv2d connected in sequence. The face decoder includes a fifth two-dimensional convolutional layer Conv2d, a first spatial attention mechanism SA, a first fusion function Concat, a fourteenth convolutional layer Conv, a second spatial attention mechanism SA, a second fusion function Concat, a fifteenth convolutional layer Conv, a third spatial attention mechanism SA, a third fusion function Concat, a sixteenth convolutional layer Conv, a fourth spatial attention mechanism SA, a fourth fusion function Concat, a seventeenth convolutional layer Conv, a fifth spatial attention mechanism SA, a fifth fusion function Concat, an eighteenth convolutional layer Conv, a sixth spatial attention mechanism SA, a sixth fusion function Concat, a nineteenth convolutional layer Conv, a first position attention mechanism CA, a seventh spatial attention mechanism SA, a seventh fusion function Concat, a twentieth convolutional layer Conv, a second position attention mechanism CA, an eighth spatial attention mechanism SA, an eighth fusion function Concat, a twenty-first convolutional layer Conv, a third position attention mechanism CA, a ninth spatial attention mechanism SA, a ninth fusion function Concat, a sixth two-dimensional convolutional layer Conv2d, and a seventh two-dimensional convolutional layer Conv2d connected in sequence.

[0058] For each group of human audio and consecutive-frame face images in the lip synchronization training set, after converting the human audio into a Mel spectrogram and inputting it into the audio encoder of the face generator for processing to output a second audio feature vector, input the consecutive-frame face images into the face encoder for processing to output a second lip feature vector. Input the second audio feature vector into the fifth two-dimensional convolution Conv2d of the face decoder for processing. Input the output of the fifth two-dimensional convolution Conv2d processing and the second lip feature vector into the first spatial attention mechanism SA of the face encoder for processing. Input the output of the first spatial attention mechanism SA and the second lip feature vector into the first fusion function Concat for processing and then input it into the fourteenth convolutional layer Conv for processing. Input the output of the fourteenth convolutional layer Conv and the output of the thirteenth convolutional layer Conv into the second spatial attention mechanism SA for processing. Input the output of the second spatial attention mechanism SA and the output of the thirteenth convolutional layer Conv into the second fusion function Concat for processing and then output it into the fifteenth convolutional layer Conv for processing. Input the output of the fifteenth convolutional layer Conv and the output of the twelfth convolutional layer Conv into the third spatial attention mechanism SA for processing. Input the output of the third spatial attention mechanism SA and the output of the twelfth convolutional layer Conv into the third fusion function Concat for processing and then output it into the sixteenth convolutional layer Conv for processing. Input the output of the sixteenth convolutional layer Conv and the output of the eleventh convolutional layer Conv into the fourth spatial attention mechanism SA for processing. Input the output of the fourth spatial attention mechanism SA and the output of the eleventh convolutional layer Conv into the fourth fusion function Concat for processing and then output it into the seventeenth convolutional layer Conv for processing. Input the output of the seventeenth convolutional layer Conv and the output of the tenth convolutional layer Conv into the fifth spatial attention mechanism SA for processing. Input the output of the fifth spatial attention mechanism SA and the output of the tenth convolutional layer Conv into the fifth fusion function Concat for processing and then output it into the eighteenth convolutional layer Conv for processing. Input the output of the eighteenth convolutional layer Conv and the output of the ninth convolutional layer Conv into the sixth spatial attention mechanism SA for processing. Input the output of the sixth spatial attention mechanism SA and the output of the ninth convolutional layer Conv into the sixth fusion function Concat for processing and then sequentially output it into the nineteenth convolutional layer Conv and the first position attention mechanism CA for processing. Input the output of the first position attention mechanism CA and the output of the eighth convolutional layer Conv into the seventh spatial attention mechanism SA for processing. Input the output of the seventh spatial attention mechanism SA and the output of the eighth convolutional layer Conv into the seventh fusion function Concat for processing and then sequentially output it into the twentieth convolutional layer Conv and the second position attention mechanism CA for processing.The output of the second position attention mechanism CA and the output of the seventh convolutional layer Conv are jointly input into the eighth spatial attention mechanism SA for processing. The output of the eighth spatial attention mechanism SA and the output of the seventh convolutional layer Conv are jointly input into the eighth fusion function Concat for processing and then sequentially output to the twenty-first convolutional layer Conv and the third position attention mechanism CA for processing. The output of the third position attention mechanism CA and the output of the second two-dimensional convolution Conv2d are jointly input into the ninth spatial attention mechanism SA for processing. The output of the ninth spatial attention mechanism SA and the output of the second two-dimensional convolution Conv2d are jointly input into the ninth fusion function Concat for processing and then sequentially output to the sixth two-dimensional convolution Conv2d and the seventh two-dimensional convolution Conv2d for processing, and the final output is used as the output of the face generator.

[0059] a) The face encoder is composed of residual convolutional layers. Its main task is to encode the face reference frame F in the face video. The reference frame F obtained from the face video needs to be concatenated with the pose frame covering the lower half of the face in the channel dimension. In particular, to make the face generation more continuous and reduce the appearance of exaggerated lip shapes and unnatural lower contours, the reference frame generally uses the first T frames of the training frames. If there are less than T frames in front of the training frames, they are randomly selected. b) The audio encoder is a standard convolutional neural network, and its task is to encode the corresponding speech segment. Its input is the mel-frequency cepstral coefficient heat map MFCC with a size of M×T×1, and the output is an audio embedding with a size of h. c) The face decoder connects the feature vector generated by the audio encoder with a size of h, and connects the fused feature vector with the face embedding to generate a joint embedding with a size of 2×h. The joint embedding is used as the input of the face decoder. The face decoder is composed of residual convolutional layers and transposed convolutional layers for upsampling.

[0060] The dimensions of the audio input to the audio encoder change from (80, 16, 1) to (80, 16, 32), (27, 16, 64), (9, 6, 128), (3, 3, 256), (3, 3, 512), (1, 1, 1024), (1, 1, 1024) in sequence; the dimensions of the image input to the face encoder change from (384, 384, 6) to (384, 384, 8), (192, 192, 16), (96, 96, 32), (48, 48, 64), (24, 24, 128), (12, 12, 256), (6, 6, 512), (3, 3, 1024), (1, 1, 1024), (1, 1, 1024) in sequence; the dimensions of the input to the face decoder change from (1, 1, 1024) to (1, 1, 1024), (1, 1, 1024), (1, 1, 2048), (3, 3, 1024), (3, 3, 1024), (3, 3, 2048), (6, 6, 1024), (6, 6, 1024), (6, 6, 1536), (12, 12, 768), (12, 12, 768), (12, 12, 1024), (24, 24, 512), (24, 24, 512), (24, 24, 640), (48, 48, 256), (48, 48, 256), (48, 48, 320), (96, 96, 128), (96, 96, 128), (96, 96, 128), (96, 96, 160), (192, 192, 80), (192, 192, 80), (192, 192, 80), (192, 192, 96), (384, 384, 64), (384, 384, 64), (384, 384, 64), (384, 384, 72), (384, 384, 32), (384, 384, 3), (384, 384, 3) in sequence. The number in the third dimension of one image is 3, representing RGB, and the number in the third dimension is 6 for two images, that is, one is the lower half of the face without the face frame and reference frame for which the corresponding lip shape needs to be generated.

[0061] Such as Figure 4As shown in the figure, the spatial attention mechanism SA has two input ends, which are the input of the face feature vector and the input of the audio feature vector respectively; in addition, it also includes an average pooling layer AvgPool, a max pooling layer MaxPool, a fusion function Concat, a two-dimensional convolution Conv2d with a convolution kernel size of 7×7, an activation function sigmoid, and an output end. After the vector is input to the input end, the face feature vector is processed by the average pooling layer AvgPool and the max pooling layer MaxPool respectively, and the processed outputs are jointly input to the fusion function Concat for splicing processing in the channel dimension. The spliced vector is then sequentially input to the two-dimensional convolution Conv2d and the activation function sigmoid for processing. The output of the activation function sigmoid is multiplied element by element with the audio feature vector and then added, and finally output to the output end. In particular, the dimension of the finally output vector is the same as that of the audio feature vector.

[0062] The spatial self-attention mechanism SSA includes an input end, an average pooling layer AvgPool, a max pooling layer MaxPool, a fusion function Concat, a two-dimensional convolution Conv2d with a convolution kernel size of 7×7, an activation function sigmoid, and an output end. After the feature vector is input to the input end, it is processed by the average pooling layer AvgPool and the max pooling layer MaxPool respectively, and then jointly input to the fusion function Concat for splicing in the channel dimension. The spliced vector is sequentially input to the two-dimensional convolution Conv2d with a convolution kernel size of 7×7 and the activation function Sigmoid for processing; the output of the Sigmoid layer is multiplied element by element with the feature vector at the input end and then added, and finally output to the output end.

[0063] The position attention mechanism CA includes an input end, two average pooling layers AvgPool, a fusion function Concat, a two-dimensional convolution Conv2d, two convolutional layers Conv, two activation functions Sigmoid, and an output end. After the feature vector is input to the input end, it passes through the two average pooling layers AvgPool in the width direction and the height direction respectively, and then the outputs are jointly input to the fusion function Concat for splicing in the channel dimension. The spliced vector is input to the two-dimensional convolution Conv2d. The output of the two-dimensional convolution Conv2d is divided into a feature vector of size H and a feature vector of size W in the width direction, and is respectively input to the two convolutional layers Conv in the H direction and the W direction. The outputs of the two convolutional layers Conv are respectively input to the respective activation functions Sigmoid. The outputs of the two activation functions Sigmoid are multiplied with the feature vector at the input end and then added to the feature vector at the input end, and finally output to the output end.

[0064] The video quality discriminator includes a number of convolutional groups connected in sequence, and each convolutional group includes a convolutional block and the activation function LeakyReLU connected in sequence. Specifically, during implementation, the video quality discriminator includes 17 convolutional groups connected in sequence. When training the generator and the discriminator, the Adam optimizer is used, and the initial learning rate is 1E–4.

[0065] The overall loss function L all is specifically as follows:

[0066] L all = (1 - w s - w g - w f )L recon + w s ·L sync + w g ·L gen + w f ·L FSIM

[0067] where w s , w g and w f respectively represent the synchronization penalty weight, the adversarial loss weight, and the FSIM loss weight, and the default settings are 0.03, 0.04, and 0.03; L recon , L sync , L gen and L FSIM respectively represent the reconstruction loss function, the discriminant loss function of the lip synchronization discriminant network after training is completed, the quality evaluation loss function of the lower half face including the lip shape of the video quality discriminator, and the quality evaluation loss index loss FSIM of the face dialogue generation model.

[0068] The discriminant loss function of the lip synchronization discriminant network is obtained according to the synchronization discriminant weight file. The generated lip shape and audio are input into the lip synchronization discriminant network loaded with the synchronization discriminant weight file to obtain the audio feature vector and the lip shape feature vector respectively, and then the discriminant loss function of the lip synchronization discriminant network is calculated by further using the similarity function.

[0069] The loss function generated by the video quality discriminator consists of two parts: a) the adversarial loss of the generator, that is, the original adversarial loss, and the generator hopes that the generated video is considered real by the discriminator; b) the adversarial loss of the discriminator, and the discriminator can correctly distinguish between real videos and generated videos.

[0070] The reconstruction loss function L recon is a loss function composed of losses obtained based on the Manhattan L1 distance between each frame of face image corresponding in the generated face video and the real face video.

[0071] The quality assessment loss function L of the lower half face including lip shape of the video quality discriminator gen Specifically as follows:

[0072]

[0073] The adversarial loss function L of the lower half face including lip shape of the video quality discriminator disc Specifically as follows:

[0074]

[0075] Among them, x i~Lg represents the i-th number of the one-dimensional vector generated by inputting the generated face video into the video quality discriminator, N represents the total number of elements in the one-dimensional vector generated by the video quality discriminator, D() represents the probability that the video quality discriminator discriminates the input as true; x i~LG represents the i-th number of the one-dimensional vector generated by inputting the real face video into the video quality discriminator.

[0076] The quality assessment loss index loss L of the face dialogue generation model FSIM Specifically as follows:

[0077]

[0078]

[0079] Among them, G g and G G respectively represent the gradient magnitudes of the corresponding video frames in the generated face video and the real face video obtained by using the Sobel operator, P g and P G respectively represent the corresponding video frame phase consistencies calculated according to the gradient magnitudes of the corresponding video frames in the generated face video and the real face video; T1 and T2 respectively represent the first and second adjustment parameters for controlling the weights of the phase consistency and the gradient magnitude, and the default settings are 0.85 and 160; the feature similarity index (FSIM, feature similarity index measure) is the latter half of the quality assessment loss index loss L FSIM . L represents the video frame in the generated face video or the real face video, Sobel x and Sobel y respectively represent the gradient maps of the Sobel operator in the horizontal X direction and the vertical Y direction of the image, ∈ represents a preset minimum value.

[0080] The FSLM loss L FSIM, the image quality evaluation of the similarity of image features between the video generated by the generator and the real video is carried out, aiming to make some detailed features of the generated video tend to be similar to the real features.

[0081] S2) Collect videos of several people speaking and divide each video into the person's audio and consecutive-frame face images, and jointly construct a lip-sync training set from each group of person's audio and consecutive-frame face images.

[0082] S3) Input the lip-sync training set into the lip-sync discriminant network for training to make it more accurate in discriminating lip-sync faces until the discriminant loss function of the lip-sync discriminant network converges, and obtain the trained lip-sync discriminant network. The convergence of the discriminant loss function of the lip-sync discriminant network is constructed according to the synchronous discriminant weight file and similarity function output by the lip-sync discriminant network during training.

[0083] S4) Input the lip-sync training set into the quality attention-based audio-driven lip shape network QAW for training, construct the overall loss function of the quality attention-based audio-driven lip shape network QAW based on the discriminant loss function, obtain the overall loss value of the quality attention-based audio-driven lip shape network QAW and use the gradient descent method to update the network parameters until the loss value of the overall loss function converges, and obtain the trained quality attention-based audio-driven lip shape network QAW.

[0084] S5) Input the text or audio to be replied into the dialogue system. For each piece of text or audio to be replied, the dialogue system processes it and outputs the reply text of the text or audio to be replied, and converts the reply text into a reply audio; input the reply audio and the face image of the person to be generated into the trained quality attention-based audio-driven lip shape network QAW. The trained quality attention-based audio-driven lip shape network QAW processes it and outputs the face video of the current person when reading the current reply audio, and finally displays it on the monitor to realize the generation of the face video.

[0085] The dialogue system includes a natural language processing module and a chatgpt module. After the text to be replied is input into the dialogue system, it is directly input into the chatgpt module to generate the reply text. After the audio to be replied is input into the dialogue system, it is first input into the natural language processing module for processing and converted into the text in the audio to be replied, and then input into the chatgpt module to generate the reply text; input each piece of reply text into the text-to-speech conversion model TTS for processing and output the reply audio. The TTS model is a deep learning-based model that can generate high-quality and natural speech.

[0086] Input the response audio and the face image of the person to be generated into the face generator of the quality attention-based audio-driven lip network QAW that has been trained. After being processed by the trained face generator, a face video of the current person when reading the current response audio is output.

[0087] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for generating face videos based on an audio-driven face dialogue generation model, characterized in that: include: S1) Establishing a lip synchronization discrimination network and a quality attention-based audio-driven lip shape network QAW; S2) collecting videos of several people speaking and dividing each video into person audio and continuous frames of face images, and constructing each group of person audio and continuous frames of face images together as a lip synchronization training set; S3) inputting the lip synchronization training set into the lip synchronization discrimination network for training until the discrimination loss function of the lip synchronization discrimination network converges, thereby obtaining a trained lip synchronization discrimination network; S4) inputting the lip synchronization training set into the quality-attention-based audio-driven lip shape network QAW for training, constructing the overall loss function of the quality-attention-based audio-driven lip shape network QAW based on the discriminant loss function, obtaining the overall loss value of the quality-attention-based audio-driven lip shape network QAW and using the gradient descent method to update the network parameters until the loss value of the overall loss function converges, thereby obtaining the trained quality-attention-based audio-driven lip shape network QAW; S5) inputting the text or audio to be replied into the dialogue system, and for each piece of text or audio to be replied, the dialogue system processes and outputs a reply text of the text or audio to be replied, and converts the reply text into a reply audio; Input the reply audio and the facial image of the person to be generated into the trained quality-attention-based audio-driven lip network QAW, and the trained quality-attention-based audio-driven lip network QAW processes and outputs the facial video of the current person when reading the current reply audio, and finally displays it on the display to achieve the generation of the facial video; In the step S1), the audio-driven lip shape network QAW based on quality attention includes a face generator and a video quality discriminator connected in sequence, the lip synchronization training set is input into the face generator, the face generator outputs the generated face video after processing, the generated face video and the real face video are input into the video quality discriminator together, the adversarial loss function and the quality assessment loss function of the video quality discriminator are obtained, and when the adversarial loss function of the video quality discriminator converges, the video quality discriminator training is completed; the reconstruction loss function is obtained according to the generated face video and the real face video, and finally the overall loss function of the audio-driven lip shape network QAW based on quality attention is constructed according to the discriminant loss function of the lip synchronization discriminant network, the quality assessment loss of the video quality discriminator, the FSIM loss function of the face dialogue generation model and the reconstruction loss function; The face generator consists of an audio encoder, a face encoder and a face decoder. The face encoder contains nine spatial attention mechanisms SA and three position attention mechanisms CA.

2. The method for generating face video based on the audio-driven face dialogue generation model according to claim 1, characterized in that: In the step S1), the lip synchronization discrimination network includes an audio encoder and an image encoder, the audio encoder includes a plurality of convolutional layers Conv and a plurality of two-dimensional convolutional layers Conv2d connected in sequence, the image encoder includes a plurality of two-dimensional convolutional layers Conv2d and a plurality of convolutional layers Conv connected in sequence, each convolutional layer Conv is composed of three two-dimensional convolutional layers Conv2d connected in sequence, and each two-dimensional convolutional layer Conv2d includes a convolutional block, a batch normalization layer BN and a ReLU activation function connected in sequence; for each group of character audio and continuous frame facial images in the lip synchronization training set, the character audio is converted into a Mel-spectrogram and then input into the audio encoder for processing to output a first audio feature vector, and the continuous frame facial images are input into the image encoder to obtain a first lip shape feature vector.

3. The method for generating face video based on the audio-driven face dialogue generation model according to claim 1, characterized in that: The audio encoder includes a first convolution layer Conv, a second convolution layer Conv, a third convolution layer Conv, a fourth convolution layer Conv, a fifth convolution layer Conv, a sixth convolution layer Conv, and a first two-dimensional convolution Conv2d connected in sequence, each convolution layer Conv is composed of three two-dimensional convolution Conv2d connected in sequence, and each two-dimensional convolution Conv2d includes a convolution block, a batch normalization layer BN and a ReLU activation function connected in sequence; the face encoder includes a second two-dimensional convolution Conv2d, a seventh convolution layer Conv, an eighth convolution layer Conv, a ninth convolution layer Conv, a first spatial self-attention mechanism SSA, a tenth convolution layer Conv, a second spatial self-attention mechanism SSA, an eleventh convolution layer Conv, a twelfth convolution layer Conv, a thirteenth convolution layer Conv, a third two-dimensional convolution Conv2d and a fourth two-dimensional convolution Conv2d connected in sequence; the face decoder includes a fifth two-dimensional convolution Conv2d, a first spatial attention mechanism SA, a first fusion function Co connected in sequence ncat, the fourteenth convolution layer Conv, the second spatial attention mechanism SA, the second fusion function Concat, the fifteenth convolution layer Conv, the third spatial attention mechanism SA, the third fusion function Concat, the sixteenth convolution layer Conv, the fourth spatial attention mechanism SA, the fourth fusion function Concat, the seventeenth convolution layer Conv, the fifth spatial attention mechanism SA, the fifth fusion function Concat, the eighteenth convolution layer Conv, the sixth spatial attention mechanism SA, the sixth fusion function Concat, the nineteenth convolution layer Conv, the first position attention mechanism CA, the seventh spatial attention mechanism SA, the seventh fusion function Concat, the twentieth convolution layer Conv, the second position attention mechanism CA, the eighth spatial attention mechanism SA, the eighth fusion function Concat, the twenty-first convolution layer Conv, the third position attention mechanism CA, the ninth spatial attention mechanism SA, the ninth fusion function Concat, the sixth two-dimensional convolution Conv2d and the seventh two-dimensional convolution Conv2d; For each group of character audio and continuous frame face images in the lip synchronization training set, the character audio is converted into a Mel-spectrogram and then input into the audio encoder of the face generator for processing to output a second audio feature vector, the continuous frame face images are input into the face encoder for processing to output a second lip shape feature vector, the second audio feature vector is input into the fifth two-dimensional convolution Conv2d of the face decoder for processing, the output of the fifth two-dimensional convolution Conv2d and the second lip shape feature vector are input into the first spatial attention mechanism SA of the face encoder for processing, the output of the first spatial attention mechanism SA and the second lip shape feature vector are input into the first fusion function Concat for processing and then input into the fourteenth convolution layer Conv for processing, the output of the fourteenth convolutional layer Conv and the output of the thirteenth convolutional layer Conv are input into the second spatial attention mechanism SA for processing, the output of the second spatial attention mechanism SA and the output of the thirteenth convolutional layer Conv are input into the second fusion function Concat for processing and then output to the fifteenth convolutional layer Conv for processing, the output of the fifteenth convolutional layer Conv and the output of the twelfth convolutional layer Conv are input into the third spatial attention mechanism SA for processing, the output of the third spatial attention mechanism SA and the output of the twelfth convolutional layer Conv are input into the third fusion function Concat for processing and then output to the sixteenth convolutional layer Conv for processing, The output of the sixteenth convolutional layer Conv and the output of the eleventh convolutional layer Conv are input into the fourth spatial attention mechanism SA for processing. The output of the fourth spatial attention mechanism SA and the output of the eleventh convolutional layer Conv are input into the fourth fusion function Concat for processing and then output to the seventeenth convolutional layer Conv for processing. The output of the seventeenth convolutional layer Conv and the output of the tenth convolutional layer Conv are input into the fifth spatial attention mechanism SA for processing. The output of the fifth spatial attention mechanism SA and the output of the tenth convolutional layer Conv are input into the fifth fusion function Concat for processing and then output to the eighteenth convolutional layer Conv for processing. The output of the eighteenth convolutional layer Conv The output of the sixth spatial attention mechanism SA and the output of the ninth convolutional layer Conv are input together to the sixth spatial attention mechanism SA for processing. The output of the sixth spatial attention mechanism SA and the output of the ninth convolutional layer Conv are input together to the sixth fusion function Concat for processing and then output to the nineteenth convolutional layer Conv and the first position attention mechanism CA for processing in sequence. The output of the first position attention mechanism CA and the output of the eighth convolutional layer Conv are input together to the seventh spatial attention mechanism SA for processing. The output of the seventh spatial attention mechanism SA and the output of the eighth convolutional layer Conv are input together to the seventh fusion function Concat for processing and then output to the twentieth convolutional layer Conv and the second position attention mechanism CA for processing.The output of the second position attention mechanism CA and the output of the seventh convolution layer Conv are input into the eighth spatial attention mechanism SA for processing. The output of the eighth spatial attention mechanism SA and the output of the seventh convolution layer Conv are input into the eighth fusion function Concat for processing and then output to the twenty-first convolution layer Conv and the third position attention mechanism CA for processing. The output of the third position attention mechanism CA and the output of the second two-dimensional convolution Conv2d are input into the ninth spatial attention mechanism SA for processing. The output of the ninth spatial attention mechanism SA and the output of the second two-dimensional convolution Conv2d are input into the ninth fusion function Concat for processing and then output to the sixth two-dimensional convolution Conv2d and the seventh two-dimensional convolution Conv2d for processing, and the final output is used as the output of the face generator. , 4. The method for generating face video based on the audio-driven face dialogue generation model according to claim 1, characterized in that: The video quality discriminator includes a plurality of convolution groups connected in sequence, and each convolution group includes convolution blocks and an activation function LeakyReLU connected in sequence.

5. The method for generating face video based on the audio-driven face dialogue generation model according to claim 1, characterized in that: The overall loss function L all The details are as follows: L all =(1-w s -w g -w f )L recon +w s ·L sync +w g ·L gen +w f ·L FSIM Among them, w s 、w g and w f They represent the synchronization penalty weight, adversarial loss weight and FSIM loss weight respectively; L recon , L sync , L gen and L FSIM They respectively represent the reconstruction loss function, the discriminant loss function of the trained lip synchronization discriminant network, the quality assessment loss function of the video quality discriminator, and the quality assessment loss index loss FSIM of the face dialogue generation model.

6. The method for generating face video based on the audio-driven face dialogue generation model according to claim 5, characterized in that: The reconstruction loss function L recon The loss function is composed of the loss obtained based on the Manhattan L1 distance between the generated face video and the corresponding face image of each frame in the real face video.

7. The method for generating face video based on the audio-driven face dialogue generation model according to claim 5, characterized in that: The quality assessment loss function L of the video quality discriminator is gen The details are as follows: The adversarial loss function L of the video quality discriminator is disc The details are as follows: Among them, x i~Lg represents the i-th number of the one-dimensional vector generated by the generated face video input to the video quality discriminator, N represents the total number of elements in the one-dimensional vector generated by the video quality discriminator, and D() represents the probability that the video quality discriminator judges the input to be true; x i~LG Represents the i-th number of the one-dimensional vector generated by the real face video input to the video quality discriminator.

8. The method for generating face video based on the audio-driven face dialogue generation model according to claim 5, characterized in that: The quality assessment loss index loss L of the face dialogue generation model FSIM The details are as follows: Among them, G g and G G They represent the gradient amplitude of the corresponding video frames in the generated face video and the real face video calculated using the Sobel operator, respectively. g and P G They respectively represent the corresponding video frame phase consistency calculated according to the gradient amplitude of the corresponding video frames in the generated face video and the real face video; T1 and T2 respectively represent the first and second mediation parameters.

9. The method for generating face video based on the audio-driven face dialogue generation model according to claim 1, characterized in that: In the step S5), the dialogue system includes a natural language processing module and a chatgpt module. After the text to be replied is input into the dialogue system, it is directly input into the chatgpt module to generate the reply text. After the audio to be replied is input into the dialogue system, it is first input into the natural language processing module for processing and then converted into the text in the audio to be replied, and then input into the chatgpt module to generate the reply text; each reply text is input into the text-to-speech conversion model TTS for processing and then the reply audio is output; The reply audio and the facial image of the character to be generated are input into the trained quality attention-based audio-driven lip network QAW face generator. After processing, the trained face generator outputs the facial video of the current character when reading the current reply audio.

Citation Information

Patent Citations

  • Generator training method of digital human generation model and digital human generation method and device

    CN117456062A

  • Human-machine interaction

    US20210280190A1