Voice-driven high-effect digital population type synthesis algorithm

By adopting a high-effect algorithm driven by voice in digital lip synthesis technology, combining multi-head self-attention module and generative adversarial network, problems such as inaccurate lip sync and inaccurate movements in the existing technology are solved, and multilingual support with high precision and low computing cost and natural facial movements are achieved.

CN120034700APending Publication Date: 2025-05-23GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510191308.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing digital lip synthesis technology has problems such as inaccurate lip synchronization, inaccurate lip movements, blurred facial images, high calculation costs, lack of multilingual support and unnatural teeth and tongue.

Method used

Using a high-effect digital verbal synthesis algorithm driven by voice, the model structure is designed through audio encoder, image encoder, audio feature selector, feature fusion module and image decoder, combined with multi-head self-attention module and generation adversarial network, many-to-many two-way fusion of audio and image features is achieved.

Benefits of technology

It improves the accuracy of lip synchronization and movement accuracy, the generated facial images are clear, supports multilingual lip drive, reduces calculation costs, and achieves the natural and consistent teeth and tongue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034700A_ABST
    Figure CN120034700A_ABST
Patent Text Reader

Abstract

The invention relates to a voice-driven high-effect digital population type synthesis algorithm. The digital population type synthesis effect is improved by introducing lip-reading expert, redesigned lip-sync expert, an innovative reference frame selection strategy, a well-designed bidirectional feature fusion module, a training loss function and other skills. And the device has the functions of controllable mouth opening amplitude and multi-language support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lip synthesis, and in particular to a high-effect digital mouth synthesis algorithm driven by speech. Background Art

[0002] Digital humans can be applied to customer service digital humans, teacher digital humans, virtual anchors, and digital immortality. They can be applied to various industries such as the film and television production industry, anchor industry, advertising industry, and education industry.

[0003] In the prior art, digital population synthesis has the following deficiencies:

[0004] (1) Inaccurate lip synchronization: The lip synchronization synthesized by existing methods is poor, such as opening the mouth prematurely.

[0005] (2) Inaccurate lip movements: The lip shapes synthesized by existing methods do not match the lip movements during pronunciation. For example, the pronunciation of some words requires pouting, but the synthesized lip shapes do not have pouting.

[0006] (3) Unable to control the mouth opening amplitude: Sometimes the input audio volume is relatively low, resulting in the digital human opening its mouth with a relatively small amplitude.

[0007] (4) The image is blurry: Existing methods are unable to synthesize fine facial skin texture, causing the image to appear blurry.

[0008] (5) High computational cost: Existing methods generally use a large number of transformers or diffusion modules, which are time-consuming and computationally intensive.

[0009] (6) Lack of multi-language support: The existing methods can only drive a model in one language, and it is impossible to share the same model in multiple languages.

[0010] (7) The existing lip-synthesis of teeth and tongue is unnatural because the existing algorithms usually only use one frame of reference image. If the teeth are not exposed in the reference image, the model needs to guess what the person's teeth look like, and there is no guarantee that the teeth between different frames are consistent.

[0011] Therefore, it is necessary to provide a voice-driven, high-efficiency digital population synthesis algorithm to improve the effect of digital population synthesis in business scenarios. Summary of the invention

[0012] The purpose of the present invention is to provide a voice-driven high-efficiency digital population synthesis algorithm to improve the effect of digital population synthesis in business scenarios.

[0013] In order to solve the problems existing in the prior art, the present invention provides a high-effect digital population synthesis algorithm driven by speech, comprising the following steps:

[0014] S1: Obtain and preprocess the dataset;

[0015] S2: Design the model structure using audio encoder, image encoder, audio feature selector, feature fusion module and image decoder;

[0016] S3: Pytorch is used as the training framework to train the digital population synthesis model, where the calculation formula of the model's loss function is as follows:

[0017] L total =λ s ×L sync +λ c ×L lip +λ r ×L rec +λ p ×L p +λ g ×L gan ;

[0018] Among them, L total is the loss function value of the model, L sync is the lip synchronization loss value, λ s is the weight of lip synchronization loss, L lip is the lip reading loss, λ c is the weight of lip reading loss, L rec is the reconstruction loss value, λ r is the weight of the reconstruction loss, L p is the perceptual loss value, λ p is the weight of perceptual loss, L gan is the loss value of the generated adversarial network, λ g is the weight of the GAN loss.

[0019] Optionally, in the voice-driven high-effect digital population synthesis algorithm, the data set is obtained as follows:

[0020] Record several hours of talking videos of people, covering different ages, genders, face shapes and skin colors, with a resolution of 720P and a frame rate of 25FPS. Cut the long videos into 5-second short video clips, extract the audio of each short video clip, and obtain a standard digital audio file with a sampling rate of 16K.

[0021] Optionally, in the voice-driven high-effect digital population synthesis algorithm, the data set is preprocessed as follows:

[0022] The coordinates of the key points of the face in each frame of the short video clip are extracted, and the coordinates of the key points are used to calculate an affine transformation matrix that can make the face stand at attention. The affine transformation matrix is ​​applied to the frame to obtain a picture with the face standing at attention, the coordinates of the face rectangular frame are detected in the picture, and a 256×256 face image is cropped from the picture using the coordinates of the face rectangular frame.

[0023] Optionally, in the voice-driven high-effect digital population synthesis algorithm,

[0024] The audio encoder inputs a standard digital audio file and outputs audio features with a tensor shape of [B,L,D], where B is the batch size, L is the time length, and D is the feature dimension.

[0025] Optionally, in the speech-driven high-effect digital population synthesis algorithm, the image encoder is stacked by several layers of two-dimensional convolution, several groups of normalization and several activation functions, the number of groups is R+1, R is the number of input reference pictures, and downsampling is performed by setting a step size of 2 in some two-dimensional convolution layers;

[0026] The input of the image encoder is R+1 images concatenated in the channel dimension, and the output is the image feature with a tensor shape of [B,H,W,C], where B represents the number of inputs, H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map.

[0027] Optionally, in the voice-driven high-effect digital population synthesis algorithm,

[0028] The audio feature selector selects some audio features from the audio features output by the audio encoder to obtain audio features with a tensor shape of [B, F, D], where F is the selection parameter.

[0029] Optionally, in the voice-driven high-effect digital population synthesis algorithm, the feature fusion module works as follows:

[0030] For the audio feature whose tensor shape is [B, F, D], the tensor of the audio feature is multiplied by the mouth opening amplitude coefficient;

[0031] Use an adapter to map audio domain features to a common space with image features;

[0032] The tensor is sinusoidally absolute position encoded;

[0033] Input the tensor encoded by the sinusoidal absolute position into a multi-head self-attention module, which is used to learn audio context information at different times. The fully connected layer in the multi-head self-attention module maps the audio features into Query, Key and Value;

[0034] Perform the attention function operation on Query, Key and Value as follows:

[0035] Among them, Q is Query, K is Key, V is Value, T is the transposition operation, and G is the number of channels for each token;

[0036] Then, residual connection addition and layer normalization are performed, and the resulting tensor is used as the query of the image-to-audio cross-attention module;

[0037] For image features, the tensor shape of [B, H, W, C] is reshaped from [B, H, W, C] to [B, H × W, C]. The changed tensor shape is used as the key and value of the image-to-audio cross-attention module to enable the image features to extract information from the audio features and realize multimodal information fusion.

[0038] Optionally, in the voice-driven high-effect digital population synthesis algorithm, the model structure of the image decoder is composed of M upsampling, M convolutions and M SPADEResnetBlocks stacked together, where M is a natural number, and is also connected with a Leaky ReLU function, a Conv2d function and a Sigmoid function.

[0039] Optionally, in the voice-driven high-efficiency digital population synthesis algorithm, pytorch is used as a training framework to train the digital population synthesis model in the following manner:

[0040] The training framework uses pytorch, the batch size is set to 16, the optimizer uses Adam, the learning rate regulator uses CosineAnnealingLR, and the initial learning rate is set to 1×10^-4;

[0041] Before formal training, the learning rate gradually increases from zero to the initial learning rate;

[0042] During training, the parameters of the lip-sync expert and lip-reading expert are frozen, and data augmentation uses a strategy of randomly changing the gamma and brightness of the image.

[0043] Optionally, in the voice-driven high-effect digital population synthesis algorithm,

[0044] The expression for the lip sync loss value is as follows:

[0045]

[0046] Where a is the audio vector of the S-frame audio Mel-spectrogram after lip-sync expert encoding, v is the image vector of the lower half of the image generated by the S-frame model after lip-sync expert encoding, S is a natural number, and ε is a small positive constant to prevent the denominator from being zero;

[0047] The expression of lip reading loss value is as follows:

[0048]

[0049] Among them, Y is the real text, is the text predicted by the lip reading expert based on the given video V, and P is the probability;

[0050] The expression of the reconstruction loss value is as follows:

[0051]

[0052] Where N is the number of elements in the tensor, i is the element index in the tensor, and v is the value of each element in the tensor;

[0053] The expression of the perceptual loss value is as follows:

[0054]

[0055] Among them, Vgg i represents the i-th layer in the VGG-19 network, Io is the frame image output by the model, Ir is the real reference frame image, and W i Represents the width of the i-th layer tensor, H i Represents the height of the i-th layer tensor, C i Represents the number of channels of the i-th layer tensor;

[0056] The expression for the loss value of the generated adversarial network is as follows:

[0057]

[0058] Among them, D j represents the GAN discriminator, E represents the expected value, Io is the frame image output by the model, and Ir is the real reference frame image.

[0059] Compared with the prior art, the present invention has the following advantages:

[0060] (1) Improved lip sync accuracy: Audio and lip sync can be synchronized more accurately.

[0061] (2) Accurate lip movements: Ability to synthesize lip movements that are consistent with pronunciation.

[0062] (3) Clear facial image: Able to synthesize clear skin texture material that is consistent with the skin in the input image.

[0063] (4) Multi-language support: The same model can support lip-syncing in multiple languages ​​at the same time, such as Chinese, English, Japanese, and French, and can easily handle multiple languages ​​mixed in the input audio.

[0064] (5) Low computational overhead: Through a carefully designed model architecture, the computational cost of the model is controlled, so that the technology can be implemented in real time on edge devices and is suitable for resource-constrained devices such as mobile phones and tablets.

[0065] (6) Broad application prospects: This algorithm can be widely used in digital humans from all walks of life (including but not limited to customer service, anchors, teachers, relatives, self-media bloggers, advertising spokespersons) and has high commercial potential.

[0066] (7) The teeth and tongue are natural and the teeth are consistent between different frames. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 The overall architecture diagram of the model provided by the embodiment of the present invention;

[0068] Figure 2 A structural diagram of a feature fusion module provided in an embodiment of the present invention;

[0069] Figure 3 A structural diagram of an audio encoder provided by an embodiment of the present invention;

[0070] Figure 4 A structural diagram of an image decoder provided by an embodiment of the present invention;

[0071] Figure 5 A structural diagram of SPADEResnetBlock provided in an embodiment of the present invention;

[0072] Figure 6 A structural diagram of a SPADE provided in an embodiment of the present invention;

[0073] Figure 7 A reference frame selection strategy flow chart provided by an embodiment of the present invention;

[0074] Figure 8 A structural diagram of syncXnet provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0075] The specific implementation of the present invention will be described in more detail below in conjunction with the schematic diagram. The advantages and features of the present invention will become clearer based on the following description. It should be noted that the drawings are all in a very simplified form and are not in exact proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.

[0076] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0077] In the prior art, digital population synthesis has many shortcomings.

[0078] In order to solve the problems existing in the prior art, the present invention provides a high-effect digital population synthesis algorithm driven by speech, comprising the following steps:

[0079] S1: Obtain and preprocess the dataset;

[0080] Specifically, S11: the method of obtaining the data set is as follows:

[0081] Record several hours of talking videos of people, covering different ages, genders, face shapes and skin colors, with a resolution of 720P and a frame rate of 25FPS. Cut the long videos into 5-second short video clips, extract the audio of each short video clip, and obtain a standard digital audio file with a sampling rate of 16K.

[0082] S12: The way to preprocess the dataset is as follows:

[0083] Use mediapipe to extract the coordinates of the key points of the face in each frame of the short video clip, use the key point coordinates to calculate the affine transformation matrix that can make the face stand at attention, apply the affine transformation matrix to the frame to obtain the picture after the face stands at attention, use yolov8-face to detect the coordinates of the face rectangular frame in the picture, and use the face rectangular frame coordinates to crop a 256×256 face image from the picture.

[0084] S2: Design the model structure using audio encoder, image encoder, audio feature selector, feature fusion module and image decoder, such as Figure 1 As shown;

[0085] Specifically, (1) the audio encoder inputs a standard digital audio file and outputs audio features with a tensor shape of [B, L, D], where B is the batch size, L is the time length, and D is the feature dimension. The audio encoder can choose a pre-trained ASR model, such as wav2vec, whisper, HuBERT, etc. By using speech recognition models of different languages, it can recognize lip-syncing in multiple languages. You can also build an audio encoder and train it from scratch, such as Figure 3 The present invention uses an audio encoder pre-trained for multi-language speech recognition, so that the algorithm of the present invention can handle input audio in multiple languages ​​and maintain a certain accuracy and stability.

[0086] (2) The image encoder is composed of several layers of two-dimensional convolution, several groups of normalization and several activation functions stacked together. The number of groups is R+1, where R is the number of input reference images. In the schematic diagram of this article, R=3. Downsampling is performed by setting the stride to 2 in some two-dimensional convolution layers.

[0087] The input of the image encoder is R+1 images concatenated in the channel dimension, and the output is the image features of the tensor shape [B,H,W,C], where B represents the number of inputs, H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map;

[0088] In one embodiment, the input of the image encoder is (R+1) 256×256×3 images concatenated in the channel dimension, with a tensor shape of [B,H,W,(R+1)×3], and the output is an image feature with a tensor shape of [B,H,W,C]. Among the (R+1) input images, R images are reference frames selected by the reference frame selection strategy specially designed in the preprocessing stage. Its function is to provide references for different mouth shapes for the model. The last image is the frame I_i with the lower half set to zero. Its function is to provide a priori of the head posture. The main task of the model is to synthesize the lower half of the face with the correct mouth shape. In the training stage, its input is S groups of continuous frames, each group of continuous frames is the same as the inference stage, and the purpose of inputting S groups is that the model will also output S images at the end for the lip-sync expert to calculate the lip sync loss.

[0089] (3) The audio feature selector selects some audio features from the audio features output by the audio encoder to obtain audio features with a tensor shape of [B, F, D], where F is a selection parameter. Generally, the selection strategy can have multiple forms.

[0090] In one embodiment, for example, different spacings and different lengths, assuming that the index of the current frame is i, and the audio feature sequence is {a_0, a_1, …, a_{i-1}, a_i, a_{i+1}, …, a_{L}}, then the selected audio feature sequence can be a sequence with a spacing of 1 and a length of 11: {a_{i-5}, …, a_{i-1}, a_i, a_{i+1}, …, a_{i+5}}, or a sequence with a spacing of 2 and a length of 11: {a_{i-10}, …, a_{i-2}, a_i, a_{i+2}, …, a_{i+10}}, and so on.

[0091] (4) Figure 2 As shown in Figure 2, the feature fusion module works as follows:

[0092] Step a: For the audio features selected to obtain a tensor shape of [B, F, D], the tensor of the audio features is multiplied by the mouth opening amplitude coefficient. The mouth opening amplitude coefficient is usually 1. When the mouth opening amplitude coefficient is greater than 1, the mouth opening amplitude of the final output of the model will be larger, otherwise the mouth opening amplitude of the final output of the model will be smaller. The present invention supports inputting a parameter of a controllable mouth opening amplitude to realize the control of the mouth opening amplitude of the lip shape.

[0093] Step b: Use an adapter to map the audio domain features to a common space with the image features. The adapter is mainly composed of several layers of fully connected layers and activation functions stacked together;

[0094] Step c: The tensor is encoded with sinusoidal absolute position, the formula is as follows:

[0095]

[0096] Among them, PE stands for position encoding, pos is the position index of the token in the sequence, i is the index of the channel dimension, and d model is the number of channels.

[0097] Step d: Input the tensor encoded by the sinusoidal absolute position into a multi-head self-attention (MHSA) module, which is used to learn audio context information at different times. The fully connected layer in the multi-head self-attention module maps the audio features into Query, Key and Value;

[0098] Step e: Perform the attention function operation on Query, Key and Value as follows:

[0099] Where Q is Query, K is Key, V is Value, T is transpose operation, and G is the number of channels for each token (for example, Q is an E×G matrix, that is, the matrix has E rows and G columns, E is the number of tokens, and G is the number of channels for each token);

[0100] Step f: Perform residual connection addition and layer normalization, and the resulting tensor is used as the query of the image-to-audio cross-attention module;

[0101] Step g: For image features, the image features with a tensor shape of [B, H, W, C] are reshaped, and the tensor shape changes from [B, H, W, C] to [B, H×W, C]. The changed tensor shape is used as the key and value of the image-to-audio cross-attention module to enable the image features to extract information from the audio features and achieve multimodal information fusion. Note that the input of this module also has a Mask, which sets the elements of the upper half of the H×W dimension of the image feature tensor to negative infinity, that is, [B,:H×W / 2,C]=-inf. The function of this module is to allow each audio feature to extract information only from the lower half of the image features of multiple reference images.

[0102] (5) The model structure of the image decoder is as follows Figure 4 As shown in Figure 1, the model structure of the image decoder is composed of M upsampling, M convolution and M SPADEResnetBlock stacks, where M is a natural number, for example, M is 5, and is also connected with a Leaky ReLU function, a Conv2d function and a Sigmoid function. The internal structure of SPADEResnetBlock is as follows: Figure 5 As shown, SPADEResnetBlock contains SPADE, and the internal structure of SPADE is as follows Figure 6 In order to reduce the amount of calculation, SPADEResnetBlock can be replaced with convolution, and the output of the image decoder is a 256×256×3 image.

[0103] The bidirectional modal fusion module proposed in the present invention performs many-to-many, bidirectional feature fusion of audio features and image features, which improves the synchronization and accuracy of lip synthesis, and also improves the naturalness and consistency of teeth and tongue. The image decoder part of the model uses the SPADE module to synthesize clear images. The model does not use a large number of transformers and diffusion, so the computational cost is low and it is developed for real-time reasoning.

[0104] S3: Pytorch is used as the training framework to train the digital population synthesis model. The batch size is set to 16, the optimizer uses Adam, the learning rate regulator uses CosineAnnealingLR, and the initial learning rate is set to 1×10^-4. Before formal training, a warm-up is performed, that is, the learning rate gradually increases from zero to the initial learning rate. The parameters of the lip-sync expert and lip-reading expert are frozen during training. Data augmentation uses a strategy of randomly changing the gamma and brightness of the image, such as Figure 7 The present invention improves the synchronization of lip synthesis by using lip-sync expert and improves the accuracy of lip synthesis by using lip-reading expert.

[0105] Among them, the calculation formula of the model's loss function is as follows:

[0106] L total =λ s ×L sync +λ c ×L lip +λ r ×L rec +λ p ×L p +λ g ×L gan ;

[0107] Among them, L total is the loss function value of the model, L sync is the lip synchronization loss value, λ s is the weight of lip synchronization loss, L lip is the lip reading loss, λ c is the weight of lip reading loss, L rec is the reconstruction loss value, λ r is the weight of the reconstruction loss, L p is the perceptual loss value, λ p is the weight of perceptual loss, L gan is the loss value of the generated adversarial network, λ g is the weight of the GAN loss.

[0108] Specifically, (1) the expression of the lip synchronization loss value is as follows:

[0109]

[0110] Where a is the audio vector of the S-frame audio Mel-spectrogram encoded by the lip-sync expert, v is the image vector of the lower half of the image generated by the S-frame model, encoded by the lip-sync expert, S is a natural number, where S can be 5, ε is a small positive constant to prevent the denominator from being zero, where ε is 10^(-6); the lip synchronization loss value is a contrast loss that allows synchronized audio features and image features to attract each other, while allowing unsynchronized audio features and image features to repel each other. Different from the commonly used syncnet as a lip-sync expert, a syncXnet that is more robust to noise is retrained here. The structure of syncXnet is as follows Figure 8 As shown, the audio encoder uses the pre-trained whisper audio encoder, and the image encoder uses the pre-trained ResNeXt, and the training method is consistent with syncnet.

[0111] (2) The expression of lip reading loss value is as follows:

[0112]

[0113] Among them, Y is the real text, is the text predicted by the lip reading expert based on the given video V, where V is the video and P is the probability;

[0114] (3) The expression of reconstruction loss value is as follows:

[0115]

[0116] Among them, N is the number of elements in the tensor, i is the element index in the tensor, and v is the value of each element in the tensor.

[0117] (4) The perceptual loss value is used to calculate the perceptual loss of the features of the VGG model at three image scales (3×H×W, 3×H / 2×W / 2, 3×H / 4×W / 4), and its expression is as follows:

[0118]

[0119] Among them, Vgg i represents the i-th layer in the VGG-19 network, Io is the frame image output by the model, Ir is the real reference frame image, and W i Represents the width of the i-th layer tensor, H i Represents the height of the i-th layer tensor, C i Represents the number of channels of the i-th layer tensor;

[0120] (5) Regarding the generative adversarial network loss, it includes the discriminator loss and the generator loss. The expression of the generative adversarial network loss value is as follows:

[0121]

[0122] The expressions of generator loss and discriminator loss are:

[0123]

[0124] Among them, D j It stands for GAN discriminator, a general term in the field of generative adversarial networks (GANs), E stands for expected value, Io is the frame image output by the model, and Ir is the real reference frame image.

[0125] The present invention designs reference frame selection strategies during training and inference respectively, so that the model can select the most appropriate reference frame during training and inference, and there is not much difference between the two.

[0126] Compared with the prior art, the present invention has the following advantages:

[0127] (1) Improved lip sync accuracy: Audio and lip sync can be synchronized more accurately.

[0128] (2) Accurate lip movements: Ability to synthesize lip movements that are consistent with pronunciation.

[0129] (3) Clear facial image: Able to synthesize clear skin texture material that is consistent with the skin in the input image.

[0130] (4) Multi-language support: The same model can support lip-syncing in multiple languages ​​at the same time, such as Chinese, English, Japanese, and French, and can easily handle multiple languages ​​mixed in the input audio.

[0131] (5) Low computational overhead: Through a carefully designed model architecture, the computational cost of the model is controlled, so that the technology can be implemented in real time on edge devices and is suitable for resource-constrained devices such as mobile phones and tablets.

[0132] (6) Broad application prospects: This algorithm can be widely used in digital humans from all walks of life (including but not limited to customer service, anchors, teachers, relatives, self-media bloggers, advertising spokespersons) and has high commercial potential.

[0133] (7) The teeth and tongue are natural and the teeth are consistent between different frames.

[0134] The above is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any technician in the relevant technical field, without departing from the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification to the technical solution and technical content disclosed in the present invention, which does not depart from the content of the technical solution of the present invention and still falls within the protection scope of the present invention.

Claims

1. A high-performance voice-driven digital population synthesis algorithm, characterized in that: The following steps are involved: S1: Obtain and preprocess the dataset; S2: Design the model structure using audio encoder, image encoder, audio feature selector, feature fusion module and image decoder; S3: Pytorch is used as the training framework to train the digital population synthesis model, where the calculation formula of the model's loss function is as follows: L total =λ s ×L sync +λ c ×L lip +λ r ×L rec +λ p ×L p +λ g ×L gan ; Among them, L total is the loss function value of the model, L sync is the lip synchronization loss value, λ s is the weight of lip synchronization loss, L lip is the lip reading loss, λ c is the weight of lip reading loss, L rec is the reconstruction loss value, λ r is the weight of the reconstruction loss, L p is the perceptual loss value, λ p is the weight of perceptual loss, L gan is the loss value of the generated adversarial network, λ g is the weight of the GAN loss.

2. The high-performance voice-driven digital population synthesis algorithm according to claim 1, characterized in that: The way to obtain the dataset is as follows: Record several hours of talking videos of people, covering different ages, genders, face shapes and skin colors, with a resolution of 720P and a frame rate of 25FPS. Cut the long videos into 5-second short video clips, extract the audio of each short video clip, and obtain a standard digital audio file with a sampling rate of 16K.

3. The high-performance voice-driven digital population synthesis algorithm as claimed in claim 2, characterized in that: The dataset is preprocessed as follows: The coordinates of the key points of the face in each frame of the short video clip are extracted, and the coordinates of the key points are used to calculate an affine transformation matrix that can make the face stand at attention. The affine transformation matrix is ​​applied to the frame to obtain a picture with the face standing at attention, the coordinates of the face rectangular frame are detected in the picture, and a 256×256 face image is cropped from the picture using the coordinates of the face rectangular frame.

4. The high-performance voice-driven digital population synthesis algorithm according to claim 1, characterized in that: The audio encoder inputs a standard digital audio file and outputs audio features with a tensor shape of [B, L, D], where B is the batch size, L is the time length, and D is the feature dimension.

5. The high-performance voice-driven digital population synthesis algorithm as claimed in claim 4, characterized in that: The image encoder is composed of several layers of two-dimensional convolution, several groups of normalization and several activation functions stacked together. The number of groups is R+1, where R is the number of input reference images. Downsampling is performed by setting the stride to 2 in some two-dimensional convolution layers. The input of the image encoder is R+1 images concatenated in the channel dimension, and the output is the image feature with a tensor shape of [B,H,W,C], where B represents the number of inputs, H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels of the feature map.

6. The high-performance voice-driven digital population synthesis algorithm according to claim 5, characterized in that: The audio feature selector selects some audio features from the audio features output by the audio encoder to obtain audio features with a tensor shape of [B, F, D], where F is the selection parameter.

7. The high-performance voice-driven digital population synthesis algorithm of claim 6, characterized in that: The feature fusion module works as follows: For the audio feature whose tensor shape is [B, F, D], the tensor of the audio feature is multiplied by the mouth opening amplitude coefficient; Use an adapter to map audio domain features to a common space with image features; The tensor is sinusoidally absolute position encoded; Input the tensor encoded by the sinusoidal absolute position into a multi-head self-attention module, which is used to learn audio context information at different times. The fully connected layer in the multi-head self-attention module maps the audio features into Query, Key and Value; Perform the attention function operation on Query, Key and Value as follows: Among them, Q is Query, K is Key, V is Value, T is the transposition operation, and G is the number of channels for each token; Then, residual connection addition and layer normalization are performed, and the resulting tensor is used as the query of the image-to-audio cross-attention module; For image features, the tensor shape of [B, H, W, C] is reshaped from [B, H, W, C] to [B, H × W, C]. The changed tensor shape is used as the key and value of the image-to-audio cross-attention module to enable the image features to extract information from the audio features and realize multimodal information fusion.

8. The high-performance voice-driven digital population synthesis algorithm according to claim 7, characterized in that: The model structure of the image decoder is composed of M upsampling, M convolutions and M SPADEResnetBlock stacks, where M is a natural number, and is also connected with Leaky ReLU function, Conv2d function and Sigmoid function.

9. The speech-driven high-effect digital population synthesis algorithm as claimed in claim 1, characterized in that: Use pytorch as the training framework to train the digital population synthesis model as follows: The training framework uses pytorch, the batch size is set to 16, the optimizer uses Adam, the learning rate regulator uses CosineAnnealingLR, and the initial learning rate is set to 1×10^-4; Before formal training, the learning rate gradually increases from zero to the initial learning rate; During training, the parameters of the lip-sync expert and lip-reading expert are frozen, and data augmentation uses a strategy of randomly changing the gamma and brightness of the image.

10. The high-performance voice-driven digital population synthesis algorithm according to claim 9, characterized in that: The expression for the lip sync loss value is as follows: Where a is the audio vector of the S-frame audio Mel-spectrogram after lip-sync expert encoding, v is the image vector of the lower half of the image generated by the S-frame model after lip-sync expert encoding, S is a natural number, and ε is a small positive constant to prevent the denominator from being zero; The expression of lip reading loss value is as follows: Among them, Y is the real text, is the text predicted by the lip reading expert based on the given video V, and P is the probability; The expression of the reconstruction loss value is as follows: Where N is the number of elements in the tensor, i is the element index in the tensor, and v is the value of each element in the tensor; The expression of the perceptual loss value is as follows: Among them, Vgg i represents the i-th layer in the VGG-19 network, Io is the frame image output by the model, Ir is the real reference frame image, and W i Represents the width of the i-th layer tensor, H i Represents the height of the i-th layer tensor, C i Represents the number of channels of the i-th layer tensor; The expression for the loss value of the generated adversarial network is as follows: Among them, D j represents the GAN discriminator, E represents the expected value, Io is the frame image output by the model, and Ir is the real reference frame image.