A method for digital human speech and lip synchronization
By using the lip synchronization model to synchronize audio and image features in digital life generation technology, the problem of low matching of digital human voice and lip shape and poor adaptability of multilinguals is solved, and a more natural and smooth digital life generation is achieved.
Patent Information
- Application Number
- CN202411960919.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The digital human voice and lip shape generated by the prior art are not matched very well and cannot adapt to multilingual environments.
The lip-synchronous model is used to synchronize the audio-encoded features and image-encoded features, and synchronous features are generated to achieve video generation through linear feature projection and feature similarity calculation.
It improves the matching degree of digital human voice and lip shape, and can adapt to multi-lingual environments to achieve more natural and smooth digital human generation.
Smart Images

Figure CN119400207B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital human generation, and particularly to a digital human speech lip synchronization method. Background Art
[0002] With the rapid development of artificial intelligence technology, 2D digital human technology has become a hot topic in multiple fields such as virtual reality, augmented reality, games, entertainment, and education. A digital human, that is, a virtual digital character, can imitate the behaviors and expressions of real people and provide an interactive user experience. In these applications, lip synchronization technology is one of the key factors to achieve natural and smooth digital humans.
[0003] Traditional 2D digital human generation algorithms usually rely on pre-recorded animations or simple deformation techniques to simulate lip shapes, but these methods are difficult to adapt to real-time changing speech signals, resulting in the phenomenon of out-of-sync audio and video when the generated digital human speaks. To solve this problem, researchers have begun to explore lip synchronization technology based on deep learning. These technologies achieve high-precision lip synchronization effects by analyzing audio signals and video frames, enabling the digital human to present lip shape changes that match the speech content when speaking.
[0004] Existing lip synchronization technologies are generally implemented based on the Wav2Lip technology, but there are problems such as difficult model convergence and inability to adapt to multilingual environments. Summary of the Invention
[0005] This application provides a digital human speech lip synchronization method to solve the problems that the generated digital human speech and lip shapes in the prior art do not match well and cannot adapt to multilingual environments.
[0006] The method includes:
[0007] Obtain a source video, preprocess the source video to obtain audio data and image data;
[0008] Use a lip synchronization model to synchronize process the audio encoding features and the image encoding features to obtain synchronization features; the audio encoding features are output by the lip synchronization model from the audio data; the image encoding features are obtained by the lip synchronization model through image partitioning, feature fusion, and feature encoding; the synchronization process includes linear feature projection and feature similarity calculation;
[0009] Generate a target video according to the synchronization features.
[0010] Preferably, the lip synchronization model includes:
[0011] An image processing module, which is configured to perform image encoding according to the image data to obtain a plurality of the image encoding features;
[0012] An audio processing module, which is configured to perform audio encoding according to the audio data to obtain a plurality of the audio encoding features;
[0013] A synchronization processing module, which is configured to perform audio - image synchronization processing according to the audio encoding features and the image encoding features to obtain the synchronization features.
[0014] Preferably, the image processing module includes:
[0015] A first grouping unit, which is configured to divide the input image data into a plurality of consecutive image sequences; each of the image sequences includes a plurality of consecutive image frames;
[0016] A feature fusion unit, which includes a plurality of fusion layers in parallel distribution. The plurality of fusion layers are configured to perform feature fusion on the image frames in the plurality of image sequences respectively to obtain a plurality of corresponding first sequences;
[0017] An image encoding unit, which includes a plurality of image encoding layers in parallel distribution. The plurality of image encoding layers correspond to the plurality of fusion layers one by one. The plurality of image encoding layers are configured to perform convolutional encoding on the plurality of first sequences respectively to obtain a plurality of corresponding the image encoding features.
[0018] Preferably, the audio processing module includes:
[0019] A second grouping unit, which is configured to divide the input audio data into a plurality of consecutive audio sequences; each of the audio sequences includes a plurality of consecutive audio sub - data, and each of the audio sub - data corresponds to one of the image frames;
[0020] An audio encoding unit, which includes a plurality of audio encoding layers in parallel distribution. The plurality of audio encoding layers are configured to perform convolutional encoding on the plurality of audio sequences respectively to obtain a plurality of corresponding the audio encoding features.
[0021] Preferably, the image encoding unit further includes a plurality of spatial attention layers in parallel distribution, and the plurality of spatial attention layers correspond to the plurality of image encoding layers one by one;
[0022] The spatial attention layer is further configured to calculate spatial attention weights according to the first sequence;
[0023] The image encoding layer is further configured to perform convolutional encoding according to the first sequence and the spatial attention weights to output the image encoding features.
[0024] Preferably, the image encoding unit further includes a channel attention layer;
[0025] The image encoding unit is further configured to divide the first sequence into a plurality of subsequences, perform feature encoding processing on each subsequence, calculate the attention weights corresponding to different subsequences through the spatial attention layer, and fuse the results of the feature encoding processing of the plurality of subsequences based on the attention weight calculation results to obtain local feature encodings;
[0026] Divide the local feature encodings into a plurality of subsequences, perform feature encoding processing on each subsequence, calculate the attention weights corresponding to different subsequences through a preset spatial attention module, and fuse the results of the feature encoding processing of the plurality of subsequences based on the attention weight calculation results to obtain local feature encodings;
[0027] Iteratively update, and independently output the local feature encodings obtained in each round to obtain a plurality of the local feature encodings;
[0028] Calculate the attention weights for the plurality of local feature encodings according to the channel attention layer, and perform feature processing based on the attention weight calculation results to obtain global feature encodings; fuse the plurality of local feature encodings and the global feature encodings to obtain the image encoding features.
[0029] Preferably, the synchronization processing module includes:
[0030] A feature projection module, the feature projection module includes a plurality of linear layers, and the plurality of linear layers correspond to the plurality of image encoding layers and the plurality of audio encoding layers respectively; the plurality of linear layers are configured to map the image encoding features and the audio encoding features into a multi-modal embedding space;
[0031] A similarity calculation module, the similarity calculation module includes a plurality of calculation layers, and the plurality of calculation layers correspond to the plurality of linear layers respectively; the plurality of calculation layers are configured to calculate the similarity between the image encoding features and the audio encoding features in the multi-modal embedding space, and fuse the image encoding features and the audio encoding features with the highest similarity to obtain the synchronization features.
[0032] Preferably, the training process of the lip synchronization model includes:
[0033] Obtain a training video, and extract audio training data and image training data according to the training video;
[0034] Perform model training on the audio training data, the image training data, and the lip synchronization model;
[0035] Use a loss function to converge the lip synchronization model until it meets the preset model requirements.
[0036] Preferably, the step of using a loss function to converge the lip synchronization model includes:
[0037] Perform contrastive loss training on the synchronization features using a contrastive loss function;
[0038] When the synchronization features meet the first preset model requirements, use a triplet loss function to perform positive sample training and negative sample training on the lip synchronization model until it meets the second preset model requirements.
[0039] Preferably, the steps of performing positive sample training and negative sample training on the lip synchronization model include:
[0040] Obtain an anchor feature, a positive feature, and a negative feature; the anchor feature is a randomly selected image feature or audio feature; the positive feature is an image feature or audio feature that matches the anchor feature; the negative feature is an image feature or audio feature that does not match the anchor feature;
[0041] Perform positive sample training based on the anchor feature and the positive feature, and perform negative sample training based on the anchor feature and the negative feature until both the positive sample training and the negative sample training meet the second preset model requirements.
[0042] Preferably, the steps of preprocessing the source video include:
[0043] Extract source images from the source video, and sequentially perform image frame processing, face detection processing, face region cropping processing, occlusion detection processing, and image covering processing on the source images to obtain the image data;
[0044] Extract source audio from the source video, and sequentially perform audio frame processing and audio denoising processing on the source audio to obtain the audio data.
[0045] As can be seen from the above, the present application provides a method for digital human speech and lip synchronization. The method includes obtaining a source video, preprocessing the source video to obtain audio data and image data; using a lip synchronization model to perform synchronization processing on the audio coding features and the image coding features to obtain synchronization features; the audio coding features are output by the lip synchronization model from the audio data; the image coding features are output by the lip synchronization model from the image data; the synchronization processing includes linear feature projection and feature similarity calculation; generating a video according to the synchronization features to obtain a target video. The present application solves the problems in the prior art that the matching degree between the speech and lip of the generated digital human is not high and it cannot adapt to a multilingual environment through the above method. Description of the Drawings
[0046] To more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of a method for digital human speech and lip synchronization of the present application;
[0048] Figure 2 It is a schematic diagram of a lip synchronization model in a method for digital human speech and lip synchronization of the present application;
[0049] Figure 3 It is a schematic diagram of an image processing module in a method for digital human speech and lip synchronization of the present application;
[0050] Figure 4 It is a schematic diagram of an audio processing module in a method for digital human speech and lip synchronization of the present application;
[0051] Figure 5 It is a schematic diagram of a synchronization processing module in a method for digital human speech and lip synchronization of the present application;
[0052] Figure 6 It is a training flowchart of a lip synchronization model in a method for digital human speech and lip synchronization of the present application;
[0053] Figure 7 It is a flowchart of the convergence mode of a lip synchronization model in a method for digital human speech and lip synchronization of the present application;
[0054] Figure 8 It is a flowchart of positive sample training and negative sample training in a method for digital human speech and lip synchronization of the present application;
[0055] Figure 9 It is a flowchart of preprocessing in a method for digital human speech and lip synchronization of the present application. Detailed implementation manners
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] In recent years, a technology called Wav2Lip has attracted wide attention. Wav2Lip is a general speaker model that can generate videos with lip-sync accuracy matching real synchronous videos. Its core architecture includes a generator and two discriminators: an expert lip-sync discriminator and a visual quality discriminator. The expert lip-sync discriminator is responsible for accurately discriminating the synchronization of sound and mouth shapes in the video, while the visual quality discriminator is used to improve the picture quality.
[0058] In practical applications, such as lip-sync for videos recorded outdoors with a hand-held device, there are still challenges. These challenges include the diverse expressions of people in the video, lighting changes, occlusion problems, and different speaking styles. To improve the robustness and accuracy of the generation network, researchers have proposed a method of using an expert discriminator to correct the generation network. This method uses a pre-trained expert-level lip-sync discriminator to penalize inaccurate generations of the generation network during training, thereby improving the lip-sync quality of the generated frames.
[0059] However, the expert lip-sync discriminator is prone to mode collapse during the training process, especially when dealing with large datasets composed of multiple languages, and it is difficult for the model to converge, which brings great difficulties to multi-language applications. In addition, since Wav2Lip is mainly trained on English datasets, its accuracy and robustness may decrease when dealing with non-English languages, which poses additional challenges for applications in multi-language environments.
[0060] Based on the above problems, the present application provides the following embodiments to solve the above problems.
[0061] Figure 1 It is a flowchart of a digital human speech lip-sync method of the present application.
[0062] See Figure 1 It can be seen that the present embodiment provides a digital human speech lip-sync method, and the method includes:
[0063] S10. Obtain the source video and preprocess the source video. Specifically, in this embodiment, since conventional video data sources are generally obtained from open-source databases and the quality of such videos varies, relevant preprocessing of the video is required before digital human generation to meet the standards for digital human generation and sequentially improve the quality of digital human generation.
[0064] Among them, a video with sound generally consists of sound and images. Through the preprocessing of this embodiment, the sound and images are segmented to obtain audio data and image data.
[0065] Figure 9 It is a flowchart of the preprocessing in a digital human voice-lip synchronization method of this application.
[0066] See Figure 9 It can be seen that, further, in some embodiments, the step of preprocessing the source video includes:
[0067] S11. Extract the source image from the source video, and sequentially perform image frame processing, face detection processing, cropping of the face region, occlusion detection processing, and image covering processing on the source image to obtain the image data;
[0068] S12. Extract the source audio from the source video, and sequentially perform audio frame processing and audio denoising processing on the source audio to obtain the audio data.
[0069] Specifically, in this embodiment, before digital human generation, video acquisition is first required to obtain continuous video data containing clear speech and lip shape changes. Subsequently, these video data are decomposed into individual frames, and at the same time, the accompanying audio in the video is also segmented into corresponding audio frames for synchronous processing.
[0070] In the video frame processing stage, first, the face detection technology is used to locate the face region in each frame, and then the face part is accurately cropped and the upper half of the face is covered because the lip shape changes mainly occur in this region. Occlusion detection is also performed on the cropped face to ensure that the face region is not occluded by other objects, thereby ensuring the accuracy of lip shape recognition.
[0071] Next, denoising processing is performed on the segmented audio frames to extract clearer and more accurate audio features, which are crucial for subsequent lip synchronization analysis. After processing the audio, the extracted audio features are matched and analyzed with the face region in the video frames. To make the training model more focused on the movement of the lips rather than other facial features, the upper half of the face is also covered.
[0072] It should be noted that steps S11 and S12 are carried out synchronously without a sequence.
[0073] The method further includes:
[0074] S20, using a lip synchronization model to perform synchronization processing on the audio coding features and the image coding features. Specifically, in this embodiment, the lip synchronization model is used to perform synchronization processing on the audio coding features and the image coding features to obtain relevant features for generating a digital human video, that is, synchronization features.
[0075] Among them, correspondingly, the audio coding features are obtained by inputting the audio data into the lip synchronization model for calculation, and the image coding features are obtained by inputting the image data into the lip synchronization model for calculation; the audio coding features and the image coding features are respectively extracted through the lip synchronization model, and the audio coding features and the image coding features are used for synchronization processing to combine the sound with the lip movement changes of the person, thereby realizing speech lip synchronization.
[0076] Among them, the synchronization processing includes linear feature projection and feature similarity calculation. Through the linear feature projection and the feature similarity calculation, the image features and the audio features are mapped together to achieve the fusion of audio and features.
[0077] The method further includes:
[0078] S30, generating a video according to the synchronization features. Specifically, after completing step S20, the synchronization features for characterizing the digital human video are obtained, and a video is generated according to the synchronization features to obtain the final target video.
[0079] Figure 2 It is a schematic diagram of a lip synchronization model in a digital human speech lip synchronization method of the present application.
[0080] See Figure 2 It can be seen that, further, in some embodiments, the lip synchronization model includes:
[0081] An image processing module, which is configured to perform image coding according to the image data. Specifically, in this embodiment, the image data in the source video is subjected to image coding through the image processing module to obtain the features corresponding to the images in the source video, where the obtained features include a plurality of image coding features.
[0082] The lip synchronization model further includes:
[0083] An audio processing module, which is configured to perform audio encoding according to the audio data. Specifically, in this embodiment, the audio data in the source video is encoded by the audio processing module to obtain the features corresponding to the audio in the source video, where the obtained features include a plurality of audio encoding features.
[0084] It should be noted that the image encoding features and the audio encoding features correspond one by one.
[0085] The lip-sync model further includes:
[0086] A synchronization processing module, which is configured to perform audio-visual synchronization processing according to the audio encoding features and the image encoding features. Specifically, in this embodiment, the synchronization processing module performs audio-visual synchronization processing on the image encoding features obtained by the image processing module and the audio encoding features obtained by the audio processing module, so as to unify the image encoding features and the audio encoding features together to obtain the synchronization features, and thereby achieve the synchronization of speech and lip movement.
[0087] Figure 3 This is a schematic diagram of the image processing module in a digital human speech-lip synchronization method of the present application.
[0088] See Figure 3 It can be seen that, further, in some embodiments, the image processing module includes:
[0089] A first grouping unit, which is configured to divide the input image data into a plurality of consecutive image sequences. Specifically, in this embodiment, in order to better synchronize speech and lip movement, a single image cannot well reflect the features in the image, so only one image cannot be used for feature extraction. In this embodiment, the first grouping unit is used to divide the image data to divide the image data into a plurality of consecutive image sequences, and each image sequence includes a plurality of consecutive image frames, thereby providing a basis for subsequent synchronization processing.
[0090] The image processing module further includes:
[0091] Feature fusion unit, the feature fusion unit includes a number of fusion layers distributed in parallel, and the number of fusion layers are configured to perform feature fusion on the image frames in a number of the image sequences respectively. Specifically, in this embodiment, the first grouping unit divides a number of the image sequences, and each of the image sequences includes a number of consecutive image frames. However, simply using each image frame cannot effectively reflect the characteristics of the lip shape of the person in the image. Therefore, it is necessary to fuse the number of image frames in each of the image sequences to obtain a feature map with the characteristics of multiple image frames. For different image sequences, a number of consecutive feature maps, that is, a number of first sequences, can be obtained.
[0092] The image processing module further includes:
[0093] Image encoding unit, the image encoding unit includes a number of image encoding layers distributed in parallel, and the number of image encoding layers correspond one-to-one with the number of fusion layers. The number of image encoding layers are configured to perform convolutional encoding on a number of the first sequences respectively. Specifically, in this embodiment, when the feature fusion unit completes the fusion of the image frames in all or part of the image sequences, feature encoding is performed on the number of first sequences after fusion to obtain a number of corresponding image encoding features.
[0094] Among them, convolutional encoding of each image sequence is performed through the image encoding layers distributed in parallel in the image encoding unit. One or more image sequences correspond to one image encoding layer, and one image encoding layer performs convolutional encoding on one or more image sequences to obtain the image encoding features corresponding to each image sequence.
[0095] Further, in some embodiments, the image encoding unit further includes a number of spatial attention layers distributed in parallel, and the number of spatial attention layers correspond one-to-one with the number of image encoding layers;
[0096] The spatial attention layer is further configured to calculate the spatial attention weights according to the first sequence;
[0097] The image encoding layer is further configured to perform convolutional encoding according to the first sequence and the spatial attention weights to output the image encoding features.
[0098] Specifically, in this embodiment, the image processing module is improved based on the residual architecture, in which the global average pooling layer is replaced by a spatial attention mechanism. The spatial attention mechanism is used to find the most important parts of the face for processing, so as to enhance the feature expression of the key area. This improvement enables the network to concentrate on processing the key area of the lip movement in the image and extract more representative visual features.
[0099] Furthermore, in some embodiments, the image encoding unit further includes a channel attention layer;
[0100] The image encoding unit is further configured to divide the first sequence into multiple subsequences, perform feature encoding processing on each subsequence, calculate the attention weights corresponding to different subsequences through the spatial attention layer, and fuse the multiple feature encoding processing results based on the attention weight calculation results to obtain a local feature encoding. Specifically, in this embodiment, through the sequence division, attention weight calculation of the first sequence, and local feature encoding according to the attention weights, a spatial attention mechanism is introduced in the calculation process of local features. The spatial attention mechanism focuses on the important points at each spatial position in the feature map, so it has a better presentation result when calculating the attention of the local area.
[0101] The image encoding unit is further configured to divide the local feature encoding into multiple subsequences, perform feature encoding processing on each subsequence, calculate the attention weights corresponding to different subsequences through a preset spatial attention module, and fuse the multiple feature encoding processing results based on the attention weight calculation results to obtain a local feature encoding. Specifically, in this embodiment, the processing of the first sequence by the image encoding unit is the same as above, which is to perform local feature analysis on the feature encoding. However, the difference from the above is that after completing the feature analysis of the local feature encoding, the obtained feature encoding needs to be iteratively updated, and cyclic local feature analysis is performed on the updated feature encoding, and the obtained local feature encoding is independently output after each round of local feature analysis.
[0102] The image encoding unit is further configured to calculate the attention weights for multiple local feature encodings according to the channel attention layer, and perform feature processing based on the attention weight calculation results to obtain a global feature encoding; fuse the multiple local feature encodings and the global feature encoding to obtain the image encoding feature. Specifically, in this embodiment, after several rounds of local feature analysis, several local feature encodings can be obtained. Then, it is necessary to use the image encoding unit to perform further feature analysis on these local feature encodings. The core purpose of this embodiment is to introduce a channel attention mechanism in the feature analysis process. The channel attention mechanism focuses on calculating the importance of different feature channels relative to the global, so it has a better presentation in the effect of calculating the global feature.
[0103] Figure 4 It is a schematic diagram of the audio processing module in a digital human speech lip synchronization method of the present application.
[0104] See Figure 4It can be seen that, further, in some embodiments, the audio processing module includes:
[0105] A second grouping unit, which is configured to divide the input audio data into several consecutive audio sequences. Specifically, in this embodiment, in order to better synchronize the speech and lip shapes, a single audio frame cannot well reflect the features in the audio, and correspondingly, the image processing module extracts features from multiple image frames. Similarly, the features cannot be extracted only using one frame of audio. In this embodiment, the second grouping unit is used to divide the audio data to divide the audio data into several consecutive image sequences, and each audio sequence includes several consecutive audio sub-data, providing an audio data basis for subsequent synchronization processing.
[0106] It should be noted that each audio sub-data corresponds to one image frame, that is, the number of image frames included in the image sequence is the same as the number of audio sub-data included in the audio sequence.
[0107] The audio processing module further includes:
[0108] An audio encoding unit, which includes several audio encoding layers distributed in parallel. The several audio encoding layers are configured to perform convolutional encoding on several audio sequences respectively. Specifically, in this embodiment, convolutional encoding of each audio sequence is performed through the audio encoding layers distributed in parallel in the audio encoding unit. One or more audio sequences correspond to one audio encoding layer, and one audio encoding layer performs convolutional encoding on one or more audio sequences, thereby obtaining the audio encoding features corresponding to each audio sequence.
[0109] The audio processing module processes audio features using a convolutional neural network (CNN). Through a series of convolutional layers, the encoder extracts the key features in the audio signal and converts them into a high-dimensional feature vector.
[0110] Figure 5 This is a schematic diagram of the synchronization processing module in a digital human speech and lip synchronization method of the present application.
[0111] See Figure 5 It can be seen that, further, in some embodiments, the synchronization processing module includes:
[0112] A feature projection module, which includes several linear layers. The several linear layers correspond one by one to several image encoding layers and several audio encoding layers respectively; the several linear layers are configured to map the image encoding features and the audio encoding features into a multi-modal embedding space;
[0113] A similarity calculation module, the similarity calculation module includes a number of calculation layers, and the number of calculation layers corresponds one-to-one with a number of the linear layers; the number of calculation layers are configured to calculate the similarity between the image encoding features and the audio encoding features in the multi-modal embedding space, and perform feature pair fusion on the image encoding features and the audio encoding features with the highest similarity to obtain the synchronization features.
[0114] Specifically, in this embodiment, after encoding, the audio encoding features and the image encoding features are sent into a linear layer by the feature projection module for projection, and mapped into a shared multi-modal embedding space, that is, two groups of high-dimensional vectors are found to represent the audio encoding features and the image encoding features respectively. Through multi-modal technology, they are projected into a shared latent representation, so that the reconstruction error of each modality on the common latent space is minimized, and these projection matrices are as sparse as possible. In this space, the network is trained by the method of contrast learning, so that the embedding vectors of synchronous audio-visual pairs are closer in the space, while asynchronous pairs are farther away.
[0115] Using optimization algorithms such as backpropagation and gradient descent, adjust the network parameters to minimize the contrast loss function and improve the accuracy of the model in recognizing lip synchronization. Finally, the network outputs embedding vectors that can represent the synchronization of lip shapes and voices, and these vectors can be used in applications such as lip synchronization discrimination, speech enhancement, or digital human animation generation.
[0116] Exemplary Embodiment 1
[0117] In this exemplary embodiment, the digital human speech lip synchronization method can determine the audio-visual synchronization between the mouth movement and the speech in a digital human broadcast or singing video, and determine whether there is an issue of out-of-sync audio and video in the video.
[0118] Exemplary Embodiment 2
[0119] In this exemplary embodiment, the digital human speech lip synchronization method can determine the human subject corresponding to the current speech in a multi-person digital human broadcast or singing video.
[0120] Figure 6 This is the training flowchart of the lip synchronization model in a digital human speech lip synchronization method of the present application.
[0121] See Figure 6 It can be seen that, further, in some embodiments, the training process of the lip synchronization model includes:
[0122] S100, obtain training videos, and extract audio training data and image training data according to the training videos;
[0123] S200, perform model training on the audio training data, the image training data, and the lip synchronization model;
[0124] S300, use the loss function to converge the lip synchronization model until it meets the preset model requirements.
[0125] Specifically, in this embodiment, before performing model training, it is necessary to first obtain the training video, use the training video to perform model training on the lip synchronization model, and use the loss function to converge the lip synchronization model during the model training process until the lip synchronization model meets the preset model requirements.
[0126] Figure 7 This is a flowchart of the convergence method of the lip synchronization model in a digital human speech lip synchronization method of the present application.
[0127] See Figure 7 It can be seen that, further, in some embodiments, the step of using the loss function to converge the lip synchronization model includes:
[0128] S310, perform contrastive loss training on the synchronization features using a contrastive loss function. Specifically, in this embodiment, the loss function includes the contrastive loss function, and the contrastive loss function is used to perform contrastive loss training on the synchronization features to achieve preliminary loss convergence.
[0129] The step of using the loss function to converge the lip synchronization model further includes:
[0130] S320, when the synchronization features meet the first preset model requirements, use a triplet loss function to perform positive sample training and negative sample training on the lip synchronization model until it meets the second preset model requirements. Specifically, in this embodiment, the loss function further includes the triplet loss function, and the triplet loss function is used to perform secondary loss convergence on the lip synchronization model that has completed contrastive loss function loss convergence to enhance the confidence of the model.
[0131] Figure 8 This is a flowchart of positive sample training and negative sample training in a digital human speech lip synchronization method of the present application.
[0132] See Figure 8 It can be seen that, further, in some embodiments, the step of performing positive sample training and negative sample training on the lip synchronization model includes:
[0133] Obtain the anchor features, positive features, and negative features; the anchor features are randomly selected image features or audio features; the positive features are image features or audio features that match the anchor features; the negative features are image features or audio features that do not match the anchor features.
[0134] Perform positive sample training based on the anchor features and the positive features, and perform negative sample training based on the anchor features and the negative features until both the positive sample training and the negative sample training meet the requirements of the second preset model.
[0135] Specifically, in this embodiment, the loss convergence using the triplet loss function mainly includes positive sample training and negative sample training. The positive sample training and negative sample training are reflected in that positive sample training is performed using the anchor features and the positive features, and negative sample training is performed using the anchor features and the negative features, so as to achieve the loss convergence of the triplet loss function.
[0136] Exemplarily, the loss convergence of the lip synchronization model can be understood as the following two stages.
[0137] The first stage: Contrastive loss training:
[0138] ;
[0139] In the initial stage of training, the model only uses the matched audio-visual data pairs for contrastive loss training. Contrastive loss is widely used in unsupervised learning. This loss function is mainly used for dimensionality reduction. That is, for samples that are originally similar, after dimensionality reduction (feature extraction), in the feature space, the two samples are still similar; while for samples that are originally dissimilar, after dimensionality reduction, in the feature space, the two samples are still dissimilar. Similarly, this loss function can also well express the matching degree of paired samples. d represents the Euclidean distance between the features of two samples, y is the label indicating whether the two groups of samples match, and m (margin) is the set threshold. If the distance exceeds m, the loss is regarded as 0. That is, if two dissimilar features are far apart, then the contrastive loss should be very low. The goal of this stage is to learn a feature space in which synchronous audio-visual feature vectors are close to each other, while asynchronous feature vectors are far apart. Most importantly, a similarity matrix is used to calculate the loss. The similarity matrix is calculated at the batch level, which means that all sample pairs within a batch are considered simultaneously. By constructing a similarity matrix, the model can compare the similarity of a sample with all other samples in the batch at the same time. This is an innovative idea in multi-modal alignment methods. This contrastive learning strategy helps the model learn to distinguish positive and negative samples because unmatched sample pairs provide rich negative sample information.
[0140] The Second Stage: Fine-tuning
[0141] Based on the first stage, the training in the second stage is fine-tuned by introducing mismatched audio-visual data pairs to improve the generalization performance of the model. The mismatched audio-visual data pairs include two cases: one is the audio-visual frames that are not temporally corresponding from the same video; the other is the random combination of audio and video frames from different videos. These mismatched data pairs are randomly selected with a certain probability to simulate the asynchronous situations that may be encountered in the real world.
[0142] ;
[0143] In the second stage, a triplet loss function is adopted to process both the matched and mismatched audio-visual data pairs simultaneously. The input is a triplet, including an Anchor example, a Positive example, and a Negative example. By optimizing the distance between the anchor feature and the positive feature to be less than the distance between the anchor feature and the negative feature, the similarity calculation between samples is realized. The triplet loss function can optimize the model's learning of positive samples for matched pairs and negative samples for mismatched pairs simultaneously, thereby further distinguishing synchronous and asynchronous data in the feature space. The design of this loss function helps the model better understand the complex relationships between audio-visual data and improve its recognition ability when facing unseen data.
[0144] Through the training of these two stages, the lip-sync network can not only effectively learn the data in the training set but also accurately judge the synchronization of new and unseen audio-visual data pairs, thus showing better robustness and adaptability in practical applications.
[0145] In actual use, the speech lip-sync network no longer participates in the training during the digital human training process. The digital human generation network uses five consecutive frames generated by the speech to form the input of the speech lip-sync network. After passing through the speech lip-sync network, the corresponding image multi-modal latent space vector and speech multi-modal latent space vector are obtained. By calculating the cosine similarity of these two vectors, the images with inaccurate mouth shapes generated by the generation network can be penalized to generate digital human videos with more accurate mouth shapes.
[0146] This embodiment has the following advantages:
[0147] Adopting a deep learning network structure with multi-modal feature alignment, the method for video-driven digital human speech lip-sync is realized. Under real-time conditions, it judges whether the consecutive video frames are synchronized with the audio.
[0148] Through the two-tower design and two-stage training method, more accurate audio-visual synchronization judgment is realized, and it also has good performance for videos recorded in the wild, with better generalization.
[0149] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed.
Claims
1. A digital human voice lip synchronization method, characterized in that: The method comprises: Acquire a source video, and preprocess the source video to obtain audio data and image data; The lip synchronization model is used to synchronously process the audio coding features and the image coding features to obtain synchronization features; the audio coding features are output by the audio data through the lip synchronization model; the image coding features are obtained by image segmentation, feature fusion and feature coding by the lip synchronization model; the synchronization processing includes linear feature projection and feature similarity calculation; Generating a video according to the synchronization feature to obtain a target video; The lip sync model includes: An image processing module, wherein the image processing module is configured to perform image coding according to the image data to obtain a plurality of image coding features; An audio processing module, wherein the audio processing module is configured to perform audio encoding according to the audio data to obtain a plurality of audio encoding features; A synchronization processing module, wherein the synchronization processing module is configured to perform audio and image synchronization processing according to the audio coding feature and the image coding feature to obtain the synchronization feature; The image processing module comprises: a first grouping unit, wherein the first grouping unit is configured to divide the input image data into a plurality of continuous image sequences; each of the image sequences includes a plurality of continuous image frames; A feature fusion unit, wherein the feature fusion unit includes a plurality of fusion layers that are distributed in parallel, and the plurality of fusion layers are configured to perform feature fusion on the image frames in the plurality of image sequences respectively to obtain a plurality of corresponding first sequences; An image coding unit, wherein the image coding unit comprises a plurality of image coding layers using parallel distribution, wherein the plurality of image coding layers correspond one-to-one to the plurality of fusion layers, and the plurality of image coding layers are configured to respectively perform convolution coding on a plurality of the first sequences to obtain a plurality of corresponding image coding features.
2. A digital human voice lip synchronization method according to claim 1, characterized in that: The audio processing module comprises: a second grouping unit, wherein the second grouping unit is configured to divide the input audio data into a plurality of continuous audio sequences; each of the audio sequences includes a plurality of continuous audio sub-data, and each of the audio sub-data corresponds to one of the image frames; An audio encoding unit, wherein the audio encoding unit comprises a plurality of audio encoding layers that are distributed in parallel, and the plurality of audio encoding layers are configured to respectively perform convolution encoding on a plurality of the audio sequences to obtain a plurality of corresponding audio encoding features.
3. A digital human voice lip synchronization method according to claim 1, characterized in that: The image encoding unit further comprises a plurality of spatial attention layers using parallel distribution, and the plurality of spatial attention layers correspond one to one with the plurality of image encoding layers; The spatial attention layer is further configured to calculate a spatial attention weight according to the first sequence; The image encoding layer is also configured to perform convolution encoding according to the first sequence and the spatial attention weight to output the image encoding feature.
4. A digital human voice lip synchronization method according to claim 3, characterized in that: The image encoding unit also includes a channel attention layer; The image encoding unit is further configured to divide the first sequence into a plurality of subsequences, perform feature encoding processing on each subsequence, calculate attention weights corresponding to different subsequences through the spatial attention layer, and fuse a plurality of feature encoding processing results based on the attention weight calculation results to obtain local feature encoding; The local feature code is divided into multiple subsequences, feature coding processing is performed on each subsequence, attention weights corresponding to different subsequences are calculated by a preset spatial attention module, and multiple feature coding processing results are fused based on the attention weight calculation results to obtain the local feature code; Iteratively updating, and independently outputting the local feature codes obtained in each round to obtain multiple local feature codes; Attention weights are calculated for the multiple local feature codes according to the channel attention layer, and feature processing is performed based on the attention weight calculation results to obtain a global feature code; the multiple local feature codes are fused with the global feature code to obtain the image coding feature.
5. A digital human voice lip synchronization method according to claim 2, characterized in that: The synchronization processing module comprises: A feature projection module, wherein the feature projection module includes a plurality of linear layers, wherein the plurality of linear layers correspond one-to-one to the plurality of image coding layers and the plurality of audio coding layers, respectively; the plurality of linear layers are configured to map the image coding features and the audio coding features into a multimodal embedding space; A similarity calculation module, wherein the similarity calculation module includes a plurality of calculation layers, and the plurality of calculation layers correspond one-to-one to the plurality of linear layers; the plurality of calculation layers are configured to calculate the similarity of the image coding features and the audio coding features in the multimodal embedding space, and perform feature fusion on the image coding features and the audio coding features with the highest similarity to obtain the synchronization features.
6. A digital human voice lip synchronization method according to claim 1, characterized in that: The training process of the lip synchronization model includes: Obtaining a training video, and extracting audio training data and image training data based on the training video; Performing model training on the audio training data, the image training data and the lip synchronization model; The lip synchronization model is converged using a loss function until it meets the preset model requirements.
7. A digital human voice lip synchronization method according to claim 6, characterized in that: The step of using the loss function to converge the lip synchronization model comprises: Performing contrast loss training on the synchronization features using a contrast loss function; When the synchronization feature meets the requirements of the first preset model, the lip synchronization model is trained with positive samples and negative samples using a ternary loss function until it meets the requirements of the second preset model.
8. A digital human voice lip synchronization method according to claim 7, characterized in that: The step of performing positive sample training and negative sample training on the lip synchronization model comprises: Acquire an anchor feature, a positive feature, and a negative feature; the anchor feature is a randomly selected image feature or audio feature; the positive feature is an image feature or audio feature that matches the anchor feature; the negative feature is an image feature or audio feature that does not match the anchor feature; Positive sample training is performed according to the anchor feature and the positive feature, and negative sample training is performed according to the anchor feature and the negative feature, until both the positive sample training and the negative sample training meet the second preset model requirement.
9. A digital human voice lip synchronization method according to claim 1, characterized in that: The step of preprocessing the source video comprises: Extracting a source image according to the source video, and sequentially performing image framing processing, face detection processing, face region cropping processing, occlusion detection processing, and image covering processing on the source image to obtain the image data; The source audio is extracted according to the source video, and the source audio is subjected to audio frame processing and audio denoising processing in sequence to obtain the audio data.
Citation Information
Patent Citations
Real-time multi-language processing live broadcast method and system based on deep learning
CN117253486A