Digital human voice lip synchronization method and device based on feature perception and conditional adversarial network
By combining feature perception and conditional adversarial network methods, the feature extraction capability of the Wav2lip model is improved, and the shortcomings of the existing sound and lips synchronization methods in complex scenarios are solved, and high-quality digital human sound and lips synchronization effect is achieved.
Patent Information
- Application Number
- CN202510455067.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-24
AI Technical Summary
The existing lip synchronization methods have problems such as insufficient feature extraction capabilities, insufficient naturalness and coordination of the generated results when dealing with complex scenarios, making it difficult to achieve high-quality digital human voice synchronization effects.
Using a method based on feature perception and conditional adversarial network, the feature extraction capability of the Wav2lip model is improved, accurate alignment and fusion of audio and video features is achieved, and high-quality sound and lip synchronous video is generated.
It improves the feature extraction and mapping ability of digital human voice lip synchronization, enhances the authenticity and nature of the generated results, and realizes high-quality voice lip synchronization effect, suitable for complex scenes and multi-angle facial expressions.
Smart Images

Figure CN120201226A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision and deep learning, and particularly relates to a digital human audio-visual lip synchronization method and device based on feature perception and conditional adversarial network. Background Art
[0002] Digital human technology is an important research direction in the fields of modern artificial intelligence and computer vision. Among them, audio-visual lip synchronization, as one of the key technologies of digital humans, is of great significance for realizing natural and smooth human-computer interaction. The audio-visual lip synchronization technology requires that the generated facial expressions of digital humans, especially the lip shape changes, need to be highly synchronized and coordinated with the speech content.
[0003] Currently, the mainstream audio-visual lip synchronization methods are mainly based on deep learning technology. Among them, Wav2lip, as a classic audio-visual lip synchronization model, extracts features through an audio encoder and a video encoder to realize the mapping relationship between audio and video features. However, the traditional Wav2lip model still has some limitations in dealing with complex scenarios: First, the feature extraction ability of the model is limited, and it is difficult to capture the fine changes in facial details; Second, the generated lip movements lack naturalness and often show mechanical and rigid characteristics; Finally, the overall coordination of facial expressions is insufficient, affecting the realism of digital humans.
[0004] In recent years, feature perception and conditional adversarial network have made remarkable progress in the field of computer vision. Feature perception technology can deeply understand the semantic information and detailed features of images through multi-level feature extraction and fusion, providing the possibility for precise control of facial expressions. The conditional adversarial network, through the adversarial learning mechanism of the generator and the discriminator, continuously improves the authenticity and naturalness of the generated results. In tasks such as image generation and style transfer, these two technologies have shown excellent performance. However, there is currently a lack of a solution that effectively combines these two technologies and applies them to the audio-visual lip synchronization task, especially in dealing with fast speech, multi-angle facial expressions, and maintaining detailed features, etc., which still needs to be broken through.
[0005] Currently, the existing audio-visual lip synchronization methods still face challenges in the following aspects. For example, the feature expression ability is insufficient, and it is difficult to accurately capture the subtle correspondence between audio and video; the temporal consistency of the generated results is poor, and the transition between adjacent frames is not smooth enough; the diversity of facial expressions is insufficient, and it is difficult to adapt to different speaking styles and emotional expressions; there is a lack of an effective quality assessment mechanism, and it is difficult to accurately control the quality of the generated results.
[0006] Therefore, how to combine the advantages of feature perception and conditional adversarial network, improve the performance of the traditional Wav2lip model, enhance the feature extraction and mapping capabilities, enhance the authenticity and naturalness of the generated results, and achieve high-quality digital human audio-visual lip synchronization has become a technical problem to be solved urgently. This is not only of great significance to improving the overall level of digital human technology but also will provide strong support for the popularization of digital humans in various application fields. Summary of the Invention
[0007] Object of the Invention: Aiming at the above problems, the present invention proposes a digital human audio-visual lip synchronization method and device based on feature perception and conditional adversarial network. By improving the feature extraction ability of the Wav2lip model and combining a deep feature perception module and a conditional adversarial network, high-quality audio-visual lip synchronization effects are achieved.
[0008] Technical Solution: The present invention proposes a digital human audio-visual lip synchronization method based on feature perception and conditional adversarial network, including the following steps:
[0009] Step 1: Collect digital human audio-visual data and preprocess it to obtain a preprocessed audio and video dataset.
[0010] Step 2: Extract audio and video features through a multimodal encoder, perform cross-modal feature fusion using a deep feature perception module, design a cascaded feature optimization network to optimize the fused features, and generate enhanced audio-visual lip synchronization feature X.
[0011] Step 3: Use a conditional generative adversarial network to perform accurate alignment training of audio and video features, and evaluate the authenticity of the generated audio-visual lip synchronization effect through a discriminator to generate adversarially optimized alignment feature A.
[0012] Step 4: Fuse the output results of the deep feature perception module and the conditional adversarial network, perform hierarchical parsing on the fused features, decode the fused features through a three-level decoding architecture, perform local optimization adjustment on abnormal regions, and synthesize the optimized audio-visual lip synchronization video into the original video to generate the final high-resolution audio-visual lip synchronization video.
[0013] Further, the specific method of Step 1 is as follows:
[0014] Step 1.1: Collect audio-visual data of digital human products, including original audio and standard facial expression videos, and the same video data should include facial expression changes from the front, side, and multiple angles.
[0015] Step 1.2: Perform rotation transformation, scaling transformation, and flipping transformation on the collected video data.
[0016] Step 1.3: Normalize the audio data, perform time synchronization and alignment on the processed audio-visual data, establish a mapping relationship, and form a standardized training dataset.
[0017] Furthermore, the specific method of step 2 is as follows:
[0018] Step 2.1: Construct a multi-modal encoder including an audio encoder and a video encoder, input the preprocessed audio and video data into the audio encoder and the video encoder respectively, output the acoustic feature matrix A and the facial feature tensor V, and process the acoustic feature matrix A and the facial feature tensor V through a fully connected layer or a pooling layer to obtain the audio feature vector p and the video feature vector q;
[0019] Step 2.2: Construct a deep feature perception module DFPM, use a feature extraction network with a three-layer 3D convolutional structure to extract features from the acoustic feature matrix A and the facial feature tensor V, and use a feature fusion network with an attention mechanism to adaptively weight and fuse the extracted audio features {F1_a, F2_a, F3_a} and video features {F1_v, F2_v, F3_v}, and output the fused feature tensor F;
[0020] Step 2.3: Construct a feature optimization network, input the fused feature tensor F output by the deep feature perception module DFPM into a three-layer fully connected layer (256-512-256) for optimization, add a BatchNorm layer and a ReLU activation function between each layer, and finally use the Tanh function for normalization to output the optimized feature vector group H;
[0021] Step 2.4: Concatenate the initial audio feature vector p and the video feature vector q obtained in step 2.1 with the optimized feature H output in step 2.3 in the channel dimension to form a mixed feature group M∈R (d1+d2+d3) ; Subsequently, use a gated attention mechanism to assign weights to the features of each modality in M, and the calculation method is:
[0022] g = σ(W g ·M + b g )
[0023] M' = G⊙M
[0024] where g is used to calculate the weights of the audio and video modalities, G is the weighted fusion of the original features, W g is a learnable parameter matrix, σ is the Sigmoid function, and ⊙ represents element-wise multiplication; finally, add the weighted feature M' and the fused feature of DFPM residually to output the enhanced audio-lip synchronization feature X;
[0025] Step 2.5: Input the enhanced audio-visual synchronization feature X into the decoder network to generate the final audio-visual synchronization video sequence. The decoder network adopts a combination of deconvolution layers and upsampling layers to gradually restore the spatial details of the video frames.
[0026] Furthermore, when using the Deep Feature Perception Module (DFPM) to output the fused feature tensor, the following operations are also included:
[0027] Step 2.2.1: Design a feature correlation metric function: For the audio feature vector p and the video feature vector q, define the correlation weight ω as:
[0028] ω = σ(α · exp(-||p - q|| 2 ))
[0029] where σ is the sigmoid function, α is a learnable scaling parameter, and ||p - q|| represents the Euclidean distance between the feature vectors;
[0030] Step 2.2.2: Design a feature perception loss function, which consists of three parts:
[0031] Content loss L _content used to ensure the visual quality of the generated video, and its expression is:
[0032] L _content = ||φ(G(x)) - φ(y)||2 2
[0033] where G(x) is the video frame output by the generator, y is the corresponding real video frame, and finally the content loss is obtained by calculating the difference between the generated image G(x) and the target image y in the feature space φ;
[0034] Temporal consistency loss L _temp used to maintain the continuity of the video sequence, and its expression is:
[0035] L _temp = ∑||f t - f t+1 ||1
[0036] where f t 、f t+1 represent the fused features at the t-th time step and the (t + 1)-th time step;
[0037] Synchronization loss L _sync used to enhance the cross-modal alignment of audio-visual features, and its expression is:
[0038] L _sync = 1 - ω = 1 - σ(α · exp(-||p - q|| 2 ))
[0039] Among them, ω is the correlation weight;
[0040] The final comprehensive loss function is:
[0041] L _total = λ1L _content + λ2L _temp + λ3L _sync
[0042] Among them, λ1, λ2, and λ3 are the weight coefficients of each loss term, used to balance the contributions of different loss terms, and are set to 1.0, 0.5, and 0.2 respectively. Their optimal ratios are determined through experiments.
[0043] Furthermore, the specific method of step 3 is as follows:
[0044] Step 3.1: Construct a lip-sync training system based on the conditional generative adversarial network GAN, including a generator G and a discriminator D; the generator G uses a multi-scale Transformer architecture to process audio-visual features, and the discriminator D is a multi-branch convolutional network, which respectively evaluates global content consistency, local detail authenticity, and cross-modal synchronization. The input of the lip-sync training system based on the conditional generative adversarial network GAN is the enhanced lip-sync feature X, and feature space alignment is achieved through adversarial training;
[0045] Step 3.2: Divide the enhanced lip-sync feature X into N×N patch blocks, and perform feature extraction through L-layer Transformer encoding blocks. The encoding process is expressed as:
[0046] H i = TransformerBlock(H i-1 )
[0047] Among them, H0 is the patch embedding after linear projection, H0 ∈ R P×d , P is the number of patches, d = 768 is the feature dimension. Each encoding block contains multi-head self-attention and an MLP expansion layer, representing the processing process of the i-th layer Transformer encoding block on the input feature H i-1 , and outputs the updated feature H i ;
[0048] Step 3.3: Select 4 feature groups from the encoder output, and select the output H {6k} of the corresponding Transformer block for each group, where k ∈ {4, 8, 12, 16} as the key feature vector F k ∈ R P×d ;
[0049] Step 3.4: The key feature vector Fk Input the bidirectional multi-scale feature aggregation module BiMLA. Through top-down and bottom-up dual-path processing, output 4 coarse-grained features G k ∈R P / 4×d , and fuse them through addition between the paths. The formula is:
[0050] G k = Upsample(F k ) + Downsample(F k+1 )
[0051] Among them, Upsample is the upsampling operation, which transmits F k from the low layer to the high layer; Downsample is the downsampling operation, which transmits the high-layer features of F k+1 from the high layer to the low layer, and transmits both high-layer semantics and low-layer details at the same time;
[0052] Step 3.5: Divide the input feature map into 8 local regions, each region is 32×32 pixels, and each region is encoded by 6 layers of Transformer to obtain local features m = 8, and the window attention is used in the encoding process;
[0053] Step 3.6: Select the output of the middle layer of the local feature L j as the input of BiMLA to generate 8 fine-grained features Calculation formula:
[0054] S j = Conv 1×1 (Concat[L_{i,3}, L_{i,4}])
[0055] Among them, L_{i,3} represents the output feature after the i-th local region is encoded by the 3rd layer of Transformer, L_{i,4} represents the output feature after the i-th local region is encoded by the 4th layer of Transformer, Conv 1×1 is the channel adjustment convolution, and Concat is used to splice two features along the channel dimension;
[0056] Step 3.7: Integrate the global feature G k and the local feature S j through the feature fusion module FFM. First, calculate the fusion weight γ of the two, and then combine the features according to the weight:
[0057] γ = σ(Conv 1×1 (Concat[G k , S j ))
[0058] A = γ·Gk +(1 - γ)·S j + Residual(G k , S j )
[0059] Meanwhile, a residual connection is added, and the final output alignment feature A is consistent with the encoder dimension;
[0060] Step 3.8: Input the alignment feature A generated in Step 3.7 into the discriminator D for authenticity evaluation, and minimize the distribution difference between the generated feature and the real audio-visual synchronization feature through adversarial training, comprehensively considering content similarity, temporal consistency, and feature synchronization;
[0061] Step 3.9: Input the alignment feature A optimized by adversarial training into the cascaded upsampling decoder, gradually restore the spatial resolution through 4 layers of transposed convolution, and fuse the audio features extracted by the multi-modal encoder as conditional parameters for each layer to generate the final high-quality audio-visual synchronization video sequence.
[0062] Furthermore, the specific method of Step 4 is as follows:
[0063] Step 4.1: Fuse the enhanced audio-visual synchronization feature X output by the deep feature perception module and the alignment feature A generated by the conditional adversarial network through a dynamic weight mechanism. The fusion process is expressed as:
[0064] F _final = β·X + (1 - β)·A + Residual(X)
[0065] where β is the dynamic fusion weight, and its value range is [0.3, 0.7];
[0066] Step 4.2: Perform hierarchical parsing on the fused feature F through the multi-scale feature pyramid network FPN _final and use the local attention module LAM to generate the attention weight map of the facial region, focusing on the feature representation of the key facial regions;
[0067] Step 4.3: Adopt a three-level decoding architecture to process the fused feature F _final : The basic decoder gradually restores the video resolution through a 5-layer transposed convolution network; the temporal optimizer combines 3D convolution and LSTM units to ensure the natural transition of lip movements between consecutive frames; the region refiner focuses on enhancing the features of the lip region based on the facial attention map;
[0068] Step 4.4: Define and calculate the key region evaluation indicators including lip shape accuracy, movement naturalness, and expression coordination, and generate a facial feature analysis report;
[0069] Step 4.5: Apply spatio-temporal filtering to the generated result to eliminate jitter, perform color and illumination compensation, and locally optimize and adjust the abnormal area according to the evaluation index;
[0070] Step 4.6: Synthesize the optimized audio-visual lip-sync video into the original video to generate the final high-resolution audio-visual lip-sync video.
[0071] The present invention also discloses a digital human audio-visual lip-sync device based on feature perception and conditional adversarial network, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the digital human audio-visual lip-sync method based on the above-mentioned feature perception and conditional adversarial network.
[0072] Beneficial effects:
[0073] The present invention adopts deep learning technology, combines a feature perception module and a conditional adversarial network, and realizes a more accurate digital human audio-visual lip-sync effect. Specifically, by introducing a deep feature perception module to optimize the Wav2lip model, more accurate extraction and processing of facial detail features are realized. In addition, a multi-scale feature pyramid and an attention mechanism are added to the feature extraction network, which improves the model's recognition ability for complex facial expressions and enhances the adaptability to different speaking styles and emotional expressions. Combining the conditional generative adversarial network for feature alignment optimization further improves the accuracy of audio-visual synchronization. Finally, through feature fusion technology, the output results of the feature perception module and the conditional adversarial network are combined to achieve a more accurate audio-visual lip-sync effect, and at the same time, the accuracy of facial details is maintained through the deep feature perception loss function, which has a better processing effect on complex expression changes, such as fast speech and subtle lip shape changes. Brief description of the drawings
[0074] Figure 1 It is the overall flowchart of the digital human audio-visual lip-sync method and device based on feature perception and conditional adversarial network;
[0075] Figure 2 It is the data flow diagram during training;
[0076] Figure 3 It is the overall model diagram of wav2lip;
[0077] Figure 4 It is the model diagram of Conditional GAN. Detailed implementation manners
[0078] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art fall within the scope defined by the appended claims of this application.
[0079] The present invention discloses a digital human audio-visual lip synchronization method and device based on feature perception and conditional adversarial network. The digital human audio-visual lip synchronization method based on feature perception and conditional adversarial network includes the following steps:
[0080] Step 1: Collect digital human audio-visual data and preprocess it to obtain an original data set.
[0081] Step 1.1: Collect audio-visual data of digital human products, including original audio and standard facial expression videos, and the same video data should include frontal, side, and multi-angle expression changes.
[0082] Step 1.2: Perform technical means such as rotation transformation, scaling transformation, and flipping transformation on the collected video data.
[0083] Step 1.3: Normalize the audio data, synchronize and align the processed audio-visual data in time, establish a mapping relationship, and form a standardized training data set.
[0084] Step 2: Improve the feature extraction effect of audio-visual data by optimizing the deep feature perception module, obtain more accurate audio-visual features with deep semantic information, and provide high-quality input data for subsequent audio-visual lip synchronization training.
[0085] Step 2.1: Construct a multi-modal encoder, including an audio encoder and a video encoder. See Figure 2 : The audio encoder extracts acoustic features through short-time Fourier transform (STFT) and Mel spectrogram conversion, and outputs an acoustic feature matrix A; the video encoder uses 3D convolution combined with temporal attention and spatial attention mechanisms to extract a facial feature tensor V containing facial expression and lip shape features; the preprocessed audio and video data are respectively input into the audio encoder and the video encoder, and the acoustic feature matrix A is processed through a fully connected layer or a pooling layer to obtain an audio feature vector p, and at the same time, a video feature vector q is extracted from the facial feature tensor V to establish a preliminary audio-visual lip synchronization feature mapping.
[0086] Step 2.2: Construct the Deep Feature Perception Module (DFPM). Use a feature extraction network with a three-layer 3D convolutional structure to extract features from the acoustic feature matrix A and the facial feature tensor V. Use a feature fusion network with an attention mechanism to adaptively weight and fuse the extracted audio features {F1_a, F2_a, F3_a} and video features {F1_v, F2_v, F3_v}, and output the fused feature tensor F.
[0087] Step 2.3: Design a feature correlation measurement function: For the audio feature vector p and the video feature vector q, define the correlation weight ω as:
[0088] ω = σ(α·exp(-||p - q|| 2 ))
[0089] where σ is the sigmoid function, α is a learnable scaling parameter, and ||p - q|| represents the Euclidean distance between the feature vectors.
[0090] Step 2.4: Based on the fused features output by the Deep Feature Perception Module (DFPM), design a cascaded feature optimization network, whose structure includes a primary optimization layer, an intermediate optimization layer, and an output layer. The primary optimization layer is a 256-dimensional fully connected layer that receives the fused features output by the DFPM and performs non-linear transformation through the BatchNorm layer and the ReLU activation function; the intermediate optimization layer is a 512-dimensional fully connected layer that further extracts high-order cross-modal features and uses a Dropout layer (probability 0.3) to prevent overfitting; the output layer is a 256-dimensional fully connected layer that normalizes the features to the interval [-1, 1] through the Tanh function. Input the fused feature tensor F output by the Deep Feature Perception Module (DFPM) into a three-layer fully connected layer (256 - 512 - 256) for optimization, add a BatchNorm layer and a ReLU activation function between each layer, and finally use the Tanh function to normalize the output to obtain the optimized feature vector group H. This step eliminates the redundant information between multi-modal features and enhances the explicit expression of the audio-lip correlation features through hierarchical feature compression and non-linear mapping.
[0091] Step 2.5: Design a feature perception loss function. The loss function consists of three parts and ensures the authenticity, temporal coherence, and cross-modal synchronization of the audio-lip synchronization effect through multi-dimensional constraints:
[0092] Content loss L _content is used to ensure the visual quality of the generated video, and its expression is:
[0093] L _content = ||φ(G(x)) - φ(y)||2 2
[0094] Among them, G(x) is the video frame output by the generator, and y is the corresponding real video frame. This loss ensures that the generated facial expressions and lip movements conform to natural laws by comparing the differences in deep visual features.
[0095] Temporal consistency loss L _temp is used to maintain the continuity of the video sequence, and its expression is:
[0096] L _temp = ∑||f t - f t+1 ||1
[0097] where f t represents the fused feature at the t-th time step. This loss eliminates the jitter phenomenon between video frames by constraining the feature differences between adjacent frames, ensuring a smooth transition of lip movements.
[0098] Synchronization loss L _sync is used to enhance the cross-modal alignment of audio-visual features, and its expression is:
[0099] L _sync = 1 - ω = 1 - σ(α·exp(-||p - q|| 2 ))
[0100] where ω is the correlation weight defined in step 2.3. This loss precisely matches the speech content with the lip movements by maximizing the inter-modal feature correlation.
[0101] The final comprehensive loss function is:
[0102] L _total = λ1L _content + λ2L _temp + λ3L _sync
[0103] where λ1, λ2, and λ3 are the weight coefficients of each loss term, used to balance the contributions of different loss terms, and are set to 1.0, 0.5, and 0.2 respectively. Their optimal ratios are determined through experiments.
[0104] Step 2.6: Concatenate the initial audio feature vector p, video feature vector q obtained in step 2.1 with the optimized feature H output in step 2.4 along the channel dimension to form a mixed feature group M ∈ R (d1+d2+d3) . Subsequently, the gated attention mechanism is used to allocate weights to the features of each modality in M, and the calculation method is:
[0105] g = σ(W_g·M + b_g)
[0106] M' = G⊙M
[0107] Among them, g is used to calculate the weights of the audio and video modalities, G is the weighted fusion of the original features, W g is a learnable parameter matrix, σ is the Sigmoid function, and ⊙ represents element-wise multiplication; finally, the weighted feature M' and the fusion feature of DFPM are added residually to output the enhanced audio-visual lip-sync feature X;
[0108] Step 2.7: Input the enhanced audio-visual lip-sync feature X into the decoder network to generate the final audio-visual lip-sync video sequence. The decoder uses a combination of deconvolution layers and upsampling layers to gradually restore the spatial details of the video frames, ensuring the authenticity and smoothness of the generated results.
[0109] Step 3: Through the newly added conditional generative adversarial network, achieve accurate alignment training of audio-visual features, and use the discriminator to evaluate the authenticity of the generated audio-visual lip-sync effect, so as to obtain a more natural and efficient audio-visual lip-sync effect and improve the accuracy and authenticity of the overall system.
[0110] Step 3.1: Construct a conditional generative adversarial network GAN to achieve alignment mapping of audio-visual features through the generator G and the discriminator D, including global and local feature extraction and optimization;
[0111] Step 3.1: Construct an audio-visual lip-sync training system based on the conditional generative adversarial network GAN, including the generator G and the discriminator D. The generator G uses a multi-scale Transformer architecture to process audio-visual features, and the discriminator D is a multi-branch convolutional network, which respectively evaluates global content consistency, local detail authenticity, and cross-modal synchronization. The input X is the enhanced audio-visual lip-sync feature X output in Step 2.7, and feature space alignment is achieved through adversarial training.
[0112] Step 3.2: Divide the input feature sequence X into N×N patch blocks (by default, N = 16), and perform feature extraction through L layers of Transformer encoding blocks (L = 24). The encoding process is expressed as:
[0113] H i = TransformerBlock(H i-1 )
[0114] Among them, H0 is the patch embedding after linear projection, H0 ∈ R P×d , P is the number of patches, d = 768 is the feature dimension. Each encoding block contains multi-head self-attention and an MLP expansion layer, representing the processing process of the i-th layer of Transformer encoding block for the input feature H i-1 , and outputs the updated feature H i . Each encoding block contains multi-head self-attention (number of heads = 12) and an MLP expansion layer (expansion ratio = 4).
[0115] Step 3.3: Select K = 4 groups of features from the encoder output, and for each group, select the output H of the corresponding Transformer block {6k} (k ∈ {4, 8, 12, 16}) as the key feature vector F k ∈ R P×d . This design retains features at different abstraction levels, where low-level features capture the details of mouth movements and high-level features model the semantic associations of speech lip shapes.
[0116] Step 3.4: Input the key feature vector F k into the bidirectional multi-scale feature aggregation module BiMLA. Through top-down and bottom-up dual-path processing, output 4 coarse-grained features G k ∈ R P / 4×d . The fusion between paths is through addition, and the formula is:
[0117] G k = Upsample(F k ) + Downsample(F k+1 )
[0118] where Upsample is the upsampling operation, which transfers F k from low level to high level; Downsample is the downsampling operation, which transfers the high-level features of F k+1 from high level to low level, while transferring high-level semantics and low-level details.
[0119] Step 3.5: Divide the input feature map into K = 8 local regions (each region is 32×32 pixels), and each region is encoded by M = 6 layers of Transformer to obtain local features m = 8. The window attention (window size = 4×4) is used in the encoding process to reduce the computational amount while maintaining the sensitivity to local details.
[0120] Step 3.6: Select the output of the middle layer of the local feature L j as the input of BiMLA to generate 8 fine-grained features Calculation formula:
[0121] S j = Conv 1×1 (Concat[L_{i,3}, L_{i,4}])
[0122] where L_{i,3} represents the output feature of the i-th local region after being encoded by the 3rd layer of Transformer, L_{i,4} represents the output feature of the i-th local region after being encoded by the 4th layer of Transformer, and Conv 1×1It is a channel adjustment convolution, and Concat is used to splice two features along the channel dimension.
[0123] Step 3.7: Integrate the global feature G k and the local feature S j through the Feature Fusion Module (FFM). First, calculate the fusion weight γ of the two, and then combine the features according to the weight:
[0124] γ = σ(Conv 1×1 (Concat[G k , S j ))
[0125] A = γ·G k + (1 - γ)·S j + Residual(G k , S j )
[0126] At the same time, add a residual connection, and the final output aligned feature A is consistent with the encoder dimension.
[0127] Step 3.8: Input the aligned feature A generated in Step 3.7 into the discriminator D for authenticity evaluation, and minimize the distribution difference between the generated feature and the real audio-visual synchronization feature through adversarial training, taking into account content similarity, temporal consistency, and feature synchronization.
[0128] Step 3.9: Input the feature A optimized by the adversarial training into the cascaded upsampling decoder, and gradually restore the spatial resolution through 4 layers of transposed convolution. Each layer fuses the audio features extracted in Step 2.1 as conditional parameters to generate the final high-quality audio-visual synchronization video sequence, ensuring the authenticity and smoothness of the generated result.
[0129] Step 4: Integrate the output results of the feature perception module and the conditional adversarial network to generate the final audio-visual synchronization effect, and at the same time annotate and analyze the features of the key facial regions.
[0130] Step 4.1: Fuse the enhanced audio-visual synchronization feature X output by the deep feature perception module and the aligned feature A generated by the conditional adversarial network through a dynamic weight mechanism. The fusion process is expressed as:
[0131] F _final = β·X + (1 - β)·A + Residual(X)
[0132] where β is the dynamic fusion weight, and its value range is [0.3, 0.7].
[0133] Step 4.2: Hierarchically analyze the fused features through a multi-scale feature pyramid network (FPN), and use a local attention module (LAM) to generate an attention weight map for the facial region, focusing on the feature representation of key facial areas;
[0134] Step 4.3: Adopt a three-level decoding architecture to process the fused feature F _final : The basic decoder gradually upsamples through a 5-layer transposed convolution network to restore the video resolution; the temporal optimizer combines 3D convolution and LSTM units to ensure the natural transition of lip movements between consecutive frames; the region refiner enhances the features of the lip region based on the facial attention map. The decoding process synchronously optimizes three objectives: visual reconstruction quality, motion coherence, and audio-lip synchronization, accurately restoring the lip shape details while ensuring the naturalness of the entire face.
[0135] Step 4.4: Define and calculate key region evaluation metrics including lip shape accuracy, motion naturalness, and expression coordination, and generate a facial feature analysis report to guide subsequent optimization.
[0136] Step 4.5: Apply spatio-temporal filtering to the generated results to eliminate jitter, perform color and illumination compensation, and locally optimize and adjust abnormal regions according to the evaluation metrics to improve the quality of the final audio-lip synchronization effect.
[0137] Step 4.6: Synthesize the optimized audio-lip synchronized video into the original video to ensure the natural and smooth transition of audio and facial expressions. Generate the final high-resolution audio-lip synchronized video to ensure that the output quality is suitable for various application scenarios, such as virtual reality (VR), video production, and real-time translation.
[0138] The present invention also discloses a digital human audio-lip synchronization device based on feature perception and conditional adversarial network, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the above digital human audio-lip synchronization method based on feature perception and conditional adversarial network.
[0139] The following conducts relevant experiments on the digital human audio-lip synchronization method based on feature perception and conditional adversarial network of the present application. Use the same audio and video to conduct a comparison experiment on audio-lip synchronization accuracy and cross-angle adaptability using traditional 3DMM method, LSTM sequence modeling, Wav2Lip (GAN), and the method of this patent respectively. The experimental results are shown in Tables 1 and 2:
[0140] Table 1 Comparison of audio-lip synchronization accuracy of different methods
[0141] Method LMD (mm) SyncConfidenceScore PSNR (dB) AUE (degrees) Traditional 3DMM Method 4.82 0.67 28.5 8.7 LSTM Sequence Modeling 3.91 0.73 30.1 6.9 Wav2Lip (GAN) 2.45 0.86 32.8 5.2 The Method of This Patent 1.78 0.94 35.6 3.8
[0142] Table 2 Comparison of cross-angle adaptability of different methods
[0143] Method Frontal LMD Profile 45° LMD Profile 90° LMD Synchronization Success Rate Wav2Lip 2.45 3.87 6.21 82% SyncNet 2.61 4.12 7.05 78% The Method of This Patent 1.78 2.95 4.33 93%
[0144] As can be seen from the above experimental results, the present invention has higher accuracy in audio-lip synchronization compared to several other conventional methods, and also has better adaptability in cross-angle adaptability compared to the other two methods.
[0145] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A digital human lip synchronization method based on feature perception and conditional adversarial network, characterized in that: The steps include: Step 1: Collect digital human audio and video data and preprocess them to obtain preprocessed audio and video data sets; Step 2: Extract audio and video features through a multimodal encoder, use a deep feature perception module to fuse cross-modal features, design a cascade feature optimization network to optimize fusion features, and generate enhanced lip synchronization features X; Step 3: Use the conditional generative adversarial network to perform accurate alignment training of audio and video features, and use the discriminator to evaluate the authenticity of the generated lip synchronization effect to generate the alignment feature A after adversarial optimization; Step 4: Fuse the output results of the deep feature perception module and the conditional adversarial network, perform hierarchical analysis on the fused features, decode the fused features through a three-level decoding architecture, perform local optimization adjustments on the abnormal areas, and synthesize the optimized lip-sync video into the original video to generate the final high-resolution lip-sync video.
2. The digital human voice lip synchronization method based on feature perception and conditional adversarial network according to claim 1, characterized in that: The specific method of step 1 is: Step 1.1: Collect the audio and video data of the digital human product, including the original audio and standard facial expression video. The same video data should include expression changes from the front, side and multiple angles; Step 1.2: Perform rotation transformation, scaling transformation, and flip transformation on the collected video data; Step 1.3: Normalize the audio data, synchronize and align the processed audio and video data, establish a mapping relationship, and form a standardized training data set.
3. The digital human voice lip synchronization method based on feature perception and conditional adversarial network according to claim 1, characterized in that: The specific method of step 2 is: Step 2.1: Construct a multimodal encoder including an audio encoder and a video encoder, input the preprocessed audio and video data into the audio encoder and the video encoder respectively, output the acoustic feature matrix A and the facial feature tensor V, and process the acoustic feature matrix A and the facial feature tensor V through a fully connected layer or a pooling layer to obtain an audio feature vector p and a video feature vector q; Step 2.2: Construct a deep feature perception module DFPM, use a three-layer 3D convolutional structure feature extraction network to extract features from the acoustic feature matrix A and the facial feature tensor V, use an attention mechanism feature fusion network to perform adaptive weighted fusion on the extracted audio features {F1_a, F2_a, F3_a} and video features {F1_v, F2_v, F3_v}, and output the fused feature tensor F; Step 2.3: Build a feature optimization network, input the fused feature tensor F output by the deep feature perception module DFPM into three fully connected layers (256-512-256) for optimization, add a BatchNorm layer and a ReLU activation function between each layer, and finally use the Tanh function to normalize and output the optimized feature vector group H; Step 2.4: Concatenate the initial audio feature vector p and video feature vector q obtained in step 2.1 with the optimized feature H output in step 2.3 in the channel dimension to form a mixed feature group M∈R (d1+d2+d3) ; Then the gated attention mechanism is used to assign weights to the modal features in M, calculated as: g=σ(W g ·M+b g ) M'=G⊙M Among them, g is used to calculate the weights of audio and video modalities, G is the weighted fusion of the original features, and W g is a learnable parameter matrix, σ is a Sigmoid function, and ⊙ represents element-by-element multiplication. Finally, the weighted feature M' is residually added to the fusion feature of DFPM to output the enhanced lip synchronization feature X. Step 2.5: The enhanced lip-sync feature X is input into the decoder network to generate the final lip-sync video sequence. The decoder network uses a combination of deconvolution layers and upsampling layers to gradually restore the spatial details of the video frames.
4. The digital human voice lip synchronization method based on feature perception and conditional adversarial network according to claim 3 is characterized in that: When using the deep feature perception module DFPM to output the fused feature tensor, the following operations are also included: Step 2.2.1: Design feature correlation measurement function: For audio feature vector p and video feature vector q, define the correlation weight ω as: ω=σ(α·exp(-||pq|| 2 )) Among them, σ is the sigmoid function, α is a learnable scaling parameter, and ||pq|| represents the Euclidean distance between feature vectors; Step 2.2.2: Design a feature-aware loss function, which consists of three parts: Content loss L _content Used to ensure the visual quality of the generated video, the expression is: L _content =||φ(G(x))-φ(y)||2 2 Among them, G(x) is the video frame output by the generator, y is the corresponding real video frame, and finally the content loss is obtained by calculating the difference between the generated image G(x) and the target image y in the feature space φ; Timing consistency loss L _temp Used to maintain the continuity of the video sequence, its expression is: L _temp =Σ||f t -f t+1 ||1 Among them, f t 、f t+1 Represents the fusion features of the tth time step and the t+1th time step; Synchronicity loss L _sync The cross-modal alignment used to enhance audio and video features is expressed as: L_ sync =1-ω=1-σ(α·exp(-||pq|| 2 )) Among them, ω is the correlation weight; The final comprehensive loss function is: L_ total =λ1L_ content +λ2L_ temp +λ3L_ sync Among them, λ1, λ2, and λ3 are weight coefficients of each loss term, which are used to balance the contribution of different loss terms. They are set to 1.0, 0.5, and 0.2, respectively, and their optimal ratios are determined through experiments.
5. The digital human voice lip synchronization method based on feature perception and conditional adversarial network according to claim 1, characterized in that: The specific method of step 3 is: Step 3.1: Construct a lip synchronization training system based on conditional generative adversarial network (GAN), which includes a generator G and a discriminator D. The generator G uses a multi-scale Transformer architecture to process audio and video features. The discriminator D is a multi-branch convolutional network that evaluates global content consistency, local detail authenticity, and cross-modal synchronization. The input of the lip synchronization training system based on conditional generative adversarial network (GAN) is the enhanced lip synchronization feature X, and feature space alignment is achieved through adversarial training. Step 3.2: Divide the enhanced lip synchronization feature X into N×N patch blocks, and extract features through L-layer Transformer encoding blocks. The encoding process is expressed as: H i =TransformerBlock(H i-1 ) Among them, H0 is the patch embedding after linear projection, H0∈R P×d , P is the number of patches, d = 768 is the feature dimension, each encoding block contains multi-head self-attention and MLP expansion layers, indicating that the i-th layer Transformer encoding block has an input feature H i-1 The processing process outputs the updated feature H i ; Step 3.3: Select 4 feature groups from the encoder output, each of which corresponds to the output H of the Transformer block {6k} , k∈{4,8,12,16} is used as the key feature vector F k ∈R P×d ; Step 3.4: Transform the key feature vector F k Input the bidirectional multi-scale feature aggregation module BiMLA, through top-down and bottom-up dual path processing, output 4 coarse-grained features G k ∈R P / 4×d , the paths are fused by addition, the formula is: G k =Upsample(F k )+Downsample(F k+1 ) Among them, Upsample is the upsampling operation, which converts F k The lower layer is passed to the higher layer; Downsample is a downsampling operation, which transfers F k+1 High-level features are passed to low-level layers, delivering both high-level semantics and low-level details; Step 3.5: Divide the input feature map into 8 local regions, each with 32×32 pixels, and encode each region through 6 layers of Transformer to obtain local features The encoding process uses window attention; Step 3.6: Select local features L j The intermediate layer output of is used as BiMLA input to generate 8 fine-grained features Calculation formula: S j =Conv 1×1 (Concat[L_{i,3},L_{i,4}]) Among them, L_{i,3} represents the output feature of the i-th local area after being encoded by the third layer of Transformer, L_{i,4} represents the output feature of the i-th local area after being encoded by the fourth layer of Transformer, Conv 1×1 For channel-adjusted convolution, Concat is used to concatenate two features along the channel dimension; Step 3.7: The global feature G is fused through the feature fusion module FFM k and local features S j To integrate, first calculate the fusion weight γ of the two, and then combine the features according to the weight: γ=σ(Conv 1×1 (Concat[G k ,S j ])) A=γ·G k +(1-γ)·S j +Residual(G k ,S j ) At the same time, residual connections are added, and the final output alignment feature A is consistent with the encoder dimension; Step 3.8: Input the alignment feature A generated in step 3.7 into the discriminator D for authenticity evaluation. Through adversarial training, the distribution difference between the generated features and the real lip synchronization features is minimized, and content similarity, temporal consistency and feature synchronization are comprehensively considered; Step 3.9: Input the adversarially optimized aligned feature A into the cascade upsampling decoder, and gradually restore the spatial resolution through 4 layers of deconvolution. Each layer fuses the audio features extracted by the multimodal encoder as conditional parameters to generate the final high-quality lip-sync video sequence.
6. The digital human voice lip synchronization method based on feature perception and conditional adversarial network according to claim 1, characterized in that: The specific method of step 4 is: Step 4.1: The enhanced lip synchronization feature X output by the deep feature perception module is fused with the alignment feature A generated by the conditional adversarial network through a dynamic weight mechanism. The fusion process is expressed as: F_ final =β·X+(1-β)·A+Residual(X) Among them, β is the dynamic fusion weight, and its value range is [0.3, 0.7]; Step 4.2: Fusion feature F_ through multi-scale feature pyramid network FPN final Perform hierarchical analysis and use the local attention module LAM to generate the attention weight map of the facial area, focusing on the feature representation of the key facial areas; Step 4.3: Use a three-level decoding architecture to process the fused feature F_ final :The basic decoder gradually upsamples and restores the video resolution through a 5-layer deconvolution network; the timing optimizer combines 3D convolution and LSTM units to ensure the natural transition of lip motion between consecutive frames; the region refiner focuses on enhancing the lip region features based on the facial attention map; Step 4.4: Define and calculate key area evaluation indicators including lip shape accuracy, movement naturalness and expression coordination, and generate a facial feature analysis report; Step 4.5: Apply spatiotemporal filtering to the generated results to eliminate jitter, perform color and lighting compensation, and perform local optimization adjustments on abnormal areas based on evaluation indicators; Step 4.6: Composite the optimized lip-sync video into the original video to generate the final high-resolution lip-sync video.
7. A digital human voice lip synchronization device based on feature perception and conditional adversarial network, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the digital human voice lip synchronization method based on feature perception and conditional adversarial network is implemented according to any one of claims 1 to 6.
Citation Information
Cited By
Cross-modal learning preference optimization enhanced three-dimensional face generation method and system
CN120510260A
Artificial intelligence identification method for discovering abnormal state in audio and video playing content
CN120877191A