High-resolution video redubbing generation method of mask recovery appearance

By introducing the MSR module to generate structural feature maps and the MDA module to achieve dynamic feature alignment, the problem of facial structure distortion and insufficient synchronization in facial generation of high-resolution speakers is solved, and the stability and synchronization performance of generated videos are improved.

CN120388108APending Publication Date: 2025-07-29SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510267673.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art has problems of facial structure distortion, insufficient posture adaptability and insufficient audio-mouth synchronization in the face generation of high-resolution speakers, especially in the event of extreme posture changes.

Method used

The Masked Structure Reasoning (MSR) module is used to generate structural feature maps through a multi-scale decoder, and the static structural features of the reference frame and the target frame are aligned with the gating mechanism. The Motion Detail Aligning (MDA) module is used to realize affine transformation of dynamic features through the AdaAT algorithm to ensure dynamic and audio synchronization between the mouth.

Benefits of technology

Significantly improves the stability and posture adaptability of the generation, ensuring accurate synchronization of the mouth and audio, and the generated video is better than existing methods in clarity and texture details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388108A_ABST
    Figure CN120388108A_ABST
Patent Text Reader

Abstract

The invention discloses a high-resolution video redubbing generation method of mask recovery appearance. The method comprises the following steps: (1) inputting data and preprocessing; (2) performing mask reconstruction by adopting an MAE encoder, reasoning structural features of the face, generating feature maps with different resolutions through a multi-scale decoder, aligning the features with reference frames with different input sizes in combination with a gating mechanism, and capturing static structural features of key areas such as the mouth and the eyes; (3) extracting texture features through a reference frame, and fusing the extracted texture features with the generated structural features; meanwhile, the driving audio features are combined with the fusion features through a trans-attention mechanism, and texture features after dynamic alignment are generated; adaAT algorithm is adopted to realize affine transformation of dynamic characteristics, and dynamic and audio synchronization of the mouth is ensured; (4) generating a redubbed audio and video; according to the method, the generation stability and the posture adaptability are remarkably improved, and accurate mouth and audio synchronization is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and audio - video processing, and particularly relates to a method for generating high - resolution video re - dubbed voiceovers with mask - restored appearance. Background Art

[0002] Some recent research methods have been dedicated to solving the problems of pose differences and texture details in the field of high - resolution speaker face generation, such as DINet. This method performs spatial deformation on the reference frame feature map to map the appearance features of the reference frame to the target frame, in order to achieve higher - resolution face generation and retain high - frequency texture details. This strategy solves the problems caused by pose differences between the reference frame and the target frame to a certain extent, and enhances the fineness of visual texture during the generation process. However, this method also has some potential limitations. Specifically, DINet lacks a clear constraint on the target frame structure when directly performing feature deformation, resulting in possible facial structure distortion and misalignment of key features in the generated results. In addition, simply relying on spatial deformation operations fails to effectively integrate audio - driven information, making it difficult to ensure the precise synchronization of mouth movements with the audio input. In scenarios with large pose differences, the generation quality of DINet drops significantly, and its stability and robustness still need to be improved. The problems of existing methods are: (1) Lack of structural constraints: easily leading to facial feature distortion or deformation. (2) Insufficient pose adaptability: the generated visual effects are unstable in extreme pose changes. (3) Insufficient audio - mouth synchronization: the generated mouth movements may lag or not match the speech input. Summary of the Invention

[0003] Object of the Invention: The object of the present invention is to provide a method for generating high - resolution video re - dubbed voiceovers with mask - restored appearance, to solve the pose difference problem between the reference frame and the target frame.

[0004] Technical Solution: A method for generating high - resolution video re - dubbed voiceovers with mask - restored appearance according to the present invention includes the following steps:

[0005] (1) Input data and pre - processing: source video frames and driving audio;

[0006] (2) Structure Feature Reasoning MSR module: Use a MAE encoder for mask reconstruction, reason out the structural features of the face, generate feature maps of different resolutions through a multi - scale decoder, and combine the gating mechanism to align the features with reference frames of different input sizes, capturing the static structural features of key regions such as the mouth and eyes;

[0007] (3) Dynamic Feature Alignment MDA Module: Extract texture features from reference frames and fuse them with the generated structural features; at the same time, combine the driving audio features with the fused features through a cross-attention mechanism to generate texture features after dynamic alignment; use the AdaAT algorithm to implement the affine transformation of dynamic features to ensure the synchronization of mouth dynamics and audio;

[0008] (4) Generate the re-dubbed video: Concatenate the structural features and the dynamically aligned features, and restore them to the image space through a convolutional decoder to generate the target frame.

[0009] Furthermore, in step (1), the source video frames are divided into non-overlapping image blocks of size 16×16. Each image block generates a token representation through an embedding operation, and masks the mouth region with a high ratio of 75%; the driving audio extracts semantic features through an audio encoder for subsequent dynamic adjustment.

[0010] Furthermore, in step (2), the MSR module is as follows: First, perform a partial masking operation on the source image, and only input the features of the visible region into the MAE encoder. Through a lightweight multi-scale decoder, generate structural features of different scales.

[0011] Furthermore, in step (2), 25% of the visible tokens are input into the Transformer encoder for processing, and finally generate structural features.

[0012] Furthermore, step (3) is as follows: First, extract texture features from multiple reference frames; then, concatenate the texture features of the reference frames with the static structural features output by the MSR to form preliminarily fused image features; at the same time, the audio encoder extracts the semantic features of the audio; finally, the image features and the semantic features jointly guide the deformation of the reference frame texture features to obtain the final dynamic texture features.

[0013] Introduce the adaptive affine transformation AdaAT algorithm. The audio features and the fused image features will be calculated through a cross-attention mechanism to obtain the texture deformation parameters for guiding the reference frame, and the formula is as follows:

[0014]

[0015] where, are the transformed pixel coordinates, (x c , y c ) are the original coordinates, s c is the scaling ratio, θ c is the rotation angle, is the translation amount.

[0016] Combine the texture features output by the MDA module with the structural features output by the MSR to generate the facial video.

[0017] Further, step (4) is specifically as follows: The specific process of the convolutional decoder is as follows: Connect the texture feature and the structural feature in the feature channel, and jointly use them as the feature encoding to input an upsampling convolutional network. Finally, restore the feature to the image space to obtain the final dubbed image.

[0018] A high-resolution video re-dubbing generation system for mask-restored appearance according to the present invention includes:

[0019] Input data and preprocessing module: Used to obtain the source video frame and the driving audio;

[0020] Structural feature inference MSR module: Used to perform mask reconstruction using the MAE encoder, infer the structural features of the face, generate feature maps of different resolutions through a multi-scale decoder, and align the features with the reference frames of different input sizes in combination with the gating mechanism to capture the static structural features of key areas such as the mouth and eyes;

[0021] Dynamic feature alignment MDA module: Used to fuse the texture features extracted from the reference frame with the generated structural features; at the same time, combine the driving audio features with the fused features through the cross-attention mechanism to generate the texture features after dynamic alignment; use the AdaAT algorithm to implement the affine transformation of the dynamic features to ensure that the mouth dynamics are synchronized with the audio;

[0022] Generated re-dubbed video module: Used to splice the structural features and the dynamically aligned features, and restore them to the image space through a convolutional decoder to generate the target frame.

[0023] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements a method for generating high-resolution video re-dubbing for mask-restored appearance according to any one of the above.

[0024] A storage medium according to the present invention stores a computer program, characterized in that when the computer program is executed by a processor, it implements a method for generating high-resolution video re-dubbing for mask-restored appearance according to any one of the above.

[0025] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: By introducing the MotionDetail Aligning (MDA) module, the present invention dynamically aligns the texture and structural features of the reference frame and the target frame, thereby significantly improving the stability and pose adaptability of generation. Using the Masked Structure Reasoning (MSR) module, the reasoning ability for facial structures is enhanced by partially masking the input, ensuring that the generated mouth and facial structures are consistent with the target frame. Achieving precise mouth and audio synchronization: Using an audio-driven Cross Attention network, the audio and mouth movements are precisely aligned to ensure that the generated mouth dynamics are highly synchronized with the speech input. Description of the Drawings

[0026] Figure 1 It is a network framework diagram of the present invention. Detailed Embodiments

[0027] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0028] As Figure 1 shown, an embodiment of the present invention provides a method for generating high-resolution video re-dubbing with masked recovery appearance, including the following steps:

[0029] (1) Input data and preprocessing: source video frames and driving audio; the source video frames are divided into non-overlapping image patches of size 16×16, each image patch generates a token representation through an embedding operation, and the mouth region is masked at a high ratio of 75%; the driving audio extracts semantic features through an audio encoder for subsequent dynamic adjustment.

[0030] (2) Structure Feature Reasoning MSR module: Using a MAE encoder for masked reconstruction, reasoning out the structural features of the face, generating feature maps of different resolutions through a multi-scale decoder, and combining a gating mechanism to align the features with reference frames of different input sizes, capturing static structural features of key regions such as the mouth and eyes; specifically as follows: Using a pre-trained MAE (Masked Autoencoder) model, which is responsible for performing masking and reconstruction tasks. The input image frame is divided into small patches of size 16×16, and then each patch becomes a token through an embedding operation. During training, 75% of the tokens are randomly masked (that is, "hidden"), and the network can only learn how to reason about the structure of the hidden part through the remaining 25% of the visible tokens. These visible tokens are input into the Transformer encoder for processing, and finally structural features are generated.

[0031] To generate richer feature information, we modify the decoder into a multi-scale decoder. The decoder can not only process detailed information but also generate feature maps of different resolutions. These feature maps capture the overall shape and edge details of the face from rough to fine through progressive upsampling. This process uses a gating mechanism to better match the generated structural features with input frames of different sizes. Finally, the features output by the MSR module capture the static position information of the face, providing a basis for subsequent dynamic alignment.

[0032] (3) Dynamic Feature Alignment MDA Module: Extract texture features from reference frames and fuse them with the generated structural features; at the same time, combine the driving audio features with the fused features through cross-attention mechanism to generate texture features after dynamic alignment; use the AdaAT algorithm to implement the affine transformation of dynamic features to ensure the synchronization of mouth dynamics with audio; specifically as follows: The task of the MDA module is to dynamically align the static structural features output by the MSR module and the texture features provided by the reference frames. At the same time, this module combines the information of the audio input to generate synchronized and realistic facial dynamic details.

[0033] First, extract texture features from multiple reference frames. The texture information of these reference frames is used to make up for the missing details in the source frame. Then, splice the texture features of the reference frames with the static structural features output by the MSR to form preliminarily fused image features. Next, the audio encoder extracts the semantic features of the audio, such as how the opening and closing movements of the mouth should synchronize with the audio rhythm.

[0034] To achieve dynamic feature alignment, the Adaptive Affine Transformation (AdaAT) algorithm is introduced. The audio features and the fused image features are calculated through the cross-attention mechanism to obtain the texture deformation parameters for guiding the reference frames. The formula is as follows:

[0035]

[0036] Among them, are the transformed pixel coordinates, (x c , y c ) are the original coordinates, s c is the scaling ratio, θ c is the rotation angle, is the translation amount.

[0037] In this way, the MDA module can dynamically adjust the texture features of the reference frames to the same position and shape as the target frame. Finally, the dynamic features output by the MDA module are combined with the structural features output by the MSR to generate a facial video with high-quality dynamic details.

[0038] (4) Generate the re-dubbed video: Concatenate the structural features and the dynamic alignment features, and restore them to the image space through a convolutional decoder to generate the target frame. Specifically as follows: The specific process of the convolutional decoder is as follows: Connect the texture features and the structural features in the feature channels, and jointly use them as the feature encoding to input an upsampling convolutional network. Finally, restore the features to the image space to obtain the final dubbed image.

[0039] The experimental results of the present invention on multiple public datasets verify the superiority of the technology:

[0040] Visual quality: On the HDTF and LRS2-Pose datasets, the generated videos are significantly superior to the existing methods in terms of metrics such as PSNR, SSIM, and LPIPS. Among them, the PSNR is improved by about 1.2dB, and the SSIM is increased by 4%, indicating that the videos generated by this method are superior in terms of clarity and texture detail fidelity.

[0041] Audio-visual synchronization: The performance in terms of metrics such as Lip Sync Error Distance (LSE-D) and LipLMD is comparable to that of the existing methods, proving that the synchronization performance between the mouth dynamics and the audio is well retained, with high practicality.

[0042] Pose robustness: For scenarios with a large Pose Gap, this method can better retain and restore the structural consistency of key parts, reducing distortion and blurring phenomena. In the experiment, the generation effect score under the condition of large pose changes is superior to the DINet and Wav2Lip methods, and the average performance is improved by more than 15%.

Claims

1. A method for generating high-resolution video re-dubbing that restores the apparent mask, characterized in that, It includes the following steps: (1) Input data and preprocessing: source video frames and driving audio; (2) Structure feature inference MSR module: Use the MAE encoder for mask reconstruction, infer the structural features of the face, generate feature maps of different resolutions through a multi-scale decoder, and combine the gating mechanism to align the features with reference frames of different input sizes, capturing static structural features in key areas such as the mouth and eyes; (3) Dynamic feature alignment MDA module: Extract texture features from the reference frame and fuse them with the generated structural features; at the same time, combine the driving audio features with the fused features through the cross-attention mechanism to generate texture features after dynamic alignment; Use the AdaAT algorithm to implement the affine transformation of dynamic features to ensure that the mouth dynamics are synchronized with the audio; (4) Generate the re-dubbed video: Concatenate the structural features and the dynamically aligned features, and restore them to the image space through a convolutional decoder to generate the target frame.

2. A method for generating high-resolution video re-dubbing with masked restored appearance according to claim 1, characterized in that In step (1), the source video frames are divided to generate non-overlapping image blocks of size 16×16. Each image block generates a token representation through an embedding operation, and the mouth area is masked at a high ratio of 75%; the driving audio extracts semantic features through an audio encoder for subsequent dynamic adjustment.

3. A method for generating high-resolution video re-dubbing with mask-restored appearance according to claim 1, characterized in that In step (2), the MSR module is as follows: First, perform a partial masking operation on the source image, and only input the features of the visible area into the MAE encoder. Through a lightweight multi-scale decoder, generate structural features of different scales.

4. A method for generating high-resolution video re-dubbing with mask-restored appearance, according to claim 3, characterized in that In step (2), 25% of the visible tokens are input into the Transformer encoder for processing, and finally structural features are generated.

5. A method for generating high-resolution video re-dubbing with masked restored appearance according to claim 1, wherein Step (3) is as follows: First, extract texture features from multiple reference frames; then, concatenate the texture features of the reference frame with the static structural features output by the MSR to form a preliminarily fused image feature; at the same time, the audio encoder extracts the semantic features of the audio; finally, the image feature and the semantic feature jointly guide the deformation of the reference frame texture feature to obtain the final dynamic texture feature. Introduce the adaptive affine transformation AdaAT algorithm. The audio features and the fused image features will be calculated through the cross-attention mechanism to obtain the texture deformation parameters for guiding the reference frame, and the formula is as follows: Among them, are the transformed pixel coordinates, (x c , y c ) are the original coordinates, s c is the scaling ratio, θ c is the rotation angle, is the translation amount. Combine the texture features output by the MDA module with the structural features output by the MSR to generate a facial video.

6. A method for generating high-resolution video re-dubbing with masked restored appearance according to claim 1, characterized in that, Step (4) is as follows: The specific process of the convolutional decoder is as follows: Connect the texture features and the structural features in the feature channel, and jointly use them as the feature encoding to input an upsampling convolutional network. Finally, restore the features to the image space to obtain the final dubbed image.

7. A high-resolution video re-dubbing generation system for mask-restored appearance, characterized in that, It includes: Input data and preprocessing module: used to obtain source video frames and driving audio; Structure feature inference MSR module: used to use the MAE encoder for mask reconstruction, infer the structural features of the face, generate feature maps of different resolutions through a multi-scale decoder, and combine the gating mechanism to align the features with reference frames of different input sizes, capturing static structural features in key areas such as the mouth and eyes; Dynamic Feature Alignment MDA Module: It is used to fuse the texture features extracted from the reference frame with the generated structural features; at the same time, the driving audio features are combined with the fused features through the cross-attention mechanism to generate the dynamically aligned texture features; the AdaAT algorithm is used to implement the affine transformation of the dynamic features to ensure the synchronization of the mouth dynamics and the audio. Generated Dubbed Video Module: It is used to splice the structural features and the dynamically aligned features, and restore them to the image space through a convolutional decoder to generate the target frame.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements a high-resolution video dubbed generation method for mask restoration appearance according to any one of claims 1-6.

9. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a high-resolution video dubbed generation method for mask restoration appearance according to any one of claims 1-6.