Video mouth shape matching correction method for cross-language dubbing
Through multimodal deep learning and facial key point detection technology, combined with StyleGAN model and synchronous regularization, the problem of mismatch between the pronunciation of the translation target language in cross-language dubbing and the lip shape of the original video characters is solved, and the timing coherence and natural transition of lip movements are achieved.
Patent Information
- Application Number
- CN202510736834.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the process of cross-language dubbing of videos, traditional methods cannot solve the problem of mismatching the pronunciation of the translation target language and the lip shape of the original video characters, while artificial intelligence methods are prone to problems such as inability to match the pronunciation and lip shape of individual words and stiff lip movements.
The phoneme-lip-type dynamic mapping technology based on multimodal deep learning is adopted, combined with facial key point detection and StyleGAN model, and the dynamic mask generation technology and 3D facial mesh predictor are used to achieve accurate mapping of cross-language audio features to target lip-type actions, and the synchronization of audio and video is forced to be constrained through synchronization regularization technology.
The timing coherence and natural transition effect of lip movements in cross-language dubbing is achieved, avoiding the breakage or mechanical feeling of the movement, and ensuring the synchronization and naturalness of the translated lip shape and the original video character's lip shape.
Smart Images

Figure CN120259139A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and specifically refers to a method for correcting video lip synchronization for cross-language dubbing. Background Art
[0002] When performing voice conversion in scenarios such as video production, film and television translation, and internationalization of multimedia content, not only the voice content needs to be accurately converted, but also the lip movements after dubbing should look synchronized and natural with the voice. It is necessary to analyze the lip movement characteristics in the original video through technical means and readjust the lip movements according to the new target language voice to make them visually match as much as possible. In the traditional manual cross-language video dubbing process, the playback speed of the original video remains unchanged, and the matching of the dubbing speed of the target language of translation with the playback speed of the original video is solved by the dubbing actor adjusting the speed, and the problem of matching the pronunciation of the target language of translation with the lip movements of the characters in the original video cannot be solved; in the artificial intelligence cross-language video dubbing process, the speed problem is solved by adjusting the video speed and the dubbing speed of the target language of translation, and a pronunciation lip model is established by deep learning the matching of the pronunciation and lip shapes of all words in each language. However, the drawback of this method is that there is no corresponding lip shape for words that have not been learned, and problems such as the inability to match the pronunciation and lip shape of individual words and rigid lip movements are likely to occur in actual use. Summary of the Invention
[0003] In view of the above situation, to overcome the defects of the prior art, the present invention provides a method for video lip synchronization correction for cross-lingual dubbing. In the traditional manual cross-lingual video dubbing process, the dubbing actor adjusts the speaking speed to match the playback speed of the original video, but it cannot solve the problem that the pronunciation of the target language of translation does not match the lip movements of the characters in the original video. This solution adopts a phoneme-lip dynamic mapping technology based on multi-modal deep learning. First, the speech of the original video is converted across languages, and then the facial key point detection technology is used to accurately locate the pixel area of the lips. The multi-modal encoder of the StyleGAN model encodes the translated audio features and the temporal information of the lip movements into a style code sequence, and separates the mesh vertices of the lip area through the dynamic mask generation technology, realizing the accurate mapping from cross-lingual audio features to target lip movements, and solving the core problem that the phoneme duration and pronunciation method do not match the original lip shape due to language differences; in the process of artificial intelligence cross-lingual video dubbing, the speech speed problem is solved by adjusting the video speed and the dubbing speed of the target language of translation, and a pronunciation-lip shape model is established for the pronunciation and lip shape matching of all words in each language. However, in actual use, it is easy to have problems such as the pronunciation and lip shape of individual words not matching and the lip movements being rigid. This solution extracts the facial dynamic parameters through a 3D facial mesh predictor, combines the moving average latent smoothing technology, and the model automatically learns the gradual change process of the opening amplitude according to the weight coefficients of adjacent frames, so that the generated lip movements can integrate the facial structure features and the action patterns of historical frames. At the same time, the synchronous regularization technology is introduced to force the synchronization of the generated video and audio, realizing the temporal coherence and natural transition effect of the lip movements, and avoiding the action breakage or mechanical feeling caused by frame-by-frame correction.
[0004] The technical solution adopted by the present invention is as follows: The present invention provides a method for video lip synchronization correction for cross-lingual dubbing, specifically including the following steps:
[0005] Step S1: Speech and lip data collection, collecting language data in multiple languages and the corresponding lip movement data, marking the lip features corresponding to each phoneme, and jointly establishing the mapping relationship between phonemes and lips with the reference style code of the lips;
[0006] Step S2: Speech recognition and translation, using the existing sequence-to-sequence model based on deep neural network to recognize and translate the video speech;
[0007] Step S3: Facial key point detection, cutting the original video into video frames as the target video, and recognizing the facial features of the characters in the video frames based on the deep learning model to output the segmented pixels of the facial features;
[0008] Step S4: Lip shape matching and correction. Identify the lip region based on the pixel segmentation of facial features. Combine the mapping relationship between phonemes and lip shapes, and use the StyleGAN model to correct the lip movements frame by frame to generate a realistic video.
[0009] Further, Step S3: Facial key point detection, specifically including the following steps:
[0010] Step S31: Collect face images, use a mask to label the facial feature categories, and form a training image set.
[0011] Step S32: Construct and initialize a convolutional neural network as a facial feature segmentation model. The convolutional neural network uses MobileNetV2 as the backbone, inserts an auxiliary head to enhance the gradient backpropagation during training, and integrates an ASPP (Atrous Spatial Pyramid Pooling) module as the decoding head to capture semantic information at various scales.
[0012] Step S33: Input the face images in the training image set into the facial feature segmentation model for training. Define the loss functions of the backbone and the auxiliary backbone as cross-entropy loss and dice loss respectively. The used formulas are as follows: ; ;
[0013] In the formula, is the cross-entropy loss, is the dice loss, is the number of samples in the training image set, is the traversal of , is the total number of categories, is the traversal of , is the true facial feature category in the training image set, is the facial feature category predicted by the facial feature segmentation model, is the sum of the exponential functions of the predicted values of each facial feature category, is the number of pixels in each facial feature category, is the traversal of , is the pixel value of the predicted facial feature category, is the pixel value of the true facial feature category in the training image set;
[0014] Step S34: The loss function of the convolutional neural network is the weighted sum of the cross-entropy loss and the dice loss. After training is completed, output the facial feature segmentation pixels.
[0015] Further, Step S4: Lip shape matching and correction, specifically including the following steps:
[0016] Step S41: Generate a dynamic mask. Input the segmented facial feature pixels into a 3D facial mesh predictor, and use the 3D facial mesh predictor to extract facial mesh vertices frame by frame to obtain facial 3D parameters. The facial 3D parameters include expression parameters, translation parameters, and rotation parameters. Based on the facial 3D parameters, adjust the facial mesh vertices to simulate opening and closing mouth actions, only retain the mesh vertices corresponding to the lip region, and project them onto the 2D image plane to generate a dynamic mask video frame aligned with the head pose;
[0017] Step S42: Construct and initialize a pre-trained StyleGAN model, use a style-based generator as an image decoder, and construct a multi-modal encoder. The multi-modal encoder includes a facial encoder, a reference encoder, and an audio encoder. The facial encoder processes the mask video frame, retains the facial structure, and obtains facial encoding features. The reference encoder maps a single video frame to a two-dimensional style code to encode the identity features of the target person in the video frame. The audio encoder maps the translated audio segment to a two-dimensional style code sequence to encode the temporal information of the lip movement;
[0018] Step S43: Obtain the fused features. Introduce skip connections in the image decoder. The image decoder is composed of decoder blocks. Each decoder block modulates the convolutional weights with the style code, predicts a 1-channel spatial mask, and fuses the facial encoding features to obtain the fused features. The formula used is as follows: ;
[0019] In the formula, is the fused feature, is the 1-channel spatial mask predicted by the -th layer decoder block at time , is the element-wise multiplication, is the dynamic mask video frame input at time , is the 2D spatial feature extracted by the facial encoder at the -th layer decoder block, is the feature map generated by the decoder at the -th layer decoder block;
[0020] Step S44: Moving average latent smoothing. Apply weighted moving average and one-dimensional convolution operations to the two-dimensional style code sequence generated by the audio encoder to ensure smooth lip movement. The formula used is as follows: ;
[0021] In the formula, is the -th layer at time The smoothed latent code, is the component of the reference style code of the mouth shape at the layer, is a one-dimensional convolutional operation for learning the local action patterns between adjacent video frames, is the window average moving weight, is to perform a weighted sum of the latent codes of adjacent video frames ( );
[0022] Step S45: Synchronization regularization. For the video frames of the target person, freeze the encoder weights, fine-tune the decoder parameters, introduce a synchronization regularization term in the fine-tuning loss function, use the mapping relationship between phonemes and mouth shapes to drive the generation of videos, calculate the synchronization loss of the video and audio, and impose a mandatory constraint on the synchronization of the video and audio. The formula used is as follows: ;
[0023] In the formula, is the synchronization loss function, is the visual feature of the video frame in the generated video, is the audio feature of the audio. The visual feature and the audio feature are extracted by a convolutional neural network;
[0024] Step S46: Define the loss function. Define the total loss function of the StyleGAN model as the weighted sum of the perceptual loss and the synchronization loss. The perceptual loss is the perceptual difference between the generated frame and the real frame. The formula used is as follows: ;
[0025] In the formula, is the total loss function of the StyleGAN model, is the perceptual loss function, and are the weights of the perceptual loss function and the synchronization loss function respectively, is the number of feature layers, is the traversal of , is the feature extractor, is the real video.
[0026] The beneficial effects achieved by the present invention using the above scheme are as follows:
[0027] (1) In the traditional manual cross - language dubbing process of videos, relying on voice actors to adjust the speaking speed to match the playback speed of the original video cannot solve the problem that the pronunciation of the target language in translation does not match the lip movements of the characters in the original video. This solution adopts a phoneme - lip dynamic mapping technology based on multi - modal deep learning. First, it performs cross - language conversion on the speech of the original video, then uses facial key - point detection technology to accurately locate the pixel area of the lips. Through the multi - modal encoder of the StyleGAN model, the translated audio features and the temporal information of lip movements are encoded into a sequence of style codes, and through dynamic mask generation technology, the grid vertices of the lip area are separated, realizing an accurate mapping from cross - language audio features to target lip movements, and solving the core problem of the mismatch between phoneme duration, pronunciation method and the original lip shape caused by language differences.
[0028] (2) In the process of artificial intelligence cross - language video dubbing, the speed problem is solved by adjusting the video speed and the speaking speed of the target - language dubbing. An articulation - lip - shape model is established for the pronunciation and lip - shape matching of all words in each language. However, in actual use, problems such as the inability to match the pronunciation and lip - shape of individual words and rigid lip movements are likely to occur. This solution extracts facial dynamic parameters through a 3D facial mesh predictor, combines moving - average latent smoothing technology. The model automatically learns the gradual change process of the opening amplitude of the mouth according to the weight coefficients of adjacent frames, enabling the generated lip movements to integrate facial structure features and historical frame action patterns. At the same time, a synchronization regularization technology is introduced to forcibly constrain the synchronization of the generated video and audio, achieving the temporal coherence and natural transition effect of lip movements, and avoiding action breaks or mechanical sensations caused by frame - by - frame correction. Brief Description of the Drawings
[0029] Figure 1 It is a schematic flow chart of a method for video lip - shape matching and correction for cross - language dubbing provided by the present invention;
[0030] Figure 2 It is a schematic flow chart of lip - shape matching and correction provided by the present invention.
[0031] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. Detailed Embodiments
[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0033] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.
[0034] Example 1. Refer to Figure 1 , a method for video mouth shape matching and correction for cross-lingual dubbing provided by the present invention, the steps of the method for video mouth shape matching and correction for cross-lingual dubbing include:
[0035] Step S1: Speech and mouth shape data collection. Collect language data in multiple languages and corresponding mouth shape change data, mark the mouth shape features corresponding to each phoneme, and jointly establish the mapping relationship between phonemes and mouth shapes with the reference style code of the mouth shape.
[0036] Step S2: Speech recognition and translation. Use the existing sequence-to-sequence model based on a deep neural network to recognize and translate the video speech.
[0037] Step S3: Facial key point detection. Cut the original video into video frames as the target video, recognize the facial features of the person in the video frames based on a deep learning model, and output the facial feature segmentation pixels.
[0038] Step S4: Mouth shape matching and correction. Identify the lip area according to the facial feature segmentation pixels, combine the mapping relationship between phonemes and mouth shapes, and use the StyleGAN model to correct the lip movements frame by frame to generate a real video.
[0039] Example 2. Refer to Figure 1 , based on the above example, Step S3: Facial key point detection specifically includes the following steps:
[0040] Step S31: Collect face images, and use a mask to mark the facial feature categories to form a training image set.
[0041] Step S32: Construct and initialize a convolutional neural network as the facial feature segmentation model. The convolutional neural network uses MobileNetV2 as the backbone, inserts an auxiliary head for enhancing the gradient backpropagation during training, and simultaneously integrates an ASPP (Atrous Spatial Pyramid Pooling) module as the decoder head for capturing semantic information at various scales.
[0042] Step S33: Input the face images in the training image set into the facial feature segmentation model for training. Define the loss functions of the backbone and the auxiliary backbone as the cross-entropy loss and the dice loss respectively. The used formulas are as follows: ; ;
[0043] In the formula, is the cross-entropy loss, is the dice loss, is the number of samples in the training image set, is the traversal of ; is the total number of categories, is the traversal of ; is the true facial feature category in the training image set, is the facial feature category predicted by the facial feature segmentation model, is the sum of the exponential functions of the predicted values of each facial feature category, is the number of pixels in each facial feature category, is the traversal of ; is the pixel value of the predicted facial feature category, is the pixel value of the true facial feature category in the training image set;
[0044] Step S34: The loss function of the convolutional neural network is the weighted sum of the cross-entropy loss and the dice loss. After training, output the facial feature segmentation pixels.
[0045] Example 3. Refer to Figure 1 and Figure 2 , this example is based on the above example. Step S4: Mouth shape matching and correction, which specifically includes the following steps:
[0046] Step S41: Generate a dynamic mask. Input the facial feature segmentation pixels into the 3D facial mesh predictor. Use the 3D facial mesh predictor to extract the facial mesh vertices frame by frame, obtain the facial 3D parameters. The facial 3D parameters include expression parameters, translation parameters, and rotation parameters. Adjust the facial mesh vertices based on the facial 3D parameters to simulate the opening and closing of the mouth, and only retain the mesh vertices corresponding to the lip area, and project them onto the 2D image plane to generate a dynamic mask video frame aligned with the head pose;
[0047] Step S42: Construct and initialize a pre-trained StyleGAN model. Use a style-based generator as an image decoder, and construct a multi-modal encoder. The multi-modal encoder includes a face encoder, a reference encoder, and an audio encoder. The face encoder processes the masked video frames, preserves the facial structure, and obtains facial encoding features. The reference encoder maps a single video frame to a two-dimensional style code, encoding the identity features of the target person in the video frame. The audio encoder maps the translated audio segments to a sequence of two-dimensional style codes, encoding the temporal information of the lip movements;
[0048] Step S43: Obtain the fused features. Introduce skip connections in the image decoder. The image decoder is composed of decoder blocks. Each decoder block modulates the convolutional weights with the style code, predicts a 1-channel spatial mask, and fuses the facial encoding features to obtain the fused features. The formula used is as follows: ;
[0049] In the formula, is the fused feature, is the 1-channel spatial mask predicted by the -th decoder block at time , is the element-wise multiplication, is the dynamic masked video frame input at time , is the 2D spatial feature extracted by the face encoder at the -th decoder block, is the feature map generated by the decoder at the -th decoder block;
[0050] Step S44: Moving average latent smoothing. Apply weighted moving average and one-dimensional convolution operations to the sequence of two-dimensional style codes generated by the audio encoder to ensure smooth lip movements. The formula used is as follows: ;
[0051] In the formula, is the smoothed latent code at the -th layer at time , is the component of the reference style code of the mouth shape at the -th layer, is the one-dimensional convolution operation used to learn the local action patterns between adjacent video frames, is the window average moving weight, is the weighted sum of the latent codes of adjacent video frames ( );
[0052] Step S45: Synchronization regularization. For the video frames of the target person, freeze the encoder weights, fine-tune the decoder parameters, introduce a synchronization regularization term in the fine-tuning loss function, use the mapping relationship between phonemes and lip shapes to drive the generation of videos, calculate the synchronization loss between the video and audio, and impose a mandatory constraint on the synchronization of the video and audio. The formula used is as follows: ;
[0053] In the formula, is the synchronization loss function, is the visual feature of the video frame in the generated video, is the audio feature, and the visual feature and audio feature are extracted by a convolutional neural network;
[0054] Step S46: Define the loss function. Define the total loss function of the StyleGAN model as the weighted sum of the perceptual loss and the synchronization loss. The perceptual loss is the perceptual difference between the generated frame and the real frame. The formula used is as follows: ;
[0055] In the formula, is the total loss function of the StyleGAN model, is the perceptual loss function, and are the weights of the perceptual loss function and the synchronization loss function respectively, is the number of feature layers, is the traversal of , is the feature extractor, is the real video.
[0056] Example 4. This example is based on the above example. In this solution, the learning rate and decay weight of the five sense organs segmentation model described in Example 2 are set to 0.001 and 0.0001, and the loss function is the weighted sum of the cross-entropy loss and the dice loss, and their weight coefficients are 1 and 3 respectively;
[0057] In Example 3, when calculating the total loss function of the StyleGAN model, is the number of feature layers, which is set to 3 in this solution.
[0058] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.
[0059] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
[0060] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.
Claims
1. A method for correcting video lip synchronization for cross-language dubbing, characterized in that, It includes the following steps: Step S1: Speech and lip movement data collection. Collect language data in multiple languages and corresponding lip movement change data, mark the lip movement characteristics corresponding to each phoneme, and jointly establish the mapping relationship between phonemes and lip movements with the reference style code of lip movements; Step S2: Speech recognition and translation. Use the existing sequence-to-sequence model based on deep neural network to recognize and translate the video speech; Step S3: Facial key point detection. Cut the original video into video frames as the target video, recognize the facial features of the people in the video frames based on the deep learning model, and output the facial feature segmentation pixels; Step S4: Lip movement matching and correction. Identify the lip area according to the facial feature segmentation pixels, combine the mapping relationship between phonemes and lip movements, and use the StyleGAN model to correct the lip movements frame by frame to generate a real video.
2. A method for correcting video lip synchronization for cross - language dubbing according to claim 1, characterized in that, Step S3: Facial key point detection, specifically including the following steps: Step S31: Collect face images, use a mask to mark the facial feature categories and then form a training image set; Step S32: Build and initialize a convolutional neural network as the facial feature segmentation model. The convolutional neural network uses MobileNetV2 as the backbone, inserts an auxiliary head to enhance the gradient backpropagation during training, and integrates the ASPP (Atrous Spatial Pyramid Pooling) module as the decoding head; Step S33: Input the face images in the training image set into the facial feature segmentation model for training, and define the loss functions of the backbone and the auxiliary backbone as cross-entropy loss and dice loss respectively; Step S34: The loss function of the convolutional neural network is the weighted sum of cross-entropy loss and dice loss. After training is completed, output the facial feature segmentation pixels.
3. A method for correcting video lip synchronization for cross-language dubbing according to claim 1, characterized in that Step S4: Lip movement matching and correction, specifically including the following steps: Step S41: Generate a dynamic mask. Input the facial feature segmentation pixels into a 3D facial mesh predictor, use the 3D facial mesh predictor to extract the facial mesh vertices frame by frame, obtain the facial 3D parameters, adjust the facial mesh vertices based on the facial 3D parameters to simulate the opening and closing of the mouth, only retain the mesh vertices corresponding to the lip area, and project them onto the 2D image plane to generate a dynamic mask video frame aligned with the head pose; Step S42: Build and initialize a pre-trained StyleGAN model, use the style-based generator as the image decoder, build a multi-modal encoder. The multi-modal encoder includes a facial encoder, a reference encoder, and an audio encoder. The facial encoder processes the masked video frame, retains the facial structure, and obtains the facial encoding features. The reference encoder maps a single video frame into a two-dimensional style code to encode the identity features of the target person in the video frame. The audio encoder maps the translated audio segment into a two-dimensional style code sequence to encode the temporal information of the lip movements; Step S43: Obtain the fused features. Introduce skip connections in the image decoder. The image decoder is composed of decoder blocks. Each decoder block modulates the convolutional weights with the style code, predicts a 1-channel spatial mask, and fuses the facial encoding features to obtain the fused features; Step S44: Moving average potential smoothing. Apply weighted moving average and one-dimensional convolution operations to the two-dimensional style code sequence generated by the audio encoder to ensure smooth lip movements. Step S45: Synchronization regularization. For the video frames of the target person, freeze the encoder weights, fine-tune the decoder parameters, introduce a synchronization regularization term into the fine-tuning loss function, use the mapping relationship between phonemes and lip shapes to drive the generation of videos, calculate the synchronization loss between audio and video, and impose a mandatory constraint on the synchronization of video and audio. Step S46: Define the loss function. Define the total loss function of the StyleGAN model as the weighted sum of the perceptual loss and the synchronization loss, where the perceptual loss is the perceptual difference between the generated frame and the real frame.
Citation Information
Patent Citations
Voice face driving method based on vision element correction
CN117079663A
Lightweight personalized face visual dubbing method
CN118250411A
Mouth shape generation method and device based on deep learning and storage medium
CN118782082A
Cited By
Real person voice mouth shape animation generation method and system, electronic equipment and storage medium
CN121392080A
Real-person voice lip animation generation method and system, electronic device, and storage medium
CN121392080B
An audio-visual lip-synching video generation method, system and device based on a discrete code prediction model
CN122511287A