Short play character video translation method based on diffusion model

Through the short drama character video translation method based on the diffusion model, combined with input from multiple modalities, efficient lip alignment and face-changing operations are achieved, solving the problems of low efficiency and poor fluency in traditional technologies, and improving the adaptability of user experience and application scenarios.

CN119942398AInactive Publication Date: 2025-05-06BEIJING ZHONGKE JINCAI TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411849011.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional face swapping and lip alignment technologies are inefficient and poorly fluent when processing complex videos and real-time applications, making it difficult to achieve a high-quality user experience.

Method used

Using a short drama character video translation method based on diffusion model, the details of faces are captured through the GPEN model, CRNet enhances image brightness and contrast, decouples the network to separate identity and attribute features, and processes audio signals through AudioNet to achieve end-to-end training of multimodal fusion.

Benefits of technology

It realizes efficient completion of lip alignment and face change operations in a single process. The generated face images are real and natural, and the lip synchronization effect is accurate, improving the adaptability of user experience and application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942398A_ABST
    Figure CN119942398A_ABST
Patent Text Reader

Abstract

The invention discloses a diffusion model-based short episode video translation method, which comprises the following steps of: firstly, segmenting pictures in a video frame according to a fixed size, forming a batch with an original image, and sending the batch into a face detection model for detection; capturing detail features of the source face through a GPEN model, and enhancing the detail features of the source face; enhancing the brightness and contrast of the target image through CRNet; the detail features of the source face are effectively separated through a decoupling network; inputting the audio signal into an Audio Net network, and converting the audio signal into feature representation after noise reduction; face changing and mouth shape alignment tasks are combined through a multi-modal fusion mechanism, and end-to-end training is carried out. The invention provides a set of complete processing flow, which covers face detection to image enhancement, identity information extraction, audio feature processing and final face change and mouth shape alignment model training, and ensures that a natural and smooth video translation effect is generated under multi-modal input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a method for translating short play character videos based on a diffusion model. Background Art

[0002] With the rapid advancement of computer vision and generative AI technologies, digital humans, face-changing and lip-syncing technologies have been widely used in many fields. These technologies not only play an important role in the entertainment industry, such as film special effects production, virtual character creation and game development, but also show great potential in social media, virtual reality (VR), augmented reality (AR), online education and other fields. For example, in online education, the creation of virtual teachers and lip-syncing technology can provide a more realistic teaching experience, while on social media, face-changing technology is widely used in user-generated content (UGC) to further enhance interactivity and personalized expression.

[0003] However, traditional face-swapping and lip-alignment technologies face many challenges. These technologies usually require multiple complex processing steps, such as facial recognition, feature extraction, facial reconstruction, texture mapping, and lip-matching. Each step involves different algorithms and model processing, which makes the whole process cumbersome and time-consuming, and easily limited by the performance bottleneck of the algorithm. Especially in real-time application scenarios (such as social media video filters or real-time virtual meetings), existing technologies are difficult to achieve efficient and smooth processing effects, and the user experience is greatly affected. Summary of the invention

[0004] The purpose of the present invention is to provide a method for translating short play character videos based on a diffusion model, thereby solving the above-mentioned problems existing in the prior art.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A method for translating short play character videos based on a diffusion model comprises the following steps:

[0007] S100, firstly, the images in the video frame are segmented into a fixed size, and are grouped together with the original images into a batch and sent to the face detection model for detection, and the detected face model is used as the source face;

[0008] S200, capturing detailed features of the source face through the GPEN model, and enhancing the detailed features of the source face;

[0009] S300, enhance the brightness and contrast of the target image through CRNet;

[0010] S400, effectively separating the detail features of the source face through a decoupling network;

[0011] S500, inputting the audio signal into the AudioNet network to convert it into a feature representation after noise reduction, and extracting audio features related to the speaker through the AudioNet network;

[0012] S600, through a multimodal fusion mechanism, the face-changing and lip-alignment tasks are combined for end-to-end training, ensuring the accuracy and consistency of face-changing and lip-alignment in the final generated video.

[0013] Preferably, the specific method of step S100 is:

[0014] S110, Image Enhancement: The input image is enhanced by CRNet to improve image quality and detail features;

[0015] S120, batch processing: combining multiple enhanced images into a batch for subsequent processing;

[0016] S130, face detection: use the face detection model to perform face detection on the images in the batch, identify and crop the face images.

[0017] Preferably, the specific method of step S200 is:

[0018] Face enhancement: The cropped face image is further enhanced by the GPEN model to improve the quality and detail features of the face image.

[0019] Preferably, the detail features include: identity features of the source face and attribute features of the target face.

[0020] Preferably, the face detection model is: YOLOv8 model.

[0021] Preferably, AudioNet includes three identical Transform structures, each Transform structure including:

[0022] Norm Linear module,The Norm Linear module first normalizes the input features to improve the stability and convergence speed of the model. Secondly, it maps the features to a new space through linear transformation to enhance the expressiveness of the features;

[0023] Attention module: used to capture important information in input features and improve the model's attention to key information by weighting the influence of different features;

[0024] Residual Connection module: It is used to add the input features with the features after Transform processing to help the model better learn the relationship between features, avoid the gradient vanishing problem, and promote the flow of information.

[0025] According to another aspect of an embodiment of the present invention, a server is further provided, including:

[0026] AttEnc and IdeEnc;

[0027] AttEnc is used to extract attribute information from audio signals;

[0028] IdeEnc is used to extract the identity features of the target face in the audio signal;

[0029] VAE decoder, used to convert latent features into visual face images;

[0030] AudioNet, which is responsible for converting audio signals into audio embeddings;

[0031] AdaLN layer, used to enhance the robustness and generation capability of the model.

[0032] Preferably, AttEnc and IdeEnc are positionally added to obtain a comprehensive Latent.

[0033] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, wherein the program executes any one of the methods.

[0034] According to another aspect of an embodiment of the present invention, a system for generating emotional images is also provided, comprising: a processor, a video acquisition device, and an audio acquisition device, wherein the processor communicates with the video acquisition device and the audio acquisition device respectively, and the processor is used to execute any one of the methods of claims 1 to 6.

[0035] The beneficial effects of the present invention are:

[0036] The present invention discloses a method for translating short play characters' videos based on a diffusion model. By combining inputs of multiple modalities, a unified face-changing and lip-alignment framework is successfully constructed, and an end-to-end processing method is realized, which has the following significant effects:

[0037] The operation is completed in one step: The present invention can realize lip alignment and face-changing operations simultaneously in a single process, which greatly simplifies the complexity of the traditional method which requires multiple steps and improves the processing efficiency and convenience.

[0038] High-quality generation effect: Through deep learning of audio and facial images, the facial images generated by the present invention are visually realistic and natural, and the lip synchronization effect is accurate, which can effectively show the expressions and lip shape changes consistent with the audio signals.

[0039] Application of multimodal input: The present invention is powerful and can handle multiple input modalities (such as audio and images). It shows good adaptability and is suitable for different application scenarios, such as film and television production, social media, virtual reality, etc.

[0040] Improve user experience: Through efficient and accurate face-changing and lip-alignment operations, users will have a smoother and more natural visual experience, which enhances the realism of the interaction and increases the fun and sense of participation in the application scenario.

[0041] Universality of technology: This framework is not only applicable to specific face swapping cases, but also has scalability and universality. It can be integrated into a variety of human-computer interaction and entertainment applications, opening up new opportunities for innovative development in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flow chart of a method for translating short play characters' videos based on a diffusion model of the present invention;

[0043] Figure 2 It is a schematic diagram of the decoupling network structure of the present invention;

[0044] Figure 3 It is a schematic diagram of the AudioNet network structure of the present invention;

[0045] Figure 4 This is a network architecture diagram for realizing face swapping and accurate lip alignment in the present invention;

[0046] Figure 5 is a flow chart of face detection processing of the present invention;

[0047] Figure 6 is a flow chart of another embodiment of the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation methods described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] Reference Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6 The method for translating short play characters' videos based on a diffusion model includes the following steps:

[0050] S100, firstly, the images in the video frame are segmented into a fixed size, and are grouped together with the original images into a batch and sent to the face detection model for detection, and the detected face model is used as the source face.

[0051] Face detection and image preprocessing:

[0052] When performing face detection, the images in the video frame are first divided into fixed sizes and sent to the face detection model as a batch together with the original image. This batch processing method can improve detection efficiency and accuracy, especially in the process of video stream processing, ensuring that faces of different sizes can be accurately detected.

[0053] S200: Capture detailed features of the source face through the GPEN model, and enhance the detailed features of the source face.

[0054] Face enhancement module:

[0055] The present invention uses the GPEN (Generative Prior Enhanced Network) model to enhance the source face. The model can capture high-precision details of the source face, especially identity features, and enhance the similarity of the face-swapped image. By introducing generative prior knowledge, GPEN significantly improves the facial details and natural expression of the face-swapped image, which helps to maintain the similarity between the source face and the target face.

[0056] S300, enhance the brightness and contrast of the target image through CRNet.

[0057] Target image enhancement:

[0058] For scenes with low light, this paper uses CRNet (Contrastive Learning-based Image Enhancement Network) to enhance the target image. By improving the brightness and contrast of the image, the visibility of facial features in low light environments is enhanced, thereby improving the face recall rate under complex lighting conditions, ensuring that the face swap operation can be carried out smoothly under different lighting conditions.

[0059] S400, effectively separating the detail features of the source face through a decoupling network;

[0060] Identity and attribute decoupled network:

[0061] In order to obtain the identity features of the source face and the attribute information (such as expression, posture, etc.) of the target face respectively during face swapping, the present invention designs a decoupling network. Through this network, the identity features of the source face and the attribute features of the target face can be effectively separated, so that the result of face swapping retains the identity features of the source face and maintains the consistency of the expression and posture of the target face.

[0062] S500: Input the audio signal into the AudioNet network to convert it into a feature representation after noise reduction, and extract audio features related to the speaker through the AudioNet network.

[0063] Audio feature processing (AudioNet):

[0064] For audio signals, the present invention designs an AudioNet network to convert the input audio signal into a feature representation after noise reduction, which is convenient for subsequent fusion with visual features. AudioNet can extract audio features related to the speaker, ensure that the voice information in the audio is accurately reflected when generating lip shapes, and achieve lip shape alignment.

[0065] S600, through a multimodal fusion mechanism, the face-changing and lip-alignment tasks are combined for end-to-end training, ensuring the accuracy and consistency of face-changing and lip-alignment in the final generated video.

[0066] Face swapping and lip alignment model training:

[0067] During the training process of the entire model, the present invention combines the face swapping and lip alignment tasks for end-to-end training. Through the multimodal fusion mechanism (audio and visual features), the model can learn how to naturally perform face swapping and lip synchronization during the training phase, ensuring the accuracy and consistency of face swapping and lip alignment in the final generated video.

[0068] Speaker localization and audio segmentation in the inference phase:

[0069] During the inference process, if multiple faces appear in the video, the present invention first locates the speaker through facial key point detection and identifies the start and end time of his speech. Combined with the video shot segmentation, the continuity of the speaker across shots is ensured. In addition, the input audio is segmented according to the time interval obtained by the facial key points to ensure that the audio signals in different time periods can be accurately aligned with the corresponding facial lip shapes.

[0070] Through the above steps, the present invention realizes efficient video translation based on the diffusion model, which can not only accurately complete the face-changing task of multiple faces, but also achieve accurate synchronization of the speaker's voice and lip shape when there is audio input. This method greatly simplifies the processing flow while ensuring the generation effect, and provides strong technical support for related applications (such as film and television production, virtual meetings, social media content creation, etc.).

[0071] In this embodiment, training a model can decouple attribute and identity information. The network structure is as follows Figure 2 As shown:

[0072] In the present invention, AttEnc and AttDec represent attribute codec and identity codec, respectively, while IdeEnc and IdeDec represent identification codec, respectively. We design a codec framework based on the variational autoencoder (VAE) structure to achieve efficient modeling of input data.

[0073] By introducing the adversarial training mechanism, we set up an adversarial loss function. The core idea of ​​this process is to continuously adjust the parameters of the encoder and decoder so that the model can achieve a decoupled state when processing attribute and identity information. Such decoupling can not only improve the quality of generated data, but also enhance the adaptability and flexibility of the model between different tasks. Specifically, our loss function design is inspired by the FaceShifter article, which fully considers how to effectively separate attribute information from identity information in adversarial training.

[0074] In the actual inference phase, the model workflow is further simplified. At this time, only the encoder needs to be used for feature extraction without calling the decoder. This is because the main function of the decoder is to assist in generating high-quality reconstructed samples during the training phase. It learns to generate outputs that match the input, thereby promoting the optimization of the encoder parameters. The training process of the decoder ensures that the model can accurately reproduce samples belonging to a specific identity but with different attributes during the generation phase, thereby achieving accurate attribute editing and identity preservation.

[0075] In another embodiment, an AudioNet is designed to convert audio signals into features in the noise reduction network input space, which consists of a pre-trained Whisper Encoder and three Transform structures. The specific network structure is as follows Figure 3 As shown:

[0076] AudioNet is an innovative neural network architecture designed to convert audio signals into features in the denoising network input space. The network mainly consists of a pre-trained Whisper Encoder and three Transform structures, which can effectively extract important features from audio signals and provide high-quality input for subsequent denoising processing.

[0077] Detailed description of network structure

[0078] 1. Whisper Encoder:

[0079] Whisper Encoder is the core part of AudioNet, responsible for converting raw audio signals into high-dimensional feature representations. The encoder is pre-trained to capture speech features and background noise information in audio signals, providing rich contextual information for subsequent processing.

[0080] 2. Transform structure:

[0081] AudioNet contains three identical Transform structures, each of which consists of the following parts:

[0082] Norm+Linear:

[0083] The module first normalizes the input features to improve the stability and convergence speed of the model. Then, it maps the features to a new space through linear transformation to enhance the expressiveness of the features.

[0084] Attention:

[0085] The Attention mechanism is used to capture important information in the input features and improve the model's attention to key information by weighting the influence of different features. This process helps the model focus on the most relevant parts when processing complex audio signals.

[0086] Residual Connection(⊕):

[0087] The residual connection adds the input features to the features after Transform processing, helping the model to better learn the relationship between features, avoid the gradient vanishing problem, and promote the flow of information.

[0088] 3. Repeating Structure:

[0089] Due to the repeated use of the Transform structure (X3), AudioNet is able to extract and process audio features at multiple levels, enhancing the expressiveness and robustness of the model.

[0090] In another embodiment, the present invention designs an innovative network architecture to achieve accurate alignment of lip movements while face swapping. The network combines multiple pre-trained models and custom structures to ensure that in the audio-driven face swapping process, the generated face image can naturally match the audio content. Figure 4 As shown:

[0091] 1. Pre-trained model:

[0092] AttEnc (attribute encoder) and IdeEnc (identity encoder) are the pre-trained models in step 1, which are used to extract attribute information in the audio signal and identity features of the target face respectively.

[0093] The VAE decoder is an open source pre-trained model responsible for converting the final latent features into a visual face image.

[0094] 2. Training phase:

[0095] During the training phase, except for the parameters of AudioNet, all other parameters need to be updated and trained to ensure that the model can adapt to the specific face-changing task.

[0096] 3. Feature extraction and fusion:

[0097] Attribute encoding and identity encoding:

[0098] AttEnc and IdeEnc extract the feature codes of audio and face respectively. By adding the two codes, a comprehensive latent feature (Latent) is obtained.

[0099] Identify Unet:

[0100] After being processed by Identify Unet, the Latent features are further refined and features at different levels are extracted to enhance the details and accuracy of the generated images.

[0101] 4. Audio embedding and noise processing:

[0102] AudioNet:

[0103] AudioNet is responsible for converting audio signals into audio embeddings, providing a basis for subsequent feature fusion.

[0104] AdaLN layer:

[0105] After the audio encoding and identity encoding pass through the AdaLN layer, they are added with noise to enhance the robustness and generation ability of the model.

[0106] 5. Feature merging and generation:

[0107] Denoising Unet and Identify Unet:

[0108] The different levels of features of Denoising Unet and Identify Unet are merged to generate the final Latent features.

[0109] VAE decoding:

[0110] Finally, the Latent features are converted into face images with face swapping and lip alignment through the VAE decoder, ensuring that the generated images are natural and consistent with the audio content.

[0111] The face detection process is as follows:

[0112] Preferably, the specific method of step S100 is:

[0113] S110, Image Enhancement: The input image is enhanced by CRNet to improve image quality and detail features;

[0114] S120, batch processing: combining multiple enhanced images into a batch for subsequent processing;

[0115] S130, face detection: use the face detection model to perform face detection on the images in the batch, identify and crop the face images.

[0116] Preferably, the specific method of step S200 is:

[0117] S140, face enhancement: the cropped face image is further enhanced by the GPEN model to improve the quality and detail features of the face image.

[0118] The process is processed through multiple steps to ensure that a high-quality face image is finally obtained, which is suitable for subsequent applications and analysis. Wherein Face Detect is any model that can realize face detection, and the yolo8 model is used in the present invention.

[0119] Preferably, the detail features include: identity features of the source face and attribute features of the target face.

[0120] Preferably, the face detection model is: YOLOv8 model.

[0121] Preferably, AudioNet includes three identical Transform structures, each Transform structure including:

[0122] Norm Linear module,The Norm Linear module first normalizes the input features to improve the stability and convergence speed of the model. Secondly, it maps the features to a new space through linear transformation to enhance the expressiveness of the features;

[0123] Attention module: used to capture important information in input features and improve the model's attention to key information by weighting the influence of different features;

[0124] Residual Connection module: It is used to add the input features with the features after Transform processing to help the model better learn the relationship between features, avoid the gradient vanishing problem, and promote the flow of information.

[0125] According to another aspect of an embodiment of the present invention, a server is further provided, including:

[0126] AttEnc and IdeEnc;

[0127] AttEnc is used to extract attribute information from audio signals;

[0128] IdeEnc is used to extract the identity features of the target face in the audio signal;

[0129] VAE decoder, used to convert latent features into visual face images;

[0130] AudioNet, which is responsible for converting audio signals into audio embeddings;

[0131] AdaLN layer, used to enhance the robustness and generation capability of the model.

[0132] Preferably, AttEnc and IdeEnc are positionally added to obtain a comprehensive Latent.

[0133] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, wherein the program executes any one of the methods.

[0134] According to another aspect of an embodiment of the present invention, a system for generating emotional images is also provided, comprising: a processor, a video acquisition device, and an audio acquisition device, wherein the processor communicates with the video acquisition device and the audio acquisition device respectively, and the processor is used to execute any one of the methods of claims 1 to 6.

[0135] The beneficial effects of the present invention are:

[0136] By combining inputs from multiple modalities, the present invention successfully constructs a unified face-swapping and lip-alignment framework, realizing an end-to-end processing approach with the following significant effects:

[0137] The operation is completed in one step: The present invention can realize lip alignment and face-changing operations simultaneously in a single process, which greatly simplifies the complexity of the traditional method which requires multiple steps and improves the processing efficiency and convenience.

[0138] High-quality generation effect: Through deep learning of audio and facial images, the facial images generated by the present invention are visually realistic and natural, and the lip synchronization effect is accurate, which can effectively show the expressions and lip shape changes consistent with the audio signals.

[0139] Application of multimodal input: The present invention is powerful and can handle multiple input modalities (such as audio and images). It shows good adaptability and is suitable for different application scenarios, such as film and television production, social media, virtual reality, etc.

[0140] Improve user experience: Through efficient and accurate face-changing and lip-alignment operations, users will have a smoother and more natural visual experience, which enhances the realism of the interaction and increases the fun and participation of the application scenarios.

[0141] Universality of technology: This framework is not only applicable to specific face swapping cases, but also has scalability and universality. It can be integrated into a variety of human-computer interaction and entertainment applications, opening up new opportunities for innovative development in related fields.

[0142] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be considered as the scope of protection of the present invention.

Claims

1. A method for translating short play character videos based on a diffusion model, characterized in that: The following steps are involved: S100, firstly, the images in the video frame are segmented into a fixed size, and are grouped together with the original images into a batch and sent to the face detection model for detection, and the detected face model is used as the source face; S200, capturing detailed features of the source face through a GPEN model, and enhancing the detailed features of the source face; S300, enhance the brightness and contrast of the target image through CRNet; S400, effectively separating the detail features of the source face through a decoupling network; S500, inputting the audio signal into the AudioNet network to convert it into a feature representation after noise reduction, and extracting audio features related to the speaker through the AudioNet network; S600, through a multimodal fusion mechanism, the face-changing and lip-alignment tasks are combined for end-to-end training, ensuring the accuracy and consistency of face-changing and lip-alignment in the final generated video.

2. The method for translating short play characters' videos based on a diffusion model according to claim 1, characterized in that: The specific method of step S100 is: S110, image enhancement: the input image is enhanced by the CRNet to improve the image quality and the detail features; S120, batch processing: combining multiple enhanced images into a batch for subsequent processing; S130, face detection: use the face detection model to perform face detection on the images in the batch, identify and crop the face images.

3. The method for translating short play characters based on a diffusion model according to claim 1, characterized in that: The specific method of step S200 is: Face enhancement: The cropped face image is further enhanced by the GPEN model to improve the quality and detail features of the face image.

4. The method for translating short play characters based on a diffusion model according to claim 3 is characterized in that: The detailed features include: identity features of the source face and attribute features of the target face.

5. The method for translating short play characters' videos based on a diffusion model according to claim 4 is characterized in that: The face detection model is: YOLOv8 model.

6. The method for translating short play characters' videos based on a diffusion model according to claim 5 is characterized in that: The AudioNet includes three identical Transform structures, each of which includes: Norm Linear module: The Norm Linear module first normalizes the input features to improve the stability and convergence speed of the model, and secondly maps the features to a new space through linear transformation to enhance the expressiveness of the features; Attention module: used to capture important information in input features and improve the model's attention to key information by weighting the influence of different features; Residual Connection module: It is used to add the input features with the features after Transform processing to help the model better learn the relationship between features, avoid the gradient vanishing problem, and promote the flow of information.

7. A server, characterized in that: include: AttEnc and IdeEnc; The AttEnc is used to extract attribute information from the audio signal; The IdeEnc is used to extract the identity features of the target face in the audio signal; VAE decoder, used to convert latent features into visual face images; AudioNet, which is responsible for converting audio signals into audio embeddings; AdaLN layer, used to enhance the robustness and generation capability of the model.

8. The server according to claim 7, characterized in that: The AttEnc and the IdeEnc are added together to obtain a comprehensive Latent.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method of any one of claims 1 to 6.

10. A system for generating emotional images, characterized in that: include: A processor, a video acquisition device, and an audio acquisition device, wherein the processor communicates with the video acquisition device and the audio acquisition device respectively, and the processor is used to execute the method described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Lip shape synchronization face forgery generation method and system based on image completion

    CN114663962A

  • Image processing method and device, electronic equipment, computer readable storage medium and computer program product

    CN118230081A

  • Speaker video synthesis method, system and device and storage medium

    CN118379401A