Video generation method and device based on lip sound synchronization, equipment and medium

By using a lip-sync video generation method, a model is trained using audio and facial reference images. By combining masking and audio features, high-quality lip-sync videos are generated, solving the problems of visual artifacts and unnatural motion in existing technologies and achieving higher synchronization and naturalness.

CN121099151APending Publication Date: 2025-12-09BEIJING XIAOBING YUEDONG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511126609.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing techniques that enhance temporal coherence by using pre-trained perceptual models or additional sequence discriminators can easily lead to subtle visual artifacts and unnatural mouth movements, affecting the synchronicity and naturalness of lip-synced videos.

Method used

A lip-sync video generation method is adopted. By acquiring audio data and facial reference images, a lip-sync video generation model is trained for forward noise addition and reverse noise reduction. A loss function is constructed by combining a preset mask and audio features to ensure that the model learns the correlation between audio and lip shape. Noise reduction is performed through a backdiffusion unit and an attention mechanism to generate high-quality lip-sync videos.

Benefits of technology

It improves the accuracy and synchronization of lip-sync generation, adapts to different audio styles and character characteristics, reduces the risk of overfitting, reduces reliance on large-scale labeled data, and generates videos with higher naturalness and synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121099151A_ABST
    Figure CN121099151A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device based on lip sound synchronization, equipment and a medium. The method comprises the following steps: acquiring audio data and a face reference image; inputting audio data and the face reference image into a lip synchronization video generation model to obtain a lip synchronization video output by the lip synchronization video generation model; wherein the lip synchronous video generation model is obtained by training according to an audio sample, a face reference sample and a corresponding lip synchronous video sample, and the lip synchronous video generation model is used for carrying out forward noise addition on the input lip synchronous video sample and generating a lip synchronous video according to a noise sequence obtained by forward noise addition and the lip synchronous video sample. And carrying out reverse denoising training by utilizing a preset mask and combining the lip synchronous video sample and the audio sample. According to the method, the accuracy of mouth shape generation is improved, the synchronism and naturalness of the generated video are ensured, different audio styles and character features are adapted, and the over-fitting risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and in particular to a video generation method and device based on lip-sync, equipment and medium. BACKGROUND

[0002] In recent years, audio-driven facial animation technology has made significant progress, mainly benefiting from the application of generative adversarial networks (GANs) and diffusion models. However, the closely related lip-sync technology has relatively lagged behind, despite its important application value in automatic dubbing, virtual avatars, etc. The core goal of audio-driven facial animation methods is to generate high-quality talking head videos that not only maintain the identity features of the input face but also ensure precise synchronization between lip movements and input audio.

[0003] Early methods based on generative adversarial networks (GANs) mainly improve lip-sync accuracy by introducing temporal constraints and expert discriminators, focusing on the animation generation of speaker facial expressions. Subsequent research adds head pose modeling to these methods to generate more realistic animations, but is prone to artifacts and unnatural movements.

[0004] However, methods that enhance temporal coherence by pre-training perception models or additional sequence discriminators can only indirectly control inter-frame consistency, often leading to subtle visual artifacts and unnatural mouth movements, thereby reducing realism and limiting practical applicability. SUMMARY

[0005] The present application provides a video generation method and device based on lip-sync, which solves the defect that the method of enhancing temporal coherence by pre-training perception models or additional sequence discriminators in the prior art is prone to subtle visual artifacts and unnatural mouth movements, improves the accuracy of lip shape generation, ensures the synchronization and naturalness of the generated video, adapts to different audio styles and character features, and reduces the risk of overfitting.

[0006] The present application provides a video generation method based on lip-sync, comprising: obtaining audio data and a face reference image; inputting the audio data and the face reference image into a lip-sync video generation model to obtain a lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained according to an audio sample, a face reference sample, and a corresponding lip-sync video sample, and the lip-sync video generation model is used for forward noise addition on the input lip-sync video sample, and according to the noise sequence obtained by the forward noise addition and the lip-sync video sample, utilizes a preset mask to combine the lip-sync video sample and the audio sample to obtain a backward denoising training result.

[0007] According to the video generation method based on lip synchronization provided by the application, before audio data and a face reference image are input into a lip synchronization video generation model, the method comprises the following steps: obtaining an audio sample, a face reference sample and a corresponding lip synchronization video sample; inputting the audio sample, the face reference sample and the lip synchronization video sample into a to-be-trained model to encode the face reference sample and the lip synchronization video sample, and forwardly adding noise to the encoded video frame sequence; obtaining a first mask result by using a preset mask according to the noise sequence and the encoded video frame sequence, and obtaining a noise fusion video frame sequence by combining the noise sequence; performing audio feature extraction on the audio sample to use the audio feature to perform reverse denoising on the noise fusion video frame sequence, obtain a denoised image frame sequence and decode the denoised image frame sequence to obtain a lip synchronization prediction video; constructing a first loss function according to the denoised image frame sequence and the encoded video frame sequence, constructing a second loss function according to the lip synchronization prediction video and the lip synchronization video sample, obtaining a total loss function according to the first loss function and the second loss function, and determining that the total loss function converges when a preset maximum iteration number is reached to obtain a lip synchronization video generation model for generating a lip synchronization video.

[0008] According to the video generation method based on lip synchronization provided by the application, the lip synchronization video generation model comprises an encoder, a forward noise adding layer, a mask layer, a feature fusion layer, an audio feature extraction layer, a time embedding layer, a reverse denoising layer and a decoder; the lip synchronization prediction video is obtained by: inputting the face reference sample and the lip synchronization video sample into the encoding layer to encode the face reference sample and the lip synchronization video sample to obtain an encoded video frame sequence; inputting the encoded video frame sequence into the forward noise adding layer to perform forward noise adding to obtain a noise sequence; inputting the noise sequence and the encoded video frame sequence into the mask layer to use a preset mask to mask the encoded video frame sequence and use a mask obtained by performing an inversion operation on the preset mask to mask the noise sequence, and splicing the corresponding mask results to obtain a first mask result; inputting the first mask result and the noise sequence into the feature fusion layer to splice them to obtain a noise fusion video frame sequence; inputting the audio sample into the audio feature extraction layer to perform audio feature extraction to obtain an audio feature; inputting the audio feature into the time embedding layer to combine time step embedding to obtain a time step embedded audio feature; inputting the noise fusion video frame sequence, the time step embedded audio feature and the audio feature into the reverse denoising layer to perform reverse denoising to obtain a denoised image frame sequence; inputting the denoised image frame sequence into the decoder to perform decoding to obtain a lip synchronization prediction video.

[0009] According to the present invention, a video generation method based on lip-sync is provided, which encodes a face reference sample and a lip-sync video sample to obtain an encoded video frame sequence. The method includes: extracting features from the lip-sync video sample based on a preset interval of frames to obtain a keyframe sequence; interpolating between adjacent keyframes according to the keyframe sequence at a preset interval of frames to obtain an interpolated feature sequence; repeatedly extracting features from the face reference sample based on the number of frames of the lip-sync video sample to obtain a reference feature sequence; and performing gating processing on the reference feature sequence and the interpolated feature sequence to obtain the encoded video frame sequence.

[0010] According to the present invention, a video generation method based on lip-sync provides an inverse denoising layer comprising at least two connected first inverse diffusion units and at least two connected second inverse diffusion units. The first inverse diffusion unit at the tail end of the at least two connected first inverse diffusion units is connected to the second inverse diffusion unit at the head end of the at least two connected second inverse diffusion units. The method inputs a noisy fused video frame sequence, time-step embedded audio features, and audio features into the inverse denoising layer for inverse denoising to obtain a denoised image frame sequence. This includes: inputting the noisy fused video frame sequence, time-step embedded audio features, and audio features into the head first inverse diffusion unit, so that the input noisy fused video frame sequence or the first inverse diffusion unit passes through at least two connected first inverse diffusion units sequentially, and the input noisy fused video frame sequence is denoised according to the time-step embedded audio features and audio features, respectively. The initial denoised image frame sequence output by the first back diffusion unit is subjected to inverse denoising to obtain the initial denoised image frame sequence output by the first back diffusion unit at the tail. The initial denoised image frame sequence is then input into the second back diffusion unit at the head, and sequentially passed through at least two connected second back diffusion units. Based on the embedded audio features and audio features at the time step, the input initial denoised image frame sequence or the denoised image frame sequence output by the first second back diffusion unit is subjected to inverse denoising to obtain the denoised image frame sequence output by the second back diffusion unit at the tail. The output of the second back diffusion unit at the tail is used as the input of the first back diffusion layer at the head, and the output of the first back diffusion layer at the tail is used as the input of the second back diffusion layer at the head, and iterative denoising is performed until a preset number of iterations is reached to obtain the final denoised image frame sequence.

[0011] According to the present invention, a video generation method based on lip-sync provides a first backdiffusion unit comprising a first residual block, a first audio cross-attention sublayer, and a first temporal attention sublayer. The method inputs a noisy fused video frame sequence, time-step embedded audio features, and audio features into the head first backdiffusion unit, so that the input noisy fused video frame sequence or the preliminary denoised image frame sequence output by the previous first backdiffusion unit is reverse-denoised according to the time-step embedded audio features and audio features, respectively, to obtain the preliminary denoised image frame sequence output by the tail first backdiffusion unit. This includes inputting the time-step embedded audio features and the noisy fused video frame sequence or the preliminary denoised image frame sequence output by the previous first backdiffusion unit into the first residual block of the corresponding first backdiffusion unit to perform reverse denoising on the noisy fused video frame sequence or the preliminary denoised image frame sequence output by the previous first backdiffusion unit. Feature extraction is performed on the video frame sequence or the preliminary denoised image frame sequence output by the first backdiffusion unit, and the time-step embedded audio features are fused to obtain the first denoised frame sequence output by the first residual block; the first denoised frame sequence and audio features are input into the first audio cross-attention sublayer of the corresponding first backdiffusion unit to determine the first audio cross-attention weight between the audio features and the first denoised frame sequence based on the audio cross-attention mechanism, and combined with the first denoised frame sequence to obtain the second denoised frame sequence output by the first audio cross-attention sublayer; the second denoised frame sequence is input into the first temporal attention sublayer of the corresponding first backdiffusion unit to determine the temporal attention weight between different frames in the second denoised frame sequence based on the temporal attention mechanism, and combined with the second denoised frame sequence to obtain the preliminary denoised image sequence output by the corresponding first backdiffusion unit; The second backdiffusion unit includes a second residual block, a second audio cross-attention sublayer, and a second temporal attention sublayer. A preliminary denoised image frame sequence is input into the head second backdiffusion unit, and sequentially passed through at least two connected second backdiffusion units. Based on the time-step embedded audio features and audio features, the input preliminary denoised image frame sequence or the denoised image frame sequence output by the previous second backdiffusion unit is subjected to reverse denoising to obtain the denoised image frame sequence output by the tail second backdiffusion unit. This includes: inputting the time-step embedded audio features and the preliminary denoised image frame sequence or the denoised image frame sequence output by the previous second backdiffusion unit into the second residual block of the corresponding second backdiffusion unit, to perform reverse denoising on the input preliminary denoised image frame sequence or the denoised image frame sequence output by the previous second backdiffusion unit based on the time-step embedded audio features and audio features. Feature extraction is performed on the noisy image frame sequence, and the time-step embedded audio features are fused to obtain the third denoised frame sequence output by the second residual block. The third denoised frame sequence and the audio features are input into the second audio cross-attention sublayer of the corresponding second backdiffusion unit to determine the second audio cross-attention weight between the audio features and the third denoised frame sequence based on the audio cross-attention mechanism. Combined with the third denoised frame sequence, the fourth denoised frame sequence output by the second audio cross-attention sublayer is obtained. The fourth denoised frame sequence is input into the second temporal attention sublayer of the corresponding second backdiffusion unit to determine the temporal attention weight between different frames in the fourth denoised frame sequence based on the temporal attention mechanism. Combined with the fourth denoised frame sequence, the denoised image frame sequence output by the corresponding second backdiffusion unit is obtained.

[0012] According to the present invention, a video generation method based on lip-sync is provided, which obtains a total loss function based on a first loss function and a second loss function, including: occluding object samples of face reference samples using a zero-sample video segmentation model based on lip-sync video samples to generate an occlusion mask; performing a logical NOT operation on the occlusion mask and a logical intersection operation on it with a preset mask to obtain a combined mask; and constructing a total loss function based on the combined mask, the first loss function, and the second loss function.

[0013] This invention also provides a lip-sync video generation device, comprising: a data acquisition module for acquiring audio data and a facial reference image; and a video generation module for inputting the audio data and the facial reference image into a lip-sync video generation model to obtain a lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained based on audio samples, facial reference samples, and corresponding lip-sync video samples. The lip-sync video generation model is used to perform forward noise addition on the input lip-sync video samples, and then, based on the noise sequence obtained by forward noise addition and the lip-sync video samples, uses a preset mask to perform reverse denoising training in combination with the lip-sync video samples and audio samples.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video generation method based on lip-sync as described above.

[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the lip-sync-based video generation method as described above.

[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the video generation method based on lip-sync as described above.

[0017] The present invention provides a lip-sync video generation method, apparatus, device, and medium that simultaneously provides audio and facial reference images to ensure that the model can learn the correlation between audio and lip movements while maintaining the personalized features of the face. This allows the model to generate lip-sync videos using a lip-sync video generation model, supporting the generation of lip-sync videos for different individuals and avoiding the lip-sync mismatch problem that may occur with general models. Furthermore, the model progressively adds noise to the input lip-sync video samples to disrupt them, simulating the challenges of the real generation process and enhancing the model's robustness. Combined with audio as a condition, the model is guided to restore realistic lip-sync dynamics during denoising, ensuring the synchronization and naturalness of the generated video. Preset masks limit the key areas of focus for the model, such as the lips, to avoid interference from irrelevant facial features and improve the accuracy of lip-sync generation. Joint training with noise sequences and masks enables the model to adapt to different audio styles and individual characteristics, reducing the risk of overfitting and allowing it to learn more generalized lip-sync rules from limited samples, thus reducing dependence on large-scale labeled data. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the video generation method based on lip-sync provided by the present invention; Figure 2 This is a schematic diagram of the architecture of the lip-synchronized video generation model provided by the present invention; Figure 3 This is a schematic diagram of the process for obtaining an encoded video frame sequence provided by the present invention; Figure 4 This is a schematic diagram of the process for obtaining the combined mask provided by the present invention; Figure 5 This is a schematic diagram of the structure of the video generation device based on lip-sync provided by the present invention; Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0021] Figure 1 This is a flowchart illustrating the video generation method based on lip-sync provided by the present invention, as shown below. Figure 1 As shown, the method includes: S11, acquire audio data and facial reference image; S12, the audio data and the face reference image are input into the lip-synchronized video generation model to obtain the lip-synchronized video output by the lip-synchronized video generation model; wherein, the lip-synchronized video generation model is trained based on the audio samples, face reference samples and corresponding lip-synchronized video samples. The lip-synchronized video generation model is used to perform forward noise addition on the input lip-synchronized video samples, and to perform reverse noise reduction training based on the noise sequence obtained by forward noise addition and the lip-synchronized video samples, using a preset mask, combined with the lip-synchronized video samples and audio samples.

[0022] It should be noted that the step number "S1N" in this manual does not represent the order of the lip-sync-based video generation methods. The following details will explain further. Figure 2 This invention describes a video generation method based on lip-sync.

[0023] Step S11: Obtain audio data and a facial reference image.

[0024] It should be noted that the audio data is used to represent the speech information of the lip-synced video to be generated; the face reference image can be selected according to the actual video generation requirements. For example, when it is necessary to strictly maintain the consistency of the person's identity, the face reference image and the real face image need to come from the same person; when it is only necessary to produce lip movements that conform to the audio, the face reference image does not need to be bound to a specific person and can be any face, such as a standardized neutral expression face, without further restrictions here.

[0025] Step S12: Input the audio data and the face reference image into the lip-synchronized video generation model to obtain the lip-synchronized video output by the lip-synchronized video generation model; wherein, the lip-synchronized video generation model is trained based on the audio samples, the face reference samples and the corresponding lip-synchronized video samples. The lip-synchronized video generation model is used to perform forward noise addition on the input lip-synchronized video samples, and based on the noise sequence obtained by forward noise addition and the lip-synchronized video samples, it is trained by using a preset mask and combining the lip-synchronized video samples and audio samples for reverse noise reduction.

[0026] In an optional embodiment, before inputting the audio data and face reference image into the lip-sync video generation model, the method includes: acquiring audio samples, face reference samples, and corresponding lip-sync video samples; inputting the audio samples, face reference samples, and lip-sync video samples into the model to be trained to encode the face reference samples and lip-sync video samples, and performing forward noise addition on the encoded video frame sequence; based on the noise sequence obtained from the forward noise addition and the encoded video frame sequence, using a preset mask, obtaining a first mask result, and combining it with the lip noise sequence to obtain a noise-fused video frame. The sequence of audio samples is used to extract audio features, which are then used to reverse denoise the fused video frame sequence to obtain a denoised image frame sequence. This sequence is then decoded to obtain a lip-synchronized prediction video. Based on the denoised image frame sequence and the encoded video frame sequence, a first loss function is constructed, and based on the lip-synchronized prediction video and lip-synchronized video samples, a second loss function is constructed. Based on the first and second loss functions, a total loss function is obtained. When the preset maximum number of iterations is reached, the model converges based on the total loss function, resulting in a lip-synchronized video generation model for generating lip-synchronized videos.

[0027] It should be noted that by encoding face reference samples and lip-sync video samples, high-dimensional image and video data are compressed into low-dimensional representations (latent representations), reducing computational complexity while retaining key information, allowing the model to process data more efficiently. Forward noise addition to the encoded video frame sequence helps the model learn the inverse process from noise to real data, thus enabling better recovery of real data during the generation stage. Furthermore, pre-set masks protect key regions from excessive noise interference, and a noise-fused video frame sequence is generated by combining the noise sequence to ensure that the generated video frames are close to the real data in terms of noise level. The approach achieves a balance by preserving the structure of the real data while introducing sufficient noise for the model to learn. It also incorporates reverse denoising using audio features extracted from audio samples to ensure that the generated video frames are highly synchronized with the audio content, achieving a lip-sync effect. Furthermore, by combining a first loss function to measure the difference between the denoised image frames and the original encoded video frames, the model can learn a more accurate denoising process, thereby generating higher-quality video frames. Finally, by using a second loss function to measure the difference between the generated lip-synced video and the real lip-synced video, the model can learn a more accurate lip-sync effect, ensuring a high degree of match between the generated video and the audio.

[0028] Specifically, refer to Figure 2 The lip-synchronized video generation model includes an encoder, a forward noise layer, a masking layer, a feature fusion layer, an audio feature extraction layer, a temporal embedding layer, a reverse denoising layer, and a decoder. To obtain the predicted lip-synchronized video, the model includes: inputting face reference samples and lip-synchronized video samples into the encoding layer to encode them, obtaining an encoded video frame sequence; inputting the encoded video frame sequence into the forward noise layer for forward noise addition, obtaining a noise sequence; and inputting the noise sequence and the encoded video frame sequence into the masking layer to mask the encoded video frame sequence using a preset mask, and using a mask obtained by inverting the preset mask. The noise sequence is masked, and the corresponding mask results are concatenated to obtain the first mask result. The first mask result and the noise sequence are input into the feature fusion layer for concatenation to obtain the noisy fused video frame sequence. Audio samples are input into the audio feature extraction layer for audio feature extraction to obtain audio features. The audio features are input into the temporal embedding layer to combine with temporal embedding to obtain temporal embedded audio features. The noisy fused video frame sequence, temporal embedded audio features, and audio features are input into the inverse denoising layer for inverse denoising to obtain the denoised image frame sequence. The denoised image frame sequence is input into the decoder for decoding to obtain the lip-sync prediction video.

[0029] It's important to note that the encoder compresses the input face image and synchronized lip-sync video into a low-dimensional latent representation to retain the most crucial information from the input data, reducing the amount of data and computational complexity in subsequent processing. A forward noise-adding layer adds noise to the encoded video frame sequence, allowing the model to learn how to progressively denoise pure noise back to the real data. This enables the model to learn the complete distribution of data in the latent space. Furthermore, masking operations in the masking layer emphasize information in specific regions, and the masked results are concatenated to more precisely control which parts the model should focus on or recover in subsequent steps. Finally, a feature fusion layer combines the masked information with the noise sequence, enhancing the model's denoising capabilities. This approach utilizes both the original structural information and the masked portions, providing a more powerful feature representation for subsequent inverse denoising. It combines audio features with temporal embedding vectors representing time positions, enabling the model to understand the changes and positions of audio features on the time axis. This helps the model distinguish phonemes and lip movements at different time points. Thus, the inverse denoising layer integrates the noise-fused video frame sequence (visual reference), time-step embedded audio features (temporal information), and audio features (driving information) for inverse denoising, achieving the synergistic effect of multimodal information to generate lip movements synchronized with the audio. Finally, the decoder converts the potential video frame sequence output by the inverse denoising layer back into the original spatial video.

[0030] It should be added that the encoder and decoder can be Variational Autoencoder (VAE) encoder and decoder, and the inverse denoising layer can be a denoising network based on U-Net. The specific choice can be made according to the actual design requirements, and no further restrictions are made here.

[0031] Additionally, the first mask result is represented as: in, This represents the result of the first mask; Indicates the default mask; This represents the mask after inverting the preset mask. This represents element-wise multiplication. Represents a noise sequence The t-th frame of the noisy video image; Represents the sequence of encoded video frames The t-th frame of the video image.

[0032] Furthermore, the preset mask can be generated by detecting facial key points on face reference samples or lip-synchronized video samples to reduce error propagation, and combined with preset mask rules. The preset mask rules can be designed according to the actual design goals. For example, if the goal is to ensure that the newly generated lip region does not reuse (or "leak") the original mouth shape features that contradict the original audio, the preset mask rules are as follows: isolate the lower half of the face and extend slightly above the nose to cover the upper cheek area that may transmit lip movement information, while fully preserving facial identity features. At the same time, the mask also extends to the lower edge of the image to effectively prevent feature leakage caused by jaw movement, so as to fully cover the lower half of the face while preserving necessary contextual information (such as cheek and chin movements). This makes it easier to generate lips in reverse denoising, relying not only on audio but also combining facial contextual information to avoid the unnatural problem of "isolated lips". The study found that the mask achieves the best balance between the following two types of traditional masks: full lower face mask: may obscure too much contextual information, leading to loss of identity features and abnormal facial continuity; pure mouth mask: because residual mouth movement or lighting information is visible, it is easy to cause leakage of lower facial expressions.

[0033] Furthermore, when extracting audio features, the HuberT audio encoder can be used to convert audio samples into corresponding learning representations. After obtaining the audio features, before combining them with time-step embedding, a multilayer perceptron (MLP) is used to process the audio features to extract high-level features. Based on the extracted high-level features, combined with time-step embedding, time-step embedded audio features are obtained. This facilitates the reverse denoising layer's dual fusion mechanism based on time-step embedded audio features and audio features, significantly enhancing the alignment capability between video frames and audio frames, thereby improving lip synchronization.

[0034] To go further, refer to Figure 3 The process involves encoding face reference samples and lip-synchronized video samples to obtain an encoded video frame sequence. This includes: extracting features from the lip-synchronized video samples based on a preset frame interval to obtain a keyframe sequence; performing interpolation between adjacent keyframes according to the keyframe sequence at a preset frame interval to obtain an interpolated feature sequence; repeatedly extracting features from the face reference samples based on the number of frames in the lip-synchronized video samples to obtain a reference feature sequence; and performing gating processing on the reference feature sequence and the interpolated feature sequence to obtain the encoded video frame sequence.

[0035] It should be noted that the interpolation feature sequence is represented as: in, Represents the interpolation feature sequence; () represents two adjacent keyframes in a keyframe sequence; The learnable embedding vector representing the missing frame. Through ( The interpolation is repeated S times to achieve a smooth and time-consistent motion effect, where S represents the preset frame interval. Multi-frame feature extraction of the lip-sync video samples is performed based on the preset frame interval to avoid reliance on a single reference image and enhance robustness.

[0036] In an optional embodiment, the inverse denoising layer includes at least two connected first inverse diffusion units and at least two connected second inverse diffusion units. The first inverse diffusion unit at the tail end of the at least two connected first inverse diffusion units is connected to the second inverse diffusion unit at the head end of the at least two connected second inverse diffusion units. The noisy fused video frame sequence, time-step embedded audio features, and audio features are input into the inverse denoising layer for inverse denoising to obtain a denoised image frame sequence. This includes: inputting the noisy fused video frame sequence, time-step embedded audio features, and audio features into the head first inverse diffusion unit, so that the input noisy fused video frame sequence or the first inverse diffusion unit passes through at least two connected first inverse diffusion units sequentially, and the input noisy fused video frame sequence is denoised according to the time-step embedded audio features and audio features, respectively. The output of the preliminary denoised image frame sequence is subjected to inverse denoising to obtain the preliminary denoised image frame sequence output by the first inverse diffusion unit at the tail. The preliminary denoised image frame sequence is then input into the second inverse diffusion unit at the head, and passes through at least two connected second inverse diffusion units in sequence. Based on the embedded audio features and audio features at the time step, the input preliminary denoised image frame sequence or the denoised image frame sequence output by the first inverse diffusion unit is subjected to inverse denoising to obtain the denoised image frame sequence output by the second inverse diffusion unit at the tail. The output of the second inverse diffusion unit at the tail is used as the input of the first inverse diffusion layer at the head, and the output of the first inverse diffusion layer at the tail is used as the input of the second inverse diffusion layer at the head, and iterative denoising is performed until a preset number of iterations is reached to obtain the final denoised image frame sequence.

[0037] Specifically, the first backdiffusion unit includes a first residual block, a first audio cross-attention sublayer, and a first temporal attention sublayer. The noisy fused video frame sequence, time-step embedded audio features, and audio features are input into the head first backdiffusion unit. The noisy fused video frame sequence or the preliminary denoised image frame sequence output by the previous first backdiffusion unit is then processed sequentially through at least two connected first backdiffusion units. Based on the time-step embedded audio features and audio features, the input noisy fused video frame sequence or the preliminary denoised image frame sequence output by the previous first backdiffusion unit is subjected to reverse denoising, resulting in the preliminary denoised image frame sequence output by the tail first backdiffusion unit. This includes inputting the time-step embedded audio features and the noisy fused video frame sequence or the preliminary denoised image frame sequence output by the previous first backdiffusion unit into the first residual block of the corresponding first backdiffusion unit, to process the noisy fused video frame sequence or the preliminary denoised image frame sequence output by the previous first backdiffusion unit. Feature extraction is performed on the initial denoised image frame sequence output by a backdiffusion unit, and the time-step embedded audio features are fused to obtain the first denoised frame sequence output by the first residual block. The first denoised frame sequence and audio features are input into the first audio cross-attention sublayer of the corresponding first backdiffusion unit to determine the first audio cross-attention weight between the audio features and the first denoised frame sequence based on the audio cross-attention mechanism. Combined with the first denoised frame sequence, the second denoised frame sequence output by the first audio cross-attention sublayer is obtained. The second denoised frame sequence is input into the first temporal attention sublayer of the corresponding first backdiffusion unit to determine the temporal attention weight between different frames in the second denoised frame sequence based on the temporal attention mechanism. Combined with the second denoised frame sequence, the initial denoised image sequence output by the corresponding first backdiffusion unit is obtained.

[0038] Additionally, the second backdiffusion unit includes a second residual block, a second audio cross-attention sublayer, and a second temporal attention sublayer. The initial denoised image frame sequence is input into the head second backdiffusion unit, and sequentially passed through at least two connected second backdiffusion units. Based on the time-step embedded audio features and audio features, the input initial denoised image frame sequence or the denoised image frame sequence output by the previous second backdiffusion unit is subjected to reverse denoising to obtain the denoised image frame sequence output by the tail second backdiffusion unit. This includes: inputting the time-step embedded audio features and the initial denoised image frame sequence or the denoised image frame sequence output by the previous second backdiffusion unit into the second residual block of the corresponding second backdiffusion unit, to perform reverse denoising on the time-step embedded audio features or the denoised image frame sequence output by the previous second backdiffusion unit. Feature extraction is performed on the denoised image frame sequence, and the time-step embedded audio features are fused to obtain the third denoised frame sequence output by the second residual block. The third denoised frame sequence and audio features are input into the second audio cross-attention sublayer of the corresponding second backdiffusion unit to determine the second audio cross-attention weight between the audio features and the third denoised frame sequence based on the audio cross-attention mechanism. Combined with the third denoised frame sequence, the fourth denoised frame sequence output by the second audio cross-attention sublayer is obtained. The fourth denoised frame sequence is input into the second temporal attention sublayer of the corresponding second backdiffusion unit to determine the temporal attention weight between different frames in the fourth denoised frame sequence based on the temporal attention mechanism. Combined with the fourth denoised frame sequence, the denoised image frame sequence output by the corresponding second backdiffusion unit is obtained.

[0039] In an optional embodiment, the total loss function is obtained based on the first loss function and the second loss function, including: using a zero-shot video segmentation model to occlude objects on face reference samples based on lip-sync video samples to generate an occlusion mask; performing a logical NOT operation on the occlusion mask and a logical intersection operation on it in combination with a preset mask to obtain a combined mask; and constructing the total loss function based on the combined mask, the first loss function, and the second loss function.

[0040] It should be noted that the process for obtaining the combined mask can be found by referring to [reference needed]. Figure 4 The combined mask is represented as: in, Indicates a combined mask; Indicates the default mask; Represents the logical intersection operation; Indicates the logical NOT operation; This represents the occlusion mask. Further, occluded objects are segmented using a zero-shot video segmentation model to generate occlusion masks. By excluding occluded regions, the preset mask M is optimized. This ensures seamless reconstruction of the mouth region while preserving the occluded object, and only the generated region participates in the loss calculation, thus guaranteeing visual continuity. It avoids the situation where the model incorrectly generates a mouth on the occluded object, leading to boundary artifacts, when the occluded region overlaps with the mouth mask region. This ensures stable effectiveness when handling various types of occlusion.

[0041] It should be added that the total loss function is expressed as: in, Represents the total loss function; Indicates a combined mask; This represents the weighting factor that depends on time step t, where T represents the total number of time steps. By employing a fixed strategy, the stability of training can be improved, avoiding the negative impact of dynamic adjustments. Introducing additional variance; Represents the first loss function; This represents the second loss function; Indicates the weighting coefficient; The expected value of the loss on the training data can be represented by the distribution of the input data ( The average is obtained; This indicates a preset weighting function, based on noise level. The loss weights are dynamically adjusted, and high-frequency details (high noise) are usually given higher weights. This can be selected according to actual design requirements, such as signal-to-weight ratio, etc. No further restrictions are made here. Indicates the model to be trained. Indicates the parameters corresponding to the model; This represents the denoised image at frame t in the denoised image frame sequence; The conditional inputs to the model include audio features and time-step embedded audio features; Indicates noise level; Represents the sequence of encoded video frames The t-th frame of the video image; This represents the predicted image of the t-th frame in the lip-sync prediction video output by the decoder. Indicates decoder; This represents the t-th frame of a synchronized lip-sync video sample. This indicates a denoiser; the inverse denoising layer can use the denoiser for inverse denoising. , , and Indicates dependence on noise level The scaling factors can dynamically adjust the intensity and impact of noise at different stages of the denoising process, thereby effectively improving the computational efficiency and robustness of the network during the diffusion process. Indicates input data, including and , .

[0042] In summary, this invention provides both audio and facial reference images simultaneously, ensuring the model can learn the correlation between audio and lip movements while maintaining the individualized features of the face. This allows the model to generate synchronized lip-sync videos, supporting the generation of synchronized lip-sync videos for different individuals and avoiding the lip-sync mismatch problem that may occur with general models. Furthermore, the model progressively adds noise to the input synchronized lip-sync video samples to disrupt them, simulating the challenges of the real generation process and enhancing the model's robustness. Combined with audio as a condition, the model is guided to restore realistic lip-sync dynamics during denoising, ensuring the synchronization and naturalness of the generated video. Preset masks limit the model's focus to key areas, such as the lips, to avoid interference from irrelevant facial features and improve the accuracy of lip-sync generation. Joint training with noise sequences and masks enables the model to adapt to different audio styles and individual characteristics, reducing the risk of overfitting. It can also learn more generalized lip-sync rules from limited samples, reducing reliance on large-scale labeled data.

[0043] The lip-sync video generation device provided by the present invention is described below. The lip-sync video generation device described below and the lip-sync video generation method described above can be referred to in correspondence.

[0044] Figure 5 A schematic diagram of a video generation device based on lip-sync is shown. The device includes: Data acquisition module 51 acquires audio data and facial reference images; The video generation module 52 inputs audio data and a face reference image into the lip-synchronized video generation model to obtain the lip-synchronized video output by the lip-synchronized video generation model. The lip-synchronized video generation model is trained based on audio samples, face reference samples, and corresponding lip-synchronized video samples. The lip-synchronized video generation model is used to perform forward noise addition on the input lip-synchronized video samples, and then perform reverse denoising training based on the noise sequence obtained by forward noise addition and the lip-synchronized video samples, using a preset mask, combined with the lip-synchronized video samples and audio samples.

[0045] It should be noted that the specific principles of the embodiments of the present invention are the same as those of the method embodiments described above. For details, please refer to the method embodiments above. More detailed explanations will not be repeated here.

[0046] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640. The processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a lip-sync video generation method. This method includes: acquiring audio data and a face reference image; inputting the audio data and the face reference image into a lip-sync video generation model to obtain a lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained based on audio samples, face reference samples, and corresponding lip-sync video samples. The lip-sync video generation model is used to perform forward noise addition on the input lip-sync video samples, and then, based on the noise sequence obtained from the forward noise addition and the lip-sync video samples, uses a preset mask and combines the lip-sync video samples and audio samples for inverse denoising training.

[0047] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0048] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the lip-sync video generation method provided by the above methods. The method includes: acquiring audio data and a face reference image; inputting the audio data and the face reference image into a lip-sync video generation model to obtain a lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained based on audio samples, face reference samples, and corresponding lip-sync video samples. The lip-sync video generation model is used to perform forward noise addition on the input lip-sync video samples, and to perform reverse denoising training based on the noise sequence obtained by forward noise addition and the lip-sync video samples, using a preset mask, combined with the lip-sync video samples and audio samples.

[0049] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the lip-sync video generation method provided by the above methods. The method includes: acquiring audio data and a face reference image; inputting the audio data and the face reference image into a lip-sync video generation model to obtain a lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained based on audio samples, face reference samples, and corresponding lip-sync video samples. The lip-sync video generation model is used to perform forward noise addition on the input lip-sync video samples, and to perform reverse denoising training based on the noise sequence obtained by forward noise addition and the lip-sync video samples, using a preset mask, combined with the lip-sync video samples and audio samples.

[0050] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0051] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video generation method based on lip-sync, characterized in that, include: Acquire audio data and facial reference images; The audio data and the face reference image are input into the lip-synchronized video generation model to obtain the lip-synchronized video output by the lip-synchronized video generation model. The lip-synchronized video generation model is trained based on audio samples, face reference samples, and corresponding lip-synchronized video samples. The lip-synchronized video generation model is used to perform forward noise addition on the input lip-synchronized video samples, and then perform reverse denoising training based on the noise sequence obtained by forward noise addition and the lip-synchronized video samples, using a preset mask, combined with the lip-synchronized video samples and the audio samples.

2. The video generation method based on lip-sync according to claim 1, characterized in that, Before inputting the audio data and the facial reference image into the lip-sync video generation model, the process includes: Acquire audio samples, face reference samples, and corresponding lip-sync video samples; The audio samples, the face reference samples, and the lip-sync video samples are input into the model to be trained to encode the face reference samples and the lip-sync video samples. The encoded video frame sequence is then forward-denoised. Based on the noise sequence obtained by forward-denoising and the encoded video frame sequence, a first mask result is obtained using a preset mask. Combined with the noise sequence, a noisy fused video frame sequence is obtained. Audio features are extracted from the audio samples to reverse-denoise the noisy fused video frame sequence, resulting in a denoised image frame sequence, which is then decoded to obtain the lip-sync prediction video. A first loss function is constructed based on the denoised image frame sequence and the encoded video frame sequence, and a second loss function is constructed based on the lip-synchronized prediction video and the lip-synchronized video samples. A total loss function is obtained based on the first loss function and the second loss function. When the preset maximum number of iterations is reached, the model converges based on the total loss function, and a lip-synchronized video generation model for generating lip-synchronized videos is obtained.

3. The video generation method based on lip-sync according to claim 2, characterized in that, The lip-synchronized video generation model includes an encoder, a forward noise layer, a mask layer, a feature fusion layer, an audio feature extraction layer, a temporal embedding layer, a reverse noise reduction layer, and a decoder. The obtained lip-synchronized prediction video includes: The face reference sample and the lip-synchronized video sample are input into the coding layer to encode the face reference sample and the lip-synchronized video sample to obtain a sequence of encoded video frames. The encoded video frame sequence is input into the forward noise layer for forward noise addition to obtain a noise sequence; The noise sequence and the encoded video frame sequence are input into the masking layer to mask the encoded video frame sequence using a preset mask and to mask the noise sequence using a mask obtained by inverting the preset mask. The corresponding masking results are then concatenated to obtain the first masking result. The first mask result and the noise sequence are input into the feature fusion layer for concatenation to obtain a noisy fused video frame sequence. The audio sample is input into the audio feature extraction layer for audio feature extraction to obtain audio features; The audio features are input into the temporal embedding layer to combine with temporal step embedding to obtain temporal step embedded audio features; The noise-fused video frame sequence, the time-step embedded audio features, and the audio features are input into the inverse denoising layer for inverse denoising to obtain a denoised image frame sequence. The denoised image frame sequence is input into the decoder for decoding to obtain the lip-sync prediction video.

4. The video generation method based on lip-sync according to claim 3, characterized in that, The face reference sample and the lip-sync video sample are encoded to obtain an encoded video frame sequence, including: Based on a preset interval of frames, feature extraction is performed on the lip-synchronized video samples to obtain a keyframe sequence; Based on the keyframe sequence, interpolation is performed between adjacent keyframes according to the preset interval number of frames to obtain an interpolation feature sequence. Based on the number of frames of the lip-synchronized video sample, feature extraction is repeatedly performed on the face reference sample to obtain a reference feature sequence. The reference feature sequence and the interpolated feature sequence are subjected to gating processing to obtain the encoded video frame sequence.

5. The video generation method based on lip-sync according to claim 3, characterized in that, The inverse denoising layer includes at least two connected first inverse diffusion units and at least two connected second inverse diffusion units. The tail-end first inverse diffusion unit of the at least two connected first inverse diffusion units is connected to the head-end second inverse diffusion unit of the at least two connected second inverse diffusion units. The noisy fused video frame sequence, the time-step embedded audio features, and the audio features are input into the inverse denoising layer for inverse denoising to obtain a denoised image frame sequence, including: The noise-added fused video frame sequence, the time-step embedded audio features, and the audio features are input into the head first back diffusion unit, so that they pass through at least two connected first back diffusion units in sequence. Based on the time-step embedded audio features and the audio features, the input noise-added fused video frame sequence or the preliminary denoised image frame sequence output by the first first back diffusion unit is reversed to obtain the preliminary denoised image frame sequence output by the tail first back diffusion unit. The initial denoised image frame sequence is input into the head second back diffusion unit, and then passes through at least two connected second back diffusion units in sequence. Based on the time step embedded audio features and the audio features, the input initial denoised image frame sequence or the denoised image frame sequence output by the first second back diffusion unit is denoised in reverse, so as to obtain the denoised image frame sequence output by the tail second back diffusion unit. The output of the second back diffusion unit at the tail is used as the input of the first back diffusion layer at the head, and the output of the first back diffusion layer at the tail is used as the input of the second back diffusion layer at the head. Iterative denoising is performed until the preset number of iterations is reached, and the final denoised image frame sequence is obtained.

6. The video generation method based on lip-sync according to claim 5, characterized in that, The first back-diffusion unit includes a first residual block, a first audio cross-attention sublayer, and a first temporal attention sublayer; The noisy fused video frame sequence, the time-step embedded audio features, and the audio features are input into the head first back-diffusion unit, so that they pass sequentially through at least two connected first back-diffusion units. Based on the time-step embedded audio features and the audio features, the input noisy fused video frame sequence or the preliminary denoised image frame sequence output by the first first back-diffusion unit is subjected to reverse denoising, resulting in the preliminary denoised image frame sequence output by the tail first back-diffusion unit, including: The time-step embedded audio features and the noisy fused video frame sequence or the preliminary denoised image frame sequence output by the first back diffusion unit are input into the first residual block of the corresponding first back diffusion unit to extract features from the noisy fused video frame sequence or the preliminary denoised image frame sequence output by the first back diffusion unit, and fuse the time-step embedded audio features to obtain the first denoised frame sequence output by the first residual block. The first denoised frame sequence and the audio features are input into the first audio cross-attention sub-layer of the corresponding first back diffusion unit, so as to determine the first audio cross-attention weight between the audio features and the first denoised frame sequence based on the audio cross-attention mechanism, and combine the first denoised frame sequence to obtain the second denoised frame sequence output by the first audio cross-attention sub-layer. The second denoised frame sequence is input into the first temporal attention sublayer of the corresponding first back diffusion unit to determine the temporal attention weights between different frames in the second denoised frame sequence based on the temporal attention mechanism, and combined with the second denoised frame sequence to obtain the preliminary denoised image sequence output by the corresponding first back diffusion unit. The second backdiffusion unit includes a second residual block, a second audio cross-attention sublayer, and a second temporal attention sublayer. The initial denoised image frame sequence is input into the head second backdiffusion unit, and sequentially passes through at least two connected second backdiffusion units. Based on the embedded audio features at the time step and the audio features, the input initial denoised image frame sequence or the denoised image frame sequence output by the previous second backdiffusion unit is subjected to reverse denoising to obtain the denoised image frame sequence output by the tail second backdiffusion unit, including: The time-step embedded audio features and the preliminary denoised image frame sequence or the denoised image frame sequence output by the first second back diffusion unit are input into the second residual block of the corresponding second back diffusion unit to extract features from the time-step embedded audio features or the denoised image frame sequence output by the first second back diffusion unit, and the time-step embedded audio features are fused to obtain the third denoised frame sequence output by the second residual block. The third denoised frame sequence and the audio features are input into the second audio cross-attention sub-layer of the corresponding second back diffusion unit, so as to determine the second audio cross-attention weight between the audio features and the third denoised frame sequence based on the audio cross-attention mechanism, and combine the third denoised frame sequence to obtain the fourth denoised frame sequence output by the second audio cross-attention sub-layer. The fourth denoised frame sequence is input into the second temporal attention sublayer of the corresponding second backdiffusion unit to determine the temporal attention weights between different frames in the fourth denoised frame sequence based on the temporal attention mechanism, and combined with the fourth denoised frame sequence, the denoised image frame sequence output by the corresponding second backdiffusion unit is obtained.

7. The video generation method based on lip-sync according to claim 2, characterized in that, Based on the first loss function and the second loss function, the total loss function is obtained, including: Based on the lip-sync video samples, the face reference samples are occluded using a zero-sample video segmentation model to generate an occlusion mask. Perform a logical NOT operation on the occlusion mask and a logical intersection operation on the preset mask to obtain a combined mask; Based on the combined mask, the first loss function, and the second loss function, a total loss function is constructed.

8. A video generation device based on lip-sync, characterized in that, include: The data acquisition module acquires audio data and facial reference images; The video generation module inputs the audio data and the face reference image into the lip-synchronized video generation model to obtain the lip-synchronized video output by the lip-synchronized video generation model. The lip-synchronized video generation model is trained based on audio samples, face reference samples, and corresponding lip-synchronized video samples. The lip-synchronized video generation model is used to perform forward noise addition on the input lip-synchronized video samples, and then, based on the noise sequence obtained from the forward noise addition and the lip-synchronized video samples, performs reverse denoising training using a preset mask, combining the lip-synchronized video samples and the audio samples.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the lip-sync-based video generation method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the lip-sync-based video generation method as described in any one of claims 1 to 7.