Audio-driven lip synchronization method and device, equipment and medium
By training a lip-sync video generation model and combining the features of audio and facial reference samples, the problem of limited audio type diversity in existing technologies is solved, and accurate synchronization and rich detail of lip movement generation are achieved, enhancing the realism and personalization of the video.
Patent Information
- Application Number
- CN202510771256.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-05
AI Technical Summary
The audio-driven lip synchronization task in the existing technology relies on expert experience, which limits the diversity of input audio types and results in deficiencies in the visual effects and synchronization of the generated videos.
By combining multimodal features and denoising mechanisms, the lip-sync video generation model is trained using audio samples and facial reference samples, and forward denoising and reverse denoising are performed to generate accurately synchronized lip movements and rich facial details.
The generated video can maintain the facial features and identity of the original person, enhance the realism and personalization, have strong generalization ability, and can generate videos that meet specific requirements.
Smart Images

Figure CN120602741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an audio-driven lip synchronization method, device, equipment and medium. Background Art
[0002] The audio-driven lip synchronization task aims to generate synchronized lip movements in a video based on the input audio while maintaining the identity and appearance consistency of the speaker. It has important applications in video dubbing, virtual avatars, and live streaming platforms.
[0003] In the field of video generation, audio-driven lip synchronization tasks mainly utilize GAN-based methods, which introduce lip experts to achieve strong audio-visual alignment in pixel space. However, the realistic priors relied on by such experts limit the diversity of input audio types. Summary of the Invention
[0004] The present invention provides an audio-driven lip synchronization method, apparatus, device, and medium to address the drawback of existing technologies that rely on expert experience and thus limit the diversity of input audio types. By combining multimodal features and a denoising mechanism, a lip synchronization video is generated with precise lip movement synchronization, rich facial details, and a natural overall visual effect.
[0005] The present invention provides an audio-driven lip synchronization method, comprising: obtaining a real face image, audio data, and a face reference image; inputting the real face image, audio data, and the face reference image into a lip synchronization video generation model to obtain a lip synchronization video output by the lip synchronization video generation model; wherein the lip synchronization video generation model is trained based on audio samples, face reference samples, and corresponding lip synchronization video samples, and the lip synchronization video generation model is obtained by forward denoising the input lip synchronization video samples, and performing reverse denoising training on the noise sequence obtained by the forward denoising in combination with the audio samples and the face reference samples.
[0006] According to the present invention, an audio-driven lip synchronization method is provided. Before inputting a real face image, audio data and a face reference image into a lip synchronization video generation model, the method includes: obtaining audio samples, face reference samples and corresponding lip synchronization video samples; inputting the audio samples, face reference samples and lip synchronization video samples into the model to be trained, adding corresponding noise to each frame image in the lip synchronization video sample to obtain a noise sequence, performing feature extraction on the input audio samples and face reference samples respectively, and performing reverse denoising on the noisy image sequence after forward noise addition by combining the extracted audio sample features and face reference sample features to obtain a video frame sequence; constructing a loss function based on the video frame sequence and the lip synchronization video samples, and obtaining a lip synchronization video generation model for generating lip synchronization videos based on the convergence of the loss function.
[0007] According to an audio-driven lip synchronization method provided by the present invention, the model to be trained includes a mask layer, an encoding layer, a feature extraction layer, a reference feature extraction layer, a forward diffusion layer, a first backward diffusion layer, a second backward diffusion layer and a decoding layer; audio samples, face reference samples and lip synchronization video samples are input into the model to be trained, and corresponding noise is added to each frame image in the lip synchronization video sample to obtain a noise sequence, and feature extraction is performed on the input audio samples and face reference samples respectively, so as to combine the extracted audio sample features and face reference sample features to perform the forward noise-added feature extraction on the model. The noisy image sequence is subjected to reverse denoising to obtain a video frame sequence, including: inputting the lip synchronization video sample into the mask layer to adaptively mask each frame image in the lip synchronization video sample, and obtaining an image mask sequence output by the mask layer; inputting the image mask sequence into the encoding layer for encoding, and obtaining an image encoding sequence output by the encoding layer; inputting the audio sample into the feature extraction layer for feature extraction, and obtaining the audio sample feature output by the feature extraction layer; inputting the human face reference sample into the reference feature extraction layer for feature extraction, and obtaining the human face reference sample output by the reference feature extraction layer. face reference sample features; inputting the image coding sequence into the forward diffusion layer to compress the image coding sequence into a low-dimensional latent space to obtain a low-dimensional coding feature sequence, and adding noise to each frame of the low-dimensional coding feature sequence to obtain a noise sequence; inputting the audio sample features, the face reference sample features and the noise sequence into the first back diffusion layer to perform preliminary denoising and alignment on the noise sequence based on the audio sample features and the face reference sample features, thereby obtaining a preliminary denoised and aligned image sequence output by the first back diffusion layer; inputting the audio sample features and the preliminary denoised and aligned image sequence into the second back diffusion layer to denoise the preliminary denoised and aligned image sequence based on the audio sample features, thereby obtaining a denoised image sequence output by the second back diffusion layer; combining the output of the second back diffusion layer with the face reference sample features and the audio sample features as the input of the first back diffusion layer, and combining the output of the first back diffusion layer with the audio sample features as the input of the second back diffusion layer, performing iterative denoising until a preset number of iterations is reached, and inputting the final denoised image sequence as the denoised video frame sequence into the decoding layer for decoding to obtain a lip synchronization video.
[0008] According to an audio-driven lip synchronization method provided by the present invention, the feature extraction layer includes a first extraction sublayer, a second extraction sublayer, a third extraction sublayer and a fourth extraction sublayer, and the face reference sample features include a first feature, a second feature, a third feature and a fourth feature; the face reference sample is input into the reference feature extraction layer for feature extraction to obtain the face reference sample features, including: inputting the face reference sample into the first extraction sublayer for feature extraction to obtain the first feature output by the first extraction sublayer; inputting the first feature into the second feature sublayer for feature extraction to obtain the second feature output by the second feature extraction sublayer; inputting the second feature into the third feature extraction sublayer for feature extraction to obtain the third feature output by the third feature extraction sublayer; inputting the third feature into the fourth feature extraction sublayer for feature extraction to obtain the fourth feature output by the fourth feature extraction sublayer.
[0009] According to an audio-driven lip synchronization method provided by the present invention, the first back diffusion layer includes a first residual network sublayer, a first self-attention sublayer, a first audio cross-attention sublayer and a first time attention sublayer; the audio sample features, the face reference sample features and the noise sequence are input into the first back diffusion layer to perform preliminary denoising and alignment on the noise sequence according to the audio sample features and the face reference sample features, so as to obtain a preliminary denoised and aligned image sequence output by the first back diffusion layer, including: inputting the noise sequence, the first feature and the audio sample features into the first residual network sublayer to perform feature extraction on the noise sequence, and fusing the audio sample features and the first feature to obtain a first denoised image sequence output by the first residual network sublayer; inputting the first denoised image sequence, the second feature and the audio sample features into the first self-attention sublayer to determine the first denoised image sequence based on the self-attention mechanism. The correlation between different frames is used to obtain the first self-attention weight, and combined with the audio sample features and the second feature, the second denoised image sequence output by the first self-attention sub-layer is obtained; the second denoised image sequence, the third feature and the audio sample feature are input into the first audio cross-attention sub-layer to determine the correlation between the audio sample features and the second denoised image sequence based on the audio cross-attention mechanism, to obtain the first audio cross-attention weight, and combined with the third feature, to obtain the third denoised image sequence output by the first audio cross-attention sub-layer; the third denoised image sequence, the fourth feature and the audio sample feature are input into the first time attention sub-layer to determine the temporal correlation between different frames in the third denoised image sequence based on the time attention mechanism, to determine the first time attention weight, and combined with the audio sample features and the fourth feature, to obtain the preliminary denoised and aligned image sequence output by the first time attention sub-layer.
[0010] According to an audio-driven lip synchronization method provided by the present invention, the second back diffusion layer includes a second temporal attention sublayer, a second audio cross-attention sublayer, a second self-attention sublayer, and a second residual network sublayer; the audio sample features and the preliminary denoised and aligned image sequence are input into the second back diffusion layer to denoise the preliminary denoised and aligned image sequence according to the audio sample features to obtain a denoised image sequence output by the second back diffusion layer, including: inputting the preliminary denoised and aligned image sequence and the audio sample features into the second temporal attention sublayer to determine the temporal correlation between different frames in the input preliminary denoised and aligned image sequence based on the temporal attention mechanism to obtain a second temporal attention weight, and fusing the audio sample features to obtain a fourth denoised image sequence output by the second temporal attention sublayer; The four denoised image sequences and audio sample features are input into the second audio cross-attention sub-layer to determine the correlation between the audio sample features and the fourth denoised image sequence based on the audio cross-attention mechanism, and align the audio sample features with the corresponding visual representation to obtain the fifth denoised image sequence; the fifth denoised image sequence and audio sample features are input into the second self-attention sub-layer to determine the correlation between different frames in the fifth denoised image sequence based on the self-attention mechanism, obtain the second self-attention weight, and combine with the audio sample features to obtain the sixth denoised image sequence output by the second self-attention sub-layer; the sixth denoised image sequence and audio sample features are input into the second residual network sub-layer to perform feature extraction on the sixth denoised image sequence, and fuse the audio sample features to obtain the denoised image sequence output by the second residual network sub-layer.
[0011] According to an audio-driven lip synchronization method provided by the present invention, each frame image in the lip synchronization video sample is adaptively masked, including: when it is determined that there is a video frame with head deflection in the lip synchronization video sample, selecting a corresponding preset head posture mask; when it is determined that there is a video frame with facial occlusion in the lip synchronization video sample, selecting a corresponding preset perceptual occlusion mask; when it is determined that there is a video frame in the lip synchronization video sample whose voice intensity change compared to the previous video frame is greater than a first preset threshold, selecting a corresponding preset audio energy driving mask; when it is determined that there is a video frame in the lip synchronization video sample whose expression change degree compared to the previous video frame is greater than a second preset threshold, selecting a corresponding preset expression mask; based on the selected mask, temporally smoothing the boundary value of the selected mask, and using the processed mask to adaptively mask the corresponding frame image in the lip synchronization video sample.
[0012] The present invention also provides an audio-driven lip synchronization device, comprising: a data acquisition module, which acquires a real face image, audio data and a face reference image; a video generation module, which inputs the real face image, audio data and the face reference image into a lip synchronization video generation model to obtain a lip synchronization video output by the lip synchronization video generation model; wherein the lip synchronization video generation model is trained based on audio samples, face reference samples and corresponding lip synchronization video samples, and the lip synchronization video generation model is obtained by forward denoising the input lip synchronization video samples, and performing reverse denoising training on the noise sequence obtained by the forward denoising in combination with the audio samples and the face reference samples.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the audio-driven lip synchronization method described above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements any of the above-described audio-driven lip synchronization methods.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned audio-driven lip synchronization methods.
[0016] The audio-driven lip synchronization method, device, equipment and medium provided by the present invention positively add noise to real lip synchronization video samples and use them as a training basis, so that the finally generated lip synchronization video can maintain the facial features, expressions and identity of the original character, greatly enhancing the realism and personalization of the video, and using facial reference samples and audio samples as conditions to guide the denoising process, ensuring that during the denoising process, the audio samples provide a basis for driving lip shape changes, and the facial reference samples provide more specific style, lighting or posture guidance, which helps to generate videos that better meet specific requirements, ensuring that the finally generated video is not only lip-synchronized, but also maintains the specific character identity and visual style represented by the reference image, and has strong generalization ability for wild videos and animated characters, so that the lip synchronization video generation model generates lip synchronization videos with precise synchronization of lip movements, rich facial details and natural overall visual effects based on the acquired real face images, audio data and face reference images. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 1 is a flow chart of the audio-driven lip synchronization method provided by the present invention; Figure 2 Schematic diagram of the architecture of the lip-sync video generation model provided by the present invention; Figure 3 1 is a schematic structural diagram of an audio-driven lip synchronization device provided by the present invention; Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0020] Figure 1 FIG is a flow chart of the audio-driven lip synchronization method provided by the present invention, such as Figure 1 As shown, the method includes: S11, obtaining a real face image, audio data, and a reference face image; S12, inputting the real face image, audio data and face reference image into the lip-sync video generation model to obtain the lip-sync video output by the lip-sync video generation model; wherein, the lip-sync video generation model is trained based on the audio samples, face reference samples and corresponding lip-sync video samples, and the lip-sync video generation model is obtained by forward denoising the input lip-sync video samples, and combining the audio samples and face reference samples to perform reverse denoising training on the noise sequence obtained by the forward denoising.
[0021] It should be noted that the step numbers "S1N" in this manual do not represent the order of the audio-driven lip synchronization method. Figure 2 The audio driven lip synchronization method of the present invention is described.
[0022] Step S11, obtaining a real face image, audio data and a reference face image.
[0023] It should be noted that the audio data is used to represent the voice information of the lip-synced video to be generated; the facial reference image can be selected according to the actual video generation requirements. For example, when it is necessary to strictly maintain the consistency of the character's identity, the facial reference image and the real face image need to come from the same person; when only lip movements that conform to the general audio need to be produced, the facial reference image does not need to be bound to a specific person and can be any face, such as a standardized neutral expression face, and no further restrictions are made here.
[0024] Step S12: input the real face image, audio data and face reference image into the lip-sync video generation model to obtain the lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained based on the audio samples, face reference samples and corresponding lip-sync video samples, and the lip-sync video generation model is obtained by forward denoising the input lip-sync video samples, and performing reverse denoising training on the noise sequence obtained by the forward denoising in combination with the audio samples and face reference samples.
[0025] In an optional embodiment, before inputting the real face image, audio data and face reference image into the lip-sync video generation model, it includes: obtaining audio samples, face reference samples and corresponding lip-sync video samples; inputting the audio samples, face reference samples and lip-sync video samples into the model to be trained, adding corresponding noise to each frame image in the lip-sync video sample to obtain a noise sequence, performing feature extraction on the input audio samples and face reference samples respectively, and performing reverse denoising on the noisy image sequence after forward noisy addition by combining the extracted audio sample features and face reference sample features to obtain a video frame sequence; constructing a loss function according to the video frame sequence and the lip-sync video samples, and obtaining a lip-sync video generation model for generating lip-sync videos based on the convergence of the loss function.
[0026] It should be noted that when adding noise to each frame of the lip-sync video sample, the noise to be added can be determined for each frame according to a predefined noise schedule, and then added to the corresponding frame according to the noise schedule. The predefined noise schedule can be set according to actual design requirements, such as linear, exponential, or cosine scheduling, and is not further limited here. In addition, by gradually and controllably adding noise to each frame of the lip-sync video sample along a predetermined timeline to generate a sequence from clear to noisy, the model can learn how to remove noise by combining the audio samples and facial reference samples with the noisy sequence, thereby helping the model generate lip-synced video.
[0027] Specifically, refer to Figure 2The model to be trained includes a mask layer, a coding layer, a feature extraction layer, a reference feature extraction layer, a forward diffusion layer, a first reverse diffusion layer, a second reverse diffusion layer and a decoding layer; the audio samples, the face reference samples and the lip-sync video samples are input into the model to be trained, so as to add corresponding noise to each frame image in the lip-sync video sample to obtain a noise sequence, and feature extraction is performed on the input audio samples and the face reference samples respectively, so as to combine the extracted audio sample features and the face reference sample features to perform reverse denoising on the noisy image sequence after forward noisy addition to obtain a video frame sequence, including: inputting the lip-sync video samples into the mask layer to adaptively mask each frame image in the lip-sync video sample to obtain an image mask sequence output by the mask layer; inputting the image mask sequence into the coding layer for encoding to obtain an image coding sequence output by the coding layer; inputting the audio samples into the feature extraction layer for feature extraction to obtain the audio sample features output by the feature extraction layer; inputting the face reference samples into the reference feature extraction layer for feature extraction to obtain the face reference sample features output by the reference feature extraction layer; and inputting the image The coding sequence is input into a forward diffusion layer to compress the image coding sequence into a low-dimensional latent space to obtain a low-dimensional coding feature sequence, and noise is added to each frame of the low-dimensional coding feature sequence to obtain a noise sequence. The audio sample features, the face reference sample features, and the noise sequence are input into a first back diffusion layer to perform preliminary denoising and alignment on the noise sequence based on the audio sample features and the face reference sample features, thereby obtaining a preliminary denoised and aligned image sequence output by the first back diffusion layer. The audio sample features and the preliminary denoised and aligned image sequence are input into a second back diffusion layer to denoise the preliminary denoised and aligned image sequence based on the audio sample features, thereby obtaining a denoised image sequence output by the second back diffusion layer. The output of the second back diffusion layer is combined with the face reference sample features and the audio sample features as input to the first back diffusion layer, and the output of the first back diffusion layer is combined with the audio sample features as input to the second back diffusion layer for iterative denoising until a preset number of iterations is reached. The final denoised image sequence is input into a decoding layer as a denoised video frame sequence for decoding to obtain a lip-synced video.
[0028] Furthermore, the feature extraction layer includes a first extraction sublayer, a second extraction sublayer, a third extraction sublayer and a fourth extraction sublayer, and the face reference sample features include a first feature, a second feature, a third feature and a fourth feature; the face reference sample is input into the reference feature extraction layer for feature extraction to obtain the face reference sample features, including: inputting the face reference sample into the first extraction sublayer for feature extraction to obtain the first feature output by the first extraction sublayer; inputting the first feature into the second feature sublayer for feature extraction to obtain the second feature output by the second feature extraction sublayer; inputting the second feature into the third feature extraction sublayer for feature extraction to obtain the third feature output by the third feature extraction sublayer; inputting the third feature into the fourth feature extraction sublayer for feature extraction to obtain the fourth feature output by the fourth feature extraction sublayer.
[0029] Correspondingly, the first back diffusion layer includes a first residual network sublayer, a first self-attention sublayer, a first audio cross-attention sublayer and a first temporal attention sublayer; the audio sample features, the face reference sample features and the noise sequence are input into the first back diffusion layer to perform preliminary denoising and alignment on the noise sequence according to the audio sample features and the face reference sample features, thereby obtaining a preliminary denoised and aligned image sequence output by the first back diffusion layer, including: inputting the noise sequence, the first feature and the audio sample feature into the first residual network sublayer to perform feature extraction on the noise sequence, and fusing the audio sample feature and the first feature to obtain a first denoised image sequence output by the first residual network sublayer; inputting the first denoised image sequence, the second feature and the audio sample feature into the first self-attention sublayer to determine the correlation between different frames in the first denoised image sequence based on the self-attention mechanism. , obtain the first self-attention weight, and combine the audio sample features and the second features to obtain the second denoised image sequence output by the first self-attention sub-layer; input the second denoised image sequence, the third feature and the audio sample features into the first audio cross-attention sub-layer to determine the correlation between the audio sample features and the second denoised image sequence based on the audio cross-attention mechanism, obtain the first audio cross-attention weight, and combine the third feature to obtain the third denoised image sequence output by the first audio cross-attention sub-layer; input the third denoised image sequence, the fourth feature and the audio sample features into the first temporal attention sub-layer to determine the temporal correlation between different frames in the third denoised image sequence based on the temporal attention mechanism, determine the first temporal attention weight, and combine the audio sample features and the fourth feature to obtain the preliminary denoised and aligned image sequence output by the first temporal attention sub-layer.
[0030] It should be noted that the first residual network sublayer usually contains multiple convolutional layers, normalization layers, and activation functions. During the processing, the first feature and the audio sample feature are used as conditional information, and a specific fusion method, such as 1×1 convolution, channel splicing and convolution, is used to guide the first residual network sublayer to learn features that conform to the first feature and the audio sample feature.
[0031] The first self-attention sublayer calculates the correlation between different frames in the first denoised image sequence to capture the spatial / temporal dependencies within the image or sequence. After obtaining the first self-attention weight, the audio sample features and the second features are fused to guide the self-attention mechanism to focus on the areas or time points related to the reference face and audio.
[0032] The first audio cross-attention sublayer calculates the correlation between the audio sample features and the second denoised image sequence by taking the second denoised image sequence as the query and the audio sample features as the key and value, so as to align the visual representation with the audio content (such as phonemes and pitch) to facilitate lip synchronization. After calculating the first audio cross-attention weight, the third feature is fused as conditional information to further ensure that the alignment result conforms to the lip shape and facial features of the reference face.
[0033] The first temporal attention sublayer determines the temporal correlation between different frames in the third denoised image sequence to capture the temporal dynamics and consistency in the video sequence. After calculating the first temporal attention weight, the audio sample features and the fourth feature conditional information are fused. The fourth feature can help keep the temporal dynamics consistent with the natural movement of the reference face, and the audio sample features can help align the temporal dynamics with the audio rhythm.
[0034] In addition, the second back diffusion layer includes a second temporal attention sublayer, a second audio cross attention sublayer, a second self-attention sublayer and a second residual network sublayer; the audio sample features and the preliminary denoised and aligned image sequence are input into the second back diffusion layer to denoise the preliminary denoised and aligned image sequence according to the audio sample features to obtain a denoised image sequence output by the second back diffusion layer, including: inputting the preliminary denoised and aligned image sequence and the audio sample features into the second temporal attention sublayer to determine the temporal correlation between different frames in the input preliminary denoised and aligned image sequence based on the temporal attention mechanism to obtain a second temporal attention weight, and fusing the audio sample features to obtain a fourth denoised image sequence output by the second temporal attention sublayer; the fourth denoised image sequence and the audio The sample features are input into the second audio cross-attention sublayer to determine the correlation between the audio sample features and the fourth denoised image sequence based on the audio cross-attention mechanism, and the audio sample features are aligned with the corresponding visual representation to obtain the fifth denoised image sequence; the fifth denoised image sequence and the audio sample features are input into the second self-attention sublayer to determine the correlation between different frames in the fifth denoised image sequence based on the self-attention mechanism, obtain the second self-attention weight, and combine the audio sample features to obtain the sixth denoised image sequence output by the second self-attention sublayer; the sixth denoised image sequence and the audio sample features are input into the second residual network sublayer to perform feature extraction on the sixth denoised image sequence, and fuse the audio sample features to obtain the denoised image sequence output by the second residual network sublayer.
[0035] It should be noted that since the preliminary denoised and aligned image sequence is the output of the first back diffusion layer, the second temporal attention sublayer is used to calculate the correlation between the preliminary denoised and aligned image sequence at different time steps to more finely adjust the temporal consistency, and use the output of the first back diffusion layer as a stronger context. After obtaining the second temporal attention weight, the audio sample features are fused as conditional information to further ensure that the temporal dynamics are aligned with the audio rhythm.
[0036] Since the fourth denoised image sequence has been processed by the temporal attention of the first and second back-diffusion layers, the correlation between the audio sample features and the fourth denoised image sequence is determined by the second audio cross-attention sub-layer to more accurately align the visual representation (especially the lip shape) with the audio content.
[0037] Since the fifth denoised image sequence already contains more refined temporal consistency and audio alignment information, the correlation between different frames in the fifth denoised image sequence is calculated through the second self-attention sub-layer, so that the self-attention mechanism can further capture and enhance the details and spatial / temporal dependencies within the image or sequence.
[0038] In addition, the second residual network sublayer usually contains multiple convolutional layers, normalization layers, and activation functions to perform the final feature conversion and fusion.
[0039] It should be added that the loss function is expressed as: Among them, L represents a function; Represents the potential noise representation corresponding to the first frame of the image coding sequence; , Represents the t-th frame image x in the image coding sequence t The potential noise is represented by n, which represents the noise variable. , N represents normal distribution, represents the noise level, i.e. the mean of the normal distribution, and t represents the time step; represents weight; Represents the features of the reference face sample; Represents audio sample features; Represents an image coding sequence; 、 、 and represents the preprocessing hyperparameters; The network representing the model to be trained; represents the model parameters, express The denoiser is trained using the Denoising Score Matching (DSM) objective. It combines features from reference facial samples, audio samples, and image encoding sequences to avoid the need for multi-stage training or additional intermediate representations. While ensuring broad generalization, it fully leverages the universal visual priors of lip-sync video sample diffusion to directly drive lip synchronization, enabling the generation of talking heads across various styles. This approach inherits the structure and parameters of the pre-trained stable video diffusion model as much as possible. Furthermore, by directly adopting the DSM objective, the need for adversarial loss or oral expert supervision is eliminated, effectively balancing different conditions in the latent space. While preserving identity, it achieves natural lip movement and consistent tooth rendering, as well as greater motion dynamics. Through a unified conditional fusion scheme, it effectively balances audio and visual conditions and generates high-quality results without the need for additional supervisory signals, enabling the model to precisely control the generation of appearance, motion, and specific regions without the need for additional supervisory signals or intermediate representations.
[0040] In addition, the reference feature extraction layer, ID-Guider, is an efficient encoding module composed of convolutional layers. Before the face reference sample is input to the reference feature extraction layer, the image encoding sequence is concatenated with the face reference sample as a single channel. It is then processed through a simple pure convolutional downsampler with a channel dimension of [64, 32, 128, 64] to align the input with the denoising UNet, thereby facilitating identity preservation through the face reference sample features during the subsequent denoising process. Furthermore, compared to ReferenceNet, computationally intensive modules such as 3D convolutions are removed, and only 2D ResBlock modules consistent with stable video diffusion are retained to ensure proper residual integration into the denoising UNet. The upsampling layers are also removed. By removing timestep-dependent modules, ID-Guider does not need to be recomputed during inference. As a result, ID-Guider retains only 98M parameters, a parameter reduction of over 90%, significantly improving computational efficiency and reducing the influence of visual information in the face reference sample, while maintaining strong identity preservation capabilities.
[0041] The feature extraction layer uses Whisper as a feature extractor to compare each video frame with a window (x t-k ,...,x t ,...,x t+k ) to capture the temporal context and enhance the control of motion generation by weakly correlated audio sample features, where x t Represents the audio sample features of step t, k determines the temporal context range, and zero padding is applied to the boundary frames without additional processing to enhance the influence of the audio signal, thereby facilitating the spatial configuration of the lips driven by the audio sample features in the subsequent denoising process.
[0042] In an optional embodiment, each frame image in the synchronized video sample is adaptively masked, including: when it is determined that there is a video frame with head deflection in the lip-sync video sample, selecting a corresponding preset head posture mask; when it is determined that there is a video frame with face occlusion in the lip-sync video sample, selecting a corresponding preset perceptual occlusion mask; when it is determined that there is a video frame in the lip-sync video sample whose voice intensity change compared to the previous video frame is greater than a first preset threshold, selecting a corresponding preset audio energy drive mask; when it is determined that there is a video frame in the lip-sync video sample whose expression change degree compared to the previous video frame is greater than a second preset threshold, selecting a corresponding preset expression mask; based on the selected mask, temporally smoothing the boundary value of the selected mask, and using the processed mask to adaptively mask the corresponding frame image in the lip-sync video sample.
[0043] It should be added that the lip editing area is adjusted through adaptive masking and the lip movement is eliminated from the dependence of the surrounding pixel movement pattern. Each mask can be set according to the actual design requirements. For example, the occlusion mask can be selected according to the corresponding shape of the part to be occluded. For example, when occluding the lip posture, a rectangular mask can be selected to avoid the potential guidance of the mask shape, ensure that the pixels around the lips are masked, and eliminate the influence of the jaw muscle movement pattern on the lip movement. In addition, the temporal smoothing process is expressed as: in, Represents the boundary coordinates of the mask after temporal smoothing; Indicates the boundary coordinates of the mask corresponding to the t-th frame video image, Indicates the horizontal (width) boundary value of the mask (such as the left or right boundary coordinates). Indicates the vertical (height) boundary value of the mask (such as the upper or lower boundary coordinates); Represents the smoothing parameter. The specific value can be set according to prior experience and experimental results. For example, it can be 0.75. No further limitation is made here. A balance is provided between temporal consistency and local accuracy to ensure that temporal smoothing can enhance the stability of mask motion and further prevent the model from inferring lip motion based on landmark motion patterns.
[0044] In summary, the embodiment of the present invention positively adds noise to real lip-sync video samples and uses them as a training basis, so that the finally generated lip-sync video can maintain the facial features, expressions and identity of the original character, greatly enhancing the realism and personalization of the video, and uses facial reference samples and audio samples as conditions to guide the denoising process, ensuring that during the denoising process, the audio samples provide a basis for driving lip shape changes, and the facial reference samples provide more specific style, lighting or posture guidance, which helps to generate videos that better meet specific requirements, ensuring that the finally generated video is not only lip-synced, but also maintains the specific character identity and visual style represented by the reference image, and has strong generalization capabilities for wild videos and animated characters, so that the lip-sync video generation model generates lip-sync videos with precise synchronization of lip movements, rich facial details, and natural overall visual effects based on the acquired real face images, audio data and face reference images.
[0045] The audio-driven lip synchronization device provided by the present invention is described below. The audio-driven lip synchronization device described below and the audio-driven lip synchronization method described above can be referred to in correspondence with each other.
[0046] Figure 3A schematic structural diagram of an audio-driven lip synchronization device is shown, the device comprising: The data acquisition module 31 acquires a real face image, audio data and a reference face image; The video generation module 32 inputs the lip-sync video samples, audio samples and face reference samples into the lip-sync video generation model to obtain the lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained based on the lip-sync video samples, audio samples, face reference samples and corresponding lip-sync video samples; the lip-sync video generation model is used to perform forward noise addition on the input lip-sync video samples, and reverse denoise training on the noise sequence obtained by the forward noise addition in combination with the audio samples and face reference samples.
[0047] In an optional embodiment, the device further includes: a sample acquisition module, which acquires audio samples, facial reference samples and corresponding lip-sync video samples before inputting the real face image, audio data and facial reference image into the lip-sync video generation model; a training module, which inputs the audio samples, facial reference samples and lip-sync video samples into the model to be trained, so as to add corresponding noise to each frame image in the lip-sync video sample to obtain a noise sequence, and performs feature extraction on the input audio samples and facial reference samples respectively, and performs reverse denoising on the noisy image sequence after forward noise addition in combination with the extracted audio sample features and facial reference sample features to obtain a video frame sequence; constructs a loss function according to the video frame sequence and the lip-sync video samples, and obtains a lip-sync video generation model for generating lip-sync videos based on the convergence of the loss function.
[0048] Specifically, the model to be trained includes a mask layer, a coding layer, a feature extraction layer, a reference feature extraction layer, a forward diffusion layer, a first reverse diffusion layer, a second reverse diffusion layer and a decoding layer; the training module includes: a mask unit, which inputs the lip synchronization video sample into the mask layer to adaptively mask each frame image in the lip synchronization video sample to obtain an image mask sequence output by the mask layer; an encoding unit, which inputs the image mask sequence into the coding layer to encode it separately to obtain an image coding sequence output by the coding layer; a feature extraction unit, which inputs the audio sample into the feature extraction layer to extract features to obtain audio sample features output by the feature extraction layer; a reference feature extraction unit, which inputs the face reference sample into the reference feature extraction layer to extract features to obtain face reference sample features output by the reference feature extraction layer; a forward diffusion unit, which inputs the image coding sequence into the forward diffusion layer to compress the image coding sequence into a low-dimensional latent space to obtain a low-dimensional coding feature sequence, and adds noise to each frame of the low-dimensional coding feature sequence. , obtaining a noise sequence; a first back diffusion unit, inputting the audio sample features, the face reference sample features, and the noise sequence into the first back diffusion layer, performing preliminary denoising and alignment on the noise sequence according to the audio sample features and the face reference sample features, and obtaining a preliminary denoised and aligned image sequence output by the first back diffusion layer; a second back diffusion unit, inputting the audio sample features and the preliminary denoised and aligned image sequence into the second back diffusion layer, performing denoising on the preliminary denoised and aligned image sequence according to the audio sample features, and obtaining a denoised image sequence output by the second back diffusion layer; an iterative denoising unit, combining the output of the second back diffusion layer with the face reference sample features and the audio sample features as input to the first back diffusion layer, and combining the output of the first back diffusion layer with the audio sample features as input to the second back diffusion layer, and performing iterative denoising until a preset number of iterations is reached; a decoding unit, inputting the final denoised image sequence as a denoised video frame sequence into the decoding layer for decoding, and obtaining a lip-synced video.
[0049] Furthermore, the feature extraction layer includes a first extraction sublayer, a second extraction sublayer, a third extraction sublayer and a fourth extraction sublayer, and the face reference sample features include a first feature, a second feature, a third feature and a fourth feature; the reference feature extraction unit includes: a first extraction sublayer, which inputs the face reference sample into the first extraction sublayer for feature extraction, and obtains the first feature output by the first extraction sublayer; a second extraction sublayer, which inputs the first feature into the second feature sublayer for feature extraction, and obtains the second feature output by the second feature extraction sublayer; a third extraction sublayer, which inputs the second feature into the third feature extraction sublayer for feature extraction, and obtains the third feature output by the third feature extraction sublayer; and a fourth extraction subunit, which inputs the third feature into the fourth feature extraction sublayer for feature extraction, and obtains the fourth feature output by the fourth feature extraction sublayer.
[0050] Correspondingly, the first back diffusion layer includes a first residual network sublayer, a first self-attention sublayer, a first audio cross-attention sublayer and a first time attention sublayer; the first back diffusion unit includes: a first residual network sublayer, which inputs the noise sequence, the first feature and the audio sample feature into the first residual network sublayer to extract features of the noise sequence, and fuses the audio sample feature and the first feature to obtain a first denoised image sequence output by the first residual network sublayer; a first self-attention sublayer, which inputs the first denoised image sequence, the second feature and the audio sample feature into the first self-attention sublayer to determine the correlation between different frames in the first denoised image sequence based on the self-attention mechanism, obtain a first self-attention weight, and combine the audio sample feature and the second feature to obtain a second denoised image sequence output by the first self-attention sublayer. image sequence; a first audio cross-attention sub-unit, inputting the second denoised image sequence, the third feature and the audio sample feature into the first audio cross-attention sub-layer to determine the correlation between the audio sample feature and the second denoised image sequence based on the audio cross-attention mechanism, obtain the first audio cross-attention weight, and combine the third feature to obtain the third denoised image sequence output by the first audio cross-attention sub-layer; a first time attention sub-unit, inputting the third denoised image sequence, the fourth feature and the audio sample feature into the first time attention sub-layer to determine the temporal correlation between different frames in the third denoised image sequence based on the time attention mechanism, determine the first time attention weight, and combine the audio sample feature and the fourth feature to obtain the preliminary denoised and aligned image sequence output by the first time attention sub-layer.
[0051] In addition, the second reverse diffusion unit includes: a second time attention sub-unit, which inputs the preliminary denoised and aligned image sequence and audio sample features into the second time attention sub-layer, and determines the temporal correlation between different frames in the input preliminary denoised and aligned image sequence based on the time attention mechanism, obtains a second time attention weight, and fuses the audio sample features to obtain a fourth denoised image sequence output by the second time attention sub-layer; a second audio cross attention sub-unit, which inputs the fourth denoised image sequence and audio sample features into the second audio cross attention sub-layer, and determines the correlation between the audio sample features and the fourth denoised image sequence based on the audio cross attention mechanism. The audio sample features are aligned with the corresponding visual representations to obtain a fifth denoised image sequence; the second self-attention sub-unit inputs the fifth denoised image sequence and the audio sample features into the second self-attention sub-layer to determine the correlation between different frames in the fifth denoised image sequence based on the self-attention mechanism, obtain the second self-attention weight, and combine the audio sample features to obtain the sixth denoised image sequence output by the second self-attention sub-layer; the second residual network sub-unit inputs the sixth denoised image sequence and the audio sample features into the second residual network sub-layer to perform feature extraction on the sixth denoised image sequence, and fuse the audio sample features to obtain the denoised image sequence output by the second residual network sub-layer.
[0052] In an optional embodiment, the mask unit includes: when determining that there are video frames with head deflection in the lip-sync video sample, selecting a corresponding preset head posture mask; when determining that there are video frames with facial occlusion in the lip-sync video sample, selecting a corresponding preset perceptual occlusion mask; when determining that there are video frames in the lip-sync video sample whose voice intensity change compared to the previous video frame is greater than a first preset threshold, selecting a corresponding preset audio energy drive mask; when determining that there are video frames in the lip-sync video sample whose expression change degree compared to the previous video frame is greater than a second preset threshold, selecting a corresponding preset expression mask; according to the selected mask, temporally smoothing the boundary value of the selected mask, and using the processed mask to adaptively mask the corresponding frame image in the lip-sync video sample.
[0053] In summary, the embodiment of the present invention positively adds noise to real lip-sync video samples and uses them as a training basis, so that the finally generated lip-sync video can maintain the facial features, expressions and identity of the original character, greatly enhancing the realism and personalization of the video, and uses facial reference samples and audio samples as conditions to guide the denoising process, ensuring that during the denoising process, the audio samples provide a basis for driving lip shape changes, and the facial reference samples provide more specific style, lighting or posture guidance, which helps to generate videos that better meet specific requirements, ensuring that the finally generated video is not only lip-synced, but also maintains the specific character identity and visual style represented by the reference image, and has strong generalization capabilities for wild videos and animated characters, so that the lip-sync video generation model generates lip-sync videos with precise synchronization of lip movements, rich facial details, and natural overall visual effects based on the acquired real face images, audio data and face reference images.
[0054] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440. The processor 410, the communications interface 420, and the memory 430 communicate with each other via the communications bus 440. The processor 410 may call logic instructions in the memory 430 to execute an audio-driven lip synchronization method, which includes: obtaining a real face image, audio data, and a reference face image; inputting the real face image, audio data, and reference face image into a lip synchronization video generation model to obtain a lip synchronization video output by the lip synchronization video generation model; wherein the lip synchronization video generation model is trained based on audio samples, reference face samples, and corresponding lip synchronization video samples; and the lip synchronization video generation model is trained based on forward denoising of the input lip synchronization video samples and reverse denoising of the noise sequence obtained by the forward denoising in combination with the audio samples and reference face samples.
[0055] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0056] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the audio-driven lip synchronization method provided by the above methods, which includes: obtaining a real face image, audio data and a face reference image; inputting the real face image, audio data and the face reference image into a lip synchronization video generation model to obtain a lip synchronization video output by the lip synchronization video generation model; wherein the lip synchronization video generation model is trained based on audio samples, face reference samples and corresponding lip synchronization video samples, and the lip synchronization video generation model is obtained by forward denoising the input lip synchronization video samples, and combining the audio samples and face reference samples to reverse denoise the noise sequence obtained by the forward denoising.
[0057] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the audio-driven lip synchronization method provided by the above-mentioned methods, the method comprising: obtaining a real face image, audio data and a face reference image; inputting the real face image, audio data and the face reference image into a lip synchronization video generation model to obtain a lip synchronization video output by the lip synchronization video generation model; wherein the lip synchronization video generation model is trained based on audio samples, face reference samples and corresponding lip synchronization video samples, and the lip synchronization video generation model is obtained by forward denoising the input lip synchronization video samples, and performing reverse denoising training on the noise sequence obtained by the forward denoising in combination with the audio samples and the face reference samples.
[0058] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0059] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An audio-driven lip synchronization method, characterized in that: include: Obtain real face images, audio data, and face reference images; The real face image, the audio data and the face reference image are input into a lip-sync video generation model to obtain a lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained based on audio samples, face reference samples and corresponding lip-sync video samples, and the lip-sync video generation model is obtained by forward denoising the input lip-sync video samples and performing reverse denoising training on the noise sequence obtained by forward denoising in combination with the audio samples and the face reference samples.
2. The audio-driven lip synchronization method according to claim 1, wherein: Before inputting the real face image, the audio data, and the reference face image into the lip-sync video generation model, the method includes: Obtain audio samples, face reference samples, and corresponding lip-sync video samples; Inputting the audio sample, the facial reference sample, and the lip-sync video sample into the to-be-trained model, adding corresponding noise to each frame of the lip-sync video sample to obtain a noise sequence, performing feature extraction on the input audio sample and the facial reference sample, and performing reverse denoising on the noisy image sequence after forward noisy addition by combining the extracted audio sample features and facial reference sample features to obtain a video frame sequence; A loss function is constructed according to the video frame sequence and the lip synchronization video sample, and a lip synchronization video generation model for generating a lip synchronization video is obtained based on the convergence of the loss function.
3. The audio-driven lip synchronization method according to claim 2, wherein: The model to be trained includes a mask layer, an encoding layer, a feature extraction layer, a reference feature extraction layer, a forward diffusion layer, a first reverse diffusion layer, a second reverse diffusion layer, and a decoding layer; the audio sample, the face reference sample, and the lip-sync video sample are input into the model to be trained, and corresponding noise is added to each frame image in the lip-sync video sample to obtain a noise sequence; feature extraction is performed on the input audio sample and the face reference sample respectively, and reverse denoising is performed on the noisy image sequence after forward noise addition by combining the extracted audio sample features and face reference sample features to obtain a video frame sequence, including: Inputting the lip-sync video sample into the mask layer to adaptively mask each frame of the lip-sync video sample, thereby obtaining an image mask sequence output by the mask layer; Inputting the image mask sequence into the coding layer for encoding respectively, to obtain the image coding sequence output by the coding layer; Inputting the audio sample into the feature extraction layer for feature extraction, and obtaining audio sample features output by the feature extraction layer; Inputting the human face reference sample into the reference feature extraction layer for feature extraction, and obtaining the human face reference sample features output by the reference feature extraction layer; Inputting the image coding sequence into the forward diffusion layer to compress the image coding sequence into a low-dimensional latent space to obtain a low-dimensional coding feature sequence, and adding noise to each frame of the low-dimensional coding feature sequence to obtain a noise sequence; Inputting the audio sample features, the facial reference sample features, and the noise sequence into the first back diffusion layer to perform preliminary denoising and alignment on the noise sequence based on the audio sample features and the facial reference sample features, thereby obtaining a preliminary denoised and aligned image sequence output by the first back diffusion layer; Inputting the audio sample features and the preliminary denoised and aligned image sequence into the second back diffusion layer to denoise the preliminary denoised and aligned image sequence according to the audio sample features, thereby obtaining a denoised image sequence output by the second back diffusion layer; The output of the second back diffusion layer is combined with the facial reference sample features and the audio sample features as the input of the first back diffusion layer, and the output of the first back diffusion layer is combined with the audio sample features as the input of the second back diffusion layer. Iterative denoising is performed until a preset number of iterations is reached. The resulting denoised image sequence is input as a denoised video frame sequence to the decoding layer for decoding to obtain a lip-synced video.
4. The audio-driven lip synchronization method according to claim 3, wherein: The feature extraction layer includes a first extraction sublayer, a second extraction sublayer, a third extraction sublayer and a fourth extraction sublayer, and the face reference sample features include a first feature, a second feature, a third feature and a fourth feature; Inputting the human face reference sample into the reference feature extraction layer to perform feature extraction to obtain human face reference sample features, including: Inputting the face reference sample into the first extraction sublayer to perform feature extraction, thereby obtaining a first feature output by the first extraction sublayer; Inputting the first feature into the second feature sublayer for feature extraction, and obtaining a second feature output by the second feature extraction sublayer; Inputting the second feature into the third feature extraction sublayer for feature extraction, to obtain a third feature output by the third feature extraction sublayer; The third feature is input into the fourth feature extraction sublayer for feature extraction, to obtain a fourth feature output by the fourth feature extraction sublayer.
5. The audio-driven lip synchronization method according to claim 4, wherein: The first back diffusion layer includes a first residual network sublayer, a first self-attention sublayer, a first audio cross-attention sublayer, and a first temporal attention sublayer; the audio sample features, the face reference sample features, and the noise sequence are input into the first back diffusion layer to perform preliminary denoising and alignment on the noise sequence based on the audio sample features and the face reference sample features, thereby obtaining a preliminary denoised and aligned image sequence output by the first back diffusion layer, including: Inputting the noise sequence, the first feature, and the audio sample feature into the first residual network sublayer to extract features from the noise sequence, and fusing the audio sample feature with the first feature to obtain a first denoised image sequence output by the first residual network sublayer; Inputting the first denoised image sequence, the second feature, and the audio sample feature into the first self-attention sub-layer to determine the correlation between different frames in the first denoised image sequence based on the self-attention mechanism, obtain first self-attention weights, and combining the audio sample feature and the second feature to obtain a second denoised image sequence output by the first self-attention sub-layer; Inputting the second denoised image sequence, the third feature, and the audio sample feature into the first audio cross-attention sub-layer to determine the correlation between the audio sample feature and the second denoised image sequence based on the audio cross-attention mechanism, obtain a first audio cross-attention weight, and combining the third feature to obtain a third denoised image sequence output by the first audio cross-attention sub-layer; The third denoised image sequence, the fourth feature and the audio sample feature are input into the first temporal attention sublayer to determine the temporal correlation between different frames in the third denoised image sequence based on the temporal attention mechanism, determine the first temporal attention weight, and combine the audio sample feature and the fourth feature to obtain a preliminary denoised and aligned image sequence output by the first temporal attention sublayer.
6. The audio-driven lip synchronization method according to claim 3, wherein: The second back diffusion layer includes a second temporal attention sublayer, a second audio cross attention sublayer, a second self-attention sublayer, and a second residual network sublayer; the audio sample features and the preliminary denoised and aligned image sequence are input into the second back diffusion layer to denoise the preliminary denoised and aligned image sequence according to the audio sample features, thereby obtaining a denoised image sequence output by the second back diffusion layer, including: Inputting the preliminary denoised and aligned image sequence and the audio sample features into the second temporal attention sublayer, determining the temporal correlation between different frames in the input preliminary denoised and aligned image sequence based on the temporal attention mechanism, obtaining a second temporal attention weight, and fusing the audio sample features to obtain a fourth denoised image sequence output by the second temporal attention sublayer; Inputting the fourth denoised image sequence and the audio sample features into the second audio cross-attention sub-layer to determine the correlation between the audio sample features and the fourth denoised image sequence based on the audio cross-attention mechanism, and aligning the audio sample features with corresponding visual representations to obtain a fifth denoised image sequence; Inputting the fifth denoised image sequence and the audio sample features into the second self-attention sub-layer to determine the correlation between different frames in the fifth denoised image sequence based on the self-attention mechanism, obtain second self-attention weights, and combining them with the audio sample features to obtain a sixth denoised image sequence output by the second self-attention sub-layer; The sixth denoised image sequence and the audio sample features are input into the second residual network sublayer to perform feature extraction on the sixth denoised image sequence, and the audio sample features are fused to obtain a denoised image sequence output by the second residual network sublayer.
7. The audio-driven lip synchronization method according to claim 2, wherein: Adaptively masking each frame of the lip-sync video sample includes: When determining that a video frame with head deflection exists in the lip synchronization video sample, selecting a corresponding preset head posture mask; When determining that a video frame with facial occlusion exists in the lip-sync video sample, selecting a corresponding preset perceptual occlusion mask; When it is determined that a video frame in the lip-sync video sample has a speech intensity change greater than a first preset threshold compared to a previous video frame, selecting a corresponding preset audio energy driving mask; When it is determined that there is a video frame in the lip-sync video sample whose expression change degree compared with the previous video frame is greater than a second preset threshold, selecting a corresponding preset expression mask; According to the selected mask, a temporal smoothing process is performed on the boundary value of the selected mask, and the processed mask is used to adaptively mask the corresponding frame image in the lip synchronization video sample.
8. An audio-driven lip synchronization device, characterized in that: include: Data acquisition module, which obtains real face images, audio data and face reference images; The video generation module inputs the real face image, the audio data and the face reference image into a lip-sync video generation model to obtain a lip-sync video output by the lip-sync video generation model; wherein the lip-sync video generation model is trained based on audio samples, face reference samples and corresponding lip-sync video samples, and the lip-sync video generation model is obtained by forward denoising the input lip-sync video samples and performing reverse denoising training on the noise sequence obtained by the forward denoising in combination with the audio samples and the face reference samples.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the audio-driven lip synchronization method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the audio-driven lip synchronization method according to any one of claims 1 to 7 is implemented.