Lip drive method and apparatus, electronic device, and storage medium

By using a diffusion model and elliptic Gaussian spatiotemporal adaptive weighting technology, the influence of audio conditions on video latent vectors is dynamically adjusted, solving the problems of mouth region generation distortion and inaccurate synchronization in lip-driven videos, and achieving high-quality and efficient lip-driven video generation.

CN122134891APending Publication Date: 2026-06-02BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2026-02-26
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as distortion in the mouth area, unnatural lip closure, and low generation efficiency when generating lip-driven videos. This is especially true in e-commerce and virtual live-streaming sales scenarios, where it is difficult to achieve high-quality and efficient audio-visual synchronization.

Method used

By employing a diffusion model combined with elliptic Gaussian spatiotemporal adaptive weighting technology, the influence intensity of audio conditions on video latent vectors is dynamically adjusted. Through latent vector encoding, decoding, and transformation, lip-driven video synchronized with audio is generated.

Benefits of technology

It improves the generation quality and efficiency of lip-driven video, achieves precise synchronization of the mouth area, solves the problems of mouth corner adhesion and unnatural lip closure in traditional methods, and improves audio-visual synchronization accuracy and generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134891A_ABST
    Figure CN122134891A_ABST
Patent Text Reader

Abstract

This disclosure provides a lip-syncing method, apparatus, electronic device, and storage medium, relating to the field of artificial intelligence technology, particularly computer vision, augmented reality, virtual reality, and deep learning, and applicable to scenarios such as metaverse and virtual digital humans. The method includes: encoding a template video and target audio into a latent space to obtain a first latent vector and a second latent vector; dynamically adjusting the influence intensity of the second latent vector on the first latent vector according to an elliptic Gaussian spatiotemporal adaptive weight during denoising using a diffusion model to obtain a denoised latent vector; decoding and transforming the denoised latent vector to obtain a transformed image; and generating a lip-syncing video based on the transformed image and the target audio. This disclosure improves the generation quality and efficiency of lip-syncing video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of computer vision, augmented reality, virtual reality, and deep learning, and can be applied to scenarios such as metaverse and virtual digital humans. Specifically, it relates to a lip-syncing method, device, electronic device, and storage medium. Background Technology

[0002] With the rapid development of AIGC (Artificial Intelligence Generated Content) technology in the field of video generation, "fixed high-quality video templates + secondary lip-motion driving" has become the mainstream solution that balances generation efficiency, cost control, and commercial feasibility. This solution retains the high-quality image of the template video and only reconstructs or replaces the mouth area to synchronize it with the target speech, effectively reducing the randomness and computational consumption of end-to-end video generation. It is especially suitable for e-commerce, virtual anchor live streaming, and other scenarios. Summary of the Invention

[0003] This disclosure presents a lip-driven method, apparatus, electronic device, and storage medium.

[0004] According to a first aspect of this disclosure, a lip-sync driving method is provided, comprising: encoding a template video and a target audio into a latent space respectively to obtain a first latent vector and a second latent vector; during denoising using a diffusion model, dynamically adjusting the influence intensity of the second latent vector on the first latent vector according to an elliptic Gaussian spatiotemporal adaptive weight to obtain a denoised latent vector; decoding and transforming the denoised latent vector to obtain a transformed image; and generating a lip-sync driven video based on the transformed image and the target audio.

[0005] According to a second aspect of this disclosure, a lip-syncing device is provided, comprising: an encoding module configured to encode a template video and a target audio into a latent space respectively to obtain a first latent vector and a second latent vector; a denoising module configured to dynamically adjust the influence intensity of the second latent vector on the first latent vector according to an elliptic Gaussian spatiotemporal adaptive weight during denoising using a diffusion model to obtain a denoised latent vector; a decoding module configured to decode and transform the denoised latent vector to obtain a transformed image; and a generation module configured to generate a lip-syncing video based on the transformed image and the target audio.

[0006] According to a third aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect.

[0007] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform a method as described in any implementation of the first aspect.

[0008] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is an exemplary system architecture diagram to which this disclosure can be applied; Figure 2 This is a flowchart of a first embodiment of the lip-driving method according to the present disclosure; Figure 3 This is a flowchart of a second embodiment of the lip-driving method according to the present disclosure; Figure 4 This is a flowchart of a third embodiment of the lip-driving method according to the present disclosure; Figure 5 yes Figure 4 A flowchart of an embodiment of step 406; Figure 6 This is a flowchart of the fourth embodiment of the lip-driving method according to the present disclosure; Figure 7 This is a flowchart of the fifth embodiment of the lip-driving method according to the present disclosure; Figure 8 This is a flowchart of the sixth embodiment of the lip-driving method according to the present disclosure; Figure 9 This is a schematic diagram of a structure of an embodiment of the lip-driven device according to the present disclosure; Figure 10 This is a block diagram of an electronic device used to implement the lip-driving method of the embodiments of this disclosure. Detailed Implementation

[0011] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0012] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0013] Figure 1 An exemplary system frame 100 is shown, to which embodiments of the lip-driving method or lip-driving device of this disclosure may be applied.

[0014] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0015] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include cloud storage applications and instant messaging applications.

[0016] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.

[0017] Server 105 can provide various services through its built-in applications. For example, users can operate the server through applications on terminal devices 101, 102, and 103 and send requests to server 105 to generate lip-sync video. Server 105 can receive and process these requests, performing the following steps: encoding the template video and target audio into the latent space to obtain a first latent vector and a second latent vector; dynamically adjusting the influence of the second latent vector on the first latent vector using an elliptic Gaussian spatiotemporal adaptive weighting system during denoising to obtain a denoised latent vector; decoding and transforming the denoised latent vector to obtain a transformed image; and generating a lip-sync video based on the transformed image and the target audio.

[0018] It should be noted that the lip-sync driving method provided in this embodiment is generally executed by server 105, and correspondingly, the lip-sync driving device is generally located in server 105.

[0019] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0020] Continue to refer to Figure 2 The diagram illustrates a flow 200 of a first embodiment of a lip-driving method according to the present disclosure. The lip-driving method includes the following steps: Step 201: Encode the template video and the target audio into the latent space respectively to obtain the first latent vector and the second latent vector.

[0021] In this embodiment, the execution body of the lip-driven method (e.g. Figure 1 The server 105 shown encodes the template video and the target audio into the latent space, respectively, to obtain a first latent vector and a second latent vector. The template video is a pre-shot, high-quality video that retains only the mouth area as a variable region. The target audio is the lip-driven audio, which can be an audio file of any language; this embodiment does not impose specific limitations on this. In other words, the execution entity reconstructs the speaker's mouth area in the template video based on the target audio to synchronize it with the target audio.

[0022] For template videos, the aforementioned execution entity performs face detection and facial landmark localization frame by frame, and determines the affine transformation matrix from the original image to the standardized face space based on the facial landmarks. This standardized face space is typically 512. 512 pixels. Then, the video frame images in the template video are mapped to the normalized face space according to the affine transformation matrix, thus obtaining the mapped aligned frame images. That is, each video frame image of the template video is mapped to the normalized face space to obtain the corresponding aligned frame image for each video frame image. Afterwards, the above-mentioned execution entity uses an encoder to encode the obtained aligned frame images to obtain the corresponding first latent vector. For example, a VAE encoder (Variational Autoencoder Encoder) can be used to encode the aligned frame images into the latent space to obtain the first latent vector. Through affine transformation, the face can be corrected to a horizontally centered state, eliminating the interference of rotation and scaling.

[0023] The aforementioned execution entity also utilizes an audio feature extraction network to extract features from the target audio, obtaining temporal audio features. These temporal audio features are then encoded using an encoder to obtain a second latent vector. For example, a VAE encoder can be used to encode the temporal audio features into the latent space to obtain the second latent vector.

[0024] It's important to note that latent space is a low-dimensional, dense, and structured abstract feature space obtained by a model through learning the data distribution and compressing and mapping high-dimensional raw data (such as images, audio, and video). In this space, each point (latent vector) represents the core semantic information of the original data, rather than surface pixels or waveforms, facilitating efficient generation, editing, and computation by the model. If the data in the latent space is expressed using a mathematical expression, the variables within it are generally referred to as variables in the latent space, or simply latent variables.

[0025] Step 202: In the process of denoising using the diffusion model, the influence intensity of the second latent vector on the first latent vector is dynamically adjusted according to the elliptic Gaussian spatiotemporal adaptive weight to obtain the denoised latent vector.

[0026] In this embodiment, during the denoising process using the diffusion model, the execution entity dynamically adjusts the influence intensity of the second latent vector on the first latent vector according to the elliptic Gaussian spatiotemporal adaptive weight to obtain the denoised latent vector.

[0027] The diffusion model is a generative artificial intelligence model based on stochastic processes and inverse generation. Its core is to learn the true distribution of data by simulating the process of "gradually adding noise (forward diffusion)" and "gradually removing noise (backward diffusion)," and finally generate content that is highly similar to real data from pure noise.

[0028] As an example, the diffusion model in this embodiment is LatentSync, which is an end-to-end video lip-syncing AI (Artificial Intelligence) framework based on a latent diffusion model under audio conditions.

[0029] In this embodiment, the denoising process using the diffusion model is the reverse diffusion process of the diffusion model. In this process, elliptic Gaussian spatiotemporal adaptive CFG (Classifier-Free Guidance) weights are introduced to dynamically adjust the guidance intensity of audio conditions on different spatial locations and time steps. That is, the influence intensity (guidance intensity) of the second latent vector (audio conditions) on the first latent vector (mouth region) is dynamically adjusted according to the elliptic Gaussian spatiotemporal adaptive weights, thereby gradually stripping away noise and obtaining a mouth feature latent vector synchronized with the audio, which is the denoised latent vector.

[0030] Traditional LatentSync often treats the impact of audio on the image uniformly during generation, which can lead to severe distortion in the generated mouth. This embodiment, however, introduces elliptic Gaussian spatiotemporally adaptive CFG weights at each time step and spatial location during the LatentSync diffusion generation process. This adjusts the influence of audio conditions on the current latent vector update, resulting in a denoised latent vector. The elliptic Gaussian spatiotemporally adaptive CFG weights are generated based on elliptic Gaussian spatial weights and temporal scheduling weights. The elliptic Gaussian spatial weights are modeled using an elliptic Gaussian distribution and used as spatial modulation factors for the CFG weights, thus achieving pixel-by-pixel guidance intensity control. The temporal scheduling weights employ configurable scheduling strategies (e.g., linear scheduling, cosine scheduling, step scheduling, etc.) based on the denoising progress. In other words, the execution entity dynamically adjusts the audio guidance intensity according to the denoising progress bar (time step).

[0031] Step 203: Decode and transform the denoised latent vectors to obtain the transformed image.

[0032] In this embodiment, the execution entity decodes and transforms the denoised latent vectors to obtain a transformed image. Specifically, the execution entity first uses a VAE decoder to decode the denoised latent vectors, obtaining an aligned image in the normalized face space. This aligned image differs from the video frame image in the template video only in the mouth region. Then, the execution entity uses the affine transformation matrix generated in the preceding steps to map the aligned image back to the original coordinate system, thus obtaining a mapped image frame. Finally, the execution entity stitches and merges the mapped image frame with the background of the video frame image in the template video to obtain the transformed image.

[0033] Step 204: Generate lip-driven video based on the transformed image and target audio.

[0034] In this embodiment, the execution entity generates a lip-sync video based on the transformed image and target audio. The execution entity synthesizes the transformed image according to the frame rate of the template video to obtain the target video. Here, the frame rate can be 30fps (frames per second). The execution entity follows the original 30fps frame rate of the template video, synthesizing the ordered image frame sequence into a continuous audio-free video, i.e., the target video, thus ensuring the smoothness of video playback matches the playback rhythm of the original template video. Then, the execution entity synthesizes the target video with the target audio to obtain the lip-sync video. Specifically, the execution entity can employ audio-video synchronization fusion technology, based on the principle of precise timestamp alignment, to synthesize the audio-free target video and the target audio into the final lip-sync video, thereby ensuring frame-by-frame synchronization between lip movements and audio pronunciation.

[0035] The lip-guided video method disclosed herein first encodes the template video and target audio into a latent space, obtaining a first latent vector and a second latent vector. Then, during denoising using a diffusion model, the influence intensity of the second latent vector on the first latent vector is dynamically adjusted according to elliptic Gaussian spatiotemporal adaptive weights, resulting in a denoised latent vector. Next, the denoised latent vector is decoded and transformed to obtain a transformed image. Finally, a lip-guided video is generated based on the transformed image and the target audio. This method achieves precise alignment of visual and audio features in the latent space by encoding the template video and target audio into first and second latent vectors of the same dimension, improving the synchronization accuracy of audio and video. Furthermore, by dynamically adjusting the influence intensity of the second latent vector (audio) on the first latent vector (video) using elliptic Gaussian spatiotemporal adaptive weights, the audio guidance accurately focuses on the central region of the mouth, solving problems such as lip corner adhesion and unnatural lip closure in traditional methods, thus improving the generation quality and efficiency of lip-guided videos.

[0036] Furthermore, the collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, involved in the technical solutions disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0037] Continue to refer to Figure 3 , Figure 3 A flow 300 of a second embodiment of the lip-driving method according to this disclosure is shown. The lip-driving method includes the following steps: Step 301: Map the video frame images in the template video to the standardized face space to obtain the mapped aligned frame images.

[0038] In this embodiment, the execution body of the lip-driven method (e.g. Figure 1 The server 105 shown maps video frame images from the template video to a standardized face space, obtaining mapped aligned frame images. For the template video, the aforementioned execution entity performs face detection and facial landmark localization frame by frame, and determines the affine transformation matrix from the original image to the standardized face space based on the facial landmarks. The standardized face space is typically 512. 512 pixels. Then, the video frame images in the template video are mapped to the normalized face space according to the affine transformation matrix, thus obtaining the mapped aligned frame images.

[0039] In some optional implementations of this embodiment, step 301 includes: Step 3011: Calculate the mapping relationship between the original coordinate system and the standardized face space corresponding to the video frame image.

[0040] For a video frame image, the aforementioned execution entity first determines the coordinates of each key point (or reference point) in the image in the original coordinate system, denoted as the first coordinate; and determines the coordinates of each key point (or reference point) in the standardized face space, denoted as the second coordinate. Then, based on the correspondence between the first and second coordinates, the mapping relationship between the original coordinate system and the standardized face space is determined.

[0041] In some optional implementations of this embodiment, step 3011 includes: detecting the face region in the video frame image and extracting the face key points in the face region; calculating the affine transformation matrix of the face key points from the original coordinate system to the standardized face space, and determining the mapping relationship based on the affine transformation matrix.

[0042] In this implementation, the aforementioned execution entity inputs the original video frames into a face detector to locate the bounding boxes of face regions in the video frame images and remove background interference. Then, the face regions are input into a keypoint detection model to extract the original pixel coordinates of all keypoints and filter out the actual coordinates of the core reference points. Next, according to a preset rule for a 512×512 pixel standardized space, the standard coordinates of the core reference points are defined. These coordinates remain fixed, ensuring that the face pose and position are consistent after all video frames are mapped to the standardized space. Finally, the affine transformation matrix is ​​solved according to a preset formula. Specifically, the affine transformation matrix can be solved based on the coordinates of the keypoints in the standardized face space and their coordinates in the original coordinate system. This matrix represents the unique mapping relationship between the original coordinate system and the standardized face space, describing the coordinate transformation rules from any pixel in the original video frame to the standardized space. Thus, based on the determined affine transformation matrix, the video frame images are quickly and accurately mapped from the original coordinate system to the standardized face space, improving the accuracy of the mapping.

[0043] Step 3012: Map the video frame image to the standardized face space according to the mapping relationship to obtain the mapped aligned frame image.

[0044] Since the original coordinate system and the standardized face space have a unique mapping relationship, and this mapping relationship can characterize the coordinate transformation rules from any pixel in the original video frame to the standardized space, the aforementioned execution entity can map the video frame image to the standardized face space according to this mapping relationship, thereby obtaining the mapped aligned frame image. The aforementioned execution entity will perform the mapping operation on each video frame image, thereby obtaining the mapped aligned frame image corresponding to each video frame image.

[0045] Since the calculation of the mapping relationship and the generation of the alignment frame for a single 1920×1080 pixel video takes less than 20 milliseconds, the above steps improve the calculation efficiency of the mapping relationship and the mapping efficiency. It supports real-time preprocessing of 30fps video and meets the needs of batch video production.

[0046] Step 302: Encode the mapped aligned frame image to obtain the first latent vector.

[0047] In this embodiment, the execution entity encodes the mapped alignment frame image to obtain a first latent vector. The execution entity uses an encoder to encode the obtained alignment frame image, thereby obtaining the corresponding first latent vector. For example, a VAE encoder can be used to encode the alignment frame image into the latent space to obtain the first latent vector.

[0048] Step 303: Extract temporal audio features from the target audio and encode the temporal audio features to obtain the second latent vector.

[0049] In this embodiment, the execution entity extracts temporal audio features from the target audio and encodes these features to obtain a second latent vector. The execution entity also utilizes an audio feature extraction network to extract features from the target audio, obtaining temporal audio features. These features are then encoded using an encoder to obtain the second latent vector. For example, a VAE encoder can be used to encode the temporal audio features into the latent space to obtain the second latent vector.

[0050] Encoding can reduce the size of an image (e.g., from 512×512 pixels to 64×64 pixels) while preserving core information such as the image's structure, texture, and lighting.

[0051] Step 304: In the process of denoising using the diffusion model, the influence intensity of the second latent vector on the first latent vector is dynamically adjusted according to the elliptic Gaussian spatiotemporal adaptive weight to obtain the denoised latent vector.

[0052] Step 305: Decode and transform the denoised latent vectors to obtain the transformed image.

[0053] Step 306: Generate lip-driven video based on the transformed image and target audio.

[0054] Steps 304-306 are basically the same as steps 202-204 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 202-204, which will not be repeated here.

[0055] from Figure 3 It can be seen from this that, with Figure 2 Compared to the corresponding embodiments, the lip-driven method in this embodiment emphasizes the step of encoding the template video and the target audio. By compressing the high-resolution standardized face image into a low-dimensional first latent vector and encoding the temporal audio features into a second latent vector, the computational load is reduced by several orders of magnitude while accurately preserving the core semantic features of the face (including face contour, skin texture, basic mouth shape, lighting features, etc.) and discarding redundant pixel details. In other words, the computational load of encoding is reduced while the accuracy of features is improved.

[0056] Continue to refer to Figure 4 , Figure 4 A flow 400 of a third embodiment of the lip-driving method according to this disclosure is shown. The lip-driving method includes the following steps: Step 401: Map the video frame images in the template video to the standardized face space to obtain the mapped aligned frame images.

[0057] Step 402: Encode the mapped aligned frame image to obtain the first latent vector.

[0058] Step 403: Extract temporal audio features from the target audio and encode the temporal audio features to obtain the second latent vector.

[0059] Steps 401-403 are basically the same as steps 301-303 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 301-303, which will not be repeated here.

[0060] Step 404: In the process of denoising using the diffusion model, for the aligned frame image, random Gaussian noise with the same dimension as the first latent vector is generated, and the first latent vector and the random Gaussian noise are used as the initial input of the diffusion model.

[0061] In this embodiment, during the denoising process using the diffusion model, for each aligned frame image, the execution entity of the lip-sync driving method (e.g., Figure 1 The server 105 shown generates random Gaussian noise with the same dimension as the first latent vector, and uses the first latent vector and the random Gaussian noise as the initial input to the diffusion model. That is, the execution entity first extracts the dimensional information of the first latent vector to obtain a dimensional tuple, then calls a preset function to generate random Gaussian noise following a standard normal distribution, using the extracted dimensional tuple as input. The generated random Gaussian noise can then be verified and normalized to ensure it maintains the same numerical range as the first latent vector, avoiding distortion of the initial input features of the diffusion model due to differences in numerical scale. Finally, the first latent vector and the normalized random Gaussian noise are fused to construct the initial input tensor for inverse denoising of the diffusion model, providing basic features for subsequent time-step denoising guided by elliptic Gaussian spatiotemporal adaptive CFG.

[0062] Step 405: Bind the second latent vector to the identifier of the current time step, and use the bound second latent vector as the conditional input of the diffusion model.

[0063] In this embodiment, the aforementioned execution entity determines the identification information of the current time step, then binds the second latent vector with the identification information of the current time step, loads the bound fused features into the conditional input layer of the diffusion model, and inputs them in conjunction with the initial input tensor (first latent vector + random Gaussian noise) and elliptic Gaussian spatiotemporal adaptive CFG weights of the diffusion model, providing conditional guidance for the noise residual prediction of the UNet network (a type of convolutional neural network).

[0064] Step 406: Input the initial input and conditional input into the diffusion model. During the denoising process using the diffusion model, dynamically adjust the influence intensity of the second latent vector on the first latent vector according to the elliptic Gaussian spatiotemporal adaptive weight to obtain the denoised latent vector.

[0065] In this embodiment, the aforementioned execution entity introduces elliptic Gaussian spatiotemporal adaptive CFG weights to dynamically adjust the guidance intensity of audio conditions on different spatial locations and time steps. Specifically, the conditional input of the diffusion model (second latent vector + current time step identifier), the initial input tensor of the diffusion model (first latent vector + random Gaussian noise), and the elliptic Gaussian spatiotemporal adaptive CFG weights are used as co-inputs. Based on the elliptic Gaussian spatiotemporal adaptive weights, the influence intensity (guidance intensity) of the second latent vector (audio conditions) on the first latent vector (mouth region) is dynamically adjusted, thereby gradually stripping away noise to obtain a mouth feature latent vector synchronized with the audio, i.e., the denoised latent vector.

[0066] Step 407: Decode and transform the denoised latent vectors to obtain the transformed image.

[0067] Step 408: Generate lip-driven video based on the transformed image and target audio.

[0068] Steps 407-408 are basically the same as steps 203-204 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 203-204, which will not be repeated here.

[0069] from Figure 4 It can be seen from this that, with Figure 3 Compared to the corresponding embodiments, the lip-guided method in this embodiment emphasizes the step of obtaining the denoised latent vector. This is achieved by generating random Gaussian noise that perfectly matches the dimension of the first latent vector and fusing the two as the initial input to the diffusion model. This avoids excessive noise masking of core visual features such as facial contours, skin texture, and basic mouth shape in the first latent vector. Simultaneously, the second latent vector (audio latent vector) is bound to the current time step identifier as a conditional input, enabling the audio guidance feature to carry temporal process information of diffusion denoising, thereby improving the matching degree between the audio guidance and denoising processes.

[0070] Continue to refer to Figure 5 , Figure 5 It shows Figure 4 The process 500 of the embodiment of step 406 includes: Step 501: For each time step, calculate the elliptic Gaussian spatiotemporal adaptive weights corresponding to the aligned frame image.

[0071] For each time step, the aforementioned execution entity calculates the elliptic Gaussian spatiotemporal adaptive weights corresponding to the aligned frame image. These elliptic Gaussian spatiotemporal adaptive weights are generated based on elliptic Gaussian spatial weights and temporal scheduling weights. Specifically, for each denoising time step, the execution entity calculates the spatiotemporal CFG weights pixel-by-pixel. These spatiotemporal CFG weights are generated based on elliptic Gaussian spatial weights and temporal scheduling weights, and are used to dynamically adjust the influence of audio conditions. The elliptic Gaussian spatial weights are modeled using an elliptic Gaussian distribution and are used as spatial modulation factors for the CFG weights, thereby achieving pixel-by-pixel guidance intensity control. The temporal scheduling weights employ configurable scheduling strategies (e.g., linear scheduling, cosine scheduling, step scheduling, etc.) based on the denoising process.

[0072] In some optional implementations of this embodiment, the elliptical Gaussian space weights are generated through the following steps: For the aligned frame image, the width and height of the target lip region are determined by facial key point detection; the horizontal and vertical standard deviations are dynamically adjusted according to the width and height so that the elliptical region covers the target lip region; for a point in the standardized face space, the horizontal and vertical offsets of the point are calculated; the elliptical Gaussian space weights are calculated according to the horizontal standard deviation, vertical standard deviation, horizontal offset, vertical offset, and a preset weight calculation formula.

[0073] In this implementation, for each aligned frame image, the aforementioned execution entity constructs an elliptical Gaussian distribution of spatial weights with the mouth region as the core, so that the audio guidance intensity decays from the center of the mouth to the periphery, avoiding distortion in non-mouth regions (such as eyes, cheeks, and background) due to audio guidance.

[0074] Specifically, the aforementioned execution entity acquires key points such as the corners of the mouth and the cupid's bow through facial key point detection (e.g., 68-point detection) and fits the coordinates of the mouth center. Then, it dynamically adjusts the horizontal and vertical standard deviations based on the mouth width and height to ensure that the ellipse accurately covers the mouth area. For any point (x, y) in the standardized face space, the aforementioned execution entity calculates the horizontal and vertical offsets. Finally, based on the horizontal and vertical standard deviations, the horizontal and vertical offsets, and the preset weight calculation formula, the weights of the elliptical Gaussian space are calculated.

[0075] The center of the mouth has a weight close to 1, providing the strongest audio guidance and ensuring precise synchronization between lip movements and audio. The cheeks, chin, and other transitional areas have weights of 0.3 to 0.7, achieving a smooth decay of guidance intensity. The eyes, background, and other areas have weights close to 0, forcibly preserving the original image structure and texture.

[0076] Elliptical Gaussian spatial weights are obtained by modeling elliptical Gaussian distributions. These weights are used as spatial modulation factors for CFG weights, thereby achieving pixel-by-pixel guidance intensity control.

[0077] In some optional implementations of this embodiment, the time scheduling weight is determined based on the total number of time steps, the current time step, and a predefined normalized time.

[0078] In this implementation, the aforementioned execution entity dynamically adjusts the audio guidance intensity according to the denoising process. In the early stages of denoising, the focus is on "mouth structure shaping" and a high guidance intensity is used; in the later stages of denoising, the focus is on "detail and texture preservation" and a low guidance intensity is used to avoid excessive guidance that could damage the image quality.

[0079] Let the total number of diffusion steps be K, and the current time step be k (1~K). Define the normalized time τ=k / K (τ∈[0,1]). The smaller τ is, the earlier the denoising begins. Then the time weight g(τ) satisfies: (1) Early denoising (smaller τ): g(τ) is larger, used to quickly determine the overall structure and posture of the mouth; (2) Late denoising (larger τ): g(τ) is smaller, to reduce the disturbance of audio to details and textures, so as to preserve the details of the face and background.

[0080] Supported scheduling strategies are: Linear scheduling: g_linear(τ) = 1 - τ, where the guiding strength decreases linearly with time step, is computationally simple, and adaptable to most scenarios; Cosine scheduling: g_cos(τ) = (1 + cos(π / 2)) / (τ - τ ... The guidance intensity decays gradually in the early stage and decreases rapidly in the later stage, making it suitable for scenarios with high detail requirements. Tiered scheduling controls the intensity in three segments: 0.8~1.0 for τ∈[0,τ1] (high guidance), 0.4~0.7 for τ∈(τ1,τ2] (medium guidance), and 0.1~0.3 for τ∈(τ2,1] (low guidance). This improves the accuracy of lip-sync.

[0081] In some optional implementations of this embodiment, the above method further includes: calculating the product of the elliptic Gaussian space weights and the time scheduling weights, and using the product as the elliptic Gaussian spatiotemporal adaptive weights.

[0082] In this implementation, the aforementioned execution entity multiplies the elliptic Gaussian space weights with the time scheduling weights to obtain the pixel-by-pixel, time-step-by-time spatiotemporal CFG weights wk(x,y). In implementation, the effective CFG at a given pixel location can be defined as the product of CFG_max and wk(x,y), where CFG_max is the global maximum guiding strength.

[0083] By adjusting the influence of the calculated elliptic Gaussian spatiotemporal adaptive weights on the current latent vector update, the accuracy of lip-sync is improved, the stability of non-mouth regions is enhanced, and the naturalness of detail textures is increased, making it suitable for various complex scenarios.

[0084] Step 502: Input the first latent vector, the second latent vector, and the elliptic Gaussian spatiotemporal adaptive weights into the convolutional neural network to predict the current noise residual.

[0085] The aforementioned execution entity preprocesses the first latent vector, the second latent vector, and the elliptic Gaussian spatiotemporal adaptive weights to complete batch dimension expansion and feature normalization calibration, adapting to the input format requirements of the UNet convolutional neural network. Then, the execution entity performs multi-scale feature learning on the modulated fused features through UNet's encoder downsampling, decoder upsampling, and convolutional feature extraction, ultimately outputting the noise residual at the current time step with the same dimension as the first latent vector.

[0086] Step 503: Update and optimize the first latent vector based on the current noise residual and the preset update formula.

[0087] The aforementioned execution entity performs pixel-by-pixel multiplication of the predicted current noise residual with the elliptical Gaussian spatiotemporal adaptive weights to obtain the spatially modulated noise residual, thereby achieving effective noise stripping only in the mouth region, while significantly suppressing noise residuals in non-mouth regions. Then, it calculates the core noise component to be stripped from the first latent vector at the current time step and quantifies the noise stripping magnitude. Next, it subtracts the core noise stripping term from the first latent vector to be updated and multiplies it by the basic scaling factor to complete the inverse denoising update. Finally, it adds a small amount of random Gaussian noise to the basic update term to ensure generation diversity, resulting in the updated and optimized first latent vector.

[0088] Step 504: In response to determining the preset total number of time steps completed, output the optimized first latent vector and use the optimized first latent vector as the denoised latent vector.

[0089] If the aforementioned execution entity determines that the preset total number of time steps has been completed, it will output the first latent vector optimized for the current time step, and use the output optimized first latent vector as the denoised latent vector.

[0090] By collaboratively inputting the first latent vector (visual), the second latent vector (audio temporal), and elliptic Gaussian spatiotemporally adaptive weights (temporal modulation) into the convolutional neural network, deep fusion of visual features, audio-driven features, and modulation weights is achieved, improving the accuracy of denoising and lip-sync. Since the optimized latent vector output after iterative denoising only updates the features of the mouth region in sync with the audio, the visual features of non-mouth regions and the original template are preserved to the greatest extent. Therefore, distortion and blurring problems are avoided, improving the accuracy of the denoised latent vector.

[0091] Continue to refer to Figure 6 , Figure 6A flow 600 of a fourth embodiment of the lip-driving method according to this disclosure is shown. The lip-driving method includes the following steps: Step 601: Map the video frame images in the template video to the standardized face space to obtain the mapped aligned frame images.

[0092] Step 602: Encode the mapped aligned frame image to obtain the first latent vector.

[0093] Step 603: Extract temporal audio features from the target audio and encode the temporal audio features to obtain the second latent vector.

[0094] Steps 601-603 are basically the same as steps 301-303 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 301-303, which will not be repeated here.

[0095] Step 604: In the process of denoising using the diffusion model, for the aligned frame image, random Gaussian noise with the same dimension as the first latent vector is generated, and the first latent vector and the random Gaussian noise are used as the initial input of the diffusion model.

[0096] Step 605: Bind the second latent vector to the identifier of the current time step, and use the bound second latent vector as the conditional input of the diffusion model.

[0097] Step 606: Input the initial input and conditional input into the diffusion model. During the denoising process using the diffusion model, dynamically adjust the influence intensity of the second latent vector on the first latent vector according to the elliptic Gaussian spatiotemporal adaptive weight to obtain the denoised latent vector.

[0098] Steps 604-606 are basically the same as steps 404-406 in the previous embodiment. For specific implementation methods, please refer to the above description of steps 404-406, which will not be repeated here.

[0099] Step 607: Decode the denoised latent vectors to obtain the aligned image in the standardized face space.

[0100] In this embodiment, the execution body of the lip-driven method (e.g. Figure 1 The server 105 shown uses a VAE decoder to decode the denoised latent vectors, thereby obtaining an aligned image in the normalized face space. The denoised latent vectors are then fed into the VAE decoder, and the VAE encoder restores them to a high-resolution aligned face image. At this point, the mouth movements are synchronized with the audio.

[0101] Step 608: Map the aligned image to the original coordinate system according to the affine transformation matrix to obtain the mapped image frame.

[0102] In this embodiment, the execution entity will again utilize the affine transformation matrix between the original coordinate system and the standardized face space to map the aligned image in the standardized face space back to the original coordinate system, thereby obtaining the mapped image frame. That is, the execution entity will use the affine transformation matrix calculated in the preprocessing stage to map the generated aligned face image containing the new lip shape back to the coordinate system of the original video frame, thereby obtaining the mapped image frame.

[0103] Step 609: The mapped image frame is stitched together with the background in the template video to obtain the transformed image.

[0104] In this embodiment, the execution entity will stitch the mapped image frame (face image) with the background in the video frame image of the template video to obtain the transformed image.

[0105] Step 610: Generate lip-driven video based on the transformed image and target audio.

[0106] Step 610 is basically the same as step 204 in the aforementioned embodiment. For the specific implementation method, please refer to the aforementioned description of step 204, which will not be repeated here.

[0107] from Figure 6 It can be seen from this that, with Figure 4 Compared to the corresponding embodiments, the lip-shaped driving method in this embodiment emphasizes the steps of decoding and transforming the denoised latent vectors. By decoding the denoised latent vectors into a standardized face space, the resulting aligned image only completes feature updates in the mouth region that are precisely synchronized with the target audio. The non-mouth regions retain the original texture, lighting, and contour features of the template video, thereby improving the quality of the generated lip-shaped image and enhancing the matching degree between the generated lip-shaped image and the audio.

[0108] Continue to refer to Figure 7 , Figure 7 A flow 700 of a fifth embodiment of the lip-driving method according to this disclosure is shown. The lip-driving method includes the following steps: Step 701: Map the video frame images in the template video to the standardized face space to obtain the mapped aligned frame images.

[0109] Step 702: Encode the mapped aligned frame image to obtain the first latent vector.

[0110] Step 703: Extract temporal audio features from the target audio and encode the temporal audio features to obtain the second latent vector.

[0111] Steps 701-703 are basically the same as steps 301-303 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 301-303, which will not be repeated here.

[0112] Step 704: In the process of denoising using the diffusion model, for the aligned frame image, random Gaussian noise with the same dimension as the first latent vector is generated, and the first latent vector and the random Gaussian noise are used as the initial input of the diffusion model.

[0113] Step 705: Bind the second latent vector to the identifier of the current time step, and use the bound second latent vector as the conditional input of the diffusion model.

[0114] Step 706: Input the initial input and conditional input into the diffusion model. During the denoising process using the diffusion model, dynamically adjust the influence intensity of the second latent vector on the first latent vector according to the elliptic Gaussian spatiotemporal adaptive weight to obtain the denoised latent vector.

[0115] Steps 704-706 are basically the same as steps 404-406 in the previous embodiment. For specific implementation methods, please refer to the above description of steps 404-406, which will not be repeated here.

[0116] Step 707: Decode the denoised latent vectors to obtain the aligned image in the standardized face space.

[0117] Step 708: Map the aligned image to the original coordinate system according to the affine transformation matrix to obtain the mapped image frame.

[0118] Steps 707-708 are basically the same as steps 607-608 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 607-608, which will not be repeated here.

[0119] Step 709: In response to the identification of occlusion in the target region of the video frame image, generate the occlusion mask corresponding to the target region.

[0120] In this embodiment, the execution body of the lip-driven method (e.g. Figure 1 If the server 105 (shown) identifies an occlusion in the target region of a video frame image, where the target region is the mouth (or lips), it will use semantic segmentation or depth estimation methods to identify the occluder (such as a finger, microphone, etc.) in front of the mouth in the video frame image and generate a mask M_t^(0) for the target region (corner of the mouth region). Here, the occluded region is segmented in the standard space, rather than in the original image space. Because the face position and scale are standardized: small faces are enlarged, the relative proportion of the occluder is more reasonable, and the segmentation is more accurate; in side-view and low-light scenes, the structure of the face and the background is more consistent, and the segmentation network is easier to converge.

[0121] In addition, the aforementioned implementing entities will fine-tune the diffusion model in advance using e-commerce data. Specifically, the implementing entities will first collect a large number of e-commerce broadcast videos, including product displays, live-streaming sales, and explanatory short videos, covering various lighting, occlusion, makeup, and poses. Then, using the latent space diffusion training framework consistent with LatentSync, the e-commerce data will be mixed with the original general data or trained in stages. Finally, global pixel reconstruction loss L_pixel and perceptual loss L_perc will be used to adapt the model to the composition and style of e-commerce scenarios.

[0122] Step 710: Perform multi-stage morphological processing on the occlusion mask to obtain the processed mask; use the processed mask to perform weighted fusion of the aligned image and the video frame image to obtain the target image frame.

[0123] After obtaining the coarse occlusion mask M_t^(0), the above-mentioned execution entity will perform multi-stage morphological processing on it, including: a) Morphological closing operation: perform dilation operation and erosion (closing operation) on the mask, fill small holes, connect adjacent occlusion areas, and obtain M_t^(1); b) Contour filling: detect the contour of each occlusion area in M_t^(1), fill the holes inside the contour, eliminate unnecessary internal blanks, and obtain a more connected occlusion area M_t^(2); c) CLAHE illumination preprocessing: apply adaptive histogram equalization (CLAHE) to the brightness channel of the standardized face image to enhance the segmentation network to distinguish occlusions from skin areas more easily in low light or strong backlight conditions. CLAHE can be used to enhance the illumination robustness of the training and inference stages.

[0124] Furthermore, to eliminate the jitter caused by frame-by-frame segmentation, this stage smooths and fuses the occlusion mask in the time dimension. Specifically, a) Sliding window averaging or EMA smoothing: Option 1: Perform pixel-by-pixel sliding average on the mask within a time window of length W; Option 2: Use exponential moving average (EMA); b) Edge feathering fusion: Apply Gaussian blur or morphological blur to the edge regions of the smoothed mask to obtain a soft mask M_soft_t. When fusing the mouth region, the soft mask is used to perform weighted fusion of the generated mouth image and the original image to obtain the target image frame.

[0125] Step 711: The target image frame is stitched together with the background in the template video to obtain the transformed image.

[0126] In this embodiment, the execution entity stitches the target image frame after occlusion processing with the background in the video frame image of the template video to obtain the transformed image. When pasting the generated mouth back into the original image, the execution entity uses the mask to perform "selective fusion," so that the generated lips are only placed in the unoccluded skin area, while the original occlusion texture is completely preserved. This means that if a user runs their hand over the mouth, the generated lips will naturally appear behind the gap between the fingers, instead of being mistakenly drawn on the fingers.

[0127] Step 712: Generate lip-driven video based on the transformed image and target audio.

[0128] Step 712 is basically the same as step 204 in the aforementioned embodiment. For the specific implementation method, please refer to the aforementioned description of step 204, which will not be repeated here.

[0129] from Figure 7 It can be seen from this that, with Figure 6 Compared to the corresponding embodiments, the lip-sync driving method in this embodiment emphasizes the step of processing the occluded area, thereby stably preserving the shape and texture of the occluded object under various occlusion scenarios, making the occlusion boundary and the generated mouth area transition smoothly without obvious hard edges or flickering, and improving the stability and visual appeal of lip-sync driven video under occlusion conditions.

[0130] Continue to refer to Figure 8 , Figure 8 A flow 800 of a sixth embodiment of the lip-driving method according to the present disclosure is shown. The lip-driving method includes the following steps: Step 801: Encode the template video and the target audio into the latent space respectively to obtain the first latent vector and the second latent vector.

[0131] Step 802: In the process of denoising using the diffusion model, the influence intensity of the second latent vector on the first latent vector is dynamically adjusted according to the elliptic Gaussian spatiotemporal adaptive weight to obtain the denoised latent vector.

[0132] Step 803: Decode and transform the denoised latent vectors to obtain the transformed image.

[0133] Steps 801-803 are basically the same as steps 201-203 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 201-203, which will not be repeated here.

[0134] Step 804: Synthesize the transformed image according to the frame rate of the template video to obtain the target video.

[0135] In this embodiment, the execution body of the lip-driven method (e.g. Figure 1The server 105 shown will synthesize the transformed image according to the frame rate of the template video to obtain the target video. The frame rate here can be 30fps. The above-mentioned execution entity will follow the original frame rate of 30fps of the template video and synthesize the ordered image frame sequence into a continuous video without audio, that is, synthesize it into the target video, thereby ensuring that the smoothness of video playback is consistent with the playback rhythm of the original template video.

[0136] Step 805: Combine the target video and target audio to obtain the lip-driven video.

[0137] In this embodiment, the aforementioned execution entity synthesizes the target video and target audio to obtain a lip-sync driven video. Specifically, the execution entity can employ audio-video synchronization fusion technology, based on the principle of precise timestamp alignment, to synthesize the target video without audio and the target audio into the final lip-sync driven video, thereby ensuring frame-by-frame synchronization between lip movements and audio pronunciation.

[0138] from Figure 8 It can be seen from this that, with Figure 2 Compared to the corresponding embodiments, the lip-sync driving method in this embodiment emphasizes the step of generating a lip-sync driven video. This method synthesizes the transformed image frames according to the original frame rate of the template video, ensuring that the generated target video is completely consistent with the original template video in terms of playback rhythm and frame rate smoothness. This avoids problems such as screen stuttering and fast / slow motion caused by frame rate mismatch. At the same time, the unified frame rate synthesis rule makes the lip-sync changes between the entire video frame smooth and natural, without frame skipping or inter-frame flickering, restoring the original visual playback experience of the template video and improving the smoothness of visual playback.

[0139] The lip-guided method disclosed herein serves video production scenarios related to translation and e-commerce, with typical applications including but not limited to: (1) Cross-language e-commerce product explanation videos (single virtual anchor) Merchants provide original Chinese videos or text, which the system translates into the target language (such as English, Spanish, etc.) and generates the target audio via TTS (Text To Speech). Using the lip-syncing method provided in this disclosure, the system synchronizes the lip movements of the template virtual anchor in the video with the target audio, outputting multilingual, highly synchronized audio, and stable video even in occluded scenes.

[0140] (2) Automatic generation of intros for AIGC e-commerce short dramas by two people

[0141] Multiple two-person character video templates are pre-designed, retaining each character's mouth as a variable area. Based on the script and multi-character TTS, the lip-driven method disclosed herein is applied to the videos of the two characters. When characters occlude each other or cover products, multi-stage occlusion processing ensures natural occlusion and smooth edges, thereby automatically generating high-quality two-person short drama intros and lowering the creative production threshold.

[0142] (3) Live streaming assistance and replication

[0143] Using historical live stream clips as templates, and through overseas translation and the lip-syncing technology of this invention, multilingual "pseudo-live stream" clips are automatically generated. Even when the streamer is holding a microphone or covering their mouth with a product, the obstruction is naturally preserved, while lip-syncing for multiple languages ​​is achieved simultaneously.

[0144] (4) Other applicable fields

[0145] This includes scenarios that require synchronization with voice, such as multilingual online education videos, multilingual versions of the narrator in corporate promotional videos, digital human customer service, and virtual front desks.

[0146] Further reference Figure 9 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a lip-shaped driving device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0147] like Figure 9 As shown, the lip-sync driving device 900 of this embodiment includes: an encoding module 901, a denoising module 902, a decoding module 903, and a generation module 904. The encoding module 901 is configured to encode the template video and the target audio into the latent space, respectively, to obtain a first latent vector and a second latent vector. The denoising module 902 is configured to dynamically adjust the influence intensity of the second latent vector on the first latent vector according to an elliptic Gaussian spatiotemporal adaptive weight during denoising using a diffusion model, to obtain a denoised latent vector. The decoding module 903 is configured to decode and transform the denoised latent vector to obtain a transformed image. The generation module 904 is configured to generate a lip-sync driven video based on the transformed image and the target audio.

[0148] In this embodiment, the specific processing of the encoding module 901, the noise reduction module 902, the decoding module 903, and the generation module 904 in the lip-sync driving device 900, and the resulting technical effects, can be found in reference to [reference needed]. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.

[0149] In some optional implementations of this embodiment, the encoding module 901 includes: a first mapping submodule, configured to map video frame images in the template video to a standardized face space to obtain a mapped aligned frame image; a first encoding submodule, configured to encode the mapped aligned frame image to obtain a first latent vector; and a second encoding submodule, configured to extract temporal audio features from the target audio and encode the temporal audio features to obtain a second latent vector.

[0150] In some optional implementations of this embodiment, the first mapping submodule includes: a calculation unit configured to calculate the mapping relationship between the original coordinate system corresponding to the video frame image and the standardized face space; and a mapping unit configured to map the video frame image to the standardized face space according to the mapping relationship to obtain the mapped aligned frame image.

[0151] In some optional implementations of this embodiment, the computing unit is further configured to: detect the face region in the video frame image and extract the facial key points in the face region; calculate the affine transformation matrix of the facial key points from the original coordinate system to the standardized face space, and determine the mapping relationship based on the affine transformation matrix.

[0152] In some optional implementations of this embodiment, the denoising module 902 includes: a generation submodule configured to generate random Gaussian noise with the same dimension as the first latent vector for the aligned frame image, and use the first latent vector and the random Gaussian noise as the initial input of the diffusion model; a binding submodule configured to bind the second latent vector to the identifier of the current time step, and use the bound second latent vector as the conditional input of the diffusion model; and a denoising submodule configured to input the initial input and the conditional input into the diffusion model, and dynamically adjust the influence intensity of the second latent vector on the first latent vector according to the elliptic Gaussian spatiotemporal adaptive weights during the denoising process using the diffusion model to obtain the denoised latent vector.

[0153] In some optional implementations of this embodiment, the denoising submodule is further configured to: for each time step, calculate the elliptic Gaussian spatiotemporal adaptive weights corresponding to the aligned frame image, wherein the elliptic Gaussian spatiotemporal adaptive weights are generated based on the elliptic Gaussian spatial weights and the time scheduling weights; input the first latent vector, the second latent vector, and the elliptic Gaussian spatiotemporal adaptive weights into the convolutional neural network to predict the current noise residual; update and optimize the first latent vector according to the current noise residual and the preset update formula; in response to determining the preset total number of time steps completed, output the optimized first latent vector, and use the optimized first latent vector as the denoised latent vector.

[0154] In some optional implementations of this embodiment, the lip-shaped driving device 900 further includes: a first weight generation module, configured to determine the width and height of the target lip region for an aligned frame image by detecting facial key points; dynamically adjust the horizontal and vertical standard deviations according to the width and height so that the elliptical region covers the target lip region; calculate the horizontal and vertical offsets of a point in the standardized face space; and calculate the elliptical Gaussian space weights according to the horizontal standard deviation, vertical standard deviation, horizontal offset, vertical offset, and a preset weight calculation formula.

[0155] In some optional implementations of this embodiment, the lip-shaped driving device 900 further includes a second weight generation module, configured to determine the time scheduling weight based on the total number of time steps, the current time step, and a predefined normalized time.

[0156] In some optional implementations of this embodiment, the lip-shaped driving device 900 further includes: a third weight generation module, configured to calculate the product of the elliptic Gaussian space weight and the time scheduling weight, and use the product as the elliptic Gaussian spatiotemporal adaptive weight.

[0157] In some optional implementations of this embodiment, the decoding module 903 includes: a decoding submodule configured to decode the denoised latent vectors to obtain an aligned image in the standardized face space; a second mapping submodule configured to map the aligned image to the original coordinate system according to the affine transformation matrix to obtain a mapped image frame; and a stitching submodule configured to stitch the mapped image frame with the background in the template video to obtain a transformed image.

[0158] In some optional implementations of this embodiment, the lip-sync driving device 900 further includes: a mask generation module configured to generate an occlusion mask corresponding to the target region in response to recognizing that an occlusion occurs in the target region of the video frame image; a processing module configured to perform multi-stage morphological processing on the occlusion mask to obtain a processed mask; a fusion module configured to use the processed mask to perform weighted fusion of the aligned image and the video frame image to obtain a target image frame; and a stitching submodule further configured to stitch the target image frame with the background in the template video to obtain a transformed image.

[0159] In some optional implementations of this embodiment, the generation module 904 is further configured to: synthesize the transformed image according to the frame rate of the template video to obtain the target video; and synthesize the target video with the target audio to obtain the lip-sync driven video.

[0160] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0161] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0162] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0163] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0164] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0165] The computing unit 1001 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the lip-sync method. For example, in some embodiments, the lip-sync method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the lip-sync method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform the lip-driven method by any other suitable means (e.g., by means of firmware).

[0166] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0167] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0168] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0169] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0170] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0171] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0172] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0173] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A lip-guided method, comprising: The template video and the target audio are encoded into the latent space respectively to obtain the first latent vector and the second latent vector; In the process of denoising using the diffusion model, the influence intensity of the second latent vector on the first latent vector is dynamically adjusted according to the elliptic Gaussian spatiotemporal adaptive weight to obtain the denoised latent vector. The denoised latent vectors are decoded and transformed to obtain the transformed image; A lip-driven video is generated based on the transformed image and the target audio.

2. The method according to claim 1, wherein, The step of encoding the template video and the target audio into the latent space respectively to obtain the first latent vector and the second latent vector includes: The video frame images in the template video are mapped to the standardized face space to obtain the mapped aligned frame images; The mapped aligned frame image is encoded to obtain the first latent vector; Temporal audio features are extracted from the target audio, and the temporal audio features are encoded to obtain the second latent vector.

3. The method according to claim 2, wherein, The step of mapping video frame images from the template video to a standardized face space to obtain mapped aligned frame images includes: Calculate the mapping relationship between the original coordinate system corresponding to the video frame image and the standardized face space; The video frame image is mapped to the standardized face space according to the mapping relationship to obtain the mapped aligned frame image.

4. The method according to claim 3, wherein, The calculation of the mapping relationship between the original coordinate system corresponding to the video frame image and the standardized face space includes: Detect the face region in the video frame image and extract the facial key points in the face region; Calculate the affine transformation matrix of the facial key points from the original coordinate system to the standardized face space, and determine the mapping relationship based on the affine transformation matrix.

5. The method according to claim 2, wherein, In the process of denoising using a diffusion model, the influence intensity of the second latent vector on the first latent vector is dynamically adjusted according to the elliptic Gaussian spatiotemporal adaptive weights to obtain the denoised latent vector, including: For the aligned frame image, generate random Gaussian noise with the same dimension as the first latent vector, and use the first latent vector and the random Gaussian noise as the initial input of the diffusion model; The second latent vector is bound to the identifier of the current time step, and the bound second latent vector is used as the conditional input of the diffusion model. The initial input and the conditional input are input into the diffusion model. During the denoising process using the diffusion model, the influence intensity of the second latent vector on the first latent vector is dynamically adjusted according to the elliptic Gaussian spatiotemporal adaptive weights to obtain the denoised latent vector.

6. The method according to claim 5, wherein, In the process of denoising using the diffusion model, the influence intensity of the second latent vector on the first latent vector is dynamically adjusted according to the elliptic Gaussian spatiotemporal adaptive weights to obtain the denoised latent vector, including: For each time step, the elliptic Gaussian spatiotemporal adaptive weight corresponding to the aligned frame image is calculated, wherein the elliptic Gaussian spatiotemporal adaptive weight is generated based on the elliptic Gaussian spatial weight and the time scheduling weight. The first latent vector, the second latent vector, and the elliptic Gaussian spatiotemporal adaptive weights are input into a convolutional neural network to predict the current noise residual. The first latent vector is updated and optimized based on the current noise residual and the preset update formula; In response to determining the preset total number of time steps completed, the optimized first latent vector is output, and the optimized first latent vector is used as the denoised latent vector.

7. The method according to claim 6, wherein, The elliptic Gaussian space weights are generated through the following steps: For the aligned frame image, the width and height of the target lip region are determined by facial key point detection; The horizontal and vertical standard deviations are dynamically adjusted based on the width and height to ensure that the elliptical region covers the target lip region. For a point in the standardized face space, calculate the horizontal and vertical offset of that point; The elliptic Gaussian space weights are calculated based on the horizontal standard deviation, the vertical standard deviation, the horizontal offset, the vertical offset, and a preset weight calculation formula.

8. The method according to claim 6, wherein, The time scheduling weight is determined based on the total number of time steps, the current time step, and the predefined normalized time.

9. The method according to claim 6, wherein, The method further includes: Calculate the product of the elliptic Gaussian space weights and the time scheduling weights, and use the product as the elliptic Gaussian spatiotemporal adaptive weights.

10. The method according to claim 4, wherein, The step of decoding and transforming the denoised latent vectors to obtain the transformed image includes: Decode the denoised latent vectors to obtain the aligned image in the standardized face space; The aligned image is mapped to the original coordinate system according to the affine transformation matrix to obtain the mapped image frame; The mapped image frame is stitched together with the background in the template video to obtain the transformed image.

11. The method according to claim 10, wherein, The method further includes: In response to the identification of occlusion in a target region in the video frame image, an occlusion mask corresponding to the target region is generated; The occlusion mask is subjected to multi-stage morphological processing to obtain the processed mask; The aligned image and the video frame image are weighted and fused using the processed mask to obtain the target image frame; and The step of concatenating the mapped image frame with the background in the template video to obtain the transformed image includes: The target image frame is stitched together with the background in the template video to obtain the transformed image.

12. The method according to any one of claims 1-11, wherein, The step of generating lip-driven video based on the transformed image and the target audio includes: The transformed image is synthesized based on the frame rate of the template video to obtain the target video; The target video and the target audio are combined to obtain the lip-sync driven video.

13. A lip-shaped driving device, comprising: The encoding module is configured to encode the template video and the target audio into the latent space respectively, to obtain the first latent vector and the second latent vector; The denoising module is configured to dynamically adjust the influence intensity of the second latent vector on the first latent vector according to the elliptic Gaussian spatiotemporal adaptive weight during the denoising process using the diffusion model, so as to obtain the denoised latent vector. The decoding module is configured to decode and transform the denoised latent vectors to obtain the transformed image. The generation module is configured to generate lip-driven video based on the transformed image and the target audio.

14. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.

15. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method of any one of claims 1-12.

16. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-12.