A Lip Syncing Method and System Based on Multiple Reference Frames and Style Control
By extracting the lip movement style characteristics and reference key points of multiple reference frames in the speaker video, combining the audio to key point module and the key point to the video module, the problem of ignoring the speaker style and single reference frame information in the prior art is solved, and the style controllable and high-fidelity lip-synchronous video generation is achieved.
Patent Information
- Application Number
- CN202510274090.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The existing lip-sync method ignores the speaking style characteristics of the speaker in the original video, and only uses single or average multi-reference frame information for lip-sync generation, resulting in unsatisfactory generation.
A lip synchronous method based on multi-reference frames and style controllable is proposed. By obtaining speaker video data and driving audio data, lip motion style features and reference key points of multi-reference frames are extracted, and combined with audio to key point module and key point to video module, to generate lip synchronous video with controllable style and high fidelity.
It realizes the lip-synchronous video generation with controllable style and high fidelity, which can effectively utilize multi-reference frame information and speaker style characteristics, and the generated video is closer to the style and quality of the original video.
Smart Images

Figure CN119815096B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video generation, and in particular, to a lip synchronization method and system based on multi-reference frames and style controllability. Background Art
[0002] At present, the task of speaker video generation has attracted wide attention and become an important field in artificial intelligence. The lip synchronization task is a common task in this field, aiming to modify the lip shape of a speaker in a given audio-visual video through a specified driving audio so that the lip shape is synchronized with the driving audio. Lip synchronization methods and systems have been widely used in many application scenarios such as video dubbing, cross-language education, and e-commerce live streaming.
[0003] The existing mainstream lip synchronization methods and technologies are mainly two-stage key point methods, that is, using the key points extracted by pre-trained tools as intermediate representations, and dividing them into two generation sub-tasks: audio-to-key point and key-to-video, and training two sub-models respectively. Compared with tasks without speaker style input such as single-image driving, there is naturally richer information in the lip synchronization task, such as the speaking style characteristics of the speaker and the texture information brought by more reference frames. However, the existing methods do not fully utilize this information, and there are mainly two problems: one is ignoring the speaking style characteristics of the speaker in the original video; the other is only using single reference frame information for lip synchronization generation, or simply averaging the multi-reference frame information. Summary of the Invention
[0004] To overcome the above problems, the present invention proposes a lip synchronization method and system based on multi-reference frames and style controllability, which fully utilizes the speaking style characteristic information and multi-reference frame information existing in the original video to achieve style-controllable and high-fidelity lip synchronization video generation.
[0005] The specific technical solution adopted by the present invention is as follows:
[0006] In the first aspect, the present invention proposes a lip synchronization method based on multi-reference frames and style controllability, including the following steps:
[0007] S1, obtaining speaker video data and driving audio data, and extracting the driving audio features of each audio frame and the driving key points of each video frame;
[0008] S2, randomly selecting two groups of video frames from the video as the first multi-reference map and the second multi-reference map respectively, and regarding the driving key points of the reference map as reference key points;
[0009] S3, extracting the lip movement style features of the speaker from the video, combining the driving audio features and the reference key points of the first multi-reference map, and predicting fine key points using the audio-to-key point module;
[0010] S4. Combine the second multi-reference graph and its reference key points, and the fine key points predicted by the audio-to-key-points module, and use the key-points-to-video module to generate lip-sync frames;
[0011] S5. Calculate the loss based on the prediction results of the audio-to-key-points module and the generation results of the key-points-to-video module, and update the modules;
[0012] S6. Given the driving audio and the original video, extract the lip motion style features of the speaker from the original video or directly given the lip motion style features of the speaker, and use the trained audio-to-key-points module and key-points-to-video module to generate a synthetic video that is lip-synced with the given driving audio, completing the lip-sync task.
[0013] Further, the first multi-reference graph and the second multi-reference graph are independent of each other.
[0014] Further, the driving key points include 3D driving key points and 2D driving key points, and the reference key points include 3D reference key points and 2D reference key points; the fine key points predicted by the audio-to-key-points module are fine 2D driving key points.
[0015] Further, the lip motion style features are scalars, which are composed of lip opening amplitude and lip speed; the method for extracting the lip motion style features of the speaker from the video is as follows:
[0016] Screen the lip key points from the 3D driving key points of the speaker in each frame image of the video, including the upper lip key points and the lower lip key points;
[0017] Use the average value of the difference in the y coordinates of the upper lip key points and the lower lip key points to represent the lip opening amplitude;
[0018] Use the first-order average difference in time series of the absolute value of the difference in the y coordinates of the upper lip key points and the lower lip key points to represent the lip speed.
[0019] Further, the audio-to-key-points module includes:
[0020] A sparse 3D driving key point predictor, which takes the driving audio features of each audio frame, the lip motion style features of the speaker, and the 3D reference key points of the first multi-reference graph as inputs, and outputs the predicted sparse 3D driving key points;
[0021] A fine 2D driving key point predictor, which takes the 2D reference key points of the first multi-reference graph and the sparse 3D driving key points as inputs, and outputs the predicted fine 2D driving key points.
[0022] Furthermore, the sparse 3D driving key point predictor includes a reference key point encoder, a lip opening embedding layer, a lip velocity embedding layer, an audio encoder, and a driving key point decoder. The calculation process includes:
[0023] The reference key point encoder encodes the 3D reference key points of the first multi-reference graph to obtain the encoded 3D reference key point features.
[0024] The lip opening embedding layer and the lip velocity embedding layer respectively encode and embed the lip opening and lip velocity in the lip movement style features of the speaker to obtain the lip movement style embedding features composed of lip opening embedding and lip velocity embedding.
[0025] The audio encoder encodes the driving audio features of each audio frame to obtain the audio encoded features.
[0026] The encoded 3D reference key point features and the lip movement style embedding features are extended to the same dimension as the audio encoded features, and then connected with the audio encoded features to obtain the sparse 3D driving key point features. The number of frames of the sparse 3D driving key point features is the same as the number of audio frames.
[0027] The sparse 3D driving key point features are decoded by the driving key point decoder to obtain the predicted sparse 3D driving key points.
[0028] Furthermore, the fine 2D driving key point predictor includes a reference key point encoder, a linear layer, and a driving key point decoder. The calculation process includes:
[0029] The reference key point encoder encodes the 2D reference key points of the first multi-reference graph to obtain the encoded 2D reference key point features.
[0030] The predicted sparse 3D driving key points are embedded into sparse 3D driving key point embeddings by the linear layer.
[0031] The encoded 2D reference key point features are extended to the same dimension as the sparse 3D driving key point embeddings, and then connected with the sparse 3D driving key point embeddings to obtain the fine 2D driving key point features. The number of frames of the fine 2D driving key point features is the same as the number of audio frames.
[0032] The fine 2D driving key point features are decoded by the driving key point decoder to obtain the predicted fine 2D driving key points.
[0033] Furthermore, step S4 includes:
[0034] S4-1, according to the 2D reference key points of each frame of the second reference graph and the fine 2D driving key points of the current frame, warp each frame of the second reference graph feature map to generate multiple warped feature maps.
[0035] S4-2. Calculate the multi-scale aggregation ratio of the multi-distortion feature map using the fine 2D driving key points of the current frame, the 2D reference key points of all the second reference images, and the multi-distortion feature map. The multi-scale aggregation ratio includes the frame-scale ratio and the pixel-scale ratio. Aggregate the texture information of the multi-distortion feature map using the multi-scale aggregation ratio to obtain an aggregated feature map.
[0036] S4-3. Decode the aggregated feature map to obtain a generated image. Generate a smooth lower face mask using the driving key points of the original frame image. Fuse the speaker's lower face of the generated image and the original frame image using the smooth lower face mask to obtain the lip-sync frame image of the current frame.
[0037] S4-4. Repeat S4-1 to S4-3. Use the fine 2D driving key points of the next frame to generate the lip-sync frame image of the next frame until all the frames of the fine 2D driving key points are traversed.
[0038] In a second aspect, the present invention provides a lip-sync system based on multi-reference frames and style controllability for implementing the above-mentioned lip-sync method based on multi-reference frames and style controllability.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] The present invention realizes style-controllable lip sync based on multi-reference frames. By given a driving audio and an original video, the lip movement style features of the speaker can be extracted from the original video, or the lip movement style features of the speaker can be directly given, supporting any explicit or implicit speaker style specification. By the audio-to-key-points module from sparse to fine, the task difficulty of the one-step generation result is reduced, making the training effect more stable. By introducing the multi-scale aggregation ratio in the key-points-to-video module to generate lip sync, the multi-reference frame information can be effectively utilized to generate a video frame with higher fidelity. In summary, style-controllable and high-fidelity lip-sync video generation is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic diagram of a lip-sync method based on multi-reference frames and style controllability shown in an embodiment of the present invention;
[0042] Figure 2 is a schematic diagram of the structure of a sparse 3D driving key-point predictor shown in an embodiment of the present invention;
[0043] Figure 3 is a schematic diagram of the structure of a fine 2D driving key-point predictor shown in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The present invention will be further described and explained below in conjunction with specific embodiments. The described embodiments are merely illustrative of the present disclosure and do not delimit the scope of limitation. The technical features of each embodiment of the present invention can be combined correspondingly on the premise of not conflicting with each other.
[0045] The accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0046] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.
[0047] See Figure 1 , a lip synchronization method based on multiple reference frames and style controllability proposed by the present invention mainly includes the following steps:
[0048] Step 1, obtain the audio-visual data of the speaker, extract the 3D driving key points of the speaker's face in the video frame and the camera estimation matrix, and project the 3D driving key points using the camera estimation matrix to obtain 2D driving key points; randomly select some video frames from them to form the first multi-reference graph, and use the 3D driving key points and 2D driving key points of the selected video frames as the 3D reference key points and 2D reference key points of the first multi-reference graph;
[0049] And, extract the driving audio features of each audio frame.
[0050] Step 2, calculate the lip movement style features of the speaker in the video based on the 3D driving key points of all video frames.
[0051] Step 3, combine the lip movement style features of the speaker, the driving audio features of each audio frame, the 3D reference key points and 2D reference key points of the first multi-reference graph, and use the audio-to-key point module to predict the fine 2D driving key points; the number of frames corresponding to the fine 2D driving key points is the same as and corresponds one by one to the number of audio frames.
[0052] Step 4, randomly select some video frames from the original video frames to form the second multi-reference graph, encode them to obtain the second multi-reference graph feature map, and use the same method to obtain the 2D reference key points of the second multi-reference graph.
[0053] Step 5: Combine the second reference image and its 2D reference key points, and the refined 2D driving key points predicted by the audio-to-key points module, and use the key points-to-video module to generate lip-sync frames.
[0054] Step 6: Train the audio-to-key points module and the key points-to-video module.
[0055] Step 7: Given the driving audio and the original video, extract the lip movement style features of the speaker in the video from the original video or directly specify the lip movement style features of the speaker, and use the trained audio-to-key points module and key points-to-video module to generate a synthetic video that is lip-synced with the driving audio, thus completing the lip-sync task.
[0056] In the above Step 1, the 2D and 3D key point information of the face in each video image of the continuous video frames is extracted through a pre-trained face key point detection tool. An optional method is as follows: Use a pre-trained 3D face key point detection tool including but not limited to Mediapipe to perform 3D key point detection on the face part of each video image to obtain the 3D driving key point coordinates and the camera estimation matrix; perform 2D projection on the 3D driving key point coordinates through the camera estimation matrix to obtain the 2D driving key point coordinates. Here, each video frame image corresponds to a set of 3D driving key points and a set of 2D driving key points.
[0057] Similarly, the driving audio features of each frame of the continuous audio frames are extracted through a pre-trained audio feature extraction tool. An optional pre-trained audio feature extraction tool is HuBERT.
[0058] The present invention also introduces multiple reference frames, which are extracted from the video frames in the input audio-visual data, and the extraction method is random extraction. The 3D driving key points and 2D driving key points corresponding to the selected video frames are used as the 3D reference key points and 2D reference key points of the first multi-reference image. Here, each reference image corresponds to a set of 3D reference key points and a set of 2D reference key points.
[0059] In the above Step 2, an optional method for calculating the lip movement style features based on the 3D driving key points is as follows:
[0060] Use the lip key points in the 3D driving key points of the speaker in the video to extract and estimate the quantified lip movement style features, including two speaking style attributes: lip opening amplitude and lip speed. Specifically, use the average value of the difference in the y coordinates of the upper lip and lower lip key points to measure the lip opening amplitude, and use the first-order average difference in time series of the absolute value of the difference in the y coordinates of the upper lip and lower lip key points to measure the lip speed, that is:
[0061]
[0062]
[0063] Among them, and represent the lip opening amplitude and lip speed, is the difference in the y - coordinates of the upper - lip and lower - lip key points at the current time step in the given video, and is the total number of time steps in the given video. The lip opening amplitude and lip speed together constitute the speaking - style attribute of the speaker. The lip - movement style features calculated here correspond to the entire audio - video segment.
[0064] During model training, the real lip - movement style features are extracted as input according to the above - mentioned method; during model inference, the user can use the lip - movement style features extracted from the original video or freely specify the values of the lip - movement style features. When the lip - movement style features extracted from the original video are used in the model inference stage, the lip movement style in the generated lip - synchronized video is consistent with that in the input video; since the lip - movement style features obtained through the above calculation are a scalar, the style of the generated lip - synchronized video can be modified by the values of the lip - movement style features freely specified by the user.
[0065] In step 3 above, the audio - to - key - point module includes a sparse 3D - driven key - point predictor and a fine 2D - driven key - point predictor. Among them, the sparse 3D - driven key - point predictor takes the driving audio features of each audio frame, the lip - movement style features of the speaker, and the 3D reference key points of the first multi - reference map as input and outputs the predicted sparse 3D - driven key points; the fine 2D - driven key - point predictor takes the 2D reference key points of the first multi - reference map and the sparse 3D - driven key points as input and outputs the predicted fine 2D - driven key points.
[0066] Here, in the sparse 3D - driven key - point predictor, the 3D reference key points of the first multi - reference map are encoded and extended to the same dimension as the driving audio features of each audio frame. Therefore, the number of frames of the predicted sparse 3D - driven key points is the same as the number of audio frames, that is, each audio frame corresponds to one frame of sparse 3D - driven key points, and the so - called one frame of sparse 3D - driven key points refers to a set of key points corresponding to one frame of video image. Similarly, in the fine 2D - driven key - point predictor, the 2D reference key points of the first multi - reference map are encoded and extended to the same dimension as that after the sparse 3D - driven key - point encoding embedding. Therefore, the number of frames of the predicted fine 2D - driven key points is also the same as the number of audio frames, that is, each audio frame corresponds to one frame of fine 2D - driven key points, and the so - called one frame of fine 2D - driven key points refers to a set of key points corresponding to one frame of video image.
[0067] Through the multi - level model structure from sparse to fine, the task difficulty of the model's one - step generation result is reduced, making the training effect more stable.
[0068] As Figure 2 shown, in a specific implementation of the present invention, the sparse 3D driving key point predictor includes a reference key point encoder, a lip opening amplitude embedding layer, a lip speed embedding layer, an audio encoder, and a driving key point decoder; the calculation process of the sparse 3D driving key point predictor is as follows:
[0069] 3-1a) The reference key point encoder encodes the 3D reference key points of the first multi-reference graph to obtain the encoded 3D reference key point features;
[0070] 3-2a) The lip opening amplitude embedding layer and the lip speed embedding layer respectively encode and embed the lip opening amplitude and lip speed in the lip movement style features of the speaker to obtain the lip movement style embedding features composed of the lip opening amplitude embedding and the lip speed embedding;
[0071] In this embodiment, the lip opening amplitude embedding is obtained by multiplying a scalar by a learnable vector, that is, f_amp = s_amp × e_amp, where s_amp is the scalar value of the lip opening amplitude, e_amp is the learnable vector of the lip opening amplitude embedding layer, and f_amp is the lip opening amplitude embedding; the calculation method of the lip speed embedding is the same and will not be elaborated.
[0072] 3-3a) The audio encoder encodes the driving audio features of each audio frame to obtain the audio encoded features;
[0073] 3-4a) The encoded 3D reference key point features and the lip movement style embedding features are extended to the same length as the audio encoded features, and then connected with the audio encoded features to obtain the sparse 3D driving key point features; here, the number of frames of the sparse 3D driving key point features is the same as the number of audio frames; here, the extension method adopts the repeated padding method, which is a well-known technique in the art;
[0074] 3-5a) The sparse 3D driving key point features are decoded by the driving key point decoder to obtain the predicted sparse 3D driving key points. The sparse 3D driving key points mainly include the lip key points and a sparse part of the face contour points.
[0075] In this embodiment, the key point decoder can adopt a conventional decoder structure, for example, it can be a stack of N residual modules. Each residual module sequentially includes a layer normalization layer, a 1D convolutional layer, a PReLu activation function, a final 1D convolutional layer, and a residual connection from the input to the output.
[0076] As Figure 3As shown, in a specific implementation of the present invention, the fine 2D driving key point predictor includes a reference key point encoder, a linear layer, and a driving key point decoder; the calculation process of the fine 2D driving key point predictor is as follows:
[0077] 3-1b) The reference key point encoder encodes the 2D reference key points of the first multi-reference graph to obtain the encoded 2D reference key point features;
[0078] 3-2b) The predicted sparse 3D driving key points are embedded into sparse 3D driving key point embeddings by the linear layer;
[0079] 3-3b) The encoded 2D reference key point features are extended to the same dimension as the sparse 3D driving key point embeddings and connected to the sparse 3D driving key point embeddings to obtain fine 2D driving key point features; the number of frames of the fine 2D driving key point features is the same as the number of audio frames;
[0080] 3-4b) The fine 2D driving key point features are decoded by the driving key point decoder to obtain the predicted fine 2D driving key points.
[0081] There is no association between the second multi-reference graph in step 4 above and the first multi-reference graph in step 1, and the two can be different. The method for obtaining the 2D reference key points of the second multi-reference graph is the same as that of the first multi-reference Figure 1 Similarly, directly use the 2D driving key points corresponding to the selected video frames as the 2D reference key points of the second multi-reference graph. Here, each reference graph corresponds to a set of 2D reference key points.
[0082] In step 5 above, an optional implementation process is as follows:
[0083] 5-1) According to the 2D reference key points of each frame of the second reference graph and the fine 2D driving key points of the current frame, warp the feature map of each frame of the second reference graph into a feature map consistent with the driving key points to generate a warped feature map of multiple reference graphs, that is, a multi-warped feature map.
[0084] In a specific implementation of the present invention, a method for obtaining the multi-warped feature map is: including but not limited to implementing it using the dense action hourglass network disclosed in the paper "Thin-Plate Spline Motion Model for Image Animation", taking the 2D reference key points and their feature maps of each frame of the reference graph and any frame of fine 2D driving key points as inputs, and taking the warped feature map of each frame of the reference graph as the output. For convenience, the warped feature map of multiple reference graphs is simply referred to as the multi-warped feature map.
[0085] 5-2) Calculate the multi-scale aggregation ratio of the multi-distortion feature map by using the fine 2D driving key points of the current frame, the 2D reference key points of all the second reference images, and the multi-distortion feature map. The multi-scale aggregation ratio includes a frame-scale ratio and a pixel-scale ratio; aggregate the texture information of the multi-distortion feature map by using the multi-scale aggregation ratio to obtain an aggregated feature map.
[0086] In a specific implementation of the present invention, the calculation process of the frame-scale ratio includes:
[0087] Calculate the query matrix , key-value matrix , and value matrix corresponding to each frame of the second reference image by using the selected fine 2D driving key points and the 2D reference key points of all the second reference images:
[0088]
[0089]
[0090]
[0091] Among them, represents three linear transformation matrices, and respectively represent the selected fine 2D driving key points and the 2D reference key points of the i-th frame of the second reference image;
[0092] Calculate the cross-attention value of each frame of the reference image according to the query matrix, key matrix, and value matrix, and obtain the frame-scale ratio after softmax processing. Here, is a set of numbers between 0 and 1, which is consistent with the number of frames of the second reference image.
[0093] The pixel-scale ratio is obtained by calculating the class cross-attention of adaptive instance normalization (AdaIN). The process includes:
[0094] Calculate the query matrix , key-value matrix , and value matrix corresponding to each frame of the reference image:
[0095]
[0096]
[0097]
[0098] Among them, is a modulation convolutional layer, is the distorted feature map of the i-th frame of the reference image, is a convolution kernel, which is obtained from reference key points.
[0099] Expand and to the same shape as , and calculate the cross-attention value of each frame of the reference image according to the expanded query matrix, key matrix, and value matrix. After softmax processing, obtain the pixel-scale ratio . Here, is a set of numbers between 0 and 1, which is consistent with the number of frames of the second reference image. This ratio has different values among different pixels, reflecting the local similarity between the driving key points and the pixel region at the pixel-level fine-grained scale.
[0100] The process of aggregating the texture information of the multi-warped feature maps using the multi-scale aggregation ratio to obtain the aggregated feature map is as follows:
[0101]
[0102] Among them, represents the mixed multi-scale aggregation ratio, which is consistent with the number of frames of the second reference image; represents the frame-scale ratio, represents the pixel-scale ratio, represents the pre-set frame weight, and in this embodiment, is used.
[0103] The information of multiple feature maps is aggregated into the information of a single feature map by the multi-scale attention mixer to reduce the subsequent calculation amount. The aggregation of the multi-warped feature maps is to perform weighted averaging on the pixels of each feature map according to a specific ratio, and this aggregation ratio is obtained by mixing the frame-scale ratio and the pixel-scale ratio, so it is called multi-scale.
[0104] 5-3) Decode the aggregated feature map to obtain the generated image, and fuse the lower half face of the speaker in the generated image with the original frame image to obtain the lip-sync frame image of the current frame;
[0105] In a specific implementation of the present invention, the fusion method of the lower half face of the speaker in the generated image and the original frame image is as follows:
[0106] Use the key points related to the lower half face in the original video image to form a convex hull, and the convex hull area is regarded as the hard lower half face mask . In the boundary area of the convex hull, use a Gaussian model with a certain intensity for feathering to obtain the smooth lower half face mask . Finally, use the smooth lower half face mask to mix the original frame image and the generated image:
[0107]
[0108] Among them, is the mixed result image, i.e., the final lip-sync frame image; and are the generated image and the original image respectively.
[0109] 5-4) Repeat steps 5-1) to 5-3), and use the next-frame fine 2D driving key points to generate the next-frame lip-sync frame image until all the frames of the fine 2D driving key points are traversed.
[0110] In the above step 6, an optional way to train the audio-to-key-points module and the key-points-to-video module is as follows:
[0111] For the training of the audio-to-key-points module, the calculation formula of the loss function is:
[0112]
[0113] where is the loss function of the audio-to-key-points module, is the L1 loss between the true value of the face 2D driving key points and the predicted fine 2D driving key points, is the L1 loss between the true value of the face 3D driving key points and the predicted sparse 3D driving key points, is the L1 loss between the true value and the predicted value of the lip movement style feature.
[0114] Here, the true values of the 3D driving key points and the 2D driving key points are the 3D driving key points and the 2D driving key points of each video frame image obtained through step 1; as previously introduced, the number of frames of the predicted sparse 3D driving key points and the predicted fine 2D driving key points is the same as the number of audio frames, and the audio and video in the training stage correspond to each other; the true value of the lip movement style feature is calculated through step 2, and the predicted value of the lip movement style feature is calculated using the same method as step 2 with the predicted sparse 3D driving key points.
[0115] For the training of the key-points-to-video module, the loss function includes two parts: the whole image and the lip region.
[0116] For the whole image, the calculation formula of the loss function is:
[0117]
[0118] where is the whole-image loss function of the key-points-to-video module, is the mean squared error loss between the generated image and the original frame image, is the mean squared error loss between the generated image and the original frame image in the perceptual space of the pre-trained model VGG-19.
[0119] Introduce additional loss enhancement for the lip region. Extract the lip region from the full image, and sample the same method to calculate the mean square error loss between the lip region of the generated image and the lip region of the original frame image. It is the mean square error loss between the lip region of the generated image and the lip region of the original frame image in the perceptual space of the pre-trained model VGG-19. Denote the lip region loss function from the key points to the video module as .
[0120] The complete training loss function from the key points to the video module is:
[0121]
[0122] where is the loss function from the key points to the video module.
[0123] In the above step 7, by given a driving audio and an original video, extract the lip movement style features of the speaker in the original video or directly specify the lip movement style features of the speaker, and use the trained audio-to-key points module and key points-to-video module to generate a synthetic video synchronized with the lip shape of the driving audio to complete the lip synchronization task.
[0124] Since the lip movement style feature is a scalar, the style of the generated lip-synchronized video can be modified by the value of the lip movement style feature freely specified by the user.
[0125] The above method will be applied to the following embodiments to demonstrate the technical effects of the present invention. The specific steps in the embodiments will not be elaborated.
[0126] The present invention conducts experiments on the mainstream lip synchronization task test set HDTF (High-Definition Talking Face). 20 random samples in the test set are selected as the test sample set for the comparative experiment. In order to objectively evaluate the performance of the present invention, the present invention is not trained on the HDTF dataset distribution.
[0127] This comparative test adopted the mainstream quantitative metrics for the lip-sync task, including metrics in multiple dimensions such as image quality, key-point reconstruction accuracy, and person identity maintenance. Specifically, it included PSNR, SSIM, LPIPS, FID, CSIM, and LipLDM. Peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) were used to evaluate the similarity between the generated video and the ground-truth video. Perceptual path length (LPIPS) and Frechet Inception distance (FID) were used to evaluate the distance in the latent feature space, which represents the generated visual quality. To better evaluate the ability to maintain identity, the cosine similarity (CSIM) of the features extracted by ArcFace was also evaluated. The lip key-point distance (LipLDM) was calculated to measure audio-visual lip synchronization. It should be noted that the lip key-points of the generated results and the real video extracted by Mediapipe were used here for the calculation of LipLDM, and the distance metric used was the L1 distance.
[0128] The baseline models for comparison included MakeItTalk, Wav2Lip, Wav2Lip-GAN, PC-AVS, VideoRetalking, IP-LAP, and DINet. Similar to the method of the present invention, IP-LAP and DINet utilized multiple reference frames to enhance the prior knowledge of the face. The official open-source implementations and pre-trained models of these baselines were used during inference. Some baseline models such as PC-AVS only generated cropped or resized speaker videos. For these methods, the preprocessed source video was regarded as the ground truth.
[0129] According to the steps described in the specific implementation manner, the experimental results obtained are shown in Table 1.
[0130] Table 1: Lip-sync comparative test results obtained by the present invention for the HDTF dataset
[0131]
[0132] As can be seen from Table 1, overall, the present invention trained on publicly accessible datasets achieved better performance in evaluation results such as PSNR, SSIM, LPIPS, FID, and CSIM, as shown in Table 1. IP-LAP uses an audio-to-keypoint network trained on two large datasets, so it performs better on LipLDM. Due to the design of the keypoint-to-video module and the inference reference selection strategy, the FID and LPIPS scores of the present invention are significantly better than previous methods. For example, the FID of the present invention is about 1.45, while the FIDs of the baseline models are all above 4.9, which indicates the high visual quality of the keypoint-to-video module. In addition, the CSIM of the present invention reaches a level higher than 0.95, while the best baseline model is only about 0.92, showing the high identity preservation ability of the audio-to-keypoint module of the present invention.
[0133] Based on the same inventive concept, in this embodiment, a lip synchronization system based on multiple reference frames and style controllability is further provided, and this system is used to implement the above embodiment. The following terms such as "module" and "unit" can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.
[0134] In this embodiment, a lip synchronization system based on multiple reference frames and style controllability includes:
[0135] An audio-visual data preprocessing module, which is used to obtain the speaker video data and the driving audio data, extract the driving audio features of each audio frame and the driving keypoints of each video frame; randomly select two groups of video frames from the video as the first multi-reference map and the second multi-reference map respectively, and regard the driving keypoints of the reference map as reference keypoints;
[0136] An audio-to-keypoint module, which is used to extract the lip movement style features of the speaker from the video, and combine the driving audio features and the reference keypoints of the first multi-reference map to predict fine keypoints;
[0137] A keypoint-to-video module, which is used to combine the second multi-reference map and its reference keypoints, and the fine keypoints predicted by the audio-to-keypoint module to generate lip synchronization frames;
[0138] A training module, which is used to calculate the loss according to the prediction results of the audio-to-keypoint module and the generation results of the keypoint-to-video module and update the module;
[0139] The lip synchronization generation module is used to extract the speaker's lip movement style features from the original video or directly give the speaker's lip movement style features for the given driving audio and original video, and use the trained audio to key point module and key point to video module to generate a synthetic video that is lip-synchronized with the given driving audio to complete the lip synchronization task.
[0140] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be repeated here. The system embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0141] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.
[0142] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. A lip synchronization method based on multiple reference frames and controllable style, characterized in that: The following steps are involved: S1, obtaining speaker video data and driving audio data, extracting driving audio features of each audio frame and driving key points of each video frame; The driving key points include 3D driving key points and 2D driving key points; S2, randomly selecting two groups of video frames from the video as the first multi-reference image and the second multi-reference image, and taking the driving key points of the reference image as the reference key points; The reference key points include 3D reference key points and 2D reference key points; S3, extracts the speaker's lip movement style features from the video, combines the driving audio features and the reference key points of the first multi-reference graph, and uses the sparse to fine audio to key point module to predict the fine key points; The refined key points predicted by the audio to key point module are refined 2D driven key points; The lip movement style feature is a scalar, which is composed of lip width and lip speed. The method for extracting the speaker's lip movement style feature from the video is as follows: Select lip key points from the 3D driving key points of the speaker in each frame of the video, including upper lip key points and lower lip key points; The average value of the y-coordinate difference between the upper lip key points and the lower lip key points is used to represent the lip width; The lip velocity is represented by the time series first-order mean difference of the absolute value of the y-coordinate difference between the upper lip key points and the lower lip key points; S4, combining the second multi-reference image and its reference key points, and the refined key points predicted by the audio-to-key point module, and introducing a multi-scale aggregation ratio in the key point-to-video module to generate a lip sync frame; Step S4 includes: S4-1, warping the feature map of the second reference image of each frame according to the 2D reference key points of the second reference image of each frame and the fine 2D driving key points of the current frame to generate a multi-warped feature map; S4-2, using the fine 2D driving key points of the current frame, the 2D reference key points of all the second reference images, and the multi-distortion feature map, calculating a multi-scale aggregation ratio of the multi-distortion feature map, wherein the multi-scale aggregation ratio includes a frame scale ratio and a pixel scale ratio; and using the multi-scale aggregation ratio to aggregate texture information of the multi-distortion feature map to obtain an aggregated feature map; S4-3, decoding the aggregated feature map to obtain a generated image, using the driving key points of the original frame image to generate a smooth lower half face mask, and using the smooth lower half face mask to fuse the lower half face of the speaker in the generated image with the original frame image to obtain a lip sync frame image of the current frame; S4-4, repeating S4-1 to S4-3, using the next frame of fine 2D driving key points to generate the next frame of lip sync frame image, until all the frames of fine 2D driving key points are traversed; S5, calculating the loss and updating the module according to the prediction results of the audio to key point module and the generation results of the key point to video module; S6, given the driving audio and the original video, extract the speaker's lip movement style features from the original video or directly give the speaker's lip movement style features, use the trained audio to key point module and key point to video module to generate a synthetic video that is lip-synced with the given driving audio, and complete the lip synchronization task.
2. The lip synchronization method based on multiple reference frames and controllable style according to claim 1, characterized in that: The audio to key point module includes: a sparse 3D driving keypoint predictor, which takes the driving audio features of each audio frame, the speaker's lip movement style features, and the 3D reference keypoints of the first multi-reference graph as input, and outputs predicted sparse 3D driving keypoints; A refined 2D driving key point predictor takes the 2D reference key points of the first multiple reference images and the sparse 3D driving key points as inputs, and outputs predicted refined 2D driving key points.
3. The lip synchronization method based on multiple reference frames and controllable style according to claim 2, characterized in that: The sparse 3D driving key point predictor includes a reference key point encoder, a lip width embedding layer, a lip velocity embedding layer, an audio encoder and a driving key point decoder. The calculation process includes: The reference key point encoder encodes the 3D reference key points of the first multi-reference image to obtain encoded 3D reference key point features; The lip width embedding layer and the lip velocity embedding layer encode and embed the lip width and lip velocity of the speaker's lip movement style features respectively, and obtain the lip movement style embedding features composed of lip width embedding and lip velocity embedding; The audio encoder encodes the driving audio feature of each audio frame to obtain an audio coding feature; The encoded 3D reference key point features and lip motion style embedding features are expanded to the same dimension as the audio encoding features, and then connected with the audio encoding features to obtain sparse 3D driving key point features; the number of frames of the sparse 3D driving key point features is consistent with the number of audio frames; The sparse 3D driving keypoint features are decoded by the driving keypoint decoder to obtain the predicted sparse 3D driving keypoints.
4. The lip synchronization method based on multiple reference frames and controllable style according to claim 2, characterized in that: The refined 2D driving key point predictor includes a reference key point encoder, a linear layer and a driving key point decoder, and the calculation process includes: The reference key point encoder encodes the 2D reference key points of the first multi-reference image to obtain encoded 2D reference key point features; The predicted sparse 3D driving keypoints are embedded into sparse 3D driving keypoint embeddings by a linear layer; The encoded 2D reference key point features are expanded to the same dimension as the sparse 3D driving key point embedding, and then concatenated with the sparse 3D driving key point embedding to obtain refined 2D driving key point features; the number of frames of the refined 2D driving key point features is consistent with the number of audio frames; The refined 2D driving keypoint features are decoded by the driving keypoint decoder to obtain the predicted refined 2D driving keypoints.
5. The lip synchronization method based on multiple reference frames and controllable style according to claim 1, characterized in that: The frame scale ratio is obtained by calculating the cross-attention value between the selected fine 2D driving key points and the 2D reference key points of the second reference image; the pixel scale ratio is obtained by calculating the adaptive instance-normalized class cross-attention between the selected fine 2D driving key points and the 2D reference key points of the second reference image.
6. The lip synchronization method based on multiple reference frames and controllable style according to claim 1, characterized in that: The training loss of the audio to keypoint module includes the loss between the true value of the 2D driving keypoint and the predicted fine 2D driving keypoint, the loss between the true value of the 3D driving keypoint and the predicted sparse 3D driving keypoint, and the loss between the lip movement style features calculated using the true value of the 3D driving keypoint and the predicted sparse 3D driving keypoint respectively; The training loss of the key point to video module includes two parts: full image loss and lip area loss, including the mean square error loss between the generated image and the original frame image, the mean square error loss between the image features of the generated image and the original frame image, the mean square error loss between the lip area of the generated image and the lip area of the original frame image, and the mean square error loss between the image features of the lip area of the generated image and the lip area of the original frame image.
7. A lip synchronization system based on multiple reference frames and controllable style, used to implement the method of claim 1; characterized in that: The system comprises: An audio and video data preprocessing module is used to obtain speaker video data and driving audio data, extract driving audio features of each audio frame and driving key points of each video frame; randomly select two groups of video frames from the video as the first multi-reference image and the second multi-reference image, and regard the driving key points of the reference image as reference key points; An audio-to-keypoint module is used to extract the speaker's lip movement style features from the video, and predict fine keypoints by combining the driving audio features and the reference keypoints of the first multi-reference graph; a keypoint-to-video module for combining the second multiple reference images and their reference keypoints, and the refined keypoints predicted by the audio-to-keypoint module, to generate a lip sync frame; A training module, which is used to calculate the loss and update the module based on the prediction results of the audio to key point module and the generation results of the key point to video module; The lip synchronization generation module is used to extract the speaker's lip movement style features from the original video or directly give the speaker's lip movement style features for the given driving audio and original video, and use the trained audio to key point module and key point to video module to generate a synthetic video that is lip-synchronized with the given driving audio to complete the lip synchronization task.
Citation Information
Patent Citations
Virtual human image video generation method, system and device and storage medium
CN113192161A
Digital character audio and video data processing system and method and live broadcast system
CN119562142A