Speaker video generation method jointly driven by key point and expression action unit
Through the speaker video generation method driven by key points and expression action units, combined with time consistency-driven video quality enhancement and video super-resolution processing, the accuracy and quality problems of speaker video generation in low bandwidth scenarios are solved, and high-quality and time consistency video reconstruction is achieved.
Patent Information
- Application Number
- CN202510126961.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-06-03
AI Technical Summary
The existing speaker video generation technology is difficult to achieve high-quality and accurate posture and expression reconstruction in low bandwidth scenarios, and the traditional method relies on a single driver source, resulting in unsatisfactory results in accuracy and quality.
A method of video generation of speakers driven by key points and expression action units is proposed. By combining a small amount of multiple driving information, the accuracy of generated video expressions and posture generation is improved, and the video quality driven by time consistency is enhanced by the network and the video super-resolution processing of prior drives to improve picture perception quality and time consistency.
It realizes the generation of reconstructed videos with more accurate posture and expression in low bandwidth scenarios, significantly improving the quality and time consistency of the video, reducing the amount of data of driver information, and is suitable for various application scenarios such as video conferencing and virtual reality.
Smart Images

Figure CN120088375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a method for generating a speaker video jointly driven by key points and expression action units. Background Art
[0002] The technology of generating Talking Head videos is a highly potential innovative technology in the fields of computer vision and multimedia, which brings new possibilities to many application scenarios. This technology takes a deep neural network as the core, and through learning the human facial feature representation and dynamic patterns from a large number of real human speaking video data, it realizes generating a speaker video with realistic dynamic expressions and mouth shapes from a single portrait image according to various driving information.
[0003] The selection of driving information is crucial for the generation of speaker avatars. Some methods use facial coordinates as driving information. Facial landmark points can concisely and effectively express rough facial movements, but there are certain limitations in using sparse facial landmark points. Due to their sparsity, they can only provide limited information and are difficult to comprehensively and accurately express rich and diverse expression details. When sparse facial landmark points are used as the driving source to drive video generation, the generated video performs poorly in terms of reconstruction fidelity and cannot achieve the desired effect. This is mainly because sparse facial landmark points are difficult to capture key information such as subtle muscle movements, complex emotional changes, and various expression combinations in facial expressions, thus affecting the quality and realism of the finally generated video, unable to provide sufficient information support for the subsequent generation process, making the generated video not delicate and accurate enough in expression, difficult to completely reproduce the original expression information, resulting in large information loss and deviation in reconstructing the facial image or video, and thus limiting its performance in high-quality speaker avatar generation tasks. Inspired by this, many subsequent works use dense motion flows for speaker avatar generation, such as incorporating the motion field into the generation of dynamic human behaviors.
[0004] Some methods select speech audio as the driving information source. From the perspective of advantages, this method can efficiently capture lip movements and rhythms, presenting a natural and realistic effect, and has a positive role in generating natural and vivid speaker videos. However, speech that cannot contain posture information will bring certain challenges to the posture generation link, making it difficult to effectively control the posture generation process. In addition, the noise problem in the audio will have an adverse effect on the quality of the final generated result and interfere with the quality of the generated video. In existing research, although some audio-driven methods have shown certain advantages, they also face the above-mentioned problems of posture drive and noise interference, and further optimization and improvement are needed to improve the performance of speaker avatar generation technology in speech drive. Some methods use the embedded information obtained from the video as the driving source. This type of driving source shows unique advantages. It allows the network to learn adaptively, and then it can mine and master the most relevant potential features.
[0005] The most advanced speaker generation methods are described from two dimensions, namely 2D-based methods and 3D-based methods. Among them, 2D-based methods aim to learn and obtain 2D representation information of the face by analyzing the input image or video. Such representation covers key elements such as appearance and action. In practice, these methods often rely on the powerful capabilities of generative adversarial networks (GANs) to achieve realistic and diverse speaker video generation. Compared with 3D graphics-based models, 2D-based methods show certain advantages in the synthesis of details such as hair and beards, and can process these details more finely, so that they show a high degree of detail restoration in the generated results. However, 2D-based methods face many challenges in generating natural and consistent facial expressions. In addition, 2D-based methods are prone to artifacts, and the quality of the generated images will be reduced in the case of significant actions, which undoubtedly has a negative impact on the quality of the final generated speaker video. 3D-based methods focus on learning 3D representations of the face from the input image, and their representation content mainly involves key elements such as geometry and texture. In order to achieve the generation of speaking videos, such methods usually rely on Neural Radiance Fields (NeRF) or other related 3D models, and use the powerful 3D structure and appearance capture capabilities of these models to generate more natural and realistic speaking videos. Despite the above advantages, 3D-based methods also have their own limitations, the most notable of which is their high computational cost, which may impose a heavy burden on computing resources and affect their promotion and use in practical applications.
[0006] Among them, through the adversarial training of the generator and discriminator, the generative adversarial network can enable the generator to generate highly realistic video frames, while the autoencoder architecture can perform encoding and decoding operations on video frames, facilitating the operation of latent space vectors to generate the required videos. In terms of multi-modal information fusion, combined with speech driving, speech features are combined with facial information, and a deep neural network is used to map speech signals into facial actions, providing rich clues for video generation; text-driven video generation fuses the semantic vectors obtained by converting natural language processing technology with facial features, and can generate videos that match the text, showing advantages in fields such as virtual customer service. In the part of motion modeling and pose estimation, modeling based on expression action units can better control and generate facial expressions, while head pose estimation can ensure the natural and coherent poses of the generated videos.
[0007] Referring to the OSFV method, it is based on the GAN generative adversarial network and uses 3D key point information to drive the generation of speaker videos. In this method, in the driving information extraction stage, the encoder group extracts the standard identity key point information, head rotation and translation values, and expression estimation values in the driving frame and the original frame, and integrates the calculations into the driving frame key points (K driving) and the original frame key points (K source). This part is collectively referred to as key point extraction, and the key point extraction in the method of the present invention refers to this method. Subsequently, in the driving generation stage, the input original frame, the original frame key points, and the driving frame key points are input. The 3D optical flow field of the driving frame key points and the original frame key points is calculated, the appearance features of the original frame are warped, and the result frame consistent with the pose and expression of the driving frame is decoded, realizing the generation of speaker videos driven by key points. However, since it is difficult to balance the generation of pose and fine expressions using a single driving source, there is still room for improvement in its reconstruction accuracy, especially for the reconstruction of fine expressions. In addition, artifacts, blurred lines, chromaticity and brightness loss, and frame-to-frame non-smoothness are prone to occur during reconstruction, and its visual perception quality is still insufficient.
[0008] Therefore, in a bandwidth-constrained environment, traditional video conferencing transmission methods face great difficulties, because the bandwidth required to transmit high-resolution, high-frame-rate complete video data is large, which can easily lead to problems such as video freeze, delay, and image quality degradation. And the problem of relying on speaker video generation technology. Moreover, most of the existing methods currently focus on optimizing the generation quality, but there are relatively few targeted studies in low-bandwidth scenarios. Among them, some methods use optical flow as the driving source. Although they show certain advantages in some aspects, they are not suitable for low-bandwidth scenarios due to factors such as their large data volume. In addition, most methods often rely on only a single driving source for driving. The limitations of this single driving source lead to unsatisfactory generation results in accuracy and quality, which is difficult to meet actual needs. Therefore, it is necessary to provide a method for generating speaker video generation by multi-information joint driving, which can extract driving information from the original video frame, and only transmit the driving information of the first frame and subsequent frames, so as to reconstruct the speaker video at the receiving end, and is suitable for low-bandwidth scenarios, so as to achieve the purpose of using less driving information and generating a more accurate reconstructed video with more accurate posture and expression, so as to provide a more effective, accurate and high-quality technical solution for speaker video generation in a low-bandwidth environment. Summary of the invention
[0009] In view of this, the purpose of the present invention is to propose a method for generating a speaker video jointly driven by key points and expression action units, which effectively improves the accuracy of generating video expressions and postures by combining a small amount of multiple driving information. As the first stage of the overall method of the present invention, it addresses the problem that a single driving source is difficult to achieve accurate posture and expression reconstruction, and solves the problem of accurate and fast driven generation of speaker videos.
[0010] The present invention also uses a video quality enhancement network driven by temporal consistency to perform global and local enhancement using spatiotemporal information, effectively improving the perceived quality of the picture and reducing quality fluctuations. At the same time, through innovative spatiotemporal joint adversarial learning and background fusion methods to process background jitter, the temporal consistency of the video is effectively improved. As the second stage of the overall method of the present invention, it solves the problem of low perceived quality of the generated video and improves the overall perceived quality of the speaker video.
[0011] In view of the problems of low video resolution and poor picture quality, the method of the present invention performs video super-resolution processing on the generated results, solving the problem of high-definition and high-quality programs. By integrating the three decoupled tasks of drive generation, quality enhancement, and video super-resolution into a comprehensive speaker video generation method, it realizes low-resolution recording, low-bandwidth transmission, and high-quality display, enabling more precise drive control, providing better and more accurate drive support for speaker avatar generation, showing greater potential in realizing natural and realistic avatar generation, and is expected to play an important role in future related research and practical applications.
[0012] To achieve the above object, the present invention provides the following technical solutions:
[0013] Based on the above object, in a first aspect, the present invention provides a method for generating a speaker video jointly driven by key points and expression action units, including the following steps:
[0014] 1.1 Keypoint and Au jointly drive generation stage (KAG):
[0015] Extract the key points and Action Units (AUs) information in the driving frame, and drive the original frame to generate a target frame of posture and expression through motion flow field estimation and dynamic feature transformation;
[0016] 1.2 Temporal-consistency-driven enhancement stage (TDE):
[0017] Perform global enhancement and key region enhancement on the initially generated frames, and perform quality correction and enhancement on the results generated frame by frame through multi-frame joint adversarial training and background fusion to obtain an enhanced video;
[0018] 1.3 Prior-driven video super-resolution stage (PDS):
[0019] Generate preliminary video frames in the low-resolution space, perform super-resolution processing on the enhanced video, and utilize facial prior information to improve the clarity of the video and the high-frequency part of the facial details.
[0020] As a further solution of the present invention, the joint driving information of key points and expression action units is extracted from the original frame (s) and the driving frame (d) using a key point extractor (K) implemented based on a convolutional neural network (CNN) tThe 3D key point information in ( ) and calculate the 3D motion flow information from the original frame to the driving frame through the motion field estimation network (M).
[0021] As a further solution of the present invention, extract the 3D appearance features of the original frame through the appearance feature extractor (F), and perform feature domain warping on the original frame features based on the motion flow information; input the AU information, key point information, and warped appearance features in the driving frame into the Dynamic Feature Transformation module to finely adjust the warped appearance features; use the Generative Adversarial Network (GAN) framework to generate the target frame (y with precise pose and expression), where w represents the warping generation stage and t represents the t-th frame. w,t )
[0022] As a further solution of the present invention, the joint driving generation stage of key points and expression action units includes the extraction of driving information, including the following steps:
[0023] Extract the 3D key point information and expression action unit (AUs) information in the driving frame (d). The expression action unit information vector is used to represent most facial expressions and is independent of individual identity;
[0024] Use the 3D key point detection network (K) based on the Convolutional Neural Network (CNN) to extract the 3D key points in the original frame (s) and the driving frame (d), and calculate the motion flow field (w) based on these key points, so as to realize the pose adjustment of the original frame;
[0025] Use the AU detector (A) to downsample and encode the driving frame image based on the multi-layer 2D convolutional network layer, and finally obtain the driving AU vector to represent the facial expression information in the driving frame and further fine-tune the expression.
[0026] As a further solution of the present invention, the joint driving generation stage of key points and expression action units includes driving and generation, including the following steps:
[0027] Use the appearance feature extractor (F) to extract the 3D appearance features (f from the source image (s), and warp and deform the features (f of the original frame with the motion flow field (w) through the motion flow field estimator (M) to generate features with a new pose; s ) s )
[0028] Input the AU information, key point information, and the deformed original frame features in the driving frame into the Spatial Feature Transform (SFT) module, and finely adjust the features through the Spatial Feature Transform layer to generate an accurate expression that conforms to the driving information;
[0029] Finally, the adjusted features are input into the generator to generate the target frame (y w,t ).
[0030] As a further solution of the present invention, the joint driving generation stage of key points and expression action units includes the training of the generation stage, which includes the following steps:
[0031] During the training process, two frames of images are randomly sampled from the same video: one frame is used as the source image (s), and the other frame is used as the driving image (d);
[0032] The following loss function is used for training:
[0033] Reconstruction loss: Calculate the pixel-level difference between the driving frame (d t ) and the generated frame (y w,t ), and optimize the reconstruction accuracy of the generated image;
[0034] Adversarial loss: Adopt a multi-scale discriminator for adversarial training to generate images highly similar to real images;
[0035] Perceptual loss: Use the VGG19 perceptual network to calculate the perceptual loss between the generated frame (y w,t ) and the driving image (d t ), and constrain the semantic and spatial consistency of the generated image;
[0036] Head pose loss: Calculate the difference in the head rotation angle between the generated frame and the driving frame, reduce the error between the pose prediction and the true value, and optimize the accuracy of the key point detector (K);
[0037] Expression loss: Calculate the mean square error loss between the AUs vectors of the driving frame (d t ) and the generated frame (y w,t ), constrain the accuracy of expression generation, and ensure the generation of more realistic facial details.
[0038] As a further solution of the present invention, the AU detector (A) realizes downsampling encoding through a multi-layer 2D convolutional neural network and generates the AUs vector in the driving frame (d); the appearance feature extractor (F) uses a 3D convolutional neural network to extract the appearance features of the original frame (s), and the extracted appearance features include the appearance, identity information, and background information of the person; the dynamic feature transformation module (SFT) performs feature adjustment based on the Spatial Feature Transform layer, and dynamically adjusts the intermediate features using key point and AUs information to generate accurate facial expressions; the generator generates the target frame (y w,t) where the key points and AUs information are input into the generator as prior conditions; the reconstruction loss, adversarial loss, perceptual loss, head pose loss, and expression loss jointly optimize the generator to improve the accuracy and naturalness of video generation; the head pose loss calculates the difference in head rotation angles between the generated frame and the driving frame to reduce the difference in head pose between the generated frame and the driving frame.
[0039] As a further solution of the present invention, the video quality enhancement stage driven by temporal consistency includes global enhancement, key region enhancement, multi-frame joint adversarial training, and background fusion. The steps are as follows:
[0040] For the target frame y w,t , take the adjacent generated frames y w,t-1 and y w,t+1 as auxiliary frames. After concatenating these three frames in the channel dimension, input them into the global enhancement network to repair the artifact regions in the target frame using the inter-frame information provided by the auxiliary frames, and improve the picture quality and temporal consistency;
[0041] Use a pre-trained lightweight facial key point extraction module to detect the key points in the mouth region, input the mouth region into the key region enhancement network to generate the enhanced mouth image I c and the mask M, and seamlessly integrate the enhanced mouth region into the overall image through the fusion formula;
[0042] Construct a temporal consistency discriminator T, input the spatio-temporal joint features of three consecutive frames into the discriminator, calculate the temporal consistency loss L t , and optimize the inter-frame consistency and stability of the generated video through adversarial training;
[0043] Extract the intersection region mask M of the background of the original frame s and the background of the enhanced generated image, and replace the background of the original frame with the background of the enhanced image to suppress the background jitter and flicker problems.
[0044] As a further solution of the present invention, the global enhancement network includes a downsampling module and multiple Res-Net modules, which are used to extract inter-frame joint features and capture image texture details through skip connections and residual learning mechanisms.
[0045] As a further solution of the present invention, the key region enhancement network includes:
[0046] A generation branch for generating the enhanced mouth image I c ;
[0047] A mask branch for generating the mask M to guide the generator to focus on specific regions and ensure smooth edges.
[0048] As a further solution of the present invention, the temporal consistency discriminator includes:
[0049] A 3D convolutional layer for extracting spatio-temporal joint features;
[0050] A feature processing block including a 2D convolutional layer, a normalization layer, and an activation layer for refining features;
[0051] A linear layer for calculating the temporal consistency loss.
[0052] As a further solution of the present invention, the background fusion network B is used to extract the mask M of the intersection area of the background of the original frame s and the background of the intermediate generated image that has undergone global and local enhancement, and the fusion is performed through the following formula: The background of the intermediate generated image is replaced with the background of the original frame s. For each frame in the generated sequence, the background is replaced with the background in the original frame, so as to obtain a video sequence with a fixed background, which can effectively suppress the background jitter and flicker problems.
[0053]
[0054] Replace the background of the intermediate generated image with the background of the original frame s For each frame in the generated sequence, replace the background with the background in the original frame, so as to obtain a video sequence with a fixed background, which can effectively suppress the background jitter and flicker problems.
[0055] As a further solution of the present invention, in the prior-driven video super-resolution stage (PDS), video generation is performed in the low-resolution space, and GFPGAN is used to perform super-resolution processing on the low-resolution output; the rich facial prior knowledge in the pre-trained StyleGAN2 is utilized to enhance the high-frequency part of the face details; before performing super-resolution processing, through the time-consistency-driven video quality enhancement stage, the initial generation result is corrected and enhanced.
[0056] In another aspect of the present invention, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the computer program is executed by the processor, it executes any one of the above-mentioned speaker video generation methods jointly driven by key points and expression action units according to the present invention.
[0057] In still another aspect of the present invention, a computer-readable storage medium is further provided, storing computer program instructions, and when the computer program instructions are executed, any one of the above-mentioned speaker video generation methods jointly driven by key points and expression action units according to the present invention is implemented.
[0058] Compared with the prior art, a speaker video generation method jointly driven by key points and expression action units proposed by the present invention has the following beneficial effects:
[0059] 1. Beneficial effects of the joint drive of key points and expression action units:
[0060] More accurate pose and expression generation: By jointly driving key points with expression action units (AUs), the present invention can more precisely control the generation of poses and expressions. Experimental results show that the present invention is superior to existing methods in terms of average landmark distance (ALD) and average action unit distance (AUD), demonstrating its advantage in pose and expression accuracy.
[0061] Decoupling pose and expression information: The present invention decouples expression information from pose driving and extracts AUs in a supervised manner, which not only improves the accuracy of expression generation but also enables key points to focus more on pose guidance, thereby enhancing the pose and expression accuracy of the overall generation result.
[0062] Reducing the amount of driving information: Compared with methods that rely on depth images or other complex driving information, the present invention only needs to transmit 15 three-dimensional key points and 17 AU values per frame, significantly reducing the data volume of driving information and being applicable to low-bandwidth scenarios.
[0063] 2. Beneficial effects of the time-consistency-driven video quality enhancement (TDE) stage:
[0064] Improving reconstruction fidelity: After being processed by the TDE stage, the generated video is significantly improved in terms of metrics such as L1 Loss, LPIPS, PSNR, and SSIM, indicating that it has a higher similarity to real images in terms of lighting, color, and structural texture details.
[0065] Reducing artifacts and blurring: The TDE stage significantly reduces the problems of artifacts, blurring, and color loss in the generated video through global enhancement and key-region enhancement, especially in large-scale motion scenarios, with a more realistic visual effect.
[0066] Enhancing video temporal consistency: Through multi-frame joint adversarial training (MFA), the TDE stage optimizes the spatio-temporal relationship between frames, significantly improving the temporal consistency of the video and reducing inter-frame jitter.
[0067] Improving background stability: The background fusion module effectively suppresses background jitter and flickering problems, enhancing the overall visual experience.
[0068] 3. Beneficial effects of key-region enhancement:
[0069] Improving detail texture: The key-region enhancement module focuses on important regions such as the mouth, significantly improving the generation quality of these regions, especially in details such as teeth and lips, with a more realistic generation effect.
[0070] Laying a foundation for super-resolution processing: After key-region enhancement, the generated video is clearer in terms of details and texture, providing high-quality input for subsequent super-resolution processing.
[0071] 4. Beneficial effects of the Multi-frame Continuous Adversarial (MFA) scheme:
[0072] Optimizing inter-frame consistency: The MFA scheme optimizes the spatio-temporal relationship between frames through adversarial learning, significantly reducing the inter-frame L1 loss and enhancing the temporal consistency of the video without affecting the fidelity of the overall pose.
[0073] Avoiding error accumulation: Compared with methods that rely on the enhanced results of the previous frame, the MFA scheme avoids the accumulation of generated errors, ensuring the stability and smoothness of the video sequence.
[0074] 5. Beneficial effects of the Background Blending (BB) scheme
[0075] Improving background stability: The background blending module significantly reduces background jitter and flickering problems by reusing the background of the original frame, enhancing the visual stability of the video.
[0076] Enhancing fidelity: Experiments on the HDTF dataset show that the background blending module significantly reduces the L1 loss and improves the fidelity of the results, especially in terms of the color and exposure of the background.
[0077] 6. Beneficial effects of the Prior-driven Video Super-Resolution (PDS) module
[0078] Enhancing resolution and detail performance: The PDS module uses facial prior knowledge to upscale low-resolution videos to high-resolution, significantly enhancing the detail performance of the videos and improving the visual quality.
[0079] Maintaining pose and expression accuracy: While upscaling the resolution, the PDS module maintains the accuracy of the pose and expression, further optimizing the visual experience of the viewers.
[0080] In summary, the present invention significantly improves the quality and temporal consistency of speaker video generation through joint driving by key points and expression action units, time-consistency-driven video quality enhancement, and prior-driven video super-resolution processing. Compared with the prior art, the present invention shows higher efficiency and better visual effects in low-bandwidth scenarios, while reducing the data volume of the driving information, and is applicable to various application scenarios such as video conferencing, virtual reality, and telemedicine.
[0081] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings
[0082] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the related art, the following will briefly introduce the drawings required for the description of the exemplary embodiments or the related art. The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0083] Figure 1 It is a flowchart of a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0084] Figure 2 It is a schematic diagram of the processing process in three stages in a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0085] Figure 3 It is a schematic diagram of the display of the generation results in three stages in a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0086] Figure 4 It is a flowchart of the stage of jointly driving key points and expression action units to generate in a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0087] Figure 5 It is a schematic diagram of the structure of the AU detector in a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0088] Figure 6 It is a flowchart of the dynamic feature transformation module in a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0089] Figure 7 It is a flowchart of the video quality enhancement network driven by temporal consistency in a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0090] Figure 8 It is a schematic diagram of the structure of the temporal consistency discriminator in a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0091] Figure 9 It is a schematic diagram of the background fusion network in a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention.
[0092] Figure 10 It is a schematic diagram of the qualitative comparison between a method for generating a speaker video by jointly driving key points and expression action units according to an embodiment of the present invention and the mainstream methods.
[0093] Figure 11 This is a schematic diagram for comparing the image results of the TDE module in the method for generating a speaker video jointly driven by key points and expression action units according to an embodiment of the present invention.
[0094] Figure 12 This is a schematic diagram for comparing the image results of the key area module in the method for generating a speaker video jointly driven by key points and expression action units according to an embodiment of the present invention.
[0095] Figure 13 This is a schematic diagram for comparing the image results of the PDS module in the method for generating a speaker video jointly driven by key points and expression action units according to an embodiment of the present invention. Detailed implementation manners
[0096] Next, in combination with the accompanying drawings and specific implementation manners, the present application will be further described. It should be noted that, on the premise of no conflict, the following-described embodiments or technical features can be arbitrarily combined to form new embodiments.
[0097] To make the purpose, technical solution and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail with reference to specific embodiments and the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0098] A method for generating a speaker video jointly driven by multiple information applicable to a low-bandwidth scenario according to the present invention will use less driving information to generate a reconstructed video with relatively accurate postures and expressions, and at the same time highly focus on the quality of the generated video, aiming to provide a more effective, accurate and high-quality technical solution for generating a speaker video in a low-bandwidth environment. The method for generating a speaker video jointly driven by key points and expression action units provided by the present invention solves the following problems:
[0099] (1) Aiming at the problem that it is difficult to achieve accurate posture and expression reconstruction with a single driving source, the present invention proposes a method for generating a speaker video jointly driven by key points and expression action units. By jointly using a small amount of multiple driving information, the accuracy of generating the expressions and postures of the generated video is effectively improved. As the first stage of the overall method of the present invention, it solves the problem of accurately and quickly driving the generation of a speaker video.
[0100] (2) Aiming at the problem of low perceptual quality in the generated video, the present invention proposes a video quality enhancement network driven by temporal consistency, which uses spatio-temporal information for global and local enhancement, effectively improving the perceptual quality of the picture and reducing quality fluctuations. At the same time, through innovative spatio-temporal joint adversarial learning and processing background jitter through background fusion method, the temporal consistency of the video is effectively improved. As the second stage of the overall method of the present invention, it solves the problem of improving the overall perceptual quality of the speaker video.
[0101] (3) Aiming at the problems of low video resolution and poor picture quality, the present invention performs video super-resolution processing on the generated results, and solves the problem of high-definition and high-quality programs. The present invention integrates the three decoupled tasks of driving generation, quality enhancement, and video super-resolution into a comprehensive speaker video generation method, so as to achieve low-resolution recording, low-bandwidth transmission, and high-quality display.
[0102] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0103] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all the contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.
[0104] Next, in conjunction with the accompanying drawings, some embodiments of the present application will be described in detail. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0105] See Figure 1 and Figure 2 As shown, the embodiments of the present invention provide a method for generating a speaker video jointly driven by key points and expression action units, including a stage of jointly driving generation by key points and expression action units, a stage of video quality enhancement driven by temporal consistency, and a stage of video super-resolution. In the driving generation stage of this method, the driving information of the video frame will be extracted, and a reconstructed video frame with accurate posture and expression will be generated through the driving information. The quality enhancement stage is used to improve the quality of the initial generated frame, improve the image quality and the temporal consistency of the video. The video super-resolution stage will perform super-resolution processing on the enhanced video to achieve the best visual effect. The generation effects of the three stages are as shown in Figure 3 As shown, the three stages gradually enhance the video quality and visual effect of the generated results.
[0106] In this embodiment, as shown in Figure 1 the method for generating a speaker video jointly driven by key points and expression action units includes the following steps:
[0107] Step S10: Extract the key points and expression action unit information in the driving frame, and drive the original frame to generate a target frame of the pose and expression through motion flow field estimation and dynamic feature transformation;
[0108] Step S20: Perform global enhancement and key area enhancement on the initially generated frame, and perform quality correction and enhancement on the results generated frame by frame through multi-frame joint adversarial training and background fusion to obtain an enhanced video;
[0109] Step S30: Generate preliminary video frames in the low-resolution space, perform super-resolution processing on the enhanced video, and utilize facial prior information to improve the clarity of the video and the high-frequency part of the facial details.
[0110] Among them, in the joint driving generation stage of key points and expression action units (Keypoint and Au jointly drive generation, KAG), the goal of this stage is to extract key points and expression action units (Action Units, abbreviated as AUs) from the driving frame d t to drive the original frame s to generate the target frame y w,t (where the subscript w represents the warping generation stage, and t represents the t-th frame). The joint driving information of the present invention is extracted by a key point extractor K (Keypoints detector) and an AU encoder A (AU detector) based on a convolutional neural network (CNN). The present invention first extracts the 3D key point information of the original frame and the driving frame through the key point extractor K, and sends it into the motion field estimation network M to obtain the 3D motion flow information from the original frame to the driving frame. Through the motion flow information, the present invention warps the 3D features of the original frame extracted from the appearance feature extractor F in the feature domain. Subsequently, the present invention extracts the AUs information of the driving frame through the AU encoder A, and incorporates the AUs and key points of the driving frame, as well as the warped appearance features, into the dynamic feature transformation module (Dynamic Feature Transformation) for further fine adjustment of the warped appearance features. Finally, the adjusted appearance features are input into the generator (Generator) to generate the target frame. This part is trained using the framework of a generative adversarial network.
[0111] In this embodiment, the joint driving and generating stage of key points and expression action units includes the extraction of driving information, driving and generating, and the training of the generating stage. Among them, when extracting driving information, key driving information is required to generate a speaker's video from a single image. The present invention divides the driving information into pose information and fine expression information and decouples them, so as to be able to generate accurate poses and detailed expressions simultaneously. 3D key points can accurately represent the spatial positions of human feature points, so the present invention uses them to drive the changes of poses and rough expressions. The AU vector can represent most facial expressions and is independent of individual identities, and is used to further fine-tune expressions after the initial rough pose adjustment, so as to improve the visual quality of expressions.
[0112] In terms of key point extraction, the present invention uses a 3D key point detection network K (refer to the key point extraction network in the baseline OSVF method, such as Figure 1 ) to extract the key points x of the original frame s s,k ∈R 15×3 and the key points x of the driving frame d d,k ∈R 15×3 , for motion field calculation. In terms of AU extraction, the structure of the AU detector A is as Figure 5 shown, and the input driving image is downsampled and encoded based on a multi-layer 2D convolutional network layer, and finally the driving AU vector x is obtained d,a ∈R 17 . The AUs are decoupled from identities and poses, so they can drive the source image more freely to generate accurate expressions and improve the visual quality of expressions. Since the AU encoder is supervised trained, effective and reliable expression guidance information can be obtained through it.
[0113] The FACS Facial Action Coding System divides the human face into 42 muscle regions according to anatomy, corresponding to 42 expression action units (AUs). Different AU values can represent different muscle activity levels. According to the FACS system, the combination of 17 AU units is sufficient to represent most facial expressions. In terms of key point selection, experiments show that 15 3D key points are sufficient to ensure high-precision pose generation while minimizing the additional bandwidth and computational cost caused by the increase in the number of key points.
[0114] In this embodiment, when driving and generating, the present invention uses an appearance feature extractor F to extract 3D features f from the source image s s , and these features contain information such as the appearance, identity, and background of the person. These features are extracted in three-dimensional space to achieve better deformation in the spatial domain. Warping in these 3D features can achieve free control of head movement. After extracting the driving information, the present invention encodes, deforms, and generates the source image according to the driving data.
[0115] As shown Figure 4 in the figure, the present invention inputs the key point x s,k , x d,k and the source 3D feature f s into the motion flow field estimator M to calculate the 3D motion flow field w, and uses w to warp and deform f s so as to realize the change of the posture from the original frame to the target frame. However, simply deforming the features is not sufficient to guide the generation of natural expressions, especially the generation of teeth and wrinkles. The present invention unfolds the deformed features from 3D to 2D and embeds the target AUs information x s,k into the features. The Aus information x d,a of the driving frame, the key point x d,k , and the warped appearance features are fed into the dynamic feature transformation module for further delicate adjustment to obtain a more accurate expression. Finally, the present invention inputs the transformed features into the generator to generate the target frame y w,t .
[0116] As shown Figure 6 in the figure, the present invention takes the key point and AUs as prior conditions, and dynamically adjusts the intermediate features through the Spatial Feature Transform (SFT) layer, so that it can generate more accurate expressions according to the provided expression and posture information. The present invention expands the key point x d,k ∈R 15×3 and the Aus information x d,a ∈R 17 into embeddings of dimensions R 45×h×w and R 17×h×w , and concatenates them with the warped appearance feature f s ∈R c ×h×w in the channel dimension. The SFT layer generates element-wise scaling and offset parameters using the intermediate features of the previous layer and the prior conditions, and performs element-wise affine transformation on the features. In this way, the network can dynamically make more accurate adjustments to the intermediate features according to the driving information.
[0117] In this embodiment, during the training of the generation stage, the present invention randomly samples two frames of images from the same video: one frame is used as the source image s, and the other frame is used as the driving image d. Its loss function is defined as follows:
[0118] L = λ c L c + λ adv L adv + λ p L p + λ h Lh +λ au L au
[0119] Among them, λ c , λ adv , λ p , λ h , λ au is the loss weight.
[0120] During the training process, two frames of images are randomly sampled from the same video: one frame is used as the source image (s), and the other frame is used as the driving image (d), and the following loss function is used for training:
[0121] (1)L c is the reconstruction loss. Calculate the pixel-level difference between the driving frame d t and the generated frame y w,t to improve the reconstruction accuracy of the generated image.
[0122] (2)L adv is the adversarial loss. The present invention adopts a multi-scale discriminator for adversarial training, and generates an image highly similar to the real image by minimizing the adversarial loss.
[0123] (3)L p is the perceptual loss. The present invention uses the VGG19 perceptual network to calculate the perceptual loss between the generated frame y w,t and the driving image d t , which constrains the generator to generate an image similar to d t in terms of semantics and space. At the same time, the perceptual loss has been proven to be helpful for generating clear and human vision-consistent images.
[0124] (4)L h is the head pose. The present invention calculates the difference in the head rotation angle between the generated frame and the driving frame as the loss, and this loss calculates and reduces the difference between the predicted head pose information and the true value. By minimizing this loss, it is beneficial for the key point detector K to generate accurate head key points, thereby generating accurate motion optical flow, and ultimately contributing to a more accurate generated result of the head pose.
[0125] (5)L au is the expression loss. The present invention obtains the Aus vector values of the driving frame d t and the generated frame y w,t through the AU detector, and calculates the mean square error loss between them. A smaller L au indicates accurate expression generation. By using AUs as the supervision signal, the present invention can constrain the network to generate more accurate expression results and achieve more realistic facial details, such as teeth and smile textures, etc.
[0126] Among them, the AU detector (A) realizes downsampling encoding through a multi-layer 2D convolutional neural network and generates the AU vector in the driving frame (d); the appearance feature extractor (F) uses a 3D convolutional neural network to extract the appearance features of the original frame (s), and the extracted appearance features include the appearance, identity information, and background information of the person; the dynamic feature transformation module (SFT) performs feature adjustment based on the Spatial Feature Transform layer, and dynamically adjusts the intermediate features using key point and AU information to generate accurate facial expressions; the generator generates the target frame (y w,t ) through a conditional generative adversarial network (cGAN) architecture, where key point and AU information are input into the generator as prior conditions; the reconstruction loss, adversarial loss, perceptual loss, head pose loss, and expression loss jointly optimize the generator to improve the accuracy and naturalness of video generation; the head pose loss calculates the difference in head rotation angles between the generated frame and the driving frame to reduce the difference in the head pose of the generated frame and the driving frame.
[0127] In this embodiment, the video quality enhancement stage driven by temporal consistency includes global enhancement, key region enhancement, multi-frame joint adversarial training, and background fusion. In the first stage, the present invention drives the generation by jointly using key points and expression action units (AUs) to obtain the distorted frame y w,t . However, due to its focus on generating accurate poses and expressions, y w,t is lacking in picture quality and detail texture, and there are problems such as texture blurring, contrast reduction, and light loss. When large-scale movements occur, it is also prone to unavoidable artifact problems. In addition, the generation in the first stage is processed frame by frame, which leads to quality fluctuations and inter-frame jitter between the generated frames {y w,1 , y w,2 , ……y w,t}. These problems reduce the overall visual experience. To solve this problem, the present invention proposes a temporal consistency-based driven video quality enhancement network. The schematic diagram of the network framework is as Figure 7 shown.
[0128] In this embodiment, spatio-temporal information is used for global enhancement (Global enhancement). Among them, for the frame y w,t , the present invention uses the adjacent generated frames y w,t-1 and y w,t+1 in the first stage as auxiliary frames. After channel concatenation of these three frames, they are input into the global enhancement network. The auxiliary frames provide rich inter-frame information, which can effectively enhance the picture quality of the frame y w,t and ensure better temporal consistency. Due to various reasons, when the result frame y generated in the first stagew,t Poor generation in some regions y w,t The generation of the front and back frames is better. At this time, the auxiliary frame y w,t-1 and y w,t+1 can provide a reference for the regions with poor generation in y w,t to help repair these regions, ultimately improving the global video quality and minimizing the inter-frame quality fluctuations. Using the frame y w,t+1 as an auxiliary frame for enhancement will introduce a one-frame delay during the enhancement process, but this delay has little impact on real-time performance.
[0129] After the auxiliary frame branch enters the network, the inter-frame joint features that help repair artifacts and enhance inter-frame consistency will be extracted through downsampling and multiple Res-Nets. Among them, Res-Net can further arrange the inter-frame information through skip connections and residual learning mechanisms, and capture and enhance the texture details of the image more deeply.
[0130] The present invention proposes a content-adaptive attention-based fusion method for fusing inter-frame joint features into the features of the frame y w,t in the feature domain. The high-quality, artifact-free regional features from the auxiliary frame are incorporated into the features of the frame y w,t as feature blocks, aiming to repair the regions in y w,t that are prone to generating artifacts and enhance the global generation quality. The generated feature mask M is used for feature fusion of the features of y w,t and the inter-frame joint features. This method also promotes inter-frame information exchange, thereby improving the inter-frame consistency in the generation results. The global quality enhancement network can sharpen blurred lines, eliminate the blurred artifacts caused by distortion generation, restore the lost lighting details and contrast, and significantly improve the visual quality.
[0131] In this embodiment, critical area enhancement
[0132] In the task of generating a speaker, the mouth area is usually particularly difficult to generate. The quality of the mouth area significantly affects the overall visual perception. However, many previous works equally process the global area, often ignoring the importance of the mouth area, resulting in poor generation quality in the mouth area. To solve this problem, the present invention focuses more on the mouth area and performs further enhancement to improve its quality.
[0133] After global quality enhancement is completed, the present invention uses a pre-trained lightweight facial key-point extraction module to detect the key points at both ends of the mouth in the image. This enables the present invention to extract the mouth region, and then inputs it into a dedicated key-region enhancement network based on content-adaptive attention. At the end of the network, two branch networks are generated: one generates the enhanced mouth image I c , and the other generates the mask M. Then, according to the following formula, I c , M, and the unenhanced mouth image I c are fused:
[0134] I out = M × I s +(1 - M) × I c
[0135] The fused image I out Subsequently, it will cover the original area. Among them, the mask M guides the generator to focus on specific areas, such as generating teeth and lips with poor quality. At the same time, M ensures that the edges of the cropped image are smoothly and naturally fused with the original image, so that the enhanced mouth region can be seamlessly integrated into the overall image without obvious blocky artifacts.
[0136] Since there are fewer pixels in the mouth region and a lightweight model is used for this part, there is no need to worry about increasing the burden on the operation of the global network.
[0137] In this embodiment, during multi-frame joint adversarial training, although the above method can effectively generate and enhance video frames, due to frame-by-frame processing and no constraint on the inter-frame consistency of the training loss, it is easy to cause inter-frame unevenness problems. In addition, since the task of the present invention hopes to generate in real time, it is not desirable to process the entire video sequence simultaneously. If the present invention previously relied on the enhanced frame of the previous frame to guide the enhancement of the current frame to enhance temporal consistency, the problem of cumulative generation errors would occur. The present invention proposes an innovative multi-frame joint adversarial training strategy. This strategy enables the enhancement network to effectively learn the latent information between adjacent frames through adversarial loss constraints, thereby generating smooth and stable-quality video frames.
[0138] Such as Figure 8As shown, the temporal consistency discriminator T mainly consists of three parts. The initial features of three consecutive frames (single frame 3×H×W) are concatenated in the channel dimension to obtain spatio-temporal joint features of 9×H×W, and the spatio-temporal information between frames is extracted through a 3D convolutional layer (Cond3d). Subsequently, the spatio-temporal features between frames will be fed into 7 feature processing blocks for feature refinement. A single processing block includes a 2D convolutional layer, a normalization layer, and an activation layer (Conv2d+BatchNorm+ReLU). Finally, the features will pass through a linear layer (Linear) to calculate the temporal consistency loss L t .
[0139] When training the video quality enhancement network driven by temporal consistency, the present invention processes 3 consecutive frames simultaneously, and performs global enhancement and key region enhancement respectively. The 3 generated result frames are then concatenated as spatio-temporal sequence features and input into the temporal consistency discriminator T to calculate the temporal consistency loss L t . The present invention adopts a generative adversarial training method. Through adversarial training of the generated three consecutive frames and the corresponding real three frames, the network is prompted to learn towards generating as realistic images as possible. This method enables the enhancement network to effectively capture the potential relationships in spatio-temporal information, ensuring that the generated sequence is highly similar to the real video sequence in terms of quality, texture details, and spatio-temporal consistency, thereby achieving the best video generation effect.
[0140] In this embodiment, for the background fusion network (Background fusion), many methods for generating videos of multiple speakers have the problem of background jitter, which will significantly reduce the visual experience. For this task, the present invention discovers that the original image s provides reliable background prior information. The present invention uses the background fusion network B to extract the mask M of the intersection region of the background of the original frame s and the background of the intermediate generated image after global and local enhancement and performs fusion through the following formula:
[0141]
[0142] As Figure 9 shown, the present invention can replace the background of the intermediate generated image with the background of the original frame s. For each frame in the generated sequence, the present invention replaces its background with the background in the original frame, thereby obtaining a video sequence with a fixed background. In this way, the present invention can effectively suppress the problems of background jitter and flicker.
[0143] After global enhancement, local enhancement, and background fusion, the result frame y initially generated in the first stage w,tThe picture quality and video time consistency have been significantly enhanced, effectively reducing problems such as artifacts and blurring, enhancing the detail texture quality, and achieving better visual quality.
[0144] In this embodiment, during the training of the enhancement stage, the time-consistency-driven video quality enhancement module is trained by minimizing the following loss function:
[0145] L = λ c L c + λ adv L adv + λ p L p + λ m L m + λ t L t + λ col L col
[0146] Among them, L c is the L 1 loss, L adv is the ordinary image adversarial loss, and L p is the perceptual loss. L m is the adversarial loss in the mouth region. Additionally:
[0147] (1) L t is the time consistency loss. In the present invention, the time consistency discriminator T is adversarially trained with the generated consecutive frames and the corresponding consecutive real frames. During the training of the generator, the generated result will be calculated by the time consistency discriminator T to obtain the time consistency score, which can be used as the time consistency loss L t . A lower L t value can generate a smoother and more stable frame sequence. In addition, T promotes information sharing between adjacent frames during training, thus helping to generate more realistic results.
[0148] (2) L m is the mouth region adversarial loss. During the training process, the present invention designs a discriminator specifically for the mouth region to ensure that the generated mouth features are more realistic and rich in details. The mouth region image generated by the key region enhancement network will pass through the mouth region discriminator to calculate the authenticity score of the mouth region, that is, the mouth region adversarial loss L m .
[0149] (3) L col is the color consistency loss. In the present invention, the color consistency is evaluated by calculating the average value difference of each RGB channel between adjacent generated frames. This method ensures that the colors between frames are more uniform and effectively alleviates the flicker problem caused by significant chromaticity and brightness changes.
[0150] In this embodiment, in the prior-driven video super-resolution (PDS) stage, considering the common face resolution limitation in video conferencing, the aforementioned generation stage is carried out in the low-resolution space (256×256 in the experiment). To achieve high-definition and high-quality results, the present invention uses GFPGAN to perform super-resolution processing on the low-resolution output (from 256×256 to 512×512), and GFPGAN utilizes the prior knowledge of the face. The pre-trained StyleGAN2 embedded in GFPGAN contains rich face prior information, which can significantly enhance the high-frequency part of the facial details in face super-resolution.
[0151] Since there are problems such as artifacts, blurriness, and loss of lighting details in the images initially generated in the first stage of the present invention, that is, the joint driving generation stage of key points and expression action units, directly performing super-resolution on the results y of the first stage w,t will cause the image to be super-resolved in the wrong direction, or cause irreversible damage to the image quality, such as incorrect line enhancement of artifacts, or ignoring the lighting information lost during the distortion process, etc. In addition, directly applying super-resolution to the blurred frame sequence will deteriorate the temporal consistency of the generated frames. Therefore, the present invention first applies temporally consistent-driven video quality enhancement (stage two) to correct and enhance the video quality of the initially generated results, and then performs prior-based video super-resolution processing, which can maximize the video quality of the final video sequence super-resolution results. This method can be selected according to the requirements for the result quality and computing resources of the present invention. Even if the present invention only performs generation through the first and second stages without super-resolution processing, it can still obtain visually satisfactory results.
[0152] Finally, after three stages of driven production, the speaker video generation method of the present invention supports low-resolution recording, low-bandwidth transmission, and high-quality presentation, and will provide a feasible solution for video conferencing scenarios under low-bandwidth conditions.
[0153] Due to the resolution limitation of the comparative method, the present invention only uses the results generated through the first step (KAG) and the second step (TDE) as the basis for quantitative comparison.
[0154] Dataset: The model of the present invention and other comparative models were trained on the VFHQ dataset, which contains more than 16,000 high-quality video clips from different interview scenarios. The present invention used approximately 15,000 videos as the training set and selected two frames from each video as the training pairs. During testing, the present invention used the first frame as the source image and the next 200 frames as the driving images. In addition, to evaluate the generalization ability of the model, the present invention also conducted tests on the HDTF dataset.
[0155] Evaluation Metrics: To evaluate the authenticity of the generated images, the present invention adopted a variety of evaluation metrics. These metrics include the mean absolute error of pixel deviation (L1 Loss), structural similarity (SSIM), peak signal-to-noise ratio (PSNR), and learned perceptual image patch similarity (LPIPS), which are used to measure the differences in color, illumination, structure, and texture between the generated images and the driving images. In addition, the present invention also used the average landmark distance (ALD) and the average facial action unit distance (average AUs distance - AUD) to evaluate the accuracy of pose and expression.
[0156]
[0157] Table 1. Quantitative Comparison between the Method of the Present Invention and Mainstream Methods
[0158] The work of the present invention mainly focuses on the video conferencing scenario and compares the results of reconstructing the same identity. The present invention uses multiple metrics to quantitatively evaluate the method of the present invention and the most mainstream and latest methods on two datasets. The results are shown in Table 1, clearly indicating that the method of the present invention is superior to other methods in most metrics. The higher ALD and AUD of the present invention indicate that the method of the present invention can more accurately control pose and expression. As Figure 6 shown, the method of the present invention generates more accurate expressions, verifying the effectiveness of the joint driving of key points and expression action units. In contrast, methods such as FOMM and OSFV that rely solely on key points lack accurate expression guidance, resulting in lower expression accuracy. The method of the present invention decouples the expression information from the pose driving, extracts the expression action units (AUs) in a supervised manner, not only improves the accuracy of expression extraction and generation, but also makes the key points pay more attention to the guidance of the pose itself, thereby enhancing the accuracy of the overall generated result in terms of pose and expression.
[0159] In terms of L1 Loss, LPIPS, PSNR, and SSIM, the method of the present invention outperforms other methods in most cases, indicating that the results of the present invention have a higher similarity to the driving image in terms of illumination, color, and structural texture details. In summary, the method of the present invention performs better in terms of reconstruction consistency and authenticity. It should be noted that among the comparison methods, Dagan not only uses key points but also uses depth images as driving information. The depth image provides richer spatial motion details, which is also the reason why Dagan's ALD on the VFHQ dataset and PSNR on the HDFT dataset exceed the method of the present invention. However, in low-bandwidth video conferencing scenarios, it is not practical to transmit depth images for driving because the depth image data volume is large and requires a large bandwidth for transmission. The method of the present invention only needs to transmit 15 three-dimensional key points and 17 AU values per frame for reconstruction. Similarly, Metaportrait uses far more facial landmarks than the number of combined driving information in the method of the present invention, which is also the reason for its higher ALD on the VFHQ dataset. Since OSFV lacks clear expression constraints, and the present invention makes up for this by imposing constraints through AU loss, the method of the present invention performs better in terms of the accuracy of expression generation. DPE drives generation by extracting latent encodings, but the data volume of this latent encoding is still larger than that of the method of the present invention, and the method of the present invention surpasses DPE in terms of both accuracy and video quality. Compared with methods with comparable driving information data volume, the reconstruction quality of the method of the present invention is the best.
[0160] Figure 10 demonstrates the advantages of the method of the present invention in terms of visual effects. Compared with other methods, the results of the present invention retain clearer texture structures and more realistic details, such as teeth. In scenarios involving large movements, the method of the present invention introduces the fewest artifacts, as shown in Figure 6 the second row of the hair part of. In addition, the results of the present invention have the smallest differences in overall color and illumination distribution, the best restoration of the original video reconstruction, and the best visual quality. Therefore, it can be concluded that the method of the present invention can generate the best visual effects with the least amount of driving information while maintaining the highest pose and expression accuracy.
[0161] In this embodiment, the beneficial effects of the combined driving of key points and expression action units: As can be seen from Table 1, the method of the present invention has better ALD and AUD indicators, proving that the combined driving of key points and expression action units can provide more effective guidance for pose and expression generation, making the generated results have more accurate poses and expressions.
[0162] In this embodiment, the beneficial effects of the time-consistency-driven video quality enhancement (TDE) stage: As can be seen from Table 2, after the first-stage key point and expression action unit joint driving of the present invention and then the time-consistency-driven video quality enhancement, the indicators representing the reconstruction fidelity, such as L1, LPIPS, PSNR, and SSIM, have been improved. It is proved that the TDE stage has improved the reconstruction accuracy and fidelity of the generated images. TDE has restored the colors and lighting lost during the generation process in the first stage. Although LPIPs on the VFHQ dataset has slightly decreased, this is inevitable when improving image clarity.
[0163]
[0164] Table 2. Ablation experiments for the TDE module
[0165] In terms of image quality, Figure 11 shows the overall reconstruction effect under large-scale motion. Before TDE, when there is a large movement between the original frame and the driving frame characters, there are obvious artifacts, color loss, and blurred lines in the reconstructed image. However, after applying TDE, these artifacts are significantly reduced, the texture details are clearer, and the visual effect is more realistic.
[0166] In this embodiment, the beneficial effects of key region enhancement in the TDE stage: Figure 12 shows the impact of key region enhancement on the mouth and texture details. After enhancement, the reconstructed image shows more realistic and clearer details, especially in the tooth area. This improvement has laid a solid foundation for the next-stage super-resolution process. Generally speaking, the results after TDE processing show better visual effects.
[0167] In this embodiment, the beneficial effects of the multi-frame continuous adversarial (MFA) scheme in the TDE stage: The present invention also tested the impact of multi-frame continuous adversarial (MFA) on the inter-frame L1 loss (IF L1). Theoretically speaking, when the result reconstruction accuracy and image fidelity remain unchanged, a smaller inter-frame L1 loss indicates better temporal consistency and smoother video effects of the video. As shown in Table 3, after applying MFA, other indicators related to the structural and color fidelity have not changed significantly. However, the inter-frame L1 loss has decreased, indicating that MFA has improved the inter-frame temporal consistency by optimizing the inter-frame spatio-temporal relationship in the generated result sequence without affecting the fidelity of the overall pose.
[0168]
[0169] Table 3. Ablation experiments for the MFA scheme
[0170] In this embodiment, the beneficial effects of the background blending (BB) scheme in the TDE stage: Considering that most of the video backgrounds in the HDTF dataset are static, the present invention evaluated the effect of the background blending module on this dataset. As shown in Table 4, after applying the background blending module, the LPIPS, PSNR, and SSIM metrics, which mainly evaluate structural similarity, did not change significantly. However, the L1 loss decreased significantly, indicating that the background blending module improved the fidelity of the results, especially in terms of the color and exposure of the background, while maintaining the integrity of the overall structure. In addition, from the comparison results, the present invention observed that the video sequences processed by applying the background blending module showed better background stability, effectively reducing background jitter and flicker.
[0171] This improvement significantly enhances the overall visual experience.
[0172]
[0173] Table 4. Ablation experiments for the BB scheme
[0174] In this embodiment, the beneficial effects of the prior-driven video super-resolution (PDS) module: Finally, the present invention conducted an ablation study on prior-based super-resolution (PDS). As Figure 13 shown, while maintaining the accuracy of poses and expressions, PDS improved the resolution and detail performance of the generated results, presenting clearer and more delicate images. This step will further improve the visual experience of the audience.
[0175] In summary, the present invention integrates the three independent tasks of drive generation, quality enhancement, and video super-resolution into a comprehensive speaker avatar generation method, thereby achieving low-resolution recording, low-bandwidth transmission, and high-quality display. The three stages are as follows: 1. The key point and expression action unit joint drive generation stage; 2. The time-consistency-driven video quality enhancement stage; 3. The video super-resolution stage. In the drive generation stage, the key points and expression action units of the video frames are extracted, and a reconstructed video frame with accurate poses and expressions is generated through the drive information. The quality enhancement stage is used to improve the quality of the initially generated frames, enhancing the image quality and video time consistency. The video super-resolution stage will perform super-resolution processing on the enhanced video to achieve the best visual effect.
[0176] The present invention proposes a speaker video generation network based on the joint drive of key points and expression action units, which can dynamically integrate various drive information to improve the accuracy of expression and pose generation. The focus of this part of the scheme is to use the joint drive of key points and expression action units, and through a dynamic feature transformation module for detailed feature adjustment, so that the generated results achieve better pose and expression accuracy.
[0177] The present invention proposes a temporal consistency enhancement network. To achieve the best video quality, the present invention utilizes spatio-temporal information for global and local enhancement, effectively improving the perceived quality of the picture and reducing quality fluctuations. At the same time, through innovative spatio-temporal joint adversarial learning and by processing background jitter through a background fusion method, the temporal consistency of the video is effectively improved. The network framework part of the temporal consistency-driven video quality enhancement method of the present invention proposes a comprehensive picture quality enhancement scheme from global to local, and utilizes the spatio-temporal information between frames to improve the picture quality, reduce artifacts, blurring, chrominance and luminance loss, etc. in the driving generation process, and enhance the detail texture. A multi-frame joint adversarial training scheme is proposed to prompt the network to learn the potential inter-frame information in the auxiliary frames, thereby improving the temporal consistency of the resulting video. A background fusion module is proposed, which greatly improves the background stability of the resulting video by reusing the background of the original frame and reduces background jitter and flicker.
[0178] The present invention proposes a method for generating a speaker video jointly driven by key points and expression action units, aiming to reconstruct a video sequence with accurate postures and expressions, high picture quality, and strong temporal consistency through a small amount of driving information. For requirements such as video conferencing in low-bandwidth scenarios, the method of the present invention can reconstruct high-quality video content at the receiving end by only transmitting a small amount of driving information (such as the driving data of the first frame and subsequent frames), rather than a complete high-resolution video frame, thus significantly reducing the required bandwidth. The method is jointly driven by key points and expression action units, improving the accuracy of expressions and postures in the generated video and ensuring video quality and smoothness.
[0179] In low-bandwidth video conferencing scenarios, the facial expressions and posture expressions of participants usually need to be accurately transmitted to improve the communication effect. However, traditional video transmission methods with high resolution and high frame rate consume a large amount of bandwidth and are prone to video stuttering, latency, or image quality degradation. The present invention can effectively solve the problems of background jitter and picture distortion while ensuring video quality through global enhancement of spatio-temporal information, detail enhancement of key regions, and innovative spatio-temporal joint adversarial learning.
[0180] The application of this method not only improves the quality of video conferencing in bandwidth-constrained environments, but its potential applications include video conferencing and remote collaboration, low-bandwidth video streaming, virtual / augmented reality (VR / AR), telemedicine and diagnosis, intelligent customer service and virtual assistants, intelligent monitoring and security, personalized video generation for social media, and film and television production and post-editing, etc. By generating accurate facial expression and posture information, the present invention provides an efficient, accurate and low-bandwidth video generation technology, which is widely applicable to various practical scenarios with bandwidth constraints.
[0181] In the second aspect of the embodiments of the present invention, a computer device is further provided, which includes a memory and a processor. A computer program is stored in the memory, and when the computer program is executed by the processor, the method of any one of the above embodiments is implemented.
[0182] The computer device includes a processor and a memory, and may further include an input system and an output system. The processor, the memory, the input system, and the output system may be connected through a bus or other means. The input system can receive input digital or character information, and generate a signal input related to the migration of the speaker video generation jointly driven by the key points and the expression action units. The output system may include a display device such as a display screen.
[0183] As a non-volatile computer-readable storage medium, the memory can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for generating a speaker video jointly driven by key points and expression action units in the embodiments of the present application. The memory can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by using the method for generating a speaker video jointly driven by key points and expression action units, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the local module through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0184] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data. The processors of multiple computer devices in this embodiment of the computer device execute various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory, that is, implement the steps of the method for generating a speaker video jointly driven by key points and expression action units in the above method embodiments.
[0185] It should be understood that, without conflict, all the embodiments, features, and advantages described above for the method of generating a speaker video by jointly driving key points and expression action units according to the present invention are equally applicable to the method of generating and storing a speaker video by jointly driving key points and expression action units according to the present invention.
[0186] Those skilled in the art will also understand that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the functions of various illustrative components, blocks, modules, circuits, and steps have been described generally. Whether this function is implemented as software or hardware depends on the particular application and the design constraints imposed on the overall system. The functions that can be implemented in various ways for each specific application by those skilled in the art, but such implementation decisions should not be construed as causing a departure from the scope of the disclosure of the embodiments of the present invention.
[0187] Finally, it should be noted that the computer-readable storage medium herein (e.g., a memory) can be a volatile memory or a non-volatile memory, or can include both volatile memory and non-volatile memory. By way of example and not limitation, non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM), which can serve as an external cache memory. By way of example and not limitation, RAM can be obtained in various forms, such as synchronous RAM (DRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The storage devices of the disclosed aspects are intended to include, but are not limited to, these and other suitable types of memories.
[0188] The various exemplary logic blocks, modules, and circuits described in connection with the disclosure herein can be implemented or performed using the following components designed to perform the functions herein: a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of such components. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP and / or any other such configuration.
[0189] The above are exemplary embodiments of the disclosure of the present invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments of the disclosure of the present invention as defined by the claims. The functions, steps, and / or actions of the method claims according to the disclosed embodiments herein need not be performed in any particular order. In addition, although the elements of the embodiments of the disclosure of the present invention may be described or claimed in individual form, they may also be understood as plural unless explicitly limited to the singular.
[0190] It should be understood that, as used herein, unless the context clearly supports the exception, the singular form "a" is also intended to include the plural form. It should also be understood that the "and / or" used herein refers to any and all possible combinations of one or more of the associated listed items. The serial numbers of the disclosed embodiments of the present invention above are for description only and do not represent the superiority or inferiority of the embodiments.
[0191] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the disclosure of the embodiments of the present invention (including the claims) is limited to these examples; under the concept of the embodiments of the present invention, the technical features between the above embodiments or different embodiments can also be combined, and there are many other variations in different aspects of the embodiments of the present invention as above, which are not provided in detail for the sake of brevity. Therefore, any omission, modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present invention shall be included within the protection scope of the embodiments of the present invention.
Claims
1. A method for generating a speaker video driven jointly by key points and expression action units, characterized in that: The method comprises the following steps: Key point and expression action unit joint drive generation stage: Extract key points and expression action unit information in the driving frame, and drive the original frame to generate the target frame of posture and expression through motion flow field estimation and dynamic feature transformation; Temporal consistency driven video quality enhancement stage: Perform global enhancement and key area enhancement on the initial generated frames, and perform quality correction and enhancement on the frame-by-frame generated results through multi-frame joint adversarial training and background fusion to obtain the enhanced video; Prior-driven video super-resolution stages: Generate preliminary video frames in low-resolution space, perform super-resolution processing on the enhanced video, and use facial prior information to improve the clarity of the video and the high-frequency parts of facial details.
2. The method for generating a speaker video driven by a key point and an expression action unit as claimed in claim 1, characterized in that: The key point and expression action unit joint driving information is obtained by using a key point extractor based on a convolutional neural network to extract the 3D key point information in the original frame and the driving frame, and calculating the 3D motion flow information from the original frame to the driving frame through a motion field estimation network.
3. The method for generating a speaker video driven by a key point and an expression action unit as claimed in claim 2, characterized in that: Also includes: The 3D appearance features of the original frame are extracted by the appearance feature extractor, and the original frame features are warped in the feature domain based on the motion flow information; The AUs information, key point information and warped appearance features in the driving frame are input into the dynamic feature transformation module to fine-tune the warped appearance features; Generate target frames with accurate pose and expression using a generative adversarial network framework.
4. The method for generating a speaker video driven by a key point and an expression action unit as claimed in claim 3, characterized in that: The joint drive generation phase of key points and expression action units includes the extraction of drive information, including the following steps: Extract 3D key point information and expression action unit information in the driving frame. The expression action unit information vector is used to represent most facial expressions and is independent of individual identity. Use a 3D key point detection network based on a convolutional neural network to extract 3D key points in the original frame and the driving frame, and calculate the motion flow field based on these key points to achieve posture adjustment of the original frame; Using the AU detector, the driving frame image is downsampled and encoded based on a multi-layer 2D convolutional network layer, and finally the driving AU vector is obtained to represent the facial expression information in the driving frame and further fine-tune the expression.
5. The method for generating a speaker video driven by a combination of key points and expression action units as claimed in claim 4, characterized in that: The joint driving generation phase of key points and expression action units includes driving and generation, including the following steps: Use the appearance feature extractor to extract 3D appearance features from the source image, and use the motion flow field estimator to warp the features of the original frame with the motion flow field to generate features with a new posture; The AU information, key point information and deformed original frame features in the driving frame are input into the dynamic feature transformation module, and the features are fine-tuned through the spatial feature transformation layer to generate accurate expressions that conform to the driving information; The adjusted features are fed into the generator to generate the target frame.
6. The method for generating a speaker video driven by a combination of key points and expression action units as claimed in claim 5, characterized in that: The key point and expression action unit jointly drive the generation phase, including the training of the generation phase, including the following steps: During the training process, two frames are randomly sampled from the same video: one frame is used as the source image and the other frame is used as the driving image; The following loss function is used for training: Reconstruction loss: Calculate the pixel-level difference between the driving frame and the generated frame to optimize the reconstruction accuracy of the generated image; Adversarial loss: Use a multi-scale discriminator for adversarial training to generate images that are highly similar to real images; Perceptual loss: Use the VGG19 perception network to calculate the perceptual loss between the generated frame and the driving image, constraining the semantic and spatial consistency of the generated image; Head posture loss: Calculate the difference in head rotation angle between the generated frame and the driving frame, reduce the error between the posture prediction and the true value, and optimize the accuracy of the key point detector; Expression loss: Calculate the mean square error loss between the AUs vectors of the driving frame and the generated frame, constrain the accuracy of expression generation, and generate facial details.
7. The method for generating a speaker video driven by a key point and an expression action unit as claimed in claim 6, characterized in that: The temporal consistency driven video quality enhancement stage includes global enhancement, key area enhancement, multi-frame joint adversarial training and background fusion. The steps are as follows: For the target frame, the adjacent generated frames are used as auxiliary frames. These three frames are input into the global enhancement network after channel splicing, and the inter-frame information provided by the auxiliary frames is used to repair the artifact area in the target frame; Use the pre-trained lightweight facial key point extraction module to detect the key points of the mouth area, input the mouth area into the key area enhancement network, generate the enhanced mouth image and mask, and seamlessly integrate the enhanced mouth area into the overall image through the fusion formula; Construct a temporal consistency discriminator, input the spatiotemporal joint features of three consecutive frames into the discriminator, calculate the temporal consistency loss, and optimize the inter-frame consistency and stability of the generated video through adversarial training; The intersection area mask of the background of the original frame and the enhanced generated image background is extracted, and the background of the original frame is replaced with the enhanced image background to suppress background jitter and flicker.
8. The method for generating a speaker video driven by a combination of key points and expression action units as claimed in claim 7, characterized in that: The key area enhancement network includes: Generate a branch to generate an enhanced mouth image; The mask branch is used to generate a mask to guide the generator to focus on specific areas and ensure smooth edges.
9. The method for generating a speaker video driven by a combination of key points and expression action units as claimed in claim 8, characterized in that: The temporal consistency discriminator comprises: 3D convolutional layer, used to extract spatiotemporal joint features; Feature processing block, including 2D convolution layer, normalization layer and activation layer, is used to refine features; Linear layer for computing temporal consistency loss.
10. The method for generating a speaker video driven by a combination of key points and expression action units as claimed in claim 9, characterized in that: Use the background fusion network B to extract the background of the original frame s and the intermediate generated image that has been globally and locally enhanced. The mask M of the intersection area of the background is fused by the following formula: Replace the intermediate generated image with the background of the original frame s For each frame in the generated sequence, the background is replaced with the background in the original frame to obtain a video sequence with a fixed background.
Citation Information
Cited By
Digital face landmark point generation method and digital face animation generation method
CN121259141A