Conversation head animation synthesis method and related device
By combining a multi-frame diffusion model and a super-resolution model, high-quality and efficient conversational head animation videos are generated, solving the problems of low generation quality, poor temporal consistency, and insufficient lip-sync in existing technologies. This technology is applicable to fields such as digital humans and virtual anchors.
Patent Information
- Application Number
- CN202511997469.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-27
- Publication Date
- 2026-04-10
AI Technical Summary
Existing methods for synthesizing conversational head animations suffer from problems such as low generation quality, poor temporal consistency, insufficient lip-phonetic synchronization, and low generation efficiency.
A multi-frame diffusion model is used to generate low-resolution videos, and a super-resolution model is used for feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual feature fusion to generate high-resolution conversational head animation videos.
It improves generation efficiency, ensures consistency of identity, temporal continuity, and lip-sync in generated videos, while reducing computational costs, meeting the needs of high-quality applications such as digital humans and virtual anchors.
Smart Images

Figure CN121837463A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing and computer vision technology, and in particular to a method and related apparatus for synthesizing conversational head animations. Background Technology
[0002] Voice-driven talking head animation synthesis refers to generating high-definition videos that are synchronized with the voice content, have the same identity, and natural expressions based on the input voice signal and a reference face image. It is widely used in fields such as virtual anchors, digital humans, online education, film and television production, and intelligent voice assistants.
[0003] Existing methods for synthesizing conversational head animations, such as image stitching, keypoint-driven methods, and Generative Adversarial Network (GAN) methods, suffer from problems such as low generation quality, poor temporal consistency, insufficient lip-phonetic synchronization, and low generation efficiency. Therefore, there is an urgent need for a new method for synthesizing conversational head animations to solve the above technical problems. Summary of the Invention
[0004] This application provides a method and related apparatus for synthesizing conversational head animations. It rapidly generates low-resolution videos with preliminary temporal consistency through a multi-frame diffusion model, and then generates animated videos that meet the requirements of identity consistency, temporal continuity, lip-sync, and high resolution through a super-resolution model. While ensuring the quality of the generated videos, it significantly improves the generation efficiency, reduces the computational cost, and meets the needs of high-quality applications such as digital humans and virtual anchors.
[0005] In a first aspect, embodiments of this application provide a method for synthesizing conversational head animation. The method includes: acquiring a reference face image and a target audio sequence; extracting features from the reference face image to obtain reference face image features; extracting features from the target audio sequence to obtain an audio feature vector; inputting the reference face image features and the audio feature vector into a multi-frame diffusion model to generate a continuous low-resolution face video frame sequence; inputting the low-resolution face video frame sequence, the reference face image features, and the audio feature vector into a super-resolution model, wherein the super-resolution model is used to perform feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual-feature fusion branching steps to generate a high-resolution face video frame sequence; and outputting a high-resolution conversational head animation video conforming to a target scene based on the high-resolution face video frame sequence.
[0006] Secondly, embodiments of this application provide a conversational head animation synthesis apparatus, the apparatus comprising: an acquisition unit for acquiring a reference face image and a target audio sequence; a processing unit for extracting features from the reference face image to obtain reference face image features, and extracting features from the target audio sequence to obtain an audio feature vector; inputting the reference face image features and the audio feature vector into a multi-frame diffusion model to generate a continuous low-resolution face video frame sequence; inputting the low-resolution face video frame sequence, the reference face image features, and the audio feature vector into a super-resolution model, the super-resolution model being used to perform feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual-feature fusion branching steps to generate a high-resolution face video frame sequence; and an output unit for outputting a high-resolution conversational head animation video conforming to the target scene based on the high-resolution face video frame sequence.
[0007] Thirdly, embodiments of this application provide a server, including a processor and a memory, wherein the memory stores a computer program, and when the processor invokes the computer program in the memory, it executes the step instructions of the method as described in any one of the first aspects.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of any of the possible methods in the first aspect.
[0009] As can be seen, in this embodiment, a reference face image and a target audio sequence are first acquired. Feature extraction is performed on the reference face image to obtain its features, and feature extraction is performed on the target audio sequence to obtain its audio feature vector. Next, the reference face image features and the audio feature vector are input into a multi-frame diffusion model to generate a continuous low-resolution face video frame sequence. Then, the low-resolution face video frame sequence, the reference face image features, and the audio feature vector are input into a super-resolution model. The super-resolution model performs feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual-feature fusion branching steps to generate a high-resolution face video frame sequence. Finally, a high-resolution conversational head animation video conforming to the target scene is output based on the high-resolution face video frame sequence. Compared to existing technologies, this application rapidly generates low-resolution videos with preliminary temporal consistency through a multi-frame diffusion model, and then generates animation videos that meet the requirements of identity consistency, temporal continuity, lip-sync, and high resolution through a super-resolution model. While ensuring the quality of the generated video, it significantly improves generation efficiency, reduces computational costs, and meets the high-quality application needs of digital humans, virtual anchors, and other applications. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a method for synthesizing conversational head animations, provided as an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the super-resolution model provided in the embodiments of this application; Figure 3 A diagram illustrating the synthesized head animation model for a conversation, provided in an embodiment of this application. Figure 4 A schematic diagram illustrating the comparative experimental results of different methods on the CREMA-D dataset provided in this application embodiment; Figure 5 This application provides a schematic diagram of the structure of a server according to an embodiment of the present application. Figure 6 A functional unit structure block diagram of a conversation head animation synthesis device provided in this application embodiment; Figure 7 This is a schematic diagram illustrating the comparative experimental results of different methods on the HDTF dataset provided in the embodiments of this application. Detailed Implementation
[0012] The following detailed description illustrates the specific implementation method: Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for synthesizing conversational head animations, as provided in an embodiment of this application. Figure 1 As shown, the method includes the following steps S101-S104: Step S101: Obtain a reference face image and a target audio sequence; extract features from the reference face image to obtain reference face image features; and extract features from the target audio sequence to obtain an audio feature vector.
[0013] The reference face image must contain clear facial identity information of the target person (such as facial structure, facial features, etc.) to ensure the consistency of identity in the generated video; the target audio sequence must contain complete speech content and timing information to drive the motion animation of the face (especially the mouth).
[0014] When extracting features from the reference face image and the target audio sequence, a pre-trained face feature extraction network (such as ArcFace) can be used to extract features from the reference face image to obtain 512-dimensional reference face image features; a pre-trained audio feature extraction network (such as Mel spectrum encoder) can be used to process the target audio sequence to generate frame-level audio feature vectors. For example, the sampling rate is set to 16kHz and the number of Mel filter banks is 80.
[0015] In some embodiments, the audio feature vector extraction process employs a sliding window overlap strategy, which sets a preset length of sliding window to perform frame-by-frame processing on the target audio sequence, so that the audio feature vectors corresponding to adjacent low-resolution face video frames have overlapping regions, thereby strengthening the temporal correlation between frames.
[0016] Step S102: Input the reference face image features and the audio feature vector into the multi-frame diffusion model to generate a continuous low-resolution face video frame sequence.
[0017] In some embodiments, the multi-frame diffusion model is built based on an improved U-Net architecture, which includes a downsampling module group, an intermediate feature fusion module, and an upsampling module group connected in sequence. Each of the downsampling module group, the intermediate feature fusion module, and the upsampling module group embeds a spatiotemporal convolutional unit. Each spatiotemporal convolutional unit consists of a two-dimensional spatial convolutional layer and a one-dimensional temporal convolutional layer connected in series, and both the two-dimensional spatial convolutional layer and the one-dimensional temporal convolutional layer have residual connection structures. The two-dimensional spatial convolutional layer is used to extract the image spatial structure features of a single-frame face, and these features characterize the spatial positional relationship between local regions within a single-frame face and the overall face. The one-dimensional temporal convolutional layer is used to model the inter-frame temporal evolution features between adjacent frames, and these features characterize the continuous relationship of face motion between adjacent frames. The spatiotemporal convolutional unit outputs a first feature map that fuses the image spatial structure features and the inter-frame temporal evolution features.
[0018] In this embodiment, the multi-frame diffusion model is built on an improved U-Net architecture. This improved U-Net architecture adopts the classic three-level structure of U-Net: "encoder-intermediate fusion-decoder". Specifically, it includes a downsampling module group (encoder end), an intermediate feature fusion module (bottleneck layer), and an upsampling module group (decoder end) connected in series. The specific settings of each module are as follows (the basic convolution, normalization, activation, and other units of the prior art are only briefly described): The downsampling module group consists of multiple consecutive downsampling modules. Each downsampling module includes convolutional layers for spatial resolution dimensionality reduction, as well as conventional normalization layers and activation function layers, achieving deep feature mining and dimensionality compression. It should be noted that the core improvement of the downsampling module group is not the basic dimensionality reduction structure, but rather the embedding of spatiotemporal convolutional units after the convolutional layers of each module to achieve spatiotemporal feature fusion during the dimensionality reduction process.
[0019] Intermediate Feature Fusion Module: As the bottleneck layer of the improved U-Net architecture, its core function is to perform global spatiotemporal fusion on the deepest features output by the downsampling module group, enhancing the global consistency between inter-frame temporal dependencies and single-frame spatial correlations. This module internally sets up one spatiotemporal convolutional unit and one dual attention fusion unit, with the spatiotemporal convolutional unit being the core improved component, used to receive features from the downsampling module group and complete the initial global spatiotemporal fusion.
[0020] Upsampling module group: Symmetrically arranged with the downsampling module group, it contains multiple consecutive upsampling modules. Each upsampling module includes a deconvolution layer for spatial resolution restoration, as well as a regular normalization layer and activation function layer, realizing the spatial resolution restoration and channel number adjustment of the feature map. Similar to the downsampling module group, its core improvement is the embedding of spatiotemporal convolutional units after the deconvolution layers of each module, realizing the supplementation of spatiotemporal features during the dimensionality upsampling process.
[0021] The downsampling module group, intermediate feature fusion module, and upsampling module group all embed spatiotemporal convolutional units with the same structure. The embedding positions are "after the convolutional layer of the downsampling module", "after the input of the intermediate feature fusion module", and "after the deconvolutional layer of the upsampling module". This ensures that spatiotemporal feature fusion can be performed after each feature dimension transformation (dimensionality reduction / dimensionality increase). This is also the core improvement that distinguishes it from the traditional U-Net architecture.
[0022] In this embodiment, the spatiotemporal convolutional unit is composed of a two-dimensional spatial convolutional layer and a one-dimensional temporal convolutional layer connected in series, and both the two-dimensional spatial convolutional layer and the one-dimensional temporal convolutional layer are provided with residual connection structures. The specific structure and implementation logic are as follows: Two-dimensional spatial convolutional layer: Employs a 2D convolutional kernel of preset size. The stride and padding are set according to the resolution requirements of the input feature map to ensure that the spatial resolution of the feature map remains unchanged after convolution. The input of this layer is a stack of multi-frame features with a time dimension (dimension T×C×W×H, where T is the number of frames, C is the number of channels, and W×H is the resolution of a single frame) output from the previous module. Its core function is to extract the spatial structure features of a single-frame face. Specifically, the image spatial structure features are used to characterize the spatial relationship between local regions within a single-frame face and the overall face, such as the relative positions of the eyes and nose, and the contour relationship between the mouth and cheeks. Core spatial features are enhanced by the weight allocation of the convolutional kernel to strengthen the spatial feature expression of key areas such as facial features and suppress background noise interference.
[0023] One-dimensional temporal convolutional layer: Employs a 1D convolutional kernel of a preset size, with stride and padding settings to ensure the feature length in the temporal dimension remains constant. The input to this layer is the feature map output from the two-dimensional spatial convolutional layer. Its core function is to model the inter-frame temporal evolution features between adjacent frames. Specifically, these inter-frame temporal evolution features characterize the continuous correlation of facial motion between adjacent frames, such as the gradual opening and closing of the mouth and the transition trend of head posture. Through convolutional operations in the temporal dimension, it models and transmits inter-frame motion information, avoiding abrupt changes in motion between frames.
[0024] Residual connection structure: Independent residual connections are set for the two-dimensional spatial convolutional layer and the one-dimensional temporal convolutional layer, that is, the input feature map of each layer is directly added to the feature map after convolution. Residual connections are an existing technology, and their function is to avoid gradient vanishing during the training process of deep networks, ensure the effective transmission of spatial structural features of images and temporal evolution features between frames, and improve the training stability of the model.
[0025] Output of the spatiotemporal convolutional unit: After extracting spatial features through a two-dimensional spatial convolutional layer and modeling temporal features through a one-dimensional temporal convolutional layer, the spatiotemporal convolutional unit outputs a first feature map that integrates the image spatial structure features and the inter-frame temporal evolution features. The dimension of the first feature map is consistent with the dimension of the feature map input to the spatiotemporal convolutional unit, and it can be directly input to the next level module for subsequent processing.
[0026] Based on the improved U-Net architecture, the embedding logic and workflow of the spatiotemporal convolutional unit are as follows: Embedding and operation in the downsampling module group: A spatiotemporal convolutional unit is connected after the dimensionality reduction convolutional layer of each downsampling module; First, the convolutional layer of the downsampling module performs conventional dimensionality reduction processing on the input feature map to obtain a dimensionality-reduced feature map; then, the dimensionality-reduced feature map is input into a two-dimensional spatial convolutional layer to extract the spatial structural features of a single frame face; then, the spatial features are input into a one-dimensional temporal convolutional layer to model the inter-frame temporal evolution features; finally, the features before and after processing are fused through residual connections to obtain the first feature map, which serves as the input of the next downsampling module, simultaneously completing feature dimensionality reduction and spatiotemporal feature fusion.
[0027] Embedding and operation in the intermediate feature fusion module: The spatiotemporal convolutional unit is directly connected to the input of the intermediate feature fusion module; the input is the deepest dimensionality-reduced feature map output by the last module of the downsampling module group; the spatial structure features of each frame in the deep feature are extracted by the two-dimensional spatial convolutional layer of the spatiotemporal convolutional unit, and the one-dimensional temporal convolutional layer models the temporal evolution features between multiple frames, and outputs the fused first feature map, which is directly input into the dual attention fusion unit of the intermediate feature fusion module for further global fusion.
[0028] Embedding and operation in the upsampling module group: After the dimensionality-upgrading deconvolution layer of each upsampling module, a spatiotemporal convolution unit is connected; first, the deconvolution layer of the upsampling module performs conventional dimensionality-upgrading processing on the input feature map to obtain an up-dimensional feature map; then, the up-dimensional feature map is input to the spatiotemporal convolution unit, which sequentially extracts spatial structure features and models temporal evolution features, and outputs the first feature map through residual connections; this first feature map serves as the input of the next upsampling module, simultaneously completing feature dimensionality upgrading and spatiotemporal feature supplementation.
[0029] In some embodiments, the multi-frame diffusion model employs a mean prediction strategy during the training phase. That is, the model's learning objective is to predict the mean of the denoised image, rather than the noise term in traditional DDPM. Specifically, a training objective function for the multi-frame diffusion model is constructed. The core output of this function is the mean vector μ_θ(x_t,t,c) of the denoised high-resolution face video frames, where x_t is the noisy image at time t, t is the diffusion step number, and c is the conditional input (including audio feature vectors and reference face image features). The model learns this mean vector to directly approximate the distribution center of the denoised true image, rather than indirectly reconstructing the image by predicting noise. During training, Gaussian noise is added to the sequence of real high-resolution face video frames according to the noise addition process of traditional diffusion models, resulting in noisy images x_t with different diffusion step numbers. x_t, the diffusion step number t, and the conditional input c are simultaneously input into the multi-frame diffusion model, and the model outputs the predicted denoised mean μ_θ. The model parameters are updated by calculating the deviation between the predicted mean and the real noise-free image. Compared to traditional DDPM noise prediction, this strategy reduces the accumulation of prediction bias and accelerates convergence. The output layer of the multi-frame diffusion model is fine-tuned by removing the output channel corresponding to traditional noise prediction and adding a mean vector output channel to ensure that the output dimension matches the mean vector dimension of the denoised image. The remaining network layer structure remains consistent with the improved U-Net architecture defined in this embodiment.
[0030] As can be seen, this embodiment, by embedding spatiotemporal convolutional units in each core module of the improved U-Net architecture, and through the concatenation of two-dimensional spatial convolutional layers and one-dimensional temporal convolutional layers and residual connection design, can effectively extract the spatial structural features of a single-frame face and the temporal evolution features between frames. The output first feature map has both spatial rationality and temporal continuity. When generating low-resolution face video frame sequences based on the multi-frame diffusion model of this architecture, it can effectively avoid problems such as misalignment of facial features in a single frame and motion stuttering between frames, ensuring that the generated low-resolution frame sequence has good spatial structural integrity and temporal continuity, providing a high-quality feature foundation for the subsequent super-resolution reconstruction stage, and ensuring the generation quality of the final conversational head animation.
[0031] In some embodiments, the deep layer of the downsampling module group, the intermediate feature fusion module, and the shallow layer of the upsampling module group all integrate a dual attention fusion unit. The dual attention fusion unit includes a self-attention subunit and a cross-attention subunit connected in a serial order, and both the self-attention subunit and the cross-attention subunit are configured with residual connection structures. The input of the self-attention subunit is the first feature map output by the previous stage spatiotemporal convolution unit in its module, used to calculate the global attention weights within the first feature map of a single frame, and outputs a second feature map after self-attention weighting. The global attention weights include a first spatial attention weight and a first temporal attention weight. Inter-temporal attention weights are used to assign weights to the image spatial structure features in the first feature map, and the first temporal attention weights are used to assign weights to the inter-frame temporal evolution features in the first feature map. The input of the cross-attention sub-unit includes the second feature map and the audio feature vector, and is used to establish an association mapping between facial visual features and speech temporal features based on the second feature map and the audio feature vector, and output a third feature map that fuses audio information. The facial visual features include the image spatial structure features, the inter-frame temporal evolution features and the global attention weights. The third feature map is decoded by the upsampling module group and restored to the low-resolution facial video frame sequence.
[0032] In this embodiment, the dual attention fusion unit adopts a serial connection method of "self-attention sub-unit - cross-attention sub-unit", and each sub-unit is independently configured with a residual connection structure. The residual connection is a prior art technique, and its core function is to avoid gradient vanishing during deep network training and ensure the effective transmission of feature information; its specific implementation details will not be elaborated here.
[0033] The overall working logic of the dual attention fusion unit is as follows: First, the spatiotemporal features in the first feature map are intrinsically enhanced through the self-attention sub-unit, and key spatiotemporal information is filtered and highlighted; then, the cross-attention sub-unit is used to establish cross-modal association between the enhanced spatiotemporal features and the audio feature vector, so as to achieve accurate fusion of audio information, and finally output a third feature map that has both spatiotemporal rationality and lip-phonetic association, forming a progressive processing flow of "spatiotemporal feature enhancement - cross-modal feature fusion".
[0034] The core function of the self-attention sub-unit is to calculate the global attention weights within the first feature map of a single frame, accurately assign weights to the image spatial structure features and inter-frame temporal evolution features in the first feature map, and output the second feature map after self-attention weighting.
[0035] The specific implementation process is as follows: The first step is feature format adaptation: The first input feature map has dimensions of T×C×W×H (where T is the number of frames, C is the number of feature channels, and W×H is the resolution of a single frame). To adapt to the self-attention calculation logic, it is first converted to the "pixel-feature" format through a dimension transformation operation, that is, the dimension is adjusted to N×D (N=T×W×H, representing the total number of pixels in all frames; D=C, representing the feature dimension of each pixel).
[0036] The second step is attention vector generation: the adjusted features are linearly transformed by three independent 1×1 convolutional layers to generate query vector Q, key vector K and value vector V, all of which have a dimension of N×D_k (D_k is the dimension of the attention head, which is preset according to the actual model training requirements).
[0037] The third step is to calculate the global attention weights: The global attention weights are obtained by calculating the similarity between the query vector Q and the key vector K. These global attention weights specifically include the first spatial attention weight and the first temporal attention weight. ① First spatial attention weight: This is obtained by calculating the similarity between Q and K corresponding to different pixels within the same frame. Its core function is to assign weights to the spatial structural features of the image in the first feature map. For example, higher weights are assigned to the spatial structural features of key areas such as facial features (mouth, eyes, nose), while lower weights are assigned to the spatial features of non-key areas such as cheek edges and background, thereby enhancing the rationality of the spatial structure of the face in a single frame.
[0038] ② First-time attention weight: This is obtained by calculating the similarity between Q and K for corresponding pixels in different frames. Its core function is to assign weights to the inter-frame temporal evolution features in the first feature map. For example, higher weights are assigned to continuous temporal evolution features such as gradual opening and closing of the mouth and smooth transition of head posture in adjacent frames, while lower weights are assigned to abnormal temporal features such as sudden changes in motion and posture jumps between frames, thereby strengthening the continuity of inter-frame temporal sequence.
[0039] The fourth step is weighted feature generation and output: The calculated global attention weights and value vector V are weighted and summed to obtain the preliminary self-attention weighted features; then, the preliminary weighted features are fused with the input features of the self-attention sub-unit (N×D format features after dimensionality transformation) through residual connection; finally, the dimensionality inverse transformation operation is used to restore it to T×C×W×H format, and the second feature map is output.
[0040] The core function of the cross-attention subunit is to establish a mapping between facial visual features and speech temporal features based on the second feature map and audio feature vectors, thereby achieving audio information fusion and outputting a third feature map of fused audio information. Specifically, the facial visual features include the image spatial structure features and inter-frame temporal evolution features in the first feature map, as well as the global attention weights calculated by the self-attention subunit.
[0041] The specific implementation process is as follows: The first step is input feature preprocessing: On the one hand, the second feature map output by the self-attention subunit is subjected to the same dimensional transformation as the above steps to obtain visual features in N×D format, which serve as the query input (query vector Q_vis) for cross-attention; on the other hand, the input audio feature vector (with dimensions T×D_a, where D_a is the audio feature dimension) is subjected to dimensional expansion and linear transformation to ensure that its time dimension is consistent with the time dimension T of the second feature map, and the key vector K_audio and value vector V_audio (both with dimensions T×D_k) for cross-attention are generated.
[0042] The second step is cross-modal association weight calculation: The similarity between the query vector Q_vis and the key vector K_audio is calculated to obtain the cross-modal attention weight. This weight is used to quantify the association strength between the facial visual features of each pixel in each frame and the audio features at the corresponding time. For example, when the audio feature vector represents "vowel pronunciation," the association weight of the visual features of the mouth region in the corresponding frame will be significantly increased, ensuring that mouth movements match the pronunciation state.
[0043] The third step is to generate and output fused features: the cross-modal attention weights and the value vector V_audio are weighted and summed to obtain the preliminary features of the fused audio information; the preliminary features are fused with the input visual features (the second feature map in N×D format) of the cross-attention subunit through residual connection, and then restored to the T×C×W×H format through inverse dimensionality transformation to output the third feature map of the fused audio information.
[0044] The dual-attention fusion unit is precisely integrated into three key locations in the improved U-Net architecture. The integration logic at each location is adapted to the core functions of the corresponding module, ensuring the targeted and progressive nature of feature processing, as detailed below: (1) Deep integration of the downsampling module group: The core function of the downsampling module group is feature dimensionality reduction and deep spatiotemporal feature mining. Its deep layer has completed multiple rounds of dimensionality reduction and spatiotemporal fusion. At this time, the dual attention fusion unit is integrated to accurately assign weights to deep spatiotemporal features, initially screen key spatiotemporal information and fuse audio features, laying the foundation for the global processing of the subsequent intermediate feature fusion module. The integration position is after the spatiotemporal convolution unit of the last module of the downsampling module group. The third feature map output at this position is directly input into the intermediate feature fusion module.
[0045] (2) Integration of intermediate feature fusion module: The intermediate feature fusion module is the bottleneck layer of the improved U-Net architecture. Its core function is to globally integrate deep features. Here, a dual attention fusion unit is integrated, which can enhance global deep spatiotemporal features and deeply fuse audio information to enhance global spatiotemporal dependence and cross-modal consistency. The integration position is after the spatiotemporal convolution unit inside the intermediate feature fusion module, and the third feature map output is input to the upsampling module group.
[0046] (3) Shallow integration of the upsampling module group: The core function of the upsampling module group is feature dimensionality enhancement and resolution restoration. Its shallow layer has completed multiple rounds of dimensionality enhancement and temporal detail supplementation. At this time, the dual attention fusion unit is integrated to optimize the spatiotemporal consistency of the dimensionality-enhanced features and further calibrate the correlation accuracy between audio and visual features to ensure the lip-sync of the output features. The integration position is after the spatiotemporal convolution unit of the first module of the upsampling module group. The third feature map output at this position participates in the subsequent dimensionality enhancement and feature splicing processing. Finally, after complete decoding by the upsampling module group, it is restored to a low-resolution face video frame sequence.
[0047] As can be seen, this embodiment effectively achieves the synergistic advancement of spatiotemporal feature enhancement and cross-modal fusion through the serial structure design and precise integration logic of the dual attention fusion unit. Specifically, the self-attention subunit significantly improves the spatiotemporal consistency of features, avoiding problems such as misalignment of facial spatial structure in a single frame and abrupt motion changes between frames; the cross-attention subunit establishes a precise correlation between facial visual features and audio features, ensuring the synchronization of the generated low-resolution facial video frame sequence with the target audio. Simultaneously, the configuration of residual connections enhances the stability of model training, and the integration logic at each key position adapts to the functional requirements of the corresponding modules, ensuring the progressiveness and effectiveness of feature processing. This provides a high-quality feature foundation with both spatiotemporal continuity and lip-sync for the subsequent super-resolution reconstruction stage, thereby ensuring the quality of the final conversational head animation.
[0048] In some embodiments, the multi-frame diffusion model employs a DDIM (Denoising Diffusion Implicit Models) accelerated sampling strategy during the sampling phase, with the number of sampling steps set to 10 to 20, reducing generation time and improving real-time performance. Specifically, the core parameters for the sampling phase are set as follows: the number of sampling steps S∈[10,20] (which can be dynamically adjusted according to real-time requirements, such as 10 steps for low-latency scenarios and 20 steps for high-quality scenarios), and a fixed value for the sampling variance (different from the random variance of DDPM) to ensure the determinism and efficiency of the sampling process. Based on the multi-frame diffusion model architecture, the sampling process is executed according to the implicit sampling logic of DDIM: starting from random noise, the model predicts the denoised mean vector, and denoises are progressively reversed using a recursive formula; in each sampling step, only the sampling result of the previous step and the mean vector output by the model are needed, without calculating complex noise distributions, significantly simplifying the sampling computation; after 10-20 recursive steps, a low-resolution face video frame sequence is directly output, completing the sampling generation. Compared to the traditional DDPM sampling process that requires thousands of steps, the 10-20 steps of DDIM sampling in this embodiment can shorten the generation time of low-resolution face video frame sequences by an order of magnitude, ensuring the real-time performance of subsequent super-resolution reconstruction and video output, and adapting to the real-time synthesis requirements of conversational head animation.
[0049] Step S103: Input the low-resolution face video frame sequence, the reference face image features, and the audio feature vector into the super-resolution model. The super-resolution model is used to perform feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual feature fusion branching steps to generate a high-resolution face video frame sequence.
[0050] In some embodiments, the super-resolution model includes a feature extractor, a Transformer encoder, and a reconstruction network connected in sequence; the feature extractor is composed of multiple residual convolutional modules connected in series, used to extract deep features from the low-resolution face video frame sequence; the Transformer encoder has a built-in spatiotemporal self-attention mechanism, used to calculate the second spatial attention weight of the deep features in a single frame and the second temporal attention weight between adjacent frames, to obtain encoded features containing spatiotemporal dependency information; the second spatial attention weight is used to assign weights to the image spatial structure features in the deep features, and the second temporal attention weight is used to assign weights to the inter-frame temporal evolution features in the deep features; the reconstruction network is used to map the encoded features to high-resolution image features and output the high-resolution face video frame sequence.
[0051] In this embodiment, the super-resolution model is an end-to-end deep learning architecture, with three core modules connected sequentially in the order of "feature extractor → Transformer encoder → reconstruction network". Specifically, the feature extractor converts low-resolution face video frame sequences into deep feature representations; the Transformer encoder mines the spatiotemporal dependencies in the deep features, generating encoded features containing spatiotemporal dependency information; and the reconstruction network inversely maps the encoded features containing spatiotemporal dependency information to high-resolution image features, ultimately outputting a high-resolution face video frame sequence. The input and output dimensions of each module are precisely matched to ensure the coherence of feature transfer.
[0052] The core function of the feature extractor is to extract deep features from low-resolution face video frame sequences. It consists of multiple residual convolutional modules connected in series. The key role of these residual convolutional modules is to avoid the loss of shallow features and enhance the expressive power of deep features (residual connections are existing technology and will not be detailed here). After the low-resolution face video frame sequence is input into the feature extractor, it undergoes layer-by-layer fusion processing by multiple residual convolutional modules, gradually removing redundant information and strengthening key features, ultimately outputting deep features. These deep features contain the image spatial structure features and inter-frame temporal evolution features of the low-resolution face video frame sequence, providing high-quality input for the subsequent spatiotemporal modeling of the Transformer encoder.
[0053] The core function of the Transformer encoder is its built-in spatiotemporal self-attention mechanism. This mechanism calculates the second spatial attention weights for single frames and the second temporal attention weights between adjacent frames for deep features, uncovering spatiotemporal dependencies within these features and outputting encoded features containing spatiotemporal dependency information. The Transformer encoder consists of multiple identical stacked encoding layers. Each layer contains a spatiotemporal self-attention sublayer and a feedforward network sublayer. Both sublayers are configured with residual connections and layer normalization (existing technology; this description only illustrates the function and does not elaborate on details) to ensure stable training and effective feature propagation. The input deep features are in feature map format. To adapt to the computational logic of the spatiotemporal self-attention mechanism, dimensionality reorganization and linear transformation are performed first, converting the feature maps into sequential feature forms, which serve as the input to the spatiotemporal self-attention sublayer.
[0054] Among them, the spatiotemporal self-attention sublayer achieves joint modeling of spatiotemporal dependencies by computing the second spatial attention weight and the second temporal attention weight in parallel: a. Second Spatial Attention Weight Calculation: For the sequence features corresponding to each frame, the second spatial attention weight of a single frame is obtained by calculating the similarity of the attention vectors. This second spatial attention weight is used to assign weights to the image spatial structure features in the deep features, thereby enhancing the rationality of the spatial structure of a single frame.
[0055] b. Calculation of Second Temporal Attention Weights: For feature vectors at corresponding positions in different frames, the second temporal attention weights between adjacent frames are obtained by calculating the similarity of attention vectors. These second temporal attention weights are used to assign weights to the inter-frame temporal evolution features in deep features, thereby enhancing the temporal continuity between frames.
[0056] After weighted fusion of the second spatial attention weight and the second temporal attention weight, the output features of a single coding layer are obtained through residual connection, layer normalization and feedforward network sub-layer processing; after stacking multiple coding layers, the output features contain spatiotemporal dependent information.
[0057] The core function of the reconstruction network is to map encoded features containing spatiotemporal dependency information into high-resolution image features, ultimately outputting a high-resolution face video frame sequence. The specific process is as follows: First, the encoded features containing spatiotemporal dependency information are inversely recombined in terms of dimensions to restore the feature map format. Second, the feature map resolution is progressively increased to the target high-resolution size through a transposed convolutional module, while adjusting the number of feature channels. Third, the number of feature channels is adjusted to the number of channels corresponding to the RGB image format through the output convolutional layer, obtaining high-resolution image features. Finally, the high-resolution image features are arranged sequentially along the time dimension to obtain a continuous high-resolution face video frame sequence, completing the mapping from encoded features containing spatiotemporal dependency information to a high-resolution face video frame sequence.
[0058] In summary, the collaborative workflow of the super-resolution model in this embodiment is as follows: low-resolution face video frame sequence → feature extractor (multiple residual convolutional modules cascaded) → deep features (including image spatial structure features and inter-frame temporal evolution features) → Transformer encoder (with built-in spatiotemporal self-attention mechanism) → calculation of the second spatial attention weight of a single frame and the second temporal attention weight between adjacent frames → encoded features containing spatiotemporal dependency information → reconstruction network → high-resolution image features → high-resolution face video frame sequence. Throughout the process, each module progresses layer by layer and collaborates precisely. The feature extractor provides a high-quality foundation for spatiotemporal modeling, the Transformer encoder strengthens spatiotemporal dependency by calculating the second spatial attention weight and the second temporal attention weight to ensure sequence consistency, and the reconstruction network achieves precise resolution enhancement, ensuring that the output high-resolution face video frame sequence possesses both clear spatial details and coherent temporal motion.
[0059] As can be seen, the super-resolution model in this embodiment effectively solves the problems of spatial detail loss and temporal consistency disruption during the mapping process from low-resolution face video frame sequences to high-resolution face video frame sequences through the architecture design and collaborative process of "feature extractor-Transformer encoder-reconstruction network". The series design of multiple residual convolutional modules in the feature extractor ensures the integrity of deep features; the spatiotemporal self-attention mechanism of the Transformer encoder effectively mines the spatiotemporal dependencies in deep features by accurately calculating the second spatial attention weight and the second temporal attention weight, avoiding inter-frame motion abrupt changes and single-frame structural misalignment; the reconstruction network achieves a smooth resolution improvement and detail supplementation. The final output high-resolution face video frame sequence has clear facial texture, reasonable facial feature structure and coherent motion state, providing a high-quality foundation for subsequent generation of conversation head animation videos that conform to the target scene.
[0060] In some embodiments, the super-resolution model further integrates a bidirectional optical flow alignment unit, which is deployed between the Transformer encoder and the reconstruction network to calculate the forward and backward optical flow between adjacent frames, construct inter-frame motion constraint relationships, perform alignment optimization on the encoded features, and output temporally aligned encoded features.
[0061] In this embodiment, the optimized super-resolution model adopts a sequential architecture of "feature extractor → Transformer encoder → bidirectional optical flow alignment unit → reconstruction network". The core division of labor of each module remains as follows: the feature extractor is responsible for extracting deep features from the low-resolution face video frame sequence; the Transformer encoder is responsible for mining the spatiotemporal dependencies in the deep features and generating encoded features containing spatiotemporal dependency information; the bidirectional optical flow alignment unit is responsible for temporal alignment optimization of the encoded features containing spatiotemporal dependency information; the reconstruction network is responsible for mapping the temporally aligned encoded features to high-resolution image features, and finally outputting a high-resolution face video frame sequence. The newly added bidirectional optical flow alignment unit does not change the core structure of the original modules, but only realizes the temporal optimization function through precise feature reception and output.
[0062] The core function of the bidirectional optical flow alignment unit is to calculate the forward and backward optical flow between adjacent frames, construct the inter-frame motion constraint relationship, perform alignment optimization on the coding features containing spatiotemporal dependency information, and output the temporally aligned coding features.
[0063] In the specific implementation, the input encoded features containing spatiotemporal dependency information have a dimension of T×N×128 (T is the number of frames, N is the sequence length after flattening the features of a single frame, and 128 is the number of feature channels). To adapt to the optical flow calculation logic, the encoded features are first reverse-reorganized in terms of dimension, restoring them from the sequence feature format to the feature map format (dimension of T×H_l×W_l×128, where H_l×W_l is the low-resolution size of a single frame), resulting in a sequence of encoded feature maps.
[0064] For the sequence of encoded feature maps, select pairs of encoded feature maps from adjacent frames in chronological order (frame t and frame t+1, where t ranges from 1 to T-1), and calculate the forward optical flow and backward optical flow respectively: ① Forward optical flow: Characterizes the direction and amplitude of motion of a pixel in the coded feature map of frame t to the corresponding pixel in the coded feature map of frame t+1, and is used to describe the motion trend from the previous frame to the next frame; ② Backward optical flow: Characterizes the direction and amplitude of motion of a pixel in the coded feature map of frame t+1 to the corresponding pixel in the coded feature map of frame t, and is used to describe the motion trend from the next frame to the previous frame.
[0065] By co-computing bidirectional optical flow, complete inter-frame motion constraints can be constructed, avoiding motion estimation bias caused by unidirectional optical flow calculation.
[0066] Based on the calculated forward and backward optical flows, the sequence of encoded feature maps is optimized for temporal alignment. Specifically, using the coded feature map of frame t as a reference, the pixel positions of the coded feature map of frame t+1 are corrected according to the forward optical flow; simultaneously, using the coded feature map of frame t+1 as a reference, the pixel positions of the coded feature map of frame t are verified according to the backward optical flow, ensuring that the spatial positions and motion trends of corresponding pixels in adjacent frame encoded feature maps match.
[0067] For feature regions that have been verified to have motion offset, the pixel feature values are adjusted by interpolation to achieve feature alignment; for feature regions with consistent motion trends, the original features are kept unchanged, and finally, a temporally aligned encoded feature map sequence is obtained.
[0068] The temporally aligned encoded feature map sequence is re-dimensioned to restore it to the sequence feature format (dimension T×N×128), and the temporally aligned encoded features are directly output to the reconstruction network.
[0069] In summary, in this embodiment, the collaborative workflow of the optimized super-resolution model is as follows: low-resolution face video frame sequence → feature extractor (multiple residual convolutional modules cascaded) → deep features (including image spatial structure features and inter-frame temporal evolution features) → Transformer encoder (with built-in spatiotemporal self-attention mechanism) → calculation of the second spatial attention weight of a single frame and the second temporal attention weight between adjacent frames → encoded features containing spatiotemporal dependency information → bidirectional optical flow alignment unit → calculation of forward and backward optical flow → construction of inter-frame motion constraint relationships → temporally aligned encoded features → reconstruction network → high-resolution image features → high-resolution face video frame sequence. Throughout the process, the bidirectional optical flow alignment unit receives the output features from the Transformer encoder and further enhances inter-frame temporal consistency through bidirectional optical flow calculation and alignment optimization, achieving a progressive processing of "spatiotemporal dependency modeling - temporal alignment optimization," providing more temporally coherent input features for the reconstruction network.
[0070] In some embodiments, the reconstruction network incorporates a dual-feature fusion branch unit, including an audio feature fusion branch and an identity feature fusion branch; the audio feature fusion branch is used to receive the audio feature vector and fuse the audio feature vector with the encoded feature channel by channel; the identity feature fusion branch is used to receive the reference face image feature and concatenate the reference face image feature with the encoded feature channel by channel; the dual-feature fusion branch unit is used to output the encoded feature after fusing audio information and identity information.
[0071] In this embodiment, the optimized reconstruction network adopts an internal architecture of a "dual-feature fusion branch unit resolution enhancement module," where the dual-feature fusion branch unit is a newly added core module, and the core function of the resolution enhancement module (mapping encoded features to high-resolution image features) remains unchanged. The overall architecture of the entire super-resolution model is a "feature extractor." Transformer encoder [Bidirectional Optical Flow Alignment Unit] Reconstructing the network (dual feature fusion branch unit) The resolution enhancement module precisely matches the input and output dimensions of each module. The newly added dual-feature fusion branch unit does not interfere with the core working logic of the original module, but only enhances the information integrity of the encoded features through multi-source feature fusion.
[0072] The core function of the dual-feature fusion branch unit is to fuse audio feature vectors and reference face image features into encoded features through parallel audio feature fusion branches and identity feature fusion branches, respectively, ultimately outputting encoded features that fuse audio and identity information. The basic implementation methods of channel concatenation and channel-by-channel fusion are existing technologies; only their application logic in this unit is described here, without elaborating on specific technical details.
[0073] The input to the dual-feature fusion branch unit consists of three parts, and dimensional adaptation is required to ensure the effectiveness of the fusion: ①Core input: Encoded features from the Transformer encoder (or bidirectional optical flow alignment unit), with dimensions of T×N×128 (T is the number of frames, N is the sequence length after flattening the features of a single frame, and 128 is the number of feature channels), denoted as F_enc; ② Audio input: The preprocessed audio feature vector has an original dimension of T×D_a (D_a is the audio feature dimension). The number of channels D_a is mapped to 128 through linear transformation to obtain the audio feature vector F_audio (dimension T×1×128) that matches F_enc. ③Identity Input: The preprocessed reference face image features have an original dimension of 1×D_id (D_id is the identity feature dimension). The time dimension is expanded to T through a broadcast mechanism, and then the number of channels D_id is mapped to 128 through a linear transformation to obtain the reference face image features F_id (dimension T×1×128) that match F_enc.
[0074] The core function of the audio feature fusion branch is to receive the adapted audio feature vector F_audio and fuse it with the encoded feature F_enc channel by channel. This strengthens the correlation between the encoded features and the audio information, ensuring the lip-sync of the subsequently generated video. Specifically, the encoded feature F_enc (T×N×128) and the audio feature vector F_audio (T×1×128) are added channel by channel. That is, each channel feature of F_enc is summed element-wise with the corresponding channel feature of F_audio to obtain the intermediate feature F_enc_audio (dimension T×N×128) of the fused audio information. This channel-by-channel fusion method ensures that the features of each pixel in the encoded features carry the audio information at the corresponding moment, guaranteeing accurate matching between the features of key areas such as the mouth and the audio pronunciation state.
[0075] The core function of the identity feature fusion branch is to receive the adapted reference face image feature F_id and concatenate it with the encoded feature F_enc through a channel concatenation process. This strengthens the identity attributes of the encoded feature and ensures that the face identity in the subsequently generated video is consistent with the reference face. Specifically, the encoded feature F_enc (T×N×128) and the reference face image feature F_id (T×1×128) are concatenated along the channel dimension. After concatenation, the feature dimension becomes T×N×256, yielding the intermediate feature F_enc_id with fused identity information. Channel concatenation completely preserves the identity feature information of the reference face, ensuring that the encoded feature consistently maintains the core identity attributes of the reference face, such as facial contours and features, during subsequent reconstruction, thus preventing the loss or distortion of identity information.
[0076] The dual-feature fusion branch unit performs weighted fusion of the intermediate features (F_enc_audio, F_enc_id) output by the two branches using preset fusion weights. Specifically, the number of channels of F_enc_id is first compressed from 256 to 128 through a 1×1 convolutional layer, so that the dimensions of the two intermediate features are unified to T×N×128. Then, the two are weighted and summed according to preset weights (adaptively learned during training) to obtain the encoded feature F_enc_fusion (dimension T×N×128) after fusing audio information and identity information, and outputs it to the resolution enhancement module of the reconstruction network.
[0077] In summary, see Figure 2 , Figure 2 This is a schematic diagram of the structure of the super-resolution model provided in the embodiments of this application. The workflow of the optimized super-resolution model is as follows: Figure 2As shown, the low-resolution face video frame sequence is first input into the "feature extractor" module (composed of multiple residual convolutional modules) to extract deep visual features of the frames. Simultaneously, the features undergo "positional encoding" to preserve the temporal positional information of the frames, providing a foundation for subsequent spatiotemporal dependency modeling. The extracted and encoded features are then input into the core sub-modules of the Transformer encoder, including a self-attention module and an optical flow module. The self-attention module captures the spatiotemporal dependencies within the features, strengthening the consistency between intra-frame facial structures (such as the position of facial features) and inter-frame motion trends (such as head rotation and mouth opening / closing). The optical flow module calculates the forward and backward optical flow between adjacent frames, constructing inter-frame motion constraints to further optimize temporal continuity and avoid inter-frame motion stuttering or misalignment. The features after spatiotemporal modeling are then input into the "reconstruction module" along with the input "audio features." The reconstruction module fuses the audio features through a feedforward network layer, specifically optimizing details in the mouth region to match mouth movements with the rhythm of audio pronunciation. After concatenating the features output by the reconstruction module with the features of the reference image, the details are first optimized by the convolution module, and then the features are upsampled and magnified to the target resolution. At the same time, the upsampled features are added and fused with relevant features to finally obtain a complete high-resolution face video frame sequence containing facial details of the target person and matching background, thus completing the entire super-resolution model processing flow.
[0078] As can be seen, this embodiment achieves the synchronous fusion of audio feature vectors and reference face image features into coded features by incorporating a dual-feature fusion branch unit into the reconstruction network. The audio feature fusion branch strengthens the correlation between coded features and audio, ensuring the lip-sync of the subsequently generated video; the identity feature fusion branch fully preserves the identity attributes of the reference face, ensuring the consistency of identity between the generated face and the reference face. The high-resolution face video frame sequence output by the optimized super-resolution model not only possesses clear spatial details and coherent temporal motion, but also accurately matches the target audio rhythm and continues the core features of the reference face, providing a crucial guarantee for generating high-quality conversational head animation videos that conform to the target scene.
[0079] Step S104: Output a high-resolution conversation head animation video that matches the target scene based on the high-resolution face video frame sequence.
[0080] In some embodiments, the step of outputting a high-resolution conversational head animation video that conforms to the target scene based on the high-resolution face video frame sequence includes: integrating the high-resolution face video frame sequence in chronological order to form a video stream; performing quality verification on the video stream to confirm that the video stream meets the requirements of identity consistency, temporal continuity, lip-sync, and high resolution; and outputting the high-resolution conversational head video.
[0081] The specific process for identity consistency verification is as follows: Facial features are extracted from multiple randomly selected keyframes in the video stream (e.g., one frame out of every 10 frames) and compared with the features of a reference facial image. The core criterion is that the matching degree between the facial features of all keyframes and the features of the reference facial image meets a preset threshold, ensuring that the core identity features of the face in the video stream, such as facial contours and textures, are consistent with the reference face, without identity distortion or confusion. If the threshold is not met, the process is returned to the super-resolution model for re-optimization.
[0082] The specific process of temporal continuity verification is as follows: The motion state between frames of the video stream is detected, with the core standard being: smooth transitions in facial movements (such as head rotation and mouth opening / closing) between adjacent frames, without obvious jumps or stuttering. By analyzing the pixel change amplitude and motion trajectory of adjacent frames, it is determined whether the temporal continuity meets the standard; if abrupt motion changes occur, the process returns to the bidirectional optical flow alignment unit for re-performing temporal alignment optimization.
[0083] The specific process of lip-phonetic synchronization verification is as follows: The audio stream in the video stream is compared synchronously with the mouth movement state of the corresponding frame. The core standard is that the time node of the mouth opening and closing action precisely matches the time node of the pronounced syllable in the audio stream, without significant delay or advance. By correlating the audio pronunciation rhythm with the movement trajectory of key mouth feature points, it is determined whether the lip-phonetic synchronization meets the standard; if it does not meet the standard, the process returns to the dual-feature fusion branch unit to re-optimize the audio feature fusion logic.
[0084] The specific process for verifying high resolution requirements is as follows: The resolution parameters of the video stream (such as pixel size and sharpness) are detected. The core standard is that the resolution of a single frame of the video stream is not lower than a preset high resolution threshold (such as 1920×1080), and the facial textures within the frame are clear, without blurring or noise. Verification is completed using resolution detection tools and image sharpness evaluation algorithms. If the requirements are not met, the process returns to the reconstruction network of the super-resolution model for re-resolution upscaling.
[0085] If the video stream passes all the quality checks mentioned above, it proceeds to the next output step. If any item fails to meet the standard, the problematic step is located based on the check results, and the corresponding front-end module is returned to be optimized again. After optimization, the quality check of this step is executed again until all items meet the standard.
[0086] Once the video stream passes quality verification, its format is fine-tuned according to the application requirements of the target scenario (such as online meetings, video playback, virtual live streaming, etc.). This includes video encoding format optimization, bitrate adjustment, and audio sampling rate adaptation to ensure the video can be played normally in the target application scenario and occupies a reasonable amount of storage space. The format-adapted video stream is then stored to a preset path or directly pushed to the target output terminal (such as a monitor, server, live streaming platform, etc.) to complete the final output of the high-resolution conversation head animation video. During the output process, the core parameters of the video (resolution, frame rate, duration, identity information, etc.) are recorded simultaneously to form an output log for subsequent traceability and management.
[0087] As can be seen, this embodiment ensures the reliability of the final high-resolution conversation head animation video through standardized timing integration and multi-dimensional quality verification processes. The timing integration stage guarantees the temporal consistency between the frame sequence and the audio stream; the multi-dimensional quality verification stage can accurately filter out video streams that do not meet the requirements, and improve the overall synthesis quality through backtracking optimization; the output format adaptation stage improves the video's scene compatibility. The final output high-resolution conversation head animation video can fully meet the core requirements of the target scene for identity consistency, temporal continuity, lip-sync, and high resolution, and has good practicality and adaptability.
[0088] As can be seen, in this embodiment, a reference face image and a target audio sequence are first acquired. Feature extraction is performed on the reference face image to obtain its features, and feature extraction is performed on the target audio sequence to obtain its audio feature vector. Next, the reference face image features and the audio feature vector are input into a multi-frame diffusion model to generate a continuous low-resolution face video frame sequence. Then, the low-resolution face video frame sequence, the reference face image features, and the audio feature vector are input into a super-resolution model. The super-resolution model performs feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual-feature fusion branching steps to generate a high-resolution face video frame sequence. Finally, a high-resolution conversational head animation video conforming to the target scene is output based on the high-resolution face video frame sequence. Compared to existing technologies, this application rapidly generates low-resolution videos with preliminary temporal consistency through a multi-frame diffusion model, and then generates animation videos that meet the requirements of identity consistency, temporal continuity, lip-sync, and high resolution through a super-resolution model. While ensuring the quality of the generated video, it significantly improves generation efficiency, reduces computational costs, and meets the high-quality application needs of digital humans, virtual anchors, and other applications.
[0089] In some embodiments, the method further includes a model training phase, in which a multi-dimensional loss function is constructed, including pixel difference loss, perceptual consistency loss, and lip-sync loss, for joint training and model parameter optimization of the multi-frame diffusion model and the super-resolution model; wherein the pixel-level MSE loss is used to measure the pixel difference between the generated image and the real image, the perceptual loss is calculated based on the feature map of the pre-trained VGG19 network to ensure visual perceptual consistency; the lip-sync loss is used to calculate the matching degree between the mouth region features and the audio feature vector, the mouth region features being extracted by a pre-trained facial keypoint detection network.
[0090] The multi-dimensional loss function L_total is a weighted sum of the sub-losses, with the formula L_total=α·L_pix+β·L_perc+γ·L_lip, where α, β, and γ are preset weight coefficients (adaptively adjusted during training to satisfy α+β+γ=1), which balance the influence weight of each sub-loss. This loss function is applied to the backpropagation process of model training, and the parameters of the multi-frame diffusion model and the super-resolution model are co-optimized by minimizing L_total.
[0091] The pixel difference loss (L_pix) employs pixel-level MSE loss, which is primarily used to measure the pixel difference between the generated image (low-resolution frames output by the multi-frame diffusion model and high-resolution frames output by the super-resolution model) and the real image. Specifically, it calculates the mean of the squared differences between the grayscale values (or RGB values) of corresponding pixels in the generated image and the real image, and uses this loss to constrain the pixel-level accuracy of the generated image.
[0092] The perceptual consistency loss (L_perc) is calculated based on the feature maps of the pre-trained VGG19 network. Its core function is to ensure visual perceptual consistency between the generated image and the real image. The pre-trained VGG19 network is existing technology; here, only the application logic is explained: the generated image and the real image are input into the pre-trained VGG19 network respectively, and feature maps of a certain intermediate layer (such as the conv4_3 layer) are extracted. The Euclidean distance between the two sets of feature maps is calculated as the perceptual consistency loss value. This loss constrains the high-level visual semantics of the generated image to be consistent with the real image.
[0093] The lip-pronunciation synchronization loss (L_lip) is used to calculate the matching degree between the mouth region features and the audio feature vector, ensuring the lip-pronunciation synchronization of the generated video. Specifically, it is implemented as follows: First, a pre-trained facial landmark detection network (existing technology) extracts key feature points (such as the corners of the mouth and the cupid's bow) from the generated image to construct a mouth region feature vector; then, the cosine similarity between this mouth region feature vector and the corresponding audio feature vector is calculated; 1 is subtracted from this similarity value to obtain the lip-pronunciation synchronization loss value, which constrains the matching of mouth movements with the audio pronunciation rhythm.
[0094] Steps S1-S6 are an embodiment of another conversation head animation synthesis method provided by the embodiments of this application, including the following steps: S1: Obtain the reference face image and the target audio sequence.
[0095] S2: Perform preliminary processing on the reference face image and the target audio sequence to obtain the processed dataset: use the designed super-resolution model to obtain 256×256 video frames from 64×64 video frames, generate 5 video frames at once in a single training or sampling, and set the length of the denoising step T to 200.
[0096] During the sampling process, 10 sampling steps were used on the CREMA-D dataset, while 20 sampling steps were used on the HDTF dataset, to adapt to the characteristics of different datasets and balance generation speed and quality. Using an A40 GPU with 40GB of memory to train the model, training the diffusion model on the CREMA-D dataset took approximately 50 hours, and training the super-resolution model also took approximately 50 hours. For the HDTF dataset, achieving satisfactory results required approximately 70 hours and 80 hours respectively. It is evident that directly inputting the super-resolution model results in a longer training time.
[0097] S3: To improve motion consistency and generation speed, this application uses a multi-frame diffusion model in the pixel domain of low-resolution images. Furthermore, the multi-frame diffusion model is used to predict the mean instead of following the prediction variance of DDPM; this design improves the performance of the multi-frame diffusion model in low-resolution scenarios. Simultaneously, this application adopts a U-Net architecture, which uses spatiotemporal convolution and spatiotemporal attention layers. Compared to computationally expensive 3D convolution and 3D attention layers, this design significantly reduces the model's parameter size and computational complexity. Specifically, a spatiotemporal convolution layer consists of one 2D convolution layer and one 1D convolution layer. The 2D convolution focuses on spatial domain information, while the 1D spatiotemporal convolution focuses on temporal domain information.
[0098] The multi-frame diffusion model here requires information from the input audio. Therefore, in addition to the self-attention layer mentioned above, this application also adds a cross-attention layer to receive audio features extracted from the pre-trained audio encoder.
[0099] In this application, both the self-attention layer and the cross-attention layer contain convolutional layers, location embedding layers, and feedforward network layers. The self-attention layer takes the output of the previous network layer as input, while the cross-attention layer takes the output of the previous network layer and audio features as input. This application replaces the original convolutional layers of self-attention and cross-attention with spatiotemporal convolutional layers.
[0100] Specifically, in U-Net, this application adds self-attention layers to all downsampling, intermediate, and upsampling modules. Conversely, to reduce computational resource consumption, this application only adds cross-attention layers to the first downsampling module, all intermediate modules, and the last upsampling module. When obtaining audio features, this application uses a sliding window technique, where one video frame corresponds to multiple audio features, and the audio features of adjacent video frames overlap. This allows the model to obtain more information from the audio, thereby enhancing the continuity between generated video frames.
[0101] To obtain higher resolution video, a super-resolution model was designed, further improving the resolution of each frame based on the video frames generated by the multi-frame diffusion model described above. Furthermore, the super-resolution model can supplement details and slightly adjust the mouth region, making the generated video frames more closely match the input audio. Borrowing from the Transformer structure of VSRTransformer, some modules and structures were modified and added to implement the super-resolution model presented here. This Transformer consists of a feature extractor, a Transformer encoder, and a reconstruction network module. The method adds a feedforward network layer to the reconstruction network module to receive additional input audio features, and adds a convolutional module after the reconstruction network module to incorporate information from the reference image.
[0102] Specifically, given a video sequence, a series of residual modules are first used to extract video features. Then, the Transformer encoder maps the obtained features to a series of sequential representations, including spatiotemporal convolutional self-attention modules and feedforward layers with bidirectional optical flow design. Furthermore, the method in this chapter adds several feedforward network layers to the reconstruction module to incorporate audio features. A i As a condition, after processing and inference, the reconstruction network module reconstructs high-resolution video frames based on the obtained features and additional audio features from the input. Finally, the model uses the reference image... x refThe output of the reconstruction network module is added, connected in the channel dimension, additional information from the reference image is introduced, and processed using a convolutional block containing a one-dimensional temporal convolutional layer and a two-dimensional spatial convolutional layer to obtain the final video frame.
[0103] S4: Train the conversation head animation synthesis model based on the diffusion model and Transformer to obtain the trained conversation head animation synthesis model based on the diffusion model and Transformer. During the training phase, the learning rate was set to 1e-5, and the Adam algorithm was selected as the optimization algorithm. After 15,000 epochs, the learning rate decayed to 0.
[0104] Then, the model was trained on an experimental platform with 24GB of memory and an NVIDIA RTX 2080 Ti, using PyTorch as the backend.
[0105] S5: Input the processed audio sequence into the trained conversation head animation synthesis model based on the diffusion model and Transformer to obtain the result; S6: The synthesis results are evaluated using peak signal-to-noise ratio (PSNR), structural similarity, perceptual distance, and frame rate as evaluation metrics. PSNR and SSIM are two traditional image quality metrics.
[0106] Please see Figure 3 , Figure 3 The diagram shows a synthesis model of the conversation head animation provided in the embodiments of this application, such as... Figure 3 As shown, the conversation head animation synthesis method of this application first fuses three types of visual inputs: a reference image containing the identity information of the target person, a noisy image required by the diffusion model, and motion frames that assist in motion trends. At the same time, it combines the audio features extracted and preprocessed by the pre-trained audio encoder and inputs them into a diffusion model with a built-in self-attention module (covering the entire U-Net module to capture spatiotemporal dependencies) and a cross-attention module (deployed only in specific modules to fuse audio features). Then, the diffusion model generates a sequence of low-resolution face video frames containing multiple consecutive frames based on the mean prediction training strategy and the DDIM accelerated sampling strategy. Finally, the low-resolution face video frame sequence is input into the Transformer super-resolution model, and audio features are input again to optimize mouth details. The Transformer super-resolution model, through feature extraction, spatiotemporal convolutional self-attention, and bidirectional optical flow feedforward layer processing, fuses the identity information of the reference image and outputs high-resolution conversation head animation frames, completing the entire process from multi-source input to the final synthesis result. This embodiment takes into account generation efficiency, temporal continuity, and lip-phonetic synchronization.
[0107] The above approach has the following beneficial effects: 1. This application employs a two-stage cascaded architecture of "low-resolution multi-frame diffusion + Transformer super-resolution" to decouple the generation task into "first, fast consistency determination, then pixel-by-pixel refinement." The diffusion side can output multiple consecutive low-resolution video blocks with a single sampling, significantly reducing the number of sampling steps compared to traditional "frame-by-frame denoising" methods. The Transformer side utilizes spatiotemporal self-attention and bidirectional optical flow constraints to perform parallel upsampling to the target resolution, resulting in fast single-frame processing speed and rapid generation of the entire video. The generation speed is significantly improved compared to Diffused-Head and on par with GAN solutions, but with significantly superior visual quality.
[0108] 2. This application introduces "audio-image cross-attention" and "sliding window audio overlap" mechanisms within the diffusion model, enabling each frame to "predict" the preceding and following speech content during the sampling stage. Combined with motion prior guidance, this effectively suppresses lip tremors and head flicker caused by frame-by-frame independent sampling. In the super-resolution stage, the application utilizes a three-branch approach of "optical flow-alignment-fusion" to incorporate the motion fields of adjacent frames as explicit constraints into the loss function, significantly optimizing the temporal continuity of the synthesized video and outperforming the existing best GAN scheme.
[0109] 3. This application reconstructs the diffusion training objective through the "mean prediction + DDIM acceleration" strategy, abandons the traditional noise regression, and directly predicts the mean of the denoised image, which greatly improves the convergence speed and can achieve the image fidelity of more training rounds in fewer training rounds. The super-resolution loss is combined with MSE, VGG perceptual loss and lip-sync loss, and the weights of multiple tasks are automatically balanced, which significantly improves the image fidelity. While maintaining the same level of lip-sync performance as Wav2Lip, the overall image naturalness and identity consistency are significantly better than the pure lip-sync patch method.
[0110] To verify the effectiveness and superiority of the proposed model in video generation tasks, a qualitative comparison was made through the visual effects of experiments, and the method of this application was comprehensively evaluated by combining quantitative indicators.
[0111] Datasets. This application uses two conversation head datasets, CREMA-D and HDTF, to validate the effectiveness of the proposed model. CREMA-D is a dataset containing 7442 original segments from 91 speakers. These videos are from 48 men and 43 women, aged 20 to 74, representing diverse ethnicities and backgrounds. This application randomly selected 81 speakers for training and used the remaining 10 speakers for testing. HDTF contains conversation head videos at 720p or 1080p resolution from over 300 identities, totaling 16 hours in length. This paper randomly selected 150 videos for training and used the remaining data as the test set. During the experiments, the HDTF videos were cropped and downsampled to 256×256 resolution for both training and sampling.
[0112] Evaluation Metrics. This application evaluates the proposed method using both visual results and quantitative metrics. Five metrics are used to assess the quality of the generated video: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Learned Perceptual Patch Similarity (LPIPS), Video Frame Distance (FVD), and Lip-Sync Score (Sync). PSNR and SSIM are two traditional metrics for measuring image quality, representing pixel-wise consistency and structural similarity, respectively. LPIPS is a deep learning-based perceptual similarity measure, which is a more human-like evaluation metric. FVD is a video evaluation metric used to assess the distance between the generated video and the real video. The Sync score assesses the synchronization between the generated video and the audio, and this metric can intuitively reflect lip-sync. Lip-sync is crucial for audio-driven conversational head animation synthesis tasks.
[0113] Implementation Details. This application's method first resizes the input image to 64×64, then conducts experiments related to the diffusion model. Next, the designed super-resolution model is used to obtain 256×256 video frames from the 64×64 video frames. In the experiments, five video frames are generated at once in a single training or sampling run. During training, the length of the denoising step T is set to 200. During sampling, 10 sampling steps are used on the CREMA-D dataset, and 20 sampling steps are used on the HDTF dataset. This method uses an A40 GPU with 40GB of memory to train the model. On the CREMA-D dataset, training the diffusion model takes approximately 50 hours, and training the super-resolution model also takes approximately 50 hours. However, for the HDTF dataset, achieving satisfactory results requires approximately 70 hours and 80 hours respectively.
[0114] I. Comparative Experiment Qualitative comparison. The methods described in this application will be compared with four other methods (MakeItTalk, Wav2Lip, Sadtalker, and Diffused-Head) and ground-truth video frames of actual conversation heads. The results of all methods are as follows: Figure 4 and Figure 7 As shown.
[0115] The MakeItTalk method utilizes facial markers as an intermediate representation. It predicts the location of facial markers from speech and translates them using an image-to-image translation network to generate corresponding animated images. While it implicitly controls head pose through facial markers, it doesn't explicitly model them, leading to inconsistencies between lip movements and audio. In contrast, the Wav2Lip method uses a GAN-based model. This model takes an upper half of the face image and a speech segment as input and generates a lower half of the face image. It doesn't use an intermediate representation but relies on an external lip-sync expert model to control the model's output. However, the head pose in the video generated using the Wav2Lip method can be completely fixed because it's controlled by only a single reference image, and only the upper half of the head's movement is provided by that reference image. Another method, Sadtalker, utilizes realistic 3D motion coefficients generated by a model to drive the generation of talk head videos from a single image, producing high-quality talk head videos. Notably, all experiments with Sadtalker used a GAN-based face enhancer, an auxiliary module that significantly improves image quality. However, this method also has significant drawbacks. This method focuses only on lip and eye movements, neglecting facial expressions in other areas, resulting in a stiff and unnatural appearance. Diffused-Head is a conversation head generation method based on a diffusion model. It generates one video frame at a time, corresponding to a given identity, while the given reference image remains fixed throughout the generation process. This method uses a pre-trained audio encoder to extract speech information and inputs the extracted speech features into the model. While this method can generate relatively realistic conversation heads, its drawbacks include long generation times and inconsistencies.
[0116] This application first uses qualitative methods to compare the effects of different methods with the method proposed in this application. It can be clearly seen that... Figure 4 and Figure 7In particular, in the areas marked by the box in the figure, the consistency between the method in this application and the ground-truth of the actually filmed conversational head video frames is superior to the other methods mentioned above. Furthermore, the method in this application incorporates more eye and head movements, making the generated results more realistic and natural. Qualitative comparisons show that the method in this application, using a diffusion model for low-resolution generation and a Transformer for image super-resolution, achieves satisfactory visual results.
[0117] Quantitative Comparison. Next, this application uses quantitative methods to compare the performance differences between different methods and the method proposed in this application. For quantitative evaluation, a video segment is generated from the CREMA-D and HDTF datasets, given a first frame and an audio sequence. This application calculates five metrics: PSNR, SSIM, LPIPS, FVD, and Sync, to compare the degree of matching between the generated results and the real reference video. The quantitative comparison results are shown in Tables 1 and 2. It should be emphasized that the pre-trained weights given by the Diffused-Head method were only trained on the CREMA-D dataset; therefore, the method in this application cannot be compared with the proposed method on the HDTF dataset. Table 1 is as follows: Table 1. Quantitative comparison results of different methods on the CREMA-D dataset.
[0118] Table 2 is as follows: Table 2 Quantitative comparison results of different methods on the HDTF dataset
[0119] The experiments show that the method presented in this application outperforms MakItTalk, Wav2Lip, and Diffused-Head on all image quality metrics. Specifically, the method's PSNR score is on average about 2.5 points higher than Wav2lip and significantly higher than the other two methods. Furthermore, the method achieves the highest SSIM and LPIPS scores. Experimental results demonstrate that the method can generate high-quality conversational head videos. Moreover, due to the randomness inherent in the diffusion model, the model can generate videos with more facial expressions and movements. While the method does not outperform Sadtalker on the FVD metrics on both datasets or the LPIPS metrics on the HDTF dataset, it still demonstrates high competitiveness on these two metrics; and it performs better on both PSNR and SSIM metrics. Although the method does not outperform Wav2lip and Sadtalker on the Sync metric, it still outperforms the other two methods and maintains a high level of lip-congruence consistency. A possible explanation for this is that the high randomness of the diffusion model leads to a relatively low Sync score for consistency.
[0120] The MakeItTalk method produces videos that lack dynamism and perform poorly in terms of image quality. The Wav2lip method achieves the best lip-congruity, but its generated videos, apart from the mouth area, show no movement and sometimes exhibit strange distortions at the edges, making them look unrealistic. The Sadtalker method achieves the highest FVD score and also scores highly on other metrics. However, this method requires face cropping and cannot drive other parts besides the face. Furthermore, its 3D-based motion sometimes leads to abnormal head postures, making the generated videos unrealistic. The Diffused-Head method, like the method in this application, uses a diffusion model to synthesize conversational head videos, offering more variation compared to the previous three methods; however, this method generates higher-quality videos and has a significant advantage in generation speed. In addition, the results generated by this method include variations in facial expressions and postures, making the videos more realistic and natural.
[0121] II. Ablation Experiment This application designed several ablation experiments to demonstrate the rationality and effectiveness of the model design. This application used PSNR, SSIM, and Sync metrics to quantitatively compare videos generated by different ablation methods; the results are shown in Table 3. Table 3 Quantitative comparison results of ablation experiments
[0122] Ablation experiments on the predicted mean value of the diffusion model. This application investigates the performance of the diffusion model in predicting the mean value. Preliminary experiments trained two diffusion models: one predicting the mean value during diffusion, and the other predicting noise during diffusion, all other conditions being identical. Video frames generated by these two models were compared using PSNR, SSIM, and Sync metrics. Quantitative comparison results are shown in items 1 and 5 of Table 3.
[0123] As can be seen, the diffusion model predicting the mean outperforms the diffusion model predicting noise on all image quality metrics. The explanation in this application is that, given limited time and training conditions, the former converges faster and requires less data to produce satisfactory results. Therefore, this application selects the diffusion model that predicts the mean during the diffusion process in subsequent experiments.
[0124] Ablation experiments of the cross-attention layer. This application investigates the effect of the cross-attention layer. In the model of this application, the cross-attention layer is used to integrate features of video frames and audio. To evaluate the design of this application, a diffusion model without the cross-attention layer was trained here, with all other structures and conditions being completely consistent. Quantitative comparison results are shown in items 2 and 5 of Table 3.
[0125] Here it can be seen that the cross-attention layer added in this application improves image quality, especially the Sync metric. This is likely because the cross-attention layer can fully blend audio and image features, generating video frames that better match the input audio.
[0126] Ablation experiments on the multi-frame diffusion model. To evaluate the effectiveness of the proposed multi-frame diffusion model, a diffusion model without using a multi-frame structure was trained while keeping other conditions the same. This model takes a reference image, two motion frames, and audio features as input, and generates only one video frame at a time.
[0127] As shown in items 3 and 5 of Table 3, the reference model still maintains high generation quality and lip-sync. However, even under these conditions, the method in this application still outperforms it in all metrics. This is likely because the method in this application requires fewer iterations during generation, and the multi-frame design mitigates image distortion caused by iteration during generation. The multi-frame diffusion method also maintains good inter-frame consistency. Most importantly, the multi-frame diffusion method can generate multiple video frames at once, meaning that the method in this application can achieve a faster generation speed when generating videos of the same length.
[0128] Ablation experiments of the super-resolution Transformer. This application evaluates the effectiveness of the designed super-resolution Transformer. In this comparative experiment, the unmodified VSRTransformer was directly used as the super-resolution module for training. Quantitative comparison results are shown in items four and five of Table 3.
[0129] The method presented in this application outperforms methods using only the original VSRTransformer in all metrics. As can be seen, the improvements to the super-resolution Transformer in this application significantly enhance image quality and lip-sync. This application incorporates more audio and image information, which helps generate superior results.
[0130] Please see Figure 5 , Figure 5 This application provides a schematic diagram of the structure of a server, as shown in the embodiment of the present application. Figure 5 As shown, the server 10 includes a processor 101, a memory 103, a communication interface 102, and a computer program 1031, which is stored in the memory 103 and configured to be executed by the processor 101. The program includes a method for performing a conversation head animation synthesis method as described in the above embodiments.
[0131] The above primarily describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the server includes the corresponding hardware structure and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0132] This application embodiment can divide the server into functional units according to the above method example. For example, each function can be divided into different functional units, or two or more functions can be integrated into one processing module. The integrated unit can be implemented in hardware or as a software program module. It should be noted that the unit division in this application embodiment is illustrative and only represents a logical functional division, while other division methods may be used in actual implementation.
[0133] In the case of using integrated units, please refer to Figure 6 , Figure 6This application provides a functional unit structural block diagram of a conversation head animation synthesis device 6, which includes: Acquisition unit 601 is used to acquire reference face images and target audio sequences; Processing unit 602 is configured to extract features from the reference face image to obtain reference face image features, and extract features from the target audio sequence to obtain audio feature vectors; input the reference face image features and the audio feature vectors into a multi-frame diffusion model to generate a continuous low-resolution face video frame sequence; input the low-resolution face video frame sequence, the reference face image features, and the audio feature vectors into a super-resolution model, wherein the super-resolution model is used to perform feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual feature fusion branching steps to generate a high-resolution face video frame sequence; The output unit 603 is used to output a high-resolution conversation head animation video that conforms to the target scene based on the high-resolution face video frame sequence.
[0134] As can be seen, in this embodiment, a reference face image and a target audio sequence are first acquired. Feature extraction is performed on the reference face image to obtain its features, and feature extraction is performed on the target audio sequence to obtain its audio feature vector. Next, the reference face image features and the audio feature vector are input into a multi-frame diffusion model to generate a continuous low-resolution face video frame sequence. Then, the low-resolution face video frame sequence, the reference face image features, and the audio feature vector are input into a super-resolution model. The super-resolution model performs feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual-feature fusion branching steps to generate a high-resolution face video frame sequence. Finally, a high-resolution conversational head animation video conforming to the target scene is output based on the high-resolution face video frame sequence. Compared to existing technologies, this application rapidly generates low-resolution videos with preliminary temporal consistency through a multi-frame diffusion model, and then generates animation videos that meet the requirements of identity consistency, temporal continuity, lip-sync, and high resolution through a super-resolution model. While ensuring the quality of the generated video, it significantly improves generation efficiency, reduces computational costs, and meets the high-quality application needs of digital humans, virtual anchors, and other applications.
[0135] In some embodiments, the multi-frame diffusion model is built based on an improved U-Net architecture, which includes a downsampling module group, an intermediate feature fusion module, and an upsampling module group connected in sequence. Each of the downsampling module group, the intermediate feature fusion module, and the upsampling module group embeds a spatiotemporal convolutional unit. Each spatiotemporal convolutional unit consists of a two-dimensional spatial convolutional layer and a one-dimensional temporal convolutional layer connected in series, and both the two-dimensional spatial convolutional layer and the one-dimensional temporal convolutional layer have residual connection structures. The two-dimensional spatial convolutional layer is used to extract the image spatial structure features of a single-frame face, and these features characterize the spatial positional relationship between local regions within a single-frame face and the overall face. The one-dimensional temporal convolutional layer is used to model the inter-frame temporal evolution features between adjacent frames, and these features characterize the continuous relationship of face motion between adjacent frames. The spatiotemporal convolutional unit outputs a first feature map that fuses the image spatial structure features and the inter-frame temporal evolution features.
[0136] In some embodiments, the deep layer of the downsampling module group, the intermediate feature fusion module, and the shallow layer of the upsampling module group all integrate a dual attention fusion unit. The dual attention fusion unit includes a self-attention subunit and a cross-attention subunit connected in a serial order, and both the self-attention subunit and the cross-attention subunit are configured with residual connection structures. The input of the self-attention subunit is the first feature map output by the previous stage spatiotemporal convolution unit in its module, used to calculate the global attention weights within the first feature map of a single frame, and outputs a second feature map after self-attention weighting. The global attention weights include a first spatial attention weight and a first temporal attention weight. Inter-temporal attention weights are used to assign weights to the image spatial structure features in the first feature map, and the first temporal attention weights are used to assign weights to the inter-frame temporal evolution features in the first feature map. The input of the cross-attention sub-unit includes the second feature map and the audio feature vector, and is used to establish an association mapping between facial visual features and speech temporal features based on the second feature map and the audio feature vector, and output a third feature map that fuses audio information. The facial visual features include the image spatial structure features, the inter-frame temporal evolution features and the global attention weights. The third feature map is decoded by the upsampling module group and restored to the low-resolution facial video frame sequence.
[0137] In some embodiments, the super-resolution model includes a feature extractor, a Transformer encoder, and a reconstruction network connected in sequence; the feature extractor is composed of multiple residual convolutional modules connected in series, used to extract deep features from the low-resolution face video frame sequence; the Transformer encoder has a built-in spatiotemporal self-attention mechanism, used to calculate the second spatial attention weight of the deep features in a single frame and the second temporal attention weight between adjacent frames, to obtain encoded features containing spatiotemporal dependency information; the second spatial attention weight is used to assign weights to the image spatial structure features in the deep features, and the second temporal attention weight is used to assign weights to the inter-frame temporal evolution features in the deep features; the reconstruction network is used to map the encoded features to high-resolution image features and output the high-resolution face video frame sequence.
[0138] In some embodiments, the super-resolution model further integrates a bidirectional optical flow alignment unit, which is deployed between the Transformer encoder and the reconstruction network to calculate the forward and backward optical flow between adjacent frames, construct inter-frame motion constraint relationships, perform alignment optimization on the encoded features, and output temporally aligned encoded features.
[0139] In some embodiments, the reconstruction network incorporates a dual-feature fusion branch unit, including an audio feature fusion branch and an identity feature fusion branch; the audio feature fusion branch is used to receive the audio feature vector and fuse the audio feature vector with the encoded feature channel by channel; the identity feature fusion branch is used to receive the reference face image feature and concatenate the reference face image feature with the encoded feature channel by channel; the dual-feature fusion branch unit is used to output the encoded feature after fusing audio information and identity information.
[0140] In some embodiments, the processing unit 602, in outputting a high-resolution conversational head animation video conforming to the target scene based on the high-resolution face video frame sequence, is further configured to: integrate the high-resolution face video frame sequence in chronological order to form a video stream; perform quality verification on the video stream to confirm that the video stream meets the requirements of identity consistency, temporal continuity, lip-sync, and high resolution; and output the high-resolution conversational head video.
[0141] This application provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implement the steps of any possible embodiment of the method.
[0142] The above descriptions are merely embodiments of the present invention, and common knowledge such as specific structures and / or characteristics in the solutions are not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the structure of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
Claims
1. A method for synthesizing conversational head animation, characterized in that, The method includes: A reference face image and a target audio sequence are acquired. Feature extraction is performed on the reference face image to obtain reference face image features, and feature extraction is performed on the target audio sequence to obtain an audio feature vector. The reference face image features and the audio feature vector are input into a multi-frame diffusion model to generate a continuous low-resolution face video frame sequence. The low-resolution face video frame sequence, the reference face image features, and the audio feature vector are input into the super-resolution model. The super-resolution model is used to perform feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual feature fusion branching steps to generate a high-resolution face video frame sequence. Output a high-resolution conversational head animation video that matches the target scene based on the high-resolution face video frame sequence.
2. The method according to claim 1, characterized in that, The multi-frame diffusion model is built on an improved U-Net architecture, which includes a downsampling module group, an intermediate feature fusion module, and an upsampling module group connected in sequence. The downsampling module group, the intermediate feature fusion module, and the upsampling module group all embed spatiotemporal convolutional units. The spatiotemporal convolutional unit is composed of a two-dimensional spatial convolutional layer and a one-dimensional temporal convolutional layer connected in series. Both the two-dimensional spatial convolutional layer and the one-dimensional temporal convolutional layer are provided with residual connection structures. The two-dimensional spatial convolutional layer is used to extract the image spatial structure features of a single frame face. The image spatial structure features are used to characterize the spatial positional relationship between the local region inside the single frame face and the overall face. The one-dimensional temporal convolutional layer is used to model the inter-frame temporal evolution features between adjacent frames. The inter-frame temporal evolution features are used to characterize the continuous relationship of face motion between adjacent frames. The spatiotemporal convolutional unit is used to output a first feature map that fuses the image spatial structure features and the inter-frame temporal evolution features.
3. The method according to claim 2, characterized in that, The deep layer of the downsampling module group, the intermediate feature fusion module, and the shallow layer of the upsampling module group all integrate a dual attention fusion unit. The dual attention fusion unit includes a self-attention subunit and a cross-attention subunit connected in a serial order, and both the self-attention subunit and the cross-attention subunit are configured with a residual connection structure. The input of the self-attention sub-unit is the first feature map output by the previous spatiotemporal convolution unit in its module, which is used to calculate the global attention weights within the first feature map of a single frame and output the second feature map after self-attention weighting; the global attention weights include a first spatial attention weight and a first temporal attention weight, the first spatial attention weight is used to assign weights to the image spatial structure features in the first feature map, and the first temporal attention weight is used to assign weights to the inter-frame temporal evolution features in the first feature map. The input of the cross-attention subunit includes the second feature map and the audio feature vector, which is used to establish an association mapping between facial visual features and speech temporal features based on the second feature map and the audio feature vector, and output a third feature map that fuses audio information. The facial visual features include the image spatial structure features, the inter-frame temporal evolution features and the global attention weights. The third feature map is decoded by the upsampling module group and restored to the low-resolution facial video frame sequence.
4. The method according to claim 3, characterized in that, The super-resolution model includes a feature extractor, a Transformer encoder, and a reconstruction network connected in sequence. The feature extractor is composed of multiple residual convolutional modules connected in series, and is used to extract deep features from the low-resolution face video frame sequence. The Transformer encoder incorporates a spatiotemporal self-attention mechanism to calculate the second spatial attention weight of a single frame and the second temporal attention weight between adjacent frames for the deep features, thereby obtaining encoded features containing spatiotemporal dependency information. The second spatial attention weight is used to assign weights to the image spatial structure features in the deep features, and the second temporal attention weight is used to assign weights to the inter-frame temporal evolution features in the deep features; The reconstruction network is used to map the encoded features to high-resolution image features and output the high-resolution face video frame sequence.
5. The method according to claim 4, characterized in that, The super-resolution model also integrates a bidirectional optical flow alignment unit, which is deployed between the Transformer encoder and the reconstruction network. This unit is used to calculate the forward and backward optical flow between adjacent frames, construct inter-frame motion constraint relationships, perform alignment optimization on the encoded features, and output temporally aligned encoded features.
6. The method according to claim 4 or 5, characterized in that, The reconstruction network incorporates a dual-feature fusion branch unit, including an audio feature fusion branch and an identity feature fusion branch; The audio feature fusion branch is used to receive the audio feature vector and fuse the audio feature vector with the encoded features channel by channel; The identity feature fusion branch is used to receive the reference face image features and perform channel concatenation of the reference face image features and the encoded features; The dual-feature fusion branch unit is used to output the coded features after fusing audio information and identity information.
7. The method according to claim 1, characterized in that, The step of outputting a high-resolution conversation head animation video that matches the target scene based on the high-resolution face video frame sequence includes: The high-resolution face video frame sequence is integrated in chronological order to form a video stream; The video stream is subjected to quality verification to confirm that it meets the requirements of identity consistency, temporal continuity, lip-sync, and high resolution. Output the high-resolution head video of the conversation.
8. A device for synthesizing animated head movements during conversation, characterized in that, The device includes: The acquisition unit is used to acquire reference face images and target audio sequences; The processing unit is configured to extract features from the reference face image to obtain reference face image features, and extract features from the target audio sequence to obtain audio feature vectors; input the reference face image features and the audio feature vectors into a multi-frame diffusion model to generate a continuous low-resolution face video frame sequence; input the low-resolution face video frame sequence, the reference face image features, and the audio feature vectors into a super-resolution model, wherein the super-resolution model is used to perform feature extraction, spatiotemporal attention modeling, bidirectional optical flow alignment, and dual feature fusion branching steps to generate a high-resolution face video frame sequence; The output unit is used to output a high-resolution conversational head animation video that conforms to the target scene based on the high-resolution face video frame sequence.
9. A server, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the processor executes the step instructions of the method as described in any one of claims 1-7 when it invokes the computer program in the memory.
10. A computer-readable storage medium, characterized in that, It stores a computer program / instruction thereon, which, when executed by a processor, implements the steps of the method as described in any one of claims 1-7.