Portrait image reconstruction method, portrait video generation method and electronic equipment

By acquiring multiple reference images and speaker information for deformation alignment and pixel-level fusion, the problems of motion inconsistency and texture distortion in portrait video generation are solved, and stable visual effects and high fidelity are achieved in different dynamic scenes.

CN120472057APending Publication Date: 2025-08-12AISPEECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510533659.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing portrait video generation methods are prone to inconsistency in motion, image texture distortion and artifacts when facial expression changes and head posture changes, which affect the naturalness and sense of reality of the generated results.

Method used

By acquiring the portrait source image and multiple reference images, the speaker's perspective, head posture and facial expression information is used for deformation alignment, combined with image fusion technology to generate portrait target images, and using implicit key points and gated networks for pixel-level fusion, optimizing texture coverage and visual consistency.

Benefits of technology

The portrait images and videos generated in different dynamic scenes maintain a stable visual effect, improving the naturalness and realistic sense of virtual characters, and solving the problems of image texture distortion and inconsistent facial movement during viewing angle switching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472057A_ABST
    Figure CN120472057A_ABST
Patent Text Reader

Abstract

The invention discloses a portrait image reconstruction method, a portrait video generation method and electronic equipment, and relates to the technical field of virtual digital faces, the method comprises the following steps: obtaining a portrait source image, a portrait target image and at least one portrait reference image, each portrait reference image having corresponding speaker reference information; extracting a first target portrait motion feature corresponding to the portrait target image, and performing deformation alignment on the portrait source image and each portrait reference image according to the first target portrait motion feature to obtain a corresponding source image deformation texture feature and each reference image deformation texture feature; and reconstructing a portrait target image according to the source image deformation texture features and the reference image deformation texture features. Therefore, the reference information of different visual angles, head postures and facial expressions is integrated, the alignment texture coverage of the source image is optimized, and the naturalness and vividness of virtual character generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of virtual digital faces, and in particular to a portrait image reconstruction method, a portrait video generation method, and an electronic device. Background Art

[0002] With the continuous advancement of computer graphics, deep learning and facial recognition technology, virtual digital humans (Avatar) and face synthesis technology have been widely used in many fields, such as virtual reality, augmented reality, digital entertainment, personalized advertising and virtual anchors.

[0003] Portrait video generation is an important computer vision application. Producing high-quality virtual human portrait videos, especially those that maintain realistic visual effects despite varying head poses, facial expressions, and perspectives, has become a core research topic. However, existing portrait video generation methods still face several technical bottlenecks, particularly in motion consistency, texture coverage, generation realism, and perspective adaptability.

[0004] Specifically, traditional portrait video generation methods often rely on a single reference frame (such as the first frame or a specific frame) as the basis for generation. As facial expressions and head postures change, the generated videos often exhibit motion inconsistencies. This is particularly evident when rapid motion or large changes in expression occur, causing the generated portrait video or image frames to flicker or exhibit artifacts. This affects the viewing experience, making the generated results appear unnatural, and also compromising the realism of the virtual characters in dynamic scenes.

[0005] To address the above issues, the industry has not yet proposed a better technical solution. Summary of the Invention

[0006] The present application provides a portrait image reconstruction method, a portrait video generation method and an electronic device, which are used to at least solve the problems of image texture distortion, inconsistent facial movement or artifacts when the perspective of the portrait images or portrait videos generated in the current related technology is switched.

[0007] In a first aspect, an embodiment of the present application provides a portrait image reconstruction method, comprising: acquiring a portrait source image, a portrait target image, and at least one portrait reference image; each of the portrait reference images has corresponding speaker reference information, and the speaker reference information includes at least one of the following: speaker perspective, head posture, and facial expression; extracting a first target portrait motion feature corresponding to the portrait target image, and deforming and aligning the portrait source image and each of the portrait reference images according to the first target portrait motion feature to obtain corresponding source image deformation texture features and each reference image deformation texture feature; reconstructing the portrait target image based on the source image deformation texture features and each of the reference image deformation texture features.

[0008] In a second aspect, an embodiment of the present application provides a portrait video generation method, comprising: obtaining a target portrait motion feature sequence and a first video frame sequence corresponding to a first portrait video; determining a portrait source frame and at least one portrait reference frame from the first video frame sequence; each of the portrait reference frames has corresponding speaker reference information, and the speaker reference information includes at least one of the following: speaker perspective, head posture and facial expression; for each second target portrait motion feature in the target portrait motion feature sequence, deforming and aligning the portrait source frame and each portrait reference frame according to the second target portrait motion feature to obtain corresponding source frame deformation texture features and each reference frame deformation texture features, and generating corresponding portrait video frames based on the source frame deformation texture features and each reference frame deformation texture features; sequentially combining the generated portrait video frames according to the target portrait motion feature sequence to generate a corresponding second portrait video.

[0009] In a third aspect, an embodiment of the present application provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the portrait image reconstruction method or the portrait video generation method of any embodiment of the present application.

[0010] In a fourth aspect, an embodiment of the present application provides a storage medium on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the portrait image reconstruction method or the portrait video generation method of any embodiment of the present application are implemented.

[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the portrait image reconstruction method or the portrait video generation method of any embodiment of the present application.

[0012] The beneficial effects of the embodiments of the present application are:

[0013] Through the embodiments of the present application, at least one reference image is introduced on the basis of the portrait source image. During the image reconstruction process, the speaker information (such as the speaker's perspective, head posture and facial expression) of multiple reference images is utilized. Through the texture alignment assistance of multiple reference images for the target portrait motion characteristics, the reference information of different perspectives, head postures and facial expressions is integrated to optimize the aligned texture coverage of the source image, so that the reconstructed portrait image can still present a stable visual effect in different dynamic scenes, is not affected by changes in perspective, and improves the naturalness and realism of the generated virtual character. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0015] Figure 1 A flowchart showing an example of a portrait image reconstruction method according to an embodiment of the present application is shown;

[0016] Figure 2 A flowchart showing an example of a method for generating a portrait video according to an embodiment of the present application is shown;

[0017] Figure 3 A schematic diagram showing an effect of an example of the RAGTalker framework provided in an embodiment of the present application is shown;

[0018] Figure 4 A schematic diagram of the architecture of an example of a rendering module of RAGTalker according to an embodiment of the present application is shown;

[0019] Figure 5 A schematic diagram showing the comparative effects of an example of an implicit alignment strategy and an explicit alignment strategy for texture fusion is shown;

[0020] Figure 6 A schematic diagram illustrating an example of the architecture of the motion sequence generation part of RAGTalker is shown;

[0021] Figure 7 A schematic diagram showing the effects of an example of a key frame selection and motion-guided gating mechanism according to an embodiment of the present application is shown;

[0022] Figure 8 A schematic diagram showing the qualitative comparison effect of an example of a reference frame selection strategy in motion-driven speaker synthesis;

[0023] Figure 9 A schematic diagram showing an example of comparative effects of different video-driven speaking face generation methods on lip synchronization accuracy;

[0024] Figure 10 This is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION

[0025] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] It should be noted that speaking video generation aims to synthesize realistic speaking videos in which the speech content can be modified or regenerated while maintaining the speaker's identity, context, and visual realism. However, methods based on single-frame references often fail when encountering large pose changes or significant head motion, while methods based purely on diffusion models are computationally expensive and inconvenient for practical applications.

[0027] Specifically, talking head video generation (Talking Head Video Generation) is gaining increasing attention in applications such as multimedia production, dubbing, virtual conferencing, and cross-language content adaptation. Its core task is to synthesize or modify a target person's speaking video so that new or modified speech content (such as text or audio) can be seamlessly integrated while preserving the speaker's identity, background environment, and overall visual realism.

[0028] Early research focused on synchronizing sound and lip movements through mouth retargeting, but this approach was limited to local lip control. As demand increased, subsequent work expanded to full-head talking head synthesis, which can control and preserve more facial and head movements. However, most of these methods have two limitations: (i) they rely only on a single frame reference, and texture blurring or loss often occurs when the head undergoes significant posture changes or position shifts; (ii) while solutions based on large-scale diffusion models can leverage rich "world knowledge" (such as reasonable background generation and consistent multi-view appearance), their high computational cost limits real-time applications.

[0029] Note that the blurry or missing textures that appear in single-frame methods are often found in other frames of the original video (occluded or cropped due to different poses or angles). However, when the generated frame differs significantly from the reference frame in pose (for example, when a person turns to a side face that has not been seen before), single-frame methods have difficulty filling in these areas, thus introducing noticeable artifacts.

[0030] In view of this, Figure 1 A flowchart of an example of a portrait image reconstruction method according to an embodiment of the present application is shown.

[0031] Regarding the executor of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities. With the assistance of multiple reference images, it maintains good adaptability under different perspectives, so that the generated portrait video can still present a stable visual effect in different dynamic scenes, solving the challenge of distorted results or artifacts generated at multiple angles, and ensuring the visual consistency and authenticity of the virtual character under multiple perspectives.

[0032] In some examples, it can be integrated into an electronic device or terminal through software, hardware, or a combination of software and hardware, and the type of terminal or electronic device can be diverse, such as a mobile phone, tablet computer, or desktop computer, etc.

[0033] like Figure 1 As shown, in step S110 , a portrait source image, a portrait target image and at least one portrait reference image are acquired, and each portrait reference image has corresponding speaker reference information.

[0034] Here, the speaker identities corresponding to the portrait source image, the portrait target image, and the portrait reference image should be consistent. In addition, the speaker reference information includes at least one of the following: speaker perspective, head posture, and facial expression. The speaker perspective expresses the angle of the speaker's face relative to the camera, the head posture describes the dynamic changes such as rotation and tilt of the head, and facial expressions (such as smiling, frowning, opening the mouth, etc.), thereby expressing the portrait image of the speaker in a dynamic scene. It should be understood that the data source types of the portrait source image, the portrait target image, and the portrait reference image can be diverse, such as a portrait image library or a video frame, which is not limited here.

[0035] In step S120, a first target portrait motion feature corresponding to the portrait target image is extracted, and the portrait source image and each portrait reference image are deformed and aligned according to the first target portrait motion feature to obtain corresponding source image deformation texture features and each reference image deformation texture features.

[0036] In some embodiments, it is first necessary to analyze the facial motion of the target image, extract its spatially varying features (such as the rotation angle of the head, changes in facial expression, etc.), and use them as target portrait motion features. For example, a motion feature encoder based on computer vision and deep learning technology is used to extract and encode information such as head rotation, changes in facial expression, and displacement of key points in the target image.

[0037] Based on these extracted motion features, the portrait source image and reference images are then deformed and aligned using various unrestricted image deformation algorithms. For example, the facial structures of the source image and each reference image are adjusted to align with the motion features of the target image using 3D face modeling, optical flow methods, or deep learning models. Thus, the source image and each reference image are deformed and aligned using the same motion features of the portrait target image, ensuring that the motion state of all deformed images matches the motion features of the target image.

[0038] In step S130 , a portrait target image is reconstructed based on the deformed texture features of the source image and the deformed texture features of each reference image.

[0039] In some embodiments, a portrait target image is generated by image fusion technology based on the deformed texture features of the source image and the deformed texture features of the reference image. Specifically, these texture features are combined using an image fusion algorithm (such as a weighted average method or an image synthesis method based on deep learning) to generate a final portrait target image. Exemplarily, the fusion algorithm automatically adjusts the matching degree between the texture and the image content based on the speaker's perspective, head posture, and facial expression information of each reference image, so that the texture details of the image remain consistent during dynamic changes. By fusing multi-dimensional speaker reference information, the generated portrait target image can present a smooth and natural effect in multiple different postures and expressions.

[0040] In the embodiment of the present application, the portrait motion feature can be represented by display or latent space, which is not limited here.

[0041] It's important to note that early approaches to speaker motion modeling relied on explicit representations, such as two-dimensional facial landmarks or three-dimensional parametric models (e.g., 3DMM). Recent research has introduced motion latent space representations, which offer more flexible modeling of facial dynamics. However, these latent representations lack explicit depth constraints, making it difficult to accurately distinguish between tooth positions or internal oral structures. To address this, implicit keypoints can be used to perform explicit spatial decomposition, encoding depth information into the keypoints, thereby robustly handling complex pose variations.

[0042] In some examples of the embodiments of the present application, portrait motion features adopt a structured motion representation based on implicit key points, which encode multidimensional information of facial key points, and the multidimensional information includes at least one of the following: expression offset information, head posture information, translation information and scaling information.

[0043] More specifically, the implicit key point can be expressed as follows:

[0044] x=s·(x c R+δ)+t, Formula (1)

[0045] in, represents N implicit key points, x c To standardize facial key points, R represents rotation, δ is the deformation related to expression, t is the translation vector, and s is the additional scaling factor introduced by Live Portrait generation.

[0046] As a further optimized implementation, the motion features of the source image and the reference image can be extracted, and the difference information between the motion features can be calculated, thereby improving the accuracy, detail expression and naturalness of the generated portrait image.

[0047] Specifically, the source portrait motion features corresponding to the portrait source image are extracted, and the reference portrait motion features corresponding to each portrait reference image are extracted. Then, the motion feature difference information of the source portrait motion features and each reference portrait motion feature relative to the first target portrait motion feature is calculated. The motion feature difference information can reflect the deviation between the motion features of the source image and the reference image and the target image. Furthermore, based on the motion feature difference information, the source image deformed texture features and the each reference image deformed texture features are pixel-wise fused to obtain a fused deformed texture feature, and the portrait target image is reconstructed based on the fused deformed texture feature. Thus, the source image deformed texture features and the each reference image deformed texture features are pixel-wise fused under the guidance of the motion feature difference. During the image reconstruction process, each image can be appropriately adjusted according to its motion feature difference relative to the target image, so that the texture features of the source image and the reference image can transition naturally, eliminating the sense of fragmentation or artifacts caused by the motion difference between different images during texture fusion.

[0048] Regarding the details of pixel-level fusion, in some examples of the embodiments of the present application, motion feature difference information, source image deformation texture features and each reference image deformation texture features are input into a gating network, so that the gating network determines the gating scores corresponding to each reference image deformation texture feature based on the motion feature difference information.

[0049] Specifically, the Gating Network is a deep learning-based model that learns how to adaptively adjust the contribution of different reference images (via gating scores) based on input features (in this scenario, motion feature difference information). For example, the gating network processes the motion feature difference information, determines the weight of each reference image's deformed texture features in the final image synthesis, and generates corresponding gating scores. For example, the contribution of reference images with larger changes in motion features relative to the target portrait should be reduced, thereby dynamically adjusting the contributions of different reference images.

[0050] Furthermore, based on each gated score, the corresponding key frames are fused with the source image deformation texture feature in a pixel-level weighted manner to obtain the fused deformation texture feature.

[0051] Specifically, the deformed texture features of each reference image can be weighted according to its corresponding gating score. Reference images with higher gating scores have a greater impact on the final generated image; conversely, reference images with lower gating scores contribute less to the fusion result, achieving adaptive weighted fusion of different reference images. Furthermore, during pixel-level fusion, the texture features of each reference image are weighted averaged with those of the source image at the pixel level. The weighted average is calculated for each pixel using the gating score, and the features of the source and reference images are fused together to generate the fused deformed texture features.

[0052] Through the embodiments of this application, a gated network learns the differences in motion characteristics between the source and reference images and generates corresponding gated scores, dynamically adjusting the reference image's contribution to the final image based on different motion states. Furthermore, a pixel-level weighted fusion method is used to finely fuse the texture features of the source and reference images, resulting in a portrait image that exhibits greater consistency, detail accuracy, and natural transitions despite dynamic changes.

[0053] It should be noted that in the embodiment of the present application, at least one portrait reference image is introduced in the fusion of the self-supervised source-target pipeline, but it should be noted that the newly added reference texture may interfere with the learned feature representation, sometimes causing flickering or artifacts.

[0054] In view of this, in some examples of the embodiments of the present application, the loss function for the reconstruction of the portrait target image includes a feature matching loss term in addition to the image reconstruction loss term. Specifically, the image reconstruction loss term is defined based on the pixel-level difference between the fused deformed texture features and the portrait target image. That is, during the image reconstruction process, the pixel values of the target image and the fused deformed texture features are compared, and the reconstruction quality of the target image is optimized by minimizing this difference. In addition, the feature matching loss term is defined based on the pixel-level difference between the fused deformed texture features and the deformed texture features of the source image, which encourages the fused features to be as close to the source features as possible, ensuring that the fused image can maintain the original details and style of the source image, and reducing artifacts or flickering introduced by the reference image.

[0055] Figure 2 A flowchart of an example of a method for generating a portrait video according to an embodiment of the present application is shown.

[0056] like Figure 2 As shown, in step S210, a target portrait motion feature sequence and a first video frame sequence corresponding to a first portrait video are obtained.

[0057] Here, the target portrait motion feature sequence can be a plurality of second target portrait motion features comprising a sequential combination, so as to constrain the dynamic portrait changes of the video to be synthesized. The second target portrait motion feature can adopt various non-restrictive motion representation types. For more details, please refer to the above combined Figure 1 The description of the motion characteristics of the first target portrait is not repeated here.

[0058] Regarding the input mode of the target portrait motion feature sequence, it can be diverse, such as input audio, input text or input video, and automatically identified and extracted from the input audio, input text or input video by the motion representation generator.

[0059] Regarding the first video frame sequence, it can be obtained by extracting multiple video frames (eg, all frames) from the first portrait video and sequentially combining them. Here, the first portrait video serves as a source video and provides an important portrait reference for subsequent portrait video reconstruction.

[0060] In step S220 , a portrait source frame and at least one portrait reference frame are determined from the first video frame sequence, and each portrait reference frame has corresponding speaker reference information.

[0061] Here, the speaker reference information includes at least one of the following: speaker perspective, head posture and facial expression. More details can be found in the combined Figure 1 The description in will not be repeated here.

[0062] Specifically, the portrait source frame is a source frame selected from the first video frame sequence and serves as the basic image input for the entire video generation process, providing a portrait texture benchmark. The portrait reference frame can provide speaker reference information for different speaker perspectives, head postures, or facial expressions for the target video, helping to provide accurate texture optimization assistance during the motion alignment process of source frame reconstruction. The selection of the portrait source frame and the portrait reference frame can be diverse, such as randomly selecting them from the first portrait video.

[0063] In step S230, for each second target portrait motion feature in the target portrait motion feature sequence, the portrait source frame and each portrait reference frame are deformed and aligned according to the second target portrait motion feature to obtain corresponding source frame deformation texture features and each reference frame deformation texture features, and corresponding portrait video frames are generated based on the source frame deformation texture features and each reference frame deformation texture features.

[0064] Specifically, each second target portrait motion feature describes the head posture and expression changes of the target person at that moment. The deformation network is used to perform deformation alignment on the portrait source frame and the reference frame based on these target portrait motion features. For details on the texture deformation alignment of the portrait source frame and the reference frame based on the target portrait motion features, please refer to and combine Figure 1 The description of the texture deformation alignment of the portrait source image and the portrait reference image is not repeated here.

[0065] In step S240 , the generated portrait video frames are sequentially combined according to the target portrait motion feature sequence to generate a corresponding second portrait video.

[0066] Specifically, the generated portrait video frames are grouped together into a video sequence based on the order of the second target portrait motion features specified in the target portrait motion feature sequence. By sequentially combining each frame, the resulting portrait video fully displays the dynamic changes of the target portrait, ensuring video coherence and integrity.

[0067] Regarding the implementation details of step S210, in some examples of the embodiments of the present application, by obtaining user input audio and determining the target portrait motion feature sequence corresponding to the user input audio based on a motion generator, the motion generator adopts a diffusion Transformer model.

[0068] Specifically, the user's audio input (such as voice or speech content) is first processed by a speech encoder, which can use a deep neural network, such as a convolutional neural network (CNN) or a long short-term memory network (LSTM), to extract key information from the audio signal, which may include information about pitch, speaking speed, volume, etc. This information can describe the rhythm, prosody, and details related to lip shape and facial expressions in the audio content.

[0069] In order to align the audio signal with the facial motion features of the target portrait, the Motion Diffusion Transformer (Motion DiT) is used here. It is a Transformer architecture based on a diffusion model that can convert audio signals into continuous facial motion representations. In addition, the model not only generates corresponding facial motion based on the audio input, but also combines historical motion information (for example, facial motion features of the previous frame of video) to generate a smooth and continuous motion sequence.

[0070] For example, audio is converted into an audio feature vector (such as Mel-Frequency Cepstral Coefficients (MFCCs) or other audio feature representations) through a speech encoder. The audio feature vector can contain the spectral characteristics of the speech and can reflect the content, speaking rate, pitch, etc. The audio feature vector is then combined with initial motion cues derived from a single input image or a previous video frame (the portrait video frame generated at the previous moment). The initial motion cues provide a starting state for the facial motion of the target portrait and can be the facial expression, head posture, or lip movements of the previous frame. The audio signal is then combined with the initial motion cues through the Motion Diffusion Transformer model to generate a coherent target portrait motion feature sequence. This sequence contains portrait motion information aligned with the audio content (such as lip synchronization, facial expression changes, head posture adjustments, etc.). The diffusion process of the Motion Diffusion Transformer ensures the smoothness and naturalness of the motion sequence, ensuring that the facial movements are highly consistent with the rhythm, rhythm, and intonation of the audio.

[0071] Through the embodiments of the present application, the motion diffusion transformer ensures the coherence and smoothness of the generated target portrait motion feature sequence by combining audio information and historical motion information, effectively solving the problem of instability or jerky transition of the generated image caused by sudden changes in audio (such as emotional fluctuations and changes in tone).

[0072] Regarding the details of obtaining each portrait reference frame, in some examples of the embodiments of the present application, the facial landmark vectors corresponding to each first video frame in the first video frame sequence are determined; each facial landmark vector is clustered to obtain at least one corresponding cluster; and a representative frame is selected from each cluster as the corresponding portrait reference frame.

[0073] More specifically, in facial landmark detection, a facial key point detection algorithm (such as Dlib's 68-point model or a deep learning facial detection model) is used to process each video frame and extract facial key points, which may include the coordinates of the eyes, eyebrows, nose, mouth, chin, etc. These coordinate points form a facial landmark vector, which represents the spatial distribution of each key position of the face in the current video frame. It should be understood that this facial landmark vector can be a vector representation different from the portrait motion representation, and the portrait motion representation can adopt a structured motion representation based on implicit key points.

[0074] Furthermore, once facial landmark vectors are extracted from each frame of video, the next step is to cluster these facial landmark vectors. For example, various clustering algorithms (such as K-means, DBSCAN, or hierarchical clustering) can be used to cluster facial landmark vectors, grouping frames with similar facial postures or expressions into the same group to ensure that each cluster represents a similar type of facial dynamics. Through clustering, representative dynamic features (such as the frame closest to the cluster center) can be extracted from a large number of facial landmark vectors and used as portrait reference frames to effectively distinguish different facial dynamics in the video (such as stillness, smiling, anger, etc.), providing diverse and non-redundant portrait reference information under different dynamic changes.

[0075] In some examples of the embodiments of the present application, a method for generating a speaker video (RAGTalker) based on key frame retrieval enhancement and motion guided texture fusion is also provided.

[0076] RAGTalker is a framework for generating talking heads through keyframe retrieval enhancement. By strategically sampling and aligning diverse reference frames from the original video, it effectively supplements the texture details missing or under-represented in a single frame reference, especially the background. Specifically, this framework consists of three core modules: (1) Retrieval module: robustly identifies keyframes in different head poses and backgrounds through clustering and alignment techniques; (2) Enhancement module: adaptively fuses fine-grained appearance details from multiple keyframes using a lightweight motion-guided gating mechanism to improve visual consistency; (3) Generation module: efficiently synthesizes natural facial motion from audio input using a motion diffusion Transformer, utilizing a compact motion representation to reduce computational complexity. Extensive experiments demonstrate that this method outperforms existing methods in terms of visual consistency, temporal coherence, and generalization across different individuals, while significantly reducing computational burden.

[0077] I. Introduction

[0078] Figure 3A schematic diagram showing the effect of an example of the RAGTalker framework provided according to an embodiment of the present application is shown.

[0079] like Figure 3 As shown in the figure below, the RAGTalker framework draws on the concept of Retrieval-Augmented Generation (RAG), a technique commonly used in language models. This paradigm strategically retrieves multiple reference frames covering a variety of facial poses and background scenes. As shown in the comparison at the bottom of the figure, compared to using only a single reference frame, "using keyframes" effectively expands the texture coverage, significantly improving the visual consistency and realism of the generated talking video.

[0080] By comparing the results, we can see that the baseline method based on a single frame reference will cause background blur when the posture changes greatly; however, after introducing additional reference frames that capture different backgrounds (such as frame 2), it can provide more comprehensive texture clues, making the final synthesis result closer to the real appearance of the source video.

[0081] Furthermore, inspired by the Retrieval-Augmented Generation (RAG) technique used in large language models, this paper proposes applying a similar retrieval-augmented approach to speaker video generation. The key difference is that this paper specifically retrieves and reuses multiple frames from the original video to enhance the texture representation of specific regions.

[0082] However, simply collecting multi-frame references does not guarantee efficient texture reuse: each keyframe may come from a different perspective, with different head poses and background layouts. To efficiently reuse textures across frames, accurate alignment is crucial. Face alignment methods based on affine transformation can only achieve coarse alignment, while this paper requires fine, patch-level alignment at the motion-aware level. Otherwise, the system will have to rely on complex self-attention or temporal attention networks to fuse misaligned frames, which incurs huge computational overhead. In contrast, mapping each keyframe to the same coordinate space through warping allows the most appropriate local patches to be precisely aligned in space, enabling efficient patch-based fusion.

[0083] In summary, the multi-frame RAG paradigm raises two core issues: first, how can we ensure that the selected reference frames adequately cover the various head poses and movements in the source video? Second, how can we effectively transfer and align the diverse texture information from multiple frames to the target frame without excessively increasing the computational burden?

[0084] To address the above challenges, this paper proposes RAGTalker, which consists of three key modules:

[0085] 1. Keyframe Clustering and Alignment (KFCA): Retrieve compact and diverse source frames and transform them into a unified coordinate space, providing a robust multi-view texture reference for subsequent processing.

[0086] 2. Motion-guided Texture Fusion (MTF): Adaptively integrates the most appropriate local details across aligned frames to enhance visual consistency and minimize artifacts.

[0087] 3. Motion Diffusion Transformer (Motion DiT) Generator: Efficiently models facial dynamics in motion space to balance computational efficiency with high-fidelity generation.

[0088] The main contributions of this paper can be summarized as follows:

[0089] - Retrieval: Propose a strategy to select compact and diverse keyframes to cover a wide range of pose, expression, and background variations;

[0090] - Enhancement: Introducing a deformation-based fine alignment strategy and combining it with a motion-guided fusion module to accurately transfer local appearance details across multiple frames;

[0091] -Generation: Leveraging a diffusive Transformer operating in motion space to efficiently model motion and achieve coherent and natural video synthesis.

[0092] II. Related Work

[0093] A. Speaker lip movement redirection

[0094] Mouth re-targeting, also known as the lip synchronization task, aims to generate a video of mouth movements synchronized with a given audio track while preserving the original facial identity and expression. Such methods usually apply a mask to the mouth or related areas, and then resynthesize the masked area through a renderer. Wav2Lip pioneered the masking of the lower half of the face and introduced an attention-based synchronization mechanism. Subsequent research has improved many details on this basis, such as optimizing masking strategies, refining lip synchronization loss functions, and enhancing rendering processes. Although these methods can achieve impressive lip synchronization effects, there is still a noticeable inconsistency between the edited mouth area and the overall video picture.

[0095] A major difficulty lies in the visibility of mask boundaries, requiring complex fusion techniques to ensure seamless integration. Another challenge is alignment with non-verbal movements—for example, when saying "no," if the original video includes a nod, simple lip editing can result in unnatural or conflicting facial movements. Our method overcomes these limitations by extending the editing scope from the mouth to the entire face and head pose.

[0096] B. Speaker Texture Modeling

[0097] The feature representation of a speaking video usually consists of two parts: (1) texture information, including appearance, background, and fine-grained visual details, which is usually represented as a high-dimensional tensor of dimension C×W×H, where C is the number of channels, W and H are the spatial width and height respectively; (2) motion information, which is much more compact, usually only a few hundred dimensions or even lower. Due to the high-dimensional nature of texture information, direct modeling and manipulation face challenges such as computational complexity and GPU memory limitations. Early methods alleviated this problem by limiting the receptive field to a short time window of a few hundred milliseconds (corresponding to a few frames in the video), but such a short time window may not be sufficient to capture certain speech patterns.

[0098] To overcome this limitation, subsequent research proposed a decoupled architecture that separates texture and motion features. This approach first extracts motion parameters from a high-dimensional space, allowing the synthesis of several minutes of video using only a few GB of memory. However, these methods are limited in their image manipulation capabilities, with most only able to perform simple transformations and struggling to preserve the fine-grained details in the original frames.

[0099] With the advent of stable diffusion and video diffusion models, pixel-level head texture modeling and manipulation have become possible, enabling comprehensive facial appearance modeling. However, diffusion methods typically require models with billions of parameters and 20–50 iterative denoising steps, resulting in high computational costs and hindering large-scale applications.

[0100] In this paper's method, the texture-motion decoupling strategy is also adopted. However, unlike the diffusion renderer, this paper introduces a lightweight texture fusion module, which achieves high-quality synthesis without the need for a large diffusion model, combining efficiency and scalability.

[0101] III.RAGTalker Framework

[0102] Inspired by the theory of Retrieval-Augmented Generation (RAG), which is widely used in language models to enhance generation performance by targeted retrieval of external information, this paper proposes a retrieval-augmented method specifically tailored for talking head video generation. The fundamental idea of RAG is to retrieve and integrate relevant contextual information to improve the richness and accuracy of the generated output. In the context of talking head video synthesis, this paper enhances visual realism by strategically retrieving multiple frames from the original video, providing strong texture coverage and ensuring consistent appearance under various head poses and motions.

[0103] Specifically, the RAGTalker framework consists of two main stages. The first stage is the rendering stage, which uses a self-supervised learning method to transfer the texture information of the source frame to the target frame. This process achieves two main goals: first, it trains a motion extraction module that can capture motion information; second, it uses the extracted motion and texture information to render realistic talking videos. In this stage, two additional modules are introduced to enhance rendering quality and efficiency: Keyframe Clustering and Alignment (KFCA) and Motion-guided 2D Texture Fusion (MTF), which effectively achieve the retrieval and enhancement goals. The second stage is the motion generation stage, which synthesizes motion representations from the audio signal. To address practical limitations on inference time, the audio input is segmented into fixed-length blocks and processed sequentially to maintain smooth transitions between blocks. This stage mainly handles the generation part.

[0104] Figure 4 A schematic diagram of the architecture of an example of a rendering module of RAGTalker according to an embodiment of the present application is shown.

[0105] like Figure 4 As shown in Figure 2, RAGTalker rendering consists of two core processes: retrieval and augmentation, which are implemented by the key frame clustering and alignment (KFCA) and motion-guided texture fusion (MTF) modules respectively. s , target frame I t And the set of key frames that maximize the diversity of poses and expressions selected by clustering Texture and motion features are extracted and aligned to the target motion. The MTF module adaptively fuses these aligned textures based on the motion description. The resulting fused texture significantly improves synthesis quality. This design reduces reliance on large-scale generative models, ensuring identity consistency and greater visual realism.

[0106] A. Basic Concept: Self-Supervised Deformation

[0107] To clarify the basic concepts, this paper first introduces the core self-supervised deformation method of the framework. Consider two frames randomly extracted from the video: the source reference frame I s , providing reference texture and appearance information; target frame I t , providing the required motion information (posture, expression, scale and translation). Texture features F are extracted from each frame, while motion features M encode the corresponding posture and expression parameters. In addition, the learned deformation module W transforms the texture features F of the source frame into s The motion feature M of the target frame t Align and generate aligned texture features T s→t :

[0108] T s→t =W(F s ,M s ,M t ), Formula (2)

[0109] Crucially, this explicit deformation alignment is learned through a self-supervised training strategy. The network optimizes itself by randomly selecting pairs of video frames and reconstructing the texture of a single frame under constraints, without the need for explicit annotations. This significantly improves generalization and computational efficiency. Building on this self-supervised foundation, this article will detail the components of the RAGTalker framework, beginning with the keyframe retrieval and alignment process performed by our Keyframe Clustering and Alignment (KFCA) module.

[0110] B. Keyframe Clustering and Alignment (KFCA)

[0111] Since the single source frame I s It may not be possible to capture all variations in the speaker's appearance (such as diverse postures and expressions), so this paper introduces a keyframe clustering and alignment module. Its purpose is to select a minimal but comprehensive set of keyframes from the source video. And align the texture features of these frames to the target motion. Similar to the key frames in video compression, which serve as anchor frames for reconstructing intermediate content, this paper defines key frames for facial texture fusion to extract key information from the source video. It is worth noting that the key frames in this paper are not only intended to supplement the texture and perspective of the newly generated frames, but also to minimize the number of reference frames while maximizing semantic coverage, which means that a small number of frames should fully cover the possible perspectives and facial expressions of the speaker. Specifically, this module includes three steps, such as Figure 4 As shown in the lower left part of , the specific description is as follows:

[0112] 1) Keyframe Clustering: In order to identify a compact and diverse set of keyframes, this paper uses a clustering method based on facial landmarks. For each frame in the source video, facial landmarks (e.g., 68 two-dimensional coordinates) are first extracted* (*: these facial landmarks are only used for keyframe selection and are not used in subsequent motion generation or texture rendering. In addition, landmark extraction and clustering are only performed once for each video). These landmarks encode rough but critical information about facial geometry, including mouth shape, eye openness, and head posture (yaw, pitch, and roll).

[0113] We then cluster these frames in landmark space by defining a distance metric that measures the similarity between facial landmarks. Frames with small landmark distances are clustered together, while frames with larger distances are placed into separate clusters. We select N cluster centers, each serving as a keyframe that encapsulates a unique pose or expression. This approach allows us to obtain a balanced subset that adequately covers the visual variability of speakers while minimizing redundancy.

[0114] 2) Motion and texture extraction:

[0115] In identifying key frames Afterwards, each frame is passed through two dedicated encoders:

[0116] -Texture Encoder E tex : Extract fine-grained appearance features (including background information) and generate unaligned texture features F k .

[0117] -Motion Encoder E mot : Encode facial dynamics (e.g., expression shift, head rotation, and scaling) into a compact motion latent representation M k .

[0118] These two features (F k ,M k ) provides an absolute description of the appearance and motion for the kth keyframe.

[0119] 3) Deformation alignment:

[0120] F k Capturing absolute, unaligned appearance information, directly fusing these features will require a large number of parameters. Therefore, this paper uses a pre-trained network (assuming a preliminary alignment network is pre-trained) to convert these features into aligned representations. Specifically, this step is achieved by using the motion parameters M of the target frame t , align all keyframes to the target frame I t . The deformation alignment module is based on (F k ,M k ,Mt ) is input, output:

[0121] T k→t =W(F k ,M k ,M t ), Formula (3)

[0122] Where T k→t is the aligned texture of the kth keyframe. Repeat this process for all keyframes to obtain N aligned textures that are consistent with the target pose and expression.

[0123] Through this process, multiple spatially aligned keyframe features are obtained, which facilitates efficient fusion in subsequent modules.

[0124] C. Motion-guided texture fusion (MTF)

[0125] Given N aligned textures T k→t (plus source aligned T s→t ), this paper proposes a motion-guided texture fusion (MTF) module to adaptively select relevant features from each reference frame. Different from simple averaging or computationally expensive spatiotemporal attention methods, we use a gating-based soft fusion mechanism driven by motion differences, such as Figure 4 shown in the central part of the .

[0126] First, a motion difference descriptor is calculated to quantify the motion of each key frame compared to the source frame I s , target frame I t and reference frame I k The relationship between:

[0127] △ k =MLP((M t -M s ),(M t -M k ), Formula (4)

[0128] This motion difference descriptor captures the relative changes in expression or head pose, indicating the relevance of a given keyframe texture to the target frame.

[0129] Next, align each reference texture T k→t and source aligned texture T s→t Input to the gating network:

[0130] G k =GatingNet(concat(T k→t ,T s→t ),△k ), Formula (5)

[0131] Among them G k ∈[0,1] H′×W′ is a spatial map indicating how much texture is borrowed from the kth keyframe at each pixel. These maps are then normalized by performing a softmax over all keyframes to obtain the fused texture It is a weighted combination of keyframe textures:

[0132]

[0133] In the final fusion process, With T s→t To combine:

[0134]

[0135] in

[0136]

[0137] is the mean gated response. In fact, the pixel areas with higher gated values are more likely to use the fused reference Other areas retain the source texture.

[0138] Finally, when keyframe fusion is introduced into a previously trained self-supervised source-target pipeline, the newly added reference texture may interfere with the learned feature representation, sometimes causing flickering or artifacts. To alleviate this problem, a feature matching loss is introduced:

[0139] L fm =||T out -T s→t ||1, Formula (9)

[0140] This loss encourages the fused features to be as close as possible to the source features. This regularization stabilizes the training process and further improves temporal consistency.

[0141] Figure 5 A schematic diagram showing the comparative effects of an example of an implicit alignment strategy and an explicit alignment strategy for texture fusion is shown.

[0142] like Figure 5 As shown in Figure 2, given misaligned textures as input, implicit attention-based methods rely on computationally expensive spatiotemporal attention modules to implicitly align and fuse the textures. In contrast, our approach explicitly aligns the textures through a lightweight deformation operation, enabling efficient motion-guided fusion and producing high-quality aligned textures as output.

[0143] Specifically, implicit methods rely on computationally expensive spatiotemporal attention modules for texture alignment and fusion, which often imposes a significant computational burden. In contrast, our approach explicitly aligns textures through an efficient deformation mechanism and performs motion-guided fusion. This allows for efficient texture alignment, produces high-quality texture output, and is computationally more economical than implicit methods.

[0144] D. Motion Diffusion Transformer (Motion DiT)

[0145] Inspired by the two-stage approach, a compact motion generation method is also adopted in the audio-driven stage, aiming to decouple texture and motion and generate only motion information. To facilitate the generation of extended video sequences and maintain continuity, this paper incorporates historical frame information from the previous segment. The detailed process is as follows:

[0146] a) Overview:

[0147] set up

[0148] M={M1,M2,…,M T}, Formula (10)

[0149] represents the motion potential representation of T frames, where each Including expression parameters, head posture parameters, a scaling factor and translation parameters. Here, these potential representations are divided into N consecutive blocks:

[0150] M (1) ,M (2) ,…,M (N) , Formula (11)

[0151] Each M (n) covers a continuous range of frames. The goal is to generate each block in a temporally consistent manner, conditioned on its local audio features c (n) and the final latent representation of the previous paragraph To further maintain the continuity within and between blocks, relative position coding is used after the audio signal is appended.

[0152] b) Diffusion model formula:

[0153] A Diffusion Transformer (DiT) is constructed based on the DDPM architecture. During training, the block M is corrupted by adding Gaussian noise in the diffusion step t. (n) :

[0154]

[0155] where α t is the noise variance schedule, ∈~N(0,I). Network ∈ θIt is then trained to predict this noise:

[0156]

[0157] By conditioning on the local audio feature c (n) and the final latent representation of the previous paragraph The model learns to generate smooth transitions between segments.

[0158] c) Reasoning:

[0159] Figure 6 This figure shows an example architecture diagram of the motion sequence generation part of RAGTalker. During inference, motion sequences are generated block by block. Specifically, the initial frame of the first block is provided by the input image as a visual cue, and the initial frame of subsequent blocks is set to the final frame of the previous block to maintain temporal continuity. Each block is then generated through a diffusion-based sampling process, starting with random noise and gradually denoising it. Based on local audio features and the provided initial frame, consistent and natural motion is synthesized across the entire sequence.

[0160] like Figure 6 As shown, the audio is first processed through a speech encoder and then combined with initial motion cues derived from a single input image or the final frame of the previous segment. The Motion Diffusion Transformer (Motion DiT) generates coherent motion sequences that are rendered as real video frames by referencing keyframes in the original video. This approach ensures smooth and natural lip sync and facial movements that align with the audio input.

[0161] IV. Experiment

[0162] A. Experimental Setup

[0163] 1) Dataset:

[0164] Rendering dataset

[0165] This paper uses publicly available datasets, including video datasets such as VoxCeleb, VFHQ, CC, MultiTalk, and CelebV-HQ. These datasets primarily contain individuals of European descent. To achieve better balance, they were supplemented with an additional 150 hours of internally collected data, focusing on Asian faces. Most of the datasets were sourced from online platforms. To ensure consistency, the original videos were re-downloaded and processed using a standardized pipeline, including scene detection, video quality assessment, and facial occlusion detection. The resulting dataset contains approximately 255,000 video clips totaling 1,067 hours, with an average clip length of 15 seconds. Furthermore, static image datasets including CelebA-HQ, CelebAMask, AAHQ, and FFHQ were incorporated, totaling 151,000 still images. In total, if all frames in this stage were converted to images, the total number of images used in the rendering process would be approximately 96 million. All images and videos were resized to 512×512 for rendering training.

[0166] Audio to Latent Dataset

[0167] A small, manually annotated subset was carefully selected, taking into account the temporal continuity of the videos and the presence of audio tracks. This refined dataset excluded low-quality samples, such as those with facial occlusions, heavily edited passages, or inconsistent lip sync. Ultimately, a high-quality selection totaling 100 hours of cross-lingual video was compiled.

[0168] Evaluation dataset

[0169] To evaluate performance on Indo-European language and video reconstruction tasks, we used HDTF as the test set, following Dinet. To evaluate the performance of the algorithm on non-speaker video for non-Indo-European languages, we also tested it on the Multilingual Non-Indo-European Talking Head Evaluation Corpus (MNTE). MNTE covers a variety of commonly used non-Indo-European languages, including Mandarin, Korean, Japanese, Arabic, Swahili, and Turkish.

[0170] 2) Evaluation Setup

[0171] We comprehensively evaluate the capabilities of our model through facial reproduction and audio-driven speaker video generation tasks, closely simulating real-world application scenarios.

[0172] To evaluate rendering performance, a facial reproduction task is performed. Starting from the first frame of the video and the subsequent motion-driven signal, our model reconstructs the entire video sequence. The synthesized results are compared with real videos, allowing us to quantify the accuracy and fidelity of the rendering. To maintain consistency in subject identity, the reproduction is restricted to scenes involving the same identity. Our algorithm is benchmarked against existing technologies, including FOMM, DPE, MTIA, Vid2Vid, LIA, FADM, AniTalker, LivePortrait, EMOPortrait, and VQTalker. These methods typically adopt a self-supervised strategy to learn motion representations by transferring motion between images.

[0173] To further validate the synthesis capabilities, we performed an audio-driven speaking video generation task. Starting from initial static frames, we synthesized video sequences guided by audio input. Our model was compared with state-of-the-art methods, including SadTalker, EAT, PDFGC, AniTalker, EDTalker, EchoMimic, and VQTalker. Unlike methods that rely solely on lip sync, these methods generate complete facial animations, including speech-driven movements and natural head movements.

[0174] 3) Evaluation Metrics

[0175] This paper uses a series of evaluation criteria to measure the quality and similarity of generated images and videos in different scenarios. These metrics include:

[0176] - Image similarity metrics: Structural Similarity Index (SSIM) and Learning Perceptual Image Patch Similarity (LPIPS), which quantify the structural and perceptual similarity between generated images and real images;

[0177] -Facial fidelity metrics: Cosine Similarity (CSIM), which evaluates facial similarity; Landmark Distance (LMD), which measures the accuracy of key facial locations;

[0178] -Image quality metric: Cumulative Probability of Blur Detection (CPBD). To ensure fair comparison, the resolution is normalized to 256×256.

[0179] 4) Model configuration:

[0180] Clustering

[0181] For clustering-based keyframe selection, this paper uses K-means clustering, with L2 distance as the similarity metric, unless otherwise specified. Given a set of extracted frame features, K-means is applied to partition them into N clusters, each representing a unique facial expression or head pose. The cluster centroids serve as anchor points to ensure coverage of a wide range of poses and expressions.

[0182] To select representative frames, we identify the closest frame to each cluster center. Specifically, for each cluster, we compute the Euclidean distance (L2 norm) between each frame and the corresponding cluster center. The frame with the smallest distance is selected as the keyframe. If the number of available frames is less than the number of clusters, all frames are selected to ensure sufficient coverage. By default, three keyframes are selected, which means that three additional reference frames are added for texture fusion. This selection strategy ensures that the selected frames capture a wide range of expressions and viewpoints while minimizing redundancy, thereby improving the quality and stability of the synthesis results.

[0183] To train the facial renderer, this paper adopts an image encoder-decoder architecture based on LivePortrait. The texture encoder uses 3D ResBlocks to extract 3D appearance feature volumes, and the motion encoder is based on ConvNeXtV2-Tiny. The deformation module adopts the hourglass network structure and is further optimized to learn the deformation sampling process. The final rendering stage uses a decoder based on SPADE and integrates a texture fusion module to improve synthesis quality. This paper replicates the original training pipeline and combines it with the R3GAN training strategy to stabilize the GAN training process. The total number of parameters in the appearance feature extractor, motion feature extractor, deformation module, SPADE generator, and MTF module is 836K, 28.1M, 45.5M, 55.4M, and 1.4M, respectively, for a total of approximately 131.2M parameters. Compared to the original LivePortrait, this represents an increase of 1.4M parameters.

[0184] On this basis, this paper adopts an audio-to-motion module based on the Diffusion Transformer (DiT). To enhance feature integration, a 4-layer Transformer-based network is used, equipped with relative position encoding, so that speech-driven motion synthesis can more effectively model long-distance dependencies. In addition, this paper integrates a pre-trained WhisperBase speech encoder as an audio feature extractor. To ensure synchronization with the video frame rate, a downsampling layer reduces the audio sampling rate from 50Hz to 25Hz, so that the motion is generated accurately and time-aligned. The total number of parameters at this stage is 49M. The default block size is 3 seconds (75 frames).

[0185] 5) Training process and hardware:

[0186] Training Process

[0187] We adopt a staged training strategy, where the MTF module is initially untrained. First, a basic deformable module is trained to learn texture alignment. This process begins with a self-supervised learning phase, training a source-to-target image transformation model, following previous work. Once the deformable module is trained, reference frame texture information is introduced by randomly sampling 1–5 frames from the video and training the MTF module to learn texture fusion. To ensure adaptability, no clustering operation is applied during training, allowing the network to autonomously decide how to utilize the available reference frames in different scenarios.

[0188] In the second stage, the motion encoder from the previous stage extracts motion features and trains a motion generator. Random regions within the video frame are occluded, and the loss is computed only in these occluded regions, forcing the network to learn motion completion. This approach treats synthesis as an integrated task, ensuring consistency and fluidity across the entire video sequence.

[0189] Hardware and Convergence Time

[0190] All experiments were conducted on an NVIDIA L20 server equipped with eight GPUs, each with 48GB of memory. Under default experimental conditions, the training phase of the rendering module took approximately two weeks, using bf16 mixed precision training with a batch size of 4. The MSE loss for image reconstruction converged at approximately 0.003, followed by the fine-tuning phase of the texture fusion module, which took an additional day. For the DiT model, training used a batch size of 16 and reached full convergence within three days. The noise loss stabilized at approximately 0.002. All evaluations were performed on V100 GPUs.

[0191] Figure 7 A schematic diagram showing the effects of an example of a key frame selection and motion-guided gating mechanism according to an embodiment of the present application is shown.

[0192] like Figure 7 As shown in Figure 3, the adaptive contribution of each keyframe in different regions of the synthesized image is illustrated by visualizing the gating map. When synthesizing a specific view, the gating heatmap corresponding to that view is red, indicating that the relevant region of the frame is actively used. Figure 7 An example of perspective compensation is shown in Figure 7 See the provided demo page for more detailed texture supplement information.

[0193] Figure 7 The upper row shows the selected keyframes Designed to maximize expression and pose variety. Figure 7 The left column shows the source frame I s and target frames {t1, t2, t3}. Figure 7 The heatmap on the right side represents the gating scores, where the red areas indicate that the corresponding keyframes have stronger contributions during the fusion process. This adaptive weighting ensures high-fidelity synthesis and temporal consistency.

[0194] C. Face Reproduction Experiment

[0195] To evaluate the rendering capabilities of the model provided by the embodiments of this application, a facial reconstruction experiment was conducted. Specifically, only the first frame of the video was driven using a motion sequence, and an attempt was made to reconstruct the entire sequence. During the evaluation process, quantitative metrics were calculated between the generated frames and the original video to assess the model's ability to accurately redraw each frame. This experiment provided insights into how the model preserves fine-grained details and reconstructs the original content with high fidelity. The experimental results are shown in Table I.

[0196] Table I. Face reconstruction on HDTF. Resolution (Reso.) indicates the image resolution (width and height), and the bold and underlined values represent the best and second-best results, respectively. Symbols Means: Only the driver module (stage 1) is enabled; the splicing and redirection modules are disabled.

[0197]

[0198] The results show that our method outperforms existing methods in terms of pixel-level reconstruction (SSIM), perceptual similarity (LPIPS), and identity preservation. This improvement can be attributed to the effective integration of reference frames, which enhances texture fusion and ensures greater consistency in complex areas such as background texture and eye details. However, it is observed that the CPBD (clarity) score is lower compared to some baseline methods. It is speculated that this may be because our rendering module does not combine super-resolution technology or multi-scale fusion, resulting in the loss of high-frequency details. Therefore, our method fails to achieve the clarity scores of methods such as EMOPortrait (which combines a super-resolution module) or VQTalker (which adopts a multi-scale texture fusion method and combines hierarchical skip connections).

[0199] Figure 8 A diagram showing a qualitative comparison of an example of reference frame selection strategies for motion-driven speaker synthesis. Compared to direct deformation or random selection, clustering-based keyframe selection improves texture diversity, reduces artifacts, and ensures more stable motion. The inset shows the selected reference frames.

[0200] Figure 8 The qualitative visualizations in further illustrate the advantages of our approach.

[0201] (a) Texture-free fusion: This model directly deforms the portrait frame to match the target motion, resulting in noticeable artifacts and unnatural texture distortions in areas with significant motion, such as the forehead and cheeks.

[0202] (b) Random Frame: A reference frame is randomly selected from the video (see inset). Although this method reduces some artifacts, the inconsistent texture fusion causes temporal flickering and misalignment, especially in the forehead and hair regions.

[0203] (c) Cluster-based frames: Keyframes are selected by clustering to maximize texture diversity and motion coverage. This results in smoother texture transitions, better detail preservation, and fewer fusion artifacts, especially highlighted in the dashed area. The selected reference frames (bottom right inset) provide a more representative set of textures, resulting in more stable and realistic synthesis. This experiment highlights the importance of selecting diverse and relevant reference frames for texture fusion, demonstrating that cluster-based selection significantly improves synthesis quality. The first three rows demonstrate the superiority of background reconstruction, while the last two rows highlight the improvement in eye consistency.

[0204] D. Audio driver experiment

[0205] To evaluate the audio generation capabilities of our DiT module, we compared it with state-of-the-art methods for one-shot speaker video generation. Specifically, we reconstructed the entire video sequence using the initial video frame, audio signals, and pose information. Experiments were conducted on the HDTF dataset (Indo-European languages) and the MNTE dataset (non-Indo-European languages). Quantitative results are summarized in Table II.

[0206] The results demonstrate strong generalization capabilities. Even without training or fine-tuning on the HDTF or MNTE datasets, our method performs well on multiple evaluation metrics, including pixel-level similarity (PSNR, SSIM), facial similarity (CSIM), and landmark accuracy (LMD). However, consistent with the findings on the face reproduction task, our method performs slightly worse on clarity (CPBD).

[0207] In multilingual scenarios, our approach demonstrates robust generalization across multiple language families. Notably, it excels at preserving facial identity, particularly on the MNTE dataset, which typically contains more significant head pose variations than HDTF. This advantage is attributed to our multi-reference frame fusion strategy, which provides enhanced texture coverage and thus improved identity consistency.

[0208] Figure 9The figure shows a comparison of different video-driven speaking face generation methods in terms of lip synchronization accuracy. Given a portrait image, driving posture and audio input, the lip movements generated by this method are more accurate and natural, corresponding to the pronunciation, and have better results than EDTalk, AniTalker and EchoMimic. Here, the method provided in this paper is Figure 9 The lip sync visualization is shown in Figure 3. The examples show that our model effectively captures a wider range of mouth shapes, especially in syllables with significant pronunciation changes (such as the rounded lip shape in the syllable "ou").

[0209] E. Reasoning Speed Experiment

[0210] In this section, runtime speed and memory usage are evaluated. EchoMimic, a representative open-source inference framework based on Stable Diffusion, contains 2,202M parameters (including VAE, reference network, and denoising Unet). In comparison, our method - which includes a motion DiT, LivePortrait, and an MTF module (1.4M) - has a total parameter count of 131M. For fairness, no optimization acceleration was applied to any method. As shown in Table III, our framework is about 17 times smaller than EchoMimic in terms of the number of parameters, runs 7 times faster, and uses about half the GPU memory of the latter (although EchoMimic processes 14 frames at a time, while our method processes 75 frames). These results demonstrate the clear advantages of our method in speed and memory efficiency.

[0211] Table II Quantitative comparison with speech-driven baseline methods on the HDTF (Indo-European languages) and MNTE (non-Indo-European languages) datasets. The bold and underlined values indicate the best and second-best results, respectively.

[0212]

[0213] Table III Inference speed test on V100. #Parameters: model parameters.

[0214] FPS: frames per second. GPU MEM: maximum GPU memory usage.

[0215]

[0216] Table IV. User studies on lip sync (LS), visual appeal (VA), pose following (PF), and naturalness (N).

[0217]

[0218] F. User Research

[0219] We conducted a user study with 20 participants, evaluating videos from the MNTE dataset using the following metrics: lip sync (LS), visual appeal (VA), pose following (PF), and naturalness (N). The results show that RAGTalker achieves significant improvements in LS and PF, reflecting its effectiveness in capturing precise lip movements and accurately controlling head pose. This improvement is primarily due to our decoupled implicit keypoint strategy, which simplifies learning complex facial movements.

[0220] In terms of VA and N, our RAGTalker is slightly inferior to EchoMimic, mainly because stable diffusion-based methods (such as EchoMimic) inherently generate more expressive visual content. However, our method prioritizes consistency and better preserves the original background and large pose details. Considering that EchoMimic's model size is approximately ten times that of ours (in terms of number of parameters), RAGTalker demonstrates superior efficiency, providing comparable visual quality and naturalness, while emphasizing reconstruction fidelity.

[0221] G. Ablation studies

[0222] 1) Impact of Reference Frame Selection: To evaluate the impact of reference frame selection on synthesis quality, we compare different numbers of reference frames as well as two selection strategies: random sampling (randomly selecting reference frames from the source video) and clustering-based selection (using a clustering algorithm to select keyframes to maximize pose and expression diversity).

[0223] Table V. Ablation study on reference frame selection strategies. We compare the effects of using different numbers of additional reference frames and evaluate two selection strategies: random sampling and clustering-based selection.

[0224]

[0225] Table V shows that increasing the number of reference frames improves synthesis performance, as evidenced by higher PSNR, SSIM, and CSIM scores, and lower LPIPS values. This suggests that more reference frames provide richer texture priors, leading to more accurate reconstructions. However, clustering-based selection consistently outperforms random selection with the same number of reference frames, demonstrating that selecting diverse and representative keyframes is more effective than relying on random selection. Beyond three reference frames, improvements become limited, indicating that excessive redundancy does not further improve synthesis quality.

[0226] These findings highlight the importance of selecting diverse and informative reference frames. Our clustering-based selection strategy enhances synthesis by ensuring better motion and texture coverage and reducing unnecessary redundancy, ultimately leading to more stable and visually consistent results.

[0227] 2) Impact of key components:

[0228] To evaluate the impact of different components in our framework, we conduct a series of ablation experiments, removing or modifying key modules and analyzing their impact on the synthesis quality.

[0229] No motor differences (w / o MD)

[0230] This variant removes the motion difference module responsible for capturing keypoint-based motion changes between frames. The model directly concatenates the source and reference frame features and relies solely on the gating network for fusion. Results show that without explicit motion difference modeling, the model struggles to accurately reconstruct motion and expression, demonstrating the importance of motion difference descriptors in improving synthesis quality.

[0231] Average feature fusion (Avg Fusion)

[0232] In this variant, we replace the motion-guided gated fusion with a simple averaging operation, averaging the aligned multi-frame features, eliminating the effects of motion difference modeling and the gating mechanism. Results show that this simple fusion strategy fails to properly account for the importance of different texture regions, resulting in poor motion and expression consistency.

[0233] No feature matching (w / o FM)

[0234] Feature matching ensures that textures are spatially aligned before fusion. After removing this module, textures become misaligned, increasing artifacts and reducing synthesis quality. These findings suggest that our model benefits from minimal spatial perturbations because it is trained via an alignment-based feature extraction process.

[0235] Hard Fusion

[0236] Instead of using soft gating, this variant applies a hard mask-based selection, where a binary mask determines which regions retain features from the source frame. Results show that explicit masking reduces synthesis quality, suggesting that flexible soft gating methods are more effective in achieving smooth and natural texture fusion.

[0237] Complete model (Ours)

[0238] Our complete framework integrates motion difference modeling, feature matching, and soft-gating-based fusion. As shown in Table VI, it achieves the highest PSNR, SSIM, and CSIM scores, demonstrating the contribution of each component to the overall synthesis quality.

[0239] 3) Impact of clustering methods:

[0240] In this study, we cluster video frames to select representative reference frames. In addition to the traditional K-Means algorithm, we also tested two alternative clustering methods: Gaussian mixture model (GMM) and hierarchical clustering (AC). Specifically, we replaced the K-Means clustering step in the pipeline with these methods, while keeping the other data preprocessing and generation modules unchanged.

[0241] Table VI shows an ablation study of the impact of key components in our framework. The effects of motion difference (MD), gating mechanism (GATE), feature matching (FM), and final fusion (FF) are evaluated using PSNR, SSIM, and CSIM metrics. Higher values indicate better reconstruction quality.

[0242]

[0243] As shown in Table VI, the final results of the three clustering algorithms show minimal differences, with all evaluation metrics being highly similar. This demonstrates that our overall approach is robust to the choice of clustering algorithm, ensuring stable frame selection and consistent synthesis quality. Even when applying different clustering strategies, the final output remains comparable, indicating that the specific clustering method has a negligible impact on overall performance.

[0244] V. Conclusion

[0245] This paper proposes RAGTalker, a retrieval-augmented framework that robustly generates visually consistent and temporally coherent talking heads in video by strategically sampling and aligning diverse reference frames. Extensive experiments demonstrate that our approach effectively addresses the texture completion problem, significantly outperforming existing diffusion-based methods in terms of visual consistency, strong generalization, and computational complexity. Our approach represents a significant step forward in efficient, realistic, and practical speaker synthesis, opening up possibilities for a variety of real-world multimedia applications.

[0246] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0247] In some embodiments, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores one or more programs including execution instructions, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the steps of any of the above-mentioned portrait image reconstruction methods or portrait video generation methods of the present application.

[0248] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer performs the steps of any one of the above-mentioned portrait image reconstruction methods or portrait video generation methods.

[0249] In some embodiments, an embodiment of the present application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the portrait image reconstruction method or the portrait video generation method.

[0250] Figure 10 is a schematic diagram of the hardware structure of an electronic device for executing a portrait image reconstruction method or a portrait video generation method provided by another embodiment of the present application, such as Figure 10 As shown, the device includes:

[0251] One or more processors 1010 and memory 1020, Figure 10 A processor 1010 is taken as an example.

[0252] The apparatus for performing the portrait image reconstruction method or the portrait video generation method may further include: an input device 1030 and an output device 1040 .

[0253] The processor 1010, the memory 1020, the input device 1030 and the output device 1040 may be connected via a bus or other means. Figure 10 The bus connection is taken as an example.

[0254] Memory 1020, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the portrait image reconstruction method or portrait video generation method in the embodiments of the present application. Processor 1010 executes the non-volatile software programs, instructions, and modules stored in memory 1020 to execute various server functional applications and data processing, thereby implementing the portrait image reconstruction method or portrait video generation method in the aforementioned method embodiments.

[0255] The memory 1020 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 1020 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 1020 may optionally include a memory remotely located relative to the processor 1010, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0256] The input device 1030 may receive input digital or character information and generate signals related to user settings and function control of the electronic device. The output device 1040 may include a display device such as a display screen.

[0257] The one or more modules are stored in the memory 1020 , and when executed by the one or more processors 1010 , perform the portrait image reconstruction method or the portrait video generation method in any of the above method embodiments.

[0258] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.

[0259] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to:

[0260] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and their primary purpose is to provide voice and data communications. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.

[0261] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers and have computing and processing capabilities, and generally also have mobile Internet access. These terminals include PDAs, MIDs, and UMPCs.

[0262] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players, handheld game consoles, e-books, smart toys, and portable car navigation devices.

[0263] (4) Other onboard electronic devices with data interaction functions, such as onboard computer devices installed in vehicles.

[0264] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0265] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.

[0266] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A portrait image reconstruction method, comprising: acquiring a portrait source image, a portrait target image, and at least one portrait reference image; Each of the portrait reference images has corresponding speaker reference information, and the speaker reference information includes at least one of the following: speaker perspective, head posture, and facial expression; extracting a first target portrait motion feature corresponding to the portrait target image, and deformably aligning the portrait source image and each of the portrait reference images according to the first target portrait motion feature to obtain corresponding source image deformation texture features and each of the reference image deformation texture features; The portrait target image is reconstructed according to the source image deformation texture feature and each of the reference image deformation texture features.

2. The method according to claim 1, further comprising: extracting source portrait motion features corresponding to the portrait source image, and extracting reference portrait motion features corresponding to each of the portrait reference images; Calculating motion feature difference information of the source portrait motion feature and each of the reference portrait motion features relative to the first target portrait motion feature; The step of reconstructing the portrait target image according to the deformation texture feature of the source image and the deformation texture features of each of the reference images includes: Based on the motion feature difference information, the source image deformation texture feature and each of the reference image deformation texture features are pixel-wise fused to obtain a fused deformation texture feature; The portrait target image is reconstructed according to the fused deformation texture features.

3. The method according to claim 2, wherein: The pixel-level fusion of the source image deformation texture feature and each of the reference image deformation texture features based on the motion feature difference information to obtain a fused deformation texture feature includes: Inputting the motion feature difference information, the source image deformation texture feature, and each of the reference image deformation texture features into a gating network, so that the gating network determines a gating score corresponding to each of the reference image deformation texture features according to the motion feature difference information; Based on the gated scores, pixel-level weighted fusion is performed on the corresponding key frames and the source image deformation texture feature to obtain a fused deformation texture feature.

4. The method according to claim 2, wherein: The loss function for reconstructing the portrait target image includes an image reconstruction loss term and a feature matching loss term; wherein the feature matching loss term is defined based on the pixel-level difference between the fused deformed texture feature and the deformed texture feature of the source image, and the image reconstruction loss term is defined based on the pixel-level difference between the fused deformed texture feature and the portrait target image.

5. The method according to any one of claims 1 to 4, wherein The portrait motion feature adopts a structured motion representation based on implicit key points, which encodes multi-dimensional information of facial key points. The multi-dimensional information includes at least one of the following: expression offset information, head posture information, translation information and scaling information.

6. A method for generating a portrait video, comprising: Obtaining a target portrait motion feature sequence and a first video frame sequence corresponding to the first portrait video; Determining a portrait source frame and at least one portrait reference frame from the first video frame sequence; each of the portrait reference frames has corresponding speaker reference information, the speaker reference information including at least one of the following: speaker perspective, head posture, and facial expression; For each second target portrait motion feature in the target portrait motion feature sequence, deformably align the portrait source frame and each portrait reference frame according to the second target portrait motion feature to obtain corresponding source frame deformed texture features and each reference frame deformed texture features, and generate corresponding portrait video frames based on the source frame deformed texture features and each reference frame deformed texture features; The generated portrait video frames are sequentially combined according to the target portrait motion feature sequence to generate a corresponding second portrait video.

7. The method according to claim 6, wherein: The step of obtaining a target portrait motion feature sequence includes: User input audio is obtained, and a target portrait motion feature sequence corresponding to the user input audio is determined based on a motion generator; the motion generator adopts a diffusion Transformer model.

8. The method according to claim 6, wherein: Acquisition of each portrait reference frame includes: Determining facial landmark vectors corresponding to respective first video frames in the first video frame sequence; performing clustering processing on each of the facial landmark vectors to obtain at least one corresponding cluster; A representative frame is selected from each of the clusters as a corresponding portrait reference frame.

9. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

10. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Single-image 3D portrait generation method and system based on mixed prior and noise resampling

    CN121305004A

  • Human portrait three-dimensional asset generation method and system

    CN121392159A