High-fidelity expression portrait video generation method based on personalized representation
By employing a personalized representation approach, utilizing the SMPL-X model and an identity-adaptive facial expression transfer module, the problems of stiff facial expressions and identity drift in existing technologies have been solved, achieving high-fidelity, controllable dynamic portrait video generation with cinematic rendering quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-28
Smart Images

Figure CN121937598A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and graphics technology, and in particular to a method for generating high-fidelity facial expression portrait videos based on personalized representations. Background Technology
[0002] With the rapid development of computer vision and computer graphics technologies, AI-based digital human image generation technology is increasingly widely used in film and television special effects production, virtual reality, augmented reality, gaming, and metaverse virtual avatar driving. In particular, generating dynamic human images with cinematic quality, vivid expressions, and consistent identity has become a core technological focus for both academia and industry. In these applications, precisely controlling facial expressions and head posture while perfectly preserving the target subject's identity is a key indicator of the quality of the generation technology. However, existing dynamic human image generation technologies still face significant challenges in pursuing high fidelity and precise control. To drive dynamic expressions from static human images, existing methods typically rely on intermediate motion representations as control signals, but these representations have significant shortcomings in expressive power, decoupling effect, and transfer adaptability.
[0003] Existing mainstream methods use 2D facial landmarks or 3D parametric models (such as 3DMM, SMPL, and FLAME) as explicit motion control signals. However, these representation methods have inherent limitations. 2D landmark signals are too sparse, lacking the spatial details needed to describe facial geometry, making it difficult to define micro-expressions or specific object identity features, and exhibiting extremely poor stability under large pose changes. On the other hand, 3D parametric models are essentially low-rank linear approximations of facial or human body geometry. Their predefined hybrid shape subspace is not only dimensionally limited but also unable to represent high-frequency nonlinear skin deformations such as wrinkles and dimples that change with facial expressions. This low-dimensional template limits the upper limit of the model's expression, resulting in videos that often have stiff expressions, lack details, or produce unnatural geometric deformations when attempting to express exaggerated expressions, failing to meet the requirements of high-fidelity cinematic production.
[0004] Another approach attempts to learn implicit latent motion features from the driving video using a motion extractor and inject them into the generative network. While this method preserves more motion information to some extent, its core flaw lies in the entanglement of facial expressions and identity. The learned latent features often fail to completely separate the driver's identity attributes, causing the driver's facial features to inevitably leak into the target image during generation, resulting in "identity drift" or "identity confusion." This incomplete decoupling severely undermines the consistency of the generated results, especially in cross-face reenactment tasks. The generated video often no longer resembles the original target person but becomes a hybrid of the driver and the target person, thus affecting the realism of the generated image.
[0005] In cross-identity facial expression-driven tasks, existing methods often ignore the facial differences between the driving and target individuals, directly applying the driving person's expression parameters or offsets to the target person's geometric model. However, individual facial details are highly personalized, and direct transfer can lead to severe geometric incompatibility. For example, directly transferring the deep wrinkles of an elderly person to a child's smooth face, or forcibly applying specific muscle deformations to mismatched skeletal structures, will produce visual artifacts or distortions that do not conform to anatomical rules. Existing technologies lack a mechanism that can adaptively map driving expressions onto the target identity-specific geometry, limiting the naturalness and realism of expression transfer.
[0006] In view of this, the present invention is hereby proposed. Summary of the Invention
[0007] The purpose of this invention is to provide a high-fidelity facial expression portrait video generation method based on personalized representation, thereby solving the aforementioned technical problems existing in the prior art. The method of this invention does not rely on a large amount of specific training data; it can generate cinematic-quality portrait videos through only a single optimization of network parameters, significantly improving the realism and controllability of the generated results.
[0008] The objective of this invention is achieved through the following technical solution: A method for generating high-fidelity facial expression portrait videos based on personalized representations, the method comprising: Step 1: Perform initial 3D reconstruction based on the expressive skinned multi-person linear model SMPL-X for the input reference object image and driving video; Step 2: Construct a personalized head representation of the reference object. By optimizing the static offset field and the dynamic offset field, the head geometry of the reference object is decoupled into identity-specific static geometry and expression-related dynamic details. Step 3: Based on the decoupling principle of personalized head representation, spatial constraints and temporal regularization are applied to optimize the personalized head representation; Step 4: Construct an identity-adaptive expression transfer module. Based on the static geometry of the driving signal and the reference object, the driving signal is mapped to the geometric deformation specific to the target identity to achieve dynamic expression offset across identities. Step 5: Using the identity-adaptive expression transfer module, the expression of the driving video is transferred to the personalized head representation of the reference object, resulting in a 3D mesh sequence with the target expression while maintaining the identity of the reference object. Step 6: Render the geometric condition map based on the generated 3D mesh sequence, and use it as a control signal to input the video diffusion model to generate a high-fidelity facial expression portrait video.
[0009] Compared with existing technologies, the beneficial effects of this invention are: 1) High fidelity and detail capture: By introducing static and dynamic offset fields, the low-rank limitation of the traditional SMPL-X model is overcome, enabling the reconstruction of high-frequency geometric details such as wrinkles and micro-expressions. 2) High decoupling of identity and expression: Through the identity-adaptive expression transfer module, the geometric incompatibility problem when driven by different identities is solved, effectively preventing the leakage of the driver's features to the generated object. 3) Stability and consistency of generated videos: By using explicit personalized 3D head representation as strong geometric guidance and combining it with the generation capability of the diffusion model, temporally stable, identity-consistent, and expressively vivid video generation is achieved. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating the high-fidelity facial expression portrait video generation method based on personalized representation provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the optimization of personalized head representation as described in an embodiment of the present invention; Figure 3 This is a schematic diagram of the identity adaptive facial expression migration module according to an embodiment of the present invention. Detailed Implementation
[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them, and do not constitute a limitation on the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the protection scope of the present invention.
[0013] First, the following explanations are provided for the terms that may be used in this article: The term "and / or" means that either or both can be achieved simultaneously. For example, X and / or Y means that it includes both "X" or "Y" as well as the three cases of "X and Y".
[0014] The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.
[0015] The technical solution provided by this invention will be described in detail below. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they shall be performed according to conventional conditions in the art or conditions recommended by the manufacturer. Reagents or instruments used in the embodiments of this invention whose manufacturers are not specified are all conventional products that can be purchased commercially.
[0016] like Figure 1 The diagram shown is a flowchart illustrating a high-fidelity facial expression portrait video generation method based on personalized representation provided in an embodiment of the present invention. The method includes: Step 1: For the input reference object image and driving video, perform initial 3D reconstruction based on the Expressive Skinned Multi-Person Linear Model (SMPL-X); In this step, specifically, it involves using an image of the input reference object. and drive video sequence Features are extracted, and the shape parameters of SMPL-X are estimated from the reference image and driving video frame using the SMPL-X reconstruction algorithm. Facial expression coefficient Mandibular posture parameters Perform initial 3D reconstruction to obtain the basic mesh. ; Because the original SMPL-X is too sparse for capturing personalized geometry and subtle expressions, centroid interpolation is used. , base mesh Mapped to a high-resolution dense mesh This provides a geometric basis for subsequent high-frequency detail modeling.
[0017] Step 2: Construct a personalized head representation of the reference object. By optimizing the static offset field and the dynamic offset field, the head geometry of the reference object is decoupled into identity-specific static geometry and expression-related dynamic details. In this step, the personalized head representation of the reference object is achieved by adding a decoupled offset field to the dense SMPL-X mesh. For the i-th frame, the personalized head representation... Defined as: (1) in, This is the upsampled SMPL-X dense mesh; It is a static offset field used to capture static geometric features unique to the reference object that do not change with facial expressions, such as specific head shape, hairstyle outline, etc. It is a dynamic offset field used to capture non-linear skin deformations that change over time and are related to specific facial expressions, such as wrinkles and dimples; For dynamic offset field Apply a minimum magnitude penalty so that it only represents a tiny zero-mean deviation relative to the static geometric features, ensuring It only includes facial expressions and does not include identity structure.
[0018] Through the explicit decoupling design described above, the embodiments of the present invention can endow the model with extremely strong expressive power while maintaining the stability of the identity structure.
[0019] Step 3: Based on the decoupling principle of personalized head representation, spatial constraints and temporal regularization are applied to optimize the personalized head representation; In this step, such as Figure 2 The diagram illustrates the optimization process of the personalized head representation according to an embodiment of the present invention. It shows the construction process from SMPL-X to a detailed mesh containing static / dynamic offset fields, the application of spatial constraints and temporal regularization to the personalized head representation, and the loss function of the optimization process including keypoint loss. Normal loss Deep loss Regularization loss Total loss The formula is as follows: (2) Among them, key point losses The geometry is constrained by calculating the distance between the projected 3D keypoints and the 2D keypoints detected in the image, as shown in the following formula: (3) in, It is a projection operation; These are 3D key points selected from personalized head representations; These are 2D key points detected from the image; These are camera parameters; Normal loss and depth loss The formula used to capture high-frequency geometric details and perform pixel-level supervision is as follows: (4) (5) in, and These are the normal map and depth map generated by the differentiable renderer, respectively. and These are the ground normal map and depth map estimated from the input image, respectively. To ensure the effectiveness of decoupling and the physical rationality of the geometry, regularization loss is used. Includes the following items: 1) Expression coefficient regularization term : SMPL-X expression coefficients for each frame Apply Constraints are represented as: (6) 2) Displacement amplitude penalty item To force the dynamic offset field to represent only small dynamic deviations, the dynamic offset field at each grid vertex is... Apply the minimum amplitude penalty, the formula is: (7) capital letters This represents the total number of vertices in the grid; the lowercase 'v' represents one of the grid vertices. Displacement amplitude penalty This causes the optimizer to interpret all static and shared geometric details as static offset fields. Thus Represents the true geometric mean; 3) Laplace smoothing term To suppress high-frequency artifacts and ensure surface smoothness, the static offset field... and dynamic offset field All are subject to Laplace smoothing constraints, denoted as .
[0020] Step 4: Construct an identity-adaptive expression transfer module. Based on the static geometry of the driving signal and the reference object, the driving signal is mapped to the geometric deformation specific to the target identity to achieve dynamic expression offset across identities. In this step, in cross-identity driven scenarios, directly applying the driver's facial expression offset can lead to geometric feature incompatibility, such as imposing adult wrinkle textures on a child's face. Therefore, this invention constructs an identity-adaptive facial expression transfer module, such as... Figure 3 The diagram shown is a schematic of the design of the identity-adaptive expression transfer module according to an embodiment of the present invention. It illustrates how to predict dynamic offsets based on target geometric features and driving encoders. This identity-adaptive expression transfer module uses a multilayer perceptron (MLP) as its main structure. The constructed identity-adaptive expression transfer module includes a driving signal encoder and a detail prediction network, wherein: Drive signal encoder The SMPL-X expression coefficients driving the video and mandibular posture parameters Encode as driving feature vector , represented as: (8) Detail Prediction Network Personalized neutral geometric features of the reference object and driving feature vector As input, predict dynamic offset fields adapted to the identity of the reference object. : (9) in, , This is the upsampled SMPL-X dense mesh. This is a static offset field; The identity-adaptive expression transfer module learns the mapping relationship between the geometric features of the reference object and the driving signal to achieve adaptive transfer of expression details between different identities, thus avoiding identity leakage and artifacts.
[0021] Step 5: Using the identity-adaptive expression transfer module, the expression of the driving video is transferred to the personalized head representation of the reference object, resulting in a 3D mesh sequence with the target expression while maintaining the identity of the reference object. In this step, the resulting three-dimensional mesh sequence Represented as: (10).
[0022] Step 6: Render the geometric condition map based on the generated 3D mesh sequence, and use it as a control signal to input the video diffusion model to generate a high-fidelity facial expression portrait video.
[0023] In this step, the pose-based mesh for each frame is first obtained through a linear blend skinning (LBS) operation. , is represented as: (11) The 3D mesh sequence obtained in step 5, which contains the target expression and retains the identity of the reference object, is represented as follows: ; Then, the pose mesh for each frame is generated using a differentiable renderer. Rendered as a sequence of geometric normal maps , is represented as: (12) Where F represents the total length of the normal graph sequence; D represents any frame in the sequence; Indicates a differentiable renderer; This represents the camera parameters, which are used to project the pose-based mesh into camera space and perform differentiable rendering. This geometric normal diagram sequence It contains precise 3D geometry and micro-expression details, but strips away texture and color information; Sequence of geometric normal maps As a spatial geometric control signal, combined with the image appearance features of the reference object, it is input into the video diffusion model; the video diffusion model adopts a diffusion transformer (DiT) architecture; The video diffusion model in the geometric normal map sequence Guided by this process, noise is gradually removed to generate continuous video frames, resulting in a high-fidelity portrait video that combines consistency with the identity of the reference object with rich facial expression details.
[0024] In practice, because the normal map provides extremely accurate 3D geometric guidance (including micro-expression and wrinkle details) while stripping away the texture color information of the driver, the generated result has both extremely high expression fidelity and perfectly maintains the identity characteristics of the reference object.
[0025] In summary, the method described in this invention solves the problems of identity drift and stiff facial expressions in traditional methods by combining explicit personalized geometric representation with a powerful generative diffusion model, and achieves high-fidelity, controllable and highly generalizable digital portrait video generation.
[0026] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method.
[0027] This invention also provides a computer storage medium storing a plurality of instructions adapted for loading and executing the method by a processor.
[0028] In summary, the method described in this embodiment of the invention has the following significant advantages compared to traditional methods for generating and driving human portrait videos: (1) Deep decoupling and high-fidelity reconstruction of identity features and facial expression dynamics: This invention breaks the limitation of traditional parametric models (such as 3DMM and SMPL-X) that treat the face as a low-rank linear subspace, and can decouple and model the static identity geometry and dynamic facial expression details of the 3D face separately. By introducing jointly optimized static and dynamic offset fields, this invention not only retains the prior advantages of the SMPL-X topology, but also endows the model with the ability to capture high-frequency nonlinear skin deformations (such as dynamic wrinkles, dimples and other micro-expressions). Compared with the shortcomings of existing methods that are difficult to distinguish between identity structure and facial expression changes, resulting in missing details or identity confusion in the generated results, this invention achieves accurate independent representation of identity and expression, thereby supporting personalized reconstruction with extremely high fidelity; (2) Identity-Adaptive Cross-Object Expression Transfer Capability: With the help of the identity-adaptive expression transfer module, this invention can effectively solve the "geometric incompatibility" problem caused by differences in facial geometry when driving across identities. This module maps the driving signal to a dynamic offset that conforms to the target's anatomical structure by using the personalized neutral geometric features of the target object as a condition, thus ensuring the anatomical rationality of expression transfer. Compared with traditional methods that directly copy expression parameters, resulting in the incorrect transfer of wrinkles from the elderly to children's faces or producing unnatural facial distortions, this invention achieves vivid expression driving while minimizing identity leakage and ensuring the purity of the target identity.
[0029] (3) Detailed capture of high-frequency micro-expressions and complex movements: By combining explicit personalized head representation with high-resolution normal map guidance, this invention can accurately capture and reproduce subtle facial movements that cannot be expressed by sparse keypoint methods. Whether it is a subtle twitch at the corner of the eye or a complex muscle change at the corner of the mouth, the dynamic offset field of this invention can be accurately described by nonlinear deformation. Compared with the shortcomings of existing methods based on implicit features, which are prone to losing high-frequency details or causing blurred expressions, the videos generated by this invention have stronger expressiveness and richer detail levels, significantly improving the visual realism.
[0030] (4) Temporal stability and cinematic quality of generated videos: This invention effectively constrains temporal jitter during the generation process by inputting high-precision 3D geometric signals (normal maps) as strong conditions into the video diffusion model (DiT), ensuring smooth transitions between video frames. Compared to the facial distortion or texture flickering that easily occurs in long video generation by pure 2D generation methods, this invention utilizes the stability of explicit 3D geometry to achieve excellent temporal consistency in both self-replay and cross-replay tasks, and achieves cinematic rendering quality thanks to the generation capabilities of the video diffusion model.
[0031] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.
Claims
1. A method for generating high-fidelity facial expression portrait videos based on personalized representation, characterized in that, The method includes: Step 1: Perform initial 3D reconstruction based on the expressive skinned multi-person linear model SMPL-X for the input reference object image and driving video; Step 2: Construct a personalized head representation of the reference object. By optimizing the static offset field and the dynamic offset field, the head geometry of the reference object is decoupled into identity-specific static geometry and expression-related dynamic details. Step 3: Based on the decoupling principle of personalized head representation, spatial constraints and temporal regularization are applied to optimize the personalized head representation; Step 4: Construct an identity-adaptive expression transfer module. Based on the static geometry of the driving signal and the reference object, the driving signal is mapped to the geometric deformation specific to the target identity to achieve dynamic expression offset across identities. Step 5: Using the identity-adaptive expression transfer module, the expression of the driving video is transferred to the personalized head representation of the reference object, resulting in a 3D mesh sequence with the target expression while maintaining the identity of the reference object. Step 6: Render the geometric condition map based on the generated 3D mesh sequence, and use it as a control signal to input the video diffusion model to generate a high-fidelity facial expression portrait video.
2. The method for generating high-fidelity facial expression portrait videos based on personalized representation according to claim 1, characterized in that, In step 1, specifically, the image of the input reference object... and drive video sequence Features are extracted, and the shape parameters of SMPL-X are estimated from the reference image and driving video frame using the SMPL-X reconstruction algorithm. Facial expression coefficient Mandibular posture parameters Perform initial 3D reconstruction to obtain the basic mesh. ; Then through centroid interpolation operation , base mesh Mapped to a high-resolution dense mesh This provides a geometric basis for subsequent high-frequency detail modeling.
3. The method for generating high-fidelity facial expression portrait videos based on personalized representation according to claim 2, characterized in that, In step 2, the personalized head representation of the reference object is achieved by adding a decoupled offset field to the dense SMPL-X mesh. For the i-th frame, the personalized head representation... Defined as: (1) in, This is the upsampled SMPL-X dense mesh; It is a static offset field used to capture static geometric features unique to the reference object that do not change with facial expressions; It is a dynamic offset field used to capture non-linear skin deformations that vary over time and are related to specific facial expressions; For dynamic offset field Apply a minimum magnitude penalty so that it only represents a tiny zero-mean deviation relative to the static geometric features, ensuring It only includes facial expressions and does not include identity structure.
4. The method for generating high-fidelity facial expression portrait videos based on personalized representation according to claim 1, characterized in that, In step 3, spatial constraints and temporal regularization are applied to the personalized head representation, and the loss function of the optimization process includes keypoint loss. Normal loss Deep loss Regularization loss Total loss The formula is as follows: (2) Among them, key point losses The geometry is constrained by calculating the distance between the projected 3D keypoints and the 2D keypoints detected in the image, as shown in the following formula: (3) in, It is a projection operation; These are 3D key points selected from personalized head representations; These are 2D key points detected from an image; These are camera parameters; Normal loss and depth loss The formula used to capture high-frequency geometric details and perform pixel-level supervision is as follows: (4) (5) in, and These are the normal map and depth map generated by the differentiable renderer, respectively. and These are the ground normal map and depth map estimated from the input image, respectively. Regularization loss Includes the following items: 1) Expression coefficient regularization term : SMPL-X expression coefficients for each frame Apply Constraints are represented as: (6) 2) Displacement amplitude penalty item Dynamic offset field at each grid vertex Apply the minimum amplitude penalty, the formula is: (7) capital letters This represents the total number of vertices in the grid; the lowercase 'v' represents one of the grid vertices. Displacement amplitude penalty This causes the optimizer to interpret all static and shared geometric details as static offset fields. Thus Represents the true geometric mean; 3) Laplace smoothing term : For static offset field and dynamic offset field All are subject to Laplace smoothing constraints, denoted as .
5. The method for generating high-fidelity facial expression portrait videos based on personalized representation according to claim 2, characterized in that, In step 4, the constructed identity-adaptive expression transfer module includes a driving signal encoder and a detail prediction network, wherein: Drive signal encoder The SMPL-X expression coefficients driving the video and mandibular posture parameters Encoding as driving feature vectors , represented as: (8) Detail Prediction Network Personalized neutral geometric features of the reference object and driving feature vector As input, predict dynamic offset fields adapted to the identity of the reference object. : (9) in, , This is the upsampled SMPL-X dense mesh. This is a static offset field; The identity-adaptive expression transfer module learns the mapping relationship between the geometric features of a reference object and the driving signal to achieve adaptive transfer of expression details between different identities.
6. The method for generating high-fidelity facial expression portrait videos based on personalized representation according to claim 5, characterized in that, In step 6, the pose mesh for each frame is first obtained through a linear blending skinning (LBS) operation. , represented as: (11) The 3D mesh sequence obtained in step 5, which contains the target expression and retains the identity of the reference object, is represented as follows: ; Then, the pose mesh for each frame is generated using a differentiable renderer. Rendered as a sequence of geometric normal maps , represented as: (12) Where F represents the total length of the normal graph sequence; D represents any frame in the sequence; Indicates a differentiable renderer; This represents the camera parameters, which are used to project the pose-based mesh into camera space and perform differentiable rendering. This geometric normal diagram sequence It contains precise 3D geometry and micro-expression details, but strips away texture and color information; Sequence of geometric normal maps As a spatial geometric control signal, combined with the image appearance features of the reference object, it is input into the video diffusion model; the video diffusion model adopts a diffusion transformer architecture; The video diffusion model in the geometric normal map sequence Guided by this process, noise is gradually removed to generate continuous video frames, resulting in a high-fidelity portrait video that combines consistency with the identity of the reference object with rich facial expression details.
7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 6.
8. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the method of any one of claims 1 to 6.