A head portrait expression transfer method based on parameter fitting and gradient migration

By combining parameter fitting and gradient transfer methods, the problems of global consistency and detail restoration in facial expression transfer in 3D facial animation are solved, achieving high-precision and natural facial expression transfer effects, which are suitable for efficient facial expression-driven and interactive applications with multiple characters and multiple scenarios.

CN120976381BActive Publication Date: 2026-01-27SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511510273.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-27
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-precision, natural facial expression transfer while maintaining consistency in character identity, especially in 3D facial animation, where it is difficult to balance global consistency with detailed reproduction.

Method used

A method combining parameter fitting and gradient transfer is adopted. A continuous 3D mesh sequence of the source character is generated by driving the audio signal. The control parameters are fitted using the FLAME model. The target character is reconstructed by combining a 3D Gaussian head model. Subtle facial expression changes are transferred through linear blending skinning and gradient optimization, and the Gaussian point cloud model is updated.

Benefits of technology

It achieves high-precision facial expression transfer with low computational overhead, maintains the identity consistency and geometric stability of the target character, and compensates for subtle facial expression changes lost during parameter fitting, generating natural and realistic facial expressions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976381B_ABST
    Figure CN120976381B_ABST
Patent Text Reader

Abstract

The application provides a head portrait expression migration method based on parameter fitting and gradient migration, and relates to the technical field of three-dimensional face animation, and specifically comprises the following steps: inputting an audio signal into an audio driving model to generate a continuous three-dimensional grid sequence of a source character, obtaining a control parameter set, and extracting a mouth movement parameter from the control parameter set; reconstructing a target character head portrait based on a three-dimensional Gaussian head model, generating a parameter fitting grid sequence of the target character through linear mixed skinning, and obtaining a rough expression migration result; comparing an original grid of the source character with the parameter fitting grid to calculate the deformation gradient of the two; migrating the deformation gradient of the source character to the deformation gradient of the target character through optimization, thereby obtaining a fine grid sequence of the target character, and updating the Gaussian point cloud model of the target character to obtain a high-precision expression migration result. The technical scheme of the application overcomes the problem in the prior art that it is difficult to balance global consistency and detail restoration in head portrait expression migration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D facial animation, specifically to a method for transferring facial expressions from avatars based on parameter fitting and gradient transfer. Background Technology

[0002] With the rapid development of virtual humans, digital interaction, and immersive human-computer communication, 3D facial animation technology has gradually become an important research direction in this field. Existing methods typically utilize audio signals to drive dynamic changes in a 3D facial mesh, thereby generating a sequence of expressions synchronized with speech. These methods can achieve automatic synthesis of facial movements to a certain extent, but because they are directly based on mesh representation, they still have shortcomings in capturing subtle changes in facial expressions and maintaining character consistency.

[0003] In recent years, parametric face modeling techniques have been widely used in face modeling and animation, controlling expressions, poses, and identities through low-dimensional parameters. However, parameter differences between different characters often lead to inaccurate expression transfer, making it difficult to achieve natural expression mapping while preserving individual characteristics. Meanwhile, while relying solely on parametric models can ensure consistent global expression modeling, its expressive power is limited, easily losing high-frequency details such as the corners of the mouth and wrinkles; directly relying on the original mesh for transfer makes it difficult to maintain consistency and controllability between characters.

[0004] Meanwhile, 3D avatar modeling methods based on Gaussian representation have attracted widespread attention due to their high rendering efficiency and realistic visual effects. These methods typically achieve controllable reconstruction and rendering of avatars by binding Gaussian kernels to 3D meshes. However, directly using audio signals to drive the Gaussian representation model often requires a large amount of paired data for training, resulting in high training costs and difficulty in widespread adoption. In contrast, using existing mesh-driven methods for expression transfer can map the expressions of a source character to a target character at a lower cost, but relying on a single method makes it difficult to balance global consistency and detail reproduction.

[0005] Therefore, there is a need for a parameter fitting and gradient transfer-based avatar expression transfer method that can achieve high-precision and natural expression transfer while maintaining the consistency of the character's identity. Summary of the Invention

[0006] The main objective of this invention is to provide a method for avatar expression transfer based on parameter fitting and gradient transfer, so as to solve the problem that avatar expression transfer in the prior art is difficult to balance global consistency and detail restoration.

[0007] To achieve the above objectives, this invention provides a method for avatar expression transfer based on parameter fitting and gradient transfer, specifically including the following steps:

[0008] S1, input the audio signal into the audio driving model to generate a continuous three-dimensional mesh sequence of the source character, and fit it based on the FLAME model to obtain a set of control parameters, and extract the mouth movement parameters from the set of control parameters.

[0009] S2 reconstructs the target character's head based on a 3D Gaussian head model, combines the mouth motion parameters extracted from the source character with the initial parameters of the target character, and generates a parameter fitting mesh sequence for the target character through linear blending skinning to obtain a rough expression transfer result.

[0010] S3 calculates the deformation gradient of the source character's original mesh and the parameter-fitted mesh by comparing the two. This allows for the extraction of subtle facial expressions that were not captured by parameter fitting.

[0011] S4. The deformation gradient of the source character is transferred to the deformation gradient of the target character through optimization, thereby obtaining a fine mesh sequence of the target character and updating the Gaussian point cloud model of the target character to obtain a high-precision expression transfer result.

[0012] Furthermore, step S1 specifically includes the following steps:

[0013] S1.1 Select a character from the dataset as the source character, and obtain a frame of the source character with no expression and no pose change as the initial mesh. .

[0014] S1.2, Input an audio signal Call the audio driver model, based on the audio signal. Output source character A continuous three-dimensional mesh sequence driven by audio:

[0015] ;

[0016] in, For frame number, Indicates the source role in the Frame grid, It is an audio-driven model.

[0017] S1.3, will As the source character's original mesh sequence, with the initial mesh FLAME model fitting was performed separately to obtain the face control parameters corresponding to the grid:

[0018] ;

[0019] ;

[0020] in, Indicates source role In the Frame control parameters; This represents the initial control parameters of the source character.

[0021] S1.4, from and Parameters related to mouth movements were extracted separately and denoted as follows: and ,in, express Parameters related to mouth movements, express Parameters related to mouth movements.

[0022] S1.5, normalize the mouth parameters for each frame:

[0023] ;

[0024] in, Indicates the source role in the The increment of mouth motion parameters relative to a neutral expression frame includes: jaw motion parameters and overall rotation parameters.

[0025] Furthermore, step S2 specifically includes the following steps:

[0026] S2.1, Using a 3D Gaussian head model, train and reconstruct a Gaussian point cloud model of the target character. .

[0027] S2.2, will Corresponding 3D network As the target character, a single frame of mesh data with no facial expression or pose change is obtained from the target character and used as the initial mesh. And obtain the corresponding FLAME control parameters from the dataset, that is, the initial control parameters of the target character. .

[0028] S2.3, Increment the mouth motion parameters obtained from the source character. Initial control parameters of the target character Combined, the target role is obtained. In the Frame control parameters :

[0029] ;

[0030] in, This represents the superposition of mouth movement parameters.

[0031] S2.4, based on the linear blending skinning method, adjusts the target character control parameters. Parameter fitting mesh sequence mapped to target role ;in, For target role In the The parameters of the frame are fitted to the grid.

[0032] Furthermore, step S3 specifically includes the following steps:

[0033] S3.1, Increment the mouth motion parameters obtained from the source character. Initial control parameters of the source character Combining these, we obtain the source character in the first place. Frame control parameters :

[0034] ;

[0035] in, This represents the superposition of mouth movement parameters.

[0036] S3.2, based on the linear hybrid skinning method, will Parameter fitting mesh sequence mapped to source role ;in, For the source character In the The parameters of the frame are fitted to the grid.

[0037] S3.3, Calculate the original mesh sequence of the source role. and The gradient difference between them.

[0038] Furthermore, step S3.3 specifically includes:

[0039] Set the character grid by Composed of triangular facets, for the first... The first frame There are three triangles, and the edge vector matrix of the original mesh is denoted as... The edge vector matrix of the parameter-fitted mesh is denoted as The edge vector matrix is ​​obtained from the grid vertices; deformation gradient Defined as:

[0040] ;

[0041] in, Indicates the source role in the The first frame of the original mesh relative to the parameter-fitted mesh The local deformation gradient of a triangle.

[0042] Furthermore, step S4 specifically includes the following steps:

[0043] S4.1, in order to transform the gradient of the source character Migrate to the target role using the gradient transfer method and randomly fix a constraint vertex. As a constraint, the deformation gradient of the target character is obtained through optimization. ; Set the character grid as follows: Composed of triangular facets, for the first... The first frame There are three triangles, and the edge vector matrix of the mesh fitted to the target character parameters is denoted as... The set of grid vertices is denoted as ;No. The optimization objective for the frame is:

[0044] ;

[0045] in, Indicates the source role in the The first frame The deformation gradient of a triangle, It is a constant. For constrained vertices, It is a minimum value function. This represents the Frobenius norm.

[0046] S4.2, based on the deformation gradient of the target character To obtain the fine mesh sequence of the target character ;in, For target role A finely detailed grid.

[0047] S4.3, Based on the target character's fine mesh sequence Update the Gaussian point cloud model of the target character. The relevant parameters are used to achieve facial expression transfer from the source character to the target character.

[0048] The present invention has the following beneficial effects:

[0049] This invention combines parameter fitting and gradient transfer to achieve high-precision facial expression transfer from a source character to a target character based on Gaussian representation without additional training. It can not only compensate for subtle facial expression changes lost during parameter fitting, ensuring the naturalness and realism of facial expressions, but also maintain the identity consistency and geometric stability of the target character. At the same time, it has low computational overhead and good real-time performance and versatility, making it suitable for efficient facial expression-driven and interactive applications with multiple characters and multiple scenarios. Attached Figure Description

[0050] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0051] Figure 1 A flowchart of an image expression transfer method based on parameter fitting and gradient transfer according to the present invention is shown.

[0052] Figure 2 The diagram shows the effect of the method proposed in this invention for avatar expression transfer. Detailed Implementation

[0053] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0054] like Figure 1 The illustrated method for avatar expression transfer based on parameter fitting and gradient transfer specifically includes the following steps:

[0055] S1. Input the audio signal into the audio driving model (such as Imitator) to generate a continuous three-dimensional mesh sequence of the source character, and fit it based on the FLAME model to obtain a set of control parameters. Extract the mouth movement parameters from the set of control parameters.

[0056] S2 reconstructs the target character's head based on a 3D Gaussian head model, combines the mouth motion parameters extracted from the source character with the initial parameters of the target character, and generates a parameter fitting mesh sequence for the target character through linear blending skinning to obtain a rough expression transfer result.

[0057] S3 calculates the deformation gradient of the source character's original mesh and the parameter-fitted mesh by comparing the two. This allows for the extraction of subtle facial expressions that were not captured by parameter fitting.

[0058] S4. The deformation gradient of the source character is transferred to the deformation gradient of the target character through optimization, thereby obtaining a fine mesh sequence of the target character and updating the Gaussian point cloud model of the target character to obtain a high-precision expression transfer result.

[0059] Specifically, step S1 includes the following steps:

[0060] S1.1 Select a character from the dataset (e.g., the vocaset dataset) as the source character, and obtain a frame of the source character with no expression and no pose change as the initial mesh. The mesh is generated by the parametric face model FLAME, which can drive the dynamic deformation of the mesh by controlling parameters.

[0061] S1.2, Input an audio signal It calls an existing audio driver model (such as Imitator) based on the audio signal. Output source character A continuous three-dimensional mesh sequence driven by audio:

[0062] ;

[0063] in, For frame number, Indicates the source role in the Frame grid, It is an audio-driven model.

[0064] S1.3, will As the source character's original mesh sequence, with the initial mesh FLAME model fitting was performed separately to obtain the face control parameters corresponding to the grid:

[0065] ;

[0066] ;

[0067] in, Indicates source role In the Frame control parameters; This represents the initial control parameters of the source character.

[0068] S1.4, from and Parameters related to mouth movements were extracted separately and denoted as follows: and ,in, express Parameters related to mouth movements, express Parameters related to mouth movements.

[0069] S1.5, In order to eliminate individual static differences, the mouth parameters of each frame are normalized:

[0070] ;

[0071] in, Indicates the source role in the The increment of mouth motion parameters relative to a neutral expression frame includes: jaw motion parameters and overall rotation parameters.

[0072] Specifically, step S2 includes the following steps:

[0073] S2.1, using existing 3D Gaussian head models (such as Gaussian Avatars), train and reconstruct a Gaussian point cloud model of the target character. This model achieves a driveable high-fidelity Gaussian avatar by binding Gaussian points to a 3D mesh.

[0074] S2.2, will Corresponding 3D network As the target character, a single frame of mesh data with no facial expression or pose change is obtained from the target character and used as the initial mesh. And obtain the corresponding FLAME control parameters from the dataset, that is, the initial control parameters of the target character. .

[0075] S2.3, Increment the mouth motion parameters obtained from the source character. Initial control parameters of the target character Combined, the target role is obtained. In the Frame control parameters :

[0076] ;

[0077] in, This represents the superposition of mouth movement parameters.

[0078] S2.4, based on the Linear Blend Skinning (LBS) method, adjusts the target character control parameters. Parameter fitting mesh sequence mapped to target role ;in, For target role In the The parameters of the frame are fitted to the grid.

[0079] Specifically, step S3 includes the following steps:

[0080] S3.1, Increment the mouth motion parameters obtained from the source character. Initial control parameters of the source character Combining these, we obtain the source character in the first place. Frame control parameters :

[0081] ;

[0082] in, This represents the superposition of mouth movement parameters.

[0083] S3.2, based on the linear hybrid skinning method, will Parameter fitting mesh sequence mapped to source role ;in, For the source character In the The parameters of the frame are fitted to the grid.

[0084] S3.3, Calculate the original mesh sequence of the source role. and The gradient difference between them.

[0085] Specifically, step S3.3 is as follows:

[0086] Set the character grid by Composed of triangular facets, for the first... The first frame There are three triangles, and the edge vector matrix of the original mesh is denoted as... The edge vector matrix of the parameter-fitted mesh is denoted as The edge vector matrix is ​​obtained from the grid vertices; deformation gradient Defined as:

[0087] ;

[0088] in, Indicates the source role in the The first frame of the original mesh relative to the parameter-fitted mesh The local deformation gradient of each triangle depicts subtle changes in the original expression that were not fully captured by the parameter fitting.

[0089] Specifically, step S4 includes the following steps:

[0090] S4.1, in order to transform the gradient of the source character Migrate to the target role using the gradient transfer method and randomly fix a constraint vertex. As a constraint, the deformation gradient of the target character is obtained through optimization. ; Set the character grid as follows: Composed of triangular facets, for the first... The first frame There are three triangles, and the edge vector matrix of the mesh fitted to the target character parameters is denoted as... The set of grid vertices is denoted as ;No. The optimization objective for the frame is:

[0091] ;

[0092] in, Indicates the source role in the The first frame The deformation gradient of a triangle, It is a constant. For constrained vertices, It is a minimum value function. This represents the Frobenius norm; the constraint is used to fix a vertex to avoid non-uniqueness of the solution.

[0093] S4.2, based on the deformation gradient of the target character To obtain the fine mesh sequence of the target character ;in, For target role A finely detailed grid.

[0094] S4.3, Based on the target character's fine mesh sequence Update the Gaussian point cloud model of the target character. The relevant parameters are used to achieve facial expression transfer from the source character to the target character.

[0095] To verify the feasibility and superiority of this invention, the following comparative experiment was conducted. The speech-driven source characters (characters 0731 and 0809 in the VOCASET dataset) generated by the Imitator model were selected as the source character inputs, and 20 audio segments were extracted from each character, for a total of 40 audio segments.

[0096] In the experiment, the facial expression information of the source character was mapped to the target character using two transfer methods: (1) coefficient transfer method, and (2) the parameter fitting and gradient transfer combination method proposed in this invention. Simultaneously, three audio-driven models, VOCA, Faceformer, and Imitator, were compared in the experiment. By analyzing the performance of these methods in lip movement accuracy and speech synchronization, the advantages of this invention can be effectively verified. The comparison results are shown in Table 1.

[0097] Table 1 Comparison of Audio-Driven Model Metrics

[0098]

[0099] Table 1 lists the different methods used for lip shaping. error( ) and lip synchronicity index ( On the surface: lip shape The error metric quantifies the spatial error between the lip vertex position in the generated mesh and the reference mesh. A lower metric indicates more accurate lip shape generation. The lip-sync metric measures the degree of synchronization between the generated lip shape and the input audio. A lower metric indicates more accurate alignment of lip movements with speech, resulting in a more natural lip shape. Experimental results show that the gradient transfer and gradient transfer combination method proposed in this invention achieves better lip shape generation. It outperforms the comparison methods in both error and lip synchronization metrics. Specifically, the lip error is as low as 0.0889, significantly better than the imitator and coefficient transfer methods, and can more accurately reproduce lip dynamics. Figure 2 Qualitative results further demonstrate that this method can more naturally preserve the subtle lip dynamics of the source character, and the generated target character has a clearer mouth shape, smoother boundaries, and a mouth opening range that is closer to the source character.

[0100] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A method for transferring facial expressions from avatars based on parameter fitting and gradient transfer, characterized in that, Specifically, the steps include the following: S1, input the audio signal into the audio driving model to generate a continuous three-dimensional mesh sequence of the source character, and fit it based on the FLAME model to obtain a set of control parameters, and extract the mouth movement parameters from the set of control parameters; S2, based on the 3D Gaussian head model, reconstruct the target character's head image, combine the mouth motion parameters extracted from the source character with the initial parameters of the target character, and generate the parameter fitting mesh sequence of the target character through linear blending skinning to obtain a rough expression transfer result; S3 calculates the deformation gradient of the source character's original mesh and the parameter-fitted mesh by comparing the two. This allows for the extraction of subtle facial expressions that were not captured by parameter fitting. S4. The deformation gradient of the source character is transferred to the deformation gradient of the target character through optimization, thereby obtaining a fine mesh sequence of the target character, and updating the Gaussian point cloud model of the target character to obtain a high-precision expression transfer result. Step S3 specifically includes the following steps: S3.1, Increment the mouth motion parameters obtained from the source character. Initial control parameters of the source character Combining these, we obtain the source character in the first place. Frame control parameters : ; in, This represents the superposition of mouth movement parameters; S3.2, based on the linear hybrid skinning method, will Parameter fitting mesh sequence mapped to source role ;in, For the source character In the Frame parameter fitting grid; S3.3, Calculate the original mesh sequence of the source role. and The gradient difference between them; Step S3.3 specifically includes: Set the character grid by Composed of triangular facets, for the first... The first frame There are three triangles, and the edge vector matrix of the original mesh is denoted as... The edge vector matrix of the parameter-fitted mesh is denoted as The edge vector matrix is ​​obtained from the grid vertices; deformation gradient Defined as: ; in, Indicates the source role in the The first frame of the original mesh relative to the parameter-fitted mesh The local deformation gradient of a triangle.

2. The avatar expression transfer method based on parameter fitting and gradient transfer according to claim 1, characterized in that, Step S1 specifically includes the following steps: S1.1 Select a character from the dataset as the source character, and obtain a frame of the source character with no expression and no pose change as the initial mesh. ; S1.2, Input an audio signal Call the audio driver model, based on the audio signal. Output source character A continuous three-dimensional mesh sequence driven by audio: ; in, For frame number, Indicates the source role in the Frame grid, An audio-driven model; S1.3, will As the source character's original mesh sequence, with the initial mesh FLAME model fitting was performed separately to obtain the face control parameters corresponding to the grid: ; ; in, Indicates source role In the Frame control parameters; Indicates the initial control parameters of the source character; S1.4, from and Parameters related to mouth movements were extracted separately and denoted as follows: and ,in, express Parameters related to mouth movements, express Parameters related to mouth movements; S1.5, normalize the mouth parameters for each frame: ; in, Indicates the source role in the The increment of mouth motion parameters relative to a neutral expression frame includes: jaw motion parameters and overall rotation parameters.

3. The avatar expression transfer method based on parameter fitting and gradient transfer according to claim 1, characterized in that, Step S2 specifically includes the following steps: S2.1, Using a 3D Gaussian head model, train and reconstruct a Gaussian point cloud model of the target character. ; S2.2, will Corresponding 3D network As the target character, a single frame of mesh data with no facial expression or pose change is obtained from the target character and used as the initial mesh. And obtain the corresponding FLAME control parameters from the dataset, that is, the initial control parameters of the target character. ; S2.3, Increment the mouth motion parameters obtained from the source character. Initial control parameters of the target character Combined, the target role is obtained. In the Frame control parameters : ; in, This represents the superposition of mouth movement parameters; S2.4, based on the linear blending skinning method, adjusts the target character control parameters. Parameter fitting mesh sequence mapped to target role ;in, For target role In the The parameters of the frame are fitted to the grid.

4. The avatar expression transfer method based on parameter fitting and gradient transfer according to claim 1, characterized in that, Step S4 specifically includes the following steps: S4.1, in order to transform the gradient of the source character Migrate to the target role using the gradient transfer method and randomly fix a constraint vertex. As a constraint, the deformation gradient of the target character is obtained through optimization. ; Set the character grid as follows: Composed of triangular facets, for the first... The first frame There are three triangles, and the edge vector matrix of the mesh fitted to the target character parameters is denoted as... The set of grid vertices is denoted as ;No. The optimization objective for the frame is: ; in, Indicates the source role in the The first frame The deformation gradient of a triangle, It is a constant. For constrained vertices, It is a minimum value function. Denotes the Frobenius norm; S4.2, based on the deformation gradient of the target character To obtain the fine mesh sequence of the target character ;in, For target role Fine mesh; S4.3, Based on the target character's fine mesh sequence Update the Gaussian point cloud model of the target character. The relevant parameters are used to achieve facial expression transfer from the source character to the target character.

Citation Information

Patent Citations

  • Method for optimizing human face joint driving model based on sparse expression

    CN118691719A

  • Virtual population type generation training method and system, medium and program product

    CN118820787A