Gaussian mixture shape method applicable to dynamic modeling of head of human body
Patent Information
- Application Number
- PCT/CN2024/080237
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2025-10-02
AI Technical Summary
Existing technologies in dynamic modeling of the human head lack the ability to reproduce high-frequency details and highlights, and have high computational requirements, making them difficult to be widely used by ordinary users.
Using the Gaussian mixed shape method, through training image acquisition and preprocessing, we established the head Gaussian model and the oral Gaussian model, optimized the Gaussian parameters, generated highly realistic human head animation, and used the Gaussian splashing technology to draw head images under new perspectives and expressions in real time.
It significantly improves facial details and highlight reproduction to state-of-the-art levels and runs more than five times faster, making it suitable for applications such as film, animation production, and online games.
Smart Images

Figure CN2024080237_02102025_PF_FP_ABST
Abstract
Description
A Gaussian mixture shape method for dynamic modeling of the human head Technical Field
[0001] The present invention relates to a parameterized human head model, head geometry and appearance modeling, facial motion capture and real-time animation technology, and in particular to a Gaussian mixed shape method suitable for dynamic modeling of the human head. Background Art
[0002] Researchers have proposed various models for representing the human head. Early work used explicit 3D meshes to reconstruct geometry and appearance from images. Blanz and Vetter (Blanz, V. and Vetter, T., 2023. A morphable model for the synthesis of 3D faces. In Proceedings of the 26th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH 1999) pioneered the 3DMM (3D Morphable Model) approach, which used a low-dimensional linear subspace to model facial geometry and appearance. There are many follow-up works along this direction, such as the complete human head model (Ploumpis, S., Ververas, E., O'Sullivan, E., Moschoglou, S., Wang, H., Pears, N., Smith, WA, Gecer, B. and Zafeiriou, S., 2020. Towards a complete 3D morphable model of the human head. IEEE transactions on pattern analysis and machine intelligence, 43(11), pp. 4142-4160.), and the deep nonlinear model (Tran, L. and Liu, X., 2018. Nonlinear 3d face morphable model. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 7346-7355).). The three-dimensional mesh representation is also used to construct riggable heads to generate human head animation.In order to generate high-definition animations, researchers further proposed image-based dynamic human head models, which include complete control of hair and headwear (Cao, C., Wu, H., Weng, Y., Shao, T. and Zhou, K., 2016. Real-time facial animation with image-based dynamic avatars. ACM Transactions on Graphics, 35(4).), or additionally introduced detail correction (Feng, Y., Feng, H., Black, MJ and Bolkart, T., 2021. Learning an animatable detailed 3D face model from in-the-wild images. ACM Transactions on Graphics (ToG), 40(4), pp. 1-13.).
[0003] To achieve highly realistic rendering, recent methods use neural radiance fields (Mildenhall, B., Srinivasan, PP, Tancik, M., Barron, JT, Ramamoorthi, R. and Ng, R., 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1), pp. 99-106.) to implicitly represent the human head and have achieved impressive results. For example, i3DMM (Yenamandra, T., Tewari, A., Bernard, F., Seidel, HP, Elgharib, M., Cremers, D. and Theobalt, C., 2021. i3dmm: Deep implicit 3d morphable model of human heads. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 12803-12813)) proposed the first neural implicit function of the human head based on a 3D deformable model. HeadNerf (Hong, Y., Peng, B., Xiao, H., Liu, L. and Zhang, J., 2022. Headnerf: A real-time nerf-based parametric head model. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 20374-20384)) introduces a parametric human head model based on neural radiation fields, integrating neural radiation fields into the parametric representation of the head.The current state-of-the-art INSTA (Zielonka, W., Bolkart, T. and Thies, J., 2023. Instant volumetric head avatars. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 4574-4584).) uses fast neural graphics primitives (InstantNGP) (Müller, T., Evans, A., Schied, C. and Keller, A., 2022. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4), pp. 1-15.) to model a dynamic neural radiation field around a parameterized human head model, capable of reconstructing the head in less than 10 minutes. PointAvatar (Zheng, Y., Yifan, W., Wetzstein, G., Black, MJ and Hilliges, O., 2023. Pointavatar: Deformable point-based head avatars from videos. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 21057-21067).) proposed a point cloud-based representation method, which learned a deformation field from the expression vector of the human head model (FLAME, Faces Learned with an Articulated Model and Expressions) (Li, T., Bolkart, T., Black, MJ, Li, H. and Romero, J., 2017. Learning a model of facial shape and expression from 4D scans. ACM Trans. Graph., 36(6), pp. 194-1.) to drive the point cloud.NeRFBlendshape (Gao, X., Zhong, C., Xiang, J., Hong, Y., Guo, Y. and Zhang, J., 2022. Reconstructing personalized semantic facial nerf models from monocular video. ACM Transactions on Graphics (TOG), 41(6), pp. 1-12.) combines multi-resolution voxel fields with expression coefficients to construct a neural radiance field-based blend shape model for semantic animation control and realistic rendering. Although some of these methods claim to be able to achieve interactive frame rates, they use high-end or professional computing equipment and cannot be truly promoted to ordinary users. In addition, these methods have some high-frequency loss phenomena, such as the loss of wrinkles and facial highlights.
[0004] Summary of the Invention
[0005] This invention addresses the shortcomings of existing technologies and proposes a Gaussian blend shape method suitable for dynamic modeling of the human head. This method uses a head video as input, optimizes the Gaussian blend shape, and, given new expression, joint, and posture parameters, drives the character to synthesize highly realistic human head animations in real time. Compared to existing methods, this method significantly improves facial detail and highlight reproduction, reaching a state-of-the-art level. It can be widely used in applications such as film and animation production, online games, and remote conferencing, and has high practical value.
[0006] The present invention is achieved through the following technical solution, which specifically includes the following steps:
[0007] A Gaussian mixture shape method suitable for dynamic modeling of a human head comprises the following steps:
[0008] (1) Training image acquisition and preprocessing: Use a monocular camera to shoot a short video of a person speaking, process the video data and obtain: a neutral expression mesh, a set of base expression meshes, a head foreground image of each frame, a head foreground mask, camera parameters, joint and posture parameters, and expression coefficients;
[0009] (2) Gaussian mixture shape training: using step (1), a head Gaussian model and an oral Gaussian model are established and their parameters are initialized, and all parameters are optimized according to the image loss; the head Gaussian model is composed of a neutral Gaussian model and a base expression Gaussian model;
[0010] (3) Human head animation generation: Based on the human head motion parameters provided by the user, the neutral Gaussian model and the base expression Gaussian model in the Gaussian mixture shape optimized in step (2) are linearly mixed and combined with the oral Gaussian model to generate a Gaussian model corresponding to the motion parameters, and the head animation under the new perspective and expression is generated in real time.
[0011] Furthermore, the training image acquisition and preprocessing in step (1) includes the following sub-steps:
[0012] (1.1) Training image acquisition: Use a monocular camera to shoot a head video as training data. The ambient lighting should remain constant during the shooting process, and the subject should maintain normal speaking movements.
[0013] (1.2) Training image preprocessing: Extract consecutive frames from the video to obtain a continuous image sequence, perform appropriate cropping and scaling on the image sequence, and use a face tracker to obtain a neutral expression grid, a set of base expression grids, camera parameters, joint and posture parameters, and expression coefficients for each frame from the image sequence; use a segmentation algorithm to mask out the head area from the image to obtain a head foreground image and a head foreground mask.
[0014] Furthermore, the training of the Gaussian mixture shape in step (2) includes the following sub-steps:
[0015] (2.1) Parameter initialization: The neutral expression mesh is converted into a neutral Gaussian model through sampling; the deformation gradient is extracted from the base expression mesh and applied to the neutral Gaussian model to generate the base expression Gaussian model; a mesh slice representing the teeth is generated at the corresponding position of the mouth in the neutral expression mesh, and the mesh slice representing the teeth is converted into the oral Gaussian model through sampling;
[0016] (2.2) Gaussian mixture shape consistency variable substitution: Introduce intermediate variables, replace the Gaussian attribute increment of the base expression Gaussian model with the expression of the intermediate variable, and convert the optimized Gaussian attribute increment into the optimized intermediate variable;
[0017] The substitution represents the Gaussian attribute increment as the sum of the product of the initialization result and the intermediate variable in step (2.1) and a heuristic function, wherein the heuristic function is a linear function with a threshold cutoff, and the input of the heuristic function is the displacement amplitude of the position deformation of the neutral expression grid of the Gaussian nearest neighbor to the base expression grid;
[0018] (2.3) Pre-calculating the 3D grid: Preset the 3D grid, search for the nearest neighbor triangle on the neutral expression grid from each grid point, and store the corresponding base expression grid displacement amplitude and linear blend skin weight on the grid point;
[0019] (2.4) Define the loss function: The loss function includes image loss, opacity loss and oral regularization constraint, and the mathematical expression is as follows: L = λ1L rgb +λ2L α +λ3L reg ; L rgb =(1-λ)L1+λL D-SSIM ;
[0020] Among them, L represents the loss function, L rgb represents the image loss, L α Indicates opacity loss, L reg represents the oral regularization constraint; λ1, λ2, λ3, λ are the loss mixture weights; L1 represents the mean absolute error, L D-SSIM Indicates structural differences; represents the opacity image, represents the head foreground mask, M is the total number of image pixels; SDF represents the signed distance, V represents the predefined oral volume, x i represents the position of the oral Gaussian, and N′ represents the total number of oral Gaussians.
[0021] Furthermore, the step (2.2) is specifically as follows: for each Gaussian G i , let ΔG i,k Represents the base expression Gaussian model B k The "increment" of the Gaussian property relative to the neutral Gaussian model B0, that is, ΔG i,k =G i,k -G i,0 ; By introducing intermediate variables ΔG i,k Expressed as:
[0022] in is the ΔG calculated in the initialization phase described in Section 2.1 i,k The initial value of is considered as a constant during optimization; |d i,k | is the closest to G i The surface position of the M0 mesh is deformed from M0 to M k The displacement amplitude that occurs, that is, the displacement amplitude of the base expression grid; is the intermediate parameter to be optimized, and its initial value is set to all 0; the linear function f(·) transforms M k The maximum amplitude between M0 and M1 is mapped to 1, and the amplitude threshold 0.00001 is mapped to 0; and ΔG is not directly optimized. i,k , but optimize After the optimization is completed, Substitute into the formula and calculate ΔGi,k , resulting in an expression blend shape.
[0023] Furthermore, the Gaussian mixture shape in step (2) is composed of a group of Gaussian models, wherein the group of Gaussian models includes a head Gaussian model and an oral Gaussian model, and the head Gaussian model is used to generate a human head model representing arbitrary posture and arbitrary expression by linearly mixing a neutral Gaussian model and a base expression Gaussian model, and performing a linear mixed skinning transformation on the mixed model and the oral Gaussian model.
[0024] Furthermore, the human head motion parameters in step (3) include joint and posture parameters, facial expression coefficients and camera parameters.
[0025] Furthermore, the optimized Gaussian mixed shape expression in step (3) and the human head motion parameters given by the user, the Gaussian model in the linear mixed Gaussian mixed shape generates a Gaussian model corresponding to the motion parameters, and the Gaussian splashing technology is used to draw the head image and animation under the new perspective and expression in real time.
[0026] The beneficial effects of the present invention are as follows:
[0027] This method uses a set of Gaussian models to linearly blend a neutral Gaussian model with a base expression Gaussian model, then performs a linear blend skinning transformation on the resulting model and the oral Gaussian model to generate a human head model representing any pose and expression. This method not only preserves more high-frequency details and more accurately reproduces highlights, such as clearer eyebrows and beard details, accurate facial wrinkles, and sharper nose and eye reflections, resulting in overall improved image and animation rendering quality, but also runs at speeds over five times faster than current state-of-the-art rendering methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] FIG1 shows the intermediate and final results of synthesizing the first individual head animation using the method of the present invention; FIG1 is the original image captured, FIG1 is the preprocessed head foreground image, FIG1 is the head foreground mask, and FIG1 is the final result synthesized under the new expression and viewing angle;
[0029] FIG2 shows the intermediate and final results of synthesizing a second individual head animation using the method of the present invention; FIGA is the original image captured, FIGB is the preprocessed head foreground image, FIGC is the head foreground mask, and FIGD is the final result synthesized under the new expression and viewing angle;
[0030] FIG3 shows the intermediate and final results of synthesizing a third individual's head animation using the method of the present invention; ...
[0031] Figure 4 shows the intermediate and final results of synthesizing the fourth individual head animation using the method of the present invention; wherein, Figure A is the original image captured, Figure B is the preprocessed head foreground image, Figure C is the head foreground mask, and Figure D is the final result synthesized under the new expression and perspective. DETAILED DESCRIPTION
[0032] The core of the present invention lies in a Gaussian blend shape expression and construction method suitable for modeling human head motion. Given a video of a human head speaking naturally, a Gaussian blend shape is optimized, which can be used to generate human head images and animations in real time under new perspectives and expressions.
[0033] The present invention provides a Gaussian mixture shape expression suitable for human head motion modeling, namely a Gaussian mixture shape; the Gaussian mixture shape is composed of a head Gaussian model B ψ and oral Gaussian model B m Furthermore, the head Gaussian model B ψ It consists of a neutral Gaussian model B0 and 50 base expression Gaussian models {B1, B2, ..., B 50 The above models are composed of a set of three-dimensional Gaussians, each of which has the following basic properties: position x∈R 3 , opacity α∈R, rotation q (quaternion representation), size s∈R 3 and spherical harmonic coefficients SH∈R 16×3 ; Each Gaussian of the neutral Gaussian model B0 is also equipped with a linear blend skin weight w for joint and posture control; Gaussian of the neutral Gaussian model B0 and each base expression Gaussian model B k There is a one-to-one correspondence between the Gaussians in . Define the base expression Gaussian model B k The difference between the Gaussian attributes of the neutral Gaussian model B0 is the "increment" ΔB k , ΔB k =B k -B0, such as traditional mesh blend shapes, the Gaussian model of a human head with any expression can be calculated as follows:
[0034] where {ψ k} is the expression coefficient of each base expression.
[0035] In addition to facial expression control, the present invention provides joint and posture parameters Θ for controlling the movement of the head, chin, eyeballs, and eyelids. These parameters are transformed by linear blend skinning (LBS) to transform the head Gaussian model (i.e., Gaussian attributes). The specific expression is:
[0036] Where T j (Θ)∈R 4×4 Represents the motion transformation matrix of each joint (head, chin, left eyeball, right eyeball, left eyelid, right eyelid), which is controlled by the joint and posture parameters Θ, w j ∈R N is the weight of the linear blend skinning, which determines the extent to which each Gaussian is affected by the motion of each joint, and N is the total number of Gaussians; T∈R N×4×4 is the transformation after mixing, which specifies the transformation applied to each Gaussian in the model; trans is the operator that transforms the Gaussian attribute using the transformation matrix, and has different operations for different Gaussian attributes. The specific expression is as follows: * =Tx; α * =α; q * =Rq; s * =s; SH * =rot(R,SH);
[0037] Where R=proj(T) represents extracting the rotation transformation component R from the affine transformation T, and rot(R,SH) represents rotating the spherical harmonics using the rotation transformation R and reprojecting them to obtain new spherical harmonic coefficients.
[0038] The oral cavity Gaussian model B ψ It is also composed of three-dimensional Gaussian (with the same five basic Gaussian properties: position, opacity, rotation, size and spherical harmonic coefficients), and is specifically used to represent the interior of the mouth, such as teeth, tongue, etc. The oral Gaussian model is divided into two parts, the upper oral model follows the movement of the head joint, and the lower oral model follows the movement of the jaw joint. The oral Gaussian model after the joint movement transformation is recorded as
[0039] In summary, the Gaussian model obtained after transformation is Use Gaussian splatting technology to draw high-realistic images in real time.
[0040] The various steps of the invention are described in detail below with reference to Figures 1-4.
[0041] Based on the Gaussian mixture shape method for constructing a human head, the present invention proposes a Gaussian mixture shape method suitable for dynamic modeling of the human head. The method includes the following steps: training image acquisition and preprocessing, Gaussian mixture shape training, and human head animation generation. Specifically, it includes the following steps:
[0042] Collection and preprocessing of training images
[0043] 1.1 Training Image Collection
[0044] The method uses a common monocular camera to shoot a 2-3 minute video of a human head as training data. The ambient lighting must remain roughly constant during the shooting process, and the subject must speak naturally.
[0045] 1.2 Training Image Preprocessing
[0046] As shown in Figure 1-4 (A), the present invention extracts images frame by frame from the video to obtain a continuous image sequence of 3000-5000 images, and performs appropriate cropping and scaling on the image sequence. A face tracker is used to extract FLAME (face learned with an articulated model and expressions) neutral expression grids, 50 base expression grids, camera parameters C, joint and posture parameters Θ, expression coefficients {ψ k}. Use the video-based human segmentation algorithm and face semantic segmentation algorithm (segment the head area to obtain the head foreground image As shown in Figure 1-4 (B), a head foreground mask (alpha mask) is generated for each image. As shown in Figure 1-4 (C).
[0047] Gaussian mixture shape training
[0048] 2.1 Parameter Initialization
[0049] The present invention respectively performs neutral Gaussian model B0, 50 base expression Gaussian models {B k} and oral Gaussian model B m Initialize. For the neutral Gaussian model B0, use Poisson disk sampling to uniformly sample points on the FLAME neutral expression mesh M0 and use it as the initial position of the Gaussian. Initialize the Gaussian to an isotropic Gaussian, specifically: initialize the size s to the average distance to the three nearest neighbor Gaussian centers, and initialize the rotation q to a unit quaternion. The opacity α is initialized to 0.1, the 0th order of the SH coefficient is initialized to a random number uniformly sampled from the interval [0,1 / 255.), and the non-0th order is initialized to 0. For each Gaussian of the neutral expression mesh B0, search for the nearest triangle on M0, and calculate the linear blend skin weight w of the Gaussian based on the linear blend skin weight interpolation on the triangle vertices.
[0050] For the 50-basis expression Gaussian model {B k}, extract the deformation gradient from the FLAME base expression grid, that is, the FLAME neutral expression grid M0 to each base expression grid M k The deformation gradient is applied to the neutral Gaussian model B0 to obtain the base expression Gaussian model {B k}, as initialization. Specifically, for each Gaussian G of the neutral Gaussian model B0 i,0 , search for the triangle nearest to Gauss on M0, and calculate the triangle deformation from M0 to M k The affine transformation T is applied to the Gaussian G i,0 At position x, the base expression Gaussian model B is generated k Each Gaussian G i,k Similarly, extract the rotation component R in T and apply it to the Gaussian G i,0 The rotation q and SH (similar to the operation in the aforementioned linear blend skinning) to generate the Gaussian G i,k Finally, let G i,k The size s and opacity α are simply equal to G i,0 The corresponding parameters are used to complete the initialization.
[0051] For oral Gaussian model B m Two mesh slices are predefined, corresponding to the upper and lower teeth, respectively. The oral Gaussian model is initialized using a similar method to the neutral Gaussian model, converting the mesh slices into a set of Gaussians. For the neutral and base expression Gaussian models, the number of initialized Gaussians is 70k, and for the oral Gaussian model, the number of initialized Gaussians is 14k (7k for each of the upper and lower parts).
[0052] 2.2 Gaussian Mixture Shape Consistency Variable Substitution
[0053] The present invention adopts a variable substitution technique to optimize the base expression Gaussian model B k And the corresponding FLAME base expression grid M k Maintain semantic consistency to avoid overfitting. Specifically, for each Gaussian G i , let ΔG i,k Represents the base expression Gaussian model B k The "increment" (i.e. ΔG) of the Gaussian property relative to the neutral Gaussian model B0 i,k =G i,k -G i,0 ), by introducing intermediate variables ΔG i,k Expressed as:
[0054] in is the ΔG calculated in the initialization phase described in Section 2.1 i,k The initial value of d is considered as a constant during optimization. i,k | is the closest to G i The surface position of the M0 mesh is deformed from M0 to M k The displacement amplitude that occurs (the displacement amplitude of the base expression grid), is the intermediate parameter to be optimized, and its initial value is set to all 0. The linear function f(·) transforms M k The maximum amplitude between ΔG and M0 is mapped to 1, and the amplitude threshold 0.00001 is mapped to 0. The present invention does not directly optimize ΔG i,k , but optimize After the optimization is completed, Substitute into the formula and calculate ΔG i,k , resulting in an expression blend shape.
[0055] 2.3 Parameter Optimization
[0056] The present invention jointly optimizes B0, {ΔB k} and B m :First, generate a Gaussian model of each frame based on the parameters of the frame, and draw the rendered image and opacity image. Calculate the loss of the rendered image and the head foreground image respectively, calculate the loss of the opacity image and the head foreground mask, calculate the regularization constraint loss of the oral Gaussian model, calculate the weighted sum of the above losses and the gradient of the parameters to be optimized, and use the Adam optimizer to optimize the parameters.
[0057] The training process is as follows: randomly extract a frame i from the training image and reconstruct the Gaussian model of the frame. According to the aforementioned Gaussian mixture shape expression, the expression coefficient {ψ k} linear mixture B0 and {ΔB k}Get the head Gaussian model B ψ , and then use the joint and posture parameters Θ of the frame to apply linear blending skinning to obtain the transformed head Gaussian model And the transformed oral Gaussian model Using Gaussian splattering technique and Draw into an image The entire rendering pipeline is differentiable so that the gradient back propagation algorithm can be used to update B0, {ΔB k} and B m parameter.
[0058] The loss function is defined as follows: L = λ1L rgb +λ2L α +λ3L reg ;
[0059] Among them L rgb is the image loss, L α is the opacity loss, L reg is the oral regularization constraint. The mixed weights of different losses are λ1=1,λ2=10,λ3=100. The image loss is expressed as and The calculated L1 (mean absolute error) and D-SSIM (Structural dissimilarity) loss functions are combined: L rgb =(1-λ)L1+λL D-SSIM ;
[0060] Where λ = 0.2.
[0061] Opacity loss is defined as follows:
[0062] in is the opacity image, is the head foreground mask mentioned above, and M is the total number of image pixels. and The SH coefficients are replaced with all 1s for the 0th order and all 0s for the non-0th order, and the opacity image is obtained by drawing with the Gaussian splashing technique.
[0063] Oral regularization constraint L reg The Gaussian of the oral Gaussian model is constrained to a predefined oral volume. Specifically, the signed distance of each Gaussian to the oral volume boundary is calculated, and the L2 loss is applied to pull the Gaussian outside the volume back:
[0064] Where N′ represents the total number of oral Gaussians, x i represents the position of the Gaussian, SDF represents the signed distance, and V represents the predefined cylindrical oral volume.
[0065] Optimization details: 40k iterations, for B0 and B m The learning rate setting is: 1.6×10 -4 , 5×10 -2 , 1×10 -3 , 5×10 -3 , 2.5×10 -3 They correspond to position, opacity, rotation, size and spherical harmonic coefficients respectively, where the learning rate of position is increased from the initial 1.6×10 -4 The exponential decay ends at 1.6×10 -6 For {ΔB k The learning rate setting is: 3.2×10 -7 , 5×10-5 , 1×10 -4 , 5×10 -4 , 1.25×10 -3 Corresponding to position, opacity, rotation, size and spherical harmonic coefficients respectively. The present invention uses an adaptive strategy to dynamically add and delete Gaussians, and performs an addition and deletion operation every 100 iterations from the 500th to the 15kth iteration. The addition operation accumulates the gradient of each position in the screen space exceeding τ pos = 0.001 is divided into two, where the Gaussian with a size less than or equal to 0.01 is simply copied in the original position. For the size greater than 0.01, the original Gaussian function is used as the probability density function for sampling to generate two new Gaussians, inheriting the original Gaussian parameters, but dividing the size s by 1.6 and deleting the original Gaussian; the deletion operation will reduce the opacity less than ∈ α = 0.005 Gaussian deletion. The present invention pre-calculates a 256 3 A 3D grid with a resolution of 1000x1000, used to quickly calculate the weight w of the linear blend skin and the displacement amplitude of the base expression grid |d i,k |, pre-calculate and store the values at the grid vertex positions, and when querying, access the eight grid vertices nearest to the query point and perform trilinear interpolation.
[0066] Human head animation generation. The trained Gaussian mixture shapes can be used to synthesize new images and animations under given perspectives and expressions. The user only needs to provide the human head motion parameters: joint and posture parameters Θ and expression coefficients {ψ k} and camera parameters C, the Gaussian model of each frame can be generated Gaussian splattering is then used to create highly realistic images and animations, as shown in Figure 1-4 (D). The parameters can be manually edited by the user or acquired from any human head video using a face tracker.
[0067] Implementation Examples
[0068] The inventors used a Nikon D850 SLR camera mounted on a tripod to capture data in an indoor environment, capturing a 3-minute video of a human head at 25 frames per second at a 1080p resolution. The subjects were asked to emotionally read a text while moving their heads within an appropriate range to capture images from various perspectives required for modeling. The inventors implemented the image preprocessing and model training of the present invention on a server equipped with an AMD EPYC 7B12 CPU and an NVidia A800 GPU, and implemented the real-time animation generation of the present invention on a desktop computer equipped with an Intel Core i7-13700KF CPU and an NVidia 4090 GPU. The inventors used all the parameter values listed in the specific implementation plan and obtained all the experimental results shown in Figures 1-4. The present invention can smoothly draw the subject's head animation at new perspectives and expressions at a frame rate of 370 frames per second, and allows the user to rotate the viewing angle and edit facial expressions in real time. For a 512x512 resolution input training image, running the face tracker and image segmentation preprocessing on a single NVIDIA A800 GPU takes 12 hours, and training takes 25 minutes. At runtime, for each frame, it takes approximately 2 milliseconds to generate the Gaussian model and 0.7 milliseconds to render the Gaussian model.
[0069] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed in this application.
[0070] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.
Claims
1. A Gaussian mixture shape method suitable for dynamic modeling of the human head, characterized in that: The following steps are involved: (1) Training image acquisition and preprocessing: Use a monocular camera to shoot a short video of a person speaking, process the video data and obtain: a neutral expression mesh, a set of base expression meshes, a head foreground image of each frame, a head foreground mask, camera parameters, joint and posture parameters, and expression coefficients; (2) Gaussian mixture shape training: using step (1), a head Gaussian model and an oral Gaussian model are established and their parameters are initialized, and all parameters are optimized according to the image loss; the head Gaussian model is composed of a neutral Gaussian model and a base expression Gaussian model; (3) Human head animation generation: Based on the human head motion parameters provided by the user, the neutral Gaussian model and the base expression Gaussian model in the Gaussian mixture shape optimized in step (2) are linearly mixed and combined with the oral Gaussian model to generate a Gaussian model corresponding to the motion parameters, and the head animation under the new perspective and expression is generated in real time.
2. The Gaussian mixture shape method for dynamic modeling of a human head according to claim 1, characterized in that: The training image acquisition and preprocessing in step (1) includes the following sub-steps: (1.1) Training image acquisition: Use a monocular camera to shoot a head video as training data. The ambient lighting should remain constant during the shooting process, and the subject should maintain normal speaking movements. (1.2) Training image preprocessing: Extract consecutive frames from the video to obtain a continuous image sequence, perform appropriate cropping and scaling on the image sequence, and use a face tracker to obtain a neutral expression grid, a set of base expression grids, camera parameters, joint and posture parameters, and expression coefficients for each frame from the image sequence; use a segmentation algorithm to mask out the head area in the image to obtain a head foreground image and a head foreground mask.
3. The Gaussian mixture shape method for dynamic modeling of a human head according to claim 1, characterized in that: The training of the Gaussian mixture shape in step (2) includes the following sub-steps: (2.1) Parameter initialization: The neutral expression mesh is converted into a neutral Gaussian model through sampling; the deformation gradient is extracted from the base expression mesh and applied to the neutral Gaussian model to generate the base expression Gaussian model; a mesh slice representing the teeth is generated at the corresponding position of the mouth in the neutral expression mesh, and the mesh slice representing the teeth is converted into the oral Gaussian model through sampling; (2.2) Gaussian mixture shape consistency variable substitution: Introduce intermediate variables, replace the Gaussian attribute increment of the base expression Gaussian model with the expression of the intermediate variable, and convert the optimized Gaussian attribute increment into the optimized intermediate variable; The substitution represents the Gaussian attribute increment as the sum of the product of the initialization result and the intermediate variable in step (2.1) and a heuristic function, wherein the heuristic function is a linear function with a threshold cutoff, and the input of the heuristic function is the displacement amplitude of the position deformation of the neutral expression grid of the Gaussian nearest neighbor to the base expression grid; (2.3) Pre-calculating the 3D grid: Preset the 3D grid, search for the nearest neighbor triangle on the neutral expression grid from each grid point, and store the corresponding base expression grid displacement amplitude and linear blend skin weight on the grid point; (2.4) Define the loss function: The loss function includes image loss, opacity loss and oral regularization constraint. The mathematical expression is as follows: L=λ1L rgb +λ2L α +λ3L reg ; L rgb =(1-λ)L1+λL D-SSIM ; Among them, L represents the loss function, L rgb represents the image loss, L α Indicates opacity loss, L reg represents the oral regularization constraint; λ1, λ2, λ3, λ are the loss mixture weights; L1 represents the mean absolute error, L D-SSIM Indicates structural differences; represents the opacity image, represents the head foreground mask, M is the total number of image pixels; SDF represents the signed distance, V represents the predefined oral volume, x i represents the position of the oral Gaussian, and N′ represents the total number of oral Gaussians.
4. The Gaussian mixture shape method for dynamic modeling of a human head according to claim 3, characterized in that: The step (2.2) is specifically as follows: for each Gaussian G i , let ΔG i,k Represents the base expression Gaussian model B k The "increment" of the Gaussian property relative to the neutral Gaussian model B0, that is, ΔG i,k =G i,k -G i,0 ; By introducing intermediate variables ΔG i,k Expressed as: in is the ΔG calculated in the initialization phase described in Section 2.1 i,k The initial value of is considered as a constant during optimization; |d i,k | is the closest to G i The surface position of the M0 mesh is deformed from M0 to M k The displacement amplitude that occurs, that is, the displacement amplitude of the base expression grid; is the intermediate parameter to be optimized, and its initial value is set to all 0; the linear function f(·) transforms M k The maximum amplitude between M0 and M1 is mapped to 1, and the amplitude threshold 0.00001 is mapped to 0; and ΔG is not directly optimized. i,k , but optimize After the optimization is completed, Substitute into the formula and calculate ΔG i,k , resulting in an expression blend shape.
5. The Gaussian mixture shape method suitable for dynamic modeling of a human head according to claim 1, characterized in that: The Gaussian mixture shape in step (2) is composed of a group of Gaussian models, wherein the group of Gaussian models includes a head Gaussian model and an oral Gaussian model. The head Gaussian model is used to generate a human head model representing arbitrary posture and arbitrary expression by linearly mixing a neutral Gaussian model and a base expression Gaussian model, and performing a linear mixed skinning transformation on the mixed model and the oral Gaussian model.
6. The Gaussian mixture shape method for dynamic modeling of a human head according to claim 1, characterized in that: The human head motion parameters in step (3) include joint and posture parameters, facial expression coefficients and camera parameters.
7. The Gaussian mixture shape method for dynamic modeling of a human head according to claim 1, characterized in that: The optimized Gaussian mixed shape expression in step (3) and the human head motion parameters given by the user, the Gaussian model in the linear mixed Gaussian mixed shape generates a Gaussian model corresponding to the motion parameters, and uses Gaussian splattering technology to draw head images and animations under new perspectives and expressions in real time.