Dynamic human modeling method based on three-dimensional gaussians

Through a dynamic human body modeling method based on three-dimensional Gaussian, a multi-view video sequence is used to learn three-dimensional Gaussian expression, and combined with a multi-layer perceptron and Gaussian splashing method for rendering, the problem of slow rendering speed and jitter in the existing technology of new human body postures and new perspectives is solved, real-time and high-quality virtual human body rendering is achieved.

WO2025118224A1PCT designated stage expired Publication Date: 2025-06-12ZHEJIANG UNIV

Patent Information

Application Number
PCT/CN2023/137041
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

When rendering dynamic human bodies, it is difficult to achieve rapid rendering of new human poses and high-quality perspective transformation. The training and rendering speed based on implicit neural fields is slow, and the generated video results have jitter problems.

Method used

The dynamic human body modeling method based on three-dimensional Gaussian is adopted to learn three-dimensional Gaussian expression through multi-view human body action video sequences, and Gaussian correction and transformation are used to perform Gaussian correction and transformation, and real-time rendering is achieved in combination with the Gaussian splashing method.

Benefits of technology

It realizes real-time rendering of highly realistic virtual human bodies at any perspective, including the pleated effect of clothes, and achieves real-time rendering speed in new perspectives and new body posture synthesis tasks, avoiding jitter problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023137041_12062025_PF_FP_ABST
    Figure CN2023137041_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a dynamic human modeling method based on three-dimensional Gaussians. In the present method, a multi-view human video is given firstly, and in order to correct and transform the dynamic human appearance, a latent code and a set of blend weights are added to each Gaussian; and the latent code and a target pose are inputted to an MLP to correct the Gaussian in a canonical space, so as to capture an appearance change under the target pose. The corrected Gaussian is transformed into the target pose by means of the blend weights thereof in a linear blend skinning (LBS) manner. Finally, a realistic human image at a new viewing angle or a new pose can be rendered in real time by means of Gaussian splatting. Compared with an existing method based on an implicit neural radiance field (NeRF), the three-dimensional Gaussian representation of the present invention can better capture high-frequency details while achieving optimal rendering performance.
Need to check novelty before this filing date? Find Prior Art

Description

A dynamic human body modeling method based on three-dimensional Gaussian Technical Field

[0001] The present invention relates to the field of dynamic virtual human body modeling, and in particular to a dynamic human body modeling method based on three-dimensional Gaussian. Background Art

[0002] Virtual humans have always been a hot topic of research. The main problem to be solved is to realistically render animated human bodies, including clothing and other accessories, given human pose parameters and camera observation angles. Some methods attempt to construct statistically based 3D mesh models (Hanbyul Joo, Tomas Simon, and Yaser Sheikh. Total capture: A 3D deformation model for tracking faces, hands, and bodies. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8320–8329, 2018.) to model human body geometry. To colorize the human body, traditional methods obtain texture and material information by scanning the human body (Brett Allen, Brian Curless, and Zoran Popovic. The space′ of human body shapes: reconstruction and parameterization from range scans. ACM transactions on graphics (TOG), 22(3): 587–594, 2003.). For highly deformable parts, such as loose clothing, methods such as physics-based simulation, blending within a database, or modeling in a deformable space have improved rendering fidelity to a certain extent. Over the past three years, numerous approaches have utilized neural representations to render dynamic scenes or virtual humans, including voxel-based, neural texture, and implicit neural field approaches. Animatable NeRF uses the SMPL (Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multiperson linear model. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 851–866. 2023.) model to establish point correspondences between arbitrary poses and static poses. It then learns a pose-dependent latent code for each frame to model pose-dependent variations. However, for any new pose, the latent code must be relearned and optimized, which reduces the model's generalization ability. To model more local details, a complete dynamic human figure is constructed using many small, local implicit neural fields.However, these implicit neural field-based methods are slow in both training and rendering speed. (Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. Fourier plenoctrees for dynamic radiance field rendering in real-time. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 13524–13534, 2022) Dynamic scenes are modeled with a set of dynamic octrees and compressed in the time domain using Fourier transforms, which achieves a rendering speed of over 100 frames per second. However, it can only replay the video and cannot render new human poses. Only neural rendering is used to quickly render the UV map of the human body, and some textures are generated using the given target human pose, and finally texture mapping is used to obtain the final result. This method can render high-quality images containing high-frequency details, and can render new perspective images at real-time rates. However, since the generated texture cannot remain stable in the time domain, the generated video results have serious jitter problems.

[0003] Summary of the Invention

[0004] To address the shortcomings of existing methods, this paper designs a dynamic human body modeling method based on 3D Gaussian. This method uses multi-view human motion video sequences to learn a 3D Gaussian representation. Given the viewpoint and body posture parameters, it can render realistic virtual humans in real time.

[0005] The specific technical solutions of the present invention are as follows:

[0006] The present invention provides a three-dimensional Gaussian-based dynamic human body modeling method, comprising the following steps:

[0007] (1) Establishing a 3D Gaussian representation of the human body: Given a multi-view human body video, first learn several 3D Gaussians at a static human body posture to represent the average human body model in all postures; each 3D Gaussian includes five basic parameters: position, opacity, rotation, scale and a set of spherical harmonic coefficients; and each Gaussian also contains two additional parameters: a latent code and a set of mixing weights; the latent code is used as an embedding vector of the appearance residual related to the human body posture, and the mixing weights are used for linear mixing skinning to transform the Gaussian from the normative space to the target posture space; the spherical harmonic coefficients are the colors of the corresponding Gaussian in different viewing directions;

[0008] (2) Gaussian correction: For a given target human posture, a multi-layer perceptron is used to correct the Gaussian established in step (1);

[0009] (3) Gaussian transformation: Using the given target human body posture, calculate the transformation matrix of each joint point of the human body; for the modified Gaussian obtained in step (2), use its mixing weight combined with the linear mixed skinning method to mix the transformation matrix of the joint point to obtain the required transformation matrix, decompose the required transformation matrix into three transformations of rotation, displacement and scaling, and then apply rotation, displacement and scaling transformation to the modified Gaussian to transform it to the target human body posture;

[0010] (4) Parameter optimization: The Gaussian transformed in step (3) is splashed using the Gaussian splashing method to obtain a rendering image. Then, all the parameters in step (1), the multi-layer perceptron in step (2), and the human posture parameters are jointly optimized using the image loss function, contour loss, and mixed weight loss to obtain the optimized three-dimensional Gaussian model of the human body.

[0011] (5) Real-time human animation rendering at any viewpoint: Given the viewpoint to be rendered and the target human posture, the optimized three-dimensional Gaussian model obtained in step (4) is used to obtain the final human rendering image using Gaussian splattering.

[0012] Furthermore, the three-dimensional Gaussian expression is specifically as follows: using a learnable latent code attached to each Gaussian, inputting it into a multi-layer perceptron, modifying the Gaussian in the canonical space to reflect the changes brought about by the human body posture, and then using a set of mixing weights on each Gaussian to transform the Gaussian to the target posture space through a linear mixed skinning method, and finally using Gaussian splashing for real-time rendering.

[0013] Furthermore, the step (1) mainly includes the following sub-steps:

[0014] (1.1) Processing a multi-view video sequence to obtain camera parameters, human body mask parameters, and human body pose parameters; the human body pose parameters are expressed using a multi-person linear skin model, which includes 6890 vertices and 24 joint positions;

[0015] (1.2) In the static posture of the human body, the vertex position of the SMPL model is used to initialize the Gaussian position; the zero-order spherical harmonic coefficient of each Gaussian is initialized to pure white color, and the high-order is initialized to 0; the hidden code is initialized using the position code; the mixing weight is initialized using the mixing weight obtained by projecting the point position onto the surface of the SMPL model and performing centroid interpolation.

[0016] Furthermore, the step (2) mainly includes the following sub-steps:

[0017] (2.1) For a given target human pose, the human pose parameters (rotation of each joint) and the latent code on the Gaussian are input into a 4-layer MLP with a width of 128 per layer, and the correction amount to the basic properties of the Gaussian is output;

[0018] (2.2) Use the corrections from step (2.1) to correct the opacity, scaling, rotation, and zero-order spherical harmonic coefficients of the Gaussian.

[0019] Furthermore, the step (3) mainly includes the following sub-steps:

[0020] (3.1) Based on the modified Gaussian, use an MLP to modify the mixture weight of each Gaussian and save the modified mixture weight;

[0021] (3.2) Using the linear blending skinning method, first use the position and rotation of the joint point (human body posture) to obtain the transformation matrix of each joint, and use the modified blending weight to calculate the transformation matrix required for each Gaussian;

[0022] (3.3) Decompose the transformation matrix of each Gaussian into the transformation of rotation, scaling, and displacement matrices, apply these transformation matrices to the four parameters of the Gaussian position, scaling, rotation, and spherical harmonic coefficients, and transform the Gaussian to the target human posture space.

[0023] Furthermore, the step (4) mainly includes the following sub-steps:

[0024] (4.1) Based on the transformed Gaussian, a rendering image is obtained using the Gaussian splashing method;

[0025] (4.2) To optimize the parameters, the image loss is calculated based on the rendered image obtained in step (4.1) and the true value, including the mean absolute error loss L1 and the structural similarity loss D-SSIM. The contour loss and the mixed weight loss are introduced. The contour loss restricts the Gaussian position to the human body and improves the simulation ability of clothing movement. The mixed weight loss ensures that the mixed weights of all points in the corresponding area of ​​a Gaussian are close, so that the Gaussian is transformed as a whole.

[0026] (4.3) A two-stage training process, the initial stage and the fine-tuning stage, is used to prevent the optimization from falling into overfitting. The iterations in the initial stage only optimize the basic Gaussian properties and human posture parameters, that is, the parts that are irrelevant to the human posture are optimized, providing a good basis for the fine-tuning stage. In the fine-tuning stage, the MLP is enabled, and the parameters optimized in the first stage are continued to be optimized, thereby fitting the details of the changes related to the human posture. Finally, the optimized three-dimensional Gaussian model is obtained.

[0027] Furthermore, the step (5) mainly includes the following sub-steps:

[0028] (5.1) Based on the obtained optimized 3D Gaussian model and MLP internal neuron parameters, for a given action sequence and viewpoint, the transformed Gaussian is used and rendered using the Gaussian splattering method;

[0029] (5.2) When only the viewing angle is changed, save the corrected and transformed Gaussian and render it directly when the viewing angle is adjusted.

[0030] Furthermore, if the human body posture is kept unchanged and only the viewing angle is changed in step (5), the multi-layer perceptron output of step (2) is saved to further accelerate the rendering.

[0031] Rendering can obtain the most optimized human body modeling.

[0032] The beneficial effects of the present invention are as follows:

[0033] Using a trained 3D Gaussian model from multi-view video sequences, a given target human pose can be rendered from any viewpoint to produce highly realistic images or videos of the human body, including high-frequency details such as clothing wrinkles. Due to the efficiency of this invention, both new viewpoint synthesis and new human pose synthesis can be achieved at real-time speeds, allowing for real-time observation of the virtual human from various viewpoints and the ability to manipulate the virtual human body into new movements. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] FIG1 is a flow chart of the present invention;

[0035] FIG2 is a Gaussian rendering result diagram of the present invention in a static posture (in a standard space);

[0036] Figure 3 is a rendering result after Gaussian correction with given human body posture parameters;

[0037] FIG4 is a diagram showing the final result of transforming the Gaussian after obtaining the transformation matrix using the linear mixed skinning method and the Gaussian mixing weight. DETAILED DESCRIPTION

[0038] The core of the present invention is to use multi-view video sequences to train a three-dimensional Gaussian model. For a given target human posture, highly realistic human body pictures or videos with high-frequency details, including clothing wrinkle effects, can be rendered in real time at any perspective.

[0039] As shown in Figure 1, the first stage of the present invention: three-dimensional Gaussian expression

[0040] The present invention learns a set of three-dimensional Gaussian expressions {G1, G2, .., G n}, each Gaussian is associated with a location x i , opacity α i , Zoom i , rotate r i With a set of spherical harmonic coefficients SH i And other basic properties, and are also associated with the learnable hidden code f i With the mixing weight w i The basic properties of Gaussian represent the average human appearance in various human postures, as shown in Figure 2. The hidden code f i , is embedded as the posture-related appearance residual and input into a multi-layer perceptron together with the target human posture parameters to modify the canonical space to reflect the appearance changes in this posture.

[0041] Specifically, given the latent code f i and the target human pose ( represents the rotation of K joints), and the Gaussian correction model is F a , then: {Δα i ,Δs i ,Δr i ,ΔSH i0}=F a (Θ,f i )

[0042] This modifies the basic properties of Gauss:

[0043] The result after Gaussian correction can be seen in Figure 3. After Gaussian correction, the Gaussian needs to be transformed into the target human pose space. This is done using the Skinned Multi-Person Linear (SMPL) model. Specifically, the multi-person linear skinning model uses a skeletal model. Given the joint rotation parameters Θ of the target pose, the transformation matrix P corresponding to each joint can be calculated using the joint point position J. k , for each Gaussian, use its corresponding mixing weight w i ={w i1 ,w i2 ,…,w ik}, the corresponding transformation matrix can be calculated:

[0044] During the training phase, for each Gaussian mixing weight, the nearest point to that point is found on the SMPL model grid in the static pose in the canonical space, and the mixing weight of that point is interpolated using the mixing weights stored on the model vertices through centroid interpolation, and this result is used as the Gaussian mixing weight. Since the position of the Gaussian is constantly being optimized and changed during training, in order to improve the efficiency of training, the mixing weight can be calculated in advance at the grid position in space before the training starts. Then, when the Gaussian position is given, it is only necessary to interpolate using the mixing weights of the grid point. The interpolated mixing weight is recorded as In order to further obtain more accurate mixing weights, an MLP is used to predict the residual of the mixing weights and correct them, namely F Δw :x→Δw(x), the final mixing weight is After training, the blending weights are saved on each Gaussian to speed up rendering.

[0045] In order to transform the Gaussian to the current human posture space, the transformation matrix P i Will be decomposed into the scaling matrix S i , the rotation matrix R i and the displacement matrix T i , shear deformation is ignored in this step, and the Gaussian will be transformed as follows:

[0046] Where ⊙ represents the multiplication of corresponding elements, SH_Rotation represents the rotation of the spherical harmonic function, and a new set of spherical harmonic coefficients can be obtained after rotation. For spherical harmonic rotations, if direct calculations based on the spherical harmonic rotation formula are too inefficient, a better approach is to inversely rotate the viewing direction. Specifically, first calculate the direction vector from the camera center to the Gaussian center as the viewing direction, then apply the inverse of the rotation matrix R to this direction. The resulting direction vector is used to calculate the spherical harmonic function values.

[0047] Finally, for the transformed Gaussian We apply the Gaussian splash method and render under given camera parameters to obtain the final image I render , as shown in Figure 4.

[0048] The second stage of the present invention: model training and loss function

[0049] In order to train the 3D Gaussian model, the camera parameters and human body posture parameters provided in the dataset are used to render the image, and the loss is calculated with the true value provided in the dataset to update the basic parameters, hidden code and F of the Gaussian. a With F Δw Two MLP parameters. Since the human pose parameters provided by the dataset are not accurate enough, the human pose parameters Θ,J are also optimized during training.

[0050] The first loss function is the image loss function, which is the L1 distance between the rendered image and the real image. In order to avoid jagged edges and to avoid generating a large number of small Gaussians to fit the hard image boundaries, a boundary mask M is used. b (Except for the 5 pixels near the human body boundary, the rest are 1) multiply the true value and the rendering image respectively, and then calculate the L1 distance as the loss value L rgb :

[0051] Where F is the number of images in a batch, and ⊙ represents the multiplication of corresponding elements; For the rendering, is the true value graph.

[0052] In order to avoid the influence of the background, the contour loss function is used to limit the position of the Gaussian to the range of the human body. The specific method is: set all the colors of the Gaussian to white and then render once to obtain the cumulative opacity map I opacity , and calculate the L2 loss L with the human mask provided by the dataset α :

[0053] In the process of transforming the Gaussian to the target human posture, since the Gaussian is not a point but covers a range, a necessary assumption is that the mixing weights of the points within the Gaussian coverage range must be consistent, so that one Gaussian can be transformed using one transformation matrix. If this requirement cannot be met and the transformation is still performed using one transformation matrix, problems such as the Gaussian protruding from the joint after the transformation will often occur. In order to reduce the error after the transformation, the mixing weight loss is used to encourage the mixing weights of the points inside a Gaussian to be basically consistent. Specifically, for a Gaussian G i , which includes three scaling components s = {s1, s2, s3} corresponding to the three axis directions {a1, a2, a3}, and six points can be obtained using these three axis directions From this, the mixed weights of these six points can be calculated The mixture weight loss is defined as the standard deviation L of the six mixture weights w :

[0054] where n is the total number of Gaussians.

[0055] D-SSIM structural similarity loss is also introduced to encourage the rendered image to be close to the ground truth in terms of metrics such as brightness and darkness contrast.

[0056] The final loss function is defined as: L = λ1L rgb +λ2L α +λ3L w +λ4L D-SSIM ;

[0057] Among them, λ1=0.8, λ2=10, λ3=0.2, λ4=0.2 are the weights corresponding to the four loss functions respectively.

[0058] In addition, we look for it in the present invention, model F a ,F Δw They are all four-layer MLPs, using ReLU as the activation function, F a ,F Δw The hidden layer widths are 128 and 32 respectively, the dimension of the hidden code is 9, and the position code is used as its initialization value.

[0059] The training is divided into two stages, including the initial stage (the first 5000 iterations) and the fine-tuning stage. In the initialization stage, the MLP F a ,F Δw When disabled, only basic Gaussian properties and human pose parameters are optimized, which is equivalent to optimizing the average human model in various poses. During the fine-tuning phase, both MLPs are enabled, the latent code is also optimized, and the optimization of basic Gaussian properties and human pose parameters is continued.

[0060] To prevent the optimization from falling into a local optimum, the opacity of the Gaussian is reset every certain number of rounds, starting with the 3000th iteration in the initial stage and then every 6000 iterations thereafter. For changes in the number of Gaussians (adding or removing Gaussians), the gradient of the image loss and contour loss backpropagated to each Gaussian is used as the criterion for adding Gaussians. If the gradient exceeds a certain threshold, a new Gaussian is added. The basic Gaussian properties, rather than the modified Gaussian properties, are used to determine whether to split or remove Gaussians, ensuring their consistency across all human poses.

[0061] Example

[0062] The inventors implemented an embodiment of the present invention on a desktop computer equipped with an i7-13700KF CPU, 32GB of RAM, and an NVIDIA RTX 4090 graphics processor. Training converged within 80 minutes. On the novel perspective synthesis task, the method achieved an average rendering speed of 114 frames per second, and on humans synthesized in novel poses, it achieved a rendering speed of 66 frames per second, fully meeting the requirements of real-time rendering.

[0063] The inventors tested this method on various datasets. They demonstrated that it can achieve high-quality, realistic results when synthesizing human images from novel perspectives or poses, including muscle changes associated with movement and the deformation and wrinkling of clothing with posture. Compared to previous methods, this method not only offers advantages in image quality, accurately capturing the subtle changes in human movement, but also produces temporally stable, jitter-free video results. Furthermore, rendering can reach real-time speeds, supporting real-time observation of human movement from any perspective.

[0064] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed in this application.

[0065] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. A dynamic human body modeling method based on three-dimensional Gaussian, characterized in that, it includes the following steps: (1) Establish a three-dimensional Gaussian representation of the human body: Given a multi-view human body video, first learn several three-dimensional Gaussians in a static human body posture to represent the average human body model in all postures; Each three-dimensional Gaussian includes five basic parameters: position, opacity, rotation, scaling, and a set of spherical harmonic coefficients; And each Gaussian also contains two additional parameters: a latent code and a set of mixing weights; The latent code is used as an appearance residual embedding vector related to the human body posture, and the mixing weights are used to linearly mix the skinning to transform the Gaussian from the canonical space to the target posture space; The spherical harmonic coefficients are the colors of the corresponding Gaussian in different viewing directions. (2) Gaussian correction: For a given target human body posture, use a multi-layer perceptron to correct the Gaussian established in step (1). (3) Gaussian transformation: Use the given target human body posture to calculate the transformation matrix of each joint point of the human body; For the corrected Gaussian obtained in step (2), use its mixing weights and the linear skinning method to mix the required transformation matrix by the transformation matrix of the joint points, decompose the required transformation matrix into three transformations of rotation, displacement, and scaling, and then apply rotation, displacement, and scaling transformations to the corrected Gaussian to transform it to the target human body posture. (4) Parameter optimization: Use the Gaussian splashing method to splash the Gaussian transformed in step (3) to obtain a rendering image, and then use the image loss function, contour loss, and mixing weight loss to jointly optimize all the parameters in step (1), the multi-layer perceptron in step (2), and the human body posture parameters to obtain an optimized three-dimensional Gaussian model of the human body. (5) Real-time human body animation rendering from arbitrary viewpoints: Given the viewpoint to be rendered and the target human body posture, the optimized three-dimensional Gaussian model obtained in step (4) uses Gaussian splashing to obtain the final human body rendering image.

2. The dynamic human body modeling method based on three-dimensional Gaussian according to claim 1, characterized in that, the three-dimensional Gaussian representation is specifically: using the learnable latent code attached to each Gaussian, input it into a multi-layer perceptron to correct the Gaussian in the canonical space to reflect the changes brought by the human body posture, and then use a set of mixing weights on each Gaussian to transform the Gaussian to the target posture space by the method of linear skinning, and finally use Gaussian splashing for real-time rendering.

3. The dynamic human body modeling method based on three-dimensional Gaussian according to claim 1, characterized in that, step (1) mainly includes the following sub-steps: (1.1) Process the multi-view video sequence to obtain camera parameters, human body mask parameters, and human body posture parameters; The human body posture parameters are expressed using a multi-person linear skinning model, which includes 6890 vertices and 24 joint point positions. (1.2) Under the static human body posture, initialize the Gaussian position using the vertex positions of the SMPL model; the zero - order of the spherical harmonic coefficients of each Gaussian is initialized to pure white color, and the higher - order ones are initialized to 0; the latent code is initialized using positional encoding; the mixing weights are initialized using the mixing weights obtained by projecting the point positions onto the surface of the SMPL model and performing barycentric interpolation.

4. The dynamic human body modeling method based on three - dimensional Gaussian according to claim 1, wherein, the step (2) mainly includes the following sub - steps: (2.1) For a given target human body posture, input the human body posture parameters (rotations of each joint point) and the latent code on the Gaussian into a 4 - layer MLP with a width of 128 for each layer, and output the correction amount for the basic attributes of the Gaussian. (2.2) Use the correction amount in step (2.1) to correct the opacity, scaling, rotation, and zero - order spherical harmonic coefficients of the Gaussian.

5. The dynamic human body modeling method based on three - dimensional Gaussian according to claim 1, wherein, the step (3) mainly includes the following sub - steps: (3.1) Based on the corrected Gaussian, use an MLP to correct the mixing weights of each Gaussian, and save the corrected mixing weights. (3.2) Using the linear blend skinning method, first obtain the transformation matrix of each joint using the positions and rotations of the joint points (human body posture), and use the corrected mixing weights to calculate the transformation matrix required for each Gaussian. (3.3) Decompose the transformation matrix of each Gaussian into transformation matrices of rotation, scaling, and displacement, and apply these transformation matrices to the four parameters of the Gaussian's position, scaling, rotation, and spherical harmonic coefficients to transform the Gaussian into the target human body posture space.

6. The dynamic human body modeling method based on three - dimensional Gaussian according to claim 1, wherein, the step (4) mainly includes the following sub - steps: (4.1) Based on the obtained transformed Gaussian, use the Gaussian splashing method to obtain the rendered image. (4.2) To optimize the parameters, calculate the image loss based on the rendered image obtained in step (4.1) and the ground truth, including two parts of loss: the mean absolute error loss L1 and the structural similarity loss D - SSIM; and introduce the contour loss and the mixing weight loss. The contour loss restricts the Gaussian positions within the human body range and enhances the simulation ability of clothing movement. The mixing weight loss ensures that the mixing weights of all points within the corresponding area of a Gaussian are close, causing the Gaussian to be transformed as a whole. (4.3) Use two - stage training in the initial stage and the fine - tuning stage to prevent overfitting in optimization; in the initial stage of iteration, only optimize the basic Gaussian attributes and human body posture parameters, that is, optimize the parts unrelated to the human body posture, providing a good basis point for the fine - tuning stage. In the fine - tuning stage, enable the MLP and continue to optimize while maintaining the parameters optimized in the first stage, so as to fit the details related to the changes in the human body posture; finally, obtain the optimized three - dimensional Gaussian model.

7. The dynamic human body modeling method based on three - dimensional Gaussian according to claim 1, wherein, the step (5) mainly includes the following sub - steps: (5.1) Based on the obtained optimized three-dimensional Gaussian model and the internal neuron parameters of the MLP, for a given action sequence and viewing angle, use the obtained transformed Gaussian and perform rendering using the Gaussian splashing method; (5.2) When only the viewing angle is changed, save the corrected and transformed Gaussian, and directly perform rendering when the viewing angle is adjusted.

8. The dynamic human body modeling method based on three-dimensional Gaussian according to claim 1, characterized in that, if the human body posture remains unchanged in step (5) and only the viewing angle is changed, then the output of the multi-layer perceptron in step (2) is saved to further accelerate rendering, and the optimized human body modeling can be obtained.

Citation Information

Patent Citations

  • Color image processing method and device based on three-dimensional Gaussian cloud transformation

    CN103390276A

  • Dynamic human body three-dimensional reconstruction and visual angle synthesis method

    CN112465955A

  • Drivable implicit three-dimensional human body representation method

    CN113112592A

  • Human body posture prediction method based on Gaussian process regression and progressive filtering

    CN115050095A

  • Methods and apparatus for orientation keypoints for complete 3D human pose computerized estimation

    US20210097718A1

Cited By

  • Personnel positioning method and system based on 3D Gaussian splash model and video fusion

    CN120318327A

  • Three-dimensional Gaussian visual positioning method for sparse visual angle

    CN120339380A

  • Multi-modal 4D content generation method and system based on alignment

    CN120672972A

  • Modeling and rendering method and device based on layered three-dimensional Gaussian representation and medium

    CN120747311A

  • A modeling and rendering method, device and medium based on layered three-dimensional Gaussian representation

    CN120747311B