3D human avatar method and system capable of being driven in real time and generated based on 3d gaussian patterning

The method addresses the challenges of high training data requirements and slow rendering by using adaptive Gaussian ellipsoids and improved deformation models for real-time 3D human avatar generation, achieving reduced costs and high-fidelity rendering.

CN120318382APending Publication Date: 2025-07-15NANJING YINGQI INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510239129.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When generating high-conservation real-time 3D human avatars, the prior art requires a large amount of data set training, and it is difficult to achieve synchronous rendering of facial expressions and body movements, resulting in high training costs, poor rendering effects, and easy to produce artifacts when 3D observations are missing.

Method used

A small amount of human body data is collected based on a single camera, and the posture and expression information are predicted through the smplx and Flame models, and an adaptive 3D Gaussian ellipsoidal distribution is constructed. Combined with the improved 3d gaussian splatting model, a non-rigid deformation module and a timing color MLP correction module are added to optimize the Gaussian point distribution and rendering process.

Benefits of technology

It realizes low-cost and efficient 3D human avatar generation, improves the synchronous rendering quality of facial expression details and body movements, reduces artifacts, and improves the high fidelity and consistency of rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318382A_ABST
    Figure CN120318382A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D human avatar method and a 3D human avatar system which are generated based on 3d gaussian patterning and can be driven in real time. The 3D human avatar system comprises a data acquisition module, a Gaussian initialization module and a training rendering module, the data acquisition module acquires a small amount of human body data based on a single camera, predicts human body posture information based on an smplx model and expression information based on a Flame model from two-dimensional human body image data, and makes labels; a Gaussian initialization module carries out adaptive Gaussian initialization on a human body smplx model grid in any frame label to obtain a 3D Gaussian ellipsoid; and the training rendering module is used for training the 3D Gaussian ellipsoid initialized in the step 2 based on an improved 3d Gaussian sputtering model, and rendering a 3D human avatar capable of being driven in real time. According to the method, the topological consistency of the Gaussian point and the human body smplx model can be highly maintained, artifacts are reduced, and high fidelity is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method and system for a real-time drivable 3D human avatar generated based on 3D Gaussian splatting. Background Art

[0002] The task of creating realistic animated objects has always been of utmost importance in 3D computer vision. In recent years, due to the emergence of virtual and augmented reality devices, high-fidelity human avatars have been widely used in real-time applications, and digital human avatars are the key to online interaction between people. Traditional methods usually rely on dense, synchronized multi-view inputs, which may not be easily obtained in more practical scenarios. To achieve state-of-the-art rendering quality, existing methods rely on training neural radiance fields (NeRF) combined with explicit body joints or human-related encoding methods to achieve high-quality novel view / novel pose image synthesis, but usually require several days of training and are very slow during inference, making it almost impossible to achieve real-time interactive rendering frame rates.

[0003] Although existing Gaussian distribution-based methods can be driven in real-time, their common drawbacks are that they require a large amount of dataset training, and they require a large number of Gaussian distributions to represent a human avatar. For high-frequency details in animation, such as facial expressions and other regions, a large number of Gaussians may be required to look realistic enough. Previous methods required up to 200,000 or even 300,000 Gaussian points to train an avatar, increasing the training cost. And the lack of 3D observations may lead to significant ambiguities in body parts not observed in the video, which may result in obvious artifacts under new movements. Most 3D human avatars modeled from randomly captured videos only support body movements, without facial expressions and finger animations, and most datasets also do not have detailed facial expressions, reducing the high fidelity and driving effect of the character model. The pixel colors of the rendered human avatars also largely depend on local deformations. For example, local fine wrinkles on clothes can cause self-overlapping, seriously affecting shadows and deteriorating the rendering effect. Summary of the Invention

[0004] Objective of the Invention: The present invention provides a method and system for a real-time drivable 3D human avatar generated based on 3D Gaussian splatting, which can highly maintain the topological consistency between Gaussian points and the human SMPLX model, reduce artifacts, and improve high fidelity.

[0005] Technical Solution: A method for a real-time drivable 3D human avatar generated based on 3D Gaussian splatting according to the present invention includes the following steps:

[0006] Step 1: Based on a single camera to collect a small amount of human data, predict human pose information based on the SMPLX model and expression information based on the FLAME model from the two-dimensional human image data, and make labels;

[0007] Step 2: Perform adaptive Gaussian initialization on the human SMPLX model mesh in any frame of the label to obtain a 3D Gaussian ellipsoid;

[0008] Step 3: Based on the improved 3D Gaussian splatting model, train the 3D Gaussian ellipsoid initialized in Step 2, and render a real-time drivable 3D human avatar.

[0009] Furthermore, in Step 1, based on a single camera to collect a small amount of human data, predicting human pose information based on the SMPLX model and expression information based on the FLAME model from the two-dimensional human image data, and making labels specifically includes the following steps:

[0010] Step 11: Use a single camera to shoot human data; fix the camera position, let the person being photographed perform simple actions in the frame, rotate a full circle, and take pictures, and calculate the camera parameters and extract the frame of the picture;

[0011] Step 12: Predict SMPLX parameters and offsets; for each person in the photographed frame, use the SMPLX regression model to fit the SMPLX parameters of the human body, including shape, pose, root node, translation, and offsets of the patches;

[0012] Step 13: Predict FLAME parameters and offsets; for each person in the photographed frame, intercept the face part, and use the fine-tuned DECA face model to predict the expression parameters, the pose of the chin joint, and the offsets of the expression;

[0013] Step 14: Make labels for the predicted SMPLX, FLAME parameters and offsets to construct a dataset; replace the expression parameters and the pose of the chin joint in SMPLX with the expression parameters and the pose of the chin joint in FLAME, and add the offsets of the expression, and make these parameters and images correspond one by one as a label file, construct a training set, a validation set, and a test set, and randomly select 80% of the dataset as the training set, 10% of the dataset as the validation set, and 10% of the dataset as the test set.

[0014] Furthermore, in Step 2, performing adaptive Gaussian initialization on the human SMPLX model mesh in any frame of the label to obtain a 3D Gaussian ellipsoid specifically includes the following steps:

[0015] Step 21: Based on the gender of the person in the image, select a standard SMPLX model, load the selected label file, generate an SMPLX human model and obtain the patches of the human model;

[0016] Step 22: Construct an adaptive 3D Gaussian ellipsoid distribution model based on the patch size and the adjacent patch sizes.

[0017] Step 23: Perform an adaptive 3D Gaussian ellipsoid distribution on the patches of the smplx human model in Step 21 using the distribution model constructed in Step 22.

[0018] Further, in Step 22, when constructing the adaptive 3D Gaussian ellipsoid distribution model based on the patch size and the adjacent patch sizes, calculate the area size of the current patch, and calculate the patch sizes of the three adjacent faces. Set a threshold and the initial number of Gaussian ellipsoids for the patch. When the size multiple of the four patches exceeds the threshold, re - distribute the number of Gaussian ellipsoids according to the area weight, but each patch retains at least 1 Gaussian ellipsoid. After re - distribution, the large patch is re - divided equally according to the threshold, and Gaussian ellipsoids are distributed on the divided patches; if the size multiple of the current four patches is less than the threshold, perform the above operations on the adjacent patches. At the same time, to improve efficiency, directly perform a separate 3D Gaussian ellipsoid distribution on the patches that are particularly lower than the threshold until all patches are distributed. The distribution of 3D Gaussian ellipsoids by 3D gaussian splatting is random. Adaptive distribution of 3D Gaussian ellipsoids according to the patch size and the adjacent patch sizes can make the positions of 3D Gaussian ellipsoids relatively rich, have more Gaussian representations for each patch, optimize the distribution of Gaussian points, reduce the training time, and improve the training effect.

[0019] Further, in Step 23, the specific steps of performing an adaptive 3D Gaussian ellipsoid distribution on the patches of the smplx human model in Step 21 using the distribution model constructed in Step 22 are as follows:

[0020] Step 231: Read the patches of the smplx human model obtained in Step 21 and mark the indexes.

[0021] Step 232: Perform 3D Gaussian distribution on the patches of the smplx model using the adaptive 3D Gaussian ellipsoid distribution model.

[0022] Step 233: Bind the distributed 3D Gaussian ellipsoids to the patch indexes.

[0023] Further, in Step 3, based on the improved 3D gaussian splatting model, train the 3D Gaussian ellipsoids initialized in Step 2 and render a real - time drivable 3D human avatar. The specific steps are as follows:

[0024] Step 31: Construct an improved 3D Gaussian splatting model by improving the non-rigid deformation module, improving the loss function, and adding a temporal color MLP correction module; effectively enhancing the expression ability of Gaussian ellipsoids, reducing spikes and artifacts, and improving similarity, fidelity, and perceptual quality;

[0025] Step 32: Train the improved 3D Gaussian splatting model in Step 31 and render a human body model.

[0026] Furthermore, in Step 31, the non-rigid deformation module is improved: a multi-scale geometric information module is added to the model, and the position attributes of the input 3D Gaussian ellipsoids are encoded as spatial feature vectors using a multi-layer hash grid. The spatial feature vectors output by the hash grid are concatenated with the SMPLX pose features and shape parameter features, and the concatenated features are passed through a lightweight MLP network to learn the deformation offset and dynamically deform the initialized 3D Gaussian ellipsoids, which can accurately reflect the local deformations in different poses. The specific formula is as follows:

[0027]

[0028] Where δx represents the position offset, which is used to adjust the position offset of the center point of the 3D Gaussian ellipsoid; δd represents the normal offset, which controls the offset amount of the 3D Gaussian ellipsoid in the normal direction of the patch; δs represents the scale offset, which controls the size change of the 3D Gaussian ellipsoid; δq represents the rotation offset, which adjusts the rotation angle of the 3D Gaussian ellipsoid; z represents the local pose feature, which is used for the color module; represents the MLP decoding module, and γ(x cv ) represents the position encoding of the Gaussian ellipsoid position attribute x c ; S p represents the feature encoding of the pose parameters and shape parameters.

[0029] Furthermore, the loss function is improved: an additional geometric consistency regularization constraint is added to the deformation parameters to make the deformation result more physically reasonable; a point-to-patch distance loss is added, and the distance from the center point p of the Gaussian ellipsoid to the bound patch is calculated, and the minimum distance from the point to the triangular face is used to define it. The patch normal vector is n, and one of the vertices of the patch is v1, then the point-to-face distance loss is defined as:

[0030]

[0031] L distanc+ = Dist(p)

[0032] Add the projection loss of the point inside the patch to ensure that the projection of the center point of the Gaussian ellipsoid falls within the bound patch. Assume the center point of the Gaussian ellipsoid is p, the three vertices of the plane are v1, v2, v3, and any point on the plane is p , , the normal vector is n, and the projection on the plane is P p. / , and the specific formula is as follows:

[0033] p , = λ1v1 + λ # v # + λ3v3

[0034] λ1 + λ # + λ3 = 1

[0035]

[0036] Judge whether the barycentric coordinates (μ, ν, ω) of the projection point P p. / satisfy that all values are within [0, 1]. The specific formula is as follows:

[0037] v 81 = v1 - v #

[0038] v 8# = v # - v8

[0039] v 8p = P p. / - v8

[0040]

[0041] ω = 1 - μ - ν

[0042] If μ, ν, ω ≥ 0 and μ + ν + ω = 1, then the point P p. / is inside or on the boundary of the triangle; otherwise, the point is outside the triangle; the loss function is specifically as follows:

[0043]

[0044] Temporal color MLP correction module. Add the time parameter t to the color MLP module to incorporate the time dimension into the model, enabling it to correct color information that changes over time. This is very useful for color correction in video sequences or dynamic scenes because it allows the model to capture and learn the changing patterns of colors over time, helping to enhance the high fidelity of the rendering effect.

[0045] Furthermore, in step 32, train the improved 3D Gaussian splatting model in step 31 and render the human model, which specifically includes the following steps:

[0046] Step 321: Set the training parameters and use the Adam optimization algorithm for training. The initial learning rates are set as follows: for position Ir_position = 0.00016, opacity Ir_opacity = 0.05, scale Ir_scale = 0.005, rotation Ir_rot = 0.005, pose Ir_pose = 0.0002, and color Ir_color = 0.005. Set the final position learning rate to Ir_position = 0.0000016, and the number of training epochs Epoch = 10000.

[0047] Step 322: Feed the dataset constructed in Step 14 into the improved 3d gaussian splatting model in Step 31 for training and rendering.

[0048] Step 323: Based on the rendering effects of the model on the training set and the validation set, cross-validate the changes in the peak signal-to-noise ratio, structural similarity index, and perceptual image patch similarity metrics, as well as the loss change trend. Adjust the learning rate and the number of iterations until the changes in the metrics and the loss gradually tend to a stable state, and determine the final learning rate and the number of iterations.

[0049] Step 324: Based on the determined learning rate and the number of iterations, complete the training and rendering of the 3d gaussian splatting model to obtain a well-performing and real-time drivable human avatar.

[0050] Correspondingly, a real-time drivable 3D human avatar system generated based on 3d gaussian splatting includes: a data acquisition module, a Gaussian initialization module, and a training and rendering module. The data acquisition module collects a small amount of human body data based on a single camera, predicts the human body pose information based on the smplx model and the expression information based on the Flame model from the two-dimensional human body image data, and creates labels. The Gaussian initialization module adaptively initializes the Gaussian of the human smplx model mesh in any frame label to obtain a 3D Gaussian ellipsoid. The training and rendering module trains the 3D Gaussian ellipsoid initialized in Step 2 based on the improved 3d gaussian splatting model and renders a real-time drivable 3D human avatar.

[0051] Advantages: Compared with the prior art, the present invention has the following remarkable advantages: Only a small amount of data sets need to be captured for training, greatly reducing the time cost of making data sets, and specifically predicting expression labels for the face, increasing facial expression details; When initializing the Gaussian, the adaptive distribution Gaussian is used to effectively cover the human body area, which can effectively reduce the number of Gaussian points and effectively reduce the training cost; At the same time, the topological consistency between the Gaussian points and the human SMPLX model can be highly maintained, reducing artifacts and improving high fidelity; The improved non-rigid deformation module adds a multi-scale geometric information module to the model, uses a multi-layer hash grid to encode the position attributes of the input 3D Gaussian ellipsoid into spatial feature vectors, and splices the spatial feature vectors output by the hash grid with the SMPLX pose features and shape parameter features. The spliced features are used to learn the deformation offset through a lightweight MLP network to dynamically deform the initialized 3D Gaussian ellipsoid, and by improving the loss function, the 3D Gaussian ellipsoids are constrained to accurately reflect the local deformations in different poses, which helps to maintain the geometric consistency and real deformation of the avatar, especially in dynamic and changing poses; Predicting multiple offsets allows our human model to better render clothes and hair; The temporal color MLP correction module learns the latent lighting features of each frame to compensate for the different ambient light effects between frames caused by the overall movement. The real-time drivable 3D human avatar technology based on 3D Gaussian splatting combines the advantages of high fidelity, high efficiency, and low cost, providing a more feasible and reliable solution for the problems of 3D human avatar reconstruction and rendering, and providing a new 3D human avatar generation solution for the field of 3D Gaussian splatting reconstruction and rendering. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of the process for generating a real-time drivable 3D human avatar according to the present invention.

[0053] Figure 2 It is a schematic diagram of the process for making a human body data set according to the present invention.

[0054] Figure 3 It is an example diagram of a human body data set according to the present invention.

[0055] Figure 4 It is a schematic diagram of the effect of adaptive Gaussian initialization of 3D Gaussian ellipsoids according to the present invention.

[0056] Figure 5 It is an example diagram for comparing the rendering effects of real pictures and human body models according to the present invention.

[0057] Figure 6 It is an example diagram of the real-time driving effect according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] As Figure 1 shown, a real-time drivable 3D human avatar method based on 3D Gaussian splatting includes the following steps:

[0059] Step 1: Collect a small amount of human body data based on a single camera. As Figure 2 shown, predict the human body pose information based on the SMPLX model and the facial expression information based on the FLAME model from the two-dimensional human body image data, generate labels and create a dataset, which specifically includes the following steps:

[0060] Step 11: Take pictures of human body data with a single camera. The steps are as follows:

[0061] Step 111: Fix the position of the camera. The person to be photographed makes simple movements in the frame, rotates a circle, takes pictures, calculates the camera parameters, and extracts the frame of the picture;

[0062] Step 12: Predict the SMPLX parameters and offsets. The steps are as follows:

[0063] Step 121: For the person in each frame of the photographed picture, use the existing SMPLX regression model to fit the SMPLX parameters of the human body, including shape, pose, root node, translation, and the offset of the patch.

[0064] Step 13: Predict the FLAME parameters and offsets. The steps are as follows:

[0065] Step 131: For the person in each frame of the photographed picture, intercept the face part, and use the fine-tuned DECA face model to predict the expression parameters, the pose of the chin joint, and the offset of the expression.

[0066] Step 14: Make labels for the predicted SMPLX, FLAME parameters and offsets to construct a dataset. The steps are as follows:

[0067] Step 141: Replace the expression parameters and the pose of the chin joint of the SMPLX with the expression parameters and the pose of the chin joint in the FLAME, and add the offset of the expression. Corresponding these parameters and images one by one to make a label file. The example picture is as Figure 3 shown.

[0068] Step 142: Construct a training set, a validation set and a test set. Randomly select 80% of the dataset as the training set, 10% of the dataset as the validation set, and 10% of the dataset as the test set.

[0069] Step 2: Perform grid adaptive Gaussian initialization on the SMPLX model mesh of the human body in any frame label to obtain a 3D Gaussian ellipsoid. The specific steps are as follows:

[0070] Step 21: Based on the gender of the image person, select a standard SMPLX model, load the selected label file, generate an SMPLX human model, and obtain the patches of the human model.

[0071] Step 22: Construct an adaptive 3D Gaussian ellipsoid distribution algorithm based on the patch size and the size of adjacent patches. Generally, the distribution of 3D Gaussian splatting for 3D Gaussian ellipsoids is random. By adaptively distributing 3D Gaussian ellipsoids according to the patch size and the size of adjacent patches, the positions of 3D Gaussian ellipsoids can be relatively rich, there can be more Gaussian representations for each patch, which can optimize the distribution of Gaussian points, reduce the training time, and improve the training effect.

[0072] Adaptive distribution algorithm according to the patch size and the size of adjacent patches: Calculate the area size of the current patch, and calculate the patch sizes of the three adjacent faces. Set a threshold and the initial number of Gaussian ellipsoids for the patch. When the size multiples of the four patches exceed the threshold, reallocate the number of Gaussian ellipsoids according to the area weight, but each patch should retain at least 1 Gaussian ellipsoid. After reallocation, the large patch is re - equally divided according to the threshold, and Gaussian ellipsoids are distributed on the divided patches. If the size multiples of the current four patches are less than the threshold, perform the above operations on the adjacent patches. At the same time, to improve efficiency, directly perform a separate 3D Gaussian ellipsoid distribution on the patches that are particularly lower than the threshold until all patches are distributed. The effect is as Figure 4 shown.

[0073] Step 23: Perform an adaptive 3D Gaussian ellipsoid distribution on the patches of the SMPLX human model in Step 21 using the distribution algorithm constructed in Step 22. The specific steps are as follows:

[0074] Step 231: Read the patches of the SMPLX human model obtained in Step 21 and mark the indexes.

[0075] Step 232: Perform 3D Gaussian distribution on the patches of the SMPLX model using the adaptive 3D Gaussian ellipsoid distribution algorithm.

[0076] Step 233: Bind the distributed 3D Gaussian ellipsoids to the patch indexes.

[0077] Step 3: Based on the improved 3D Gaussian splatting model, train the 3D Gaussian ellipsoids initialized in Step 2 and render a real - time drivable 3D human avatar. The specific steps are as follows:

[0078] Step 31: Construct an improved 3D Gaussian splatting model: It is improved by improving the non-rigid deformation module, improving the loss function, and adding a temporal color MLP correction module, effectively enhancing the expressive ability of Gaussian ellipsoids, reducing spikes and artifacts, and improving similarity, fidelity, and perceptual quality;

[0079] Improve the non-rigid deformation module: Add a multi-scale geometric information module to the model. Use a multi-layer hash grid to encode the position attributes of the input 3D Gaussian ellipsoids into spatial feature vectors, and concatenate the spatial feature vectors output by the hash grid with the smplx pose features and shape parameter features. Pass the concatenated features through a lightweight MLP network to learn the deformation offset and perform dynamic deformation on the initialized 3D Gaussian ellipsoids so that they can accurately reflect the local deformations in different poses. The specific formula is as follows:

[0080]

[0081] Where δx represents the position offset, used to adjust the position offset of the center point of the 3D Gaussian ellipsoid; δd represents the normal offset, controlling the offset of the 3D Gaussian ellipsoid in the normal direction of the patch; δs represents the scale offset, controlling the size change of the 3D Gaussian ellipsoid; δq represents the rotation offset, adjusting the rotation angle of the 3D Gaussian ellipsoid; z represents the local pose feature, used for the color module; represents the MLP decoding module, γ(x c ) represents the position encoding of the Gaussian ellipsoid position attribute x c ; S p represents the feature encoding of the pose parameters and shape parameters.

[0082] Improve the loss function: Add an additional geometric consistency regularization constraint to the deformation parameters to make the deformation result more physically reasonable. Add the distance loss from the point to the patch, calculate the distance from the center point p of the Gaussian ellipsoid to the bound patch, and define it using the minimum distance from the point to the triangular face. The patch normal vector is n, and one of the vertices of the patch is v1, then the distance loss from the point to the face is defined as:

[0083]

[0084] L distance = Dist(p)

[0085] Add the projection loss of the point within the patch to ensure that the projection of the center point of the Gaussian ellipsoid falls within the bound patch. Assume the center point of the Gaussian ellipsoid is p, assume the three vertices of the plane are v1, v2, v3, and any point on the plane is p , , the normal vector is n, and the projection on the plane is P p. / , and the specific formula is as follows:

[0086] p , = λ1v1 + λ # v # + λ3v3

[0087] λ1 + λ # + λ3 = 1

[0088]

[0089] where 0 ≤ λ1, λ2, λ3 ≤ 1.

[0090] Determine whether all values of the barycentric coordinates (μ, ν, ω) of the projection point P p. / are within [0, 1]. The specific formula is as follows:

[0091] v 81 = v1 - v #

[0092] v 8# = v # - v8

[0093] v 8p = P p. / - v8

[0094]

[0095] ω = 1 - μ - ν

[0096] If μ, ν, ω ≥ 0 and μ + ν + ω = 1, then the point P p. / is inside or on the boundary of the triangle. Otherwise, the point is outside the triangle. The loss function is as follows:

[0097]

[0098] Temporal Color MLP Correction Module: Add the time parameter t to the Color MLP module to incorporate the time dimension into the model, enabling it to correct color information that changes over time. This is very useful for color correction in video sequences or dynamic scenes because it allows the model to capture and learn the changing patterns of colors over time, which helps to improve the high-fidelity of the rendering effect.

[0099] Step 32: Train the improved 3D Gaussian splatting model in Step 31 and render the human model. The specific steps are as follows:

[0100] Step 321: Set the training parameters and use the Adam optimization algorithm for training. The initial learning rates are set as follows: for position Ir_position = 0.00016, for opacity Ir_opacity = 0.05, for scale Ir_scale = 0.005, for rotation Ir_rot = 0.005, for pose Ir_pose = 0.0002, and for color Ir_color = 0.005. Set the final position learning rate to Ir_position = 0.0000016, and the number of training epochs Epoch = 10000;

[0101] Step 322: Feed the dataset constructed in Step 14 into the improved 3D Gaussian splatting model in Step 31 for training and rendering. The rendering effect diagram is as follows Figure 5 shown. The left side of the picture is an image from the training set, and the right side is the human body rendering model corresponding to this image;

[0102] Step 323: According to the rendering effects of the model on the training set and the validation set, cross-validate the changes in the peak signal-to-noise ratio, structural similarity index, and perceptual image patch similarity metrics, as well as the loss change trend. Adjust the learning rate and the number of iterations until the metric changes and the loss changes gradually tend to a stable state, and determine the final learning rate and the number of iterations;

[0103] Step 324: According to the determined learning rate and the number of iterations, complete the training and rendering of the 3D Gaussian splatting model to obtain a well-performing real-time drivable human body model. The driving effect is as follows Figure 6 shown. Obtain the pose parameters of the human body movement from the picture or video source, and use the pose parameters to drive the trained human body model. At this time, the human body model is consistent with the human body movement posture in the picture.

Claims

1. A real-time drivable 3D human avatar method generated based on 3D Gaussian splatting, characterized in that, Including the following steps: Step 1: Based on a single camera to collect a small amount of human data, predict the human pose information based on the smplx model and the expression information based on the Flame model from the two-dimensional human image data, and make labels; Step 2: For any frame of the label, perform adaptive Gaussian initialization on the human smplx model grid to obtain a 3D Gaussian ellipsoid; Step 3: Based on the improved 3d gaussian splatting model, train the 3D Gaussian ellipsoid initialized in Step 2, and render a real-time drivable 3D human avatar.

2. The real-time drivable 3D human avatar generation method based on 3D Gaussian splatting according to claim 1, characterized in that, In Step 1, based on a single camera to collect a small amount of human data, predicting the human pose information based on the smplx model and the expression information based on the Flame model from the two-dimensional human image data, and making labels specifically includes the following steps: Step 11: Use a single camera to take pictures of human data; fix the position of the camera, let the person being photographed perform simple actions in the picture, rotate a circle, and take pictures, and calculate the camera parameters and extract the picture frames; Step 12: Predict the smplx parameters and offsets; for the people in each frame of the pictures taken, use the smplx regression model to fit the smplx parameters of the human body, including shape, pose, root node, translation, and offsets of the patches; Step 13: Predict the Flame parameters and offsets; for the people in each frame of the pictures taken, intercept the face part, and use the fine-tuned DECA face model to predict the expression parameters, the pose of the chin joint, and the offsets of the expressions; Step 14: Make labels for the predicted smplx, Flame parameters and offsets to construct a dataset; replace the expression parameters and the pose of the chin joint in smplx with the expression parameters and the pose of the chin joint in Flame, and add the offsets of the expressions, and make these parameters and images correspond one by one to make a label file, construct a training set, a validation set and a test set, and randomly select 80% of the dataset as the training set, 10% of the dataset as the validation set, and 10% of the dataset as the test set.

3. The real-time drivable 3D human avatar method generated based on 3D Gaussian splatting according to claim 1, characterized in that, In Step 2, performing adaptive Gaussian initialization on the human smplx model grid for any frame of the label to obtain a 3D Gaussian ellipsoid specifically includes the following steps: Step 21: Based on the gender of the person in the image, select a standard smplx model, load the selected label file, generate a smplx human model and obtain the patches of the human model; Step 22: Construct a distribution model of an adaptive 3D Gaussian ellipsoid based on the patch size and the size of adjacent patches; Step 23: For the patches of the smplx human model in Step 21, perform an adaptive 3D Gaussian ellipsoid distribution using the distribution model constructed in Step 22.

4. The real-time drivable 3D human avatar method generated based on 3D Gaussian splatting according to claim 3, wherein In step 22, construct an adaptive 3D Gaussian ellipsoid distribution model based on the patch size and the adjacent patch sizes, calculate the area size of the current patch, and calculate the patch sizes of the three adjacent faces. Set a threshold and the initial number of Gaussian ellipsoids for the patch. When the size multiples of the four patches exceed the threshold, reallocate the number of Gaussian ellipsoids according to the area weight, but each patch should retain at least 1 Gaussian ellipsoid. After reallocation, the large patch re - equally divides the patch according to the threshold, and distributes Gaussian ellipsoids on the divided patches; if the size multiples of the current four patches are less than the threshold, perform the above operations on the adjacent patches. Meanwhile, to improve efficiency, directly perform a separate 3D Gaussian ellipsoid distribution on the patches that are particularly lower than the threshold until all patches are distributed.

5. The real-time drivable 3D human avatar method generated based on 3D Gaussian splatting according to claim 3, wherein In step 23, for the patches of the smplx human model in step 21, the adaptive 3D Gaussian ellipsoid distribution using the distribution model constructed in step 22 specifically includes the following steps: Step 231, read the patches of the smplx human model obtained in step 21 and mark the indexes. Step 232, perform 3D Gaussian distribution on the patches of the smplx model using the adaptive 3D Gaussian ellipsoid distribution model. Step 233, bind the distributed 3D Gaussian ellipsoids to the patch indexes.

6. The real-time drivable 3D human avatar generation method based on 3D Gaussian splatting according to claim 1, wherein, In step 3, based on the improved 3D Gaussian splatting model, train the 3D Gaussian ellipsoids initialized in step 2 and render a real - time drivable 3D human avatar, which specifically includes the following steps: Step 31, construct an improved 3D Gaussian splatting model by improving the non - rigid deformation module, improving the loss function, and adding a temporal color MLP correction module. Step 32, train the improved 3D Gaussian splatting model in step 31 and render the human model.

7. The method for generating a real-time drivable 3D human avatar based on 3D Gaussian splatting according to claim 6, wherein In step 31, improve the non - rigid deformation module: add a multi - scale geometric information module to the model, use a multi - layer hash grid to encode the position attributes of the input 3D Gaussian ellipsoids into spatial feature vectors, and concatenate the spatial feature vectors output by the hash grid with the smplx pose features and shape parameter features. Pass the concatenated features through a lightweight MLP network to learn the deformation offset and perform dynamic deformation on the initialized 3D Gaussian ellipsoids, which can accurately reflect the local deformation in different postures. The specific formula is as follows: Among them, δx represents the position offset, which is used to adjust the position offset of the center point of the 3D Gaussian ellipsoid; δd represents the normal offset, which controls the offset of the 3D Gaussian ellipsoid in the normal direction of the patch; δs represents the scale offset, which controls the size change of the 3D Gaussian ellipsoid; δq represents the rotation offset, which adjusts the rotation angle of the 3D Gaussian ellipsoid; z represents the local pose feature, which is used for the color module. denotes the MLP decoding module, γ(x c ) denotes the position encoding of the Gaussian ellipsoid position attribute x c ; S p denotes the feature encoding of the pose parameter and the shape parameter.

8. The method for generating a real-time drivable 3D human avatar based on 3D Gaussian splatting according to claim 6, wherein, Improve the loss function: add an additional geometric consistency regularization constraint to the deformation parameters to make the deformation result more physically reasonable; add a point - to - patch distance loss, calculate the distance from the center point p of the Gaussian ellipsoid to the bound patch, and use the minimum distance from the point to the triangular face to define it. The normal vector of the patch is n, and one vertex of the patch is v1, then the point - to - face distance loss is defined as: L %is(anc+ = Dist(p) Add the projection loss of the point inside the patch to ensure that the projection of the center point of the Gaussian ellipsoid falls within the bound patch. Assume that the center point of the Gaussian ellipsoid is p, and assume that the three vertices of the plane are v1, v2, and v3, and any point on the plane is p , , the normal vector is n, and the projection on the plane is P p. / , and the specific formula is as follows: p , = λ1v1 + λ # v # + λ3v3 λ1 + λ # + λ3 = 1 where \(0\leq\lambda_1,\lambda_2,\lambda_3\leq1\); Determine the projection point P p. / Check whether the barycentric coordinates (μ, ν, ω) satisfy that all values are within [0, 1]. The specific formula is as follows: v 81 = v1 - v # v 8# = v # - v8 v 8p = P p. / - v8 \(\omega = 1-\mu-\nu\) If μ, ν, ω ≥ 0 and μ + ν + ω = 1, then the point P p. / is inside or on the boundary of the triangle; otherwise, the point is outside the triangle; the loss function is specifically as follows: The temporal color MLP correction module adds a parameter \(t\) for the time moment in the color MLP module, incorporates the time dimension into the model, and enables it to correct the color information that changes over time.

9. The real-time drivable 3D human avatar method generated based on 3D Gaussian splatting according to claim 6, wherein, In step 32, train the improved 3D Gaussian splatting model in step 31 and render the human model, which specifically includes the following steps: Step 321: Set the training parameters and use the Adam optimization algorithm for training. The initial learning rates are set as follows: position Ir_position = 0.00016, opacity Ir_opacity = 0.05, scale Ir_scale = 0.005, rotation Ir_rot = 0.005, pose Ir_pose = 0.0002, color Ir_color = 0.

005. Set the final position learning rate to Ir_position = 0.0000016, and the number of training epochs Epoch = 10000; Step 322: Feed the dataset constructed in step 14 into the improved 3D Gaussian splatting model in step 31 for training and rendering; Step 323: According to the rendering effects of the model on the training set and the validation set, cross-validate the changes in the peak signal-to-noise ratio, structural similarity index, and perceptual image patch similarity metrics, as well as the loss change trend. Adjust the learning rate and the number of iterations until the changes in the metrics and the loss gradually tend to a stable state, and determine the final learning rate and the number of iterations; Step 324: According to the determined learning rate and the number of iterations, complete the training and rendering based on the 3D Gaussian splatting model to obtain a well-performing and real-time drivable human avatar.

10. A system for a real-time drivable 3D human avatar method generated based on 3D Gaussian splatting as claimed in claim 1, characterized in that, Including: A data acquisition module, a Gaussian initialization module, and a training and rendering module; The data acquisition module acquires a small amount of human body data based on a single camera, predicts the human body pose information based on the SMPLX model and the expression information based on the FLAME model from the two-dimensional human body image data, and makes labels; the Gaussian initialization module adaptively initializes the Gaussian of the human SMPLX model mesh in any frame label to obtain a 3D Gaussian ellipsoid; The training and rendering module trains the initialized 3D Gaussian ellipsoid based on the improved 3D Gaussian splatting model and renders a real-time drivable 3D human avatar.

Citation Information

Cited By

  • Newton-Simpson-based text 3D rapid generation method and system

    CN121392084A