Dressing human body geometry and texture parameterization reconstruction method based on 3D Gaussian
The integration of 3D Gaussian models with tailored parameterization and neural networks addresses the challenge of reconstructing clothing geometry and texture in dynamic human models, achieving realistic and efficient rendering of dressed humans.
Patent Information
- Application Number
- CN202510400728.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art is difficult to effectively reconstruct the geometry and texture of clothing, resulting in a large difference between the reconstructed clothing geometry and actual clothing, blurred textures or reduced accuracy, especially in dynamic scenes.
Combining the 3D Gaussian method with the human parametric model SMPL and the clothing parameterized model TailorNet, the properties of the 3D Gaussian ball are generated and optimized by calculating the masks of the background, human body, tops, and bottoms. Using improved rendering methods and neural network training, the rendering effects of each part are independently optimized, geometric constraints and dynamic learning rates are added, and the position and properties of the 3D Gaussian ball are optimized.
Realize realistic clothing geometry and texture reconstruction, improve reconstruction accuracy, reduce video memory usage, support online fitting and other applications, and adapt to clothing changes.
Smart Images

Figure CN120318394A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human motion and appearance reconstruction, and in particular to a method for parametric reconstruction of the geometry and texture of a clothed human body based on 3D Gaussian. Background Art
[0002] With the popularization and improvement of photo and video acquisition technologies, many new topics and application fields based on image and video analysis have emerged. In recent years, a series of concerns have begun to focus on image and video content, triggering extensive research. Regarding video content, exploring the reconstruction of dynamic three-dimensional human mesh models and textures from it has also received increasing attention. The technologies for human motion and appearance reconstruction have extensive applications in various fields, including human-computer interaction, individual or group behavior analysis, autonomous driving, robot-assisted care, game and movie production, and various aspects of human life such as virtual or augmented reality. These application fields together construct a framework for the metaverse, laying a foundation for achieving richer experiences. The present invention takes this as the application background and has urgent practical significance.
[0003] Recently, some works have attempted to combine human parametric models with 3D Gaussian methods, overcoming the problem that previous 3D Gaussian methods could only reconstruct static scenes, and at the same time having richer texture and geometric expression capabilities than those reconstructed by human parametric models. However, it is still difficult to replace the industry mainstream human body and its clothing modeling and animation solutions based on motion capture and scanning. One of the limitations is that in existing works, basic models such as the SMPL model are used, which have limited expression capabilities, especially unable to express the geometry and texture of clothing, resulting in a large difference between the reconstructed clothing geometry and the actual clothing, and also causing the reconstructed clothing texture to be blurred or the accuracy to decrease. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide a method for parametric reconstruction of the geometry and texture of a clothed human body based on 3D Gaussian, which can generate a 3D Gaussian sphere approximating the clothed human body, especially the geometry and texture of the clothing folds, and achieve realistic rendering of the clothed human body.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: A method for parametric reconstruction of the geometry and texture of a clothed human body based on 3D Gaussian, comprising the following steps:
[0006] 1) Calculate the masks of the background, human body, upper garment, and lower garment for each frame of the RGB video of the clothed human body, and estimate the pose parameters, shape parameters, three-dimensional meshes, global positions, and camera parameters of the human body and clothing for each frame based on the human parametric model SMPL and the clothing parametric model TailorNet;
[0007] 2) Subdivide the 3D meshes of the human body and the clothing obtained in step 1). After obtaining the subdivided meshes of the human body, upper garment, and lower garment, generate initial 3D Gaussian spheres at each vertex on the subdivided meshes. The attributes of the 3D Gaussian spheres include position μ, rotation q, scaling vector s, opacity α, and color spherical harmonic coefficients SH.
[0008] 3) Input the position μ of each 3D Gaussian sphere into the Gaussian estimation network to obtain the attributes of the 3D Gaussian sphere in the standard pose. Use the linear skinning method to transform the attributes of the 3D Gaussian sphere in the standard pose to the attributes of the 3D Gaussian sphere in the current pose. Then, render the attributes of the 3D Gaussian sphere in the current pose using the improved 3D Gaussian rendering method. Finally, compare the rendering result with the RGB video frame of the dressed human body, calculate the loss, and finally perform backpropagation to train the Gaussian estimation network, optimize the position μ of the 3D Gaussian sphere, and optimize the pose parameters and shape parameters of the human body and the clothing. Among them, the improvement of the 3D Gaussian rendering method lies in: ① Render the rendering results of the 3D Gaussian spheres belonging to the upper garment, lower garment, and human body separately and optimize them independently to improve the rendering effect of each part; ② Adopt a dynamic learning rate for different parts and alternately render the strategy of 3D Gaussian spheres with transparency and opacity to train, reducing the burrs and blurring results in the rendering; ③ Add geometric constraints to each 3D Gaussian sphere to keep it near the 3D mesh of the dressed human body, and use the strategy of dynamically splitting and removing 3D Gaussian spheres to optimize the attribute parameters and rendering effect of the 3D Gaussian spheres.
[0009] 4) Given a motion sequence, combine the attributes of the 3D Gaussian sphere in the standard pose of the Gaussian estimation network trained in step 3). Use the linear skinning method to transform the position μ and rotation q of the 3D Gaussian sphere in the standard pose to the current pose, and obtain the attributes of the 3D Gaussian sphere in the current pose. Finally, use the Gaussian splashing method to render the dressed human body rendering result of the given motion sequence.
[0010] Furthermore, in step 1), the specific implementation method of calculating the masks of the background, human body, upper garment, and lower garment is as follows: First, use a neural network model based on UNet to extract a rough clothing mask, calculate its two-dimensional bounding box, and then use the bounding box as a prompt box for the segmentation network model Segment Anything to perform segmentation, generate a fine segmentation mask, and finally separate the masks of the background, upper garment, lower garment, and human body into four parts.
[0011] The specific implementation method of estimating the pose parameters, shape parameters, 3D meshes, global positions, and camera parameters of the human body and the clothing for each frame is as follows: For the video of the dressed human body, use the human body estimation method ROMP to parse out the shape and pose parameters of its human body parametric model SMPL and the shape parameters of the clothing parametric model TailorNet, and obtain the 3D mesh.
[0012] Furthermore, in step 2), the 3D mesh of the human body and the 3D mesh of the clothing are subdivided using the loop method to obtain a subdivided mesh, and an initial 3D Gaussian sphere is generated at each vertex on the subdivided mesh. The specific method for the attributes of the 3D Gaussian sphere is as follows: the position μ is set to the coordinates of the vertex; the rotation q and the scaling vector s are calculated using an improved local frame algorithm based on the projection plane constraint. For each vertex p, calculate the normal vectors of all the triangles to which it belongs, and the normal vector of vertex p is obtained as N, where n is the number of triangles to which vertex p belongs, and N i is the normal vector of the i-th triangle to which vertex p belongs. A normal plane is established at vertex p along the normal vector N, and all the points adjacent to vertex p are projected onto this normal plane to obtain a set of projection points P'. Using PCA analysis, the two main axes are analyzed as p1 and p2. The rotation q is set to {p1, p2, N}, and the scaling vector s is set to {||p1||, ||p2||, min(||p1||, ||p2||) * 0.001}, where ||·|| represents the vector norm operation; the opacity α of the human body part is initialized to 0.5, and the upper and lower clothing parts are initialized to 1; the color spherical harmonic coefficients SH are initialized using random numbers.
[0013] Furthermore, in step 3), the Gaussian estimation network consists of a feature plane, a Gaussian geometry estimation network, and a Gaussian color estimation network. After the position μ of the 3D Gaussian sphere is input into the feature plane, the feature plane outputs a hidden vector z. After the hidden vector z is input into the Gaussian geometry estimation network, the Gaussian geometry estimation network outputs the predicted rotation q and scaling vector s. At the same time, the hidden vector z is input into the Gaussian color estimation network, and the Gaussian geometry estimation network outputs the opacity α and the color spherical harmonic coefficients SH. Among them, the Gaussian geometry estimation network and the Gaussian color estimation network are MLP neural networks;
[0014] The specific implementation method of rendering using the improved 3D Gaussian rendering method is: I = 3DGS(U, π K , π E ), where I represents the rendered image, 3DGS represents 3D Gaussian splash, and π K represents the camera internal parameters, defining the projection matrix of the camera; π E represents the camera external parameters, describing the position and orientation of the camera;
[0015] The specific implementation method of comparing the rendering result with the ground truth video frame is: optimize the pixel-level error between the rendered image I and the input real image I gt : I pix = ||I - I gt ||, where I gt represents the input real image, and I pixRepresents the absolute value of the pixel difference, which is used to guide the properties of the 3D Gaussian sphere; the rendering process is achieved through differentiable rasterization technology, and through the backpropagation mechanism, the properties of the 3D Gaussian sphere are continuously optimized to improve the quality of the rendered image and its matching degree with the input image.
[0016] Furthermore, in step 4), using the trained Gaussian estimation network, for a given motion sequence, it can drive the rendering result of the dressed human body generated with respect to this motion sequence. The process is as follows: input the position of the 3D Gaussian sphere in the optimized standard pose into the Gaussian estimation network to obtain the properties of the complete 3D Gaussian sphere in the standard pose, and then combine the motion sequence parameters and use the linear skinning method to obtain the properties of the 3D Gaussian sphere in the current pose frame by frame. Finally, perform 3D Gaussian splashing to obtain the rendering result of the dressed human body under the given motion sequence.
[0017] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0018] 1. The present invention uses the vertex positions generated by the physics-based clothing parameterization model as the initial values and geometric constraints of the 3D Gaussian sphere, and can reconstruct the dressed human body with clothing having wrinkle details.
[0019] 2. The neural network training of the present invention uses the rendering loss, which can further improve the geometric accuracy of the reconstructed dressed human body.
[0020] 3. The present invention proposes a method for adaptively updating the number of 3D Gaussian spheres, which can use a relatively small number of 3D Gaussian spheres to render results with similar accuracy, saving video memory occupancy and reducing training time.
[0021] 4. The present invention trains and optimizes the human body and clothing independently, can perform rendering after changing clothing, and can be applied to fields such as online virtual fitting. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the logical flow of the method of the present invention.
[0023] Figure 2 、 Figure 3 、 Figure 4 It is a schematic diagram of the dressed human body in different poses rendered by the present invention. The left picture is a real video frame, and the right picture is the rendering result of the present invention.
[0024] Figure 5 It is a schematic diagram of the Gaussian estimation network designed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0026] As Figures 1 to 5 shown, this embodiment discloses a method for parametric reconstruction of the geometry and texture of a clothed human body based on 3D Gaussian. Based on the human parametric model SMPL and the clothing parametric model TailorNet, the human body and clothing meshes of each frame of the RGB video are generated. After the meshes are subdivided, the 3D Gaussian sphere is initialized using the subdivision surface. The Gaussian attribute prediction network is used to estimate the 3D Gaussian sphere attributes, and then the linear skinning method is used to transform to the current frame 3D Gaussian sphere. The Gaussian splashing method is used for rendering. The rendered result is compared with the RGB video frame of the clothed human body, the loss is calculated, the Gaussian estimation network is trained, and the 3D Gaussian sphere attributes are optimized to achieve the reconstruction of the clothed human body, including the following steps:
[0027] 1) Calculate the masks of the background, human body, upper clothing, and lower clothing for each frame of the RGB video of the clothed human body, and estimate the pose parameters, shape parameters, three-dimensional meshes, global positions, and camera parameters of the human body and clothing for each frame.
[0028] The specific implementation method for calculating the masks of the background, human body, upper clothing, and lower clothing is as follows: First, a rough clothing mask is extracted using a neural network model based on UNet, and its two-dimensional bounding box is calculated. Then, the bounding box is used as a prompt box for the segmentation network model Segment Anything for segmentation to generate a fine segmentation mask, and finally, the masks of the four parts of the background, upper clothing, lower clothing, and human body are separated.
[0029] The specific implementation method for estimating the pose parameters, shape parameters, three-dimensional meshes, global positions, and camera parameters of the human body and clothing for each frame is as follows: For the video of the clothed human body, the human estimation method ROMP is used to parse the shape and pose parameters of the human parametric model SMPL and the shape parameters of the clothing parametric model TailorNet, and the three-dimensional mesh is obtained.
[0030] 2) The three-dimensional meshes of the human body and the clothing obtained are subdivided using the loop method to obtain subdivision meshes. Initial 3D Gaussian spheres are generated at each vertex on the subdivision meshes. The attributes of the 3D Gaussian spheres include position μ, rotation q, scaling vector s, opacity α, and color spherical harmonic coefficients SH. The specific method for the attributes of the 3D Gaussian spheres is as follows: The position μ is set to the coordinates of the vertex; the rotation q and the scaling vector s are calculated using an improved local frame algorithm based on the projection plane constraint: For each vertex p, calculate the normal vectors of all the triangles to which it belongs, and obtain the normal vector N of the vertex p, where n is the number of triangles to which the vertex p belongs, N iLet \(N\) be the normal of the \(i\)-th triangle to which the vertex \(p\) belongs. A normal plane is established at the vertex \(p\) along the normal \(N\). All points adjacent to the vertex \(p\) are projected onto this normal plane to obtain a set of projection points \(P'\). Using PCA analysis, the two main axes are analyzed as \(p1\) and \(p2\). The rotation \(q\) is set as \(\{p1, p2, N\}\), and the scaling vector \(s\) is set as \(\{\|p1\|, \|p2\|, \min(\|p1\|, \|p2\|) * 0.001\}\), where \(\|\cdot\|\) represents the vector norm operation; the opacity \(\alpha\) of the human body part is initialized to \(0.5\), and the clothing parts of the upper and lower garments are initialized to \(1\); the spherical harmonic coefficients \(SH\) of the color are initialized using random numbers.
[0031] 3) Input the position \(\mu\) of each 3D Gaussian sphere into the Gaussian estimation network, as Figure 5 shown, to obtain the 3D Gaussian sphere attributes in the standard pose. Use the linear skinning method to transform the 3D Gaussian sphere attributes in the standard pose to the 3D Gaussian sphere attributes in the current pose, and then use the improved 3D Gaussian rendering method to render the obtained 3D Gaussian sphere attributes in the current pose. Finally, compare the rendering result with the ground truth video frame, calculate the loss, update the network weights, and update the parameters. As Figure 2 、 Figure 3 、 Figure 4 shown, which are the reconstructed results of the clothed human body in different poses respectively.
[0032] Among them, the Gaussian estimation network consists of a feature plane, a Gaussian geometry estimation network, and a Gaussian color estimation network. After inputting the position \(\mu\) of the 3D Gaussian sphere into the feature plane, the feature plane will output a latent vector \(z\); after inputting the latent vector \(z\) into the Gaussian geometry estimation network, the Gaussian geometry estimation network will output the predicted rotation \(q\) and scaling vector \(s\); at the same time, input the latent vector \(z\) into the Gaussian color estimation network, and the Gaussian geometry estimation network will output the opacity \(\alpha\) and the spherical harmonic coefficients \(SH\) of the color. The Gaussian geometry estimation network and the Gaussian color estimation network are MLP neural networks.
[0033] The specific implementation method of rendering using the improved 3D Gaussian rendering method is: \(I = 3DGS(U, \pi\) K , \pi\) E ), where \(I\) represents the rendered image, \(3DGS\) represents 3D Gaussian splash, \(\pi\) K represents the camera internal parameters, defining the projection matrix of the camera; \(\pi\) E represents the camera external parameters, describing the position and orientation of the camera.
[0034] The specific implementation method of comparing the rendering result with the ground truth video frame is: optimize the pixel-level error between the rendered image \(I\) and the input real image \(I\) gt : \(I\) pix =\(\|I - I\) gt \|, where \(I\) gtRepresents the real image of the input, I pix Represents the absolute value of the pixel difference, which is used to guide the attributes of the 3D Gaussian sphere. The rendering process is implemented through differentiable rasterization technology. Through the backpropagation mechanism, the attributes of the 3D Gaussian sphere are continuously optimized to improve the quality of the rendered image and its matching degree with the input image.
[0035] 4) Using the trained Gaussian estimation network, for a given motion sequence, it can drive the rendering result of the generated dressed human body with respect to this motion sequence. The process is as follows: Input the position of the 3D Gaussian sphere in the optimized standard pose into the Gaussian estimation network to obtain the attributes of the 3D Gaussian sphere in the complete standard pose. Then, combined with the motion sequence parameters, use the linear skinning method to obtain the attributes of the 3D Gaussian sphere in the current pose frame by frame. Finally, perform 3D Gaussian splashing to obtain the rendering result of the dressed human body under the given motion sequence.
[0036] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for parametric reconstruction of the geometry and texture of a dressed human body based on 3D Gaussian, characterized in that, The steps include the following: 1) Calculate the masks of the background, the human body, the upper garment, and the lower garment for each frame of the RGB video of the dressed human body. Based on the human parametric model SMPL and the garment parametric model TailorNet, estimate the pose parameters, shape parameters, 3D meshes, global positions, and camera parameters of the human body and the garment for each frame; 2) Subdivide the 3D meshes of the human body and the garment obtained in step 1). After obtaining the subdivided meshes of the human body, the upper garment, and the lower garment, generate initial 3D Gaussian spheres at each vertex on the subdivided meshes. The attributes of the 3D Gaussian spheres include position μ, rotation q, scaling vector s, opacity α, and color spherical harmonic coefficients SH; 3) Input the position μ of each 3D Gaussian sphere into the Gaussian estimation network to obtain the attributes of the 3D Gaussian spheres in the standard pose. Use the linear skinning method to transform the attributes of the 3D Gaussian spheres in the standard pose to the attributes of the 3D Gaussian spheres in the current pose. Then, render the attributes of the 3D Gaussian spheres in the current pose using the improved 3D Gaussian rendering method. Finally, compare the rendering result with the frame of the RGB video of the dressed human body, calculate the loss, and finally perform backpropagation to train the Gaussian estimation network, optimize the position μ of the 3D Gaussian spheres, and optimize the pose parameters and shape parameters of the human body and the garment. Among them, the improvement of the 3D Gaussian rendering method lies in: ① Render the rendering results of the 3D Gaussian spheres belonging to the upper garment, the lower garment, and the human body separately and optimize them independently to improve the rendering effect of each part; ② Adopt a dynamic learning rate for different parts and train using the strategy of alternately rendering 3D Gaussian spheres with transparency and opacity to reduce the burrs and blurred results in the rendering; ③ Add geometric constraints to each 3D Gaussian sphere to keep it near the 3D mesh of the dressed human body, and use the strategy of dynamically splitting and removing 3D Gaussian spheres to optimize the attribute parameters and rendering effect of the 3D Gaussian spheres; 4) Given a motion sequence, combine the attributes of the 3D Gaussian spheres in the standard pose of the Gaussian estimation network trained in step 3). Use the linear skinning method to transform the position μ and rotation q of the 3D Gaussian spheres in the standard pose to the current pose, obtain the attributes of the 3D Gaussian spheres in the current pose, and finally use the Gaussian splashing method to render the dressed human body rendering result of the given motion sequence.
2. The method for parametric reconstruction of the geometry and texture of a dressed human body based on 3D Gaussian according to claim 1, wherein: In step 1), the specific implementation method for calculating the masks of the background, the human body, the upper garment, and the lower garment is as follows: First, use a neural network model based on UNet to extract a rough garment mask, calculate its two-dimensional bounding box, and then use the bounding box as a prompt box for the segmentation network model Segment Anything for segmentation to generate a fine segmentation mask. Finally, separate the masks of the background, the upper garment, the lower garment, and the human body into four parts; The specific implementation method for estimating the pose parameters, shape parameters, 3D meshes, global positions, and camera parameters of the human body and the garment for each frame is as follows: For the video of the dressed human body, use the human estimation method ROMP to parse the shape and pose parameters of its human parametric model SMPL and the shape parameters of the garment parametric model TailorNet, and obtain the 3D mesh.
3. The method for parametric reconstruction of the geometry and texture of a dressed human body based on 3D Gaussian according to claim 2, wherein: In step 2), the 3D mesh of the human body and the 3D mesh of the clothing are subdivided using the loop method to obtain a subdivided mesh, and an initial 3D Gaussian sphere is generated at each vertex on the subdivided mesh. The specific manner of the attributes of the 3D Gaussian sphere is as follows: the position μ is set to the coordinates of the vertex; the rotation q and the scaling vector s are calculated using an improved local frame algorithm based on the projection plane constraint. For each vertex p, calculate the normal vectors of all the triangles to which it belongs, and the normal vector of vertex p is obtained as N, where n is the number of triangles to which vertex p belongs, and N i is the normal vector of the i-th triangle to which vertex p belongs. A normal plane is established at vertex p along the normal vector N, and all the points adjacent to vertex p are projected onto this normal plane to obtain a set of projection points P'. Using PCA analysis, the two main axes are analyzed as p1 and p2. The rotation q is set to {p1, p2, N}, and the scaling vector s is set to {||p1||, ||p2||, min(||p1||, ||p2||) * 0.001}, where ||·|| represents the vector norm operation; the opacity α of the human body part is initialized to 0.5, and the upper and lower clothing parts are initialized to 1; the color spherical harmonic coefficients SH are initialized using random numbers.
4. The method for parametric reconstruction of geometric and texture parameters of a dressed human body based on 3D Gaussian according to claim 3, characterized in that: In step 3), the Gaussian estimation network consists of a feature plane, a Gaussian geometry estimation network, and a Gaussian color estimation network; after inputting the position μ of the 3D Gaussian sphere into the feature plane, the feature plane outputs a latent vector z; after inputting the latent vector z into the Gaussian geometry estimation network, the Gaussian geometry estimation network outputs a predicted rotation q and a scaling vector s; simultaneously inputting the latent vector z into the Gaussian color estimation network, the Gaussian geometry estimation network outputs an opacity α and color spherical harmonic coefficients SH; among them, the Gaussian geometry estimation network and the Gaussian color estimation network are MLP neural networks; The specific implementation method of rendering with the improved 3D Gaussian rendering method is: I = 3DGS(U, π K , π E ), where I represents the rendered image, 3DGS represents 3D Gaussian splash, and π K represents the camera internal parameters, defining the projection matrix of the camera; π E represents the camera external parameters, describing the position and orientation of the camera; The specific implementation method of comparing the rendering result with the ground-truth video frame is as follows: optimize the pixel-level error between the rendered image I and the input real image I gt : I pix = ||I - I gt ||, where I gt represents the input real image, and I pix represents the absolute value of the pixel difference, which is used to guide the attributes of the 3D Gaussian sphere; the rendering process is implemented by differentiable rasterization technology, and the attributes of the 3D Gaussian sphere are continuously optimized through the backpropagation mechanism, which is used to improve the quality of the rendered image and its matching degree with the input image.
5. The method for parametric reconstruction of the geometry and texture of a dressed human body based on 3D Gaussian according to claim 4, wherein: In step 4), using the trained Gaussian estimation network, for a given motion sequence, it can drive the rendering result of the dressed human body generated with respect to this motion sequence. The process is as follows: input the position of the 3D Gaussian sphere in the optimized standard pose into the Gaussian estimation network to obtain the complete attributes of the 3D Gaussian sphere in the standard pose, then combine the motion sequence parameters and use the linear skinning method to obtain the attributes of the 3D Gaussian sphere in the current pose frame by frame, and finally perform 3D Gaussian splashing to obtain the rendering result of the dressed human body under the given motion sequence.
Citation Information
Cited By
Digital human generation method and related device
CN121810883A
Double-flow collaborative reconstruction method and system based on template frame initialization
CN122289507A
A dual-stream collaborative reconstruction method and system based on template frame initialization
CN122289507B