Single-view human body three-dimensional reconstruction method based on Gaussian surface elements
Through the single-view human body three-dimensional reconstruction method based on Gaussian surface elements, the problems of low reconstruction accuracy and poor quality in the prior art are solved, and high-quality three-dimensional human body reconstruction is achieved, reducing production costs and improving efficiency.
Patent Information
- Application Number
- CN202411982759.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing three-dimensional human body reconstruction methods have problems with low reconstruction accuracy and poor quality, especially in single-view conditions, it is difficult to generate high-quality geometric reconstruction.
A single-view human body three-dimensional reconstruction method based on Gaussian surface elements is adopted to collect image data from multiple human body data sets for pre-processing, and a 3D Gaussian attribute prediction model of the human body is constructed, including a hierarchical feature extraction module, a feature fusion network, and a 3D Gaussian function decoder, for training and three-dimensional reconstruction.
The reconstruction geometric quality is improved, and the reconstruction speed and quality are taken into account, which greatly reduces the cost of virtual digital human production, and improves the efficiency and accuracy of building dynamic human models.
Smart Images

Figure CN119991937A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human body three-dimensional reconstruction, and in particular to a single-view human body three-dimensional reconstruction method based on Gaussian facets. Background Art
[0002] In recent years, technologies such as the metaverse and virtual reality / augmented reality have developed rapidly. Virtual anchors, digital avatars, film and game production have shown great development potential, and the market size of related industries has steadily increased. As an important part of the virtual world and the main body of virtual interaction, the three-dimensional digital human, and the closely related fields of human three-dimensional reconstruction and digital human driving have received extensive attention from the industry and academia, and have become a common research hotspot in computer vision, computer graphics and other fields. As an emerging three-dimensional human reconstruction method, human 3D Gaussian Splatting (Gaussian splatting, a 3D rendering method) represents 3D scenes through a large number of 3D Gaussian functions (splats), which have parameters such as position and direction. The parameters are optimized by using a method similar to training neural networks, which has advantages in rendering speed and scene representation accuracy. This enables high-quality synthesis of new perspectives and new postures, bringing new ideas and breakthroughs to three-dimensional human reconstruction.
[0003] At present, the main methods for obtaining commercial high-quality human models are: art production, dedicated hardware scanning, multi-view 3D reconstruction, etc. These methods can obtain detailed human models and even surface material information, but the cost is high. In the field of 3D scene reconstruction and rendering, many methods have their own advantages and disadvantages. Traditional multi-view stereo (MVS) technology relies on photometric consistency to reconstruct geometric representations across views, but it is difficult to accurately capture complete geometric shapes due to fuzzy correspondence. With the development of deep learning technology, various sparse view reconstruction methods have emerged in large numbers, and methods based on parametric models and implicit function fields have developed rapidly. However, there are still some problems with these methods: implicit function field methods tend to reconstruct static human bodies and are difficult to be directly used in commercial animation production pipelines. Therefore, they are prone to poor reconstruction texture and low precision; and the representation ability of parametric models is limited, and only naked models can be obtained, which is prone to poor reconstruction effects.
[0004] Therefore, the prior art needs to be improved. Summary of the invention
[0005] The technical problem to be solved by the present invention is that, in view of the defects of the prior art, the present invention provides a single-view human body three-dimensional reconstruction method based on Gaussian facets to solve the problems of low reconstruction accuracy and poor quality in the existing three-dimensional human body reconstruction methods.
[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0007] In a first aspect, the present invention provides a single-view human body 3D reconstruction method based on Gaussian facets, comprising:
[0008] Collect image data from multiple human data sets, and pre-process them using targeted pre-processing strategies according to data characteristics corresponding to each image data to obtain a processed human data set;
[0009] Constructing a 3D Gaussian attribute prediction model for a human body, wherein the 3D Gaussian attribute prediction model for a human body comprises: a hierarchical feature extraction module, a feature fusion network, and a 3D Gaussian function decoder;
[0010] Training the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain a trained human body 3D Gaussian attribute prediction model;
[0011] Based on the trained human body 3D Gaussian attribute prediction model, the input human body picture is three-dimensionally reconstructed, and a reconstructed human body three-dimensional image is output.
[0012] In one implementation, the collecting of image data from a plurality of human body data sets, and preprocessing using a targeted preprocessing strategy according to data characteristics corresponding to each image data to obtain a processed human body data set include:
[0013] Collecting corresponding image data from multiple human body datasets;
[0014] For the image data of the human body scan dataset, the camera pose is fixed, and the rendered image, image mask, depth map, normal map and corresponding SMPL human body template parameters are obtained from the human body scan model to obtain the corresponding processed human body dataset.
[0015] In one implementation, for the image data of the human body scan data set, fixing the camera pose, obtaining the rendered image, image mask, depth map, normal map and corresponding SMPL human body template parameters from the human body scan model, includes:
[0016] The camera posture is fixed to obtain a fixed number of frames of random sampling of the image, and the spherical harmonic coefficients of the radiation transmission of the human body scanning model are calculated, and the rendering image is obtained by scanning according to the fixed number of frames and the spherical harmonic coefficients;
[0017] Based on the human posture estimation algorithm, the human body key points of the human body scanning model are estimated, and the estimated human body key points are input into the MLP neural network to predict and obtain corresponding SMPL parameters;
[0018] The vertices of the human body scan model are transformed and projected onto the camera plane, and the depth values of the projection points are calculated, and the depth values are stored in a two-dimensional array with the same resolution as the projection plane to obtain the depth map;
[0019] The vertex normals in the human body scan model are calculated, and the vertex normals are converted into a texture space, and the normal map is obtained after interpolation and mapping processing.
[0020] In one implementation, the training of the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain the trained human body 3D Gaussian attribute prediction model includes:
[0021] Randomly selecting images from different perspectives as sample data based on the processed human body data set;
[0022] Inputting the view image of the sample data and the corresponding camera parameters and sampling points into the human body 3D Gaussian attribute prediction model;
[0023] Extracting features from different dimensions using the hierarchical feature extraction module, fusing the extracted features through the feature fusion network, predicting Gaussian patch parameters based on the 3D Gaussian function decoder, and generating a predicted image of the target view based on volume rendering;
[0024] The photometric loss, mask loss, structural similarity index loss, and depth-normal consistency loss are calculated, and the model parameters are optimized according to the calculated losses to obtain the trained human 3D Gaussian attribute prediction model.
[0025] In one implementation, the three-dimensional reconstruction of the input human body image based on the trained human body 3D Gaussian attribute prediction model includes:
[0026] The human body picture is input, and the parameters of the explicit human model are estimated using a trained regression network. The distance field method is used to generate the surface distance field of the explicit human template mesh, and points that evenly cover the area near the surface are obtained to obtain an initial point cloud.
[0027] In one implementation, the three-dimensional reconstruction of the input human body image based on the trained human body 3D Gaussian attribute prediction model includes:
[0028] Compressing the input human body image into potential features through a neural network, and globally encoding it through a three-plane method;
[0029] Projecting the vertices of the explicit human model into the two-dimensional feature map of the input human body image, extracting features of each point, and performing sparse three-dimensional convolution processing after voxelization;
[0030] Encode the corresponding colors of the point cloud, concatenate them with the processed features, enhance the local features, and project the point cloud features using geometric perception coding, input them into the tansformer for decoding to obtain Gaussian facets;
[0031] Generate images, depth maps, and normal maps based on the predicted Gaussian surfaces;
[0032] The generated depth map and normal map are used as input, and a screened Poisson reconstruction algorithm is used to build a point cloud model based on the generated depth map. By solving the Poisson equation to fit the function gradient to the input normal field, the conversion from discrete point cloud to continuous surface representation is achieved, and the human body mesh is reconstructed.
[0033] In one implementation, the projecting of point cloud features by using geometric perception coding and combining multiple types of feature decoding to obtain Gaussian facets include:
[0034] Based on the three-plane feature representation method, feature query is performed. The given position in the point cloud is projected onto each plane, and the interpolation feature is obtained from the corresponding plane using the trilinear interpolation function. All interpolation features are sequentially spliced to obtain the final feature.
[0035] Inputting the given position in the point cloud and the final feature into the MLP neural network, decoding to obtain 3D Gaussian attributes; wherein the 3D Gaussian attributes include: position offset, opacity, anisotropic covariance and spherical harmonic coefficients;
[0036] The Gaussian surface element is constructed from the 3D Gaussian points according to the 3D Gaussian attributes.
[0037] In a second aspect, the present invention provides a single-view human body 3D reconstruction system based on Gaussian facets, comprising:
[0038] A preprocessing module is used to collect image data from multiple human body data sets, and to perform preprocessing using a targeted preprocessing strategy according to data characteristics corresponding to each image data, so as to obtain a processed human body data set;
[0039] A model building module is used to build a 3D Gaussian attribute prediction model for a human body, wherein the 3D Gaussian attribute prediction model for a human body includes: a hierarchical feature extraction module, a feature fusion network, and a 3D Gaussian function decoder;
[0040] A model training module, used for training the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain a trained human body 3D Gaussian attribute prediction model;
[0041] The human body three-dimensional image reconstruction module is used to perform three-dimensional reconstruction on the input human body picture based on the trained human body 3D Gaussian attribute prediction model, and output a reconstructed human body three-dimensional image.
[0042] In a third aspect, the present invention provides a terminal comprising: a processor and a memory, wherein the memory stores a single-view human body three-dimensional reconstruction program based on Gaussian surfaces, and when the single-view human body three-dimensional reconstruction program based on Gaussian surfaces is executed by the processor, it is used to implement the operation of the single-view human body three-dimensional reconstruction method based on Gaussian surfaces as described in the first aspect.
[0043] In a fourth aspect, the present invention further provides a medium, which is a computer-readable storage medium, and which stores a single-view human body three-dimensional reconstruction program based on Gaussian surfaces. When the single-view human body three-dimensional reconstruction program based on Gaussian surfaces is executed by a processor, it is used to implement the operation of the single-view human body three-dimensional reconstruction method based on Gaussian surfaces as described in the first aspect.
[0044] The present invention adopts the above technical solution to achieve the following effects:
[0045] The present invention collects image data from multiple human data sets, and can adopt targeted preprocessing strategies for preprocessing according to the data characteristics corresponding to each image data; and by constructing a human 3D Gaussian attribute prediction model, the human 3D Gaussian attribute prediction model can be trained based on the processed human data set, so as to perform three-dimensional reconstruction on the input human body picture based on the trained human 3D Gaussian attribute prediction model, and output a reconstructed three-dimensional image of the human body. The reconstruction result obtained by the present invention combines the optimization flexibility of 3D Gaussian points and the surface alignment of facets, improves the quality of reconstruction geometry, takes into account the reconstruction speed and quality, greatly reduces the cost of virtual digital human production, and improves the efficiency of accurately constructing dynamic human body models. The generated human body model has significant value in the fields of virtual anchors, digital avatars, movies and game production. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.
[0047] Figure 1 It is a flow chart of the single-view human body 3D reconstruction method based on Gaussian facets in the present invention.
[0048] Figure 2 It is the overall block diagram of the human body 3D Gaussian attribute prediction model in the present invention.
[0049] Figure 3 It is a functional principle diagram of a terminal in one implementation of the present invention.
[0050] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0052] Exemplary Methods
[0053] At present, the three-dimensional reconstruction methods of the human body mainly include reconstruction methods based on implicit functions and reconstruction methods based on explicit shapes; among them, the reconstruction methods based on implicit functions have flexible topological structures through implicit representations (such as occupancy and signed distance fields), and can effectively describe three-dimensional clothed people in various scenes, including loose clothing and complex postures. A series of studies have focused on regressing implicit surfaces directly from a single input image. Others add three-dimensional human bodies before enhancing the 2D feature extraction and 3D feature reconstruction process. These methods lack information from other perspectives or prior knowledge (such as diffusion models), resulting in unsatisfactory textures. Some methods use diffusion models for mesh drawing, but they decline as mesh reconstruction is inaccurate.
[0054] Explicit shape-based reconstruction methods use parametric body models to estimate the shape and pose of a 3D human body. To incorporate clothing into the 3D model, these methods usually use 3D clothing offsets or use adjustable clothing templates on a basic body shape. Explicit shape methods may be limited by topological constraints, which becomes apparent when dealing with different and complex clothing styles in the real world, such as dresses and skirts.
[0055] Human NeRF model (3D human reconstruction) and human 3D Gaussian Splatting (Gaussian splatting, a 3D rendering method) can synthesize high-fidelity new views or 3D human poses given multiple views or monocular human videos. Although these methods have achieved impressive results, they usually require a lot of time and dense viewpoints. To address this problem, people are increasingly interested in generalizable human NeRF and human 3D Gaussian Splatting. These methods require fewer viewpoints and only one inference to achieve the goal. Although NeRF can achieve photorealistic view synthesis, it is easy to introduce high-frequency noise when extracting isosurfaces based on density value heuristic thresholds. Point rendering technology represents geometry with topological structure-free samples, such as 3D Gaussian sputtering (3DGS), which uses Gaussian points to represent scenes for fast reconstruction and real-time rendering, but it is difficult to generate high-quality geometric reconstruction. This is due to the non-zero thickness of Gaussian points, normal blur and sharp edge modeling deviations. Although there are methods to introduce regularization terms to alleviate the thickness problem, the quality of the reconstructed surface is still poor. Therefore, there is an urgent need to develop high-quality surface reconstruction technology.
[0056] In response to the above technical problems, a single-view human body 3D reconstruction method based on Gaussian facets is provided in an embodiment of the present invention. The method mainly collects image data from multiple human body data sets, and can adopt targeted preprocessing strategies for preprocessing according to the data characteristics corresponding to each image data; and by constructing a human body 3D Gaussian attribute prediction model, the human body 3D Gaussian attribute prediction model can be trained based on the processed human body data set, so as to perform 3D reconstruction on the input human body picture based on the trained human body 3D Gaussian attribute prediction model, and output a reconstructed human body 3D image. The reconstruction result obtained in the embodiment of the present invention combines the optimization flexibility of 3D Gaussian points and the surface alignment of facets, improves the quality of reconstruction geometry, takes into account the reconstruction speed and quality, greatly reduces the cost of virtual digital human production, and improves the efficiency of accurately constructing dynamic human body models. The generated human body model has significant value in the fields of virtual anchors, digital avatars, movies and game production.
[0057] like Figure 1 As shown, an embodiment of the present invention provides a single-view human body 3D reconstruction method based on Gaussian facets, comprising the following steps:
[0058] Step S100 , collecting image data from a plurality of human body data sets, and preprocessing the image data using a targeted preprocessing strategy according to data characteristics corresponding to each image data, to obtain a processed human body data set.
[0059] In this embodiment, a hierarchical multi-scale 3D transformer feature fusion network is proposed for predicting Gaussian attributes. This is a new hybrid representation that uses explicit and implicit representations to quickly and high-quality single-view reconstruction. At the same time, 2D Gaussian facets are used for single-view human reconstruction, based on the human display representation method and the MLP network human reconstruction method, while giving play to the prior information of the human body parameterized model SMPL (explicit human model). The obtained reconstruction results combine the flexibility of 3D Gaussian point optimization and the surface alignment of facets to improve the quality of reconstruction geometry, take into account the reconstruction speed and quality, overcome the limitations of existing technologies, and make it possible to automatically reconstruct hyper-realistic and drivable digital humans. The method provided in this embodiment has a high degree of automation and does not require manual intervention throughout the process. The obtained reconstruction results are compatible with mainstream commercial animation production software / pipeline, which will greatly reduce the cost of virtual digital human production and improve the efficiency of accurately constructing dynamic human models. The generated human body model has significant value in the fields of virtual anchors, digital avatars, movies and game production.
[0060] In order to achieve the above-mentioned purpose of the invention, a method for training a human 3D reconstruction model is first proposed in this embodiment. In the training process, image data needs to be collected from multiple human data sets and preprocessed according to the characteristics of different data sets. For example, for the THuman2.0 data set, the camera pose is fixed to render the image, depth map, and normal map, and the SMPL template parameters are estimated at the same time, and the distance field is used to sample the initial point cloud on the template mesh.
[0061] Specifically, in an implementation of this embodiment, step S100 includes the following steps:
[0062] Step S101, collecting corresponding image data from multiple human body data sets;
[0063] Step S102, for the image data of the human body scan data set, fix the camera posture, obtain the rendered image, image mask, depth map, normal map and corresponding SMPL human body template parameters from the human body scan model, and obtain the corresponding processed human body data set.
[0064] In this embodiment, image data is collected from multiple large-scale human data sets (e.g., THuman2.0, RenderPeople, ZJU_MoCap, and HuMMan data sets). For different data sets, targeted preprocessing strategies are adopted according to their own characteristics. Taking the THuman2.0 data set as an example, for the 3D scan model of each subject, it is necessary to fix the camera pose and render images, image masks, depth maps, and normal maps from the human scan model.
[0065] Specifically, in an implementation of this embodiment, step S102 includes the following steps:
[0066] Step S102a, fixing the camera posture to obtain a fixed number of frames of random sampling of the image, and calculating the spherical harmonic coefficients of the radiation transmission of the human body scanning model, and scanning according to the fixed number of frames and the spherical harmonic coefficients to obtain the rendered image;
[0067] Step S102b, estimating the human body key points of the human body scanning model based on a human body posture estimation algorithm, and inputting the estimated human body key points into an MLP neural network for prediction to obtain corresponding SMPL parameters;
[0068] Step S102c, transforming and projecting the vertices of the human body scan model onto the camera plane, calculating the depth values of the projection points, and storing the depth values into a two-dimensional array having the same resolution as the projection plane to obtain the depth map;
[0069] Step S102d, calculating the vertex normals in the human body scan model, and converting the vertex normals into texture space, and obtaining the normal map through interpolation and mapping processing.
[0070] Specifically, in this embodiment, the camera posture is fixed to obtain a fixed number of frames (for example, 72 frames) of random sampling of the image. For the human body scanning model, the spherical harmonic coefficients of the pre-obtained radiation transfer (PRT) are calculated, and the rendered image is obtained based on the fixed number of frames and the spherical harmonic coefficients. In short, PRT is used to consider accurate light transmission (including ambient occlusion) without affecting the online rendering time, which significantly improves the authenticity of the photo compared to ordinary sperm harmonic rendering using surface normals. In the preprocessing stage, a normalization operation is performed on the image to normalize the pixel values to a specific interval (for example, [0,1]) to unify the data scale and improve the stability of model training; cropping and alignment operations are performed according to the human body posture and image content to ensure that the human body is located in the core area of the image and the posture is standardized, providing a standard data format for subsequent model processing.
[0071] In order to obtain the human SMPL template parameters, in this embodiment, OpenPose (i.e., a human posture estimation algorithm, including a human posture estimation library) is also used to estimate the key points of the human body, and the key point information is provided to the MLP neural network for prediction to obtain the SMPL template parameters.
[0072] In order to obtain a depth map from the human body scan model, in this embodiment, the vertices of the human body scan model are converted from the model space to the camera space and projected to the camera plane, and the depth values of the projection points are calculated (for example, according to the z coordinate conversion of the vertex in the camera space under perspective projection). Finally, the depth values are stored in a two-dimensional array with the same resolution as the projection plane to obtain a depth map.
[0073] In order to obtain a normal map from a human body scan model, in this embodiment, the normal of each triangular face in the human body scan model is calculated, and the face normals of the shared vertices are added and normalized to obtain the vertex normal. Then the vertex normal is converted to the texture space (using UV coordinates), stored in the texture by interpolation, and the normal vector is mapped from [-1, 1] to [0, 1] for storage, and finally a normal map is obtained.
[0074] In this embodiment, through the above preprocessing process, a rendered image, an image mask, a depth map and a normal map are obtained from the human body scan model of the THuman2.0 data set, thereby obtaining a corresponding processed human body data set.
[0075] like Figure 1 As shown, an embodiment of the present invention provides a single-view human body 3D reconstruction method based on Gaussian facets, comprising the following steps:
[0076] Step S200, constructing a human body 3D Gaussian attribute prediction model, wherein the human body 3D Gaussian attribute prediction model comprises: a hierarchical feature extraction module, a feature fusion network and a 3D Gaussian function decoder.
[0077] In this embodiment, after collecting image data and performing data preprocessing, a 3D Gaussian attribute prediction model of the human body is constructed. This process includes global feature extraction, voxel feature fusion, Triplane-based feature query and Gaussian attribute decoding, as well as position correction and local feature fusion operations, so as to achieve the purpose of effectively encoding and decoding three-dimensional Gaussian distribution and improving reconstruction accuracy.
[0078] Specifically, the human body 3D Gaussian Splatting model (i.e., human body 3D Gaussian attribute prediction model) constructed in this embodiment mainly includes a hierarchical feature extraction module, a feature fusion network, and a 3D Gaussian function decoder.
[0079] Global Feature Extraction: Capturing the global structure and overall appearance is crucial for recovering 3D splats from a single view. In this embodiment, the entire human image is compressed into a compact latent code for global encoding, which helps to encode this global information. A two-dimensional encoder is used to compress the input image into a compact latent code. In order to effectively decode the three-dimensional representation, a three-plane representation is used, which plays an important role in missing information completion.
[0080] Given an image with camera parameters, they are first encoded into a set of potential features using a pre-trained model. In this embodiment, encoding is performed based on a transformer three-plane encoder, which encodes an implicit feature field that can encode three-dimensional Gaussian properties. The three planes include three axially arranged orthogonal feature planes {Txy, Txz, Tyz}. For any position x, the corresponding feature vector can be queried to project from the three planes to the axis-aligned feature planes, and the three trilinear interpolation features are connected as the final feature interp(Txy,pxy)⊕interp(Txz,pxz)⊕interp(Tyz,pyz), where interp and ⊕ represent trilinear interpolation and connection operations, and p represents the projection position on each plane.
[0081] Fusion of Voxel-wise Features: For single image human 3D Gaussian Splatting, it is important to recover both global structure and local details from the input image, which can be fused via an underlying explicit human model, i.e., the SMPL model. The SMPL vertices are first projected into the 2D feature map of the input image and per-point features are extracted. One problem with the above feature extraction process in a single human image input setting is that only half of the SMPL vertices are visible from the input view. Visible vertex features are extracted and voxelized into a sparse 3D volume tensor, which is further processed using sparse 3D convolutions.
[0082] Pixel feature fusion: pixel-aligned features. Local feature enhancement is performed by convolution of point-level feature space. However, due to the limited grid resolution and voxel resolution of SMPL, it may suffer from severe information loss. In order to compensate for the problem of fine-grained local information loss, this embodiment further extracts pixel-aligned features by projecting the 3D point xc into the input view.
[0083] Finally, in this embodiment, geometry-aware coding is used to project point cloud features into three-plane potential initial position embeddings, and point cloud, three-plane features, and image features are used to decode the three-dimensional Gaussian distribution for new view rendering.
[0084] Specifically, Triplane feature query and Gaussian attribute decoding: For a given position x∈R in the point cloud 3 , querying features from Triplane T is the core step. Triplane consists of T xy , T xz , T yz The three axis-aligned orthogonal feature planes are formed by projecting the position x onto each plane (e.g., p xy is x in T xyPlane projection), and then use the trilinear interpolation function interpl to obtain the interpolation features from the corresponding plane, and finally splice the three interpolation features in sequence ( Operation) to obtain the final feature f t This process is based on strict mathematical interpolation and splicing logic, fully mining the feature information of each plane of the Triplane, and providing rich context for Gaussian attribute decoding. t Input MLPφ g Decode 3D Gaussian attribute (Δx′, α, s, q, sh) = φ g (x, f), where Δx′ is the position offset, α is the opacity, s and q define the anisotropic covariance, and sh is the spherical harmonic coefficient. MLP is trained with a large amount of data to learn the complex mapping relationship from features to Gaussian attributes, ensuring the accuracy and rationality of attribute decoding.
[0085] Position correction and local feature fusion improve accuracy: Considering that surface points may not be the best choice for 3D Gaussian representation, additional prediction of position offset Δx, the new position x = x + Δx can optimize Gaussian distribution positioning. At the same time, in order to strengthen the connection between the reconstruction effect and the input image, the projection perception condition is introduced. Based on the input camera posture π and point cloud P, the local projection feature f is calculated using the projection function P l =P(π,P), and with the Triplane characteristic f t Stitching. Local features include RGB colors to provide color information, masks to distinguish foreground and background, and two-dimensional distance transformations to refine spatial relationships. The stitching operation integrates multi-source features to improve the matching degree between 3D Gaussian attributes and local details of the image. For example, in the reconstruction of objects with complex textures, the Gaussian attributes corresponding to subtle texture changes can be accurately restored to avoid reconstruction distortion.
[0086] In this embodiment, Gaussian surfels representation is used in single view reconstruction, combining 3D Gaussian point optimization flexibility and surfel surface alignment to improve the reconstruction geometry quality, balance reconstruction speed and quality, and overcome the limitations of existing technologies.
[0087] Construct Gaussian facets from 3D Gaussian points, and assume that the covariance matrix of the 3D Gaussian points is Set its z scale to 0 and flatten the 3D Gaussian distribution ellipsoidal shape into a 2D ellipse. For example, the original 3D Gaussian distribution is:
[0088] in,
[0089] After transformation In this representation, the normal of the Gaussian face can be directly calculated as n i =R(r i)[:,2], and each Gaussian face is truncated into a 2D ellipse to clarify the direction for subsequent optimization, overcome the 3D Gaussian point normal ambiguity, and improve the optimization stability and surface alignment possibility. This design provides clear guidance for the optimizer. By using the local z-axis as the normal direction, the optimization stability and surface alignment ability are greatly improved.
[0090] like Figure 1 As shown, an embodiment of the present invention provides a single-view human body 3D reconstruction method based on Gaussian facets, comprising the following steps:
[0091] Step S300: training the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain a trained human body 3D Gaussian attribute prediction model.
[0092] Specifically, in an implementation of this embodiment, step S300 includes the following steps:
[0093] Step S301, randomly selecting images from different perspectives as sample data based on the processed human body data set;
[0094] Step S302, inputting the view image of the sample data and the corresponding camera parameters and sampling points into the human body 3D Gaussian attribute prediction model;
[0095] Step S303, extracting features from different dimensions using the hierarchical feature extraction module, fusing the extracted features through the feature fusion network, predicting Gaussian patch parameters based on the 3D Gaussian function decoder, and generating a predicted image of the target view based on volume rendering;
[0096] Step S304, calculate the photometric loss, mask loss, structural similarity index loss and depth-normal consistency loss, and optimize the model parameters according to the calculated losses to obtain the trained human 3D Gaussian attribute prediction model.
[0097] In this embodiment, the specific training process of the human body 3D Gaussian attribute prediction model is:
[0098] Data sampling: In each training iteration, for the same actor, randomly sample target and input view image pairs from preprocessed large-scale human datasets (e.g., THuman, RenderPeople, ZJU_MoCap, and HuMMan). From the 2,000 training set subjects in the THuman2.0 dataset, for each subject's 72 frames, randomly select images from different perspectives to form image pairs, thereby enriching the diversity of training data, improving the model's ability to understand and reconstruct human images from different perspectives, and enhancing generalization performance.
[0099] Point cloud initialization: Due to the lack of multi-view information, a "shell" structure sampling method based on the surface distance field and the human body SMPL template is proposed in this embodiment to provide the initial point cloud information required for single-view reconstruction. The "shell" structure greatly improves the sampling efficiency.
[0100] Generate distance field: First, calculate the shortest distance from each vertex on the Mesh surface to any point in space. Start with the known boundary (Mesh surface) and gradually expand outward to calculate the distance. It is similar to the process of wavefront propagation, with the Mesh surface as the initial wavefront, propagating outward at a certain speed (for example, unit speed), recording the time when each point is reached by the wavefront, and this time can be converted into distance.
[0101] Sampling point generation: Once the distance field is constructed, sampling can be performed within a certain distance range from the Mesh surface. For example, a distance threshold is set and all points in the distance field with distance values within the interval are found as sampling points. This method can accurately control the distance from the sampling point to the Mesh surface and can evenly cover the area near the surface.
[0102] Forward propagation: The input view image and its corresponding camera parameters, as well as the sampling points are input into the model. Subsequently, features are extracted from different dimensions through the hierarchical feature extraction module of the model. These features include global features, point-level features, and pixel alignment features. These features are fused through a feature fusion transformer (i.e., a feature fusion network), and finally the fused features are input into a 3D Gaussian function decoder to predict the spherical harmonics, opacity, covariance and other parameters of the Gaussian patch; in this embodiment, a predicted image of the target view is generated based on volume rendering, completing a forward propagation process to achieve end-to-end mapping from the input image to the predicted image.
[0103] Loss calculation and back propagation: photometric loss L p : Same as in 3DGS, based on predicted images And the corresponding real target image C(r). The calculation basis is the difference between the two in the pixel color dimension, through the formula In this formula, A set of pixels representing a Gaussian projection. The core purpose of introducing this loss function is to drive the model to accurately learn color and texture information, measure image quality differences at the pixel level, and drive Gaussian surface optimization to fit the input image features.
[0104] At the same time, the mask loss is calculated by closely combining the human region mask The calculation formula is: In this formula, represents the predicted cumulative volume density, while M(r) is the true binary mask label.
[0105] Further application of SSIM loss The similarity between the predicted image and the real image is measured from the unique perspective of structural similarity. SSIM, or structural similarity index, is essentially to encourage the model to focus not only on pixel-level differences during the learning process.
[0106] Depth-normal consistency loss L c :according to Calculate, use functions V(·) and N(·) to convert pixels and depth into 3D points and calculate normals, and force rendering depth and Rendering Normal When one of the center depth or normal is accurate, it can assist in correcting the other parameter, solving the gradient vanishing and depth-normal ambiguity problems in Gaussian facet optimization, ensuring the correct optimization direction, and improving the quality of the reconstructed surface.
[0107] In this embodiment, through the above human body 3D Gaussian attribute prediction model training process, a trained human body 3D Gaussian attribute prediction model is obtained. The reconstruction result obtained by this model combines the 3D Gaussian point optimization flexibility and the facet surface alignment to improve the reconstruction geometry quality.
[0108] like Figure 1 As shown, the embodiment of the present invention provides a single-view human body 3D reconstruction method based on Gaussian surface elements, which also includes the following steps:
[0109] Step S400, performing three-dimensional reconstruction on the input human body picture based on the trained human body 3D Gaussian attribute prediction model, and outputting a reconstructed human body three-dimensional image.
[0110] Based on the human body 3D Gaussian attribute prediction model trained by the above human body 3D reconstruction model training method, this embodiment also proposes a human body 3D reconstruction method, which is implemented based on the framework of the trained human body 3D Gaussian attribute prediction model, such as Figure 2 As shown in FIG, the method firstly needs to input a human body picture, then use the trained regression network to estimate the SMPL model parameters, and use the distance field method to generate the SMPL template Mesh surface distance field, so as to obtain points that can evenly cover the area near the surface as the initial point cloud.
[0111] In one implementation of this embodiment, the three-dimensional reconstruction of the input human body image based on the trained human body 3D Gaussian attribute prediction model includes: inputting the human body image, estimating explicit human model parameters using a trained regression network, and generating an explicit human template mesh surface distance field through a distance field method, obtaining points that evenly cover the area near the surface, and obtaining an initial point cloud.
[0112] In this embodiment, a human body image is input and the parameters of the SMPL model are estimated through a regression network. The regression network is trained to learn the mapping from image features to SMPL shape and posture parameters.
[0113] After that, the obtained SMPL parameters are used to generate the initial point cloud: First, the distance field method is used to generate the SMPL template Mesh surface distance field. Sampling is performed within a certain distance range from the Mesh surface. For example, a distance threshold is set to find all points in the distance field whose distance values are within the interval. The points in the area that can evenly cover the surface are obtained as the initial point cloud.
[0114] After obtaining the initial point cloud data, this embodiment compresses the human body image into a compact latent code for global encoding, encodes the image with camera parameters into a latent feature tag with the help of the pre-trained ViT model, and then encodes the implicit feature domain based on the transformer three-plane decoder to encode the three-dimensional Gaussian attributes. At the same time, with the help of the SMPL model, its vertices are projected into the two-dimensional feature map of the input image to extract the features of each point, and sparse three-dimensional convolution is used after voxelization. In addition, pixel alignment feature processing and point-level feature space convolution are used to enhance local features and compensate for information missing problems. Geometric perception coding is also used to project point cloud features, and finally multiple types of features are combined to decode Gaussian face element attributes.
[0115] Specifically, in an implementation of this embodiment, step S400 includes the following steps:
[0116] Step S401, compressing the input human body image into potential features through a neural network, and performing global encoding through a three-plane method;
[0117] Step S402, projecting the vertices of the explicit human model into the two-dimensional feature map of the input human body image, extracting the features of each point, and performing sparse three-dimensional convolution processing after voxelization;
[0118] Step S403, encode the corresponding color of the point cloud, splice it with the processed features, enhance the local features, and project the point cloud features using geometric perception coding, input it into the tansformer for decoding to obtain Gaussian facets.
[0119] In this embodiment, the image input human body 3D Gaussian Splatting model mainly includes a hierarchical feature extraction module, a feature fusion network and a 3D Gaussian function decoder.
[0120] Global feature extraction: The entire human image is compressed into a compact latent code for global encoding. Given an image with camera parameters, they are first encoded into a set of latent feature markers using a pre-trained ViT model. A transformer-based three-plane encoder is used to encode an implicit feature domain that encodes three-dimensional Gaussian properties. The three-plane T includes three axially arranged orthogonal feature planes {Txy, Txz, Tyz}.
[0121] Fusion of voxel features: Through an underlying SMPL model, the SMPL vertices can be projected into the 2D feature map of the input image, the features of each point can be extracted, and the voxels can be converted into a sparse 3D volume tensor, which can then be further processed using sparse 3D convolution.
[0122] Pixel feature fusion: pixel alignment features. Point-level feature spatial convolution is used to enhance local features and compensate for the problem of fine-grained local information loss. In this embodiment, the pixel alignment features are further extracted by projecting the 3D point xc into the input view.
[0123] Finally, this embodiment uses geometry-aware coding to project point cloud features into three-plane potential initial position embeddings, and uses point cloud, three-plane features, and image features to decode the three-dimensional Gaussian distribution for new view rendering.
[0124] Specifically, in an implementation of this embodiment, step S403 includes the following steps:
[0125] Step S403a, performing feature query based on the three-plane feature representation method, projecting the given position in the point cloud onto each plane, and obtaining interpolation features from the corresponding plane using a trilinear interpolation function, and sequentially concatenating all interpolation features to obtain the final feature;
[0126] Step S403b, inputting the given position in the point cloud and the final feature into the MLP neural network, decoding to obtain 3D Gaussian attributes; wherein the 3D Gaussian attributes include: position offset, opacity, anisotropic covariance and spherical harmonic coefficients;
[0127] Step S403c: constructing the Gaussian facet from the 3D Gaussian points according to the 3D Gaussian attributes.
[0128] In this embodiment, Triplane feature query and Gaussian attribute decoding: For a given position x∈R in the point cloud 3 , querying features from Triplane T is the core step. Triplane consists of T xy、 T xz , T yz The three axis-aligned orthogonal feature planes are formed by projecting the position x onto each plane (e.g., pxy is x in T xy Plane projection), and then use the trilinear interpolation function interpl to obtain the interpolation features from the corresponding plane, and finally splice the three interpolation features in sequence ( Operation) to obtain the final feature f t This process is based on strict mathematical interpolation and splicing logic, fully mining the feature information of each plane of the Triplane, and providing rich context for Gaussian attribute decoding. t Input MLPφ g Decode 3D Gaussian attribute (Δx′, α, s, q, sh) = φ g (x,f), where Δx′ is the position offset, α is the opacity, s and q define the anisotropic covariance, and sh is the spherical harmonic coefficient. MLP has learned the complex mapping relationship from features to Gaussian attributes through a large amount of data training, ensuring the accuracy and rationality of attribute decoding. At the same time, the position offset Δx is additionally predicted, and the new position x=x+Δx can optimize the Gaussian distribution positioning.
[0129] Construct Gaussian facets from 3D Gaussian points, and assume that the covariance matrix of the 3D Gaussian points is Set its z scale to 0 and flatten the 3D Gaussian distribution ellipsoidal shape into a 2D ellipse. For example, the original 3D Gaussian distribution is:
[0130] in,
[0131] After transformation In this representation, the normal of the Gaussian face can be directly calculated as n i =R(r i )[:,2], and each Gaussian face is truncated into a 2D ellipse to clarify the direction for subsequent optimization, overcome the normal ambiguity of the 3D Gaussian point, and improve the optimization stability and surface alignment possibility.
[0132] In this embodiment, through the processing of the above hierarchical feature extraction module, feature fusion network and 3D Gaussian function decoder, Gaussian face elements are decoded; based on the Gaussian face elements, new perspective generation and human body mesh extraction are performed to obtain a high-quality single-view human body three-dimensional reconstructed image.
[0133] Specifically, in an implementation of this embodiment, step S400 further includes the following steps:
[0134] Step S404, generating an image, a depth map, and a normal map based on the predicted Gaussian surface elements;
[0135] Step S405, using the generated depth map and normal map as input, using the filtered Poisson reconstruction algorithm, building a point cloud model based on the generated depth map, and realizing the conversion from discrete point cloud to continuous surface representation by solving the Poisson equation fitting function gradient to the input normal field, and reconstructing the human body mesh.
[0136] In this embodiment, the predicted Gaussian surface elements are mathematically calculated and fused to generate images and geometric information.
[0137] During the rendering process, for each pixel u in the image, its color is determined by the weighted contribution of the surrounding Gaussian pixels, achieved through alpha blending. By formula Calculate, where α i =G ′ (u;u i ,∑ i ′ ) i This G ′ is the Gaussian function reparameterized by 3D Gaussian in 2D ray space, that is:
[0138]
[0139] in w k is the view transformation matrix of input image k, J k This is an affine approximation of the projective transformation. This calculation determines the pixel color based on the Gaussian bin locations, covariances, and other properties so that the rendered image reflects the appearance of the scene.
[0140] depth With normal The calculation is similar, the formulas are:
[0141] Depth calculation takes into account the 2D elliptical characteristics of Gaussian surface elements and accurately calculates pixel depth based on the intersection of rays and ellipses;
[0142] formula middle is the inverse Jacobian matrix that maps the key image space pixels to Gaussian bin tangent planes.
[0143] After obtaining the depth map, due to the errors in the rendered depth map (especially at depth discontinuities), the volume cutting technology is used for optimization. A voxel grid is constructed within the bounding box of the target object, and the Gaussian ellipse is traversed to calculate the weighted opacity of its intersection with the voxel (according to the Gaussian function). If it is far from the surface, it is pruned to remove the wrong 3D points and improve the quality of the depth map.
[0144] Use Poisson reconstruction to obtain the human body mesh: take the processed depth map and the corresponding normal map as input, adopt the screening Poisson reconstruction algorithm, build a point cloud model based on the depth map, regard the point cloud as a sampling of the zero level set of the indicator function, and solve the Poisson equation to fit the function gradient to the input normal field to achieve the conversion from discrete point cloud to continuous surface representation, and reconstruct a high-quality surface mesh.
[0145] In this embodiment, a hierarchical multi-scale 3D transformer feature fusion network is proposed for predicting Gaussian attributes. This is a new hybrid representation that uses explicit and implicit representations for fast and high-quality single-view reconstruction. In addition, in this embodiment, 2D Gaussian facets are used for single-view human reconstruction to solve the 3D Gaussian depth distortion problem and facilitate normal supervision. The reconstruction result combines the flexibility of 3D Gaussian point optimization with the facet surface alignment to improve the quality of reconstruction geometry and balance reconstruction speed and quality.
[0146] This embodiment achieves the following technical effects through the above technical solution:
[0147] This embodiment collects image data from multiple human data sets, and can adopt targeted preprocessing strategies for preprocessing according to the data characteristics corresponding to each image data; and by constructing a human 3D Gaussian attribute prediction model, the human 3D Gaussian attribute prediction model can be trained based on the processed human data set, so as to perform three-dimensional reconstruction of the input human body picture based on the trained human 3D Gaussian attribute prediction model, and output a reconstructed three-dimensional image of the human body. The reconstruction result obtained by the present invention combines the optimization flexibility of 3D Gaussian points and the surface alignment of face elements, improves the quality of reconstruction geometry, takes into account the reconstruction speed and quality, greatly reduces the cost of virtual digital human production, and improves the efficiency of accurately constructing dynamic human body models. The generated human body model has significant value in the fields of virtual anchors, digital avatars, movies and game production.
[0148] Exemplary Devices
[0149] Based on the above embodiments, the present invention further provides a single-view human body 3D reconstruction system based on Gaussian facets, comprising:
[0150] A preprocessing module is used to collect image data from multiple human body data sets, and to perform preprocessing using a targeted preprocessing strategy according to data characteristics corresponding to each image data, so as to obtain a processed human body data set;
[0151] A model building module is used to build a 3D Gaussian attribute prediction model for a human body, wherein the 3D Gaussian attribute prediction model for a human body includes: a hierarchical feature extraction module, a feature fusion network, and a 3D Gaussian function decoder;
[0152] A model training module, used for training the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain a trained human body 3D Gaussian attribute prediction model;
[0153] The human body three-dimensional image reconstruction module is used to perform three-dimensional reconstruction on the input human body picture based on the trained human body 3D Gaussian attribute prediction model, and output a reconstructed human body three-dimensional image.
[0154] This embodiment achieves the following technical effects through the above technical solution:
[0155] This embodiment collects image data from multiple human data sets, and can adopt targeted preprocessing strategies for preprocessing according to the data characteristics corresponding to each image data; and by constructing a human 3D Gaussian attribute prediction model, the human 3D Gaussian attribute prediction model can be trained based on the processed human data set, so as to perform three-dimensional reconstruction of the input human body picture based on the trained human 3D Gaussian attribute prediction model, and output a reconstructed three-dimensional image of the human body. The reconstruction result obtained by the present invention combines the optimization flexibility of 3D Gaussian points and the surface alignment of face elements, improves the quality of reconstruction geometry, takes into account the reconstruction speed and quality, greatly reduces the cost of virtual digital human production, and improves the efficiency of accurately constructing dynamic human body models. The generated human body model has significant value in the fields of virtual anchors, digital avatars, movies and game production.
[0156] Based on the above embodiment, the present invention further provides a terminal, whose principle block diagram can be as follows: Figure 3 shown.
[0157] The terminal includes: a processor, a memory, an interface, a display screen and a communication module connected through a system bus; wherein the processor of the terminal is used to provide computing and control capabilities; the memory of the terminal includes a storage medium and an internal memory; the storage medium stores an operating system and a computer program; the internal memory provides an environment for the operation of the operating system and the computer program in the storage medium; the interface is used to connect to external devices; the display screen is used to display corresponding information; and the communication module is used to communicate with a cloud server or other devices.
[0158] When the computer program is executed by a processor, it is used to implement the operation of a single-view human body three-dimensional reconstruction method based on Gaussian surface elements.
[0159] It can be understood by those skilled in the art that Figure 3 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal to which the scheme of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0160] In one embodiment, a terminal is provided, which includes: a processor and a memory, wherein the memory stores a single-view human body three-dimensional reconstruction program based on Gaussian surface elements, and the single-view human body three-dimensional reconstruction program based on Gaussian surface elements is used to implement the operation of the above-mentioned single-view human body three-dimensional reconstruction method based on Gaussian surface elements when executed by the processor.
[0161] In one embodiment, a storage medium is provided, wherein the storage medium stores a single-view human body three-dimensional reconstruction program based on Gaussian surface elements, and when the single-view human body three-dimensional reconstruction program based on Gaussian surface elements is executed by a processor, it is used to implement the operation of the single-view human body three-dimensional reconstruction method based on Gaussian surface elements as described above.
[0162] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and volatile memory.
[0163] In summary, the present invention provides a single-view human body 3D reconstruction method based on Gaussian facets, including: collecting image data from multiple human body data sets, preprocessing with targeted preprocessing strategies according to the data characteristics corresponding to each image data, and obtaining a processed human body data set; constructing a human body 3D Gaussian attribute prediction model, the human body 3D Gaussian attribute prediction model including: a hierarchical feature extraction module, a feature fusion network, and a 3D Gaussian function decoder; training the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain a trained human body 3D Gaussian attribute prediction model; performing 3D reconstruction on an input human body picture based on the trained human body 3D Gaussian attribute prediction model, and outputting a reconstructed human body 3D image. The present invention greatly reduces the cost of making virtual digital humans and improves the efficiency and accuracy of building dynamic human body models.
[0164] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A single-view human body 3D reconstruction method based on Gaussian facets, characterized in that: include: Collect image data from multiple human data sets, and pre-process them using targeted pre-processing strategies according to data characteristics corresponding to each image data to obtain a processed human data set; Constructing a 3D Gaussian attribute prediction model for a human body, wherein the 3D Gaussian attribute prediction model for a human body comprises: a hierarchical feature extraction module, a feature fusion network, and a 3D Gaussian function decoder; Training the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain a trained human body 3D Gaussian attribute prediction model; Based on the trained human body 3D Gaussian attribute prediction model, the input human body picture is three-dimensionally reconstructed, and a reconstructed human body three-dimensional image is output.
2. The single-view human body 3D reconstruction method based on Gaussian surface elements according to claim 1, characterized in that: The method of collecting image data from a plurality of human body data sets and preprocessing the image data using a targeted preprocessing strategy according to data characteristics corresponding to each image data to obtain a processed human body data set includes: Collecting corresponding image data from multiple human body datasets; For the image data of the human body scan dataset, the camera pose is fixed, and the rendered image, image mask, depth map, normal map and corresponding SMPL human body template parameters are obtained from the human body scan model to obtain the corresponding processed human body dataset.
3. The single-view human body 3D reconstruction method based on Gaussian surface elements according to claim 2, characterized in that: For the image data of the human body scan data set, the camera pose is fixed, and the rendered image, image mask, depth map, normal map and corresponding SMPL human body template parameters are obtained from the human body scan model, including: The camera posture is fixed to obtain a fixed number of frames of random sampling of the image, and the spherical harmonic coefficients of the radiation transmission of the human body scanning model are calculated, and the rendering image is obtained by scanning according to the fixed number of frames and the spherical harmonic coefficients; Based on the human posture estimation algorithm, the human body key points of the human body scanning model are estimated, and the estimated human body key points are input into the MLP neural network to predict and obtain corresponding SMPL parameters; The vertices of the human body scan model are transformed and projected onto the camera plane, and the depth values of the projection points are calculated, and the depth values are stored in a two-dimensional array with the same resolution as the projection plane to obtain the depth map; The vertex normals in the human body scan model are calculated, and the vertex normals are converted into a texture space, and the normal map is obtained after interpolation and mapping processing.
4. The single-view human body 3D reconstruction method based on Gaussian surface elements according to claim 1, characterized in that: The step of training the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain a trained human body 3D Gaussian attribute prediction model comprises: Randomly selecting images from different perspectives as sample data based on the processed human body data set; Inputting the view image of the sample data and the corresponding camera parameters and sampling points into the human body 3D Gaussian attribute prediction model; Extracting features from different dimensions using the hierarchical feature extraction module, fusing the extracted features through the feature fusion network, predicting Gaussian patch parameters based on the 3D Gaussian function decoder, and generating a predicted image of the target view based on volume rendering; The photometric loss, mask loss, structural similarity index loss, and depth-normal consistency loss are calculated, and the model parameters are optimized according to the calculated losses to obtain the trained human 3D Gaussian attribute prediction model.
5. The single-view human body 3D reconstruction method based on Gaussian surface elements according to claim 1, characterized in that: The three-dimensional reconstruction of the input human body image based on the trained human body 3D Gaussian attribute prediction model comprises: The human body picture is input, and the parameters of the explicit human model are estimated by using a trained regression network. The distance field method is used to generate the distance field of the explicit human template mesh surface, and points that evenly cover the area near the surface are obtained to obtain an initial point cloud.
6. The single-view human body 3D reconstruction method based on Gaussian surface elements according to claim 1, characterized in that: The three-dimensional reconstruction of the input human body image based on the trained human body 3D Gaussian attribute prediction model includes: Compressing the input human body image into potential features through a neural network, and globally encoding it through a three-plane method; Projecting the vertices of the explicit human model into the two-dimensional feature map of the input human body image, extracting features of each point, and performing sparse three-dimensional convolution processing after voxelization; Encode the corresponding colors of the point cloud, concatenate them with the processed features, enhance the local features, and project the point cloud features using geometric perception coding, input them into the tansformer for decoding to obtain Gaussian facets; Generate images, depth maps, and normal maps based on the predicted Gaussian surfaces; The generated depth map and normal map are used as input, and a screened Poisson reconstruction algorithm is used to build a point cloud model based on the generated depth map. The Poisson equation is solved to fit the function gradient to the input normal field to achieve the conversion from discrete point cloud to continuous surface representation, and the human body mesh is reconstructed.
7. The single-view human body 3D reconstruction method based on Gaussian surface elements according to claim 6, characterized in that: The method of projecting point cloud features by using geometric perception coding and decoding multiple types of features to obtain Gaussian facets includes: Based on the three-plane feature representation method, feature query is performed. The given position in the point cloud is projected onto each plane, and the interpolation feature is obtained from the corresponding plane using the trilinear interpolation function. All interpolation features are sequentially concatenated to obtain the final feature. Inputting the given position in the point cloud and the final feature into the MLP neural network, decoding to obtain 3D Gaussian attributes; wherein the 3D Gaussian attributes include: position offset, opacity, anisotropic covariance and spherical harmonic coefficients; The Gaussian surface element is constructed from the 3D Gaussian points according to the 3D Gaussian attributes.
8. A single-view human body 3D reconstruction system based on Gaussian facets, characterized in that: include: A preprocessing module is used to collect image data from multiple human body data sets, and to perform preprocessing using a targeted preprocessing strategy according to data characteristics corresponding to each image data, so as to obtain a processed human body data set; A model building module is used to build a 3D Gaussian attribute prediction model for a human body, wherein the 3D Gaussian attribute prediction model for a human body includes: a hierarchical feature extraction module, a feature fusion network, and a 3D Gaussian function decoder; A model training module, used for training the human body 3D Gaussian attribute prediction model based on the processed human body data set to obtain a trained human body 3D Gaussian attribute prediction model; The human body three-dimensional image reconstruction module is used to perform three-dimensional reconstruction on the input human body picture based on the trained human body 3D Gaussian attribute prediction model, and output a reconstructed human body three-dimensional image.
9. A terminal, characterized in that: include: A processor and a memory, wherein the memory stores a single-view human body 3D reconstruction program based on Gaussian surfaces, and when the single-view human body 3D reconstruction program based on Gaussian surfaces is executed by the processor, it is used to implement the operation of the single-view human body 3D reconstruction method based on Gaussian surfaces as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a single-view human body three-dimensional reconstruction program based on Gaussian surfaces, and when the single-view human body three-dimensional reconstruction program based on Gaussian surfaces is executed by a processor, it is used to implement the operation of the single-view human body three-dimensional reconstruction method based on Gaussian surfaces as described in any one of claims 1-7.
Citation Information
Patent Citations
Human body new viewpoint rendering method based on pixel alignment 3D Gaussian point cloud representation
CN118212337A
Human body three-dimensional imaging method and system
US20160300383A1
Multi-distribution entropy modeling of latent features in image and video coding using neural networks
US20240163485A1
Cited By
VR-based garment rendering method and system
CN120318392A
Four-dimensional content synthesis method and device, electronic equipment and storage medium
CN120912732A
Three-dimensional Gaussian modeling method and system based on Poincare sphere and three-plane representation
CN122066910A
A 3D Gaussian Modeling Method and System Based on Poincaré Sphere and Three-Plane Representation
CN122066910B
Three-dimensional human body reconstruction method and system based on three-dimensional Gaussian splashing
CN122156415A