Sparse visual angle three-dimensional reconstruction method based on 3DGS
By using a sparse viewpoint 3D reconstruction method based on 3DGS, high-quality Gaussian field models are generated using FPN, 2D U-Net and MLP networks. This solves the problems of traditional methods in terms of illumination variation, viewpoint variation and texture complexity, and achieves efficient and low-cost dynamic scene reconstruction.
Patent Information
- Application Number
- CN202511079130.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-02
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional sparse-view 3D reconstruction methods perform poorly in terms of lighting variations, viewpoint variations, and texture complexity. They require expensive equipment and struggle to handle dynamic scenes, resulting in poor reconstruction quality and high costs.
A sparse-view 3D reconstruction method based on 3DGS is adopted. High-level semantic features are extracted through FPN, 2D U-Net and MLP networks, and a high-quality Gaussian field model is generated by combining the initial point cloud coordinates. The Gaussian distribution parameters are optimized through an adaptive mechanism to improve the reconstruction accuracy and efficiency.
It achieves high-quality 3D reconstruction under sparse perspective, reduces equipment requirements and operational complexity, is suitable for rendering real-world scenes, and improves reconstruction efficiency and accuracy.
Smart Images

Figure CN120807805A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer vision and three-dimensional reconstruction technology, and relates to a sparse-view three-dimensional reconstruction method based on deep learning, which is suitable for the construction of virtual environments or augmented reality scenes, or unmanned aerial vehicle aerial photography and map construction under the condition of limited shooting angles. BACKGROUND
[0002] Three-dimensional reconstruction technology is a technology for generating three-dimensional models from two-dimensional images or other data sources. Three-dimensional reconstruction technology can convert complex real-world scenes into interactive digital spaces, providing users with richer experiences and interaction capabilities.
[0003] Sparse-view three-dimensional reconstruction technology has wide application prospects in many fields, such as medical imaging, industrial manufacturing, and visual navigation. In medical imaging, sparse-view CT reconstruction technology can reduce radiation exposure and improve reconstruction accuracy. In industrial manufacturing, sparse-view three-dimensional reconstruction technology can be used to efficiently obtain accurate three-dimensional models from a small number of images, thereby optimizing the design and production process.
[0004] Traditional sparse-view three-dimensional reconstruction methods rely on image feature matching, but this approach performs poorly in terms of lighting changes, viewing angle changes, and texture complexity. For example, artificially defined geometric features such as SIFT are sensitive to lighting and viewing angle changes, and cannot effectively match weak texture or repetitive texture areas. In addition, traditional methods often require expensive and complex equipment, making the data collection process time-consuming, labor-intensive, and costly. Furthermore, obtaining high-quality three-dimensional models relies on high-quality manual data collection, which increases the difficulty and cost of operation. Traditional methods are mainly suitable for modeling static objects, and are difficult to model dynamic scenes. Dynamic objects require a dense camera array to capture enough viewing angle information, which increases the demand for equipment and operational complexity. Deep learning models can overcome these limitations through end-to-end learning mechanisms, improving reconstruction accuracy.
[0005] Three-dimensional Gaussian Splatting (3DGS) emerges as a promising alternative, providing an explicit representation for three-dimensional scene synthesis that bypasses the computational burden of NeRF voxel rendering. By leveraging Gaussian primitives, 3DGS facilitates faster and more efficient scene reconstruction and novel view synthesis. The explicit nature of Gaussian representation allows for direct sampling and manipulation in space, improving rendering efficiency and flexibility for three-dimensional scene editing. 3DGS excels in novel view synthesis, combining depth priors with generative models to reduce background collapse and lens floaters, thereby improving rendering quality.
[0006] In the aspect of sparse-view 3D reconstruction, the main challenges faced by 3DGS include initialization failure, overfitting to input images, and missing details, etc. These problems lead to poor and incomplete reconstruction results generated under sparse-view. For example, due to insufficient overlap between input images, traditional structure from motion (SfM) techniques cannot successfully handle sparse-view scenes, thereby affecting the initialization of 3DGS. SUMMARY
[0007] Based on the above problems, the present application provides a sparse-view 3D reconstruction method based on 3DGS, which can obtain high 3D scene reconstruction quality, while being simple to calculate and easy to implement in engineering.
[0008] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:
[0009] Step 1, input sparse-view images to obtain initial point cloud coordinates. Wherein, the depth, pose and confidence of the input sparse-view images are predicted, 3D point clouds are generated by back projection, the point clouds of different views are converted to the same coordinate system according to multi-view fusion, and the initial point cloud is obtained by filtering low-confidence points;
[0010] Step 2, the initial point cloud generated in step 1 is down-sampled by farthest point sampling, which can significantly reduce the calculation amount while preserving the geometric features;
[0011] Step 3, the input images are input into the Gaussian parameter extraction network to obtain high-quality Gaussian ellipsoid parameters.
[0012] Step 3.1, high-level semantic features of multi-view are extracted, and high-level features containing rich semantics are extracted from sparse multi-view images through multi-scale feature fusion. The features of different scales of each picture are extracted by using a CNN backbone network, and these features are fused through the top-down path of the feature extraction network to enhance the semantic expression ability of high-level features;
[0013] Step 3.2, the high-level semantic features generated in step 3.1 are transformed into the same coordinate system by camera pose, and then multi-view features are obtained by pooling; Figure One
[0014] Step 3.3, the features obtained in step 3.2 are enhanced by 2D U-Net to enhance feature association. The encoding-decoding structure and skip connection of U-Net are used to fuse multi-view features and strengthen the semantic consistency across views; Figure One
[0015] Step 3.4, the enhanced features in step 3.3 are decoded by using an MLP decoder to obtain high-quality scale, rotation, spherical harmonic, and opacity parameters;
[0016] Step 4, fuse the Gaussian ellipsoid parameters obtained in step 3 with the point cloud coordinates generated in step 2 to obtain a complete Gaussian field model, with one point representing a 3D Gaussian sphere.
[0017] Step 5, convert the Gaussian field model obtained in step 4 to a 2D image plane, sort the Gaussian distributions according to depth information to ensure correct occlusion relationships. Mix the sorted Gaussian distributions to generate the final 2D image. Compare the generated 2D image with the real image to calculate the image loss (such as L1 loss or SSIM loss) to measure the reconstruction quality of the model.
[0018] Step 6, calculate the gradient according to the loss function obtained in step 5 and update the parameters of the Gaussian distribution, such as position, covariance, opacity, etc. According to the gradient information, the Gaussian distribution is pruned (remove Gaussians with smaller contributions) or densified (add new Gaussians) to enhance the detailed representation of the scene. This adaptive mechanism helps to improve the reconstruction accuracy while keeping the model simple.
[0019] Step 7, after training, input any camera shooting pose in the constructed 3D Gaussian model, and render the output scene optical information, depth information under the corresponding view angle, to describe the visual texture information and related scene structure information of the entire scene.
[0020] Beneficial effects
[0021] The present application uses three-dimensional Gaussian splashing to realize three-dimensional reconstruction model under sparse view angle condition, which has less constraint condition for scene imaging, does not need large-scale three-dimensional reconstruction dataset of a large number of similar scenes, and has lower requirement for the number of view angles under the condition of the current scene to be reconstructed. It is more suitable for user-friendly three-dimensional scene rendering in real situation.
[0022] In the present application, a Gaussian field initialization model is proposed. The present application uses FPN, 2D U-Net and MLP to constitute a Gaussian parameter initial extraction network, which integrates high-level and low-level semantic information and enhances the understanding of multi-scale features by the network. Then, the complete Gaussian field model is obtained by combining the initial point cloud coordinates, so as to improve the quality of the Gaussian ellipsoid. Compared with traditional multi-view three-dimensional reconstruction technology, the reconstruction efficiency of the present application is higher. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 The flowchart of the sparse view three-dimensional reconstruction method based on 3DGS of the present application.
[0024] Figure 2 This is the flow chart of the Gaussian field initialization model of the present invention. DETAILED DESCRIPTION
[0025] The present invention will be further described below with reference to the accompanying drawings and examples. The embodiments described are some, but not all, of the present invention. The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but rather merely represents selected embodiments of the present invention.
[0026] like Figure 1 As shown, the specific steps of the present invention are as follows: First, a high-quality Gaussian field model is obtained based on the Gaussian field initialization model. The specific structure is as follows: Figure Two As shown in the figure, the initial point cloud coordinates are calculated by DUSt3R, and the farthest point sampling is used to reduce the amount of calculation; FPN extracts high-level semantic features of multiple views, and uses the pose to transform the multi-view features to the same coordinate system, and pools the multi-view features. Figure One The algorithm then uses a 2D U-net to enhance multi-view feature associations, decodes and obtains high-quality scale, rotation, spherical harmonics, and opacity parameters, and fuses the point cloud coordinates to obtain a complete Gaussian field model. Three-dimensional Gaussian splattering is then used as the main framework of the algorithm to project and rasterize the 3D Gaussian sphere at different viewpoints, and adaptively clone and prune the 3D Gaussian spheres with large gradient variations. The output is the optical and depth information of the rendered image at each viewpoint. Then, during gradient backpropagation, the L1 norm loss is used to supervise the absolute value deviation of the pixel values between the rendered RGB image and the true RGB value.
[0027] Specifically, the method comprises the following steps:
[0028] Step 1: Input sparse view images to obtain the initial point cloud coordinates. The depth, pose, and confidence of the input sparse view images are jointly predicted, and a 3D point cloud is generated by backprojection: for each pixel p(u,v) in the image, its 3D coordinates are calculated:
[0029]
[0030] Internal parameter processing: If the internal parameters are known, the K matrix is used directly. If the internal parameters are unknown, DUST3R predicts the focal length f and the principal point (cx, cy) to construct the K matrix.
[0031] According to multi-view fusion, the point clouds of different views are converted to the same coordinate system, and the low-confidence points are filtered to obtain the initial point cloud; the point clouds of different views are converted to the same coordinate system (with the first frame as a reference):
[0032]
[0033] Fusion of overlapping points based on confidence weights:
[0034]
[0035] Through uncertainty modeling and joint optimization strategies, the initial point cloud can be robustly generated even if the input image is sparse, providing reliable input for subsequent fine reconstruction.
[0036] Step 2: Based on the initial point cloud generated in step 1, the initial point cloud is downsampled using the farthest point sampling (FPS), and the point cloud P={P1,P2,...,P N} select M points (M≪N) so that the sampling points cover the original geometry as much as possible. Original point cloud , target point number M, randomly select an initial point Join the sampling set S. For each step (from 1 to M−1): Calculate the distance matrix: for all points , calculate the minimum distance to the current sample S:
[0037]
[0038] Select the farthest point: The largest point joins, when Stop when
[0039] Step 3: A Backbone network extracts multi-scale features from the input sparse viewport image. Each stage outputs feature maps of different scales: C2 local structural features: [N, 256, H / 4, W / 4] (high resolution, low-level details), C3 component-level features: [N, 512, H / 8, W / 8], C4 object-level semantic features: [N, 1024, H / 16, W / 16], and C5 scene-level global semantic features: [N, 2048, H / 32, W / 32] (low resolution, high-level semantics). These four layers correspond to the P2, P3, P4, and P5 input layers of the FPN. P5 is upsampled by a factor of 2 and added to the principal elements of C4 to generate P4. This process is repeated for P2, forming a pyramid structure. These features are fused through the top-down path of the FPN to enhance the semantic expressiveness of high-level features. This process combines low-level geometric details with high-level semantic information to adapt to scale variations across different viewports, and compensates for geometric ambiguity caused by insufficient viewports through high-level semantic features.
[0040] Step 4: Project the 2D image features generated in step 3 into a unified 3D space coordinate system and generate a multi-view feature map through pooling. Figure One Consistent 3D feature representation; for each pixel (u, v) on the feature map, calculate its 3D world coordinates :
[0041]
[0042]
[0043] Each 3D point Associate a feature vector . Defines a voxel grid in 3D space , where D, H, W are voxel resolutions, for each 3D point , find the corresponding voxel index :
[0044]
[0045] The features Accumulate to the voxel. Then use the maximum pooling to retain the most significant features and enhance high-frequency details. Given the features of N views And corresponding poses, multi-view pooling can be expressed as:
[0046]
[0047] in Is a view Projection to voxels A collection of pixels.
[0048] Step 5: The features obtained in step 4 are enhanced through 2D U-Net to enhance feature association, and the encoding-decoding structure and skip connection of U-Net are used to fuse multi-view Figure One Consistent features and strengthen semantic consistency across perspectives. The encoder consists of multiple convolutional layers and pooling layers, which gradually reduce the resolution and extract high-level semantic information. Each stage contains two 3×3 convolutions + ReLU, followed by 2×2 maximum pooling. As the depth increases, the receptive field expands, and the network can learn more global features. The decoder consists of transposed convolution and convolution layers, which gradually restore the resolution while fusing the low-level detail information of the encoder. Transposed convolution is used for upsampling, and the corresponding layer features of the encoder are spliced with the decoder features through jump connections to preserve spatial details. Jump connections can solve the problem of spatial information loss caused by downsampling, allowing the decoder to recover more accurate features;
[0049] Step 6: Use the MLP decoder to decode the features enhanced in step 5 to obtain high-quality scale, rotation, spherical harmonics, and opacity parameters;
[0050] Step 7: Fuse the parameters obtained in step 6 with the point cloud coordinates generated in step 2 to obtain a complete Gaussian field model, where one point represents a 3D Gaussian sphere.
[0051] Step 8, convert the Gaussian field model obtained in step 7 to the 2D image plane, sort the Gaussian distributions according to the depth information to ensure correct occlusion relationship. Mix the sorted Gaussian distributions to generate the final 2D image. Compare the generated 2D image with the real image, calculate the image loss (such as L1 loss or SSIM loss) to measure the reconstruction quality of the model.
[0052] Step 9, calculate the gradient according to the loss function obtained in step 8 and update the parameters of the Gaussian distribution such as position, covariance, opacity, etc. According to the gradient information, the Gaussian distribution is pruned (remove the Gaussian with smaller contribution) or densified (add new Gaussian) to enhance the detailed representation of the scene. This adaptive mechanism helps to improve the reconstruction accuracy while keeping the model simple.
[0053] Step 10, after training, the model can render the optical information and depth information of the scene at any viewing angle, describe the visual texture information of the entire scene, and related scene structure information.
[0054] Those skilled in the art can understand that the above description is only a preferred example of the application and is not intended to limit the application. Although the application has been described in detail, those skilled in the art can still modify the technical solutions described in the examples or replace some technical features with equivalent ones. Any modification, equivalent replacement, etc. within the spirit and principles of the application is included in the protection scope of the application.
Claims
1. A sparse view 3D reconstruction method based on 3DGS, characterized in that: The method comprises the following steps: Step 1: Input a sparse view image to obtain the initial point cloud coordinates. The depth, pose, and confidence of the input sparse view image are predicted, and a 3D point cloud is generated through backprojection. Based on multi-view fusion, the point clouds from different views are converted to the same coordinate system, and low-confidence points are filtered to obtain the initial point cloud. Step 2: Downsampling the initial point cloud generated in step 1 by sampling the farthest point, which can significantly reduce the amount of calculation while retaining the geometric features; Step 3: The input image is passed through the Gaussian parameter extraction network to obtain high-quality Gaussian ellipsoid parameters. Step 3.1: Extract high-level semantic features from multiple views. Through multi-scale feature fusion, high-level features with rich semantics are extracted from sparse multi-view images. A CNN backbone network is used to extract features at different scales for each image. These features are then fused from top to bottom through the feature extraction network to enhance the semantic expression capability of high-level features. Step 3.2: Based on the high-level semantic features generated in step 3.1, the multi-view features are transformed into the same coordinate system through the camera pose, and then the multi-view consistent features are obtained through pooling; Step 3.3: The features obtained in step 3.2 are passed through a 2D U-Net to enhance feature association. The U-Net's encoder-decoder structure and skip connections are used to fuse multi-view consistent features and strengthen semantic consistency across viewpoints. Step 3.4: Use the MLP decoder to decode the features enhanced in step 3.3 to obtain high-quality scale, rotation, spherical harmonics, and opacity parameters; Step 4: Fuse the Gaussian ellipsoid parameters obtained in step 3 with the point cloud coordinates generated in step 2 to obtain a complete Gaussian field model, where one point represents one 3D Gaussian sphere. Step 5: Convert the Gaussian field model obtained in Step 4 to a 2D image plane. Sort the Gaussian distributions based on depth information to ensure correct occlusion relationships. The sorted Gaussian distributions are blended to generate the final 2D image. The resulting 2D image is compared with the ground truth image, and an image loss (such as L1 loss or SSIM loss) is calculated to measure the model's reconstruction quality. Step 6: Calculate the gradient based on the loss function obtained in Step 5 and update the parameters of the Gaussian distribution, such as position, covariance, and opacity. Based on the gradient information, the Gaussian distribution is pruned (removing Gaussians with less contribution) or densified (adding new Gaussians) to enhance the representation of scene details. This adaptive mechanism helps improve reconstruction accuracy while maintaining model simplicity. Step 7. After training is completed, any camera shooting pose is input into the constructed 3D Gaussian model, and the scene optical information and depth information under the corresponding perspective are rendered and output to describe the visual texture information of the entire scene and related scene structure information.
2. The sparse view 3D reconstruction method based on 3DGS according to claim 1, characterized in that The method for obtaining the initial point cloud in step 1 uses DUSt3R instead of relying on COLMAP to generate the initial point cloud in the original 3DGS. DUSt3R can generate a dense and geometrically accurate initial point cloud. The traditional SfM algorithm cannot handle low-texture areas, while the initial point cloud generated by DUSt3R can cover weak-texture areas to reduce holes, and can still generate reasonable depth in dim or overexposed scenes. DUSt3R uses a point map as its core representation, which is a dense 2D field that contains information about 3D points. The point map provides a corresponding 3D point for each pixel, thereby establishing a direct correspondence between image pixels and 3D scene points. This method can output point clouds in seconds under GPU acceleration, which is significantly faster than traditional SfM.
3. The sparse view 3D reconstruction method based on 3DGS according to claim 1, characterized in that: The Gaussian parameter extraction network in step 3 replaces the original Gaussian, which directly inherits the 3D point coordinates output by the SfM to obtain the mean. This is highly dependent on the quality of the SfM point cloud. The sparsity and noise of the point cloud can lead to an uneven initial Gaussian distribution. Furthermore, the fixed spherical covariance and transparency of the 3DGS require extensive subsequent iterative optimization. The Gaussian parameter extraction network model architecture: First, high-level semantic features are extracted from sparse multi-view images using the Feature Pyramid Network (FPN). These multi-view features are then aligned to the world coordinate system through pose alignment. Cross-view information is then integrated through multi-view feature pooling. A 2D U-Net is then used to enhance the correlation and consistency between features. Finally, an MLP (Multi-Layer Perceptron) decoder is used to generate Gaussian parameters (including position, covariance, and transparency). The optimized Gaussian ellipsoid parameters have enhanced geometric representation capabilities, support anisotropic shapes (arbitrary rotation and scaling), accurately fit complex surfaces (such as thin walls and slender structures), and can dynamically adjust the Gaussian distribution to adapt to scene complexity.
4. The sparse view 3D reconstruction method based on 3DGS according to claim 1, characterized in that: Step 3.1 of the FPN multi-view semantic feature extraction process involves extracting multi-scale features from the input sparse viewport image using a Backbone network. Each stage outputs feature maps of different scales: C2 local structural features: [N, 256, H / 4, W / 4] (high resolution, low-level details), C3 component-level features: [N, 512, H / 8, W / 8], C4 object-level semantic features: [N, 1024, H / 16, W / 16], and C5 scene-level global semantic features: [N, 2048, H / 32, W / 32] (low resolution, high-level semantics). These four layers correspond to the P2, P3, P4, and P5 input layers of the FPN. P4 is generated by upsampling P5 by a factor of 2 and then added to the principal elements of C4. This is repeated to generate P2, forming a pyramid structure. This extracts semantically rich multi-scale features from sparse multi-view images, providing the foundation for subsequent cross-view fusion and Gaussian parameter decoding.
Citation Information
Cited By
Aerial photography data building height calculation method based on 3D Gaussian technology
CN121120744A
Automatic parameter identifying and optimizing system for reconstruction of Gaussian splash model
CN121353541A
Head reconstruction method and system based on diffusion model and three-dimensional Gaussian sputtering
CN121414991A
Multi-stereoscopic view volume rendering method and system based on 3D Gaussian sputtering
CN121482243A
Rapid 3DGS three-dimensional reconstruction method based on graph neural network anchor point cooperation
CN121639922A