Indoor real scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering
Through the improved 3D Gaussian sputtering method, the feedforward point cloud density-enhancing network and depth regularization loss function are used to solve the accuracy and efficiency of three-dimensional reconstruction in complex indoor scenes, and high-quality real-time three-dimensional reconstruction effect is achieved.
Patent Information
- Application Number
- CN202510350662.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing three-dimensional reconstruction technology has problems such as excessive model parameters, lack of real-time performance, low modeling accuracy and poor visual effects in complex indoor scenarios. Especially in complex indoor environments, insufficient detailed feature capture leads to a significant decline in reconstruction performance.
The improved 3D Gaussian sputtering (GD-3DGS) method is adopted to improve point cloud density by building a feed-forward point cloud density-enhancing network, combining depth regularization loss function and adaptive density control, optimize radiation field parameters, and enhance detailed characterization and structural consistency of complex scenarios.
The visual effect and modeling accuracy of three-dimensional reconstruction of indoor scenes are improved, the peak signal-to-noise ratio and structural similarity of the reconstruction view are significantly improved, floating point artifacts and fuzzy problems are solved, and real-time interactive application needs are met.
Smart Images

Figure CN120279159A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of three-dimensional reconstruction, and particularly relates to a method for indoor real-scene three-dimensional reconstruction based on improved 3D Gaussian sputtering. Background Technique
[0002] Three-dimensional reconstruction technology is an interdisciplinary field of computer vision and computer graphics, aiming to restore the three-dimensional geometric structure and surface properties of a scene through sensor data (such as images, depth information, etc.). In recent years, with the development of deep learning, multi-sensor fusion, and high-performance computing, three-dimensional reconstruction technology has moved from the laboratory to practical applications. China's "14th Five-Year Plan" clearly states that the cognition and reconstruction of three-dimensional scenes in digital city construction are listed as an important direction for national industrial development. As an important part of digital city construction, three-dimensional reconstruction technology is widely used in indoor navigation and location services, simultaneous localization and mapping (SLAM), virtual reality (VR), digital twins, etc., laying a solid foundation for the construction of smart cities and providing strong support for the construction of real-scene three-dimensional China.
[0003] The evolution of three-dimensional reconstruction technology mainly relies on the progress of sensor technology and the innovation of algorithm frameworks: on the one hand, sensors such as RGB-D cameras (such as Kinect) and lidar (LiDAR) have significantly reduced the threshold for obtaining three-dimensional data; on the other hand, the algorithm framework has experienced a transformation from traditional multi-view geometry methods (such as SFM / SLAM relying on feature matching and bundle adjustment optimization) to deep learning-driven (such as neural implicit representation NeRF, three-dimensional Gaussian sputtering 3DGS). Although the RGB-D sensor-based solution can directly obtain depth information, it faces challenges such as cumulative sensor measurement errors, insufficient reconstruction accuracy for long-distance or transparent objects, and is prone to model distortion due to feature matching failures in indoor low-texture scenes; the three-dimensional reconstruction method based on NeRF achieves good visual effects, however, its computational complexity increases sharply with the scene scale, resulting in low rendering efficiency and long training time, facing a bottleneck in real-time performance, which limits its application in interactive scenes such as VR and AR; the three-dimensional reconstruction technology based on 3DGS optimizes rendering through parallel projection, breaking through the limitation of real-time performance and achieving real-time rendering. However, in the face of the intricate indoor scene 3DGS, due to the lack of sufficient point cloud constraints, the captured detailed features are extremely sparse or even missing, and the modeling effect deteriorates sharply, accompanied by problems such as floating-point artifacts, structural blurring, and texture loss. These limitations all point to the core problems of the current technology: the quality of the reconstructed model is poor, it is difficult to balance computational efficiency and hardware cost, and there is a lack of effective solutions for global consistency optimization.
[0004] The existing technical solutions are as follows:
[0005] The three-dimensional reconstruction technology has shifted from traditional methods that rely on densely sampled images to modern techniques that employ deep learning architectures, giving rise to a boom in three-dimensional reconstruction with Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) as the mainstream. NeRF is time-consuming to train and slow to render, limiting its application in interactive scenarios. On the other hand, 3DGS optimizes rendering through parallel projection, avoiding complex ray tracing operations and supporting anisotropic splatting, achieving real-time rendering while ensuring the modeling quality, making it the preferred algorithm in real-time interactive applications. However, the three-dimensional reconstruction task of indoor scenes still faces many challenges due to its complex background environment, topological structure, dense detailed textures, varying numbers of objects, and mutual occlusion. To improve the modeling quality and visual effects of three-dimensional reconstruction of indoor scenes, many studies have carried out active explorations.
[0006] For example, Ref. [1] proposed a partitioned joint bilateral filtering algorithm. By analyzing the left-right differences between the depth image and the grayscale image to determine pixel confidence, and performing interpolation on the missing pixels in the depth image, it achieved partitioned joint bilateral filtering of the image based on confidence. As a result, the noise was effectively smoothed, and the accuracy of 3D point cloud mapping and pose estimation in indoor scenes was improved. However, this method increased the number of model parameters and could not perform real-time reconstruction. Ref. [2] used an RGB-D sensor for 3D reconstruction of indoor complex scenes, proposed a depth sensor noise model estimation algorithm, a depth SLAM algorithm based on the fusion of dense points and planar features, and a sensor tracking algorithm based on the sampling integrity evaluation criterion, designed the volume data fusion of the sensor noise model and the region of interest, and filled the missing data using a model combining a 3D deep neural network and planar rules. A robust and high-precision 3D reconstruction system for indoor complex scenes was obtained, which can be directly applied to scenarios such as indoor service robot navigation, indoor positioning, and augmented reality. Ref. [3] proposed a 3D reconstruction simulation method for indoor scenes based on a high-low frequency feature enhancement algorithm. The wavelet transform algorithm was used to extract the high-low frequency features of a single indoor scene image, and the high-low frequency feature images were enhanced through a weighted guided filtering algorithm. A deep neural network was adopted to predict the segmentation map and depth map of the image, and the long short-term memory network was combined to estimate the planar parameters of a single indoor scene image. This method can effectively extract the high-low frequency features of a single indoor scene image, enhance the high-low frequency feature images, and improve the image clarity, but the modeling effect is still not ideal. Ref. [4] proposed a method for reconstructing a 3D Mesh model of an indoor scene based on the POCO deep neural network, improved training, and reconstruction strategies. First, the training strategy was improved, and the existing small amount of simulation scene and extremely small amount of real scene data were fully utilized to fine-tune the original model. Then, the reconstruction strategy was improved by introducing the farthest point sampling and the strategy of consistent bounding box scale. Finally, the scale of the reconstructed 3D Mesh model was restored. A good balance was achieved between the modeling speed and accuracy.
[0007] Anhui Jianzhu University. A method for architectural design and analysis based on 3D reconstruction of point clouds, Patent Application No.: CN202510026830.X. By obtaining the external point cloud data of a house, combining with the CAD design drawings (CAD images) of the house, and using the collected dataset to train an indoor scene inference network; using a generative adversarial network (GANs) to infer and generate multi-view images of the indoor scene, then generating indoor point clouds from these images, and finally fusing the indoor and outdoor point clouds, it is possible to infer and design the indoor layout of the house and establish a house model without knowing the indoor scene layout, which not only protects the privacy of the house owner but also improves the design efficiency and accuracy.
[0008] Communication University of China. Three-dimensional Reconstruction Method and System for Static Indoor Scenes Based on Neural Radiance Fields, Patent Application No.: CN202411270104.4. A three-dimensional reconstruction method and system for static indoor scenes based on neural radiance fields are provided. The method includes: obtaining data of the indoor scene to be reconstructed; dividing the space of the indoor scene to be reconstructed into two or more parameterized voxel grids based on the neural radiance field algorithm, and describing the scene information corresponding to the voxel grids through parameters to reconstruct an implicit model of the neural radiance field of the scene; based on the particle swarm optimization algorithm, iteratively evaluating the performance of the parameter combinations of the current voxel grids of the neural radiance field implicit model until obtaining the parameter combinations of the voxel grids corresponding to the globally optimal position; obtaining the trained neural radiance field implicit model. The present invention enables non-professional users to quickly and accurately complete the three-dimensional reconstruction of static indoor scenes; at the same time, it improves the efficiency and quality of three-dimensional reconstruction; it is suitable for application scenarios that require instant data and quick response.
[0009] Baolue Technology (Zhejiang) Co., Ltd., Baolue Digital Technology (Hangzhou) Co., Ltd. A Three-dimensional Scene Reconstruction Method and System Based on Slam and 3D Gaussian Fusion, Patent Application No.: CN202411857878.7. A three-dimensional scene reconstruction method and system based on Slam and 3D Gaussian fusion are provided. Indoor images are taken for the target indoor scene to obtain a dense point cloud A, and a dense point cloud B is scanned; then, the dense point cloud B is aligned to the dense point cloud A to obtain a fused point cloud; next, the fused point cloud is processed to generate multiple Gaussian points, and the point cloud data points corresponding to each Gaussian point are recorded; further, based on each point cloud data point, the normal vectors of each Gaussian point are aligned with the normal vector of the fused point cloud; at the same time, the tangential vector distance between the center of each Gaussian point and the center of the fused point cloud is restricted; finally, the normal vector distance between the center of each Gaussian point and the center of the fused point cloud is restricted to obtain the reconstructed three-dimensional scene for measuring the image distance between any two points in the three-dimensional scene and converting it into a real distance. The beneficial effect is that the present invention can improve the visual effect of the three-dimensional reconstruction scene.
[0010] China Mobile Communications Group Guangdong Co., Ltd., CMCC Bay Area (Guangdong) Innovation Research Institute Co., Ltd., Tsinghua University Shenzhen International Graduate School, etc. An indoor three-dimensional scene reconstruction method and system based on implicit encoding and geometric prior, Patent Application No.: CN202411332432.2, discloses an indoor three-dimensional scene reconstruction method and system based on implicit encoding and geometric prior. The method of the present invention includes performing data processing on the image data to be reconstructed to obtain corresponding preprocessed data; calculating the rays in the world coordinate system corresponding to the pixels to be trained on the image data according to the internal and external parameters of the camera; inputting the three-dimensional coordinates encoding, observation angle encoding of the sampling points on the rays, and the implicit appearance encoding of the corresponding image into the neural radiance field model to obtain the predicted density value and color value, and predicting the pixel value using volume rendering based on the signed distance function to obtain a trained neural radiance field model; obtaining the standard implicit appearance encoding based on the density value and color value, and inputting the data including at least the standard implicit appearance encoding into the trained neural radiance field model to obtain the scene geometric information. The present invention can reduce the dependence of the existing reconstruction scheme on the three-dimensional consistency of the collected data, and improve the robustness and representation ability of the algorithm in complex scenarios.
[0011] The 54th Research Institute of China Electronics Technology Group Corporation, Shandong Jianzhu University. An indoor three-dimensional reconstruction method considering indoor curved surface structures
[0012] Patent Application No.: CN202411748650.4, discloses an indoor three-dimensional reconstruction method considering indoor curved surface structures, which relates to the technical field of cartography. After inputting the indoor three-dimensional point cloud, the present invention first divides the indoor three-dimensional point cloud into curved surface points and plane supervoxel sets based on supervoxels; for the curved surface points and plane supervoxel sets, model fitting is performed respectively to extract indoor plane and curved surface models, and the indoor two-dimensional space is divided based on the 3D-2D projection model; the indoor two-dimensional floor plan is constructed based on overlay analysis and Markov field; and finally, the indoor three-dimensional model is established based on constrained Delaunay triangulation. The present invention performs indoor three-dimensional reconstruction based on three-dimensional point cloud, realizes the three-dimensional reconstruction of indoor curved surface structures, especially cylindrical curved surface structures, and can quickly, effectively and accurately realize indoor scene reconstruction to meet the requirements of complex indoor scene navigation and modeling.
[0013] Although the above methods have made certain progress in indoor scene three-dimensional reconstruction, most methods still have problems such as excessive model parameter quantity, lack of real-time performance, low modeling accuracy and poor visual effect in complex scenes.
[0014] In summary, in the field of 3D reconstruction of complex indoor scenes, due to the numerous and mutually occluding scene objects and the complex spatial topological structure, it is extremely sparse or even missing to capture texture details and scale structure features. The loss of such tiny detail features has an adverse impact on the integrity of scene modeling. Although current research enhances the model's ability to capture key information through techniques such as point cloud segmentation, voxel representation, and implicit encoding, and improves the model's sensitivity to spatial position, channels, and semantic information, the above existing methods face dual technical bottlenecks in complex indoor scenes: one is the insufficient positioning accuracy of the spatial structure of the region of interest (ROI), and the other is the insufficient representation ability of the detail feature information of the target region. As a result, the reconstruction performance of the 3D modeling system in complex indoor environments has decreased significantly, specifically manifested as secondary problems such as geometric structure distortion and texture detail loss. Limited by the insufficient computational efficiency and reconstruction accuracy, the current technology is difficult to meet the millisecond-level response and dynamic update requirements of the 3D reconstruction system in real-time interactive scenarios.
[0015] [1] Xiao Zhiyuan, Li Hongwei, Zhang Bin, et al. A Partitioned Joint Bilateral Filtering Method for 3D Reconstruction of Indoor Scenes [J]. Journal of Geomatics Science and Technology, 2024, 40(05): 505 - 510 + 540.
[0016] [2] Research on 3D Reconstruction Technology of Complex Indoor Scenes Combining Geometry and Semantics. Beijing, Institute of Automation, Chinese Academy of Sciences, 2022 - 12 - 01.
[0017] [3] Zheng Xiaoqian. Simulation of 3D Reconstruction of Indoor Scenes Based on High - Low Frequency Feature Enhancement Algorithm [J]. Journal of Ningbo University of Technology, 2024, 36(03): 117 - 123.
[0018] [4] Song Peiyan, Ye Qin, Zeng Liang, et al. A POCO Reconstruction Method for Indoor Scene 3D Mesh Models with Significantly Different Point Cloud Densities [J]. Bulletin of Surveying and Mapping, 2025, (02): 7 - 12 + 47. Summary of the Invention
[0019] To solve the above problems, the present invention proposes: an indoor real - scene 3D reconstruction method based on improved 3D Gaussian sputtering, comprising the following steps:
[0020] S1. Obtain the 3D reconstruction dataset of the indoor scene captured in the real scene and perform pre - processing;
[0021] S2. Construct a feed - forward point cloud densification network to obtain a denser point cloud with higher resolution and camera positions;
[0022] S3. Improve the 3DGS algorithm model and construct the GD - 3DGS algorithm model;
[0023] S4. Train the indoor scene three-dimensional reconstruction network based on GD-3DGS;
[0024] S5. Input the test set for testing and evaluation.
[0025] Furthermore, in the step S1, for dataset preprocessing: Obtain the indoor datasets captured from real scenes: Tanksand temples and Mip-NeRF360. These two datasets cover multiple groups of indoor scenes with different styles.
[0026] Furthermore, in the step S2, constructing the dense point cloud thickening network includes the following steps:
[0027] 1) The network takes the SfM point cloud data as input. The point cloud is represented as a set of three-dimensional points, and each point contains its coordinate values on the x, y, and z axes. Represent the number of points as N. Then the network input is a two-dimensional matrix with a dimension of N×3, and each row corresponds to the three-dimensional coordinates of a point;
[0028] 2) Then perform coordinate transformation through a 3×3 transformation matrix, and the transformation matrix is learned by the T-Net spatial transformation network;
[0029] 3) Then, through a multi-layer perceptron network MLP with shared weights, learn the features of each point, and the feature dimension is N×64;
[0030] 4) After the above operations, use a T-Net and MLP network to perform feature extraction again to obtain a feature matrix with a dimension of N×1024;
[0031] 5) Based on the per-point features obtained above, perform a max-pooling operation along the feature dimension direction on the feature matrix to aggregate the global features, and obtain the global features of the input point cloud data, with a feature dimension of 1×1024;
[0032] 6) Finally, predict the overall global geometry of the point cloud through the feature pyramid decoder FPN, complement the missing point cloud, and generate the initial distribution of the dense point cloud.
[0033] Furthermore, in the step S3, constructing the GD-3DGS network model includes the following steps:
[0034] S31. Construct a dense 3D Gaussian radiation field;
[0035] S32. 3D Gaussian ellipsoid projection;
[0036] S33. Image rendering based on rasterization;
[0037] S34. Construct the depth regularization loss function DR-Loss;
[0038] S35. Adaptive density control.
[0039] Furthermore, in step S31, the dense point cloud in step S2 is initialized and modeled as a 3D three-dimensional Gaussian, and a dense 3D Gaussian radiation field is constructed by superimposing these Gaussian distributions, which can more efficiently express the geometric structure and appearance information of the scene;
[0040] The 3D Gaussian function consists of position, size, and covariance matrix, and its probability density function is as shown in Equation (1); since the Gaussian function is centered at any point, the mean value is set to the zero vector for de-centralization processing; the scale coefficient is to ensure that the integral of G(x) is constant (∫G(x) = 1), and in order to freely control the size of the Gaussian distribution, this coefficient is set to 1 for normalization processing, and the final probability density is as shown in Equation (2);
[0041]
[0042] where x ∈ R 3 is any point cloud; μ ∈ R 3 is the spatial mean value, which controls the center position of the Gaussian function; Σ ∈ R 3 is the covariance matrix, which controls the size and shape of the Gaussian function;
[0043] By adjusting Σ, the anisotropy of the ellipsoid is controlled and the scene details are shaped while maintaining the differentiability of the Gaussian distribution. Σ is decomposed as:
[0044] Σ = RSSR (3)
[0045] where S ∈ R 3 is the scaling matrix; R ∈ R 4 is the quaternion rotation matrix.
[0046] Furthermore, in step S32, when rendering any pixel point, first eliminate the Gaussian functions outside the frustum, and only retain the Gaussian functions within the 3σ region;
[0047] When projecting, the view plane is divided into 16×16 grids, and the 3D Gaussian is projected onto the view plane to obtain a 2D Gaussian. The corresponding covariance matrix Σ2D is:
[0048] Σ 2D = J·W·Σ·W T ·J T (4)
[0049] where J is the Jacobian matrix of the affine approximation of the perspective projection transformation, representing the local approximation of a multivariate function at a certain point; W is the transformation matrix from the world coordinate system to the camera coordinate system, that is, the camera pose.
[0050] Furthermore, in step S3, the rasterization-based image rendering process is as follows:
[0051] Subsequently, the raster information and depth values of the two-dimensional ellipsoid are recorded; the ellipsoid is quickly depth-sorted, and a list is generated for each raster by recording the subscripts of the first and last Gaussian functions in the sorting; during the rasterization rendering process, a thread is started for each raster to load the Gaussian function data packet into the shared memory, and each pixel traverses the list from near to far, and combines alpha blending to fuse the colors and opacities of these N two-dimensional Gaussian functions to obtain the final pixel color C:
[0052]
[0053] Among them, c i and α i respectively represent the contributions of the i-th 3D Gaussian function to the color and opacity of the rendered pixel point; c i is calculated by the SH function; α i = α × Σ2D, where α is the opacity; when the opacity α of a certain pixel point reaches saturation (α = 1), the corresponding thread is stopped; when all pixels in the raster reach saturation, the processing of this raster block is terminated.
[0054] Furthermore, in the step S34, during the process of fitting the scene, 3DGS uses the structural similarity SSIM combined with the photometric error loss function LSSIM of the 1-norm to optimize the radiation field parameters, as shown in Equation (6); depth information is introduced, and the depth regularization loss function DR-Loss is proposed, as shown in Equation (7);
[0055] L SSIM = (1 - λ)L1 + λL D-SSIM (6)
[0056] L DR = L SSIM + γL Depth (7)
[0057] Among them, L1 is the photometric error between the rendered image and the real image; L D-SSIM is the structural similarity error; λ and γ are hyperparameters that control the loss weights; L Depth is the depth loss function;
[0058] L Depth optimizes the model performance by maximizing the linear correlation between the rendered value and the real value; to reduce the error accumulated by repeated alignment during the operation, the Pearson correlation coefficient PCC is selected to construct the depth loss function L Depth as shown in Formula (8);
[0059] L Depth = 1 - r pcc (8)
[0060]
[0061] Among them, D p is the depth of the rendered image; D r is the depth of the original image; σ p is the standard deviation of the rendered image; σ r is the standard deviation of the original image; The depth information of the original image is obtained using the Marigold model; and for each pixel point (x, y) in the rendered image, the gradients in the horizontal and vertical directions are calculated:
[0062] grad x = d(x + 1, y) - d(x, y) (10)
[0063] grad y = d(x, y + 1) - d(x, y) (11)
[0064] Next, the Euclidean norm is used to calculate the magnitude of the gradient vector D p in formula (12), the average value of the image gradient in formula (13) and the standard deviation σ p in formula (14)
[0065]
[0066] Furthermore, in step S35, two types of scene representation problems are solved through adaptive density control:
[0067] For the under-reconstruction area, the number and distribution density of Gaussian functions are increased by cloning, so as to more accurately capture local detail features;
[0068] For the over-reconstruction area, the Gaussian functions are redistributed through splitting operations to reduce the local density of the radiation field, achieving a balance between computational efficiency and representation accuracy, and improving the adaptability of scene representation.
[0069] The beneficial effects of the present invention are as follows: Aiming at complex indoor scenes, the classical 3D reconstruction methods have problems such as insufficient detail representation, blurred edges, and aliasing artifacts. Based on 3D Gaussian Splatting (3DGS), an indoor complex scene reconstruction algorithm is proposed. The algorithm first proposes a pre - point cloud densification network to densify and complete the sparse point cloud recovered by Structure from Motion, constructs a high - resolution enhanced 3D Gaussian radiation field, and improves the representation accuracy of the detailed parts of complex scenes; subsequently, the depth information of the scene is introduced, a depth regularization loss function is designed, and the depth cost information is used to associate objects at different depths in the scene to solve the problem of structural scale ambiguity; combined with an adaptive density control strategy, the radiation field parameters are dynamically optimized to improve the accuracy and visual effect of the 3D reconstruction model. Experiments on the Tanks and temples and Mip - NeRF360 real - world datasets show that compared with traditional 3DGS: the peak signal - to - noise ratio of the reconstructed views is increased by 6.77% and 4.79% respectively, and both the structural similarity and the learned perceptual image patch similarity are improved. It effectively solves the problems of insufficient accuracy and blurred aliasing in 3D digital image reconstruction, and provides technical support for indoor design, navigation, and digital urban construction.
[0070] An efficient 3D reconstruction method for indoor scenes, GD - 3DGS, is proposed. The method designs a feed - forward point cloud densification network based on PointNet, and the high - resolution point cloud improves the expression of the fine - grained features of the scene; with the help of the dense point cloud, the 3D Gaussian radiation field is enhanced to fit the real - world scene. In the process of cost - volume optimization, a depth regularization loss function for scale structure is proposed to constrain the relative depth consistency between adjacent pixels and global objects, eliminating the blurred artifacts caused by the depth difference of objects; combined with the adaptive density control and Gaussian pruning strategies, the radiation field parameters and density distribution are refined to achieve the indoor scene reconstruction with double improvement of visual fidelity and structural integrity.
[0071] Through a series of comparative experiments and ablation experiments, the reconstruction model is evaluated. The experimental results prove that the GD - 3DGS algorithm of the present invention shows excellent performance advantages in the 3D reconstruction of complex indoor scenes, and can effectively cope with challenges such as resistance to light changes, capturing complex topological structures and detailed texture features, and dealing with real - time interactive applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is the structural diagram of the GD - 3DGS model of the present invention;
[0073] Figure 2 It is the structural diagram of the point cloud densification network of the present invention;
[0074] Figure 3 It is the flow chart of the GD - 3DGS of the present invention;
[0075] Figure 4 Schematic diagram of the Gauss projection of the present invention;
[0076] Figure 5 Flowchart of rasterized image rendering of the present invention;
[0077] Figure 6 Flowchart of adaptive density control of the present invention;
[0078] Figure 7 Comparison result diagram of the experimental results of the "playroom" scene in Tanks and temples of the present invention;
[0079] Figure 8 Comparison result diagram of the experimental results of the "room" scene in Mip-NeRF360 of the present invention;
[0080] Figure 9 Comparison result diagram of point cloud visualization of the present invention;
[0081] Figure 10 Comparison result diagram of captured light visualization of the present invention. Detailed implementation manners
[0082] To make the technical means and achieved purposes adopted by the present invention easy to understand, the present invention will be further described below in conjunction with the detailed implementation manners. Aiming at the poor quality of the three-dimensional reconstruction model of indoor complex scenes, there are phenomena such as structural distortion and blurred artifacts. A method for three-dimensional reconstruction of indoor real scenes based on improved 3D Gaussian sputtering is as Figure 1 shown. This method includes the following steps:
[0083] S1. Obtain the three-dimensional reconstruction dataset of the indoor scene captured from the real scene and perform preprocessing;
[0084] S2. Construct a feed-forward point cloud densification network to obtain a denser point cloud with higher resolution and the camera positions;
[0085] S3. Improve the 3DGS algorithm model and construct the GD-3DGS algorithm model;
[0086] S4. Train the three-dimensional reconstruction network of the indoor scene based on GD-3DGS;
[0087] S5. Input the test set for testing and evaluation.
[0088] The present invention effectively improves the ability of the model to extract detailed features of complex scenes, reduces problems such as structural blurring, texture distortion, and floating-point black shadows, and thus significantly improves the modeling accuracy and visual effect of the three-dimensional scene model of the indoor scene.
[0089] S1. Dataset preprocessing
[0090] Obtain indoor datasets captured in real scenes: Tanks and temples and Mip-NeRF360. These two datasets cover multiple groups of indoor scenes with different styles, including bedroom rooms, kitchens, counter studies, etc., such as "playroom", "drjohnson", "counter", "kitchen", "bonsai" scenes, etc. These scenes have strong diversity and representativeness.
[0091] S2. Construct a dense point cloud densification network
[0092] 3DGS uses Structure from Motion (SfM) technology to generate an initial point cloud. The SfM point cloud is unevenly distributed and very sparse. For indoor scenes with a large number of mutual occlusions, the coverage rate of the initial point cloud is very small. If a sparse initial point cloud is used to construct a 3D Gaussian radiation field to synthesize new views, problems such as missing structures and blurred black shadows are likely to occur. To make up for the limitations of the sparse point cloud, a point cloud densification network based on PointNet is designed, as Figure 2 shown.
[0093] 1) The network takes the SfM point cloud data as input. The point cloud can be represented as a set of three-dimensional points, and each point contains its coordinate values on the x, y, and z axes. If N represents the number of points, then the network input is a two-dimensional matrix of dimension N×3, and each row corresponds to the three-dimensional coordinates of a point.
[0094] 2) Then, coordinate transformation is performed through a 3×3 transformation matrix, which is learned by the T-Net spatial transformation network, ensuring the result stability of the algorithm in the case of spatial transformation of the three-dimensional model.
[0095] 3) Then, through a multi-layer perceptron (MLP) with shared weights, the features of each point are learned, and the feature dimension is N×64;
[0096] 4) After the above operations, a T-Net and an MLP network are used again for feature extraction to obtain a feature matrix of dimension N×1024;
[0097] 5) Based on the per-point features obtained above, a max pooling operation is performed on the feature matrix along the feature dimension direction to aggregate global features, obtaining the global features of the input point cloud data, with a feature dimension of 1×1024;
[0098] 6) Finally, through a Feature Pyramid Network (FPN), the overall global geometry of the point cloud is predicted, the missing point cloud is complemented, and the initial distribution of the dense point cloud is generated.
[0099] Max-Pooling ensures that the output remains unchanged for any order of input point sets, thus ensuring that the position of the point cloud does not change during the feature extraction process. The benchmark of FPN is the fully-connected decoder, which can decode on multi-scale feature maps, effectively utilize feature information at different levels, and make the decoding calculation efficiency relatively higher; at the same time, multi-scale features can reduce errors caused by insufficient resolution. The densification network complements the sparse SfM point cloud into a denser point cloud with a higher resolution. The dense point cloud is used as the initial point cloud of GD-3DGS to guide the algorithm to accurately reconstruct the geometric scene.
[0100] S3. Construct the GD-3DGS network model
[0101] 3DGS optimizes rendering through parallel projection, avoiding complex ray tracing operations and supporting anisotropic sputtering. While ensuring the modeling quality, it realizes real-time rendering, making interactive applications possible, providing a real-time solution for indoor scene reconstruction, and broadening the application of 3D reconstruction technology. Therefore, the present invention uses 3DGS as the basic model for improvement to construct the GD-3DGS algorithm model. The algorithm process is as Figure 3 shown.
[0102] S31. Construct a dense 3D Gaussian radiation field
[0103] Initialize the dense point cloud in step S2 as 3D three-dimensional Gaussians, and construct a dense 3D Gaussian radiation field by superimposing these Gaussian distributions, which can more efficiently express the geometric structure and appearance information of the scene. The 3D Gaussian function consists of position, size, covariance matrix, etc., and its probability density function is as shown in Equation (1). Since the Gaussian function is centered at any point, the mean value is set to the zero vector for decentralization processing; the scale coefficient ensures that the integral of G(x) is constant (∫G(x) = 1). In order to freely control the size of the Gaussian distribution, this coefficient is set to 1 for normalization processing. The final probability density is as shown in Equation (2).
[0104]
[0105] where x ∈ R 3 is any point cloud; μ ∈ R 3 is the spatial mean value, controlling the center position of the Gaussian function; Σ ∈ R 3 is the covariance matrix, controlling the size and shape of the Gaussian function.
[0106] The probability density of the 3D Gaussian gradually decays from the center to the surroundings and reaches the highest value at the center position. This progressive distribution is more in line with the real scene. Geometrically, it is a three-dimensional ellipsoid centered at (x, y, z). Tens of thousands of Gaussian ellipsoids can represent any three-dimensional scene. By adjusting Σ to control the anisotropy of the ellipsoid and shape the scene details, in order to maintain the differentiability of the Gaussian distribution, Σ is decomposed:
[0107] Σ = RSSR (3)
[0108] where S ∈ R 3 is a scaling matrix; R ∈ R 4 is a quaternion rotation matrix.
[0109] In addition, the radiation field also stores the color coefficient c and the opacity α. c is represented by spherical harmonics (SH). The SH functions are a set of orthogonal functions on the unit sphere. When performing lighting transformation or reflection, SH changes with the projection direction, simulating the transformation of the lighting intensity and distribution in each direction. Supplemented with the opacity α in rendering, it can approximate the human eye vision.
[0110] S32, 3D Gaussian ellipsoid projection
[0111] When rendering any pixel point, due to the perspective difference, Gaussian functions outside the frustum can be removed first, and only the Gaussian functions within the 3σ (about 99.7% confidence interval) region are retained. This will neither affect the subsequent rendering effect nor increase the computational complexity. As Figure 4 shown, it is the projection process.
[0112] During projection, the view plane is divided into 16×16 grids. This is a non-linear approximation process. Projecting the 3D Gaussian onto the view plane to obtain a 2D Gaussian, and the corresponding covariance matrix Σ2D is:[[]]
[0113] Σ 2D = J·W·Σ·W T ·J T (4)
[0114] where J is the Jacobian matrix of the affine approximation of the perspective projection transformation, representing the local approximation of a multivariate function at a certain point; W is the transformation matrix from the world coordinate system to the camera coordinate system, that is, the camera pose.
[0115] S33, Image rendering based on rasterization
[0116] Subsequently, record the grid information and depth values of the two-dimensional ellipsoid, as Figure 5 (a). Then perform a fast depth sorting on the ellipsoid as Figure 5 (b). By recording the subscripts of the first and last Gaussian functions in the sorting, generate a list for each grid. During the rasterization rendering process Figure 5In (c), for each grid, a thread is launched to load the Gaussian function data packet into the shared memory. For each pixel, the list is traversed from near to far, and the colors and opacities of these N two-dimensional Gaussian functions are fused in combination with alpha blending, so as to achieve data sharing and processing in parallel, maximize the gain, and obtain the final pixel color C:
[0117]
[0118] Among them, c i and α i respectively represent the contributions of the i-th 3D Gaussian function to the color and opacity of the rendered pixel point; c i is calculated by the SH function; α i = α × Σ2D, where α is the opacity. When the opacity α of a certain pixel point reaches saturation (α = 1), the corresponding thread is stopped; when all pixels in the grid reach saturation, the processing of this grid block is terminated.
[0119] S34. Construct the depth regularization loss function DR-Loss
[0120] During the process of fitting the scene, 3DGS uses the Structural Similarity (SSIM) combined with the L1 norm photometric error loss function LSSIM to optimize the radiance field parameters, as shown in Equation (6). SSIM focuses on the similarity between local regions of the view and can only capture the structural information carried by neighboring pixels. In order to better maintain the global structural relationship of the scene, improve the problems of scale ambiguity and lack of geometric texture, depth information is introduced, and the depth regularization loss function (DR-Loss) is proposed, as shown in Equation (7). By strengthening the constraint on the radiance field through depth information, it guides the 3D Gaussian to learn global structural information and enhances the algorithm's ability to understand the spatial structure and geometric expression ability.
[0121] L SSIM =(1 - λ)L1 + λL D-SSIM (6)
[0122] L DR = L SSIM + γL Depth (7)
[0123] Among them, 1 is the photometric error between the rendered image and the real image; L D-SSIM is the structural similarity error; λ, γ are hyperparameters that control the loss weights; L Depth is the depth loss function.
[0124] L DepthOptimize the model performance by maximizing the linear correlation between the rendered value and the true value. However, the loss calculation requires repeated scale alignment work. These tiny physical scale position changes will cause significant parallax. To reduce the error accumulated by repeated alignment during the operation, the Pearson Correlation Coefficient (PCC) is selected to construct the depth loss function L Depth As shown in formula (8). PCC measures the linear relationship between two variables, rather than the specific numerical difference, and will not change due to the scale change of the variables. Even if there is a certain proportional difference between the rendered depth and the true depth in terms of numerical values, as long as their change trends are the same, it will not affect the degree of PCC correlation.
[0125] L Depth = 1 - r pcc (8)
[0126]
[0127] where D p is the depth of the rendered image; D r is the depth of the original image; σ p is the standard deviation of the rendered image; σ r is the standard deviation of the original image. Since the input RGB image dataset does not contain depth information, the Marigold model is used to obtain the depth information of the original image; and for each pixel point (x, y) in the rendered image, calculate the gradients in the horizontal and vertical directions:
[0128] grad x = d(x + 1, y) - d(x, y) (10)
[0129] grad y = d(x, y + 1) - d(x, y) (11)
[0130] Then, use the Euclidean norm to calculate the magnitude of the gradient vector D p as formula (12), the mean value of the image gradient as formula (13) and the standard deviation σ p as formula (14).
[0131]
[0132] DR-Loss randomly samples a subset of training samples at each optimization iteration, and synchronously adjusts parameters such as the geometric properties and radiation characteristics of the Gaussian function along the gradient descent direction to minimize the difference between the rendered radiation field distribution and the real data. It takes into account both photometric information and structural scale, accurately reflects the spatial structure while optimizing details such as texture and color, avoids the problems of blurred scene scale and geometric distortion in reconstruction, and provides a more comprehensive view quality optimization metric for 3D reconstruction, especially in indoor scene reconstruction tasks that require attention to depth consistency and structural scale.
[0133] S35, Adaptive Density Control
[0134] The core of adaptive density control is to solve two types of scene representation problems: as Figure 6 shown, for under-reconstructed regions, the number and distribution density of Gaussian functions are increased by cloning to more precisely capture local detail features; for over-reconstructed regions, the Gaussian functions are redistributed through splitting operations to reduce the local density of the radiation field, achieving a balance between computational efficiency and representation accuracy, and enhancing the adaptability of scene representation.
[0135] In addition, there is instability in adaptive density control, which will generate abnormal 3D Gaussians that deviate from the scene surface in position and do not conform to geometric constraints, etc., and Gaussian pruning is required; in particular, Gaussians with low transparency (α≥0.99) are directly deleted to save storage and computational overhead.
[0136] The adaptive density control strategy monitors the Gaussian distribution in the local area in real time, dynamically adjusts the radiation field parameters and updates step size coefficients, achieving a dynamic balance of accuracy, efficiency, and robustness. This strategy enables 3D Gaussians to surpass traditional point cloud or voxel methods and become one of the most competitive technologies in the current real-time 3D reconstruction field.
[0137] S4, Training the DG-3DGS Indoor Scene 3D Reconstruction Network
[0138] The experimental configuration is as follows: The experimental environment is configured in the Ubuntu20.04 operating system and trained by an NVIDIA GeForce RTX3080 graphics card. Among them are Python3.8, Pytorch2.0.1, CUDA11.8, and Tensorflow2.13.0.
[0139] The training-related parameter settings are as follows: Each scene is trained 30k times, with 7k iterations in the coarse-grained stage and 23k iterations in the fine-grained stage. For each frame of image, the current frame is first used for iterative calculation to fully optimize the 3D Gaussian of the current perspective. Subsequently, the first n frames overlapping with the current frame are selected from the multi-list, and one of them is randomly selected for iteration each time.
[0140] S5. Testing and Evaluation
[0141] The Peak Signal-to-Noise Ratio (PSNR), SSIM, Learned Perceptual Image Patch Similarity (LPIPS), and Frames Per Second (FPS) are used as quantitative evaluation metrics to compare the performance differences of different methods. The higher the PSNR, SSIM, and FPS, and the lower the LPIPS, the better the image quality.
[0142] II. Experimental Results of the Method
[0143] (1) Comparative Experiment of Detection Methods
[0144] To better demonstrate the superiority and effectiveness of the method proposed in the present invention, the DG-3DGS method was compared with other typical 3D reconstruction methods, including NGP, Plenoxels, NeRF, Mip-NeRF36, 2DGS, SuGaR, and 3DGS. The experiments were verified on the Tanks and temples and Mip-NeRF360 datasets. The comparative experimental results are shown in Tables 1 and 2. As can be seen from Tables 1-2, compared with the baseline algorithm 3DGS: on the Tanks and Temples dataset, the PSNR and SSIM of the reconstructed views increased by 6.77% and 3.89% respectively, and the LPIPS index decreased by 2.19%. On the Mip-NeRF360 dataset, the PSNR and SSIM increased by 4.79% and 2.34% respectively, and the LPIPS index decreased by 2.65%; achieving image quality comparable to that of the Mip-NeRF360 method specifically designed for this dataset and a faster modeling speed than it. This is closely related to the feed-forward point cloud densification network in this paper. The high-resolution point cloud can effectively express the detailed features of the scene and predict occlusion information, making the reconstructed model of higher quality.
[0145] Table 1 Experimental Results of Different Methods on Tanks and temples Table 1 Experimental results of different methods on Tanks and temples
[0146]
[0147] Note: Bold data is the best quality; italic data is the second best quality.
[0148] "↑" indicates that the higher the index, the better; "↓" indicates that the lower the index, the better.
[0149] Table 2 Experimental results of different methods on Mip-NeRF360
[0150]
[0151]
[0152] (2) Visualization analysis
[0153] To verify the 3D reconstruction effect of the GD-3DGS network model, the present invention performs inference on the test sets of the Tanks and temples and Mip-NeRF360 datasets using the INGP, Mip-NeRF360, 3DGS algorithms and the GD-3DGS algorithm. Figure 7 、 Figure 8 Shows the novel view reprojection results of the algorithm of this paper, INGP, Mip-NeRF360, and 3DGS on the "playroom" scene in Tanks and temples and the "room" scene in Mip-NeRF360. Among them, it can be clearly observed from the two figures (b) that INPG performs poorly on both datasets, with a large number of blurred distortions and floating-point black shadows; in the two figures (c) of the Mip-NeRF360 algorithm, due to spending a large amount of time on rendering, the effect is relatively good, but upon closer inspection Figure 7 (c) upper right and Figure 8 (c) lower left, upper right bounding boxes, where the light changes strongly (weak light or strong light). In contrast, the algorithm of this paper has a better optimization effect on the light and shadow colors and texture details of these regions because 3D Gaussian sputtering rendering has differentiable continuity and can naturally simulate the gradual change effect of light; comparing the 3DGS figure (d) and the algorithm figure (e) of this paper, it can be clearly seen the improvement effect of the algorithm of this paper on the missing fine-grained features such as the micro-structure of the wall surface and texture transition and artifact distortion. In addition, reconstructing high-reflective materials is one of the technical problems faced by current 3D reconstruction technologies. This special material with low texture and high reflectivity will cause ground reflection and specular reflection ( Figure 8 the TV and the floor in), in contrast, the algorithm of this paper can reconstruct these reflection details more completely. The present invention has good view reconstruction ability on the basis of resisting light transformation, with richer modeling details and better visual effects. On the one hand, it benefits from the high-resolution point cloud of the feed-forward densification network, which improves the expression of the fine-grained features of the scene. On the other hand, it is the use of a depth regularization loss function for scale structure to constrain the relative depth consistency between adjacent pixels and global objects, eliminating the blurred artifacts caused by the depth difference of objects.
[0154] To verify the improvement effect of the feed-forward point cloud densification network based on PointNet and the depth regularization loss module on the modeling accuracy and artifact blurring after adding them to the original 3DGS model, the following verification experiments were carried out with 3DGS as the baseline on Tanks and temples, gradually adding each module. In the experiment of fusing the densification network, the number of point clouds increased by about 0.9 million compared with the baseline 3DGS. Figure 9 For the visual comparison between the dense point cloud and the original sparse point cloud, it further verified the effectiveness of the point cloud densification operation visually. The densified point cloud significantly improved the expression of fine-grained features in complex indoor scenes. In the experiment of adding the depth regularization loss function, the algorithm convergence speed was improved. DR-Loss strengthened the constraint on the global object depth information and could better combat the lighting consistency, especially in the indoor environment facing light and shadow changes. Figure 10 It demonstrated the light capture effect of this module.
[0155] Key points of the present invention:
[0156] Aiming at the problems of sharp degradation of the reconstruction model accuracy, poor visual effects, accompanied by floating-point artifacts, structural blurring, etc. in the three-dimensional reconstruction of complex indoor scenes, a three-dimensional reconstruction method for indoor scenes based on 3DGS was proposed.
[0157] 1. To avoid the problem of missing global scene information due to the lack of sufficient point cloud constraints, resulting in the attenuation of modeling accuracy, a feed-forward point cloud densification network based on convolutional neural network was designed to capture detailed features and predict occlusion information, and adaptively complement the initial point cloud.
[0158] 2. A high-resolution dense 3D Gaussian radiation field was constructed, leveraging the progressive probability density distribution of the 3D Gaussian radiation field: gradually decaying from the center to the surrounding, reaching the highest value at the center position, fitting the real scene and filling in the missing details of the scene, and more efficiently expressing the geometric structure and appearance information of the scene.
[0159] 3. To better maintain the global structural relationship of the scene and improve the problems of scale blurring and geometric texture missing, depth information was introduced, and the depth regularization loss function DR-Loss was proposed. By strengthening the constraint on the radiation field through depth information, guiding the 3D Gaussian to learn global structural information, and enhancing the algorithm's understanding ability and geometric expression ability of the spatial structure.
[0160] 4. Through the adaptive density control strategy, the Gaussian distribution in the local area was monitored in real time, the radiation field parameters and the update step coefficient were dynamically adjusted, and a dynamic balance of accuracy, efficiency, and robustness was achieved.
[0161] As mentioned above, it is only the preferred specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and its concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. An indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering, characterized in that, The following steps are involved: S1, obtaining a 3D reconstruction dataset of indoor scenes shot from real scenes and performing preprocessing; S2, build a feedforward point cloud densification network to obtain higher resolution dense point clouds and camera positions; S3, based on the 3DGS algorithm model, improve and build the GD-3DGS algorithm model; S4, training the indoor scene 3D reconstruction network based on GD-3DGS; S5. Input the test set for testing and evaluation.
2. The indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering according to claim 1, wherein, In the step S1, data set preprocessing: obtaining indoor data sets shot from real scenes: Tanks and temples and Mip-NeRF360, these two data sets cover multiple groups of indoor scenes of different styles.
3. The indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering according to claim 1, wherein In step S2, constructing a dense point cloud densification network includes the following steps: 1) The network takes SfM point cloud data as input. The point cloud is represented as a set of three-dimensional points. Each point contains its coordinate values on the x, y, and z axes. N represents the number of points. Then the network input is a two-dimensional matrix of dimension N×3, where each row corresponds to the three-dimensional coordinates of a point. 2) Then the coordinates are transformed through a 3×3 transformation matrix, which is learned by the T-Net spatial transformation network; 3) Then, a multi-layer perceptron network (MLP) with shared weights is used to learn the features of each point, with a feature dimension of N×64. 4) After the above operations, a T-Net and MLP network are used to extract features again to obtain a feature matrix of N×1024 dimensions; 5) Based on the point-by-point features obtained above, perform the maximum pooling operation on the feature matrix along the feature dimension direction to aggregate the global features to obtain the global features of the input point cloud data, with a feature dimension of 1×1024; 6) Finally, the feature pyramid decoder FPN predicts the global geometric shape of the point cloud as a whole, completes the missing point cloud, and generates the initial distribution of the dense point cloud.
4. The indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering according to claim 3, wherein, In step S3, constructing the GD-3DGS network model includes the following steps: S31, construct dense 3D Gaussian radiation field; S32, 3D Gaussian ellipsoid projection; S33, image rendering based on rasterization; S34, construct deep regularization loss function DR-Loss; S35, adaptive density control.
5. The indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering according to claim 4, characterized in that, In the step S31, the dense point cloud in step S2 is initialized and modeled as a 3D Gaussian, and a dense 3D Gaussian radiation field is constructed by superimposing these Gaussian distributions to more efficiently express the geometric structure and appearance information of the scene; The 3D Gaussian function consists of position, size, and covariance matrix, and its probability density function is as shown in formula (1). Since the Gaussian function is centered at any point, the mean is set to zero vector for decentralized processing. The scale coefficient is to ensure that the integral of G(x) is constant (∫G(x) = 1). In order to freely control the size of the Gaussian distribution, the coefficient is set to 1 for normalization processing. The final probability density is as shown in formula (2). where \(x\in R\) 3 is any point cloud; \(\mu\in R\) 3 is the spatial mean, controlling the center position of the Gaussian function; \(\Sigma\in R\) 3 is the covariance matrix, controlling the size and shape of the Gaussian function; By adjusting Σ to control the anisotropy of the ellipsoid and shape the scene details, Σ is decomposed to maintain the differentiability of the Gaussian distribution: Σ=RSSR(3) where \(S\in\mathbb{R}\) 3 is a scaling matrix; \(R\in\mathbb{R}\) 4 is a quaternion rotation matrix.
6. The indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering according to claim 5, characterized in that, In the step S32, when rendering any pixel point, the Gaussian function outside the viewing cone is first eliminated, and only the Gaussian function within the 3σ region is retained; When projecting, the view plane is divided into a 16×16 grid, and the 3D Gaussian is projected onto the view plane to obtain a 2D Gaussian. The corresponding covariance matrix Σ2D is as follows: Σ 2D = J·W·Σ·W T ·J T (4) Among them, J is the Jacobian matrix of the affine approximation of the perspective projection transformation, representing the local approximation of a multivariate function at a certain point; W is the transformation matrix from the world coordinate system to the camera coordinate system, that is, the camera pose.
7. The indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering according to claim 6, wherein In the step S3, the rasterization-based image rendering process is as follows: Subsequently, the grid information and depth values of the two-dimensional ellipsoid are recorded; the ellipsoid is quickly depth-sorted, and a list is generated for each grid by recording the subscripts of the first and last Gaussian functions in the sorting; during the rasterization rendering process, a thread is started for each grid to load the Gaussian function data packet into the shared memory, and the list is traversed from near to far for each pixel, and the colors and opacities of these N two-dimensional Gaussian functions are fused by combining alpha blending to obtain the final pixel color C: where c i and α i represent the contributions of the i-th 3D Gaussian function to the color and opacity of the rendered pixel, respectively; c i is calculated by the SH function; α i = α × Σ2D, where α is the opacity; when the opacity α of a certain pixel reaches saturation (α = 1), the corresponding thread is stopped; when the opacities of all pixels in the raster reach saturation, the processing of this raster block is terminated.
8. The indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering according to claim 7, wherein, In the step S34, during the process of fitting the scene, 3DGS uses the structural similarity SSIM combined with the photometric error loss function LSSIM of the 1-norm to optimize the radiation field parameters, as shown in Equation (6); depth information is introduced, and the depth regularization loss function DR-Loss is proposed, as shown in Equation (7); L SSIM =(1 - λ)L1 + λL D-SSIM (6) L DR = L SSIM + γL Depth (7) where 1 is the photometric error between the rendered image and the ground truth image; L D-SSIM is the structural similarity error; λ, γ are hyperparameters that control the loss weights; L Depth is the depth loss function; L Depth Optimize the model performance by maximizing the linear correlation between the rendering value and the true value; reduce the error accumulated by repeated alignment during the operation, and select the Pearson correlation coefficient PCC to construct the depth loss function L Depth As shown in Equation (8); L Depth =1-r pcc (8) Among them, D p is the depth of the rendered image; D r is the depth of the original image; σ p is the standard deviation of the rendered image; σ r is the standard deviation of the original image; the depth information of the original image is obtained using the Marigold model; and for each pixel point (x, y) in the rendered image, the gradients in the horizontal and vertical directions are calculated: grad x = d(x + 1, y) - d(x, y) (10) grad y = d(x, y + 1) - d(x, y) (11) Next, the Euclidean norm is used to calculate the magnitude of the gradient vector D p in Equation (12), the mean of the image gradient D_ p in Equation (13), and the standard deviation σ p in Equation (14) 9. The indoor real-scene three-dimensional reconstruction method based on improved 3D Gaussian sputtering according to claim 4, wherein In the step S35, two types of scene representation problems are solved through adaptive density control: For under-reconstructed regions, the number and distribution density of Gaussian functions are increased by cloning, so as to more accurately capture local detail features; For over-reconstructed regions, the Gaussian functions are redistributed through splitting operations to reduce the local density of the radiation field, achieve a balance between computational efficiency and representation accuracy, and improve the adaptability of scene representation.
Citation Information
Patent Citations
A method for indoor 3D reconstruction taking into account indoor curved surface structures
CN119229059B
Indoor three-dimensional scene reconstruction method and system based on implicit coding and geometric prior
CN119251402A
A three-dimensional scene reconstruction method and system based on SLAM and 3D Gaussian fusion
CN119313843B
Static indoor scene three-dimensional reconstruction method and system based on neural radiation field
CN119399399A
Architectural design and analysis method based on point cloud 3D reconstruction
CN119416333B
Cited By
Overall dual-constraint dynamic registration navigation system for intraoperative scene
CN120451463A
Point cloud reconstruction method and system based on perception and dynamic initialization
CN120495542A
Three-dimensional scene reconstruction method based on Gaussian splashing
CN120495543A
Three-dimensional forest reconstruction method based on NeRF and GS joint optimization
CN120510307A
A 3D forest reconstruction method based on joint optimization of NeRF and GS
CN120510307B