Structured Gaussian splashing method based on image and radar data
By introducing depth prior information, virtual background wall and adaptive voxelization strategy into the three-dimensional Gaussian splashing method, the problem of insufficient reconstruction effect of three-dimensional scenes under single-view data is solved, and higher robustness, accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510051383.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
The existing three-dimensional Gaussian splattering method has insufficient reconstruction effect under single-view data, lacks robustness, is too dependent on fixed voxel structure, and does not fully utilize the depth prior information, resulting in limited reconstruction accuracy and efficiency in LiDAR point cloud sparse, background missing, and large-scale scenarios.
A structured Gaussian splashing method based on image and radar data is proposed. By introducing depth prior information, designing virtual background walls, adaptive voxelization and optimization strategies, it improves the robustness, accuracy and efficiency of 3D Gaussian scene reconstruction under single-view RGB-LiDAR data.
It effectively solves the problem of sparseness and absence of LiDAR point clouds in long-distance areas, improves the integrity of scene initialization and background rendering quality, and enhances the robustness and detail fidelity of scene reconstruction.
Smart Images

Figure CN119991902A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of three-dimensional reconstruction and three-dimensional Gaussian splashing, and relates to a structured Gaussian splashing method based on image and radar data. Background Art
[0002] Three-dimensional scene reconstruction is an important research task in computer vision and computer graphics. In recent years, 3D Gaussian Splatting (3DGS) technology has become a new generation of 3D scene representation methods due to its efficient rendering capabilities and high-quality new perspective synthesis effects. However, most existing 3DGS methods rely on multi-view data for reconstruction, and ensure accurate reconstruction of scene structure and new perspective synthesis effects through the geometric consistency provided by multi-view images. However, in practical applications, the cost of obtaining multi-view data is high and difficult to achieve, which limits the application scope of existing methods. In "Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering" published by Lu et al. at CVPR 2024, a voxel grid-based 3D Gaussian Splatting method is used to guide the growth and optimization of 3D Gaussians through voxelized structures to achieve efficient and structurally consistent 3D scene reconstruction and new perspective synthesis. However, this method lacks robustness when dealing with reconstruction tasks of single-viewpoint data, relies too much on fixed voxel structure, does not fully utilize depth prior information, and has limited reconstruction accuracy and efficiency in sparse LiDAR point clouds, missing backgrounds, and large-scale scenes. Summary of the invention
[0003] The present invention aims to improve the existing 3D Gaussian splashing method in terms of insufficient reconstruction effect when using single-viewpoint data such as images and radars, and proposes a structured Gaussian splashing method based on image and radar data. This method constructs a structured 3D Gaussian scene, introduces depth prior information, designs a virtual background wall, adaptive voxelization and optimization strategies, and improves the robustness, accuracy and efficiency of 3D Gaussian scene reconstruction under single-viewpoint RGB-LiDAR data.
[0004] The following three modules are designed in this method:
[0005] 1. Virtual background wall construction. The module constructs a virtual background wall behind the camera field of view to fill the sparsity of point clouds in distant areas. The specific method is to use camera parameters and depth information to calculate the position and size of the background wall and generate a set of virtual point clouds in this area.
[0006] 2. Adaptive 3DGS voxelization. This module analyzes the distribution density of point clouds in the scene and adaptively adjusts the voxel size. The K-nearest neighbor algorithm is used to calculate the average neighbor distance of the point cloud, which more reasonably represents point clouds with different density distributions, thereby solving the problem of reduced training speed when the scene is too dense.
[0007] 3. Voxel growth and deletion strategy guided by depth prior. This module guides the growth and optimization of voxels based on the prior information generated by the depth map. By detecting the RGB gradient changes and depth consistency of voxels in the projected image, candidate voxels are screened, and semantic and depth constraints are imposed on their neighborhoods to decide whether to expand or delete them.
[0008] Beneficial Effects
[0009] 1) By introducing a virtual background wall construction module, the problem of sparse and missing LiDAR point clouds in distant areas is effectively solved, and the integrity of scene initialization and the quality of background rendering are improved; 2) An adaptive 3DGS voxelization method is used to more reasonably represent point clouds with different density distributions, solving the problem of decreased training speed when the scene is too dense; 3) A depth prior-guided voxel growth and deletion strategy is designed, and RGB gradient, depth consistency and neighborhood information are used to guide voxel optimization, which supplements the missing geometric information under single-viewpoint data and enhances the robustness and detail fidelity of scene reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 Schematic diagram of the model framework of the method of the present invention;
[0011] Figure 2 This is the core structure used in the present invention, and is a schematic diagram of the projection calculation of the voxel growth and deletion strategy guided by the depth prior
[0012] Figure 3 The experimental results of the present invention are as follows: (a) is a real image, (b) is the new perspective synthesis result in “Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering” published by Lu et al. at CVPR in 2024, and (c) is the new perspective synthesis result of the method of the present invention; DETAILED DESCRIPTION
[0013] The present invention is based on the open source tools Torch and CUDA for deep learning, and uses the GPU processor NVIDIA GTX4090 for 3D reconstruction.
[0014] The following is a further explanation of the composition of each module in the method of the present invention, as well as the training and use process of the method model in conjunction with the accompanying drawings and specific implementation methods. It should be understood that the specific examples in the text are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0015] The module composition and process of the present invention are as follows Figure 1 As shown, it specifically includes the following modules:
[0016] 1. Virtual background wall construction module.
[0017] The proposed virtual background wall construction method aims to solve the problem that the point cloud data of the distant background area is missing due to the limited detection distance and field of view of the LiDAR sensor, making it impossible to effectively initialize the scene. The core goal of this method is to construct a virtual point cloud background wall located behind the camera imaging plane and parallel to the plane to fill the gap in the distant background point cloud. To ensure that the virtual background wall is parallel to the camera imaging plane, the scene center x is first determined by calculating the mean of all points in the scene on the three axes. scene , and then use the camera pose x provided by the dataset cam The distance d between the scene center and the camera is calculated by calculating x scene and x cam The Euclidean space distance between them. Then the normalized vector x cam -x scene As the normal vector of the virtual background wall. Next, determine the range of the background wall: use the focal length information f of the camera in the x-axis and y-axis directions x With f y , and the length and width pixel values w of the image img and width img Calculate the actual physical distance corresponding to each pixel on the camera plane, and then infer the length w of the virtual background wall located in the distance based on the principle of similar triangles wall and width wall :
[0018]
[0019] Among them, ζ is an adjustable fixed integer value, which is set to 5. After obtaining the size of the background wall, the corresponding positive and negative intervals are set based on the center, and a certain number of virtual point clouds are sampled in this area, for example, 100,000 points are sampled. In this way, a background wall that is far away from the camera can be constructed, which can effectively fill in the missing background point cloud data during rendering, thereby improving the integrity of scene initialization and rendering effect.
[0020] 2. Adaptive 3DGS voxelization.
[0021] In the adaptive voxelization method, the problem that the existing voxelized 3DGS framework mechanically sets the same voxel size for all data in the same dataset, making the scene too dense and causing the training speed to decrease, is solved. The voxel size is determined by calculating the average K-Nearest Neighbors (KNN) distance of each point in the scene. In this process, the KNN algorithm is used to measure the spatial distance between each point and its nearest K neighboring points, and these distances reflect the distribution density of the point cloud in the scene. By performing statistical analysis on the KNN distances of all points in the scene, the average KNN distance of the scene is calculated, and this distance value can well represent the degree of dispersion of the point cloud. Based on this average distance, the voxel size v of the scene is calculated to more reasonably represent point clouds with different density distributions.
[0022]
[0023] where p i is the i-th point, p ij is the jth neighboring point obtained by the KNN search of the i-th point, d(p i , p ij ) calculates the spatial distance between two points, where N is the number of points in the scene. Set K to 5. For each voxel in the voxelized scene, an RGB gradient queue and a covariance gradient queue are set up to detect the RGB change and covariance change of the image block formed after the voxel is projected onto the image. In the algorithm, the RGB gradient is obtained by calculating the RGB error between the image block corresponding to the voxel projected onto the image and the real image block at each iteration, and the covariance gradient is obtained by calculating the covariance gradient mean of the 3D Gaussian in the voxel. The cumulative value in the statistical gradient queue is used to determine whether a voxel is a candidate voxel that needs to be optimized. The judgment criteria are set to that the RGB gradient threshold of the voxel exceeds 10 and the covariance gradient threshold exceeds 8. The candidate voxel selection strategy is based on the following considerations: a larger RGB gradient usually leads to unstable rendering quality, while a larger covariance gradient often reflects a drastic change in the Gaussian shape, thereby showing unstable structural characteristics in the rendering.
[0024] 3. Voxel growth and deletion strategies guided by depth priors.
[0025] For example Figure 2 As shown, V i Represents a candidate voxel. Search the neighborhood of the candidate voxel and define its neighborhood voxel as V j The algorithm projects the candidate voxel and its neighboring voxels onto three reference images and calculates three error terms to evaluate the projection effect. irepresents the standard exponential decay error used to evaluate V i Whether it needs to be deleted. ΔE represents the RGB difference between the neighboring voxel and the candidate voxel, which is intended to compare the error between different voxel rendering errors. ΔD is used to measure the depth difference between the neighboring voxel and the candidate voxel, thereby evaluating its geometric consistency in the scene.
[0026] For each candidate voxel, it is projected onto the real image and the rendered image to obtain two image blocks. The algorithm will first calculate the standard exponential decay error between the two image blocks. If the error is greater than 0.9, the voxel will be deleted. Then the six neighboring voxels of the remaining candidate voxels are searched and projected onto the rendered image and the depth image to obtain image blocks. The RGB difference between the image blocks corresponding to the neighboring voxels and the candidate voxels and the depth difference between the image blocks corresponding to the neighboring voxels and the candidate voxels are calculated. If the depth difference is less than the voxel size and the RGB difference is greater than 0, and it is determined that it will not be added repeatedly, the neighboring voxel will be added to the scene.
[0027] The design idea of this part is: the standard exponential decay error can measure the difference between two image blocks. This error index combines color difference and structural similarity information to evaluate the rendering effect of voxels. A larger error value usually means that the projection effect of the voxel does not match the actual scene, which may affect the reconstruction accuracy. If the error value exceeds the preset higher threshold, the voxel is considered difficult to meet the reconstruction requirements and is moved into the pruning set, that is, the voxel is deleted. On the contrary, if the error value is within the threshold range, it indicates that the voxel has the potential to grow and can continue to be processed. The algorithm further compares the projected image blocks of the candidate voxel and the neighboring voxel, calculates the color error (i.e., the L1 difference in the color of the image block) and the depth difference (i.e., the difference in the mean depth in the image block) to evaluate the similarity and consistency between voxels. In the final judgment stage, the algorithm sets three conditions: whether the depth difference between the neighboring voxel and the candidate voxel is within the threshold, the rendering error of the growing voxel is larger than the rendering error of the candidate voxel, and no repeated addition of voxels. If all three conditions are met, the neighboring voxel is added to the growing set, indicating that the voxel is suitable for expansion to enhance the reconstruction quality of the scene.
[0028] Training phase.
[0029] Step 1: Preparation of target dataset.
[0030] The dataset used for training and testing in this method is the Waymo dataset. Waymo is an important dataset in the field of autonomous driving, providing sensor data in urban, highway and diversified scenarios, including high-resolution RGB images, 3D point clouds, sensor calibration and target annotation information. The camera resolution in the Waymo dataset is 1920*1280, covering five perspectives. The point cloud is collected using Velodyne VLP-16, and about 150,000 points can be collected per frame. First, the coordinate system in the dataset is transformed into the COLMAP format, and the point cloud and image are aligned to a consistent global coordinate system through a unified transformation to ensure the correct geometric relationship and rendering results in three-dimensional Gaussian splashing. First, the point cloud data and the RGB image of the corresponding frame are extracted, and the internal and external parameter matrices of the camera and LiDAR are obtained according to the calibration file to ensure data synchronization and coordinate alignment. Subsequently, the processed point cloud and image are transformed into a data structure adapted to 3DGS, including pose, density distribution and attributes that can be used for rendering, providing basic support for subsequent scene reconstruction and new perspective synthesis. Use any camera frame and point cloud in the Waymo dataset for 3D, and perform new perspective synthesis evaluation after ten frames.
[0031] Step 2: Training of the overall model.
[0032] The model first excludes the point cloud outside the camera's visible range, then calculates the adaptive voxel size, uses this voxel size to voxelize the scene, and saves the voxel center as an anchor point. The center coordinates of the point cloud within the camera's visible range are then calculated, and then the point cloud of the background wall is calculated with the camera pose, and each point in the point cloud is added to the model as an anchor point. During initialization, the entire model needs to initialize three MLPs and all anchor points. Each anchor point first initializes a 32-dimensional feature as a latent variable, and then initializes 100 three-dimensional Gaussian three-dimensional offsets about the anchor point. The model is set for 5,000 iterations for training, and the optimization target is the three MLPs of the model and the features and offsets in each anchor point.
[0033] During the training process, a deep prior-guided voxel growth and deletion strategy was added to the model to analyze the problems in the current scene in rendering images in the training set. By growing and deleting voxels in the scene, the densification of anchor points was completed to enhance the structure of the scene. At the same time, the standard exponential decay loss was used to learn the MLP and anchor point attributes during the entire training process.
[0034] Usage phase.
[0035] Construct the model structure according to the above method and prepare the data training set. When the training is completed, input the camera parameters of the new perspective into the trained model, and output the new perspective synthesis result of the scene at a given perspective.
[0036] Method testing.
[0037] The present invention is tested on the Waymo dataset and compared with the previously disclosed method Scaffold-GS mentioned in the above article to verify that the new perspective synthesis result obtained by the present invention has better rendering quality and scene details. The visualization results of the comparison are as follows: Figure 3 As shown, (a) is a real image, (b) is the new perspective synthesis result in "Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering" published by Lu et al. at CVPR in 2024, and (c) is the new perspective synthesis result of the present invention. The comparative quantitative evaluation results are shown in Table 1. The first row in the table is the evaluation result of Scaffold-GS by Lu et al., and the second row is the evaluation result of the present invention. The evaluation criteria used are the same as those of Lu et al., and the indicators used are peak signal-to-noise ratio PNSR, structural similarity SSIM and perceptual image score LPIPS.
[0038] Table 1
[0039] LPIPS↓ SSIM↑ PSNR↑ Scaffold-GS 0.370 0.921 23.473 Our Approach 0.238 0.878 23.801
[0040] 1. Comparison of visualization results.
[0041] exist Figure 3 Compared with Scaffold-GS, the rendering of the street is more complete, the structure of the object in the red box is more complete, there is less distortion, almost no obvious artifacts, and details such as wheels, windows and leaves are better preserved.
[0042] 2. Comparison of quantitative evaluation results.
[0043] In the quantitative evaluation results in Table 1, compared with the Scaffold-GS method, there are obvious improvements in all three indicators.
[0044] In summary, the present invention proposes a structured Gaussian splatting method based on image and radar data, aiming to improve the accuracy and efficiency of 3D scene reconstruction under single viewpoint data. By introducing virtual background wall, adaptive voxelization and depth prior guidance strategy, the point cloud sparsity and geometric consistency problems are effectively solved. Experimental results show that this method has significant advantages in reconstruction quality and new perspective synthesis. The effectiveness of this method is proved by testing and comparing the Scaffold-GS results.
Claims
1. A structured Gaussian splashing method based on image and radar data, characterized in that:
1. Virtual background wall construction: The module constructs a virtual background wall behind the camera field of view to fill the sparsity of point clouds in distant areas. The specific method is to use camera parameters and depth information to infer the position and size of the background wall, and generate a set of virtual point clouds in this area.
2. Adaptive 3DGS voxelization: This module adaptively adjusts the voxel size by analyzing the distribution density of the point cloud in the scene; and uses the K nearest neighbor algorithm to calculate the average neighbor distance of the point cloud; 3. Voxel growth and deletion strategy guided by depth prior: This module guides the growth and optimization of voxels based on the prior information generated by the depth map. It screens out candidate voxels by detecting the RGB gradient changes and depth consistency of voxels in the projected image, and performs semantic and depth constraints on their neighborhood to decide whether to expand or delete them.
2. The structured Gaussian splashing method based on image and radar data according to claim 1, characterized in that: The proposed virtual background wall construction method ensures that the virtual background wall is parallel to the camera imaging plane. First, the scene center x is determined by calculating the mean of all points in the scene on the three axes. scene , and then use the camera pose x provided by the dataset cam The distance d between the center of the scene and the camera is calculated by calculating x scene and x cam The Euclidean space distance between them; then the normalized vector x cam -x scene As the normal vector of the virtual background wall; Next, determine the range of the background wall: use the focal length information f of the camera in the x-axis and y-axis directions x With f y , and the length and width pixel values of the image w ing and width img Calculate the actual physical distance corresponding to each pixel on the camera plane, and then infer the length w of the virtual background wall located in the distance based on the principle of similar triangles wall and width wall : Among them, ζ is an adjustable fixed integer value, which is set to 5. After the size of the background wall is obtained, the corresponding positive and negative intervals are set based on the center, and virtual point clouds are generated by sampling in this area; 2. Adaptive 3DGS voxelization; In the adaptive voxelization method, the voxel size is determined by calculating the average K nearest neighbor KNN distance of each point in the scene; By performing statistical analysis on the KNN distances of all points in the scene, the average KNN distance of the scene is calculated. Based on this average distance, the voxel size v of the scene is calculated to more reasonably represent point clouds with different density distributions. where p i is the i-th point, p ij is the jth neighboring point obtained by the KNN search of the i-th point, d(p i ,p ij ) calculates the spatial distance between two points, where N is the number of points in the scene; K is set to 5; for each voxel in the voxelized scene, an RGB gradient queue and a covariance gradient queue are set up to detect the RGB change and covariance change of the image block formed after the voxel is projected onto the image; The RGB gradient is obtained by calculating the RGB error between the image block corresponding to the voxel projected onto the image and the real image block at each iteration, and the covariance gradient is obtained by calculating the covariance gradient mean of the 3D Gaussian in the voxel; the cumulative value in the statistical gradient queue is used to determine whether a voxel is a candidate voxel that needs to be optimized. The judgment standard is set to the voxel RGB gradient threshold exceeding 10 and the covariance gradient threshold exceeding 8; 3. Voxel growth and deletion strategies guided by depth priors; V i Represents a candidate voxel; search the neighborhood of the candidate voxel and define its neighborhood voxel as V j ; The algorithm projects the candidate voxel and its neighboring voxels onto three reference images and calculates three error terms to evaluate the projection effect; E i represents the standard exponential decay error used to evaluate V i Whether it needs to be deleted; ΔE represents the RGB difference between the adjacent voxel and the candidate voxel, which is intended to compare the error between different voxel rendering errors; ΔD is used to measure the depth difference between the neighboring voxel and the candidate voxel, thereby evaluating its geometric consistency in the scene; For each candidate voxel, it is projected onto the real image and the rendered image to obtain two image blocks. The algorithm will first calculate the standard exponential decay error between the two image blocks. If the error is greater than 0.9, the voxel will be deleted. Then, the six neighboring voxels of the remaining candidate voxels are searched and projected onto the rendered image and the depth image to obtain image blocks. The RGB difference between the image blocks corresponding to the neighboring voxels and the candidate voxels and the depth difference between the image blocks corresponding to the neighboring voxels and the candidate voxels are calculated. If the depth difference is less than the voxel size and the RGB difference is greater than 0, and it is determined that it will not be added repeatedly, the neighboring voxel will be added to the scene. The projected image blocks of the candidate voxels and the neighboring voxels are further compared to calculate the color error, i.e., the L1 difference in the color of the image block, and the depth difference, i.e., the difference in the mean depth in the image block, to evaluate the similarity and consistency between the voxels. In the final judgment stage, three conditions are set: the depth difference between the neighboring voxel and the growing voxel is within the size of a voxel, the rendering error of the growing voxel is larger than that of the candidate voxel, and the voxel is not added repeatedly. If all three conditions are met, the neighboring voxel is added to the growing set, indicating that the voxel is suitable for expansion to enhance the reconstruction quality of the scene. The dataset is used to set up more than 5,000 iterations for training. When the training is completed, the camera parameters of the new perspective are input into the trained model, and the output is the new perspective synthesis result of the scene at a given perspective.
Citation Information
Cited By
Dynamic deletion method and system for three-dimensional Gaussian spattering
CN120635289A
Dynamic scene reconstruction method and system and related equipment
CN120689555A
A dynamic scene reconstruction method, system and related devices
CN120689555B
Cloud Gaussian splashing scene automatic generation and multi-terminal distribution method
CN121982220A