Scene three-dimensional reconstruction method based on camera and laser radar fusion and related device
By adopting camera lidar fusion and 3DGS technology in the SLAM method, combining photometric and geometric losses to optimize the position, and introducing scale regularization losses, the performance problems of the SLAM method in strong light and textureless environments are solved, and accurate pose estimation and high-quality rendering are achieved.
Patent Information
- Application Number
- CN202510155379.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The existing SLAM method performs poorly in environments such as strong lighting changes and textureless areas, and the 3DGS-based method is affected by floating Gaussians in pose estimation, resulting in geometric inconsistency problems.
A three-dimensional reconstruction method based on camera lidar fusion is adopted to perform sensor tracking, map update and loop detection through 3DGS, combining photometric loss and geometric loss for posture optimization, and by introducing scale regularization loss and size loss, the introduction of floating Gaussians is avoided.
Accurate posture estimation and high-quality image rendering in environments such as strong lighting changes and textureless areas are achieved, avoiding geometric inconsistencies caused by floating Gaussians, and improving the robustness and adaptability of the system.
Smart Images

Figure CN120088403A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of scene three-dimensional reconstruction, and particularly relates to a method and related device for scene three-dimensional reconstruction based on camera-lidar fusion. Background Art
[0002] Simultaneous Localization and Mapping (SLAM) has been an active research field in the past few decades, aiming to simultaneously reconstruct an unknown environment and localize the sensor pose; as a fundamental task in 3D computer vision, SLAM is crucial for more advanced tasks such as autonomous driving, embodied intelligence, and robot navigation.
[0003] Currently, traditional SLAM methods use point clouds, meshes, and surfelClouds as map representations, and these methods demonstrate excellent real-time performance and scalability; however, in environments such as strong light changes and textureless areas, their performance is poor.
[0004] With the development of Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) in the field of novel view synthesis, these two methods have been applied to some SLAM systems, which can provide better surface modeling and more robust noise processing. Compared with NeRF, 3DGS renders images faster; in addition, NeRF uses neural networks to represent the scene, while 3DGS uses millions of Gaussian ellipsoids distributed throughout the space, and the map capacity can be easily expanded by adding more Gaussian ellipsoids, so 3DGS is more suitable for the incremental map construction process in SLAM. In summary, in SLAM, using 3DGS can achieve more accurate pose estimation and higher-quality rendering.
[0005] The application of 3DGS in visual-LiDAR SLAM systems mainly focuses on the map construction phase. Such methods first determine the sensor pose based on traditional LiDAR SLAM algorithms, then initialize a Gaussian map with LiDAR (Light Detection and Ranging) point clouds, and finally optimize the 3D Gaussian through image rendering loss. During the initialization and optimization of the Gaussian map, the sensor pose is kept fixed. This indicates that 3DGS is mainly used for map construction in visual-LiDAR SLAM systems, and its potential in the pose tracking part has not been fully exploited. Additionally, it should be emphasized that in existing 3DGS-based SLAM methods, pose estimation is affected by floating Gaussians introduced by gradient-based Gaussian densification strategies, which can lead to geometric inconsistency problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and related device for three-dimensional scene reconstruction based on camera-LiDAR fusion to solve one or more of the above-mentioned technical problems. The technical solution provided by the present invention can achieve accurate pose estimation and high-quality image rendering during three-dimensional scene reconstruction in environments such as strong light changes and textureless areas.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In the first aspect of the present invention, a method for three-dimensional scene reconstruction based on camera-LiDAR fusion is provided, including the following steps:
[0009] Based on the selected scene, obtain the relative pose between the camera and the LiDAR, as well as the camera images and LiDAR point cloud sensor data of the scene;
[0010] According to the obtained relative pose between the camera and the LiDAR, as well as the camera images and LiDAR point cloud sensor data of the scene, sequentially perform sensor tracking, map update, and loop detection based on 3DGS to obtain the three-dimensional scene reconstruction result;
[0011] Among them, in the sensor tracking stage, first initialize the poses of the camera and the LiDAR through the linear motion hypothesis, then optimize the poses by minimizing the photometric loss and geometric loss, and finally obtain the optimized poses of the camera and the LiDAR; after sensor tracking is completed, enter the map update stage. First, render the transparency map and depth map according to the current pose of the camera, add the LiDAR point clouds that meet the preset conditions to the Gaussian point cloud map, and then optimize the attributes of the Gaussian point cloud according to the photometric loss and depth loss to obtain the Gaussian map; after the map update is completed, enter the loop detection stage, optimize the Gaussian map, and use the optimized Gaussian map as the three-dimensional scene reconstruction result.
[0012] A further improvement of the 3D reconstruction method for the present invention scenario lies in that
[0013] in the step of performing pose optimization by minimizing photometric loss and geometric loss to finally obtain the optimized poses of the camera and lidar, the overall loss function L track is expressed as:
[0014]
[0015] where λ pho , λ geo are respectively the weight terms of photometric loss and geometric loss; L pho , L geo are respectively photometric loss and geometric loss.
[0016] A further improvement of the 3D reconstruction method for the present invention scenario lies in that
[0017] in the calculation expression of the overall loss function L track :
[0018]
[0019] where λ 1 is a hyperparameter; ∥*∥1 represents the L1 norm; represents the structural similarity between the rendered image and the ground truth image I;
[0020]
[0021] where ω ij is a weight term; d ij is the distance between two points, d ij =p i -p j , p i is a point in the lidar point cloud P c , p j is the nearest neighbor point of p l in the local keyframe point cloud P i ; ∑′ j is the covariance matrix corresponding to the point p j ; T c is the pose of the current lidar; ∑ i ′ is the covariance matrix corresponding to the point p i .
[0022] A further improvement of the 3D reconstruction method for the present invention scenario lies in that
[0023] In the step of optimizing the attributes of the Gaussian point cloud according to the photometric loss and the depth loss to obtain the Gaussian map, when optimizing the attributes of the Gaussian point cloud, a scale regularization loss and a size loss are also added, and the overall optimization loss function L map is calculated as follows:
[0024]
[0025] where λ d , λ reg , and λ size are the weights of the depth loss, the scale regularization loss, and the size loss respectively; L d is the depth loss; L reg is the scale regularization loss; L size is the size loss.
[0026] A further improvement of the 3D reconstruction method for the scene of the present invention lies in that
[0027] in the calculation expression of the overall optimization loss function L map :
[0028]
[0029] where I d is the depth map; x i , d i are the two-dimensional projection point obtained by projecting the point p i onto the image plane and the corresponding projection depth respectively;
[0030]
[0031] where max(s i ) and min(s i ) represent obtaining the maximum and minimum elements respectively; represents the scale of the Gaussian ellipsoid in three directions;
[0032]
[0033] where f is the focal length of the camera.
[0034] A further improvement of the 3D reconstruction method for the scene of the present invention lies in that
[0035] in the step of optimizing the attributes of the Gaussian point cloud according to the photometric loss and the depth loss to obtain the Gaussian map, the attributes include: rotation, scaling, opacity, and RGB color.
[0036] A further improvement of the 3D reconstruction method for the scene of the present invention lies in that
[0037] The steps of optimizing the Gaussian map and using the optimized Gaussian map as the result of the three-dimensional reconstruction of the scene include:
[0038] First, based on the camera images of the scene, the lidar point cloud sensor data, and the optimized poses of the camera and lidar, the cumulative errors of the poses of the camera and lidar are corrected by detecting loop frames and optimizing the pose graph to obtain the corrected poses. Then, the Gaussian map is optimized according to the corrected poses, and the optimized Gaussian map is used as the result of the three-dimensional reconstruction of the scene.
[0039] In the second aspect of the present invention, a three-dimensional reconstruction system for a scene based on camera-lidar fusion is provided, including:
[0040] A data acquisition module for obtaining the relative pose between the camera and the lidar and the camera images of the scene and the lidar point cloud sensor data based on the selected scene;
[0041] A three-dimensional reconstruction module for performing sensor tracking, map updating, and loop detection in sequence based on 3DGS according to the obtained relative pose between the camera and the lidar and the camera images of the scene and the lidar point cloud sensor data to obtain the result of the three-dimensional reconstruction of the scene;
[0042] Among them, in the sensor tracking stage, the poses of the camera and the lidar are initialized by the linear motion assumption first, and then the pose optimization is performed by minimizing the photometric loss and the geometric loss to finally obtain the optimized poses of the camera and the lidar. After the sensor tracking is completed, it enters the map updating stage. First, the transparency map and the depth map are rendered according to the current pose of the camera, and the lidar point cloud that meets the preset conditions is added to the Gaussian point cloud map. Then, the attributes of the Gaussian point cloud are optimized according to the photometric loss and the depth loss to obtain the Gaussian map. After the map updating is completed, it enters the loop detection stage, the Gaussian map is optimized, and the optimized Gaussian map is used as the result of the three-dimensional reconstruction of the scene.
[0043] In the third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for three-dimensional reconstruction of a scene based on camera-lidar fusion as described in any one of the first aspects of the present invention.
[0044] In the fourth aspect of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for three-dimensional reconstruction of a scene based on camera-lidar fusion as described in any one of the first aspects of the present invention.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] In the 3D scene reconstruction method based on camera-LiDAR fusion disclosed in the present invention, a method of simultaneously using 3DGS for map representation and camera tracking is proposed, which can be called VLGS-SLAM (Vision LiDAR GaussianSplatting SLAM). The input of the method of the present invention is strictly synchronized camera images, LiDAR point cloud data, and the relative pose between the two sensors. During the sensor tracking stage, the pose is optimized by minimizing the photometric loss and geometric loss. During the map update stage, the LiDAR point cloud is used as a Gaussian ellipsoid to enhance the accuracy of pose estimation, avoiding the geometric inconsistency problem caused by floating Gaussians, and finally accurate pose estimation and high-quality image rendering can be achieved.
[0047] In the existing 3DGS optimization process, the photometric loss causes the Gaussians in the distance to be amplified, so that these Gaussians are used to render the background of the scene. However, as the camera moves forward, these background Gaussians become closer to the camera, and then visual artifacts are generated, thus interfering with tracking and rendering. To solve the above problems, in the preferred technical solution of the present invention, a loss function is introduced to regularize the scale of the Gaussian ellipsoid, improving the robustness of pose estimation.
[0048] In the existing SLAM methods based on 3DGS, pose estimation is affected by floating Gaussians, which are introduced by the gradient-based Gaussian densification strategy. To reduce their influence, VLGS-SLAM assigns additional attributes to the LiDAR points (for example, including rotation, scaling, opacity, and RGB color), and then forms new Gaussian ellipsoids. This method can avoid introducing floating Gaussians. Brief Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art; obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a schematic flow chart of a 3D scene reconstruction method based on camera-LiDAR fusion in an embodiment of the present invention;
[0051] Figure 2 It is a schematic principle diagram of a 3D scene reconstruction method based on camera-LiDAR fusion in an embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of a 3D scene reconstruction system based on camera-LiDAR fusion in an embodiment of the present invention. Detailed Embodiments
[0053] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention; obviously, the described embodiment technical solutions are part of the embodiments of the present invention, not all of the embodiments.
[0054] Based on the technical solutions disclosed in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0055] Please refer to Figure 1 and Figure 2 A 3D scene reconstruction method based on camera-lidar fusion provided by an embodiment of the present invention includes the following steps:
[0056] Step 1, based on the selected scene, obtain the relative pose between the camera and the lidar, as well as the camera images of the scene and the lidar point cloud sensor data.
[0057] Step 2, according to the relative pose between the obtained camera and lidar, as well as the camera images of the scene and the lidar point cloud sensor data, perform sensor tracking, map update and loop detection in sequence based on 3DGS to obtain the 3D scene reconstruction result.
[0058] Among them, in the sensor tracking stage, first initialize the poses of the camera and the lidar through the linear motion hypothesis, and then optimize the poses by minimizing the photometric loss and the geometric loss to finally obtain the optimized poses of the camera and the lidar; after the sensor tracking is completed, enter the map update stage. First, render the transparency map and the depth map according to the current pose of the camera, add the lidar point cloud that meets the preset conditions to the Gaussian point cloud map, and then optimize the attributes of the Gaussian point cloud according to the photometric loss and the depth loss to obtain the Gaussian map; after the map update is completed, enter the loop detection stage, optimize the Gaussian map, and use the optimized Gaussian map as the 3D scene reconstruction result; in a further exemplary technical solution, in the loop detection stage, first, according to the camera images of the scene, the lidar point cloud sensor data, and the optimized poses of the camera and the lidar, correct the cumulative error of the poses of the camera and the lidar by detecting the loop frames and optimizing the pose graph to obtain the corrected poses, and then optimize the Gaussian map according to the corrected poses to obtain the 3D scene reconstruction result.
[0059] In the technical solution disclosed in the embodiment of the present invention, a 3D Gaussian ellipsoid is used for map representation and camera tracking, which achieves significant technical progress. The core of the method of the embodiment of the present invention is to fuse two different types of sensor data, camera images and lidar point clouds, and use the strict synchronization and relative posture information between them to achieve accurate posture estimation and high-quality image rendering. This method not only improves the accuracy of three-dimensional reconstruction, but also enhances the robustness and adaptability of the system. Further specifically explanatory, the camera image provides rich texture information and color information, which helps to build a more realistic three-dimensional scene; the lidar point cloud provides accurate depth information and geometric structure, which is the basis of three-dimensional reconstruction; the relative posture between the two sensors ensures that the camera and lidar data can be accurately fused. The present invention uses 3DGS for map representation, which can flexibly describe complex three-dimensional scenes while maintaining a high representation efficiency; combining camera images and lidar point clouds, accurate tracking of cameras and lidars is achieved through optimization algorithms, thereby obtaining continuous posture information.
[0060] Furthermore, the present invention forms a new Gaussian ellipsoid by assigning additional attributes (such as rotation, scaling, opacity and RGB color) to the LiDAR points, which effectively avoids the introduction of floating Gaussians and improves the accuracy of attitude estimation.
[0061] Explanatory, 3DGS is a three-dimensional scene expression method, the entire scene is expressed by multiple (usually millions) 3D Gaussian ellipsoids, each Gaussian ellipsoid is a three-dimensional Gaussian distribution; during the training process, the following properties of each Gaussian ellipsoid need to be optimized: center (i.e., the mean of the Gaussian distribution), scale (i.e., the variance of the Gaussian distribution), transparency and color. In the technical solution provided by the embodiment of the present invention, each Gaussian ellipsoid has the following properties: transparency o∈[0,1], color Rotation scale Mean Using these properties and the camera's pose, each Gaussian ellipsoid can be projected onto an image to render an image.
[0062] In the specific exemplary technical solution, the specific rendering process of each pixel is introduced as follows:
[0063] Assuming the current pixel is p, first arrange the Gaussian ellipsoids in space from near to far according to the distance from the camera plane, and then calculate the transparency of each Gaussian ellipsoid to the pixel. The calculation expression is:
[0064]
[0065] In the formula, α i is the transparency of the current Gaussian to pixel p; αi ′ is the transparency of the Gaussian ellipsoid itself; is the projection position of the center (mean value) of the Gaussian ellipsoid on the image plane; Σ is the variance of the projection of the Gaussian ellipsoid on the image plane;
[0066] Then, the weight of each Gaussian ellipsoid for the current pixel can be obtained, and the calculation expression is:
[0067]
[0068] In the formula, α j is the transparency corresponding to other Gaussians arranged before the current Gaussian; ω i is the weight of the current Gaussian;
[0069] Finally, the calculation expression for the rendering color of the current pixel is:
[0070]
[0071] In the formula, is the rendering color of the current pixel p; c i is the color of each Gaussian ellipsoid; M is the number of all Gaussian ellipsoids;
[0072] Repeating the above process for each pixel of the current image, the rendered image rendered by the Gaussian can be obtained Then, the rendered image and the ground truth image I are used to calculate the photometric loss, and the calculation expression is:
[0073]
[0074] In the formula, represents the structural similarity between the two images; ∥*∥1 represents the L1 norm; λ 1 is a hyperparameter used to balance the measurement of the two errors;
[0075] Finally, the loss is backpropagated, and all attributes of the Gaussian ellipsoid can be updated.
[0076] In an embodiment of the present invention, since the pose of the camera is used in the rendering process, the gradient of the loss function will also be backpropagated to the camera pose, thereby optimizing the camera pose. In the technical solution of the embodiment of the present invention, in the pose estimation stage, the attributes of the Gaussian will be fixed, and the photometric loss will only optimize the camera pose; in the mapping stage, the camera pose will be fixed, and the photometric loss will only optimize the attributes of the Gaussian. In addition, since the Gaussian ellipsoid in the method of the present invention is obtained from the lidar point cloud, its position will not be optimized, and only when the position of the lidar changes, the position of the Gaussian ellipsoid will change.
[0077] Based on the above embodiments, in the preferred embodiment of the present invention, during the sensor tracking process, first, the poses of the camera and the lidar are initialized through a linear motion assumption, and then further optimized by minimizing the photometric loss and the geometric loss. In the embodiment of the present invention, an image is rendered based on a Gaussian map using the current camera pose, and the photometric loss is calculated using the above-mentioned photometric loss calculation formula. To avoid the influence of newly observed regions, the method of the present invention also renders an opacity image and only calculates the loss corresponding to pixels with an opacity greater than a preset threshold (exemplarily, such as 0.9). Similar to other 3DGS-based SLAM methods, in the sensor tracking stage of the present invention, the SSIM loss is not used and λ 1 = 0. The plane-to-plane distance between the current LiDAR point cloud P c and the local keyframe point cloud P l is used as the geometric loss. Exemplarily, P l here is a set of LiDAR point clouds from the previous 5 keyframes.
[0078] In a specific exemplary technical solution, first, the lidar point cloud is downsampled using a voxel grid with a resolution of (0.5m, 0.5m, 0.5m); then, for each point p i in the point cloud, the k nearest neighbor points of it are calculated using the K-nearest neighbor algorithm, and the set of these points is denoted as
[0079] The covariance matrix of is:
[0080]
[0081] In the formula, p j is the nearest neighbor point of the current point p i ; is 's mean value;
[0082] Then, the covariance matrix is decomposed by SVD to obtain ∑ i = U i D i V i T ;
[0083] Next, the diagonal matrix D i is changed to Here, ε is a very small quantity, generally set to 0.001;
[0084] Through this parameterization, the final covariance matrix is ∑′ i = U i D i ′V iT ;
[0085] This covariance matrix represents that the uncertainty of the current point in the direction of the normal vector of the current plane is very low, but the uncertainty of its specific position on the current plane is relatively high; that is to say, it can be determined that this point is on the plane, but it is not certain where exactly it is on the plane.
[0086] For the LiDAR point cloud P c and the local key-frame point cloud P l the above process is repeated for both; for each point p c in P i , its nearest neighbor point p l ∈P j is found in P l , and the distance between these two points is d ij =p i -p j .
[0087] The final geometric loss can be defined as:
[0088]
[0089] where is the geometric loss; T c is the pose of the current LiDAR; ω ij is a weight term used to filter outliers; ∑′ j is the covariance matrix corresponding to the point p j ; ∑ i ′ is the covariance matrix corresponding to the point p i ; if ∥d ij ∥ 2 >0.5 then ω ij =0, otherwise ω ij =1.
[0090] In the sensor tracking stage, the total loss is:
[0091]
[0092] where λ pho and λ geo are two weight terms;
[0093] The entire tracking step is iterative and needs to be iterated 30 times to obtain better camera and LiDAR poses. In the tracking stage, the properties of the 3D Gaussian ellipsoid are fixed, and only the poses of the camera and the radar are variable.
[0094] The key-frame selection strategy in the VLGS-SLAM of the present invention is as follows: If the distance between the current frame and the previous key frame exceeds nkf meters, or the rotation angle of the camera since the last key frame exceeds r kf degrees, then this frame is regarded as a new key frame.
[0095] In the embodiments of the present invention, after the sensor tracking is completed, the poses of the camera and the lidar are fixed, and the Gaussian map corresponding to the new observation area needs to be further optimized so as to better represent the current scene.
[0096] In a specific exemplary technical solution, first, a transparency map I o and a depth map I d are rendered according to the current camera pose; then, each point p c of the current lidar point cloud P i is projected onto the image plane to obtain a two-dimensional projection point x i = [u i , v i T , and the corresponding projection depth d i ; if any of the following conditions is satisfied, p i is added to the Gaussian point cloud map:
[0097] 1) I o (x i ) < 0.9;
[0098] 2) d i < 0.8 * I d (x i ).
[0099] After adding a new Gaussian, the data of the current frame and the previous 5 frames are required to update the local map; the map update process continues for 30 iterations, and in each iteration, a frame is randomly selected and an RGB image I rgb and a depth image I d are rendered.
[0100] Then the photometric loss is calculated, and the depth loss is calculated using the following formula:
[0101]
[0102] In the formula, L d is the depth loss.
[0103] In the original 3DGS, the scale of each Gaussian ellipsoid is unconstrained, which enables the Gaussian ellipsoid to better capture scene details and render more accurate images; however, this also causes the Gaussian to appear needle-like, that is, the scale is larger in one direction and smaller in the other two directions. These needle-shaped Gaussians have a negative impact on camera tracking because their appearance changes significantly with different viewing angles, which in turn causes many artifacts in the rendered image; in addition, the Gaussians far away from the camera are usually larger in scale because they are rendered as the background, but as the camera approaches, these distant Gaussians will gradually approach the camera. Because of their larger scale, they will affect most of the pixels, which may result in poor rendering quality. In order to solve the above problems, the present invention introduces scale regularization loss and size loss. For each Gaussian ellipsoid, the present invention uses To represent its scale in three directions, the scale regularization loss can be expressed as:
[0104]
[0105] Where, L reg is the scale regularization loss; max(s i )、min(s i ) represent obtaining the largest and smallest elements respectively;
[0106] The size loss is calculated as:
[0107]
[0108] Where, L size is the size loss; f is the focal length of the camera; d i is the depth of Gauss;
[0109] The meaning of size loss is that after all Gaussian ellipsoids are projected onto the image, the radius of the projection is less than 100 pixels, so that the size of the Gaussian can be limited.
[0110] In summary, all losses in the map update phase are:
[0111]
[0112] In the formula, λ pho , d , reg , size is the weight term.
[0113] During the map update phase, only the properties of the Gaussian ellipsoid are optimized, and the pose of the sensor is fixed. Specifically, the overall process of the technical solution of the embodiment of the present invention is divided into three stages: the sensor tracking stage, the map update stage, and the loop detection stage. Among them, in the sensor tracking stage, the sensor pose is estimated based on the existing map. In this stage, the sensor pose is variable and the map is fixed. The map update stage is to update the existing map according to the existing sensor pose. Therefore, the sensor pose needs to be fixed at this time, and only the map is optimized. The Gaussian ellipsoid here is the basic unit that constitutes the map, and optimizing the properties of the Gaussian ellipsoid is optimizing the map.
[0114] In the embodiment of the present invention, in the SLAM system, cumulative error is inevitable during loop detection because the small tracking error of each frame will accumulate over time, resulting in significant drift of the trajectory eventually. The loop detection module can correct the accumulated error by detecting loop frames and pose graph optimization, thus solving this problem. The visualization results before and after loop closure are shown in the last part of Figure 2 The core part of loop detection is to identify loop frames.
[0115] In an exemplary technical solution, for the current frame (I c , P c ), the NetVLAD algorithm is used to find the 10 key frames {(I i , P i ), i = 1, 2,..., 10} that are closest to it. Then, for each LiDAR point cloud p i in these candidate loop frames, the geometric loss L geo is used to register the current point cloud. If the average point-to-point distance between the two point clouds is less than 5 cm, it is considered that the registration is good, and the current frame and the i-th frame can be considered to form a loop.
[0116] To accelerate the loop detection stage, each current frame forms at most two loop constraints, and other candidate frames will be discarded. Subsequently, the pose graph is used to optimize all the frames in the loop, and this optimization framework is implemented based on the g2o framework. In the pose graph, each vertex represents the pose of a data frame, and the edges between the vertices represent the pose from tracking and the relative pose constraints of loop closure. During the optimization process, the loop constraint is crucial because inaccurate loop closure will lead to incorrect pose optimization. To enhance robustness and accuracy, pose graph optimization is only performed when loop closure is detected in three consecutive frames. After optimization, all the key frames in the loop and the mapping loss L map are used to update the local map. After loop detection is completed, the processing of the current input frame is completed. Subsequently, the system loads the next frame in the dataset and executes the sensor tracking, map update, and loop detection steps again.
[0117] In summary, cameras and lidars are two types of sensors commonly used in 3D reconstruction. Vision methods based on cameras are vulnerable to textureless regions and lighting, while lidar-based methods tend to degrade in scenes lacking significant structural features. Therefore, the current mainstream approach is to fuse images and lidar for 3D reconstruction. Currently, the 3D Gaussian splash method is very effective in the vision field. The present invention proposes an algorithm based on 3DGS to fuse image and lidar information for pose estimation and 3D reconstruction, which can be called VLGS-SLAM. To ensure the accuracy of scene geometry and reduce pose estimation errors caused by floating Gaussians, the method of the present invention uses LiDAR points as 3D Gaussian ellipsoids; to enhance the performance of pose estimation, the present invention applies regularization to the scale attribute of the Gaussian ellipsoids, restricting the shape of each Gaussian ellipsoid; in addition, in terms of loop closure detection, the present invention combines image similarity and LiDAR point cloud distance to effectively detect and close loops.
[0118] In specific exemplary experiments, the experimental results show that the VLGS-SLAM of the present invention has reached the state-of-the-art level in the field of SLAM based on 3DGS, outperforming many traditional SLAM algorithms.
[0119] In the specific embodiments of the present invention, the experimental comparisons are as follows:
[0120] 1) Test dataset: The KITTI dataset is used for performance testing of the algorithm, and VLGS-SLAM is compared with other SLAM methods. The KITTI dataset is a publicly available large-scale outdoor dataset, which contains a total of 11 sequences with a total trajectory length of 22 kilometers.
[0121] 2) Evaluation metric: The Absolute Trajectory Error (ATE), which is commonly used in the field of SLAM, is used for evaluation. This error calculates the distance between the optical center of each camera and the ground truth, and then calculates the root mean square of all distances to finally obtain the trajectory error. This error metric is in meters, and the lower the error, the more accurate the trajectory.
[0122] 3) Comparison method: Multiple different types of SLAM methods are adopted; among them, the vision SLAM methods are ORBSLAM3 and OpenVSLAM; the lidar SLAM methods are MULLS, Fast-LIO2, and LeGO-LOAM; the image-lidar fusion SLAM method is CamVox; the SLAM methods based on 3DGS are MonoGS and SplaTAM.
[0123] 4) Result comparison: The trajectory errors of each method on the KITTI dataset are shown in Table 1; among them, X represents that the current sequence cannot be run; the black bold represents the best result of the current sequence, and the underline represents the second-best result of the current sequence.
[0124] Table 1. Comparison results of trajectory errors
[0125]
[0126] As can be seen from Table 1, the method disclosed in the embodiment of the present invention achieves the best results on 5 sequences, and is also sub-optimal on the other 6 sequences. The final average result is also the best, showing significant improvement and progress compared with the existing solutions.
[0127] The following is the device embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For the details not disclosed in the device embodiment, please refer to the method embodiment of the present invention.
[0128] Please refer to Figure 3 , in the embodiment of the present invention, a three-dimensional scene reconstruction system based on camera-lidar fusion is provided, including:
[0129] A data acquisition module, configured to obtain the relative pose between the camera and the lidar, as well as the camera image and lidar point cloud sensor data of the scene, based on the selected scene;
[0130] A three-dimensional reconstruction module, configured to perform sensor tracking, map update, and loop detection in sequence based on 3DGS according to the relative pose between the obtained camera and lidar, as well as the camera image and lidar point cloud sensor data of the scene, to obtain the three-dimensional scene reconstruction result;
[0131] Among them, in the sensor tracking stage, the poses of the camera and the lidar are first initialized through the linear motion hypothesis, and then the poses are optimized by minimizing the photometric loss and the geometric loss. Finally, the optimized poses of the camera and the lidar are obtained; after the sensor tracking is completed, it enters the map update stage. First, the transparency map and the depth map are rendered according to the current pose of the camera, and the lidar point cloud that meets the preset conditions is added to the Gaussian point cloud map. Then, the attributes of the Gaussian point cloud are optimized according to the photometric loss and the depth loss to obtain the Gaussian map; after the map update is completed, it enters the loop detection stage, and the Gaussian map is optimized, and the optimized Gaussian map is used as the three-dimensional scene reconstruction result.
[0132] In an embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used to execute the operations of the method for three-dimensional reconstruction of a scene based on camera-lidar fusion.
[0133] In an embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed Random Access Memory (RAM) or a non-volatile memory, such as at least one disk memory. The one or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the method for three-dimensional reconstruction of a scene based on camera-lidar fusion in the above embodiment.
[0134] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) that contain computer-usable program code.
[0135] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the processes and / or Figure 1 blocks or multiple blocks.
[0136] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the processes and / or Figure 1 blocks or multiple blocks.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes and / or Figure 1 blocks or multiple blocks.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A scene 3D reconstruction method based on camera-lidar fusion, characterized in that: The following steps are involved: Based on the selected scene, obtain the relative posture between the camera and the lidar, as well as the camera image and lidar point cloud sensor data of the scene; Based on the relative posture between the camera and the LiDAR, the camera image of the scene, and the LiDAR point cloud sensor data, sensor tracking, map update, and loop detection are performed in sequence based on 3DGS to obtain the 3D reconstruction result of the scene. Among them, in the sensor tracking stage, the camera and lidar postures are first initialized by the linear motion hypothesis, and then the postures are optimized by minimizing the photometric loss and geometric loss, and finally the optimized camera and lidar postures are obtained; after the sensor tracking is completed, the map update stage is entered, and the transparency map and depth map are first rendered according to the current camera posture, and the lidar point cloud that meets the preset conditions is added to the Gaussian point cloud map, and then the properties of the Gaussian point cloud are optimized according to the photometric loss and depth loss to obtain the Gaussian map; after the map update is completed, the loop detection stage is entered to optimize the Gaussian map, and the optimized Gaussian map is used as the scene 3D reconstruction result.
2. The method for scene 3D reconstruction based on camera-lidar fusion according to claim 1, characterized in that: In the step of optimizing the pose by minimizing the photometric loss and the geometric loss to finally obtain the optimized pose of the camera and the lidar, the overall loss function L is used. track It is expressed as: In the formula, λ pho , geo They are the weight terms of photometric loss and geometric loss respectively; L pho , L geo They are photometric loss and geometric loss respectively.
3. The method for scene 3D reconstruction based on camera-lidar fusion according to claim 2, characterized in that: The overall loss function L track In the calculation expression: Where λ1 is a hyperparameter; ∥*∥1 represents the L1 norm; Represents a rendered image Structural similarity with the true image I; In the formula, ω ij is a weight term; d ij is the distance between two points, d ij =p i -p j , p i is the LiDAR point cloud P c The point in the j is the local keyframe point cloud P l Medium i The nearest neighbor of j It's point p j The corresponding covariance matrix; T c is the current position of the lidar; ∑ i ′ is point p i The corresponding covariance matrix.
4. The method for scene 3D reconstruction based on camera-lidar fusion according to claim 1, characterized in that: In the step of optimizing the properties of the Gaussian point cloud according to the photometric loss and the depth loss to obtain the Gaussian map, when optimizing the properties of the Gaussian point cloud, the scale regularization loss and the size loss are also added, and the overall optimization loss function L map The calculation expression is: In the formula, λ d , reg , size are the weights of depth loss, scale regularization loss and size loss respectively; L d is the depth loss; L reg is the scale regularization loss; L size For size loss.
5. The method for scene 3D reconstruction based on camera-lidar fusion according to claim 4, characterized in that: The overall optimization loss function L map In the calculation expression: In the formula, I d is the depth map; x i d i They are point p i The two-dimensional projection points and the corresponding projection depth obtained by projecting onto the image plane; In the formula, max(s i )、min(s i ) represent obtaining the largest and smallest elements respectively; Represents the scale of the Gaussian ellipsoid in three directions; Where f is the focal length of the camera.
6. The method for scene 3D reconstruction based on camera-lidar fusion according to claim 1, characterized in that: In the step of optimizing the properties of the Gaussian point cloud according to the photometric loss and the depth loss to obtain the Gaussian map, the properties include: rotation, scaling, opacity and RGB color.
7. The method for scene 3D reconstruction based on camera-lidar fusion according to claim 1, characterized in that: The step of optimizing the Gaussian map and using the optimized Gaussian map as the result of the three-dimensional reconstruction of the scene includes: First, based on the camera image of the scene, the lidar point cloud sensor data, and the optimized camera and lidar poses, the accumulated errors of the camera and lidar poses are corrected by detecting loop frames and optimizing the pose graph to obtain the corrected pose. Then, the Gaussian map is optimized according to the corrected pose, and the optimized Gaussian map is used as the scene 3D reconstruction result.
8. A scene 3D reconstruction system based on camera-lidar fusion, characterized in that: include: A data acquisition module, used to acquire the relative posture between the camera and the lidar, as well as the camera image and lidar point cloud sensor data of the scene based on the selected scene; The 3D reconstruction module is used to perform sensor tracking, map update and loop detection in sequence based on 3DGS according to the relative posture between the camera and the lidar, as well as the camera image and lidar point cloud sensor data of the scene, to obtain the 3D reconstruction result of the scene; Among them, in the sensor tracking stage, the camera and lidar postures are first initialized by the linear motion hypothesis, and then the postures are optimized by minimizing the photometric loss and geometric loss, and finally the optimized camera and lidar postures are obtained; after the sensor tracking is completed, the map update stage is entered, and the transparency map and depth map are first rendered according to the current camera posture, and the lidar point cloud that meets the preset conditions is added to the Gaussian point cloud map, and then the properties of the Gaussian point cloud are optimized according to the photometric loss and depth loss to obtain the Gaussian map; after the map update is completed, the loop detection stage is entered to optimize the Gaussian map, and the optimized Gaussian map is used as the scene 3D reconstruction result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the scene three-dimensional reconstruction method based on camera lidar fusion as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for three-dimensional reconstruction of a scene based on camera-lidar fusion as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Three-dimensional reconstruction method, device and equipment fusing panoramic camera and laser radar
CN117351140A
Radiation field model reconstruction method and device, computer equipment and storage medium
CN117593436A
Indoor complex scene high-fidelity real-time rendering method based on three-dimensional Gaussian representation
CN118096988A
Large-scale three-dimensional scene real-time reconstruction method based on Gaussian expression
CN118314280A
Scene three-dimensional reconstruction method based on prior depth and Gaussian sputtering model fusion
CN118351252A
Cited By
High-fidelity three-dimensional reconstruction method fusing attitude prior and geometric constraint
CN120931839A
New view synthesis method based on Gaussian probability distribution and feature regularization
CN120976032A