Loopback SLAM method based on three-dimensional Gaussian sputtering and multi-camera input

By introducing multi-camera input and three-dimensional Gaussian sputtering into the SLAM method, combined with loopback detection and optimization modules, the existing SLAM method solves the problem of single sensor input and lack of loopback detection, and realizes efficient and accurate three-dimensional map reconstruction and camera posture adjustment.

CN120147533APending Publication Date: 2025-06-13HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202510225916.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing SLAM method based on three-dimensional Gaussian sputtering is mainly used on single sensor input, lacking loopback detection and optimization capabilities, and cannot fully utilize the three-dimensional Gaussian representation.

Method used

Multi-camera input combined with three-dimensional Gaussian sputtering is used to construct constraints through overlapping parts between cameras to achieve more accurate camera pose estimation and high-quality scene modeling. Using the three-dimensional Gaussian loopback detection and optimization module, the camera position is adjusted after the loopback is detected, the camera drift problem is solved, and the three-dimensional model is updated.

Benefits of technology

It realizes high efficiency and high quality positioning and map reconstruction, corrects camera pose drift, and maintains the accuracy and completeness of the three-dimensional model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147533A_ABST
    Figure CN120147533A_ABST
Patent Text Reader

Abstract

The invention discloses a loopback SLAM (Simultaneous Localization and Mapping) method based on three-dimensional Gaussian sputtering and multi-camera input. According to the invention, fusion of three-dimensional Gaussian sputtering and multi-camera input is utilized to realize autonomous panoramic data acquisition and scene reconstruction with higher acquisition efficiency; according to the method, the constraint is constructed by using the overlapped part between the cameras, so that more accurate camera pose estimation and high-quality scene modeling are realized; according to the method, timestamp attributes are added to gauss, the gauss are classified according to the timestamp attributes, and the loopback is rapidly and effectively detected through different classes of gauss proportions under a current frame camera view angle; after the loopback is detected, the camera pose is adjusted, the problem of camera drifting is solved, meanwhile, a Gaussian map can be updated according to adjustment of the camera pose, and an accurate three-dimensional model is kept. According to the two-stage binding adjustment strategy provided by the invention, the global camera pose is finely adjusted by using the multi-view rendering image loss and the pose image constraint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly aims at three-dimensional scene reconstruction. Specifically, it relates to a loop SLAM method based on three-dimensional Gaussian splatting and multi-camera input. Background Art

[0002] Dense three-dimensional reconstruction is one of the hot research topics in the fields of computer vision and computer graphics. By using data such as color, depth, texture, and position collected by various sensors such as RGB-D cameras, and adopting mathematical tools such as multi-view geometry, probability statistics, and optimization theory, three-dimensional modeling of the real physical world is carried out. And the KinectFusion, a dense three-dimensional reconstruction system based on RGB-D proposed by Newcombe et al. from Imperial College London, makes high-quality and real-time dense three-dimensional reconstruction possible.

[0003] Three-dimensional Gaussian splatting redefines the method of scene representation and rendering, integrating the differentiable rendering method of NeRF (Neural Radiance Field) and the point-based rendering method. It not only has the ability of fast rendering, but also can maintain high rendering quality, making it possible for a SLAM system with real-time high-fidelity mapping. By combining Gaussian splatting with technologies such as NeRF, researchers can improve the rendering speed while ensuring the accuracy of reconstruction, and finally achieve efficient and accurate environmental perception and map construction. This opens up a new direction for future SLAM research and applications, and promotes the development of computer vision and robotics.

[0004] However, currently, the SLAM (Simultaneous Localization and Mapping) method based on three-dimensional Gaussian splatting is mainly applied to single-sensor input, and there is basically no loop or the loop optimization method provided by traditional SLAM methods, and it cannot make full use of the three-dimensional Gaussian representation. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention proposes a loop SLAM method based on three-dimensional Gaussian splatting and multi-camera input. The present invention uses the fusion of three-dimensional Gaussian splatting and multi-camera input to achieve autonomous panoramic data acquisition and scene reconstruction with higher acquisition efficiency; uses the overlapping parts between cameras to construct constraints to achieve more accurate camera pose estimation and high-quality scene modeling; uses a loop detection and optimization module based on three-dimensional Gaussian to adjust the camera pose after detecting a loop, solve the camera drift problem, and at the same time, according to the adjustment of the camera pose, the three-dimensional model can also be adjusted to maintain an accurate three-dimensional model.

[0006] The technical solutions adopted by the present invention to solve its technical problems are as follows:

[0007] Step 1: Control the multi-camera rotation acquisition device to rotate at a constant speed, collect information through multiple sensors, and input the color and depth maps of the same moment and different perspectives collected by the three calibrated cameras for each frame.

[0008] Step 2: Initialize the Gaussian map.

[0009] Step 3: Perform camera tracking.

[0010] Step 4: Perform key frame selection.

[0011] Step 5: Perform loop detection, then optimize the pose graph and adjust the camera pose.

[0012] Step 6: Update the Gaussian parameters.

[0013] Step 7: Return to Step 3 again, and continue to update and optimize according to the color map and depth map input in the next frame until all frames are completed.

[0014] Further, the specific implementation of Step 2 is as follows:

[0015] The map scene is represented by a three-dimensional Gaussian. Each Gaussian contains position information u, color information c, radius r, rotation quaternion q, opacity α, and timestamp t. The timestamp t indicates that the Gaussian is added to the map at the t-th frame. By the difference between the Gaussian timestamp and the current latest timestamp and a preset threshold, the Gaussians can be divided into two categories: historical Gaussians and new Gaussians. Calculate the final pixel color to obtain the rendered color map C(p). The specific calculation formula is:

[0016]

[0017] where C(p) represents the calculated pixel color, n represents the number of Gaussians that can splash onto this pixel, c i represents the color information carried on each three-dimensional Gaussian, and α i represents the opacity of each three-dimensional Gaussian;

[0018] Calculate the rendered depth map D(p):

[0019]

[0020] where D(p) represents the calculated pixel depth, and d i represents the depth value of each three-dimensional Gaussian in the camera coordinates;

[0021] Use the three-dimensional Gaussian to render the visibility map S(p) to indicate whether the pixel contains information from the current map. The calculation formula of the visibility map S(p) is:

[0022]

[0023] Initialize the 3D Gaussian point position information u and color information c with the input depth map D(p) and color map C(p), and optimize the 3D Gaussian parameters through the loss function. The loss function L m is as follows:

[0024] L m = ∑ p (0.8L 1 (C(p)) + 0.2L D-SSIM + 2L 1 (D(p))) (4)

[0025] where, L 1 the loss represents the absolute difference between the rendered image and the real input image, and L D-SSIM represents the structural similarity loss;

[0026] Use the losses L m from three camera views to jointly optimize the 3D Gaussian parameters. The final Gaussian initialization loss function L map is:

[0027] L map = λL m-up + L m-mid + λL m-down (5)

[0028] where, L m-mid , L m-down , L m-up represent the losses L m from the middle camera, lower camera, and upper camera views respectively, and λ is a weight coefficient between 0 and 1.

[0029] Furthermore, the specific implementation of step 3 is as follows:

[0030] Set an initial value for the camera pose of the current frame, then fix the Gaussian parameters unchanged, and solve the inter-frame camera motion by minimizing the loss L 1 of the depth map D(p) and color map C(p). The specific loss function L t is expressed as:

[0031] L t = ∑ p (S(p) > 0.99)(L 1 (D(p)) + 0.5L 1 (C(p))) (6)

[0032] During the tracking process, not only use the middle camera to render the color map and depth map, but also use the upper camera and lower camera to render the color map and depth map to calculate the loss L 1, jointly optimize the camera pose T of the current frame cw , the final loss function is defined as:

[0033] L track = L t-up + L t-mid + L t-down (7)

[0034] Among them, L t-mid , L t-down , L t-up respectively represent the L t loss functions from the perspectives of the middle camera, the lower camera, and the upper camera; during the tracking process, when using formulas (1), (2), and (3) to calculate the rendered color map, the rendered depth map, and the visibility map S(p), only the new Gaussian is used, and the historical Gaussians with a time difference greater than the preset threshold from the current latest timestamp do not participate in the calculation.

[0035] Furthermore, step 4 is specifically implemented as follows:

[0036] Save every 5 frames as a key frame. For each key frame, store the color map and depth map input by the current three cameras; at the same time, take the set of intermediate frames (N + 2) of adjacent key frames (N, N + 5) as the random list frame set rand-list, and save the color map and depth map captured by the middle camera of these frames;

[0037] After tracking, only use the new Gaussian to render the visibility map S(p) from the perspectives of the three cameras of the current frame, and initialize a new Gaussian using the depth values and color values of the pixel regions in the visibility map S(p) that are less than the set threshold.

[0038] Furthermore, step 5 is specifically implemented as follows:

[0039] Classify the Gaussians into historical Gaussians and new Gaussians according to the timestamp attributes of the Gaussians and the current timestamp. Obtain all the Gaussians from the perspective of the current frame through the reprojection method, and calculate the ratio of historical Gaussians to new Gaussians from the perspective of the current frame. When the proportion of historical Gaussians is greater than the set loop closure threshold, it indicates that a loop closure is detected, and then the following optimizations are performed on the current frame:

[0040] 5-1. Pose graph optimization:

[0041] For the pose graph corresponding to the detected loop closure frame, each vertex in the graph is based on the camera pose T cw estimated in step 3, and there are two types of edges ε to connect these vertices;

[0042] The first type of edge is the relative transformation of the camera poses T cw of adjacent frames, that is, the inverse multiplication of the camera pose of the previous vertex of this vertex and the camera pose of this vertex;

[0043] The second type of edge is the relative transformation at the loop endpoints Re-optimize the camera pose of the current frame according to Step 3 Different from Step 3, only the historical Gaussians whose difference between the timestamp and the current latest timestamp is greater than a preset threshold participate in the calculation of rendering the color map, the depth map, and the visibility map S(p) in formulas (1), (2), and (3). The second type of edge From the camera pose of the loop endpoint F k and the camera pose of the current frame F r and inverse multiply to obtain, that is where the loop endpoint F k is obtained by analyzing the timestamps of the historical Gaussians under the perspective of the current frame F r and getting the timestamp t with the highest count. Finally, the pose graph is defined as: k

[0044]

[0045] where represents the Euclidean distance between two translation vectors t a , t b , and L SO(3) (W a , W b ) represents the distance metric between two directions in SO(3). t ij represents the translation part of the relative transformation from vertex i to vertex j, and W ij represents the rotation part of the relative transformation from vertex i to vertex j. represents the translation part of the camera pose of vertex i, represents the translation part of the camera pose of vertex j, represents the rotation part of the camera pose of vertex i, represents the rotation part of the camera pose of vertex j;

[0046] 5-2. Gaussian and Camera Pose Adjustment:

[0047] For the 3D Gaussian at timestamp t, associate it with the camera pose of the t-th frame. Thanks to this association, according to the transformation of the camera pose, adjust the position u of the 3D Gaussian at the corresponding timestamp:

[0048]

[0049] where represents the camera pose before pose graph optimization, and T opt ​Represents the camera pose after optimizing the pose graph, u * Represents the adjusted 3D Gaussian position;

[0050] 5 - 3. Fine-tune the camera poses of all frames:

[0051] First, mix the frames in the key frame set and the frames in the random list frame set rand - list, and then divide the mixed frames into K sets, each set B K contains N frames adjacent in timestamp; for each set B K , use the middle camera view of set B K to calculate the loss of each frame through formula (6), sum the N losses in set B K to jointly optimize the camera poses of these n frames, and finally achieve fine - tuning of the camera poses of the key frame set and the random list frame set rand - list. The specific multi - view rendering image loss function L local is:

[0052]

[0053] After optimizing the camera poses of the key frame set to be optimized and the random list frame set, fix these camera poses, and perform pose graph optimization according to the pose graph established in formula (8) to adjust the camera poses of all non - key frames and ordinary frames in the non - random list frame set.

[0054] Furthermore, the present invention performs misjudgment detection on the detected loopback frames, specifically as follows:

[0055] Perform color map rendering on the current frame using the new Gaussian and the historical Gaussian respectively, and then calculate the structural similarity index between the two rendered color maps. If the structural similarity index is lower than a preset threshold, it indicates that there may be a misjudgment situation.

[0056] Furthermore, step 6 is specifically as follows:

[0057] Update the Gaussian parameters through the loss function L of formula (4) m , randomly select a key frame from the current frame and the key frame set with overlapping views with the current frame for each iteration update, and use formula (5) to calculate the loss L 3 ; in addition, randomly select a key frame from all key frame sets to calculate the loss L under the middle camera view using formula (4) 4 ; at the same time, select a frame from the random list frame set rand - list without overlap with the current frame to calculate the loss L under the middle camera view using formula (4) 5 ; finally, add L 3 , L 4 and L 5Add them up as the final loss function, fix the camera pose and update the Gaussian parameters;

[0058] Furthermore, in the Gaussian parameter update process of the present invention, when calculating the rendered color image, the rendered depth image and the visibility map S(p) using formulas (1), (2), and (3), only the new Gaussian is used, and the historical Gaussians with the difference between the timestamp and the current latest timestamp greater than a preset threshold do not participate in the calculation.

[0059] The beneficial effects of the present invention are as follows:

[0060] The present invention combines the 3D Gaussian-based SLAM method and the multi-camera rotation acquisition device to achieve high-efficiency and high-quality positioning and map reconstruction. The present invention adds a timestamp attribute to the Gaussian, classifies the Gaussians according to the timestamp attribute, and quickly and effectively detects loops by the ratio of different categories of Gaussians under the current frame camera view. The present invention proposes a loop closure optimization strategy based on 3D Gaussian representation to correct the camera pose drift and update the Gaussian map, maintaining the integrity of the model. The present invention proposes a two-stage Bundle Adjustment strategy, using the constraints of the multi-view rendered image loss and the pose graph to fine-tune the global camera pose.

[0061] The present invention uses a multi-camera rotation acquisition device to drive three synchronized RGB-D cameras to rotate for data acquisition. The visual data input for each frame is the color and depth maps at the same time and different perspectives collected by the three calibrated cameras. Subsequently, camera pose tracking, addition of new Gaussians, and 3D reconstruction of the surrounding scene are performed. Among them, the optimization and reconstruction of the surrounding area are achieved by optimizing the parameters of the 3D Gaussian. During the camera tracking and scene reconstruction process, we make full use of the three camera inputs and design an effective loss function to optimize the camera pose and 3D Gaussian parameters.

[0062] During the reconstruction process, the present invention assigns timestamps to each Gaussian and divides them into historical Gaussians and new Gaussians according to their timestamps. By detecting the ratio of the two types of Gaussians under the current frame view, it is judged whether a loop is detected. To avoid misjudgment, we use the new Gaussian and the historical Gaussian to perform color image rendering on the current frame respectively, and calculate the structural similarity index (SSIM) between the two. If the index is lower than a preset threshold, it indicates that there may be a misjudgment situation.

[0063] During the reconstruction process, the present invention assigns timestamps to each Gaussian and divides them into historical Gaussians and new Gaussians according to their timestamps. By detecting the ratio of the two types of Gaussians under the current frame view, it is judged whether a loop is detected. To avoid misjudgment, we use the new Gaussian and the historical Gaussian to perform color image rendering on the current frame respectively, and calculate the structural similarity index (SSIM) between the two. If the index is lower than a preset threshold, it indicates that there may be a misjudgment situation.

[0064] After detecting the loop, the present invention uses a pose graph to adjust the camera poses of all frames, then adjusts the positions of the three-dimensional Gaussians according to the transformation of the camera pose adjustment, and finally uses local BA (bundle adjustment) and the pose graph to finely adjust the global camera pose. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a flow chart of the present invention.

[0066] Figure 2 It is a comparison chart of the results of the final rendered color map on the virtual data set of the present invention.

[0067] Figure 3 It is a comparison chart of the results of the final rendered color map on the real data set of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0068] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0069] As Figure 1 shown, a loop SLAM method based on three-dimensional Gaussian sputtering and multi-camera input includes the following steps:

[0070] Step 1: Control the multi-camera rotation acquisition device to rotate at a constant speed, collect information through multiple sensors, and input the color and depth maps of the same moment and different perspectives collected by the three calibrated cameras for each frame.

[0071] Step 2: Initialize the Gaussian map.

[0072] The present invention represents the map scene by three-dimensional Gaussians. Each Gaussian contains position information u, color information c, radius r, rotation quaternion q, opacity α, and timestamp t. The timestamp t indicates that the Gaussian is added to the map at the t-th frame. By the difference between the Gaussian timestamp and the current latest timestamp and a preset threshold, the Gaussians can be divided into two categories: historical Gaussians and new Gaussians. By training the three-dimensional Gaussian parameters, the three-dimensional scene can be realistically represented. Given the camera pose, the three-dimensional Gaussians can be projected onto the two-dimensional imaging plane, and the color information of the multiple ordered Gaussians mapped to the image pixels is weighted and superimposed to calculate the final pixel color to obtain the rendered color map. The specific calculation formula is:

[0073]

[0074] where C(p) represents the calculated pixel color, n represents the number of Gaussians that can sputter to this pixel, ci represents the color information carried by each three-dimensional Gaussian, and α i represents the opacity of each three-dimensional Gaussian.

[0075] Similarly, a depth map is rendered using a similar method:

[0076]

[0077] where D(p) represents the calculated pixel depth, and d i represents the depth value of each three-dimensional Gaussian in the camera coordinates. In addition, the present invention also uses three-dimensional Gaussian to render the visibility map S(p) to indicate whether the pixel contains information from the current map, and the calculation formula is:

[0078]

[0079] The present invention initializes the three-dimensional Gaussian point position information u and color information c through the input depth map and color map, and optimizes the three-dimensional Gaussian parameters through the loss function. The loss function L m is as follows:

[0080] L m = ∑ p (0.8L 1 (C(p)) + 0.2L D-SSIM + 2L 1 (D(p))) (4)

[0081] where, L 1 the loss represents the absolute difference between the rendered image and the real input image. L D-SSIM represents the structural similarity loss (D-SSIM), which is a variant of SSIM (Structural Similarity Index Measure) and is used to measure the similarity of color images in terms of visual quality. Due to the input of three cameras, the present invention needs to optimize more three-dimensional Gaussian parameters. Therefore, the losses L m from three camera views are used to jointly optimize the three-dimensional Gaussian parameters. The final Gaussian initialization loss function L map is:

[0082] L map = λL M-up + L m-mid + λL m-down (5)

[0083] where, L m-mid , L m-down , L m-up respectively represent the losses L m under the middle camera, lower camera, and upper camera views, and λ is a weight coefficient between 0 and 1.

[0084] Step 3: Camera tracking.

[0085] The present invention uses a constant speed assumption to set an initial value for the camera pose of the current frame, and then keeps the Gaussian parameters unchanged. By minimizing the loss L between the depth map and the color map 1 to solve for the inter-frame camera motion, with the specific loss function L t expressed as:

[0086] L t = ∑ p (S(p)>0.99)(L 1 (D(p)) + 0.5L 1 (C(p))) (6)

[0087] Since the upper camera, the lower camera and the middle camera overlap, an intra-frame constraint can be formed, which makes the optimization of the camera pose more accurate and stable. Therefore, in the tracking process of the present invention, not only the middle camera is used to render the color map and the depth map, but also the upper camera and the lower camera are used to render the color map and the depth map to calculate the loss L 1 , jointly optimizing the camera pose Tcw of the current frame. The final loss function is defined as:

[0088] L track = L t-up + L t-mid + L t-down (7)

[0089] where L t-mid , L t-down , L t-up respectively represent L t loss functions from the perspectives of the middle camera, the lower camera and the upper camera. It should be noted that during the tracking process, when calculating the rendered color map, the rendered depth map and the visibility map S(p) using formulas (1), (2) and (3), only the new Gaussian is used, and the historical Gaussians with the time difference from the current latest timestamp greater than a preset threshold do not participate in the calculation.

[0090] Step 4, Key frame selection.

[0091] To improve the tracking and rendering efficiency of the system, the present invention saves every 5 frames as a key frame. For each key frame, the color image and the depth image input by the current three cameras are stored. In addition to the saved key frames, the present invention also takes the set of intermediate frames (N + 2) of adjacent key frames (N, N + 5) as the random list frame set rand-list, and saves the color image and the depth image captured by the middle camera of these frames.

[0092] After tracking, the present invention only uses the new Gaussian to render the visibility map S(p) from the perspectives of the three cameras of the current frame, and initializes a new Gaussian using the depth values and color values of the pixel regions in the visibility map S(p) that are less than the set threshold.

[0093] Step 5, Loop Detection.

[0094] After step 3, the present invention classifies Gaussians into historical Gaussians and new Gaussians according to the timestamp attribute of Gaussians and the current timestamp, obtains all Gaussians in the current frame view through the reprojection method, calculates the ratio of historical Gaussians to new Gaussians in the current frame view, and determines whether a loop is detected based on the ratio. When the proportion of historical Gaussians is greater than the set loop threshold, it indicates that a loop is detected, and the following optimizations are made:

[0095] Pose graph optimization: For the detected loop frame, the present invention implements a lightweight pose graph to eliminate camera pose drift and obtain an accurate trajectory. In the pose graph, each vertex is based on the camera pose T estimated in step 3 cw constructed, and there are two types of edges ε to connect these vertices; the first is the relative transformation of adjacent frame camera poses T cw , that is, the inverse multiplication of the camera pose of the previous vertex of this vertex and the camera pose of this vertex; the second is the relative transformation at the loop end point Re-optimize the camera pose of a current frame according to step 3 Different from step 3, only historical Gaussians whose difference between the timestamp and the current latest timestamp is greater than a preset threshold participate in the processes of calculating and rendering the color map, depth map, and visibility map S(p) using formulas (1), (2), and (3). The second edge is obtained by the inverse multiplication of the camera pose of the loop end point F k and the camera pose of the current frame F r , that is where the loop end point F k is determined by the timestamp tk with the highest count obtained by analyzing the timestamps of historical Gaussians in the view of the current frame F r , is the camera pose at the loop end point F k . Finally, the pose graph can be defined as:

[0096]

[0097] where represents the Euclidean distance between two translation vectors t a , t b , and L SO(3) (W a , W b ) represents the distance metric between two directions in SO(3). t ij represents the translation part of the relative transformation from vertex i to vertex j, and W ij represents the rotation part of the relative transformation from vertex i to vertex j, represents the translation part of the camera pose of vertex i, represents the translational part of the camera pose of vertex j represents the rotational part of the camera pose of vertex i represents the rotational part of the camera pose of vertex j

[0098] In this way, pose graph optimization enables us to adjust the camera poses of all frames

[0099] Gaussian and camera pose adjustment: After completing the optimization of the pose graph, to correct possible map errors, the 3D Gaussian and camera poses are associated according to their respective timestamp attributes. Specifically, for the 3D Gaussian at timestamp t, it is associated with the camera pose of the t-th frame. Thanks to this association, according to the transformation of the camera pose, the position u of the 3D Gaussian at the corresponding timestamp is adjusted as follows:

[0100]

[0101] where represents the camera pose before pose graph optimization, T opt represents the camera pose after pose graph optimization, u * represents the adjusted position of the 3D Gaussian

[0102] Finally, fine-tune the camera poses of all frames to obtain more accurate and consistent camera poses aligned with the 3D Gaussian map. Specifically, first mix the frames in the key frame set and the frames in the random list frame set rand-list, and then divide the mixed frames into multiple sets, each set B K contains N frames adjacent in timestamp. For each set B K , use the middle camera view of set B K to calculate the loss of each frame through formula (6), add up the N losses in set B K and jointly optimize the camera poses of these N frames. Finally, fine-tune the camera poses of the key frame set and the random list frame set rand-list. Specifically, the multi-view rendering image loss function L local is as follows:

[0103]

[0104] where, after optimizing the camera poses of the key frame set and the random list frame set, we fix these camera poses and perform pose graph optimization according to the pose graph established in formula (8) to adjust the camera poses of all non-key frames and ordinary frames in the non-random list frame set

[0105] The present invention proposes a two-stage Bundle Adjustment strategy, that is, a strategy combining the fine-tuning stage of the camera poses of the key frame set and the random list frame set with the fine-tuning stage of the camera poses of the ordinary frames, and uses the multi-view rendering image loss L local and the constraints of the pose graph to fine-tune the global camera pose.

[0106] To avoid misjudgment, the present invention also performs color image rendering on the current frame using the new Gaussian and the historical Gaussian respectively, and then calculates the structural similarity index (SSIM) between the two rendered color images. If the structural similarity index is lower than a preset threshold, it indicates that there may be a misjudgment situation.

[0107] Step 6: Gaussian parameter update;

[0108] Similar to the initialization of the Gaussian map in Step 2, the Gaussian parameter update is also optimized through the loss function L in formula (4) m However, to prevent global forgetting, a key frame is randomly selected from the current frame and the key frame set with overlapping perspectives with the current frame for each iteration of optimization to calculate the loss L using formula (5) 3 . In addition, a key frame is randomly selected from all key frame sets to calculate the loss L at the intermediate camera perspective using formula (4) 4 . At the same time, a frame is selected from the random list frame set rand-list that has no overlap with the current frame to calculate the loss L at the intermediate camera perspective using formula (4) 5 , and finally L 3 , L 4 and L 5 are added as the final loss function, and the camera pose is fixed and the Gaussian parameters are optimized.

[0109] Similar to the tracking process, when calculating the rendered color map, the rendered depth map, and the visibility map S(p) using formulas (1), (2), and (3) in the Gaussian parameter update process, only the new Gaussian is used, and the historical Gaussian whose time difference from the current latest timestamp is greater than a preset threshold does not participate in the calculation.

[0110] Step 7: Return to Step 3 again, and continue to update and optimize according to the color map and depth map input for the next frame until all frames are completed.

[0111] Experimental data:

[0112] Table 1:

[0113]

[0114] As shown in Table 1, the comparison of the absolute trajectory error ATE (cm)↓ results of the present invention and other 3D Gaussian-based SLAM methods MonoGS, Gaussian-SLAM, and SplaTAM on the virtual dataset is presented. It can be clearly seen from Table 1 that the present invention has obvious advantages in various indicators.

[0115] Table 2:

[0116]

[0117] As shown in Table 2, the comparison of the peak signal-to-noise ratio PSNR, structural similarity SSIM, learned perceptual image patch similarity LPIPS, and L1 loss results of the rendered depth map and the real depth map of the present invention and other 3D Gaussian-based SLAM methods MonoGS, Gaussian-SLAM, and SplaTAM on the virtual dataset is presented. It can be clearly seen from Table 2 that the present invention has obvious advantages in various indicators.

[0118] As Figure 2 shown, the comparison of the final rendered color map results of the present invention and other 3D Gaussian-based SLAM methods MonoGS, Gaussian-SLAM, and SplaTAM on the virtual dataset is presented.

[0119] Table 3:

[0120]

[0121] The comparison of the peak signal-to-noise ratio PSNR, structural similarity SSIM, learned perceptual image patch similarity LPIPS, and L1 loss results of the rendered depth map and the real depth map of the present invention and other 3D Gaussian-based SLAM methods MonoGS, Gaussian-SLAM, and SplaTAM on the real dataset is presented. It can be clearly seen from Table 3 that the present invention has obvious advantages in various indicators.

[0122] As Figure 3 shown, the comparison of the final rendered color map results of the present invention and other 3D Gaussian-based SLAM methods MonoGS, Gaussian-SLAM, and SplaTAM on the real dataset is presented.

[0123] In summary, the present invention has obvious advantages in various indicators compared with the existing methods, whether in the virtual dataset or the real dataset.

Claims

1. A loop closure SLAM method based on three-dimensional Gaussian sputtering and multi-camera input, characterized in that: The steps include: Step 1: Control the multi-camera rotating acquisition device to rotate at a constant speed, collect information through multiple sensors, and input each frame of the color and depth map collected by the three calibrated cameras at the same time and from different perspectives; Step 2: Initialize the Gaussian map; Step 3: Perform camera tracking; Step 4: Select key frames; Step 5: Perform loop detection, optimize the pose graph, and adjust the camera pose; Step 6: Update Gaussian parameters; Step 7: Return to step 3 and continue updating and optimizing according to the color image and depth image input by the next frame until all frames are completed.

2. The loop-closure SLAM method based on three-dimensional Gaussian sputtering and multi-camera input according to claim 1, characterized in that: Step 2 is implemented as follows: The map scene is represented by a three-dimensional Gaussian. Each Gaussian contains location information u, color information c, radius r, rotation quaternion q, opacity α, and timestamp t. Timestamp t indicates that the Gaussian is added to the map in the tth frame. The difference between the Gaussian timestamp and the current latest timestamp and the preset threshold can be used to divide the Gaussian into two types: historical Gaussian and new Gaussian. The final pixel color is calculated to obtain the rendered color map C(p). The specific calculation formula is: Where C(p) represents the calculated pixel color, n represents the number of Gaussians that can splash onto the pixel, and c i Represents the color information carried by each three-dimensional Gaussian, α i Indicates the opacity of each 3D Gaussian; Calculate the rendered depth map D(p): Where D(p) represents the calculated pixel depth, d i Represents the depth value of each 3D Gaussian in camera coordinates; The visibility map S(p) is rendered using a three-dimensional Gaussian to indicate whether a pixel contains information from the current map. The calculation formula for the visibility map S(p) is: The three-dimensional Gaussian point position information u and color information c are initialized by inputting the depth map D(p) and color map C(p), and the three-dimensional Gaussian parameters are optimized by the loss function. The loss function L m as follows: L m =∑ p (0.8L1(C(p))+0.2L D-SSIM +2L1(D(p))) (4) Among them, L1 loss represents the absolute difference between the rendered image and the real input image, L D-SSIM represents the structural similarity loss; The loss L using three camera views m To jointly optimize the three-dimensional Gaussian parameters, the final Gaussian initialization loss function L map for: L map =λL m-up +L m-mid +λL m-down (5) Among them, L m-mid , L m-down , L m-up Respectively represent the loss L under the perspective of the middle camera, the lower camera, and the upper camera m ,λ is a weight coefficient between 0 and 1.

3. The loop-closure SLAM method based on three-dimensional Gaussian sputtering and multi-camera input according to claim 2, characterized in that: Step 3 is implemented as follows: Set an initial value for the camera pose of the current frame, then fix the Gaussian parameters unchanged, and solve the inter-frame camera motion by minimizing the loss L1 of the depth map D(p) and the color map C(p). The specific loss function L t It is expressed as: L t =∑ p (S(p)>0.99)(L1(D(p))+0.5L1(C(p))) (6) During the tracking process, not only the middle camera is used to render the color image and depth image, but also the upper camera and the lower camera are used to render the color image and depth image to calculate the loss L1, and jointly optimize the camera pose T of the current frame. cw , the final loss function is defined as: L track =L t-up +L t-mid +L t-down (7) Among them, L t-mid , L t-down , L t-up Respectively represent L from the perspective of the middle camera, the lower camera, and the upper camera t Loss function; During the tracking process, when using formulas (1), (2), and (3) to calculate the rendered color image, rendered depth image, and visibility image S(p), only the new Gaussian is used, and the historical Gaussian whose timestamp difference with the current latest timestamp is greater than a preset threshold is not involved in the calculation.

4. The loop-closure SLAM method based on three-dimensional Gaussian sputtering and multi-camera input according to claim 3, characterized in that: Step 4 is implemented as follows: Save every 5 frames as a key frame. For each key frame, store the color images and depth images input by the three current cameras. At the same time, take the set of intermediate frames (N+2) of adjacent key frames (N, N+5) as the random list frame set rand-list, and save the color images and depth images captured by the intermediate cameras of these frames. After tracking, only the new Gaussian is used to render the visibility map S(p) from the three-camera perspective of the current frame, and the new Gaussian is initialized using the depth values ​​and color values ​​of the pixel areas in the visibility map S(p) that are less than the set threshold.

5. The loop-closure SLAM method based on three-dimensional Gaussian sputtering and multi-camera input according to claim 4, characterized in that: Step 5 is implemented as follows: According to the timestamp attribute of Gaussian and the current timestamp, Gaussian is divided into historical Gaussian and new Gaussian. All Gaussian in the current frame perspective are obtained through the reprojection method, and the ratio of historical Gaussian to new Gaussian in the current frame perspective is calculated. When the proportion of historical Gaussian is greater than the set loop threshold, it means that a loop is detected, and the current frame is optimized as follows: 5-1. Pose graph optimization: For the pose graph corresponding to the detected loop frame, each vertex in the graph is the camera pose T estimated based on step 3 cw constructed, and there are two types of edges ε to connect these vertices; The first type of edge is the camera pose T of adjacent frames. cw The relative transformation is the multiplication of the camera pose of the previous vertex of the vertex and the inverse of the camera pose of the vertex; The second type of edge is the relative transformation at the loop endpoint Re-optimize the current frame camera pose according to step 3 The difference from step 3 is that only the historical Gaussian whose timestamp difference with the current latest timestamp is greater than the preset threshold participates in the process of calculating the rendered color image, rendered depth image and visibility image S(p) in formulas (1), (2) and (3). By loopback endpoint F k The camera pose With the current frame F r The camera pose Inverse multiplication gives The loop endpoint F k is obtained by analyzing the current frame F r The timestamp with the highest count is obtained from the timestamp of the historical Gaussian in the perspective of k Finally, the pose graph is defined as: in, Represents two translation vectors t a ,t b The Euclidean distance between SO(3) (W a ,W b ) represents the distance measure between two directions in SO(3), t ij represents the translation part of the relative transformation from vertex i to vertex j, W ij Represents the rotation part of the relative transformation from vertex i to vertex j, represents the translation part of the camera pose of vertex i, represents the translation part of the camera pose of vertex j, represents the rotation part of the camera pose of vertex i, represents the rotation part of the camera pose of vertex j; 5-2. Gaussian and camera pose adjustment: For the 3D Gaussian at timestamp t, it is associated with the camera pose of the tth frame. Thanks to this association, the 3D Gaussian position u at the corresponding timestamp is adjusted according to the transformation of the camera pose: in, represents the camera pose before pose graph optimization, T opt represents the camera pose after pose graph optimization, u * represents the adjusted three-dimensional Gaussian position; 5-3. Fine-tune the camera pose of all frames: First, the frames in the key frame set are mixed with the frames in the random list frame set rand-list, and then the mixed frames are divided into K sets, each set B K Contains N frames that are adjacent in timestamp; for each set B K , using set B K The intermediate camera view of the set B is calculated by formula (6) to calculate the loss of each frame. K The N losses in the sum are jointly optimized to optimize the camera poses of these N frames, and finally the camera poses of the key frame set and the random list frame set rand-list are fine-tuned. The specific multi-view rendering image loss function L local for: After the camera poses of the key frame set and the random list frame set are optimized, these camera poses are fixed, and the pose graph optimization is performed according to the pose graph established in formula (8), and the camera poses of all non-key frames and ordinary frames in the non-random list frame set are adjusted.

6. The loop-closure SLAM method based on three-dimensional Gaussian sputtering and multi-camera input according to claim 4, characterized in that: Perform misjudgment detection on the detected loop frame, as follows: The new Gaussian and the historical Gaussian are used to render the color image on the current frame respectively, and then the structural similarity index between the two rendered color images is calculated. If the structural similarity index is lower than the preset threshold, it indicates that there may be a misjudgment.

7. The loop-closure SLAM method based on three-dimensional Gaussian sputtering and multi-camera input according to claim 5, characterized in that: Step 6 is as follows: The loss function L through formula (4) m Perform Gaussian parameter update. In each iteration, randomly select a key frame from the current frame and the key frame set that overlaps with the current frame’s perspective, and use formula (5) to calculate the loss L3. In addition, randomly select a key frame from all key frame sets and use formula (4) to calculate the loss L4 under the intermediate camera perspective. At the same time, select a frame from the random list frame set rand-list that does not overlap with the current frame and use formula (4) to calculate the loss L5 under the intermediate camera perspective. Finally, add L3, L4 and L5 as the final loss function, fix the camera pose and update the Gaussian parameters.

8. The loop-closure SLAM method based on three-dimensional Gaussian sputtering and multi-camera input according to claim 7, characterized in that: When the Gaussian parameter update process uses formulas (1), (2), and (3) to calculate the rendered color image, rendered depth image, and visibility image S(p), only the new Gaussian is used, and the historical Gaussian whose timestamp difference with the current latest timestamp is greater than a preset threshold is not involved in the calculation.

Citation Information

Patent Citations

  • Laser enhanced vision three-dimensional reconstruction method and system based on Gaussian splashing

    CN119180908A

  • Multi-view automatic driving scene reconstruction method based on three-dimensional Gaussian splashing

    CN119229002A

  • Simultaneous positioning and dense three-dimensional reconstruction method

    WO2018129715A1

Cited By

  • Self-supervision three-dimensional scene mapping method based on multi-view pose self-optimization

    CN121236129A

  • A self-supervised three-dimensional scene mapping method based on multi-view pose self-optimization

    CN121236129B

  • Camera pose regression estimation system and method based on long-time-sequence arbitrary point tracking

    CN121259094A

  • Semantic Gaussian sputtering dynamic RGB-D SLAM method and system with loopback closed optimization

    CN122415741A