NeRF visual SLAM mapping method based on ORB tracking and three-plane hash coding

By adopting ORB tracking and three-plane hash coding technology in NeRF visual SLAM, the problem of high computing burden and difficult to meet real-time abilities of the visual SLAM method based on NeRF is solved, efficient and robust visual SLAM mapping is achieved, and high-quality NeRF maps are constructed.

CN120219622APending Publication Date: 2025-06-27HANGZHOU DIANZI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510288693.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The NeRF-based visual SLAM method has a heavy computing burden during ray sampling and rendering, which makes it difficult to meet the real-time requirements of operation speed. Moreover, the training and inference process of neural radiation fields have high demands on computing resources and memory, which limits its deployment in practical applications.

Method used

The NeRF visual SLAM mapping method based on ORB tracking and three-plane hash encoding is adopted to accurately track ORB feature points, and the scene is efficiently mapped using compact memory three-plane hash encoding technology to optimize the local and global features of the NeRF map.

Benefits of technology

While ensuring system robustness and real-timeness, it can build high-quality NeRF maps, providing stable and efficient visual SLAM solutions, reducing the computational cost of NeRF map optimization algorithms, and improving map construction accuracy and operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219622A_ABST
    Figure CN120219622A_ABST
Patent Text Reader

Abstract

The invention provides a NeRF visual SLAM mapping method based on ORB tracking and three-plane hash coding. Firstly, ORB feature point extraction and feature matching are carried out, key frames are obtained through screening, and an initial sparse map is constructed: local common-view key frames are screened and optimized through a local beam adjustment method, and light sampling is carried out on the key frames in a local common-view key frame window; performing local optimization on scene features corresponding to the key frames in the NeRF map by using a NeRF optimization method based on three-plane hash coding; after loopback is detected, global light beam adjustment method optimization is carried out on the global key frame, and a NeRF optimization method based on three-plane Hash coding is used for carrying out global optimization on scene features of a NeRF map so as to guarantee the global consistency of the constructed map; and finally, extracting surface information of the scene from the NeRF map, and generating a three-dimensional reconstruction model. The combination of three-plane Hash coding and a multi-layer perceptron is used as NeRF map scene representation, and scene details can be reconstructed efficiently and delicately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of Simultaneous Localization And Mapping (SLAM) and three-dimensional reconstruction, and specifically refers to a NeRF visual SLAM mapping method based on ORB tracking and three-plane hashing encoding. Background Art

[0002] Simultaneous Localization And Mapping (SLAM) is a technology that estimates the state of a robot by means of sensors equipped on the robot and constructs an environmental map model. Its principle is to use sensors such as cameras, lidar, and inertial measurement units to collect environmental information, fuse the collected information, and construct an environmental map while estimating the state of the robot in an unknown environment. Currently, SLAM is widely used in many fields such as three-dimensional reconstruction, unmanned aerial vehicles, and autonomous driving.

[0003] Neural Radiance Fields (NeRF), as an emerging computer vision technology based on deep learning, its principle is to use an implicit neural network to represent the radiation field of three-dimensional space, take images and the camera poses of the images as inputs, output the color and density of sampled points in space, and optimize the neural network to make the rendered result as consistent as possible with the input image, so as to achieve three-dimensional reconstruction of the scene and high-fidelity rendering of views. Currently, it is widely used in many fields such as novel view synthesis and three-dimensional reconstruction.

[0004] Currently, traditional visual SLAM methods can achieve real-time and robust operation in indoor scenarios by tracking the camera pose and constructing maps using methods such as feature points, point clouds, grids, or voxels. However, this map representation has obvious limitations: its expressive ability is relatively limited, making it difficult to capture the continuous geometric details and rich texture information of the scene, and thus unable to meet the requirements of high-precision scene modeling in upper-layer applications such as autonomous driving and virtual reality. In addition, the maps of traditional visual SLAM are usually stored in a discretized form, lacking a global consistent understanding of the scene, which limits its effectiveness in complex application scenarios. In contrast, NeRF-based visual SLAM methods such as NICE-SLAM and ESLAM construct a hybrid representation map by combining implicit neural networks with explicit structures such as hierarchical grids and tri-planes. This hybrid representation map can not only efficiently model the geometric structure and appearance information of the scene but also support high-quality 3D rendering, providing a richer and more continuous map expression ability for the SLAM system and enabling more upper-layer applications. Although the positioning accuracy of NeRF-based visual SLAM methods has reached the level of traditional SLAM on some public datasets, their tracking accuracy and mapping quality depend on fixed and high-frequency light sampling and rendering, and they cannot autonomously select image frames for light sampling and rendering based on the observed visual information, resulting in a heavy computational burden on the system and a running speed that is difficult to meet the real-time requirements. In addition, the training and inference processes of neural radiance fields have high requirements for computational resources and memory, further limiting their deployment in practical applications. Therefore, how to ensure the robustness and real-time performance of the system while constructing a high-quality NeRF map has become the core issue in the current research on NeRF-based visual SLAM. Summary of the Invention

[0005] The present invention aims to solve the problems existing in NeRF-based visual SLAM and proposes a NeRF visual SLAM mapping method based on ORB tracking and tri-plane hash coding. This method accurately tracks through ORB feature points and uses the memory-compact tri-plane hash coding technology to efficiently and finely map the scene. Compared with other NeRF-based visual SLAM methods, this method can not only ensure the robustness and real-time performance of the system but also construct a high-quality NeRF map, providing a stable and efficient visual SLAM solution in complex environments. To achieve the above object, the present invention provides the following technical solutions:

[0006] A NeRF visual SLAM mapping method based on ORB tracking and tri-plane hash coding, comprising the following steps:

[0007] Step 1, obtain the RGB image and depth image of the scene through an RGB-D camera;

[0008] Step 2: Use ORB-SLAM3 as the ORB tracking module to extract ORB feature points and perform feature matching on the RGB image, estimate the frame pose, screen for key frames, and convert the ORB feature points into 3D map points based on the depth information obtained from the depth image to construct an initial sparse map:

[0009] Step 3: According to the number of 3D map points in the sparse map that are commonly visible to the current key frame and other key frames, screen for locally co-visible key frames, and perform local bundle adjustment (Local BA) optimization on the locally co-visible key frames to optimize the poses of the locally co-visible key frames and the positions of the 3D map points; Define the number of 3D map points in the sparse map that are commonly visible to the current key frame and other key frames as the co-visibility;

[0010] Step 4: Sort the locally co-visible key frames according to their co-visibility with the current key frame, intercept a number of key frames with the highest co-visibility with the current key frame as the locally co-visible key frame window, and perform ray sampling on the key frames in the locally co-visible key frame window.

[0011] Use the NeRF optimization method based on tri-plane hashing encoding to locally optimize the scene features corresponding to these key frames in the NeRF map, where the tri-plane hashing encoding is used to represent the scene geometric features and appearance features;

[0012] Step 5: Use the loop closure module based on ORB-SLAM3. After detecting a loop, perform global bundle adjustment optimization on the global key frame poses; Then, use the global key frames to perform iterative optimization of the map features. During the iteration process, use the NeRF optimization method based on tri-plane hashing encoding to globally optimize the scene features of the NeRF map to ensure the global consistency of the constructed map;

[0013] Step 6: Use the Marching Cubes algorithm to extract the surface information of the scene from the NeRF map and generate a three-dimensional reconstruction model.

[0014] Preferably, step 4 specifically includes the following steps:

[0015] S4-1: Sort according to the co-visibility with the current key frame, intercept the first 20 key frames with the highest co-visibility with the current key frame as the locally co-visible key frame window, convert the poses of the key frames in the locally co-visible key frame window to the NeRF map coordinate system, and record the key frame poses;

[0016] S4-2: Perform ray sampling on the key frames in the locally co-visible key frame window. In the first round, sample N rays for the locally co-visible key frames raySampling rays, and in the second round, additionally sample the key frames recorded in step S4-1 sampling rays;

[0017] Randomly select image pixels from the key frames, and calculate the starting points and ray directions of the above two rounds of sampling rays by back-projection using the key frame poses and depth image information;

[0018] Take the sampling rays as input, and use the NeRF optimization method based on triplane hash encoding to optimize the local feature information of the NeRF map.

[0019] Preferably, in step S4-2, the NeRF optimization method based on triplane hash encoding specifically includes the following steps:

[0020] (1) For each sampling ray, sample N start sampling points along the ray direction starting from the ray origin, and uniformly sample an additional N imp sampling points within the truncation distance tr of the depth corresponding to the sampling ray;

[0021] (2) Perform triplane hash encoding and position encoding on the N = N start + N imp sampling points on each ray;

[0022] For any sampling point p, normalize the spatial coordinates of point p according to the scene boundary to obtain the normalized coordinate p nor , and project the normalized coordinates onto three aligned spatial planes to obtain the projected coordinates p nor-xy , p nor-xz , p nor-yz on the xy, xz, and yz planes, and input them into the Instant-NGP hash grid to map the coordinates and perform bilinear interpolation on the grid nodes around the mapped coordinates to obtain the features on the corresponding xy, xz, and yz planes, denoted as H xy (p), H xz (p), H yz (p);

[0023] Use two triplane hash grids with different granularity levels to encode the scene features. For the scene features of the hash grids of the three planes obtained at the rough level, use to represent; for the scene features of the hash grids of the three planes obtained at the fine level, use to represent, and obtain the rough scene feature h c (p) and the fine scene feature h f (p) according to the following formula:

[0024]

[0025] Use the One-Blob encoding as the positional encoding to obtain the positional encoding feature γ(p) of p;

[0026] (3) Take the rough scene feature h c (p), the fine scene feature h f (p), and the positional encoding feature γ(p) as inputs and input them into the decoder;

[0027] The decoder consists of a rough geometry decoder a fine geometry decoder and a color decoder f color and their structures are all two-layer multilayer perceptrons (MLPs) with an intermediate dimension of 32.

[0028] (4) Use the volume rendering method to render the sampled rays to obtain the color and depth

[0029] Convert the TSDF value φ g (p) of the sampled points output from the decoder into the volume density of the sampled points for NeRF ray rendering through the following formula:

[0030]

[0031] where β ∈ R is a learnable parameter used to control the sharpness of the surface boundary, and the Sigmoid function is defined as

[0032] Use the volume density of each sampled point on the sampled ray to calculate the rendered color of the sampled ray through integration and depth value

[0033]

[0034] where p n is the nth sampled point on the sampled ray, φ a (p n ) is the color value of the sampled point p n , and z n represents the depth value corresponding to the sampled point p n .

[0035] (5) Construct a loss function using the color and depth of the sampled ray and the depth value of the sampled point.

[0036] (6) Iteratively optimize using the loss function to train the parameters of the three - plane hash grid and the decoder network to optimize the scene representation.

[0037] Preferably, in the step (3),

[0038] the rough geometry decoder takes the position encoding γ(p) and the rough scene feature h c (p) as inputs and outputs the rough TSDF value and the rough feature vector v c ; where, TSDF refers to the truncated signed distance function;

[0039] the fine geometry decoder takes the position encoding γ(p) and the fine scene feature h f (p) as inputs and outputs the fine residual TSDF value and the fine feature vector v f ;

[0040] the color decoder f color takes the sum of the rough feature vector v c and the fine feature vector v f and the position encoding γ(p) as inputs to obtain the color value φ a (p) of the sampling point, and the TSDF value φ g (p) of the sampling point is equal to the sum of the rough TSDF value and the fine residual TSDF value .

[0041] Preferably, the step (5) specifically includes the following steps:

[0042] For the sampling points between the camera and the boundary of the truncation region, construct the forward blank interval loss function:

[0043]

[0044] In the above formula, R represents the set of sampling rays, and r represents a sampling ray. represents the set of sampling points of the sampling ray r from the camera center to the boundary of the truncation region, and the |·| operator represents calculating the number of elements in the set.

[0045] For the points where the distance from the depth value z(p) of the sampling point to the reference depth value D(r) of the sampling ray is less than the truncation distance tr, construct the truncation region center loss function and the truncation region edge loss function. The truncation region center loss function is as follows:

[0046]

[0047] Where, Denote the set of sampling points located at the center of the truncated region, i.e.,

[0048] For the sampling points located at the edge of the truncated region, construct the edge loss function of the truncated region.

[0049]

[0050] Among them, Denote the set of sampling points located at the edge of the truncated region, i.e., For the color and depth values of the sampled rays for rendering, construct the depth loss L d and the color loss function L c :

[0051]

[0052] Among them, and are the depth value and color value of the sampled ray r obtained by rendering, and D(r) and I(r) are the reference depth value and reference color value of the sampled ray r.

[0053] The final loss function is as follows:

[0054] L = λ fs L fs + λ tr-m L tr-m + λ tr-t L tr-t + λ d L d + λ c L c

[0055] Among them, λ fs , λ tr-m , λ tr-t , λ d , λ c are the preset weights of the corresponding loss terms.

[0056] The present invention has the following features and beneficial effects:

[0057] Compared with other SLAM solutions that perform tracking and mapping through NeRF rendering, this technical solution has the following characteristics and beneficial effects: Using ORB-SLAM3 as the ORB tracking module of this invention has better real-time performance; Using the combination of tri-plane hash encoding and multi-layer perceptron as the scene representation can efficiently and exquisitely reconstruct scene details; It has loop detection and global optimization functions to ensure the global consistency of the NeRF map; Using the local co-visible keyframe window as the optimized keyframe window for NeRF mapping, compared with NeRF-based visual SLAM methods such as NICE-SLAM that use video frames at fixed intervals as keyframes, it can better utilize the feature information extracted from images to screen keyframes, making the keyframes more distinguishable from each other, reducing the redundant keyframe information involved in NeRF map optimization, improving the mapping accuracy while reducing the computational complexity of the NeRF map optimization algorithm, and effectively ensuring the real-time performance of the system operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0059] Figure 1 It is the system framework of the NeRF visual SLAM mapping method based on ORB tracking and tri-plane hash encoding of the present invention.

[0060] Figure 2 It is the flowchart of the method of the present invention.

[0061] Figure 3 It is the schematic diagram of the local co-visible keyframe window in the tracking module of the present invention.

[0062] Figure 4 It is the detailed network block diagram of NeRF in the NeRF optimization method based on tri-plane hash encoding of the present invention.

[0063] Figure 5 It is the mapping effect diagram of the method of the present invention and the comparison method in the scence_0000 sequence of the Scannet dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following will specifically describe the implementation scheme of the present invention in conjunction with the drawings. It should be understood that the description here only explains the present invention and is not used to limit the present invention.

[0065] The system framework of the NeRF visual SLAM mapping method based on ORB tracking and tri-plane hashing coding proposed by the present invention is as follows Figure 1 shown. The present invention uses ORB-SLAM3 as the ORB tracking module to perform pose tracking, screen key frames, optimize the pose of key frames, screen local co-visible key frames for the current key frame, and optimize the scene representation using the NeRF optimization method based on tri-plane hashing coding. The specific method is as follows Figure 2 shown. This embodiment includes the following steps:

[0066] Step S1: Use an RGB-D camera to obtain the RGB image and depth image of the scene.

[0067] Step S2: Use ORB-SLAM3 as the ORB tracking module to extract ORB feature points and perform feature matching on the RGB image, estimate the frame pose, screen key frames, and convert the ORB feature points into 3D map points through the depth information obtained from the depth image to construct an initial sparse map:

[0068] Step S3: According to the number of 3D map points in the sparse map that are co-visible to the current key frame and other key frames, screen out local co-visible key frames, and perform local bundle adjustment optimization on the local co-visible key frames to optimize the pose of the local co-visible key frames and the positions of 3D map points. The key frame pose refers to the camera pose of the key frame.

[0069] Define the number of 3D map points in the sparse map that are co-visible to the current key frame and other key frames as the co-visibility;

[0070] The specific steps are as follows:

[0071] S3-1: According to the co-visibility relationship with the current key frame, classify the key frames that observe more than 15 identical 3D map points with the current key frame K into the first-level adjacent key frame set K L , and classify the key frames that observe more than 15 identical 3D map points with the key frames in K L into the second-level adjacent key frame set K F , Figure 3 shows the relationship between the first-level adjacent key frames and the second-level adjacent key frames. K L1 , K L2 are the first-level adjacent key frames, and they observe more than 15 identical 3D map points with the current key frame K. K F11 , K F12 observe less than 15 identical 3D map points with the current key frame K, but the common observation quantity with the first-level adjacent key frame K L1 exceeds 15, and they are the second-level adjacent key frames of the current key frame K; similarly, K F21 , KF22 The number of 3D map points jointly observed with the current key frame K is less than 15, but the number of 3D map points jointly observed with the adjacent key frame K at the same level L2 exceeds 15, and it is the second-level adjacent key frame of the current key frame K.

[0072] S3-2. Through the set K of first-level adjacent key frames L and the set K of second-level adjacent key frames F to construct the set χ of matching pairs between the 3D map points and the ORB feature points of the current key frame k Construct the reprojection loss term and use the Levenberg-Marquardt algorithm to optimize and solve. The specific optimization objective can be expressed as the following minimization problem:

[0073]

[0074] where P L is the set of 3D map points observed in K L , X i is the three-dimensional spatial position of the specific 3D map point i in P L , R l and t l are the rotation matrix and translation vector of the key frame l in K L respectively. The function ρ represents the robust Huber loss function, and its definition is where δ is a hyperparameter used to control the error break point, and E kj represents the reprojection error term constructed by the key frame k and the matched 3D map point j in χ k , which is as follows.

[0075]

[0076] x j and X i represent the image coordinates of the ORB feature points of the matched key frame and the spatial coordinates of the 3D map point j in K L and K F respectively. R k and t k are the rotation matrix and translation vector of the key frame k respectively. π is the projection function, and its role is to project the 3D map points in space onto the key frame k.

[0077] Step S4: Sort according to the co-visibility with the current key frame, intercept the first 20 key frames with the highest co-visibility with the current key frame as the local co-visibility key frame window, and perform ray sampling on the key frames therein. Use the NeRF optimization method based on tri-planar hash encoding to optimize the scene features corresponding to these key frames in the NeRF map. Among them, the tri-planar hash encoding is used to represent the scene geometry features and appearance features. The specific process is as follows:

[0078] S4-1: Sort according to the co-visibility with the current key frame, intercept the first 20 key frames with the highest co-visibility with the current key frame as the local co-visibility key frame window, transform the key frame poses in the local co-visibility key frame window into the NeRF map coordinate system, and record the transformed key frame poses. The coordinate transformation of the key frame pose is shown by the following formula:

[0079]

[0080] Among them represents the transformation matrix of key frame c1 in the NeRF map coordinate system, and represents the transformation matrix of the NeRF map initialization key frame c std The transformation matrix of represents the transformation matrix from key frame c1 to c std The transformation matrix of p T represents the transformation matrix from the OpenCV camera coordinate system to the OpenGL camera coordinate system. Specifically

[0081] S4-2: Perform ray sampling on the key frames in the local co-visibility key frame window. In the first round, sample N ray rays for the key frames in the local co-visibility key frame window. In the second round, sample an additional rays for the key frames newly recorded in step S4-1. Randomly select image pixels from the key frames, and calculate the starting points and ray directions of rays by back-projection using the key frame poses and depth image information. Use the NeRF optimization method based on tri-planar hash encoding to optimize the local information of the NeRF map with the sampled rays as the input;

[0082] The NeRF optimization method based on tri-planar hash encoding includes the following steps:

[0083] (1) Point sampling based on depth information. For each sampling ray, sample N start sampling points along the ray direction starting from the ray origin in layers, and uniformly sample an additional N imp sampling points within the truncation distance tr of the depth corresponding to the sampling ray.

[0084] (2) For each of the N = N start + N imp sampling points on each ray, perform three - plane hash encoding and position encoding.

[0085] For the sampling point p, normalize the spatial coordinates of point p according to the scene boundary to obtain the normalized coordinate p nor , and project the normalized coordinates onto three aligned spatial planes to obtain the projected coordinates p nor-xy 、p nor-xz 、p nor-yz on the xy, xz, and yz planes respectively. And input them into the Instant - NGP hash grid. The hash grid maps the coordinates and performs bilinear interpolation on the grid nodes around the mapped coordinates to obtain the features corresponding to the xy, xz, and xz planes. The features of the three planes obtained by the above process are represented by H xy (p), H xz (p), H yz (p).

[0086] The present invention uses three - plane hash grids at two granularity levels to encode scene features. For the scene features of the three - plane hash grids obtained at the rough level, use to represent; for the scene features of the three - plane hash grids obtained at the fine level, use to represent. The finally obtained rough scene feature h c (p) and fine scene feature h f (p) are as follows:

[0087]

[0088] Use One - Blob encoding as the position encoding, and the feature of encoding point p is represented by γ(p).

[0089] (3) Take the rough scene feature h c (p), the fine scene feature h f (p) and the position - encoding feature γ(p) as inputs and input them into the decoder. The structures of the encoder and decoder are as Figure 4 shown:

[0090] The decoder consists of a rough geometry decoder a fine geometry decoder and a color decoder f color and is composed of three parts. Their structures are both two - layer multi - layer perceptrons with an intermediate dimension of 32.

[0091] The rough geometry decoder combines the position encoding γ(p) and the rough scene feature h c(p) is used as input to output a rough TSDF value and a rough feature vector v c , where TSDF refers to the truncated signed distance function. Similarly, the fine geometric decoder position encoding γ(p) and the fine scene feature h f (p) are used as input to output a fine residual TSDF value and a fine feature vector v f . The rough geometric decoder The fine geometric decoder The input and output processes are specifically shown as follows:

[0092]

[0093] After that, the sum of the rough feature vector v c and the fine feature vector v f is input into the color decoder f color together with the position encoding γ(p) to obtain the color value φ a (p) of the final sampling point, and the TSDF value φ g (p) of the final sampling point is equal to the sum of the rough TSDF value and the fine residual TSDF value . Their input and output are shown as follows:

[0094]

[0095] (4) Render the color and depth

[0096] Convert the TSDF value of the sampling points output from the decoder into the volume density of the sampling points for NeRF ray rendering through the following formula:

[0097]

[0098] where β ∈ R is a learnable parameter that mainly controls the sharpness of the surface boundary, and the Sigmoid function is defined as

[0099] Use the volume density of each sampling point on the sampling ray to calculate the rendered color and depth value

[0100]

[0101] where p n is the nth sampling point on the sampling ray, and φ a (pn ) is the color value of the sampling point p n , and z n represents the depth value corresponding to the sampling point p n .

[0102] (5) Construct a loss function through the spatial information of the sampling points, the depth information of the rendering rays, and the color information of the rendering rays.

[0103] For the sampling points between the camera and the boundary of the truncation region, construct a forward blank interval loss function:

[0104]

[0105] In the above formula, R represents the set of sampling rays, and r represents a sampling ray. represents the set of sampling points of the sampling ray r from the camera center to the boundary of the truncation region, and the |·| operator represents calculating the number of elements in the set.

[0106] For the points where the distance from the depth value z(p) of the sampling point to the reference depth value D(r) of the sampling ray is less than the truncation distance tr, construct a truncation region center loss function and a truncation region edge loss function. The truncation region center loss function is as follows:

[0107]

[0108] where represents the set of sampling points located at the center of the truncation region, that is

[0109] For the sampling points located at the edge of the truncation region, construct a truncation region edge loss function.

[0110]

[0111] where represents the set of sampling points located at the edge of the truncation region, that is

[0112] For the color and depth values of the rendered sampling rays, construct a color loss function and a depth loss function:

[0113]

[0114] where and are the depth value and color value of the rendered sampling ray r, and D(r) and I(r) are the reference depth value and reference color value of the sampling ray r.

[0115] The final loss function is as follows:

[0116] L = λ fs L fs + λ tr-m L tr-m + λ tr-t L tr-t + λ d L d + λ c L c

[0117] where λ fs , λ tr-m , λ tr-t , λ d , λ c are the weights of the preset corresponding loss terms.

[0118] (6) Use the loss function for iterative optimization to train the network parameters of the three-plane hash encoder and decoder to optimize the scene representation.

[0119] Step S5: Use the loop closure module based on ORB-SLAM3. After detecting a loop, optimize the global keyframes using the global bundle adjustment method. Iterate multiple times. During each iteration, randomly select several keyframes from the global keyframes for ray sampling, and use the NeRF optimization method based on three-plane hash encoding to globally optimize the scene features of the NeRF map to ensure the global consistency of the constructed map.

[0120] The specific process is as follows:

[0121] S5-1: Use the loop closure module based on ORB-SLAM3. After detecting a loop, optimize the global keyframe poses using the global bundle adjustment method;

[0122] S5-2: Use the global keyframes for iterative optimization of the map features multiple times. In each iteration, randomly select 20 keyframes from the global keyframes, perform ray sampling on these 20 keyframes, and use the NeRF optimization method based on three-plane hash encoding in S4-2 with the sampled rays as the input to globally optimize the NeRF map to ensure the global consistency of the constructed map.

[0123] Step S6: Use the Marching Cubes algorithm to extract the surface information of the scene from the NeRF map to generate a high-quality 3D reconstruction model.

[0124] To verify the effectiveness of the present invention, the following experiments are carried out:

[0125] The present invention is evaluated on the publicly available datasets Replica, ScanNet, and TUM-RGBD datasets. For the accuracy of tracking, the present invention uses the Root Mean Square Error of the Absolute Trajectory Error (ATE RMSE) for evaluation; for the results of 3D reconstruction, the present invention, like NICE-SLAM, uses Accuracy (Acc.), Completion (Comp.), Completion Ratio (Comp.Ratio), and Depth L1 error to evaluate the results of 3D reconstruction.

[0126] Above, the Root Mean Square Error of the Absolute Trajectory Error (ATE RMSE) refers to the root mean square error of the camera pose between the true trajectory and the estimated trajectory for each frame after aligning the true trajectory and the estimated trajectory. Its calculation formula is shown as follows:

[0127]

[0128] e i =‖p est,i -p gt,i ‖

[0129] where n is the number of poses in the trajectory, and e i refers to the translational error for each frame. The ‖·‖ operator represents the Euclidean distance, and p est,i and p gt,i represent the i-th position of the estimated trajectory and the i-th position of the true trajectory, respectively.

[0130] Accuracy refers to the average distance from the sampled points in the reconstructed mesh to the nearest true point. Completion refers to the average distance from the sampled points in the true mesh to the nearest reconstructed point. Completion Ratio refers to the percentage of the completion in the reconstructed mesh that is less than 5 cm. The Depth L1 error specifically refers to the L1 loss of 1000 randomly sampled depth maps in the reconstructed mesh and the true mesh.

[0131] All experiments of the present invention are conducted on a computer with an Intel i9-13900K, RTX3090, and 32G of memory.

[0132] 1. Replica dataset:

[0133] The Replica dataset is a high-quality indoor environment dataset developed by Facebook Reality Labs (FRL), which contains 18 highly realistic indoor scenes. The scenes in its dataset are accurately modeled based on real indoor environments, with a high degree of realism. The Replica dataset also provides high-resolution RGB-D images and semantic information of various objects in the scenes, which is very helpful for various visual tasks (such as object detection, semantic segmentation, SLAM, etc.). In the present invention, the video frame sequences of room0, room1, room2, office0, office1, office2, office3, and office4 are selected for experiments, and the tracking accuracy and the quality of 3D scene reconstruction are evaluated. Each video frame sequence provides 2000 RGB images and 2000 depth images. The evaluation results of the present invention and other NeRF-based visual SLAMs on the above 8 video frame sequences of the Replica dataset are shown in Table 1:

[0134] Table 1. Results of different methods verified on the Replica dataset

[0135]

[0136]

[0137] The comparison of the verification results of the method of the present invention and the prior art on the Replica dataset is shown in Table 1.

[0138] On the Replica dataset, the evaluation of the 3D scene reconstruction quality of the method of the present invention is significantly better than that of the iMAP method and the NICE-SLAM method. Its tracking accuracy index ATE RMSE is also much lower than these two methods. Compared with the ESLAM method, the 3D scene reconstruction quality and tracking accuracy of the method of the present invention are slightly higher than those of the ESLAM. However, as shown in Table 4, the running efficiency of the method of the present invention is significantly higher than that of other NeRF-based visual SLAM methods.

[0139] 2. ScanNet dataset:

[0140] The ScanNet dataset is a large-scale RGB-D video dataset jointly developed by researchers from Stanford University, Princeton University, and the Technical University of Munich. It contains 1,513 scenes and provides 2.5 million views. These views include RGB images and depth images, along with annotations of precise 3D camera poses, surface reconstructions, and instance-level semantic segmentations, and have extensive applications in tasks such as 3D scene understanding, semantic segmentation, visual recognition, SLAM, and 3D reconstruction. The method of the present invention uses six video frame sequences, namely Sc.0000, Sc.0059, Sc.0106, Sc.0169, Sc.0181, and Sc.0207, for verification. Since the real mesh scene reconstruction provided in ScanNet is incomplete, the verification of this method for this dataset is mainly evaluated on the ATERMSE parameter. The reconstruction result of the method of the present invention on the Sc.0000 dataset is as Figure 5 shown. The evaluation results of the method of the present invention and other NeRF-based visual SLAM methods on the ScanNet dataset are shown in Table 2:

[0141] Table 2. Root Mean Square Error (cm) Results of Absolute Trajectory Errors of Different Methods on the ScanNet Dataset

[0142]

[0143]

[0144] As can be seen from Table 2, the positioning accuracy of the method of the present invention on the ScanNet dataset is slightly lower than that of ESLAM, but much higher than that of the iMAP and NICE-SLAM methods. And as can be seen from Table 4, the running speed of the method of the present invention far exceeds that of other methods. Considering the positioning accuracy and running efficiency, the experimental effect of the method of the present invention is better than that of other comparison methods.

[0145] 3. TUM RGB-D Dataset:

[0146] The TUM RGB-D dataset is provided by the Computer Vision Group of the Technical University of Munich and is widely used in the research fields such as visual SLAM and visual inertial odometry. The dataset contains RGB images and depth images recorded using a Microsoft Kinect sensor, as well as the ground truth trajectory data of the sensor. The data is recorded at a frequency of 30Hz and a resolution of 640×480. The method of the present invention is mainly verified using fr1 / desk, fr2 / xyz, and fr3 / office. Since the dataset does not provide the mesh of the actual scene, the method of the present invention only verifies the tracking accuracy and uses the ATE RMSE parameter for evaluation. The evaluation results of the method of the present invention and other NeRF-based visual SLAM methods on the TUM RGB-D dataset are shown in Table 3 as follows:

[0147] Table 3. Root Mean Square Error (cm) Results of Absolute Trajectory Errors of Different Methods on the TUM RGB-D Dataset

[0148] Methods fr1 / desk fr2 / xyz fr3 / office iMAP* 4.90 2.05 5.80 NICE-SLAM 2.85 2.39 3.02 ESLAM 2.47 1.11 2.42 The present invention 1.60 0.50 0.10

[0149] Table 4. Running Efficiencies (HZ) of Different Methods on the Replica Dataset and the ScanNet Dataset

[0150] iMAP* NICE-SLAM ESLAM The present invention Replica 0.19 0.48 5.5 10 ScanNet 0.19 0.29 2 7

[0151] The running efficiency in Table 4 refers to the number of image frames that the system can process per second for tracking and mapping when different methods are running on different datasets. It can be seen that on different datasets, the method of the present invention can run at a relatively high frame rate while ensuring the mapping quality, realizing real-time positioning and high-quality 3D scene reconstruction.

[0152] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, and the embodiments of the present invention have been shown and described. For those skilled in the art, without departing from the principles and spirits of the present invention, various changes, modifications, substitutions, and variations to these embodiments including components still fall within the protection scope of the present invention.

Claims

1. A NeRF visual SLAM mapping method based on ORB tracking and three-plane hash coding, characterized in that: The steps include: Step 1: Obtain the RGB image and depth image of the scene through the RGB-D camera; Step 2: Use ORB-SLAM3 as the ORB tracking module to extract and match ORB feature points from the RGB image, estimate the frame pose, filter out key frames, and convert the ORB feature points into 3D map points through the depth information obtained from the depth image to build the initial sparse map: Step 3, according to the number of 3D map points in the sparse map that are visible to the current keyframe and other keyframes, select the local common view keyframe, perform local bundle adjustment (Local BA) on the local common view keyframe, and optimize the local common view keyframe pose and 3D map point position; define the number of 3D map points in the sparse map that are visible to the current keyframe and other keyframes as common view; Step 4: sort the local common view key frames according to the common view degree with the current key frame, intercept several key frames with the largest common view degree with the current key frame as the local common view key frame window, and perform light sampling on the key frames in the local common view key frame window; The scene features corresponding to these keyframes in the NeRF map are locally optimized using a NeRF optimization method based on three-plane hash coding, where three-plane hash coding is used to represent scene geometry and appearance features. Step 5: After detecting the loop, the global keyframe pose is optimized by global bundle adjustment using the loop closure module based on ORB-SLAM3; then, the global keyframe is used to perform multiple iterative optimizations of map features. During the iterative process, the scene features of the NeRF map are globally optimized using the NeRF optimization method based on three-plane hash coding to ensure the global consistency of the constructed map; Step 6: Use the Marching Cubes algorithm to extract the surface information of the scene from the NeRF map and generate a 3D reconstruction model.

2. The NeRF visual SLAM mapping method based on ORB tracking and three-plane hash coding according to claim 1, characterized in that, The step 4 specifically comprises the following steps: S4-1, sorting according to the co-viewing degree with the current key frame, intercepting the first 20 key frames with the largest co-viewing degree with the current key frame as the local co-viewing key frame window, converting the key frame pose in the local co-viewing key frame window to the NeRF map coordinate system, and recording the key frame pose; S4-2, sampling the key frames in the local common view key frame window; in the first round, sampling N local common view key frames ray The second round samples the key frames recorded in step S4-1. A sampling ray; Randomly select image pixels from the key frame, and calculate the starting point and direction of the above two rounds of sampling light by back-projecting the key frame pose and depth image information; Taking the sampling rays as input, the local feature information of the NeRF map is optimized using the NeRF optimization method based on three-plane hash coding.

3. The NeRF visual SLAM mapping method based on ORB tracking and three-plane hash coding according to claim 2, characterized in that, In step S4-2, the NeRF optimization method based on three-plane hash coding specifically includes the following steps: (1) For each sampling ray, N layers of sampling are performed along the direction of the ray starting from the ray origin. start sampling points, and uniformly sample additional N points within the cutoff distance tr corresponding to the depth of the sampling ray imp sampling points; (2) For each ray, N = N start +N imp The sampling points are coded with three-plane hash and position; For any sampling point p, the spatial coordinates of point p are normalized according to the scene boundary to obtain the normalized coordinates p nor , and project the normalized coordinates onto three aligned spatial planes to obtain the projection coordinates p of the xy, xz, and yz planes nor-xy 、p nor-xz 、p nor-yz , and input them into the Instant-NGP hash grid, map the coordinates, and perform bilinear interpolation on the grid nodes around the mapped coordinates to obtain the corresponding features on the xy, xz, and yz planes, denoted as H xy (p), H xz (p), H yz (p); The scene features are encoded using three-plane hash grids with two different granularity levels. For the scene features of the three-plane hash grid obtained at the coarse level, use Represented; for the scene features of the hash grid of the three planes obtained at the fine level, use Expressed as follows, the rough scene feature h is obtained by the following formula c (p) and fine scene features h f (p): Use One-Blob encoding as the position encoding to obtain the position encoding feature γ(p) of p; (3) The rough scene feature h c (p) and fine scene features h f (p) and the position encoding feature γ(p) are taken as input and input to the decoder; The decoder is a coarse geometry decoder. Fine Geometry Decoder Color Decoder f color The structure of the two-layer multi-layer perceptron is 32 in the middle dimension. (4) Use volume rendering method to render the sampling light and obtain the color of the sampling light and depth The TSDF value φ of the sampling point output from the decoder g (p) is converted to the volume density of the sampling points for NeRF ray rendering by the following formula: Where β∈R is a learnable parameter that controls the sharpness of the surface boundary, and the Sigmoid function is defined as Use the volume density of each sampling point on the sampling light to render the sampling light color by integral calculation and depth value in p n is the nth sampling point on the sampling ray, φ a (p n ) is the sampling point p n The color value, z n represents the sampling point p n The corresponding depth value; (5) Using the color of the sampled light and depth Construct the loss function with the depth value of the sampling point; (6) The loss function is used for iterative optimization to train the three-plane hash grid and decoder network parameters to achieve optimization of the scene representation.

4. The NeRF visual SLAM mapping method based on ORB tracking and three-plane hash coding according to claim 3, characterized in that, In the step (3), The coarse geometry decoder The position encoding γ(p) and the rough scene feature h c (p) as input, output rough TSDF value and the rough eigenvector v c ; The fine geometry decoder The position encoding γ(p) and the fine scene feature h f (p) as input, output fine residual TSDF value and the refined feature vector v f ; The color decoder f color The rough feature vector v c With the refined feature vector v f The sum of the same position encoding γ(p) is used as input to obtain the color value φ of the sampling point a (p), and the TSDF value φ at the sampling point g (p) is equal to the rough TSDF value Same as the fine residual TSDF value sum.

5. The NeRF visual SLAM mapping method based on ORB tracking and three-plane hash coding according to claim 3, characterized in that, The step (5) specifically comprises the following steps: For the sampling points between the camera and the boundary of the truncated region, construct the forward blank interval loss function: In the above formula, R represents the sampling light set, and r represents the sampling light; represents the set of sampling points of the sampling ray r from the camera center to the boundary of the truncation area, and the |·| operator represents the number of elements in the calculation set; For the point where the distance from the sampling point depth value z(p) to the sampling light reference depth value D(r) is less than the truncation distance tr, the truncation area center loss function and the truncation area edge loss function are constructed. The truncation area center loss function is as follows: in, represents the set of sampling points located at the center of the truncated region, For the sampling points located at the edge of the truncated region, a truncated region edge loss function is constructed; in, Represents a set of sampling points located at the edge of the truncated area. For the color and depth values ​​of the rendered sampling light, a depth loss L is constructed. d And the color loss function L c : in, and are the depth value and color value of the sampling light r obtained by rendering, D(r) and I(r) are the reference depth value and reference color value of the sampling light r; The total loss function is constructed based on the blank interval loss function, the truncated area center loss function, the constructed truncated area edge loss function, the depth loss and the color loss function.

Citation Information

Cited By

  • End-to-end 3D environment SLAM method and apparatus, and electronic device

    CN122435188A