Dense simultaneous localization and mapping method based on voxel-based neural implicit surfaces

By adopting a method based on the implicit surface of voxel nerves in dense positioning and mapping methods, using Morton encoding and dynamic voxel block construction, the problem of large occupation of invisible areas and video memory space in the prior art is solved, and efficient and realistic visual effects and good generalization capabilities are achieved.

CN115619951BActive Publication Date: 2025-05-13ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211263616.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-16
Publication Date
2025-05-13
Estimated Expiration
2042-10-16

AI Technical Summary

Technical Problem

The existing dense positioning and mapping methods cannot effectively predict invisible areas, making it difficult to generate realistic visual effects from a new perspective, and requires a large amount of video memory space, and the generalization ability is limited, especially in unknown scenarios.

Method used

The dense synchronous positioning and mapping method based on the implicit surface of voxel nerves is adopted to dynamically update the scene information through the data area shared by the front and back ends, and the voxel block index is accelerated by Morton encoding to construct the octree structure and texture feature vector to realize dynamic voxel block construction and map expansion.

Benefits of technology

It realizes the generation of realistic visual effects from a new perspective, reduces storage requirements, accelerates the positioning and mapping process, supports dynamic map expansion and editing, and has good generalization capabilities in unknown scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115619951B_ABST
    Figure CN115619951B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for dense simultaneous positioning and mapping based on voxel neural implicit surface. The present invention decomposes a three-dimensional scene into geometric units with voxel blocks as units, and stores its internal geometry and texture information in the form of feature vectors in the voxel blocks, obtains the features of the corresponding three-dimensional points by interpolation, and obtains the signed distance field (SDF) and the corresponding color through the two parts of the geometric analysis network and the texture analysis network. On this basis, the present invention proposes cross-iterative optimization of the two processes of positioning and mapping, and transfers the map latent feature vector between the two processes by sharing variables; the present invention innovatively introduces an octree method based on Morton coding to further improve the efficiency of map updates. The present invention can render the edited surface and texture effects by interactively editing the generated voxel blocks, so that it can be applied to applications such as virtual reality and augmented reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and computer graphics, and in particular to a dense positioning and mapping method based on voxel neural implicit surface. Background Art

[0002] Dense Localization and Mapping (DSLAM) is the basis of many 3D applications. Based on the accurate map reconstructed by 3D, some interactive displays such as occlusion and collision can be completed in the virtual-real fusion scene, achieving more realistic effects in enhanced display applications.

[0003] Traditional DSLAM methods usually use feature matching-based methods and optimization methods that minimize energy functions to solve camera poses and optimize map structures. These methods usually use discrete point clouds, surface elements or continuous signed distance fields (SDFs) to represent dense maps, but the existing problems are also obvious. First, since these methods cannot predict invisible areas, they are usually unable to synthesize realistic visual effects from new perspectives. Second, these methods require a large amount of video memory space.

[0004] Methods based on deep features, such as code-slam and di-fusion, store local scene information in compressed codes and optimize these code fields through multi-view constraints to update the map. Although these methods reduce storage, they are constrained by the network's expression capabilities and pre-training scenarios, and generalization to new scenarios will cause problems.

[0005] With the rise of Neural Radiance Field (NeRF) technology, using MLP networks to store scene information and generate realistic rendering effects from various perspectives has become a new development trend. For example, the iMap method uses this kind of idea to complete the DSLAM system based on neural implicit fields. However, the problem with this system is that the entire scene is stored in a single MLP, which requires prior information about the size of the scene. This makes it impossible for this type of method to model unknown scenes. Moreover, since the scene is implicitly stored in the MLP, further operations such as editing the scene become very difficult. Summary of the invention

[0006] In order to solve the problems in the prior art, the present invention provides a method for dense synchronous positioning and mapping based on voxel neural implicit surface. The method of the present invention includes two processes: front-end tracking and back-end mapping. The information related to the scene is stored in a data area shared by the front-end and back-end, and is dynamically updated as it runs. When the system starts, the global mapping is initialized by running some mapping iterations for the first frame. The system accepts a sequence of RGBD images as input, creates voxels only for areas where depth information exists, and optimizes the surface and texture information in the voxels. The front-end tracking process aligns the surface and texture of the existing map of the current frame, and gradually optimizes the camera's pose. The back-end mapping process jointly optimizes the frame with the estimated pose with the existing map, and updates the map.

[0007] In order to achieve the above object, the present invention adopts the following technical solution:

[0008] The present invention first provides a dense synchronous localization and mapping based on voxel neural implicit surface, comprising the following steps:

[0009] Step 1: Get the RGB-D image of the first frame, back-project the depth corresponding to each pixel in the first frame image into the three-dimensional space, and obtain the initial three-dimensional point cloud in the map; set the coordinate system where the initial three-dimensional point cloud is located as the reference coordinate system, and construct multiple non-overlapping voxel blocks aligned with the coordinate axes of the reference coordinate system based on the initial three-dimensional point cloud; construct an octree structure based on these voxel blocks, and insert the Moron code corresponding to the voxel block into the octree; at the same time, assign fixed-length feature vectors to the 8 vertices of each voxel block, and these fixed-length feature vectors are used to store the geometry and texture information of the scene to be constructed;

[0010] Step 2: Randomly sample M pixels from the acquired image, generate a ray starting from the camera center corresponding to the image and passing through each pixel, and calculate the intersection of the ray and the constructed voxel block; uniformly sample the area where the ray and the voxel block intersect to obtain the sampled 3D points, obtain the feature vectors of the 8 vertices of the voxel block where the 3D point is located through the 3D coordinates of the 3D point, and obtain the feature vector corresponding to the 3D point through the feature extraction function; obtain the signed distance field (SDF) and intermediate information through the geometric parsing network, and then obtain the color of the intermediate information through the texture parsing network; calculate the spatial density value corresponding to the 3D point through the SDF, and then perform weighted accumulation of the color and depth of the 3D point on the ray through volume rendering, and finally obtain the predicted color and depth of the pixel corresponding to the ray; compare the predicted color and depth with the real color and depth, thereby optimizing the fixed-length feature vector on the voxel block vertex and the geometric parsing network and texture parsing network;

[0011] Step 3: After step 2 is completed, the tracking process is started. The tracking process is: repeat step 2 for the image obtained from the second frame, but keep the fixed-length feature vectors on the voxel block vertices and the geometric analysis network and texture analysis network unchanged, and only optimize the camera 6-DOF pose corresponding to the image. After optimization, positioning is completed, and the optimized camera 6-DOF pose and the corresponding RGB-D image are constructed into a frame and put into the candidate key frame list;

[0012] Step 4: Start the mapping process. The mapping process is as follows: get the keyframe list from step 3, traverse the candidate keyframe list, and back-project the depth corresponding to the pixel point of each frame image into the three-dimensional space according to the 6-DOF position of the camera corresponding to the image to obtain the three-dimensional point cloud corresponding to each frame; for each three-dimensional point in the three-dimensional point cloud, determine whether the three-dimensional point is included in the created voxel block. If not,

[0013] Create a new voxel block and update the octree structure in step 1, so as to achieve the purpose of dynamically creating voxel blocks and expanding the mapping area;

[0014] Select several suitable frames from the key frame list as key frames and optimize them together with the latest frame in the candidate key frame list; repeat step 2 for the images in all frames to be optimized, and optimize the 6-DOF pose of the frame while optimizing the fixed-length feature vectors on the vertices of the voxel blocks and the geometric parsing network and the texture parsing network.

[0015] Furthermore, the step 1 of constructing a plurality of non-overlapping voxel blocks aligned with the coordinate axes of the reference coordinate system based on the initial three-dimensional point cloud is specifically as follows:

[0016] The initial 3D point cloud is divided into a set of voxel blocks Divide, each voxel block has three-dimensional coordinates V k =(x, y, z); these three-dimensional coordinates are converted into 64-bit binary code information through Morton coding; each voxel block has 8 vertices, each of which contains a fixed-length feature vector The geometry and texture information of the scene to be constructed, L e is the length of the eigenvector; therefore, for any voxel V i Any three-dimensional point in Adjacent voxel blocks share the feature vectors of 4 vertices.

[0017] As a preferred solution of the present invention, the intersection of the calculated ray in step 2 and the voxel block constructed in step 1 is specifically:

[0018] The ray passing through the pixels on the image from the camera center o along the direction d is defined as r(t)=o+dt, where t is the depth along the ray direction; each ray uses the Ray-AABB intersection detection algorithm to calculate the depth of the intersection of the ray and the voxel block, thereby dividing the area on the ray that intersects with the voxel block.

[0019] Furthermore, in step 2, the feature vector corresponding to the three-dimensional point is obtained by the feature extraction function, the signed distance field (SDF) and the intermediate information are obtained by the geometric analysis network, and the color is obtained by the obtained intermediate information through the texture analysis network, specifically:

[0020] Feature extraction function Map the 3D point p to a point of length L e The eigenvector of The feature extraction function is implemented by trilinear interpolation. According to the three-dimensional coordinates of p and the relative position of p in the voxel block, the feature vectors contained in the eight vertices of the voxel block are interpolated to obtain the feature vector e of p.

[0021] A multi-layer perceptron network (MLP) is used to represent the geometric parsing network F σ and texture parsing network F c ; Geometric parsing network Generate its signed distance field through p's eigenvector e and the length is Lf geometric eigenvector The sign of σ indicates whether p is inside or outside the surface S; the surface S of the scene is extracted by:

[0022]

[0023] The operation [0] means to σ The signed distance field σ of position p is obtained from the 3D point p; the geometric feature vector f of the 3D point p, the ray direction d where p is located, and the feature vector e of p are connected as the texture parsing network F c Input, get the color c at p.

[0024] Furthermore, in step 2, the color and depth of the three-dimensional points on the ray are weighted and accumulated by volume rendering, and finally the color and depth of the pixel corresponding to the predicted ray are obtained, specifically:

[0025] Using the function φ s (σ) Convert the SDF of a 3D point p to a density, φ s (σ) is a function of the signed distance σ of point p, where

[0026]

[0027] in is the Sigmoid function, tr is the predefined cutoff distance, and the value of points close to the surface is greater than the weight of distant points;

[0028] Based on φ s (σ), normalize the density on the same ray, and normalize the density on the ray p 3D sampling points are used for volume rendering to obtain the accumulated color C(r) and depth D(r):

[0029]

[0030]

[0031] where c i is the color of point i on the ray, d i is the distance from point i on the ray to the optical center.

[0032] Compared with the prior art, the advantages of the present invention are:

[0033] 1) The present invention utilizes Morton encoding scene voxel structure to accelerate the indexing speed of voxel blocks, thereby accelerating the speed of positioning and mapping.

[0034] 2) The voxel neural implicit surface method of the present invention can construct a more complete surface structure with realistic colors, and support dynamic voxel block construction and map expansion. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic diagram of the method of the present invention;

[0036] Figure 2 It is a diagram showing the reconstruction effect of the present invention. DETAILED DESCRIPTION

[0037] The present invention is described in detail below in conjunction with the accompanying drawings. The technical features of each embodiment of the present invention can be combined accordingly without conflicting with each other.

[0038] Reference Figure 1 The present invention utilizes two processes, front-end tracking and back-end mapping. The information related to the scene is stored in a data area shared by the front-end and back-end, and is dynamically updated as it runs. The front-end tracking process aligns the surface and texture of the existing map of the current frame, and gradually optimizes the camera's posture. The back-end mapping process jointly optimizes the frame with the estimated posture with the existing map, and updates the map. The present invention is described in detail below. The dense synchronous positioning and mapping method based on voxel neural implicit surface of the present invention includes the following steps:

[0039] Step 1: Get the RGB-D image of the first frame, and back-project the depth corresponding to each pixel in the first frame image into the three-dimensional space to obtain the initial three-dimensional point cloud in the map. Set the coordinate system where the initial three-dimensional point cloud is located as the reference coordinate system, and divide the initial three-dimensional point cloud into multiple non-overlapping voxel blocks that are aligned with the coordinate axes of the reference coordinate system. Specifically, each voxel block With three-dimensional coordinates V k =(x, y, z). These three-dimensional coordinates are converted into 64-bit binary coded information through Morton coding. Each voxel block has 8 vertices, each of which contains a fixed-length feature vector The geometry and texture information of the scene to be constructed, L e is the length of the eigenvector; therefore, for any voxel V i Any three-dimensional point in Adjacent voxel blocks share the feature vectors of 4 vertices. An octree structure is constructed based on these voxel blocks, and the Morton code corresponding to the voxel block is inserted into the octree; at the same time, a fixed-length feature vector is assigned to each voxel block's 8 vertices. These fixed-length feature vectors are used to store the geometry and texture information of the scene to be constructed, such as Figure 2 As shown in the left picture.

[0040] Step 2: If Figure 1 As shown in the volume rendering section, M pixels are randomly sampled from the image, and a ray is generated from the camera center corresponding to the image and passes through each pixel, and the intersection of the ray and the voxel block constructed in step 1 is calculated. Specifically, the ray passing through the pixel on the image from the camera center o along the direction d is defined as r(t) = o + dt, where t is the depth along the ray direction; each ray is calculated through the Ray-AABB intersection detection algorithm to find the depth of the intersection of the ray and the voxel block, thereby dividing the area on the ray that intersects with the voxel block.

[0041] In the area that intersects with the voxel block, a 3D point p is sampled with uniform probability and the feature vector e of p is obtained through the feature extraction function. Specifically, the feature extraction function is defined as Map the 3D point p to a point of length L e The eigenvector of The feature extraction function is implemented through trilinear interpolation. According to the three-dimensional coordinates of p and the relative position of p in the voxel block, the feature vectors contained in the eight vertices of the voxel block are interpolated to obtain the feature vector of p.

[0042] The signed distance field (SDF) and intermediate information are obtained through the geometric analysis network, and then the intermediate information is passed through the texture analysis network to obtain the color. The spatial density value corresponding to the 3D point is calculated through the SDF, and then the color and depth of the 3D point on the ray are weighted and accumulated through volume rendering, and finally the color and depth of the pixel corresponding to the predicted ray are obtained. Specifically, the function φ is used. s (σ) Convert the SDF of a 3D point p to a density, φ s (σ) is a function of the signed distance σ of point p, where

[0043]

[0044] in is the Sigmoid function, and tr is a predefined cutoff distance. Points close to the surface have a greater weight than those far away;

[0045] Based on φ s (σ), normalize the density on the same ray, and normalize the density on the ray p The accumulated color C(r) and depth D(r) can be obtained by performing volume rendering on three-dimensional sampling points:

[0046]

[0047]

[0048] where c i is the color of point i on the ray, d i is the distance from point i on the ray to the optical center.

[0049] The predicted color and depth are compared with the actual color and depth, thereby optimizing the fixed-length feature vectors on the vertices of the voxel block and the geometry parsing network and the texture parsing network.

[0050] Step 3: If Figure 1 As shown in the tracking process, the process repeats step 2 for the image starting from the second frame, but keeps the fixed-length feature vectors on the voxel block vertices and the geometric parsing network and texture parsing network unchanged, only optimizes the camera 6-DOF pose corresponding to the image, and constructs the optimized camera 6-DOF pose and the corresponding RGBD image into a frame and puts it into the candidate key frame list, as shown in Figure 1 The shared data area is shown.

[0051] Step 4: Figure 1As shown in the mapping process, the process selects several suitable frames as key frames from the candidate key frame list in step 3, builds a new voxel block, and optimizes it together with the latest frame in the candidate key frame list. For the images in all frames to be optimized, the process of step 2 is repeated, and while optimizing the fixed-length feature vectors and the geometric parsing network and texture parsing network on the vertices of the voxel block, the 6-DOF pose of the frame is optimized. During the mapping process, the map will gradually grow until the scene to be reconstructed is included in the voxel block.

[0052] Example

[0053] The present invention has conducted experimental comparisons on public datasets such as ScanNet and Replica. Its reconstruction accuracy, completeness, and camera pose estimation accuracy are significantly improved compared to existing methods, and it is faster and requires less storage. The present invention has also conducted generalization tests on outdoor scenes and can obtain better reconstruction results.

[0054] The following table shows the positioning effect of the present invention in the Replica dataset. The table compares three measurement indicators (RMSE, mean, median) of trajectory accuracy in 8 scenes. The smaller the corresponding value, the higher the accuracy. Compared with the two existing methods (iMap* and NICE-SLAM), it can be seen that the method proposed by the present invention is superior to the existing methods in all three indicators, indicating that the present invention has better effect.

[0055]

[0056] The following table shows the reconstruction effect of the present invention in the Replica dataset, which compares the reconstruction accuracy and completeness (Acc, Comp, Comp Ratio) of 8 scenes. The smaller the values ​​corresponding to the first two indicators, the higher the reconstruction accuracy and completeness, and the higher the last indicator, the higher the completeness. Comparing the two existing methods, it can be seen that the method proposed by the present invention is superior to the existing methods in all three indicators, indicating that the present invention has better effects.

[0057]

[0058] The present invention can be used to locate and map indoor and outdoor environments, edit reconstructed scenes, and integrate virtual and real in augmented reality. The above are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by ordinary technicians in this field should be considered to be within the scope of protection of the present invention.

Claims

1. A method for dense simultaneous localization and mapping based on voxel neural implicit surfaces, characterized in that: The following steps are involved: Step 1: Get the RGB-D image of the first frame, and back-project the depth corresponding to each pixel in the first frame image into the three-dimensional space to obtain the initial three-dimensional point cloud in the map; The coordinate system where the initial 3D point cloud is located is set as the reference coordinate system, and multiple non-overlapping voxel blocks aligned with the coordinate axes of the reference coordinate system are constructed based on the initial 3D point cloud; an octree structure is constructed based on these voxel blocks, and the Morton codes corresponding to the voxel blocks are inserted into the octree; at the same time, fixed-length feature vectors are assigned to 8 vertices of each voxel block, and these fixed-length feature vectors are used to store the geometry and texture information of the scene to be constructed; Step 2: Randomly sample M pixels from the acquired image, generate a ray starting from the camera center corresponding to the image and passing through each pixel, and calculate the intersection of the ray and the constructed voxel block; uniformly sample the area where the ray and the voxel block intersect to obtain the sampled 3D point, obtain the feature vectors of the 8 vertices of the voxel block where the 3D point is located through the 3D coordinates of the 3D point, and obtain the feature vector corresponding to the 3D point through the feature extraction function; obtain the signed distance field SDF and intermediate information through the geometric analysis network, and then obtain the color of the intermediate information through the texture analysis network; calculate the spatial density value corresponding to the 3D point through the SDF, and then perform weighted accumulation of the color and depth of the 3D point on the ray through volume rendering, and finally obtain the predicted color and depth of the pixel corresponding to the ray; compare the predicted color and depth with the real color and depth, thereby optimizing the fixed-length feature vector on the voxel block vertex and the geometric analysis network and texture analysis network; Step 3: After step 2 is completed, the tracking process is started. The tracking process is: repeat step 2 for the image obtained from the second frame, but keep the fixed-length feature vectors on the voxel block vertices and the geometric analysis network and texture analysis network unchanged, and only optimize the camera 6-DOF pose corresponding to the image. After optimization, positioning is completed, and the optimized camera 6-DOF pose and the corresponding RGB-D image are constructed into a frame and put into the candidate key frame list; Step 4: Start the mapping process. The mapping process is as follows: get the keyframe list from step 3, traverse the candidate keyframe list, and back-project the depth corresponding to the pixel point of each frame image into the three-dimensional space according to the 6-DOF pose of the camera corresponding to the image to obtain the three-dimensional point cloud corresponding to each frame; for each three-dimensional point in the three-dimensional point cloud, determine whether the three-dimensional point is included in the created voxel block. If not, create a new voxel block and update the octree structure in step 1, thereby achieving the purpose of dynamically creating voxel blocks and expanding the mapping area; Select several suitable frames from the key frame list as key frames and optimize them together with the latest frame in the candidate key frame list; repeat step 2 for the images in all frames to be optimized, and optimize the 6-DOF pose of the frame while optimizing the fixed-length feature vectors on the vertices of the voxel blocks and the geometric parsing network and the texture parsing network.

2. The method for dense synchronous localization and mapping based on voxel neural implicit surface according to claim 1, characterized in that: The step 1 of constructing a plurality of non-overlapping voxel blocks aligned with the coordinate axes of the reference coordinate system based on the initial three-dimensional point cloud is specifically as follows: The initial 3D point cloud is divided into a set of voxel blocks Divide, each voxel block has three-dimensional coordinates V k =(x, y, z); these three-dimensional coordinates are converted into 64-bit binary coded information through Morton coding; each voxel block has 8 vertices, each of which contains a fixed-length feature vector The geometry and texture information of the scene to be constructed, L e is the length of the eigenvector; therefore, for any voxel V i Any three-dimensional point in Adjacent voxel blocks share the feature vectors of 4 vertices.

3. The method for dense synchronous localization and mapping based on voxel neural implicit surface according to claim 1, characterized in that: The intersection of the calculated ray in step 2 and the voxel block constructed in step 1 is specifically: The ray passing through the pixels on the image from the camera center o along the direction d is defined as r(t)=o+dt, where t is the depth along the ray direction; each ray uses the Ray-AABB intersection detection algorithm to calculate the depth of the intersection of the ray and the voxel block, thereby dividing the area on the ray that intersects with the voxel block.

4. The method for dense synchronous localization and mapping based on voxel neural implicit surface according to claim 1, characterized in that: In the step 2, the feature vector corresponding to the three-dimensional point is obtained by the feature extraction function, the signed distance field SDF and the intermediate information are obtained by the geometric analysis network, and then the intermediate information is obtained by the texture analysis network to obtain the color, specifically: Feature extraction function Map the 3D point p to a point of length L e The eigenvector of The feature extraction function is implemented by trilinear interpolation. According to the three-dimensional coordinates of p and the relative position of p in the voxel block, the feature vectors contained in the eight vertices of the voxel block are interpolated to obtain the feature vector e of p. A multi-layer perceptron network (MLP) is used to represent the geometric parsing network F σ and texture parsing network F c ; Geometric parsing network Generate its signed distance field through p's eigenvector e and length L f Geometric eigenvectors The sign of σ indicates whether p is inside or outside the surface S; the surface S of the scene is extracted by: The operation [0] means to σ The signed distance field σ of position p is obtained from the 3D point p; the geometric feature vector f of the 3D point p, the ray direction d where p is located, and the feature vector e of p are connected as the texture parsing network F c Input, get the color c at p.

5. The method for dense synchronous localization and mapping based on voxel neural implicit surface according to claim 1, characterized in that: In step 2, the weighted accumulation of the color and depth of the three-dimensional points on the ray is performed by volume rendering, and finally the predicted color and depth of the pixel corresponding to the ray is obtained, specifically: Using the function φ s (σ) Convert the SDF of a 3D point p to a density, φ s (σ) is a function of the signed distance σ of point p, where Where τ is the Sigmoid function, tr is the predefined cutoff distance, and the value of points close to the surface is greater than the weight of distant points; Based on φ s (σ), normalize the density on the same ray, and normalize the density on the ray p 3D sampling points are used for volume rendering to obtain the accumulated color C(r) and depth D(r): where c i is the color of point i on the ray, d i is the distance from point i on the ray to the optical center.