Dense RGB-D SLAM method and device based on neural point cloud
Through the dense RGB-D SLAM method based on neural point cloud, the tracking loss and accuracy reduction problems of SLAM system in complex environments are solved, efficient tracking and precise mapping are achieved, computing and storage overhead are reduced, and robustness is improved.
Patent Information
- Application Number
- CN202510737189.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
Existing SLAM methods are prone to problems such as tracking loss and drift in complex environments, resulting in reduced accuracy and increased computational and resource requirements.
A dense RGB-D SLAM method based on neural point cloud is adopted. By acquiring RGB-D data stream, extracting the global neural point cloud, optimizing the pose information of the image acquisition device, performing key frame extraction and iterative rendering optimization, and reconstructing the scene using a multi-layer perceptron and re-rendering loss.
It improves the tracking performance of the SLAM system in dynamic and complex environments, reduces computational overhead and storage costs, and improves the accuracy and robustness of point cloud representation.
Smart Images

Figure CN120672981A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a dense RGB-DSLAM method and device based on neural point cloud. Background Art
[0002] In recent years, with the development of robotics, SLAM (Simultaneous Localization and Mapping) technology has been widely used in fields such as robotics, autonomous driving, virtual reality, and drones. In robotics, SLAM not only supports the autonomous navigation of service robots in indoor environments, but is also used by industrial robots to accurately perceive their working environments and plan operational paths. In autonomous driving, SLAM technology provides key support for vehicles to achieve high-precision positioning and dynamic environmental perception on urban roads. In virtual and augmented reality, the application of SLAM technology helps achieve real-time tracking of the real world and dynamic overlay of virtual environments, thereby enhancing the user's immersive experience. In the field of drones, SLAM technology supports autonomous flight and environmental perception in complex airspace. However, despite the significant progress of SLAM technology in various application areas, existing methods still face many challenges when dealing with complex environments.
[0003] Traditional SLAM methods typically rely on feature point matching and graph optimization techniques, which make them highly efficient and low-resource in environments with limited computational resources. However, in complex scenarios such as occlusion, rapid motion, and changing illumination, traditional methods are prone to tracking loss and drift, resulting in reduced accuracy. Furthermore, as the complexity of the environment increases, the system's computational overhead and resource requirements also increase. To address these issues, deep learning-based SLAM methods proposed in recent years can significantly improve the robustness and accuracy of the system by adopting more complex scene representations and feature extraction techniques. For example, multi-resolution hashing, octree-based scene representations, and point cloud representations can more stably cope with complex environmental changes. However, while these methods improve robustness, they still lack the ability to handle details and reconstruction accuracy. Summary of the Invention
[0004] Based on this, it is necessary to address the above technical problems and provide a dense RGB-DSLAM method and device based on neural point cloud, which can improve the tracking performance of the SLAM system in dynamic and complex environments and reduce the computational overhead.
[0005] In a first aspect, the present application provides a dense RGB-D SLAM method based on neural point cloud. The method comprises:
[0006] Acquire an RGB-D data stream including several image frames through an image acquisition device, extract the image frames, and obtain a global neural point cloud;
[0007] Randomly obtain sampling points in the current image frame, project the sampling points to the world coordinate system to obtain three-dimensional points, evenly obtain a number of sampling points on the spatial line segment formed between the origin of the image acquisition device and each three-dimensional point, combine each sampling point and a number of first neighboring neural point clouds corresponding to each sampling point to perform volume rendering and obtain a first re-rendering loss to optimize the pose information of the image acquisition device;
[0008] Key frames are extracted based on the pose information between image frames, and iterative rendering optimization is performed based on the pose information of the key frames to obtain a 3D scene reconstruction image.
[0009] In one embodiment, the first neighboring neural point cloud is obtained from a neighboring search in the local neural point cloud, and the local neural point cloud set includes the neural point cloud of the new exploration area, the neighboring neural point cloud of the feature point of the current key frame, and the supplementary neural point cloud of time correlation.
[0010] In one embodiment, the time-correlated supplementary neural point cloud is obtained by filtering from the global neural point cloud using a filtering mechanism;
[0011] Among them, the screening mechanism includes:
[0012] When the number of points in the local neural point cloud is not greater than the first threshold, the number of points discarded from the global neural point cloud in the current time step is equal to the number of points discarded in the previous time step plus the part of the number of points in the neural point cloud of the newly explored area that exceeds the second threshold;
[0013] When the number of points in the local neural point cloud is greater than the first threshold, the number of points discarded from the global neural point cloud in the current time step is equal to the number of points discarded in the previous time step plus the number of points in the neural point cloud of the newly explored area and the increase or decrease in the number of points that need to be removed from the global neural point cloud each time;
[0014] When the difference between the number of points in the global neural point cloud and the number of points discarded from the global neural point cloud at the current time step is not greater than a third threshold, the number of points discarded from the global neural point cloud is equal to the difference between the number of points in the global neural point cloud and the third threshold.
[0015] In one embodiment, the probability of each neural point in the global neural point cloud being discarded is determined by the index value of each neural point, and the screening mechanism combines the number of points discarded from the global neural point cloud and the probability of each neural point in the global neural point cloud being discarded to obtain a time-correlated supplementary neural point cloud.
[0016] In one embodiment, performing iterative rendering optimization based on the pose information of the key frames includes:
[0017] Randomly obtain sampling points in the current image frame, project the sampling points to the world coordinate system to obtain 3D points, uniformly obtain several sampling points on the spatial line segment formed between the origin of the image acquisition device and each 3D point, combine each sampling point and several second neighboring neural point clouds corresponding to each sampling point to perform volume rendering and obtain a second re-rendering loss to update the reconstruction parameters;
[0018] Among them, the second neighboring neural point cloud is obtained from the global neural point cloud through neighboring search.
[0019] In one embodiment, the proximity search adopts a hybrid radius search principle that introduces a query radius and an additional radius. The hybrid radius search principle includes:
[0020] If the neighborhood color gradient of the current feature point is not less than the maximum gradient threshold, the query radius is the minimum search radius;
[0021] If the neighborhood color extraction of the current feature point is not greater than the minimum gradient threshold, the query radius is the maximum search radius;
[0022] If the neighborhood color gradient of the current feature point is less than the maximum gradient threshold and greater than the minimum gradient threshold, the query radius changes dynamically according to the neighborhood color gradient;
[0023] Adds a fixed radius.
[0024] In a second aspect, the present application also provides a dense RGB-D SLAM device based on neural point cloud. The device includes:
[0025] A data acquisition module is used to acquire an RGB-D data stream including several image frames through an image acquisition device, extract the image frames, and obtain a global neural point cloud;
[0026] A trajectory tracking module is used to randomly obtain sampling points in the current image frame, project the sampling points into the world coordinate system to obtain three-dimensional points, uniformly obtain a number of sampling points on the spatial line segment formed between the origin of the image acquisition device and each three-dimensional point, perform volume rendering on each sampling point and a number of first neighboring neural point clouds corresponding to each sampling point, and obtain a re-rendering loss. Based on the re-rendering loss, trajectory tracking is performed to obtain the pose information of the current image frame;
[0027] The mapping module is used to extract key frames based on the pose information between each image frame, and perform iterative rendering optimization according to the pose information of the key frames to obtain a 3D scene reconstruction image.
[0028] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the neural point cloud-based dense RGB-DSLAM method when executing the computer program.
[0029] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-mentioned neural point cloud-based dense RGB-DSLAM method.
[0030] In a fifth aspect, the present application further provides a computer program product, comprising a computer program that, when executed by a processor, implements the steps of the above-mentioned neural point cloud-based dense RGB-D SLAM method.
[0031] The neural point cloud-based dense RGB-D SLAM method and device acquires an RGB-D data stream consisting of several image frames through an image acquisition device, extracts the image frames, and obtains a global neural point cloud. Random sampling points are obtained in the current image frame and projected onto the world coordinate system to obtain three-dimensional points. Several sampling points are evenly obtained on the spatial line segment formed between the image acquisition device's origin and each three-dimensional point. Volume rendering is performed on each sampling point and its corresponding first neighboring neural point clouds, and a first re-rendering loss is obtained to optimize the image acquisition device's pose information. Keyframes are extracted based on the pose information between each image frame, and iterative rendering optimization is performed based on the pose information of the keyframes to obtain a three-dimensional scene reconstruction. First, the gradient information of the RGB-D image is used to more finely sample detailed areas, achieving efficient tracking and accurate mapping in complex scenes and reducing the waste of blank space in scene modeling. Then, by optimizing the distance distribution between the point cloud set and the standardized neural point cloud, the efficiency of pose optimization is improved and storage costs are reduced. The advantage of this method is that it significantly improves the representation accuracy of point clouds while reducing computational and storage overhead, and has stronger robustness in depth map environments with noise and holes. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 1 is a flow chart of a dense RGB-D SLAM method based on neural point cloud in one embodiment;
[0033] Figure 2 A comparison diagram of the number of point clouds used for trajectory tracking before and after optimization in one embodiment;
[0034] Figure 3 A comparison diagram of the time taken for trajectory tracking before and after optimization in one embodiment;
[0035] Figure 4 A comparison chart of the trajectory tracking performance of the algorithm of the present invention with Vox-Fusion, NICE-SLAM, and Point-SLAM in one embodiment;
[0036] Figure 5 A diagram comparing the image rendering performance of the algorithm of the present invention with Vox-Fusion, NICE-SLAM, and Point-SLAM in one embodiment;
[0037] Figure 6 2 is a structural diagram of an implementation of a dense RGB-D SLAM method based on neural point cloud in one embodiment. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0039] The present application embodiment provides a dense RGB-D SLAM method based on neural point cloud, such as Figure 1 As shown, the following steps are included:
[0040] Step 102: Acquire an RGB-D data stream including several image frames through an image acquisition device.
[0041] In this embodiment, the image acquisition device can be a device with an RGB-D sensor, also known as an RGB-D camera, which can simultaneously acquire a color image (RGB) and depth information (D) of a scene. By continuously capturing the scene, a number of image frames with depth information (in the form of a depth map) are obtained to form an RGB-D data stream.
[0042] In one embodiment, after acquiring the RGB-D data stream, a guided filter is introduced to perform data preprocessing. The main function is to repair large holes in the depth map of the RGB-D data stream. The formula of the guided filter is expressed as: r (u,v)=UF r (A(u,v)*I+UF r (B(u,v))), where I r (u,v) represents the restored image after filtering; UF r is the mean filter, r is the radius, here r can be set to 5; I represents the guide image, which is the grayscale image corresponding to the RGB image in the RGB-D data stream; A(u,v) and B(u,v) are the linear coefficient and deviation of the filter respectively. B(u,v)=UF r (G r )-UF r (I)*A(u,v), where u and v are the horizontal and vertical coordinates of the pixel respectively; μ is a small constant that can be 0.01 to prevent the denominator from being 0; Gr is the image to be repaired.
[0043] Step 104: extract the image frame to obtain a global neural point cloud.
[0044] The purpose of this step is to initialize the system. The global neural point cloud will be extracted from the first image frame of the RGB-D data stream, and the position of the image acquisition device corresponding to the first frame will be fixed.
[0045] When extracting neural point clouds, the image frame is first divided into 64 small blocks, and the number of points to be extracted is obtained based on the average gradient of each small block: G wi =f sm (G mi *θ), where i=1,....,64; G mi Represents the average gradient of each small block; θ is a constant to adjust the influence of different image areas; Softmax function f sm Used to convert the average gradient of each block into probability; finally get the sampling ratio G of each small block wi To ensure that the total number of samples is an integer, use the function f ceil Round up and finally get the number of points S that need to be sampled for each small piece of the picture ni , the formula is: S ni =f ceil (G wi *S n ), where S n Is the total number of points that need to be sampled in the current frame. Project the sampled points into the three-dimensional world to obtain three-dimensional coordinates, and pass the three-dimensional coordinates into two multi-layer perceptrons F g 、F c In the 3D point, the geometric features and color features are obtained. Then the coordinates, color features, and geometric features of the 3D point are saved as a neural point, and all neural points form a neural point cloud: Where N represents the number of neural points in the neural point cloud; Represented as the three-dimensional coordinates of the current neural point in the world coordinate system; are the geometric features and color features of the neural points respectively.
[0046] In this embodiment, the multi-layer perceptron F gIt includes an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is equal to the coordinate dimension of each three-dimensional point, that is, 3 neurons, corresponding to the x, y, and z coordinates respectively; multiple hidden layers can be set, and the number of neurons in each hidden layer can be adjusted according to the specific task. For example, 3 hidden layers can be set, with the number of neurons being 64, 128, and 256 respectively. The activation function of the hidden layer can choose the ReLU function; the number of neurons in the output layer is determined by the dimension of the geometric features to be extracted. For example, if 128-dimensional geometric features need to be extracted, the output layer is set to 128 neurons. Multilayer Perceptron F c The algorithm also consists of an input layer, a hidden layer, and an output layer. The input layer has three neurons, one for each of the red, green, and blue channels. Multiple hidden layers can be used, and the number of neurons and activation function are similar to those used in MLP for geometric feature extraction. The number of neurons in the output layer is determined by the dimensionality of the color features to be extracted. For example, if a 64-dimensional color feature is to be extracted, the output layer would have 64 neurons.
[0047] In this embodiment, in each iteration, the global neural point cloud is updated by extracting feature points from a certain image frame or key frame, projecting the feature points into the world, and using r add Search for the radius in the current global neural point cloud. If the search result is 0, that is, the feature point has no neighboring point cloud, then the feature point is added to the global neural point cloud.
[0048] In step 106, sampling points are randomly obtained in the current image frame, and the sampling points are projected into the world coordinate system to obtain three-dimensional points. Several sampling points are evenly obtained on the spatial line segment formed between the origin of the image acquisition device and each three-dimensional point. Volume rendering is performed on each sampling point and several first neighboring neural point clouds corresponding to each sampling point, and a first re-rendering loss is obtained to optimize the posture information of the image acquisition device.
[0049] In this embodiment, the core goal is to optimize the posture of the image acquisition device in the entire SLAM system so that the system's estimation results are as close to the actual situation as possible. The neural radiation field represents the scene through a neural network, and its training process depends on the first re-rendering loss. Among them, re-rendering refers to regenerating simulated image frames based on the currently estimated camera posture and the constructed map. The re-rendering loss is an indicator that measures the difference between the re-rendered image and the actual captured image. Common re-rendering losses include pixel-level photometric errors (such as mean square error), feature-level errors, etc. When the posture of the image acquisition device is inaccurate, the difference between the image rendered from the neural radiation field and the actual image will increase. By minimizing the re-rendering loss, the posture of the image acquisition device can be adjusted at the same time, making the posture estimation more accurate.
[0050] In a SLAM system, the trajectory tracking of the camera is essentially to continuously estimate the camera's pose (position and attitude) in each frame. The re-rendering loss can be used as an optimization objective to continuously adjust the camera's pose estimation by minimizing the re-rendering loss so that the re-rendered image is as close to the actual captured image as possible. Specifically, given an initial camera pose estimate and a constructed map, a re-rendered image can be projected from the map based on this pose. The re-rendering loss is then calculated between the re-rendered image and the actual image. Next, an optimization algorithm (such as gradient descent, Gauss-Newton method, etc.) is used to update the camera's pose so that the re-rendering loss gradually decreases. This process is repeated until the re-rendering loss reaches a smaller value or meets certain convergence conditions.
[0051] The calculation process of the first rendering loss is as follows:
[0052] First, randomly sample the entire image and project the sampled pixels into the world coordinate system. A ray is formed between the camera origin and the sampling point in the three-dimensional point in the world coordinate system. A number of (e.g., 5) sampling points are uniformly sampled on each ray. These 5 sampling points will be in a r query For the radius, search whether there are neural points around, obtain several first neighboring neural points near the sampling point, and extract the geometric features and color features of each first neighboring neural point. Interpolation is performed through the geometric features and color features of the first neighboring neural points, and the interpolated features and the features of the first neighboring neural points are input into the multi-layer perceptron to achieve smoothing of the features of the neighboring neural points, so that the changes in the features in space are more continuous and natural, reducing the adverse effects of feature mutations on the training of the multi-layer perceptron, and enhancing the generalization ability of the multi-layer perceptron. The occupancy value and color features of the sampling point are obtained by the multi-layer perceptron, which are expressed as O ij =f g (X ij, F g (X ij )),c ij =f c (x ij ,F c (X ij )), where X ij represents the jth sampling point on the line connecting the i-th 3D point, F g (X ij ) and F c (X ij ) are respectively the geometric features and color features after encoding; ij and c ij Represent the encoded occupancy value and color value respectively. Then the distance between the sampling point and the three-dimensional point is used as the weight Perform weighted summation to obtain the depth value D of the pixel corresponding to the three-dimensional pointir and color value c ir , respectively expressed as: where a ij Indicates that the light is at point X ij The possibility of termination, p ij Indicates the sampling point X ij The depth of c ij Indicates the sampling point X ij Color value. This is the volume rendering formula, which can be understood as: sampling a point in the three-dimensional world, connecting the origin and the three-dimensional point, sampling multiple points on the line, and performing weighted summation of the samples of multiple points on the line to obtain the depth and color value of the three-dimensional point. In general, the depth color value of the three-dimensional point is obtained by weighting. The depth of the point on the line is determined at the time of sampling, and the color value is obtained by the decoder above. It means the color value and depth value of the j-th sampling point on the line of the i-th 3d point are used to render. Finally, through the corresponding depth value D ir and color value c ir , calculate the first re-rendering loss L g , the formula is: Among them, |·|1 represents the 1-norm, D i and C i Represents the real depth and color values, Dir and C respectively ir represents the rendered depth and color values; Y is a hyperparameter representing the weight of the color loss, set to 0.2 here. The re-rendering loss function here is used for parameter optimization. During the optimization, the neural network automatically adjusts the parameters to minimize the loss value.
[0053] In one embodiment, a local neural point cloud is introduced into the pose optimization process, and the first neighboring neural point cloud is obtained by screening the local neural point cloud. p It consists of three parts: p ={(P new ∪P neight ∪P time )}. Among them, P new It is a new neural point that has no neighboring points in the neural point cloud stage, representing the newly explored area in the scene. The point cloud is used to reconstruct the scene. Each round of pose optimization will perform the reconstruction operation. The newly explored area is the point cloud set newly added to the global neural point cloud in each frame after reconstruction. neight As the core of the local neural point cloud, the neighborhood of the feature points extracted from the current key frame is highly consistent with the current key frame in space. That is, assuming that the current frame is the Nth frame, the feature point p is extracted in the Nth frame and projected into the 3D world for mapping (at this time, the feature points of the previous N-1 frames have already been projected into the 3D world). With p as the center, r queryis the radius, and the neighboring point cloud is searched in the global neural point cloud, which is the neighbor of the feature point p. time To supplement the neural point cloud for temporal correlation, P time It is obtained by deleting the global neural point cloud. The global neural point cloud will continue to grow as new areas are explored. In order to ensure the stability of the system and the efficiency of optimization, a complete screening mechanism is added to screen P from the global neural point cloud. time Specifically, the screening mechanism first calculates the number of neural points that need to be discarded in this time frame:
[0054]
[0055] where N ct-1 is the number of points that need to be discarded from the global neural point cloud in the previous time step; N ct N is the number of points that need to be discarded from the global neural point cloud at the current time step; pnew is the newly added point in the global neural point cloud at the current time step; N global Refers to the total number of points in the global neural point cloud; σ represents the increase or decrease in the number of points that need to be removed from the global neural point cloud each time; N local Represents the total number of points in the local neural point cloud; is the second threshold, which is a constant set to 100; N threshold is the first threshold, the upper limit of the number of points in the local neural point cloud, set to 32000; τ is the third threshold, set to 30000; N in the third row of the formula ct To get N through the first or second row ct Entered by later generations.
[0056] The above N ct The physical meaning of the formula is as follows:
[0057] 1. The first line of the formula: when N local When the number of local point clouds does not reach the threshold, we need to discard fewer points to ensure that the local point set is growing. Suppose the global point set at the last moment had 500 points, the number of points discarded at the last moment was 300, and the number of new points added to the global point set at the current moment is 150. At the same time, the local point set is 400, and the threshold is 600, which does not reach the threshold. Then 300+50=350 points should be discarded from the global set. Constant This ensures that not all newly added points are discarded each time. max(·) is used to prevent the number of discarded points from decreasing. If this is not done, the number of local points at the previous moment will be smaller than the number of local points at the current moment, which means that the number of local points will increase.
[0058] 2. The second line of the formula: When the number of local points is greater than the threshold, more global points need to be discarded. Similarly, the local point set at the previous moment will be larger than the current local point set.
[0059] 3. The third line of the formula: Ensure that the number of local points is always lower than a certain number of the global point set to ensure that too many local points are not retained.
[0060] At the same time, the probability Prob of each neural point being discarded is calculated according to the time order in which the neural points are added ire , the formula is: where p idx is the index of the neural point, the earlier the neural point p is added idx The smaller it is; ε is the adjustment of time weight, the earlier the neural point is added, the easier it is to be discarded. ire , discard N in its global neural point cloud ct neural points, and finally the set of retained neural points P is obtained time .
[0061] The number of optimized neural points is as follows Figure 2 As shown in the figure, the time taken for trajectory tracking after optimization is as follows: Figure 3 As shown. Figure 2 and Figure 3 It can be seen that in the tracking step, the number of neural point clouds involved in this embodiment is far lower than that of the traditional Point SLAM algorithm. Accordingly, the tracking time used in this embodiment is significantly lower than that of the traditional Point SLAM algorithm, which effectively improves the tracking efficiency and reduces the computational overhead.
[0062] The camera pose obtained in step 104 may contain errors, and pose optimization is to correct these errors and improve the accuracy of pose estimation.
[0063] Step 108 : extract key frames based on the pose information between the image frames, perform iterative rendering optimization based on the pose information of the key frames, and obtain a 3D scene reconstruction image.
[0064] During the tracking process, in order to avoid excessive redundant calculations, key frames need to be extracted from continuous image sequences. There are many bases for selecting key frames, such as changes in the camera's posture. When the rotation or translation of the camera exceeds a certain threshold, it means that the camera's viewing angle has changed significantly. At this time, the current frame can be used as a key frame. In other embodiments, it can also be judged based on the feature information of the image. If the number of feature matches between the current frame and the most recent key frame is lower than a certain threshold, it means that the current frame contains new scene information and can be extracted as a key frame. Key frames are like important landmarks in a map, recording key information of the scene and providing a basis for subsequent processing.
[0065] Based on the optimized camera pose and keyframe image data, the scene is reconstructed in three dimensions. Combined with Nerf technology, a neural radiance field is constructed to represent the scene. The neural radiance field is a multi-layer perceptron that receives a point in three-dimensional space and the viewing direction as input and outputs the color and density of that point. Feature information is extracted from keyframe images from different perspectives and combined with the camera pose to initialize the neural radiance field. The depth and color values of the scene are then estimated using a large number of sampling points. The neural radiance field is trained using this sampling data, and its parameters are continuously adjusted so that it can accurately represent the appearance and geometry of the scene. It's like building a three-dimensional scene model with building blocks.
[0066] This process is roughly the same as that of step 106, that is, first randomly sample the entire image and project the sampled pixels into the world coordinate system. A ray is formed between the camera origin and the sampling point in the three-dimensional point in the world coordinate system, and a number of (e.g., 5) sampling points are uniformly sampled on each ray. These 5 sampling points will be r query Take the radius as the radius, search whether there are neural points around, obtain several second neighboring neural points near the sampling point, and extract the geometric features and color features of each first neighboring neural point. Interpolate the geometric features and color features of the second neighboring neural points, and input the interpolated features and the features of the second neighboring neural points into the multi-layer perceptron to smooth the features of the neighboring neural points, so that the spatial changes of the features are more continuous and natural, reduce the adverse effects of feature mutations on the training of the multi-layer perceptron, and enhance the generalization ability of the multi-layer perceptron. The geometric features and color features of the sampling point are obtained through the multi-layer perceptron, and the geometric features and color features are encoded respectively to obtain the depth value and color value. Then, the distance from the sampling point to the three-dimensional point is used as the weight, and a weighted sum is performed to obtain the depth value and color value of the pixel point corresponding to the three-dimensional point. Finally, the second rendering loss is calculated through the corresponding depth value and color value.
[0067] The difference is that the neighbor search for the second neighboring neural point is performed in the global neural point cloud.
[0068] Step 108 provides a local neural point cloud for step 106 for trajectory tracking; step 106 provides optimized pose information for step 108; the loop is iterated until the expectation is met, which can be when the number of iterations reaches a threshold or when the second rendering loss reaches the expected effect.
[0069] like Figure 4 and 5 As shown, the rendering results of the solution provided by the invention are compared with those of the traditional solution. It can be seen that the rendering details and hole filling capabilities of the present invention are significantly better.
[0070] At the same time, unlike the traditional algorithm that uses dynamic radius in the whole system, this embodiment adopts hybrid radius search to minimize the burden of system calculation. Specifically, when adding neural points to the global neural point cloud in step 104, a fixed radius is used for addition, that is, a radius of r is added. add (u, v), and in the subsequent step 106, a dynamic query radius is used to search the neighboring point cloud, and the formula is as follows:
[0071]
[0072] r add (u,v)=0.08
[0073] Among them, G(u,v) represents the neighborhood color gradient of the current feature point (pixel point on the key frame), r query (u,v) and r add (u,v) respectively query radius and add radius, g max and g min They represent the maximum gradient threshold and minimum gradient threshold of the feature points, respectively, between 0 and 1.5, β and γ are mapping parameters automatically generated by the interpolation function interpld, r min and r max are preset fixed values, representing the minimum search radius and the maximum search radius respectively. query Represents the length of the search radius, which changes dynamically according to the size of the color gradient and is kept between 0.11 and 0.9.
[0074] In one embodiment, the algorithm implementation involved in the neural point cloud-based dense RGB-D SLAM method uses the Visual Studio Code software development tool, and the operating environment is Ubuntu 20.4.
[0075] The present invention first uses the gradient information of the RGB image to perform finer sampling of detail areas, achieving efficient tracking and precise mapping in complex scenes, and reducing the waste of blank space in scene modeling. Then, by optimizing the distance distribution of the point cloud set and the standardized neural point cloud, the efficiency of posture optimization is improved and the storage cost is reduced. At the same time, in order to solve the noise and hole problems in the depth map, an image inpainting method based on guided filtering is adopted to further improve the robustness of the system. The advantage of this method is that while reducing the computational and storage overhead, it significantly improves the representation accuracy of the point cloud and has stronger robustness in depth map environments with noise and holes.
[0076] It should be noted that the values involved in this application are all empirical values obtained through experiments.
[0077] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0078] Based on the same inventive concept, the embodiments of the present application also provide a neural point cloud-based dense RGB-D SLAM device for implementing the above-mentioned neural point cloud-based dense RGB-D SLAM method. The implementation solution provided by the device is similar to the implementation solution described in the above-mentioned method. Therefore, the specific limitations in one or more neural point cloud-based dense RGB-D SLAM device embodiments provided below can be found in the above-mentioned limitations on the neural point cloud-based dense RGB-D SLAM method, which will not be repeated here.
[0079] In one embodiment, Figure 6 As shown, a dense RGB-DSLAM device based on neural point cloud is provided, including: a data acquisition module, a preprocessing module, a trajectory tracking module, and a mapping module, wherein:
[0080] The data acquisition module uses an RGB-D sensor to collect data about the environment and obtains the depth map and RGB image of the environment through the RGB-D camera.
[0081] The preprocessing module is used to preprocess the collected data. When there are large noises and holes in the depth map of the collected data, guided filter alignment will be introduced to repair them.
[0082] The data acquisition module is also used to initialize the system using the data from the first frame. The neural point cloud is collected on the first frame and used to initialize the entire 3D scene. The pose of the first frame is fixed to prevent drift in the null space.
[0083] Subsequent data streams are first passed to the trajectory tracking module, where the re-rendering loss is used to obtain the optimized pose information for the current frame. The trajectory tracking module randomly samples points in the current image frame, projects these points into the world coordinate system to obtain 3D points, and evenly samples several points along the spatial line segment formed between the image acquisition device's origin and each 3D point. Volume rendering is performed on each sampling point and its corresponding first neighboring neural point cloud, and the first re-rendering loss is obtained to optimize the pose information of the image acquisition device.
[0084] The pose information and keyframes are passed to the mapping module, where a neural point cloud is extracted from the keyframes to represent the newly explored 3D scene. A second rendering loss is used to optimize the color and geometric features of the neural point cloud. The rendering optimization is iteratively performed to obtain a 3D scene reconstruction.
[0085] Each module in the above-mentioned neural point cloud-based dense RGB-D SLAM device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.
[0086] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in all the above method embodiments when executing the computer program.
[0087] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in all the above method embodiments are implemented.
[0088] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in all the above method embodiments when executed by a processor.
[0089] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0090] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, etc., but are not limited to these.
[0091] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0092] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A dense RGB-D SLAM method based on neural point cloud, characterized in that The method comprises: Acquire an RGB-D data stream including a plurality of image frames through an image acquisition device, extract the image frames, and obtain a global neural point cloud; Randomly acquiring sampling points in the current image frame, projecting the sampling points into a world coordinate system to acquire three-dimensional points, uniformly acquiring a number of sampling points on a spatial line segment formed between an origin of the image acquisition device and each of the three-dimensional points, performing volume rendering on each of the sampling points and a number of first neighboring neural point clouds corresponding to each of the sampling points, and acquiring a first re-rendering loss to optimize the pose information of the image acquisition device; Key frames are extracted based on the posture information between the image frames, and iterative rendering optimization is performed according to the posture information of the key frames to obtain a three-dimensional scene reconstruction image.
2. The method according to claim 1, characterized in that The first neighboring neural point cloud is obtained from a neighboring search in a local neural point cloud, and the local neural point cloud set includes a neural point cloud in a newly explored area, a neighboring neural point cloud of a feature point of the current key frame, and a supplementary neural point cloud of time correlation.
3. The method according to claim 2, characterized in that The time-correlated supplementary neural point cloud is obtained by screening from the global neural point cloud using a screening mechanism; The screening mechanism includes: When the number of points in the local neural point cloud is not greater than a first threshold, the number of points discarded from the global neural point cloud in the current time step is equal to the number of points discarded in the previous time step plus the portion of the number of points in the neural point cloud in the new exploration area that exceeds a second threshold; When the number of points in the local neural point cloud is greater than the first threshold, the number of points discarded from the global neural point cloud in the current time step is equal to the number of points discarded in the previous time step plus the number of points in the neural point cloud of the new exploration area and the increase or decrease in the number of points to be discarded from the global neural point cloud each time; When the difference between the number of points in the global neural point cloud and the number of points discarded from the global neural point cloud at the current time step is not greater than a third threshold, the number of points discarded from the global neural point cloud is equal to the difference between the number of points in the global neural point cloud and the third threshold.
4. The method according to claim 3, characterized in that The probability of each neural point in the global neural point cloud being discarded is determined by the index value of each neural point. The screening mechanism combines the number of points discarded from the global neural point cloud and the probability of each neural point in the global neural point cloud being discarded to obtain the time-correlated supplementary neural point cloud.
5. The method according to claim 1, wherein The iterative rendering optimization according to the pose information of the key frame includes: randomly acquiring sampling points in the current image frame, projecting the sampling points into a world coordinate system to acquire three-dimensional points, uniformly acquiring a number of sampling points on a spatial line segment formed between an origin of the image acquisition device and each of the three-dimensional points, performing volume rendering on each of the sampling points and a number of second neighboring neural point clouds corresponding to each of the sampling points, and acquiring a second re-rendering loss to update reconstruction parameters; The second neighboring neural point cloud is obtained by neighboring search in the global neural point cloud.
6. The method according to claim 1, characterized in that The method further includes adopting a hybrid radius search principle that introduces a query radius and an additive radius, wherein the hybrid radius search principle includes: If the neighborhood color gradient of the current feature point is not less than the maximum gradient threshold, the query radius is the minimum search radius; If the neighborhood color extraction of the current feature point is not greater than the minimum gradient threshold, the query radius is the maximum search radius; If the neighborhood color gradient of the current feature point is less than the maximum gradient threshold and greater than the minimum gradient threshold, the query radius changes dynamically according to the neighborhood color gradient; The added radius is a fixed value.
7. A dense RGB-D SLAM device based on neural point cloud, characterized in that The device comprises: A data acquisition module is used to acquire an RGB-D data stream including a plurality of image frames through an image acquisition device, extract the image frames, and obtain a global neural point cloud; a trajectory tracking module, configured to randomly acquire sampling points in the current image frame, project the sampling points into a world coordinate system to acquire three-dimensional points, uniformly acquire a number of sampling points on a spatial line segment formed between an origin of the image acquisition device and each of the three-dimensional points, perform volume rendering on each of the sampling points and a number of first neighboring neural point clouds corresponding to each of the sampling points, and acquire a re-rendering loss; perform trajectory tracking based on the re-rendering loss to acquire pose information of the current image frame; A mapping module is used to extract key frames based on the posture information between each of the image frames, and perform iterative rendering optimization according to the posture information of the key frames to obtain a three-dimensional scene reconstruction image.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Robot operation method and system based on task perception virtual visual angle re-rendering
CN121946544A
A robot operation method and system based on task-aware virtual perspective re-rendering
CN121946544B