Semantic vision SLAM (Simultaneous Localization and Mapping) method and system for indoor low-texture and dynamic environment
Through the semantic visual SLAM method combined with deep learning and traditional visual SLAM framework, the problems of low positioning accuracy and poor map quality in visual SLAM system in low texture and dynamic environments are solved, and high-precision positioning and high-quality map construction are achieved, which improves the robustness of the system and map readability.
Patent Information
- Application Number
- CN202510507931.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
AI Technical Summary
The existing visual SLAM system is difficult to effectively extract feature points in low indoor texture and dynamic environments, resulting in low positioning accuracy, poor robustness, and the generated map lacks semantic information and poor quality.
The semantic visual SLAM method is adopted, combined with deep learning and traditional visual SLAM framework, and by obtaining image texture degree and object semantic information, dynamic point culling, loopback detection and other algorithms, the removal of dynamic feature points and semantic map construction are achieved, and the feature point extraction accuracy and map quality are improved.
Achieve high-precision positioning and high-quality map construction in low texture and dynamic environments, significantly improving feature point extraction and matching accuracy, eliminating dynamic object interference, improving system robustness and map readability.
Smart Images

Figure CN120333414A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual processing, and particularly to a semantic visual SLAM method and system for indoor low-texture and dynamic environments. Background Art
[0002] Currently, visual SLAM (Simultaneous Localization and Mapping) technology has been widely applied to robot autonomous positioning and environmental mapping. However, in indoor low-texture and dynamic environments, existing visual SLAM systems still face many challenges. Taking the currently widely used algorithm ORB-SLAM2 as an example, the specific problems include:
[0003] 1. Feature point extraction problem in low-texture environments: Existing SLAM systems often cannot extract a sufficient number of feature points in low-texture environments, resulting in low positioning accuracy and poor system stability.
[0004] 2. Influence of dynamic objects: Dynamic objects (such as people, pets, etc.) often exist in indoor environments. The movement of these objects will cause a significant decrease in the positioning accuracy of the SLAM system and even lead to system crashes. Traditional loop closure detection algorithms also have difficulty dealing with image mismatches in dynamic scenes.
[0005] 3. Quality problem of map construction: The maps generated by traditional SLAM systems lack semantic information of objects in the environment, and the constructed maps have redundancy and overlap, affecting the readability and practicality of the maps. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a semantic visual SLAM method and system for indoor low-texture and dynamic environments, and this method can process data in indoor low-texture and dynamic scenes in real time.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] The semantic visual SLAM method for indoor low-texture and dynamic environments provided by the present invention includes the following steps:
[0009] 1) Obtain an image, and obtain dynamic feature points of an object according to the texture degree of the image and the semantic information of the object;
[0010] 2) Dynamic point elimination: Eliminate the dynamic feature points of the object through motion consistency checking and in combination with semantic information;
[0011] 3) Front-end tracking: After removing dynamic feature points, it enters the tracking thread for camera pose estimation. For the images entering the tracking thread, first match the static feature points with adjacent key frames to calculate the rough pose of the camera, then track the local map for pose optimization, and finally decide whether to insert a new key frame into the local mapping thread;
[0012] 4) Local mapping: After the key frame enters the local mapping thread, first remove the redundant map points of the key frame, then create new map points to restore the co-visible points, and further optimize the spatial points and camera poses using local BA to obtain the accurate camera pose. Finally, remove the redundant key frames and enter the loop detection thread;
[0013] 5) The loop detection thread calculates the similarity score for the historical key frames of the current key frame through semantic template matching to determine whether a loop is formed, corrects the pose of the images forming the loop to eliminate the cumulative drift error of the system, and finally performs global BA optimization to update the environment map to obtain the trajectory of the system's camera.
[0014] Further, in step 1), after obtaining the image, it is processed by an improved visual odometry. The specific steps are as follows:
[0015] 1) Obtain the current nth frame of RGB image;
[0016] 2) Feature point extraction: Calculate the texture degree of the current image. If it is judged as a low-texture image, extract GCNv2 feature points; if it is judged as a high-texture image, extract ORB feature points;
[0017] 3) Tracking: First judge whether the current constant velocity model is empty. If it is empty, perform feature point matching and tracking between the reference key frame and the current frame; if it is not empty, use the constant velocity motion model to perform feature point matching and tracking between the previous frame and the current frame;
[0018] 4) Tracking the local map: According to the above tracking results, optimize the pose of the current frame by minimizing the reprojection error. If the pose calculation is successful, add the pose information of the current frame to the local map of the pose trajectory;
[0019] 5) Select key frames: Judge whether the key selection conditions are met. If they are met, add the current frame to the key frame queue; if not, do not select it as a key frame.
[0020] Further, the method for removing dynamic points in step 2) is as follows:
[0021] First, obtain the object semantic information from the input image through instance segmentation. Meanwhile, extract the image features and perform a motion consistency check. Combine the results of both to detect dynamic feature points. Finally, remove the dynamic feature points and calculate the camera pose using the remaining static feature points.
[0022] Furthermore, the loop detection algorithm in step 3) specifically includes the following steps:
[0023] First, extract the semantic template of the image. Then, combine the geometric and semantic information to match and filter the semantic template, thereby calculating the image similarity score, and finally complete the loop closure detection.
[0024] Furthermore, the map in step 5) is a dense point cloud semantic map, which specifically includes the following steps:
[0025] 1) Process the image through the front-end tracking thread of visual SLAM to obtain the camera pose corresponding to each frame of the image. Then, judge whether the current frame is a key frame according to the PKS key frame selection strategy. Meanwhile, for the input image, obtain the semantic information of the image using the instance segmentation network for subsequent mapping the semantic information to the point cloud.
[0026] 2) Detect the key frames according to the dynamic point detection method based on semantic information. For the pixel points of the dynamic objects in the detected key frames, construct the point cloud.
[0027] 3) Then, based on the point cloud construction method and the PCL library, project the coordinate values of the static pixel points in the key frames and the different colors corresponding to the object semantic information detected by the semantic thread into the 3D space coordinate system to complete the point cloud construction.
[0028] 4) Perform statistical filtering on the outlier points in the point cloud constructed for the current frame. Subsequently, splice the remaining point cloud of the current frame and the pose information of the current frame into the already constructed point cloud map.
[0029] 5) Finally, obtain the global dense point cloud semantic map of the system.
[0030] Furthermore, the semantic map is constructed using an octree structure, and the dense point cloud is converted into an octree map through the octomap library.
[0031] Furthermore, the thread for obtaining the semantic information and the front-end tracking thread are executed in parallel; or the local mapping thread and the loop detection thread are executed in parallel.
[0032] Furthermore, the segmentation, feature point extraction, and dynamic point removal of the image are processed for the same image. After the motion consistency check of the image, wait for the processing result of the semantic information acquisition thread, and then remove the dynamic feature points in combination with the semantic information.
[0033] A semantic visual SLAM system for indoor low-texture and dynamic environments provided by the present invention includes
[0034] an image input module for respectively inputting the acquired images into a feature point extraction module and a semantic information acquisition module;
[0035] a specific point extraction module for respectively calculating ORB feature points and GCNv2 feature points according to the image texture degree;
[0036] a semantic information acquisition module for processing with Yolov8-Seg instance segmentation to obtain semantic information;
[0037] a dynamic point elimination module for eliminating dynamic feature points; eliminating the dynamic feature points of objects through motion consistency checking and combining semantic information;
[0038] a front-end tracking module, after eliminating dynamic feature points, enters a tracking thread for camera pose estimation; for the images entering the tracking thread, first matches static feature points with adjacent key frames, calculates the rough pose of the camera, and tracks the local map for pose optimization, and finally decides whether to insert a new key frame into the local mapping thread;
[0039] a local mapping module, after the key frame enters the local mapping thread, first eliminates redundant map points of the key frame, then creates new map points for restoring coplanar points, and further optimizes the spatial points and camera poses using local BA to obtain accurate camera poses, and finally eliminates redundant key frames and enters the loop detection thread;
[0040] a loop detection module, the loop detection thread calculates the similarity score for the historical key frames of the current key frame through semantic template matching, thereby determining whether a loop is formed, and corrects the pose of the images forming the loop to eliminate the cumulative drift error of the system, and finally performs global BA optimization to update the environment map to obtain the trajectory of the camera of the system.
[0041] Furthermore, it further includes an improved visual odometry calculation module for processing according to the following steps by an improved visual odometry calculation method after acquiring the images:
[0042] 1) Obtain the current nth frame RGB image;
[0043] 2) Feature point extraction, calculate the texture degree of the current image, if it is judged as a low-texture image, extract GCNv2 feature points; if it is judged as a high-texture image, extract ORB feature points;
[0044] 3) Tracking. First, it is judged whether the current constant velocity model is empty. If it is empty, the feature point matching and tracking between the reference key frame and the current frame are performed. If it is not empty, the feature point matching and tracking between the previous frame and the current frame are performed using the constant velocity motion model.
[0045] 4) Tracking the local map. According to the above tracking results, the pose of the current frame is optimized by minimizing the reprojection error. If the pose calculation is successful, the pose information of the current frame is added to the local map of the pose trajectory.
[0046] 5) Selecting key frames. It is judged whether the key selection condition is satisfied. If it is satisfied, the current frame is added to the key frame queue. If it is not satisfied, it is not selected as a key frame.
[0047] The beneficial effects of the present invention are as follows:
[0048] A semantic visual SLAM method and system for indoor low-texture and dynamic environments provided by the present invention integrate algorithms such as visual odometry, dynamic point removal, and loop detection into a complete visual SLAM system, realizing real-time processing of data in indoor low-texture and dynamic scenarios, and providing high-precision positioning and high-quality map construction. The entire system adds a semantic information acquisition thread and a dynamic point removal module, and improves the loop detection module. Among them, the semantic information acquisition thread runs in parallel with the front-end tracking thread, the local mapping thread, and the loop detection thread. Due to the different running speeds of the semantic thread and the dynamic point removal module, the running time of instance segmentation is usually longer than that of feature point extraction and the dynamic point removal module. To ensure that the same image is processed between the two, the dynamic point removal module will wait for the processing result of the semantic thread after performing the motion consistency check, and then combine the semantic information to remove dynamic feature points.
[0049] This method solves the problems of poor positioning accuracy, low robustness, and poor map quality in visual SLAM systems in low-texture environments and dynamic environments in the prior art. By combining deep learning and traditional visual SLAM frameworks, a visual SLAM method and system suitable for indoor low-texture and dynamic environments are constructed, improving the feature point extraction and matching accuracy in low-texture environments; removing the interference of dynamic objects to the SLAM system, improving the positioning accuracy and loop detection performance; enhancing the map construction quality, and constructing a highly readable semantic map by integrating semantic information.
[0050] This method uses the GCNv2 feature extraction network and the texture degree calculation method, significantly improving the accuracy of feature point extraction in low-texture environments. The dynamic point elimination algorithm that combines YOLOv8-Seg and geometric information effectively removes the interference of dynamic objects on the SLAM system, enhancing the system's robustness. By semantic template matching, the accuracy and efficiency of loop detection are improved, with the loop detection performance increased by more than 90% compared to traditional methods. Combining dynamic object elimination and semantic information to construct a dense point cloud and an octree semantic map enhances the readability and storage efficiency of the map, making it suitable for navigation and path planning tasks in practical applications.
[0051] Other advantages, objectives, and features of the present invention will, to some extent, be described in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the examination and research of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings
[0052] To make the objectives, technical solutions, and beneficial effects of the present invention clearer, the present invention provides the following drawings for illustration.
[0053] Figure 1 It is the improved visual SLAM algorithm framework.
[0054] Figure 2 It is the improved visual odometry algorithm framework.
[0055] Figure 3 It is the dynamic feature point elimination method.
[0056] Figure 4 It is the loop detection algorithm framework based on semantic template matching.
[0057] Figure 5 It is the dense point cloud semantic map construction process.
[0058] Figure 6 It is the operation of different algorithms in low-texture scenarios.
[0059] Figure 7 It is the comparison of feature extraction in low-texture sequences.
[0060] Figure 8 It is the result of dynamic feature point elimination in low-dynamic sequences.
[0061] Figure 9 It is the result of dynamic feature point elimination in high-dynamic sequences.
[0062] Figure 10 It is the true trajectory map of the two sequences.
[0063] Figure 11 It is a test image for the fr2_desk sequence.
[0064] Figure 12 It is a test image for the fr2_desk sequence.
[0065] Figure 13 It is a comparison of the absolute trajectory error for the fr3_s_sta sequence.
[0066] Figure 14 It is a comparison of the relative pose error for the fr3_s_sta sequence.
[0067] Figure 15 It is a comparison of the absolute trajectory error for the fr3_w_xyz sequence.
[0068] Figure 16 It is a comparison of the absolute trajectory error for the fr3_w_sta sequence.
[0069] Figure 17 It is a comparison of the absolute trajectory error for the fr3_w_hal sequence.
[0070] Figure 18 It is a comparison of the absolute trajectory error for the fr3_w_rpy sequence.
[0071] Figure 19 It is a comparison of the relative pose error for the fr3_w_xyz sequence.
[0072] Figure 20 It is a comparison of the relative pose error for the fr3_w_sta sequence.
[0073] Figure 21 It is a comparison of the relative pose error for the fr3_w_hal sequence.
[0074] Figure 22 Comparison of the relative pose error for the fr3_w_rpy sequence.
[0075] Figure 23 It is for the construction of the ORB-SLAM2 dense point cloud map.
[0076] Figure 24 It is for the construction of the dense point cloud map by the algorithm of this module.
[0077] Figure 25 It is for the construction of the ORB-SLAM2 octree map.
[0078] Figure 26 It is for the construction of the octree map by the algorithm of this section. Specific implementation manners
[0079] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.
[0080] Embodiment 1
[0081] As Figure 1 shown, the semantic visual SLAM method and system for indoor low-texture and dynamic environments provided in this embodiment have the following specific steps for the entire process:
[0082] 1) Obtain an image, and obtain the dynamic feature points of the object according to the image texture degree and object semantic information; input the RGB image obtained by the camera sensor into the system, and enter the feature point extraction module and the semantic information acquisition thread respectively. In the feature point extraction module, first calculate the image texture degree to determine the type of feature points to be extracted;
[0083] 2) Dynamic point elimination: Eliminate the object dynamic feature points through motion consistency checking and in combination with semantic information; after the feature points are extracted from the image, enter the dynamic point elimination module. First, perform motion consistency checking, judge the motion attributes of the feature points in the image through the epipolar geometry constraint, and wait for the semantic thread to obtain the semantic information of the image. Finally, complete the elimination of the dynamic feature points in combination with the semantic information;
[0084] 3) Front-end tracking. After eliminating the dynamic feature points, enter the tracking thread for camera pose estimation; for the image entering the tracking thread, first match the static feature points with the adjacent key frames, calculate the rough pose of the camera, and track the local map for pose optimization. Finally, decide whether to insert a new key frame into the local mapping thread;
[0085] 4) Local mapping. After the key frame enters the local mapping thread, first eliminate the redundant map points of the key frame, then create new map points to restore the co-visible points, and use local BA to further optimize the spatial points and camera pose to obtain the accurate camera pose. Finally, eliminate the redundant key frames and enter the loop detection thread;
[0086] 5) The loop detection thread calculates the similarity score for the historical key frames of the current key frame through semantic template matching to determine whether a loop is formed, correct the pose of the image that forms a loop to eliminate the cumulative drift error of the system, and finally perform global BA optimization to update the environment map to obtain the trajectory of the camera of the system.
[0087] The semantic visual SLAM method provided in this embodiment is an improved visual SLAM algorithm. In this algorithm, improved visual odometry, dynamic point removal, loop detection and other algorithms are integrated, so that the above algorithms are integrated into a complete visual SLAM system, which can process data in indoor low-texture and dynamic scenes in real time, and provide high-precision positioning and high-quality map construction. The whole method provided in this embodiment adds a semantic information acquisition thread and a dynamic point removal module, and improves the loop detection module. Among them, the semantic information acquisition thread runs in parallel with the front-end tracking thread, the local mapping thread and the loop detection thread. Since the running speeds of the semantic thread and the dynamic point removal module are different, the running time of instance segmentation is usually longer than that of feature point extraction and the dynamic point removal module. To ensure that the same image is processed between the two, the dynamic point removal module will wait for the processing result of the semantic thread after performing the motion consistency check, and then combine the semantic information to remove the dynamic feature points.
[0088] Semantic label: It refers to the semantic information of each object or region in the image, such as "chair", "table", etc. These labels help the algorithm understand the semantic content of the environment, thereby improving the accuracy of map construction and positioning.
[0089] Implementation process:
[0090] 1. Semantic information acquisition: Extract semantic information from RGB images through methods such as YOLOv8-Seg instance segmentation.
[0091] 2. Semantic label: Convert the extracted semantic information into labels for subsequent semantic template extraction and matching.
[0092] ORB feature points: ORB (Oriented FAST and Rotated BRIEF) is a fast and robust feature point detection and descriptor extraction method, suitable for real-time applications.
[0093] GCNv2 feature points: GCNv2 (Graph Convolutional Networks version 2) is a feature extraction method based on graph convolutional networks, which can capture complex structural information in images.
[0094] GMS feature points (Grid-based Motion Statistics): Grid-based motion statistics feature points, which extract feature points by analyzing the motion statistical information of grid regions in the image, and are suitable for feature point extraction in dynamic environments; GMS feature points represent Grid-based Motion Statistics, that is, grid-based motion statistics.
[0095] Semantic template extraction: Used to extract semantic templates from a known environment, which contain semantic and geometric information of objects in the environment. Functional role: Helps the algorithm identify and match known objects in a new environment, improving the accuracy of localization and map construction.
[0096] Semantic template matching: Used to identify and match known semantic templates in a new environment. Functional role: By matching semantic templates, the algorithm can identify objects in the environment, thereby improving the accuracy of localization and the robustness of map construction.
[0097] Image similarity calculation: Used to calculate the similarity between two images. Functional role: Helps the algorithm identify repetitive structures in the environment for loop closure detection and correction.
[0098] Update global map: Used to update and maintain global map information. Functional role: Ensures the accuracy and consistency of map information, providing support for localization and navigation.
[0099] Global BA optimization: Global Bundle Adjustment is used to optimize all points and camera poses in the global map. Functional role: Improves the accuracy and consistency of the map, reducing cumulative errors.
[0100] Loop closure correction: Used to detect and correct loop closure errors. Functional role: By identifying and correcting loop closure errors, improves the accuracy of localization and the robustness of map construction.
[0101] Loop closure detection: Used to detect loop closures in the environment. Functional role: By detecting loop closures, the algorithm can identify repetitive structures in the environment for loop closure correction.
[0102] Key frame insertion: Used to insert key frames at key positions. Functional role: By inserting key frames, the algorithm can better capture changes in the environment, improving the accuracy of localization.
[0103] Delete redundant map points: Used to delete redundant points in the map. Functional role: Reduces the complexity of the map and improves the running efficiency of the algorithm.
[0104] Create new map points: Used to create new map points in a new environment. Functional role: Expands the coverage of the map and improves the integrity of the map.
[0105] Local BA optimization: Local Bundle Adjustment is used to optimize points and camera poses in the local map. Functional role: Improves the accuracy and consistency of the local map, reducing local errors.
[0106] Delete redundant key frames: Used to delete redundant key frames. Functional role: Reduces the number of key frames and improves the running efficiency of the algorithm.
[0107] As Figure 2 shown Figure 2 To improve the visual odometry calculation method framework, in this embodiment, after obtaining the image, it is processed by improving the visual odometer. The specific steps are as follows:
[0108] 1) Obtain the current nth frame of RGB image;
[0109] 2) Feature point extraction. Calculate the texture degree of the current image. If it is determined to be a low-texture image, extract GCNv2 feature points; if it is determined to be a high-texture image, extract ORB feature points;
[0110] 3) Tracking. First, determine whether the current constant velocity model is empty. If it is empty, perform feature point matching and tracking between the reference key frame and the current frame; if it is not empty, use the constant velocity motion model to perform feature point matching and tracking between the previous frame and the current frame;
[0111] 4) Track the local map. According to the above tracking results, optimize the pose of the current frame by minimizing the reprojection error. If the pose calculation is successful, add the pose information of the current frame to the local map of the pose trajectory;
[0112] 5) Select key frames. Determine whether the key selection condition is satisfied. If it is satisfied, add the current frame to the key frame queue; if it is not satisfied, do not select it as a key frame.
[0113] The visual odometer method provided in this embodiment is integrated into the ORB-SLAM2 framework to improve the robustness of the system in a low-texture environment.
[0114] As Figure 3 shown Figure 3 For the dynamic feature point elimination method, the dynamic feature point elimination method provided in this embodiment. In a dynamic environment, the movement of dynamic objects will interfere with the positioning accuracy of the SLAM system. This method is a dynamic point elimination method based on semantic information. The specific steps are as follows:
[0115] First, obtain the object semantic information of the input image through instance segmentation, and at the same time extract the image features and perform motion consistency verification. Combine the two results to detect dynamic feature points. Finally, eliminate the dynamic feature points and calculate the camera pose using the remaining static feature points.
[0116] As Figure 4 shown Figure 4 For the loop closure detection algorithm framework based on semantic template matching, the loop closure detection algorithm provided in this embodiment improves the accuracy of the system loop closure detection. The specific steps are as follows:
[0117] First, extract the semantic template of the image, and then combine geometric and semantic information to match and filter the semantic template, so as to calculate the image similarity score, and finally complete the closed-loop detection.
[0118] The method provided in this embodiment combines a visual SLAM system of deep learning and the ORB-SLAM2 framework, which overcomes the problem of feature point extraction in low-texture environments.
[0119] As Figure 5 shown, Figure 5 For the construction process of the dense point cloud semantic map, which improves the map quality generated by the SLAM system, the method provided in this embodiment can be used for the construction of static maps in indoor environments, and convert semantic information into three-dimensional point clouds. By combining the photogrammetry key-frame selection strategy and the dynamic object removal method, a dense point cloud semantic map is constructed, and the semantic information of the object is mapped into the three-dimensional point cloud. The specific steps are as follows:
[0120] 1) Process the image through the front-end tracking thread of visual SLAM to obtain the camera pose corresponding to each frame of the image. Then, according to the PKS key-frame selection strategy, judge whether the current frame is a key frame. At the same time, for the input image, use the instance segmentation network to obtain the semantic information of the image for subsequent mapping of semantic information into the point cloud;
[0121] The PKS key-frame selection method in this embodiment mainly considers two geometric conditions. One is the appropriate angular distance between potential key frames to provide good spatial intersection, which is determined by calculating the angle between the vector from the camera to the point and the surface normal of the 3D point. The other is the good distribution of valid points within the potential key frame. A 3×3 grid is formed in each frame, and the number of valid points in each cell is counted, and the judgment is made according to the center-of-gravity balance criterion. In addition, if there is an inertial measurement unit (IMU) sensor in the video acquisition system, when the acceleration change observed by the IMU exceeds the threshold, a new key frame will also be selected. The PKS method ensures the robustness of the algorithm, generates a more complete and accurate point cloud, and improves the positioning accuracy and point cloud quality.
[0122] 2) Detect the key frames according to the dynamic point detection method based on semantic information. For the dynamic objects in the detected key frames, such as people, cats, dogs, etc., skip the pixel points on them for point cloud construction to ensure the stability and accuracy of the point cloud;
[0123] 3) Then, based on the point cloud construction method and the PCL library, project the coordinate values of the static pixel points in the key frame and the different colors corresponding to the object semantic information detected by the semantic thread into the 3D space coordinate system to complete the point cloud construction;
[0124] The PCL (Point Cloud Library) in this embodiment is an open-source point cloud processing library that provides a large number of algorithms and tools for 3D point cloud processing. Its main functions include:
[0125] 1. Storage and management of point cloud data: Provide an efficient point cloud data structure for convenient storage, access, and operation.
[0126] 2. Point cloud preprocessing: Filter the point cloud to remove noise points and outliers and improve the quality of the point cloud.
[0127] 3. Point cloud feature extraction: Extract geometric features of the point cloud, such as normal vectors, curvatures, boundaries, etc.
[0128] 4. Point cloud segmentation: Segment the point cloud into different regions or target objects.
[0129] 5. Point cloud classification and recognition: Use machine learning and deep learning algorithms to classify and recognize the point cloud.
[0130] 6. Point cloud registration and fusion: Register and fuse point cloud data obtained from different perspectives or at different times.
[0131] 7. Point cloud visualization: Provide rich visualization tools for convenient viewing and analysis of point cloud data.
[0132] 4) Perform statistical filtering on the outliers in the point cloud constructed for the current frame, and then splice the remaining point cloud of the current frame into the constructed point cloud map in combination with the pose information of the current frame;
[0133] 5) Finally, obtain the global dense point cloud semantic map of the system.
[0134] In order to improve the storage efficiency of the map in this embodiment, an octree structure is used to construct the semantic map, reducing unnecessary details and improving the readability and storage efficiency of the map. The construction of the octree map requires continuously updating the octree map according to the depth information in the RGB-D image, and filtering out the residual moving objects in the map by comparing the occupancy probability of the nodes with the threshold; when constructing the octree map, the above-mentioned dense point cloud map can be directly converted into an octree map through the octomap library.
[0135] Embodiment 2
[0136] In this embodiment, the following experiments were conducted to implement the above method. Experimental configuration: The hardware configuration for all experiments was an Intel(R) Core(TM) i5-10400F CPU, NVIDIA GeForce RTX 3060, 16GB of memory, the operating system was Ubuntu18.04, CUDA11.2, Pytorch1.10.1, and Opencv3.4.15.
[0137] The experimental process of the visual odometry based on deep learning in the method provided in this embodiment is as follows: To test the performance of the algorithm of this module in an indoor low-texture scene, in the above experimental environment, the performance of the visual odometry algorithm and the ORB-SLAM2 algorithm was compared on the low-texture sequences fr3_nos_not_far, fr3_nos_not_near, fr3_str_not_far, and fr3_str_not_near in the TUM RGB-D dataset. The evaluation criteria were the absolute trajectory error (ATE) and the relative pose error (RPE), and the root mean square error (RMSE), mean error (Mean), median error (Median), and standard deviation (S.D.) were used for statistics to better reflect the accuracy of the visual SLAM system.
[0138] As Figure 6 shown, Figure 1 For different algorithms running in a low-texture scene, as can be seen from Figure 6 , when the ORB-SLAM2 algorithm runs on a low-texture sequence, it will encounter situations where it cannot be initialized or the tracking is lost immediately after initialization, which fully shows that ORB-SLAM2 cannot run on a low-texture sequence. Therefore, it is also impossible to save the camera trajectory results for accuracy comparison. However, the algorithm of the present invention is not affected when running in a low-texture scene and will not have problems such as inability to initialize and tracking loss. (a) ORB-SLAM2 cannot be initialized, (b) ORB-SLAM2 tracking is lost, (c) the algorithm of this module runs normally.
[0139] As Figure 7 shown, Figure 7 For the comparison of feature extraction of low-texture sequences, Figure 7 in (a) the fr3_nos_not_far sequence, (b) the fr3_nos_not_near sequence, (c) the fr3_str_not_far sequence, and (d) the fr3_str_not_near sequence. Figure 7The trajectory results of the algorithm in this section on four low-texture sequences are given. The left image is the absolute trajectory error map, where the red line segments represent the error values. The longer the line segment, the larger the error. The right image is the relative pose error fluctuation map. It can be seen that the absolute trajectory error of the algorithm in this module is relatively large on the fr3_nos_not_near sequence, but the absolute trajectories of the other sequences basically coincide with the true trajectories, and the error values are relatively small. At the same time, from the relative trajectory result map, the relative trajectory error values of the algorithm in this module on the low-texture sequences are all lower than 0.7m, indicating that the algorithm in this module can meet the positioning accuracy requirements in the low-texture scenario.
[0140] Among them, Table 1 shows the specific statistical values of the absolute trajectory error and relative pose error of the algorithm in this section. Similarly, it can be seen that except for the relatively large statistical value of the absolute trajectory error on the fr3_nos_not_near sequence, exceeding 1.0m, the statistical values of the other sequences are all within 0.2m. The reason for the relatively large error in the fr3_nos_not_near sequence may be that the fr3_nos_not_near sequence is a loop sequence. The loop detection module of the algorithm in this section uses the method of ORB-SLAM2 and judges according to the pixel features of the images. However, the pixel feature similarities in the low-texture environment are very high, and the loop cannot be detected, resulting in the continuous accumulation of the error value of the system as the system runs, leading to a relatively large final error value.
[0141] Table 1 Absolute Trajectory Error and Relative Pose Error of the Algorithm in This Module
[0142]
[0143] It can be seen from the above experimental results that the algorithm in this paper effectively improves the working performance of the ORB-SLAM2 algorithm in the indoor low-texture scenario, and significantly improves the positioning accuracy and robustness of the system.
[0144] The experimental process of the dynamic visual SLAM algorithm based on semantic information in this embodiment is as follows: To evaluate the performance of the dynamic feature point elimination algorithm of this module, it was experimentally tested on the dynamic sequences in the TUM RGB-D dataset. The TUM RGB-D dataset includes several common dynamic scene sequences. In the low-dynamic sequence sitting, there are two people sitting in front of a table, with occasional local movements on the chair, such as raising the arm and waving the arm, etc.; in the high-dynamic sequence walking, there are two people sitting in front of a table, and they stand up and move around the table significantly. In both sequences, the camera has four types of movements: xyz, static, halfsphere, and rpy. Among them, xyz means the camera moves along the xyz axis; static means the camera is stationary; halfsphere means the camera moves along a hemispherical trajectory; rpy means the camera rotates along the roll-pitch-yaw axes. The sequence names are simplified to fr3_s_xyz, fr3_s_sta, fr3_s_hal, fr3_s_rpy and fr3_w_xyz, fr3_w_sta, fr3_w_hal, fr3_w_rpy.
[0145] (1) Experimental results of the low-dynamic sequence
[0146] As Figure 8 shown, Figure 8 is the result of eliminating dynamic feature points in the low-dynamic sequence. (a) fr3_s_hal sequence, (b) fr3_s_rpy sequence, (c) fr3_s_sta sequence, (d) fr3_s_xyz sequence; Figure 8 is the test result of the dynamic point elimination algorithm in this section on the low-dynamic sequence, where the green points are static feature points and the red points are dynamic feature points.
[0147] In Figure 8 the (a) fr3_s_hal sequence in, only the right hand of the person is moving in the first three images, while the left person is stationary. It can be seen that the algorithm in this section can well eliminate the feature points on the dynamic object and retain the feature points on the stationary object whose prior information is dynamic. In the last image, the person is stationary and only the camera is rotating, and the algorithm in this section can still retain the static feature points on the person. In Figure 8 the (b) fr3_s_rpy sequence in, both people's heads are moving in the first image, and the algorithm in this section eliminates all the feature points on the dynamic object. In the second image, the moving heads of the two people do not appear in the image, and the algorithm in this section determines the person as a static object, so it does not eliminate the feature points on the object. In the third image, all the feature points on the dynamic object are accurately eliminated. The last image is the last two frames of the sequence, and both people have long stopped moving, and the algorithm in this section also well retains the feature points on them. In Figure 8In the (c) fr3_s_sta sequence, except for the person on the left in the second image who is stationary, the arms of the two people in the remaining images are moving when they are communicating, which is a dynamic object. The algorithm in this section also accurately detects and completely removes the feature points of the real dynamic objects. Figure 8 In the (d)fr3_s_sta sequence in Figure 1, the first three images completely remove dynamic feature points, and retain the feature points on the static object. In the last image, the person on the left is a dynamic object, and the algorithm in this section does not completely remove the feature points on his body. The reason for this result is that the mask in the instance segmentation result of the last image does not completely cover the object on the left, resulting in some feature points not being removed. From the above experimental results, it can be seen that the algorithm in this section can combine semantic information and motion consistency to detect the number of dynamic points on the prior dynamic object, thereby judging the actual motion of the object and accurately removing the feature points on the dynamic object.
[0148] (2) High dynamic sequence experimental results
[0149] like Figure 9 As shown, Figure 9 Results of dynamic feature point removal of high dynamic sequences: (a) fr3_w_hal sequence, (b) fr3_w_rpy sequence, (c) fr3_w_sta sequence, and (d) fr3_w_xyz sequence; Figure 9 The following are the test results of the dynamic point removal algorithm in this section on a high-dynamic sequence. In the high-dynamic sequence, both people in the image are moving in a large range, including sitting down, standing up, and walking around the table. Figure 9 In the (a) fr3_w_hal sequence in Figure 1, except for the feature points on the right edge of the person in the second image, the feature points on the moving people in the remaining images were accurately removed. Figure 9 In the (b) fr3_w_rpy sequence in Figure 1, although the camera is rotating, the algorithm in this section still accurately removes the feature points of the dynamic object. Figure 9 In the (c) fr3_w_sta sequence in the figure, we can see that the edge feature points on the left side of the person in the first image are not completely removed, and some dynamic feature points are still retained on the right side of the person's legs in the fourth image. This is because there are errors in instance segmentation and the dynamic objects are not fully covered. Figure 9In the (d) fr3_w_xyz sequence, this algorithm removed the feature points on the static teddy bear doll in the left cabinet of the person on the right in the second picture. This is because the camera in this sequence was moving in the x - y - z directions, and there were errors in calculating the motion consistency check, which misjudged multiple feature points on the teddy bear doll as dynamic points, resulting in the teddy bear doll being defined as a dynamic object. However, from the overall experimental results, the algorithm in this section can effectively and accurately remove the feature points on dynamic objects and retain static feature points when facing dynamic scenes.
[0150] As Figure 10 shown, Figure 2 For the true trajectory diagrams of the two sequences, the experimental process of the loop closure detection algorithm based on semantic template matching provided in this embodiment is as follows: Select the "fr2_desk" and "fr3_long_office" sequences in the TUM RGB - D dataset to conduct a similarity score experimental comparison between the algorithm in this section and the bag - of - words (BoW) model method. Both of these sequences describe indoor multi - object scenes, including desks, chairs, books, etc. In addition, Figure 10 the true trajectory diagrams drawn from their camera pose information are given respectively, and it can be seen that there are multiple sets of loop closures in these two sequences.
[0151] As Figure 11 shown, Figure 3 For the test images of the fr2_desk sequence, six images in the fr2_desk sequence are selected as test images, and the similarity scores between the images are calculated respectively. The first image and the last image form a loop closure and have the highest similarity. At the same time, in this experiment, the dictionary of the BoW model is not retrained, and the dictionary model of ORB - SLAM2 is directly used.
[0152] The test results of the BoW method are shown in Table 2. The similarity value range between images is [0, 1]. In this range, a similarity of 1 means the images are exactly the same, and a similarity of 0 means the images are completely different. The higher the similarity value, the more similar the images are. From the results, the similarity score between Image 1 and Image 6 is 0.02574, which are the most similar images, while the similarity scores between the remaining images are all lower than 0.02. Although the similarity score calculated by the BoW method between the most similar images is also the highest, the gap between the similarity scores with other images is not obvious enough. When the total number of images increases, it is easy to detect incorrect loop closures, and the BoW method depends on the dictionary scale. The larger the scale, the better the detection effect, but it is very difficult to train the objects in the scene in advance in practical applications.
[0153] Table 2 Similarity calculation results of the BoW algorithm
[0154]
[0155] Table 3 presents the test results of the image similarity calculation method based on semantic template matching in this section. From the result, the similarity score between Image 1 and Image 6 is the highest, which is 0.34632, while the score between the dissimilar Image 1 and Image 4 is only 0.04333. The similarity scores between the remaining images are also relatively low, and there is an obvious gap in scores between similar images and dissimilar images. The similarity scores of the algorithm in this section are screened through semantic and geometric multiple matching. Therefore, the scores between highly similar images may not be very high, but there is an obvious gap compared with the scores between dissimilar images.
[0156] Table 3 Similarity calculation results of the algorithm in this section
[0157]
[0158] Six test images were also selected for experimental testing on the fr3_long_office sequence, as Figure 12 shown. Similar to the experiment on the fr2_desk sequence, Image 1 and Image 6 form a loop.
[0159] Table 4 presents the test results of the BoW algorithm on the fr3_long_office sequence. It can be seen that the highest similarity score between Image 1 and Image 6 is 0.02544, which is higher than the scores between the remaining dissimilar images. However, the similarity scores between similar images and dissimilar images still have a relatively small difference. Comparing Image 3 and Image 6, there are basically no similar objects between the images, but the similarity score between Image 3 and Image 6 is 0.01206, which is still higher than the similarity values between some images. It can be seen that since the BoW algorithm calculates based on low-level features such as image pixels and does not consider the specific object information in the image, there will be cases of descriptor mis-matching calculation, reducing the accuracy of loop detection.
[0160] Table 4 Similarity calculation results of the BoW algorithm
[0161]
[0162] Table 5 shows the test results of the algorithm in this section. It can be seen that the similarity score between the similar Image 1 and Image 6 is 0.35606, the similarity scores between the significantly dissimilar Image 1 and Image 3, and Image 3 and Image 6 are 0, and the scores between similar images are much higher than the scores between the remaining images. The algorithm in this section fully considers the object semantic information in the image and can accurately judge whether it is a similar scene based on the object information in the image.
[0163] Table 5 Similarity calculation results of the algorithm in this section
[0164]
[0165] The experimental process of the entire system integration provided in this embodiment is as follows: To verify the working performance of the integrated dynamic SLAM algorithm in an indoor dynamic scenario, dynamic sequences in the TUM RGBD dataset are selected for testing. Specifically, one low-dynamic sequence: fr3_sitting_static and four high-dynamic sequences: fr3_walking_xyz, fr3_walking_static, fr3_walking_halfsphere, fr3_walking_rpy are chosen. By comparing the absolute trajectory error and relative pose error of the algorithm provided in this embodiment, the original ORB-SLAM2, and current excellent dynamic SLAMs (DS-SLAM, DynaSLAMError! Reference source not found., and RDMO-SLAM) on the above sequences, the accuracy and robustness of the system are analyzed. Meanwhile, the average time consumption of different SLAM systems is also compared to analyze the real-time performance of the system.
[0166] As Figure 13 and Figure 14 shown, Figure 13 Comparison of absolute trajectory errors for the fr3_s_sta sequence; (a) ORB-SLAM2; (b) The algorithm in this section; Figure 14 Comparison of relative pose errors for the fr3_s_sta sequence; (a) ORB-SLAM2; (b) The algorithm in this section; Figure 13 and Figure 14 show the absolute trajectory error and relative pose error of ORB-SLAM2 and the algorithm in this section on the low-dynamic sequence fr3_s_sta. It can be seen that on the low-dynamic sequence, the errors of both are relatively small. From the perspective of relative pose error, the error values of the two algorithms are basically within 0.025m, but the error of the algorithm in this section is more stable. Because there are fewer dynamic feature points in the low-dynamic sequence, ORB-SLAM2 has already removed some dynamic matching point pairs using the RANSAC algorithm. Therefore, the system already has high accuracy in the low-dynamic scenario.
[0167] The absolute trajectory error and relative pose error under high-dynamic sequences are as follows. By comparing the four sequences in the high-dynamic scenario, it can be seen that the overall absolute trajectory error value of the algorithm in this section is smaller than that of ORB-SLAM2, and at the same time, the relative pose error value is also smaller and more stable. From the comparison chart of the absolute trajectory error, the trajectory of the algorithm in this module basically coincides with the real trajectory, while there are obvious differences between the trajectory of the ORB-SLAM2 algorithm and the real trajectory. From the comparison chart of the relative pose error, the maximum error value of the algorithm in this section on the fr3_w_rpy sequence is 0.16m, while the maximum error value of the ORB-SLAM2 algorithm on the fr3_w_hal sequence is 1.2m, and the minimum value of the maximum error value of ORB-SLAM2 reaches 0.8m, which is much higher than the error value of the algorithm in this section. Therefore, this is sufficient to show that the performance of the algorithm in this section is better than that of ORB-SLAM2 in the high-dynamic scenario.
[0168] As Figures 15 to 22 shown, Figure 15 is the comparison of the absolute trajectory error of the fr3_w_xyz sequence; (a) ORB-SLAM2; (b) the algorithm in this section; Figure 16 is the comparison of the absolute trajectory error of the fr3_w_sta sequence; (a) ORB-SLAM2; (b) the algorithm in this section; Figure 17 is the comparison of the absolute trajectory error of the fr3_w_hal sequence; (a) ORB-SLAM2; (b) the algorithm in this section; Figure 18 is the comparison of the absolute trajectory error of the fr3_w_rpy sequence; (a) ORB-SLAM2; (b) the algorithm in this section; Figure 19 is the comparison of the relative pose error of the fr3_w_xyz sequence; (a) ORB-SLAM2; (b) the algorithm in this section; Figure 20 is the comparison of the relative pose error of the fr3_w_sta sequence; (a) ORB-SLAM2; (b) the algorithm in this section; Figure 21 is the comparison of the relative pose error of the fr3_w_hal sequence; (a) ORB-SLAM2; (b) the algorithm in this section; Figure 22 is the comparison of the relative pose error of the fr3_w_rpy sequence; (a) ORB-SLAM2; (b) the algorithm in this section.
[0169] To more intuitively observe the performance of the algorithm in this section, Table 6 presents the root mean square error (RMSE), standard deviation (S.D.), and percentage of performance improvement of the absolute trajectory error values of the algorithm in this section and ORB-SLAM2. It can be seen that compared with ORB-SLAM2, the root mean square error of the algorithm in this section decreased by 63.46% and the standard deviation decreased by 58.53% in the low-dynamic sequence. In the high-dynamic sequence, compared with ORB-SLAM2, the root mean square error of the algorithm in this section decreased by an average of 96.00% and the standard deviation decreased by 91.25%. This fully demonstrates that in most dynamic scenarios, the algorithm in this section can significantly improve the positioning accuracy, especially in high-dynamic scenarios. Tables 7 and 8 present the relative pose error statistical data of the algorithm in this section and ORB-SLAM2. It can be seen that in the low-dynamic scenario, the root mean square and standard deviation of the relative displacement error of the algorithm in this section decreased by 14.44% and 41.86% compared with ORB-SLAM2, and the root mean square and standard deviation of the relative rotation error also decreased by 5.21% and 5.80%. In the high-dynamic environment, compared with ORB-SLAM2, the root mean square and standard deviation of the relative pose error of the algorithm in this section decreased by more than 90%. This shows that the algorithm in this section greatly improves the system positioning accuracy in dynamic scenarios.
[0170] Table 6 Comparison of Absolute Trajectory Errors between the Algorithm in this Section and ORB-SLAM2
[0171]
[0172] Table 7 Comparison of Relative Pose Errors (Displacement / m) between the Algorithm in this Section and ORB-SLAM2
[0173]
[0174]
[0175] Table 8 Comparison of Relative Pose Errors (Rotation / deg) between the Algorithm in this Section and ORB-SLAM2
[0176]
[0177] To further verify the performance of the algorithm in this section, the algorithm in this section was compared with other excellent current dynamic SLAM systems in terms of absolute trajectory error, relative pose error, and running time in an indoor dynamic scene sequence. Among them, the data results of other dynamic SLAM systems all come from the data given in their papers. The comparison of the statistical values of absolute trajectory error and relative pose error is shown in Tables 9, 10, and 11. It can be seen that except that the root mean square value of the absolute trajectory error of the algorithm in this section on the fr3_w_sta sequence and the standard deviation value on the fr3_w_xyz sequence are higher than those of DynaSLAM, the results on other sequences are better than those of the remaining dynamic SLAM systems. At the same time, from the results of relative pose error, except that the root mean square and standard deviation of the algorithm in this section on the fr3_w_sta sequence and the fr3_w_xyz sequence are slightly larger than those of DynaSLAM, the root mean square and standard deviation on the remaining sequences are the smallest. Since DynaSLAM uses a semantic segmentation network and combines depth image information and parallax angle to judge dynamic feature points, while the algorithm in this section judges dynamic feature points by combining semantic information and motion consistency, and the camera of the fr3_w_sta sequence is fixed, so the error of the algorithm in this section in calculating the projection matrix will be relatively large, resulting in a slightly lower positioning accuracy. However, generally speaking, the algorithm in this section has obvious advantages in terms of accuracy and robustness compared with the remaining dynamic SLAM algorithms.
[0178] Table 9 Comparison of absolute trajectory errors (m) between the algorithm in this section and other dynamic SLAMs
[0179]
[0180] Table 10 Comparison of relative pose errors (displacement / m) between the algorithm in this section and other dynamic SLAMs
[0181]
[0182] Table 11 Comparison of relative pose errors (rotation / deg) between the algorithm in this section and other dynamic SLAMs
[0183]
[0184] In addition to positioning accuracy and robustness, real-time performance is also an important indicator for judging the performance of visual SLAM systems. Table 12 shows the comparison of the running times of different visual SLAM systems. It can be seen that the average calculation time per frame of the algorithm in this section using the YOLOv8-Seg network is the least, only 15 - 25 ms. At the same time, the average calculation time per frame of the system of the algorithm in this section is 50 - 60 ms. Although it is higher than the ORB-SLAM2 and RDMO-SLAM systems, it also meets the real-time requirements of visual SLAM.
[0185] Table 12 Comparison of running times between the algorithm in this section and other SLAMs
[0186]
[0187] In summary, through the analysis of the experimental results, it can be seen that the algorithm in this section can well balance the positioning accuracy and real-time performance of the system when facing indoor dynamic scenes, and its comprehensive performance is stronger than that of ORB-SLAM2 and other dynamic SLAM systems.
[0188] The experimental process of constructing the semantic map in this embodiment is as follows: Select the low-texture sequence fr3_str_not_far, low-dynamic sequence fr3_s_sta, and high-dynamic sequences fr3_w_sta and fr3_w_xyz in the TUM RGB-D dataset. Conduct a mapping effect comparison experiment between the dense point cloud semantic map and octree semantic map construction algorithms in this section and the algorithm of ORB-SLAM2 after adding dense point cloud and octree mapping threads to verify the performance of the algorithm in this section.
[0189] (1) Point cloud map construction experiment
[0190] As Figure 23 shown, Figure 23 For the construction of the dense point cloud map of ORB-SLAM2, the mapping effect after ORB-SLAM2 adds a dense point cloud construction thread is as Figure 23 shown. ORB-SLAM2 cannot run in a low-texture environment, so there is no result of the dense point cloud map. However, from the dense point cloud map constructed from the dynamic sequence, it can be seen that there are a large number of dynamic objects in the map, and as the dynamic objects move, there are a large number of ghosts in the map, seriously affecting the readability of the map. Figure 23 In (a) belongs to the low-dynamic sequence, where only part of the human body moves. It can be seen that there are fewer ghosts in the constructed map. However, due to the large key frame pose error of ORB-SLAM2 in a dynamic environment, there are still large errors in the spliced dense point cloud. Figure 23 In (b) and (c) are high-dynamic sequences. It can be seen that there are a large number of continuous ghosts in the constructed map, and there is a large amount of overlapping redundant information in objects such as walls. In subsequent navigation and path planning tasks, these ghosts will be defaulted as obstacles, seriously affecting the use of the map.
[0191] As Figure 24 shown, Figure 24Schematic diagram of the mapping effect of the dense point cloud semantic map construction method. (a) fr3_str_not_far sequence; (b) fr3_s_sta sequence; (c) fr3_w_sta_ sequence; (d) fr3_w_xyz sequence. Object color information mapped from the semantic information of key frames to the point cloud is added to the map, and different objects detected in the image are represented by different colors. For example, pink represents a chair, and blue represents a computer, etc. Figure 24 In (a), it is a low-texture sequence. Since there are no objects in the sequence, the semantic information of the detected objects is not available, and the point cloud map does not contain semantic information either. Figure 24 In (b), it is a low-dynamic sequence. It can be seen that the person on the right has been completely removed, but at the same time, the person will occlude the scene, resulting in partial loss of scene information in the map. At the same time, during the movement of the camera, the image will be blurred, resulting in missed detection when performing instance segmentation on the person on the left. Therefore, some residual dynamic object information is also retained in the map. Figure 24 In (c) and (d), they are high-dynamic sequences. It can be seen that the algorithm in this section effectively removes the information of dynamic objects, and there are no ghosts of dynamic objects in the constructed map. Due to the occlusion of dynamic people, the mapping of objects such as chairs is incomplete, but generally, compared with the point cloud map constructed by ORB-SLAM2, there is less redundant information. For example, there is no large amount of overlap between the walls and the background, and the contours of objects such as tables are more obvious. By combining semantic information to remove dynamic objects and improving the key frame selection strategy, the constructed dense point cloud map has a higher coincidence degree and does not show a layered situation. At the same time, the map contains the semantic information of objects, which not only improves the readability and usability of the map but also lays a foundation for realizing more advanced command tasks such as human-computer interaction.
[0192] (2) Octree map construction experiment
[0193] As Figure 25 shown, Figure 25 For the octree map construction of ORB-SLAM2, (a) fr3_s_sta sequence; (b) fr3_w_sta sequence; (c) fr3_w_xyz sequence. The result of converting the ORB-SLAM2 dense point cloud map into an octree map is as Figure 25 shown. The effect is the same as that of the dense point cloud map in the previous section. It can be seen that the constructed octree map contains more noise points and ghosts. Figure 25In (b) and (c), they are high-dynamic sequences. As the dynamic object moves, the entire map scene is occupied by the ghosts generated by the dynamic object, and it is impossible to see the specific outlines of the objects in the octree map. The whole scene appears very cluttered. Moreover, due to the relatively large camera pose error in the dynamic scene, the map shows a layering phenomenon, seriously affecting the readability of the map.
[0194] As Figure 26 shown, Figure 26 For the construction of the octree map, (a) fr3_str_not_far sequence; (b) fr3_s_sta sequence; (c) fr3_w_sta_ sequence; (d) fr3_w_xyz sequence; From Figure 26 in (d), it can be seen that there are still a small number of noise points and overlapping phenomena in the octree map. This is because when the semantic network performs instance segmentation, some limbs of the moving person will be outside the mask coverage, and not all information can be completely filtered out. Although filtering processing has been carried out, some noise points will still be retained. However, from Figure 26 the overall mapping results of (a), (b) and (c), it can be seen that the octree semantic map constructed by the algorithm in this section effectively filters out the information of dynamic objects when facing low-texture and dynamic scenes, with low map redundancy and low overlap degree.
[0195] The above experimental results fully prove that the scene and object outlines in the semantic map constructed by the algorithm in this section are clear, and the semantic information is rich. Through the semantic information, it is better to distinguish the objects in the surrounding environment, greatly improving the readability and practicality of the map.
[0196] The method provided in this embodiment solves the problems of low positioning accuracy and poor robustness of traditional visual SLAM algorithms in low-texture and dynamic scenes. This method has been verified by a large number of experiments and has achieved remarkable results in the following aspects:
[0197] 1. Performance improvement in low-texture scenes. In the low-texture sequence test of the TUM RGB-D dataset, the algorithm of the present invention can run stably, while the traditional ORB-SLAM2 algorithm has problems such as inability to initialize or tracking loss; except for the fr3_nos_not_near sequence, the absolute trajectory errors of the other sequences are all within 0.2m, reflecting the stability of the system in low-texture environments; compared with ORB-SLAM2, the root mean square error (RMSE) of the absolute trajectory error has decreased by 63.46%, and the standard deviation (S.D.) has decreased by 58.53%.
[0198] 2. Dynamic scene processing ability. In low-dynamic sequences, the root mean square and standard deviation of the relative displacement error of this system decreased by 14.44% and 41.86% respectively compared with ORB-SLAM2; in high-dynamic sequences, the system performed even more prominently: the root mean square error of the absolute trajectory error decreased by an average of 96.00%, the standard deviation decreased by an average of 91.25%, and the root mean square and standard deviation of the relative pose error decreased by more than 90%; compared with other advanced dynamic SLAM systems (DS-SLAM, DynaSLAM, and RDMO-SLAM), this system achieved optimal or near-optimal performance on most sequences.
[0199] 3. Improvement in loop closure detection. The loop closure detection method based on semantic template matching significantly improved the detection accuracy; in the fr2_desk sequence test, the similarity score of this system reached 0.34632, much higher than 0.02574 of the traditional BoW method; in the fr3_long_office sequence, the system could effectively avoid false matches and improve the reliability of loop closure detection.
[0200] 4. Quality of semantic map construction. Point cloud semantic map: successfully eliminated the ghosting problem of dynamic objects, improved the readability of the map through semantic information annotation, with clearer object contours and less redundant information; octree semantic map: effectively filtered out dynamic object information, with lower map redundancy and overlap, more complete scene structure, and clear object contours.
[0201] 5. System real-time performance. Using the YOLOv8-Seg network, the single-frame semantic segmentation time is only 15 - 25 ms; the average processing time per frame of the system as a whole is 50 - 60 ms; although it is slightly slower than ORB-SLAM2 (20 - 30 ms) and RDMO-SLAM (25 - 35 ms), it is better than DS-SLAM (82.6 ms) and DynaSLAM (365.2 ms); fully meeting the real-time requirements.
[0202] 6. System limitations and future improvement directions. In the fr3_nos_not_near sequence, large errors occurred due to loop closure detection problems, and the loop closure detection algorithm needs to be further improved; there are still a small number of noise points and overlapping phenomena in some high-dynamic sequences, which can be improved by optimizing the semantic segmentation accuracy; although the system processing time meets the real-time requirements, there is still room for optimization; during the semantic map construction process, information in areas blocked by dynamic objects may be missing, and a filling strategy needs to be explored.
[0203] In summary, the method provided in this embodiment significantly improves the positioning accuracy and robustness of the SLAM system in low-texture and dynamic scenarios, while maintaining good real-time performance. By introducing deep learning and semantic information, the system not only solves the key problems of traditional visual SLAM, but also provides better support for subsequent high-level applications such as navigation planning and human-computer interaction. Future work will focus on improving the loop detection accuracy, optimizing the semantic segmentation effect, enhancing the real-time performance of the system, and improving the construction quality of the semantic map.
[0204] The above-described embodiments are merely preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.
Claims
1. A semantic visual SLAM method for indoor low-texture and dynamic environments, characterized in that: Including the following steps: 1) Obtain an image, and obtain the dynamic feature points of the object according to the image texture degree and the object semantic information; 2) Dynamic point elimination: Eliminate the dynamic feature points of the object through motion consistency checking and combining semantic information; 3) Front-end tracking. After eliminating the dynamic feature points, enter the tracking thread for camera pose estimation; for the image entering the tracking thread, first match the static feature points with the adjacent key frames to calculate the rough pose of the camera, and track the local map for pose optimization, and finally decide whether to insert a new key frame into the local mapping thread; 4) Local mapping. After the key frame enters the local mapping thread, first eliminate the redundant map points of the key frame, then create new map points to restore the co-visible points, and use local BA to further optimize the spatial points and the camera pose to obtain the accurate camera pose, and finally eliminate the redundant key frames and enter the loop detection thread; 5) The loop detection thread calculates the similarity score for the historical key frames of the current key frame through semantic template matching, so as to judge whether a loop is formed, and correct the pose of the image forming the loop to eliminate the cumulative drift error of the system, and finally perform global BA optimization to update the environmental map to obtain the trajectory of the camera of the system.
2. The semantic visual SLAM method for indoor low-texture and dynamic environments according to claim 1, wherein: In step 1), after obtaining the image, it is processed by an improved visual odometer. The specific steps are as follows: 1) Obtain the current nth frame RGB image; 2) Feature point extraction. Calculate the texture degree of the current image. If it is judged as a low-texture image, extract GCNv2 feature points; If it is judged as a high-texture image, extract ORB feature points; 3) Tracking. First judge whether the current constant velocity model is empty. If it is empty, perform feature point matching and tracking between the reference key frame and the current frame; If it is not empty, use the constant velocity motion model to perform feature point matching and tracking between the previous frame and the current frame; 4) Track the local map. According to the above tracking results, optimize the pose of the current frame by minimizing the reprojection error. If the pose calculation is successful, add the pose information of the current frame to the local map of the pose trajectory; 5) Select key frames. Judge whether the key selection condition is satisfied. If it is satisfied, add the current frame to the key frame queue; if it is not satisfied, do not select it as a key frame.
3. The semantic visual SLAM method for indoor low-texture and dynamic environments according to claim 1, wherein: The method for eliminating dynamic points in step 2) is as follows: First, obtain the object semantic information by instance segmentation of the input image, and at the same time extract the image features and perform motion consistency checking. Combine the two results to detect the dynamic feature points. Finally, eliminate the dynamic feature points, and use the static feature points after elimination to calculate the camera pose.
4. The semantic visual SLAM method for indoor low-texture and dynamic environments according to claim 1, characterized in that: The loop detection algorithm in step 3) is specifically as follows: First, extract the semantic template of the image, and then combine geometric and semantic information to perform matching and screening on the semantic template, so as to calculate the image similarity score, and finally complete the closed-loop detection.
5. The semantic visual SLAM method for indoor low-texture and dynamic environments according to claim 1, wherein: The map in step 5) is a dense point cloud semantic map, which specifically includes the following steps: 1) Process the image through the front-end tracking thread of visual SLAM to obtain the camera pose corresponding to each frame of the image. Then, judge whether the current frame is a key frame according to the PKS key frame selection strategy. At the same time, for the input image, use the instance segmentation network to obtain the semantic information of the image for subsequent mapping the semantic information to the point cloud; 2) Detect the key frames according to the dynamic point detection method based on semantic information, and construct the point cloud for the pixel points of the dynamic objects in the detected key frames; 3) Then, based on the point cloud construction method and the PCL library, project the coordinate values of the static pixel points in the key frames and the different colors corresponding to the object semantic information detected by the semantic thread into the 3D space coordinate system to complete the point cloud construction; 4) Perform statistical filtering on the outlier points in the point cloud constructed for the current frame, and then splice the remaining point cloud of the current frame and the pose information of the current frame into the constructed point cloud map; 5) Finally, obtain the global dense point cloud semantic map of the system.
6. The semantic visual SLAM method for indoor low-texture and dynamic environments according to claim 5, characterized in that: The semantic map is constructed using an octree structure, and the dense point cloud is converted into an octree map through the octomap library.
7. The semantic visual SLAM method for indoor low-texture and dynamic environments according to claim 1, characterized in that: The thread for obtaining semantic information and the front-end tracking thread are executed in parallel; or the local mapping thread and the loop detection thread are executed in parallel.
8. The semantic visual SLAM method for indoor low-texture and dynamic environments as claimed in claim 1, wherein: The segmentation, feature point extraction, and dynamic point removal of the image are processed for the same image. After the motion consistency check of the image, wait for the processing result of the semantic information acquisition thread, and then remove the dynamic feature points in combination with the semantic information.
9. A semantic visual SLAM system for indoor low-texture and dynamic environments, characterized in that: Including An image input module for inputting the acquired images into the feature point extraction module and the semantic information acquisition module respectively; A specific point extraction module for calculating ORB feature points and GCNv2 feature points according to the image texture degree respectively; A semantic information acquisition module for processing using Yolov8-Seg instance segmentation to obtain semantic information; A dynamic point removal module for removing dynamic feature points; removing the dynamic feature points of the object through motion consistency check and in combination with semantic information; A front-end tracking module. After removing the dynamic feature points, enter the tracking thread for camera pose estimation; for the image entering the tracking thread, first match the static feature points with the adjacent key frames, calculate the rough pose of the camera, and track the local map for pose optimization. Finally, decide whether to insert a new key frame into the local mapping thread; A local mapping module. After the key frame enters the local mapping thread, first remove the redundant map points of the key frame, then create new map points for restoring the co-visible points, and further optimize the spatial points and camera poses using local BA to obtain the accurate camera pose. Finally, remove the redundant key frames and enter the loop detection thread; A loop detection module. The loop detection thread calculates the similarity score for the historical key frames of the current key frame through semantic template matching to judge whether a loop is formed, corrects the pose of the image forming the loop to eliminate the cumulative drift error of the system, and finally performs global BA optimization to update the environment map to obtain the trajectory of the camera of the system.
10. The semantic visual SLAM system for indoor low-texture and dynamic environments according to claim 9, characterized in that: It also includes an improved visual odometry calculation module, which is used to process the acquired images through an improved visual odometry calculation method according to the following steps: 1) Obtain the current nth frame RGB image; 2) Feature point extraction. Calculate the texture degree of the current image. If it is determined to be a low-texture image, extract GCNv2 feature points; If it is determined to be a high-texture image, extract ORB feature points; 3) Tracking. First, determine whether the current constant velocity model is empty. If it is empty, perform feature point matching and tracking between the reference key frame and the current frame; If it is not empty, use the constant velocity motion model to perform feature point matching and tracking between the previous frame and the current frame; 4) Tracking the local map. According to the above tracking results, optimize the pose of the current frame by minimizing the reprojection error. If the pose calculation is successful, add the pose information of the current frame to the local map of the pose trajectory; 5) Select key frames. Determine whether the key selection condition is met. If it is met, add the current frame to the key frame queue; if it is not met, do not select it as a key frame.
Citation Information
Cited By
Synchronous positioning and mapping method for underwater pipeline detection
CN121115020A
Trackless navigation texture information redundancy analysis method
CN122156395A