Object-level semantic vision SLAM method based on object geometric constraint and abnormal point elimination

By introducing object geometric constraints and exception point removal technologies into the SLAM system, the problem of low recognition accuracy in complex environments is solved, and higher robustness and accuracy are achieved.

CN119992509APending Publication Date: 2025-05-13BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510013284.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The robustness and accuracy of traditional visual SLAM in complex environments, especially in the precise recognition of objects. The existing semantic visual SLAM method is not thorough enough to eliminate abnormal points, which affects the performance of the system.

Method used

An object-level semantic visual SLAM method based on object geometric constraints and exception point removal is proposed. The object detection is carried out through the YOLOv8 algorithm, and the object geometric model is established, and the exception point removal is used to use the improved Extended Isolation Forest (IEIF) algorithm to improve the accuracy and reliability of recognition.

Benefits of technology

It significantly improves the robustness and accuracy of SLAM systems in dynamic and complex environments, improves the accuracy of object recognition and the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992509A_ABST
    Figure CN119992509A_ABST
Patent Text Reader

Abstract

The invention discloses an object-level semantic vision SLAM (Simultaneous Localization and Mapping) method based on object geometric constraint and abnormal point elimination. The performance of the traditional visual SLAM in a complex environment is still insufficient, especially for accurate identification of an object. According to a current semantic vision SLAM method, such as EAO-SLAM, abnormal points are not thoroughly removed when point cloud data are processed, and the performance of a system is affected. The invention provides an object-level semantic vision SLAM (Simultaneous Localization and Mapping) method based on object geometric constraint and abnormal point elimination, and aims to improve the autonomous navigation and environment cognitive ability of a robot in a complex environment. According to the invention, geometric constraints and an optimized Isolation Forest algorithm are introduced, and the accuracy of abnormal point detection and elimination is improved, so that the robustness and adaptive capacity of the SLAM system are enhanced. Experimental results show that the method is superior to a traditional algorithm in various scenes, and the potential of the method in practical application is shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics technology, and in particular to an object-level semantic visual SLAM (Simultaneous Localization and Mapping) method based on object geometric constraints and outlier elimination. Background Art

[0002] With the rapid development of robotics, visual SLAM plays an increasingly important role in autonomous navigation and environmental cognition. However, the performance of traditional visual SLAM in complex environments is still insufficient, especially in the accurate recognition of objects. Existing semantic visual SLAM methods, such as EAO-SLAM, do not eliminate outliers thoroughly when processing point cloud data, which affects the performance of the system. The robustness and low accuracy of SLAM systems in dynamic and complex environments have not yet been solved. Summary of the invention

[0003] This paper proposes an object-level semantic visual SLAM method based on object geometric constraints and outlier removal, aiming to improve the robot's autonomous navigation and environmental cognition capabilities in complex environments. The specific steps are as follows:

[0004] Step S1: Acquire continuous images and depth information of the environment through an RGB-D camera or a stereo camera.

[0005] Step S2: Use the YOLOv8 algorithm to perform object detection on the acquired image, identify various objects in the environment, and generate corresponding two-dimensional bounding boxes.

[0006] Step S3: Assign a category label to each detected object, such as “table”, “chair”, “person”, etc., to ensure the richness of semantic information.

[0007] Step S4: Generate a dense point cloud based on the depth information to represent the three-dimensional structure in the environment. For each detected object, establish its geometric model, including the object's center of mass, boundary and size information, and establish geometric constraints. The specific steps are as follows:

[0008] Step S41: Calculate the Euclidean distance between each map point and the object's centroid to determine which points in the point cloud belong to a specific object. The distance calculation formula is:

[0009]

[0010] Among them, m is the map point, p is the center of the object, and p l is the x-axis coordinate of the center of the object, p w is the y-axis coordinate of the center of the object, p h is the z-axis coordinate of the center of the object, ml is the x-axis coordinate of the map point, m w is the y-axis coordinate of the map point, m h is the z-axis coordinate of the map point;

[0011] Step S42: When Dis(m)>ObjectMax8, where ObjectMax is the maximum range value of the current object, it indicates that the point cannot belong to the object and is ignored.

[0012] Step S43: For basic common objects, their geometric constraints are defined and limited to ensure that the characteristics of the objects can be accurately reflected during the recognition process. Taking the length as the reference value, the corresponding width and height ratios of different types of objects are set:

[0013] prop=[1,prop wl ,prop hl ] T

[0014] where prop wl is the ratio of the width to the length of the object, prop hl It is the ratio of the height to the length of the object. For different objects, we customize different ratios of width to height and length to height, that is, we set a ratio for different objects based on experience.

[0015] Step S44: Calculate the distance between the newly added map point and the center point of the object in the width direction: Add the distance between the map point and the center point of the object in the height direction: When Dis w (m)>prop wl ·s l or Dis h (m)>prop hl ·s l When, s l The length of the object indicates that the length, width and height ratio of the newly added point far exceeds the ratio range of the basic object. Therefore, it can be reasonably inferred that the point cannot be a point in the object. Based on this judgment, the system will ignore the point to further improve the accuracy and reliability of recognition.

[0016] Step S5: In the process of system recognition, there will be factors such as occlusion between objects and changes in illumination, so an outlier removal thread is introduced. The specific steps are as follows:

[0017] Step S51: Projecting the filtered point cloud data onto the geometric model of the object to perform outlier detection.

[0018] Step S52: Use the improved Extended Isolation Forest (IEIF) algorithm to detect outliers on the projected point cloud. The algorithm calculates the outlier score based on the distribution characteristics of the points. The calculation formula for the outlier score is:

[0019]

[0020] Where x∈X is a point in the point cloud, n is the number of points in the point cloud, and E(h(x)) is the average depth reached by a single data point x in all trees. c(n) is a normalization factor defined as the average depth of an unsuccessful search in a binary search tree, calculated as:

[0021]

[0022] H(n-1) is the harmonic function, and the calculation formula is:

[0023] H(n-1)≈1n(n-1)+0.5772156649

[0024] Step S53: remove points with anomaly scores higher than 0.65 to improve the quality and reliability of point cloud data.

[0025] Compared with the traditional SLAM system, the object-level visual semantic SLAM method based on object geometric constraints and outlier removal provided by the present invention solves the scale problem of common objects by introducing a geometric constraint mechanism. The outlier removal algorithm is optimized to remove outliers that do not belong to the point cloud, which significantly improves the robustness and accuracy of the SLAM system in dynamic and complex environments, and has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 The core process of the present invention

[0027] Figure 2 To remove outliers from the keyboard point cloud using IEIF and IF respectively

[0028] Figure 3 The three columns are the TUM fr1_desk sequence, TUM fr2_desk sequence, and TUM fr3_long_office_household sequence. The four rows are the original images, semi-dense graphs, EAO-SLAM, and the proposed method of different dataset sequences.

[0029] Figure 4 Comparison of camera trajectories estimated by EAO-SLAM and the present invention in the TUM fr3_office dataset sequence

[0030] Figure 5 Algorithm flow chart for geometric constraints

[0031] Figure 6 Algorithm flow chart for outlier removal DETAILED DESCRIPTION

[0032] The technical solution of the present invention is further described below in conjunction with the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention without departing from the spirit and scope of the technical solution of the present invention should be included in the protection scope of the present invention.

[0033] Based on the traditional SLAM technology, this paper provides an object-level visual semantic SLAM method based on object geometric constraints and outlier removal. The core process is as follows: Figure 1 As shown, the method includes the following contents:

[0034] 1. YOLOv8

[0035] YOLO is an object detection algorithm based on deep learning. Its core idea is to treat the object detection task as a regression problem and predict multiple bounding boxes and their corresponding class probabilities on an image simultaneously through a single neural network. Specifically, the YOLO model divides the input image into 16×16 grids, each of which is responsible for predicting a fixed number of bounding boxes and their confidences. The prediction of the bounding box can be expressed as: i =(x, y, w, h, p), where (x, y) is the coordinate of the center of the bounding box relative to the grid, w and h are the width and height of the bounding box, and p is the confidence that an object exists in the box. YOLO calculates the confidence p and category probability C of each bounding box and finally obtains the category prediction of each object.

[0036] 2. Geometric constraints

[0037] The symbols used in this chapter are as follows:

[0038] t=[t x , t y , t z ] T : The position of the object coordinate system in the world coordinate system, where t x ,t y ,t z are the position coordinates in the x, y, and z directions respectively.

[0039] s=[s l ,s w ,s h ] T : The half side length of the 3D bounding box, i.e. the scale of the object, where s l 、sw 、s h are the length, width, and height of the object respectively.

[0040] First, the data association process traverses all map points in the current frame, and by calculating the Euclidean distance between each map point and the center of the object, it selects points whose Euclidean distance is within the set threshold. This distance filtering mechanism not only considers the relative position in space, but also combines the category information of the object, thereby effectively limiting the range of valid map points. This preprocessing step is particularly important when dealing with complex scenes, because it can significantly reduce the complexity of subsequent calculations and improve the accuracy of object recognition.

[0041] When dealing with certain categories of objects, the preprocessing step of geometric constraint verification is particularly important. This process is not only the key to ensure recognition accuracy, but also the basis for improving the robustness of the system in dynamic environments.

[0042] The details of the algorithm are as follows Figure 5 As shown, first of all, it involves the limitation of the overall size of the object. The maximum boundary of the object is defined by the Euclidean distance between the vertex and the center of the object:

[0043]

[0044] The calculation of this maximum boundary provides an important geometric reference for subsequent point cloud screening. Next, we need to analyze the relationship between the newly added map points and the center of the object. The Euclidean distance from the newly added map point to the center of the object is expressed as:

[0045]

[0046] Among them, m is the map point, p is the center of the object, and p l is the x-axis coordinate of the center of the object, p w is the y-axis coordinate of the center of the object, p h is the z-axis coordinate of the center of the object, m l is the x-axis coordinate of the map point, m w is the y-axis coordinate of the map point, m h is the z-axis coordinate of the map point. By calculating this distance, we can evaluate whether the newly added map point is within the effective range of the object. Specifically, when Dis(m)>ObjectMax·8, delete the map point.

[0047] Secondly, for basic common objects, we further defined and restricted their geometric constraints to ensure that the characteristics of the objects can be accurately reflected during the recognition process. Taking length as the benchmark value, we set the corresponding ratio of width and height to length for different types of objects:

[0048] prop = [1, prop wl , prop hl ] T #(4)

[0049] where prop wl is the ratio of the width to the length of the object, prop hl It is the ratio of the height to the length of the object. For different objects, we customize different ratios of width to height and length to height, that is, we set a ratio for different objects based on experience.

[0050] First, calculate the distance between the newly added map point and the center point of the object in the width direction:

[0051]

[0052] And the distance between the newly added map point and the center point of the object in the height direction:

[0053]

[0054] By calculating these distances, we can further evaluate whether the newly added points conform to the geometric characteristics of the object. w (m)>prop wl ·s l or Dis h (m)>prop hl ·s l When, s l The length of the object indicates that the length, width and height ratio of the newly added point far exceeds the ratio range of the basic object. Therefore, it can be reasonably inferred that the point cannot be a point in the object. Based on this judgment, the system will ignore the point to further improve the accuracy and reliability of recognition.

[0055] 3. Outlier removal

[0056] In practical applications, the 2D bounding box is often inconsistent with the actual boundary of the object, which is a common problem in object detection and recognition tasks. The bounding box of an object may be deviated due to many factors, such as the camera's viewing angle, the complexity of the object's shape, and environmental interference. Therefore, in order to improve the accuracy of object recognition, it is necessary to remove abnormal points from the feature points in the point cloud to ensure the reliability and effectiveness of the point cloud used.

[0057] Geometric feature points are usually randomly distributed within the bounding box of an object, and sometimes even appear sparse. This random distribution makes the presence of outliers more significant, especially in complex environments where feature points may be affected by factors such as occlusion, reflection, or illumination changes. These outliers not only interfere with the subsequent object recognition process, but may also cause errors in the recognition results, thereby affecting the performance of the overall system.

[0058] To solve this problem, we further improved the EIF algorithm for the SLAM system and used the Improved Extended Isolation Forest (IEIF) algorithm to remove abnormal points in the projected point cloud for further processing. IEIF is a tree-based anomaly detection algorithm that can effectively identify abnormal points in the point cloud by building a random tree. The core idea is to gradually isolate the data points in the point cloud by random segmentation. Compared with the traditional Isolation Forest, IEIF performs better in processing high-dimensional point clouds and can better capture the complex structure and distribution characteristics of point clouds.

[0059] The core idea of ​​the algorithm is to recursively divide the data space into a series of isolated data points, and then regard those points that are easy to be isolated as outliers. We take the feature points within the bounding box of the object as the input of the algorithm. The algorithm first performs multiple random splits on each feature point to build multiple trees. The outlier function of each point will be recorded and used to calculate the anomaly score of the point. As shown in the algorithm, we first create t isolated trees using the point cloud of the object, and then calculate the anomaly score by calculating the path length of each point x∈X, where the anomaly score function s(x, n) is defined as follows:

[0060]

[0061] Where x∈X is a point in the point cloud, n is the number of points in the point cloud, and E(h(x)) is the average depth reached by a single data point x in all trees. c(n) is a normalization factor defined as the average depth of an unsuccessful search in a binary search tree, calculated as:

[0062]

[0063] H(n-1) is the harmonic function, and the calculation formula is:

[0064] H(n-1)≈ln+0.5772156649#(9)

[0065] Where n is the number of points used to build the tree.

[0066] like Figure 2As shown in the figure, IEIF and IF are used to remove outliers from the keyboard point cloud respectively. It is obvious that IEIF is better than the IF algorithm. Occlusion and overlap between objects may also lead to the appearance of outliers, especially in complex scenes, where a single object usually corresponds to multiple outliers. However, some of these outliers may come from other objects, which makes the removal of outliers more complicated. We regard these as outlier clusters that need to be detected and removed. Through the characteristics of IEIF, we can better identify these outliers, thereby improving the accuracy of object recognition and processing.

[0067] 4. Experiment

[0068] We tested on the open-source computer vision datasets of the Technical University of Munich, which are all indoor office image sequences containing desks, chairs, monitors, books, and other common household items. Among them, fr1_desk contains a desktop scene, which mainly captures a fixed desktop environment with some objects on the desktop (such as books, cups, etc.). fr2_desk also revolves around the desktop scene, but compared with fr1_desk, the environment of fr2_desk is more complex and contains more objects and details. fr3_long_office_household contains a longer office scene with a larger space and more environmental details, involving more dynamic changes and complex object layouts. These three datasets are designed for scenes of different complexity and dynamics, providing rich RGB-D data, which are suitable for evaluating and comparing the performance of our algorithm with existing algorithms. All experiments were conducted on a server equipped with an Intel Xeon Platinum 8255C CPU processor and an RTX2080Ti graphics card, with an operating system of Ubuntu 18.04.

[0069] Qualitative experiments such as Figure 3 As shown in Fig. 3, we have made a visual comparison of the original images and sparse semantic maps of the TUM dataset. As shown in Fig. 3, the first to third columns correspond to the TUM fr1_desk, TUM fr2_desk, and TUM fr3_long_office_household sequences respectively. There are four rows of images in each column, showing four different views of the TUM sequence: the first row is the original image, the second row is the semantic map generated by our algorithm, and the next two rows show the similarities and differences between EAO-SLAM and our algorithm in terms of the accuracy of the bounding box annotation.

[0070] The semi-dense map shows the ability of our algorithm in extracting semantic information and can effectively identify key objects in the scene. (c) Although the results of EAO-SLAM can identify objects, their bounding boxes are often not accurate enough and have certain deviations, resulting in a decrease in the match between the object model and the actual object. In contrast, our method is particularly outstanding in the accuracy of the bounding box, and the match between the generated object model and the real object is significantly improved. By combining object geometric constraints and optimized outlier removal strategies, our algorithm can more accurately locate and reconstruct objects, ensuring that the bounding box better reflects the true shape and position of the object. This accuracy not only improves the reliability of object recognition, but also provides a more solid foundation for subsequent navigation and task execution. In general, the experimental results show that our method is significantly better than the existing EAO-SLAM algorithm in object detection and reconstruction in complex environments, showing its potential in practical applications.

[0071] Next, we compared the absolute trajectory error (RMSE) of our algorithm with ORB-SLAM2 and EAO-SLAM on different TUM sequences. Table 1 shows the performance comparison of the three algorithms in multiple scenarios:

[0072] Table1: COMPARISON OF RMSE FOR ATE IN SEVERAL TUM SEQUENCES

[0073]

[0074] As can be seen from Table 1, in the fr2_desk and fr3_office sequences, our method significantly reduces the absolute trajectory error. This result not only demonstrates the robustness of our algorithm in complex environments, but also shows its effectiveness in dealing with outliers in dynamic scenes. This improvement is attributed to the geometric constraints and optimized outlier removal strategy we introduced, which enables the system to more accurately identify and locate objects, thereby reducing the accumulation of errors.

[0075] In the fr1_desk sequence, our method performs slightly worse than ORB-SLAM2 and Object-OrientedSemantic SLAM, but is still better than EAO-SLAM overall. The scene of the fr1_desk sequence is relatively simple and contains fewer features, which results in our algorithm's advantages in feature extraction and matching not being fully utilized. In feature-rich scenes such as fr2_desk and fr3_office, our method is able to better utilize geometric constraints and optimization strategies. This shows that although ORB-SLAM2 and Object-Oriented Semantic SLAM perform well in certain specific scenarios, our method has obvious advantages in dealing with complex scenes, especially in terms of object detection and reconstruction accuracy, which enables our algorithm to maintain good performance in a variety of environments.

[0076] like Figure 4 As shown in the figure, in fr3_office we compared our algorithm with EAO-SLAM. First, the RMSE FOR ATE of our algorithm is 1.0863, and the RMSE FOR ATE of EAO-SLAM is 1.2397. The dotted line in the figure is the ground truth, the blue line is the trajectory predicted by EAO-SLAM, and the orange line is the curve predicted by our algorithm. We can see that when the camera changes significantly, the error of our algorithm is smaller than that of EAO-SLAM, which means that our algorithm is more stable.

[0077] In this study, we conducted ablation experiments on geometric constraints, outlier removal, and the YOLO algorithm in the fr3_desk sequence to evaluate the impact of each component on the absolute trajectory error (ATE). The experimental results show the RMSEFOR ATE values ​​under different configurations, as shown in Table 2:

[0078] Table2: COMPARISON OF RMSE FOR ATE IN ABLATION STUDY

[0079]

[0080]

[0081] Without using any improvements, that is, the EAO-SLAM method, the RMSE is 1.2397. This result is used as a baseline, indicating that the positioning accuracy of the system is significantly reduced without any processing. When only YOLOv8 is used without enabling geometric constraints and outlier removal, the RMSE is 1.1023. This result shows that YOLOv8 is significantly better than YOLOv3 in object detection performance and can provide more accurate object positioning information, thus laying the foundation for the subsequent SLAM process. When the system only enables geometric constraints, the RMSE is 1.1259. Although geometric constraints can improve positioning accuracy to a certain extent, their effect is not significant in the absence of accurate object detection information. This shows that relying solely on geometric constraints without effective object detection may not make full use of environmental information. When only outlier removal is used, the RMSE is 1.1185. This result shows that outlier removal is effective in dealing with noisy and inconsistent data, but its improvement effect is limited in the absence of object detection information. When YOLOv8, geometric constraints, and outlier removal are enabled at the same time, the RMSE drops to 1.0863. This result shows that the combination of the three significantly improves the overall performance of the system. The high-quality object detection information provided by YOLOv8, combined with the advantages of geometric constraints and outlier removal, enables the system to more accurately locate and map the environment.

Claims

1. An object-level semantic visual SLAM method based on object geometric constraints and outlier removal, characterized in that: Includes steps: Step S1: Use the YOLOv8 algorithm to perform object detection on the acquired image, identify various objects in the environment, and generate corresponding two-dimensional bounding boxes; Step S2: assign a category label to each detected object; Step S3: Generate a dense point cloud based on the depth information to represent the three-dimensional structure in the environment; for each detected object, establish its geometric model, including the object's center of mass, boundary and size information, and establish geometric constraints. The specific steps are as follows: Step S31: Calculate the Euclidean distance between each map point and the object centroid to determine which points in the point cloud belong to a specific object; the distance calculation formula is: Among them, m is the map point, p is the center of the object, and p l is the x-axis coordinate of the center of the object, p w is the y-axis coordinate of the center of the object, p h is the z-axis coordinate of the center of the object, m l is the x-axis coordinate of the map point, m w is the y-axis coordinate of the map point, m h is the z-axis coordinate of the map point; Step S32: when Dis(m)>ObjectMax·8, where ObjectMax is the maximum range value of the current object, it indicates that the point cannot belong to the object and is ignored; Step S33: Using the length as a reference value, set the corresponding ratio between the width and height and the length for different types of objects: prop=[1,prop wl ,prop hl ] T where prop wl is the ratio of the width to the length of the object, prop hl It is the ratio of the height to the length of the object. Different width to height ratios and length to height ratios can be customized for different objects. Step S34: Calculate the distance D between the newly added map point and the center point of the object in the width direction: Add the distance between the map point and the center point of the object in the height direction: When Dis w (m)>prop wl ·s l or Dis h (m)>prop hl ·s l When, s l is the length of the object, it is inferred that the point cannot be a point in the object; Step S4: Introduce an outlier removal thread. The specific steps are as follows: Step S41: projecting the filtered point cloud data onto the geometric model of the object; Step S42: Detect outliers on the projected point cloud; the algorithm calculates the outlier score based on the distribution characteristics of the points. The calculation formula for outliers is: Where x∈X is a point in the point cloud, n is the number of points in the point cloud, E(h(x)) is the mean depth reached by a single data point x in all trees; c(n) is a normalization factor, defined as the average depth of an unsuccessful search in a binary search tree, and is calculated as: H(n-1) is the harmonic function, and the calculation formula is: H(n-1)≈ln(i)+0.5772156649Step S43: Eliminate points with anomaly scores higher than 0.65.