Dynamic Feature Point Elimination Method Combining Semantic Information and Geometric Constraints

By combining semantic information and geometric constraints in the visual SLAM system, the problems of real-time and positioning accuracy of SLAM systems in dynamic scenarios are solved, and more efficient and real-time robot positioning is achieved.

CN116310799BActive Publication Date: 2025-05-27CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310115772.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-05-27
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

The prior art ignores the real-time nature of the SLAM system in dynamic scenarios, resulting in an increase in time consumption, the system cannot run in real time, the robot consumes high power, and there is a problem of "over-exclusion" of feature points, which affects the positioning accuracy.

Method used

The dynamic feature point removal method combined with semantic information and geometric constraints is adopted. By running the lightweight YOLOv5 target detection network in parallel in the ORB-SLAM2 system framework, the image ORB feature points are extracted, and the dynamic nature of feature points is judged using the motion consistency detection algorithm for polar geometric constraints. Finally, the feature points of moving objects are accurately eliminated in the dynamic feature point removal module.

Benefits of technology

It improves the positioning accuracy of visual SLAM in dynamic scenarios, reduces the RMSE value of absolute trajectory error and relative pose error, and improves the real-time and battery life of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310799B_ABST
    Figure CN116310799B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and particularly to a method for removing dynamic feature points by combining semantic information and geometric constraints. The steps include: creating an object detection thread in the ORB-SLAM2 system framework and designing a YOLOv5 object detection network; extracting ORB feature points of the image and using an epipolar geometry constraint-based motion consistency detection algorithm to judge the dynamics of the feature points; proposing a dynamic feature point removal module that combines semantic information and geometric constraints, and finally determining the dynamic objects in the object detection bounding box, thereby realizing the effective removal of dynamic feature points. The method of the present invention can quickly identify the object categories in the scene. Through verification, it is proved that the RMSE values of the absolute trajectory error, translational and rotational relative pose errors of the method of the present invention on the TUM dataset are respectively reduced by 97.71%, 95.10% and 91.97% compared with ORB-SLAM2.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a dynamic feature point elimination method combining semantic information with geometric constraints. Background Art

[0002] Simultaneous Localization and Mapping (SLAM) technology refers to a mobile robot equipped with specific sensors starting from any position in an unknown environment, and calculating the distance and angle between the landmarks in the field of view and the robot based on the observed landmarks during the movement, thereby locating its own position and posture, and incrementally building a map of the surrounding environment based on its own positioning. SLAM technology is a key technology for robots to perceive their own state and external environment, and it is also the key and foundation for robots to complete complex tasks such as environmental perception, autonomous positioning and navigation, path planning, and human-computer interaction.

[0003] Visual SLAM is widely used because of its low cost and ability to obtain rich environmental information. It can ensure its robustness and efficiency in static environments. However, in actual dynamic scenes, due to the interference of feature points of dynamic objects in the scene, mismatching occurs when matching feature points between visual SLAM image frames, and the estimated pose information has a large error, which ultimately leads to serious deviations between the robot's motion trajectory and map construction. At present, dynamic scene visual SLAM technology at home and abroad has eliminated the feature points of moving objects to a certain extent, reduced the interference of dynamic feature points, and improved positioning accuracy. However, most technologies do not consider the feature points of potential moving objects, such as temporarily parked vehicles. If the feature points of an object are judged as dynamic feature points and eliminated based only on prior category information, there will be a problem of "over-elimination" of feature points, which is also likely to lead to low positioning accuracy of visual SLAM. In addition, the current dynamic SLAM technology ignores the real-time nature of the SLAM system, increases time consumption, causes the system to be unable to run in real time, and the robot consumes high power. Therefore, the research on dynamic feature point elimination that combines semantic information with geometric constraints has extremely important practical application value. At the same time, semantic information can also be used to construct a three-dimensional semantic map to achieve high-level perception of the robot to the scene, which is conducive to the intelligent robot to complete advanced tasks such as human-computer interaction.

[0004] The purpose of the present invention is to study a dynamic feature point elimination method suitable for visual SLAM in dynamic scenes, so as to enable the robot to accurately estimate the position and obtain accurate positioning information. The present invention is committed to carrying out research from four aspects, namely, improving lightweight target detection network, feature point extraction, motion consistency detection, and dynamic feature point elimination, in an effort to improve the positioning accuracy of dynamic scene visual SLAM to meet the positioning requirements of mobile robots. Summary of the invention

[0005] The purpose of the present invention is to provide a dynamic feature point removal method combining semantic information with geometric constraints, which is used to solve the technical problems in the prior art that the real-time performance of the SLAM system is ignored, the time consumption is increased, the system cannot run in real time, and the robot consumes high power.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The present invention provides a dynamic feature point elimination method combining semantic information with geometric constraints, comprising the following steps:

[0008] S1: Create a target detection thread in the ORB-SLAM2 system framework to run in parallel, and design a lightweight YOLOv5 target detection network to provide semantic information of objects for visual SLAM;

[0009] S2: Extract image ORB feature points and use the motion consistency detection algorithm constrained by epipolar geometry to determine the dynamics of the feature points;

[0010] S3: A dynamic feature point removal module combining semantic information and geometric constraints is proposed to finally determine the dynamic objects in the target detection bounding box and realize the effective removal of dynamic feature points;

[0011] S4: Use the retained static feature points for pose estimation to improve the SLAM positioning accuracy in dynamic scenes.

[0012] Further, S1 comprises the following steps:

[0013] S100: Improved ShuffleNetv2 network of ORB-SLAM2 system framework:

[0014] The Ghost module and SE module are introduced into the basic unit block with a step size of 1;

[0015] Add SE module to the basic unit block with a step size of 2;

[0016] And depth-wise separable convolution is added to the two basic unit blocks with a step size of 1 and a step size of 2, so as to obtain two improved basic unit blocks with different step sizes;

[0017] S101: Replace the backbone network of YOLOv5 with the improved ShuffleNetv2 network:

[0018] S102: The backbone network uses the Hard-Swish activation function instead of ReLu;

[0019] The neck uses the PAN structure to obtain feature maps of three different scales, and uses the CSP structure to connect and fuse adjacent feature maps;

[0020] S103: Use SIoU loss function during training;

[0021] S104: Perform ShuffleNet-YOLOv5 target detection model training using the data set to obtain semantic information.

[0022] Furthermore, the SIoU loss function consists of four Cost functions: Angle cost, Distance cost, Shape cost, and IoU cost;

[0023] The total loss function is:

[0024] Loss = W box L box +W cls L cls

[0025] Among them, L cls Use Focal Loss as the classification prediction loss, W box , W cls are the weights of the bounding box and classification losses, respectively, and L box is the coordinate prediction loss of the bounding box, as follows:

[0026]

[0027] Δ=∑ t=x,y (1-e -(2-Λ)ρt )

[0028]

[0029] Where Δ is the Distance cost function with the angle penalty Angle cost added, where the expression of Angle cost is as shown in the above formula Λ, Ω is the Shape cost function, and IOU is the IoU cost function.

[0030] Furthermore, the motion consistency detection algorithm using epipolar geometry constraints described in S2 to determine the dynamics of feature points includes:

[0031] S200: Extract feature points using ORB feature extraction algorithm;

[0032] S201: Assume two frames of image I 1 ,I 2 In the movement between 1 and O 2 is the optical center of the two cameras, and the spatial point P is at I 1 ,I 2 The feature matching point positions P are obtained respectively1 and P 2 ;O 1 , O 2 The plane defined by the three points O and P is called the polar plane. 1 O 2 It is called the baseline, the polar plane and the two image planes I 1 ,I 2 The intersection line between 1 , l 2 That is the polar line;

[0033] Feature point P 1 and P 2 The normalized coordinates of 1 =[u 1 ,v 1 ,1] T ,P 2 =[u 2 ,v 2 ,1] T ;

[0034] S202: Calculate polar lines:

[0035] The polar line l can be calculated by the following equation 1 :

[0036]

[0037] In the formula, the basic matrix F is calculated based on the pixel positions of the paired feature points;

[0038] If a spatial point P is a static point, then it satisfies the standard constraint formula P defined in the static scene in multi-view geometry 2 T FP 1 =0, and the feature points extracted from the dynamic target will violate the above constraints. Therefore, the correctness of the matching is distinguished according to whether the feature points violate the constraints;

[0039] S203: Determine the static and dynamic state of feature points:

[0040] The distance D from a point to the epipolar line is calculated by the following formula:

[0041]

[0042] If D is less than a preset threshold, point P is considered static, otherwise it is dynamic.

[0043] Furthermore, S3 includes:

[0044] S300: Obtain the category and bounding box result of the object according to the method described in S1, record the potential dynamic points according to the method described in S2, and propose a dynamic feature elimination module to accurately eliminate the feature points of the moving object;

[0045] S301: According to the target detection threads running in parallel to obtain the prediction results of the objects, according to the motion attributes of the objects, the objects are divided into three categories according to the category labels: high dynamic (people, vehicles, etc. have motion attributes), low dynamic (books, chairs, etc. have moving attributes) and static (tables, etc. have static attributes);

[0046] S302: Initial judgment:

[0047] If the feature points in the YOLOv5 bounding box are only located on low-dynamic or static objects, no detailed judgment processing is performed on the feature points; detailed judgment is performed only when the feature points are within the frame of high-dynamic objects;

[0048] S303: Detailed judgment:

[0049] At this time, for the feature points in the high-dynamic object frame, it is necessary to consider whether the potential feature points detected by the previous motion consistency are also in the high-dynamic object frame. If the potential feature points are in the high-dynamic object frame at the same time and the number of its feature points is greater than a certain threshold, the object is determined to be in a moving state and all feature points on the object are discarded;

[0050] S304: Estimate the camera pose using the static feature points in the scene to complete the overall positioning of the SLAM system.

[0051] The present invention has at least the following beneficial effects:

[0052] The invention firstly adds a designed lightweight YOLOv5 target detection network to the front end of the system as an independent thread for parallel processing based on the open source framework ORB-SLAM2. The method can quickly identify the object category in the scene, provide semantic information of the object category for SLAM, extract the ORB feature points of the image in the tracking thread, and use the motion consistency detection algorithm based on the epipolar geometry to record the potential dynamic points. Then, in the dynamic feature point elimination module, the objects are divided into three categories of high dynamic (people, vehicles, etc.), low dynamic (books, chairs, etc.) and static according to the category labels of the target detection results, and the number of potential dynamic points detected by the previous motion consistency falling in the segmented objects is considered, so as to accurately eliminate the feature points of the moving object and effectively retain the static feature points, and then use the retained static feature points to perform pose estimation. The RMSE values ​​of the absolute trajectory error (ATE), translation and rotation relative pose error (RPE) of the method on the TUM data set are reduced by 97.71%, 95.10% and 91.97% respectively compared with ORB-SLAM2. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.

[0054] Figure 1 It is a flow chart of the dynamic feature point elimination method combining semantic information with geometric constraints disclosed in the present invention;

[0055] Figure 2 (a) and (b) are diagrams of the improved basic unit block of ShuffleNetv2;

[0056] Figure 3 This is the improved ShuffleNet-YOLOv5 network structure diagram;

[0057] Figure 4 is the epipolar geometry constraint relationship diagram;

[0058] Figure 5 (a) and (b) are the effects of dynamic feature point removal;

[0059] Figure 6 (a) and (b) are trajectory comparison diagrams on the fr3_w_xyz sequence;

[0060] Figure 7 (a) and (b) are trajectory comparison diagrams on the fr3_w_half sequence. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0062] See also Figure 1 , is a flow chart of a dynamic feature point elimination method combining semantic information with geometric constraints disclosed in the present invention, comprising the following steps:

[0063] S1: Created a target detection thread in the ORB-SLAM2 system framework to run in parallel, and designed a lightweight YOLOv5 target detection network to provide semantic information of objects for visual SLAM;

[0064] S100: In order to make the model lighter and more accurate, the Ghost module and SE (Squeeze-and-Excitation Networks) module in the GhostNet network are introduced on the basis of the ShuffleNetv2 basic unit block with a step size of 1, and the SE module is added to the basic unit block with a step size of 2, and depthwise separable convolution is added to both unit blocks, thus obtaining two improved ShuffleNetv2 basic unit blocks with different step sizes. Figure 2 This is the improved basic unit block diagram of ShuffleNetv2.

[0065] S101: Replace the backbone network of YOLOv5 with the improved lightweight ShuffleNetv2 network.

[0066] S102: The backbone network uses the Hard-Swish activation function instead of ReLu, which is faster in calculation. The neck uses the PAN (Path Aggregation Network) structure to obtain feature maps of three different scales, and uses the CSP structure to connect and fuse adjacent feature maps. Figure 3 This is the improved ShuffleNet-YOLOv5 network structure diagram.

[0067] S103: The effectiveness of object detection depends largely on the definition of the loss function. Minimizing the loss during training can match the predicted box of the object with the corresponding real box. A new loss function SIoU is used in the training process, in which the penalty indicator is redefined taking into account the vector angle between the required regressions. The SIoU loss function consists of 4 Cost functions: Angle cost, Distance cost, Shape cost, I oU cost. The total loss function is:

[0068] Loss = W box L box +W cls L cls

[0069] Among them, L cls Use Focal Loss as the classification prediction loss, W box , W cls are the weights of the bounding box and classification losses, respectively, and L box is the coordinate prediction loss of the bounding box, as follows:

[0070]

[0071] Δ=∑ t=x,y (1-e -(2-Λ)ρt )

[0072]

[0073] Where Δ is the distance cost function with angle penalty Anglecost added, where the expression of Angle cost is as shown in the above formula Λ, Ω is the shape cost function, and IOU is the IoU cost function. SIoU loss adding angle penalty function effectively reduces the total degrees of freedom of loss, improves the training speed and inference accuracy.

[0074] S104: Train the ShuffleNet-YOLOv5 target detection model on the NVIDIA RTX8000 48G graphics card, using the Microsoft coco2017 dataset to obtain semantic label information.

[0075] S2: Extract image ORB feature points in the tracking thread at the same time, and use the motion consistency detection algorithm constrained by epipolar geometry to determine the dynamics of the feature points;

[0076] In specific implementation, the motion consistency algorithm based on epipolar geometry constraints includes the following specific steps:

[0077] S200: extracting feature points using the ORB feature extraction algorithm.

[0078] S201: Based on the ORB features extracted in S200, such as Figure 4 The epipolar geometric constraint relationship, assuming that the two frames of image I 1 ,I 2 Movement between 1 and O 2 is the optical center of the two cameras, and the spatial point P is at I 1 ,I 2 The feature matching point positions P are obtained respectively 1 and P 2 .O 1 , O 2 The plane defined by the three points O, P is called the polar plane. 1 O 2 It is called the baseline, the polar plane and the two image planes I 1 ,I 2 The intersection line between 1 , l 2 The characteristic point P 1 and P 2 The normalized coordinates of 1 =[u 1 ,v 1 ,1] T ,P 2 =[u2 ,v 2 ,1] T .

[0079] S202: Calculate the polar line. The polar line l can be calculated by the following equation 1 .

[0080]

[0081] In the formula, the basic matrix F can be calculated according to the pixel positions of the paired feature points. If the spatial point P is a static point, it satisfies the standard constraint formula P defined in the static scene in multi-view geometry. 2 T FP 1 =0, and the feature points extracted from the dynamic target will violate the above constraints. Therefore, the correctness of the matching can be distinguished based on whether the feature points violate the constraints.

[0082] S203: Determine the static and dynamic state of the feature point. 2 Not on the polar line l 2 The spatial point P may be a dynamic feature point. Calculate the distance D from the point to the epipolar line. The following formula is the epipolar line constraint formula of epipolar geometry.

[0083]

[0084] If D is less than a preset threshold, point P is considered static, otherwise it is dynamic.

[0085] S3: A dynamic feature point removal module combining semantic information and geometric constraints is proposed to finally determine the dynamic objects in the target detection bounding box, thereby achieving effective removal of dynamic feature points;

[0086] In specific implementation, dynamic feature points are eliminated based on semantic information and geometric constraints, and the category and bounding box of the object are obtained by target detection. Directly removing all feature points that fall into the bounding box is a simple way to reduce the interference of dynamic objects, but it will remove some static feature points contained in the bounding box. When dynamic objects occupy a large part of the image, direct elimination will result in too few feature points tracked by the system, causing the system to fail. In addition, in addition to eliminating the moving objects themselves, it is also necessary to judge potential moving objects, such as chairs moved by people, temporarily parked vehicles, etc. Therefore, the epipolar geometry constraint method described in claim 3 is combined to accurately judge the dynamic nature of the feature points, thereby achieving the purpose of greatly improving the positioning accuracy in a dynamic environment. To this end, after comprehensive consideration, S3 includes the following specific steps:

[0087] S300: Based on the method described in S1, the category and bounding box results of the object are obtained. Based on the method described in S2, a dynamic feature elimination module is proposed to accurately eliminate the feature points of the moving object.

[0088] S301: The target detection threads run in parallel to obtain the prediction results of the objects. According to the motion properties of the objects, the objects are divided into three categories according to the category labels: high dynamics (people, vehicles, etc. have motion properties), low dynamics (books, chairs, etc. have moved properties) and static (tables, etc. have static properties).

[0089] S302: Initial judgment. If the feature point in the YOLOv5 boundary frame is only located on a low-dynamic or static object, no detailed judgment processing is performed on the feature point. Detailed judgment is performed only when the feature point is within a high-dynamic object frame.

[0090] S303: Detailed judgment. At this time, for the feature points in the high-dynamic object frame, it is necessary to consider whether the potential feature points detected by the previous motion consistency are also in the high-dynamic object frame. If the potential feature points are in the high-dynamic object frame at the same time and the number of their feature points is greater than a certain threshold, the object is determined to be in a moving state and all feature points on the object are discarded.

[0091] S304: Finally, the camera position is estimated using the static feature points in the scene to complete the positioning.

[0092] S4: Use the retained static feature points for pose estimation to improve the SLAM positioning accuracy in dynamic scenes.

[0093] Verify the dynamic feature point elimination method of semantic information and geometric constraints disclosed in the present invention: Figure 5 (a) shows the effect of not removing dynamic feature points. Figure 5 (b) shows the effect of removing dynamic feature points. It can be seen that the feature points of dynamic objects are successfully removed. Some red potential dynamic objects in the figure are not moved, and the static feature points are effectively retained.

[0094] Verify the positioning accuracy, robustness and time efficiency of the SLAM system after removing dynamic feature points in high dynamic scenes: test the dynamic object class "walking" sequence of the TUM data set, which includes 4 groups of high dynamic scenes: fr3_walking_static, fr3_walking_xyz, fr3_walking_rpy, and fr3_walking_half, where xyz, rpy, halfsphere, and static correspond to the movement of the camera. The calculation results are evaluated by the online test tool provided by TUM to evaluate the accuracy of the SLAM system pose estimation. Figure 6 This is the visualization trajectory error result of the method of the present invention and the ORB-SLAM2 method on the fr3_walking_xyz sequence. Figure 7This is the visualized trajectory error result of the method of the present invention and the ORB-SLAM2 method on the fr3_walking_half sequence. The black line in the figure represents the true value of the camera trajectory, the blue line represents the camera trajectory of motion estimation, and the red line represents the difference between the true trajectory and the estimated trajectory. The absolute trajectory error comparison results of the method of the present invention and ORB-SLAM2 are shown in Table 1.

[0095] In the quantitative analysis, the root mean square error (RMSE) and standard deviation error (SD) of the absolute trajectory error are used as evaluation indicators. RMSE represents the deviation between the estimated value and the true value, and SD reflects the dispersion of the estimated camera trajectory. From the absolute trajectory error in Table 1, it can be seen that the RMSE and SD values ​​of the proposed method in the fr3_walking_xyz sequence are reduced by 97.71% and 97.81% respectively compared with ORB-SLAM2, and are also significantly reduced in other sequences.

[0096] Table 1. Comparison results of absolute trajectory error ATE (unit: m)

[0097]

[0098] Verify the time efficiency of the SLAM system after removing dynamic feature points in a high dynamic scene: the time required for tracking thread processing is tested on the fr3_walking_xyz sequence on the TUM data set. The time consumption data of the tracking thread can be seen from Table 2.

[0099] Table 2 Comparison of average time for tracking threads (unit: ms)

[0100]

[0101] Combined with the above, it can be seen that: in terms of algorithm time, adding new threads in the SLAM system will increase time consumption, and the main time consumption of the present invention is used for the target detection thread. The lightweight target detection model designed by the present invention has small calculation amount and high precision, and adopts multi-threaded parallel operation, which reduces part of the time consumption. The system as a whole can reach a processing speed of 20 frames / s, which has little impact on the real-time performance of the system and can meet the real-time requirements of the algorithm. The method of the present invention achieves excellent performance using only the CPU when processing high-dynamic scenes, indicating that it better meets the real-time requirements.

[0102] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.

Claims

1. A dynamic feature point elimination method combining semantic information and geometric constraints, characterized in that, it includes the following steps: S1: Create a target detection thread to run in parallel in the ORB-SLAM2 system framework, and design a lightweight YOLOv5 target detection network to provide semantic information of objects for visual SLAM; S2: Extract ORB feature points of the image, and use the motion consistency detection algorithm based on epipolar geometry constraints to judge the dynamics of the feature points; S3: Propose a dynamic feature point elimination module combining semantic information and geometric constraints, and finally determine the dynamic objects in the target detection bounding box, realizing the effective elimination of dynamic feature points; S4: Use the remaining static feature points for pose estimation to improve the SLAM positioning accuracy in dynamic scenarios; S1 includes the following steps: S100: Improve the ShuffleNetv2 network of the ORB-SLAM2 system framework: Introduce the Ghost module and the SE module into the basic unit block with a stride of 1; Add the SE module to the basic unit block with a stride of 2; And add depthwise separable convolutions to both the basic unit blocks with a stride of 1 and a stride of 2, obtaining two improved basic unit blocks with different strides; S101: Use the improved ShuffleNetv2 network to replace the backbone network of YOLOv5; S102: The backbone network uses the Hard-Swish activation function instead of ReLu; The neck uses the PAN structure to obtain feature maps of three different scales, and uses the CSP structure to perform feature connection and fusion on adjacent feature maps; S103: Use the SIoU loss function during the training process; S104: Conduct training on the ShuffleNet-YOLOv5 target detection model, use the dataset for training, and obtain semantic information; S3 includes: S300: Obtain the category and bounding box results of the object according to the method described in S1, record the potential dynamic points according to the method described in S2, and propose a dynamic feature elimination module to accurately eliminate the feature points of the moving object; S301: Obtain the prediction results of the object according to the parallel operation of the target detection thread, and divide the object into three categories: high-dynamic, low-dynamic, and static according to the class label according to the motion attributes of the object; S302: Initial judgment: If the feature points in the YOLOv5 bounding box are only located on low-dynamic or static objects, no fine judgment processing is performed on the feature points; fine judgment is only performed when the feature points are within the high-dynamic object box; S303: Fine judgment: At this time, for the feature points in the high-dynamic object box, it is necessary to consider whether the potential feature points detected by the previous motion consistency are also in the high-dynamic object box. If the potential feature points are simultaneously in the high-dynamic object box and the number of its feature points is greater than a certain threshold, then the object is determined to be in a moving state, and all feature points on the object are discarded; S304: Use the static feature points in the scene to estimate the camera pose and complete the overall positioning of the SLAM system.

2. The dynamic feature point elimination method combining semantic information and geometric constraints according to claim 1, characterized in that, The SIoU loss function consists of 4 cost functions: Angle cost, Distance cost, Shape cost, and IoU cost; The total loss function is: , Among them, use Focal Loss as the classification prediction loss, , are the weights of the bounding box and classification losses respectively, is the coordinate prediction loss of the bounding box, as follows: , In the formula, is the Distance cost function with the added Angle cost, where the expression of the Angle cost is as shown in the above formula as shown is the Shape cost function, and IOU is the IoU cost function.

3. The dynamic feature point elimination method combining semantic information and geometric constraints according to claim 1, characterized in that, In S2, the algorithm for detecting the motion consistency using the epipolar geometry constraint to determine the dynamics of feature points includes: S200: Use the ORB feature extraction algorithm to extract feature points; S201: Set two frames of images During the movement between are the optical centers of two cameras, and the spatial point P in respectively obtains the positions of feature matching points and ; , , P The plane determined by the three points is called the epipolar plane, is called the baseline, and the intersection line between the epipolar plane and the two image planes is the epipolar line; Feature points and have normalized coordinates of ; S202: Calculate the epipolar line: The epipolar line can be calculated by the following equation :[[-END]] , In the formula, the fundamental matrix F is obtained based on the pixel positions of the paired feature points; If the spatial point P is a static point, it satisfies the standard constraint formula for the definition of a static scene in multi-view geometry , while the feature points extracted from dynamic objects will violate the above constraints. Therefore, the correctness of the matching is distinguished according to whether the feature points violate the constraints; S203: Determine the static or dynamic state of the feature points: Calculate the distance D from the point to the epipolar line through the following formula: , If D is less than the preset threshold, the point P is considered static, otherwise it is dynamic.