Implementation method of indoor environment DFPM-SLAM system based on dynamic feature point matching method

By introducing dynamic feature point matching technology in the indoor environment SLAM system, the accurate elimination of dynamic feature points is solved, and the tracking accuracy and robustness of the SLAM system is improved.

CN115546478BActive Publication Date: 2025-05-13BEIJING TECH & BUSINESS UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211028073.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-25
Publication Date
2025-05-13
Estimated Expiration
2042-08-25

AI Technical Summary

Technical Problem

In dynamic environments, it is difficult for the existing technology to accurately remove dynamic objects, resulting in inaccurate construction of indoor environment semantic maps and insufficient tracking efficiency and accuracy.

Method used

A DFPM-SLAM system based on dynamic feature point matching is proposed. Through the keyframe scoring selection mechanism, switching probability concepts and fuzzy logic, the state of dynamic feature points is accurately judged and updated, thereby eliminating dynamic feature points and improving tracking efficiency and accuracy.

Benefits of technology

It realizes accurate judgment and elimination of dynamic feature points, improves the tracking accuracy and robustness of the SLAM system in the indoor environment, and meets the real-time requirements of the system in the dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546478B_ABST
    Figure CN115546478B_ABST
Patent Text Reader

Abstract

The present invention is an implementation method of an indoor environment DFPM-SLAM system based on a dynamic feature point matching method, which is used for indoor positioning and map construction. The present invention implements an indoor environment DFPM-SLAM system based on an ORB-SLAM3 framework, adds a key frame scoring selection mechanism module to a tracking subsystem, and selects key frames for image scoring based on the number of feature points contained in the image, the rotation offset, and the number of frames separated from the nearest key frame; an MR-DFPM subsystem is set, and a switching probability method based on fuzzy thinking is used to determine the state of feature points, and a switching probability update method based on Kalman filtering is used to accurately divide dynamic / static feature points to eliminate dynamic objects in the scene. The method of the present invention can accurately eliminate dynamic objects to solve the problem of tracking loss, and the absolute trajectory error and relative posture error in indoor dynamic scenes are greatly improved, which can meet the real-time requirements of system operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of indoor semantic SLAM (Simultaneous Localization and Mapping), and specifically relates to constructing a real-time semantic mapping system for indoor dynamic environments based on a dynamic feature point matching (DFPM) method, referred to as a DFPM-SLAM system. Background Art

[0002] For a robot to achieve truly autonomous movement, it needs to have a certain ability to perceive and understand the environment, which is reflected in the robot's ability to independently build the surrounding environment and obtain relative positions in the built environment map. Accurate positioning requires accurate maps, and building accurate maps requires accurate position information of the robot itself, namely SLAM technology. With the continuous development of SLAM technology, it has now developed from traditional SLAM methods to visual SLAM (V-SLAM) systems based on different cameras. However, ordinary V-SLAM systems have problems such as inaccurate dynamic / static object judgment in non-static complex environments, resulting in loss of dynamic feature point tracking and large tracking errors.

[0003] The core problem of solving real-time semantic SLAM in indoor dynamic environments is how to accurately remove dynamic objects in the scene. At present, there are many methods that can be used for real-time semantic SLAM in dynamic environments. In 2017, Runz M et al. proposed an RGB-D dense map online SLAM system - Co-Fusion, which uses a multi-model fitting method to distinguish the foreground and background of the image, so that the moving objects in the foreground image are independent of the background, thereby achieving effective tracking of the moving target (Reference 1: Rünz M, Agapito L. Co-fusion: Real-time segmentation, tracking and fusion of multiple objects [C] / / 2017 IEEE International Conference on Robotics and Automation (ICRA). Washington, USA: IEEE, 2017: 4471-4478.). AN Lifeng et al. proposed a camera pose optimization method that combines the direct method and the feature point method. By calculating the total projection error of all pixel points belonging to a certain type of object between adjacent image frames, different weights are assigned to each type of object, and semantic information is introduced to reduce the impact of dynamic objects on tracking effects (Reference 2: AN Lifeng, ZHANG Xinyu, GAO Hongbo, et al. Semanticsegmentation–aided visual odometry for urban autonomous driving [J]. International Journal of Advanced Robotic Systems. 2017, 14(5): 1-11.). In 2018, Bescós B et al. proposed DynaSLAM, which is an improvement on the ORB-SLAM2 system for dynamic scenes. It uses instance semantic segmentation method to detect and remove possible moving objects in key frames, and perform background repair at the same time (Reference 3: Bescos B, Facil JM, Civera J, et al. DynaSLAM: Tracking, mapping, and inpainting in dynamic scenes [J]. IEEE Robotics and Automation Letters. 2018, 3 (4): 4076-4083.).In 2018, DS-SLAM proposed by YU Chao et al. from Tsinghua University uses a method that combines semantic information with dynamic feature point detection to improve the accuracy of camera pose estimation by filtering out dynamic targets in key frames. However, this method takes a long time to complete semantic segmentation, so its application in practical scenarios is limited (Reference 4: YU Chao, LIU Zuxin, IU Xinjun, et al. DS-SLAM: A semantic visual SLAM towards dynamic environments [C] / / 2018 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS). Washington, USA: IEEE, 2018: 1168-1174.).

[0004] In 2019, Palazzolo E et al. proposed a dynamic SLAM model - ReFusion. They first used a pure geometric filtering method to eliminate dynamic areas, and then used the matching residuals between the observed values ​​and the estimated values ​​to detect dynamic objects in the scene and determine the dynamic areas (Reference 5: Palazzolo E, Behley J, Lottes P, et al. Refusion: 3D reconstruction in dynamic environments for RGB-D cameras exploiting residuals [C] / / 2019 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS). Washington, USA: IEEE, 2019: 7855-7862.). VDO-SLAM[6] proposed by Zhang Jun et al. is a highly robust dynamic object perception SLAM system. When the target shape is unknown, the dynamic area is determined by estimating the motion of dynamic objects in the scene (Reference 6: Zhang Jun, Henein M, Mahony R, et al. VDO-SLAM: a visual dynamic object-aware SLAM system[J / OL].arXiv:2005.11052[cs.RO], 2020.). However, most of the above dynamic object elimination methods first perform semantic segmentation based on deep learning, and then track feature points. The specific steps are roughly to first determine the dynamic feature points, then assign semantic labels to the corresponding key frames, and then track the feature points, which results in too long calculation time and cannot meet the real-time requirements of the system. Summary of the invention

[0005] The present invention aims at the problem of how to effectively remove dynamic objects in a dynamic environment so as to construct a more accurate semantic map of an indoor environment, and proposes a method for implementing an indoor environment DFPM-SLAM system based on a dynamic feature point matching method. The system implemented by the method of the present invention aims at the problem of key frame selection, state judgment and update of dynamic feature points, designs a key frame scoring selection mechanism, proposes the concept of switching probability, and introduces the concept of fuzziness to accurately judge and update the state of dynamic feature points, thereby removing dynamic feature points and improving the tracking efficiency and accuracy of the indoor semantic SLAM system.

[0006] The method for implementing the indoor environment DFPM-SLAM system based on the dynamic feature point matching method of the present invention implements the indoor environment DFPM-SLAM system based on the ORB-SLAM3 framework, comprising the following steps:

[0007] Step 1: Add a key frame scoring selection mechanism module to the tracking subsystem to select key frames;

[0008] The key frame scoring selection mechanism module first scores each frame of the input image sequence according to the three set criteria: the number of feature points contained, the rotation offset of the image, and the number of frames between the image and the nearest key frame. It then normalizes the weighted sum of the three scores and compares them with a preset threshold to select the key frame.

[0009] Step 2: Set the MR-DFPM subsystem to generate a dynamic / static point marker map of the key frame;

[0010] The MR-DFPM subsystem is used to: (1) first, use the Mask-R-CNN deep learning network to perform semantic segmentation on the key frame image, assign semantic information labels to the objects, and the category confidence of different objects obtained by semantic segmentation is the switching probability; (2) secondly, use the concept of membership in fuzzy logic to divide the feature points in the segmentation edge area into three categories according to the size of the switching probability: dynamic points, static points and fuzzy dynamic / static points; (3) then, use the Kalman filtering method to update the switching probability of the fuzzy dynamic / static point to clarify whether the feature point is a dynamic point or a static point.

[0011] Step 3: The MR-DFPM subsystem returns the dynamic / static point marker map of the key frame to the tracking subsystem, and the tracking subsystem removes the dynamic feature points.

[0012] The indoor environment DFPM-SLAM system constructed by the method of the present invention also includes a local mapping subsystem and a closed-loop detection subsystem. The tracking subsystem selects key frames, identifies ORB feature points, removes dynamic feature points therein, and then outputs them to the local mapping subsystem. The local mapping subsystem inserts the key frames selected by the tracking subsystem, removes redundant map points and key frames, and creates new map points.

[0013] The present invention realizes accurate judgment and elimination of dynamic feature points, and thus can construct an indoor environment SLAM system with better effect, which can accurately eliminate dynamic feature points and improve tracking accuracy. Compared with the prior art, the advantages and positive effects of the present invention are:

[0014] (1) The method of the present invention first provides the image sequence key frame selection criteria and the corresponding scoring mechanism to more accurately select the image key frames; through the key frame quantization scoring mechanism, the effective key frames can be selected more accurately, which is beneficial to the subsequent semantic segmentation processing.

[0015] (2) For fuzzy dynamic / static feature points, the present invention proposes a switching probability method based on fuzzy thinking to determine the feature point state through the Mask R-CNN semantic segmentation network, and adopts a switching probability update method based on Kalman filtering to achieve accurate division of dynamic / static feature points, accurate identification of dynamic objects, and elimination of dynamic objects in the scene.

[0016] (3) Experimental results show that the DFPM-SLAM system implemented by the method of the present invention can accurately eliminate dynamic objects to solve the tracking loss problem, and the absolute trajectory error and relative attitude error in indoor dynamic scenes are greatly improved, and can meet the real-time requirements of the system operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is the overall framework diagram of the indoor environment DFPM-SLAM system constructed by the method of the present invention;

[0018] Figure 2 is a schematic diagram of a feature point pair for matching two images in an embodiment of the present invention;

[0019] Figure 3 It is a schematic diagram of different feature points divided based on fuzzy logic and clearly classified;

[0020] Figure 4 It is a comparison chart of experimental results of the DFPM-SLAM system of the present invention and the ORB-SLAM3 system. DETAILED DESCRIPTION

[0021] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0022] Aiming at the problem of how to effectively remove dynamic objects in a dynamic environment in order to build a more accurate semantic map of the indoor environment, the present invention proposes a method for implementing an indoor environment DFPM-SLAM system based on a dynamic feature point matching method on the basis of simplifying ORB-SLAM3, which has an image sequence key frame selection mechanism and a feature point state switching probability mechanism in the image key frame to improve the camera posture tracking effect in indoor dynamic scenes. Figure 1 The figure shows the overall framework of the system constructed by the method of the present invention, and the implementation method is described in detail below.

[0023] Step 1: ORB-SLAM3 is used as the basic SLAM framework, which includes tracking subsystem, local mapping subsystem, MR-DFPM (Mask RCNN-Dynamic Feature Point Match) subsystem and closed-loop detection subsystem. Figure 1The present invention improves the tracking subsystem, establishes a key frame scoring selection mechanism for image sequences, and adds a MR-DFPM subsystem to eliminate the influence of dynamic objects on pose estimation as much as possible.

[0024] Tracking subsystem: It mainly realizes the ORB (Oriented Fast and Rotated Brief) feature extraction of continuous image frames, establishes a key frame scoring selection mechanism to complete the key frame selection, and combines with the MR-DFPM subsystem to accurately remove dynamic feature points, eliminate the influence of dynamic objects on pose estimation as much as possible, thereby reducing the error generated when tracking the latest frame and local map.

[0025] Local mapping subsystem: inserts keyframes selected by the tracking subsystem, removes redundant map points and keyframes and creates new map points, and uses the remaining global keyframes for local BA (Bundle Adjustment) optimization.

[0026] MR-DFPM subsystem: mainly realizes the semantic segmentation of key frames, assigns semantic information labels to objects, and then uses the switching probability to fuzzily classify the dynamic / static points (dynamic feature points or static feature points), and then accurately classifies the dynamic / static points through the switching probability update. The dynamic / static point marker map generated by the system is fed back to the tracking subsystem to help it reduce the impact of dynamic objects on tracking accuracy, more accurately estimate the camera pose, and reduce tracking errors. The specific implementation of the MR-DFPM subsystem will be described in steps 2 and 3 below.

[0027] Closed-loop detection subsystem: Eliminate accumulated errors by discovering closed loops and optimizing based on closed-loop results.

[0028] This step describes the key frame scoring selection mechanism module added to the tracking subsystem, in which the corresponding key frame screening criteria and scoring standards are set, as described in the following steps 101 to 104.

[0029] Step 101: Setting criteria 1: The key frame should ensure that the image is clear and stable and has enough significant feature points. In the experiment of the embodiment of the present invention, the number of feature points contained in a frame of image is roughly between 100-350, so the minimum number of feature points that a key frame should contain is set to N. T = 300. The number of significant feature points contained in the key frame can be adjusted according to the actual situation.

[0030] For Criterion 1, the scoring items The calculation is as follows:

[0031]

[0032] Among them, Ni is the number of feature points contained in the i-th frame, i = 1, 2, ..., N, N represents the number of image frames; N T is the threshold value of the number of feature points, which is set to 300 in the embodiment of the present invention. T , then the score of this frame is zero; if the number of feature points is greater than or equal to N T , then the score of this item is the ratio of the two.

[0033] Step 102: Setting criterion 2: The difference in feature points between adjacent key frames should be able to reflect significant rotation changes in the camera pose.

[0034] For criterion 2, a matching feature point pair in two images is denoted as (P, P'), such as Figure 2 As shown, the centroid coordinates of the grayscale pixel points corresponding to the feature points are c and c' respectively, and the matching pair after removing the centroid point is (A, A'). The present invention calculates the image frame rotation offset score according to the following ① to ③.

[0035] ① For two images, first obtain the matching feature point pairs through visual odometer, determine the respective centroid coordinates c and c', and then obtain the corresponding de-centroided feature points A and A'. In the embodiment of the present invention, the de-centroided feature points are represented as an n*n pixel matrix, where n is a positive integer.

[0036] For the current frame image, a matching pair of feature points will be found with the previous frame image, and the rotation offset will be calculated; for the most recent key frame, a matching pair of feature points will be found with the previous key frame, and the rotation offset will be calculated.

[0037] ② Calculate the rotation transformation matrix R, and then calculate the rotation offset L of the i-th (i=1, 2, ..., N) frame image according to the following formula (2): i * The rotation offset L from the most recent keyframe k * ,as follows:

[0038]

[0039] Among them, A j , A′ j They respectively represent the j-th pixel value in the grayscale pixel matrix after removing the centroid point of the two images.

[0040] ③By comparison and According to formula (3), the image frame rotation offset score is calculated as follows:

[0041]

[0042] When obtaining the first key frame, the rotation offset of the frame containing enough feature points is used as the comparison standard

[0043] Step 103: Setting criterion 3: Adjacent key frames should be separated by a certain number of image frames.

[0044] For criterion 3, the newly selected key frame must be separated from the previous key frame by a certain number of frames. The current image frame number is compared with the number of the most recent key frame in the input image sequence, and the frame separation score is calculated according to formula (4): as follows:

[0045]

[0046] Among them, i c is the current frame number, i k The image frame number of the most recent key frame.

[0047] Step 104: In order to more effectively select key frames, the scores calculated in the above three sub-steps are first normalized by Z-Score to the same range, and then different weights (α, β, γ) are assigned to the above three evaluation indicators to obtain the final score of the i-th frame image, as shown in formula (5):

[0048]

[0049] in, are the normalized scores, respectively.

[0050] In the embodiment of the present invention, the weights α, β, and γ are weight values ​​optimized by the ant colony algorithm.

[0051] The score threshold T1 is set in advance. When the score S of the i-th frame image i When it is higher than the set threshold T1, it is selected as a key frame and stored as the nearest key frame, and the score and rotation offset of the nearest key frame are stored at the same time; otherwise, the i-th frame image is not selected as a key frame, and the next frame is judged.

[0052] Step 2: Fuzzy classification of feature point states based on switching probability. The efficiency of frame-by-frame semantic segmentation of images is the key to whether the tracking subsystem can quickly and accurately determine dynamic feature points. The present invention uses the Mask-R-CNN deep learning network to perform semantic segmentation on images. The category confidence of the object obtained by semantic segmentation is called the switching probability.

[0053] The MR-DFPM subsystem performs the following steps 201 to 202 on the key frame selected by the tracking subsystem:

[0054] Step 201: Using the key frames obtained by the key frame scoring selection mechanism as input, determine the region of interest (ROI) through the Region Proposal Network (RPN) module in the Mask-R-CNN deep learning network, and feedback the output proposals into the Region of Interest Feature Aggregation (ROI Align) module for object classification and bounding box regression, and assign semantic information labels to the objects. For the indoor environment, the semantic labels of the classification categories can be divided into 10 categories, as specifically shown in Table 1.

[0055] Table 1 Object Category Classification

[0056] Serial number category Dynamic / static state 1 people dynamic 2 cat dynamic 3 dog dynamic 4 table Static 5 Chair Static 6 monitor Static 7 desk lamp Static 8 bed Static 9 cabinet Static 10 air conditioner Static

[0057] Step 202: The class confidence obtained by semantic segmentation of the above different objects is called the switching probability. Considering the problem of classification uncertainty in the edge region of dynamic objects during semantic segmentation, the concept of membership degree in fuzzy logic is introduced to divide the feature points in the segmentation edge region into three categories according to the size of the switching probability: dynamic points (D), static points (S), and fuzzy dynamic / static points (F), as shown in the left figure below. It can be seen from the figure that dynamic points and static points are mainly distributed inside dynamic and static objects, while fuzzy dynamic / static points mostly appear at the boundaries between dynamic objects and the background or static objects. Figure 3 As shown in the left figure above. It can be seen from the figure that dynamic points and static points are mainly distributed inside dynamic and static objects, while fuzzy dynamic / static points mostly appear at the boundaries between dynamic objects and the background or static objects.

[0058] The present invention pre-sets feature point state classification thresholds a and b, where 0 < b < a < 1, for classifying feature points into dynamic points, static points, or fuzzy dynamic / static points. When the switching probability is greater than a, it is classified as dynamic; when it is less than b, it is classified as static; otherwise, it is classified as a fuzzy dynamic / static point.

[0059] In the embodiment of the present invention, the switching probability is obtained according to the class confidence, and the class confidence is converted into the confidence belonging to the same state category, which is the switching probability. For example, if the confidence of segmenting a person is 0.94 and the state corresponding to "person" is dynamic, then the switching probability of the feature points corresponding to this category is 0.94, and 0.94 > a, so it is a dynamic feature point. If the confidence of segmenting a table is 0.85 and the state of "table" is static, the higher the confidence belonging to static indicates the lower the probability that it is dynamic. The switching probability of this category is 0.25, and 0.25 < b, so these feature points are static feature points. If the confidence of the segmented category is 0.5 (between a and b), it is classified as a fuzzy point.

[0060] Step 3: Further clarify the states of the feature points of fuzzy dynamic / static.

[0061] Usually in the practical application of V-SLAM system, the robot moves continuously in an unknown environment, and the dynamic objects in the scene are in a non-static state. Most of the feature points located at the junction of the dynamic object and the static background in the continuous image frame sequence are fuzzy dynamic / static. In order to further clarify the state of these feature points, the feature points belonging to the fuzzy dynamic / static category are classified as dynamic feature points or static feature points, so as to more accurately eliminate dynamic objects. The present invention adopts the Kalman filtering method to update the switching probability of the fuzzy dynamic / static feature points obtained after the semantic segmentation of the Mask R-CNN network in step 2, so as to achieve clear classification. The specific steps 301 to 303 are as follows.

[0062] Step 301: First, the preset initial probability of the present invention is described.

[0063] The switching probability of the i-th blurred dynamic / static feature point in the current key frame predicted by Mask R-CNN semantic segmentation and initial state is:

[0064] P t (a i ),a i ∈{F} (6)

[0065] Among them, the subscript t represents the key frame at the current time t, i = 1, 2, ... m, m is the number of fuzzy state feature points in the key frame; F represents the fuzzy dynamic / static feature point, a i represents the ith fuzzy dynamic / static point. t (a i ) represents the switching probability of the i-th fuzzy dynamic / static point in the current key frame. The present invention sets the initial switching probability of the first key frame P0=0.5.

[0066] Step 302: Predict the updated value of the switching probability of the fuzzy dynamic / static feature point at the current moment through the prediction equation of the Kalman filter method. The prediction equation of the Kalman filter method is as follows:

[0067] x t =Ax t-1 +Bu t-1 +w t-1 (7)

[0068] z t =Hx t +v t (8)

[0069] In the above formula, x t The estimated value P of the switching probability of the fuzzy dynamic / static feature point t x (a i ), z tThe observed value P representing the switching probability of the feature point t z (a i ).

[0070] The corresponding prediction equation for the switching probability of fuzzy dynamic / static feature points at the current moment is as follows:

[0071]

[0072] P t z (a i )=HP t x (a i )+v t (10)

[0073] Among them, A is the state transfer matrix, B is the state input matrix, u t-1 is the control input at time t-1, w t-1 ~N(0,1) is the system noise; H is the transfer matrix from state variables to observation variables, v t ~N(0,1) is the observation noise.

[0074] Step 303: Update the switching probability of the fuzzy dynamic / static feature points through the update equation of the Kalman filter method.

[0075]

[0076]

[0077] Among them, K t is the filter gain matrix, Indicates a i The updated estimated value of the switching probability; is the a priori optimal estimation result at time t; is time t The estimated covariance, that is, the uncertainty of the state; r is the covariance of the observation noise, which is set to 1 in the embodiment of the present invention.

[0078] According to the switching probability value calculated above The state of the feature point in the fuzzy state can be clearly classified as dynamic or static. In the embodiment of the present invention, the thresholds are set to 0.75 and 0.25. When the switching probability is greater than 0.75, it is classified as dynamic, and when it is less than 0.25, it is classified as static. In extremely special cases, when the updated switching probability value of the fuzzy feature point is still near 0.5, it can be considered as a dynamic feature point and removed to ensure that the SLAM algorithm will not retain any dynamic feature points in the subsequent camera posture tracking and mapping process.

[0079] like Figure 1 As shown, the MR-DFPM subsystem of the present invention generates a feature point state marking map of the key frame, in which the state of the feature points is marked, and sends the dynamic / static point marking map of the key frame to the tracking subsystem. The tracking subsystem removes the dynamic feature points therein to eliminate the influence of dynamic objects on pose estimation and reduce the error generated when tracking the latest frame and the local map.

[0080] Example:

[0081] The experimental system of this invention is Ubuntu20.04, and the graphics card is RTX3050 Laptop. The Mask R-CNN part is implemented using python3.6, and the rest is implemented using C++.

[0082] The experiment uses five sets of image sequences under different camera motion forms in the TUM dataset as experimental data to verify the method of the present invention. As shown in Table 2:

[0083] Table 2 Experimental image sequences

[0084] Image sequence name Camera movement mode Dynamic situation classification fr3_walking_xyz Along the x, y, z axis High Dynamics fr3_walking_rpy Along the r (roll), p (pitch), and y (yaw) axes High Dynamics fr3_walking_halfsphere Along a hemisphere with a diameter of 1 meter High Dynamics fr3_walking_static Manual hold High Dynamics fr3_sitting_static Manual hold Low dynamic range

[0085] In the above-mentioned walking series image sequences, there are multiple moving objects (people) with large motion amplitudes, which belong to complex high-dynamic indoor scenes; while in the sitting series image sequences, there are fewer moving objects with smaller motion amplitudes, which belong to low-dynamic indoor scenes.

[0086] In order to test the performance of the DFPM-SLAM system implemented by the present invention, it is compared with the original ORB-SLAM3 in the experiment, and quantitative evaluation is performed from two aspects: absolute trajectory error (ATE) and relative pose error (RPE). Figure 4 As shown, the absolute trajectory errors of the camera in 3D space, xyz axis and rpy (roll, pitch and yaw) axis under the operation of the DFPM-SLAM system of the present invention and the original ORB-SLAM3 system are given respectively. Figure 4 In the same figure, the solid line represents the actual trajectory of the camera, and the dotted line represents the estimated trajectory; Figure 4The first behavior is a rendering of trajectory tracking using the system of the present invention, and the second behavior is a rendering of trajectory tracking using the comparison system-original ORB-SLAM3. At the same time, in order to verify the stability of the DFPM-SLAM system, considering that the presence of dynamic objects will increase the uncertainty of the system, the present invention repeats 10 experiments under 5 groups of image sequences, and takes the middle value of the root mean square error (RMSE) and standard deviation (SD) of ATE and RPE as the final experimental result. The quantitative experimental results of ATE and RPE of DFPM-SLAM and ORB-SLAM3 systems are shown in Tables 3 and 4. In addition, the experimental maximum and minimum values ​​of the above two indicators under the DFPM-SLAM system are also given, and the improvement rate of the DFPM-SLAM system of the present invention compared to ORB-SLAM3 is compared and calculated (the improvement rate is calculated using the middle value). In the table, w represents the walking image series, s represents the sitting image series, and example fr3_walking_xyz is written as w / xyz.

[0087] Table 3. Comparison of absolute trajectory errors

[0088]

[0089] Table 4 Comparison of relative attitude errors

[0090]

[0091] As can be seen from Tables 3 and 4, the DFPM-SLAM system implemented by the method of the present invention improves the performance of most high-dynamic image sequences by an order of magnitude. In terms of ATE, the RMSE and SD improvement values ​​can reach 98.04% and 97.99% respectively. In terms of RPE, the RMSE and SD improvement values ​​can reach 95.38% and 97.48% respectively. The results show that DFPM-SLAM can significantly improve the robustness and stability of the SLAM system in a high-dynamic environment. However, in low-dynamic image sequences (such as fr3_sitting_static), the performance improvement is not obvious. The reason is that the original ORB-SLAM3 system can already achieve good performance for the processing of low-dynamic scenes, so there is little room for improvement.

[0092] Compared to other camera movements, the camera has a smaller lift when rotating on the rpy axis.

[0093] As a real-time semantic SLAM system, in practical applications, real-time performance is an important indicator to measure the performance of the SLAM system. The experiment tested the time required for the key frame selection module, semantic segmentation module and switching probability update module for single frame processing.

[0094] Table 5 Single frame processing time of key modules of the system

[0095] Keyframe Selection Semantic Segmentation Switching probability update Total time Average time (ms) 2.1957 38.9738 10.2763 61.8 Minimum time (ms) 2.0032 38.2573 10.0134 60.1 Maximum time (ms) 2.2745 39.4578 10.9803 62.4

[0096] As shown in Table 5, the system implemented by the present invention takes an average of 61.8 ms to process each frame, and the minimum time for a single frame can reach 60.1 ms. The maximum time for processing more complex key frames is 62.4 ms, which can meet real-time requirements.

[0097] The above experimental results show that the method of the present invention realizes the real-time interaction between semantic information and tracking subsystem through Mask R-CNN, effectively eliminates dynamic points in key frames, and improves the tracking success rate and positioning accuracy in indoor dynamic environments. The experimental results under the TUM data set show that compared with the ORB-SLAM3 system, the DFPM-SLAM system proposed in the present invention is more accurate and more robust.

Claims

1. A method for implementing an indoor environment DFPM-SLAM system based on a dynamic feature point matching method, characterized in that: The method implements an indoor environment DFPM-SLAM system based on the ORB-SLAM3 framework, including the following steps: Step 1: Add a key frame scoring and selection mechanism module in the tracking subsystem to select key frames; For the key frame scoring and selection mechanism module, first, according to three set criteria: the number of feature points included, the rotational offset of the image, and the number of frames separated from the nearest key frame, each frame of the input image sequence is scored. Then, the three scores are normalized and weighted and summed, and compared with a pre-set threshold to select key frames; Step 2: Set the MR-DFPM subsystem to generate a dynamic / static point marking map of key frames; The MR-DFPM subsystem is used for: (1) First, use the Mask-R-CNN deep learning network to perform semantic segmentation on the key frame image, assign semantic information labels to objects, and obtain the switching probability from the class confidence of different objects obtained by semantic segmentation; (2) Secondly, using the concept of membership degree in fuzzy logic, the feature points in the segmentation edge area are divided into three categories according to the size of the switching probability: dynamic points, static points, and fuzzy dynamic / static points; (3) Then, use the Kalman filter method to update the switching probability of fuzzy dynamic / static points to clarify whether the feature points are dynamic points or static points; Step 3: The MR-DFPM subsystem returns the dynamic / static point marking map of the key frame to the tracking subsystem, and the tracking subsystem removes the dynamic feature points therein.

2. The method according to claim 1, characterized in that In the above Step 1, the following three criteria and scoring methods are set in the key frame scoring and selection mechanism module to screen key frame criteria: (1) Setting criterion 1: The key frame should ensure that the image is clear and stable and has enough significant feature points; the way to score the i-th frame image according to criterion 1 is: setting the threshold number of feature points contained in the key frame N T , if the number of feature points contained in the image is less than N T , then the image scores according to criterion 1 is zero, otherwise, the score N i / N T , N i is the number of feature points contained in the i-th frame image; (2) Setting Criterion 2: The difference in feature points between adjacent key frames should reflect a significant rotation change in the camera pose. The way to score the i-th frame image according to Criterion 2 is: through the visual odometry, obtain the matching feature point pairs between the i-th frame image and the i-1-th frame image, and calculate the rotation offset Calculate the rotation offset of the nearest keyframe Then when Less than When the i-th frame image is scored according to criterion 2 is zero, otherwise, the score for (3) Setting criterion 3: Adjacent key frames should be separated by a certain number of image frames; the method of scoring the i-th frame image according to criterion 3 is: respectively obtaining the sequence number i of the i-th frame image and the nearest key frame in the input image sequence c 、i k , and then calculate the score of the i-th frame image according to criterion 3 3. The method according to claim 1 or 2, characterized in that: In the step 1, the score of the i-th frame image calculated according to the three criteria is and After normalization, they are Then the final score S of the i-th frame image is i as follows: Among them, α, β, and γ are the assigned weights respectively.

4. The method according to claim 1, characterized in that: In the above Step 2, the Kalman filter method is used to update the switching probability of fuzzy dynamic / static points, specifically as follows: Step 301: Let P t (a i ) represents the switching probability of the i-th fuzzy dynamic / static point in the key frame at the current time t, and the switching probability of the initial first key frame is set to 0.5; i = 1, 2, ... m, m is the number of fuzzy dynamic / static points in the key frame; Step 302: Predict the updated value of the switching probability of fuzzy dynamic / static points at the current time t through the prediction equation of the Kalman filter method, and the prediction equation is: P t z (a i )=HP t x (a i )+v t Among them, P t x (a i ) represents the estimated value of the switching probability state of the fuzzy dynamic / static point, P t z (a i ) represents the switching probability observation value of the feature point; A is the state transfer matrix, B is the state input matrix, u t-1 is the control input at time t-1, w t-1 is the system noise; H is the transfer matrix from state variables to observation variables, v t is the observation noise; Step 303: Update the switching probability of the fuzzy dynamic / static point through the update equation of the Kalman filter method to obtain the updated estimated value as follows: Among them, K t is the filter gain matrix, is the a priori optimal estimation result at time t; η t is time t The estimated covariance; r is the covariance of the observation noise, which is set to 1.

5. The method according to claim 1 or 4, characterized in that: In the above Step 2, feature point state classification thresholds a and b are pre-set, 0 < b < a < 1. When the switching probability is greater than a, it is classified as dynamic. When it is less than b, it is classified as static. Otherwise, it is classified as fuzzy dynamic / static.

6. The method according to claim 5, characterized in that In the above Step 2, the threshold a is set to 0.75 and b is set to 0.

25.

7. The method according to claim 1, 2 or 4, characterized in that: The indoor environment DFPM-SLAM system constructed by the method also includes a local mapping subsystem and a loop closing detection subsystem.