An RGB-D SLAM method for dynamic environments
By using the spatial structure and epipolar constraint consistency algorithm in the RGB-D SLAM method to screen dynamic matching pairs and combining it with the map point processing algorithm, the problem of inaccurate map construction in dynamic environments is solved and a high-precision SLAM effect is achieved.
Patent Information
- Application Number
- CN202311041273.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-18
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-08-18
AI Technical Summary
Existing RGB-D SLAM methods suffer from the accuracy and stability of camera pose estimation and map construction in dynamic environments due to moving objects. In particular, traditional methods have high computational costs or insufficient detection accuracy when dealing with dynamic objects of unknown categories.
The spatial structure comparison algorithm and epipolar constraint consistency algorithm are used to screen dynamic matching pairs, and the map point processing algorithm is combined to eliminate erroneous map points. The map construction process is optimized through tracking threads, local mapping threads, loop detection and map fusion threads, and deep classification and bundle adjustment methods are used to improve map consistency.
It improves the accuracy and stability of map construction in dynamic environments, reduces the interference of dynamic objects on camera pose estimation, enhances the consistency and completeness of the map, and reduces computational costs.
Smart Images

Figure CN117078757B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to an RGB-D SLAM method suitable for dynamic environments. Background Art
[0002] SLAM (Simultaneous Localization and Mapping) is a method that integrates computer vision and robotics technology to achieve simultaneous estimation of a robot's trajectory and a map of its surroundings in an unknown environment. SLAM technology has been widely used in fields such as self-driving cars, drones, and service robots, providing intelligent systems with the basic capabilities for real-time navigation and path planning. RGB-D SLAM is a simultaneous localization and mapping SLAM technology based on RGB-D sensors. By integrating RGB (color images) and D (depth images), it achieves three-dimensional reconstruction of the environment and precise positioning of the robot. Compared with traditional monocular or binocular SLAM, RGB-D SLAM has significant advantages in accuracy and stability, especially in indoor environments and close-range object recognition. However, because traditional SLAM algorithms are usually based on static environment assumptions, moving objects in dynamic environments (such as pedestrians and cars) have an adverse effect on the accuracy of map construction and the positioning accuracy of the robot.
[0003] In order to reduce the error of camera pose estimation caused by dynamic objects, many researchers have done a lot of research on SLAM in dynamic environments and proposed the following common methods:
[0004] One is the method based on semantic segmentation, represented by MaskFusion (an RGB-D SLAM method for moving objects based on the Mask R-CNN semantic segmentation network). This type of method uses deep learning or other machine learning technologies to perform pixel-by-pixel semantic segmentation on the input RGB image, thereby identifying dynamic objects belonging to known categories, such as people, cars, dogs, etc., and then remove the segmented dynamic objects from the RGB-D data or assign them lower weights to reduce their impact on camera pose estimation and map reconstruction; however, the segmentation effect of this method depends on the image quality and the performance of the segmentation algorithm. Its segmentation results are prone to errors and noise, and require training of a large number of data sets, with high computational costs, and cannot handle dynamic objects of unknown categories.
[0005] The second is the method based on geometric models, represented by StaticFusion (a background reconstruction method for dense RGB-D SLAM in dynamic environments based on optical flow information), SPWSLAM (a method that uses static point weighting to deal with RGB-D SLAM problems in dynamic environments) and EM-Fusion (a method that uses probabilistic data association to achieve dynamic object-level SLAM). Such methods usually use multi-view geometry or optical flow and other technologies to detect moving or non-rigid areas according to the depth or color changes in the RGB-D data, thereby segmenting dynamic objects. The segmented dynamic objects are then removed from the RGB-D data or given lower weights to reduce their impact on camera pose estimation and map reconstruction. This type of method does not require a training dataset, has a low computational cost, and can handle dynamic objects of unknown categories. However, the accuracy and robustness of dynamic object detection of such methods are poor for complex or severely occluded scenes.
[0006] The third is the method based on camera self-motion estimation, represented by Co-Fusion (an RGB-D SLAM method that uses different motion or semantic clues to segment the background and foreground), ORB-SLAM3 (a feature-based tightly coupled SLAM system), and other methods. This type of method uses camera self-motion estimation (visual odometry) as an intermediate step, and detects areas inconsistent with camera motion based on feature matching or photometric errors between adjacent frames, thereby segmenting dynamic objects. The segmented dynamic objects are then removed from the RGB-D data or given lower weights to reduce their impact on camera pose estimation and map reconstruction. This type of method can use the camera's own motion information to improve the accuracy and robustness of dynamic object detection, and can also handle dynamic objects of unknown categories, but it requires optimization of the camera's own motion estimation, has a high computational cost, and has poor dynamic object detection effects for fast or violent camera motions. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies in the prior art and provide an RGB-DSLAM method suitable for dynamic environments, which can effectively eliminate the adverse effects of moving objects in the environment on the map construction and positioning of SLAM technology.
[0008] To achieve the above object, the present invention is implemented by adopting the following technical solutions:
[0009] The present invention provides an RGB-D SLAM method suitable for dynamic environments, the method comprising:
[0010] The RGB image and depth image obtained by the RGB-D camera are used as input frames;
[0011] Start the tracking thread, based on the input frame and the atlas, use the spatial structure comparison algorithm and the epipolar constraint consistency algorithm to process the feature point matching pairs of the input frame, calculate the pose of the current frame relative to the active submap in the atlas and filter out new keyframes, create new map points and update the active submap in combination with the map point processing algorithm;
[0012] Start the local mapping thread, insert the new keyframe into the active submap, process the keyframe feature point matching pairs using the epipolar constraint consistency algorithm, create new local map points, and use the map point processing algorithm and bundle adjustment method to update and optimize the active submap;
[0013] Start the loop detection and map fusion thread to detect the common areas between the newly inserted keyframe and other keyframes of each submap in the atlas, and perform loop correction or fusion operations on the active submap based on the detection results;
[0014] After globally optimizing all active submaps after fusion and correction using the bundle adjustment method, the final pose of the camera in the current frame and the three-dimensional point map of its surrounding environment are obtained.
[0015] Optionally, the atlas includes an active submap, multiple unconnected dormant submaps, and a DBoW3 keyframe database; the DBoW3 keyframe database includes a feature point dictionary and a relocalization dataset for relocalization, loop detection, and map fusion; the active submap and the dormant submap both include a map point set, a keyframe set, a visibility graph, and a spanning tree; the tracking thread locates the input frame on the active submap, and the local mapping thread continuously optimizes and adds new keyframes to the active submap.
[0016] Optionally, the tracking thread specifically includes the following steps:
[0017] Extracting feature points of the RGB image in the input frame by a feature point extraction algorithm;
[0018] Match the feature points in the current frame with the feature points in the previous frame;
[0019] The current tracking state is determined based on the matching results. If tracking is lost, the current frame is relocated in all submaps of the atlas or the active submap is switched for tracking. If the re-tracking is successful, tracking is resumed; otherwise, the current active submap is stored as a dormant submap and a new active submap is initialized from scratch.
[0020] The depth classification algorithm is used to divide the matched feature point matching pairs into depth-valid feature point matching pairs and depth-inaccurate feature point matching pairs;
[0021] Use the spatial structure comparison algorithm to eliminate dynamic matching pairs from deep valid feature point matching pairs;
[0022] Calculate the pose of the current frame relative to the active submap based on the deep valid feature point matching pairs after removing the dynamic matching pairs;
[0023] The epipolar constraint consistency algorithm is used to eliminate dynamic matching pairs in feature point matching pairs with depth misalignment;
[0024] Based on the time interval between the current frame and the previous frame, determine whether the current frame becomes a key frame. If so, set the current frame as the new key frame.
[0025] According to the new keyframe and the calculated pose, new map points are created by combining the depth-valid feature point matching pairs and the depth-inaccurate feature point matching pairs after eliminating the dynamic matching pairs, and the new map points are processed using the map point processing algorithm to update the active sub-map.
[0026] Optionally, the local mapping thread specifically includes the following steps:
[0027] Insert a new keyframe into the active submap;
[0028] Eliminate untraceable map points in the active sub-map;
[0029] Perform feature matching on the feature points in the current new key frame and the feature points in the previous key frame to obtain a key frame feature point matching pair;
[0030] The epipolar constraint consistency algorithm is used to eliminate dynamic matching pairs in the key frame feature point matching pairs;
[0031] Create new local map points based on the eliminated keyframe feature point matching pairs;
[0032] The new local map points are processed using a map point processing algorithm, and the active submap is smoothed using the bundle adjustment method. At the same time, redundant keyframes in the active submap are deleted to update and optimize the current active submap.
[0033] Optionally, the loop detection and map fusion thread specifically includes the following steps:
[0034] When a new keyframe is inserted into the active submap, the new keyframe is checked for common areas with the other keyframes in the active submap;
[0035] If a common area is detected in the active submap, loop correction is performed on the current active submap;
[0036] If no common area is detected in the active submap, the new keyframe and the keyframes on each dormant submap are checked in turn to see if there is no common area, until a common area is detected or all dormant submaps have been traversed;
[0037] If there is a common area, the active submap is first fused with the dormant submap with the common area, and then loop correction is performed on the fused active submap.
[0038] Optionally, the key frame determination includes:
[0039] If the time interval between the current frame and the previous frame exceeds 0.5s, and the number of feature point matching pairs between the current frame and the previous frame is less than one-fourth of the number of feature point matching pairs between the previous frame and the previous frame, the current frame is determined to be a key frame.
[0040] Optionally, the method of using a depth classification algorithm to divide the feature point matching pairs obtained by matching into depth-valid feature point matching pairs and depth-inaccurate feature point matching pairs includes the following steps:
[0041] Get all feature point matching pairs of the current frame and the previous frame;
[0042] For each feature point matching pair obtained, classification is performed in the following order:
[0043]
[0044] Where,<p,q> For any feature point matching pair, p is the feature point in the previous frame, pd is the depth corresponding to the feature point p, q is the feature point matching p in the current frame, qd is the depth corresponding to the feature point q, L min is the lower limit of the depth threshold, L max is the upper depth threshold,<p,q> =True means<p,q> is a deep effective feature point matching pair,<p,q> =Falsd represents<p,q> is a depth misaligned feature point matching pair.
[0045] Optionally, the method of eliminating dynamic matching pairs from deep-valid feature point matching pairs by using a spatial structure comparison algorithm comprises the following steps:
[0046] Get deep valid feature point matching pairs<p0,q0> ,..., <p n-1 ,q n-1 >
[0047] By p0,...,p n-1 Calculate the corresponding back-projection points P0,...,P n-1 , the formula is as follows:
[0048] P i =di ·K -1 ·p i ,1 T ,0≤i≤n-1 (2)
[0049] Where, P i is the i-th feature point p i The back-projection point, n represents the number of deep valid feature point matching pairs, K is the intrinsic parameter matrix of the image sensor in the RGB-D camera, d i is the feature point p i Depth;
[0050] By q0,...,q n-1 Calculate the back-projection points Q0,...,Q n-1 , the formula is as follows:
[0051] Q i =d i ·K -1 ·q i ,1 T ,0≤i≤n-1 (3)
[0052] Where Q i q i The back projection point, K is the intrinsic parameter matrix of the image sensor in the RGB-D camera, d i is the feature point q i Depth;
[0053] According to the back-projection points P0,...,P n-1 , calculate the spatial structure matrix M ref , the formula is as follows:
[0054]
[0055]
[0056] Where, P j is the j-th feature point p j The back projection point of Represents the back-projection point P i and P j the distance between them;
[0057] According to the back-projection points Q0,...,Q n-1 , calculate the spatial structure matrix M cur , the formula is as follows:
[0058]
[0059]
[0060] Where, Represents the back-projection point Q i and Q j the distance between them;
[0061] Based on the spatial structure matrix M ref and M cur , calculate M c-r , the formula is as follows:
[0062]
[0063] Where M c-r Represents the spatial structure matrix M cur and M ref The ratio between them, M i:j Represents the back-projection point Q i and Q j The distance between and the back-projection point P i and P j The distance between The ratio between
[0064] According to M cr Calculate M′ using formula (9) c-r , the formula is as follows:
[0065]
[0066] Create an empty integer array Seq;
[0067] Use formula (10) to traverse M′ c-r All fractional expressions in the formula are as follows:
[0068] |M′ c-r , k, l-1.0|<θ (10)
[0069] Among them, M′ c-r , k,l is the matrix M′ c-r The element in the kth row and lth column, 0≤k≤n-1,0<l<n-1,k≠l; θ is the threshold. If equation (10) holds, the corresponding k and l are inserted into Seq;
[0070] Match the acquired feature points<p0,q0> ,..., <p n1 ,q n-1 The feature point matching pairs whose subscripts appear in the array Seq in > are set as dynamic matching pairs and are removed.
[0071] Optionally, the step of eliminating dynamic matching pairs using the epipolar constraint consistency algorithm includes:
[0072] Get depth misaligned feature point matching pairs<p′0,q′0> ,..., <p′ m-1 ,q′ m-1 >, m represents the number of feature point matching pairs with depth misalignment;
[0073] Get feature points p′0,...,p′ m-1 The pose T of the frame ref and the characteristic points q′0,...,q′ m-1 The pose T of the frame;
[0074] Calculate the pose T ref The relative posture ΔT between the posture T is as follows:
[0075]
[0076] Where ΔR is the relative rotation matrix, Δt is the relative translation vector;
[0077] According to the relative posture ΔT, the essential matrix ΔE between the current frame and the previous frame is calculated, and the formula is as follows:
[0078] ΔE=Δt^·ΔR (12)
[0079] Based on the essential matrix ΔE, the threshold of the matching error is calculated, and the formula is as follows:
[0080]
[0081] In the formula, x is the confidence level. The larger its value is, the lower the threshold of matching error is.
[0082] The feature point matching pairs obtained in turn by formula (14)<p′0,q0> ,..., <p′ m-1 ,q′ m-1 >Each pair of feature points in the <p′ u ,q′ u >Error u , the formula is as follows:
[0083] Error u =K -1 p′ u T ΔE K -1 q′ u (14)
[0084] Where K is the intrinsic parameter matrix of the image sensor in the RGB-D camera, u represents the u-th pair of feature point matching, 0≤u≤m-1;
[0085] If the calculated Erroru >Δx, then the feature points used for matching are matched <p′ u ,q′ u > is set as a dynamic matching pair and is removed.
[0086] Optionally, the steps of processing the new map point using the map point processing algorithm include:
[0087] Get new map point P c ;
[0088] Calculate the new map point P c Distance d from the camera's optical center p-c :
[0089] d p-c =||P c -P XYZ ||2 (15)
[0090] Where, P XYZ is the translation vector of the current frame’s pose;
[0091] Determine the new map point P c Distance d from the camera's optical center p-c Is it greater than the preset distance threshold?
[0092] If it is greater, use the DBoW3 keyframe database to calculate P c The node number NodeID in DBoW3 is recorded in P c Parameters;
[0093] Get all map points in the active submap whose node ID value is equal to NodeID <MP0,...,MP h-1 >, h represents the number of map points obtained;
[0094] Calculate P in sequence c With MP v Similarity:
[0095] sim P c ,MP v =||P c .des-MP v .des||2 (16)
[0096] Where, sim P c ,MP v It's P c With MP v The similarity between c .des is the map point P c Descriptor, MP v .des is the map point MPv Descriptor, 0≤v≤h-1;
[0097] Judgment P c With MP v The similarity sim P c ,MP v Is less than the preset similarity threshold, if so, then MP v Remove from the active submap and set P c Insert the active submap.
[0098] Compared with the prior art, the present invention has the following beneficial effects:
[0099] The present invention uses a spatial structure matrix comparison algorithm to screen out dynamic matching pairs from depth-valid feature point matching pairs, which has high efficiency and accuracy; and uses an epipolar constraint consistency algorithm to screen out dynamic matching pairs from depth-inaccurate feature point matching pairs, thereby increasing the number of map points that can be created; the present invention uses the proposed map point processing algorithm to effectively reduce erroneous map points and redundant map points in the map set by detecting the distance between the new map point and the camera and eliminating map points in the active sub-map that are similar to the new map; the method of the present invention can achieve high-precision RGB-D SLAM in a dynamic environment, effectively eliminate the interference of dynamic objects on camera pose estimation and map construction, and further improve the consistency and integrity of the map; in addition, the present invention adopts robot vision technology, which has good flexibility and applicability. According to the experimental results, the present invention has greatly improved various indicators compared with other RGB-D SLAM methods suitable for dynamic environments, and is suitable for promotion and application. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 A schematic diagram of a flow chart of an RGB-D SLAM method applicable to dynamic environments provided by an embodiment of the present invention;
[0101] Figure 2 A flow logic block diagram of the RGB-D SLAM method applicable to dynamic environments provided by an embodiment of the present invention;
[0102] Figure 3 Schematic diagram of the spatial structure comparison algorithm provided by the embodiment of the present invention
[0103] Figure 4 A schematic diagram of the epipolar constraint consistency algorithm flow provided by an embodiment of the present invention;
[0104] Figure 5 A flowchart of a map point processing algorithm provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0105] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Unless there is a conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.
[0106] The term "and / or" in this document simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " in this document generally indicates an "or" relationship between the related objects.
[0107] Reference Figure 1 As shown, the embodiment of the present invention introduces an RGB-D SLAM method suitable for dynamic environments, which specifically includes the following steps:
[0108] Step 1: Take the RGB image and depth image acquired by the RGB-D camera as input frames;
[0109] Step 2: Start the tracking thread. Based on the input frame and atlas, use the spatial structure comparison algorithm and the epipolar constraint consistency algorithm to process the feature point matching pairs of the input frame, calculate the pose of the current frame relative to the active submap in the atlas, and filter out new keyframes. At the same time, create new map points and update the active submap in combination with the map point processing algorithm.
[0110] Step 3: Start the local mapping thread, insert the new keyframe into the active submap, process the keyframe feature point matching pairs using the epipolar constraint consistency algorithm, create new local map points, and use the map point processing algorithm and bundle adjustment method to update and optimize the active submap;
[0111] Step 4: Start the loop detection and map fusion thread to detect the common areas between the newly inserted keyframe and other keyframes of each submap in the atlas. Perform loop correction or fusion operations on the active submap based on the detection results.
[0112] Step 5: After globally optimizing all active sub-maps after fusion and correction using the bundle adjustment method, the final pose of the camera in the current frame and the 3D point map of its surrounding environment are obtained.
[0113] The bundle adjustment method of the embodiment of the present invention uses a graph optimization algorithm (G2O) to perform smooth optimization processing on all active submaps.
[0114] Specifically, such as Figure 2As shown, the atlas provided in this embodiment consists of a set of disconnected submaps and a DBoW3 keyframe database. When executing step 1, it also includes loading a pre-trained DBoW3 keyframe database to initialize the atlas. The submap consists of a map point set, a keyframe set, a visibility graph, and a spanning tree. The DBoW3 keyframe database includes a feature point dictionary and a relocalization dataset for relocalization, loop detection, and map fusion. There is only one active submap in the submap, and the other submaps are dormant maps. The tracking thread locates the input frame on the map, and the local mapping thread continuously optimizes and adds new keyframes to the active submap.
[0115] As an embodiment of the invention, tracking the thread in step 2 specifically includes the following steps:
[0116] Step 2.1: extracting feature points of the RGB image in the input frame using a feature point extraction algorithm;
[0117] The feature point extraction algorithm used in this embodiment is preferably the ORB feature point extraction algorithm.
[0118] Step 2.2: Match the feature points in the current frame with the feature points in the previous frame;
[0119] Specifically, in this embodiment, the feature points extracted from the current frame are matched with the feature points extracted from the previous frame using a brute force matching method, and the brute force matching method calculates the similarity between the two descriptors according to the bi-norm.
[0120] Step 2.3: Determine the current tracking status based on the matching results. If tracking is lost, relocate the current frame in all submaps of the atlas or switch the active submap for tracking. If re-tracking is successful, resume tracking; otherwise, store the current active submap as a dormant submap and initialize a new active submap from scratch.
[0121] Step 2.4: Use the depth classification algorithm to divide the matched feature point matching pairs into depth-valid feature point matching pairs and depth-inaccurate feature point matching pairs;
[0122] Preferably, in this embodiment, step 2.4 specifically includes:
[0123] Step 2.4.1: Get all feature point matching pairs of the current frame and the previous frame;
[0124] Step 2.4.2: For each feature point matching pair obtained, classify them according to the following formula:
[0125]
[0126] Where,<p,q> For any feature point matching pair, p is the feature point in the previous frame, pd is the depth corresponding to the feature point p, q is the feature point matching p in the current frame, qd is the depth corresponding to the feature point q, L min is the lower limit of the depth threshold, L max is the upper depth threshold,<p,q> =True means<p,q> is a deep effective feature point matching pair,<p,q> =Falsd represents<p,q> is a depth misaligned feature point matching pair.
[0127] Step 2.5: Use the spatial structure comparison algorithm to eliminate dynamic matching pairs from deep valid feature point matching pairs;
[0128] Reference Figure 3 As shown, in this embodiment, the spatial structure comparison algorithm includes the following steps:
[0129] Step 2.5.1: Obtain deep valid feature point matching pairs<p0,q0> ,..., <p n-1 ,q n-1 >
[0130] Step 2.5.2: By p0, ..., p n-1 Calculate the corresponding back-projection points P0,...,P n-1 , the formula is as follows:
[0131] P i =d i ·K 1 ·p i ,1 T ,0≤i≤n-1 (2)
[0132] Where, P i is the i-th feature point p i The back-projection point, n represents the number of deep valid feature point matching pairs, K is the intrinsic parameter matrix of the image sensor in the RGB-D camera, d i is the feature point p i Depth;
[0133] Step 2.5.3: By q0,...,q n-1 Calculate the back-projection points Q0,...,Q n-1 , the formula is as follows:
[0134] Q i =d i ·K -1 ·q i ,1 T ,0≤i≤n-1 (3)
[0135] Where Q i qi The back projection point, K is the intrinsic parameter matrix of the image sensor in the RGB-D camera, d i is the feature point q i Depth;
[0136] Step 2.5.4: According to the back-projection points P0, ..., P n-1 , calculate the spatial structure matrix M ref , the formula is as follows:
[0137]
[0138]
[0139] Where, P j is the j-th feature point p j The back projection point of Represents the back-projection point P i and P j the distance between them;
[0140] Step 2.5.5: Based on the back-projection points Q0,...,Q n-1 , calculate the spatial structure matrix M cur , the formula is as follows:
[0141]
[0142]
[0143] Where, Represents the back-projection point Q i and Q j the distance between them;
[0144] Step 2.5.6: Based on the spatial structure matrix M ref and M cur , calculate M c-r , the formula is as follows:
[0145]
[0146] Where M c-r Represents the spatial structure matrix M cur and M ref The ratio between them, M i:j Represents the back-projection point Q i and Q j The distance between and the back-projection point P i and P j The distance between The ratio between
[0147] Step 2.5.7: According to Mc-r Calculate M′ using formula (9) c-r , the formula is as follows:
[0148]
[0149] Step 2.5.8: Create an empty integer array Seq;
[0150] Step 2.5.9: Use formula (10) to traverse M′ c-r All fractional expressions in the formula are as follows:
[0151] |M′ c-r k, l-1.0 (10)
[0152] Among them, M′ c-r , k,l is the matrix M′ c-r The element in the kth row and lth column, 0≤k≤n-1,0<l<n-1,k≠l; θ is the threshold. If equation (10) holds, the corresponding k and l are inserted into Seq;
[0153] Step 2.5.10: Match the acquired feature points<p0,q0> ,..., <p n-1 ,q n-1 The feature point matching pairs whose subscripts appear in the array Seq in > are set as dynamic matching pairs and are removed.
[0154] Step 2.6: Calculate the pose of the current frame relative to the active submap based on the deep valid feature point matching pairs after removing the dynamic matching pairs;
[0155] Preferably, in this embodiment, a graph optimization algorithm (g2o) is combined when calculating the pose based on the processed depth-valid feature point matching pairs to reduce interference.
[0156] Step 2.7: Use the epipolar constraint consistency algorithm to eliminate dynamic matching pairs from the feature point matching pairs with depth misalignment;
[0157] As an embodiment of the present invention, the process of the epipolar constraint consistency algorithm is referred to Figure 4 As shown, the specific steps include:
[0158] Step 2.7.1: Obtain depth misaligned feature point matching pairs<p′0,q′0> ,..., <p′ m-1 ,q′ m-1 >, m represents the number of feature point matching pairs with depth misalignment;
[0159] Step 2.7.2: Get feature points p′0,...,p′ m-1 The pose T of the frame refand the characteristic points q′0,...,q′ m-1 The pose T of the frame;
[0160] Step 2.7.3: Calculate the pose T ref The relative posture ΔT between the posture T is as follows:
[0161]
[0162] Where ΔR is the relative rotation matrix, Δt is the relative translation vector;
[0163] Step 2.7.4: Based on the relative pose ΔT, calculate the essential matrix ΔE between the current frame and the previous frame. The formula is as follows:
[0164] ΔE=Δt^·ΔR (12) Step 2.7.5: Based on the essential matrix ΔE, calculate the threshold of the matching error, which is as follows:
[0165]
[0166] In the formula, x is the confidence level. The larger its value is, the lower the threshold of matching error is.
[0167] Step 2.7.6: Match the feature points obtained in turn through formula (14)<p′0,q′0> ,..., <p′ m-1 ,q′ m-1 >Each pair of feature points in the <p′ u ,q′ u >Error u , the formula is as follows:
[0168] Error u =K -1 p′ u T ΔE K -1 q′ u (14)
[0169] Where K is the intrinsic parameter matrix of the image sensor in the RGB-D camera, u represents the u-th pair of feature point matching, 0≤u≤m-1;
[0170] Step 2.7.8: If the calculated Error u >Δx, then the feature points used for matching are matched <p′ u ,q′ u > is set as a dynamic matching pair and is removed.
[0171] Step 2.8: Based on the time interval between the current frame and the previous frame, determine whether the current frame becomes a key frame. If so, set the current frame as the new key frame.
[0172] Specifically, in this embodiment, the judgment of the key frame includes: if the time interval between the current frame and the previous frame exceeds 0.5s, and the number of feature point matching pairs between the current frame and the previous frame is less than one-fourth of the number of feature point matching pairs between the previous frame and the previous frame, then the current frame is judged as a key frame.
[0173] Step 2.9: Create new map points based on the new keyframe and the calculated pose, combined with the depth-valid feature point matching pairs and the depth-inaccurate feature point matching pairs after eliminating the dynamic matching pairs, and process the new map points using the map point processing algorithm to update the active submap.
[0174] Reference Figure 5 As shown, in this embodiment, the process of the map point processing algorithm specifically includes the following steps:
[0175] Step 2.9.1: Get new map point P c ;
[0176] Step 2.9.2: Calculate the new map point P c Distance d from the camera's optical center pc :
[0177] d p-c =||P c -P XYZ ||2 (15)
[0178] Where, P XYZ is the translation vector of the current frame’s pose;
[0179] Step 2.9.3: Determine the new map point P c Distance d from the camera's optical center pc Is it greater than the preset distance threshold?
[0180] Step 2.9.4: If it is greater than, calculate P using the DBoW3 keyframe database c The node number NodeID in DBoW3 is recorded in P c Parameters;
[0181] Step 2.9.5: Get all map points in the active submap whose node ID value is equal to NodeID <MP0,...,MP h-1 >, h represents the number of map points obtained;
[0182] Step 2.9.6: Calculate P in sequence c With MP vSimilarity:
[0183] sim P c ,MP v =||P c .des-MP v .des||2 (16)
[0184] Where, sim P c ,MP v It's P c With MP v The similarity between c .des is the map point P c Descriptor, MP v .des is the map point MP v Descriptor, 0≤v≤h-1;
[0185] Step 2.9.7: Determine P c With MP v The similarity sim P c ,MP v Is less than the preset similarity threshold, if so, then MP v Remove from the active submap and set P c Insert the active submap.
[0186] As an embodiment of the invention, the local mapping thread in step 3 specifically includes the following steps:
[0187] Step 3.1: Insert a new keyframe into the active submap;
[0188] Step 3.2: Eliminate untraceable map points in the active sub-map;
[0189] Among them, an untrackable map point refers to a map point that has not been tracked again within three frames after the map point was created.
[0190] Step 3.4: Perform feature matching on the feature points in the current new key frame and the feature points in the previous key frame to obtain a key frame feature point matching pair;
[0191] Step 3.5: Use the epipolar constraint consistency algorithm to eliminate dynamic matching pairs from the key frame feature point matching pairs;
[0192] Step 3.6: Create new local map points based on the eliminated keyframe feature point matching pairs;
[0193] Step 3.7: Use the map point processing algorithm to process the new local map points, and use the bundle adjustment method to smooth the active submap. At the same time, delete the redundant keyframes in the active submap to update and optimize the current active submap.
[0194] A redundant keyframe in an active submap refers to a keyframe in which 90% of the map points observed by the keyframe can be observed by at least three other keyframes at the same or better scale.
[0195] It should be noted that, in this embodiment, the process of using the epipolar constraint consistency algorithm to eliminate dynamic matching pairs in the key frame feature point matching pairs in step 3.5 is the same as the processing flow principle of the epipolar constraint consistency algorithm in step 2.7; the process of using the map point processing algorithm to process the new local map point in step 3.7 is the same as the processing flow principle of using the map point processing algorithm in step 2.9, and will not be repeated here.
[0196] As an embodiment of the invention, in step 4, the loop detection and map fusion thread specifically includes the following steps:
[0197] Step 4.1: When a new keyframe is inserted into the active submap, the new keyframe is checked for common areas with the other keyframes in the active submap;
[0198] Step 4.2: If a common area is detected in the active submap, loop correction is performed on the current active submap;
[0199] Step 4.3: If no common area is detected in the active submap, the new keyframe and the keyframes on each dormant submap are sequentially checked to see if there is no common area, until a common area is detected or all dormant submaps have been traversed;
[0200] Step 4.5: If there is a common area, first fuse the active submap with the dormant submap with the common area, and then perform loop correction on the fused active submap.
[0201] In this embodiment, the common area detection method is further described as follows: the ORB word bag of each key frame is calculated using the DBoW3 key frame database, potential similar key frames are found by comparing the similarities between the ORB word bags, and finally the similar key frames with the longest time interval are selected to form a loop, that is, the common area is detected.
[0202] In order to verify the performance of the method of the present invention, a comparative test was conducted between the method of the present invention and other commonly used RGB-D SLAM methods for dynamic environments: ORB-SLAM3 (ORB), StaticFusion (S1), SPWSLAM (S2), Co-Fusion (S3), MaskFusion (S4) and EM-Fusion (S5). The comparative experimental data are shown in Table 1.
[0203] Table 1
[0204]
[0205]
[0206] This example uses the TUM dataset for experiments, and selects 5 micro-dynamic datasets and 4 high-dynamic datasets. In these dynamic sequences, the camera operates in the following five different modes: (a) static: the camera remains basically still with some slight shaking; (b) XYZ axis movement (xyz): the camera moves roughly along the X, Y, and Z axes, and rotation is avoided as much as possible during this process; (c) rotation around the main axis (rpy): the camera mainly rotates around the main axis with almost no movement; (d) hemispherical motion (halfsphere): the camera moves back and forth along the longitude and latitude of the hemisphere, and the lens always points to the center of the sphere; (e) movement around the target (desk-person): the camera moves back and forth around the target. The selected micro-dynamic sequences include: fr2 / desk-person (fr2 / d / person), fr3 / sitting-static (fr3 / s / static), fr3 / sitting-xyz (fr3 / s / xyz), fr3 / sitting-rpy (fr3 / s / rpy), and fr3 / sitting-halfsphere (fr3 / s / half). The high-dynamic sequences include: fr3 / walking-static (fr3 / w / static), fr3 / walking-xyz (fr3 / w / xyz), fr3 / walking-rpy (fr3 / w / rpy), and fr3 / walking-halfsphere (fr3 / w / half). The experimental results are evaluated using Absolute Trajectory Error (ATE). The lower the ATE.RMSE value, the higher the accuracy of the SLAM method.
[0207] It can be observed from Table 1 that the accuracy of the present invention is better than other methods in both micro-dynamic and high-dynamic sequences, and the average accuracy is also better than other methods overall; MaskFusion's average ATE.RMSE in all sequences is second only to V-SLAM of the present invention, but MaskFusion requires more GPU computing resources to be able to run in real time; SPWSLAM's overall performance in all sequences is second only to MaskFusion, but SPWSLAM cannot handle the situation where the camera only performs pure rotational motion well. In the sequence fr3 / w / rpy, the average ATE.RMSE of the present invention is reduced by 63.15% compared with SPWSLAM. Obviously, compared with the existing dynamic RGB-D SLAM methods, the method of the present invention has better flexibility and applicability, and all indicators of the other RGB-D SLAM methods suitable for dynamic environments have been greatly improved. The present invention can achieve high-precision RGB-D SLAM in dynamic environments, effectively eliminate the interference of dynamic objects on camera pose estimation and map construction, further improve the consistency and integrity of the map, and has stronger competitiveness. Therefore, the method of the present invention has great practical significance and application value in the current and future field of robot vision technology.
[0208] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0209] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0210] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0211] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0212] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. An RGB-D SLAM method suitable for dynamic environments, characterized in that: The method comprises: The RGB image and depth image obtained by the RGB-D camera are used as input frames; Start the tracking thread, based on the input frame and the atlas, use the spatial structure comparison algorithm and the epipolar constraint consistency algorithm to process the feature point matching pairs of the input frame, calculate the pose of the current frame relative to the active submap in the atlas and filter out new keyframes, create new map points and update the active submap in combination with the map point processing algorithm; Start the local mapping thread, insert the new keyframe into the active submap, process the keyframe feature point matching pairs using the epipolar constraint consistency algorithm, create new local map points, and use the map point processing algorithm and bundle adjustment method to update and optimize the active submap; Start the loop detection and map fusion thread to detect the common areas between the newly inserted keyframe and other keyframes of each submap in the atlas, and perform loop correction or fusion operations on the active submap based on the detection results; After globally optimizing all active submaps after fusion and correction using the bundle adjustment method, the final pose of the camera in the current frame and the 3D point map of its surroundings are obtained; The tracking thread specifically includes the following steps: Extracting feature points of the RGB image in the input frame by a feature point extraction algorithm; Match the feature points in the current frame with the feature points in the previous frame; The current tracking state is determined based on the matching results. If tracking is lost, the current frame is relocated in all submaps of the atlas or the active submap is switched for tracking. If the re-tracking is successful, tracking is resumed; otherwise, the current active submap is stored as a dormant submap and a new active submap is initialized from scratch. The depth classification algorithm is used to divide the matched feature point matching pairs into depth-valid feature point matching pairs and depth-inaccurate feature point matching pairs; Use the spatial structure comparison algorithm to eliminate dynamic matching pairs from deep valid feature point matching pairs; Calculate the pose of the current frame relative to the active submap based on the deep valid feature point matching pairs after removing the dynamic matching pairs; The epipolar constraint consistency algorithm is used to eliminate dynamic matching pairs in feature point matching pairs with depth misalignment; Based on the time interval between the current frame and the previous frame, determine whether the current frame becomes a key frame. If so, set the current frame as the new key frame. According to the new keyframe and the calculated pose, new map points are created by combining the depth-valid feature point matching pairs and the depth-inaccurate feature point matching pairs after eliminating the dynamic matching pairs, and the new map points are processed using the map point processing algorithm to update the active sub-map.
2. The RGB-D SLAM method for dynamic environments according to claim 1, wherein The atlas includes an active submap, multiple disconnected dormant submaps, and a DBoW3 keyframe database; the DBoW3 keyframe database includes a feature point dictionary and a relocalization dataset for relocalization, loop detection, and map fusion; the active submap and dormant submap both include a map point set, a keyframe set, a visibility graph, and a spanning tree; the tracking thread locates the input frame on the active submap, and the local mapping thread continuously optimizes and adds new keyframes to the active submap.
3. The RGB-D SLAM method for dynamic environments according to claim 1, wherein: The local mapping thread specifically includes the following steps: Insert a new keyframe into the active submap; Eliminate untraceable map points in the active sub-map; Perform feature matching on the feature points in the current new key frame and the feature points in the previous key frame to obtain a key frame feature point matching pair; The epipolar constraint consistency algorithm is used to eliminate dynamic matching pairs in the key frame feature point matching pairs; Create new local map points based on the eliminated keyframe feature point matching pairs; The new local map points are processed using a map point processing algorithm, and the active submap is smoothed using the bundle adjustment method. At the same time, redundant keyframes in the active submap are deleted to update and optimize the current active submap.
4. The RGB-D SLAM method for dynamic environments according to claim 3, wherein: The loop detection and map fusion thread specifically includes the following steps: When a new keyframe is inserted into the active submap, the new keyframe is checked for common areas with the other keyframes in the active submap; If a common area is detected in the active submap, loop correction is performed on the current active submap; If no common area is detected in the active submap, the new keyframe and the keyframes on each dormant submap are checked in turn to see if there is no common area, until a common area is detected or all dormant submaps have been traversed; If there is a common area, the active submap is first fused with the dormant submap with the common area, and then loop correction is performed on the fused active submap.
5. The RGB-D SLAM method suitable for dynamic environments according to claim 1, wherein The determination of the key frame includes: If the time interval between the current frame and the previous frame exceeds 0.5s, and the number of feature point matching pairs between the current frame and the previous frame is less than one-fourth of the number of feature point matching pairs between the previous frame and the previous frame, the current frame is determined to be a key frame.
6. The RGB-D SLAM method suitable for dynamic environments according to claim 1, wherein: The method of using a depth classification algorithm to divide the feature point matching pairs obtained by matching into depth-valid feature point matching pairs and depth-inaccurate feature point matching pairs includes the following steps: Get all feature point matching pairs of the current frame and the previous frame; For each feature point matching pair obtained, classification is performed in the following order: (1) Where, For any feature point matching pair, is the feature point in the previous frame, is a feature point The corresponding depth, The current frame Matched feature points, is a feature point The corresponding depth, is the lower depth threshold, is the upper depth threshold, express is a deep effective feature point matching pair, express is a depth misaligned feature point matching pair.
7. The RGB-D SLAM method suitable for dynamic environments according to claim 1, wherein: The method of eliminating dynamic matching pairs from deep-valid feature point matching pairs by using a spatial structure comparison algorithm comprises the following steps: Get deep valid feature point matching pairs ; pass Calculate the corresponding back-projection point , the formula is as follows: (2) Where, P i is the i-th feature point p i The back-projection point, n represents the number of deep valid feature point matching pairs, K is the intrinsic parameter matrix of the image sensor in the RGB-D camera, d i is the feature point p i Depth; pass Calculate the corresponding back-projection point , the formula is as follows: (3) Where Q i q i The back projection point, K is the intrinsic parameter matrix of the image sensor in the RGB-D camera, d i is the feature point q i Depth; According to the back-projection point , calculate the spatial structure matrix M ref , the formula is as follows: (4) (5) Where, For the Feature Points The back projection point of Represents the back-projection point and the distance between them; According to the back-projection point , calculate the spatial structure matrix M cur , the formula is as follows: (6) (7) Where, Represents the back-projection point and the distance between them; Based on the spatial structure matrix M ref and M cur , calculate M c-r , the formula is as follows: (8) Where, Represents the spatial structure matrix and The ratio between Represents the back-projection point and The distance between and back-projection point and The distance between The ratio between according to Calculated by formula (9) , the formula is as follows: (9) Create an empty integer array ; Using formula (10) to traverse All fractional expressions in the formula are as follows: (10) in, is a matrix In the Rank Elements of the column, ; θ is the threshold value. If formula (10) holds, the corresponding 、 Insert into ; Match the acquired feature points The subscripts of the feature point matching pairs appear in the array The feature point matching pairs in are set as dynamic matching pairs and are eliminated.
8. The RGB-D SLAM method suitable for dynamic environments according to claim 3, wherein: The steps of eliminating dynamic matching pairs using the epipolar constraint consistency algorithm include: Get depth misaligned feature point matching pairs , The number of feature point matching pairs indicating depth misalignment; Get feature points The pose of the frame and feature points The pose of the frame ; Calculate the pose and posture The relative position between , the formula is as follows: (11) Where, is the relative rotation matrix, is the relative translation vector; According to the relative posture , calculate the essential matrix between the current frame and the previous frame , the formula is as follows: (12) Based on the essential matrix , calculate the threshold of matching error, the formula is as follows: (13) In the formula, x is the confidence level. The larger its value is, the lower the threshold of matching error is. The feature point matching pairs obtained in turn by formula (14) Each pair of feature point matching pairs The matching error in , the formula is as follows: (14) in, is the intrinsic parameter matrix of the image sensor in the RGB-D camera, Representative Matching pairs of feature points, ; If the calculated , then the feature points used are matched to Set as a dynamic matching pair and remove it.
9. The RGB-DSLAM method applicable to dynamic environments according to claim 3, characterized in that: The steps for processing new map points using the map point processing algorithm include: Get new map points ; Calculate new map points Distance from the camera's optical center : (15) Where, is the translation vector of the current frame’s pose; Determine new map points Distance from the camera's optical center Is it greater than the preset distance threshold? If it is greater, use the DBoW3 keyframe database to calculate Node number in DBoW3 , and record the node in Parameters; Get the node number value in the active submap equal to All map points , Indicates the number of map points obtained; Calculate in sequence and Similarity: (16) Where, yes and The similarity between It is a map point The descriptor of It is a map point The descriptor of ; judge and Similarity Is less than the preset similarity threshold, if so, then Remove from the active submap and Insert the active submap.
Citation Information
Patent Citations
Mobile robot semantic map construction system based on monocular vision
CN111368759A
RGB-D simultaneous localization and mapping method integrating direct method and feature method
CN111462207A