Dynamic environment vision SLAM (Simultaneous Localization and Mapping) method of instance segmentation dominant and optical flow auxiliary verification based on YOLO11s-seg and RAFT (Reversible Addition Fourier Transform) optical flow calculation
By introducing the instance segmentation-led plus optical flow assisted verification method of YOLO11s-seg and RAFT optical flow calculation in visual SLAM, the problem of traditional SLAM incorrectly matched feature points in dynamic environments is solved, and higher positioning accuracy and map construction quality are achieved.
Patent Information
- Application Number
- CN202510269140.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-10
AI Technical Summary
Traditional visual SLAM methods easily mismatch the feature points of dynamic objects as static environment feature points in dynamic environments, resulting in increased positioning errors and inaccurate map construction.
Using the method of instance segmentation-led plus optical flow assisted verification based on YOLO11s-seg and RAFT optical flow calculation, dynamic feature points are obtained through instance segmentation, and combined with optical flow analysis and strict optical flow threshold and spatial clustering methods, dynamic feature points are verified and removed, leaving static feature points for ORB-SLAM3 tracking thread.
It significantly improves the positioning accuracy and map construction quality of visual SLAM in dynamic environments, can handle dynamic feature points more accurately, and provides more reliable autonomous navigation support.
Smart Images

Figure CN120125665A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and simultaneous localization and mapping, and more particularly to a dynamic environment visual SLAM method dominated by instance segmentation and assisted by optical flow verification based on YOLO11s-seg and RAFT optical flow calculation. Background Art
[0002] Simultaneous Localization and Mapping (SLAM) technology is one of the key technologies for mobile robots to achieve autonomous navigation, aiming to enable robots to build maps in real time in unknown environments and determine their own positions in the maps. As an important branch of SLAM technology, visual SLAM has been widely used in many fields such as robotics, autonomous driving, and augmented reality by processing the image data collected by cameras, thanks to advantages such as low camera cost and rich information.
[0003] The YOLO series of models are well-known for their fast and accurate object detection capabilities and perform excellently in real-time object detection tasks. As the latest improved version of the YOLO series, YOLO11 not only has powerful object detection functions but also has excellent performance in instance segmentation, being able to effectively identify different categories of objects and their boundaries, providing rich instance segmentation information for visual SLAM. The YOLO11s-seg model has high accuracy and high processing speed in object recognition and instance segmentation due to its advantages of balancing the number of parameters and processing speed. The RAFT model is an advanced optical flow calculation model. Optical flow refers to the motion information of objects in images between two consecutive frames. The RAFT model can accurately calculate the optical flow field by deeply analyzing the image features, providing the motion information of objects for visual SLAM and accurately judging possible dynamic objects in the environment.
[0004] Although traditional SLAM research methods have achieved certain success in static environments, they still face many deficiencies in actual complex dynamic scenarios. Traditional visual SLAM methods usually perform localization and map construction based on feature point matching. Among them, ORB (Oriented FAST and Rotated BRIEF) feature points are widely used due to their high computational efficiency and good robustness. However, in some dynamic scenarios, ORB feature points are easily affected by dynamic objects and may wrongly match the feature points on dynamic objects as static environment feature points, resulting in problems such as increased positioning errors and inaccurate map construction.
[0005] In summary, it is of great practical significance to study a visual SLAM method that can remove dynamic ORB feature points in a dynamic environment and balance processing accuracy and speed. Summary of the Invention
[0006] The "Visual SLAM Method for Removing Dynamic ORB Feature Points with Instance Segmentation Dominance and Optical Flow Aided Verification Based on YOLO11s-seg and RAFT Optical Flow Calculation" proposed by the present invention aims to overcome the problem in visual SLAM that dynamic objects are misregarded as static environments for feature matching, resulting in increased positioning errors and inaccurate map construction.
[0007] By introducing the YOLO11s-seg model, it is possible to classify and segment objects in the scene using instance segmentation information to obtain possible dynamic feature points. Further combining the RAFT optical flow calculation model, optical flow analysis is performed on the extracted feature points. For those feature points with abnormal optical flow changes, strict optical flow thresholds and spatial clustering methods are used to determine whether they are dynamic feature points. Finally, the dynamic feature points obtained by YOLO11s-seg and the dynamic feature points calculated by the RAFT optical flow model are merged, and the dynamic feature points are removed from the initial ORB feature point set, leaving static feature points. The static feature points are input into the tracking thread of ORB-SLAM3, thereby effectively improving the accuracy of dynamic SLAM.
[0008] The dynamic environment visual SLAM method with instance segmentation dominance and optical flow aided verification based on YOLO11s-seg and RAFT optical flow calculation proposed by the present invention mainly includes the following steps:
[0009] S1: Use an RGB-D camera to synchronously acquire RGB color images and depth maps, and at the same time perform RGB image correction and depth map alignment according to the camera calibration internal parameters.
[0010] S2: Extract ORB feature points from the RGB image processed in S1. By performing multi-scale FAST corner detection, gray centroid direction assignment, and rotation-corrected BRIEF descriptor generation, highly discriminative ORB feature points are obtained.
[0011] S3: Feed the RGB image processed in S1 into the trained YOLO11s-seg instance segmentation model to obtain the detection boxes, label categories, and binary mask matrices of each instance of the object. According to preset dynamic label categories such as dynamic category labels of people, vehicles, etc., the initial dynamic instance mask regions are distinguished. At the same time, the initial dynamic instance mask regions are further dilated, and morphological dilation is performed on the masks to avoid the remaining of edge feature points of dynamic objects. Finally, all dynamic instance masks are logically ORed and merged into the global instance segmentation dynamic region.
[0012] S4: Input the obtained two consecutive frames of images into the RAFT model. First, the deep features of the two frames of images are extracted by the feature encoder, and at the same time, the current frame is input into the context encoder to extract the context features of the current frame. Then, the deep features of the two frames of images extracted are sent to the visual similarity module to construct the correlation volume and build the correlation pyramid. After that, the current optical flow state, hidden state, correlation features, and context features are sent to the iterative update module, and the optical flow estimation is iteratively optimized through the GRU unit to gradually approximate the true optical flow and output the dense optical flow field. Finally, calculate the optical flow magnitude, and generate the binary mask M of the optical flow dynamic region by setting a threshold. flow , distinguishing the dynamic region and the static region.
[0013] S5: According to the obtained instance segmentation mask M seg and the optical flow motion mask M flow judge the original ORB feature point set, adopting the method of instance segmentation dominance and optical flow-assisted verification. First, directly screen out most of the dynamic feature points through the instance segmentation mask obtained by high-confidence instance segmentation to obtain the instance segmentation dynamic point set. Then, the optical flow-assisted verification screens the feature point set P that is not covered by the instance segmentation but marked as dynamic by the optical flow. candidate , and use a strict optical flow threshold and spatial clustering method to filter P candidate to obtain the filtered dynamic point set. The filtered dynamic point set is further filtered by the depth map to judge the depth threshold to filter out the distant dynamic points to obtain the final optical flow dynamic point set. Finally, the obtained instance segmentation dynamic point set and the optical flow dynamic point set are merged to obtain the final dynamic point set.
[0014] S6: Use the ORB feature point set P obtained in S2 ORB to remove the final dynamic point set obtained in S5 to generate the static ORB feature point set P static , and input P static into the ORB-SLAM3 tracking thread, and perform visual SLAM in the dynamic environment through the tracking thread, local mapping thread, loop closing and map merging thread, and global BA optimization thread of ORB-SLAM3.
[0015] Preferably, the distinguishing of the initial dynamic instance mask region in step S3 specifically includes:
[0016] The processed RGB image is fed into the trained YOLO11s-seg instance segmentation model. Through the Backbone main network, feature extraction is performed to obtain multi-scale feature maps. Then, through the Neck network, feature adjustment and fusion are carried out to enhance the expression ability of features at different scales. Finally, through the Head network, the detection head predicts the bounding box, class, and confidence, and the instance mask is generated through the segmentation head. Further, according to dynamic label classes such as people, vehicles, etc., the initial dynamic instance mask M is distinguished seg-initial .
[0017] Preferably, the morphological dilation of the initial dynamic instance mask region in step S3 specifically includes:
[0018] Morphological dilation of the initial dynamic instance mask region is performed to avoid the residue of edge feature points of dynamic objects. The morphological dilation uses a 5x5 elliptical kernel to dilate the mask, and the formula is as follows:
[0019]
[0020] In the formula, kernel is a 5x5 elliptical structuring element, is the dilation operator.
[0021] After that, the dilated dynamic instance mask M seg ′ is logically OR combined into the global instance segmentation dynamic region, and finally the global instance segmentation dynamic region binary mask M seg is output, marking potential dynamic elements.
[0022] Preferably, in step S4, calculating the optical flow magnitude to generate the optical flow dynamic region binary mask M flow , distinguishing the dynamic region and the static region specifically includes:
[0023] According to the obtained dense optical flow field f final motion region extraction is performed, and the optical flow magnitude is calculated to extract the optical flow dynamic region. The formula is as follows:
[0024]
[0025] In the formula, Δx, Δy represent the optical flow displacement of the pixel point in the x direction and y direction, (u, v) represents the pixel coordinates, and τ is the threshold, which is set to 1.5.
[0026] If the optical flow displacement of a certain pixel point is greater than the threshold, it is judged as significant motion, otherwise it is static. Thus, the optical flow binary motion mask M flow is finally obtained, where the dynamic region is marked as 1.
[0027] Preferably, in step S5, most dynamic feature points are directly screened out through the instance segmentation with high confidence to obtain the instance segmentation dynamic point set. Specifically, it includes:
[0028] The instance segmentation dynamic point set is obtained according to the instance segmentation mask with high confidence, and the input is the original ORB feature point set P. ORB and the instance segmentation mask M. seg , and the determination rule is as follows:
[0029]
[0030] where M seg (x i , y i ) is the value of the instance segmentation mask at (x i , y i ). If the instance segmentation mask value of the ORB feature point is 1, it is judged as a dynamic feature point and added to the instance segmentation dynamic feature point set.
[0031] Preferably, in step S5, the optical flow assisted verification is used to screen the feature point set P that is not covered by the instance segmentation but is marked as dynamic by the optical flow. candidate , and a strict optical flow threshold and spatial clustering method are used to filter P candidate to obtain the filtered dynamic point set P. filter Specifically, it includes:
[0032] The points that are not covered by the instance segmentation but are marked as dynamic by the optical flow are screened according to the optical flow assisted verification, and the input is the original ORB feature point set P. ORB , the instance segmentation mask M. seg , the optical flow mask M. flow , and the determination rule is as follows:
[0033] P candidate = {p i ∈ P ORB | M seg (p i ) = 0 ∧ M flow (p i ) = 1}
[0034] The candidate point set P that is judged as dynamic by the optical flow but static by the instance segmentation is obtained. candidate .
[0035] Then, strict optical flow threshold verification is performed on P. candidate Given the optical flow field, first, the optical flow amplitude is calculated. Then, the threshold determination is performed, and the determination rule is as follows:
[0036]
[0037] At this time, the strict optical flow threshold τ strict is set to 3.0 pixels. If the optical flow threshold calculated from the optical flow field in the candidate points is greater than the strict optical flow threshold τ strict then it is determined as a dynamic point, and the strict optical flow threshold dynamic point set is obtained
[0038] Furthermore, for the candidate point set P candidate the remaining candidate feature points that do not pass the strict optical flow threshold test are subjected to spatial clustering verification. The DBSCAN clustering is used to determine the dynamic clusters, and the determination rules are as follows:
[0039]
[0040] At this time, τ in the formula cluster is set to 2.0 pixels. If the proportion of the optical flow amplitudes of the points in the obtained clustering cluster that are greater than τ is higher than 0.6, then this clustering cluster is determined as a dynamic cluster. Then all the dynamic clusters are merged to obtain the clustering dynamic point set
[0041] Finally, the obtained strict optical flow threshold dynamic point set and the clustering dynamic point set are merged into the filtered dynamic point set
[0042] Preferably, in step S5, the filtered dynamic point set is further filtered for distant dynamic points using the depth map to obtain the final optical flow dynamic point set Specifically, it includes:
[0043] Perform depth verification on the obtained filtered dynamic point set, input the filtered dynamic point set and the depth map to filter distant dynamic points. The determination rules are as follows:
[0044]
[0045] The depth threshold z in the formula th is set to 5.0 meters. Points less than or equal to the depth threshold z th are determined as dynamic points, and finally the optical flow dynamic point set is obtained
[0046] Finally, the obtained instance segmentation dynamic point set and the optical flow dynamic point set are merged to obtain the final dynamic point set
[0047] The beneficial effects of the present invention are:
[0048] This method of instance segmentation leading and optical flow assisted verification can handle dynamic feature points in visual SLAM more accurately, significantly improve the positioning accuracy and map construction quality of visual SLAM in dynamic environments, provide more reliable technical support for the autonomous navigation of mobile robots in complex dynamic scenarios, and has important theoretical significance and practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flowchart of a dynamic environment visual SLAM method based on YOLO11s-seg and RAFT optical flow calculation with instance segmentation leading and optical flow assisted verification provided by an embodiment of the present invention.
[0050] Figure 2 It is a schematic diagram of the YOLO11 network structure provided by an embodiment of the present invention.
[0051] Figure 3 It is a schematic diagram of the RAFT network structure provided by an embodiment of the present invention.
[0052] Figure 4 It is a detailed flowchart of the instance segmentation leading and optical flow assisted verification module provided by an embodiment of the present invention.
[0053] Figure 5 It is an overall framework diagram of applying a dynamic environment visual SLAM method based on YOLO11s-seg and RAFT optical flow calculation with instance segmentation leading and optical flow assisted verification to ORB-SLAM3 provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The following elaborates on the preferred embodiments of the present invention in conjunction with the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more definite definition of the protection scope of the present invention.
[0055] As Figure 1 shown, the present invention discloses a dynamic environment visual SLAM method based on YOLO11s-seg and RAFT optical flow calculation with instance segmentation leading and optical flow assisted verification, including the following steps:
[0056] S1: Use an RGB-D camera to synchronously acquire RGB color images and depth maps, and at the same time perform RGB image correction and depth map alignment according to the camera calibration internal parameters.
[0057] The specific embodiments are as follows:
[0058] S1.1 RGB image correction: Use the cv2.undistort() function of OpenCV, input the original image, the internal parameter matrix and the distortion coefficient D = [k1 , k 2 , p 1 , p 2 . According to the camera calibration parameters including the internal parameter matrix K, the radial distortion coefficient k 1 , k 2 , the tangential distortion coefficient p 1 , p 2 , the bilinear mapping interpolation method is used for undistortion to obtain an undistorted RGB image.
[0059] S1.2 Depth map alignment: The depth map is mapped to the RGB coordinate system through the coordinate transformation formula as follows:
[0060]
[0061] In the formula, (u depth , v depth ) is the pixel coordinate of the depth image, K depth is the internal parameter matrix of the depth camera, K color is the internal parameter matrix of the color camera, T depth2color is the external parameter matrix from the depth camera to the color camera including rotation and translation, (u color , v color ) is the pixel coordinate of the aligned color image.
[0062] After that, the holes in the mapped depth map are filled by bilinear interpolation to obtain a depth map aligned at the pixel level of the RGB image.
[0063] S2: Extract ORB feature points from the RGB image processed in S1. By performing multi-scale FAST corner detection, gray centroid direction assignment, and rotation-corrected BRIEF descriptor generation, high-distinguishability ORB feature points are obtained.
[0064] The specific embodiments are as follows:
[0065] S2.1 In step S2, the specific operation of multi-scale FAST corner detection is as follows:
[0066] First, construct an image pyramid. The number of pyramid layers is set to 4, and the scaling factor is 0.5, that is, the resolution of each layer is 1 / 2, 1 / 4, 1 / 8, 1 / 16 of the original image in turn. Multi-scale detection is used to adapt to objects at different distances. Then perform FAST corner detection. On each layer of the pyramid image, 16 pixel points are selected on the circumference with a radius of 3 centered on pixel p. If the gray value difference between 9 consecutive pixel points and p exceeds the threshold of 12, then p is determined as a corner point. At the same time, to avoid excessive corner point aggregation, retain the feature point with the highest response value and perform non-maximum suppression (NMS) operation. Calculate the maximum gray difference between consecutive points on the circumference and the center point, that is, the corner response value R, and only retain the corner point with the largest R in the 3x3 neighborhood.
[0067] In step S2.2, the specific operation of feature point direction assignment by the gray centroid method is as follows:
[0068] First, calculate the gray centroid. Taking the corner point as the center, the centroid of the neighborhood within 31x31 pixels is calculated using the following formula:
[0069]
[0070] where x and y are the coordinate offsets of the pixels in the neighborhood relative to the center point, ranging from -15 to 15 as integers, I(x, y) is the gray value at the coordinate (x, y), and m pq is the geometric moment of the image.
[0071] According to the obtained m 01 and m 10 , calculate the direction of the gray centroid through the first moment, and at the same time, quantize the direction θ into 8 directions: 0°, 45°, 90°, ……, 315°, reducing the computational complexity.
[0072] In step S2.3, the specific operation of generating the rotated and corrected BRIEF descriptor is as follows:
[0073] First, perform predefined sampling pairs, randomly generate 256 pairs of coordinates (a i , b i ) within the 31x31 neighborhood. Then, rotate the sampling point pairs according to the main direction θ, and the formula is as follows:
[0074]
[0075] where (a i , b i ) are the coordinates of the original sampling point pairs, and (a i ′, b i ′) are the coordinates of the rotated sampling point pairs.
[0076] Use the rotated coordinate pairs (a i ′, b i ′) to perform gray comparison for each pair of sampling points and generate binary bits Then, splice the binary strings, and splice the 256 comparison results into a 256-bit descriptor. Finally, obtain the feature point set P ORB ={p i |p i =(x, y, θ, d)}, where x and y are the pixel coordinates of the feature points, θ is the main direction in radians, and d is the 256-bit BRIEF descriptor. Through the above implementation, high-distinguishability ORB feature points are obtained
[0077] S3: Feed the RGB image processed in S1 into the trained YOLO11s-seg instance segmentation model to obtain the detection bounding boxes, label categories, and binary mask matrices of each instance. Distinguish the initial dynamic instance mask regions according to preset dynamic label categories such as people, vehicles, etc. At the same time, further dilate the initial dynamic instance mask regions, and perform morphological dilation on the masks to avoid the remaining of edge feature points of dynamic objects. Finally, logically OR and merge all the dynamic instance masks into the global instance segmentation dynamic region.
[0078] The specific embodiments are as follows:
[0079] In step S3.1 of S3, the specific operation of distinguishing the initial dynamic instance mask regions is as follows:
[0080] The network structure of YOLO11 is as Figure 2 shown. Feed the processed RGB image into the trained YOLO11s-seg instance segmentation model. After feature extraction by the Backbone main network, obtain multi-scale feature maps. Then, through the Neck network, perform feature adjustment and fusion to enhance the expression ability of features at different scales. Finally, through the Head network, predict the bounding boxes, categories, and confidences through the detection head, and generate instance masks through the segmentation head. Further distinguish the initial dynamic instance mask M seg-initial .
[0081] In step S3.2 of S3, the morphological dilation of the initial dynamic instance mask regions specifically includes:
[0082] Perform morphological dilation on the initial dynamic instance mask regions to avoid the remaining of edge feature points of dynamic objects. Use a 5x5 elliptical kernel for the morphological dilation operation on the masks. The formula is as follows:
[0083]
[0084] In the formula, kernel is a 5x5 elliptical structuring element, is the dilation operator.
[0085] After that, logically OR and merge the dilated dynamic instance mask M seg ′ into the global instance segmentation dynamic region. Finally, output the binary mask M seg of the global instance segmentation dynamic region to mark potential dynamic elements
[0086] S4: Input the obtained two consecutive frames of images into the RAFT model. First, the deep features of the two frames of images are extracted by the feature encoder, and at the same time, the current frame is input into the context encoder to extract the context features of the current frame. Then, the deep features of the two frames of images extracted are sent to the visual similarity module for correlation volume construction and construction of a correlation pyramid. After that, the current optical flow state, hidden state, correlation features, and context features are sent to the iterative update module, and the optical flow estimation is iteratively optimized through the GRU unit to gradually approximate the true optical flow and output a dense optical flow field. Finally, calculate the optical flow magnitude, and generate a binary mask M of the optical flow dynamic region by setting a threshold. flow , distinguishing the dynamic region and the static region.
[0087] The specific embodiments are as follows:
[0088] In step S4.1 of S4, the specific operations of extracting the deep features of the two frames of images by the feature encoder and inputting the current frame into the context encoder to extract the context features of the current frame are as follows:
[0089] The RAFT network structure is as Figure 3 shown. First, input two consecutive frames of RGB images, the current frame and the next frame, and normalize their pixel values to [-1, 1], and output an image pair (I t , I t+1 ) with the same size after normalization. Send the processed RGB images into the RAFT model, and extract the deep features of the two frames of images through the feature encoder for constructing the correlation volume. The network structure of the feature encoder consists of 6 convolutional layers, followed by the ReLU function activation for each layer, and the resolution of the output feature map is reduced to 1 / 8 of the input. The formula is as follows:
[0090]
[0091] In the formula, g 1 is the feature map of the current frame, and g 2 is the feature map of the next frame.
[0092] At the same time, input the current frame into the context encoder, whose structure is the same as that of the feature encoder but with independent parameters, to extract the context features of the current frame, and obtain the context features for initializing the hidden state of the GRU.
[0093] In step S4.2 of S4, the specific operations of sending the deep features of the two frames of images extracted to the visual similarity module for correlation volume construction and construction of a correlation pyramid are as follows:
[0094] Calculate the similarity between all pixel pairs of the two feature maps. First, expand the obtained g1 and g2 into vector forms Calculate the dot product similarity lattice Then reshape the matrix into a 4D tensor, construct a correlation pyramid, build multi-scale correlation volumes through pooling to support matching in different displacement ranges, and finally obtain a multi-scale correlation pyramid for subsequent iterative search.
[0095] In step S4.3 of S4, the optical flow estimation is iteratively optimized through GRU units to gradually approximate the true optical flow output. The specific operations are as follows:
[0096] First, according to the current optical flow estimation f k , retrieve local correlation features from the correlation pyramid, and then input the retrieved correlation features, the current optical flow f k , the context feature h 0 and the hidden feature h k into the GRU update unit. The formula is as follows:
[0097] h k+1 ,Δf = GRU(Concat(Lookup(C),f k ,h 0 ),h k )
[0098] In the formula, Lookup(C) is the retrieved correlation feature, f k is the current optical flow, h 0 is the context feature, and h k is the hidden feature.
[0099] Obtain the updated hidden state h k+1 and the optical flow residual Δf. Finally, perform optical flow update f k+1 = f k + Δf to make the optical flow estimation gradually approximate the true optical flow. Then perform an optical flow upsampling operation. First, upsample bilinearly and refine the edges using a convolutional network, i.e., convex upsampling, to upsample the low-resolution optical flow field to the original image resolution H×W to obtain a dense optical flow field f final .
[0100] In step S4.4 of S4, calculate the optical flow magnitude to generate a binary mask M flow of the optical flow dynamic region to distinguish the dynamic region and the static region. The specific operations are as follows:
[0101] Extract the motion region according to the obtained dense optical flow field f final , and calculate the optical flow magnitude to extract the optical flow dynamic region. The formula is as follows:
[0102]
[0103] Where Δx and Δy represent the optical flow displacement of the pixel point in the x and y directions, (u, v) represents the pixel coordinates, and τ is the threshold, which is set to 1.5.
[0104] If the optical flow displacement of a certain pixel point is greater than the threshold, it is judged as significant motion, otherwise it is static, so as to finally obtain the optical flow binary motion mask M flow , where the dynamic area is marked as 1.
[0105] S5: According to the obtained instance segmentation mask M seg And the optical flow motion mask M flow Judge the original ORB feature point set, adopting the method of instance segmentation dominance and optical flow assisted verification. First, directly screen out most of the dynamic feature points through the instance segmentation mask obtained by high-confidence instance segmentation to obtain the instance segmentation dynamic point set After that, the optical flow assisted verification filters the feature point set P that is not covered by the instance segmentation but is marked as dynamic by the optical flow candidate , and uses a strict optical flow threshold and spatial clustering method for P candidate To filter and obtain the filtered dynamic point set The filtered dynamic point set Is further used to judge the depth threshold by the depth map to filter out the distant dynamic points to obtain the final optical flow dynamic point set Finally, the obtained instance segmentation dynamic point set And the optical flow dynamic point set Are merged to obtain the final dynamic point set
[0106] The specific embodiments are as follows:
[0107] As Figure 4 Shown, the specific operation of the instance segmentation dominance and optical flow assisted verification module is as follows:
[0108] In step S5.1 of S5, directly screen out most of the dynamic feature points through the mask obtained by high-confidence instance segmentation to obtain the instance segmentation dynamic point set Specifically include:
[0109] Obtain the instance segmentation dynamic point set according to the high-confidence instance segmentation mask, and the input is the original ORB feature point set P ORB And the instance segmentation mask M seg , and the determination rule is as follows:
[0110]
[0111] Where M seg (x i , y i ) is the instance segmentation mask at (x i , yi ) value. If the instance segmentation mask value of the ORB feature point is 1, it is judged as a dynamic feature point and added to the instance segmentation dynamic feature point set
[0112] S5.2 In step S5, the optical flow-assisted verification is used to screen the set P of feature points that are not covered by instance segmentation but are marked as dynamic by optical flow candidate , and a strict optical flow threshold and spatial clustering method are used for P candidate to obtain the filtered dynamic point set P filter Specifically, it includes:
[0113] According to the optical flow-assisted verification, screen the points that are not covered by instance segmentation but are marked as dynamic by optical flow. The input is the original set P of ORB feature points ORB , the instance segmentation mask M seg , the optical flow mask M flow , and the determination rule is as follows:
[0114] P candidate ={p i ∈P ORB |M seg (p i ) = 0 ∧ M flow (p i ) = 1}
[0115] Get the set P of candidate points that are judged as dynamic by optical flow but judged as static by instance segmentation candidate .
[0116] Then perform strict optical flow threshold verification on P candidate . Input the optical flow field, first calculate the optical flow amplitude , and then perform threshold determination. The determination rule is as follows:
[0117] At this time, the strict optical flow threshold τ strict is set to 3.0 pixels. If the optical flow threshold calculated according to the optical flow field among the candidate points is greater than the strict optical flow threshold τ strict , it is judged as a dynamic point, and the strict optical flow threshold dynamic point set is obtained
[0118] Further, perform spatial clustering verification on the remaining candidate feature points in the candidate point set P candidate that do not pass the strict optical flow threshold test. Use DBSCAN clustering to determine the dynamic cluster. The determination rule is as follows:
[0119]
[0120] At this time, τ in the formula clusterSet to 2.0 pixels. If the proportion of the optical flow amplitude of the points in the obtained cluster that is greater than τ is higher than 0.6, then this cluster is determined to be a dynamic cluster. Then, all the dynamic clusters are merged to obtain a set of clustering dynamic points
[0121] The obtained strict optical flow threshold dynamic point set and the clustering dynamic point set are merged into a filtered dynamic point set
[0122] In step S5.3 of S5, the filtered dynamic point set Is further used to judge the depth threshold by the depth map to filter out distant dynamic points to obtain the final optical flow dynamic point set Specifically include:
[0123] Perform depth verification on the obtained filtered dynamic point set, input the filtered dynamic point set And the depth map Filter out distant dynamic points, and the judgment rule is as follows:
[0124]
[0125] The depth threshold z in the formula th Is set to 5.0 meters. Points less than or equal to the depth threshold z th Are judged as dynamic points, and finally the optical flow dynamic point set is obtained
[0126] Finally, the obtained instance segmentation dynamic point set And the optical flow dynamic point set Are merged to obtain the final dynamic point set
[0127] S6: As Figure 5 Shown, use the ORB feature point set P obtained in S2 ORB To remove the final dynamic point set obtained in S5 To generate a static ORB feature point set P static , Input P static Into the ORB-SLAM3 tracking thread, and perform dynamic environment visual SLAM through the tracking thread, local mapping thread, loop closing and map merging thread, and global BA optimization thread of ORB-SLAM3
[0128] Finally, it should be noted that the above are only preferred embodiments of the present invention for illustrative purposes and are not intended to imply that the scope of the present invention is limited to these examples; under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above, which are not provided in detail for the sake of brevity. The present invention is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A dynamic environment visual SLAM method based on instance segmentation dominated by YOLO11s-seg and RAFT optical flow calculation plus optical flow assisted verification, characterized in that: The following steps are involved: S1: Use an RGB-D camera to synchronously acquire RGB color images and depth maps, and perform RGB image correction and depth map alignment based on the camera calibration intrinsic parameters. S2: Extract ORB feature points from the RGB image processed by S1, and obtain highly discriminative ORB feature points by performing multi-scale FAST corner detection, grayscale centroid direction assignment, and rotation-corrected BRIEF descriptor generation. S3: The RGB image processed by S1 is sent to the trained YOLO11s-seg instance segmentation model to obtain the object detection box, label category and binary mask matrix of each instance. The initial dynamic instance mask area is distinguished according to the preset dynamic label categories such as human, vehicle and other dynamic category labels, and the initial dynamic instance mask area is further expanded. The mask is morphologically expanded to avoid residual edge feature points of dynamic objects. Finally, all dynamic instance masks are logically or merged into the global instance segmentation dynamic area. S4: The two consecutive frames of images are input into the RAFT model. First, the deep features of the two frames are extracted through the feature encoder. At the same time, the current frame is input into the context encoder to extract the context features of the current frame. Then, the extracted deep features of the two frames are sent to the visual similarity module to construct the relevant volume and the correlation pyramid. The current optical flow state, hidden state, relevant features and context features are further sent to the iterative update module. The optical flow estimation is iteratively optimized through the GRU unit to gradually approach the real optical flow output dense optical flow field Finally, the optical flow amplitude is calculated and the optical flow dynamic area binary mask M is generated by setting the threshold. flow , distinguish between dynamic area and static area. S5: Segment the mask M based on the obtained instance seg and the optical flow motion mask M flow The original ORB feature point set is judged by using instance segmentation-dominant and optical flow-assisted verification. First, the instance segmentation mask obtained by high-confidence instance segmentation directly selects most dynamic feature points to obtain the instance segmentation dynamic point set. Then, the optical flow assisted verification selects the feature point set P that is not covered by the instance segmentation but marked as dynamic by the optical flow. candidate , using strict optical flow threshold and spatial clustering method to candidate Filter to get the filtered dynamic point set The filtered dynamic point set The depth map is further used to determine the depth threshold to filter distant dynamic points to obtain the final optical flow dynamic point set. Finally, the obtained instance segmentation dynamic point set and optical flow dynamic point set Merge to get the final dynamic point set S6: ORB feature point set P obtained using S2 ORB The final dynamic point set obtained by S5 Eliminate and generate a static ORB feature point set P static , P static Input the ORB-SLAM3 tracking thread, and perform dynamic environment visual SLAM through ORB-SLAM3's tracking thread, local mapping thread, loop closure and map merging thread, and global BA optimization thread.
2. The method according to claim 1, characterized in that The initial dynamic instance mask area is distinguished in step S3 specifically including: The processed RGB image is sent to the trained YOLO11s-seg instance segmentation model, and the Backbone network is used for feature extraction to obtain a multi-scale feature map. Then, the Neck network is used for feature adjustment and fusion to enhance the expression of features of different scales. Finally, the Head network is used to predict the bounding box, category, and confidence through the detection head, and the instance mask is generated through the segmentation head. The initial dynamic instance mask M is further distinguished according to the dynamic label class, such as human, vehicle, and other dynamic category labels. seg-initial .
3. The method according to claim 1, characterized in that The morphological expansion of the initial dynamic instance mask region in step S3 specifically includes: For the initial dynamic instance mask M seg-initial Morphological dilation is performed to avoid residual feature points on the edges of dynamic objects. Morphological dilation uses a 5x5 elliptical kernel to dilate the mask. The formula is as follows: Where kernel is a 5x5 ellipse structure element, is the expansion operator. Then the expanded dynamic instance mask M seg ′ Logically OR and merge into the global instance segmentation dynamic region, and finally output the global instance segmentation dynamic region binary mask M seg ,Marking potential dynamic elements.
4. The method according to claim 1, characterized in that: In step S4, the optical flow amplitude is calculated to generate the optical flow dynamic area binary mask M flow , distinguishing between dynamic and static areas specifically includes: According to the obtained dense optical flow field f final Perform motion area extraction and calculate the optical flow amplitude to extract the optical flow dynamic area. The formula is as follows: Where Δx, Δy represent the optical flow displacement of the pixel in the x-direction and y-direction, (u, v) represents the pixel coordinates, and τ is the threshold value, which is set to 1.
5. If the optical flow displacement of a pixel is greater than the threshold, it is judged as significant motion, otherwise it is static, thus finally obtaining the optical flow binary motion mask M flow , where the dynamic area is marked as 1.
5. The method according to claim 1, characterized in that In step S5, the mask obtained by instance segmentation with high confidence directly selects most of the dynamic feature points to obtain the instance segmentation dynamic point set. Specifically include: According to the high-confidence instance segmentation mask, the instance segmentation dynamic point set is obtained, and the input is the original ORB feature point set P ORB and instance segmentation mask M seg , the judgment rules are as follows: Where M seg (x i ,y i ) is the instance segmentation mask in (x i ,y i ). If the instance segmentation mask value of the ORB feature point is 1, it is considered a dynamic feature point and added to the instance segmentation dynamic feature point set.
6. The method according to claim 1, characterized in that In step S5, the optical flow assisted verification selects the feature point set P that is not covered by the instance segmentation but marked as dynamic by the optical flow. candidate , and use strict optical flow threshold and spatial clustering method to candidate Filter to obtain the filtered dynamic point set P filter Specifically include: According to the optical flow assisted verification, points that are not covered by instance segmentation but marked as dynamic by optical flow are filtered, and the input is the original ORB feature point set P ORB , instance segmentation mask M seg , optical flow mask M flow , the judgment rules are as follows: P candidate ={p i ∈P ORB |M seg (p i )=0∧M flow (p i )=1} Get the set of selected points P that are judged as dynamic by optical flow but static by instance segmentation candidate . Then for P candidate Perform strict optical flow threshold verification, input the optical flow field, and first calculate the optical flow amplitude. Then the threshold is determined, and the determination rules are as follows: At this time, the strict optical flow threshold τ strict Set to 3.0 pixels, if the optical flow threshold calculated based on the optical flow field in the candidate point is greater than the strict optical flow threshold τ strict It is judged as a dynamic point, and a strict optical flow threshold dynamic point set is obtained. Further, the candidate point set P candidate The remaining candidate feature points that did not pass the strict optical flow threshold test were spatially clustered for verification, and DBSCAN clustering was used to determine the dynamic clusters. The determination rules are as follows: At this time, τ in the formula cluster Set to 2.0 pixels. If the ratio of the optical flow amplitude of the points in the cluster is greater than τ and is higher than 0.6, the cluster is determined to be a dynamic cluster. Then all dynamic clusters are merged to obtain a cluster dynamic point set. Finally, the obtained strict optical flow threshold dynamic point set and clustered dynamic point set are merged into the filtered dynamic point set.
7. The method according to claim 1, characterized in that In step S5, the filtered dynamic point set The depth map is further used to determine the depth threshold to filter distant dynamic points to obtain the final optical flow dynamic point set. Specifically include: Perform deep verification on the filtered dynamic point set obtained, and input the filtered dynamic point set and depth map Filter distant dynamic points, the judgment rules are as follows: The depth threshold z in the formula th Set to 5.0 meters, less than or equal to the depth threshold z th The points are judged as dynamic points, and finally the optical flow dynamic point set is obtained Finally, the obtained instance segmentation dynamic point set and optical flow dynamic point set Merge to get the final dynamic point set