An Optical Flow Semantic SLAM Method for Arbitrary Dynamic Objects

By adopting optical flow semantic SLAM method in the SLAM system, using deep learning and semantic segmentation technology to eliminate dynamic objects in complex scenarios, the problem of difficulty in identifying and eliminating dynamic objects in complex scenarios is solved, and more efficient feature point matching and higher positioning accuracy are achieved.

CN119399718BActive Publication Date: 2025-06-13ANHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411428026.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-06-13
Estimated Expiration
2044-10-14

AI Technical Summary

Technical Problem

Existing SLAM technology is difficult to effectively identify and eliminate unknown dynamic objects in complex scenarios, resulting in low feature point matching efficiency, reduced positioning accuracy and insufficient system robustness.

Method used

The optical flow semantic SLAM method for any dynamic object is adopted, and the images are acquired through the RGB-D depth camera, and the dynamic area mask is extracted using the deep learning-based optical flow estimation module and the significance area detection module. The dynamic area mask is eliminated by the semantic segmentation module, and the feature points are extracted and matched using a lightweight CNN architecture to improve positioning accuracy and robustness.

Benefits of technology

It effectively eliminates the interference of dynamic objects on the SLAM system, improves the efficiency and quality of feature point extraction and matching, and enhances the robustness and positioning accuracy of the system in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399718B_ABST
    Figure CN119399718B_ABST
Patent Text Reader

Abstract

The present invention proposes an optical flow semantic SLAM method for any dynamic object, including: obtaining continuous RGB images, and every eight frames of images, using a deep learning-based optical flow estimation module to estimate the optical flow of the current and the next frame of images to generate an optical flow map; using a saliency region detection module to extract the dynamic region in the optical flow map as a mask, and finding two representative points in the dynamic mask to prompt the semantic segmentation module to perform segmentation, so as to eliminate any dynamic object and remove the interference of the dynamic object to the system; using a lightweight CNN architecture to extract and match feature points from the image with the dynamic region removed, and sending the RGB image and the original depth map after feature matching points and removing the dynamic region into the ORB-SLAM2 system, and improving the accuracy of positioning and mapping by virtue of high-quality feature points. The present invention proposes a more efficient SLAM method, which can have higher positioning and mapping accuracy, and at the same time can have higher robustness in practical engineering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vision-based positioning and mapping of unmanned vehicles, and specifically to an optical flow semantic SLAM method for arbitrary dynamic objects. Background Art

[0002] SLAM (Simultaneous Localization and Mapping) is a technology for simultaneous localization and mapping. An accurate and efficient SLAM technology is a key factor to ensure precise navigation of devices such as robots and unmanned vehicles. Among them, ORB-SLAM2 is a relatively classic visual SLAM algorithm. This algorithm has high positioning and mapping accuracy in simple static scenarios with little light change, but it is not suitable for complex scenarios. For example, when the viewing angle fluctuates greatly, the dynamic blur seriously affects the quality of ORB feature points; in scenarios with dynamic objects, tracking failures or incorrect matching of feature points in dynamic regions often occur. And the current algorithms improved based on ORB-SLAM2 using deep learning still have the following defects: (1) It can only recognize pre-trained dynamic object categories. If an unknown dynamic object appears in the scene, it cannot effectively recognize the object and eliminate the dynamic feature points in its area. (2) The efficiency of feature point extraction is not high, seriously affecting the real-time performance of the system, and the rationality of feature point distribution is poor, reducing the positioning accuracy of the system. Therefore, based on the above two difficulties, a more efficient SLAM method needs to be designed to enable the unmanned vehicle to have higher positioning and mapping accuracy and to be more robust in practical engineering. Summary of the Invention

[0003] To overcome the problems faced by the above-mentioned unmanned vehicle positioning and mapping technologies based on deep learning object detection and traditional feature point extraction in complex environments, the present invention provides an optical flow semantic SLAM method for arbitrary dynamic objects. First, the unmanned vehicle obtains continuous RGB images through an RGB-D depth camera. Every eight frames of images, an optical flow estimation module based on deep learning is used to estimate the optical flow between the current and the next frame of images to generate an optical flow map. A saliency region detection module is used to extract the dynamic region in the optical flow map as a mask, and two representative points are found in the dynamic mask to prompt the semantic segmentation module for segmentation, so as to eliminate arbitrary dynamic objects and remove the interference of dynamic objects on the system. Then, a lightweight CNN architecture is used to extract and match feature points from the image with the dynamic region removed, and the feature points and the processed image are sent into the ORB-SLAM2 system to improve the positioning and mapping accuracy with high-quality feature points.

[0004] To achieve the above technical objectives, the present invention provides the following technical solutions:

[0005] An optical flow semantic SLAM method for arbitrary dynamic objects, specifically including the following steps:

[0006] S1. Obtain consecutive RGB images. Every eight frames of images, use a deep learning-based optical flow estimation module to estimate the optical flow between the current and the next frame of images, generating an optical flow map;

[0007] S2. Use a saliency region detection module to extract the dynamic region in the optical flow map as a mask, and find two representative points in the dynamic mask;

[0008] S3. According to the obtained representative points, prompt the semantic segmentation module to segment the next frame of images, realizing the removal of arbitrary dynamic objects, and obtaining the mask segmented semantically;

[0009] S4. Further extract two representative points from the mask segmented semantically in step S3, prompt the semantic segmentation module to segment the dynamic object region of the next frame of images, until eight frames later, reuse the representative points obtained by the saliency region detection module to prompt the semantic segmentation module for segmentation;

[0010] S5. Use a lightweight CNN architecture to extract and match feature points from the images with dynamic objects removed;

[0011] S6. Feed the successfully matched feature points, the RGB images after removing dynamic objects, and the original depth map into the ORB-SLAM2 system, and use the BA algorithm to minimize the reprojection error, tracking and positioning the camera pose of each frame;

[0012] S7. When a new key frame is inserted into the local map, update the key frames in the local map and the connection relationships between the key frames, at the same time perform BA optimization on the poses of the key frames and the map points in the current local map, and finally delete the redundant key frames and their corresponding map points;

[0013] S8. Use the feature point extraction module to train a bag-of-words model, and perform loop closure detection, calculate the best candidate frame at the loop closure, replace the current frame, thereby correcting the pose and the local map;

[0014] S9. Use graph optimization to optimize all key frames and map points.

[0015] Furthermore, in step S1, the optical flow estimation module adopted is a variant of RAFT.

[0016] More specifically, the specific process of the optical flow estimation described in step S1 includes:

[0017] S11. Feature extraction; Downsample the input current and next frame of images through a feature encoding network, and extract their feature maps with a resolution of 1 / 8;

[0018] S12. Visual similarity calculation: Calculate the correlation volume using the two feature maps obtained in step S11, then use the multi-level correlation of the correlation pyramid to capture larger and smaller pixel displacements, and then generate a new feature map by indexing the features in the correlation pyramid at each level;

[0019] Given the current optical flow estimate (f 1 , f 2 ), for the current original RGB image I 1 , the next-frame original RGB image I 2 , map each pixel x = (u, v) of I 1 to I 2 to obtain an estimate x' of the pixels in the next-frame image; the mapping relationship is: x' = (u + f 1 (u), v + f 2 (v)); define a local search neighborhood S(x') r around x', and the formula is:

[0020] S(x') r = {x' + Δx | Δx ∈ R 2 , ||Δx|| 1 ≤ r};

[0021] where R 2 is the set of two-dimensional integer lattices, and r is the search radius on the two-dimensional image plane;

[0022] Based on the local search neighborhood S(x') r find a feature map for each layer of the correlation pyramid, and then concatenate these feature maps;

[0023] S13. Iterative update: Concatenate the feature maps obtained from each layer of the correlation pyramid in step S12 into a total feature map and input it to an update operator with a transformer module as the core to estimate the optical flow estimation sequence {f 0 , f 1 ,..., f n} of all pixels; starting from an initial starting point f 0 , this initial starting point f 0 is directly predicted by reusing the existing context encoder and feeding it the stacked input frames; in each iteration, a flow update direction Δf is generated and added to the current estimate, that is, the formula is: f k+1 = f k + Δf k ;

[0024] The loss function in the iterative process is the mixture Laplacian loss, and the network is trained to predict the mixture parameters of the Laplacian distribution to maximize the log-likelihood of the ground truth flow; given the image pair {I 1 ,I 2} and the optical flow ground truth μ gt , the training loss is:

[0025] L prob =-logp θ (μ=μ gt ∣I 1 ,I 2 );

[0026] where p θ () represents the probability density function of the Laplacian distribution.

[0027] Furthermore, in step S2, the saliency region detection model explicitly decomposes the binary image segmentation task on high-resolution data into a localization module and a reconstruction module; and a bilateral filter is used to preserve edge information and reduce noise.

[0028] More specifically, step S2 includes the following steps:

[0029] S21. Image preprocessing; enhancing the contrast of the optical flow map and using a bilateral filter to preserve edge information and reduce noise;

[0030] S22. Constructing the localization module; using a Transformer encoder to extract features at different stages; the features extracted at different stages are respectively denoted as: The features of the first three stages are transmitted to the corresponding decoder stage with lateral connections and are stacked and cascaded in the last encoder block to generate fused features After that, the fused features are fed into the classification module, and the atrous spatial pyramid pooling (ASPP) module is continued to perform multi-context fusion, and the fused features are compressed to obtain compressed features for transfer to the reconstruction module;

[0031] S23. Constructing the reconstruction module; using the compressed features as the input; for the extracted by the i-th Transformer encoder, deformable convolutions with hierarchical receptive fields and adaptive average pooling layers are used to extract features with receptive fields of various scales; then these features extracted by different receptive fields are concatenated as and then passed through a 1×1 convolutional layer and a batch normalization layer to obtain the output features processed by the reconstruction module The output features Finally, obtain the final prediction map through a 1×1 convolutional layer;

[0032] S24. Perform bilateral reference, which is divided into internal reference and external reference. The internal reference avoids resizing the image through adaptive cropping, and the external reference uses gradient labels to obtain regions with richer gradient information;

[0033] S25. Set the objective function. Use BCE, IoU, and SSIM loss functions to perform supervision at the pixel, region, and boundary levels respectively. The formula is expressed as:

[0034]

[0035] where L pixel 、L region and L boundary are the losses at the pixel, region, and boundary levels respectively. L BCE is the cross-entropy loss for binary classification problems; L IoU is the overlap loss between the predicted segmentation result and the true segmentation label; L SSIM is the structural similarity loss; k 1 、k 2 and k 3 are the loss weights respectively.

[0036] Furthermore, in step S3, the semantic segmentation module is a variant of EfficientViT. It uses EfficientViT to replace the image encoder of the Segment Anything Model (SAM), while retaining the structure of the lightweight prompt encoder and mask decoder, denoted as EfficientViT-SAM.

[0037] More specifically, step S3 specifically includes:

[0038] S31. The input image sequentially passes through three convolutional blocks and two EfficientViT modules to obtain features in five stages. Through upsampling and addition, the fused features of the last three stages are obtained and fed into the neck including several MBConv blocks, and then fed into the SAM head;

[0039] S32. Perform training. The training includes two stages: First, use the image encoder of SAM to train the image encoder of EfficientViT-SAM, and use the L2 loss as the loss function; then train the end-to-end network of EfficientViT-SAM on the entire SA-1B dataset;

[0040] S33. Extract mask representative points: Fit the center line of the mask through polynomial fitting, and adjust the fitting degree of the curve by adjusting the order degree of the polynomial. Take two points at one-sixth and five-sixths of the center line respectively to prompt the semantic segmentation model to segment the next image, that is, use the semantic segmentation result of the previous frame image to prompt the semantic segmentation model to segment the same moving object area in the next frame image, and achieve semantic association between adjacent images.

[0041] Further, step S5 specifically includes:

[0042] S51. Separate the key point detection into an independent branch, and use 1×1 convolution to perform fast processing on an 8×8 tensor block transformed image;

[0043] S52. Design a matching refinement module, and the learning of the module predicts pixel-level offsets by only considering the nearest neighbor pairs at 1 / 8 of the original spatial resolution in the original rough-level features;

[0044] S53. Divide the image into several appropriate equal parts according to the density of feature points in the previous frame image. The greater the density, the more parts are divided, and the maximum is 8 equal parts;

[0045] S54. Select feature points with better quality in each image block and regard them as a cluster. Calculate the average quality of all feature points in the image, and screen out feature points with poor quality according to the difference between the quality of each cluster and the average quality according to N = km; where m is the quality difference, k is a proportionality constant obtained through experiments, and N is the number of feature points to be screened out;

[0046] S55. Use the calculated descriptor used to describe the pixel information around the feature points to train a bag-of-words model containing environmental semantic information, and provide richer environmental semantic vocabulary for loop detection.

[0047] Based on the above technical solutions, the present invention has the following beneficial effects:

[0048] The method proposed by the present invention uses an optical flow estimation module, a saliency region detection module, and a semantic segmentation module based on deep learning to jointly achieve the elimination of any dynamic object. And replace the original ORB feature point extraction module in the ORB_SLAM2 system with a feature extraction and matching module based on a lightweight CNN network, and combine a reasonable screening strategy to extract and match high-quality feature points from the image with the dynamic area removed, and send the feature points into the ORB-SLAM2 system to complete higher-precision positioning.

[0049] The method proposed by the present invention eliminates the influence of any dynamic object in the environment on the SLAM system. Meanwhile, in the case of drastic fluctuations in the camera view and obvious changes in illumination, it can efficiently extract high-quality feature points, enhance the reliability of the SLAM front end, and further improve the robustness of the entire system in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a block diagram of an optical flow semantic SLAM system for any dynamic object provided by the present invention;

[0051] Figure 2 It is a block diagram of dynamic object elimination provided by the present invention;

[0052] Figure 3 It is a block diagram of feature point extraction and matching provided by the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to enable those skilled in the art of the present technology to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of this application.

[0054] Although the steps in the present invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein relates to and encompasses any and all possible combinations of one or more of the associated listed items.

[0055] The objects in the complex scenarios faced by existing SLAM technologies can be divided into three states: fully dynamic, that is, the object is always in motion in the camera view, such as a person walking back and forth; semi-dynamic, that is, the object is sometimes in motion and sometimes stationary in the camera view, such as a person walking sometimes or a table and chair moving due to an external force; fully static, that is, the object is always stationary in the camera view, such as a table and chair without an external force. In this embodiment, the method proposed by the present invention is based on an optical flow estimation module, a saliency region detection module, and a semantic segmentation module to detect the moving objects captured by the camera. At the same time, as Figure 2 shown. For the elimination of dynamic objects, it does not pay attention to the object attributes, but only focuses on the current motion state of the object to increase the utilization rate of feature points in the image by the system, thereby improving the positioning accuracy. In this embodiment, the method proposed by the present invention is applied to an unmanned vehicle, and the specific process is as Figure 1As shown in the figure, an optical flow semantic SLAM method for any dynamic object proposed by the present invention is given, including the following steps:

[0056] S1. The unmanned vehicle obtains continuous RGB images through an RGB-D depth camera. Every eight frames of images, an optical flow estimation module based on deep learning is used to estimate the optical flow of the current and the next frame of images, generating an optical flow map.

[0057] As a preferred implementation manner of step S1, in this embodiment, the adopted optical flow estimation module is a variant of RAFT; optical flow estimation is a basic task in low-level vision, aiming to estimate the 2D motion of each pixel between two adjacent frames, obtaining the motion field of the corresponding pixel points. By analyzing the difference in pixel optical flow values between the moving area and the static area in the motion field, the dynamic area is identified.

[0058] The specific process of optical flow estimation is as follows:

[0059] S11. Feature extraction; the input current and next frame of images are downsampled through a feature encoding network to extract feature maps with a resolution of 1 / 8.

[0060] S12. Visual similarity calculation; the correlation volume is calculated using the two feature maps obtained in step S11, and then the multi-level correlation of the correlation pyramid is used to capture larger and smaller pixel displacements. After that, a new feature map is generated by indexing the features in each level of the correlation pyramid.

[0061] Given the current optical flow estimation value (f 1 , f 2 ), for the current original RGB image I 1 , the next frame of original RGB image I 2 , each pixel x = (u, v) of I 1 is mapped to I 2 to obtain the estimated value x' of the pixel in the next frame of image; the mapping relationship is: x' = (u + f 1 (u), v + f 2 (v)); this formula reflects the displacement information of each pixel in image I 2 relative to the corresponding pixel in image I 1 in the u and v directions.

[0062] Then, a local search neighborhood S(x') r is defined around x', and the formula is expressed as:

[0063] S(x') r = {x' + Δx | Δx ∈ R 2 , ||Δx|| 1 ≤ r};

[0064] Among them, R 2 is a set of two-dimensional integer lattice points, and r is the search radius on the two-dimensional image plane;

[0065] Based on the local search domain S(x′) r For each layer of the correlation pyramid, a feature map is found, and then these feature maps are concatenated;

[0066] S13. Iterative update; the feature maps obtained from each layer of the correlation pyramid in step S12 are concatenated into a total feature map and input to an update operator with a transformer module as the core to estimate the optical flow estimation sequence {f 0 , f 1 ,..., f n}; starting from an initial starting point f 0 , this initial starting point f 0 is directly predicted by reusing the existing context encoder and feeding it the stacked input frames; in each iteration, a flow update direction Δf is generated and added to the current estimate, that is, expressed by the formula: f k+1 = f k + Δf k ;

[0067] The loss function in the iterative process is the mixed Laplacian loss, and the network is trained to predict the mixing parameters of the Laplacian distribution to maximize the log-likelihood of the ground truth flow; given the image pair {I 1 , I 2} and the optical flow ground truth μ gt , the training loss is:

[0068] L prob = -logp θ (μ = μ gt |I 1 , I 2 );

[0069] Among them, p θ () represents the probability density function of the Laplacian distribution.

[0070] S2. Use the saliency region detection module to extract the dynamic region in the optical flow map as a mask, and find two representative points in the dynamic mask;

[0071] As a preferred implementation manner of step S2, in this embodiment, the saliency region detection model explicitly decomposes the binary image segmentation task on high-resolution data into a localization module and a reconstruction module; and a bilateral filter is used to preserve edge information and reduce noise; the specific process includes:

[0072] S21. Image preprocessing; enhancing the contrast of the optical flow map and using a bilateral filter to preserve edge information and reduce noise;

[0073] S22. Construct a positioning module; use a Transformer encoder to extract features at different stages; the features extracted at different stages are respectively denoted as: The features of the first three stages are transmitted to the corresponding decoder stage with lateral connections and are stacked and cascaded in the last encoder block to generate fused features After that, the fused features are fed into the classification module, and the Atrous Spatial Pyramid Pooling (ASPP) module is continued to perform multi-context fusion on the fused features to obtain compressed features after compression for transfer to the reconstruction module;

[0074] S23. Construct a reconstruction module; use the compressed features as the input; the compressed features contain the features extracted at different stages in step S22. Therefore, in the reconstruction module, for the features extracted by the i-th Transformer encoder deformable convolutions with hierarchical receptive fields and adaptive average pooling layers are used to extract features with various scale receptive fields; then these features extracted by different receptive fields are concatenated as and then passed through a 1×1 convolutional layer and a batch normalization layer to obtain the output features processed by the reconstruction module The output features pass through a 1×1 convolutional layer to obtain the final prediction map;

[0075] S24. Perform bilateral reference; divided into internal reference and external reference; the internal reference avoids resizing the image through adaptive cropping; the external reference uses gradient labels to obtain regions with richer gradient information;

[0076] S25. Set the objective function; use BCE, IoU, and SSIM loss functions to perform supervision at the pixel, region, and boundary levels respectively. The formula is expressed as:

[0077]

[0078] where, L pixel , and L boundary are the losses at the pixel, region, and boundary levels respectively. L BCE is the cross-entropy loss for binary classification problems; L IoU is the overlap loss between the predicted segmentation result and the true segmentation label; L SSIM is the structural similarity loss; k 1, k 2 and k 3 are the respective loss weights.

[0079] The saliency region detection module detects dynamic regions by detecting relatively salient regions in the optical flow map generated by the optical flow estimation module, that is, regions with a relatively large color difference from the surrounding background. In the case of camera view fluctuations or large changes in ambient light, the optical flow map obtained from the image through the optical flow estimation module will contain regions with incorrect estimations, which are characterized by messy light-colored blocks and seriously affect the identification of dynamic object regions in the optical flow map. When only a part of the object is moving, in the optical flow map obtained through optical flow estimation, the transition between the moving region and the static region is not obvious. However, in the present invention, the optical flow map is further processed to enhance its contrast, and a bilateral filter is used to retain edge information and reduce noise. Finally, the saliency region detection module can effectively avoid the false detection or inaccurate detection problems caused by traditional edge detection algorithms, and segment the salient regions in the optical flow map, that is, obtain the dynamic object regions.

[0080] S3. According to the obtained representative points, prompt the semantic segmentation module to segment the next frame of image, realize the removal of any dynamic object, and obtain the mask segmented semantically;

[0081] Since the optical flow estimation module and the saliency region detection module are time-consuming and to a certain extent affect the real-time performance of the SLAM system, a semantic segmentation module with better real-time performance is adopted to segment the dynamic region through the periodic prompt points of the above two modules. The semantic segmentation module designs three prompt methods for segmentation. The first is the box prompt: the object to be segmented is framed by a box. In this prompt method, the limitation of the semantic segmentation accuracy is very large. Its accuracy depends heavily on the quality of the box, that is, only when the box just encloses the region where the object is located can it be correctly and completely segmented. If the box is too large, it will result in the segmentation of redundant objects, and if the box is too small, it will result in the inability to completely segment the object. The second is the point prompt: according to the positions of one or more points, the entire region of the selected object is automatically segmented. The third is the mixed prompt of points and boxes: on the premise of the box prompt, further segmentation is performed using the point prompt. This method retains the limitations of the box prompt and also loses real-time performance. In summary, the present invention adopts the point prompt method to perform semantic segmentation on dynamic objects.

[0082] More specifically, in this embodiment, the semantic segmentation module used is a variant of EfficientViT. The image encoder of the Segment Anything Model (SAM) is replaced by EfficientViT, while retaining the structure of the lightweight prompt encoder and the mask decoder, which is denoted as EfficientViT-SAM. The specific process includes:

[0083] S31, the input image passes through three convolution blocks and two EfficientViT modules in sequence to obtain the features of five stages, and the fused features of the last three stages are obtained by upsampling and addition, and are fed to the neck including several MBConv blocks, and then fed to the SAM head;

[0084] S32, perform training; the training includes two stages: first, use SAM's image encoder to train EfficientViT-SAM's image encoder, and use L2 loss as the loss function; then train the EfficientViT-SAM end-to-end network on the entire SA-1B dataset;

[0085] S33. Extract representative points of the mask: fit the center line of the mask through a polynomial, and adjust the degree of the polynomial to adjust the fitting degree of the curve; take two points at one-sixth and five-sixths on the center line respectively to prompt the semantic segmentation model to segment the next image, that is, use the semantic segmentation result of the previous frame of the image to prompt the semantic segmentation model to segment the same moving object area in the next frame of the image, so as to realize the semantic association between adjacent images.

[0086] S4, further extracting two representative points from the mask obtained by semantic segmentation in step S3, prompting the semantic segmentation module to segment the dynamic object area of ​​the next frame image, until eight frames later, the representative points obtained by the salient area detection module are reused to prompt the semantic segmentation module to perform segmentation;

[0087] S5. Use a lightweight CNN architecture to extract and match feature points from images with dynamic objects removed;

[0088] Complex environments can have a serious impact on traditional ORB feature point extraction. For example, large camera angle fluctuations can aggravate the dynamic blur of the image, and obvious lighting changes can reduce the correlation between two frames. In order to improve the number and quality of feature points extracted in complex environments, this embodiment uses a deep learning-based feature point extraction and matching module to extract feature points from images after dynamic object removal. This module uses the Xfeat model as its core to improve the speed of feature point extraction and matching. Figure 3 As shown in the figure, this module reduces the interference of light changes and perspective fluctuations in complex environments on feature point extraction, extracts high-quality and effective feature points, and improves the positioning accuracy of the unmanned vehicle in the SLAM process.

[0089] In this embodiment, the preferred implementation of step S5 is:

[0090] S51, separate key point detection into an independent branch, using 1×1 convolution to perform fast processing on an 8×8 tensor block transformed image;

[0091] S52. Design a matching refinement module. The learning of this module predicts pixel-level offsets by only considering the nearest neighbor pairs at 1 / 8 of the original spatial resolution in the original rough-level features.

[0092] S53. Divide the image into several appropriate equal parts according to the density of feature points in the previous frame image. The greater the density, the more parts are divided, with a maximum of 8 equal parts.

[0093] S54. Select feature points with better quality in each image block and regard them as a cluster. Calculate the average quality of all feature points in the image, and filter out the feature points with poor quality according to the difference between the quality of each cluster and the average quality according to N = km; where m is the quality difference, k is a proportionality constant obtained through experiments, and N is the number of feature points to be filtered out.

[0094] S55. Use the calculated descriptors used to describe the pixel information around the feature points to train a bag-of-words model containing environmental semantic information, providing richer environmental semantic vocabulary for loop closure detection.

[0095] In the method proposed by the present invention, the above steps S1 - S5 are all processed only on RGB images. Because the depth map is used to provide depth values for the feature points extracted from the processed RGB images, the unprocessed areas in the depth map will not participate in the pose estimation process, so there is no need to process it simultaneously. In addition, the depth map can also be obtained through an RGB - D depth camera.

[0096] S6. Send the successfully matched feature points, the RGB image after removing dynamic objects, and the original depth map into the ORB - SLAM2 system, and use the BA algorithm to minimize the reprojection error to track and locate the camera pose of each frame.

[0097] S7. When a new key frame is inserted into the local map, update the key frames in the local map and the connection relationship between key frames, simultaneously perform BA optimization on the poses of key frames and map points in the current local map, and finally delete redundant key frames and their corresponding map points.

[0098] S8. Use the feature point extraction module to train a bag-of-words model and perform loop closure detection to correct the estimated pose and local map.

[0099] S9. Use graph optimization to optimize all key frames and map points.

[0100] In summary, a method for optical flow semantic SLAM for arbitrary dynamic objects proposed by the present invention is a more efficient SLAM method, which can enable the unmanned vehicle to have higher positioning and mapping accuracy, and at the same time, it has high robustness in practical engineering.

[0101] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0102] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An optical flow semantic SLAM method for arbitrary dynamic objects, characterized in that: The specific steps include: S1. Obtain continuous RGB images, and use the deep learning-based optical flow estimation module to perform optical flow estimation on the current and next frames of images every eight frames to generate an optical flow map; S2. Use the salient region detection module to extract the dynamic region in the optical flow map as a mask, and find two representative points in the dynamic mask; S3, prompting the semantic segmentation module to segment the next frame of the image according to the obtained representative points, thereby eliminating any dynamic objects and obtaining a semantically segmented mask; S4, further extracting two representative points from the mask obtained by semantic segmentation in step S3, prompting the semantic segmentation module to segment the dynamic object area of ​​the next frame image, until eight frames later, the representative points obtained by the salient area detection module are reused to prompt the semantic segmentation module to perform segmentation; S5, using a lightweight CNN architecture to extract and match feature points from the image with dynamic objects removed; step S5 specifically includes: S51, separate key point detection into an independent branch, using 1×1 convolution to perform fast processing on an 8×8 tensor block transformed image; S52, designing a matching refinement module, wherein the learning of the module predicts pixel-level offsets by considering only the nearest neighbor pairs at 1 / 8 of the original spatial resolution in the original coarse level features; S53, dividing the image into a number of appropriate equal parts according to the density of feature points of the previous frame image, the greater the density, the more parts are divided, and the maximum is 8 equal parts; S54, selecting feature points with better quality in each image block and treating them as a cluster; calculating the average quality of all feature points in the image, and filtering out feature points with poorer quality according to N=km based on the difference between the quality of each cluster and the average quality; where m is the quality difference, k is a proportional constant obtained through experiments, and N is the number of feature points to be filtered out; S55, using the calculated descriptor used to describe the pixel information around the feature point, training a bag-of-words model containing environmental semantic information to provide a richer environmental semantic vocabulary for loop detection; S6, send the successfully matched feature points, the RGB image after removing dynamic objects, and the original depth map to the ORB-SLAM2 system, and use the BA algorithm to minimize the reprojection error, track and locate the camera pose of each frame; S7, when a new key frame is inserted into the local map, the key frames in the local map and the connection relationship between the key frames are updated, and the poses and map points of the key frames in the current local map are optimized by BA, and finally the redundant key frames and their corresponding map points are deleted; S8, using the feature point extraction module to train the bag-of-words model, and perform loop detection, calculate the best candidate frame at the loop, and replace the current frame to correct the posture and local map; S9. Use graph optimization to optimize all keyframes and map points.

2. An optical flow semantic SLAM method for arbitrary dynamic objects according to claim 1, characterized in that: The optical flow estimation module used in step S1 is a new variant of RAFT.

3. An optical flow semantic SLAM method for arbitrary dynamic objects according to claim 2, characterized in that: The specific process of optical flow estimation in step S1 includes: S11, feature extraction: down-sample the current and next frame images input through the feature encoding network to extract the feature map of 1 / 8 resolution; S12, visual similarity calculation: using the two feature maps obtained in step S11 to calculate the correlation volume, and then using the multi-level correlation of the correlation pyramid to capture large and small pixel displacements, and then generating a new feature map by indexing the features in each level of the correlation pyramid; Given the current optical flow estimate , for the current original RGB image , the next frame of original RGB image ,Bundle Each pixel x =( u , v ) are mapped to In the next frame, we get an estimate of the pixels of the next frame. ; The mapping relationship is: ;exist Define a local search neighborhood around , the formula is: ; in, is a set of two-dimensional integer grid points, r is the search radius on the two-dimensional image plane; Based on local search area For each layer of the correlation pyramid, a feature map is found and then these feature maps are concatenated; S13, iterative update; the feature maps obtained from each layer of the correlation pyramid in step S12 are concatenated into a total feature map and input into the update operator with the transformer module as the core to estimate the optical flow estimation sequence of all pixels ; From an initial starting point Start, the initial starting point Direct prediction by reusing the existing context encoder and feeding it stacked input frames; at each iteration, a flow update direction is produced , and add it to the current estimate, that is, the formula is expressed as: ; The loss function in the iterative process is the mixed Laplace loss, and the network is trained to predict the mixing parameters of the Laplace distribution to maximize the log-likelihood of the ground truth flow; given an image pair and optical flow truth , the training loss is: ; in, Represents the probability density function of the Laplace distribution.

4. The optical flow semantic SLAM method for arbitrary dynamic objects according to claim 1, characterized in that: In step S2, the salient region detection module explicitly decomposes the binary image segmentation task on high-resolution data into a positioning module and a reconstruction module; and adopts a bilateral filter to retain edge information and reduce noise.

5. The optical flow semantic SLAM method for arbitrary dynamic objects according to claim 4, characterized in that: Step S2 specifically includes: S21, image preprocessing: enhancing the contrast of the optical flow map and using a bilateral filter to preserve edge information and reduce noise; S22, build a positioning module; use the Transformer encoder to extract features at different stages; the features extracted at different stages are recorded as: ; Characteristics of the first three stages are passed to the corresponding decoder stages with lateral connections and are stacked and concatenated at the last encoder block to generate fused features ; Then fusion features It is fed into the classification module, and the hollow spatial convolutional pooling pyramid ASPP module is used to perform multi-context fusion. Compression to obtain compression features , in order to transfer to the reconstruction module; S23, construct reconstruction module; compress features is the input; for the i-th Transformer encoder extracted , deformable convolution with hierarchical receptive fields and adaptive average pooling layers are used to extract features with receptive fields of various scales; these features extracted by different receptive fields are then concatenated as , and then through the 1×1 convolution layer and batch normalization layer, the output features after reconstruction module processing are obtained , the output characteristics Then pass a 1×1 convolution layer to obtain the final prediction image; S24, performing bilateral reference; divided into internal reference and external reference; the internal reference is cropped adaptively to avoid adjusting the image size; the external reference uses gradient labels to obtain areas with richer gradient information; S25. Set the objective function; use BCE, IoU and SSIM loss functions to supervise at the pixel, region and boundary levels respectively. The formula is expressed as: ; in, , and They are the losses at pixel, region and boundary levels respectively; Cross entropy loss for binary classification problems; The overlap loss between the predicted segmentation result and the true segmentation label; is the structural similarity loss; , and are the loss weights respectively.

6. The optical flow semantic SLAM method for arbitrary dynamic objects according to claim 1, characterized in that: In step S3, the semantic segmentation module is a variant of EfficientViT, which uses EfficientViT to replace the image encoder of the segmentation everything model SAM, while retaining the structure of the lightweight hint encoder and mask decoder, which is denoted as EfficientViT-SAM.

7. The optical flow semantic SLAM method for arbitrary dynamic objects according to claim 6, characterized in that: Step S3 specifically includes: S31, the input image passes through three convolution blocks and two EfficientViT modules in sequence to obtain the features of five stages, and the fused features of the last three stages are obtained by upsampling and addition, and are fed to the neck including several MBConv blocks, and then fed to the SAM head; S32, perform training; the training includes two stages: first, use SAM's image encoder to train EfficientViT-SAM's image encoder, and use L2 loss as the loss function; then train the EfficientViT-SAM end-to-end network on the entire SA-1B dataset; S33. Extract representative points of the mask: fit the center line of the mask through a polynomial, and adjust the degree of the polynomial to adjust the fitting degree of the curve; take two points at one-sixth and five-sixths on the center line respectively to prompt the semantic segmentation model to segment the next image, that is, use the semantic segmentation result of the previous frame of the image to prompt the semantic segmentation model to segment the same moving object area in the next frame of the image, so as to realize the semantic association between adjacent images.

Citation Information

Patent Citations

  • Visual SLAM method based on semantic segmentation of deep learning

    CN112132897A

  • Semantic vision SLAM method and system based on semantic segmentation and optical flow

    CN117710806A