VSLAM method combining Superpoint and target detection network

The integration of Superpoint and target detection networks in VSLAM algorithms addresses dynamic environment challenges by enhancing feature extraction and filtering, resulting in improved mapping accuracy and efficiency.

CN120318486AInactive Publication Date: 2025-07-15HUBEI UNIV OF ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264238.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional visual SLAM algorithms are difficult to process dynamic objects in dynamic environments, resulting in inaccurate matching of feature points, unstable image drift and map construction, and low accuracy.

Method used

Combining Superpoint and the improved YOLOv5 detection network, dynamic feature points are eliminated through feature point extraction, dynamic object detection and optical flow method to optimize the position pose map.

Benefits of technology

It improves the mapping accuracy and processing efficiency of visual SLAM, reduces the impact of dynamic objects on mapping, and enhances environmental perception capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318486A_ABST
    Figure CN120318486A_ABST
Patent Text Reader

Abstract

The invention relates to a VSLAM method combining Superpoint and a target detection network. The method comprises the following steps: obtaining an original image to form a data set; inputting the data set into a Superpoint network to extract feature points; an improved YOLOv5 detection model is constructed; inputting the extracted feature points into an improved YOLOv5 detection model, detecting a dynamic object in the detection image, and generating a detection frame and a detection category; removing dynamic feature points on the object in the detection frame by using an optical flow method to obtain static feature points in the image; and inputting the static feature points into the local map to optimize the pose, and obtaining an optimized pose map. According to the VSLAM method combining the Superpoint and the target detection network, the target detection algorithm is integrated, so that the influence of a dynamic object on the mapping precision is reduced, the influence of the dynamic object on the mapping precision is reduced, and the processing efficiency and precision of visual SLAM are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot mapping and positioning, and specifically relates to a VSLAM method combining Superpoint and object detection network. Background Art

[0002] Visual SLAM algorithms are widely used in robot mapping and positioning algorithms, and object detection is an image processing method with fast processing speed and high accuracy at present, which can effectively identify the types of objects in the image. By identifying dynamic objects and removing their feature points, the positioning and mapping accuracy can be enhanced, and preprocessing related to object detection is carried out in the front-end optimization of visual SLAM.

[0003] Traditional visual SLAM is often used to run in ideal environments such as static and rigid bodies, but it is prone to failure in the running scenarios of mobile robots, there are many environmental impacts and human factors interference, and it is difficult to handle these dynamic changes, resulting in inaccurate matching of feature points, image drift, unstable mapping and low accuracy. Summary of the Invention

[0004] The technical problem of the present invention is: to identify dynamic feature points on the image through a neural network, reduce the influence of dynamic objects in the environment, enhance the perception of the external environment, and improve the mapping and positioning performance of the visual SLAM algorithm.

[0005] The purpose of the present invention is to solve the above problems, and provide a VSLAM method combining Superpoint and object detection network, including the following steps: S1: Obtain the original images to form a data set; S2: Input the data set into the Superpoint network to extract feature points; S3: Build an improved YOLOv5 detection model; S4: Input the feature points extracted in step S2 into the improved YOLOv5 detection model to detect dynamic objects in the detected image, and generate detection frames and detection categories; S5: Use the optical flow method to remove the dynamic feature points on the objects within the detection frame, and obtain the static feature points in the image; S6: Input the static feature points into the local map to optimize the pose, and obtain the optimized pose map.

[0006] Further, in step S1, it includes using cameras, cameras and drones to obtain the original images.

[0007] Further, in step S2, the Superpoint network includes an encoder and a decoder; the encoder reduces the size of the original image through convolutional layers, max-pooling layers, and non-linear activation layers; the decoder inputs the feature descriptors and feature points into the SLAM algorithm.

[0008] Preferably, the decoder contains two branches. The first branch performs weighted summation and image size reduction through a convolutional module, a softmax layer, and a Reshape layer respectively, and inputs a feature point map of the same size; the second branch obtains a feature vector through bi-cubic interpolation according to the feature point positions, normalizes the channels, and outputs feature descriptors.

[0009] Preferably, in step S3, the improved YOLOv5 detection model includes a backbone network and a detection head; the backbone network includes a convolutional module Conv, a lightweight attention module GhostBottleneck, a coordinate attention mechanism module CA, and a spatial pyramid pooling SPPF connected in series in sequence; the detection head includes a convolutional module Conv, an upsampling module Upsample, a tensor concatenation Concat, and a convolutional module C3.

[0010] Further, the lightweight attention module GhostBottleneck replaces the bottleneck module in the original YOLOv5 network, generates redundant feature maps through a combined linear transformation, improves the feature expression ability and feature extraction, and the calculation formula is: ; In the formula, represents the input of the original image, represents the convolution kernel, represents the offset.

[0011] Preferably, the coordinate attention mechanism module CA divides the input feature map into two directions of horizontal coordinates and vertical coordinates respectively for global average pooling to obtain the feature map in the height direction, and the calculation formula is: ; ; In the formula, represents the height of the input data, represents the width of the input data, represents the input of channel c, represents the output of channel c.

[0012] Preferably, in step S5, it includes the following sub-steps: 1) Utilize the brightness constancy of the LK optical flow method, the characteristic that the distance of each movement is very small during the continuous movement of pixels, and the spatial consistency to solve the weighted least squares method within a small window; 2) Obtain the corresponding velocity components by solving the optical flow equation, and based on the motion change information of the object in the image; 3) Track the motion information of the dynamic feature points of the object within the detection frame and finally eliminate the dynamic feature points on the object in the image.

[0013] Furthermore, step S6 includes inputting the static feature points within the detection frame and the static feature points outside the detection frame into the local map to optimize the pose, and obtaining the optimized pose map.

[0014] Preferably, step S6 includes the following sub-steps: 1) Tracking thread: Calculate the bag-of-words vector of the feature points in the current frame, and use the bag-of-words to find the relocalization candidate key frames similar to the current frame. Then, after traversing all the candidate key frames, perform matching through the bag-of-words and add key frames as needed for the local mapping thread; 2) Local mapping thread: Generate new map points using the consensus relationship of the key frames. After searching and fusing the map poses and their locations of adjacent key frames, generate map points for local map optimization, and then optimize the bitmap points of the co-visible key frames; 3) Loop closing thread: Propagate the pose through the transformation of the current key frame, correct the poses of the key frames connected to the current frame and their map points, project all the map points in the closed-loop connected key frame group into the current key frame group for matching and fusion, and optimize the poses and map points of the key frames in the map to eliminate the cumulative trajectory error and map error.

[0015] Compared with the prior art, the beneficial effects of the present invention include: 1) The present invention provides a VSLAM method combining Superpoint and object detection network, integrating the object detection algorithm to reduce the influence of dynamic objects on the mapping accuracy and improve the processing efficiency and accuracy of visual SLAM.

[0016] 2) The present invention provides a VSLAM method combining Superpoint and object detection network. By improving the YOLOv7 network to add a mask to extract the feature points of the dynamic region, the tracking completion time of the detected object is reduced, and the running speed of the visual SLAM algorithm is improved.

[0017] 3) The present invention provides a VSLAM method combining Superpoint and object detection network. The lightweight attention mechanism backbone network is used to capture and aggregate the features of the detected object, and the computational amount of the model is reduced according to the pruning and distillation methods, enhancing the modeling ability of the network. Description of the Drawings

[0018] The present invention will be further described below with reference to the drawings and embodiments.

[0019] Figure 1 This is the overall block diagram of VSLAM combining Superpoint and the object detection network according to the embodiments of the present invention.

[0020] Figure 2 This is the structural diagram of the Superpoint feature extraction network according to the embodiments of the present invention.

[0021] Figure 3 This is the structural diagram of the improved YOLOv5s object detection network according to the embodiments of the present invention.

[0022] Figure 4 This is the comparison chart of the absolute trajectory errors between the embodiments of the present invention and ORB-SLAM2 in the fr3 / walking / halfsphere sequence.

[0023] Figure 5 This is the comparison chart of the absolute trajectory errors between the embodiments of the present invention and ORB-SLAM2 in the fr3 / walking / rpy sequence.

[0024] Figure 6 This is the comparison chart of the absolute trajectory errors between the embodiments of the present invention and ORB-SLAM2 in the fr3 / walking / static sequence.

[0025] Figure 7 This is the comparison chart of the absolute trajectory errors between the embodiments of the present invention and ORB-SLAM2 in the fr3 / walking / xyz sequence.

[0026] Figure 8 This is the comparison chart of the relative trajectory errors between the embodiments of the present invention and ORB-SLAM2 in the fr3 / walking / halfsphere sequence.

[0027] Figure 9 This is the comparison chart of the relative trajectory errors between the embodiments of the present invention and ORB-SLAM2 in the fr3 / walking / rpy sequence.

[0028] Figure 10 This is the comparison chart of the relative trajectory errors between the embodiments of the present invention and ORB-SLAM2 in the fr3 / walking / static sequence.

[0029] Figure 11 This is the comparison chart of the relative trajectory errors between the embodiments of the present invention and ORB-SLAM2 in the fr3 / walking / / xyz sequence. Specific embodiments

[0030] As Figure 1 shown, a VSLAM method combining Superpoint and the object detection network includes the following steps: S1: Obtain the original images to form a dataset; In step S1, it includes using cameras, webcams, and drones to obtain the original images.

[0031] Select the VOC2007 dataset as the training set and test set for the object detection algorithm, identify according to the common dynamic object categories in the dataset, and then use the labelme software to label the dynamic object categories in the dataset to obtain the training dataset labels, and extract the image feature information.

[0032] S2: Input the dataset into the Superpoint network to extract feature points.

[0033] As Figure 2 shown, in step S2, the Superpoint network includes an encoder and a decoder; the encoder reduces the size of the original image through convolutional layers, max pooling layers, and non-linear activation layers; the decoder inputs the feature descriptors and feature points into the SLAM algorithm.

[0034] The Superpoint network includes an encoder and a decoder; the encoder reduces the size of the original image through convolutional layers, max pooling layers, and non-linear activation layers; the decoder inputs the feature descriptors and feature points into the SLAM algorithm.

[0035] The decoder contains two branches. The first branch performs weighted sum and image size reduction respectively through a convolutional module, a softmax layer, and a Reshape layer, and inputs a feature point map of the same size; the second branch obtains a feature vector through Bi-Cubic Interpolate bilinear interpolation according to the feature point positions, and normalizes the channels to output the feature descriptors.

[0036] H, W, and D respectively represent the height, width, and dimension of the image. The input is an image with a depth of 1. The encoder composed of convolutional layers, max pooling layers, and non-linear activation layers reduces the image size. The feature point positions obtain feature vectors through Bi-Cubic Interpolate bilinear interpolation operations, and then normalize according to the channels to output the feature descriptors, and then input the feature descriptors and feature points into the SLAM algorithm.

[0037] S3: Construct an improved YOLOv5 detection model.

[0038] As Figure 3As shown, in step S3, the improved YOLOv5 detection model includes a backbone network and a detection head; the backbone network includes a convolution module Conv, a lightweight attention module GhostBottleneck, a coordinate attention mechanism module CA and a spatial pyramid pooling SPPF connected in series in sequence; the detection head includes a convolution module Conv, an upsampling module Upsample, a tensor splicing Concat and a convolution module C3.

[0039] The lightweight attention module GhostBottleneck replaces the bottleneck module in the original YOLOv5 network and generates redundant feature maps by combining linear transformations to improve feature expression and feature extraction. The calculation formula is: ; In the formula, represents the input of the original graph, ,in is the number of input channels, is the input data height, is the input data width; is the convolution operation, represents the convolution kernel, Indicates the offset, represents the generated feature map, Indicates width, Indicates height, Indicates the number of channels.

[0040] Preferably, the coordinate attention mechanism module CA divides the input feature map into two directions, horizontal coordinate and vertical coordinate, and performs global average pooling respectively to obtain the feature map in the height direction, which is expressed as: ; ; In the formula, Indicates the input data height, Indicates the input data width, represents channel c input, Represents the output of channel c.

[0041] S4: Input the feature points extracted in step S2 into the improved YOLOv5 detection model, detect dynamic objects in the detection image, and generate detection boxes and detection categories.

[0042] S5: Use the optical flow method to remove dynamic feature points on the object in the detection frame and obtain static feature points in the image.

[0043] Step S5 includes the following sub-steps: 1) Utilize the brightness constancy of the LK optical flow method, the characteristic that the movement distance of each pixel is very small during continuous movement, and the spatial consistency to solve the weighted least squares method within a small window; 2) Obtain the corresponding velocity components by solving the optical flow equation, and based on the motion change information of the objects in the image; 3) Track the motion information of the dynamic feature points of the objects within the detection frame and finally eliminate the dynamic feature points on the objects in the image.

[0044] S6: Input the static feature points into the local map to optimize the pose and obtain the optimized pose map.

[0045] Step S6 includes inputting the static feature points inside and outside the detection frame into the local map to optimize the pose and obtaining the optimized pose map.

[0046] In step S6, the following sub-steps are included: 1) Tracking thread: Calculate the bag-of-words vector of the feature points of the current frame, use the bag-of-words to find the relocalization candidate key frames similar to the current frame, and then traverse all the candidate key frames, match them through the bag-of-words, and add key frames as needed for the local mapping thread; 2) Local mapping thread: Generate new map points using the consensus relationship of the key frames, search and fuse the map poses and their locations of adjacent key frames, generate map points for local map optimization, and then optimize the bitmap points of the co-visible key frames; 3) Loop closure thread: Propagate the pose through the transformation of the current key frame, correct the poses of the key frames connected to the current frame and their map points, project all the map points in the closed-loop connected key frame group into the current key frame group for matching and fusion, and optimize the poses and map points in the key frames in the map to eliminate the cumulative trajectory error and map error.

[0047] Deploy the visual SLAM model on the Jetson TX2 development board of the mobile robot, turn on sensors such as the robot chassis and camera, run the improved YOLOv5 model and the SLAM model, and control the robot to perceive the information of the unknown environment, build a map and output the results.

[0048] To evaluate the effect of the present invention in object detection, the comparison effects of ORB-SLAM2 and the present invention on the fr3 / walking / halfsphere, fr3 / walking / rpy, fr3 / walking / static, and fr3 / walking / xyz sequences in the TUM dataset are used, as Figures 4 - 11 shown, where the left side is the experimental effect diagram of ORB-SLAM2, and the right side is the experimental effect diagram of the present invention.

[0049] In the present invention, in four sequences, through the comparison of the values of the root mean square error (RMSE) and the standard deviation (S.D.), the comparison results are shown in Table 1 and Table 2. Among them, Table 1 is the comparison table of the absolute trajectory error between the present invention and ORB-SLAM2, and Table 2 is the comparison table of the relative translational trajectory error between the present invention and ORB-SLAM2.

[0050] Table 1

[0051] Table 2

[0052] The experimental results show that in the fr3 / walking / halfsphere, fr3 / walking / rpy, fr3 / walking / static, and fr3 / walking / xyz sequences of the TUM dataset, the RMSE and S.D. generated by the evo trajectory evaluation tool of the present invention are lower than those of ORB-SLAM2, indicating that in the mapping process, the present invention can effectively reduce the influence of light changes in the environment and reduce the anomalies generated in the mapping process. Therefore, compared with ORB-SLAM2, although the present invention spends more time on preprocessing, it has a greater improvement in accuracy and mapping accuracy.

[0053] When the present invention processes the current frame in the SLAM system, the detection module performs the recognition of the next frame of image, which speeds up the running rate of the system and reduces the time required for system preprocessing. When running on the Jetson TX2 device, it can process 18 frames of images per second, and has better real-time performance compared with other SLAM systems.

[0054] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A VSLAM method combining Superpoint and object detection network, characterized in that It includes the following steps: S1: Obtain the original images to form a dataset; S2: Input the dataset into the Superpoint network to extract feature points; S3: Construct an improved YOLOv5 detection model; S4: Input the feature points extracted in step S2 into the improved YOLOv5 detection model to detect dynamic objects in the detected images, and generate detection boxes and detection categories; S5: Use the optical flow method to eliminate the dynamic feature points on the objects within the detection boxes to obtain the static feature points in the images; S6: Input the static feature points into the local map to optimize the pose and obtain the optimized pose map.

2. The VSLAM method combining Superpoint and object detection network according to claim 1, characterized in that In step S1, it includes using cameras, video cameras, and drones to obtain the original images.

3. The VSLAM method combining Superpoint and a target detection network according to claim 1, wherein In step S2, the Superpoint network includes an encoder and a decoder; the encoder reduces the size of the original image through convolutional layers, max-pooling layers, and non-linear activation layers; the decoder inputs the feature descriptors and feature points into the SLAM algorithm.

4. The VSLAM method combining Superpoint and object detection network according to claim 3, wherein The decoder contains two branches. The first branch performs weighting and image size reduction respectively through a convolutional module, a softmax layer, and a Reshape layer, and inputs a feature point map of the same size; the second branch obtains the feature vectors through bi-cubic interpolation according to the feature point positions, and normalizes the channels to output the feature descriptors.

5. The VSLAM method combining Superpoint and object detection network according to claim 1, characterized in that, In step S3, the improved YOLOv5 detection model includes a backbone network and a detection head; the backbone network includes a convolutional module Conv, a lightweight attention module GhostBottleneck, a coordinate attention mechanism module CA, and a spatial pyramid pooling SPPF connected in series in sequence; the detection head includes a convolutional module Conv, an upsampling module Upsample, a tensor concatenation Concat, and a convolutional module C3.

6. A VSLAM method combining Superpoint and object detection network according to claim 5, characterized in that, The lightweight attention module GhostBottleneck replaces the bottleneck module in the original YOLOv5 network, generates redundant feature maps through a combined linear transformation, improves the feature expression ability and feature extraction, and its calculation formula is: ; In the formula, represents the input of the original image, , where is the number of input channels, is the height of the input data, is the width of the input data; is the convolution operation, represents the convolution kernel, represents the offset, represents the generated feature map, represents the width, represents the height, represents the number of channels.

7. The VSLAM method combining Superpoint and object detection network according to claim 5, characterized in that, The coordinate attention mechanism module CA divides the input feature map into two directions of horizontal coordinates and vertical coordinates respectively for global average pooling to obtain the feature map in the height direction, and the expression is: ; ; In the formula, represents the height of the input data, represents the width of the input data, represents the input of channel c, represents the output of channel c.

8. The VSLAM method combining Superpoint and object detection network according to claim 1, characterized in that, In step S5, it includes the following sub-steps: 1) Use the brightness constancy of the LK optical flow method, the characteristic that the movement distance of each movement is very small during the continuous movement of pixels, and the spatial consistency to solve the weighted least squares method within a small window; 2) Obtain the corresponding velocity components by solving the optical flow equation and according to the movement change information of the objects in the images; 3) Track the movement information of the dynamic feature points of the objects within the detection boxes and finally eliminate the dynamic feature points on the objects in the images.

9. The VSLAM method combining Superpoint and object detection network according to claim 1, wherein Step S6 includes inputting the static feature points within the detection boxes and the static feature points outside the detection boxes into the local map to optimize the pose and obtain the optimized pose map.

10. The VSLAM method combining Superpoint and object detection network according to claim 9, characterized in that, In step S6, it includes the following sub-steps: 1) Tracking thread: Calculate the bag-of-words vector of the feature points in the current frame, and use the bag-of-words to find relocalization candidate key frames similar to the current frame. Then, after traversing all the candidate key frames, perform matching through the bag-of-words and add key frames as needed to the local mapping thread; 2) Local mapping thread: Generate new map points using the consensus relationship of key frames. After searching and fusing the map poses and their locations of adjacent key frames, generate map points for local map optimization, and then optimize the bitmap points of the co-visible key frames; 3) Loop closing thread: Propagate the pose through the transformation of the current key frame, correct the poses of the key frames connected to the current frame and their map points, project all the map points in the loop-connected key frame group into the current key frame group for matching and fusion, and optimize the poses and map points in the key frames in the graph to eliminate the cumulative trajectory error and map error.