VSLAM method for improving target detection network
By improving the object detection network, combining gradient path planning and optical flow method to eliminate dynamic feature points, the problem of low accuracy of visual SLAM algorithm in dynamic environments is solved, and more efficient mapping and positioning performance is achieved.
Patent Information
- Application Number
- CN202411836887.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-06
AI Technical Summary
The existing visual SLAM algorithms have problems with image drift and low accuracy in dynamic environments, making it difficult to effectively deal with the impact of dynamic objects on map construction and positioning.
The improved object detection network is adopted to extract image features through the RepNCSPELAN4 module, combine gradient path planning and optical flow method to remove feature points of dynamic objects in the image, and use the C2f module to enhance the nonlinear ability and representation ability of the model, and finally the improved YOLOv5 model is distilled to reduce the number of model calculation parameters.
It improves the mapping and positioning accuracy of the SLAM algorithm in a dynamic environment, reduces the system's memory usage, improves the operation efficiency, and shows good results on the VOC2007 dataset.
Smart Images

Figure CN119941844A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of robot mapping and positioning, and in particular relates to a VSLAM method for improving a target detection network. Background Art
[0002] With the development of robot mapping and positioning algorithms, SLAM algorithms are widely used. According to different sensors, they are often divided into LiDAR-based SLAM and camera-based visual SLAM. In visual SLAM, visual images are an important way for robots to obtain map information, and target detection is a high-speed and high-precision image processing method that can effectively identify the types of objects in the image. By identifying dynamic objects and removing their feature points, the positioning and mapping accuracy can be enhanced. Therefore, target detection-related preprocessing is often performed in the optimization of the visual SLAM front end.
[0003] Visual SLAM algorithm is a research hotspot in the field of robot mapping and positioning technology. Object detection is widely used in the field of image processing, and its main goal is to identify the categories of different objects in the image. By integrating the object detection algorithm into visual SLAM, the impact of dynamic objects on the accuracy of mapping can be reduced. For example, the lightweight YOLOv8 algorithm is combined with the sparse optical flow method to form a target detection thread and the ORB-SLAM3 algorithm to reduce the number of model parameters and improve the accuracy of the SLAM algorithm in dynamic scenes. (JIANG XK, YANG G, DU Y Y. Dynamic visual SLAM algorithm based on lightweightYOLOv8n[J]. Journal of Xi'an University of Posts and Telecommunications, 2024, 29(3): 75-82.) There is also the use of YOLOv7-tiny network combined with optical flow to remove dynamic feature points, and the contrast information is used to adjust the threshold of feature point extraction to improve the stability of the SLAM system, and then the running speed of the SLAM system is optimized by simplifying the local thread. (Qi Hao,Fu Yuexin,Hu Zhuhua,Wu Jiaqi,Zhao Yaochi.Lightweightsemantic VSLAM method based on adaptive thresholding[J / OL]and speedoptimization.Journal of Beijing University of Aeronautics and Astronautics.)
[0004] In addition to the target detection algorithm mentioned above, semantic segmentation or instance segmentation algorithms are often used in combination with visual SLAM algorithms to improve performance. For example, the Cross-Segnet network is used to remove feature points, and a spatiotemporal consistent mask algorithm is proposed to compare with the mask generated by the Cross-Segnet algorithm to enhance the performance of semantic segmentation and improve the processing efficiency and accuracy of the visual SLAM algorithm. (Z.Guo, N.Dong, Z.Zhang, X.Mai and D.Li, "CS-SLAM: A lightweight semantic SLAM method for dynamic scenarios," in IEEE Transactions on Cognitive and Developmental Systems, doi: 10.1109 / TCDS.2024.3462651.) By adding a mask to the YOLOv7 network to extract feature points in the dynamic area, a separate mask thread is created to improve the extraction rate of the semantic information of the visual SLAM, reduce the time the system waits for tracking to complete, and improve the running rate of the visual SLAM algorithm. (Y. Zhang and X. Wu, "SD-SLAM: Visual SLAM Combining Semantic Segmentation and Dynamic Feature Point Detection," 2023 China Automation Congress (CAC), Chongqing, China, 2023, pp. 2028-2033, doi: 10.1109 / CAC59555.2023.10451659.) Traditional visual SLAM is often used in ideal environments such as static and rigid bodies. However, the operation scenes of mobile robots will have more environmental influences and interference from human factors, which will cause SLAM algorithms to have problems such as image drift and low accuracy. By adding a neural network to the front end to identify dynamic feature points on the image, the impact of dynamic objects in the environment on positioning and mapping can be reduced, and the perception of the external environment can be enhanced. Summary of the invention
[0005] The purpose of the present invention is to propose a VSLAM method for improving the target detection network based on the disadvantage that the prior art has high requirements for edge devices. The method uses the RepNCSPELAN4 module instead of the C3 module to extract image features, uses gradient path planning to increase the width of the network to retain more image features, uses the C2f module in the Head part to further extract features, increases the nonlinear ability and representation ability of the model, and finally performs knowledge distillation on the improved YOLOv5 model to reduce the number of parameters calculated by the model and reduce the memory occupied by the system. The target detection module can effectively remove the feature points on the dynamic objects in the image, and improve the mapping and positioning performance of the SLAM algorithm.
[0006] The technical solution adopted by the present invention is: a VSLAM method for improving the target detection network, characterized in that: the method comprises the following steps:
[0007] Step 1: Obtain the original image through the camera, input the image into the gradient information target detection network for preprocessing, detect dynamic objects in the detection image, and generate detection boxes and detection categories;
[0008] Step 2, using the optical flow method to remove the dynamic feature points on the object in the detection frame in step 1, and obtain the static feature points in the image;
[0009] Step 3: Input the static feature points inside the detection frame and the static feature points outside the detection frame obtained in step 2 into the local map to optimize the pose and obtain the optimized pose graph.
[0010] Furthermore, the gradient information target detection network described in the step 1 uses YOLOv5s as the baseline model. After calculating the loss function in the deepening neural network to generate a new gradient, some information will be lost. The RepNCSPELAN4 module combines the CSP module and the ELAN module designed by gradient planning to increase the width of the network and retain the important feature information required for deep target detection. The coordinate attention mechanism CA is used to perform attention calculations on the two spatial dimensions of height and width respectively, enhance feature representation, and improve network performance. C2f replaces the commonly used C3 module to ensure lightweight while obtaining more gradient information, and then performs knowledge distillation on the improved YOLOv5 model to reduce the number of parameters calculated by the model and reduce the memory occupied by the system.
[0011] Furthermore, the gradient information target detection network is mainly divided into two parts, namely the network backbone Backbone and the detection head Head. In the Backbone, the convolution module Conv and the RepNCSPELAN4 module are used to extract image features, and the spatial pyramid pooling SPPF is used to input the downsampled feature map into the Head to classify and locate the image features through Conv, the upsampling module Upsample, the tensor splicing Concat and the improved convolution module C2f, and finally the regression operation is performed to output the classification result Detect.
[0012] Furthermore, the optical flow method described in step 2 utilizes the brightness constancy of the LK optical flow method, the characteristics of very small distance of each movement in continuous pixel movement, and spatial consistency to solve the weighted least squares method in a small window, and obtains the corresponding velocity component by solving the optical flow equation. Then, based on the motion change information of the object in the image, the motion information of the dynamic feature points of the object in the detection frame can be tracked and the dynamic feature points on the object in the image can be finally eliminated.
[0013] Furthermore, the static feature points in the detection frame and the static feature points outside the detection frame described in step 3 are input into the local map to optimize the pose, and obtain the pose graph after optimization, which specifically includes:
[0014] Tracking thread: Calculate the bag-of-words vector of the feature points of the current frame, and then use the bag-of-words to find the candidate keyframes for relocation that are similar to the current frame. After traversing all the candidate keyframes, match them through the bag-of-words. If there are enough matching points, use EPnP iteration to get the initial pose and mark the external points and select the internal points for BA optimization. Only the pose is optimized. If there are more than 50 matching points, the relocation tracking is successful, and local map tracking is performed. Keyframes are added as needed for local mapping thread.
[0015] Local map building thread: Receives keyframes input by the tracking thread, generates new map points based on the consensus relationship of keyframes, searches and fuses the map points of adjacent keyframes, creates more connections between existing keyframes, generates more reliable map points, performs local map optimization, optimizes the poses and map points of common view keyframes, makes tracking more stable, deletes redundant keyframes, and finally sends the optimized keyframes to the loopback thread;
[0016] Loop thread: After completing local mapping and global BA, perform pose propagation through the Sim(3) transformation of the current keyframe, correct the pose and map points of the keyframes connected to the current frame, project all map points in the closed-loop connected keyframe group into the current keyframe group, match and fuse them, add or replace the map points of the keyframes in the current keyframe group, optimize the poses of all keyframes in the essential graph, create a new thread for global BA optimization, optimize all keyframes and map points, and eliminate the accumulated trajectory error and map error.
[0017] Furthermore, the method specifically comprises the following steps:
[0018] Step 1: Select the VOC2007 dataset as the training set and test set of the target detection algorithm, identify the common dynamic object categories in the dataset, and then use the labelme software to label the dynamic object categories in the dataset to obtain the dataset labels required for training and extract the image feature information;
[0019] Step 2: The features in step 1 are calculated through the multi-layer Conv module in the backbone network. The gradient path planning is performed through RepNCSPELAN4 to increase the width of the network to retain more image features, and then the feature map is downsampled through the convolution layer. The convolution process is expressed as:
[0020] Y=X*f+b
[0021] By giving input data Among them, c is the number of input channels, h is the input data height, ω is the input data width, * is the convolution operation, b is the offset, and f is the convolution kernel. is the generated n-channel width and height feature map ω′ and h′ respectively;
[0022] Then, a coordinate attention layer CA is used to further extract features. The input feature maps of size (H, 1) and (1, W) are divided into two directions, horizontal coordinate and vertical coordinate, and global average pooling is performed respectively to obtain feature maps in the height direction. The process is expressed as:
[0023]
[0024] The feature map in the width direction similar to the above is expressed as:
[0025]
[0026] Among them, c is the number of input channels, h is the input data height, and w is the input data width. The above two transformations ensure that the accurate position information in one spatial direction is located to the target of interest; finally, spatial pyramid pooling SPPF is used to perform pooling operations at different spatial scales to extract features, and pooling operations at different scales are processed in parallel. Finally, features are fused and output at all scales to maintain the diversity of spatial information;
[0027] Step 3, the feature map fused in step 2 is subjected to 1 / 2 dimensionality reduction and 2-fold upsampling through a convolution module, and then concatenated with the feature map extracted in step 2 to obtain a high-dimensional feature map. After the concatenation operation, it is passed through a feature extraction module C2f to eliminate the aliasing effect that may be generated in the upsampling, and then the above upsampling process is performed again, the feature map after the two upsamplings is output once, and then concatenated with the feature map that has not been upsampled and has been upsampled once, and then output twice;
[0028] Step 4: The improved attention target detection network trained in step 3 is integrated into the local map tracking of the visual SLAM algorithm tracking thread through pruning and distillation operations, and a rotation histogram is established to detect rotation consistency. The candidate matching points are traversed to find the best matching point with the smallest distance, and the target detection network is used to calculate the histogram of the matching point rotation angle difference, and then the inconsistent matching points are eliminated;
[0029] Step 5: Deploy the visual SLAM model on the Jetson TX2 development board of the mobile robot, turn on the robot's sensors, run the improved YOLOv5 model and SLAM model, control the robot to perceive the information of the unknown environment and build a map, and finally output the results.
[0030] The present invention has the following advantages and positive effects:
[0031] (1) The present invention proposes an image feature extraction backbone network based on gradient path planning to capture and aggregate richer features to enhance the modeling ability of the network.
[0032] (2) The present invention reduces the computational complexity of the model through pruning and distillation operations, and runs more efficiently on the robot.
[0033] (3) The network model proposed in this invention can obtain good results on the VOC2007 dataset, which proves the superiority of the performance of the proposed model.
[0034] This method can be widely used in target detection, robot positioning and other related intelligent image processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 : The overall block diagram of the VSLAM method of the present invention to improve the target detection network;
[0036] Figure 2 : The present invention improves the YOLOv5s target detection network diagram;
[0037] Figure 3 (a): Absolute trajectory error diagram of the system of the present invention under the fr3 / walking / halfsphere sequence;
[0038] Figure 3 (b) is the absolute trajectory error diagram of ORB-SLAM2 in the fr3 / walking / halfsphere sequence;
[0039] Figure 3 (c): Absolute trajectory error diagram of the system of the present invention under the fr3 / walking / rpy sequence;
[0040] Figure 3 (d) is the absolute trajectory error diagram of ORB-SLAM2 in the fr3 / walking / rpy sequence;
[0041] Figure 3 (e): Absolute trajectory error diagram of the system of the present invention under fr3 / walking / static sequence;
[0042] Figure 3 (f) is the absolute trajectory error diagram of ORB-SLAM2 in the fr3 / walking / static sequence;
[0043] Figure 3 (g): Absolute trajectory error diagram of the system of the present invention under the fr3 / walking / xyz sequence;
[0044] Figure 3 (h) is the absolute trajectory error diagram of ORB-SLAM2 in the fr3 / walking / xyz sequence;
[0045] Figure 4 (a): Relative trajectory error diagram of the system of the present invention under the fr3 / walking / halfsphere sequence;
[0046] Figure 4 (b) Relative trajectory error diagram of ORB-SLAM2 in fr3 / walking / halfsphere sequence;
[0047] Figure 4 (c): Relative trajectory error diagram of the system of the present invention under the fr3 / walking / rpy sequence;
[0048] Figure 4(d) Relative trajectory error diagram of ORB-SLAM2 in fr3 / walking / rpy sequence;
[0049] Figure 4 (e): Relative trajectory error diagram of the system of the present invention under fr3 / walking / static sequence;
[0050] Figure 4 (f) Relative trajectory error diagram of ORB-SLAM2 in fr3 / walking / static sequence;
[0051] Figure 4 (g): Relative trajectory error diagram of the system of the present invention under the fr3 / walking / xyz sequence;
[0052] Figure 4 (h) is the relative trajectory error diagram of ORB-SLAM2 in the fr3 / walking / xyz sequence. DETAILED DESCRIPTION
[0053] In order to facilitate ordinary technicians in the field to understand and implement the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0054] The VSLAM method of improving the target detection network proposed by the present invention comprises the following steps:
[0055] Step 1: Obtain the original image through the camera, input the image into the gradient information target detection network for preprocessing, detect dynamic objects in the detection image, and generate detection boxes and detection categories.
[0056] Step 2: Use the optical flow method to remove the dynamic feature points on the object in the detection frame in step 1 to obtain the static feature points in the image.
[0057] Step 3: Input the static feature points inside the detection frame and the static feature points outside the detection frame obtained in step 2 into the local map to optimize the pose and obtain the optimized pose graph.
[0058] Figure 1 The overall block diagram of the VSLAM method for improving the target detection network is shown below. It is mainly divided into three parts: tracking thread, local map thread and loop detection thread. This method combines the lightweight gradient information target detection network with the optical flow method as a dynamic object removal module and integrates it into the tracking thread of the ORB-SLAM2 system as the front-end optimization of the image. The overall block diagram of the improved method is shown below. Figure 1As shown. In the initialization stage, the system inputs the image and loads the word bag according to the timestamp, extracts the ORB feature points, reads the camera calibration and other related parameters, and completes the initialization of the visual SLAM system. The system extracts the results and determines the object category through the lightweight gradient information target detection network, and uses the optical flow method to remove the feature points on the dynamic objects. The tracking thread performs local map tracking and optimizes the pose based on the extracted static feature points. The local map thread generates map points and performs local optimization by inserting key frames. The loop detection thread uses the word bag model to detect whether the camera passes through a known position, and then uses the detected loop as a constraint to further eliminate the accumulated errors and improve the accuracy of the system positioning.
[0059] In the step 1, YOLOv5s is used as the baseline model. After calculating the loss function in the deepening neural network to generate new gradients, some information will be lost. The RepNCSPELAN4 module combines the CSP module and the ELAN module designed by gradient planning to increase the width of the network and retain the important feature information required for deep target detection. The coordinate attention mechanism (CA) is used to perform attention calculations on the two spatial dimensions of height and width, enhance feature representation, and improve network performance. C2f replaces the commonly used C3 module to ensure lightweight while obtaining more gradient information. Then, the improved YOLOv5 model is subjected to knowledge distillation to reduce the number of parameters calculated by the model and reduce the memory occupied by the system. The model framework can be expressed as follows: Figure 2 .
[0060] Figure 2 It is a gradient information YOLOv5 target detection network, which is mainly divided into two parts, namely the network backbone (Backbone) and the detection head (Head). In the Backbone, the convolution module (Conv) and the RepNCSPELAN4 module are used to extract image features, and the spatial pyramid pooling (Spatial Pyramid Pooling-Fast, SPPF) is used to input the downsampled feature map into the Head, and the image features are classified and located through the Conv, upsampling module (Upsample), tensor splicing (Concat) and improved convolution module (C2f), and finally the regression operation is performed to output the classification result (Detect).
[0061] In step 2, the brightness constancy of the LK optical flow method, the characteristic that the distance of each movement is very small in the continuous movement of pixels, and the spatial consistency are used to solve the weighted least squares method in a small window. The corresponding velocity component is obtained by solving the optical flow equation, and then according to the motion change information of the object in the image, the motion information of the dynamic feature points of the object in the detection frame can be tracked and the dynamic feature points on the object in the image can be finally eliminated.
[0062] The static feature points inside the detection frame and the static feature points outside the detection frame are input into the local map to optimize the pose and obtain the optimized pose graph, which specifically includes:
[0063] Tracking thread: Calculate the bag-of-words vector of the feature points of the current frame, and then use the bag-of-words to find the candidate keyframes for relocation that are similar to the current frame. After traversing all the candidate keyframes, match them through the bag-of-words. If there are enough matching points, use EPnP iteration to get the initial pose and mark the external points and select the internal points for BA optimization. Only optimize the pose. If there are more than 50 matching points, the relocation tracking is successful, and local map tracking is performed. Keyframes are added as needed for local mapping thread.
[0064] Local mapping thread: Receives keyframes input by the tracking thread, generates new map points based on the consensus relationship of keyframes, searches and fuses the map points of adjacent keyframes, creates more connections between existing keyframes, generates more reliable map points, performs local map optimization, optimizes the poses and map points of common-view keyframes, makes tracking more stable, deletes redundant keyframes, and finally sends the optimized keyframes to the loopback thread.
[0065] Loop thread: After completing local mapping and global BA, perform pose propagation through the Sim(3) transformation of the current keyframe, correct the pose and map points of the keyframes connected to the current frame, project all map points in the closed-loop connected keyframe group into the current keyframe group, match and fuse them, add or replace the map points of the keyframes in the current keyframe group, optimize the poses of all keyframes in the essential graph, create a new thread for global BA optimization, optimize all keyframes and map points, and eliminate the accumulated trajectory error and map error.
[0066] During specific implementation, the present invention can use computer software technology to realize automatic operation process. Specific embodiment:
[0068] Step 1: Select the VOC2007 dataset as the training set and test set for the target detection algorithm, identify the common dynamic object categories in the dataset, and then use the labelme software to label the dynamic object categories in the dataset to obtain the dataset labels required for training and extract image feature information.
[0069] Step 2: The features in step 1 are calculated through the multi-layer Conv module in the backbone network. The gradient path planning is performed through RepNCSPELAN4 to increase the width of the network to retain more image features, and then the feature map is downsampled through the convolution layer. The convolution process can be expressed as:
[0070] Y=X*f+b
[0071] By giving input data Where c is the number of input channels, h is the input data height, ω is the input data width, * is the convolution operation, b is the offset, and f is the convolution kernel. The n channels with width and height ω′ and h′ are generated feature maps respectively.
[0072] A coordinate attention layer (CA) is used to further extract features. The input feature maps of size (H, 1) and (1, W) are divided into two directions, horizontal coordinate and vertical coordinate, and global average pooling is performed respectively to obtain feature maps in the height direction. The process can be expressed as:
[0073]
[0074] The feature map in the width direction similar to the above can be expressed as:
[0075]
[0076] Among them, c is the number of input channels, h is the input data height, and w is the input data width. The above two conversions can ensure that the accurate position information in one spatial direction locates the target of interest. Finally, Spatial Pyramid Pooling-Fast (SPPF) is used to extract features by performing pooling operations at different spatial scales, and pooling operations at different scales are processed in parallel. Finally, features are fused and output at all scales to maintain the diversity of spatial information.
[0077] In step 3, the feature map fused in step 2 is subjected to 1 / 2 dimensionality reduction and 2-fold upsampling through a convolution module, and then concatenated with the feature map extracted in step 2 to obtain a high-dimensional feature map. After the concatenation operation, it is passed through a feature extraction module (C2f) to eliminate the aliasing effect that may be caused by upsampling, and then the above upsampling process is performed again. The feature map after two upsamplings is output once, and then concatenated with the feature map that has not been upsampled and has been upsampled once, and then output twice.
[0078] Step 4: The improved attention target detection network trained in step 3 is integrated into the local map tracking of the visual SLAM algorithm tracking thread through pruning and distillation operations, and a rotation histogram is established to detect rotation consistency. The candidate matching points are traversed to find the best matching point with the smallest distance. The target detection network is used to calculate the histogram of the rotation angle difference of the matching points, and then the inconsistent matching points are eliminated.
[0079] Step 5: Deploy the visual SLAM model on the Jetson TX2 development board of the mobile robot, turn on the robot chassis and sensors such as the camera, run the improved YOLOv5 model and SLAM model, control the robot to perceive the information of the unknown environment and build a map, and finally output the results.
[0080] ORB-SLAM2 is one of the most widely recognized visual SLAM systems. In order to evaluate the improvement effect of the SLAM system of the present invention in a dynamic environment, the effects of ORB-SLAM2 and the present invention on the fr3 / walking / halfsphere, fr3 / walking / rpy, fr3 / walking / static and fr3 / walking / xyz sequences in the data set TUM are compared. The comparison results are shown in Figure 2. Figure 3 (a)- Figure 3 (h) and Figure 4 (a)- Figure 4 (h) shown.
[0081] The quantitative comparison results of the present invention are shown in Table 1-Table 2. The present invention provides the values of the root mean square error (RMSE) and standard deviation (SD) of these four sequences.
[0082] Table 1 Comparison of absolute trajectory errors between the system of the present invention and ORB-SLAM2
[0083]
[0084] Table 2 Comparison of relative translation trajectory errors between the present invention and ORB-SLAM2
[0085]
[0086]
[0087] Experimental platform: The visual SLAM experiment of the present invention is carried out on the Linux operating system of Jetson TX2, and Visual Studio Code and Pycharm are selected as the integrated development environment, and the model framework is implemented based on Python language and C++. The main hardware configuration of the experiment is: Ubuntu 18.04 64-bit operating system, the processor (CPU) is a quad-core ARM A75 Complex, the graphics card (GPU) model is NVIDIA Pascal architecture 256 NVIDIA CUDA cores, and the memory (RAM) is 8G. The hardware and development environment of deep learning are: CPU is AMD 5800H, GPU is GeForce 3060, RAM is 16G, VisualStudio Code, python3.8, CUDA10.1, PyTorch 1.13.0.
[0088] Experimental results: Compared with ORB-SLAM2, although it takes more time to preprocess, it has a significant improvement in precision and mapping accuracy. While the SLAM system processes the current frame, the detection module recognizes the image of the next frame, thereby speeding up the system's operating rate and reducing the time required for system preprocessing. When running on a Jetson TX2 device, it can process 18 frames of images per second. Compared with other SLAM systems, the real-time performance is better.
[0089] It should be understood that the above description of the preferred embodiment is relatively detailed, and it cannot be regarded as limiting the scope of protection of the patent of the present invention. Under the enlightenment of the present invention, ordinary technicians in this field can also make substitutions and modifications without departing from the scope protected by the patent requirements of the present invention, all of which fall within the scope of protection of the present invention. The scope of protection requested by the present invention shall be based on the attached claims.
Claims
1. A VSLAM method for improving target detection network, characterized in that: The method comprises the following steps: Step 1: Obtain the original image through the camera, input the image into the gradient information target detection network for preprocessing, detect dynamic objects in the detection image, and generate detection boxes and detection categories; Step 2, using the optical flow method to remove the dynamic feature points on the object in the detection frame in step 1, and obtain the static feature points in the image; Step 3: Input the static feature points inside the detection frame and the static feature points outside the detection frame obtained in step 2 into the local map to optimize the pose and obtain the optimized pose graph.
2. The VSLAM method of improved target detection network according to claim 1, characterized in that: The gradient information target detection network described in the step 1 uses YOLOv5s as the baseline model. After calculating the loss function in the deepening neural network to generate a new gradient, some information will be lost. The RepNCSPELAN4 module combines the CSP module and the ELAN module designed by gradient planning to increase the width of the network and retain the important feature information required for deep target detection. The coordinate attention mechanism CA is used to perform attention calculations on the two spatial dimensions of height and width respectively, enhance feature representation, and improve network performance. C2f replaces the commonly used C3 module to ensure lightweight while obtaining more gradient information, and then performs knowledge distillation on the improved YOLOv5 model to reduce the number of parameters calculated by the model and reduce the memory occupied by the system.
3. The VSLAM method of the improved target detection network according to claim 2, characterized in that: The gradient information target detection network is mainly divided into two parts, namely the network backbone Backbone and the detection head Head. In Backbone, the convolution module Conv and the RepNCSPELAN4 module are used to extract image features, and the spatial pyramid pooling SPPF is used. The downsampled feature map is input into the Head to classify and locate the image features through Conv, the upsampling module Upsample, the tensor concatenation Concat and the improved convolution module C2f, and finally the regression operation is performed to output the classification result Detect.
4. The VSLAM method of the improved target detection network according to claim 1, wherein: The optical flow method described in step 2 utilizes the brightness constancy of the LK optical flow method, the characteristics of very small distance of each movement in the continuous movement of pixels, and spatial consistency to solve the weighted least squares method in a small window, and obtains the corresponding velocity component by solving the optical flow equation. Then, based on the motion change information of the object in the image, the motion information of the dynamic feature points of the object in the detection frame can be tracked and the dynamic feature points on the object in the image can be finally eliminated.
5. The VSLAM method of improved target detection network according to claim 1, characterized in that: The static feature points in the detection frame and the static feature points outside the detection frame described in step 3 are input into the local map to optimize the pose, and obtain the pose graph after optimization, specifically including: Tracking thread: Calculate the bag-of-words vector of the feature points of the current frame, and then use the bag-of-words to find the candidate keyframes for relocation that are similar to the current frame. After traversing all the candidate keyframes, match them through the bag-of-words. If there are enough matching points, use EPnP iteration to get the initial pose and mark the external points and select the internal points for BA optimization. Only the pose is optimized. If there are more than 50 matching points, the relocation tracking is successful, and local map tracking is performed. Keyframes are added as needed for local mapping thread. Local map building thread: Receives keyframes input by the tracking thread, generates new map points based on the consensus relationship of keyframes, searches and fuses the map points of adjacent keyframes, creates more connections between existing keyframes, generates more reliable map points, performs local map optimization, optimizes the poses and map points of common view keyframes, makes tracking more stable, deletes redundant keyframes, and finally sends the optimized keyframes to the loopback thread; Loop thread: After completing local mapping and global BA, perform pose propagation through the Sim(3) transformation of the current keyframe, correct the pose and map points of the keyframes connected to the current frame, project all map points in the closed-loop connected keyframe group into the current keyframe group, match and fuse them, add or replace the map points of the keyframes in the current keyframe group, optimize the poses of all keyframes in the essential graph, create a new thread for global BA optimization, optimize all keyframes and map points, and eliminate the accumulated trajectory error and map error.
6. The VSLAM method of improved target detection network according to claim 1, characterized in that: The method specifically comprises the following steps: Step 1: Select the VOC2007 dataset as the training set and test set of the target detection algorithm, identify the common dynamic object categories in the dataset, and then use the labelme software to label the dynamic object categories in the dataset to obtain the dataset labels required for training and extract the image feature information; Step 2: The features in step 1 are calculated through the multi-layer Conv module in the backbone network. The gradient path planning is performed through RepNCSPELAN4 to increase the width of the network to retain more image features, and then the feature map is downsampled through the convolution layer. The convolution process is expressed as: Y=X*f+b By giving input data Among them, c is the number of input channels, h is the input data height, ω is the input data width, * is the convolution operation, b is the offset, and f is the convolution kernel. is the generated n-channel width and height feature map ω′ and h′ respectively; Then, a coordinate attention layer CA is used to further extract features. The input feature maps of size (H, 1) and (1, W) are divided into two directions, horizontal coordinate and vertical coordinate, and global average pooling is performed respectively to obtain feature maps in the height direction. The process is expressed as: The feature map in the width direction similar to the above is expressed as: Among them, c is the number of input channels, h is the input data height, and w is the input data width. The above two transformations ensure that the accurate position information in one spatial direction is located to the target of interest; finally, spatial pyramid pooling SPPF is used to perform pooling operations at different spatial scales to extract features, and pooling operations at different scales are processed in parallel. Finally, features are fused and output at all scales to maintain the diversity of spatial information; Step 3, the feature map fused in step 2 is subjected to 1 / 2 dimensionality reduction and 2-fold upsampling through a convolution module, and then concatenated with the feature map extracted in step 2 to obtain a high-dimensional feature map. After the concatenation operation, it is passed through a feature extraction module C2f to eliminate the aliasing effect that may be generated in the upsampling, and then the above upsampling process is performed again, the feature map after the two upsamplings is output once, and then concatenated with the feature map that has not been upsampled and has been upsampled once, and then output twice; Step 4: The improved attention target detection network trained in step 3 is integrated into the local map tracking of the visual SLAM algorithm tracking thread through pruning and distillation operations, and a rotation histogram is established to detect rotation consistency. The candidate matching points are traversed to find the best matching point with the smallest distance, and the target detection network is used to calculate the histogram of the matching point rotation angle difference, and then the inconsistent matching points are eliminated; Step 5: Deploy the visual SLAM model on the Jetson TX2 development board of the mobile robot, turn on the robot's sensors, run the improved YOLOv5 model and SLAM model, control the robot to perceive the information of the unknown environment and build a map, and finally output the results.