Dynamic Feature Point Removal Method for Visual SLAM Based on Deep Learning

By optimizing the YOLOv7 algorithm and multi-view geometry method, combining semantic segmentation and feature point clustering, the problem of waste of resources and insufficient accuracy of dynamic feature point recognition in visual SLAM system is solved, and rapid and accurate dynamic feature point removal is achieved, improving the real-time and accuracy of the SLAM system.

CN116977825BActive Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311013116.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2025-07-11
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

When handling dynamic objects, existing visual SLAM systems have problems of wasted resources and insufficient detection accuracy, resulting in positioning errors and map offsets. The existing methods cannot effectively distinguish dynamic and static feature points.

Method used

A deep learning-based method is adopted, and the optimized YOLOv7 algorithm is used to identify potential dynamic objects. Combined with semantic segmentation and multi-view geometry methods, dynamic feature points are judged through feature point clustering and reprojection depth differences, and the feature point removal strategy is optimized.

Benefits of technology

It realizes rapid and accurate identification of dynamic objects, reduces resource waste, improves the accuracy and speed of feature point removal, avoids the false detection of stationary objects, and improves the real-time and accuracy of the SLAM system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977825B_ABST
    Figure CN116977825B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for removing dynamic feature points in visual SLAM based on deep learning. Step S1: Obtain the target frame and input it into a pre-trained potential dynamic object recognition model to identify dynamic targets and static objects. Step S2: According to a preset condition, count the number of feature points in the target frame to determine whether the number of feature points exceeds a threshold. If not, go to step S4; if so, go to step S3. Step S3: Remove all the feature points within the dynamic target box. Step S4: Further refine the regions of the dynamic target and the static objects that coincide with it. Step S5: Perform K-Means clustering on the depth information of the feature points in the refined regions. Step S6: For each category after clustering, randomly select a partial number of feature points for dynamic point detection and perform feature point removal. Advantages: Ensure real-time performance while avoiding waste of feature points and accelerating the speed of dynamic point detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for removing dynamic feature points in visual SLAM based on deep learning, belonging to the fields of computer vision and robotics technology. Background Art

[0002] SLAM (Simultaneous Localization and Mapping) is a technology that realizes simultaneous localization and mapping by using sensor data. The SLAM system based on visual sensors is called visual SLAM. Compared with laser SLAM, it has the advantages of low cost, full utilization of image information, wide application range, etc. However, the operation of the visual SLAM system is based on a static assumption. The limitation of this assumption is that the movement of objects in the real environment is inevitable. For example, in scenarios where pedestrians, vehicles, etc. are moving, the movement of objects in the environment will have a huge impact on localization and mapping, resulting in problems such as positioning errors and map offsets. In addition, the static assumption may also cause the inability to correctly process dynamic objects in some cases, thus affecting the perception and control of the environment.

[0003] To address this limitation, the SLAM system usually adopts dynamic object detection and tracking technologies to effectively handle the negative impacts generated by moving objects in localization and mapping. The existing technology collects environmental image information through a depth camera, extracts ORB features from the collected RGB images, and at the same time performs object detection on them. The detected objects are classified into two categories: dynamic objects and static objects, and the feature points that only exist within the dynamic object frames are screened and removed. Then, the scene flow of the matching pairs between two adjacent frames is calculated, a Gaussian mixture model is established, and the dynamic objects and static objects in the scene are further separated to remove the remaining dynamic feature points. This method removes all the feature points within the target detection result frame, which will cause waste of feature points when the number of static feature points in the scene is small.

[0004] The existing technology also inputs the RGB image into the Mask R-CNN instance segmentation network to obtain the instance segmentation mask of prior high-probability moving objects (such as pedestrians and vehicles), and at the same time extracts ORB features on the RGB image. Then, motion consistency detection is performed to extract dynamic points, and the remaining static feature points after removing the dynamic points are input into the SLAM system for subsequent tracking and mapping. This method will cause waste of time resources for real-time detection of the input RGB image using a semantic segmentation network and motion consistency detection of all feature points within the prior moving object area.

[0005] In the prior art, there is also a method that is based on the ORB-SLAM2 system, uses Mask R-CNN to detect potentially moving objects, uses the method of multi-view geometry to detect dynamic feature points, and further obtains the dynamic region through the region growing method. However, this method selects to remove the feature points of all potentially moving objects, such as cars parked by the roadside, etc., which may lead to too few remaining static feature points and affect the camera pose estimation. In addition, this method uses the multi-view geometry method to perform dynamic point detection on the entire input image, increasing a large amount of time cost.

[0006] The defects existing in the prior art are as follows: using a single semantic segmentation or object detection technology to obtain prior dynamic information, the former consumes a large amount of resources, while the latter often fails to meet the accuracy requirements; the feature point removal strategy is not reasonable enough, which may lead to the situation that moving or stationary vehicles, objects dragged by people, etc. cannot be detected. In addition, detecting all feature points will consume a large amount of time. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a method for removing dynamic feature points in visual SLAM based on deep learning.

[0008] To solve the above technical problem, the present invention provides a method for removing dynamic feature points in visual SLAM based on deep learning, including:

[0009] Step S1: Obtain the target frame of the dynamic feature points to be removed, and input the target frame into a pre-trained potentially dynamic object recognition model to identify dynamic targets and static objects;

[0010] Step S2: According to the preset condition, count the number of feature points in the target frame. The preset condition is that the feature points to be counted are not in the dynamic target box corresponding to the dynamic target and not in the static target box corresponding to the static object with an overlapping part with the dynamic target box;

[0011] Judge whether the number of the feature points exceeds the threshold T. If it does not exceed, go to step S4; if it exceeds, go to step S3;

[0012] Step S3: Remove all the feature points within the dynamic target box;

[0013] Step S4: Use semantic segmentation to further refine the regions of the dynamic target and the static object overlapping with it for the dynamic target box and the static object box overlapping with it;

[0014] Step S5: Perform K-Means clustering on the depth information of the feature points within the region refined in step S4;

[0015] Step S6: Randomly select a partial number of feature points from each category after clustering in Step S5 and perform dynamic point detection using multi-view geometry. If more than a set proportion of the feature points in each cluster are detected as dynamic points, then remove all the feature points in that region; otherwise, only remove the clusters where more than the set proportion of feature points are dynamic points.

[0016] Further, the training of the potential dynamic object recognition model includes:

[0017] Obtain the network model of the optimized YOLOv7 algorithm to be trained;

[0018] Construct a training set using the COCO dataset;

[0019] Use the training set to train the network model of the optimized YOLOv7 algorithm to be trained, and output a model capable of recognizing potential dynamic objects as the trained potential dynamic object recognition model.

[0020] Further, the optimization process of the network model of the optimized YOLOv7 algorithm includes:

[0021] Obtain the network model of the original YOLOv7 algorithm;

[0022] Replace both the first ELAN module and the last ELAN module in the backbone of the network model of the original YOLOv7 algorithm with the C3 module in YOLOv5, and replace all ELAN-H modules in the head part with the C3 module in YOLOv5;

[0023] Replace the first ordinary convolution parallel to the MP in the first MP1 module in the head part of the network model structure of the original YOLOv7 algorithm with the CA attention mechanism.

[0024] Further, the said Step S5 includes:

[0025] Step S5-1: Denote the feature points in the refined region of Step S4 as (u1, u2,..., u n ), and denote the pixel coordinates and depth information of u i as (u i1 , u i2 ) and d i respectively, where i represents the serial number of the feature point, i = 1, 2,..., n, and n represents the total number of feature points. Select P points (C1, C2,..., C P ) as the initial center points;

[0026] Step S5-2: Calculate the distances between the feature point u i and the P center points. The distance between the feature point u i and the center point Ck Distance The expression is:

[0027]

[0028] where c k1 and c k2 are the abscissa value and ordinate value of the center point C k respectively;

[0029] Define the elements in the set S m as the set of feature points that are the closest to the center point C m compared to other center points; if the smallest value in represents that the feature point u i is the closest to the center point C m and is farther from other center points, then add u i to the set S m . Calculate the distance from each feature point to its closest center point and add it to the set corresponding to that center point, and so on, until each feature point is assigned to the corresponding set, 1 ≤ k ≤ P, 1 ≤ m ≤ P;

[0030] Step S5-3: Calculate the mean depth of all feature points in each set obtained in Step S5-2, and use the feature point with the depth closest to the mean as the new initial point;

[0031] Step S5-4: Determine whether the current sum of squared errors is less than the threshold ε. If it is less, go to Step S5-5; otherwise, go to Step S5-2;

[0032] The calculation expression of the sum of squared errors SSE is:

[0033]

[0034] where S i is one of the P sets in Step S5-2, and u m is any feature point in the set S i ;

[0035] Step S5-5: Randomly select 1 / 2 of the feature points from each of the obtained P sets as the finally selected feature points.

[0036] Furthermore, the said Step S6 includes:

[0037] Step S6-1: Perform feature point matching between the dynamic target box of the current key frame and the feature points other than the static target box that overlaps with the dynamic target box and the corresponding feature points of the previous key frame;

[0038] Step S6-2: Estimate the camera pose of the current frame using the EPnP algorithm, denoted as T CF ;

[0039] Step S6-3: Denote the set of feature points selected in Step S5 as The depth of these feature points is directly obtained by the depth camera, denoted as Project these feature points into the world coordinate system to obtain three-dimensional points Calculate according to the following formula:

[0040]

[0041] Where represents the i-th feature point selected in Step S5, represents the depth of the i-th feature point selected in Step S5, represents the three-dimensional point obtained by projecting the i-th feature point selected in Step S5 into the world coordinate system; T CF represents the pose of the current frame, u i , v i represent the pixel coordinates of the feature point, T represents the pose of the key frame, and K is the internal parameter matrix of the camera;

[0042] Step S6-4: Then project the three-dimensional points onto k key frames near the current frame. The point projected onto the j-th frame is denoted as The corresponding projected depth 1 ≤ j ≤ k and Calculate according to the following formula:

[0043]

[0044] Where, represents the pixel point of the three-dimensional point projected onto the j-th frame among the k key frames near it, represents the pixel point corresponding projected depth value, is the camera pose of the j-th frame;

[0045] Step S6-5: For each feature point calculate its average reprojection depth difference t i , if exceeds the threshold, then the feature point is considered a dynamic point;

[0046]

[0047] wherein, is a feature point reprojected to the depth value obtained by the depth camera of the j-th frame.

[0048] A computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to execute any of the methods.

[0049] A computer device, comprising,

[0050] one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods.

[0051] The beneficial effects achieved by the present invention:

[0052] 1. The present invention incorporates object detection and semantic segmentation technologies into traditional visual SLAM methods, enabling the rapid and accurate identification of potential dynamic objects. The object detection thread runs in real time throughout the process, and in certain cases, the semantic segmentation thread is enabled to finely divide the area, taking into account both real-time performance and accuracy;

[0053] 2. For the feature points in the dynamic regions obtained by object detection or semantic segmentation and the static object regions that overlap with them, the multi-view geometry method is further used to determine whether they are dynamic feature points, which can avoid the removal of stationary cars, people, etc., and the situation where a person holding a book, computer, etc. is not detected;

[0054] 3. By performing clustering analysis on the feature points within the potential dynamic object region and selecting a part of the clustered feature points for dynamic point detection, the speed of dynamic point detection can be accelerated to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a flowchart of a method for removing dynamic feature points in visual SLAM based on deep learning;

[0056] Figure 2 is an optimized network structure diagram of YOLOv7;

[0057] Figure 3 is a structure diagram of the C3 module;

[0058] Figure 4 is a structure diagram of the CA attention mechanism;

[0059] Figure 5 is the process of constructing an object detection model;

[0060] Figure 6 It is a schematic diagram of the multi-view geometry method. Specific implementation manner

[0061] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0062] Embodiment 1

[0063] As Figure 1 shown, a method for removing dynamic feature points in visual SLAM based on deep learning includes the following steps:

[0064] Step S1: Use the optimized YOLOv7 algorithm to detect the key frames of traditional SLAM to determine whether dynamic objects such as people and vehicles can be detected;

[0065] Step S2: If dynamic targets are detected in step S1, count the number of feature points outside the dynamic target box and the static object box (if any) that coincides with it, and determine whether it exceeds the threshold T. If it does not exceed, go to step S4; if it exceeds, go to step S3;

[0066] Step S3: Remove all the feature points within the dynamic target box;

[0067] Step S4: Use semantic segmentation to further refine the regions of the dynamic objects and the static objects that coincide with them for the dynamic target box and the static object box (if any) that coincides with it;

[0068] Step S5: Perform K-Means clustering on the depth information of the feature points in the region refined in step S4;

[0069] Step S6: Randomly select a partial number of feature points from each cluster after clustering in step S5 and use the multi-view geometry method to detect dynamic points. If more than 1 / 2 of the feature points in each cluster are detected as dynamic points, remove all the feature points in this region; otherwise, only remove the clusters where more than 1 / 2 of the feature points are dynamic points;

[0070] The specific steps of step S1 include:

[0071] Step S1-1: Since the COCO dataset contains more than 330,000 images, including 15 million targets and 80 target categories, and provides detailed annotations for each object, such as object category, position, size, and pose information. Therefore, the COCO dataset is selected as the training sample;

[0072] Step S1-2: Randomly divide the COCO dataset selected in step S1 according to the ratio of 8:1:1;

[0073] Step S1-3: Train the training set randomly partitioned in Step S4-2 using the optimized YOLOv7 algorithm, and output a model that can identify potential dynamic objects;

[0074] In the optimized YOLOv7 algorithm in Step S1-3, the following main modifications are made to the original YOLOv7 algorithm. The network model structure of the optimized YOLOv7 algorithm is as Figure 2 shown.

[0075] The optimization process is as follows: Obtain the network model of the original YOLOv7 algorithm;

[0076] Replace both the first ELAN module and the last ELAN module in the backbone of the network model of the original YOLOv7 algorithm with the C3 module in YOLOv5, and replace all ELAN-H modules in the head part with the C3 module in YOLOv5;

[0077] Replace the first ordinary convolution parallel to MP in the first MP1 module in the head part of the network model structure of the original YOLOv7 algorithm with the CA attention mechanism.

[0078] 1. Replace some structures in the YOLOv7 network structure with the C3 module in YOLOv5. This module is the main module for learning residual features. As Figure 3 shown, its structure is divided into two branches. One branch uses multiple Bottleneck stacks and 3 standard convolutional layers, and the other branch only passes through a basic convolutional module. Finally, the two branches are concatenated, greatly reducing the number of parameters in the original YOLOv7 network and accelerating the training and inference speed;

[0079] 2. Since the CA attention mechanism has a simple structure, can be flexibly inserted into classical mobile networks, has almost no computational overhead, and performs well in object detection tasks, it is introduced into YOLOv7. As Figure 4 shown, the CA attention mechanism regards each position in the feature map of the input image as a coordinate point, and maps these points to a low-dimensional space through coordinate embedding. Then, in this low-dimensional space, the similarity scores between each point and other points are calculated, and these scores form an attention matrix. Next, each row in the attention matrix is regarded as a weight vector, and this vector is multiplied by each position of the coordinate embedding to obtain the weighted coordinates of each position. Finally, the corresponding feature vectors are taken out at the weighted coordinates and aggregated to produce the output. Generally speaking, the CA attention mechanism uses the position information in the image feature map to calculate the similarity scores, thereby introducing a new position perception ability.

[0080] The specific steps of step S5 include:

[0081] Step S5-1: Denote the feature points in the refined area of step S4 as (u1, u2,..., u n ), and denote the pixel coordinates and depth information of u i as (u i1 , u i2 ) and d i respectively. Select P points (C1, C2,..., C P ) as the initial center points;

[0082] Step S5-2: Calculate the distances between the feature point u i and the P center points. The distance between the feature point u i and the center point C k is defined as: If

[0083]

[0084] the minimum value in is , then add u i to the set S m (1 ≤ m ≤ P), and so on, and assign each feature point to the corresponding set;

[0085] Step S5-3: Calculate the mean value of the depths of all feature points in each set obtained in step S5-2, and use the feature point with the depth closest to the mean value as the new initial point;

[0086] Step S5-4: Determine whether the current sum of squared errors is less than the threshold ε. If it is less, go to step S5-5; otherwise, go to step S5-2;

[0087] The sum of squared errors is defined as:

[0088]

[0089] Step S5-5: Randomly select 1 / 2 of the feature points from each of the obtained P sets as the finally selected feature points;

[0090] The specific steps of step S6 include:

[0091] Step S6-1: Match the feature points outside the current key-frame target box with the feature points of the previous key frame;

[0092] Step S6-2: Use the EPnP algorithm to estimate the camera pose of the current frame, denoted as T CF ;

[0093] Step S6-3. The feature points selected in Step S5 are denoted as The depths of these feature points are directly obtained by the depth camera and denoted as Project these feature points into the world coordinate system to obtain three-dimensional points Calculate according to the following formula:

[0094]

[0095] where K is the internal parameter matrix of the camera.

[0096] In Step S6-4, project the above three-dimensional points into k key frames near the current frame. Taking the j-th frame as an example, obtain the points and the corresponding projected depths and Calculate according to the following formula:

[0097]

[0098] where K is the internal parameter matrix of the camera, is the camera pose of the j-th frame.

[0099] Step S6-5. For each feature point calculate its average reprojection depth difference. If exceeds the threshold, the feature point is considered a dynamic point;

[0100]

[0101] where is the depth value obtained by reprojection of the feature point p i onto the depth camera of the j-th frame.

[0102] Example 2. Apply the present invention to remove dynamic feature points in visual SLAM for a certain indoor dynamic scene.

[0103] This example is divided into three parts in total: building the software and hardware environment, obtaining the semantic information model, and designing the dynamic point removal algorithm.

[0104] The hardware environment requires a server and a wireless router. Install the Hadoop distributed cluster architecture on the server to achieve distributed storage of data. Further, install the distributed database MongoDB in the cluster to store the map and camera pose information. The algorithm of this example is mainly improved based on ORB-SLAM3, so it is necessary to install ORB-SLAM3, ROS and their dependent libraries in the server.

[0105] The acquisition of the semantic information model can be divided into two parts: model training and model encapsulation. The steps for obtaining the object detection model are as follows Figure 5 shown. The training mainly uses the COCO dataset, which is randomly divided according to a ratio of 8:1:1. The optimized YOLOv7 algorithm is used for training to output an object detection model that can identify dynamic objects. Semantic segmentation is not the focus of this algorithm, so an open-source pre-trained model is selected. The steps for training the YOLOv7 training dataset using the optimized YOLOv7 include:

[0106] 1) Write the definition of the C3 module into the common.py file and add the C3 module to the corresponding position in the parse_model function in the yolo.py file;

[0107] 2) Write the definition of the CA attention mechanism into the common.py file and add the CA attention mechanism to the corresponding position in the parse_model function in the yolo.py file;

[0108] 3) Write the model configuration file yolov7_C3CA.yaml according to the Figure 3 network structure;

[0109] 4) Adjust the parameters according to the training results to make the training results optimal.

[0110] After that, the trained object detection model and the semantic segmentation model are encapsulated into a semantic acquisition interface. Its input is an RGB image, and the output is an object detection result map or a semantic segmentation result map;

[0111] The dynamic point removal algorithm mainly includes three parts: semantic acquisition method selection, K-Means clustering, and multi-view geometry method. After obtaining the key frame transmitted by the tracking thread, the semantic acquisition interface is called. Inside the interface, the object detection model is first used to detect dynamic objects. If no dynamic objects are detected, this call is ended; if dynamic objects are detected, the number of feature points in the area outside the target box is calculated. If it exceeds the threshold, all the feature points inside the target box are deleted; if it does not exceed the threshold, semantic segmentation is further used to perform pixel-level division on the area of the target box; the feature points in the area after semantic segmentation are divided into K categories using the K-Means clustering method, and specific steps for extracting some feature points from each category are as follows:

[0112] 1) Denote the feature points in the area after semantic segmentation as (u1, u2,.., u n ), and denote the pixel coordinates and depth information of u i as (u i1 , u i2 ) and d i , respectively. Select K points (C1, C2..., CK ) As the initial center point;

[0113] 2) Calculate the distances between each feature point and the K center points. The distance between feature point u i and center point C k is defined as: If

[0114]

[0115] the minimum value in is then add u i to set S m (1 ≤ m ≤ K), and so on, and assign each feature point to the corresponding set;

[0116] 3) Calculate the mean depth of all feature points in each set obtained in step 2), and take the feature point with the depth closest to the mean as the new initial point;

[0117] 4) Determine whether the current sum of squared errors is less than the threshold ε. If it is less, go to step 5); otherwise, go to step 2). The sum of squared errors is defined as:

[0118]

[0119] 5) Randomly select 1 / 2 of the feature points from each of the K sets obtained as the finally selected feature points;

[0120] After that, use the multi-view geometry method to perform dynamic point detection on the selected feature points. If more than 1 / 2 of the feature points in each class are detected as dynamic points, then remove all the feature points in this region; otherwise, only remove the classes with more than 1 / 2 of the feature points being dynamic points. The multi-view geometry principle is as Figure 6 shown. Project the feature points into the three-dimensional space and then re-project them onto the pixel plane of the nearby frame using the camera pose of the current frame and the feature point depth, and calculate the difference between the depth value of this point and the depth value obtained by the depth camera. If the difference exceeds the threshold, then this point is considered a dynamic feature point. The specific steps are as follows:

[0121] 1) Match the feature points outside the target box of the current key frame with the feature points of the previous key frame;

[0122] 2) Use the EPnP algorithm to estimate the camera pose of the current frame, denoted as T CF ;

[0123] 3) The feature points selected by the K-Means clustering method are denoted as The depths of these feature points are directly obtained by the depth camera, denoted as Project these feature points into the world coordinate system to obtain three-dimensional points Calculate according to the following formula:

[0124]

[0125] where K is the internal parameter matrix of the camera.

[0126] 4) Then project the above three-dimensional points onto k key frames near the current frame. Taking the j-th frame as an example, obtain the point and the corresponding projected depth and Calculate according to the following formula:

[0127]

[0128] where K is the internal parameter matrix of the camera, is the camera pose of the j-th frame.

[0129] 5) For each feature point calculate its average reprojection depth difference. If exceeds the threshold, the feature point is considered a dynamic point;

[0130]

[0131] where is the depth value obtained by reprojecting the feature point p i onto the depth camera of the j-th frame.

[0132] Embodiment 3, a computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to execute any of the methods.

[0133] Embodiment 4, a computer device, including,

[0134] one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods.

[0135] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0136] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0137] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0139] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for removing dynamic feature points in visual SLAM based on deep learning, characterized in that Including: Step S1: Obtain the target frame with dynamic feature points to be removed, input the target frame into a pre-trained potential dynamic object recognition model to identify dynamic targets and static objects; Step S2: Count the number of feature points of the target frame according to preset conditions. The preset conditions are: the feature points to be counted are not in the dynamic target box corresponding to the dynamic target, and not in the static target box corresponding to the static object that overlaps with the dynamic target box; Judge whether the number of the feature points exceeds the threshold T. If not, go to Step S4; if so, go to Step S3; Step S3: Remove all the feature points within the dynamic target box; Step S4: Use semantic segmentation for the dynamic target box and the static object box that overlaps with it to further refine the regions of the dynamic target and the static object that overlaps with it; Step S5: Perform K-Means clustering on the depth information of the feature points within the region refined in Step S4; Step S6: Randomly select a partial number of feature points from each cluster after clustering in Step S5 and use the multi-view geometry method for dynamic point detection. If more than a set proportion of the feature points in each cluster are detected as dynamic points, remove all the feature points in this region; otherwise, only remove the clusters where more than the set proportion of the feature points are dynamic points.

2. The method for removing dynamic feature points in visual SLAM based on deep learning according to claim 1, wherein The training of the potential dynamic object recognition model includes: Obtain the network model of the optimized YOLOv7 algorithm to be trained; Construct a training set using the COCO dataset; Use the training set to train the network model of the optimized YOLOv7 algorithm to be trained, and output a model capable of identifying potential dynamic objects as the trained potential dynamic object recognition model.

3. The method for removing dynamic feature points in visual SLAM based on deep learning according to claim 2, wherein The optimization process of the network model of the optimized YOLOv7 algorithm includes: Obtain the network model of the original YOLOv7 algorithm; Replace both the first ELAN module and the last ELAN module in the backbone of the network model of the original YOLOv7 algorithm with the C3 module in YOLOv5, and replace all ELAN-H modules in the head part with the C3 module in YOLOv5; Replace the first ordinary convolution parallel to MP in the first MP1 module in the head part of the network model structure of the original YOLOv7 algorithm with the CA attention mechanism.

4. The method for removing dynamic feature points in visual SLAM based on deep learning according to claim 1, wherein The said Step S5 includes: Step S5-1: Denote the feature points within the region refined in Step S4 as (u1, u2,..., u n ), and denote the pixel coordinates and depth information of u i as (u i1 , u i2 ) and d i respectively. Here, i represents the serial number of the feature point, i = 1, 2,..., n, where n represents the total number of feature points. Select P points (C1, C2,..., C P ) as the initial center points; Step S5-2, calculate the distance between the feature point u i and the P center points. The distance between the feature point u i and the center point C k is expressed as: where c k1 and c k2 are the abscissa value and ordinate value of the center point C k respectively; Define the set S m The elements in it are the set of feature points that are the closest to the center point C compared to other center points; if m the minimum value in is indicating that the feature point u i is the closest to the center point C m and is far from other center points, then add u i to the set S m , calculate the center point closest to each feature point and add it to the set corresponding to that center point, and so on, assign each feature point to the corresponding set, 1 ≤ k ≤ P, 1 ≤ m ≤ P; Step S5-3: Calculate the mean of the depths of all the feature points in each set obtained in Step S5-2, and take the feature point with the depth closest to the mean as the new initial point; Step S5-4: Judge whether the current sum of squared errors is less than the threshold ε. If so, go to Step S5-5; otherwise, go to Step S5-2; The calculation expression of the sum of squared errors SSE is: Among them, S i is one of the P sets in step S5-2, and u m is a feature point in set S i ; Step S5-5: Randomly select 1 / 2 of the number of feature points from each of the P sets obtained as the finally selected feature points.

5. The method for removing dynamic feature points in visual SLAM based on deep learning according to claim 4, characterized in that The said Step S6 includes: Step S6-1: Match the feature points outside the dynamic target box of the current key frame and the static target box that overlaps with the dynamic target box with the corresponding feature points of the previous key frame; Step S6-2: Estimate the camera pose of the current frame using the EPnP algorithm, denoted as T CF ; Step S6-3. The set of feature points selected in Step S5 is denoted as The depths of these feature points are directly obtained by the depth camera and denoted as Project these feature points into the world coordinate system to obtain three-dimensional points Calculate according to the following formula: Among them represents the i-th feature point selected in step S5, represents the depth of the i-th feature point selected in step S5, represents the three-dimensional point obtained by projecting the i-th feature point selected in step S5 into the world coordinate system; T CF represents the pose of the current frame, u i , v i represent the pixel coordinates of the feature point, T represents the pose of the key frame, and K is the internal parameter matrix of the camera; Step S6-4. Then project the three-dimensional points onto k key frames near the current frame. The point projected onto the j-th frame is denoted as The corresponding projected depth where 1 ≤ j ≤ k, and is calculated according to the following formula: Among them, represents a three-dimensional point projected onto the pixel point of the j-th frame among the nearby k key frames, represents the pixel point corresponding projected depth value, is the camera pose of the j-th frame; Step S6-5: For each feature point calculate its average reprojection depth difference t i , if exceeds the threshold, the feature point is considered a dynamic point; Among them, is the feature point reprojected to the depth value obtained by the depth camera of the j-th frame.

6. A computer-readable storage medium storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods of claims 1 to 5.

7. A computer device, characterized in that, Comprising, One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods of claims 1 to 5.

Citation Information

Patent Citations

  • Visual SLAM method based on semantic segmentation of deep learning

    CN112132897A

  • SLAM method for removing dynamic target based on RGBD sensor

    CN114283198A