Visual SLAM method under indoor dynamic scene based on deep learning

Through the instance segmentation model and geometric constraint method based on YOLOv8-seg, dynamic feature points are eliminated and dense point cloud maps are constructed, which solves the accuracy and robustness problems of traditional visual SLAM systems in dynamic scenarios, and realizes high-precision indoor dynamic scene SLAM.

CN120388139APending Publication Date: 2025-07-29SHENYANG LIGONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510476721.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Traditional visual SLAM systems have low positioning accuracy and robustness in dynamic scenarios, and the sparse point cloud maps they build cannot meet the needs of complex applications.

Method used

The lightweight instance segmentation model and geometric constraint method based on YOLOv8-seg are adopted to remove dynamic feature points, build dense point cloud maps, and optimize point cloud maps using statistical filtering and voxel filtering.

Benefits of technology

It improves the robustness of pose estimation in dynamic scenarios, builds a dense point cloud map for static scenarios, and expands the application scenarios of SLAM systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388139A_ABST
    Figure CN120388139A_ABST
Patent Text Reader

Abstract

The invention provides a visual SLAM method in an indoor dynamic scene based on deep learning, and relates to the technical field of synchronous localization and map building.The method comprises the steps that firstly, an instance segmentation model is designed based on YOLOv8-seg to lighten the segmentation model, and dynamic objects in a complex scene are detected through the instance segmentation model to assist a subsequent SLAM system; extraction of feature points in a dynamic object region is suppressed. Secondly, a geometric constraint method is used to reject dynamic points for the second time, and the influence of undetected dynamic objects on the SLAM system is weakened. And finally, designing a dense point cloud map construction module on this basis, removing a dynamic region in the scene by combining instance segmentation, optimizing the point cloud map by using statistical filtering and voxel filtering, constructing a complete dense point cloud map, and expanding the application scene of the SLAM system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of simultaneous localization and mapping, and in particular, to a visual SLAM method based on deep learning in an indoor dynamic scene. Background Art

[0002] In recent years, with the development of technology, robot technology has become increasingly mature. More and more robots replace humans to complete some simple repetitive or high-risk tasks. Mobile robots, with their characteristics of high efficiency and low cost, are currently widely used in various industries. It is clearly stated in the "14th Five-Year Plan for the Development of the Robot Industry" issued by the Ministry of Industry and Information Technology that the robot industry will be taken as the development goal for the next few decades, aiming to improve the intelligence level of robots and strive to make robots an important part of people's lives by 2035. Currently, most robots can only achieve positioning and navigation on the premise of manual control or pre-constructing corresponding map information. With the expansion of the application fields of robots, the requirements for the intelligence level of robots are constantly increasing. Mobile robots need to be able to perform some tasks on their own, which requires robots to have the ability to recognize unknown environments. Therefore, how to solve this problem and enable mobile robots to achieve positioning and autonomous movement has great research value and application prospects.

[0003] For a mobile robot to achieve autonomous movement, it needs to have a certain environmental perception ability and understanding ability. The simultaneous localization and mapping (SLAM) technology is the premise for a robot to achieve autonomous movement and navigation in an unknown environment. A mobile robot first estimates its current pose and surrounding environment information through a SLAM system, and then plans a moving route through intelligent planning. The development of SLAM technology has a history of more than 30 years. The figure of SLAM technology can be seen from household sweeping robots to the current popular autonomous driving field. How to improve the positioning and mapping capabilities of the SLAM system in complex scenarios is still a current research hotspot. Considering cost, visual SLAM has more application prospects. Traditional visual SLAM systems can achieve good results in pose estimation and mapping in static scenes. However, in actual application scenarios, there are mostly dynamic objects, and the movement of objects will have a great impact on pose estimation, ultimately leading to a decrease in the accuracy and robustness of the SLAM system. Therefore, how to remove the influence of dynamic objects in the real scene on the positioning accuracy has become the focus of the current research direction. Most SLAM systems are designed based on static scenes. When dynamic objects appear indoors, it will cause a decrease in the accuracy of the SLAM system. In addition, most visual SLAM systems construct sparse point cloud maps, which cannot be used for more complex applications. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention provides a visual SLAM method based on deep learning in an indoor dynamic scene. Based on the object detection and geometric constraint method, it reduces the problems of feature mis-matching, pose estimation cumulative error and incomplete map construction caused by the interference of moving objects in the SLAM algorithm in the dynamic scene, and improves the accuracy of the SLAM algorithm in complex scenes.

[0005] A visual SLAM method based on deep learning in an indoor dynamic scene is constructed specifically through the following methods:

[0006] Step 1: Design a lightweight instance segmentation model based on the YOLOv8-seg model;

[0007] Step 1.1: Lightweight instance segmentation model;

[0008] Select the GhostBottleneck module to replace the Bottleneck of C2f in the existing YOLOv8-seg model, and name it C2f-GhostBottleneck; the GhostBottleneck module consists of two stacked Ghost modules;

[0009] Step 1.2: Introduce the attention mechanism;

[0010] Add the SimAM attention mechanism to the neck of the YOLOv8-seg model;

[0011] Step 1.3: Train the instance segmentation model;

[0012] Train the instance segmentation model, and preprocess the existing publicly available COCO dataset. Only 3 categories are saved in the COCO dataset, namely people, cats and dogs. Select the model scale of n in YOLOv8-seg as the benchmark model. At the same time, set the training rounds of the instance segmentation model to 300 rounds, the batch size to 96, and the initial learning rate of the model to 0.01;

[0013] Step 2: Design a dynamic feature point elimination algorithm to eliminate dynamic points;

[0014] Step 2.1: Elimination of dynamic points based on instance segmentation;

[0015] Through the instance segmentation model in Step 1, people, cats and dogs in the existing publicly available COCO dataset are defined as high-dynamic objects, and the remaining categories are defined as potential dynamic objects. Use the instance segmentation model to divide the detected semantic information, and preset the area of high-dynamic objects, that is, the high-dynamic area, and identify the high-dynamic area through the mask information therein; subsequently, for the feature points obtained by ORB feature extraction in the SLAM system, the feature points located in the high-dynamic area will be regarded as abnormal points and eliminated.

[0016] Step 2.2: Suppress feature point extraction in the dynamic region;

[0017] Input the RGB image, use the instance segmentation model to identify high-dynamic objects, namely three categories of people, cats, and dogs, generate mask information, and use the mask information after instance segmentation to assist the SLAM system in extracting feature points for ORB feature extraction;

[0018] Step 2.3: Eliminate dynamic points based on the epipolar geometry constraint method;

[0019] Use the epipolar geometry constraint to determine whether the feature points extracted by ORB feature extraction are dynamic points; if the projection of the feature point in the adjacent frame deviates from the epipolar line by more than the set threshold, then this point is considered a dynamic point; if the projection of the feature point in the adjacent frame deviates from the epipolar line by less than the set threshold, then this point is considered a static point; if the feature point is a dynamic point, it is eliminated as an outlier.

[0020] Step 3: Construction of a dense point cloud map with dynamic point elimination;

[0021] Construct a dense point cloud map in the SLAM system with dynamic point elimination, use the instance segmentation method to eliminate the point cloud information in the high-dynamic region, and optimize the dense point cloud map using statistical filtering and voxel filtering;

[0022] Step 3.1: First, input the RGB-D data, perform feature extraction and instance segmentation operations on the image data respectively, detect dynamic objects through instance segmentation to eliminate dynamic points, and set the corresponding dynamic region as an outlier to be eliminated when generating the point cloud map later. At the same time, for each frame of the image, compare the pose difference and the field of view overlap with the nearest key frame. If the pose difference is large and the field of view overlap is low, it is determined as a key frame.

[0023] Step 3.2: Detect whether each frame of the image is a key frame. If it is not a key frame, wait for the key frame to be obtained again in the RGB-D image data; if it is a key frame, generate a dense point cloud map, obtain the coordinate points in the world coordinate system through the depth information and the RGB image, and generate a local point cloud map.

[0024] Step 3.3: Stitch the generated local point cloud maps, use statistical filtering to remove outliers and noise in the point cloud map, use voxel filtering to downsample the point cloud, reduce the number of point clouds on the premise of retaining the original features and the dense map effect, select appropriate experimental parameters through experiments, and then wait for the next key frame to arrive. When the image stream ends, the process ends.

[0025] The beneficial effects produced by adopting the above technical solutions are as follows:

[0026] The present invention provides a visual SLAM method based on deep learning in an indoor dynamic scene. First, it can achieve effective pose estimation in an indoor dynamic scene, improving the robustness in a dynamic scene. Second, it constructs a dense point cloud map of the static scene, eliminating the influence of dynamic objects on map construction and expanding the application scenario of the SLAM system method. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is the overall structure diagram of the instance segmentation model provided by the specific embodiment of the present invention;

[0028] Figure 2 It is the structural diagram of the visual SLAM system for dynamic point removal provided by the specific embodiment of the present invention;

[0029] Figure 3 It is the effect diagram of dynamic point removal provided by the specific embodiment of the present invention;

[0030] Figure 4 It is the geometric constraint diagram provided by the specific embodiment of the present invention;

[0031] Figure 5 It is the flow chart of dense point cloud map construction provided by the specific embodiment of the present invention;

[0032] Figure 6 It is the comparison diagram of the dynamic scene point cloud map effects provided by the specific embodiment of the present invention;

[0033] Among them, (a) is the schematic diagram of the point cloud effect corresponding to the fr3_wh motion sequence of the ORB-SLAM3 system in the TUM dataset, (b) is the schematic diagram of the point cloud effect corresponding to the fr3_wr motion sequence of the ORB-SLAM3 system in the TUM dataset, (c) is the schematic diagram of the point cloud effect corresponding to the fr3_ws motion sequence of the ORB-SLAM3 system in the TUM dataset, (d) is the schematic diagram of the point cloud effect corresponding to the fr3_wh motion sequence of this method in the TUM dataset, (e) is the schematic diagram of the point cloud effect corresponding to the fr3_wr motion sequence of this method in the TUM dataset, (f) is the schematic diagram of the point cloud effect corresponding to the fr3_ws motion sequence of this method in the TUM dataset. SPECIFIC EMBODIMENTS

[0034] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0035] A visual SLAM method based on deep learning in an indoor dynamic scene specifically includes the following steps:

[0036] Step 1: Design a lightweight instance segmentation model based on the YOLOv8-seg model;

[0037] In the SLAM system, the instance segmentation model can be used to identify dynamic objects, and then combined with an algorithm to eliminate dynamic objects, thereby improving the accuracy of map construction and pose estimation. First, YOLOv8-seg is used for instance segmentation in the SLAM system, and an instance segmentation model is designed based on YOLOv8-seg.

[0038] Step 1.1: Lightweight instance segmentation model;

[0039] Select the GhostBottleneck module to replace the Bottleneck of C2f in the existing YOLOv8-seg model, and name it C2f-GhostBottleneck to achieve the lightweight of the model; the GhostBottleneck module consists of two stacked Ghost modules, and the main difference from the Bottleneck module in C2f is that it replaces the ordinary convolution therein, so it can significantly reduce the number of model parameters and computational complexity without affecting the model's expressive ability.

[0040] Step 1.2: Introduce the attention mechanism;

[0041] Add the SimAM attention mechanism to the neck of the YOLOv8-seg model; without introducing additional parameters, dynamically allocate weights to highlight important information, suppress irrelevant information, and improve the segmentation performance of YOLOv8-seg. SimAM generates attention weights only by calculating the mean and variance, avoiding the introduction of additional convolutional layers and fully connected layers that need to be learned, thus reducing the computational complexity.

[0042] Step 1.3: Train the instance segmentation model;

[0043] Train the instance segmentation model to provide support for identifying dynamic objects when eliminating dynamic feature points subsequently. In this embodiment, the structure diagram of the final instance segmentation model is as Figure 1 shown. And preprocess the existing publicly available COCO dataset, and only save 3 categories in the COCO dataset, namely people, cats, and dogs. Finally, there are 69,713 train photos and 2,945 val photos in the training dataset. Select the model scale of n in YOLOv8-seg as the benchmark model, and at the same time set the training rounds of the instance segmentation model to 300 rounds, the batch size to 96, and the initial learning rate of the model to 0.01. Train on the processed COCO dataset on the instance segmentation model of the present invention to generate a pt file that can identify three categories of people, cats, and dogs for subsequent elimination of dynamic feature points.

[0044] Step 2: Design a dynamic feature point elimination algorithm to eliminate dynamic points;

[0045] The overall visual SLAM system structure diagram designed in this embodiment is as Figure 2 shown;

[0046] Step 2.1: Elimination of dynamic points based on instance segmentation;

[0047] Using the instance segmentation model in Step 1, humans, cats, and dogs in the existing publicly available COCO dataset are defined as high-dynamic objects, which are directly eliminated using the mask information of instance segmentation in the later stage. The remaining categories are defined as potential dynamic objects, which are eliminated using geometric constraint methods in the later stage. The instance segmentation model is used to divide the detected semantic information, and the area of high-dynamic objects is preset, which is the high-dynamic area. The high-dynamic area is identified through the mask information therein. Subsequently, for the feature points obtained by ORB feature extraction in the SLAM system, the feature points located in the high-dynamic area will be regarded as abnormal points and eliminated.

[0048] The basic effect of the dynamic point detection method based on semantic information and geometric constraints in this embodiment is as Figure 3 . First, the improved instance segmentation is used to detect the RGB-D image information, obtain the mask information and perform dilation processing. Secondly, in the initial feature extraction stage, the mask information is used and the method of extracting feature points in the dynamic area proposed later is used to reduce the feature extraction in the dynamic area, and then the feature points are extracted from the RGB-D image data. Finally, for the extracted feature points, in the Tracking stage, the dynamic points in the high-dynamic area of the mask area are directly eliminated through the mask information, and finally only the feature points after dynamic point elimination are retained. At the same time, it can be seen from the principle of ORB feature extraction that corner points can also be extracted around the mask area. Therefore, around the dynamic area, calculating the descriptors corresponding to the corner points will contain pixel information of the dynamic area, which is not conducive to subsequent feature matching. In view of this, the original image is processed using the dilation method, that is, the specific mask area is dilated, and in this way, the feature points with abnormal descriptors are reduced.

[0049] Step 2.2: Suppress feature point extraction in the dynamic area;

[0050] Input the RGB image, use the instance segmentation model to identify high-dynamic objects, that is, three categories of humans, cats, and dogs, generate mask information, and use the mask information after instance segmentation to assist the SLAM system in extracting feature points by ORB feature extraction to reduce the feature point extraction in the initial high-dynamic area.

[0051] In this embodiment, the method of Gaussian blur is used by combining the mask information obtained from instance segmentation in the early stage. By obtaining the RGB image information, instance segmentation is performed on the image information to obtain the mask information corresponding to each frame region. First, the mask information is dilated to obtain the corresponding dilated mask information. Then, in the ORB feature extraction stage, the processed mask information is used to perform Gaussian blur processing on the dynamic region of the original image. After Gaussian blur, feature extraction is performed, thereby reducing the number of feature extractions in the dynamic region.

[0052] Step 2.3: Eliminate dynamic points based on the epipolar geometry constraint method;

[0053] Use the epipolar geometry constraint to determine whether the feature points extracted by ORB feature extraction are dynamic points; if the projection of the feature point in the adjacent frame deviates from the epipolar line by more than the set threshold, then this point is considered a dynamic point; in an ideal situation, if it is a static feature point, the distance value should be 0, but in fact, due to the influence of factors such as the environment, the feature point may not necessarily be on the epipolar line, so a suitable threshold needs to be set. If the projection of the feature point in the adjacent frame deviates from the epipolar line by less than the set threshold, then this point is considered a static point; if the feature point is a dynamic point, it is eliminated as an abnormal point.

[0054] In this embodiment, the projection points of the coordinate point P in space under two camera views are p1 and p2 respectively, that is, the feature points extracted by the SLAM system at different frame numbers. e1 and e2 represent the epipoles, and the epipolar lines formed on the image planes I1 and I2 are l1 and l2 respectively. The camera centers corresponding to the two image planes are O1 and O2 respectively, and O1O2P forms an epipolar plane. In a static scene, when the feature matching of the SLAM system is correct, the feature points p1 and p2 are the projections of the point P on the image planes I1 and I2. According to the epipolar geometry theory, the matching feature points need to satisfy the geometric constraint. At this time, the epipolar line is expressed as Equation (1):

[0055]

[0056] In the formula, [a b c] T represents the a, b, and c in the parameter representation of the straight line equation ax + by + c = 0 of the epipolar line l2. F is the corresponding fundamental matrix. The coordinates of p1 on the plane I1 are (x1, y1), and its corresponding homogeneous coordinates are Assume that in a static environment, the feature point corresponding to p1 on the plane I2 is p2(x2, y2), and its corresponding homogeneous coordinates are When the pose estimation is correct, p2 must be on the epipolar line l2, that is, the distance from the point p2 to the straight line l2 is 0. When in a dynamic environment, assume that the point P moves to the point P′, and the corresponding feature point on I2 is p2′. From Figure 4It is known that the feature point p2′ is not on the epipolar line l2 and the distance from the epipolar line l2 is D, and its calculation method is shown in Equation (2):

[0057]

[0058] In the formula, represents the transpose of the homogeneous equation corresponding to the point after movement, and the finally obtained D represents the distance from the epipolar line l2.

[0059] Step 3: Construction of a dense point cloud map with dynamic point removal;

[0060] Construct a dense point cloud map in the SLAM system with dynamic point removal, use the instance segmentation method to remove the point cloud information in the high-dynamic area, and reduce the influence of the dynamic area on the construction of the dense point cloud map. And use statistical filtering and voxel filtering to optimize the dense point cloud map, and select appropriate filtering parameters through experimental analysis.

[0061] As Figure 5 shown, the steps for constructing the dense point cloud map are as follows:

[0062] Step 3.1: First, input RGB-D data, perform feature extraction and instance segmentation operations on the image data respectively, detect dynamic objects through instance segmentation to remove dynamic points, and set the corresponding dynamic area as an outlier point to be removed when generating the point cloud map later. At the same time, for each frame of image, compare the pose difference and the field of view overlap with the nearest key frame. If the pose difference is large and the field of view overlap is low, it is determined as a key frame.

[0063] Step 3.2: Detect whether each frame of image is a key frame. If it is not a key frame, wait in the RGB-D image data to obtain a key frame again; if it is a key frame, generate a dense point cloud map, and obtain the coordinate points on the world coordinate system through depth information and RGB images to generate a local point cloud map.

[0064] Step 3.3: Stitch the generated local point cloud maps, use statistical filtering to remove outliers and noise in the point cloud map, use voxel filtering to downsample the point cloud, reduce the number of point clouds on the premise of retaining the original features and the dense map effect, and select appropriate experimental parameters through experiments. Then wait for the next key frame to arrive, and end the process until the image stream ends.

[0065] After constructing the global point cloud map in the above manner, save it as a pcd file, which can be used to construct an octree map or a fence map to provide support for subsequent navigation or map reuse.

[0066] In this embodiment, a program is written in Python and C++ languages under the Linux system, and a visual SLAM system based on deep learning in an indoor dynamic scene is implemented according to the above process. A comparison between the instance segmentation model designed based on the COCO dataset and the original model is shown in Table 1. Taking the sequences in the TUM dataset as an example, a comparison of the absolute trajectory error with the original algorithm is shown in Table 2.

[0067] Table 1 Instance Segmentation Ablation Experiment

[0068]

[0069] Table 2 Comparison of Absolute Trajectory Errors in Dynamic Scenes

[0070]

[0071] Table 1 shows an improvement based on YOLOv8-seg. The improved effect is measured by the comprehensive performance of the algorithm through recall rate, accuracy, mAP50M, and Params. In the table, P and R represent precision and recall rate respectively, mAP50 is an index used to measure the detection and segmentation accuracy, and the subscript M represents the detection category and the segmentation category. After replacing all C2f modules in the instance segmentation model with the C2f-GhostBottleneck module, the number of parameters decreases by 29.41%, but mAP50 also decreases. After adding the SimAM attention mechanism, mAP50 exceeds the original model. Experiments show that the model proposed in this paper has a 29.41% decrease in the number of parameters, a 0.5% increase in mAP50B, and a 0.8% increase in mAP50M compared with the original model. When running on the GPU, the inference speed is increased by about 1ms.

[0072] Table 2 is a comparison between ORB-SLAM3 and the visual SLAM system based on deep learning in an indoor dynamic scene proposed in the present invention. Sequences in the SLAM dataset TUM dataset are selected for SLAM system experiments. The images of the motion sequences are divided into 2 groups. Those containing "sit" are low-dynamic scene data taken when a person sits down, and "walk" are high-dynamic scene data taken when a person walks. It includes low-dynamic motion sequences, that is, dynamic objects such as people do not move greatly, only slight movements occur in the hands, such as freiburg3_sitting_xyz, freiburg3_sitting_static, freiburg3_sitting_rpy, and

[0073] The freiburg3_sitting_halfsphere motion sequence. In a high-dynamic motion sequence, i.e., a dynamic object such as a person moves indoors, such as the freiburg3_walking_xyz, freiburg3_walking_static, freiburg3_walking_rpy, and

[0074] The freiburg3_walking_halfsphere motion sequence. For ease of recording, the above are simply referred to as fr3_sx, fr3_ss, fr3_sr, fr3_sh, fr3_wx, fr3_ws, fr3_wr, and fr3_wh. In a low-dynamic scenario, the proposed system and the original system have comparable accuracy in some aspects, indicating that the proposed system can maintain comparable accuracy to the original system in low-dynamic or static scenarios. The average RMSE of the absolute trajectory error of the improved SLAM system is still reduced by 11.33%. In a high-dynamic scenario, the proposed SLAM system has a significant improvement, and the RMSE of the absolute trajectory error is reduced by 82.94% compared to the original system. Obviously, in a high-dynamic scenario, the pose estimation effect of the present invention far exceeds that of the ORB-SLAM3 system.

[0075] Finally, as Figure 6 shown, (a), (b), (c) are the dense map construction effects of ORB-SLAM3, and (d), (e), (f) are the dense map construction effects of the present invention. Experiments show that the present invention can effectively construct a static dense map excluding dynamic points.

[0076] The above description is only the preferred embodiments of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A visual SLAM method for indoor dynamic scenes based on deep learning, characterized in that, Including the following steps: Step 1: Design a lightweight instance segmentation model based on the YOLOv8-seg model; Step 2: Design a dynamic feature point elimination algorithm to eliminate dynamic points; Step 3: Construction of a dense point cloud map with dynamic point elimination; Construct a dense point cloud map in the SLAM system with dynamic point elimination, use the instance segmentation method to eliminate the point cloud information in the high-dynamic area, and optimize the dense point cloud map using statistical filtering and voxel filtering.

2. The visual SLAM method based on deep learning in an indoor dynamic scene according to claim 1, characterized in that The specific steps of Step 1 include the following steps: Step 1.1: Lightweight instance segmentation model; Select the GhostBottleneck module to replace the Bottleneck of C2f in the existing YOLOv8-seg model, and name it C2f-GhostBottleneck; the GhostBottleneck module consists of two stacked Ghost modules; Step 1.2: Introduce the attention mechanism; Add the SimAM attention mechanism to the neck of the YOLOv8-seg model; Step 1.3: Train the instance segmentation model; Train the instance segmentation model, and preprocess the existing publicly available COCO dataset. Only 3 categories are saved in the COCO dataset, namely people, cats, and dogs. Select the model scale of n in YOLOv8-seg as the benchmark model. At the same time, set the training rounds of the instance segmentation model to 300 rounds, the batch size to 96, and the initial learning rate of the model to 0.

01.

3. A visual SLAM method based on deep learning in an indoor dynamic scene according to claim 1, characterized in that, The specific steps of Step 2 include the following steps: Step 2.1: Dynamic point elimination based on instance segmentation; Through the instance segmentation model in Step 1, define people, cats, and dogs in the existing publicly available COCO dataset as high-dynamic objects, and the remaining categories as potential dynamic objects. Use the instance segmentation model to divide the detected semantic information, and preset the area of high-dynamic objects, that is, the high-dynamic area, and identify the high-dynamic area through the mask information therein; subsequently, for the feature points obtained by ORB feature extraction in the SLAM system, the feature points located in the high-dynamic area will be regarded as abnormal points and eliminated. Step 2.2: Suppress feature point extraction in the dynamic area; Input the RGB image, use the instance segmentation model to identify high-dynamic objects, that is, 3 categories of people, cats, and dogs, generate mask information, and use the mask information after instance segmentation to assist the SLAM system to extract feature points through ORB feature extraction; Step 2.3: Eliminate dynamic points based on the epipolar geometry constraint method; Use the epipolar geometry constraint to judge whether the feature points extracted by ORB feature extraction are dynamic points; if the projection of the feature points in adjacent frames deviates from the epipolar line by more than the set threshold, then this point is considered a dynamic point; if the projection of the feature points in adjacent frames deviates from the epipolar line by less than the set threshold, then this point is considered a static point; if the feature point is a dynamic point, it will be eliminated as an abnormal point.

4. A visual SLAM method based on deep learning in an indoor dynamic scene according to claim 1, characterized in that, The specific steps of Step 3 include the following steps: Step 3.1: First, input RGB-D data, perform feature extraction and instance segmentation operations on the image data respectively. Detect dynamic objects through instance segmentation, remove dynamic points, and set the corresponding dynamic regions as abnormal points, which will be removed when generating the point cloud map later. At the same time, for each frame of image, compare the pose difference and field of view overlap with the nearest key frame. If the pose difference is large and the field of view overlap is low, it is determined as a key frame. Step 3.2: Detect whether each frame of image is a key frame. If it is not a key frame, wait to obtain a key frame again in the RGB-D image data. If it is a key frame, generate a dense point cloud map, obtain the coordinate points on the world coordinate system through depth information and RGB images, and generate a local point cloud map. Step 3.3: Stitch the generated local point cloud maps, use statistical filtering to remove outliers and noise in the point cloud map, use voxel filtering to downsample the point cloud, reduce the number of point clouds while retaining the original features and the effect of the dense map, select appropriate experimental parameters through experiments, and then wait for the next key frame to arrive. Until the image stream ends, the process ends.