Visual SLAM method and system based on point-line fusion in dynamic scene

By adopting a dot-line fusion method in dynamic scenarios in visual SLAM system, combined with ORB-SLAM3, improved YOLOv8 network and LSD line feature extraction algorithm, the problem of precise positioning and mapping in weak textures and dynamic environments is solved, and higher robustness and real-timeness are achieved.

CN119991806APending Publication Date: 2025-05-13ANHUI UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510098772.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to achieve precise positioning and effective mapping in weak texture and dynamic environments, especially in low-light, weak texture environments and dynamic scenes.

Method used

Using the visual SLAM method based on point-line fusion in dynamic scenarios, by building a visual SLAM framework, scene images are acquired and preprocessed, semantic information and feature points are extracted, keyframes are filtered, sparse point cloud maps are constructed, and closed-loop detection and map correction are performed. This method combines ORB-SLAM3, improved YOLOv8 network and LSD line feature extraction algorithm.

Benefits of technology

It significantly improves the robustness and real-time nature of SLAM systems in dynamic and low-textured environments, realizes more accurate positioning and map construction, and reduces the risk of trajectory deviation and system crash.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991806A_ABST
    Figure CN119991806A_ABST
Patent Text Reader

Abstract

The invention discloses a visual SLAM method and system based on point-line fusion in a dynamic scene, and belongs to the technical field of robots. Comprising the following steps: constructing a visual SLAM framework; acquiring and preprocessing a scene image; inputting the preprocessed scene image into a visual SLAM framework, and extracting semantic information; carrying out feature point extraction and fusion on the scene image to obtain a static feature point set; carrying out initialization through a visual SLAM framework according to the static feature point set, and screening a key frame scene image; constructing a sparse point cloud map according to the screened key frame scene image; compared with the prior art, closed-loop detection and map correction are carried out on a key frame scene image through a visual SLAM framework, the method has the beneficial effects that YOLOv8 and an LSD line feature extraction algorithm are improved based on ORB-SLAM 3 fusion, a semantic visual SLAM method and system are constructed, and the robustness and real-time performance of the SLAM system in a dynamic environment and a low-texture environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics, and more specifically to a visual SLAM method and system based on point-line fusion in a dynamic scene. Background Art

[0002] With the advancement of science and technology, robotics technology has developed rapidly, and mobile robots have emerged. As a collection of multiple technologies based on sensors, data processing, and information decision-making, mobile robots can autonomously perform various tasks in various environments. The application areas of mobile robots are becoming increasingly wide, covering industrial production, agriculture, aerospace, and service industries.

[0003] In order to realize the comprehensive application of mobile robots in production and life, its breakthrough development lies in the realization of automation and autonomy. This means that mobile robots need to have the ability to move purposefully in unfamiliar environments and be able to complete various complex tasks. In order to achieve this goal, the most critical problems that need to be solved are positioning and mapping. Mobile robots obtain environmental data through sensors and determine their own positions in unfamiliar environments by analyzing these data; and mobile robots obtain environmental information through sensors and identify the surrounding environment through continuous movement to build maps that can be used for navigation. During the movement process, mobile robots also need to continuously optimize location information and improve environmental maps to achieve autonomous navigation. Therefore, in order for robots to achieve autonomous navigation, they must solve the problems of positioning and mapping. In this process, in order to achieve the real-time requirements, positioning and mapping are usually carried out simultaneously, that is, the problem of simultaneous localization and mapping (SLAM) to achieve real-time requirements.

[0004] Since its birth in the 1980s, SLAM (Simultaneous Localization and Mapping) technology has become a key technology in the fields of mobile robots, autonomous driving, drones, and augmented reality. With the advancement of science and technology in these fields, SLAM technology has also experienced rapid development from theoretical exploration to practical application. Due to its important theoretical and application value, SLAM technology is regarded as the key to enabling mobile robots to autonomously explore unknown areas. The core of SLAM technology is to collect environmental information through sensors such as cameras, lidar, and inertial measurement units (IMUs), and use advanced algorithms to integrate this information to estimate the specific position of the robot in the environment and finally build an environmental map. Among them, cameras are favored for their low cost, small size, and convenience, which has promoted the rapid development of visual SLAM technology. However, visual SLAM technology also faces some challenges, especially in indoor low-light, weak-texture environments, and dynamic scenes. In indoor low-light and weak-texture environments, traditional point-feature-based SLAM algorithms may fail to build maps due to inaccurate position estimation. In dynamic scenes, traditional SLAM algorithms will cause ghosting in the constructed dense point cloud map. At the same time, if a large number of feature points distributed on dynamic objects are extracted, it will also cause large trajectory deviations in pose estimation and even cause the system to crash.

[0005] In the related technology, for example, Chinese patent CN114034299A provides a navigation system based on active laser SLAM, which includes active laser SLAM in unknown environments and autonomous navigation when the map is known. Both the active laser SLAM part and the autonomous navigation part include a mobile robot module, a laser radar module and a processor module. The processor module in the active laser SLAM part includes a laser SLAM module, an active exploration module and a path planning module. The processor module in the autonomous navigation part includes a positioning module and a path planning module. The dynamic obstacle avoidance capability of the robot is improved through the local path planning module. However, this technical solution does not provide any technical inspiration on how to extract effective feature points in a weak texture environment, and how to quickly and effectively detect and eliminate dynamic objects. Summary of the invention

[0006] 1. Technical issues to be solved

[0007] In view of the problem of how to achieve precise positioning in weak texture and dynamic environment in the prior art, the present invention provides a visual SLAM method and system based on point-line fusion in dynamic scenes, which can achieve real-time precise positioning and effective mapping in dynamic scenes.

[0008] 2. Technical solution

[0009] The purpose of the present invention is achieved through the following technical solutions.

[0010] The content of this application is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this application is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0011] Some embodiments of the present application propose a visual SLAM method and system based on point-line fusion in dynamic scenes to solve the technical problems mentioned in the above background technology section.

[0012] As the first aspect of the present application, some embodiments of the present application provide a visual SLAM method based on point-line fusion in a dynamic scene, including the following steps: building a visual SLAM framework; acquiring and preprocessing a scene image; inputting the preprocessed scene image into the visual SLAM framework to extract semantic information; and extracting and fusing feature points of the scene image to obtain a static feature point set; initializing through the visual SLAM framework and based on the static feature point set to filter key frame scene images; constructing a sparse point cloud map based on the filtered key frame scene images; and performing closed-loop detection and map correction on the key frame scene images through the visual SLAM framework.

[0013] Furthermore, the visual SLAM framework includes parallel semantic threads, tracking threads, initialization threads, local mapping threads, and loop detection threads; the semantic thread uses the YOLOv8 network with SCConv convolution modules and EMA attention mechanism added to the backbone feature extraction network.

[0014] Furthermore, the process of obtaining and preprocessing a scene image is specifically as follows: obtaining a scene image through a camera; obtaining a camera intrinsic parameter matrix; projecting a spatial point onto a normalized plane to obtain a normalized coordinate point corresponding to the spatial point; dedistorting the normalized coordinate point to obtain a dedistorted normalized coordinate point; and projecting the dedistorted normalized coordinate point onto a pixel plane of the camera to obtain a pixel coordinate.

[0015] Furthermore, the preprocessed scene image is input into the visual SLAM framework, and the target area of ​​the input scene image is segmented through the semantic thread to generate an object mask for distinguishing dynamic objects from static objects.

[0016] Furthermore, the semantic thread processes the input scene image frame by frame through the ORB algorithm to extract feature points, and identifies the object category and position in the scene image through the detection head of the YOLOv8 network to segment the target area.

[0017] Furthermore, the preprocessed scene image is input into the visual SLAM framework, and the feature points of the mixed dynamic feature points and static feature points of the scene image are extracted frame by frame through the tracking thread; the object mask generated by the semantic recognition thread is input into the tracking thread, and the feature points of the dynamic objects are filtered in combination with the object mask, and the feature points of the static objects are retained to obtain a set of static feature points.

[0018] Furthermore, the process of screening key frame scene images is specifically as follows: based on the feature point matching relationship between the current frame scene image and the previous frame scene image, the pose of the current frame scene image is estimated; the current frame scene image is matched with the key frame scene image in the local map to adjust the pose of the current frame scene image; and the key frame scene image is screened out based on the image quality of the current frame scene image and the degree of overlap with the existing key frame scene images.

[0019] Furthermore, the process of constructing a sparse point cloud map is as follows: putting the key frame scene images into the same queue, sequentially removing the map points that do not meet the requirements in all the key frame scene images in the queue, and generating new map points between the common view key frame scene images; fusing the repeated map points in the current frame scene image and the adjacent frame scene images; projecting the straight line segments corresponding to the line features onto the sparse point cloud map, and generating additional map points near the projection position.

[0020] Furthermore, the similarity between key-frame scene images is calculated using feature descriptors through the closed-loop detection thread. When the similarity score exceeds the set threshold, a closed-loop event is determined to have occurred. Sim3 similarity transformation is then performed to map the key-frame scene images and correct the drift error of the sparse point cloud map.

[0021] As the second aspect of the present application, some embodiments of the present application provide a system for a visual SLAM method based on point-line fusion in a dynamic scene, including a building module: building a visual SLAM framework; an acquisition module: acquiring and preprocessing scene images; an extraction module: inputting the preprocessed scene images into the visual SLAM framework to extract semantic information; and extracting and fusing feature points of the scene images to obtain a static feature point set; a screening module: initializing through the visual SLAM framework and based on the static feature point set to screen key frame scene images; a map module: constructing a sparse point cloud map based on the screened key frame scene images; a detection module: performing closed-loop detection and map correction on key frame scene images through the visual SLAM framework.

[0022] 3. Beneficial effects

[0023] Compared with the prior art, the advantages of the present invention are: based on ORB-SLAM3 fusion improved YOLOv8 and improved LSD line feature extraction algorithm, a semantic visual SLAM method and system are successfully constructed, thereby effectively improving the robustness and real-time performance of the SLAM system in dynamic environments and low-texture environments. In the TUM public data set and real environment, the visual SLAM system based on point-line fusion in dynamic scenes of this scheme is tested and compared with ORB-SLAM3, Dyna-SLAM, DS-SLAM and other systems in high-dynamic and low-dynamic scenes respectively. The results show that the visual SLAM method and system based on point-line fusion in dynamic scenes of the present invention have obvious advantages in accuracy and real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Flow chart of a visual SLAM method based on point-line fusion in a dynamic scene in one embodiment of the present invention;

[0025] Figure 2 A depth map obtained by using YOLOv8 semantic segmentation in one embodiment of the present invention;

[0026] Figure 3 Schematic diagram of the tracking process of the visual SLAM method based on point-line fusion in a dynamic scene in one embodiment of the present invention, wherein green marks are feature points and red marks are line features;

[0027] Figure 4 This is an effect diagram of a dense point cloud map without the "ghosting" phenomenon when YOLOv8 semantic segmentation is not used in one embodiment of the present invention;

[0028] Figure 5 This is an effect diagram of a dense point cloud map without "ghosting" phenomenon, in which a point cloud map obtained after dynamic objects are removed through YOLOv8 semantic segmentation in one embodiment of the present invention. DETAILED DESCRIPTION

[0029] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0030] like Figures 1 to 5 As shown, a visual SLAM method based on point-line fusion in a dynamic scene of the present invention comprises the following steps:

[0031] S1. Build a visual SLAM framework

[0032] The constructed visual SLAM framework includes semantic thread, tracking thread, initialization thread, local mapping thread and loop detection thread.

[0033] Specifically, the visual SLAM framework in this embodiment is based on the ORB-SLAM3 system, which contains multi-core threads to ensure that the system can process visual information in a dynamic environment in real time and accurately. Among them, the semantic thread, tracking thread, local mapping thread and loop detection thread are parallel threads.

[0034] The semantic thread is used to identify the object categories in the scene, and adopts the YOLOv8 network obtained by adding the SCConv convolution module and the EMA attention mechanism to the backbone feature extraction network. The YOLOv8 network is a target detection model. In this embodiment, the SCConv convolution module (a spatial compression convolution module) is introduced to enhance its feature extraction capability, and the EMA (exponential moving average) attention mechanism is added to more finely capture the key features in the image, thereby improving the accuracy and robustness of target detection.

[0035] Specifically, the tracking thread is used to track the motion trajectory of the camera in real time and can calculate the camera's position change; the initialization thread is responsible for collecting data to initialize the SLAM system, including calibration of the camera's internal parameters and generation of initial map points; the local mapping thread is used to further process data and build a local map; the loop detection thread is responsible for detecting whether the camera has returned to the previously visited location, that is, detecting loops. By comparing the similarity between the current view and the historical view, loop detection can help correct accumulated errors and achieve global map consistency. Since the semantic thread, tracking thread, local mapping thread, and loop detection thread run in parallel, the efficiency and real-time performance of the SLAM system are ensured, and each thread focuses on its specific task and communicates and works together.

[0036] S2. Obtain scene images and preprocess them

[0037] The camera on the mobile robot is calibrated, and the scene image is acquired through the camera and preprocessed for de-distortion.

[0038] Specifically, a camera calibration method is used to calibrate the camera on the mobile robot to obtain the camera parameters; the scene image acquired by the camera is dedistorted to obtain a preprocessed scene image for input into the visual SLAM framework.

[0039] Specifically, a camera calibration method is used to calibrate the camera to obtain the camera parameters.

[0040] In a specific embodiment, the camera calibration method is the Zhang Zhengyou calibration method. The camera parameters include the camera's intrinsic parameters and extrinsic parameters. The intrinsic parameters describe the camera's inherent characteristics, such as focal length, principal point coordinates, and pixel size, while the extrinsic parameters describe the camera's position and posture in the world coordinate system. In addition, the points that exist in the actual three-dimensional space are spatial points. These points are projected onto the camera's imaging plane through the camera's lens to form pixel points in the image. In order to perform de-distortion processing, it is necessary to obtain the exact position of these spatial points on the camera's imaging plane, that is, the normalized coordinate points [x, y] T .

[0041] Specifically, the specific steps of preprocessing the scene image acquired by the camera to dedistort are as follows:

[0042] First, obtain the camera intrinsic parameter matrix K. The camera intrinsic parameter matrix K is used to express the parameters of the camera imaging process. The camera intrinsic parameter matrix includes focal length, principal point coordinates and pixel size. These parameters determine the correspondence between each pixel in the image and the corresponding point in the actual scene. The camera intrinsic parameter matrix is ​​usually composed of focal length (f x ,f y ), principal point coordinates (c x ,c y ) and pixel size (d x ,d y )constitute.

[0043] Project the spatial point onto the normalized plane to obtain the normalized coordinate point [x, y] corresponding to the spatial point T , that is, the corresponding coordinates obtained after projecting the spatial point onto the normalized plane. Specifically, the intrinsic parameter matrix can be expressed as:

[0044]

[0045] f x ,f y are the focal lengths of the camera in the x and y directions, respectively, and c x ,c y are the coordinates of the optical center of the camera along the x-direction and the y-direction on the image plane, respectively.

[0046] After obtaining the intrinsic parameter matrix K, the scene image acquired by the camera is dedistorted. Distortion includes radial distortion and tangential distortion. Radial distortion causes the points in the image to shift toward the center or edge, while tangential distortion is caused by the non-parallelism of the lens and the imaging plane. Therefore, the scene image acquired by the camera is dedistorted to eliminate the image distortion caused by lens distortion.

[0047] Specifically, for the normalized coordinate point [x, y] T Dedistortion is performed to obtain the normalized coordinate point [xdistorted ,y distorted ] Τ , the process is expressed as follows:

[0048]

[0049] Among them, [k1, k2, k3, p1, p2] are the distortion coefficients of the camera lens, which are used to describe different types of distortion. More specifically, [k1, k2, k3] are the radial distortion coefficients, [p1, p2] are the tangential distortion coefficients, and r is the distance from the normalized coordinate point to the origin of the coordinate system;

[0050] Finally, the normalized coordinate point [x distorted ,y distorted ] Τ Project it onto the camera's pixel plane and get the pixel coordinates [u,v] T , the process is expressed as:

[0051]

[0052] At this point, the preprocessing process of dedistorting the scene image acquired by the camera is completed.

[0053] S3. Semantic information extraction and feature extraction and fusion

[0054] The scene image preprocessed by step S2 is input into the visual SLAM framework, and the input scene image is processed frame by frame through the semantic thread to extract semantic information and generate object masks. All ORB feature points of the scene image are extracted through the tracking thread, and the dynamic feature points in all ORB feature points are removed in combination with the object mask generated by the semantic thread. The static feature point set is obtained based on the remaining static feature points. Line features are then extracted and combined with the static feature point set to form a rich environment description.

[0055] The semantic thread is used to quickly segment the semantic information of objects in the input scene image through the YOLOv8 network and generate object masks, such as a walking person or a computer on a desktop. The object mask can assist the subsequent point feature extraction to determine whether the feature point in the image is a dynamic feature point.

[0056] Specifically, each frame of the input scene image is processed through the YOLOv8 network, ORB feature points are extracted from the current frame scene image, and the YOLOv8 detection head is applied to the feature map to identify the object category and its position in the image. On the basis of target detection, the target area is further finely segmented; then, the detection results are post-processed to improve the application effect of the mask, including smoothing edges and filling small holes to ensure that the generated object mask can accurately cover the target object and reduce the influence of noise and artifacts. The object mask finally generated is input into the tracking thread of the visual SLAM framework. These masks are used to distinguish between dynamic objects and static objects, so that these object masks are subsequently used in the tracking thread to filter out the feature points on the dynamic objects, and only the feature points on the static objects are retained for pose estimation and map construction.

[0057] By extracting semantic information, we can quickly identify the semantic information of objects in scene images, thereby achieving more accurate and robust positioning and mapping in dynamic scenes.

[0058] In a specific embodiment, the process of feature extraction and fusion of scene images aims to extract feature points from the scene images, combine with the object mask generated by the semantic thread, remove dynamic feature points, obtain a set of static feature points, and then combine with the extracted line features. The specific process of feature extraction and fusion of scene images is as follows:

[0059] In the tracking thread of the visual SLAM framework, the ORB (Oriented FAST and Rotated BRIEF) algorithm is used to extract ORB feature points from the input scene image. The extracted ORB feature points have rotation invariance and scale invariance, and are suitable for various lighting and perspective changes. Through the ORB algorithm, a series of significant feature points in the image and their corresponding feature descriptors can be obtained.

[0060] Specifically, the tracking thread is used to extract all ORB feature points of the input scene image, and then remove the dynamic feature points from all ORB feature points in combination with the object segmentation mask. The remaining ORB feature points form a static feature point set, which is then fused with the line features extracted by the LSD line feature extraction method. In the tracking thread, the ORB feature points and LSD line features of the scene image are extracted frame by frame. At this time, the ORB feature points are a mixture of dynamic feature points and static feature points; then, combined with the object mask input by the semantic recognition thread, the dynamic feature points are removed through the dynamic feature point screening algorithm.

[0061] In a specific embodiment, in combination with the object mask generated by the semantic thread, the ORB feature points located on the dynamic object are removed, that is, the dynamic feature points among all ORB feature points are removed, and the remaining feature points are regarded as static feature points. The static feature points are retained to obtain the current static feature point set.

[0062] According to the object segmentation mask provided by the semantic thread, we can know which ORB feature points are located on dynamic objects. Specifically, we introduce a dynamic feature point screening method to subdivide the object category into high dynamic objects, medium dynamic objects and low dynamic objects according to the object's motion probability in the scene. By removing feature points on dynamic objects, their influence on subsequent pose estimation is eliminated, thereby significantly improving the robustness of the SLAM system.

[0063] Specifically, objects in the scene that can move autonomously are defined as high-dynamic objects, such as people or animals such as dogs; objects in the scene that may move with the movement of people are defined as medium-dynamic objects, such as chairs, books, tea cups, etc.; objects that are stationary most of the time in the scene are defined as low-dynamic objects, such as computers and tables, and these objects can be directly identified as static objects.

[0064] More specifically, all ORB feature points are judged: if there are and only ORB feature points in the object mask of the high-dynamic object that do not belong to other object masks, then the ORB feature points in the object mask of the high-dynamic object are identified as dynamic feature points and removed from the ORB feature points, and the remaining ORB feature points remain unchanged. Then, combined with the object mask input by the semantic recognition thread, after completing the removal of dynamic feature points of all frame scene images, the remaining ORB feature points are static feature points, and a static feature point set is constructed based on these static feature points.

[0065] In addition, line features are also important information for describing the environment. Detecting straight line segments in an image is particularly important for mapping structured environments. Combining the extracted line features with a set of static feature points can form a richer description of the environment.

[0066] In this embodiment, the LSD (Line Segment Detector) algorithm is used to extract straight line segments in the image, which can detect long straight line segments in the image and obtain a series of significant straight line segments in the scene image and their corresponding parameters (such as starting point, end point and direction, etc.). After extracting the line features, the line features are fused with the static feature point set to form a feature set containing point features and line features, which is used in the subsequent SLAM algorithm to improve the positioning accuracy and the accuracy of map construction.

[0067] S4. Pose Estimation and Optimization

[0068] Through the initialization thread of the visual SLAM framework, the camera is initialized using the feature set to filter out key frame scene images. Specifically, when obtaining the initial frame scene image, the initial pose estimation and optimization of the camera are performed to provide a benchmark for the pose estimation of subsequent frame scene images.

[0069] During the tracking process of the tracking thread, the SLAM system gradually builds a local map based on the feature points detected in the continuous frame scene images and the relative position relationship between them. The local map is a dynamically updated structure consisting of a series of key frame scene images and the constraint relationships between them.

[0070] In a specific embodiment, the process of pose estimation and optimization includes pose estimation and pose optimization based on the previous frame scene image, tracking the local map, and screening the key frame scene image. The specific steps include:

[0071] Based on the previous frame scene image, the feature point matching relationship between the current frame scene image and the previous frame scene image is used to perform a preliminary pose estimation on the current frame scene image, and the preliminary pose of the current frame scene image is optimized to improve the pose accuracy;

[0072] Using the optimized pose, the current frame scene image is matched with the key frame scene image in the local map, and the pose of the current frame scene image is further adjusted and optimized to reduce the cumulative error;

[0073] While tracking the local map, several key-frame scene images are screened out for subsequent mapping and loop closure detection based on the image quality of the current frame scene image and the overlap with the existing key-frame scene images.

[0074] Among them, image quality includes clarity and the number of feature points. The default number of feature points of a keyframe scene image is 1000; the number of inliers of a keyframe scene image must exceed the set minimum threshold of 15. The number of inliers refers to the number of feature point pairs that are considered to be correctly matched between the current frame scene image and the reference keyframe scene image after feature point matching and geometric verification. In addition, the overlap standard is that the number of inliers of a keyframe scene image must exceed 90% of the number of inliers of a reference keyframe scene image to ensure that the overlap between keyframes is moderate. Thus, the keyframe scene image is screened out based on the image quality of the current frame scene image and the overlap with the existing keyframe scene image.

[0075] Specifically, the camera is initialized using the current static feature point set, and the pose of the current frame scene image is estimated and optimized based on the previous frame scene image; then, the optimized pose is used to track the local map, and the key frame scene images are selected based on the quality of the current frame scene image (such as the richness of feature points, image clarity, etc.). These key frame scene images are not only used for the current pose optimization, but also serve as an important basis for building a global map and performing closed loop detection in subsequent steps.

[0076] S5. Local Mapping

[0077] In the local mapping thread of the visual SLAM framework, a sparse point cloud map is constructed based on the filtered key frame scene images. This step is designed to connect front-end tracking and back-end optimization to provide basic data for the construction of the global map.

[0078] In this embodiment, the local mapping thread is used to obtain a sparse point cloud map based on the filtered key frame scene images; the closed-loop detection thread is used to perform closed-loop correction on the key frame scene images and correct the drift error of the sparse point cloud map.

[0079] Based on the key frame scene images filtered by the tracking thread, the local mapping thread constructs a sparse point cloud map. When constructing a sparse point cloud map, not only point features are considered, but also line features are incorporated. Point features mainly come from feature extraction algorithms such as ORB, while line features are obtained through line feature extraction algorithms such as LSD. By combining point features and line features, the geometric structure in the scene can be described more accurately, improving the accuracy and robustness of the map.

[0080] Specifically, the steps to construct a sparse point cloud map include:

[0081] Put the key frame scene images from the tracking thread into the same queue and process the key frame scene images in the queue in order. For each key frame scene image, first remove the map points that do not meet the requirements, which may be caused by mismatching or noise; then generate new map points between the common key frame scene images.

[0082] Among them, the map points that do not meet the requirements specifically refer to mismatched map points or map points that are greater than the corresponding critical threshold of 5.991 when using the chi-square test with a degree of freedom of 2 in the reprojection error. In addition, the co-viewed key frame scene images refer to the key frame scene images that can observe the same map points. If two or more key frame scene images can observe the same map points, they are considered to be co-viewed. The co-viewing relationship is very important in SLAM, which can help establish connections between key frames to build more accurate maps. Specifically, when constructing a sparse point cloud map, a key frame scene image is used as the center to find other key frame scene images that are co-viewed with it. This is usually achieved by comparing the matching of feature points between key frame scene images.

[0083] After processing all the keyframe scene images in the queue, the generated map points are fused and optimized to remove redundant map points (i.e., map points that appear repeatedly in multiple keyframe scene images), that is, the repeated map points in the current frame scene image and the adjacent frame scene image are fused; and the local map is optimized by BA (Bundle Adjustment). BA optimization is a nonlinear least squares problem that minimizes the reprojection error by adjusting the camera pose and the position of the map points, thereby improving the accuracy of the map.

[0084] While constructing the sparse point cloud map, the line features are integrated into it. The straight line segments corresponding to the line features are projected onto the sparse point cloud map, and together with the point feature points, a sparse map combining points and lines is formed. Finally, the local mapping thread outputs the sparse point cloud map, and the local map is used to stitch into a global map.

[0085] Through the above process, the local mapping thread can construct a sparse point cloud map based on the filtered key frame scene images, and integrate line features into it to form a sparse map combining points and lines.

[0086] S6, closed loop detection and map correction

[0087] Through the closed-loop detection thread of the visual SLAM framework, closed-loop detection and map correction are performed on the key frame scene images, and the drift error of the sparse point cloud map is corrected. The closed-loop detection thread is used to detect whether the camera has returned to a previously visited location, so as to use this information to correct and optimize the current posture and map, eliminate accumulated errors, and maintain the global consistency of the map.

[0088] Specifically, by comparing the similarity between the current keyframe scene image and the previous keyframe scene image (such as using a bag-of-words model or deep learning feature matching), a closed-loop event is detected; once a closed-loop event is detected, a graph optimization algorithm is used to correct the global map to eliminate accumulated errors and maintain the global consistency of the map. A closed-loop event refers to the process in which the SLAM system detects a location that the robot or camera has been to, and uses this information to correct and optimize the current posture and map.

[0089] In a specific embodiment, in the closed-loop detection thread, detection is performed by using a feature descriptor Bag of Words (BoW), and a Sim3 similarity transformation is performed. Finally, the key frame scene image is map-corrected to correct the drift error of the sparse point cloud map. Specifically, the key frame scene image is described using a feature descriptor to quickly compare the similarity between different key frame scene images, compare the description vector between the current key frame scene image and the previous key frame scene image, and calculate the similarity score between them. If the similarity score exceeds the set threshold, it is considered that a closed-loop event may have occurred. In order to ensure the accuracy of the closed-loop detection, further verification is performed by performing a Sim3 similarity transformation to calculate the geometric consistency between the current key frame scene image and the candidate closed-loop key frame scene image. If the geometric consistency meets the requirements, the closed-loop event is confirmed. Once a closed-loop event is detected, the global map is corrected. A graph optimization algorithm can be used to represent the key frame scene image and the map point as nodes and edges in the graph, and the global error function is minimized by adjusting the position of the node (i.e., the pose of the key frame) and the weight of the edge (i.e., the observation error). In the optimization process, the accumulated errors are gradually eliminated to make the global map more accurate and consistent, including adjusting the pose of the key frame scene image, updating the position of the map points, and optimizing the camera internal parameters. Finally, the optimized global map is updated to the SLAM system for subsequent navigation and positioning tasks.

[0090] like Figure 2 As shown, the YOLOv8 network used by the semantic segmentation thread of this embodiment, the YOLOv8 network includes a backbone feature extraction network Backbone and a detection head, which can accurately classify each pixel in the image. Among them, the backbone feature extraction network Backbone specifically includes: a convolution layer Conv, a C2f module and a spatial pyramid pooling fast (SPPF) module. The convolution layer Conv is a basic component in the YOLOv8 semantic segmentation network, which is used to preliminarily extract local features of the input image. The SPPF module is used to capture feature information of different scales and enhance the robustness of the model. Compared with the traditional spatial pyramid pooling (SPP) method, the SPPF module has faster calculation speed and fewer parameters.

[0091] Specifically, the C2f module consists of a convolution layer, an SCConv convolution module, and a channel fusion operation. The SCConv convolution module is an efficient convolution architecture unit that aims to reduce model parameters and computational complexity while improving feature representation capabilities. In the traditional solution, the Bottleneck module, as the core component of Cf2, is mainly used to fuse information between channels; in this embodiment, unlike the traditional solution that uses the Bottleneck module, this solution uses the SCConv convolution module. As a plug-and-play architecture unit, the SCConv convolution module can replace the standard convolution in various convolutional neural networks. After inserting the SCConv module, it can not only reduce the number of model parameters and FLOPs, but also enhance the ability of feature representation. Therefore, the SCConv convolution module is used to reduce the redundancy and parameter sharing of the convolution kernel, thereby improving computational efficiency. The channel fusion operation fuses the output feature maps of different convolutional layers by splicing or element-by-element addition to obtain richer feature information.

[0092] More specifically, the detection head is responsible for the final pixel-level classification task, which can accurately classify each pixel in the image and achieve high-quality semantic segmentation. The detection head specifically includes:

[0093] Upsampling: Upsampling upsamples the low-resolution feature map from the backbone feature extraction network Backbone to a higher resolution for more refined semantic segmentation. Upsampling is usually performed using the nearest neighbor interpolation method to expand each pixel of the low-resolution feature map into multiple pixels to generate a high-resolution feature map.

[0094] Feature fusion: The upsampled feature map is concatenated with the high-resolution feature map from Backbone to combine the deep and shallow feature information, enhance the YOLOv8 semantic segmentation model's ability to capture details, and improve the accuracy of semantic segmentation.

[0095] C2f module: In the detection head, the C2f module further processes the input feature map through convolutional layers, SCConv convolutional modules, and channel fusion operations to extract finer feature information. The C2f module is an improved building block in the YOLOv8 network, which reduces redundant parameters and improves computational efficiency through more efficient structural design.

[0096] Convolutional layer: further extracts features, improves the generalization ability and nonlinear expression ability of the model, and prepares for the final classification task. The convolutional layer includes batch normalization and activation function. It calculates the dot product between the convolution kernel and the local area of ​​the feature map by sliding the convolution kernel on the input feature map, thereby generating a new feature map.

[0097] Detection layer: The final semantic segmentation map is generated in the YOLOv8 network. The detection layer usually contains one or more convolutional layers, and the number of output channels of these convolutional layers matches the number of categories. Each channel corresponds to the prediction of a category, and the output of the detection layer is a probability map of the same size as the input image. The value of each pixel represents the probability that the pixel belongs to each category.

[0098] Loss function: During the training process, the loss function is used to calculate the difference between the output of the detection head and the true label for back propagation and model optimization. The YOLOv8 network uses a specific loss function to handle semantic segmentation tasks, such as the cross entropy loss function or the Dice loss function. By minimizing the loss function, the model's prediction results are closer to the true label, thereby improving the accuracy of semantic segmentation.

[0099] Post-processing: After the network outputs the detection results, post-processing is usually required, such as applying a threshold to filter low-confidence predictions and using non-maximum suppression (NMS) to remove overlapping predictions. After the network outputs the detection results, post-processing operations are required to filter low-confidence predictions and remove overlapping predictions. Post-processing operations include applying a threshold to filter low-confidence predictions and using a non-maximum suppression (NMS) algorithm to remove overlapping predictions. Among them, the threshold setting can be adjusted according to actual needs, and the NMS algorithm removes overlapping predictions by calculating the intersection-over-union ratio between prediction boxes. Post-processing operations can further improve the accuracy and robustness of semantic segmentation, making the output results of the model more accurate and reliable.

[0100] In order to further illustrate that the visual SLAM method based on point-line fusion in dynamic scenes provided by this embodiment achieves high positioning accuracy and real-time performance in dynamic scenes, further verification is performed, and the process is as follows:

[0101] The Evo evaluation tool is used to compare the camera pose obtained by the visual SLAM method based on point-line fusion in the dynamic scene of this embodiment with the real camera pose data provided in the TUM data set. Among them, the evaluation indicators include: absolute trajectory error (ATE), root mean square error (RMSE) and standard deviation (SD). ATE reflects the difference in absolute distance between the estimated motion trajectory and the real motion trajectory, which can evaluate the global consistency; RMSE describes the deviation between the estimated motion trajectory and the real motion trajectory, so the smaller its value, the closer the estimated trajectory is to the real value; SD reflects the discrete degree of the motion trajectory estimated by the visual SLAM method.

[0102] The visual SLAM method based on point-line fusion in dynamic scenes of this embodiment is tested on high dynamic sequences (w_xyz, w_static) and low dynamic sequences (s_static) in the TUM dataset. The test results are shown in Table 1.

[0103] Table 1 Evaluation of absolute trajectory error (ATE)

[0104]

[0105] As can be seen from Table 1, the visual SLAM method based on point-line fusion in the dynamic scene of this embodiment has an execution time of 20 milliseconds to 30 milliseconds per frame on the TUM data set, and the running speed reaches the real-time running speed of ORB-SLAM3, realizing the real-time operation of visual SLAM. It can be seen from the experimental results that the visual SLAM method based on point-line fusion in the dynamic scene of this embodiment achieves higher positioning accuracy and real-time performance in the dynamic scene. This is due to the visual feature extraction method of point-line fusion and the efficient closed-loop detection and map correction algorithm.

[0106] In a specific embodiment, a visual SLAM system based on point-line fusion in a dynamic scene includes a building module: building a visual SLAM framework;

[0107] Acquisition module: acquire scene images and preprocess them;

[0108] Extraction module: input the preprocessed scene image into the visual SLAM framework to extract semantic information; extract and fuse feature points of the scene image to obtain a set of static feature points;

[0109] Screening module: Initializes the key frame scene images through the visual SLAM framework and based on the static feature point set;

[0110] Map module: construct sparse point cloud map based on the filtered key frame scene images;

[0111] Detection module: Perform closed-loop detection and map correction on key frame scene images through the visual SLAM framework.

[0112] The technical solution of the present invention is based on ORB-SLAM3 fusion improved YOLOv8 network and improved LSD line feature extraction algorithm, and successfully constructs a visual SLAM method and system based on point-line fusion in dynamic scenes, thereby effectively improving the robustness and real-time performance of the SLAM system in dynamic environments and low-texture environments.

[0113] In the TUM public data set and real environment, the visual SLAM system based on point-line fusion in dynamic scenes of the present technical solution was tested and compared with ORB-SLAM3, Dyna-SLAM, DS-SLAM and other systems in high-dynamic and low-dynamic scenes respectively. The results show that the visual SLAM method and system based on point-line fusion in dynamic scenes of the present invention have obvious advantages in accuracy and real-time performance.

[0114] The above schematically describes the invention and its implementation methods, which is not restrictive. Without departing from the spirit or basic features of the invention, the invention can be implemented in other specific forms. What is shown in the accompanying drawings is only one of the implementation methods of the invention. The actual structure is not limited thereto, and any figure mark in the claims should not limit the claims involved. Therefore, if a person of ordinary skill in the art is inspired by it, without departing from the purpose of the invention, a structural method and an embodiment similar to the technical solution are designed without creativity, which should all fall within the scope of protection of this patent. In addition, the word "including" does not exclude other elements or steps, and the word "one" before an element does not exclude the inclusion of "multiple" elements. The multiple elements stated in the product claim can also be implemented by one element through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any specific order.

Claims

1. A visual SLAM method based on point-line fusion in a dynamic scene, comprising the following steps: Build a visual SLAM framework; Acquire scene images and preprocess them; The preprocessed scene image is input into the visual SLAM framework to extract semantic information; and feature points of the scene image are extracted and fused to obtain a set of static feature points; Initialize the key frame scene image through the visual SLAM framework and according to the static feature point set; Construct a sparse point cloud map based on the filtered key frame scene images; The key frame scene images are subjected to closed-loop detection and map correction through the visual SLAM framework.

2. The visual SLAM method based on point-line fusion in dynamic scenes according to claim 1, characterized in that: The visual SLAM framework includes parallel semantic threads, tracking threads, initialization threads, local mapping threads, and loop detection threads; the semantic thread uses the YOLOv8 network with SCConv convolution modules and EMA attention mechanism added to the backbone feature extraction network.

3. The visual SLAM method based on point-line fusion in dynamic scenes according to claim 1, characterized in that: The process of acquiring and preprocessing the scene image is as follows: acquiring the scene image through the camera; acquiring the camera internal parameter matrix; projecting the spatial point onto the normalized plane to obtain the normalized coordinate point corresponding to the spatial point; dedistorting the normalized coordinate point to obtain the dedistorted normalized coordinate point; Project the dedistorted normalized coordinate points onto the pixel plane of the camera to obtain the pixel coordinates.

4. The visual SLAM method based on point-line fusion in dynamic scenes according to claim 2, characterized in that: The preprocessed scene image is input into the visual SLAM framework, and the target area of ​​the input scene image is segmented through the semantic thread to generate an object mask for distinguishing dynamic objects from static objects.

5. The visual SLAM method based on point-line fusion in dynamic scenes according to claim 4, characterized in that: The semantic thread processes the input scene image frame by frame through the ORB algorithm to extract feature points, and identifies the object category and position in the scene image through the detection head of the YOLOv8 network to segment the target area.

6. The visual SLAM method based on point-line fusion in dynamic scenes according to claim 4, characterized in that: The preprocessed scene image is input into the visual SLAM framework, and the feature points of the mixed dynamic feature points and static feature points of the scene image are extracted frame by frame through the tracking thread; the object mask generated by the semantic recognition thread is input into the tracking thread, and the feature points of the dynamic objects are filtered in combination with the object mask, and the feature points of the static objects are retained to obtain a set of static feature points.

7. The visual SLAM method based on point-line fusion in dynamic scenes according to claim 1, characterized in that: The process of selecting the key frame scene image is specifically as follows: according to the matching relationship between the feature points of the current frame scene image and the previous frame scene image, the pose of the current frame scene image is estimated; Matching the current frame scene image with the key frame scene image in the local map to adjust the pose of the current frame scene image; The key frame scene images are screened out according to the image quality of the current frame scene image and the degree of overlap with the existing key frame scene images.

8. The visual SLAM method based on point-line fusion in dynamic scenes according to claim 1, characterized in that: The specific process of constructing a sparse point cloud map is as follows: put the key frame scene images into the same queue, remove the map points that do not meet the requirements in all the key frame scene images in the queue in order, and generate new map points between the common view key frame scene images; fuse the repeated map points in the current frame scene image and the adjacent frame scene images; project the straight line segments corresponding to the line features onto the sparse point cloud map, and generate additional map points near the projection position.

9. The visual SLAM method based on point-line fusion in dynamic scenes according to claim 1, characterized in that: The similarity between key-frame scene images is calculated using feature descriptors through the closed-loop detection thread. When the similarity score exceeds the set threshold, a closed-loop event is determined to have occurred. Sim3 similarity transformation is then performed to map the key-frame scene images and correct the drift error of the sparse point cloud map.

10. A system based on the visual SLAM method based on point-line fusion in dynamic scenes according to any one of claims 1 to 9, characterized in that: Including building modules: building a visual SLAM framework; Acquisition module: acquire scene images and preprocess them; Extraction module: input the preprocessed scene image into the visual SLAM framework to extract semantic information; extract and fuse feature points of the scene image to obtain a set of static feature points; Screening module: Initializes the key frame scene images through the visual SLAM framework and based on the static feature point set; Map module: construct sparse point cloud map based on the filtered key frame scene images; Detection module: Perform closed-loop detection and map correction on key frame scene images through the visual SLAM framework.

Citation Information

Patent Citations

  • Navigation system based on active laser SLAM

    CN114034299A

  • Semantic SLAM method based on indoor dynamic scene

    CN119048926A

  • YOLOv8-based point-line fusion visual SLAM (Simultaneous Localization and Mapping) method in indoor dynamic scene

    CN119164383A

Cited By

  • SLAM (Simultaneous Localization and Mapping) method for guiding adaptive dynamic feature fusion based on degradation weight

    CN120976691A