Dynamic scene SLAM (Simultaneous Localization and Mapping) method paying attention to fine target

The dynamic scene SLAM method based on multi-branch feature extraction and semantic-geometric consistency verification solves the problems of insufficient fine target recognition and key frame insertion lag in dynamic scenes of visual SLAM systems, and achieves high-precision positioning and map construction.

CN120635146APending Publication Date: 2025-09-12ZHONGBEI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510720596.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing visual SLAM systems suffer from insufficient precision in fine target recognition, delayed keyframe insertion, and incomplete keyframe quality assessment in dynamic scenes, which leads to decreased positioning accuracy and distorted map construction.

Method used

A dynamic scene SLAM method focusing on fine targets is adopted. Fine targets are detected through a multi-branch feature extraction module. Dynamic feature points are eliminated by combining semantic-geometric consistency verification. Keyframes are actively inserted using a keyframe detection strategy. The quality of keyframes is evaluated through multi-dimensional indicators to construct a dense map.

Benefits of technology

It significantly improves the robustness and reconstruction accuracy of the SLAM system in dynamic environments, effectively eliminates interference from dynamic objects, and improves positioning accuracy and map construction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635146A_ABST
    Figure CN120635146A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic scene SLAM method paying attention to a fine target, and belongs to the technical field of artificial intelligence. Aiming at the problems of insufficient fine target identification precision, passive lag of key frame insertion opportunity, incomplete key frame quality evaluation and the like of an existing visual SLAM system in a dynamic scene, a multi-branch feature extraction network comprising a low-dimensional feature branch and a high-dimensional feature branch is constructed; precise recognition of a fine target is realized through a feature pyramid fusion strategy; secondly, a key frame pre-judgment mechanism based on motion prediction and scene understanding is researched, and key frame collection is actively triggered before the tracking quality is obviously reduced by analyzing the continuity of a camera motion track and the significance of scene change; and finally, establishing a key frame quality evaluation system to screen an optimal key frame. According to the method, the robustness and scene adaptability of the visual SLAM system in the dynamic environment are effectively improved through the synergistic effect of multi-level feature fusion, active key frame management and comprehensive quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a dynamic scene SLAM method focusing on fine targets. Background Art

[0002] With the rapid development of robotics and computer vision, visual SLAM systems have become a core technology for autonomous navigation and environmental perception, playing a key role in fields such as autonomous driving, augmented reality, and service robotics. However, traditional visual SLAM systems, primarily based on the assumption of a static environment, often face severe interference from dynamic objects in practical applications, resulting in reduced positioning accuracy and distorted map construction. This problem urgently needs to be addressed.

[0003] In dynamic scene processing, early research focused on eliminating mismatched points through geometric consistency checks (such as RANSAC), but these methods have limited effectiveness against continuously moving dynamic objects. In recent years, deep learning-based object detection techniques (such as YOLO and Faster R-CNN) have been introduced into SLAM systems to identify dynamic objects using semantic information. Zhang et al. proposed a dynamic feature masking method that combines detection results with geometric constraints to filter dynamic points. Chen et al. developed an attention-based dynamic object segmentation network, improving the ability to recognize occluded objects. Wang et al. constructed a semantic-geometric joint optimization framework to enhance system robustness through multimodal information fusion.

[0004] Despite recent progress, key challenges remain: First, existing target detection modules lack accurate recognition of small and occluded objects in complex scenarios, leading to missed detection of dynamic features. Second, keyframe insertion strategies are often delayed and often triggered only after tracking quality has significantly degraded. Finally, residual interference from dynamic objects can affect keyframe selection, reducing the quality of subsequent mapping and loop closure detection. Therefore, developing a visual SLAM system that can accurately perceive dynamic targets, deeply fuse multi-source information, and ensure keyframe quality is crucial for improving localization and mapping performance in dynamic environments. Summary of the Invention

[0005] Aiming at the problems of insufficient fine target recognition accuracy, passive lag in key frame insertion timing and incomplete key frame quality assessment in existing visual SLAM systems in dynamic scenes, the present invention provides a dynamic scene SLAM method that focuses on fine targets.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] A dynamic scene SLAM method focusing on fine targets, the method comprising the following steps:

[0008] Step 1: Acquire dynamic scene data through sensors and preprocess it to obtain depth information and image information, and train the target detection model at the same time;

[0009] Step 2: Input the dynamic scene data pre-processed in step 1 into the trained target detection model and enter the visual SLAM process, which includes a target detection thread, a tracking thread, a local mapping thread, a loop detection thread, and a mapping thread;

[0010] The target detection thread focuses on fine target information that is easily lost, and uses a multi-branch feature extraction module to balance fine target and context information;

[0011] The tracking thread integrates the semantic information and geometric information transmitted by the target detection thread to jointly eliminate dynamic feature points, and determines the timing of key frame insertion through the key frame detection strategy to improve positioning accuracy;

[0012] The local mapping thread obtains high-quality key frames through a key frame selection strategy to improve the accuracy of map construction;

[0013] The loop detection thread roughly matches historical key frames through the bag-of-words model and confirms the loop through geometric verification and precise estimation;

[0014] The mapping thread removes dynamic targets on key frames through the combined results of target detection and geometric constraints, and only retains information related to static objects to obtain a dense map.

[0015] Furthermore, the target detection thread focuses on fine target information that is easily lost, and uses a multi-branch feature extraction module to balance fine target and context information. The specific operations are as follows:

[0016] The target detection network adds a multi-branch feature extraction module MBFE after the backbone network EfficientRep. After the backbone network extracts coarse features, the multi-branch feature extraction module MBFE extracts fine features from local, global and sequential convolution branches, balancing fine targets and contextual information, thereby improving the accuracy of fine target detection;

[0017] The multi-branch feature extraction module MBFE first expands and reconstructs the convolved feature tensor Split into a set of spatially continuous patches, whose shape is , by controlling the patch parameters The size of the patch parameter is used to distinguish between local and global branches. It is achieved by aggregating and shifting non-overlapping patches in the spatial dimension; then the channel averaging operation is performed, and its shape becomes ; And use FFN (Feed-Forward Network) linear calculation, then use the activation function to obtain the probability distribution in the spatial dimension and adjust the weight accordingly;

[0018] In addition, the multi-branch feature extraction module MBFE uses the self-attention module to achieve adaptive data enhancement. The attention module contains a series of efficient channel attention and spatial attention components, which effectively improves the efficiency of the network and reduces the loss of fine target information.

[0019] Furthermore, the tracking thread integrates the semantic information and geometric information transmitted by the target detection thread to jointly eliminate dynamic feature points, and determines the timing of key frame insertion through the key frame detection strategy to improve the positioning accuracy. The specific operations are as follows:

[0020] Semantic information provides prior knowledge for epipolar geometry, which in turn provides key support for semantic information. After matching the feature points of the two frames, the distance between the feature points of the current frame and their corresponding epipolar lines is calculated based on the basic matrix. The larger the distance, the greater the possibility that the feature point is a dynamic point:

[0021] Dynamic feature point elimination strategy: The camera is an important sensor in visual SLAM. The camera coordinate system and the image coordinate system are converted through the classic pinhole camera model to provide strong data support for the SLAM system. The pinhole camera model is based on the principle of pinhole imaging. It shines light through a pinhole onto a plane, and uses light to project a point in three-dimensional space onto a two-dimensional plane to form an image. The pinhole model simply explains the process of camera imaging. First, the pinhole camera model is modeled, where For the heart of light, is the center of the imaging plane, and the focal length of the camera is , is a three-dimensional point in space, assuming in the camera coordinate system The coordinates are , Its coordinates in the image coordinate system are .

[0022] Then in the camera model, the principle of similar triangles based on pinhole imaging is as follows:

[0023]

[0024] After arranging the camera imaging plane The coordinates are expressed as follows:

[0025]

[0026] Convert the image coordinates to pixel coordinates, and coordinate Convert to pixel coordinates in pixel coordinate system The image coordinate system and the pixel coordinate system are transformed by a certain scaling and translation. is the origin of the pixel coordinate system, and the coordinate axes are , assuming that On-axis scaling times, in Axis scaled times, the origin is shifted , so the pixel coordinate system is expressed as follows:

[0027]

[0028] Substitute into the above formula, express ,Will express , and the following expression is obtained:

[0029]

[0030] Further arrangement is expressed as follows:

[0031] Usually the intermediate matrix It is called the camera's internal parameters, which are set directly by the factory and cannot be modified later. In addition to the internal parameters, the camera also has corresponding external parameters. The external parameters are represented by the rotation and translation of the camera. It represents the rotation matrix and translation vectors , then the pinhole camera model can be expressed as:

[0032] in represents the external parameters of the camera.

[0033] Select images based on the pinhole camera imaging model, and Represents two adjacent frames of images, called image plane; and represents the center of the camera, and For and The corresponding feature points in is a three-dimensional point in space; 、 and The plane formed is called the polar plane. and Line and image plane and The intersection of and is the pole, the polar plane and the image plane 、 Intersecting lines 、 For the polar line, use 、 Represents a three-dimensional point in space Projection to the image plane The feature points may be On the line connecting the two, corresponding to the image plane On the polar line On, when is a static feature point, which is on the image plane The projection point on the epipolar line On, and Moved, its image plane Projection points and epilines on In the actual observation process, there may be noise. If it is a static feature point, like a plane The distance between the feature point and the epipolar line is less than the threshold , set the feature point and The pixel homogeneous coordinates of are: , ,in , is the coordinate value of the feature point in the image pixel coordinate system, which is represented by the basic matrix Computational image plane The polar line is represented as follows:

[0034]

[0035] Among them, the parameters , , Represent the three different coefficients of the polar equation in the image homogeneous coordinate representation, and Together they determine the direction of the epipolar lines and simultaneously affect their position in the image plane. It plays the role of translation in the epipolar equation, determining the position offset of the epipolar line relative to the coordinate origin in the image plane; Indicates transposing the matrix;

[0036]

[0037] The epipolar geometry constraint is expressed as:

[0038]

[0039] Feature Points To the pole The distance is expressed as:

[0040]

[0041] Ideally, static feature points should fall on the epipolar line. The value is 0. In actual operation, it is difficult to avoid the appearance of noise, so when The value is greater than 0 but less than the empirical threshold When the feature point is a static feature point, if the object moves along the polar line, the distance between the feature point and the polar line is It may also be greater than 0 and less than the empirical threshold , in this case, the point will be mistakenly considered to be a static feature point. In order to avoid such errors, the present invention considers adding depth projection error as a constraint condition.

[0042] Feature points of the current frame According to the pinhole camera motion, project it to the previous key frame and get the projection point and projection depth . Get the previous keyframe projection point directly from the depth image The true depth The depth projection error is expressed as Ideally, for a static feature point, the depth projection error should be close to zero.

[0043] Construct error vector by combining epipolar geometry constraint error and depth projection error . Assume that these errors are independently distributed and obey Gaussian distribution.

[0044] Inspired by ORB-SLAM2, the covariance of the features extracted at the nth level of the pyramid is expressed as:

[0045]

[0046] Among them, S is the reduced scale of each layer of the image pyramid, assuming that the standard deviation of the 0th layer is Take 1. Since the error vector E is a 2-dimensional vector, setting the threshold will cause problems. The inner product r of the calculated vector is converted into a scalar to solve this problem. So the feature point Chi-square value of Expressed as:

[0047]

[0048] in for point The present invention uses RGB-D images, where a binocular camera is simulated, with a degree of freedom of 3 and a threshold value set to The chi-square value is greater than the threshold The points are regarded as dynamic feature points.

[0049] Keyframe detection strategy: In general SLAM algorithms, the insertion position of the keyframe is detected by predicting the change trend of the map points observed during tracking, ensuring that the keyframe is inserted before any problems occur in the tracking. At the same time, in actual applications, the image data stream may experience frame extraction due to hardware reasons, resulting in inconsistent time differences between frames. Using reprojection technology combined with the constant acceleration motion model of the SLAM system, by observing the change trend of the map points in the tracking frame and the frame before the tracking frame, the number of map points in the next frame of the current tracking frame is predicted to ensure that the keyframe is inserted before any problems occur in the tracking. Reprojection technology can flexibly convert real three-dimensional points and image planes, effectively showing the relationship between two adjacent frames, and further indicating spatial transformation and camera motion.

[0050] Assume that the current SLAM system motion is a uniformly accelerated motion, and its motion model is expressed as follows:

[0051]

[0052] in Express posture, Indicates the speed of movement, represents the acceleration of motion, Indicates the time interval, Indicates the Process noise of the frame.

[0053] When the SLAM system tracks the Frame, according to the motion model, by the Frame pose prediction Frame pose. Using reprojection technology, the Frame map points are projected onto On the image plane corresponding to the frame pose, the Frame map point coordinates are given by According to the reprojection technique, it is expressed as follows:

[0054]

[0055] in, Indicates the depth value of the map point in the i-th frame in the camera coordinate system, It represents the camera's pose transformation matrix under the k-th frame. Represents the camera internal parameters; 、 、 Represents the three-dimensional coordinate components of the map point in the world coordinate system;

[0056] The first Frame map points are projected onto On the frame, The coordinates of the map points on the frame are given by According to the above reprojection technology, the following expression is obtained:

[0057]

[0058] According to the formula, the map point is judged. The point that appears on the image plane is the Count the map points on the frame and predict the The number of map points in the frame, whether the map point is in the The judgment is made on the frame image plane, and the judgment condition is expressed as follows:

[0059]

[0060] in, Indicates the The number of map points in the frame; if the conditions are met, add the first In the frame map point set, the number of map points is finally counted.

[0061] Furthermore, the local mapping thread obtains high-quality key frames through a key frame selection strategy to improve the accuracy of map construction. The specific operations are:

[0062] When detecting new keyframes to be inserted, it is necessary to select frames with better quality, package them, and insert them into the keyframe processing queue. Based on the original keyframe judgment, the judgment of its own quality is added, and the number of dynamic feature points is used as the frame quality judgment standard. The fewer the number of dynamic feature points, the higher the quality of the frame. The number of dynamic feature points of the keyframe is obtained by comparing the number of feature points before and after removing the dynamic feature points, and the ratio is set. It is expressed as follows:

[0063]

[0064] at this time, is the number of feature points after removing dynamic feature points, is the number of all feature points before removing dynamic feature points, Indicates the proportion of dynamic feature points to all feature points. According to experience, set the adaptive threshold. When it is between 0.2 and 0.5, the key frame quality is good and can be used in subsequent calculations.

[0065] At this time, the keyframe detection strategy is: (1) The current frame is more than 0.25s away from the previous keyframe; (2) The local mapping thread is waiting for the insertion of a new keyframe; (3) The number of valid map points observed by the frame to be tracked is less than 25% of the reference keyframe; (4) The map points tracked by the frame to be tracked are less than 75% of the reference keyframe; (5) In monocular / binocular / RGBD+IMU mode, and in the IMU fixed time period, the time between the current frame and the previous keyframe exceeds 0.5 seconds; (6) In monocular+IMU mode, the number of matched points in the current frame is between 15 and 75 or is in the RECENTLY_LOST state; (7) The number of near points in the map of the next frame predicted is less than 100; When (4) is satisfied and any one of (1), (2), (3) is satisfied, or when (5), (6), (7) are satisfied alone, a new keyframe is inserted to ensure smooth tracking.

[0066] Furthermore, the loop detection thread roughly matches historical key frames through the bag-of-words model and confirms the loop through geometric verification and precise estimation. The specific operations are:

[0067] Through a two-stage process of coarse bag-of-words matching and fine estimation with geometric verification, efficient and accurate scene re-identification is achieved. First, the current frame image features are quantified into word frequency vectors using an offline visual dictionary, allowing for rapid retrieval of similar historical keyframes. Subsequently, multi-level geometric verification is performed to eliminate mismatches, ultimately triggering pose graph optimization to correct accumulated errors, significantly improving the long-term stability of the SLAM system.

[0068] Bag-of-words model rough matching: Hierarchical K-means clustering is used to transform ORB feature descriptors Divided into visual dictionary trees, the first layer performs K-means clustering to obtain , each layer continues to subdivide until the leaf node, and finally generates a dictionary containing M visual words , when searching online, calculate the current frame With candidate frames The word frequency vector similarity is as follows:

[0069]

[0070] in, , Represent word discrimination and accelerate the selection of Top-K candidate frames through inverted index;

[0071] Geometric verification of the precise estimate, using two-level constraints:

[0072] 2D-2D epipolar geometry verification, solving the fundamental matrix through RANSAC iteration , the minimum Sampson distance is expressed as follows:

[0073] in, and Represents a pair of matching points in two different images that satisfy the condition The points of are interior points, and the point set with the most interior points is found;

[0074] 3D-2D pose verification, pose solution through EPnP , and check the reprojection error as shown below:

[0075]

[0076] in, Represents the coordinates of the three-dimensional space in the world coordinate system, express Corresponding to the two-dimensional point coordinates on the image; screening meets The final loop frame is added to the pose graph for global optimization.

[0077] Furthermore, the mapping thread removes dynamic targets on key frames through the combined results of target detection and geometric constraints, and only retains information related to static objects to obtain a dense map. The specific operations are:

[0078] After determining the keyframes and the time to insert them, the pose information and depth information are passed to the dense point cloud construction thread. In order to prevent ghosting, dynamic targets on the keyframes are removed through the combined results of target detection and geometric constraints, and only information related to static objects is passed in. First, a set of static keyframes filtered by the visual SLAM system is generated. , using multi-view geometric constraints and depth estimation to generate the initial point cloud ; Then, the point cloud quality is optimized through a multi-level filtering algorithm, and finally a TSDF fusion algorithm is used to construct a dense map, which significantly improves the geometric accuracy of the map while ensuring real-time performance.

[0079] In the point cloud generation stage, the SLAM system selects the points that meet the common view condition. The key frame pair is expressed as follows by calculating the disparity using the semi-global matching algorithm:

[0080]

[0081] in To match the cost, , is the penalty coefficient; using the camera projection model Convert the disparity map into a 3D point cloud, where represents the projection function, is the camera pose;

[0082] After that, the system performs voxel grid filtering, statistical outlier removal, and bilateral filtering in sequence, specifically:

[0083] Voxel grid filtering: Divide the space into a voxel grid with a side length of l, and retain the centroid of each voxel ;

[0084] Statistical outlier removal: Calculate the average distance of the k neighbors of point p , eliminating the satisfied points, where , are the mean and standard deviation of neighborhood distances, respectively;

[0085] Bilateral filtering: Preserving edge sharpness while reducing noise:

[0086] , in , are the spatial and normal weight functions respectively;

[0087] Finally, a dense map with a complete topological structure is generated by TSDF fusion and Poisson reconstruction. Specifically, the multi-frame observations are weighted and fused into a global truncated signed distance field, which is expressed as follows:

[0088]

[0089] in represents the truncated signed distance, when Appear in the field of vision, satisfied , in other cases ; represents the weight function, when hour, ; In other cases ;

[0090] The TSDF value and cumulative weight of each frame observation are dynamically updated as follows:

[0091]

[0092]

[0093] in, Indicates that after the current frame observation, point The updated truncated signed distance function (TSDF) value, Indicates that after the current frame observation, the point The cumulative weight of Indicates the time point of the previous frame The truncated signed distance function (TSDF) value of Indicates the previous frame time point The cumulative weight reflects the cumulative contribution of the previous frame and all previous frames to the observation of this point;

[0094] Finally, the Poisson equation is solved by Poisson reconstruction , generating a continuous surface.

[0095] Compared with the prior art, the present invention has the following advantages:

[0096] The method of the present invention is applicable to visual SLAM systems in dynamic scenes. By constructing a multi-branch feature extraction target detection network, and under the premise of filtering random motion noise interference, a dynamic target elimination mechanism is modeled based on semantic-geometric consistency joint verification to eliminate dynamic object interference. At the same time, a multi-branch feature extraction network based on deep learning is used to accurately identify dynamic targets. The camera motion trajectory prediction model is combined to actively predict the timing of key frame insertion. A key frame quality evaluation system is constructed through multi-dimensional indicators such as the number of dynamic feature points and feature distribution uniformity. Sensor data is comprehensively processed at multiple levels, significantly improving the robustness and reconstruction accuracy of the SLAM system in dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 It is the overall flow chart of the present invention;

[0098] Figure 2 It is a network structure diagram of the multi-branch feature extraction module MBFE in the target detection model of the present invention;

[0099] Figure 3 is a schematic diagram of a pinhole camera model of the present invention;

[0100] Figure 4 is a schematic diagram of the relationship between the image coordinate system and the pixel coordinate system of the present invention;

[0101] Figure 5 This is a schematic diagram of geometric information removal of dynamic feature points in the present invention;

[0102] Figure 6 Schematic diagram of the reprojection technique of the present invention. DETAILED DESCRIPTION

[0103] To gain a deeper understanding of the present invention, we will provide a comprehensive and detailed description thereof. However, the present invention has various implementations and is not limited to the specific examples listed herein. These examples are presented to enhance a comprehensive understanding of the present disclosure.

[0104] A dynamic scene SLAM method focusing on fine targets, the method comprising the following steps:

[0105] Step 1: Acquire dynamic scene data through sensors and preprocess it to obtain depth information and image information, and train the target detection model at the same time;

[0106] Step 2: Input the dynamic scene data pre-processed in step 1 into the trained target detection model and enter the visual SLAM process, which includes a target detection thread, a tracking thread, a local mapping thread, a loop detection thread, and a mapping thread;

[0107] The target detection thread focuses on fine target information that is easily lost, and uses a multi-branch feature extraction module to balance fine target and context information;

[0108] The tracking thread integrates the semantic information and geometric information transmitted by the target detection thread to jointly eliminate dynamic feature points, and determines the timing of key frame insertion through the key frame detection strategy to improve positioning accuracy;

[0109] The local mapping thread obtains high-quality key frames through a key frame selection strategy to improve the accuracy of map construction;

[0110] The loop detection thread roughly matches historical key frames through the bag-of-words model and confirms the loop through geometric verification and precise estimation;

[0111] The mapping thread removes dynamic targets on key frames through the combined results of target detection and geometric constraints, and only retains information related to static objects to obtain a dense map.

[0112] Furthermore, the target detection thread focuses on fine target information that is easily lost, and uses a multi-branch feature extraction module to balance fine target and context information. The specific operations are as follows:

[0113] The target detection network adds a multi-branch feature extraction module MBFE after the backbone network EfficientRep. After the backbone network extracts coarse features, the multi-branch feature extraction module MBFE extracts fine features from local, global and sequential convolution branches, balancing fine targets and contextual information, thereby improving the accuracy of fine target detection;

[0114] The multi-branch feature extraction module MBFE first expands and reconstructs the convolved feature tensor Split into a set of spatially continuous patches, whose shape is , by controlling the patch parameters The size of the patch parameter is used to distinguish between local and global branches. It is achieved by aggregating and shifting non-overlapping patches in the spatial dimension; then the channel averaging operation is performed, and its shape becomes ; And use FFN (Feed-Forward Network) linear calculation, then use the activation function to obtain the probability distribution in the spatial dimension and adjust the weight accordingly;

[0115] In addition, the multi-branch feature extraction module MBFE uses the self-attention module to achieve adaptive data enhancement. The attention module contains a series of efficient channel attention and spatial attention components, which effectively improves the efficiency of the network and reduces the loss of fine target information.

[0116] Furthermore, the tracking thread integrates the semantic information and geometric information transmitted by the target detection thread to jointly eliminate dynamic feature points, and determines the timing of key frame insertion through the key frame detection strategy to improve the positioning accuracy. The specific operations are as follows:

[0117] Semantic information provides prior knowledge for epipolar geometry, which in turn provides key support for semantic information. After matching the feature points of the two frames, the distance between the feature points of the current frame and their corresponding epipolar lines is calculated based on the basic matrix. The larger the distance, the greater the possibility that the feature point is a dynamic point:

[0118] Dynamic feature point elimination strategy: The camera is an important sensor in visual SLAM. The camera coordinate system and the image coordinate system are converted through the classic pinhole camera model to provide strong data support for the SLAM system. The pinhole camera model is based on the principle of pinhole imaging. It shines light through a pinhole onto a plane, and uses light to project a point in three-dimensional space onto a two-dimensional plane to form an image. The pinhole model simply explains the process of camera imaging. First, the pinhole camera model is modeled, where For the heart of light, is the center of the imaging plane, and the focal length of the camera is , is a three-dimensional point in space, assuming in the camera coordinate system The coordinates are , Its coordinates in the image coordinate system are .

[0119] Then in the camera model, the principle of similar triangles based on pinhole imaging is as follows:

[0120]

[0121] After arranging the camera imaging plane The coordinates are expressed as follows:

[0122]

[0123] Convert the image coordinates to pixel coordinates, and coordinate Convert to pixel coordinates in pixel coordinate system The image coordinate system and the pixel coordinate system are transformed by a certain scaling and translation. is the origin of the pixel coordinate system, and the coordinate axes are , assuming that On-axis scaling times, in Axis scaled times, the origin is shifted , so the pixel coordinate system is expressed as follows:

[0124]

[0125] Substitute into the above formula, express ,Will express , and the following expression is obtained:

[0126]

[0127] Further arrangement is expressed as follows:

[0128] Usually the intermediate matrix It is called the camera's internal parameters, which are set directly by the factory and cannot be modified later. In addition to the internal parameters, the camera also has corresponding external parameters. The external parameters are represented by the rotation and translation of the camera. It represents the rotation matrix and translation vectors , then the pinhole camera model can be expressed as:

[0129] in represents the external parameters of the camera.

[0130] Select images based on the pinhole camera imaging model, and Represents two adjacent frames of images, called image plane; and represents the center of the camera, and For and The corresponding feature points in is a three-dimensional point in space; 、 and The plane formed is called the polar plane. and Line and image plane and The intersection of and is the pole, the polar plane and the image plane 、 Intersecting lines 、 For the polar line, use 、 Represents a three-dimensional point in space Projection to the image plane The feature points may be On the line connecting the two, corresponding to the image plane On the polar line On, when is a static feature point, which is on the image plane The projection point on the epipolar line On, and Moved, its image plane Projection points and epilines on In the actual observation process, there may be noise. If it is a static feature point, like a plane The distance between the feature point and the epipolar line is less than the threshold , set the feature point and The pixel homogeneous coordinates of are: , ,in , is the coordinate value of the feature point in the image pixel coordinate system, and the image plane is calculated by the basic matrix F The polar line is represented as follows:

[0131]

[0132] Among them, the parameters , , Represent the three different coefficients of the polar equation in the image homogeneous coordinate representation, and Together they determine the direction of the epipolar line and also affect its position in the image plane. It plays the role of translation in the epipolar equation, determining the position offset of the epipolar line relative to the coordinate origin in the image plane; Indicates transposing the matrix;

[0133] The polar equation is expressed as:

[0134]

[0135] The epipolar geometry constraint is expressed as:

[0136]

[0137] Feature Points To the pole The distance is expressed as:

[0138]

[0139] Ideally, static feature points should fall on the epipolar line. The value is 0. In actual operation, it is difficult to avoid the appearance of noise, so when The value is greater than 0 but less than the empirical threshold When the feature point is a static feature point, if the object moves along the polar line, the distance between the feature point and the polar line is It may also be greater than 0 and less than the empirical threshold , in this case, the point will be mistakenly considered to be a static feature point. In order to avoid such errors, the present invention considers adding depth projection error as a constraint condition.

[0140] Feature points of the current frame According to the pinhole camera motion, project it to the previous key frame and get the projection point and projection depth . Get the previous keyframe projection point directly from the depth image The true depth The depth projection error is expressed as Ideally, for a static feature point, the depth projection error should be close to zero.

[0141] Construct error vector by combining epipolar geometry constraint error and depth projection error . Assume that these errors are independently distributed and obey Gaussian distribution.

[0142] Inspired by ORB-SLAM2, the covariance of the features extracted at the nth level of the pyramid is expressed as:

[0143]

[0144] Among them, S is the reduced scale of each layer of the image pyramid, assuming that the standard deviation of the 0th layer is Take 1. Since the error vector E is a 2-dimensional vector, setting the threshold will cause problems. The inner product r of the calculated vector is converted into a scalar to solve this problem. So the feature point Chi-square value of Expressed as:

[0145]

[0146] in for point The present invention uses RGB-D images, where a binocular camera is simulated, with a degree of freedom of 3 and a threshold value set to The chi-square value is greater than the threshold The points are regarded as dynamic feature points.

[0147] Keyframe detection strategy: In general SLAM algorithms, the insertion position of the keyframe is detected by predicting the change trend of the map points observed during tracking, ensuring that the keyframe is inserted before any problems occur in the tracking. At the same time, in actual applications, the image data stream may experience frame extraction due to hardware reasons, resulting in inconsistent time differences between frames. Using reprojection technology combined with the constant acceleration motion model of the SLAM system, by observing the change trend of the map points in the tracking frame and the frame before the tracking frame, the number of map points in the next frame of the current tracking frame is predicted to ensure that the keyframe is inserted before any problems occur in the tracking. Reprojection technology can flexibly convert real three-dimensional points and image planes, effectively showing the relationship between two adjacent frames, and further indicating spatial transformation and camera motion.

[0148] Assume that the current SLAM system motion is a uniformly accelerated motion, and its motion model is expressed as follows:

[0149]

[0150] in Express posture, Indicates the speed of movement, represents the acceleration of motion, Indicates the time interval, Indicates the Process noise of the frame.

[0151] When the SLAM system tracks the Frame, according to the motion model, by the Frame pose prediction Frame pose. Using reprojection technology, the Frame map points are projected onto On the image plane corresponding to the frame pose, the Frame map point coordinates are given by According to the reprojection technique, it is expressed as follows:

[0152]

[0153] in, Indicates the depth value of the map point in the i-th frame in the camera coordinate system, It represents the camera's pose transformation matrix under the k-th frame. Represents the camera internal parameters; 、 、 Represents the three-dimensional coordinate components of the map point in the world coordinate system;

[0154] The first Frame map points are projected onto On the frame, The coordinates of the map points on the frame are given by According to the above reprojection technology, the following expression is obtained:

[0155]

[0156] According to the formula, the map point is judged. The point that appears on the image plane is the Count the map points on the frame and predict the The number of map points in the frame, whether the map point is in the The judgment is made on the frame image plane, and the judgment condition is expressed as follows:

[0157]

[0158] in, Indicates the The number of map points in the frame; if the conditions are met, add the first In the frame map point set, the number of map points is finally counted.

[0159] Furthermore, the local mapping thread obtains high-quality key frames through a key frame selection strategy to improve the accuracy of map construction. The specific operations are:

[0160] When detecting new keyframes to be inserted, it is necessary to select frames with better quality, package them, and insert them into the keyframe processing queue. Based on the original keyframe judgment, the judgment of its own quality is added, and the number of dynamic feature points is used as the frame quality judgment standard. The fewer the number of dynamic feature points, the higher the quality of the frame. The number of dynamic feature points of the keyframe is obtained by comparing the number of feature points before and after removing the dynamic feature points, and the ratio is set. It is expressed as follows:

[0161]

[0162] at this time, is the number of feature points after removing dynamic feature points, is the number of all feature points before removing dynamic feature points, Indicates the proportion of dynamic feature points to all feature points. According to experience, set the adaptive threshold. When it is between 0.2 and 0.5, the key frame quality is good and can be used in subsequent calculations.

[0163] At this time, the keyframe detection strategy is: (1) The current frame is more than 0.25s away from the previous keyframe; (2) The local mapping thread is waiting for the insertion of a new keyframe; (3) The number of valid map points observed by the frame to be tracked is less than 25% of the reference keyframe; (4) The map points tracked by the frame to be tracked are less than 75% of the reference keyframe; (5) In monocular / binocular / RGBD+IMU mode, and in the IMU fixed time period, the time between the current frame and the previous keyframe exceeds 0.5 seconds; (6) In monocular+IMU mode, the number of matched points in the current frame is between 15 and 75 or is in the RECENTLY_LOST state; (7) The number of near points in the map of the next frame predicted is less than 100; When (4) is satisfied and any one of (1), (2), (3) is satisfied, or when (5), (6), (7) are satisfied alone, a new keyframe is inserted to ensure smooth tracking.

[0164] Furthermore, the loop detection thread roughly matches historical key frames through the bag-of-words model and confirms the loop through geometric verification and precise estimation. The specific operations are:

[0165] Through a two-stage process of coarse bag-of-words matching and fine estimation with geometric verification, efficient and accurate scene re-identification is achieved. First, the current frame image features are quantified into word frequency vectors using an offline visual dictionary, allowing for rapid retrieval of similar historical keyframes. Subsequently, multi-level geometric verification is performed to eliminate mismatches, ultimately triggering pose graph optimization to correct accumulated errors, significantly improving the long-term stability of the SLAM system.

[0166] Bag-of-words model rough matching: Hierarchical K-means clustering is used to transform ORB feature descriptors Divided into visual dictionary trees, the first layer performs K-means clustering to obtain , each layer continues to subdivide until the leaf node, and finally generates a dictionary containing M visual words , when searching online, calculate the current frame With candidate frames The word frequency vector similarity is as follows:

[0167]

[0168] in, , Represent word discrimination and accelerate the selection of Top-K candidate frames through inverted index;

[0169] Geometric verification of the precise estimate, using two-level constraints:

[0170] 2D-2D epipolar geometry verification, solving the fundamental matrix through RANSAC iteration , the minimum Sampson distance is expressed as follows:

[0171]

[0172] in, and Represents a pair of matching points in two different images that satisfy the condition The points of are interior points, and the point set with the most interior points is found;

[0173] 3D-2D pose verification, pose solution through EPnP , and check the reprojection error as shown below:

[0174] in, Represents the coordinates of the three-dimensional space in the world coordinate system, express Corresponding to the two-dimensional point coordinates on the image; screening meets The final loop frame is added to the pose graph for global optimization.

[0175] Furthermore, the mapping thread removes dynamic targets on key frames through the combined results of target detection and geometric constraints, and only retains information related to static objects to obtain a dense map. The specific operations are:

[0176] After determining the keyframes and the time to insert them, the pose information and depth information are passed to the dense point cloud construction thread. In order to prevent ghosting, dynamic targets on the keyframes are removed through the combined results of target detection and geometric constraints, and only information related to static objects is passed in. First, a set of static keyframes filtered by the visual SLAM system is generated. , using multi-view geometric constraints and depth estimation to generate the initial point cloud ; Then, the point cloud quality is optimized through a multi-level filtering algorithm, and finally a TSDF fusion algorithm is used to construct a dense map, which significantly improves the geometric accuracy of the map while ensuring real-time performance.

[0177] In the point cloud generation stage, the SLAM system selects the points that meet the common view condition. The key frame pair is expressed as follows by calculating the disparity using the semi-global matching algorithm:

[0178]

[0179] in To match the cost, , is the penalty coefficient; using the camera projection model Convert the disparity map into a 3D point cloud, where represents the projection function, is the camera pose;

[0180] After that, the system performs voxel grid filtering, statistical outlier removal, and bilateral filtering in sequence, specifically:

[0181] Voxel grid filtering: Divide the space into a voxel grid with a side length of l, and retain the centroid of each voxel ;

[0182] Statistical outlier removal: Calculate the average distance of the k neighbors of point p , eliminating the satisfied points, where , are the mean and standard deviation of neighborhood distances, respectively;

[0183] Bilateral filtering: Preserving edge sharpness while reducing noise:

[0184] , in , are the spatial and normal weight functions respectively;

[0185] Finally, a dense map with a complete topological structure is generated by TSDF fusion and Poisson reconstruction. Specifically, the multi-frame observations are weighted and fused into a global truncated signed distance field, which is expressed as follows:

[0186]

[0187] in represents the truncated signed distance, when Appear in the field of vision, satisfied , in other cases ; represents the weight function, when hour, ; In other cases ;

[0188] The TSDF value and cumulative weight of each frame observation are dynamically updated as follows:

[0189]

[0190]

[0191] in, Indicates that after the current frame observation, point The updated truncated signed distance function (TSDF) value, Indicates that after the current frame observation, the point The cumulative weight of Indicates the time point of the previous frame The truncated signed distance function (TSDF) value of Indicates the previous frame time point The cumulative weight reflects the cumulative contribution of the previous frame and all previous frames to the observation of this point;

[0192] Finally, the Poisson equation is solved by Poisson reconstruction , generating a continuous surface.

[0193] The present invention and other established SLAM algorithms were used to process the TUM dataset and the results were compared. The absolute trajectory error and the relative trajectory error representing the rotation and translation directions of the SLAM algorithm were used as evaluation metrics, and their root mean square error and standard deviation were calculated to assess the algorithm's accuracy. The experimental results comparing DS-SLAM, SG-SLAM, and ORB-SLAM3 with the present invention using different sequences in the TUM dataset, fr3_walking_xyz, fr3_walking_static, fr3_walking_half, fr3_sitting_xyz, and fr3_sitting_static, are shown in Tables 1, 2, and 3 below. Table 1 shows the relative trajectory error values ​​in the rotation direction, Table 2 shows the relative trajectory error values ​​in the translation direction, and Table 3 shows the absolute trajectory error values.

[0194] Table 1 Rotational RPE

[0195]

[0196] Table 2 Translational RPE

[0197]

[0198] Table 3 ATE

[0199]

[0200] The three tables show that the proposed SLAM algorithm has lower root mean square error and variance than other algorithms. Except when processing relatively static sequences such as the fr3_walking_static sequence, the translational relative trajectory error is slightly inferior to the classic static ORB-SLAM3, but it still outperforms currently established SLAM systems for handling dynamic scenes. Overall, the proposed algorithm has lower error when handling fine targets in dynamic scenes, resulting in higher positioning and mapping accuracy for the SLAM system.

[0201] Any matters not described in detail in this specification are prior art known to those skilled in the art. Although the above description of the present invention is based on specific embodiments to facilitate understanding of the present invention by those skilled in the art, it should be understood that the present invention is not limited to the scope of the specific embodiments. As long as various modifications are within the spirit and scope of the present invention as defined and determined by the appended claims, such modifications will be obvious to those skilled in the art, and all inventions and creations utilizing the concepts of the present invention are protected.

Claims

1. A dynamic scene SLAM method focusing on fine targets, characterized in that: The method comprises the following steps: Step 1: Acquire dynamic scene data through sensors and preprocess it to obtain depth information and image information, and train the target detection model at the same time; Step 2: Input the dynamic scene data pre-processed in step 1 into the trained target detection network and enter the visual SLAM process, which includes a target detection thread, a tracking thread, a local mapping thread, a loop detection thread, and a mapping thread; The target detection thread focuses on fine target information that is easily lost, and uses a multi-branch feature extraction module to balance fine target and context information; The tracking thread integrates the semantic information and geometric information transmitted by the target detection thread, uses the pinhole camera model to jointly eliminate dynamic feature points, and uses the key frame detection strategy to determine the timing of key frame insertion to improve positioning accuracy; The local mapping thread obtains high-quality key frames through a key frame selection strategy to improve the accuracy of map construction; The loop detection thread roughly matches historical key frames through the bag-of-words model and confirms the loop through geometric verification and precise estimation; The mapping thread removes dynamic targets on key frames through the combined results of target detection and geometric constraints, and only retains information related to static objects to obtain a dense map.

2. A dynamic scene SLAM method focusing on fine targets according to claim 1, characterized in that, The target detection thread focuses on fine target information that is easily lost, and uses a multi-branch feature extraction module to balance fine target and context information. The specific operations are as follows: The target detection network adds a multi-branch feature extraction module MBFE after the backbone network EfficientRep. After the backbone network extracts coarse features, the multi-branch feature extraction module MBFE extracts fine features from local, global and sequential convolution branches, balancing fine targets and contextual information, and improving the accuracy of fine target detection. The multi-branch feature extraction module MBFE first expands and reconstructs the convolved feature tensor Split into a set of spatially continuous patches, whose shape is , by controlling the patch parameters The size of the patch parameter is used to distinguish between local and global branches. It is achieved by aggregating and shifting non-overlapping patches in the spatial dimension; then the channel averaging operation is performed, and its shape becomes ; and use FFN linear calculation, then use the activation function to obtain the probability distribution in the spatial dimension and adjust the weights accordingly; In addition, the multi-branch feature extraction module MBFE uses the self-attention module to achieve adaptive data enhancement. The self-attention module contains channel attention and spatial attention components to improve the efficiency of the network, reduce the loss of fine target information, and achieve a balance between fine targets and contextual information.

3. A dynamic scene SLAM method focusing on fine targets according to claim 2, characterized in that, The tracking thread integrates the semantic information and geometric information from the target detection thread, uses the pinhole camera model to jointly eliminate dynamic feature points, and uses the key frame detection strategy to determine the timing of key frame insertion to improve positioning accuracy. The specific operations are as follows: Dynamic feature point removal strategy: Select images based on the pinhole camera model, and Represents two adjacent frames of images, called image plane; and represents the center of the camera, and For and The corresponding feature points in is a three-dimensional point in space; 、 and The plane formed is called the polar plane. and Line and image plane and The intersection of and is the pole, the polar plane and the image plane 、 Intersecting lines 、 For the polar line, use 、 Represents a three-dimensional point in space Projection to the image plane The feature points may be On the line connecting the two, corresponding to the image plane On the polar line On, when is a static feature point, which is on the image plane The projection point on the epipolar line On, and Moved, its image plane Projection points and epilines on A distance is generated on the plane. When it is a static feature point, the image plane The distance between the feature point and the epipolar line is less than the threshold , set the feature point and The pixel homogeneous coordinates of are: , ,in , is the coordinate value of the feature point in the image pixel coordinate system, and the image plane is calculated by the basic matrix F The polar line is represented as follows: , where the parameter , , Represent the three different coefficients of the polar equation in the image homogeneous coordinate representation, and Together they determine the direction of the epipolar line and also affect its position in the image plane. It plays the role of translation in the epipolar equation, determining the position offset of the epipolar line relative to the coordinate origin in the image plane; Indicates transposing the matrix; The polar equation is expressed as: , the epipolar geometry constraint is expressed as: , feature points To the pole The distance is expressed as: , the feature points of the current frame According to the pinhole camera motion, project it to the previous key frame and get the projection point and projection depth ; Get the previous keyframe projection point directly from the depth image obtained by the sensor The true depth ; The depth projection error is expressed as ; Construct error vector by combining epipolar geometry constraint error and depth projection error , assuming that the error vectors are independently distributed and obey Gaussian distribution; Key frame detection strategy: Assume that the current SLAM system motion is a uniformly accelerated motion, and its motion model is expressed as follows: ,in Indicates the position of the movement, Indicates the speed of movement, represents the acceleration of motion, Indicates the time interval, Indicates the The process noise of the frame; when the SLAM system tracks the Frame, according to the motion model, by the Frame pose prediction Frame pose, using reprojection technology, the first Frame map points are projected onto On the image plane corresponding to the frame pose, the Frame map point coordinates are given by According to the reprojection technique, it is expressed as follows: ,in, Indicates the depth value of the map point in the i-th frame in the camera coordinate system, It represents the camera's pose transformation matrix under the k-th frame. Represents the camera internal parameters; 、 、 Represents the three-dimensional coordinate components of the map point in the world coordinate system; The first Frame map points are projected onto On the frame, The coordinates of the map points on the frame are given by According to the above reprojection technology, the following expression is obtained: , judge the map point according to the formula, and the point that appears on the image plane is Count the map points on the frame and predict the The number of map points in the frame, whether the map point is in the The judgment is made on the frame image plane, and the judgment condition is expressed as follows: ,in, Indicates the The number of map points in the frame; if the conditions are met, add the first In the frame map point set, the number of map points is finally counted.

4. A dynamic scene SLAM method focusing on fine targets according to claim 3, characterized in that, The local mapping thread obtains high-quality key frames through the key frame selection strategy to improve the accuracy of map construction. The specific operations are: The number of dynamic feature points of the key frame is obtained by comparing the number of feature points before and after removing the dynamic feature points, and the ratio is set. It is expressed as follows: ,in, is the number of feature points after removing dynamic feature points, is the number of all feature points before removing dynamic feature points, Indicates the ratio of dynamic feature points to all feature points.

5. A dynamic scene SLAM method focusing on fine targets according to claim 4, characterized in that, The keyframe detection strategy is as follows: (1) the current frame is more than 0.25s away from the previous keyframe; (2) the local mapping thread is waiting for the insertion of a new keyframe; (3) the number of valid map points observed by the frame to be tracked is less than 25% of the reference keyframe; (4) the map points tracked by the frame to be tracked are less than 75% of the reference keyframe; (5) in monocular / binocular / RGBD+IMU mode, and in the IMU fixed time period, the time between the current frame and the previous keyframe exceeds 0.5 seconds; (6) in monocular+IMU mode, the number of matched points in the current frame is between 15 and 75 or is in the RECENTLY_LOST state; (7) the number of near points in the map of the next frame predicted is less than 100; when (4) is satisfied and any one of (1), (2), (3) is satisfied, or when (5), (6), (7) are satisfied alone, a new keyframe is inserted to ensure smooth tracking.

6. A dynamic scene SLAM method focusing on fine targets according to claim 5, characterized in that, The loop detection thread roughly matches historical key frames through the bag-of-words model and confirms the loop through geometric verification and precise estimation. The specific operations are as follows: Bag-of-words model rough matching: Hierarchical K-means clustering is used to transform ORB feature descriptors Divided into visual dictionary trees, the first layer performs K-means clustering to obtain , each layer continues to subdivide until the leaf node, and finally generates a dictionary containing M visual words , when searching online, calculate the current frame With candidate frames The word frequency vector similarity is as follows: ,in, , Represent word discrimination and accelerate the selection of Top-K candidate frames through inverted index; Geometric verification of the precise estimate, using two-level constraints: 2D-2D epipolar geometry verification, solving the fundamental matrix through RANSAC iteration , the minimum Sampson distance is expressed as follows: ,in, and Represents a pair of matching points in two different images that satisfy the condition The points of are interior points, and the point set with the most interior points is found; 3D-2D pose verification, pose solution through EPnP , and check the reprojection error as shown below: ,in, Represents the coordinates of the three-dimensional space in the world coordinate system, express Corresponding to the two-dimensional point coordinates on the image; screening meets The final loop frame is added to the pose graph for global optimization.

7. A dynamic scene SLAM method focusing on fine targets according to claim 6, characterized in that, The mapping thread removes dynamic targets on key frames through the combined results of target detection and geometric constraints, and only retains information related to static objects to obtain a dense map. The specific operations are: First, based on the static key frame set selected by the visual SLAM system , using multi-view geometric constraints and depth estimation to generate the initial point cloud ; Then, a multi-level filtering algorithm is used to optimize the point cloud quality, and finally a TSDF fusion algorithm is used to construct a dense map; In the point cloud generation stage, the SLAM system selects the points that meet the common view condition. The key frame pair is expressed as follows by calculating the disparity using the semi-global matching algorithm: ,in To match the cost, , is the penalty coefficient; using the camera projection model Convert the disparity map into a 3D point cloud, where represents the projection function, is the camera pose; After that, the system performs voxel grid filtering, statistical outlier removal, and bilateral filtering in sequence, specifically: Voxel grid filtering: Divide the space into a voxel grid with a side length of l, and retain the centroid of each voxel ; Statistical outlier removal: Calculate the average distance of the k neighbors of point p , eliminating the satisfied points, where , are the mean and standard deviation of neighborhood distances, respectively; Bilateral filtering: Preserving edge sharpness while reducing noise: , in , are the spatial and normal weight functions respectively; Finally, a dense map with a complete topological structure is generated by fusion of truncated signed distance function TSDF and Poisson reconstruction. Specifically, the multi-frame observations are weighted and fused into a global truncated signed distance field, which is expressed as follows: ,in represents the truncated signed distance, when Appear in the field of vision, satisfied , in other cases ; represents the weight function, when hour, ; In other cases ; The TSDF value and cumulative weight of each frame observation are dynamically updated as follows: ,, ,in, Indicates that after the current frame observation, point The updated truncated signed distance function TSDF value, Indicates that after the current frame observation, the point The cumulative weight of Indicates the time point of the previous frame The truncated signed distance function TSDF value; Indicates the previous frame time point The cumulative weight reflects the cumulative contribution of the previous frame and all previous frames to the observation of this point; Finally, the Poisson equation is solved by Poisson reconstruction , generating a continuous surface.

Citation Information

Patent Citations

  • Visual inertia indoor dynamic environment positioning system integrating attention target detection and geometric constraint

    CN116124144A

  • Visual SLAM method suitable for dynamic environment

    CN116429087A

  • Visual SLAM method based on dynamic target tracking and feature point filtering

    CN116977408A

  • RGB-D visual SLAM method for indoor dynamic scene

    CN119295721A

  • Dynamic environment dense point cloud SLAM method and system based on YOLOv11 and ORB-SLAM3

    CN119540942A