SLAM dynamic point semantic filtering method based on DDMA-SAM

By employing a semantic filtering method based on DDMA-SAM, which combines semantic and geometric constraints, the robustness and efficiency issues of visual SLAM in dynamic environments are addressed. This approach achieves high-precision dynamic point removal and pose estimation, thereby improving the system's positioning accuracy and stability.

CN120931922APending Publication Date: 2025-11-11BEIJING INST OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511030198.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing visual SLAM algorithms are not robust to fast-moving targets in dynamic environments, and are limited by the huge computational cost of semantic segmentation models, making it difficult to simultaneously achieve system accuracy and operational efficiency.

Method used

A semantic filtering method based on DDMA-SAM is adopted. By combining structural compression and multi-scale feature aggregation with semantic and geometric constraints, a dynamic point removal mechanism is constructed, including a decoupled distillation strategy, a lightweight adaptive extraction module, and a multi-path diversion feature matching module, to achieve efficient dynamic point removal.

Benefits of technology

It significantly improves the localization robustness and practicality of the visual SLAM system in dynamic environments, reduces the number of parameters, improves segmentation accuracy and inference speed, and enhances the system's adaptability and real-time processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931922A_ABST
    Figure CN120931922A_ABST
Patent Text Reader

Abstract

The invention discloses an SLAM dynamic point semantic filtering method based on DDMA-SAM, and belongs to the technical field of synchronous positioning and mapping. A decoupling distillation mechanism is introduced, an image encoder in an original SAM model is subjected to lightweight optimization, a DDMA-SAM semantic network integrating a multi-scale aggregation detection module and an efficient mask decoding module is constructed, and the SLAM dynamic point semantic filtering method based on DDMA-SAM is obtained. And the segmentation performance is improved while the model parameters are greatly compressed. Based on the semantic network, providing a semantic prior and geometric consistency combined-driven double filtering strategy; based on a semantic mask and a confidence threshold, carrying out preliminary dynamic point identification; in combination with the epipolar geometric constraint and the triangulation reprojection error of random sampling consistency estimation, fine elimination of dynamic feature points is realized, and only static points are reserved to participate in camera pose estimation. According to the method, the real-time performance of the system is kept, and meanwhile, the mapping quality and the track stability in a dynamic scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of synchronous localization and mapping technology, specifically relating to a dynamic point semantic filtering method for SLAM based on DDMA-SAM. Background Technology

[0002] With the rapid growth in demand for autonomous environmental perception and high-precision positioning in fields such as autonomous driving, mobile robotics, and augmented reality, Simultaneous Localization and Mapping (SLAM) technology has become a core supporting capability in intelligent mobile systems. Visual SLAM, as one of the mainstream implementation forms, has been widely applied in complex scenarios such as urban autonomous driving, indoor navigation, and unmanned delivery due to its advantages such as low sensor cost and rich information acquisition.

[0003] However, most existing mainstream visual SLAM algorithms are based on the assumption of a "static environment," ignoring dynamic objects (such as pedestrians and vehicles) that are common in the real world. When dynamic targets are misidentified as static feature points and used in pose estimation, it can easily lead to feature drift and map error accumulation, resulting in a decrease in system positioning accuracy or even tracking failure, which severely limits its stable operation in real dynamic environments.

[0004] To address the aforementioned issues, researchers have proposed a dynamic point elimination strategy. However, existing technologies are not robust to fast-moving targets and are prone to false detections. Furthermore, due to the large computational overhead of semantic segmentation models and the difficulty in real-time deployment, it is still difficult to simultaneously balance system accuracy and operational efficiency.

[0005] In recent years, with the introduction of the general image segmentation model Segment Anything Model (SAM), image segmentation accuracy has been significantly improved. However, its large Transformer encoder structure limits its application on edge computing platforms. Therefore, there is an urgent need for a dynamic point removal method that simultaneously possesses lightweight, high segmentation accuracy, and real-time filtering capabilities to meet the robust operation requirements of SLAM systems in highly dynamic scenarios. Summary of the Invention

[0006] In view of this, the purpose of this invention is to provide a dynamic point semantic filtering method based on the DDMA-SAM semantic segmentation network. This method improves segmentation efficiency through structural compression and multi-scale feature aggregation, and constructs a high-precision dynamic point removal mechanism by combining semantic and geometric constraints, thereby effectively enhancing the localization robustness and practicality of the visual SLAM system in dynamic environments.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A dynamic point semantic filtering method for SLAM based on DDMA-SAM includes the following steps:

[0009] S1. Construct the DDMA-SAM semantic network: including the image encoder, cue encoding module, and mask decoding module in the original SAM model, and integrate the detection and segmentation modules; use a decoupled distillation strategy to compress the image encoder; the DDMA-SAM semantic network is used to segment the input image and output the semantic mask of dynamic targets in the image;

[0010] S2. Preliminary filtering of dynamic points based on semantic segmentation results: The pixel position of the feature point is compared with the semantic mask of the dynamic target in the image using the mask matching method. If the pixel position of the feature point falls within the semantic mask area of ​​the dynamic target, it is marked as a dynamic point and removed.

[0011] S3. Secondary dynamic point filtering based on geometric consistency constraints: For the feature points filtered in S2, dynamic points are further filtered out by combining epipolar geometric constraints to obtain static feature points that meet geometric and projection constraints.

[0012] As a further preferred embodiment of the present invention, it also includes S4, static feature point driven SLAM pose estimation and mapping: using static feature points after removing dynamic points to perform inter-frame matching, pose estimation and map construction.

[0013] As a further preferred embodiment of the present invention, the decoupled distillation strategy in S1 includes:

[0014] Decoupling distillation distills a lightweight image encoder from the original image encoder. It decomposes the traditional decoupling distillation loss into target class KL loss and non-target class KL loss. The target class KL loss is used to measure the difference between the student model's prediction on the target class and the teacher model, while the non-target class KL loss is used to measure the difference between the student model's prediction on the non-target class and the teacher model.

[0015] The total loss is defined by weighted combination of target class KL loss and non-target class KL loss as follows:

[0016] L total =α·L TCKD +β·L NCKD

[0017] In the formula, α and β are hyperparameters used to balance the loss contributions of the target class and non-target classes, and L TCKD For the target class KL loss, L NCKD For non-target class KL loss.

[0018] As a further preferred embodiment of the present invention, the detection and segmentation module in S1 includes: introducing a lightweight adaptive extraction (LAE) module and a multi-path splitting feature matching (MSFM) module to extract multi-scale features and perform feature aggregation, the generated prompt information is used as the prompt input of the SAM model, and then the SAM model is used for segmentation, and the semantic mask of the dynamic target in the image is output after segmentation.

[0019] As a further preferred embodiment of the present invention, the lightweight adaptive extraction LAE module performs the following operations: the input channels are divided into N groups for grouped convolution, each lightweight adaptive extraction unit performs a four-fold downsampling, after saving the height and width information to the channels, the dimension of the feature map is expanded from four dimensions to five dimensions, the adaptive extraction path achieves information exchange through average pooling and convolution operations, the feature map is recombined according to multiple pixels, and the weights of each are calculated through a normalized exponential function, while the dimension is converted to five dimensions.

[0020] As a further preferred embodiment of the present invention, the Multipath Flow Matching (MSFM) module performs the following operations:

[0021] The MSFM module extracts information from the input feature vector in three dimensions. Let the average pooling operation be P. avg Global average pooling operation P gavg From input features F in Extract output features F' out :

[0022] F' out =Res(Align(F h ,F w ,F c ))

[0023] In the formula: F h and F w The height and width information are calculated from the input features:

[0024]

[0025] in, The height of the input image. The width of the input image;

[0026] Height and width information are integrated into the channel dimension for global information fusion, thereby capturing the common features of the ROI region and its neighborhood. Spatial information is set as auxiliary weights and passed back to the spatial vector through multiplication. Let w be the weight corresponding to the width and height. h and w w , represented as:

[0027] w h ,ww =Split(Sigmoid(Concat(F h ,F w ))

[0028] Channel information is calculated through post-processing. for:

[0029]

[0030] The final output F is obtained by branching the channel information. out The calculation formula is as follows:

[0031]

[0032] The original features are concatenated and a 1×1 convolution is performed to obtain the output result.

[0033] As a further preferred embodiment of the present invention, S2 specifically includes the following steps:

[0034] The mask matching method is used to compare the pixel position of a feature point with the semantic mask of the dynamic target in the image. If the pixel position of a feature point falls within the semantic mask region of the dynamic target in the image, it is marked as a dynamic point and removed, forming the initial set of dynamic regions.

[0035] As a further preferred embodiment of the present invention, in S2, for each semantic mask region, a confidence threshold for dynamic objects is set based on the confidence score, and only feature points on dynamic objects with a confidence score higher than the confidence threshold are removed; the semantic mask region of dynamic targets in the output image is dilated to ensure that dynamic points in the edge region can also be identified.

[0036] As a further preferred embodiment of the present invention, S3 includes the following steps:

[0037] First, the essential matrix E and the fundamental matrix F are robustly estimated from the matched feature point pairs using the RANSAC algorithm. Then, the point pairs are judged to meet the geometric consistency by using epipolar geometric constraints, and dynamic points that do not meet the requirements are eliminated.

[0038] Next, the three-dimensional coordinates of the remaining feature points are recovered by triangulation to verify their spatial consistency;

[0039] Finally, based on the projection error, further detection is performed by calculating the reprojection error of the feature points and eliminating dynamic points whose errors exceed a set threshold, thereby retaining static feature points that conform to geometric and projection constraints.

[0040] As a further preferred embodiment of the present invention, S3 specifically includes the following steps:

[0041] First, the consistency of point pairs is determined by epipolar geometric constraints. Given the fundamental matrix F, corresponding points p1 and p2 in two adjacent frames should satisfy the epipolar constraints:

[0042]

[0043] If point pairs p1 and p2 satisfy this relationship, it means they may come from the same static point; otherwise, they are considered dynamic points and are removed.

[0044] Next, triangulation is used to recover the 3D points, and the essential matrix E is used to recover the camera pose between the two frames. Then, point pairs p1 and p2 are triangulated to recover point P in 3D space.

[0045] P=λ1K -1 p1=λ2K -1 (Rp2+t)

[0046] In the formula: K is the camera intrinsic parameter matrix, λ1 and λ2 are depth values, R is the camera rotation matrix, which represents the rotation relationship from the second frame camera coordinate system to the first frame coordinate system, and t is the camera translation vector, which represents the translation from the second frame coordinate system to the first frame coordinate system.

[0047] If the depth value of point P is positive after triangulation, it means that the point is in front of the camera and conforms to the geometric relationship; otherwise, point pair p1 and p2 are considered to be mismatched or from dynamic objects and are removed.

[0048] Finally, based on the projection error, the dynamic point is calculated, the triangulated 3D point P is projected back to the two frames of images, and the reprojection error is calculated:

[0049]

[0050] In the formula: and It is the point P projected onto the first and second frame images;

[0051] If the error ε1 or ε2 exceeds the set threshold, it indicates that the point pair comes from a dynamic object and should be discarded.

[0052] The beneficial effects of this invention are as follows:

[0053] 1) Balancing segmentation performance and inference efficiency: This invention employs a decoupled distillation mechanism to compress the encoder of the original SAM model and introduces a multi-scale aggregation detection module to construct a lightweight DDMA-SAM semantic segmentation network, effectively alleviating the problems of large parameter count and high computational latency in traditional models. Compared to the original SAM, the DDMA-SAM model reduces parameters by 98.2%, improves mIoU by 1.2%, and increases inference speed by 4.2 times, significantly enhancing its deployment adaptability and real-time processing capabilities on resource-constrained platforms.

[0054] 2) Higher accuracy and stronger robustness in dynamic point filtering: This method designs a dual dynamic point filtering strategy combining semantic prior and geometric consistency constraints. First, significant dynamic targets are removed using semantic segmentation masks. Then, geometric consistency detection is performed by estimating the fundamental matrix using RANSAC and applying epipolar constraints to further identify potential dynamic points. The combination of these two approaches effectively avoids both false and false removals. The dynamic point removal rate on the TUM Dynamic Objects dataset reaches 92.7%, significantly higher than the performance of single semantic (85.4%) or geometric methods (71.2%).

[0055] 3) More accurate and stable SLAM trajectory estimation: By filtering out dynamic points, only static points are retained for camera pose estimation and map construction, effectively reducing the impact of dynamic interference on feature matching and improving the overall robustness of the system. Experiments show that in highly dynamic scenes, the absolute trajectory error (ATE) and relative pose error (RPE) of the algorithm of this invention are reduced by 94.27% and 49.52% respectively compared with ORB-SLAM3, and by 22.26% and 12.57% respectively compared with DynaSLAM, demonstrating higher accuracy and trajectory stability in complex dynamic environments;

[0056] 4) Suitable for complex application scenarios such as high dynamics and edge computing: The DDMA-SAM semantic network has the characteristics of lightweight structure, sensitive feature perception and high operating efficiency. With the help of dual filtering strategy, it can be widely used in tasks such as mobile robots and autonomous driving that have high requirements for positioning accuracy and real-time computing, and has good engineering deployment value.

[0057] Other advantages, objectives, and features of the invention will be set forth in the following description and will be apparent to those skilled in the art in some respects, or may be learned by practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0058] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration:

[0059] Figure 1 This is a structural diagram of the overall method of the present invention;

[0060] Figure 2 This is a flowchart of the decoupled distillation process of the present invention;

[0061] Figure 3 This is an optimization of the algorithm structure based on the decoupled distillation method of the present invention;

[0062] Figure 4This invention provides a lightweight detection module based on multi-scale feature aggregation.

[0063] Figure 5 This is a network structure diagram of the lightweight adaptive extraction module of the present invention;

[0064] Figure 6 This is a network structure diagram of the multi-path flow splitting feature matching module of the present invention;

[0065] Figure 7 This is a diagram of the DDMA-SAM semantic segmentation network structure of the present invention;

[0066] Figure 8 These are the experimental results of dynamic point removal in this invention;

[0067] Figure 9 This is a diagram showing the result of the semantic segmentation network of this invention. Detailed Implementation

[0068] Explanation of abbreviations:

[0069] RANSAC: Random Sample Consensus

[0070] ATE: Absolute Trajectory Error

[0071] RPE: Relative Trajectory Error

[0072] DDMA-SAM: Decoupled Distillation and Multi-scale Aggregation for Segment Anything Model.

[0073] like Figures 1-9 As shown, this invention proposes a dynamic point semantic filtering method for SLAM based on DDMA-SAM, aiming to improve the localization robustness and mapping accuracy of visual SLAM systems in dynamic environments. It has the advantages of lightweight model, strong real-time performance, and wide adaptability. The overall structure of the method includes the following steps:

[0074] Step 1: Construction and Inference of DDMA-SAM Semantic Network

[0075] This step is used to build a lightweight and efficient semantic segmentation model and output a mask of dynamic targets in the image.

[0076] Model structure optimization

[0077] This invention optimizes the original SAM model. In the original SAM model, the cue encoder and mask decoder have lightweight structures with fewer than 4 million parameters, while the image encoder, based on ViT-H / 16, has over 600 million parameters, limiting its real-time performance. Therefore, this invention employs a decoupled distillation strategy to compress the image encoder while preserving the original cue encoding and decoding module structures. This strategy directly applies to the image embedding layer, avoiding instability caused by randomness in the decoding stage, and relies only on the mean squared error loss function, eliminating the need to combine focus loss and Dice loss, further reducing computational overhead.

[0078] Decoupled distillation process as follows Figure 2 As shown, given an input image x, the teacher model and the student model output score values ​​z respectively. T and z S The probability distribution p is calculated using a normalized exponential function. T and p S Assuming the target class is t, the output distribution can be divided into target class probability and non-target class probability. Among them, the target class probability... and The probability corresponding to class t; the probability of non-target class. and This corresponds to all categories except the target class.

[0079] The algorithm structure optimization process is as follows: Figure 3 As shown. Decoupling distillation directly distills the lightweight image encoder from the original image encoder, decomposing the traditional decoupling distillation loss into target class KL loss (TCKD) and non-target class KL loss (NCKD). The target class KL loss is used to measure the difference between the student model's prediction on the target class and the teacher model, and its calculation formula is as follows:

[0080] The total loss is defined by weighted combination of TCKD and NCKD as follows:

[0081] L total =α·L TCKD +β·L NCKD

[0082] In the formula, α and β are hyperparameters used to balance the loss contributions of the target class and non-target classes.

[0083] The total loss is calculated to monitor the training results of the neural network.

[0084] Integration of detection and segmentation modules

[0085] In image recognition tasks, relying on single-scale feature representation may lead to the loss of local detail information, thus affecting the overall accuracy of the model. Therefore, some scholars have proposed incorporating multi-scale feature aggregation into the detection model, but the drawback is high computational cost.

[0086] To address this issue, a lightweight detection module based on multi-scale feature aggregation is proposed, with the network structure as follows: Figure 4 As shown, a Lightweight Adaptive Extraction (LAE) module and a Multi-Path Splitting Feature Matching (MSFM) module are introduced to quickly extract multi-scale features and perform feature aggregation. The generated prompt information serves as the input prompt for the SAM. Compared with traditional convolutional methods, the Lightweight Adaptive Extraction module can reduce the number of parameters and computational cost, while extracting more semantically informative features, thereby improving the detection capability in dynamic environments.

[0087] This invention is achieved through Figure 1 The semantic segmentation thread in the image performs segmentation, specifically by using the Lightweight Adaptive Extraction (LAE) module and the Multi-Path Feature Matching (MSFM) module to quickly extract multi-scale features and perform feature aggregation. The generated prompt information is used as the prompt input of SAM, and then the SAM model is used to achieve segmentation. The segmentation result is the semantic mask of dynamic targets in the image.

[0088] The structure of the lightweight adaptive extraction module LAE is as follows: Figure 5 As shown, dividing the input channels into N groups for grouped convolution reduces the total number of parameters to 1 / N of that of traditional convolution. Each lightweight adaptive extraction unit achieves a four-fold downsampling. After preserving the height and width information in the channels, the feature map dimension is expanded from four to five dimensions to mitigate the loss of edge information during downsampling. The adaptive extraction path achieves information exchange through average pooling and convolution operations, recombining the feature map based on the four pixels in the upper left corner, and calculating the weights of each using a normalized exponential function, while simultaneously converting the dimension to five dimensions. In the n-dimensional dimension, the adaptive weights are fused with another branch.

[0089] To effectively fuse high-level spatial information with low-level visual information, this invention proposes a multi-path splitting feature matching (MSFM) module, which performs comprehensive analysis of spatial and channel information on features from low to high levels. Its structure is as follows: Figure 6 As shown, the MSFM module's structure applies the splitting concept, where the MatchNeck block enhances the model's ability to express features through splitting operations. The splitting process first separates the information flow using a splitting operator, and the preserved original features can be used for residual connections.

[0090] The MSFM module extracts information from the input feature vector in three dimensions. Let the average pooling operation be P.avg Global average pooling operation P gavg From input features F in Extract output features F' out :

[0091] F' out =Res(Align(F h ,F w ,F c ))

[0092] In the formula: F h and F w The height and width information are calculated from the input features:

[0093]

[0094] in, The height of the input image. The width of the input image;

[0095] Next, height and width information are integrated into the channel dimension to achieve global information fusion, thereby capturing the common features of the ROI region and its neighborhood and promoting global information interaction. Spatial information is set as auxiliary weights and passed back to the spatial vector through multiplication. Let w be the weights corresponding to width and height. h and w w , can be represented as:

[0096] w h ,w w =Split(Sigmoid(Concat(F h ,F w ))

[0097] Next, the channel information can be calculated through post-processing. for:

[0098]

[0099] Then, the final output F is obtained by branching the channel information. out The calculation formula is as follows:

[0100]

[0101] Finally, the original features are concatenated and a 1×1 convolution is performed to obtain the output result. The proposed DDMA-SAM semantic segmentation network structure is as follows: Figure 7 As shown.

[0102] Step 2: Preliminary filtering of dynamic points based on semantic segmentation results

[0103] The semantic segmentation module first segments the input image and outputs the semantic mask, category label, and confidence information of the target object. Let the input image be I, the semantic segmentation model outputs a mask M, where the value of each pixel (x,y) is its semantic label, M(x,y)=1 represents a dynamic region, and M(x,y)=0 represents a static region.

[0104] In the tracking thread, the system needs to determine which feature points might belong to dynamic objects. A mask matching method is used to compare the positions of the feature points with the dynamic mask generated by the semantic segmentation module: if the pixel position of a feature point falls within the dynamic mask region, it is marked as a dynamic point and removed. The initial set D of dynamic regions can be represented as:

[0105] D = {(x,y) | M(x,y) = 1}

[0106] That is, all pixels with a mask value of 1. To more accurately remove dynamic feature points during feature extraction, the algorithm also incorporates the following improvements:

[0107] (1) Confidence filtering: For each semantic mask region, a confidence threshold for dynamic objects is set based on the confidence score, and feature points on dynamic objects with high confidence are removed to avoid accidental deletion.

[0108] (2) Region expansion strategy: To address the potential problem of missed detection of dynamic points at the edge of the mask, the dynamic mask region output by semantic segmentation is appropriately expanded to ensure that dynamic points in the edge region can also be identified.

[0109] Step 3: Quadratic dynamic point filtering based on geometric consistency constraints

[0110] For feature points filtered after semantic segmentation, dynamic points are further filtered out by combining epipolar geometric constraints. First, the RANSAC algorithm is used to robustly estimate the essential matrix E and the fundamental matrix F from the matched feature point pairs. The epipolar geometric constraints are used to determine whether the point pairs meet geometric consistency, and dynamic points that do not meet the constraints are eliminated. Next, the three-dimensional coordinates of the remaining feature points are recovered by triangulation to verify their spatial consistency. Finally, further detection is performed based on projection error. The reprojection error of the feature points is calculated, and dynamic points with errors exceeding a set threshold are eliminated, thereby retaining static feature points that meet both geometric and projection constraints.

[0111] First, the consistency of point pairs is determined by epipolar geometric constraints. Given the fundamental matrix F, corresponding points p1 and p2 in two adjacent frames should satisfy the epipolar constraints:

[0112]

[0113] If point pairs p1 and p2 satisfy this relationship, it indicates they may originate from the same static point; otherwise, they are considered dynamic points. Next, triangulation is used to recover the 3D point. The camera pose between two frames is recovered using the essential matrix E, and then point pairs p1 and p2 are triangulated to recover point P in 3D space:

[0114] P=λ1K -1 p1=λ2K -1 (Rp2+t)

[0115] In the formula: K is the camera intrinsic parameter matrix, λ1 and λ2 are depth values, R is the camera rotation matrix, which represents the rotation relationship from the second frame camera coordinate system to the first frame coordinate system, and t is the camera translation vector, which represents the translation from the second frame coordinate system to the first frame coordinate system.

[0116] If the depth value of point P after triangulation is positive, it indicates that the point is in front of the camera and conforms to geometric relationships; otherwise, point pair p1 and p2 are considered mismatched or originate from a dynamic object. Finally, dynamic points are calculated based on projection errors. The triangulated 3D point P is projected back to the two frames of images, and the reprojection error is calculated:

[0117]

[0118] In the formula: and It is the point P projected onto the first and second frame images.

[0119] If the error ε1 or ε2 exceeds the set threshold, it indicates that the point pair may originate from a dynamic object and should be removed. Experiments have shown that setting the threshold to 1 effectively eliminates feature point pairs.

[0120] Step 4: Static point-driven SLAM pose estimation and mapping

[0121] The DDMA-SAM semantic segmentation network and dynamic point filtering module are integrated into the ORB-SLAM3 system as independent threads to ensure that the processing of the input image and semantic mask inference are executed in parallel without interfering with the tracking and mapping process of the main SLAM thread. The overall process is as follows: Figure 1 As shown, by using static feature points after removing dynamic points to perform inter-frame matching, pose estimation, and map construction, only stable environmental information is retained, which effectively improves map accuracy and system robustness, and adapts to complex dynamic scenes. Figure 1 The local mapping thread, semantic fast relocation thread, loop closure detection thread, and lightweight sparse point cloud construction thread are all existing technologies in the ORB-SLAM3 system, and will not be elaborated on further here.

[0122] Table 1 shows the experimental results of the three algorithms. On the w_xyz sequence, ATE reduces performance by 95.56% and 25.81% compared to ORB-SLAM3 and DynaSLAM, respectively.

[0123] Table 1. Test results of absolute trajectory error and relative trajectory error

[0124]

[0125] Based on traditional visual SLAM systems, this method proposes a DDMA-SAM semantic filtering algorithm for dynamic environments. By optimizing the Segment Anything Model (SAM) through a dual-module decoupling structure, a high-precision, low-latency dynamic point removal strategy is achieved to improve the localization robustness and mapping accuracy of SLAM systems in complex dynamic scenes.

[0126] Specifically, the algorithm introduces a decoupled distillation mechanism to perform lightweight optimization of the image encoder in the original SAM model, constructing a DDMA-SAM semantic network that integrates a multi-scale aggregation detection module and an efficient mask decoding module, thereby improving segmentation performance while significantly compressing model parameters. Based on this semantic network, a dual filtering strategy driven by semantic prior and geometric consistency is proposed: on the one hand, preliminary dynamic point identification is performed based on semantic masks and confidence thresholds; on the other hand, by combining epipolar geometric constraints and triangulation reprojection errors estimated by Random Sample Consensus Estimation (RANSAC), fine-grained removal of dynamic feature points is achieved, retaining only static points for camera pose estimation.

[0127] Meanwhile, the algorithm embeds a lightweight segmentation module and a dynamic point discrimination module in the ORB-SLAM3 visual SLAM backbone framework as independent threads. While maintaining the real-time performance of the system, it effectively improves the mapping quality and trajectory stability in dynamic scenes, and has significant practical application value.

[0128] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. A dynamic point semantic filtering method for SLAM based on DDMA-SAM, characterized in that, Includes the following steps: S1. Construct the DDMA-SAM semantic network: including the image encoder, cue encoding module, and mask decoding module in the original SAM model, and integrate the detection and segmentation modules; use a decoupled distillation strategy to compress the image encoder; the DDMA-SAM semantic network is used to segment the input image and output the semantic mask of dynamic targets in the image; S2. Preliminary filtering of dynamic points based on semantic segmentation results: The pixel position of the feature point is compared with the semantic mask of the dynamic target in the image using the mask matching method. If the pixel position of the feature point falls within the semantic mask area of ​​the dynamic target, it is marked as a dynamic point and removed. S3. Secondary dynamic point filtering based on geometric consistency constraints: For the feature points filtered in S2, dynamic points are further filtered out by combining epipolar geometric constraints to obtain static feature points that meet geometric and projection constraints.

2. The SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 1, characterized in that: It also includes S4, static feature point driven SLAM pose estimation and mapping: using static feature points after removing dynamic points to perform inter-frame matching, pose estimation and map construction.

3. The SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 1, characterized in that: The decoupled distillation strategy described in S1 includes: Decoupling distillation distills a lightweight image encoder from the original image encoder. It decomposes the traditional decoupling distillation loss into target class KL loss and non-target class KL loss. The target class KL loss is used to measure the difference between the student model's prediction on the target class and the teacher model, while the non-target class KL loss is used to measure the difference between the student model's prediction on the non-target class and the teacher model. The total loss is defined as a weighted combination of the target class KL loss and the non-target class KL loss: L total =α·L TCKD +β·L NCKD In the formula, α and β are hyperparameters used to balance the loss contributions of the target class and non-target classes, and L TCKD For the target class KL loss, L NCKD For non-target class KL loss.

4. The SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 1, characterized in that: The detection and segmentation module described in S1 includes: introducing a lightweight adaptive extraction (LAE) module and a multi-path splitting feature matching (MSFM) module to extract multi-scale features and perform feature aggregation. The generated prompt information is used as the prompt input of the SAM model, and then the SAM model is used for segmentation. After segmentation, the semantic mask of the dynamic target in the image is output.

5. The SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 4, characterized in that: The lightweight adaptive extraction LAE module performs the following operations: the input channels are divided into N groups for grouped convolution, each lightweight adaptive extraction unit performs a four-fold downsampling, after saving the height and width information to the channels, the dimension of the feature map is expanded from four-dimensional to five-dimensional, the adaptive extraction path achieves information exchange through average pooling and convolution operations, the feature map is recombined according to multiple pixels, and the weights of each are calculated through a normalized exponential function, while the dimension is converted to five-dimensional.

6. The SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 4, characterized in that: The Multipath Flow Feature Matching (MSFM) module performs the following operations: The MSFM module extracts information from the input feature vector in three dimensions. Let the average pooling operation be P. avg Global average pooling operation P gavg From input features F in Extract output features F ' out : F ' out =Res(Align(F h ,F w ,F c )) In the formula: F h and F w The height and width information are calculated from the input features: in, The height of the input image. The width of the input image; Height and width information are integrated into the channel dimension for global information fusion, thereby capturing the common features of the ROI region and its neighborhood. Spatial information is set as auxiliary weights and passed back to the spatial vector through multiplication. Let w be the weight corresponding to the width and height. h and w w , represented as: w h ,w w =Split(Sigmoid(Concat(F h ,F w )) Channel information is calculated through post-processing. for: The final output F is obtained by branching the channel information. out The calculation formula is as follows: The original features are concatenated and a 1×1 convolution is performed to obtain the output result.

7. The SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 1, characterized in that: S2 specifically includes the following steps: The mask matching method is used to compare the pixel position of a feature point with the semantic mask of the dynamic target in the image. If the pixel position of a feature point falls within the semantic mask area of ​​the dynamic target in the image, it is marked as a dynamic point and removed.

8. The SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 7, characterized in that: In S2, for each semantic mask region, a confidence threshold for dynamic objects is set based on the confidence score, and only feature points on dynamic objects with a confidence score higher than the confidence threshold are removed; and the semantic mask region of dynamic targets in the output image is dilated to ensure that dynamic points in edge regions can also be identified.

9. A SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 1, characterized in that: S3 includes the following steps: First, the essential matrix E and the fundamental matrix F are robustly estimated from the matched feature point pairs using the RANSAC algorithm. Then, the point pairs are judged to meet the geometric consistency by using epipolar geometric constraints, and dynamic points that do not meet the requirements are eliminated. Next, the three-dimensional coordinates of the remaining feature points are recovered by triangulation to verify their spatial consistency; Finally, based on the projection error, further detection is performed by calculating the reprojection error of the feature points and eliminating dynamic points whose errors exceed a set threshold, thereby retaining static feature points that conform to geometric and projection constraints.

10. A SLAM dynamic point semantic filtering method based on DDMA-SAM according to claim 9, characterized in that: S3 specifically includes the following steps: First, the consistency of point pairs is determined by epipolar geometric constraints. Given the fundamental matrix F, corresponding points p1 and p2 in two adjacent frames should satisfy the epipolar constraints: If point pairs p1 and p2 satisfy this relationship, it means they come from the same static point; otherwise, they are considered dynamic points and are removed. Next, triangulation is used to recover the 3D points, and the essential matrix E is used to recover the camera pose between the two frames. Then, point pairs p1 and p2 are triangulated to recover point P in 3D space. P=λ1K -1 p1=λ2K -1 (Rp2+t) In the formula: K is the camera intrinsic parameter matrix, λ1 and λ2 are depth values, R is the camera rotation matrix, which represents the rotation relationship from the second frame camera coordinate system to the first frame coordinate system, and t is the camera translation vector, which represents the translation from the second frame coordinate system to the first frame coordinate system. If the depth value of point P is positive after triangulation, it means that the point is in front of the camera and conforms to the geometric relationship; otherwise, point pair p1 and p2 are considered to be mismatched or from dynamic objects and are removed. Finally, based on the projection error, the dynamic point is calculated, the triangulated 3D point P is projected back to the two frames of images, and the reprojection error is calculated: In the formula: and It is the point P projected onto the first and second frame images; If the error ε1 or ε2 exceeds the set threshold, it indicates that the point pair comes from a dynamic object and should be discarded.

Citation Information

Cited By

  • Unmanned aerial vehicle vision radar tight coupling positioning method, device, equipment, medium and product

    CN121432459A

  • Dynamic scene vision SLAM (Simultaneous Localization and Mapping) method based on Transform and multi-modal fusion

    CN121564719A