SLAM (Simultaneous Localization and Mapping) method for guiding adaptive dynamic feature fusion based on degradation weight
By calculating the degradation weights of dynamic features and the visual residual weights, a new feature set is constructed for coarse and fine matching, which solves the problem of pose estimation degradation in highly dynamic environments in SLAM technology and achieves high-precision and high-reliability pose estimation and loop closure detection.
Patent Information
- Application Number
- CN202510967464.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-18
AI Technical Summary
Existing SLAM technology suffers from pose estimation degradation and even system crashes in highly dynamic environments due to the lack of static features.
By calculating the degradation weights of dynamic features, and utilizing IMU pre-integration and visual residual weighting, a new feature set is constructed for coarse and fine matching. Combined with loop closure detection and global optimization, the accuracy and robustness of pose estimation are improved.
In highly dynamic environments, it improves the accuracy of pose estimation and the reliability of loop closure detection, avoids system crashes caused by missing features, and enhances the stability and accuracy of the system.
Smart Images

Figure CN120976691A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of image data processing, and further relates to a simultaneous localization and mapping (SLAM) method based on degenerate weight guided adaptive dynamic feature fusion. BACKGROUND
[0002] Simultaneous Localization and Mapping (SLAM) refers to the positioning and attitude determination of a mobile device such as a robot through a sensor, and the completion of map construction of the surrounding environment. As a core branch of the SLAM technology system, visual-inertial SLAM exhibits excellent navigation and positioning performance by fusing the complementary advantages of visual sensors and inertial measurement units (IMUs). At present, the mainstream visual SLAM method mainly constructs an environment map in a static environment or a dynamic environment. Based on the assumption of a static environment, the feature point-based SLAM method relies on the matching and tracking of static features. In real-world scenarios, dynamic objects such as pedestrians and vehicles are often included, which produce moving feature points in images. Since traditional SLAM algorithms cannot effectively distinguish between these dynamic feature points and static background feature points, they are prone to introduce false observations in the matching and tracking process, leading to an increase in cumulative errors and a decrease in system accuracy.
[0003] A SLAM dynamic disturbance suppression method based on fuzzy processing and target detection is disclosed in the patent document "A SLAM dynamic disturbance suppression method based on fuzzy processing and target detection" (application number 202510082525.2, publication number CN 119515730A) applied by Nankai University. The implementation steps of the method include: obtaining a blur remover; applying the blur remover to the image for blur processing and extracting ORB feature points; obtaining a dynamic target detector; detecting the clear image through the dynamic target detector to obtain a target detection box; classifying the motion state of the object in the clear image according to the target detection box and the ORB feature points, and obtaining a static feature point set using the overlapping box strategy and epipolar geometry constraint; and performing system pose estimation and trajectory prediction according to the static feature point set. Although this method reduces the adverse effects of dynamic targets on other parts of the system by deblurring, the target detection function can accurately identify dynamic targets, and combined with the overlapping box strategy and epipolar geometry constraint, it can completely remove dynamic feature points, thus improving the overall performance of the system. However, this method still has the disadvantage that moving objects will frequently occupy the field of view of the vision sensor, causing a serious lack of static features and affecting the function of the vision unit, and even possibly causing the system to crash.
[0004] A SLAM method in a dynamic scene based on point-line fusion is disclosed in the patent document "A SLAM method in a dynamic scene based on point-line fusion" (application number CN 202510098772.1, publication number CN 119991806 A) applied by Anhui University of Technology. The implementation steps of the method are as follows: building a visual SLAM framework; obtaining scene images and preprocessing; inputting the preprocessed scene images into the visual SLAM framework to extract semantic information; extracting and fusing feature points from the scene images to obtain a static feature point set; initializing the visual SLAM framework based on the static feature point set and filtering key frame scene images; constructing a sparse point cloud map based on the filtered key frame scene images; and performing loop detection and map correction on the key frame scene images through the visual SLAM framework. The invention improves the robustness and real-time performance of the SLAM system in dynamic environments and low-texture environments based on the fusion of ORB-SLAM3, improved YOLOv8, and improved LSD line feature extraction algorithm. However, this method still has the disadvantage that in the process of removing dynamic points, it often relies on prior knowledge of dynamic objects, which reduces the running efficiency of the SLAM system and makes it difficult to adapt to complex and variable environments. SUMMARY
[0005] The purpose of the present application is to solve the problem of the existing SLAM technology that the pose estimation is degraded and even the system collapses due to the lack of static features in a highly dynamic environment.
[0006] To achieve the above purpose, the main steps of the present application include the following:
[0007] Step 1, the front end of the SLAM system extracts and tracks the feature points of the image frame, and aligns the measurement data of the IMU with the image frame based on the time stamp;
[0008] Step 2, the relative pose change between adjacent frames is calculated by using IMU pre-integration, and dynamic feature detection is performed according to the motion direction consistency and the difference in displacement;
[0009] Step 3, the dynamic degree of the observation frame after dynamic feature detection is calculated, and the observation frame whose dynamic degree is greater than the set threshold ω is determined as an extreme dynamic frame;
[0010] Step 4, the degradation weight of the extreme dynamic frame is calculated, and different visual residuals are dynamically weighted;
[0011] Step 5, based on the degradation weight, the FAST corner points of the current frame are extracted and then coarse matching is performed;
[0012] Step 6, fine matching is performed on the static feature point set divided by the front-end VIO;
[0013] Step 7, the accuracy of loop detection is checked;
[0014] Step 8, based on the loop constraint, the pose graph is adjusted and global optimization is performed.
[0015] Further, the step of extracting and tracking the feature points of the image frame is as follows:
[0016] First, the Shi-Tomasi corner detection algorithm is used to extract the feature points in each image frame;
[0017] Second, the KLT optical flow method is used for tracking, and the corresponding position of each feature point in the next image frame is tracked forward-backward consistently, and the feature points that are successfully tracked are stored and the feature points that fail to be tracked are removed.
[0018] Further, the step of dynamic feature detection according to the motion direction consistency and the difference in displacement is as follows:
[0019] First, the actual motion direction of each feature point in each image frame of the sliding window between two frames is calculated according to the following formula:
[0020]
[0021] wherein, represents the actual motion direction of the pth feature point from the kth frame to the k+1th frame, [u k ,v k ] represents the observed value of the pth feature point on the kth image frame, [u k+1 ,v k+1 ] represents the observed value of the pth feature point on the k+1th image frame;
[0022] In the second step, the motion direction of the feature point is predicted based on the IMU data according to the following formula:
[0023]
[0024] wherein, represents the motion direction of the pth feature point from the kth frame to the k+1th frame in the sliding window through the re-projection of the IMU data, represents the value of the pth feature point after re-projection on the k+1th frame;
[0025] In the third step, it is judged whether the predicted pose satisfies the following motion direction consistency condition, if yes, the fourth step is executed, otherwise, the feature point is considered as a static feature point.
[0026] The motion direction consistency condition is:
[0027]
[0028] wherein, coth() represents the hyperbolic cotangent operator, ||·|| represents the vector modulus, θ min represents the minimum motion angle of the pth feature point in the sliding window;
[0029] In the fourth step, the displacement d of the feature point is calculated according to the following formula:
[0030]
[0031] wherein, p k+1 = [u k+1 ,v k+1 ],
[0032] In the fifth step, it is judged whether the displacement of the feature point satisfies the following formula, if yes, the feature point is determined as a dynamic feature point in the k+1th frame I k+1 ; otherwise, the feature point is a static feature point in the frame I k+1 .
[0033]
[0034] wherein d min , d max respectively represent the minimum displacement amount and the maximum displacement amount of the pth feature point in the sliding window.
[0035] Further, the dynamic degree is obtained by the following formula:
[0036]
[0037] wherein r k represents the dynamic degree of the kth frame I k in the sliding window, E k represents the set of dynamic feature points in the frame I k , C k represents the set of static feature points in I k which have not been dynamically transformed, D k represents the set of static feature points in I k which have been dynamically transformed, and |·| represents the modulo operation.
[0038] Further, the degradation weight of the extreme dynamic frame is obtained by the following formula:
[0039]
[0040] wherein η k represents the degradation weight of the kth extreme dynamic frame I k in the sliding window, ε represents the lower limit of the weight, μ represents the coefficient for controlling the decay speed of the weight, and e () represents the exponential function with the natural constant e as the base.
[0041] Compared with the prior art, the present application has the following advantages:
[0042] First, the present application calculates the dynamic degree of the key frame by using the dynamic feature, defines the degradation weight of the dynamic feature, dynamically weights different visual residuals, and incorporates the dynamic feature into the pose estimation, thereby overcoming the problem of the prior art that directly excludes the features related to dynamic objects and only retains the static area for positioning, which reduces the precision and robustness of the pose estimation due to the lack of features in the highly dynamic environment, so that the present application has the advantage of high pose estimation precision in the highly dynamic environment.
[0043] Second, the present application reextracts the FAST corner points of the frame, constructs a new feature set for coarse matching, and overcomes the problem of the prior art that the loop detection in the highly dynamic environment cannot guarantee the recall rate by only using static feature points, so that the loop detection of the present application has the advantage of high reliability. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 is a flowchart of the present application;
[0045] Figure 2 is a simulation comparison chart of camera trajectory tracking accuracy of the present application and prior art;
[0046] Figure 3 is a simulation comparison chart of loop detection results of the present application and prior art. DETAILED DESCRIPTION
[0047] The present application is further described in detail below with reference to the accompanying drawings and examples.
[0048] Reference Figure 1 , the implementation steps of the embodiments of the present application are further described in detail.
[0049] Step 1, the front end of the SLAM system extracts and tracks the feature points of the image frame, and aligns the measurement data of the IMU with the image frame based on the timestamp.
[0050] The feature extraction extracts feature points in each image frame through the Shi-Tomasi corner detection algorithm.
[0051] Step 2, the relative pose change between adjacent frames is calculated using IMU pre-integration, and dynamic feature detection is performed according to the consistency of the motion direction and the difference of the displacement.
[0052] Step 2.1, the actual motion direction of each feature point in each image frame of the sliding window between two frames is calculated using the following formula:
[0053]
[0054] wherein, represents the actual motion direction of the pth feature point from the kth frame to the k+1th frame, [u k ,v k ] represents the observation value of the pth feature point in the kth image frame, [u k+1 ,v k+1 ] represents the observation value of the pth feature point in the k+1th image frame.
[0055] Step 2.2, the motion direction of the feature point predicted based on the IMU data is calculated using the following formula:
[0056]
[0057] wherein, represents the motion direction of the pth feature point in the sliding window from the kth frame to the k+1th frame obtained by reprojecting the IMU data, represents the value of the pth feature point after reprojecting in the k+1th frame.
[0058] Step 2.2, judge whether the predicted pose satisfies the motion direction consistency condition as follows, if not, the feature point is a static feature point, whose angle difference is only the re-projection error; the feature point satisfying the condition is a dynamic feature point, which does not move in the expected direction between frames.
[0059] The motion direction consistency condition is as follows:
[0060]
[0061] Wherein, coth() represents the hyperbolic cotangent operator, ||·|| represents the vector modulus operation, θ min represents the minimum motion angle of the pth feature point in the sliding window.
[0062] Step 2.3, calculate the displacement d of the feature point by using the following formula:
[0063]
[0064] Wherein, p k+1 = [u k+1 , v k+1 ],
[0065] Step 2.4, judge whether the displacement of the feature point satisfies the following formula, if yes, it is determined that the feature point is a dynamic feature point in the k+1th frame I k+1 ; otherwise, the feature point is a static feature point in the frame I k+1 :
[0066]
[0067] Wherein, d min , d max respectively represent the minimum displacement and the maximum displacement of the pth feature point in the sliding window. The feature point with a motion degree less than d min may be a dynamic feature point, and its error comes from the re-projection error; the feature point with a motion degree greater than d max may be an outlier.
[0068] Step 3, calculate the dynamic degree of the observation frame after dynamic feature detection, and determine the observation frame with a dynamic degree greater than the set threshold ω as an extreme dynamic frame.
[0069] The formula for calculating the dynamic degree is as follows:
[0070]
[0071] Wherein, r k represents the kth frame Ik The degree of dynamism, E k Indicates frame I k The set of dynamic feature points, C k Indicate I k The set of static feature points in D that have not undergone dynamic transformation k Indicate I k The set of static feature points after dynamic transformation, where |·| represents the modulo operation.
[0072] Step 4: Calculate the degradation weights for extreme dynamic frames and dynamically weight different visual residuals.
[0073] Step 4.1, use the following formula to calculate the degradation weight of extreme dynamic frames:
[0074]
[0075] Where, η k Represents the k-th extreme dynamic frame I in the sliding window. k The degradation weights, where ε represents the lower limit of the weights to prevent the visual residual term from being completely abandoned during optimization if the weights are 0, and μ represents a coefficient controlling the rate of weight decay. () This represents an exponential function with the natural constant e as its base.
[0076] Calculating degradation weights ensures that the system does not simply ignore dynamic feature points, leading to failure in extreme dynamic scenarios, while also preventing dynamic features from having an excessive impact on the system.
[0077] Step 4.2: Dynamically weight the different visual residuals.
[0078] Step 4.2.1, calculate the degradation weight for each feature point:
[0079] η p,q =min(η) p ,η q )
[0080] Where, η p ,η p Representing the p-th frame I p , frame q q The degradation weights, min() represents the minimum value operation.
[0081] The degradation weight of each feature point is calculated to illustrate that the visual residual is based on the observation of the feature in two frames. The dynamic range of the two frames will affect the calculation of the residual, so the minimum value of the degradation weight of the two frames is used.
[0082] Step 4.2.2, according to the following formula, apply the static feature residual e in the sliding window that has not undergone dynamic transformation.C Weighting is performed as follows:
[0083]
[0084] wherein, P l Based on the degradation weight of the feature point calculated from the i-th frame and the j-th frame, P denotes the re-projection error of P l in the j-th frame. X denotes the state vector to be optimized.
[0085] Step 4.2.2, the dynamically transformed static feature residual e D is weighted as follows:
[0086]
[0087] wherein, P r Based on the degradation weight of the feature point calculated from the i-th frame and the j-th frame, P denotes the re-projection error of P r in the j-th frame.
[0088] Step 4.2.3, the dynamic feature residual e E is weighted as follows:
[0089] η k,k+1 e E (z h ,X)
[0090] wherein, η k,k+1 P h Based on the degradation weight of the feature point constructed from adjacent frames k, k+1, z h denotes the re-projection error of P h on k, k+1. k, k+1 denotes the motion of the feature point observed on frames I k ,I k+1 .
[0091] Step 5, based on the degradation weight, coarse matching is performed after extracting the FAST corner points in the current frame.
[0092] Step 5.1, based on the degradation weight, the feature point set based on the FAST corner points on the current frame I k is extracted by the Shi-Tomasi corner detection algorithm.
[0093] Step 5.2, based on the Bag of Words model BoW, the frame I kThe extracted feature point set is transformed into a supplementary bag-of-words vector.
[0094] Step 5.3, calculate according to the following formula. Match score:
[0095]
[0096] Among them, f k,d express With the d-th historical frame I d Supplementary bag-of-words vectors The matching score.
[0097] Step 5.4, select those greater than the minimum matching threshold f. min A matching score f of 0.15 k,d It is added to the coarse matching candidate heap, which is a max-heap with a capacity of 5.
[0098] Step 5.5: Determine the top f of the candidate heap in the coarse-matching heap. k,d′ Is it greater than the maximum matching score f? max If so, then frame I is considered... k With historical frame I d’ Coarse matching successful. In an embodiment of the present invention, f max =0.5.
[0099] Step 6: Perform fine matching on the static feature point set divided by the front-end VIO.
[0100] Step 6.1, based on BoW, transfer frame I k The static feature point set is transformed into a static feature bag-of-words vector.
[0101] Step 6.2, calculate according to the following formula. Match score:
[0102]
[0103] Among them, f k,d express With the d-th historical frame I d static feature bag-of-words vectors The matching score vector.
[0104] Step 6.3, f k,d Add the result to a fine-grained candidate heap; this fine-grained candidate heap is a max-heap with a size of 5; take the top of the fine-grained candidate heap, and denote the matching score as f. k,d′ .
[0105] Step 6.4, determine f k,d′Is it greater than the static feature matching threshold? If so, then frame I is considered k and historical frame I d’ The fine-grained matching was successful. In embodiments of the present invention, the static feature matching threshold...
[0106] Step 7: Verify the accuracy of the loop closure detection.
[0107] Step 7.1, extract frame I respectively k With historical frame I d’ For each static feature point in the dataset, a BRIEF descriptor is obtained.
[0108] Step 7.2, calculate frame I k The a-th BRIEF descriptor extracted Its historical frame I d’ The b-th descriptor extracted Hamming distance;
[0109] Step 7.3: Determine if the Hamming distance is less than 80. If so, consider the descriptor... With descriptor If a match is successful, a descriptor pair is formed; otherwise, the descriptor matching fails.
[0110] Step 7.4: Calculate whether the number of descriptor pairs is greater than 25. If yes, obtain a set of descriptor pairs; otherwise, consider frame I as... k With historical frame I d’ The cyclic relationship is inaccurate.
[0111] Step 7.5: Construct and solve the PnP problem, and optimize it using the RANSAC algorithm to obtain frame I. k With historical frame I d’ The relative pose transformation between them.
[0112] Step 7.6, calculate frame I of relative pose transformation. k With historical frame I d’ The relative distance between them and the deviation of the heading angle.
[0113] Step 7.7: For cases where the relative distance and heading angle deviations satisfy the laparoscopy constraint conditions, establish frame I. k With historical frame I d’ The closure constraint.
[0114] The closure constraint condition refers to the situation where the following two conditions are met simultaneously:
[0115] Condition 1: Relative distance < 25;
[0116] Condition 2: Heading angle deviation < 30°.
[0117] Step 8, adjust the pose graph and perform global optimization based on the loop constraint.
[0118] Step 8.1, in the constructed pose graph, based on the established loop constraint, connect frame I k , d ′ In the corresponding node N k , d’ of the pose graph;
[0119] Step 8.2, construct the optimization objective as follows:
[0120]
[0121] wherein, represents the camera coordinate set of the node in the pose graph, and represents the heading angle set of the node in the pose graph; S represents the set of sequential edges in the pose graph, and r m,n represents the residual term constructed based on the mth node N m and the nth node N n , L represents the set of loop edges in the pose graph, and r x,y represents the residual constructed based on the xth node N x and the yth node N y .
[0122] The effect of the present application can be further proved by the following simulation.
[0123] 1. Simulation experiment conditions.
[0124] The software platform of the simulation experiment of the present application is: Ubuntu 18.04LTS, 64-bit operating system and Melodic version of ROS (Robot Operating System).
[0125] The data set used in simulation experiment 1 of the present application is the data sequence containing extreme dynamic scene segments (07, 11, 12, 14, 18) in ADVIO.
[0126] The data set used in simulation experiment 2 of the present application is the data sequence containing extreme dynamic scene segments and loops (03, 05, 06) in ADVIO.
[0127] 2. Simulation content and result analysis.
[0128] The simulation experiment of the present application has two.
[0129] 2.1, simulation experiment 1 is to evaluate the precision of the degradation weight strategy.
[0130] The simulation experiment 1 of the application is to use the visual odometry (DFF-VIO dw ) in the method of the application and a prior art to respectively perform pose tracking on the data sequence of simulation experiment 1, obtain the trajectory on the ADVIO part sequence, and draw the tracking trajectory curve as shown in Figure 2 .
[0131] The prior art used in simulation experiment 1 refers to the visual odometry DFF-VIO without introducing degenerative weights provided in the patent application document “monocular visual inertial odometry method based on dynamic feature tight coupling” (application number: 2024107283467, application publication number: CN118640927A) applied by Xi'an University of Electronic Science and Technology.
[0132] 2.2, simulation experiment 2 is to evaluate the loop detection accuracy.
[0133] The simulation experiment 2 of the application is to use the loop detection module (DFF-VIO dw ) in the method of the application and a prior art to respectively perform loop detection in the data sequence of simulation experiment 2, and draw the PR curve as shown in Figure 3 .
[0134] The prior art used in simulation experiment 2 refers to the loop detection module BoW-Loop provided in the patent application document “monocular visual inertial odometry method based on dynamic feature tight coupling” (application number: 2024107283467, application publication number: CN118640927A) applied by Xi'an University of Electronic Science and Technology.
[0135] The effect of the application will be further described below in combination with simulation graphs.
[0136] Figure 2 The complete trajectory comparison graph and the local two-dimensional enlarged graph of tracking sequence 7, 11 are shown in Figure 2 . The x-axis represents the position coordinate value of the camera moving along the x-axis in the two-dimensional space, the y-axis represents the position coordinate value of the camera moving along the y-axis in the two-dimensional space, and the z-axis represents the position coordinate value of the camera moving along the z-axis in the two-dimensional space, with the unit of meter (m). The two-dimensional local trajectory enlarged graph contains y, z axes and x, y axes, with the unit of meter (m). Figure 2 The red solid line in dw represents the camera pose trajectory curve tracked by DFF-VIO in the application, and the blue solid line represents the camera pose trajectory curve tracked by the prior art.
[0137] From Figure 2As can be seen, the camera pose trajectory curve tracked by the present invention is close to the real camera pose trajectory curve and is superior to the camera pose trajectory curve tracked based on the existing technology.
[0138] Figure 3 This is a PR curve diagram showing the loop closure detection for sequences 03, 05, and 06. Figure 3 In the coordinate system, the vertical axis represents precision, while the horizontal axis represents recall. Figure 3 The red curve in the image represents the PR curve for loop closure detection based on DFF-Loop. Figure 3 The blue curve in the figure represents the PR curve plotted based on loop closure detection using the existing method 1.
[0139] from Figure 3 As can be seen, DFF-Loop achieves a better balance between precision and recall.
Claims
1. A SLAM method based on degenerate weight-guided adaptive dynamic feature fusion, characterized in that, The steps of this method include the following: Step 1: The front end of the SLAM system extracts and tracks features from the image frames, and aligns the IMU measurement data with the image frames based on the timestamps. Step 2: Calculate the relative pose change between adjacent frames using IMU pre-integration, and perform dynamic feature detection based on the consistency of motion direction and the difference in displacement. Step 3: Calculate the dynamic degree of the observed frames after dynamic feature detection, and determine the observed frames with a dynamic degree greater than the set threshold ω as extreme dynamic frames. Step 4: Calculate the degradation weights for extreme dynamic frames and dynamically weight different visual residuals; Step 5: Based on the degradation weight, extract the FAST corner points of the current frame and then perform coarse matching; Step 6: Perform fine-grained matching on the static feature point set divided by the front-end VIO; Step 7: Verify the accuracy of the loop closure detection; Step 8: Based on the loop closure constraint, adjust the pose graph and perform global optimization.
2. The SLAM method according to claim 1, characterized in that, The steps for feature extraction and tracking of the image frame described in step 1 are as follows: The first step is to extract feature points from each image frame using the Shi-Tomasi corner detection algorithm; The second step is to use the KLT optical flow method to track the forward-backward consistent feature points at the corresponding positions of each feature point in each image frame in the next image frame, store the successfully tracked feature points, and remove the untracked feature points.
3. The SLAM method according to claim 2, characterized in that, The dynamic feature detection steps described in step 2 based on the consistency of motion direction and the difference in displacement are as follows: The first step is to calculate the actual motion direction of each feature point in each image frame of the sliding window between two frames, according to the following formula: in, Represents the actual motion direction of the p-th feature point from the k-th frame to the (k+1)-th frame, [u k ,v k ] represents the observation value of the p-th feature point in the k-th image frame, [u k+1 ,v k+1 ] represents the observation value of the p-th feature point in the (k+1)-th image frame; The second step is to predict the motion direction of feature points based on IMU data, according to the following formula: in, This represents the motion direction from frame k to frame (k+1) obtained by reprojecting the p-th feature point in the sliding window using IMU data. This represents the value of the p-th feature point after reprojection in the (k+1)-th frame; The third step is to determine whether the predicted pose satisfies the following motion direction consistency condition. If so, proceed to the fourth step; otherwise, the feature point is considered a static feature point. The condition for consistency of motion direction is: Where coth() represents the hyperbolic cotangent operator, ||·|| represents the vector modulus, and θ min This represents the minimum motion angle of the p-th feature point in the sliding window; Fourth step, calculate the displacement d of the feature point according to the following formula: Where, p k+1 =[u k+1 ,v k+1 ], Fifth, determine whether the displacement of the feature point satisfies the following formula. If it does, then determine that the feature point is in the (k+1)th frame I. k+1 The above is a dynamic feature point; otherwise, the feature point is in frame I. k+1 These are static feature points; Where, d min d max These represent the minimum and maximum displacements of the p-th feature point in the sliding window, respectively.
4. The SLAM method according to claim 3, characterized in that, The degree of dynamism mentioned in step 3 is obtained by the following formula: Where, r k Represents the k-th frame I in the sliding window. k The degree of dynamism, E k Indicates frame I k The set of dynamic feature points, C k Indicate I k The set of static feature points in D that have not undergone dynamic transformation k Indicate I k The set of static feature points after dynamic transformation, where |·| represents the modulo operation.
5. The SLAM method according to claim 4, characterized in that, The degradation weight of the extreme dynamic frame mentioned in step 4 is obtained by the following formula: Where, η k Represents the k-th extreme dynamic frame I in the sliding window. k The degradation weights, where ε represents the lower limit of the weights and μ represents the coefficient controlling the rate of weight decay, e () This represents an exponential function with the natural constant e as its base.
6. The SLAM method according to claim 5, characterized in that, The step of dynamically weighting different visual residuals described in step 4 is as follows: The first step is to calculate the degradation weight η for each feature point according to the following formula. p,q : or p,q =min(η p ,or q ) Where, η p ,η p Representing the p-th frame I p , frame q q The degradation weights, min() represents the minimum value operation; The second step is to apply the following formula to the static feature residuals e in the sliding window that have not undergone dynamic transformation. C Weighting: in, Represents the l-th feature point P l Based on the The degradation weights of feature points calculated in frame j and frame j, p l The reprojection error in the j-th frame, where X represents the state vector to be optimized; The third step is to apply the following formula to the static feature residual e in the sliding window after dynamic transformation. D Weighting: in, P represents the r-th feature point. r Based on the The degradation weights of feature points calculated in frame j and frame j, p r Reprojection error in frame j Fourth step, according to the following formula, calculate the dynamic feature residual e in the sliding window. E Weighting: η k,k+1 e E (z h ,X) Where, η k,k+1 P represents the h-th feature point. h The degradation weights of feature points constructed based on adjacent frames k, k+1, z h P represents h Reprojection error on k, k+1.
7. The SLAM method according to claim 6, characterized in that, The steps described in step 5, which involve extracting FAST corner points of the current frame based on degenerate weights and then performing coarse matching, are as follows: The first step is to extract FAST corner points in the current frame I using the Shi-Tomasi corner detection algorithm, based on degradation weights. k The set of feature points on; The second step is to use the Bag-of-Words (BoW) visual model to divide frame I... k The extracted feature point set is transformed into a supplementary bag-of-words vector. Third step, calculate according to the following formula. Match score: Among them, f k,d express With the d-th historical frame I d Supplementary bag-of-words vectors The matching score; The fourth step is to match values greater than the minimum matching threshold f. min Matching score f k,d Add to the coarse matching candidate heap; Step 5: Determine the top f of the candidate heap in the coarse-matching heap. k,d′ Is it greater than the maximum matching score f? max If so, then frame I is considered... k With historical frame I d’ Coarse match successful.
8. The SLAM method according to claim 7, characterized in that, The steps for fine-tuning the static feature point set partitioned by the front-end VIO in step 6 are as follows: The first step, based on BoW, is to convert frame I... k The static feature point set is transformed into a static feature bag-of-words vector. The second step is to calculate according to the following formula. Match score: Among them, f k,d express With the d-th historical frame I d static feature bag-of-words vectors The matching score vector; The third step is to f k,d Add it to the fine matching candidate heap; The fourth step is to select the root of the fine-match candidate heap, and the matching score is denoted as f. k,d′ Determine f k,d′ Is it greater than the static feature matching threshold? If so, then frame I is considered k and historical frame I d’ The fine match was successful.
9. The SLAM method according to claim 8, characterized in that, The steps for verifying the accuracy of the loop closure detection described in step 7 are as follows: The first step is to extract frame I separately. k With historical frame I d’ For each static feature point in the dataset, a BRIEF descriptor is obtained; The second step is to calculate frame I. k The a-th BRIEF descriptor extracted Its historical frame I d’ The b-th descriptor extracted Hamming distance; The third step is to determine if the Hamming distance is less than 80. If it is, then the descriptor is considered to be... With descriptor If a match is successful, a descriptor pair is formed; otherwise, the descriptor matching fails. The fourth step is to calculate whether the number of descriptor pairs is greater than 25. If so, a set of descriptor pairs is obtained; otherwise, frame I is considered unqualified. k With historical frame I d’ The loop relationship is inaccurate; The fifth step involves constructing and solving the PnP problem, and then optimizing it using the RANSAC algorithm to obtain frame I. k With historical frame I d’ Relative pose transformation between them; Step 6: Calculate frame I of relative pose transformation. k With historical frame I d’ The relative distance and heading angle deviation between them; Step 7: For cases where the relative distance and heading angle deviations satisfy the laparoscopy constraint conditions, establish frame I. k With historical frame I d’ The loop constraint; the loop constraint condition refers to the situation where the following two conditions are met simultaneously: condition 1, relative distance < 25, condition 2, heading angle deviation < 30.
10. The SLAM method according to claim 9, characterized in that, The steps in step 8, which involve adjusting the pose graph and performing global optimization based on loop closure constraints, are as follows: The first step is to connect frames I with closure edges in the constructed pose graph based on the established closure constraints. k I d ′ The corresponding node N in the pose graph k N d’ ; The second step is to construct the optimization objective as follows: in, Let represent the set of camera coordinates of the nodes in the pose graph, ψ represent the set of heading angles of the nodes in the pose graph, S represent the set of sequential edges in the pose graph, and r represent the set of camera coordinates of the nodes in the pose graph. m,n Represents the relationship based on the m-th node N m With the nth node N n The constructed residual term, where L represents the set of loop edges in the pose graph, r x,y Represents the x-th node N x With the y-th node N y The constructed residual.
Citation Information
Patent Citations
Monocular vision inertial odometer method based on dynamic characteristic tight coupling
CN118640927A
SLAM dynamic disturbance suppression method based on fuzzy processing and target detection
CN119515730A
Visual SLAM method and system based on point-line fusion in dynamic scene
CN119991806A