Dynamic scene visual SLAM optimization method based on semantic and geometric constraints
By using a combination of semantic segmentation and optical flow constraints in the SLAM system to remove dynamic feature points and select descriptors through inter-frame rotation estimation, the problems of robustness and accuracy of SLAM system in a dynamic environment are solved, and higher system robustness and accuracy are achieved.
Patent Information
- Application Number
- CN202510033075.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Dynamic objects will lead to a decrease in the robustness and accuracy of the SLAM system. Traditional feature extraction methods are poorly robust in the lighting environment, and the feature extraction network based on deep learning is poorly rotated.
The synchronous positioning and graph building method based on semantic and optical flow constraints is adopted, and the semantic information of the object is obtained through the semantic segmentation module, combined with Lucas-Kanade optical flow tracking and motion consistency detection, dynamic feature points are removed, and a better descriptor is selected through the inter-frame rotation estimation method.
It improves the robustness and accuracy of the SLAM system, reduces the interference of dynamic objects to the system, and enhances the performance of the system in complex environments.
Smart Images

Figure CN119963833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot positioning and navigation, and more specifically to a dynamic scene visual SLAM optimization method based on semantics and geometric constraints. Background Art
[0002] Simultaneous localization and mapping (SLAM) is essential for robot vision and helps in camera pose estimation and mapping of unknown environments. Visual SLAM system uses camera as sensor input to extract image information for positioning and mapping. It can locate its own position in an unknown environment in real time and build a 3D map of the environment at the same time. It is a very critical technology in the field of computer vision and robotics and has a wide range of applications.
[0003] Most of the current SLAM frameworks are based on static assumptions. However, in actual scenes, moving objects are inevitable, which will limit the application of SLAM systems in actual scenes. Most SLAM methods are based on feature point methods. However, feature-based visual SLAM systems are seriously affected by the quality of feature extraction. First, the traditional extracted features have poor robustness to environments with changing lighting conditions, and the global features extracted using bag of words (BOW) will destroy the spatial information in the scene and reduce the closed-loop performance; second, dynamic feature points in dynamic environments will significantly interfere with the accuracy and robustness of the system. Some excellent works have begun to use deep learning to solve these potential problems. Some studies use deep learning to extract feature points and obtain better tracking accuracy. Some people use semantic information derived from deep learning to remove dynamic feature points to improve the robustness of the system, but the CNN-based feature extraction network will cause poor rotation robustness. In addition, in Crowd-SLAM, the author found that removing too many feature points will reduce the accuracy. Therefore, improving the poor system robustness caused by deep learning and reducing the interference of dynamic objects on the system are of great significance to improving the robustness of the SLAM system. Summary of the invention
[0004] The technical problem to be solved by the present invention is: how to solve the problem that dynamic objects will cause the robustness and accuracy of the SLAM system to decrease, the traditional feature extraction method has poor robustness in an environment with changing lighting conditions, and the feature extraction network based on deep learning has poor rotation robustness.
[0005] To solve the above technical problems, the present invention provides the following technical solutions: a synchronous positioning and mapping method based on semantic and optical flow constraints, which performs semantic segmentation on the input image through a semantic segmentation module to obtain the semantic information of the initial object, and then tracks the feature points through Lucas-Kanade optical flow and performs motion consistency detection to accurately remove dynamic feature points, and determines the rotation angle between consecutive frames through an inter-frame rotation estimation method, and selects a better descriptor to improve the robustness of the SLAM system. It includes the following steps:
[0006] Step 1: Obtain the RGB image and depth image of the image frame through the RGB camera;
[0007] Step 2: The RGB image obtained in step 1 is input into the semantic segmentation module and the HF-Net feature extraction module respectively;
[0008] Step 3: The RGB image obtained in step 1 is semantically segmented through the semantic segmentation network YOLACT++ to obtain the semantic information of the feature points;
[0009] Step 4: The RGB image obtained in step 1 is converted into a grayscale image, and an image pyramid is constructed. HF-Net extracts local features at each layer and generates HFNet descriptors;
[0010] Step 5: Use Lucas-Kanade optical flow to track feature points and perform motion consistency check to make an initial assessment of the motion state of feature points;
[0011] Step 6: Combine the semantic information of the object and the initial evaluation results of motion consistency detection to accurately remove dynamic feature points;
[0012] Step 7: Estimate the inter-frame rotation angle through the inter-frame rotation estimation method, and decide whether to use the HFNet descriptor or recalculate the ORB descriptor;
[0013] Step 8: Calculate the Euclidean distance between the new key frame and the key frames stored in the key frame library, and select the key frame with the smallest Euclidean distance as the candidate key frame;
[0014] Step 9: Perform subsequent local mapping and loop closure detection on the image frames with dynamic points removed.
[0015] Furthermore, the semantic segmentation network in step 3 uses YOLACT++, which adds deformable convolutions to the YOLACT framework; it uses more reasonable proportions, allocation strategies and anchors, so that each anchor can be better allocated to the correct target, and the anchors are used more densely, which can improve the detection and segmentation capabilities of small target objects. On the basis of reducing the FC layer, the Squeeze-and-Excitation mechanism (SE) is introduced to ensure the detection accuracy while improving the calculation accuracy. The input image is semantically segmented through the YOLACT++ semantic segmentation network to obtain the semantic information of the object.
[0016] Furthermore, in step 5, Lucas-Kanade optical flow is used to track feature points for motion consistency detection. By preliminarily judging the motion state of feature points in the image, first, the LK optical flow method is used to track the feature points of the previous frame extracted by HFNet, and the optical flow vector is obtained according to the result of the current frame. The points successfully tracked in the previous frame and the current frame are marked as P 1 and P 2 , they are denoted as P 1 =[u 1 v 1 1] P 2 =[u 2 v 2 1], where u and v represent pixel coordinates. The RANSAC algorithm is used to filter out abnormal optical flow values and calculate the epipolar line L 1 :
[0017]
[0018] Where, L 1 Represents pixel P 1 The corresponding epipolar line in the previous frame image, represents the three components of the polar line expressed in vector form, F is the basic matrix between the previous frame and the current frame, P 1 Represents the matching point in the previous frame corresponding to the point successfully tracked in the current frame image. Represents the homogeneous coordinates of the successfully tracked point in the previous frame.
[0019] Calculate the distance D between the pixel point in the current frame and its corresponding epipolar line,
[0020]
[0021] Where D represents P 2 The distance to its corresponding limit, P 2 Represents the points in the current frame image that match the previous frame image, P1 represents the corresponding matching point in the previous frame image, F is the basic matrix between the previous frame and the current frame, X represents the first dimension parameter of the polar line vector, and Y represents the second dimension parameter of the polar line vector. If the D value is greater than the threshold δ 1 , the threshold should be set according to the lighting level of the environment, and it is considered to be a potential dynamic point.
[0022] Furthermore, in step 6, Lucas-Kanade optical flow is used to track feature points, perform motion consistency detection, and semantic information obtained by the semantic segmentation network YOLACT++ to accurately remove dynamic feature points. The specific steps are:
[0023] After preliminary screening of the dynamic feature points and static feature points of the current frame using the optical flow method and motion consistency detection, the SLAM system divides the target detection frame into 9 areas. The overall target detection frame is called the mother detection frame, and the 9 divided areas are called sub-detection frames. Combined with the semantic information of the feature points, the ratio of dynamic feature points to static feature points is greater than the threshold δ 2 The initial threshold is set to 0.5. Each box is evaluated and classified as a dynamic sub-detection box or a static sub-detection box. If the system identifies it as a dynamic sub-detection box, in order to ensure that no potential dynamic areas are missed in the mask, the adjacent sub-detection boxes are adjusted to a dynamic state. If most of the sub-detection boxes of the target detection are dynamic sub-detection boxes, the objects in the target detection box are considered to be highly dynamic objects and should be removed.
[0024] Furthermore, in step 7, the inter-frame rotation angle is estimated by the inter-frame rotation estimation method to decide whether to use the local descriptor of HFNet or recalculate the ORB descriptor; the specific steps are:
[0025] The optical flow vector obtained in step 5 is used to optimize the inter-frame rotation estimation using the least squares method. First, the perpendicular bisector of each optical flow vector is calculated.
[0026] ax+by+c=0
[0027] Where a, b, c represent the coefficients corresponding to the line, and x and y represent the corresponding points on the optical flow vector line, including the two endpoints of the optical flow vector.
[0028] The center of the image is selected as the starting point for optimization, because most inter-frame rotations usually occur near the center of the image, which can greatly reduce the optimization time. Construct and solve the optimization function f(x,y),
[0029]
[0030] In order to save iteration time, BFGS optimization is used, and the gradient of the objective function is calculated as follows:
[0031]
[0032] In the formula, represents the partial derivative of f(x,y) in the x direction, represents the partial derivative of f(x,y) in the y direction,
[0033]
[0034]
[0035] Each update can get the optimal point P. Each update of P is performed in the specified direction, so that f(x, y) can converge quickly. The updated f(x, y) is as follows:
[0036] f(x k +α k p k,x ,y k +α k p k,y ).
[0037] In the formula, p k,x and p k,y Represents the iteration direction, which is determined by the following formula:
[0038]
[0039] Hessian Matrix It reflects the local curvature information of the objective function near the nearest iteration point and provides a more accurate descent direction. The formula is calculated by the following formula:
[0040]
[0041] In the formula, s k Represents the change vector, y k represents the gradient change vector, ρ k Represents the update amplitude.
[0042] s k =(x k+1 -x k ,y k+1 -y k )
[0043]
[0044]
[0045] After iterative optimization f(x k ,y k) The optimal point P is taken as the frame rotation center. If the motion between two frames is relatively close to one frame rotation, then the distance from point P to all optical flow vector endpoints should be the same. There will be a certain distance difference D between the distance from one endpoint P1 of the optical flow vector to P and the distance from one endpoint P2 of the optical flow vector to OP. If the distance difference D is too large, the vector will be filtered out.
[0046] ||(P-P1)|| 2 -||(P-P2)|| 2 ≤D
[0047] Where P represents the frame rotation center, and P1 and P2 represent the two endpoints of the optical flow vector. If too many vectors are discarded, it means that the motion between the two frames does not obviously involve frame rotation. The remaining vectors are then used to estimate the rotational motion between the current frame and the previous frame from the angles from both ends to the frame rotation center.
[0048]
[0049] Where P 1i and P 2i Represent the two endpoints of each optical flow vector, respectively, and the threshold angle is set to 20 degrees. If the angle is greater than the threshold angle, the system believes that the descriptor generated by HFNet should not be used due to potential frame rotation. Instead, it recalculates the ORB descriptor of the feature points between the current frame and the previous frame, and uses the same method as ORB-SLAM3 for subsequent matching. On the contrary, if the angle is smaller, the descriptor in HFNet is used.
[0050] Furthermore, in step 7, when the descriptor in HFNet is used, BOW is still used to accelerate the matching, but the distance calculation is changed from Hamming distance to Euclidean norm, as shown below:
[0051] Gdist=||gdes1-gdes2|| 2
[0052] In the formula, gdes1 and gdes2 represent the descriptors to be matched; therefore, when the system estimates that the scene has rotated, it selects the HFNet descriptor for the feature points on the previously extracted keyframe according to the estimated angle, or re-extracts the ORB descriptor, and selects the matching calculation method based on whether the rotation is detected. This effectively combines the rotation robustness of traditional extraction methods with the accuracy advantages of deep learning, allowing the system to operate normally in situations where obvious frame rotation may cause matching performance degradation and tracking failure.
[0053] Furthermore, in step 8, the Euclidean distance between the new key frame and the key frames stored in the key frame library is calculated, and the key frame with the smallest Euclidean distance is selected as the candidate key frame. Specifically:
[0054] When the local mapping thread receives a new keyframe, it calculates its global descriptor vector, denoted as gobaldes, and saves the global descriptor of the current keyframe to the keyframe library. Then it calculates the Euclidean norm (2-norm) between the global descriptor vectors of the current keyframe and other keyframes in the library:
[0055] Gdist=||gdes1-gdes2|| 2
[0056] Where gdes1 and gdes2 represent the global descriptors extracted by HFNet, and each descriptor consists of 4096 floating point numbers. After obtaining the global descriptor of the keyframe, the system calculates the distance to all global descriptors stored in the keyframe library. The smaller Gdist is, the higher the similarity between the two frames, which increases the possibility of closed loop detection. The system selects the frame with the highest similarity as the candidate frame based on the similarity of all descriptors in the keyframe library.
[0057] Furthermore, in step 9, only image frames with static features are used for subsequent local mapping and loop closure detection.
[0058] Compared with the existing known technologies, the beneficial effects of the present invention are as follows: the present invention ensures the accurate removal of dynamic feature points by combining geometric consistency detection and semantic information to remove dynamic feature points, thereby improving the accuracy and robustness of the operation of the visual SLAM system. A new feature extraction and matching mechanism is added, and HFNet is used for local feature extraction and global feature extraction, which can significantly improve the tracking and closed-loop detection performance of the visual SLAM system. The inter-frame rotation angle is estimated by the inter-frame rotation estimation method to determine whether to use the HFNet descriptor or re-extract the ORB descriptor, which effectively combines the rotation robustness of the traditional extraction method and the accuracy advantage of deep learning, so that the deep learning-based system can operate normally when obvious frame rotation may cause a decrease in matching performance and tracking failure. The robustness of the system in complex environments is guaranteed. In addition, in loop detection, the global descriptor generated by HFNet is used instead of the original BOW method to solve the problem that the BOW-based method has limited description capabilities for the scene and is prone to losing spatial information about the depicted object. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is the overall flow chart of the present invention;
[0060] Figure 2 This is a diagram of the example segmentation network framework of the present invention;
[0061] Figure 3 This is a flow chart of removing dynamic feature points of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the specification. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] like Figure 1 An embodiment of a dynamic scene visual SLAM optimization method based on semantic and geometric constraints is shown, comprising the following steps:
[0064] Step 1: This example first obtains RGB images and depth images through the RGBD camera to provide image frames for the HF-Net network and the YOLACT++ segmentation network. The entire method uses two branches for parallel processing. Branch 1 performs semantic segmentation on the RGB image to obtain object semantic information, and branch 2 extracts features through HF-Net and generates local and global descriptors.
[0065] Step 2: If Figure 2 The figure shows the YOLACT network structure. In order to improve the detection capability of small target objects, this embodiment uses YOLACT++ as the segmentation network. The input image is semantically segmented through the YOLACT++ semantic segmentation network to obtain the object semantic information;
[0066] Step 3: Convert the input RGB image to a grayscale image, construct an image pyramid, and HF-Net extracts local features at each layer and generates HFNet descriptors;
[0067] Step 4: Use Lucas-Kanade optical flow to track feature points and perform motion consistency detection to preliminarily determine the motion state of feature points in the image.
[0068] S4.1: Get the optical flow vector based on the result of the current frame, and mark the points successfully tracked in the previous frame and the current frame as P 1 and P 2 , they are denoted as P 1 =[u 1 v 1 1]P 2 =[u 2 v 2 1], where u and v represent pixel coordinates. The RANSAC algorithm is used to filter out abnormal optical flow values, and the epipolar line L 1 The calculation method is:
[0069]
[0070] Where, L 1 Represents pixel P 1 The corresponding epipolar line in the previous frame image, represents the three components of the polar line expressed in vector form, F is the basic matrix between the previous frame and the current frame, P 1 Represents the matching point in the previous frame corresponding to the point successfully tracked in the current frame image. Represents the homogeneous coordinates of the successfully tracked point in the previous frame.
[0071] S4.2: Calculate the distance D between the pixel point in the current frame and its corresponding epipolar line,
[0072]
[0073] Where D represents P 2 The distance to its corresponding limit, P 2 Represents the points in the current frame image that match the previous frame image, P 1 represents the corresponding matching point in the previous frame image, F is the basic matrix between the previous frame and the current frame, X represents the first dimension parameter of the polar line vector, and Y represents the second dimension parameter of the polar line vector. If the D value is greater than the threshold δ 1 , the threshold should be set according to the lighting level of the environment, and it is considered to be a potential dynamic point.
[0074] Step 5: If Figure 3 The figure shows the flow chart of removing dynamic feature points. First, the dynamic feature points and static feature points of the current frame are preliminarily screened using the optical flow-polar line method. The target detection frame is divided into 9 areas. The overall target detection frame is called the mother detection frame, and the 9 divided areas are called sub-detection frames. Combined with the semantic information of the feature points, the ratio of dynamic feature points to static feature points is greater than the threshold δ. 2 , the initial threshold is set to 0.5, each box is evaluated and classified as a dynamic sub-detection box or a static sub-detection box. If the system identifies it as a dynamic sub-detection box, in order to ensure that no potential dynamic areas are missed in the mask, the adjacent sub-detection boxes are adjusted to a dynamic state. If most of the target detection boxes of the target detection are dynamic sub-detection boxes, the objects in the target detection box are considered to be highly dynamic objects and should be completely removed.
[0075] Step 6: Estimate the inter-frame rotation angle through the inter-frame rotation estimation method, and decide whether to use the local descriptor of HFNet or recalculate the ORB descriptor;
[0076] S6.1: Compute the perpendicular bisector of each optical flow vector
[0077] ax+by+c=0
[0078] Where a, b, c represent the coefficients corresponding to the line, and x and y represent the corresponding points on the optical flow vector line, including the two endpoints of the optical flow vector.
[0079] S6.2: Select the center of the image as the starting point for optimization, construct and solve the optimization function f(x,y),
[0080]
[0081] S6.3: Use BFGS optimization to calculate the gradient of the objective function:
[0082]
[0083] Each update can get the optimal point P. Each update of P is performed in the specified direction, so that f(x, y) can converge quickly. The updated f(x, y) is as follows:
[0084] f(x k +α k p k,x ,y k +α k p k,y ).
[0085] In the formula, p k,x and p k,y Represents the iteration direction, which is determined by the following formula:
[0086]
[0087] Hessian Matrix It reflects the local curvature information of the objective function near the nearest iteration point and provides a more accurate descent direction. The formula is calculated by the following formula:
[0088]
[0089] In the formula, s k Represents the change vector, y k represents the gradient change vector, ρ k Represents the update amplitude.
[0090] s k =(x k+1 -x k ,y k+1 -y k )
[0091]
[0092]
[0093] S6.4: After iterative optimization f(x k ,y k ) The optimal point P is taken as the frame rotation center. If the motion between two frames is relatively close to one frame rotation, then the distance from point P to all optical flow vector endpoints should be the same. There will be a certain distance difference D between the distance from one endpoint P1 of the optical flow vector to P and the distance from one endpoint P2 of the optical flow vector to OP. If the distance difference D is too large, the vector will be filtered out.
[0094] ||(P-P1)|| 2 -||(P-P2)|| 2 ≤D
[0095] Where P represents the frame rotation center, and P1 and P2 represent the two endpoints of the optical flow vector. If too many vectors are discarded, it means that there is no obvious frame rotation in the motion between the two frames.
[0096] S6.5: Estimate the rotational motion between the current frame and the previous frame by the angles of the remaining vector from both ends to the frame rotation center.
[0097]
[0098] Where P 1i and P 2i Represent the two endpoints of each optical flow vector respectively; the threshold angle is set to 20. If the angle is greater than the threshold angle, the system believes that the descriptor generated by HFNet should not be used due to potential frame rotation. Instead, it recalculates the ORB descriptor of the feature points between the current frame and the previous frame, and uses the same method as ORB-SLAM3 for subsequent matching. On the contrary, if the angle is smaller, the descriptor in HFNet is used.
[0099] Step 7: Select the HFNet descriptor for the feature points on the previously extracted keyframe based on the estimated angle, or re-extract the ORB descriptor, and select the matching calculation method based on whether rotation is detected.
[0100] Step 8: When the local mapping thread receives a new keyframe, it calculates its global descriptor vector, denoted as gobaldes, and saves the global descriptor of the current keyframe into the keyframe library.
[0101] Step 9: Then calculate the Euclidean norm (2-norm) between the global descriptor vectors of the current keyframe and other keyframes in the library:
[0102] Gdist = |gdes1-gdes2|| 2
[0103] Where gdes1 and gdes2 represent the global descriptors extracted by HFNet, and each descriptor consists of 4096 floating-point numbers.
[0104] Step 10: After obtaining the global descriptor of the keyframe, calculate the distance to all global descriptors stored in the keyframe library. The smaller Gdist is, the higher the similarity between the two frames is, which increases the possibility of closed loop detection. The system selects the frame with the highest similarity as the candidate frame based on the similarity of all descriptors in the keyframe library.
[0105] Step 11: Use image frames containing only static features for subsequent local mapping and loop closure detection.
[0106] The above method effectively eliminates dynamic feature points by combining semantic information and geometric detection, determines the rotation angle between consecutive frames through the frame rotation estimation method, and selects more appropriate local descriptors to improve the robustness of the SLAM system and meet the needs of visual SLAM in actual scenarios.
[0107] In summary, the dynamic SLAM method based on instance segmentation and deep features provided by the present invention comprises the following steps: on the basis of ORB-SLAM3, a new feature point and descriptor extraction module, a semantic segmentation module and a geometric detection module are integrated, the feature point and descriptor extraction module uses HFNet instead of traditional feature extraction to generate local and global descriptors, the input image is semantically segmented by the semantic segmentation module to obtain an initial mask, and then the feature points are tracked by Lucas-Kanade optical flow, and a mobile consistency check is performed to accurately remove dynamic feature points, the rotation angle between consecutive frames is determined by a frame rotation estimation method, and a better local descriptor is selected to improve the robustness of the SLAM system. The present invention uses images collected by an RGB-D camera, and after semantic segmentation, optical flow tracking and mobile consistency check, and inter-frame rotation estimation, dynamic feature points can be accurately removed, and the SLAM system also has good robustness in scenes with large illumination changes and moving objects, and can be applied to most practical scenes.
[0108] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not intended to limit the implementation methods of the present invention. For ordinary technicians in the relevant field, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A dynamic scene visual SLAM optimization method based on semantic and geometric constraints, characterized in that: The following steps are involved: Step 1: Obtain the RGB image and depth image of the image frame through the RGB camera; Step 2: The RGB image obtained in step 1 is input into the semantic segmentation module and the HF-Net feature extraction module respectively; Step 3: The RGB image obtained in step 1 is semantically segmented through the semantic segmentation network YOLACT++ to obtain the semantic information of the feature points; Step 4: The RGB image obtained in step 1 is converted into a grayscale image, and an image pyramid is constructed. HF-Net extracts local features at each layer and generates HFNet descriptors; Step 5: Use Lucas-Kanade optical flow to track feature points and perform motion consistency check to make an initial assessment of the motion state of feature points; Step 6: Combine the semantic information of the object and the initial evaluation results of motion consistency detection to accurately remove dynamic feature points; Step 7: Estimate the inter-frame rotation angle through geometric consistency constraints and optical flow vectors, and decide whether to use the HFNet descriptor or recalculate the ORB descriptor; Step 8: Calculate the Euclidean distance between the new keyframe and the keyframes stored in the keyframe library. Select the key frame with the smallest Euclidean distance as the candidate key frame; Step 9: Perform subsequent local mapping and loop closure detection on the image frames with dynamic points removed.
2. A dynamic scene visual SLAM optimization method based on semantic and geometric constraints according to claim 1, characterized in that: The semantic segmentation network uses YOLACT++. YOLACT++ adds deformable convolution on the basis of the YOLACT framework, adopts a more reasonable ratio, allocation strategy and anchor, so that each anchor can be better assigned to the correct target, and the anchor is used more densely, which can improve the detection and segmentation capabilities of small target objects. On the basis of reducing the FC layer, the compression-excitation mechanism (Squeeze-and-Excitation, SE) is introduced. While ensuring the detection accuracy, the calculation accuracy is improved. The input image is semantically segmented through the YOLACT++ semantic segmentation network to obtain the semantic information of the object.
3. A dynamic scene visual SLAM optimization method based on semantic and geometric constraints according to claim 1, characterized in that: The Lucas-Kanade optical flow is used to track feature points and perform motion consistency checks to make an initial assessment of the motion state of feature points. The specific steps are as follows: First, the LK optical flow method is used to track the feature points of the previous frame extracted by HFNet, and the optical flow vector is obtained according to the results of the current frame. The points successfully tracked in the previous frame and the current frame are marked as P1 and P2, which are recorded as P1 = [u1 v1 1] P2 = [u2 v2 1], where u and v represent pixel coordinates. The RANSAC algorithm is used to filter out abnormal optical flow values and calculate the epipolar line L1: Where L1 represents the epipolar line corresponding to the pixel point P1 in the previous frame image. represents the three components of the epipolar line expressed in vector form, F is the basic matrix between the previous frame and the current frame, P1 represents the matching point corresponding to the successfully tracked point in the current frame image in the previous frame image, Represents the homogeneous coordinates of the successfully tracked point in the previous frame of the image; calculate the distance D between the pixel point and its corresponding epipolar line in the current frame: Where D represents the distance from P2 to its corresponding limit, P2 represents the point in the current frame image that matches the previous frame image, P1 represents the corresponding matching point in the previous frame image, F is the basic matrix between the previous frame and the current frame, X represents the first dimension parameter of the polar line vector, and Y represents the second dimension parameter of the polar line vector. If the D value is too large, it is considered to be a potential dynamic point.
4. A dynamic scene visual SLAM optimization method based on semantic and geometric constraints according to claim 1, characterized in that: After preliminary screening of the dynamic feature points and static feature points of the current frame using the optical flow method and motion consistency detection, the SLAM system divides the target detection frame into 9 areas. The overall target detection frame is called the parent detection frame, and the 9 divided areas are called sub-detection frames. Combined with semantic information, based on whether the ratio of dynamic feature points to static feature points is greater than the threshold δ1, the initial threshold is set to 0.5, and each frame is evaluated and classified as a dynamic sub-detection frame or a static sub-detection frame. If the system identifies it as a dynamic sub-detection frame, in order to ensure that no potential dynamic areas are missed in the mask, the adjacent sub-detection frames must be adjusted to a dynamic state. If most of the target detection frames of the target detection are dynamic sub-detection frames, the objects in the target detection frame are considered to be highly dynamic objects and must be eliminated.
5. A dynamic scene visual SLAM optimization method based on semantic and geometric constraints according to claim 1, characterized in that: Estimate the inter-frame rotation angle through the inter-frame rotation estimation method, and decide whether to use the local descriptor of HFNet or recalculate the ORB descriptor; the specific steps are: use the optical flow vector obtained in step 5 to optimize the inter-frame rotation estimation using the least squares method. First, calculate the perpendicular bisector of each optical flow vector ax+by+c=0 In the formula, a, b, c represent the coefficients corresponding to the line, x and y represent the corresponding points on the optical flow vector line, including the two endpoints of the optical flow vector; Select the center of the image as the starting point for optimization, construct and solve the optimization function f(x,y), In order to save iteration time, BFGS optimization is used, and the gradient of the objective function is calculated as follows: In the formula, represents the partial derivative of f(x,y) in the x direction, represents the partial derivative of f(x,y) in the y direction, Each update can get the optimal point P. Each update of P is performed in the specified direction, so that f(x, y) can converge quickly. The updated f(x, y) is as follows: f(x k +a k p k,x ,y k +a k p k,y ). In the formula, p k,x and p k,y Represents the iteration direction, which is determined by the following formula: Hessian Matrix It reflects the local curvature information of the objective function near the nearest iteration point and provides a more accurate descent direction. The formula is calculated by the following formula: In the formula, s k Represents the change vector, y k represents the gradient change vector, ρ k Represents the update amplitude; s k (x) k+1 -x k ,and k+1 -and k ) After iterative optimization f(x k ,y k ) Get the optimal point P as the frame rotation center. If the movement between two frames is relatively close to one frame rotation, then the distance from point P to all optical flow vector endpoints should be the same; the distance from one endpoint P1 of the optical flow vector to P and the distance from one endpoint P2 of the optical flow vector to OP will have a certain distance difference D. If the distance difference D is too large, this vector will be filtered out; ||(P-P1)||2-||(P-P2)||2≤D Where P represents the frame rotation center, P1 and P2 represent the two endpoints of the optical flow vector; therefore, some optical flow vectors can be filtered. If too many vectors are discarded, it means that there is no obvious frame rotation in the motion between the two frames. Then the remaining vectors are used to estimate the rotation motion between the current frame and the previous frame from the angles from the two ends to the frame rotation center. Where P 1i and P 2i Represent the two endpoints of each optical flow vector respectively; the threshold angle is set to 20 degrees. If the angle is greater than the threshold angle, the system believes that the descriptor generated by HFNet should not be used due to potential frame rotation. Instead, it recalculates the ORB descriptor of the feature points between the current frame and the previous frame and uses the same method as ORB-SLAM3 for subsequent matching. On the contrary, if the angle is smaller, the descriptor in HFNet is used.
6. A dynamic scene visual SLAM optimization method based on semantic and geometric constraints according to claim 1, characterized in that: When using the descriptor in HFNet, BOW is still used to accelerate matching, but the distance calculation is changed from Hamming distance to Euclidean norm, as shown below: Gdist=||gdes1-gdes2||2 Where gdes1 and gdes2 represent the descriptors to be matched; therefore, when the system estimates that the scene has been rotated, the HFNet descriptor is selected for the feature points on the previously extracted keyframe according to the estimated angle, or the ORB descriptor is re-extracted, and the matching calculation method is selected according to whether the rotation is detected; this effectively combines the rotation robustness of traditional feature extraction methods with the accuracy advantages of deep learning, enabling deep learning-based systems to operate normally in situations where obvious frame rotation may lead to degraded matching performance and tracking loss.
7. A dynamic scene visual SLAM optimization method based on semantic and geometric constraints according to claim 1, characterized in that: ORB-SLAM3 uses the bag-of-words (BOW) method to perform closed-loop detection by replacing the original BOW method with the global descriptor generated by HFNet. However, the BOW-based method has limited ability to describe the scene and is prone to losing spatial information about the depicted object. When the local mapping thread receives a new keyframe, it calculates its global descriptor vector, expressed as gobaldes, and saves the global descriptor of the current keyframe to the keyframe library. Then, the Euclidean norm (2-norm) between the global descriptor vectors of the current keyframe and other keyframes in the library is calculated: Gdist=||gdes1-gdes2||2 Where gdes1 and gdes2 represent the global descriptors extracted by HFNet, each descriptor consists of 4096 floating-point numbers. After obtaining the global descriptor of the key frame, the distance to all global descriptors stored in the key frame library is calculated; the smaller Gdist is, the higher the similarity between the two frames, thereby increasing the possibility of closed-loop detection. The system selects the frame with the highest similarity as the candidate frame based on the similarity of all descriptors in the key frame library.
Citation Information
Patent Citations
Dynamic SLAM system based on RGBD and encoder fusion
CN110458863A
Visual positioning method in dynamic obstacle interference environment
CN118470289A
Method and device for selecting keyframe based on motion state
US20220398845A1