A method for enhancing the accuracy and robustness of a SLAM system in dynamic scenarios
By identifying and evaluating the scene flow consistency of dynamic objects, and using the reprojection error to construct a factor graph model to optimize the camera position, the problem of inaccurate and insufficient positioning of dynamic SLAM systems in dynamic environments is solved, and higher system accuracy and stability are achieved.
Patent Information
- Application Number
- CN202310408084.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-04-17
AI Technical Summary
The existing dynamic SLAM technology lacks effective evaluation of dynamic objects in dynamic environments, resulting in inaccurate positioning and insufficient robustness. Traditional methods usually eliminate or add all dynamic objects, affecting the accuracy and robustness of the system.
Through the instance segmentation network, identify dynamic objects, calculate scene flow and cluster, and use the reprojection error construction factor graph model to jointly optimize the camera position and dynamic object position pose to evaluate the consistency of scene flow to improve system accuracy and robustness.
It improves the positioning accuracy and robustness of SLAM system in dynamic scenarios, reduces the impact of dynamic objects on the system, and enhances the stability of the system.
Smart Images

Figure CN116429089B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image and graphics, and relates to a simultaneous localization and mapping technology, specifically to a method for enhancing the accuracy and robustness of a SLAM system in a dynamic scenario. Background Art
[0002] At present, mobile robots have been widely used in fields such as autonomous driving, logistics distribution, and environmental monitoring. The realization of these applications depends on the support of underlying technologies, and the navigation technology is precisely the key technology for robot movement. In order to realize the navigation of mobile robots in a dynamic scenario, the simultaneous localization and mapping technology (SLAM, Simultaneously positioning and mapping) in a dynamic environment has become a research field that has received much attention.
[0003] Traditional SLAM technology relies on the assumption of a static environment and ignores the existence of dynamic obstacles in the environment. Therefore, there are problems such as inaccurate positioning and ghosting in mapping in a dynamic environment. To solve these problems, dynamic SLAM technology has emerged. In dynamic SLAM technology, the robot can perceive dynamic obstacles in the environment, model and track them, and update the state estimation of the robot and map construction according to the motion state of the obstacles. At present, dynamic SLAM lacks means for evaluating dynamic objects, and usually eliminates all dynamic objects or adds all dynamic objects to joint optimization, resulting in the accuracy and robustness of the system being affected. In order to obtain more information from dynamic objects in the scene while ensuring that the positioning accuracy of the system is not affected by them, by evaluating the strength of the scene flow consistency of dynamic objects, the dynamic objects with stronger scene flow consistency are jointly optimized with the camera pose to complete system positioning and dynamic object tracking, improving the robustness and stability of the SLAM system in a dynamic scenario. Summary of the Invention
[0004] To solve the above problems, the present invention discloses a method for enhancing the accuracy and robustness of a SLAM system in a dynamic scenario, evaluating dynamic objects and solving the camera pose and tracking dynamic objects, so that the SLAM system in a dynamic scenario can operate efficiently, accurately, and stably.
[0005] To achieve the above object, the technical solution of the present invention is as follows:
[0006] A method for enhancing the accuracy and robustness of a SLAM system in a dynamic scenario, comprising the following steps:
[0007] (1) Using an instance segmentation network to identify object categories, and solving the scene flow for the points in the object regions belonging to the dynamic category according to the depth image and the optical flow image;
[0008] (2) Determine whether an object is a dynamic object based on the proportion of the magnitude of the scene flow vector in the object region of the dynamic category, automatically cluster the scene flow of the dynamic object and calculate the energy value, and determine the strength of the scene flow consistency of the dynamic object based on the energy value;
[0009] (3) Use the reprojection error to construct a factor graph model to solve and optimize the pose of the camera and the dynamic object with strong scene flow consistency.
[0010] Furthermore, in step (1), an instance segmentation network is used to identify the object category, and the scene flow is solved for the points in the object region belonging to the dynamic category according to the depth image and the optical flow image. The specific steps are as follows:
[0011] (1.1) Use the instance segmentation network for the RGB image input at time t to obtain a mask image, thereby obtaining one or more object regions of the dynamic category
[0012] (1.2) Obtain the matching information for the pixels in the background and the object regions of non-dynamic categories according to the optical flow, and calculate the initial pose T of the camera t ;
[0013] (1.3) For each object region of the dynamic category in the pixel p i Calculate the scene flow vector set according to the depth image and the optical flow image Scene flow vector is:
[0014]
[0015] where is the position of the 3D point of pixel p i in the world coordinate system at time t, is the position of the 3D point of pixel p i in the world coordinate system at time t-1.
[0016] Furthermore, in step (2), determine whether an object is a dynamic object based on the proportion of the magnitude of the scene flow vector in the object region of the dynamic category, automatically cluster the scene flow of the dynamic object and calculate the energy value, and determine the strength of the scene flow consistency of the dynamic object based on the energy value. The specific steps are as follows:
[0017] (2.1) For all the scene flow vectors in the scene flow vector set of each object region of the dynamic category calculate the magnitude. When the proportion of the scene flow vectors with a magnitude greater than the threshold T d = 0.12 exceeds 70%, determine that the object is a dynamic object;
[0018] (2.2) Convert the scene flow vectors in the set of scene flow vectors to the spherical coordinate system, and jointly form data points with the starting points of their scene flow vectors and their three-dimensional coordinates Then use the automatic clustering algorithm DBSCAN to determine the clustering cluster C i to obtain the clustering cluster set of the dynamic object O j
[0019] (2.3) Calculate the energy value E for the obtained clustering cluster set C. If the energy value is less than the threshold ( is the number of pixels in the dynamic object area, is the number of clusters), it is considered that the scene flow consistency of the dynamic object is strong and can be jointly optimized with the camera pose. The energy value E is:
[0020]
[0021] where is the variance of each cluster C in the cluster set i
[0022] Furthermore, in step (3), a factor graph model is constructed using the reprojection error to solve and optimize the camera pose and the pose of the dynamic object with strong scene flow consistency. The specific steps are as follows:
[0023] (3.1) Construct the reprojection error of static points
[0024]
[0025] where π(·) is the mapping from 3D points to pixel points, is a pixel point on the image I t at time t, P i is the corresponding 3D point, T t is the initial pose of the camera at time t, K is the camera internal parameter, and Z is the pixel depth;
[0026] (3.2) Construct the reprojection error of dynamic object points
[0027]
[0028] where is the pose of the dynamic object at time t, is the image I t at time t with strong scene flow consistency of the dynamic object O j A pixel point within the mask is the corresponding 3D point;
[0029] (3.3) Construct the velocity error of the dynamic object Assume that the dynamic object moves at a constant speed:
[0030]
[0031] where is the velocity and angular velocity of the dynamic object at time t-, is the velocity and angular velocity of the dynamic object at time t;
[0032] (3.4) Construct the fitting error of the pose and velocity of the dynamic object
[0033]
[0034] where is the Lie algebra of the pose which is the pose of the dynamic object and is the increment from time t- to time t
[0035]
[0036] (3.5) Use the above reprojection error to construct a factor graph model to jointly solve and optimize the camera pose and the pose of the dynamic object:
[0037]
[0038] where is the set of all variables to be optimized, λ * is the optimal value obtained by solving the variables to be optimized, ∑ is the information matrix, and H is the Huber robust kernel function.
[0039] The beneficial effects of the present invention are:
[0040] Aiming at the problem that the current SLAM system in a dynamic scene lacks an efficient dynamic object evaluation for dynamic objects, the present invention proposes a method for determining the object dynamics and scene flow consistency. First, an instance segmentation network is used to identify the object category, and the scene flow of the dynamic objects of the possible moving categories is solved according to the depth image and the optical flow image; then, it is judged whether the object is moving according to the modulus value of the scene flow within the dynamic object mask, the scene flow of the moving dynamic objects is automatically clustered and the energy value is obtained, and the strength of the scene flow consistency of the dynamic objects is determined based on the energy value; finally, a factor graph model is constructed using the reprojection error to solve and optimize the camera pose and the pose of the dynamic objects with strong scene flow consistency.
[0041] In view of the problem of low accuracy and robustness of the SLAM system in dynamic scenarios, the present invention proposes a method for enhancing the accuracy and robustness of the SLAM system in dynamic scenarios, reducing the impact of dynamic objects on the positioning accuracy of the SLAM system, and improving the robustness and stability of the system, which can be widely applied to scenarios such as autonomous driving and robot navigation. Description of the Drawings
[0042] Figure 1 It is a system structure diagram of a method for enhancing the accuracy and robustness of the SLAM system in dynamic scenarios;
[0043] Figure 2 It is a flowchart of a method for enhancing the accuracy and robustness of the SLAM system in dynamic scenarios;
[0044] Figure 3 It is a factor graph model of a method for enhancing the accuracy and robustness of the SLAM system in dynamic scenarios. Detailed Embodiment
[0045] The present invention will be further clarified below in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0046] As shown in the figure, a method for enhancing the accuracy and robustness of the SLAM system in dynamic scenarios according to the present invention includes the following steps:
[0047] Step S1: Use an instance segmentation network to identify object categories, and solve the scene flow for the points in the object regions belonging to the dynamic category according to the depth image and the optical flow image. Specifically, it includes:
[0048] S1.1: Use the instance segmentation network on the RGB image input at time t to obtain a mask image, thereby obtaining one or more object regions of the dynamic category
[0049] S1.2: Obtain matching information for the pixels in the background and non-dynamic object regions according to the optical flow, and calculate the initial pose T of the camera t ;
[0050] S1.3: For each object region of the dynamic category in the pixel p i Calculate the scene flow vector set according to the depth image and the optical flow image Scene flow vector is:
[0051]
[0052] Where is the pixel p iThe position of the 3D point at time t in the world coordinate system is the pixel p i The position of the 3D point at time t-1 in the world coordinate system. The solution formulas for P and P′ are as follows:
[0053]
[0054] where (u′, v′) is the pixel in the image at time t-1, (flow u′ , flow v′ ) is its optical flow vector, Z′, Z are the depths of the pixels, R′, t′ are the rotation matrix and translation matrix of the camera at time t-1, R, t are the pose of the camera at the moment, and K is the internal parameter matrix of the camera.
[0055] Step S2: Determine whether the object is a dynamic object according to the proportion of the modulus value of the scene flow vector in the object area of the dynamic category, automatically cluster the scene flow of the dynamic object and calculate the energy value, and determine the strength of the scene flow consistency of the dynamic object based on the energy value. Specifically, it includes:
[0056] S2.1. Calculate the modulus value of all the scene flow vectors in the scene flow vector set of each object area of the dynamic category. When the proportion of the scene flow vectors with modulus values greater than the threshold T d =0.12 exceeds 70%, it is determined that the object is a dynamic object;
[0057] S2.2. Convert the scene flow vectors in the scene flow vector set to the spherical coordinate system, and jointly form data points with the starting point of its scene flow vector Then use the automatic clustering algorithm DBSCAN to determine the clustering cluster C i , and obtain the clustering cluster set j of the dynamic object O The distance D between two data points in the clustering is defined as:
[0058]
[0059] where S SF , S Dis are the proportional amplification factors, is the vector distance of the scene flow vector, is the distance between the starting points;
[0060] S2.3. Calculate the energy value E for the obtained clustering cluster set C. If the energy value is less than the threshold ( is the number of pixels in the dynamic object area, If the number of clusters is
[0061]
[0062] where is the set of clusters and the variance of each cluster C i in the set.
[0063] Step S3: Use the reprojection error to construct a factor graph model to solve and optimize the camera pose and the pose of the dynamic object with strong scene flow consistency. The specific steps are as follows:
[0064] S3.1. Construct the reprojection error of static points
[0065]
[0066] where π(·) is the mapping from a 3D point to a pixel point, is a pixel point on the image I t at time t, P i is the corresponding 3D point, T t is the camera pose at time t, K is the camera intrinsic parameter, and Z is the pixel depth;
[0067] S3.2. Construct the reprojection error of dynamic object points
[0068]
[0069] where is the pose of the dynamic object at time t, is the dynamic object O t with strong scene flow consistency on the image I j at time t within a pixel point of the mask, is the corresponding 3D point;
[0070] S3.3. Construct the dynamic object velocity error Assume that the dynamic object moves at a constant speed:
[0071]
[0072] where is the velocity and angular velocity of the dynamic object at time t - 1, is the velocity and angular velocity of the dynamic object at time t;
[0073] S3.4. Constructing the fitting error of the position and velocity of dynamic objects
[0074]
[0075] in For posture The Lie algebra, the increment from time t-1 to time t
[0076] S3.5. Use the above reprojection error to construct a factor graph model to jointly solve and optimize the camera pose and dynamic object pose:
[0077]
[0078] in is the set of all variables to be optimized, λ * is the optimal value of the variable to be optimized, ∑ is the information matrix, and H is the Huber robust kernel function. The final constructed factor graph model is as follows: Figure 3 As shown in the figure, the graph is the vertices of the optimization variables, and the line segments are the observed data. During the optimization process, the variable to be optimized will get closer and closer to the optimal value, and the value after the optimization is completed is regarded as the optimal value.
[0079] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications all fall within the protection scope of the claims of the present invention.
Claims
1. A method for enhancing the accuracy and robustness of a SLAM system in dynamic scenarios, characterized in that, It includes the following steps: (1) Use an instance segmentation network to identify object categories, and solve the scene flow for the points in the object regions belonging to the dynamic category based on the depth image and the optical flow image; (2) Determine whether the object is a dynamic object according to the proportion of the scene flow vector modulus value in the object region of the dynamic category, automatically cluster the scene flow of the dynamic object and calculate the energy value, and determine the strength of the scene flow consistency of the dynamic object based on the energy value; (3) Use the reprojection error to construct a factor graph model to solve and optimize the poses of the camera and the dynamic objects with strong scene flow consistency; The specific steps are as follows: (3.1) Reprojection error of constructing static points where π(·) is the mapping from 3D points to pixel points, is a pixel point on the image I at time t, t P i is the corresponding 3D point, t T is the initial pose of the camera at time t, K is the camera intrinsic parameter, and Z is the pixel depth; (3.2) Construct the reprojection error of dynamic object points Among them is the pose of the dynamic object at time t, is the image I at time t t on the dynamic object O with strong scene flow consistency j a pixel point within the mask, is the corresponding 3D point; (3.3) Construct the velocity error of the dynamic object Assume that the dynamic object moves at a constant speed: wherein is the velocity and angular velocity of the dynamic object at time t-1, is the velocity and angular velocity of the dynamic object at time t; (3.4) Fitting errors of the pose and velocity of the dynamic object wherein is the pose of the dynamic object and is the Lie algebra, and the increment from time t-1 to time t (3.5) Use the above reprojection error to construct a factor graph model to jointly solve and optimize the camera pose and the dynamic object pose: where is the set of all variables to be optimized, λ* is the optimal value obtained by solving the variables to be optimized, ∑ is the information matrix, and H is the Huber robust kernel function.
2. The method for enhancing the accuracy and robustness of a SLAM system in a dynamic scenario according to claim 1, wherein In the step (1), use an instance segmentation network to identify object categories, and solve the scene flow for the points in the object regions belonging to the dynamic category based on the depth image and the optical flow image. The specific steps are as follows: (1.1) Use the instance segmentation network for the RGB image input at time t to obtain a mask image, thereby obtaining the object regions of one or more dynamic categories. (1.2) Obtain matching information for pixels within the object regions of the background and non-dynamic categories based on optical flow, and calculate the initial pose T of the camera t ; (1.3) For each object region of each dynamic category for the pixel p i Calculate a set of scene flow vectors based on the depth image and the optical flow image scene flow vector is: Among them is the position of pixel p i at time t of the 3D point in the world coordinate system, is the position of pixel p i at time t-1 of the 3D point in the world coordinate system.
3. A method for enhancing the accuracy and robustness of a SLAM system in a dynamic scenario according to claim 1, characterized in that, In the step (2), determine whether the object is a dynamic object according to the proportion of the scene flow vector modulus value in the object region of the dynamic category, automatically cluster the scene flow of the dynamic object and calculate the energy value, and determine the strength of the scene flow consistency of the dynamic object based on the energy value. The specific steps are as follows: (2.1) The set of scene flow vectors for the object regions of each dynamic category for all the scene flow vectors calculate the modulus value. When the proportion of scene flow vectors with a modulus value greater than the threshold T d = 0.12 exceeds 70%, the object is determined to be a dynamic object; (2.2) Convert the scene flow vectors in the set of scene flow vectors to the spherical coordinate system, and jointly form data points with the starting points of their scene flow vectors and their three-dimensional coordinates Then use the automatic clustering algorithm DBSCAN to determine the clustering cluster C and obtain the clustering cluster set of the dynamic object O i j (2.3) For the obtained clustering clusters Calculate the energy value E. If the energy value is less than the threshold where is the number of pixels in the dynamic object area, is the number of clusters, it is considered that the scene flow consistency of the dynamic object is relatively strong and can be jointly optimized with the camera pose. The energy value E is: wherein is the clustering variance of each cluster C i in
Citation Information
Patent Citations
Robustness pose estimation method based on instance segmentation in dynamic scene
CN113362358A
Camera pose optimization method based on Manhattan world hypothesis and factor graph
CN114241050A