A Dynamic SLAM Method Based on Reprojection Error and Depth Estimation
By extracting the ORB feature points and depth generation network of RGB images, combining instance segmentation and optical flow network, reprojection error and depth residuals are calculated, and dynamic and static targets are identified, the problem of low recognition accuracy of SLAM algorithm in dynamic environments is solved, and positioning accuracy and operation stability are improved.
Patent Information
- Application Number
- CN202211265048.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-17
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-10-17
AI Technical Summary
The existing SLAM algorithm has low dynamic target recognition accuracy in dynamic environments, which affects the positioning and map building accuracy, especially in scenarios where there are many dynamic objects and fast movement speeds, which are difficult to meet the requirements.
By extracting the ORB feature points of the RGB image, combining the deep generation network, instance segmentation network and optical flow network, the reprojection error and depth residuals are calculated, dynamic and static targets are identified, and beam adjustment optimization is performed to estimate the dynamic target pose in real time.
It improves the recognition accuracy and trajectory accuracy of dynamic targets, reduces the number of false positive samples, and improves the positioning accuracy and operation stability of the camera in dynamic scenarios.
Smart Images

Figure CN115619826B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robotics, and mainly includes a dynamic SLAM method based on reprojection error and depth estimation. Background Art
[0002] There are mainly two implementation methods for the odometer module in the visual SLAM algorithm framework: the feature point method and the direct method. The feature point method optimizes the camera motion by minimizing the reprojection error according to the projection positions of spatial points in adjacent frame images, which requires the computer to complete the feature point matching between adjacent frames; the direct method avoids this process and solves it by minimizing the photometric error. Klein et al. proposed the algorithm PTAM, which first introduced the concept of key frames, greatly reducing the computational amount. At the same time, it adopts a multi-threaded mode to process tracking and mapping in parallel, laying a foundation for the formation of the visual SLAM framework. In addition, it first uses non-linear optimization to replace the past filtering method. This algorithm is an important achievement in the SLAM field; Raúl et al. proposed the algorithm ORB-SLAM based on PTAM. This is a SLAM algorithm for monocular cameras, which extracts ORB feature points, realizes pose tracking through epipolar geometry, and introduces a loop detection module to eliminate cumulative errors; subsequently, Raúl updated the algorithm to make it have better support for binocular cameras and RGB-D; Kerl et al. proposed a SLAM algorithm DVO-SLAM based on the direct method. For the RGB image and depth map collected by RGB-D, the algorithm will first set a motion transformation amount, then select points with obvious gradient changes from the previous frame and key frames as reference points, and find corresponding points in the new frame image according to the motion transformation, and optimize the motion transformation amount through a residual function minimization; Jakob et al. proposed the algorithm LSD-SLAM, which introduced normalized variance into the photometric error cost function to reduce the uncertainty impact caused by depth estimation and image noise. The above methods are all based on the static assumption of the environment. However, most actual environments will have dynamic objects, which will interfere with the entire SLAM process. Especially for the direct method, due to the lack of feature extraction, when there are dynamic objects in the environment, its positioning and mapping accuracy will be affected. The feature point method has relatively better robustness, but it is still difficult to meet the requirements in scenarios with many dynamic objects and high moving speeds.
[0003] Yu et al. proposed the algorithm DS-SLAM by combining deep learning. First, the potential dynamic objects in the image are recognized through the semantic segmentation network SegNet. Subsequently, the outliers are recognized by detecting the adjacent two-frame images through epipolar geometry. If the number of inliers and outliers of a potential dynamic object is too large, then this object is considered a dynamic object and is removed. However, when the object moves along the epipolar direction, it is difficult for the algorithm to recognize the outliers. Bescós et al. used two methods simultaneously to judge the dynamic regions in the image. First, Mask-CNN is used for semantic segmentation to filter out all potential dynamic objects in the environment. For objects without prior knowledge and unable to be segmented, the multi-view geometry method is used to judge the depth change of the pixel points in the adjacent frame images to detect the outliers. However, this algorithm also filters out some static objects, resulting in information loss. Zhang et al. proposed the algorithm VDO-SLAM. This method filters out potential dynamic objects through the semantic segmentation network and uses the static background to complete the SLAM process. In addition, the authors also identify the dynamic objects in the environment through the scene flow method and perform real-time estimation of their poses. A global optimization of the static map, camera pose, and dynamic objects is achieved through a factor graph model. This algorithm can basically meet the functional requirements in large-scale outdoor scenes, but the accuracy and recall rate in the dynamic object recognition process need to be improved.
[0004] Currently, the SLAM algorithm framework for dynamic scenes has been basically mature. Among them, the strategy of directly filtering out all potential dynamic objects will lose a large amount of information. In scenes where the static background texture is not rich or the number of objects is large, the tracking accuracy will be affected. An effective way to solve this problem is to use static objects as the input of the SLAM solution process, and this strategy will highly depend on the correct recognition of dynamic objects. In addition, the recognition accuracy is also a key factor affecting the tracking results of dynamic objects. Therefore, how to improve the recognition accuracy of dynamic objects is the key research issue. Summary of the Invention
[0005] Aiming at the above deficiencies in the prior art, a dynamic SLAM method based on reprojection error and depth estimation provided by the present invention solves the problems of low recognition accuracy of dynamic objects and great difficulty in tracking dynamic objects.
[0006] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows:
[0007] S1. Obtain the RGB image and extract the ORB feature points of the RGB image;
[0008] S2. Preprocess the RGB image to obtain the static background, object labels, the coordinates of the pixel points restored to the world coordinate system, and the pixel points after inter-frame matching;
[0009] S3. Calculate the dynamic and static targets based on the coordinates of the ORB feature points, static background, target labels, and pixel points of the RGB image restored to the world coordinate system and the pixel points after inter-frame matching, and obtain the dynamic and static targets according to the reprojection error and depth residual;
[0010] S4. Complete SLAM based on the dynamic target, static background, and static target, and estimate the pose of the dynamic target in real time;
[0011] S5. Optimize the pose of the dynamic target through bundle adjustment.
[0012] Furthermore, the preprocessing of the RGB image is completed through the following three networks:
[0013] Depth generation network: Extract the left-eye and right-eye image features through one CNN channel of the PSMNet network respectively, collect context information through the SPP module, connect the left-eye and right-eye feature maps into a cost volume function, input it into a 3D CNN for regularization, generate a depth map through disparity regression, and restore all pixel points in the RGB image to the world coordinate system to obtain the coordinates of the pixel points restored to the world coordinate system;
[0014] Instance segmentation network: Introduce the RPN module and ROI Align module in the Mask R-CNN network to segment the RGB image, and divide the segmented RGB image into a static background and a target mask;
[0015] Optical flow network: Perform inter-frame matching on randomly sampled pixel points in the RGB image through the PWC-Net network to obtain the pixel points after inter-frame matching.
[0016] Furthermore, the specific implementation method of step S3 is as follows:
[0017] When the camera and the dynamic target move in different directions:
[0018] S3-1. According to the formula:
[0019] T i = T i-1 X i = T i-1 X i-1
[0020] Obtain the current pose T i of the camera; where the transformation matrix of the previous frame of the camera is X i-1 , the pose is T i-1 , and i is the i-th ORB feature point of the target in the previous frame;
[0021] S3-2. According to the formula:
[0022]
[0023] obtain the homogeneous coordinates in the world coordinate system where k represents the serial number of the ORB feature point; i - 1 indicates that this point is obtained by inverse projection from the (i - 1)-th frame image;
[0024] S3-3. According to the formula:
[0025]
[0026] obtain the reprojection error e r ; where n is the number of target sampling points; is the pixel position of the target sampling point in the current frame; ||·|| 2 is the two-norm;
[0027] S3-4. When the reprojection error is less than the set threshold, the target is static; when the reprojection error is greater than the set threshold, the target is dynamic;
[0028] When the camera and the dynamic target run in the same direction:
[0029] S3-5. According to the formula:
[0030]
[0031] obtain the position in the camera coordinate system of the current frame where j - 1 indicates that this point is obtained by inverse projection from the (j - 1)-th frame image; j represents the j-th ORB feature point of the target in the previous frame; k’ represents the serial number of the ORB feature point;
[0032] S3-6. According to the formula:
[0033]
[0034] obtain the pixel position of the target found by the optical flow network in the current frame and inverse project it into the three-dimensional world to obtain the position of the target in the camera coordinate system of the current frame where, is the pixel position of the target sampling point in the current frame;
[0035] S3-7. According to the formula:
[0036]
[0037] obtain the depth residual e d ; where z(·) represents taking the value on the z-axis of the homogeneous coordinates; n’ represents the number of target sampling points;
[0038] S3-8. When the depth residual exceeds the set threshold, the target is a dynamic target; when the depth residual is lower than the set threshold, the target is a static target.
[0039] Further, S4-1. Input the static background and the static target into the visual odometry module to obtain the camera pose and ORB feature points.
[0040] S4-2. Input the dynamic target into the target pose estimation module to obtain the current pose of the dynamic target.
[0041] Further, the specific implementation method of step S4-2 is as follows:
[0042] S4-2-1. According to the formula:
[0043] m a =X o m a '
[0044]
[0045] Obtain the dynamic target transformation matrix X o ; where, m a =(x a ,y a ,z a ) T is the homogeneous coordinate of the a-th ORB feature point of the dynamic target in the previous frame; the current camera pose is T c ∈SE(3); the homogeneous coordinate p a =(u a ,v a ,1) T ; m a ' is the homogeneous coordinate of the ORB feature point m a in the world coordinate system of the current frame;
[0046] S4-2-2. Obtain the pose of the dynamic target in the current frame according to the target transformation matrix and the homogeneous coordinate of the dynamic target in the previous frame, that is, the current pose of the dynamic target.
[0047] Provide an electronic device, the device includes:
[0048] A memory storing executable instructions; and
[0049] A processor configured to execute the executable instructions in the memory to implement a dynamic SLAM method based on reprojection error and depth estimation.
[0050] The beneficial effects of the present invention are as follows: the accuracy of identifying dynamic targets and the accuracy of the obtained trajectory of the present invention are high; when there are many static targets and they occupy a large area on the image plane, the present invention can extract more static feature points, significantly improving the recognition accuracy; the change range and maximum error of the relative pose error of the present invention are smaller, and the offset is smaller and more stable during operation; the present invention reduces the number of false positive samples, improves the recognition precision rate, optimizes the pose error, and improves the positioning accuracy of the camera in a dynamic scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flowchart of the present invention;
[0052] Figure 2 is a comparative diagram of relative pose errors, where, Figure 2 (a) is the relative pose error diagram of ORB-SLAM2, Figure 2 (b) is the relative pose error diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0053] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0054] As Figure 1 shown, S1, obtain an RGB image and extract ORB feature points of the RGB image;
[0055] S2, preprocess the RGB image to obtain a static background, a target label, the coordinates of pixel points restored to the world coordinate system, and pixel points after inter-frame matching;
[0056] S3, calculate based on the ORB feature points, static background, target label, the coordinates of pixel points restored to the world coordinate system, and pixel points after inter-frame matching of the RGB image, and obtain dynamic targets and static targets according to the reprojection error and depth residual;
[0057] S4, complete SLAM based on the dynamic target, static background, and static target, and estimate the pose of the dynamic target in real time;
[0058] S5, perform bundle adjustment optimization on the pose of the dynamic target.
[0059] Furthermore, the preprocessing of the RGB image is completed through the following three networks:
[0060] The depth generation network extracts the left-eye and right-eye image features respectively through a CNN channel of the PSMNet network, collects context information through the SPP module, connects the left-eye and right-eye feature maps into a cost volume function, inputs it into a 3D CNN for regularization, generates a depth map through disparity regression, and restores all pixel points in the RGB image to the world coordinate system according to the depth map to obtain the coordinates of the pixel points restored to the world coordinate system;
[0061] The instance segmentation network introduces the RPN module and the ROI Align module in the Mask R-CNN network to segment the RGB image, and divides the segmented RGB image into a static background and a target mask;
[0062] The optical flow network performs inter-frame matching on randomly sampled pixel points in the RGB image through the PWC-Net network to obtain the pixel points after inter-frame matching.
[0063] Furthermore, the specific implementation method of step S3 is as follows:
[0064] When the camera and the dynamic target move in different directions:
[0065] S3-1. According to the formula:
[0066] T i =T i-1 X i =T i-1 X i-1
[0067] Obtain the current pose T of the camera i ; where the transformation matrix of the previous-frame camera is X i-1 , and the pose is T i-1 , and i is the i-th ORB feature point of the target in the previous frame;
[0068] S3-2. According to the formula:
[0069]
[0070] Obtain The homogeneous coordinates in the world coordinate system where k represents the serial number of the ORB feature point; i-1 indicates that this point is obtained by inverse projection from the (i-1)-th frame image;
[0071] S3-3. According to the formula:
[0072]
[0073] Obtain the reprojection error e r ; where n is the number of target sampling points; is the pixel position of the target sampling point in the current frame; ||·|| 2 is the two - norm;;
[0074] S3 - 4. When the reprojection error is less than the set threshold, the target is static; when the reprojection error is greater than the set threshold, the target is dynamic;
[0075] When the camera and the dynamic target are moving in the same direction:
[0076] S3 - 5. According to the formula:
[0077]
[0078] obtain the position in the camera coordinate system of the current frame where j - 1 indicates that this point is obtained by inverse projection from the (j - 1)-th frame image; j represents the j-th ORB feature point of the target in the previous frame; k’ represents the serial number of the ORB feature point;
[0079] S3 - 6. According to the formula:
[0080]
[0081] obtain the pixel position of the target in the current frame found by the optical flow network and inverse project it into the three - dimensional world to obtain the position of the target in the camera coordinate system of the current frame where, is the pixel position of the target sampling point in the current frame;
[0082] S3 - 7. According to the formula:
[0083]
[0084] obtain the depth residual e d ; where, z(·) represents taking the value on the z - axis of the homogeneous coordinate; n’ represents the number of target sampling points;
[0085] Furthermore, S4 - 1. Input the static background and static targets into the visual odometry module to obtain the camera pose and ORB feature points;
[0086] S4 - 2. Input the dynamic target into the target pose estimation module to obtain the current pose of the dynamic target.
[0087] Furthermore, the specific implementation method of step S4 - 2 is as follows:
[0088] S4 - 2 - 1. According to the formula:
[0089] m a =X o m a '
[0090]
[0091] Obtain the dynamic target transformation matrix X o ; where, m a =(x a , y a , z a ) T is the homogeneous coordinate of the a-th ORB feature point of the dynamic target in the previous frame; the current pose of the camera is T c ∈ SE(3); the homogeneous coordinate p a =(u a , v a , 1) T ; m a ' is the homogeneous coordinate of the ORB feature point m a in the world coordinate system of the current frame;
[0092] S4-2-2. Obtain the pose of the dynamic target in the current frame according to the target transformation matrix and the homogeneous coordinate of the dynamic target in the previous frame, that is, the current pose of the dynamic target.
[0093] As Figure 2 shown, compared with ORB-SLAM2, the variation range and maximum error of the relative pose error of the present invention are smaller, indicating that ORB-SLAM2 runs more stably in a dynamic environment. At the same time, it can be noted that ORB-SLAM2 has a larger error at the initial stage of the sequence and gradually flattens out as time progresses, which means that ORB-SLAM2 has a large deviation during the initialization process, resulting in a larger error in the finally generated trajectory, that is, the ATE value is larger.
[0094] In an embodiment of the present invention, the visual odometry module uses the same design as ORB-SLAM 2 to solve the camera pose.
[0095] The present invention uses the absolute trajectory error and the relative pose error as evaluation indicators, and experiments are carried out on some video sequences that meet the conditions. The common feature of these sequences is that there are a large number of dynamic targets in the environment and the camera is in a moving state. According to the formula:
[0096]
[0097]
[0098] Obtain the absolute trajectory error ATE and relative pose error RPE of each dynamic target pose; where, T esti,μ represents the measured value of the μ-th dynamic target pose, T gt,μdenote the true value of the $\mu$-th pose; $N$ represents the total number of frames; $gt$ represents the true pose value; $\Delta t$ represents the time difference; represents the transformation matrix of the pose of the dynamic target from the $\mu$-th moment to the $\mu + \Delta t$-th moment; $v$ represents converting the skew-symmetric matrix into the corresponding vector; The performance under the Tracking dataset is shown in Table 1. The parts marked with "-" indicate that the algorithm cannot complete the tracking in this sequence.
[0099] Table 1
[0100]
[0101] It can be obtained from Table 1 that the trajectory accuracy obtained by the present invention is generally higher, but the performance is inferior to that of ORB-SLAM2 in some sequences. Although sequence 0000 is an outdoor scene, there are few dynamic targets and the camera moves a short distance. It belongs to a small-scale and low-dynamic scene among all sequences. ORB-SLAM2 can still maintain good performance in such scenes. At the same time, the small-scale scene means that the dynamic target is closer to the camera and occupies more pixels on the image plane, resulting in a smaller static background area that the present invention can refer to, and finally the performance is slightly inferior to ORB-SLAM2; The data of sequence 0020 was collected on a congested highway. This scene belongs to a large-scale scene, where vehicles occupy most of the field of view, and most vehicles are moving slowly. At the same time, the texture information in the area outside the vehicles is not rich. In this environment, the number of static feature points that VDO-SLAM can refer to is small and the distribution is relatively concentrated. Therefore, the performance effect is the worst. Although the present invention retains the feature points of some static targets, the lack of a large area also makes the effect not good, while ORB-SLAM2 can still maintain good performance in such low-dynamic scenes. In the remaining sequences, the present invention can maintain a high positioning accuracy. Compared with the original ORB-SLAM2, the present invention has a relatively significant improvement, and the higher the dynamic degree, the more obvious the improvement. It should be noted that in different sequences, the improvement of the present invention relative to VDO-SLAM varies. When the proportion of dynamic targets in the environment is relatively large, the improvement is not obvious. For example, sequence 0003 is a vehicle overtaking scene on the road. There are many dynamic targets and almost no static targets in this sequence. In this scene, the present invention has to filter out almost all vehicles, which is similar to the filtering strategy effect of VDO-SLAM. In a scene with a relatively small proportion of dynamic targets, such as sequence 0007, there are many static targets and they occupy a large area on the image plane. At this time, the present invention can extract more static feature points, so the algorithm can have a more obvious improvement.
[0102] According to the formula:
[0103]
[0104] The precision P and recall rate R are obtained; wherein, each frame of the sample is segmented to obtain true positive samples TP, false positive samples FP, and false negative samples FN;
[0105] According to the precision and recall rate, the accuracy of dynamic target recognition of the present invention and VDO-SLAM is compared, and the results are shown in Table 2, where Precision is the precision rate and Recall is the recall rate.
[0106] Table 2
[0107]
[0108] The sequences selected in the experiment all contain a sufficient number of positive and negative samples. From the data, it is not difficult to see that while the recall rate of the present invention is basically the same as that of VDO-SLAM, the number of false positive samples is significantly reduced, and the recognition precision is improved. Combining Table 1, it can be found that the improvement of the precision rate indirectly improves the trajectory accuracy and has a certain optimization effect on the pose error.
[0109] In the experiment, sequence 0000 belongs to a small-scale scenario where the object is close to the camera, enabling stable dynamic target recognition and tracking; sequence 0011 is a following-vehicle driving scenario with a large number of dynamic targets and a large scene scale. However, due to the close distance between the leading vehicle and the camera and the occlusion of some targets at a relatively far distance, the number of positive samples is reduced, so the recall rate is relatively high.
[0110] When static targets have large inter-frame differences due to reasons such as light changes, perspective changes, and image blurring, they are easily misrecognized as dynamic targets. In the experiment, such false positive detections mainly occur in the two side regions of the image. This is because the input of the reprojection error and depth estimation is the three-dimensional coordinates of the point set, and the solution of the three-dimensional coordinates depends on the depth map generated by PWC-Net. At the same time, the prediction of depth in the edge region of the image is discrete and has a large error. Therefore, the error result is easily exceeded the threshold, resulting in static targets being misrecognized as dynamic targets. In sequences with a small proportion of static targets, the number of false positive detections decreases, so the precision rate increases.
[0111] The accuracy of recognizing dynamic targets and the obtained trajectory accuracy of the present invention are high; when there are many static targets and they occupy a large area on the image plane, the present invention can extract more static feature points, significantly improving the recognition accuracy; the change range and maximum error of the relative pose error of the present invention are smaller, with a smaller offset and more stable during operation; the present invention reduces the number of false positive samples, improves the recognition precision rate, has an optimization effect on the pose error, and improves the positioning accuracy of the camera in a dynamic scene.
Claims
1. A dynamic SLAM method based on reprojection error and depth estimation, characterized in that It includes the following steps: S1. Obtain an RGB image and extract the ORB feature points of the RGB image; S2. Preprocess the RGB image to obtain a static background, target labels, the coordinates of pixel points restored to the world coordinate system, and pixel points after inter-frame matching; S3. Calculate based on the ORB feature points, static background, target labels, the coordinates of pixel points restored to the world coordinate system, and pixel points after inter-frame matching of the RGB image, and obtain dynamic targets and static targets according to the reprojection error and depth residual. The specific implementation method is as follows: When the camera and the dynamic target move in different directions: S3-1. According to the formula: T i = T i-1 X i = T i-1 X i-1 Obtain the current pose T of the camera i ; where the transformation matrix of the previous-frame camera is X i-1 , the pose is T i-1 , and i is the i-th ORB feature point of the target in the previous frame; S3-2. According to the formula: Obtain Homogeneous coordinates in the world coordinate system where k represents the serial number of the ORB feature point; i - 1 indicates that this point is obtained by back-projection from the (i - 1)-th frame image S3-3. According to the formula: Obtain the reprojection error e r ; where n is the number of target sampling points; is the pixel position of the target sampling point in the current frame; ||·||2 is the two-norm; S3-4. When the reprojection error is less than the set threshold, the target is static; when the reprojection error is greater than the set threshold, the target is dynamic; When the camera and the dynamic target move in the same direction: S3-5. According to the formula: Obtain the position in the camera coordinate system of the current frame Among them, j - 1 indicates that this point is obtained by inverse projection from the (j - 1)-th frame image; j represents the j-th ORB feature point of the target in the previous frame; k' represents the serial number of the ORB feature point S3-6. According to the formula: Obtain the pixel position of the target in the current frame found by the optical flow network and back-project it into the three-dimensional world to obtain the position of the target in the camera coordinate system of the current frame wherein, is the pixel position of the target sampling point in the current frame; S3-7. According to the formula: Obtain the depth residual e d ; where z(·) represents taking the value on the z-axis of the homogeneous coordinates; n' represents the number of target sampling points; S3-8. When the depth residual exceeds the set threshold, the target is a dynamic target; when the depth residual is lower than the set threshold, the target is a static target; S4. Complete SLAM based on the dynamic target, static background, and static target, and estimate the pose of the dynamic target in real time; S5. Optimize the pose of the dynamic target by bundle adjustment.
2. A dynamic SLAM method based on reprojection error and depth estimation according to claim 1, characterized in that The preprocessing of the RGB image is completed through the following three networks: Depth generation network: Extract the left-eye and right-eye image features through one CNN channel of the PSMNet network respectively, collect context information through the SPP module, connect the left-eye and right-eye feature maps into a cost volume function, input it into a 3D CNN for regularization, generate a depth map through disparity regression, and restore all pixel points in the RGB image to the world coordinate system according to the depth map to obtain the coordinates of pixel points restored to the world coordinate system; Instance segmentation network: Introduce the RPN module and ROI Align module into the Mask R-CNN network to segment the RGB image, and divide the segmented RGB image into a static background and a target mask; Optical flow network: Perform inter-frame matching on randomly sampled pixel points in the RGB image through the PWC-Net network to obtain pixel points after inter-frame matching.
3. A dynamic SLAM method based on reprojection error and depth estimation according to claim 2, characterized in that The specific implementation method of step S4 is as follows: S4-1. Input the static background and static target into the visual odometry module to obtain the camera pose and ORB feature points; S4-2. Input the dynamic target into the target pose estimation module to obtain the current pose of the dynamic target.
4. A dynamic SLAM method based on reprojection error and depth estimation according to claim 3, characterized in that, The specific implementation method of step S4-2 is as follows: S4-2-1. According to the formula: m a = X o m a ' Obtain the dynamic target transformation matrix X o ; where m a =(x a , y a , z a ) Τ is the homogeneous coordinate of the 3D point of the a-th ORB feature point of the dynamic target in the previous frame, which is a four-dimensional vector and the last dimension is 1; the current pose of the camera is T c ∈ SE(3); the homogeneous coordinate p a =(u a , v a , 1) Τ is the projection of the a-th ORB feature point on the normalized plane of the current frame camera; m a ' is the homogeneous coordinate of the ORB feature point m a in the world coordinate system of the current frame; S4-2-2. Obtain the pose of the dynamic target in the current frame based on the target transformation matrix and the homogeneous coordinates of the dynamic target in the previous frame, that is, the current pose of the dynamic target.
5. An electronic device, characterized in that, The device includes: A memory storing executable instructions; and A processor configured to execute the executable instructions in the memory to implement the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Autonomous pose measurement method based on SLAM technology
CN112902953A
Image processing method and related equipment
CN114170290A