Mobile robot pose estimation method based on mixed point and twin line feature reprojection joint optimization
By combining the joint optimization method of hybrid point and twin line feature reprojection, the accuracy and robustness issues of trajectory estimation and self-localization of the RGB-DVPE system in complex environments are solved, and accurate estimation of the robot's pose is achieved.
Patent Information
- Application Number
- CN202510695240.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-10-03
AI Technical Summary
Existing RGB-DVPE systems based on point-line feature fusion have difficulty achieving accurate trajectory estimation and robust self-localization in complex environments with low texture, low structure, occlusion, perspective and illumination changes.
A joint optimization method of hybrid point and twin line feature reprojection is adopted. Through ORB point features and LSD line features extraction, matching and description, combined with RGB-D depth information, a hybrid point feature and twin line feature reprojection error function is constructed, and joint optimization is performed to estimate the pose of the mobile robot.
The optimal pose estimation is achieved by fusing and complementing the advantages of point and line features in complex environments, ensuring the trajectory estimation accuracy and self-positioning robustness of the visual navigation robot.
Smart Images

Figure CN120747210A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot positioning and navigation, and in particular relates to a mobile robot pose estimation method based on joint optimization of hybrid point and twin line feature reprojection. Background Art
[0002] Pose estimation is a key technology for leveraging natural features to achieve robot positioning in visual SLAM (VSLAM) systems. Its performance directly affects the accuracy, efficiency, and robustness of visually navigated robotic systems. Typically, pose estimation techniques in VSLAM focus on estimating the vehicle's own motion from continuous camera images. Compared to traditional wheel odometry or inertial measurement unit solutions, this vision-based pose estimation (VPE) technology offers significant advantages in scalability. Based on optimization error, existing VPE solutions can be roughly divided into two types: direct photometric-based methods and indirect feature-based methods.
[0003] Direct methods are based on the strong assumption of illumination invariance and directly estimate the pose by minimizing the error in image pixel intensity. Compared with indirect methods, direct methods are more sensitive to changes in viewpoint and illumination, which makes them disadvantageous in eliminating cumulative errors. Indirect methods can be divided into monocular VPE, binocular VPE, and RGB-D VPE. These three systems are developed using similar and identical schemes and mainly consist of four main threads: feature extraction, feature matching, feature tracking, and local mapping. Due to the emergence of low-cost depth cameras with pixel-by-pixel depth measurement (such as Kinect and RealSense), RGB-D based VPE systems have attracted more and more attention and research, which has effectively solved the existing scale problem.
[0004] However, in scenes with low artificial texture, extreme illumination variations, and extreme viewpoint changes that lack unique point features, pose estimation accuracy often degrades and high performance cannot be maintained. In addition, line features may be partially occluded, their endpoints are not always repeatable across views, and their appearance can vary significantly when the camera viewpoint changes, making it more difficult to describe line features in an image than point features. Therefore, current RGB-DVPE systems based on the fusion of point and line features have difficulty achieving accurate trajectory estimation and robust self-localization in complex workspaces with challenging factors such as low texture, low structure, occlusion, viewpoint, and illumination variations in mobile robot navigation applications. Summary of the Invention
[0005] In response to the above-mentioned deficiencies in the existing technologies, the present invention provides a mobile robot pose estimation method based on the joint optimization of hybrid point and twin line feature reprojection, aiming to achieve optimal pose estimation with the complementary advantages of point and line feature fusion, and to ensure the trajectory estimation accuracy and self-positioning robustness of visual navigation robots in various challenging and complex environments.
[0006] In order to achieve the above technical objectives, the present invention provides the following technical solutions:
[0007] A mobile robot pose estimation method based on joint optimization of hybrid point and twin line feature reprojection includes the following steps:
[0008] S1. Collect RGB-D images as input visual images through the on-board camera; for the current input RGB image, use the ORB point feature extraction algorithm and the LSD line feature extraction algorithm to extract ORB points and LSD line features from two consecutive navigation images;
[0009] S2. Based on the rBRIEF point feature description algorithm and the LBD line feature description algorithm, the point and line feature descriptors are calculated respectively, and the point and line feature data are associated to obtain the ORB point feature matching pairs and the LSD line feature matching pairs;
[0010] S3. Verify the depth information of the RGB-D input and divide the obtained ORB point feature matching pairs into 3D-2D matching pairs without depth values and 3D-3D matching pairs with depth values;
[0011] S4, perform depth measurement on the LSD line feature matching pair and construct its virtual right eye line;
[0012] S5. Fuse 3D-2D and 3D-3D ORB point feature matching pairs to construct a hybrid point feature reprojection error function; integrate the virtual right eye line and consider both RGB and depth cues to construct a twin line feature reprojection error function;
[0013] S6. Based on the reprojection error function of hybrid point features and twin line features, a joint unified error optimization model is constructed to simultaneously correct the point and line reprojection errors and estimate the optimal pose of the mobile robot.
[0014] Furthermore, in step S2, the specific process of obtaining the ORB point feature matching pair is as follows:
[0015] The ORB feature extraction method is used to extract the target image I at time t+1. t+1 and the reference image I at time t t The feature sets of the extracted points are recorded as:
[0016]
[0017] in, Represents the set of feature points extracted from the target image, a in total; Represents the set of feature points extracted from the reference image, a total of b;
[0018] Based on a simple brute force matching algorithm, the correspondence of each pair of ORB feature points is quickly checked. A standard ratio test threshold is set. When the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than the standard ratio test threshold, it is considered to be a correctly matched ORB feature matching pair and is retained.
[0019] Furthermore, in step S2, the specific process of obtaining the LSD line feature matching pair is as follows:
[0020] The LSD line detector is used to obtain the reference image I t and target image I t+1 The two sets of extracted line feature sets are recorded as:
[0021]
[0022] Among them, M t and M t+1 Represents image I t and I t+1 The number of LSD line features extracted from ;( t g i,0 , t g i,1 ) represents line features t l i endpoints;
[0023] The LBD descriptors are constructed for the target image LSD line features and the corresponding candidate reference image LSD line features respectively, and the Euclidean distance between the target LSD line and the reference candidate LSD line descriptors is calculated. If the minimum Euclidean distance d between the two is less than the set threshold T d , then perform LSD line feature data association to obtain a pair of LSD line feature matching pairs.
[0024] Furthermore, step S3 is specifically as follows:
[0025] Given a reference frame π r and a current frame π c , the set of matching point pairs obtained by point feature association is recorded as:
[0026]
[0027] in, Represents a pair of matching points; N P Indicates the number of matching point pairs;
[0028] Given an RGB-D input frame FID , the input frame includes the reference frame and the current frame, and contains the aligned depth image and RGB image; record I(p i ) and D(p i ) are the i-th pixel p i Grayscale value and depth value;
[0029] Based on the current frame π c RGB image I t Observed matching points Define the set of 2D matching points without depth value as: The set of 3D matching points with depth values is: Among them, N PW +N PWO =N P .
[0030] Furthermore, step S4 is specifically as follows:
[0031] Set the current frame π c As the left eye frame O L , based on π c The depth value in constructs the corresponding virtual right eye frame O R Specifically:
[0032] Assume p i =(u i , v i ) T is the left eye frame O L The i-th left eye pixel in the image is based on the x-direction focal length f x and baseline B, which corresponds to the virtual right eye frame O R The virtual right eye pixel on in
[0033] The corresponding virtual right eye pixel is obtained by detecting each pixel on the line feature in the left eye frame, that is, the corresponding virtual right eye pixel is obtained. L The virtual right eye line of the detected line feature.
[0034] Furthermore, in step S5, the specific process of constructing the mixed point feature reprojection error function is as follows:
[0035] Define N PWO The reprojection error e1 of a 2D matching point without depth value is:
[0036]
[0037] Where ξ represents the camera motion, which is a six-dimensional vector of Lie algebra; K is the 3×3 camera internal matrix, which is obtained in advance through camera calibration; is the world coordinate system πw A three-dimensional observation point in ; exp(ξ^) is expressed as r to π c The pose transformation matrix is expressed as follows:
[0038]
[0039] Define N with depth value PW The reprojection error e2 of the 3D matching points is:
[0040]
[0041] The calculation formula for the optimal camera pose based on the mixed point feature reprojection error is:
[0042]
[0043] Furthermore, in step S5, the specific process of constructing the twin line feature reprojection error function is as follows:
[0044] definition w L b =( w G b,0 , w G b,1 ) is the world frame π w 3D line features in, where ( w G b,0 , w G b,1 )express w L b A pair of 3D endpoints of a line feature; given a reference frame π r and the current frame π c , the observed two-dimensional line feature matching pair set is expressed as:
[0045] ML={[ r l b =( r g b,0 , r g b,1 ), c l b =( c g b,0 , c g b,1 )], b=1,…,N ML};
[0046] in,[ r l b , c l b ] represents a pair of matching line features; N MLIndicates the number of line feature matching pairs; ( r g b,0 , r g b,1 )and( c g b,0 , c g b,1 ) represent the reference frame π r Line features observed in r l b The current frame π c Line features observed in c l b 2D endpoints of;
[0047] Reuse ( c g′ b,0 , c g′ b,1 ) represents the c Medium heavy projection line feature c l′ b 2D endpoints; based on the standard pinhole camera model, we get:
[0048]
[0049] set up corresponds to the two-dimensional endpoints The observation line features are aligned with the secondary coordinates of c l b Normalized line coefficient c η b The calculation formula is:
[0050]
[0051] Then use the sum of the distances from the point to the line to represent the endpoints of the reprojected line ( c g′ b,0 , c g′ b,1 ) and observation line characteristics c l b The line reprojection error between (e3+e4), that is:
[0052]
[0053] Among them, d b ( c g′ b,0 , c l b ),d b ( c g′ b,1 , c l b) represent the endpoints of the reprojection line c g′ b,0 、 c g′ b,1 To observation line features c l b distance;
[0054] For the current frame π c The observed N ML Line features, total line reprojection error E L It is obtained by accumulating the distance from the point to the line. The formula is expressed as:
[0055]
[0056] Line features c l b The endpoint ( c g b,0 , c g b,1 )and c l′ b The endpoint ( c g′ b,0 , c g′ b,1 ) as the starting position, and use the binary search algorithm to find the pixel p whose depth value is less than the set depth threshold Ω and greater than 0. i As a line c l b and c l′ b The new endpoint of ; if the depth value does not meet the conditions, no virtual right eye line is generated for the line feature; the final set of virtual right eye lines is:
[0057] VRE={[ R l b =( R g b,0 , R g b,1 ), R l′ b =( R g′ b,0 , R g′ b,1 )]};
[0058] Where b = 1, 2, ..., N VRE , N VRE Indicates the number of virtual right eye lines; R l b and R l′ b Represent the virtual right eye frame O RThe virtual right eye line observed in the left eye frame O and the reprojected virtual right eye line, which correspond to the left eye frame O L Line features in c l b and c l′ b ;( R g b,0 , R g b,1 )and( R g′ b,0 , R g′ b,1 ) represent lines R l b and R l′ b 2D endpoints of;
[0059] Then calculate the virtual right eye line R l b The normalization coefficient of :
[0060]
[0061] Based on the depth measurement, the reprojection error E of the virtual right eye line is obtained R The calculation formula is:
[0062]
[0063] Combining the left eye line reprojection error and the virtual right eye line reprojection error, we get the twin line feature reprojection error function E l,cr (ξ), then the optimal camera pose based on the twin line feature reprojection error function is:
[0064]
[0065] Furthermore, step S6 is specifically as follows:
[0066] Perform joint optimization of point and line feature reprojection to obtain the optimal camera pose calculation formula:
[0067]
[0068] The solution is solved by iteratively following the Gauss-Newton optimization algorithm in the manifold tangent space se(3); first calculate the pose update:
[0069] Δξ=-(J T WJ) -1 J T WE cr ;
[0070] Among them, the error vector E cr Contains all reprojection errors Ep,cr and E l,cr ; W and J are E cr The diagonal weight matrix and Jacobian matrix of ;
[0071] Then the camera pose is iteratively updated, and the formula is expressed as:
[0072]
[0073] Iterate and update until convergence to output the optimal robot pose ξ.
[0074] Based on the above technical solution, the present invention has at least the following beneficial effects:
[0075] The method proposed in the present invention first constructs the correspondence between point and line features, which can explore the similarity of geometric constraints between adjacent frames; secondly, by fusing 3D-2D and 3D-3D matching point pairs, a hybrid point feature reprojection error is constructed to solve the problem of low estimation accuracy of some point pairs due to the lack of 3D depth values; at the same time, the virtual right eye line generated by depth measurement is used to model the line feature twin reprojection error, so as to enhance the pose constraint in the optimization process; finally, a unified joint error optimization model is constructed based on the point and line feature reprojection error function, and the optimal solution is calculated through iterative Gauss-Newton optimization; the method proposed in the present invention can simultaneously correct the point and line reprojection errors, realize the optimal pose estimation with the complementary advantages of point and line feature fusion, and ensure the trajectory estimation accuracy and self-positioning robustness of the visual navigation robot in various challenging and complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 This is the overall flow chart of the mobile robot pose estimation method based on joint optimization of hybrid point and twin line feature reprojection proposed in the present invention;
[0077] Figure 2 Schematic diagram of modeling the mixed point feature reprojection error in the method proposed in the present invention;
[0078] Figure 3 Schematic diagram of twin line feature reprojection error modeling in the method proposed in this invention;
[0079] Figure 4 Comparison of the absolute trajectory error results of the method proposed in this invention and the existing method. DETAILED DESCRIPTION
[0080] In order to make the above-mentioned objects, features and advantages of the present invention more clearly understood, the following Figure 1-4 The present invention is further described in detail with specific implementation methods, so that the application can fully understand how to use technical means to solve technical problems and achieve technical effects and implement them accordingly.
[0081] Those skilled in the art will appreciate that all or part of the steps in the above-mentioned embodiment methods can be accomplished by instructing the relevant hardware through a program. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0082] like Figure 1 As shown in FIG, the present invention proposes a mobile robot pose estimation method based on joint optimization of hybrid point and twin line feature reprojection, which specifically includes:
[0083] S1. Collect RGB-D images as input visual images through the on-board camera; for the current input RGB image, use the ORB point feature extraction algorithm and the LSD line feature extraction algorithm to extract ORB points and LSD line features from two consecutive navigation images;
[0084] In this embodiment, the vehicle-mounted camera is installed on a mobile robot to capture images of the natural environment in an environment without any landmarks. Therefore, the posture of the robot is assumed to be consistent with the posture of the vehicle-mounted camera in this application.
[0085] S2. Based on the rBRIEF point feature description algorithm and the LBD line feature description algorithm, the point and line feature descriptors are calculated respectively, and the point and line feature data are associated to obtain the ORB point feature matching pairs and the LSD line feature matching pairs;
[0086] As a preferred embodiment, in step S2, the specific process of obtaining the ORB point feature matching pair is as follows:
[0087] The ORB feature extraction method is used to extract the target image I at time t+1. t+1 and the reference image I at time t t The feature sets of the extracted points are recorded as:
[0088]
[0089] in, Represents the set of feature points extracted from the target image, a in total; Represents the set of feature points extracted from the reference image, a total of b;
[0090] Based on a simple brute force matching algorithm, the correspondence of each pair of ORB feature points is quickly checked. A standard ratio test threshold is set. When the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than the standard ratio test threshold, it is considered to be a correctly matched ORB feature matching pair and is retained.
[0091] After constructing the ORB point feature correspondence, the LSD line feature extraction algorithm is used to extract the feature from the reference image I t
[0092] and target image I t+1 Extract two sets of line feature sets, specifically:
[0093] The LSD line detector is used to obtain the reference image I t and target image I t+1 The two sets of extracted line feature sets are recorded as:
[0094]
[0095] Among them, M t and M t+1 Represents image I t and I t+1 The number of LSD line features extracted from ;( t g i,0 , t g i,1 ) represents line features t l i endpoints;
[0096] The LBD descriptors are constructed for the target image LSD line features and the corresponding candidate reference image LSD line features respectively, and the Euclidean distance between the target LSD line and the reference candidate LSD line descriptors is calculated. If the minimum Euclidean distance d between the two is less than the set threshold T d , then perform LSD line feature data association to obtain a pair of LSD line feature matching pairs.
[0097] S3. Verify the depth information of the RGB-D input and divide the obtained ORB point feature matching pairs into 3D-2D matching pairs without depth values and 3D-3D matching pairs with depth values;
[0098] As a preferred implementation mode, step S3 specifically includes:
[0099] like Figure 2 As shown, given a reference frame π r and a current frame π c , the set of matching point pairs obtained by point feature association is recorded as:
[0100]
[0101] in, Represents a pair of matching points; N P Indicates the number of matching point pairs;
[0102] Given an RGB-D input frame F ID , the input frame includes the reference frame and the current frame, and contains the aligned depth image and RGB image; record I(p i ) and D(p i ) are the i-th pixel p i Grayscale value and depth value;
[0103] Based on the current frame π c RGB image I t Observed matching points Define the set of 2D matching points without depth value as: The set of 3D matching points with depth values is: Among them, N PW +N PWO =N P .
[0104] S4, perform depth measurement on the LSD line feature matching pair and construct its virtual right eye line;
[0105] As a preferred implementation, in this embodiment, step S4 is specifically as follows:
[0106] like Figure 3 As shown, the current frame π c As the left eye frame O L , based on π c The depth value in constructs the corresponding virtual right eye frame O R Specifically:
[0107] Assume p i =(u i , v i ) T is the left eye frame O L The i-th left eye pixel in the image is based on the x-direction focal length f x and baseline B, which corresponds to the virtual right eye frame O R The virtual right eye pixel on in
[0108] The corresponding virtual right eye pixel is obtained by detecting each pixel on the line feature in the left eye frame, that is, the corresponding virtual right eye pixel is obtained. LThe virtual right eyeline is constructed based on the detected line features. Existing classic line reprojection errors only consider RGB cues, which can lead to insufficient robustness and accuracy in challenging scenes with varying lighting and occlusions. Therefore, this application utilizes depth measurements to construct a virtual right eyeline, thereby strengthening pose constraints during the optimization process and achieving accurate and reliable pose estimation.
[0109] S5. Fuse 3D-2D and 3D-3D ORB point feature matching pairs to construct a hybrid point feature reprojection error function; the existing pose estimation technology based on point feature reprojection error optimization is implemented by calculating 3D-3D point pairs based on iterative closest point (ICP) or 3D-2D point pairs using perspective-n-point (PnP). However, since some point feature pairs extracted from low-texture and illumination-changing environments lack depth values, their accuracy needs to be improved. In order to alleviate this problem, the present invention designs a hybrid point feature reprojection error function that uses valid 3D-3D and 3D-2D matching pairs at the same time; the calculation process of the hybrid point feature reprojection error function in this application is specifically as follows:
[0110] like Figure 2 As shown, define N PWO The reprojection error e1 of a 2D matching point without depth value is:
[0111]
[0112] Where ξ represents the camera motion, which is a six-dimensional vector of Lie algebra; K is the 3×3 camera internal matrix, which is obtained in advance through camera calibration; and Represents the world frame π w 3D reference point features, green points and yellow dots Represent the observed matching points with and without depth values, respectively. The blue points and Respectively and In π c Reprojected 2D points is the world coordinate system π w A three-dimensional observation point in ; exp(ξ^) is expressed as r to π c The pose transformation matrix is expressed as follows:
[0113]
[0114] Define N with depth value PWThe reprojection error e2 of the 3D matching points is:
[0115]
[0116] The calculation formula for the optimal camera pose based on the mixed point feature reprojection error is:
[0117]
[0118] So far, the calculation method of the mixed point feature reprojection error function has been obtained. Next, this application integrates the virtual right eye line and considers RGB and depth clues at the same time to construct a twin line feature reprojection error function; in the currently widely used model based on line feature reprojection optimization, its error function minimizes the distance between the reprojection endpoint and the observation line. This is an effective pose estimation method based on line features, but because it only considers RGB clues, its accuracy will still decrease under changes in lighting and perspective and occlusion. To this end, the present invention constructs a twin line feature reprojection error function, which constructs a virtual right eye line by utilizing depth observation (i.e., depth measurement) to enhance pose constraints during the optimization process, thereby achieving accurate and reliable VPE, such as Figure 3 As shown, the specific process is:
[0119] definition w L b =( w G b,0 , w G b,1 ) is the world frame π w 3D line features in, where ( w G b,0 , w G b,1 )express w L b A pair of 3D endpoints of a line feature; given a reference frame π r and the current frame π c , the observed two-dimensional line feature matching pair set is expressed as:
[0120] ML={[ r l b =( r g b,0 , r g b,1 ), c l b =( c g b,0 , c g b,1 )], b=1,…,N ML};
[0121] in,[r l b , c l b ] represents a pair of matching line features; N ML Indicates the number of line feature matching pairs; ( r g b,0 , r g b,1 )and( c g b,0 , c g b,1 ) represent the reference frame π r Line features observed in r l b The current frame π c Line features observed in c l b 2D endpoints of;
[0122] Reuse ( c g′ b,0 , c g′ b,1 ) represents the c Medium heavy projection line feature c l′ b 2D endpoints; based on the standard pinhole camera model, we get:
[0123]
[0124] set up corresponds to the two-dimensional endpoints The observation line features are aligned with the secondary coordinates of c l b Normalized line coefficient c η b The calculation formula is:
[0125]
[0126] Then use the sum of the distances from the point to the line to represent the endpoints of the reprojected line ( c g′ b,0 , c g′ b,1 ) and observation line characteristics c l b The line reprojection error between (e3+e4), that is:
[0127]
[0128] Among them, d b ( c g′ b,0 , c l b ),db ( c g′ b,1 , c l b ) represent the endpoints of the reprojection line c g′ b,0 、 c g′ b,1 To observation line features c l b distance;
[0129] For the current frame π c The observed N ML Line features, total line reprojection error E L It is obtained by accumulating the distance from the point to the line. The formula is expressed as:
[0130]
[0131] Line features c l b The endpoint ( c g b,0 , c g b,1 )and c l′ b The endpoint ( c g′ b,0 , c g′ b,1 ) as the starting position, and use the binary search algorithm to find the pixel p whose depth value is less than the set depth threshold Ω and greater than 0. i As a line c l b and c l′ b The new endpoint of ; if the depth value does not meet the conditions, no virtual right eye line is generated for the line feature; the final set of virtual right eye lines is:
[0132] VRE={[ R l b =( R g b,0 , R g b,1 ), R l′ b =( R g′ b,0 , R g′ b,1 )]};
[0133] Where b = 1, 2, ..., N VRE , N VRE Indicates the number of virtual right eye lines; R l band R l′ b Represent the virtual right eye frame O R The virtual right eye line observed in the left eye frame O and the reprojected virtual right eye line, which correspond to the left eye frame O L Line features in c l b and c l′ b ;( R g b,0 , R g b,1 )and( R g′ b,0 , R g′ b,1 ) represent lines R l b and R l′ b 2D endpoints of;
[0134] Then calculate the virtual right eye line R l b The normalization coefficient of :
[0135]
[0136] Based on the depth measurement, the reprojection error E of the virtual right eye line is obtained R The calculation formula is:
[0137]
[0138] Combining the left eye line reprojection error and the virtual right eye line reprojection error, we get the twin line feature reprojection error function E l,cr (ξ), then the optimal camera pose based on the twin line feature reprojection error function is:
[0139]
[0140] S6. Based on the reprojection error function of hybrid point features and twin line features, a joint unified error optimization model is constructed to simultaneously correct the point and line reprojection errors and estimate the optimal pose of the mobile robot.
[0141] Through the previous steps, the reprojection errors of the above point features and line features are obtained. The present invention finally establishes a new unified error optimization model, which integrates hybrid point reprojection and twin line reprojection. That is, as a preferred embodiment, step S6 is specifically as follows:
[0142] Perform joint optimization of point and line feature reprojection to obtain the optimal camera pose calculation formula:
[0143]
[0144] The solution is solved by iteratively following the Gauss-Newton optimization algorithm in the manifold tangent space se(3); first calculate the pose update:
[0145] Δξ=-(J T WJ) -1 J T WE cr ;
[0146] Among them, the error vector E cr Contains all reprojection errors E p,cr and E l,cr ; W and J are E cr The diagonal weight matrix and Jacobian matrix of ;
[0147] Then the camera pose is iteratively updated, and the formula is expressed as:
[0148]
[0149] Iterate and update until convergence to output the optimal robot pose ξ.
[0150] The entire process of the method proposed in the present invention has been introduced so far; this embodiment also provides a comparison of estimation results with existing traditional methods. Figure 4 The following figure shows the experimental results of motion estimation using the proposed method on the fr3 / s / xyz sequence of the RGB-D TUM dataset, along with a comparison of motion estimation using the ORB-SLAM3 algorithm without loop closure detection. This comparison demonstrates that the proposed method, based on joint optimization of point and line feature reprojection, can accurately estimate robot trajectory in complex and challenging environments.
[0151] In summary, the method proposed in the present invention can simultaneously correct the point-line reprojection error, achieve the optimal pose estimation with the complementary advantages of point-line feature fusion, and ensure the trajectory estimation accuracy and self-positioning robustness of the visual navigation robot in various challenging and complex environments.
[0152] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.
[0153] The logic and / or steps represented in the flowchart or otherwise described herein may be considered, for example, as an ordered list of executable instructions for implementing logical functions, and may be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device).
[0154] The above embodiments provide a detailed introduction to the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A mobile robot pose estimation method based on joint optimization of hybrid point and twin line feature reprojection, characterized by: The specific steps include: S1. Collect RGB-D images as input visual images through the on-board camera; for the current input RGB image, use the ORB point feature extraction algorithm and the LSD line feature extraction algorithm to extract ORB points and LSD line features from two consecutive navigation images; S2. Based on the rBRIEF point feature description algorithm and the LBD line feature description algorithm, the point and line feature descriptors are calculated respectively, and the point and line feature data are associated to obtain the ORB point feature matching pairs and the LSD line feature matching pairs; S3. Verify the depth information of the RGB-D input and divide the obtained ORB point feature matching pairs into 3D-2D matching pairs without depth values and 3D-3D matching pairs with depth values; S4, perform depth measurement on the LSD line feature matching pair and construct its virtual right eye line; S5. Fuse 3D-2D and 3D-3D ORB point feature matching pairs to construct a hybrid point feature reprojection error function; integrate the virtual right eye line and consider both RGB and depth cues to construct a twin line feature reprojection error function; S6. Based on the reprojection error function of hybrid point features and twin line features, a joint unified error optimization model is constructed to simultaneously correct the point and line reprojection errors and estimate the optimal pose of the mobile robot.
2. The method for mobile robot pose estimation based on joint optimization of hybrid point and twin line feature reprojection according to claim 1, characterized in that: In step S2, the specific process of obtaining the ORB point feature matching pair is as follows: The ORB feature extraction method is used to extract the target image I at time t+1. t+1 and the reference image I at time t t The feature sets of the extracted points are recorded as: in, Represents the set of feature points extracted from the target image, a in total; Represents the set of feature points extracted from the reference image, a total of b; Based on a simple brute force matching algorithm, the correspondence of each pair of ORB feature points is quickly checked. A standard ratio test threshold is set. When the ratio of the nearest neighbor distance to the second nearest neighbor distance is less than the standard ratio test threshold, it is considered to be a correctly matched ORB feature matching pair and is retained.
3. The method for mobile robot pose estimation based on joint optimization of hybrid point and twin line feature reprojection according to claim 1, characterized in that: In step S2, the specific process of obtaining the LSD line feature matching pair is as follows: The LSD line detector is used to obtain the reference image I t and target image I t+1 The two sets of extracted line feature sets are recorded as: Among them, M t and M t+1 Represents image I t and I t+1 The number of LSD line features extracted from ;( t g i,0 , t g i,1 ) represents line features t l i endpoints; The LBD descriptors are constructed for the target image LSD line features and the corresponding candidate reference image LSD line features respectively, and the Euclidean distance between the target LSD line and the reference candidate LSD line descriptors is calculated. If the minimum Euclidean distance d between the two is less than the set threshold T d , then perform LSD line feature data association to obtain a pair of LSD line feature matching pairs.
4. The method for mobile robot pose estimation based on joint optimization of hybrid point and twin line feature reprojection according to claim 1, characterized in that: Step S3 is specifically as follows: Given a reference frame π r and a current frame π c , the set of matching point pairs obtained by point feature association is recorded as: in, Represents a pair of matching points; N P Indicates the number of matching point pairs; Given an RGB-D input frame F ID , the input frame includes the reference frame and the current frame, and contains the aligned depth image and RGB image; record I(p i ) and D(p i ) are the i-th pixel p i Grayscale value and depth value; Based on the current frame π c RGB image I t Observed matching points Define the set of 2D matching points without depth value as: The set of 3D matching points with depth values is: Among them, N PW +N PWO =N P .
5. The method for mobile robot pose estimation based on joint optimization of hybrid point and twin line feature reprojection according to claim 1, characterized in that: Step S4 is specifically as follows: Set the current frame π c As the left eye frame O L , based on π c The depth value in constructs the corresponding virtual right eye frame O R Specifically: Assume p i =(u i ,v i ) T is the left eye frame O L The i-th left eye pixel in the image is based on the x-direction focal length f x and baseline B, which corresponds to the virtual right eye frame O R The virtual right eye pixel on in The corresponding virtual right eye pixel is obtained by detecting each pixel on the line feature in the left eye frame, that is, the corresponding virtual right eye pixel is obtained. L The virtual right eye line of the detected line feature.
6. The method for mobile robot pose estimation based on joint optimization of hybrid point and twin line feature reprojection according to claim 4, characterized in that: In step S5, the specific process of constructing the mixed point feature reprojection error function is as follows: Define N PWO The reprojection error e1 of a 2D matching point without depth value is: Among them, s f represents the camera scale factor; ξ represents the camera motion, which is a six-dimensional vector of Lie algebra; K is the 3×3 camera internal matrix, which is obtained in advance through camera calibration; is the world coordinate system π w A three-dimensional observation point in ; exp(ξ^) is expressed as r to π c The pose transformation matrix is expressed as follows: Where t is a 3×1 translation vector; R is a 3×3 rotation matrix; a Lie algebra element ξ∈se(3) can be mapped to T by exponential mapping crw ∈SE(3); Define N with depth value PW The reprojection error e2 of the 3D matching points is: The calculation formula for the optimal camera pose based on the mixed point feature reprojection error is:
7. The method for mobile robot pose estimation based on joint optimization of hybrid point and twin line feature reprojection according to claim 5, characterized in that: In step S5, the specific process of constructing the twin line feature reprojection error function is as follows: definition w L b =( w G b,0 , w G b,1 ) is the world frame π w 3D line features in, where ( w G b,0 , w G b,1 )express w L b A pair of 3D endpoints of a line feature; given a reference frame π r and the current frame π c , the observed two-dimensional line feature matching pair set is expressed as: ML={[ r l b =( r g b,0 , r g b,1 ), c l b =( c g b,0 , c g b,1 )],b=1,…,N ML }; in,[ r l b , c l b ] represents a pair of matching line features; N ML Indicates the number of line feature matching pairs; ( r g b,0 , r g b,1 )and( c g b,0 , c g b,1 ) represent the reference frame π c Line features observed in r l b The current frame π c Line features observed in c l b 2D endpoints of; Reuse ( c g′ b,0 , c g′ b,1 ) represents the c Medium heavy projection line feature c l′ b 2D endpoints; based on the standard pinhole camera model, we get: set up corresponds to the two-dimensional endpoints The observation line features are aligned with the secondary coordinates of c l b Normalized line coefficient c η b The calculation formula is: Then use the sum of the distances from the point to the line to represent the endpoints of the reprojected line ( c g′ b,0 , c g′ b,1 ) and observation line characteristics c l b The line reprojection error between (e3+e4), that is: Among them, d b ( c g′ b,0 , c l b ),d b ( c g′ b,1 , c l b ) represent the endpoints of the reprojection line c g′ b,0 、 c g′ b,1 To observation line features c l b distance; For the current frame π c The observed N ML Line features, total line reprojection error E L It is obtained by accumulating the distance from the point to the line. The formula is expressed as: Line features c l b The endpoint ( c g b,0 , c g b,1 )and c l′ b The endpoint ( c g′ b,0 , c g′ b,1 ) as the starting position, and use the binary search algorithm to find the pixel p whose depth value is less than the set depth threshold Ω and greater than 0. i As a line c l b and c l′ b The new endpoint of ; if the depth value does not meet the conditions, no virtual right eye line is generated for the line feature; the final set of virtual right eye lines is: VRE={[ R l b =( R g b,0 , R g b,1 ), R l′ b =( R g′ b,0 , R g′ b,1 )]}; Where b = 1, 2, ..., N VRE , N VRE Indicates the number of virtual right eye lines; R l b and R l′ b Represent the virtual right eye frame O R The virtual right eye line observed in the left eye frame O and the reprojected virtual right eye line, which correspond to the left eye frame O L Line features in c l b and c l′ b ;( R g b,0 , R g b,1 )and( R g′ b,0 , R g′ b,1 ) represent lines R l b and R l′ b 2D endpoints of; Then calculate the virtual right eye line R l b The normalization coefficient of : Based on the depth measurement, the reprojection error E of the virtual right eye line is obtained R The calculation formula is: Combining the left eye line reprojection error and the virtual right eye line reprojection error, we get the twin line feature reprojection error function E l,cr (ξ), then the optimal camera pose based on the twin line feature reprojection error function is:
8. The method for mobile robot pose estimation based on joint optimization of hybrid point and twin line feature reprojection according to claim 1, characterized in that: Step S6 is specifically as follows: Perform joint optimization of point and line feature reprojection to obtain the optimal camera pose calculation formula: The solution is solved by iteratively following the Gauss-Newton optimization algorithm in the manifold tangent space se(3); first calculate the pose update: Δξ=-(J T WJ) -1 J T WE cr ; Among them, the error vector E cr Contains all reprojection errors E p,cr and E l,cr ; W and J are E cr The diagonal weight matrix and Jacobian matrix of ; Then the camera pose is iteratively updated, and the formula is expressed as: Iterate and update until convergence to output the optimal robot pose ξ.
Citation Information
Patent Citations
Robot scene self-adaptive pose estimation method based on RGB-D camera
CN110223348A
Forest complex environment digital twin system based on space-time multiple dimensions and construction method thereof
CN117671175A
Cited By
Layered fusion-based laser radar SLAM map construction method and system
CN121784768A
Multi-modal feature combined intraoperative medical instrument tracking method and system
CN122163321A