Visual inertial odometer method based on point-line feature optimization
By optimizing the point and line features of the visual inertial odometry system, including line segment selection, fusion, and optical flow tracing, and combining IMU information for system pose estimation, the problem of low accuracy and efficiency of line feature extraction in complex environments by visual inertial odometry is solved, and higher positioning accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510996677.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
Existing visual inertial odometry systems suffer from low accuracy in line feature extraction and line segment matching in environments with weak texture, rapid movement, or large changes in lighting, resulting in low positioning accuracy.
A visual inertial odometry method based on point and line feature optimization is adopted. By filtering, fusing and weighting the line segments according to length thresholds, combined with optical flow tracing, an optimized state variable is constructed and the system pose is estimated by combining IMU information.
It improves the quality and efficiency of line segment feature extraction, enhances the robustness and positioning accuracy of the system, adapts to a larger range of motion, and solves the problem of low positioning accuracy caused by low feature extraction accuracy in traditional methods.
Smart Images

Figure CN120876539A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and inertial navigation technology, specifically relating to a visual inertial odometry method based on point and line feature optimization. Background Technology
[0002] In recent years, with the rapid development of robotics and computer technology, visual inertial odometry (VIO), which combines visual information and inertial measurement unit (IMU), has received widespread attention and application in fields such as robot navigation and localization, autonomous driving, and augmented reality.
[0003] Traditional VIO systems primarily rely on point features for pose estimation and map construction. Accurate point features effectively ensure the accuracy and robustness of VIO systems, especially in textured and relatively stable environments. However, in environments with weak texture, motion blur, or significant lighting variations, the number of extracted point features decreases significantly, leading to unstable feature matching and consequently affecting the positioning accuracy and robustness of the VIO system. In these cases, line features, due to their inherent illumination and rotation invariance and rich structural information, become an effective supplementary method. Introducing line features can effectively improve positioning accuracy and robustness in weakly textured environments and construct more intuitive map scenes. However, VIO systems incorporating line features still have unresolved issues. Currently, line segment detection methods such as LSD and EDLines are commonly used for line feature extraction. LSD, an image gradient-based line segment detection algorithm, has high computational complexity, limiting its performance in VIO systems with high real-time requirements. EDLines, an edge-drawing-based line segment detection algorithm, significantly improves computational efficiency but is susceptible to weak edges and textured backgrounds, leading to unstable detection results. Compared to the methods mentioned above, ELSED can extract line segments more quickly and accurately in complex environments, thus meeting the requirements for higher precision and efficiency. However, the line segments extracted by the ELSED method still contain a large number of redundant line segments, which reduces the efficiency of subsequent line segment tracking and the accuracy of VIO system positioning. In addition, in traditional line feature tracking methods, LBD descriptors are widely used. They construct descriptors by analyzing the gradient information around the line segment to achieve robust line segment matching, but their computational complexity is high, limiting their efficiency in real-time applications.
[0004] In summary, existing line feature tracking methods suffer from low efficiency and are difficult to apply in real time due to the high computational complexity of line segment matching. Therefore, in order to further improve the performance of visual inertial odometry using point and line features, a new visual inertial odometry method is proposed for complex environments such as weak texture, rapid motion, and changing lighting. This is of great significance for expanding the application of visual inertial odometry in complex environments. Summary of the Invention
[0005] The purpose of this invention is to solve the problems of low line feature extraction accuracy, low line segment matching efficiency and low positioning accuracy in existing VIO systems, and to propose a visual inertial odometry method based on point and line feature optimization.
[0006] The technical solution adopted by this invention to solve the above-mentioned technical problems is: a visual inertial odometry method based on point and line feature optimization, the method specifically includes the following steps:
[0007] Step 1: Extract line segments from the input image and filter the extracted line segments according to the length threshold to obtain the filtered line segments;
[0008] Step 2: Merge the line segments selected in Step 1 to obtain the merged line segments;
[0009] Step 3: Select key points from each fused line segment, and use the selected key points to perform optical flow tracing of the line segments to obtain the position of each fused line segment in the next frame image;
[0010] Step 4: Use the tracked line segments as line features, and then perform pose estimation of the visual inertial odometry system based on the line features.
[0011] Furthermore, the extracted line segments are filtered based on a length threshold to obtain the filtered line segments; the specific process is as follows:
[0012] Step 11: Set the line segment length threshold L min :
[0013] L min =[k·min(W,H)] (1)
[0014] Where W and H represent the width and height of the input image, respectively, k is the scale factor, and [·] represents the rounding up function;
[0015] Steps 1 and 2: Compare the length of each extracted line segment with the threshold L. min Compare the results and retain those with a length greater than or equal to the threshold L. min Line segments with a length less than the threshold L are removed. min The line segment.
[0016] Furthermore, the specific process of step two is as follows:
[0017] Step 2: Randomly select any line segment L from the line segments selected in Step 1. old Initialize set B 0 Only line segment L is included. old Initialize set A 0 =C / L old C is the set of all line segments selected in step one, and A is the set of all line segments selected in step one. 0 This means removing L from set C. old The resulting set, and the set A 0 The line segments in the diagram are randomly numbered;
[0018] Step 22: Initialize the iteration count l = 0;
[0019] Steps two and three: From set A l Randomly select a line segment and denote the selected line segment as L. new and line segment L new From set A l Remove from the middle to get the updated set A. l+1 ;
[0020] A weighted scoring mechanism is used to assign this line segment to set B. l Select the optimal line segment L best ;
[0021] Step 24: Determine line segment L new With line segment L best Does the fusion condition meet?
[0022] If the fusion condition is met, then for line segment L... new With line segment L best After merging, use line segment L new With line segment L best The fusion result replaces set B l line segment L in best The updated set B is obtained. l+1 ;
[0023] If the fusion conditions are not met, then directly merge line segment L. new Add to set B l In the process, we obtain the updated set B. l+1 ;
[0024] Step 25: Determine set A l+1 Is it empty?
[0025] If set A l+1 If the set is empty, stop iterating and set B is empty.l+1 The line segments in the text are the merged line segments;
[0026] If set A l+1 If it is not empty, then let l = l + 1 and return to execute steps two and three.
[0027] Furthermore, the weighted scoring mechanism is used to assign a score to the line segment from set B. l Select the optimal line segment L best Specifically:
[0028] For line segment L new Define set B l Candidate line segment L i The scoring function S(L) new ,L i )for:
[0029]
[0030] Where Δθ is the difference in direction angle, θ th The set direction angle difference threshold, where α and β are coefficients, d min Represents line segment L new With line segment L best The minimum Euclidean distance between the endpoints, d th The set distance threshold;
[0031] Set B l The candidate line segment with the lowest score function value is taken as line segment L. new The optimal line segment L best .
[0032] Furthermore, the direction angle difference Δθ is:
[0033] Δθ=|θ new -θ i | (3)
[0034] Where |·| represents taking the absolute value, θ new For line segment L new Direction angle, θ i Candidate line segment L i The direction angle.
[0035] Furthermore, the specific process of step two and four is as follows:
[0036] For line segment L new and line segment L best Construct a structure containing line segment L new endpoints and line segment L best The set of endpoints
[0037] Calculate line segment L new and line segment L best The weighted principal direction vector v main :
[0038] v main =w1v new +w2v best (4)
[0039] Among them, v new Represents line segment L new The direction vector, v best Represents line segment L best The direction vectors, w1 and w2 are weights;
[0040] set Each endpoint in the vector is along the principal direction vector v. main projection:
[0041]
[0042] Where, p min p represents the minimum point obtained by projection. max This represents the point where the projection yields the maximum value.
[0043] Construct a new line segment L from the minimum point and the maximum point obtained by projection. fused =[p min ,p max ] Calculate the new line segment L fused With line segment L best Direction angle difference:
[0044] If the new line segment L fused With line segment L best The difference in direction angles exceeds the angle threshold θ th Then line segment L new and line segment L best The fusion conditions are not met;
[0045] If the new line segment L fused With line segment L best The difference in direction angles does not exceed the angle threshold θ th Then line segment L new and line segment L best The fusion conditions are met, and the new line segment L... fused That is, line segment L new and line segment L best The merging line segments.
[0046] Furthermore, the specific process of step three is as follows:
[0047] Step 3: For any line segment L after merging, denote the coordinates of the two endpoints of line segment L as p. s =(u s ,v s ) and p e =(u e ,v e If the direction vector of line segment L is d = p, then the direction vector of line segment L is d = p. e -p s ;
[0048] And set the number of sampling points N on line segment L;
[0049] Step 3.2, the coordinates q of the i-th sampling point on line segment L. i for:
[0050]
[0051] The gradient of the i-th sampling point is calculated using the Sobel operator.
[0052]
[0053] Where (x,y) represents the coordinates of the i-th sampling point, I y-1,x+1 I represents the grayscale value of a pixel (x+1, y-1) in the input image. y-1,x-1 I represents the grayscale value of pixel (x-1, y-1) in the input image. y,x+1 I represents the grayscale value of pixel (x+1, y) in the input image. y,x-1 I represents the gray value of pixel (x-1, y) in the input image. y+1,x+1 I represents the grayscale value of pixel (x+1, y+1) in the input image. y+1,x-1 I represents the grayscale value of pixel (x-1, y+1) in the input image. y+1,x I represents the grayscale value of pixel (x, y+1) in the input image. y-1,x This represents the gray value of a pixel (x, y-1) in the input image;
[0054] Step 3: Perform the first screening of the sampling points based on the gradient of each sampling point to obtain the remaining sampling points after the first screening;
[0055] The specific process of step 33 is as follows:
[0056] Calculate the cosine of the angle between the gradient at the i-th sampling point and the line segment direction vector d, cosθ:
[0057]
[0058] Where, d eLet d be the unit vector of the direction vector of line segment L. e =d / ||d||, where ||d|| represents the magnitude of the direction vector d;
[0059] If cosθ is less than the set threshold, the i-th sampling point is retained; otherwise, the i-th sampling point is discarded.
[0060] Steps 3 and 4: Based on the gradient, perform a second screening on the remaining sampling points after the first screening to obtain the key points of line segment L;
[0061] The specific process of steps three and four is as follows:
[0062] Divide line segment L evenly into K sub-intervals. The range of each interval is then:
[0063]
[0064] Among them, Ω k Let k represent the k-th subinterval, where k = 0, 1, ..., K-1;
[0065] Select the sampling point with the largest gradient magnitude in each interval and use the selected sampling point as the key point;
[0066] Step 35: Use the selected key points to perform line feature optical flow tracing to obtain the position of each fused line segment in the next frame image.
[0067] Furthermore, the number of sampling points N is:
[0068]
[0069] Where M is the length of line segment L.
[0070] Furthermore, the specific process of step three is as follows:
[0071] Step 3.51, based on the assumption of grayscale invariance, we get:
[0072] I(u+du, v+dv, t+dt)=I(u,v,t) (11)
[0073] Where I(u,v,t) represents the gray value of key point (u,v) in the input image at time t, I(u+du,v+dv,t+dt) represents the gray value of pixel (u+du,v+dv) in the input image at time t+dt, and (u+du,v+dv) is the pixel corresponding to key point (u,v) in the input image at time t+dt.
[0074] Step 352: Perform a Taylor expansion on the left side of equation (11) to obtain:
[0075]
[0076] in, Let represent the gradient of the input image in the u direction at time t. Let represent the gradient of the input image in the v direction at time t. This represents the change in the grayscale value of the input image at time t over time.
[0077] make Substituting equation (12) into equation (11) and transforming it, we get:
[0078]
[0079] For the set of key points on line segment L {(u1,v1),(u2,v2),…,(u…}, m ,v m The position of each key point in the next frame image (u) n ′,v′ n )for:
[0080]
[0081] Where g1 represents the translation of the line segment in the v direction, g2 represents the translation of the line segment in the u direction, g3 represents the rotation angle of the position of the line segment L in the next frame relative to the line segment L in the current frame, and n = 1, 2, ..., m;
[0082] Then, according to equation (14), we get:
[0083]
[0084] Let G1 = dg1 / dt, G2 = dg2 / dt, G3 = dg3 / dt, and substitute equation (15) and G1, G2 and G3 into equation (13) to obtain:
[0085]
[0086] Each key point (u) on line segment L n ,v n All satisfy equation (16). The parameters G1, G2 and G3 are solved by the Gauss-Newton method to obtain the motion estimate of the entire line segment L and complete the optical flow tracing of line segment L.
[0087] Furthermore, the specific process of step four is as follows:
[0088] Step 4.1 Define the state variable χ, which contains IMU information, point feature information, and line feature information:
[0089]
[0090] Where, n i n is the number of keyframes within the sliding window. p n is the total number of landmarks observed within the sliding window. l x is the total number of road signs observed within the sliding window. k It is the IMU information in the world coordinate system corresponding to the k-th keyframe within the sliding window. Indicates location, Indicates speed, Indicates posture, b a Indicates acceleration bias, b g Indicates the gyroscope bias, λ i′ This represents the inverse depth of the feature of the i′-th landmark, where i′ = 1, 2, ..., n p , The four-parameter orthogonal representation of the feature of the j′-th road sign, where j′=1,2,…,n l ;
[0091] Step 42: Establish the objective function for optimization:
[0092]
[0093] Among them, I prior e represents the prior residual of marginalization imu e represents the measurement residual portion of the IMU information. point e represents the reprojection error of point features. line This represents the reprojection error portion of the line feature;
[0094]
[0095] Where ρ is the kernel function for suppressing outliers, and r p It is the marginalized residual vector, H p It is a marginalized information matrix. The set of keyframes within the sliding window. This represents the pre-integrated measurement value between the k-th keyframe and the (k+1)-th keyframe. It is the IMU pre-integration residual. It is the covariance matrix of the IMU pre-integration error term from the k-th keyframe to the (k+1)-th keyframe. The set of visual point features observed in the sliding window. The set of visual line features observed within the sliding window. This represents the observation value of the i′-th point feature in the k′-th keyframe within the sliding window. This represents the observation value of the j′-th line feature in the k″-th keyframe. Represents the residual of the feature at point i′. Represents the residual of the j′-th linear feature. This represents the covariance matrix of the i′-th feature point at the k′-th keyframe within the sliding window. This represents the covariance matrix of the j′-th line feature in the k″-th keyframe;
[0096] Step 43: Solve the objective function to obtain the state variable χ; then obtain the pose estimation result of the visual inertial odometry system based on the state variable.
[0097] The beneficial effects of this invention are:
[0098] This invention employs the ELSED algorithm for initial line segment extraction, then fuses the extracted line segments by constructing geometric constraints, thereby improving the quality and efficiency of line segment feature extraction. By using optical flow to track line segment features, the efficiency of line segment feature matching is improved, avoiding the computational overhead of traditional feature matching, and enabling adaptation to a larger range of motion, further enhancing the robustness of visual inertial odometry. The tracked line features are combined with original point features and acceleration and angular velocity information provided by the IMU. A sliding window optimization strategy is used to construct optimized state variables including camera pose, velocity, bias, inverse depth of point features, and line feature parameters. An objective function is constructed, consisting of point feature reprojection error, line feature reprojection error, IMU measurement residual, and marginalized prior residual. A robust kernel function is used to adjust for the influence of outliers. By solving the objective function, the optimal system state estimate is obtained, achieving high-precision positioning and solving the problem of low positioning accuracy or even positioning failure caused by low feature extraction accuracy in traditional methods. Attached Figure Description
[0099] Figure 1 This is a comparison chart of the effects of different line segment extraction methods;
[0100] Figure 2 This is a diagram illustrating the line segment filtering and fusion effect of the present invention;
[0101] Figure 3 This is an experimental result diagram of the MH_05 sequence;
[0102] Figure 4 This is the trajectory diagram of the MH_05 sequence;
[0103] Figure 5 This is a comparison chart of trajectory error indices for the MH_05 sequence;
[0104] Figure 6 This is a flowchart of a visual inertial odometry method based on point and line feature optimization according to the present invention. Detailed Implementation
[0105] Specific implementation method one: Combining Figure 6 This embodiment describes a visual inertial odometry method based on point and line feature optimization. The method specifically includes the following steps:
[0106] Step 1: Extract line segments from the input image and filter the extracted line segments according to the length threshold to obtain the filtered line segments;
[0107] Step 2: Merge the line segments selected in Step 1 to obtain the merged line segments;
[0108] Step 3: Select key points from each fused line segment, and use the selected key points to perform optical flow tracing of the line segments to obtain the position of each fused line segment in the next frame image;
[0109] Step 4: Use the tracked line segments as line features, and then perform pose estimation of the visual inertial odometry system based on the line features.
[0110] Specific Implementation Method Two: This implementation method further defines Specific Implementation Method One. The extracted line segments are filtered based on a length threshold to obtain the filtered line segments. The specific process is as follows:
[0111] Step 11: Set the line segment length threshold L min :
[0112] L min =[k·min(W,H)] (1)
[0113] Where W and H represent the width and height of the input image, respectively, k is the scale factor (set to 0.083 based on experience), and [·] represents the rounding up function;
[0114] Steps 1 and 2: Compare the length of each extracted line segment with the threshold L. min Compare the results and retain those with a length greater than or equal to the threshold L. min Line segments with a length less than the threshold L are removed. min The line segment.
[0115] The other steps and parameters are the same as in Specific Implementation Method 1.
[0116] This invention uses the traditional ELSED line segment detection method to extract line segments from the input image. The main extraction steps are as follows:
[0117] First: Image preprocessing. The input image is Gaussian smoothed to reduce noise. The Sobel gradient of each pixel in the processed image is calculated to obtain edge information, and the gradient direction is quantized into horizontal and vertical to optimize subsequent calculations. Simultaneously, a set gradient threshold is used to remove low-contrast regions, ensuring that only significant edges are detected.
[0118] Secondly, anchor point extraction. Local maxima of gradient magnitude are selected in the image as anchor points. A step-size scanning strategy is employed to ensure that detection is performed only in necessary regions to improve efficiency.
[0119] Again: Line segment growth. Enhanced edge drawing is used to connect pixels, and the growth process terminates when the image boundary is reached, no usable pixels are available, or a significant gradient discontinuity is detected.
[0120] Finally: line segment verification. Calculate the gradient direction error to remove falsely detected line segments.
[0121] This invention employs the traditional ELSED line segment detection method to quickly and accurately extract line segments in complex environments. Then, it filters the line segments extracted based on the traditional ELSED line segment detection method according to a length threshold, thereby meeting the requirements of high precision and high efficiency. The line segment extraction and filtering method of this invention is uniformly named the improved ELSED algorithm. Figure 1 The image shows a comparison of line segment extraction using the traditional ELSED algorithm, LSD algorithm, and EDLines algorithm. It can be seen that the line segments extracted using the traditional ELSED algorithm contain fewer messy short lines, resulting in better extraction performance. Table 1 shows a comparison of the effects of different methods.
[0122] Table 1. Line segment extraction results
[0123]
[0124]
[0125] As shown in Table 1, the traditional ELSED algorithm extracts approximately 25% fewer line segments compared to other methods, increases the average length of line segments by approximately 22%, and also reduces the time required compared to the EDLines algorithm.
[0126] Specific Implementation Method Three: This implementation method is a further limitation of Specific Implementation Method Two. The specific process of step two is as follows:
[0127] Step 2: Randomly select any line segment L from the line segments selected in Step 1. old Initialize set B 0 Only line segment L is included. old Initialize set A 0 =C / L oldC is the set of all line segments selected in step one, and A is the set of all line segments selected in step one. 0 This means removing L from set C. old The resulting set, and the set A 0 The line segments in the diagram are randomly numbered;
[0128] Step 22: Initialize the iteration count l = 0;
[0129] Steps two and three: From set A l Randomly select a line segment and denote the selected line segment as L. new and line segment L new From set A l Remove from the middle to get the updated set A. l+1 ;
[0130] A weighted scoring mechanism is used to assign this line segment to set B. l Select the optimal line segment L best ;
[0131] Step 24: Determine line segment L new With line segment L best Does the fusion condition meet?
[0132] If the fusion condition is met, then for line segment L... new With line segment L best After merging, use line segment L new With line segment L best The fusion result replaces set B l line segment L in best The updated set B is obtained. l+1 ;
[0133] If the fusion conditions are not met, then directly merge line segment L. new Add to set B l In the process, we obtain the updated set B. l+1 ;
[0134] Step 25: Determine set A l+1 Is it empty?
[0135] If set A l+1 If the set is empty, stop iterating and set B is empty. l+1 The line segments in the text are the merged line segments;
[0136] If set A l+1 If it is not empty, then let l = l + 1 and return to execute steps two and three.
[0137] The other steps and parameters are the same as in Specific Implementation Method Two.
[0138] Specific Implementation Method Four: This implementation method is a further limitation of Specific Implementation Method Three. The weighted scoring mechanism is used to assign the line segment to set B. l Select the optimal line segment L best Specifically:
[0139] For line segment L new Define set B l Candidate line segment L i The scoring function S(L) new ,L i )for:
[0140]
[0141] Where Δθ is the difference in direction angle, θ th The set direction angle difference threshold, where α and β are coefficients, d min Represents line segment L new With line segment L best The minimum Euclidean distance between the endpoints (for line segment L) new Given one endpoint, calculate the distance between that endpoint and line segment L. best The distance between each endpoint for line segment L new The other endpoint, then calculate the distance between that endpoint and line segment L. best The distance between each endpoint (the minimum Euclidean distance is the minimum of the four calculated Euclidean distances), d th The set distance threshold;
[0142] Set B l The candidate line segment with the lowest score function value is taken as line segment L. new The optimal line segment L best .
[0143] The other steps and parameters are the same as in Specific Implementation Method 3.
[0144] α and β are used to balance angle and distance constraints; in this invention, the coefficients are α = 0.6 and β = 0.4. The threshold values are set to θ. th =2°, d th =20. Normalization is used to ensure the scoring function value is in the [0,1] interval. The optimal candidate line segment L is selected based on the scoring function value. best Perform line segment merging.
[0145] Specific Implementation Method Five: This implementation method further defines Specific Implementation Method Four, wherein the direction angle difference Δθ is:
[0146] Δθ=|θ new -θ i | (3)
[0147] Where |·| represents taking the absolute value, θ new For line segment L new Direction angle, θ i Candidate line segment L i The direction angle.
[0148] The other steps and parameters are the same as in Specific Implementation Method Four.
[0149] Specific Implementation Method Six: This implementation method is a further limitation of Specific Implementation Method Five. The specific process of step two and four is as follows:
[0150] For line segment L new and line segment L best Construct a structure containing line segment L new endpoints and line segment L best The set of endpoints
[0151] Calculate line segment L new and line segment L best The weighted principal direction vector v main :
[0152] v main =w1v new +w2v best (4)
[0153] Among them, v new Represents line segment L new The direction vector, v best Represents line segment L best The direction vector, w1 and w2 are weights; the weights w1 and w2 are related to the line segment L. new and line segment L best The length is proportional to the length of the line segment, and w1+w2=1, ensuring the estimation of the dominant direction of the long line segment.
[0154] set Each endpoint in the vector is along the principal direction vector v. main projection:
[0155]
[0156] Where, p min p represents the minimum point obtained by projection. max This represents the point where the projection yields the maximum value.
[0157] Construct a new line segment L from the minimum point and the maximum point obtained by projection. fused =[p min ,p max ] Calculate the new line segment L fused With line segment Lbest Direction angle difference:
[0158] If the new line segment L fused With line segment L best The difference in direction angles exceeds the angle threshold θ th Then line segment L new and line segment L best The fusion conditions are not met;
[0159] If the new line segment L fused With line segment L best The difference in direction angles does not exceed the angle threshold θ th Then line segment L new and line segment L best The fusion conditions are met, and the new line segment L... fused That is, line segment L new and line segment L best The merging line segments.
[0160] The other steps and parameters are the same as in Specific Implementation Method 5.
[0161] This implementation verifies the new line segment L through a direction angle consistency check. fused The rationality is that if the difference in direction angle between the new line segment and the original line segment exceeds the angle threshold θ. th If the result is not found, the merging process is rejected to avoid generating erroneous line segments and ensure the accuracy of the merged line segments. Figure 2 As shown, the fused line features are of moderate quantity, evenly distributed, and relatively long. Specifically, Figure 2 The number of line segments after fusion is reduced to 50, and the average length of line segments is increased to 106.2, effectively reducing redundant line segments while retaining the main features.
[0162] Specific Implementation Method Seven: This implementation method is a further limitation of Specific Implementation Method Six. The specific process of step three is as follows:
[0163] Step 3: For any line segment L after merging, denote the coordinates of the two endpoints of line segment L as p. s =(u s ,v s ) and p e =(u e ,v e If the direction vector of line segment L is d = p, then the direction vector of line segment L is d = p. e -p s ;
[0164] And set the number of sampling points N on line segment L;
[0165] Step 3.2, the coordinates q of the i-th sampling point on line segment L. i for:
[0166]
[0167] The gradient of the i-th sampling point is calculated using the Sobel operator.
[0168]
[0169] Where (x,y) represents the coordinates of the i-th sampling point, I y-1,x+1 I represents the grayscale value of a pixel (x+1, y-1) in the input image. y-1,x-1 I represents the grayscale value of pixel (x-1, y-1) in the input image. y,x+1 I represents the grayscale value of pixel (x+1, y) in the input image. y,x-1 I represents the gray value of pixel (x-1, y) in the input image. y+1,x+1 I represents the grayscale value of pixel (x+1, y+1) in the input image. y+1,x-1 I represents the grayscale value of pixel (x-1, y+1) in the input image. y+1,x I represents the grayscale value of pixel (x, y+1) in the input image. y-1,x This represents the gray value of a pixel (x, y-1) in the input image;
[0170] Step 3: Perform the first screening of the sampling points based on the gradient of each sampling point to obtain the remaining sampling points after the first screening;
[0171] The specific process of step 33 is as follows:
[0172] Calculate the cosine of the angle between the gradient at the i-th sampling point and the line segment direction vector d, cosθ:
[0173]
[0174] Where, d e Let d be the unit vector of the direction vector of line segment L. e =d / ||d||, where ||d|| represents the magnitude of the direction vector d;
[0175] If cosθ is less than the set threshold, the i-th sampling point is retained; otherwise, the i-th sampling point is discarded.
[0176] In the first screening, this invention retains only sampling points with smaller cosθ to ensure a high degree of orthogonality between the key points and the target line direction, thereby enhancing stability.
[0177] Steps 3 and 4: Based on the gradient, perform a second screening on the remaining sampling points after the first screening to obtain the key points of line segment L;
[0178] To filter out weak gradient points and retain only the gradient magnitude ||g i For larger candidate points, the specific processes of steps three and four are as follows:
[0179] To ensure the uniform distribution of key points on the target line, the line segment L is uniformly divided into K sub-intervals. In this invention, K = 5, so the range of each interval is:
[0180]
[0181] Among them, Ω k Let k represent the k-th subinterval, where k = 0, 1, ..., K-1;
[0182] Select the sampling point with the largest gradient magnitude in each interval and use the selected sampling point as the key point;
[0183] Step 35: Use the selected key points to perform line feature optical flow tracing (based on the assumption of grayscale invariance, that is, the grayscale value of the same pixel remains unchanged between consecutive time series image frames) to obtain the position of each fused line segment in the next frame image.
[0184] The other steps and parameters are the same as in Specific Implementation Method Six.
[0185] This implementation extracts key points for optical flow tracing by fusing gradient response and piecewise optimization. Using the extracted key points for optical flow tracing can improve the accuracy of optical flow tracing.
[0186] Specific Implementation Method Eight: This implementation method is a further limitation of Specific Implementation Method Seven, wherein the number of sampling points N is:
[0187]
[0188] Where M is the length of line segment L.
[0189] The other steps and parameters are the same as in Specific Implementation Method Seven.
[0190] The sampling point setting method in this embodiment can ensure that the sampling density adapts to segments of different lengths.
[0191] Specific Implementation Method Nine: This implementation method is a further limitation of Specific Implementation Method Eight. The specific process of step three and five is as follows:
[0192] Step 3.51, based on the assumption of grayscale invariance, we get:
[0193] I(u+du, v+dv, t+dt)=I(u,v,t) (11)
[0194] Where I(u,v,t) represents the gray value of key point (u,v) in the input image at time t, I(u+du,v+dv,t+dt) represents the gray value of pixel (u+du,v+dv) in the input image at time t+dt, and (u+du,v+dv) is the pixel corresponding to key point (u,v) in the input image at time t+dt.
[0195] Step 352: Perform a Taylor expansion on the left-hand side of equation (11) and approximate it by ignoring higher-order terms to obtain:
[0196]
[0197] in, Let represent the gradient of the input image in the u direction at time t. Let represent the gradient of the input image in the v direction at time t. This represents the change in the grayscale value of the input image at time t over time.
[0198] make Substituting equation (12) into equation (11) and transforming it, we get:
[0199]
[0200] The above formula applies to every pixel in the image, so it can be used at multiple points on a line segment, while these points also satisfy the collinearity constraint.
[0201] For the set of key points on line segment L {(u1,v1),(u2,v2),...,(u...}, ...,(u...} m ,v m The position of each key point in the next frame image (u) n ′,v′ n )for:
[0202]
[0203] Where g1 represents the translation of the line segment in the v direction, g2 represents the translation of the line segment in the u direction, g3 represents the rotation angle of the position of the line segment L in the next frame relative to the line segment L in the current frame, and n = 1, 2, ..., m;
[0204] After rearranging and differentiating equation (14), we get the following from equation (14):
[0205]
[0206] Let G1 = dg1 / dt, G2 = dg2 / dt, G3 = dg3 / dt, and substitute equation (15) and G1, G2 and G3 into equation (13) to obtain:
[0207]
[0208] Each key point (u) on line segment L n ,v n All satisfy equation (16). The parameters G1, G2 and G3 are solved by the Gauss-Newton method to obtain the motion estimate of the entire line segment L and complete the optical flow tracing of line segment L.
[0209] The other steps and parameters are the same as in Specific Implementation Method 8.
[0210] Based on the line segments in the current frame image and formula (16), parameters G1, G2, and G3 can be solved. Then, by combining parameters G1, G2, G3, and formula (14), the coordinates of the key points on the line segments in the current frame image corresponding to the next frame image can be obtained, and the tracking results of the line segments in the next frame image can be obtained. By analogy, the line features in the next frame image can be tracked to update the key frame image in the sliding window and obtain the real-time positioning results.
[0211] Specific Implementation Method Ten: This implementation method further defines Specific Implementation Method Nine. It uses the matched line segments obtained through tracking as line features and obtains point features based on traditional methods. Simultaneously, it combines visual information and IMU information to achieve system pose estimation. The specific process of step four is as follows:
[0212] Step 4.1 Define the state variable χ, which contains IMU information, point feature information, and line feature information:
[0213]
[0214] Where, n i n is the number of keyframes within the sliding window. p n is the total number of landmarks observed within the sliding window. l It is the total number of road signs observed within the sliding window (i.e., the total number of line features), x k It is the IMU information in the world coordinate system corresponding to the k-th keyframe within the sliding window. Indicates location, Indicates speed, Indicates posture, b a Indicates acceleration bias, b g Indicates the gyroscope bias, λ i′ This represents the inverse depth of the feature of the i′-th landmark, where i′ = 1, 2, ..., n p , The four-parameter orthogonal representation of the feature of the j′-th road sign, where j′=1,2,…,n l ;
[0215] Step 42: Establish the objective function for optimization:
[0216]
[0217] Among them, I prior e represents the prior residual of marginalization imu e represents the measurement residual portion of the IMU information. point e represents the reprojection error of point features. line This represents the reprojection error portion of the line feature;
[0218]
[0219] Where ρ is the robust kernel function for suppressing outliers, and r p It is the marginalized residual vector, H p It is a marginalized information matrix. The set of keyframes within the sliding window. This represents the pre-integrated measurement value between the k-th keyframe and the (k+1)-th keyframe. It is the IMU pre-integration residual. It is the covariance matrix of the IMU pre-integration error term from the k-th keyframe to the (k+1)-th keyframe. The set of visual point features observed in the sliding window. The set of visual line features observed within the sliding window. This represents the observation value of the i′-th feature in the k′-th keyframe within the sliding window (k′ represents the keyframe where the i′-th feature can be observed). This represents the observation value of the j′-th line feature in the k″-th keyframe (k″ represents the keyframe where the j′-th line feature can be observed). Represents the residual of the feature at point i′. Represents the residual of the j′-th linear feature. This represents the covariance matrix of the i′-th feature point at the k′-th keyframe within the sliding window. This represents the covariance matrix of the j′-th line feature in the k″-th keyframe;
[0220] Step 43: Solve the objective function to obtain the state variable χ; then obtain the pose estimation result of the visual inertial odometry system based on the state variable.
[0221] The other steps and parameters are the same as in Specific Implementation Method Nine.
[0222] The kernel functions that can be used in this implementation include, but are not limited to, the Huber kernel function and the Cauchy kernel function.
[0223] Experimental Section
[0224] Experiments were conducted using the representative open-source EuRoC dataset, collected from a hexacopter UAV platform. This dataset includes visual images, synchronized IMU measurement data, and high-precision motion and structure ground truth. The dataset contains factory environment sequences and indoor room environments, with sequence difficulty increasing from "easy" to "medium" to "difficult." The environment gradually changes from stable motion and ample lighting to complex environments with intense motion and fluctuating lighting. To better demonstrate the advantages of the proposed method under complex conditions, two complex environment sequences of "medium" and "difficult" difficulty were selected for testing, and the results were compared with VINS-Mono (using only point features) and PL-VINS (using traditional line feature matching methods).
[0225] To ensure fairness in the comparison, all methods were tested on a unified platform, and the experimental environment used is shown in Table 2.
[0226] Table 2 Experimental Environment
[0227] system CPU RAM Ubuntu 20.04 i5-12500H 16 GB
[0228] Different methods were run to obtain estimated trajectories for each method. The EVO tool was used for result analysis, and the positioning accuracy was evaluated based on the root mean square error (RMSE) of the absolute trajectory error (ATE). ATE measures trajectory accuracy by calculating the Euclidean distance between the estimated trajectory and the true trajectory at each timestamp, while RMSE quantifies the overall error by calculating the average of the squared differences between the estimated and true positions. Figure 3 The image shows the experimental results of using the method of this invention on the MH_05 sequence. Figure 4 This is a comparison chart of the experimental results of the MH_05 sequence with the reference trajectory. Figure 5 This chart compares the trajectory error indices of different methods for the MH_05 sequence. Figure 4 As shown, the motion trajectory of the method of the present invention basically coincides with the true trajectory value, and the trajectory error is very small. Figure 5 As shown, the method of this invention achieves lower values for the maximum, average, median, and standard deviation of ATE, outperforming existing methods. Table 3 presents a comparison of RSME values for each sequence:
[0229] Table 3 Comparison of RSME values for each sequence
[0230] sequence VINS-mono PL-VINS This invention MH_03_medium 0.063 0.060 0.057 MH_04_difficult 0.124 0.111 0.102 MH_05_difficult 0.151 0.138 0.073 V1_02_medium 0.057 0.038 0.039 V1_03_difficult 0.194 0.142 0.099 V2_02_medium 0.093 0.083 0.069 V2_03_difficult 0.293 0.145 0.110 average value 0.139 0.102 0.078
[0231] As shown in Table 3, the method of this invention maintained the lowest RMSE value in almost all "medium" and "difficult" sequences, with an average reduction of 23.5%, which is better than VINS-mono and PL-VINS. In particular, in the difficult sequence, the RMSE reduction averaged 33.8% in the MH_05_dif sequence, demonstrating the robustness of the method of this invention in complex environments and its effective improvement in positioning accuracy.
[0232] In summary, to address the issue of insufficient positioning accuracy in traditional VIO systems under complex environments, this invention proposes a visual inertial odometry method that combines improved ELSED line feature extraction (traditional ELSED line feature extraction with the line segment selection and fusion method of this invention) and optical flow. This method can improve the quality and efficiency of line feature extraction; and by tracking line features using optical flow, it can adapt to a wider range of motion. Results show that the positioning accuracy of this method is superior to existing VIO systems in most complex environments, with a significantly reduced average positioning error. This is of great significance for expanding the application of visual inertial odometry in complex environments.
[0233] The above examples of the present invention are merely illustrative of the computational model and process of the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is impossible to exhaustively list all possible implementations here. Any obvious variations or modifications derived from the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A visual inertial odometry method based on point and line feature optimization, characterized in that, The method specifically includes the following steps: Step 1: Extract line segments from the input image and filter the extracted line segments according to the length threshold to obtain the filtered line segments; Step 2: Merge the line segments selected in Step 1 to obtain the merged line segments; Step 3: Select key points from each fused line segment, and use the selected key points to perform optical flow tracing of the line segments to obtain the position of each fused line segment in the next frame image; Step 4: Use the tracked line segments as line features, and then perform pose estimation of the visual inertial odometry system based on the line features.
2. The visual inertial odometry method based on point and line feature optimization according to claim 1, characterized in that, The extracted line segments are filtered based on a length threshold to obtain the filtered line segments; the specific process is as follows: Step 11: Set the line segment length threshold L min : L min =[k·min(W,H)] (1) Where W and H represent the width and height of the input image, respectively, k is the scale factor, and [·] represents the rounding up function; Steps 1 and 2: Compare the length of each extracted line segment with the threshold L. min Compare the results and retain those with a length greater than or equal to the threshold L. min Line segments with a length less than the threshold L are removed. min The line segment.
3. The visual inertial odometry method based on point and line feature optimization according to claim 2, characterized in that, The specific process of step two is as follows: Step 2: Randomly select any line segment L from the line segments selected in Step 1. old Initialize set B 0 Only line segment L is included. old Initialize set A 0 =C / L old C is the set of all line segments selected in step one, and A is the set of all line segments selected in step one. 0 This means removing L from set C. old The resulting set, and the set A 0 The line segments in the diagram are randomly numbered; Step 22: Initialize the iteration count l = 0; Steps two and three: From set A l Randomly select a line segment and denote the selected line segment as L. new and line segment L new From set A l Remove from the middle to get the updated set A. l+1 ; A weighted scoring mechanism is used to assign this line segment to set B. l Select the optimal line segment L best ; Step 24: Determine line segment L new With line segment L best Does the fusion condition meet? If the fusion condition is met, then for line segment L... new With line segment L best After merging, use line segment L new With line segment L best The fusion result replaces set B l line segment L in best The updated set B is obtained. l+1 ; If the fusion conditions are not met, then directly merge line segment L. new Add to set B l In the process, we obtain the updated set B. l+1 ; Step 25: Determine set A l+1 Is it empty? If set A l+1 If the set is empty, stop iterating and set B is empty. l+1 The line segments in the text are the merged line segments; If set A l+1 If it is not empty, then let l = l + 1 and return to execute steps two and three.
4. The visual inertial odometry method based on point and line feature optimization according to claim 3, characterized in that, The weighted scoring mechanism is used to assign this line segment to set B. l Select the optimal line segment L best Specifically: For line segment L new Define set B l Candidate line segment L i The scoring function S(L) new ,L i )for: Where Δθ is the difference in direction angle, θ th The set direction angle difference threshold, where α and β are coefficients, d min Represents line segment L new With line segment L best The minimum Euclidean distance between the endpoints, d th The set distance threshold; Set B l The candidate line segment with the lowest score function value is taken as line segment L. new The optimal line segment L best .
5. The visual inertial odometry method based on point and line feature optimization according to claim 4, characterized in that, The direction angle difference Δθ is: Δθ=|θ new -θ i | (3) Where |·| represents taking the absolute value, θ new For line segment L new Direction angle, θ i Candidate line segment L i The direction angle.
6. The visual inertial odometry method based on point and line feature optimization according to claim 5, characterized in that, The specific process of step two or four is as follows: For line segment L new and line segment L best Construct a structure containing line segment L new endpoints and line segment L best The set of endpoints Calculate line segment L new and line segment L best The weighted principal direction vector v main : in main =w1v new +w2v best (4) Among them, v new Represents line segment L new The direction vector, v best Represents line segment L best The direction vectors, w1 and w2 are weights; set Each endpoint in the vector is along the principal direction vector v. main projection: Where, p min p represents the minimum point obtained by projection. max This represents the point where the projection yields the maximum value. Construct a new line segment L from the minimum point and the maximum point obtained by projection. fused =[p min ,p max ] Calculate the new line segment L fused With line segment L best Direction angle difference: If the new line segment L fused With line segment L best The difference in direction angles exceeds the angle threshold θ th Then line segment L new and line segment L best The fusion conditions are not met; If the new line segment L fused With line segment L best The difference in direction angles does not exceed the angle threshold θ th Then line segment L new and line segment L best The fusion conditions are met, and the new line segment L... fused That is, line segment L new and line segment L best The merging line segments.
7. The visual inertial odometry method based on point and line feature optimization according to claim 6, characterized in that, The specific process of step three is as follows: Step 3: For any line segment L after merging, denote the coordinates of the two endpoints of line segment L as p. s =(u s ,v s ) and p e =(u e ,v e If the direction vector of line segment L is d = p, then the direction vector of line segment L is d = p. e -p s ; And set the number of sampling points N on line segment L; Step 3.2, the coordinates q of the i-th sampling point on line segment L. i for: The gradient g at the i-th sampling point is calculated using the Sobel operator. i =(g x ,g y ) T : Where (x,y) represents the coordinates of the i-th sampling point, I y-1,x+1 I represents the grayscale value of a pixel (x+1, y-1) in the input image. y-1,x-1 I represents the grayscale value of pixel (x-1, y-1) in the input image. y,x+1 I represents the grayscale value of pixel (x+1, y) in the input image. y,x-1 I represents the gray value of pixel (x-1, y) in the input image. y+1,x+1 I represents the grayscale value of pixel (x+1, y+1) in the input image. y+1,x-1 I represents the grayscale value of pixel (x-1, y+1) in the input image. y+1,x I represents the grayscale value of pixel (x, y+1) in the input image. y-1,x This represents the gray value of a pixel (x, y-1) in the input image; Step 3: Perform the first screening of the sampling points based on the gradient of each sampling point to obtain the remaining sampling points after the first screening; The specific process of step 33 is as follows: Calculate the cosine of the angle between the gradient at the i-th sampling point and the line segment direction vector d, cosθ: Where, d e Let d be the unit vector of the direction vector of line segment L. e =d / ||d||, where ||d|| represents the magnitude of the direction vector d; If cosθ is less than the set threshold, the i-th sampling point is retained; otherwise, the i-th sampling point is discarded. Steps 3 and 4: Based on the gradient, perform a second screening on the remaining sampling points after the first screening to obtain the key points of line segment L; The specific process of steps three and four is as follows: Divide line segment L evenly into K sub-intervals. The range of each interval is then: Among them, Ω k Let k represent the k-th subinterval, where k = 0, 1, ..., K-1; Select the sampling point with the largest gradient magnitude in each interval and use the selected sampling point as the key point; Step 35: Use the selected key points to perform line feature optical flow tracing to obtain the position of each fused line segment in the next frame image.
8. The visual inertial odometry method based on point and line feature optimization according to claim 7, characterized in that, The number of sampling points N is: Where M is the length of line segment L.
9. The visual inertial odometry method based on point and line feature optimization according to claim 8, characterized in that, The specific process of step three is as follows: Step 3.51, based on the assumption of grayscale invariance, we get: I(u+du,v+dv,t+dt)=I(u,v,t) (11) Where I(u,v,t) represents the gray value of key point (u,v) in the input image at time t, I(u+du,v+dv,t+dt) represents the gray value of pixel (u+du,v+dv) in the input image at time t+dt, and (u+du,v+dv) is the pixel corresponding to key point (u,v) in the input image at time t+dt. Step 352: Perform a Taylor expansion on the left side of equation (11) to obtain: in, Let represent the gradient of the input image in the u direction at time t. Let represent the gradient of the input image in the v direction at time t. This represents the change in the grayscale value of the input image at time t over time. make Substituting equation (12) into equation (11) and transforming it, we get: For the set of key points on line segment L {(u1,v1),(u2,v2),…,(u…}, m ,v m The position of each key point in the next frame image (u) n ′,v′ n )for: Where g1 represents the translation of the line segment in the v direction, g2 represents the translation of the line segment in the u direction, g3 represents the rotation angle of the position of the line segment L in the next frame relative to the line segment L in the current frame, and n = 1, 2, ..., m; Then, according to equation (14), we get: Let G1 = dg1 / dt, G2 = dg2 / dt, G3 = dg3 / dt, and substitute equation (15) and G1, G2 and G3 into equation (13) to obtain: Each key point (u) on line segment L n ,v n All satisfy equation (16). The parameters G1, G2 and G3 are solved by the Gauss-Newton method to obtain the motion estimate of the entire line segment L and complete the optical flow tracing of line segment L.
10. A visual inertial odometry method based on point and line feature optimization according to claim 9, characterized in that, The specific process of step four is as follows: Step 4.1 Define the state variable χ, which contains IMU information, point feature information, and line feature information: Where, n i n is the number of keyframes within the sliding window. p n is the total number of landmarks observed within the sliding window. l x is the total number of road signs observed within the sliding window. k It is the IMU information in the world coordinate system corresponding to the k-th keyframe within the sliding window. Indicates location, Indicates speed, Indicates posture, b a Indicates acceleration bias, b g Indicates the gyroscope bias, λ i′ This represents the inverse depth of the feature of the i′-th landmark, where i′ = 1, 2, ..., n p , The four-parameter orthogonal representation of the feature of the j′-th road sign, where j′=1,2,…,n l ; Step 42: Establish the objective function for optimization: Among them, I prior e represents the prior residual of marginalization imu e represents the measurement residual portion of the IMU information. point e represents the reprojection error of point features. line This represents the reprojection error portion of the line feature; Where ρ is the kernel function for suppressing outliers, and r p It is the marginalized residual vector, H p It is a marginalized information matrix. The set of keyframes within the sliding window. This represents the pre-integrated measurement value between the k-th keyframe and the (k+1)-th keyframe. It is the IMU pre-integration residual. It is the covariance matrix of the IMU pre-integration error term from the k-th keyframe to the (k+1)-th keyframe. The set of visual point features observed in the sliding window. The set of visual line features observed within the sliding window. This represents the observation value of the i′-th point feature in the k′-th keyframe within the sliding window. This represents the observation value of the j′-th line feature in the k″-th keyframe. Represents the residual of the feature at point i′. Represents the residual of the j′-th linear feature. This represents the covariance matrix of the i′-th feature point at the k′-th keyframe within the sliding window. This represents the covariance matrix of the j′-th line feature in the k″-th keyframe; Step 43: Solve the objective function to obtain the state variable χ; then obtain the pose estimation result of the visual inertial odometry system based on the state variable.