Infrared inertial odometer method for unmanned ground vehicle

By extracting and tracking line features in infrared images, combining the camera-ground geometric model and sliding window optimization, the problems of reduced accuracy of infrared SLAM in harsh environments and inaccurate scale of ground unmanned vehicles are solved, and an efficient and accurate infrared inertial odometry is achieved.

CN120628136APending Publication Date: 2025-09-12Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510758142.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The performance of existing infrared SLAM systems degrades in harsh environments such as large changes in lighting conditions, insufficient lighting, or smoke. In addition, the VINS model introduces unobservable directions when ground unmanned vehicles move at constant acceleration or speed, affecting the accuracy of pose estimation. In particular, the accuracy is low over large scales. Existing methods also require the integration of additional sensors, which increases expenses and computing costs.

Method used

The ELSED algorithm is used to extract line features from infrared images. The three-degree-of-freedom line feature optical flow tracking algorithm is combined to track line features between consecutive frames. A camera-ground geometric model is established, and BEV images are generated through IPM transformation. Planar constrained ground line features are added, and IMU data is optimized. The metric scale is restored using the camera-ground geometric relationship, and multiple factors are optimized within the sliding window.

Benefits of technology

It improves the efficiency and accuracy of infrared image calculations, overcomes the scale degradation problem of ground unmanned vehicles in large-scale environments, achieves all-weather operation, and improves accuracy by 67.88%-96.73%, which is a significant improvement compared to traditional algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120628136A_ABST
    Figure CN120628136A_ABST
Patent Text Reader

Abstract

The invention provides an infrared inertial odometer method for a ground unmanned vehicle. The infrared inertial odometer method comprises the following steps: extracting point features and line features in an original infrared image and tracking the point features and the line features; the method comprises the following steps: establishing a camera-ground geometric model, performing IPM transformation on an original infrared image to generate a BEV image, predicting and tracking the features of ground points in the BEV image, and remapping the tracked ground points to the original infrared image through the inverse process of the IPM; based on a camera-ground geometric model, judging whether the line features are ground line features or not, and if yes, adding plane constraints to the ground line features; iMU data between two frames of an original infrared image is pre-integrated, and an IMU pre-integration factor, a point feature re-projection factor, a line feature re-projection factor, a ground point feature constraint factor and a ground line feature constraint factor are optimized in real time based on a sliding window. The method improves the calculation efficiency and precision, and achieves the accurate monocular TIO under the condition that the excitation of the ground unmanned vehicle is insufficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an infrared inertial odometer method for ground unmanned vehicles. Background Art

[0002] Unmanned systems have been widely used in recent years, and accurate positioning is a prerequisite for their operation. Simultaneous localization and mapping (SLAM) technology provides an effective solution for positioning in unknown environments without relying on prior information. Visual SLAM has been widely used in small robots due to the advantages of cameras, such as low cost, small size, and low power consumption. Furthermore, cameras and IMUs have distinct complementary characteristics and are typically fused to form a visual-inertial navigation system. The IMU's short-term state prediction can assist in visual feature tracking and, to a certain extent, restore scale information that monocular cameras lack, while visual tracking can correct IMU deviations.

[0003] However, the Visual-Inertial Navigation System (VINS) has poor anti-interference capabilities. Under harsh conditions such as large variations in lighting conditions, insufficient lighting, or smoke, VINS performance can severely degrade or even fail to operate properly. Infrared cameras are independent of scene lighting conditions and offer advantages such as night vision, smoke, and fog resistance, enabling all-weather operation. Long-wave infrared (LWIR) cameras offer advantages over medium- and short-wave infrared cameras. Uncooled infrared cameras also offer advantages such as small size and low energy consumption. Their potential for application in challenging environments has garnered attention, demonstrating excellent performance in scenarios such as underground exploration, fire rescue, and post-disaster reconstruction. However, infrared images suffer from low contrast, high noise, significant grayscale variations between consecutive frames, and the increased challenges of feature extraction and matching. Directly utilizing traditional visible light SLAM algorithms significantly reduces performance, limiting their application.

[0004] Most existing infrared SLAM systems utilize only traditional or trained point features. However, infrared cameras have poor texture quality, especially in structured scenes. The number of point features extracted from weakly textured infrared images is severely insufficient, and tracking is often challenging. Recent studies have shown that infrared images, based on the thermal radiation of objects, are suitable for extracting features such as edges and lines. Furthermore, line features exhibit high grayscale invariance and are relatively stable across wide viewing angles, compensating for the significant grayscale variations between consecutive frames experienced by infrared cameras. However, the extraction and tracking of line features is currently time-consuming and unsuitable for scenarios with limited computing power. Furthermore, line features have four degrees of freedom, making them more susceptible to degradation than point features. However, lines on the ground only have two degrees of freedom. Treating ground lines as general lines increases redundant parameters, hindering convergence in the optimization process.

[0005] On the other hand, the application of monocular VIO (Visual-Inertial Odometry) in ground-based mobile vehicles also faces challenges. When the vehicle moves in a straight line or in a circle with constant acceleration or velocity, it introduces additional unobservable directions to the VINS model. Ground-based vehicles are generally unable to fully excite the IMU during motion, which significantly affects the pose estimation accuracy of the VINS, especially at scale, resulting in lower accuracy when operating over a large range. This often requires the integration of additional sensors such as lidar, GNSS, and wheel speedometers. However, this approach introduces additional costs and computational overhead. Summary of the Invention

[0006] In view of this, an embodiment of the present application provides an infrared inertial odometry method for ground unmanned vehicles, which improves computing efficiency and accuracy to achieve accurate monocular TIO (Thermal-Inertial Odometry) when the ground unmanned vehicle is insufficiently excited.

[0007] The present application provides the following technical solution: an infrared inertial odometer method for a ground unmanned vehicle, comprising:

[0008] Extracting point features and line features from an original infrared image of the target vehicle's external environment, and tracking the point features and line features between consecutive frames; wherein the line features in the original infrared image are extracted using an ELSED algorithm, and the line features between consecutive frames are tracked using a three-degree-of-freedom line feature optical flow tracking algorithm;

[0009] Establishing a camera-ground geometric model, based on the camera-ground geometric model, transforming the original infrared image through IPM to generate a BEV image, predicting and tracking the features of the ground points in the BEV image according to the point features, and remapping the tracked ground points to the original infrared image through the inverse process of IPM;

[0010] Based on the camera-ground geometric model, determining whether the extracted line feature is a ground line feature, and if so, adding a plane constraint to the ground line feature to constrain the ground line feature to a set solution space;

[0011] The IMU data between the two frames of the original infrared image are pre-integrated, and the IMU pre-integration factor, point feature reprojection factor, line feature reprojection factor, ground point feature constraint factor and ground line feature constraint factor are optimized in real time based on a sliding window.

[0012] According to one embodiment of the present application, a three-degree-of-freedom line feature optical flow tracking algorithm is used to track line features between consecutive frames, including:

[0013] The grayscale of the original infrared image between consecutive frames is set to be unchanged, and the pixel motion in the rectangular window around the line feature is set to be consistent, and the line feature optical flow tracking is performed using an image pyramid.

[0014] According to an embodiment of the present application, determining whether the extracted line feature is a ground line feature includes:

[0015] Set a matching line angle difference threshold and a midpoint pixel distance threshold, and compare the angle difference and midpoint pixel distance of the matching line extracted after projecting the homography matrix between the ground line features of two adjacent frames with the matching line angle difference threshold and the midpoint pixel distance threshold respectively to determine whether the line feature is a ground line feature.

[0016] According to an embodiment of the present application, the matching line angle difference threshold is 0.1 radians, and the midpoint pixel distance threshold is 4 pixels.

[0017] According to one embodiment of the present application, the method further includes:

[0018] The features of ground points are predicted and tracked in the BEV image, and outliers of ground points are removed after tracking using bidirectional iterative optical flow and a fundamental matrix-based RANSAC method. Finally, the tracked ground points are remapped to the original infrared image through the inverse process of IPM.

[0019] According to an embodiment of the present application, the method further includes: initializing the camera-ground parameters of the camera-ground geometric model, and after the initialization is completed, generating a BEV image by performing an IPM transformation on the original infrared image.

[0020] According to an embodiment of the present application, the Shi-Tomasi corner detection algorithm is used to extract point features in the original infrared image, and the LK optical flow method is used to track the point features between consecutive frames.

[0021] An infrared inertial odometer method for ground unmanned vehicles according to an embodiment of the present invention, compared with the prior art, the beneficial effects that can be achieved by at least one of the above-mentioned technical solutions adopted in the embodiment of this specification include at least the following: First, in order to make up for the defect of poor texture of infrared images, the embodiment of the present invention screens out unstable point features and supplements the extraction of line features, and improves the tracking method of line features to improve computational efficiency. For ground line features, plane constraints are added to loosely constrain them to the ground plane. In order to overcome the degradation problem of ground mobile carriers in large-scale environments, the metric scale is restored based on the camera-ground geometric relationship, enabling the ground unmanned platform to work all day in large-scale environments. The specific performance is as follows:

[0022] 1. This embodiment of the present invention takes into account the characteristics of infrared images and screens point features to avoid the influence of low-precision point features. It also extracts line features and improves traditional line feature extraction and tracking methods, enhancing computational efficiency and accuracy. For ground line features, plane constraints are added to avoid the influence of redundant parameters.

[0023] 2. This embodiment of the present invention overcomes the scale degradation of unmanned ground platforms by taking into account the natural camera-ground geometry and restoring the metric scale without the need for additional sensor fusion. Furthermore, the camera-ground geometry model can be calibrated online without requiring accurate prior information.

[0024] 3. The algorithm of this embodiment of the present invention was compared with the open-source algorithms VINS-Fusion, PL-VINS, and EPLF-VINS. The MIE-TIO system backend proposed in this embodiment of the present invention achieves higher accuracy and real-time performance by constructing a tightly coupled graph optimization model consisting of points, lines, IMU pre-integration, and ground point and line constraints through a sliding window. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0026] Figure 1 Schematic diagram of an infrared inertial odometer method for a ground unmanned vehicle according to an embodiment of the present invention;

[0027] FIG2 is a schematic diagram of Plücker coordinates of a line feature according to an embodiment of the present invention; FIG2(a) shows Plücker coordinates; FIG2(b) shows initialization of a line feature;

[0028] Figure 3 3. Schematic diagram of three-degree-of-freedom line feature optical flow tracking according to an embodiment of the present invention;

[0029] Figure 4 Schematic diagram of a camera-ground geometric model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0030] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0031] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, in the absence of conflict, the features in the following embodiments and embodiments can be combined with each other. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative work are within the scope of protection of this application.

[0032] like Figure 1 As shown, an embodiment of the present invention provides an infrared inertial odometer method for a ground unmanned vehicle, comprising:

[0033] S101. Extracting point features and line features from an original infrared image of the target vehicle's external environment, and tracking the point features and line features between consecutive frames; wherein the ELSED algorithm is used to extract line features from the original infrared image, and a three-degree-of-freedom line feature optical flow tracking algorithm is used to track line features between consecutive frames;

[0034] S102. Establish a camera-ground geometry model. Based on the camera-ground geometry model, transform the original infrared image through IPM to generate a BEV image. Based on the point features, predict and track the features of the ground points in the BEV image. Remap the tracked ground points to the original infrared image through the inverse process of IPM.

[0035] S103. Based on the camera-ground geometric model, determine whether the extracted line feature is a ground line feature. If so, add a plane constraint to the ground line feature to constrain the ground line feature to a set solution space;

[0036] S104. Pre-integrate the IMU data between the two frames of the original infrared image, and optimize the IMU pre-integration factor, point feature reprojection factor, line feature reprojection factor, ground point feature constraint factor and ground line feature constraint factor in real time based on the sliding window.

[0037] The embodiment of the present invention mainly includes three main parts: measurement preprocessing, initialization, and sliding window optimization. First, the IMU data between two frames is pre-integrated. Common point features and line features are extracted from the original infrared image using the Shi-Tomasi and ELSED methods, respectively. Feature tracking from the BEV image can avoid distortion caused by the camera projection model and enable efficient and accurate tracking. Therefore, the embodiment of the present invention extracts and tracks ground point features based on the BEV image generated by the IPM. Then, ground line features are selected according to the camera-ground constraints provided by the ground points. Point features and line features are tracked using the traditional KLT optical flow method and the proposed three-degree-of-freedom optical flow method, respectively. Secondly, the thermal-inertial odometry (TIO) is initialized using the initialization procedure proposed in VINS-MONO (an open source visual inertial navigation system). Once a sufficient number of ground landmarks below the uncertainty threshold are identified, the camera-ground parameters are estimated. Finally, pose, 3D points, road landmarks, and IMU bias are estimated by minimizing the IMU residuals, point and line residuals, and residuals of ground point-line constraints within the sliding window of TIO.

[0038] The specific implementation mainly includes the following points:

[0039] 1. Line feature extraction and tracking

[0040] (1) Geometric representation of spatial lines

[0041] This embodiment adopts two line feature representation methods, Plücker coordinates and orthogonal representation. Plücker coordinates have a linearized form and are often used for landmark initialization, coordinate transformation, and line projection, while orthogonal representation has no redundant parameters and is suitable for nonlinear optimization processes.

[0042] As shown in Figure 2(a), the Plücker coordinates of the space line in the world system are where v w is the direction vector of the space line, n w is the normal vector of the plane determined by the space line and the origin O, and the subscript w represents the world system.

[0043] As shown in Figure 2(b), when the same road sign is observed in two frames, the Plücker matrix of the road sign can be calculated:

[0044]

[0045] Where π is the plane determined by the origin and the line endpoints, and the Plücker coordinates of the space line can be obtained from the Plücker matrix.

[0046] Constraints included in Plücker coordinates This is still an over-parameterized representation method, so in the optimization process, this embodiment adopts a 4-parameter minimized orthogonal representation.

[0047] The orthogonal representation of the space line (U, W)∈SO(3)×SO(2) can be expressed as the matrix [n w |v w ] Perform QR decomposition to obtain:

[0048]

[0049] Where U and W represent the three-dimensional and two-dimensional rotation matrices, respectively. The conversion from Plücker coordinates to orthogonal representation can also be done without QR decomposition by the following formula:

[0050]

[0051] Let R(θ) = U, R(θ) = W, and combine equations (2) and (3) to obtain:

[0052]

[0053] Where, It represents the rotation of a straight line in space around the three axes x, y, and z, and the scalar θ represents the distance from the straight line in space to the origin.

[0054] Four-parameter orthogonal representation That is, the increment defined on its tangent space, using an unconstrained orthogonal representation during the optimization process After optimizing the space lines using the orthogonal representation, the corresponding Plücker coordinates can be calculated:

[0055]

[0056] Where ω1, ω2, u1 and u2 are calculated by substituting the updated orthogonal representation into equation (4).

[0057] (2) Line feature extraction

[0058] Currently, line segment detectors (LSDs) are commonly used to detect line features in visual SLAM. However, the original LSD line feature detection algorithm was not specifically tailored for SLAM applications, resulting in relatively high computational costs. Furthermore, LSDs detect many short lines, which are unstable and difficult to match, contributing little to accuracy improvements.

[0059] To improve efficiency and accuracy, this example uses the ELSED algorithm. ELSED employs a local growth method, characterized by high speed, high accuracy, and strong repeatability. To maintain efficient line feature matching and tracking, ELSED parameters are fine-tuned to adapt to the infrared image and eliminate short, invalid line features. Finally, if two line segments are close in distance and have minimal directional differences, they are merged.

[0060] (3) Line feature tracking based on optical flow

[0061] The traditional method of line feature matching is based on line band descriptor (LBD). The method based on descriptor has a slow calculation speed, while SLAM has high real-time requirements. Drawing on the LK optical flow method of traditional point features, since the point has only 2 degrees of freedom of axial movement on the image, the line feature not only has 2 degrees of freedom of translation, but also has angle changes. Therefore, this embodiment proposes a three-degree-of-freedom line feature optical flow tracking algorithm for tracking line features between consecutive frames. The principle of this method is the same as the LK optical flow method. It is still based on the grayscale invariant assumption, but when calculating, it is no longer assumed that the pixels of a square window have the same movement, but it is assumed that the pixels in the rectangular window around the line feature have the same movement. In practical applications, in order to speed up calculation efficiency, some points can be taken on the line feature to participate in the calculation.

[0062] For a point (u, v) in the image at time t, based on the grayscale invariance assumption, at time t+dt we can get:

[0063] I(u+du,v+dv,t+dt)=I(u,v,t) (6)

[0064] Perform Taylor expansion on the left side of the above equation and retain the first-order term:

[0065]

[0066] Based on the grayscale invariance assumption, we can get:

[0067]

[0068] Will and I u , I v , Denoted as I t , and divide both sides of the above equation by dt:

[0069]

[0070] Where, I u , I vRepresents the grayscale gradient of the image in the u and v directions at that point, I t Indicates the change of image grayscale over time.

[0071] like Figure 3 As shown, take some point sets on the line features:

[0072] l t ={(u s ,v s ),(u1,v1),...,(u n ,v n ),...,(u e ,v e )} (10)

[0073] Where, (u s ,v s ) and (u e ,v e ) represent the starting point and end point of the line feature respectively.

[0074] At time t+dt, there is a line feature l t+dt , there is a corresponding relationship between the line features of consecutive frames:

[0075]

[0076] Where, l n Representative line feature l t Two points on (u s ,v s ) and (u n ,v n )’s Euclidean distance, l′ n Indicates the corresponding l t+dt x1 and x2 represent the position changes of the starting point of the line feature in the u and v directions respectively, and x3 represents the rotation change of the line feature around the starting point.

[0077] Considering that the line features usually do not change much between two adjacent frames, we can assume that n ≈l′ n , and the angle change x3 is usually small, then sinx3≈x3, cosx3≈1, and equation (11) can be simplified to:

[0078]

[0079] Divide both sides of the above equation by dt:

[0080]

[0081] Any point on the line feature should satisfy Equation (9), and the points on the line should also satisfy the collinearity constraint shown in Equation (13). Let Substituting formula (13) into formula (9) yields:

[0082]

[0083] Each point (u n ,v n ) should all meet the above requirements, resulting in an overdetermined equation for the unknown parameter X, which can be iteratively solved using the Gauss-Newton method. To improve the ability to track line features during rapid camera motion, this embodiment employs an image pyramid for line optical flow tracking. During each pyramid tracking layer, all line features are tracked in parallel to shorten computation time. Furthermore, the endpoints of spatial lines may disappear due to occlusion in consecutive frames, so the region near the line feature endpoints is not used during tracking. Thanks to the use of a faster and more accurate line feature extraction algorithm, this embodiment does not require line extension using grayscale consistency after line feature tracking. Instead, line features are extracted using the ELSED algorithm in adjacent frames. The distance between the tracked point and the line is then used to determine whether the line features in the adjacent frames are successfully matched. This results in faster computational efficiency and more accurate matching results. In practice, optical flow tracking can be difficult because points on line features lack good corner properties. Therefore, after tracking is complete, the tracked point is determined to be closest to the line feature. If the distance exceeds a certain percentage and does not exceed a set threshold, the line features of two consecutive frames are considered to be successfully matched. Finally, in continuous frame tracking, line features usually do not change much between two adjacent frames, so the matching results need to be set to ensure that the angle difference between the two matching lines does not exceed 0.1 radians.

[0084] (4) Ground line processing

[0085] Orthogonality indicates that a 3D line has four degrees of freedom, but when the line lies on the ground plane, it only has two degrees of freedom. Redundant parameters increase the uncertainty of the ground line solution and hinder optimization convergence. To constrain ground line features to the correct solution space, this embodiment adds a plane constraint to line features determined to be on the ground.

[0086] The ground plane equation in the camera system:

[0087] (R(α,θ) T p f )2-h=0 (14)

[0088] Before adding ground constraints to line features, it is necessary to first determine whether they are ground lines. The image generated by IPM is a partial image of the original image. Extracting line features on this image may cause the line features to be truncated. Therefore, this embodiment directly determines whether the line features extracted from the original image are located on the ground. Suppose a pair of matching lines in the two frames before and after are l i and l i+1 , if they are located on the ground plane, then the projection relationship is:

[0089]

[0090] Where, and The relative rotation and translation of the camera system in two adjacent frames are respectively, which can be predicted by the IMU. H is the homography matrix between the ground line features in two adjacent frames. K is the camera intrinsic parameter matrix.

[0091] Theoretically, the matching line pairs on the ground should satisfy Equation (15), and the line pairs projected by the homography matrix The angle difference and midpoint pixel distance from the matching line extracted from the i+1 frame determine whether the feature is a ground line. In experiments, this condition is satisfied when the angle difference between the projected line and the matching line is less than 0.1 radians and the midpoint pixel distance is less than 4 pixels. Furthermore, considering factors such as the experimental environment and camera installation position, a region of interest (ROI) can be pre-defined in the image. This determination is only performed when the line feature is located or partially located within the ROI, eliminating the need to perform this determination for all line features, thus improving computational efficiency.

[0092] 2. Camera-Ground Constraints

[0093] (1) Camera-ground geometry model

[0094] There is a unique geometric relationship between the onboard camera and the ground plane. In the camera system, the ground plane can be defined by its normal vector and distance. Figure 4 As shown in , the local ground plane is parameterized by the perpendicular distance from the camera to the ground plane and a two-step rotation that aligns the XZ plane of the camera system parallel to the ground plane. Therefore, a point lying on the ground plane satisfies the following equation in the camera system.

[0095] (R(α,θ) T p f )2-h=0 (16)

[0096]

[0097] Where (·)2 represents the second row of the third-order matrix, and α and θ represent the angle values ​​of the secondary rotation, that is, the roll angle and pitch angle of the camera.

[0098] These ground points can effectively utilize the prior information of camera height. Therefore, based on the camera projection model, the inverse depth of the ground point features can be instantly recovered.

[0099]

[0100] Where u f is the normalized image coordinate, λ f Represents the inverse depth of landmark point f.

[0101] (2) Ground point processing

[0102] Ground points move quickly and disappear quickly in perspective images, making them difficult to track. Therefore, an additional parallel front-end is designed to track ground point features in bird's-eye view (BEV) images. Based on formula (19), the relationship between perspective images and bird's-eye view images can be derived:

[0103]

[0104] Where u g =[x g y g 1] T represents the normalized image coordinates of the landmark f on the BEV image.

[0105] Tracking ground points is difficult because they have weak textures and move rapidly in images. However, accurately predicting the features of ground points makes tracking possible. Based on knowledge of the geometric relationship between the camera and the ground, the relative pose derived from the IMU can be used to predict the position of ground points.

[0106]

[0107] Where, and are the position of the ground point in the previous frame and the predicted position of the ground point in the current frame, respectively.

[0108] According to formulas (21) and (22), the positions of ground points can be predicted in the BEV image. This allows the search area of ​​optical flow tracking to be limited to a few pixels, significantly improving the tracking performance. Then, bidirectional iterative optical flow and a RANSAC method based on the fundamental matrix are used to remove outliers after tracking. Finally, the tracked ground points are remapped to the perspective image through the inverse process of IPM so that they can be processed together with normal features in the back end.

[0109] (3) Camera-ground model initialization

[0110] Monocular visual-inertial odometry (VIO) is scale-aware when stimulated by an inertial measurement unit (IMU). Therefore, camera-ground parameters can be initialized online. Before camera-ground parameter initialization, ground points can only be processed on perspective images. These ground points are constrained to a region of interest (ROI), which is defined by the IMU-camera extrinsics and a rough camera height. Once a sufficient number of ground landmarks below an uncertainty threshold are identified, the camera-ground parameters can be estimated.

[0111]

[0112] Where, and are the estimated camera height, roll angle, and pitch angle, respectively. and represents the relative rotation and translation estimated by TIO.

[0113] After initialization, ground points are processed in the Bird's Eye View (BEV) image. Meanwhile, ground point constraints are introduced into the estimator to restore the metric scale of the Visual Inertial Odometry (VIO) and further optimize the camera-ground parameters.

[0114] (4) Graph Optimization Model

[0115] In the embodiment of the present invention, the variables in the sliding window are as follows:

[0116]

[0117] Where, and Represents the position, attitude and velocity of IMU respectively; b a and b g are the zero bias of the accelerometer and gyroscope respectively; x k represents the state of the world system at time k, χ is the state vector contained in the sliding window; n, m and l represent the window size, the number of point and line features, respectively.

[0118] The optimization of the objective function is done by minimizing the measurement residual within the window and the prior term:

[0119]

[0120] Where r p -H p χ represents the prior information obtained by marginalization, r I is the IMU pre-integration residual, and They represent the reprojection residuals of point and line features respectively, and ρ represents the Huber robust kernel function.

[0121] (1) Calculate the pre-integration residual of the IMU observations between two consecutive frames within the sliding window.

[0122]

[0123] Where, The world system of the kth image is loaded into the system b k The rotation matrix of and Represents the position, velocity and attitude of IMU pre-integration between consecutive frames respectively; is the three-dimensional error state corresponding to the quaternion; []xyz represents the vector part of the extracted quaternion.

[0124] (2) The residual of the point feature is defined by the reprojection error. For the mth landmark point P m , P m From the first observed i-th camera coordinate system c i , converted to c j Normalized camera coordinates under , calculate the error between the projection point and the matching feature point, the degree of freedom of the point feature residual is 2, and project the residual to c j Matching feature points The corresponding tangent plane, b1 and b2 are any two orthogonal bases on the tangent plane, then the point feature residual is:

[0125]

[0126]

[0127] Where, and P m First time in c i The pixel coordinates of the observation in ; and For the corresponding j The pixel coordinates observed in c; b1 and b2 are j matching point Corresponding to any two orthogonal bases on the tangent plane.

[0128] (3) Plücker coordinates have the characteristic of linearity when performing coordinate transformation, and are suitable for use in calculating line feature reprojection residuals:

[0129]

[0130] Where, is the transformation matrix from the world system to the camera system.

[0131] The space lines of the camera frame can be projected onto the image plane:

[0132] l=[l1,l2,l3]=K L n c (29)

[0133] Where K L Represents the line projection matrix constructed from the camera intrinsic parameters.

[0134] The reprojection error of a spatial line is defined by the distance from the midpoint of the matching line to the projected line. Due to reasons such as occlusion of line features and misdetection during operation, the detected line feature endpoints may not be consistent with the spatial line. Only the part of the matching line segment that is perpendicular to the projected line can provide an effective constraint. Therefore, this embodiment defines the distance from the midpoint of the matching line to the projected line. For c j The observed line L i have:

[0135]

[0136] Where, Represents the homogeneous coordinates of the midpoint of the detected matching line segment.

[0137] The objective function is iteratively minimized by the Gauss-Newton method using an orthogonal representation with 4 parameters. After that, the four parameters of the line feature can be updated according to formulas (2)-(4):

[0138]

[0139] Where I represents the identity matrix.

[0140] (4) When the initial frame coincides with the target frame, the camera-ground constraint residual is:

[0141]

[0142] Otherwise:

[0143]

[0144] Where, the ground point is first observed in the ith frame image, and the target projection frame is the jth frame. After the camera-ground parameters converge, the constraints can provide metric-scale geometric information to the estimator.

[0145] (5) When the line feature is located on the ground, the following plane constraints exist:

[0146]

[0147] Where, is the homogeneous coordinate of the ground plane equation, which can be calculated by equation (14). The plane constraint consists of two terms, where the first term represents that the line feature direction vector is perpendicular to the ground normal vector, and the second term indicates that the foot of the perpendicular from the origin to the line satisfies the plane equation, that is, it constrains the distance from the origin to the line.

[0148] Based on the unique thermal imaging characteristics of infrared cameras, embodiments of the present invention propose a point-line combined infrared inertial odometry (TIO) for ground unmanned vehicles. First, based on more advanced line feature extraction and tracking methods, a real-time point-line combined TIO is implemented. Second, a parallel ground point feature tracking front-end is developed, utilizing ground point constraints to overcome the scale degradation problem of monocular TIO. Furthermore, the redundant modeling of ground line features is resolved based on the camera-ground geometric relationship. Finally, a common back-end optimizes the IMU pre-integration factor, point feature reprojection factor, line feature reprojection factor, camera-ground constraint factor, and ground line feature constraint factor in real time using a sliding window. The proposed method in embodiments of the present invention was validated on a measured dataset and compared with currently leading open-source visual-inertial SLAM algorithms VINS-Fusion, PL-VINS, and EPLF-VINS. Experimental results demonstrate that MIE-TIO can achieve accurate scale recovery even with insufficient excitation for ground unmanned vehicles. The algorithm of the present invention achieves accuracies that are 67.88% to 96.73% higher than those of VINS-Fusion, PL-VINS, and 73.91% to 98.78% higher than those of EPLF-VINS. It can achieve a relative translation error of 3% to 9%. Compared to traditional algorithms, the algorithm of the present invention demonstrates superior performance, enabling accurate monocular TIO even when ground unmanned vehicles are under-excited. Its accuracy and robustness are significantly improved compared to those of VINS-Fusion, PL-VINS, and EPLF-VINS.

[0149] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An infrared inertial odometer method for ground unmanned vehicles, characterized in that: include: Extracting point features and line features from an original infrared image of the target vehicle's external environment, and tracking the point features and line features between consecutive frames; wherein the line features in the original infrared image are extracted using an ELSED algorithm, and the line features between consecutive frames are tracked using a three-degree-of-freedom line feature optical flow tracking algorithm; Establishing a camera-ground geometric model, based on the camera-ground geometric model, transforming the original infrared image through IPM to generate a BEV image, predicting and tracking the features of the ground points in the BEV image according to the point features, and remapping the tracked ground points to the original infrared image through the inverse process of IPM; Based on the camera-ground geometric model, determining whether the extracted line feature is a ground line feature, and if so, adding a plane constraint to the ground line feature to constrain the ground line feature to a set solution space; The IMU data between the two frames of the original infrared image are pre-integrated, and the IMU pre-integration factor, point feature reprojection factor, line feature reprojection factor, ground point feature constraint factor and ground line feature constraint factor are optimized in real time based on a sliding window.

2. The infrared inertial odometer method for ground unmanned vehicles according to claim 1, characterized in that: A three-degree-of-freedom line feature optical flow tracking algorithm is used to track line features between consecutive frames, including: The grayscale of the original infrared image between consecutive frames is set to be unchanged, and the pixel motion in the rectangular window around the line feature is set to be consistent, and the line feature optical flow tracking is performed using an image pyramid.

3. The infrared inertial odometer method for ground unmanned vehicles according to claim 1, characterized in that: Determining whether the extracted line feature is a ground line feature includes: Set a matching line angle difference threshold and a midpoint pixel distance threshold, and compare the angle difference and midpoint pixel distance of the matching line extracted after projecting the homography matrix between the ground line features of two adjacent frames with the matching line angle difference threshold and the midpoint pixel distance threshold respectively to determine whether the line feature is a ground line feature.

4. The infrared inertial odometer method for ground unmanned vehicles according to claim 3, characterized in that: The matching line angle difference threshold is 0.1 radians, and the midpoint pixel distance threshold is 4 pixels.

5. The infrared inertial odometer method for ground unmanned vehicles according to claim 1, characterized in that: The method further comprises: The features of ground points are predicted and tracked in the BEV image, and outliers of ground points are removed after tracking using bidirectional iterative optical flow and a fundamental matrix-based RANSAC method. Finally, the tracked ground points are remapped to the original infrared image through the inverse process of IPM.

6. The infrared inertial odometer method for ground unmanned vehicles according to claim 1, characterized in that: The method further includes initializing camera-ground parameters of the camera-ground geometric model, and after the initialization is completed, performing IPM transformation on the original infrared image to generate a BEV image.

7. The infrared inertial odometer method for ground unmanned vehicles according to claim 1, characterized in that: The Shi-Tomasi corner detection algorithm is used to extract point features in the original infrared image, and the LK optical flow method is used to track point features between consecutive frames.