A monocular thermal imaging simultaneous localization and mapping method and system

By introducing real-time denoising, static feature extraction and object detection neural networks into the monocular thermal imaging SLAM system, the problem of insufficient positioning accuracy and robustness of thermal imaging SLAM in dynamic environments is solved, and a high-precision and robust thermal imaging SLAM system is realized.

CN117036404BActive Publication Date: 2025-06-13BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310995674.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-09
Publication Date
2025-06-13
Estimated Expiration
2043-08-09

AI Technical Summary

Technical Problem

The existing thermal imaging SLAM methods have shortcomings in positioning accuracy and robustness, especially in dynamic environments, where the low signal-to-noise ratio and weak texture characteristics of thermal infrared images lead to low positioning accuracy, and traditional methods have problems such as false deletion and data loss when processing dynamic targets.

Method used

A method and system for simultaneous positioning and mapping of monocular thermal imaging is proposed. Through real-time noise denoising module, static feature extraction module, initialization module, feature tracking module, local mapping module and loop detection module, combined with object detection neural network, mobile object tracking method and instance segmentation technology, high-precision and robust positioning are achieved.

Benefits of technology

High-precision and robust positioning are achieved in the dynamic environment of visual degradation, improving the signal-to-noise ratio of thermal infrared images, broadening the application boundaries of visual SLAM systems in thermal infrared images, and demonstrating excellent accuracy and robustness on multiple real-world dynamic environmental data sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117036404B_ABST
    Figure CN117036404B_ABST
Patent Text Reader

Abstract

The present invention relates to a monocular thermal imaging simultaneous localization and mapping method and system. The system includes a collector and a processor. The collector acquires a thermal infrared video sequence image that is not interrupted by non-uniformity correction. The processor includes a denoising module, a feature extraction module, an initialization module, a feature tracking module, a local mapping module, and a loop detection module. The present invention improves the signal-to-noise ratio of thermal infrared images, ensures that thermal information is maximally retained in different environments, and broadens the application boundary of visual SLAM methods in thermal infrared images. The present invention also uses multiple strategies to eliminate the interference of dynamic targets on the SLAM system, can achieve high robustness in complex dynamic environments, and demonstrates high positioning accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of infrared thermal imaging, and in particular to a monocular thermal imaging simultaneous localization and mapping method and system. Background Art

[0002] Although China's Beidou Navigation Satellite System has become increasingly complete, in places such as indoors, tunnels, and urban canyons, or when satellite signals are interfered with or even denied, the positioning and navigation performance of vehicles and unmanned vehicles (or drones) using the Beidou system is restricted or even lost. Therefore, using the laser point cloud obtained by lidar and the image information obtained by an imaging system for positioning and navigation has become a research hotspot, which are respectively called laser simultaneous localization and mapping (abbreviated as laser SLAM) and visual simultaneous localization and mapping (abbreviated as visual SLAM).

[0003] Currently, a large number of visual SLAM systems use visible light imaging detectors as imaging detectors. Compared with visible light imaging systems, thermal imaging systems (or infrared thermal imagers) have the advantages of working day and night, being free from glare interference, and being able to penetrate smoke, haze, light rain, and light snow. However, compared with visible light images, thermal infrared images have insurmountable weaknesses, such as low signal-to-noise ratio, few scene details, weak texture features, and low spatial resolution (currently, the spatial resolution of the most advanced infrared focal plane detector is only one million pixels, but the spatial resolution of visible light imaging detectors used in daily life, such as mobile phones, is tens of millions of pixels, which is dozens of times or even hundreds of times that of infrared focal plane detectors), resulting in low positioning accuracy of existing thermal imaging SLAM methods. Secondly, the thermal imaging system needs to perform non-uniformity correction. A common non-uniformity correction method is to use a baffle to block the optical path for several seconds, which will cause data loss in a dynamic environment, thus seriously affecting the continuity of positioning.

[0004] In addition, traditional SLAM uses the random sample consensus method to eliminate incorrect data associations caused by dynamic targets, but when the proportion of dynamic targets in the entire image is relatively large, this method will fail. Another method is to use a convolutional neural network (CNN) to detect all possible moving targets in the image and delete them all, but this method also deletes some static targets by mistake. When traditional SLAM is used for thermal infrared images with weak textures, there are fewer features available for tracking, seriously affecting the effect of SLAM.

[0005] Therefore, there are many technical difficulties in using a thermal imaging system for visual SLAM, which is also the reason why infrared SLAM has not been widely developed. Summary of the Invention

[0006] To solve these problems, based on a full analysis of the low signal-to-noise ratio and weak texture characteristics of thermal infrared images, the present invention proposes a monocular thermal imaging simultaneous localization and mapping method and system, which can perform high-precision and robust localization in a dynamically degraded visual environment.

[0007] The technical solution of the present invention is specifically as follows:

[0008] A monocular thermal imaging simultaneous localization and mapping system includes a collector and a processor; the processor includes a denoising module, a static feature extraction module, an initialization module, a feature tracking module, a local mapping module, and a loop detection module; where:

[0009] The collector acquires thermal infrared video sequence images that are not interrupted by non-uniformity correction;

[0010] The denoising module performs real-time denoising on the acquired thermal infrared video sequence images;

[0011] The static feature extraction module extracts the key points and descriptors of point features and the key line segments and descriptors of line features in the image respectively; uses an object detection neural network for object detection, adopts a moving object tracking method to improve accuracy, uses an instance segmentation method to reduce bounding boxes, realizes pixel-level segmentation of moving objects, and then uses epipolar geometric constraints to screen out static features for subsequent tracking;

[0012] The initialization module performs the following processing: restores the poses of two adjacent frames based on the static point features and static line features extracted in the static feature extraction module;

[0013] The feature tracking module performs the following processing: map point and map line tracking; key frame selection;

[0014] In map point and map line tracking, data association is performed successively using inter-frame tracking and local map tracking to update the pose of the current frame;

[0015] In key frame selection, for a thermal infrared image, the key frame with the most common observations is used as the reference key frame;

[0016] The local mapping module updates the connection relationship between key frames through the co-visibility relationship of map points and map lines, and eliminates poor-quality map points and map lines according to the observation situation;

[0017] The loop detection module uses image appearance-based similarity to determine whether a loop occurs, and uses the descriptors of point features and the descriptors of line features to train a bag-of-words model based on thermal infrared images offline through clustering.

[0018] Furthermore, in real-time scene denoising:

[0019] First, perform scene-based non-uniformity correction on the thermal infrared image; at the same time, when information is lost during the compression conversion of the thermal infrared image, use the adaptive histogram equalization method to retain the original information of the image to the greatest extent; then remove the random noise of the thermal infrared image.

[0020] Further, for random noise, use FFDNet based on convolutional neural network to denoise the thermal infrared image;

[0021] Reshape the original image into multiple downsampled sub-images, and connect them with an adjustable noise level map and input them into the convolutional neural network to perform convolution, and finally generate a denoised image of the same size as the original input image.

[0022] Further, when extracting line segment features, filter out extremely short invalid line features; connect the segmented long line segments; if the lengths and distances of two line segments are very close and the direction difference is small enough, merge such line segments; use descriptors to describe the new line segments, and add geometric constraints to remove possible abnormal matches. The geometric constraints include:

[0023] (1) The length difference and angle difference of the matching line segment pair are less than a certain value;

[0024] (2) The distance between the matching line segment pair is less than a certain value;

[0025] (3) The descriptor distance of the matching line segment pair is less than a certain value;

[0026] In the moving object tracking method, the target state detected by the target detection neural network is defined as:

[0027]

[0028] Among them, x, y represent the center coordinates of the detected target, and a, p represent the area and aspect ratio of the bounding box, represents the center coordinates and area of the predicted target in the next frame; during the tracking process, the motion prediction and data association of the target are performed through the Kalman filter and the Hungarian algorithm, and the intersection over union distance of the detection boxes is used to calculate the cost matrix.

[0029] Further, in the epipolar geometry constraint, judge the dynamic situation of the key points in the prior target area in the thermal infrared image; if there are more than a certain number of dynamic key points in the object area, it is considered that the object is dynamic:

[0030] First, obtain the detection and segmentation results of the target; second, obtain the key-point matching results of consecutive frames using coarse-to-fine optical flow; third, use the Random Sample Consensus (RANSAC) algorithm with the most points in the non-target area to solve for the fundamental matrix, then calculate the epipolar lines using the fundamental matrix, and finally determine that the key points with a distance greater than the threshold from the epipolar lines are moving.

[0031] Furthermore, the initialization module processes as follows:

[0032] For a pair of parallel 3D lines, their 2D projections \(l\) 1 , \(l\) 2 intersect at the vanishing point \(l\) 1 × \(l\) 2 , and its normalized direction

[0033] Similarly, the normalized direction of the vanishing point of their 2D projections \(m\) 1 , \(m\) 2 in the second frame image captured by the collector In addition, assume that the normalized directions of the vanishing points of the 2D projections of another pair of parallel 3D lines in the first frame image and the second frame image are \(v\) p3 and \(v\) p4 , respectively. Then, for these two pairs of parallel 3D lines, the rotation matrix between the two images ideally satisfies the following formula:

[0034]

[0035] where \(g\) 1 and \(g\) 2 are constants; based on the above formula, solve for the rotation matrix \(R\) in the case where \(\|g\) 1 \| 2 + \(\|g\) 2 \| 2 is minimized;

[0036] Consider two feature 2D points \(p\) 1 , \(p\) 2 detected in the first image using a point feature detector, and after feature matching, determine that they correspond to points \(q\) 1 , \(q\) 2 in the second image. The actual translation vector is defined as:

[0037]

[0038] At the same time, parallelly use the initialization method in ORB-SLAM2, and use the Random Sample Consensus (RANSAC) method in the initialization method to sample point pairs and line pairs to select the best results.

[0039] Further, in map point and map line tracking, first, assuming the system moves at a constant speed, pose estimation is performed through the projection matching relationship between adjacent frames; project the map points and map lines onto the current frame, and search for point features and line features that meet the requirements within a given range. For the matching of point features, in addition to requiring the shortest descriptor distance, rotational consistency also needs to be satisfied. If the number of effectively matched features after the above search is still less than the given threshold, the search area is expanded until the requirements are met; the above method obtains the initial value of the current pose. Next, use the 3D-2D projection relationship to optimize the pose of the current frame, that is, make the reprojection error of 3D points and 3D lines minimum through bundle adjustment; for the inter-frame pose estimation of the constant velocity model, use the rotation matrix R * and translation vector t * of the current frame as the state variables to be optimized to construct a graph optimization model, and use the Levenberg-Marquardt method to minimize the following cost function for iterative solution:

[0040]

[0041] where ρ p and ρ l are Huber robust cost functions, and the Huber function is introduced to reduce the abnormal terms of the cost function; are the minimum reprojection errors of points and lines respectively;

[0042] ∑ p and ∑ l represent the observation covariance matrices of points and lines;

[0043] χ c represents the set of matching pairs between consecutive image frames in the video sequence. Considering the computational complexity, directly use the Jacobian matrix for solution; after performing pose optimization, remove the outliers and outer lines in the map; if the number of successfully matched map points and map lines is greater than a certain value, it is considered successful.

[0044] Adopt the matching strategy between images; convert the descriptors of features into bag-of-words vectors to accelerate the matching between the current frame and the reference key frame, and the reference key frame is the key frame with the highest co-visibility with the current frame; if the number of matched features still does not meet the standard, then choose to use the nearest neighbor matching algorithm, and use the random sample consensus method to calculate the homography matrix between images to obtain enough correct matching features; finally, project the map points and map lines onto the current frame, and use the pose of the previous frame as the initial value to optimize the pose according to the equation;

[0045] If the above methods all fail to track and result in positioning failure, repositioning is required; first, convert the current frame into a bag of words, find a candidate key frame group similar to the current frame in the key frame database, and select candidate key frames that meet the requirements from them; once the matching requirements are met, estimate the camera pose of the current frame by solving the PnP problem and optimize the pose according to the equation; if the number of inliers after optimization is too small, project the un-matched map points and map lines in the key frame onto the current frame by projection to generate new matching relationships; according to the results of the projection matching, optimize the pose again; as long as one candidate key frame is successfully repositioned, the remaining candidate key frames will no longer be considered, otherwise, repeat for the next frame until repositioning is successful;

[0046] In key frame selection, a frame is selected as a key frame under any of the following circumstances:

[0047] (1) Some frames have passed since the last global repositioning;

[0048] (2) Some frames have passed since the insertion of the last key frame or the local map construction thread is idle;

[0049] (3) The number of map points and map lines tracked by the current frame is less than a certain proportion of the number of map points and map lines tracked by the reference key frame;

[0050] (4) There is a certain degree of pose transformation from the last key frame;

[0051] (5) The current frame has tracked at least a certain number of feature points and spatial lines.

[0052] Furthermore, the local mapping module uses the current key frame and its adjacent co-visible key frames to generate new map points and map lines through triangulation to ensure more stable tracking; finally, check and fuse duplicate map points and map lines, and when the number of key frames in the local map is greater than a certain number, perform local bundle adjustment according to the following equation to adjust the camera pose in the local map Map points and map lines

[0053]

[0054] where, and represent the set of matching pairs of points and lines in the local map;

[0055] After optimization and adjustment, determine whether more than a certain percentage of the map points and map lines tracked by the key frame can be tracked by other key frames, and accordingly remove redundant key frames, and then add the current frame to the loop closure detection queue.

[0056] Furthermore, in the loop closure detection module, first calculate the similarity scores between the current key frame and the bag-of-words vectors of each co-visible key frame. The similarity is defined as:

[0057]

[0058] where p and l respectively represent weight coefficients, and s p (η a , η b ) represents the point feature similarity between images, and s l (η a , η b ) represents the line feature similarity between images; by finding the set of loop closure candidate frames among all key frames, determine whether a loop closure is successful; if successful, adjust the pose, map points, and lines through the solved similarity transformation, and finally perform global bundle adjustment to achieve optimality.

[0059] The present invention also relates to a monocular thermal imaging simultaneous localization and mapping method applicable to the above system, including: collecting thermal infrared images, image denoising, feature extraction, initialization, feature tracking, local mapping, and loop closure detection.

[0060] In the present invention, first, use an unobstructed uncooled infrared focal plane detector assembly for scene-based non-uniformity correction, and then perform thermal infrared image denoising. This not only avoids data interruption of the thermal imager but also significantly improves the quality of thermal infrared images and broadens the boundary of applying thermal infrared images in the visual SLAM system.

[0061] Secondly, combine epipolar constraint and semantic segmentation to reduce the interference of dynamic objects and improve the robustness of thermal SLAM in dynamic scenes.

[0062] Thirdly, overcome the disadvantage of poor spatial texture distribution in thermal infrared images.

[0063] We conducted experiments on more than 180,000 real-world thermal infrared images in small-scale indoor and outdoor sequences and large-scale driving sequences. Our results show that the system of the present invention can achieve camera localization and sparse structure map reconstruction in the face of visual degradation and moving objects in a dynamic environment without relying on any other sensors. In indoor and outdoor sequences, the positioning accuracy measured by the relative position error is less than 0.1 m of the ground truth trajectory.

[0064] In addition, the system of the present invention is the only system that successfully tracks all sequences and has higher positioning accuracy and unparalleled robustness compared with the current state-of-the-art monocular SLAM systems. Therefore, this system can be used as a brand-new positioning solution to replace expensive commercial navigation systems, especially in challenging urban dynamic environments with changing lighting and low visibility in the air.

[0065] The contributions of the present invention are as follows in four aspects:

[0066] The present invention proposes real-time comprehensive denoising based on the environment, avoiding data interruption during the operation of NUC of the thermal imager, improving the signal-to-noise ratio of thermal infrared images, ensuring that thermal information is maximally retained in different environments, and broadening the application boundary of visual SLAM methods in thermal infrared images;

[0067] The present invention combines semantic segmentation with epipolar constraint, eliminating the interference of dynamic objects on the SLAM system and enabling high robustness in complex dynamic environments;

[0068] Aiming at the disadvantage of weak texture in thermal infrared images, the present invention designs and implements a complete monocular thermal SLAM system using point and line features, including initialization, tracking, mapping, loop detection, and global optimization;

[0069] The system of the present invention is compared with existing state-of-the-art monocular SLAM systems on multiple real-world dynamic environment data sequences, demonstrating excellent accuracy and robustness. In addition, the system of the present invention is the only system that can completely track all data sequences. Brief Description of the Drawings

[0070] Figure 1 Neural network structure for thermal infrared image denoising according to an embodiment of the present invention;

[0071] Figure 2 Epipolar constraint of points in a dynamic environment according to an embodiment of the present invention;

[0072] Figure 3 Minimum reprojection error of points and lines according to an embodiment of the present invention;

[0073] Figure 4 System block diagram according to an embodiment of the present invention;

[0074] Figure 5 Framework for dynamic target area segmentation according to an embodiment of the present invention;

[0075] Figure 6 Projection of parallel 3D lines onto the image plane and its vanishing point direction according to an embodiment of the present invention;

[0076] Figure 7 Illustration of local map optimization, including camera pose, map points, map lines, and visual measurements;

[0077] Figure 8 Calibration of the thermal imager using a checkerboard calibration device made of special materials according to an embodiment of the present invention;

[0078] Figure 9Screenshots of the thermal infrared image sequences collected for the embodiments of the present invention; from left to right are the outdoor night sequence, the night driving sequence, the low-light indoor sequence, and the night light rain driving sequence;

[0079] Figure 10 Are the original image and the processed thermal infrared image; (1) On the left and right are the original thermal infrared image and the thermal infrared image after performing CLAHE respectively; (2)-(3) On the left is the original thermal infrared image, in the middle is the high-quality denoised thermal infrared image from the database, and on the right is the thermal infrared image after using FFDNet;

[0080] Figure 11 (1) of is point feature matching; (2) is line feature matching; the top row represents the feature matching of the original thermal infrared image, and the bottom row represents the feature matching of the denoised thermal infrared image; the green connecting lines between the images represent feature matching;

[0081] Figure 12 Is dynamic target detection; this sequence contains five representative frames, illustrating the detection and tracking process of multiple pedestrians; top row: single-frame detection method, the undetected people are circled in red; bottom row: the proposed enhanced detection method;

[0082] Figure 13 Examples of filtering static features for tracking in a driving scenario. The first column shows the detection and segmentation of moving objects; the second and third columns show the initial point features and line features respectively; the fourth column shows the tracking of optical flow, where the key points that do not satisfy the epipolar constraint are purple, and the rest are green; the fifth and sixth columns are the static point features and line features for tracking after filtering respectively;

[0083] Figure 14 Are the results of different initializations; (a) is the proposed initialization method. The left column represents the initialization map, and the right column of colored line segments represents different parallel line segments detected in the thermal infrared image; (b) is the traditional initialization method; the left column represents the map where initialization fails, and the right column of green line segments represents the tracking of point features;

[0084] Figure 15 (a)-(c) are to compare the trajectories generated by MonoThermal-SLAM, DSO, SVO2.0, and ORB-SLAM3 on the Indoor 1 3 dataset with the trajectory generated by LeGO-LOAM (gray dashed line); (d) is a detailed comparison of the Indoor 2 trajectory with the ground truth (gray dashed line) on the x, y, and z axes;

[0085] Figure 16(a)-(f) compare the trajectories generated by MonoThermal-SLAM, DSO, SVO2.0, and ORB-SLAM3 on the Outdoor 1 6 dataset with the trajectory generated by LeGO-LOAM (gray dashed line);

[0086] Figure 17 (a) shows the mapping effect of MonoThermal-SLAM; (b) is a scene snapshot of the Outdoor 1 sequence; (c) is the mapping effect of LeGO-LOAM; (d) shows the dynamic objects in the Outdoor 1 sequence;

[0087] Figure 18 Compare the trajectories of the six driving sequence datasets obtained by MonoThermal-SLAM, DSO, SVO2.0, and ORB-SLAM3 with the trajectory obtained by RTK-GNSS as the ground truth (gray dashed line).

[0088] (a) in Figure 19 shows the trajectory and global map of the MonoThermal-SLAM system in the Driving5 sequence. The top row shows the scene features and mapping effect. The bird's-eye view shows a high similarity with the trajectory, and the detected trajectory loop closes on the right. Figures 19(b) and (c) show the comparison of the tracking and position estimation performance of MonoThermal-SLAM, DSO, SVO2.0, ORB-SLAM3, and RTK-GNSS in the Driving 5 sequence. Detailed implementation manner

[0089] Next, in combination with the embodiments of the present invention, the technical solutions in this embodiment will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0090] The monocular thermal imaging simultaneous localization and mapping system in this embodiment includes a collector and a processor. The processor includes a denoising module, a feature extraction module, an initialization module, a feature tracking module, a local mapping module, and a loop detection module. Among them: The collector uses a non-shutter thermal imaging system to collect thermal infrared video sequence images that are not interrupted by non-uniformity correction, as Figure 4 shown.

[0091] The first step of the system in this embodiment is to remove the random noise of the thermal infrared image in real time. We proposed a comprehensive denoising chain from the perspectives of the detector module and the processing algorithm respectively. In view of the information loss in the compression conversion of the thermal infrared image, we first maximize the retention of the original information of the image through the adaptive histogram equalization transformation.

[0092] For random noise, we consider denoising with minimal loss of the original information of the image. The traditional BM3D is time-consuming and memory-intensive. Considering the real-time requirements of the SLAM system, we chose FFDNet based on DnCNN to denoise the thermal infrared image. It defines the denoising model between the input noise observation σ and the expected output ξ as:

[0093]

[0094] where M is the noise level map related to the noise level, which is used as the input together with σ, and Θ is the model parameter.

[0095] As Figure 1 shown, it reshapes the original image into multiple downsampled sub-images, connects them with the adjustable noise level map and inputs them into the convolutional neural network to perform operations such as convolution, Rectified Linear Unit, and Batch Normalization, and finally generates a denoised image of the same size as the original input image. It can not only use a single network to process a wide range of noise levels, but also specify the non-uniformity noise level map to remove spatially varying noise, which makes it faster without sacrificing denoising performance.

[0096] In the feature extraction module, it includes the following aspects:

[0097] 1. Feature extraction

[0098] After removing all the noise from the thermal infrared image, we first extract features. Based on ORB-SLAM2, we chose ORB (Oriented FAST and Rotated BRIEF) features, which are excellent in detecting key points and have high computational and matching efficiency. At the same time, we make the ORB feature points distribute as evenly as possible in the image during the feature extraction stage.

[0099] For line segment extraction, we use ELSED (Enhanced line SEgment drawing). Due to the adoption of the local segment growth algorithm, it is fast, has high precision and repeatability, and can also be replaced by other line feature extraction algorithms, such as LSD (Line Segment Detector). To ensure the matching and tracking efficiency of line features, we adjusted the parameters of ELSED to be applicable to thermal infrared images and filtered out extremely short invalid line features. Connect the segmented long line segments. If the lengths and distances of two line segments are very close and the direction difference is small enough, we merge such line segments. Finally, we use LBD (Line Band Descriptor) to describe the new line segments (any line segment description algorithm can be used), and add geometric constraints to remove possible abnormal matches. In practice, a pair of successfully matched line segments should meet the following conditions:

[0100] (1) The length difference and angle difference of the matched line segment pair are less than a certain value;

[0101] (2) The distance between the matched line segment pair is less than a certain value;

[0102] (3) The distance of the LBD descriptors of the matched line segment pair is less than a certain value.

[0103] 2. Semantics

[0104] Moving objects, such as walking people and driving cars, especially in urban scenes, will deceive the feature association of visual SLAM systems, thus destroying the quality of state estimation and even causing the system to fail. Thermal infrared images have advantages in detecting moving objects because the movement of objects generates heat, resulting in a temperature higher than the ambient temperature. On the one hand, this embodiment is based on the advantages of thermal infrared images, and on the other hand, it balances accuracy and real-time performance. We first use the object detection neural network YOLOv7 for object detection, but its accuracy is still insufficient to meet the requirements of SLAM. Therefore, we adopt the moving object tracking method (SORT). In addition, an instance segmentation method is used to reduce the bounding boxes and achieve pixel-level segmentation of moving objects.

[0105] YOLOv7 is trained on the CVC-14 dataset and the thermal infrared image dataset we collected. This dataset contains a total of 20,000 thermal infrared images for object detection and 10,000 thermal infrared images for semantic segmentation. This dataset can detect and segment 7 common types of moving objects, including people, cars, trucks, buses, dogs, bicycles, and motorcycles. Figure 5 is a framework for dynamic target region segmentation, showing examples of thermal infrared images and their corresponding ground truth annotations.

[0106] In the SORT method, the target state detected by YOLOv7 is defined as:

[0107]

[0108] Among them, x and y represent the center coordinates of the detected target, and a and p represent the area and aspect ratio of the bounding box. It represents the center coordinates and the area of the bounding box of the predicted target in the next frame. During the tracking process, we use the Kalman filter and the Hungarian algorithm for target motion prediction and data association, and calculate the cost matrix using the intersection over union distance.

[0109] In instance segmentation, for moving objects, if the proportion of the area occupied by the detection box in the image is less than 1%, we consider their impact on SLAM to be limited, so we directly ignore them. In addition, we only segment the bounding box area instead of the entire image, which will reduce the computational complexity and improve the running speed of SLAM. After segmentation, the epipolar constraint is used to determine the dynamic situation of the target, and all dynamic points and line features are removed.

[0110] In the epipolar constraint, x 1 , x 2 represent the matching key point pairs of consecutive frames, and their homogeneous coordinate forms are: X 1 =(μ 1 , v 1 , 1) T , X 2 =(μ 2 , v 2 , 1) T Among them, μ i , v i represent the coordinates of the key point in the image frame where it is located. From the epipolar constraint, we know that:

[0111]

[0112] where F is the fundamental matrix. Then the epipolar line L 2 can be calculated by this formula:

[0113] L 2 =[a, b, v] T =KX 1

[0114] Among them, a, b, and c are the representations of the line vectors. Then the distance e from the matching key point to the epipolar line:

[0115]

[0116] If point c is a static point, then point x 2 and the epipolar line L 2The distance e < ε; conversely, if the point P moves to the point P′, then the matching point x′ and the epipolar line L 2 have a distance e > ε. By deleting the feature points of the dynamic target, the outliers affecting the pose estimation can be eliminated.

[0117] In the initialization module:

[0118] After feature selection, it is first necessary to initialize the thermal environment map. A common way to initialize a monocular camera is to estimate the essential matrix or the homography matrix through the matching point pairs between two frames. This may be difficult to initialize successfully for low-texture thermal infrared images due to insufficient feature points.

[0119] Therefore, we propose a SLAM map initialization solution based on the combination of points and lines. Our method is a supplement for scenes with weak texture but rich structural features, such as urban environments in thermal infrared images.

[0120] Usually, the initialization of line segments uses the trifocal tensor to estimate the relative pose relying on three images. Different from this, as Figure 6 shown, our initialization method only requires two views to perform pose estimation. We assume that there are multiple pairs of parallel 3D lines in the environment, which is common in urban environments. For a pair of parallel 3D lines L 1 , L 2 , its 2D projections l 1 , l 2 in the first frame image captured by the collector intersect at the vanishing point l 1 ×l 2 , and its normalized direction Similarly, the normalized direction of the vanishing point of its 2D projections m 1 , m 2 in the second frame image captured by the collector

[0121] Similarly, the normalized direction of the vanishing point of its 2D projections m 1 , m 2 in the second frame image captured by the collector In addition, assume that the normalized directions of the vanishing points of the 2D projections of another pair of parallel 3D lines in the first frame image and the second frame image are v p3 and v p4 , respectively. Then, for these two pairs of parallel 3D lines, the rotation matrix between the two images ideally satisfies the following formula:

[0122]

[0123] where a 1 , a 2 = ±1 and g 1 , g2 = 0. In reality, due to the existence of noise, the above equation cannot be strictly satisfied, but the 3x3 matrix B can be decomposed by singular value decomposition (SVD):

[0124]

[0125] to obtain the rotation matrix under the condition of minimizing ‖g 1 ‖ 2 +‖g 2 ‖ 2

[0126]

[0127] Since a 1 , a 2 is unknown, we use geometric criteria to obtain the correct rotation matrix. Although the thermal infrared image has weak texture, some feature points can still be detected and matched, and the translation vector can be estimated from the correspondence of two points. Consider two 2D points p 1 , p 2 , after feature matching, it is determined that they correspond to q 1 , q 2 in the second image. Ideally, it satisfies:

[0128] (Rp i ×q i ) T t = 0

[0129] Therefore, the actual translation vector f can be defined as:

[0130]

[0131] To ensure the robustness of initialization, we parallelly use the initialization method in ORB-SLAM2 (an existing method) and simultaneously use the random sample consensus method to sample point pairs and line pairs to select the best result.

[0132] In the feature tracking module:

[0133] 1. Map point and map line tracking

[0134] After the map initialization is completed, we use inter-frame tracking and local map tracking in sequence for data association to update the current frame pose. Our method draws on the idea of ORB-SLAM2 and makes some adjustments for thermal infrared images.

[0135] ​First, assume that the system moves at a constant speed and estimate the pose through the projection matching relationship between adjacent frames. Project the map points and map lines to the current frame, and search for point features and line features that meet the requirements within a given range. In addition to requiring the shortest descriptor distance, the matching of point features also requires rotation consistency. The matching requirements of line features are as described in feature extraction.

[0136] If the number of effectively matched features is still less than the given threshold after the above search, the search area is expanded until the requirement is met. Through the above method, we obtain the initial value of the current pose. Next, we use the 3D-2D projection relationship to optimize the current frame pose, that is, through the bundle adjustment method, the reprojection error of 3D points and 3D lines is minimized, as shown in Figure 3 shown.

[0137] For the inter-frame pose estimation of the constant velocity model, we use the rotation matrix R of the current frame * ∈SO(3)and translation vector As the state variable to be optimized, a graph optimization model is constructed, and the Levenberg-Marquardt method is used to minimize the following cost function and solve it iteratively:

[0138]

[0139] Among them, ρ p and ρ l is the Huber robustness cost function. Introducing the Huber function can reduce the abnormal terms of the cost function. p and∑ l Represents the observation covariance matrix of points and lines.

[0140] It represents the set of matching pairs between consecutive frames. Considering the computational complexity, we directly use the Jacobian matrix to solve. After performing pose optimization, remove the external points and lines in the map. If the number of successfully matched map points and map lines is greater than a certain value, it is considered successful.

[0141] Although the tracking method based on the constant velocity model is more efficient, there may be too little data association between adjacent frames due to the violent camera motion. At this time, we adopt an image-to-image matching strategy. We convert the feature descriptor into a bag-of-words vector to accelerate the matching of the current frame with the reference keyframe, which is the keyframe with the highest degree of co-viewing with the current frame. If the number of matched features still does not meet the standard, the matching method using the nearest neighbor matching algorithm is selected, and the homography matrix between images is calculated using the random sampling consistency method to obtain enough correct matching features. Finally, the map points and map lines are projected to the current frame, and the pose of the previous frame is used as the initial value, and the pose is optimized according to the equation.

[0142] If the above methods all fail to track and lead to positioning failure, repositioning is required at this time. First, convert the current frame into a bag of words, find a group of candidate key frames similar to the current frame in the key frame database, and select candidate key frames that meet the requirements from them.

[0143] Once the matching requirements are met, estimate the camera pose of the current frame by solving the PnP problem and optimize the pose according to the equation. If the number of inliers after optimization is too small, project the unmapped map points and map lines in the key frame into the current frame by projection to generate new matching relationships. According to the results of the projection matching, optimize the pose again. It should be noted that as long as one candidate key frame is successfully repositioned, the remaining candidate key frames will no longer be considered. Otherwise, repeat the next frame until the repositioning is successful.

[0144] The above method only considers limited inter-frame information and inevitably generates errors. Therefore, after the above method, we will use local map tracking to improve the system accuracy. The local map consists of key frames that have a co-visibility relationship with the current frame and the point and line features observed by these key frames. Similar to inter-frame tracking, local map tracking first performs reprojection matching on the point and line features in the environmental map, and then optimizes the pose by bundle adjustment, updates the observed degree of the map points and map lines of the current frame, and determines whether the tracking is successful based on the number of map points and map lines tracked in the current frame.

[0145] 2. Key Frame Selection Strategy

[0146] Considering the robustness and real-time performance of the system, we select some representative images as key frames. If the key frame selection is too loose, redundant information will be introduced, resulting in an excessive computational load for backend optimization. On the contrary, if it is too strict, tracking failure may occur due to difficult feature matching. Due to the weak texture characteristics of thermal infrared images, we will insert key frames more frequently than some SLAM systems. For a frame of thermal infrared image, the key frame with the most common observations with it is used as the reference key frame, and it is selected as a key frame in any of the following situations:

[0147] (1) 15 frames have passed since the last global repositioning;

[0148] (2) 15 frames have passed since the insertion of the last key frame or the local map construction thread is idle;

[0149] (3) The number of map points and map lines tracked in the current frame is less than 90% of the number of map points and map lines tracked in the reference key frame;

[0150] (4) There is a certain degree of pose transformation from the last key frame;

[0151] (5) The current frame has tracked at least 40 feature points and 15 spatial lines.

[0152] Note: The above specific values are selected according to specific needs.

[0153] In the local mapping module:

[0154] As Figure 8 shown, after determining that the current frame is a key frame, the connection relationship between key frames is updated through the co-visibility relationship of map points and map lines, and poor-quality map points and map lines are removed according to the observation situation. We also use the current key frame and its adjacent co-visible key frames to generate new map points and map lines through triangulation to ensure more stable tracking. Finally, duplicate map points and map lines are checked and fused. When the number of key frames in the local map is greater than 3, local bundle adjustment is performed according to the equation to adjust the camera pose in the local map. Map points and map lines

[0155]

[0156] where and represent the set of matching pairs of points and lines in the local map. After optimization and adjustment, it is judged whether more than 90% (the specific percentage can be selected according to specific needs) of the map points and map lines tracked by the key frame can be tracked by other key frames, and redundant key frames are removed accordingly. Then the current frame is added to the loop closure detection queue.

[0157] In the loop closure detection module:

[0158] Although the local mapping thread can reduce errors to a certain extent, to construct a globally consistent trajectory and map, cumulative errors need to be eliminated through loop closure detection. Loop closure detection uses the similarity of image appearance to determine whether a loop occurs. We retrained the bag-of-words model based on thermal infrared images through clustering using the descriptors of ORB (Oriented FAST and Rotated BRIEF) feature points and the LBD (LineBand Descriptor) line feature descriptor offline. In loop closure detection, we first calculate the similarity score of the bag-of-words vectors of the current key frame and each co-visible key. Here, we define the similarity as:

[0159]

[0160] p and l represent weight coefficients, s p (η a , η b ) represents the point feature similarity between images, s l (η a , ηb ) represents the line feature similarity between images. By finding the set of loop closure candidate frames among all key frames, it is determined whether the loop closure is successful. If successful, the pose, map points, and lines are adjusted through the solved similarity transformation, optimized using the essential graph, and finally global bundle adjustment is performed to achieve the optimum.

[0161] In this embodiment, for the epipolar constraint of point features in a dynamic environment:

[0162] Moving objects in a dynamic environment do not satisfy the epipolar constraint. Based on this, we judge the dynamic situation of key points in the prior target area in the thermal infrared image. If there are more than a certain number of dynamic key points within the object area, the object is considered dynamic.

[0163] First, the detection and segmentation results of the target are obtained through the first three steps of the semantic part (target detection, target tracking, and semantic segmentation).

[0164] Secondly, the key point matching results of consecutive frames are obtained using coarse-to-fine optical flow.

[0165] Thirdly, the fundamental matrix F is solved using the RANSAC algorithm with the most points in the non-target area, avoiding the interference of moving objects. Then the epipolar lines are calculated using the fundamental matrix F, and finally, the key points with a distance greater than the threshold from the epipolar lines are determined to be moving.

[0166] Let x 1 , x 2 represent the matching key point pairs of consecutive frames, and their homogeneous coordinate forms are: X 1 =(μ 1 , v 1 , 1) T , X 2 =(μ 2 , ν 2 , 1) T , where μ i , ν i represent the coordinates of the key point in the corresponding image frame. From the epipolar constraint, we know that:

[0167]

[0168] where F is the fundamental matrix. Then the epipolar line L 2 can be calculated by the following formula:

[0169] L 2 =[a, b, c] T =FX 1

[0170] where a, b, c are the representations of the line vector. Then the distance e from the matching key point to the epipolar line:

[0171]

[0172] As Figure 3 shown, if point P is a static point, then the distance e between point x 2 and the polar line L 2 is e < ε; conversely, if point P moves to point P', then the distance e between the matching point x' and the polar line L 2 is e > ε. By deleting the feature points of the dynamic target, the outliers affecting the pose estimation can be eliminated.

[0173] In this embodiment, the geometric representation of lines in the thermal infrared image:

[0174] The extraction and matching of point features require an image with rich texture information. To overcome the weak texture and low contrast of the thermal infrared image, we simultaneously use line features that can provide additional structural information. A 3D line has 4 degrees of freedom, and its representation methods include: the connection of two endpoints or the intersection of two planes (eight parameters), Plücker coordinates (six parameters), and orthonormal representation (four parameters). Considering the viewpoint change and line segment occlusion during the SLAM operation, it is difficult to extract the accurate endpoints of the line from the image. In this embodiment, we regard the 3D line in space as infinitely long, use Plücker coordinates for the transformation and projection of the 3D line, and use orthonormal representation for the backend optimization.

[0175] This embodiment uses the minimum four optimization parameters to update the orthonormal representation of the 3D line, where the vector is used to update U: For That is:

[0176]

[0177] Among them, is the SO(3) matrix, representing a 3D rotation around the x-axis by angle, Similarly. And the scalar θ is used to update W: W ← WR(θ).

[0178] Similarly, the orthonormal representation can be easily converted to Plücker coordinates:

[0179]

[0180] where u i represents the i-th column of the matrix U.

[0181] In this embodiment, the joint bundle adjustment of points and lines:

[0182] The system of this embodiment executes the beam adjustment method for graph optimization. Different from the SLAM method that only uses point features, the method of this embodiment uses both point and line features and applies them to thermal infrared images, which results in different reprojection errors and different Jacobian matrices.

[0183] We designed a joint beam adjustment using the nonlinear optimization algorithm library g2o to optimize the camera pose, 3D map points, and 3D map lines to minimize the reprojection error. As Figure 3 shown, we define the projected points and lines projected onto the pixel plane as p′ and l′ = [l 1 , l 2 , l 3 T , and the matching points and line segments are p and l respectively. Then the reprojection errors of the point and line features are:

[0184] e p = p - p′

[0185]

[0186] where p 1 = [u 1 , v 1 , 1] and p 2 = [u 2 , v 2 , 1] are the two endpoints of the matching line segment l. Let f x and f y be the internal parameter matrices of the camera, and the spatial point P c = [X c , Y c , Z c T , and ζ is the camera pose represented by Lie algebra: exp(ξ ∧ ) = [R | t]. Then the Jacobian matrices Jp ζ and Jl ζ of the point and line reprojection errors with respect to ζ are:

[0187]

[0188]

[0189] Let e 1 = p 1 T l′ and e 2 = p 2 T l′, then we have:

[0190]

[0191] ​​For the projection of a straight line, there is where then:

[0192]

[0193]

[0194] Similarly, the Jacobian matrix for map points and map lines is:

[0195]

[0196]

[0197] where:

[0198]

[0199]

[0200] The bundle adjustment of points and lines is crucial for the operation of SLAM.

[0201] As a specific application:

[0202] 1. Experimental data

[0203] To evaluate the performance of MonoThermal - SLAM, we collected data in a real - world dynamic environment to verify the accuracy and robustness of our system.

[0204] The main sensor devices in this embodiment include a 3D LiDAR, a thermal infrared camera, and a high - precision RTK - GNSS receiver, and their specifications are shown in Table 1:

[0205] Table 1

[0206]

[0207] Using a non - shutter thermal infrared camera and scene - based non - uniformity correction, this ensures that the output will not be frequently frozen, and at the same time, a calibration board was used to calibrate the thermal imager.

[0208] As Figure 8 shown, in this embodiment, an existing calibration board was used to calibrate the thermal imager. This calibration board is a low - emissivity aviation aluminum plate with a square checkerboard pattern coated with a high - emissivity black coating.

[0209] We collected a total of 16 scene sequences, including 3 indoor sequences and 6 outdoor sequences on a small scale. The 7 large-scale driving sequences were collected by fixing the experimental equipment on a vehicle. This includes many challenging scenarios, including moving dynamic objects (pedestrians, vehicles), environmental illumination changes (low illumination, darkness), weather changes (light rain), environmental visibility changes (slight smoke), as Figure 9 shown.

[0210] Since it is difficult to obtain very accurate ground truth trajectory values in these challenging environments, in the indoor and outdoor sequences, we used the trajectory obtained by the well-known LeGO-LOAM algorithm as a reference for comparison. The reason for choosing the LeGO-LOAM algorithm is that it can not only obtain extremely accurate pose estimation, but also be more robust in dynamic environments. In the driving sequences, we used the measured trajectory of the high-precision real-time kinematic global satellite navigation system receiver in the fixed state as the ground truth for comparison.

[0211] Due to the lack of availability of SLAM code specifically designed for monocular thermal imagers, we compared the proposed method with SVO2.0 (monocular mode), ORB-SLAM3 (monocular mode), and DSO, which are the most advanced SLAM methods currently designed for monocular cameras. The systems in this embodiment all run on the Ubuntu 20.04 platform, and all tests were performed using a computer equipped with an Intel i7-9700 CPU and an NVIDIA GeForce RTX 2080Ti GPU.

[0212] 2. Real-time denoising performance

[0213] Although random noise usually follows a Poisson distribution, we still use the additive white Gaussian noise (AWGN) model to simulate it. The reason is that it is easy to convert a random variable with a Poisson distribution into a random variable with an approximate standard Gaussian distribution by applying a variance-stabilizing transformation, and this assumption is simple and reasonable, making the denoising problem easier to solve.

[0214] We collected 5000 thermal infrared images of different scenes. The training dataset contains input-output pairs where, o i is obtained by adding AWGN to image I i , and N i is the noise level map. The trained FFDNet has the ability to handle spatially variant noise. We applied the trained FFDNet to a real thermal infrared denoising open-source database for comparison. This database provides paired thermal infrared imaging data, where the low-quality thermal infrared images are the original thermal infrared data that has been preliminarily non-uniformly corrected and is severely affected by noise; the high-quality denoised thermal infrared images are processed by complex denoising methods.Figure 10 It shows the comparison between the original image and the denoised image. After denoising, the thermal infrared image not only has an improved signal-to-noise ratio but also clear textures. In addition, when processing a thermal infrared image with a resolution of 640x512 using FFDNet accelerated by GPU, it only takes about 10 milliseconds, while the traditional BM3D method takes several seconds or even dozens of seconds.

[0215] 3. Point and Line Feature Detection and Matching

[0216] To verify the performance improvement of the proposed method for point and line feature detection and matching, we tested the original thermal infrared images and denoised thermal infrared images of the dataset respectively. The two images were both set with the same point and line feature parameters. Figure 11 It shows the feature matching results after removing outliers. It can be found that due to the low contrast and low signal-to-noise ratio of the original thermal infrared image, very few features are successfully matched, and it can be inferred that it is very difficult to use the original image for tracking. While the denoised thermal infrared image is much more robust, and sufficient point and line features can be seen with good matching performance.

[0217] 4. Resistance to Dynamic Object Interference

[0218] To filter movable objects, we use the proposed method to detect dynamic objects in each frame. Figure 12 It shows an example of object detection. It can be seen that due to the lack of continuous frame information in traditional single-frame object detection, missed detections of objects are very common, while our method has greatly improved the detection rate of dynamic objects. At the same time, we verified the screening of dynamic objects in the system of this embodiment through the epipolar constraint through experiments. Figure 13 It shows an example that affects the accuracy of the system of this embodiment in a driving scenario with dynamic objects. Through the epipolar constraint, we screened out static point features and static line features for tracking. Through our method, resistance to dynamic object interference is achieved, and the robustness of the SLAM system in a dynamic environment is improved.

[0219] 5. Localization Accuracy Test

[0220] 5.1 Map Initialization

[0221] To measure the effect of the proposed initialization method, we conducted experiments to compare the proposed method with classical initialization. The traditional initialization method uses the method proposed in ORB-SLAM2, selects the fundamental matrix or the homography matrix for pose estimation, and reconstructs the map. Note that we tested the initialization using 800 point features and 100 line features per frame, which is basically the proposed number of features. The experimental results of different initialization tests are as Figure 14As shown, the results indicate that the initialization method proposed in this embodiment is superior to the traditional initialization method. In addition, the initialization method we proposed generates a structural line map. The traditional method only relies on feature points, while the weak-texture thermal infrared images lack sufficient and repeatable feature points, making it difficult to initialize the map after multiple runs. Figure 15 Examples of localization and mapping are also shown, which indicates that the initial map generated by the proposed initialization is faster than the traditional initialization. Especially for thermal scenes with rich structural information, this method is highly robust.

[0222] 5.2 Indoor and Outdoor Experiments

[0223] We tested the localization accuracy of several SLAM systems on sequences Indoor1 - 3 and Outdoor1 - 6. The proposed system extracts 1000 point features and 100 line features for each thermal infrared image. The parameters of the ORB - SLAM2 system are the same as the point feature parameters of the proposed system, and SVO2.0 and DSO use the officially recommended parameters. It should be noted that according to our experiments, the original thermal infrared images without denoising are frequently tracked unsuccessfully when applied to the above visual SLAM systems. Therefore, we uniformly use the thermal infrared images after real - time full - noise removal as the input. Since the scale is unavailable due to the tracking trajectory of the monocular camera, we align the trajectory with the ground truth through similarity transformation and calculate the root - mean - square error (RMSE) of the relative position error. The localization accuracy is evaluated according to the difference between the SLAM trajectory and the ground truth. In addition, we also analyzed the comparison between the tracking distance and the ground truth distance to show the robustness of the system.

[0224] Table 2

[0225]

[0226]

[0227] Comparison of the localization accuracy and tracking path length of MonoThermal SLAM, ORB - SLAM3, SVO2.0, and DSO on indoor and outdoor sequences. Systems with a tracking distance less than one - third of the ground truth distance are considered tracking failures and are indicated by "-" in the table.

[0228] Table 3

[0229]

[0230] Tables 2 and 3 show the positioning accuracy, tracking path length, and characteristics of each SLAM system on the Indoor and Outdoor sequences. Among them, systems with a tracking distance less than one-third of the ground truth distance are considered tracking failures and are indicated by "-" in the table. It can be seen that the proposed system has significant advantages compared with other methods, and the tracking trajectory is as Figure 16 shown. Due to the weak texture characteristics of thermal infrared images, ORB-SLAM2, which only uses point features, encounters challenges in feature tracking. Especially when there are dynamic target interferences in the sequence or there are not enough robust feature points in the scene, ORB-SLAM2 is prone to tracking loss during the experiment. Moreover, the path distance tracked by ORB-SLAM2 is significantly less than our method. An important reason is the frequent initialization failures. Based on the direct method, DSO improves the robustness of the system when facing moving targets in weak texture scenes by minimizing the photometric error. However, the sudden scene photometric change in the Outdoor1 sequence breaks the brightness constancy assumption, resulting in incorrect data association and tracking failure. Figure 17 Describes the tracking trajectory of the Outdoor1 sequence and the mapping effect of MonoThermal-SLAM. Based on the hybrid method, SVO2.0 occupies low computing resources and shows the advantages of feature point methods and direct methods in individual scenes. However, it has the shortest tracking distance in most sequences. Due to the interference of fast motion and dynamic scenes, it is prone to tracking loss.

[0231] In summary, due to the effective improvement of thermal infrared image quality by the real-time full noise removal link, the application boundary of the visual SLAM system in thermal infrared images is broadened. Benefiting from the high robustness of point and line features, and the epipolar constraint avoiding incorrect data association, MonoThermal-SLAM shows absolute advantages, especially in dynamic indoor and outdoor environments with unparalleled robustness. Our system has achieved success in all sequences and achieved a high positioning accuracy with an RPE below 0.1m, while ORB-SLAM2, SVO2.0, and DSO are unable to completely track all sequences.

[0232] 5.3 Driving Experiments and Loop Closure

[0233] By comparing with RTK-GNSS, the positioning accuracy and robustness of the proposed system in this embodiment are further evaluated in large-scale driving scenarios. Different from the RPE used in the outdoor and outdoor sequence evaluations, we used the RMSE of the absolute trajectory error (ATE) metric to compare the performance of each SLAM system, as shown in Tables 4 and 5:

[0234] Table 4

[0235]

[0236]

[0237] Comparison of the positioning accuracy and tracking path length of MonoThermal SLAM, ORB-SLAM3, SVO2.0, and DSO on the Driving sequence. Systems with a tracking distance less than one-third of the ground truth distance are considered tracking failures and are indicated by "-" in the table.

[0238] Table 5

[0239]

[0240] Figure 18 Shows the tracking trajectories of multiple driving sequences. Among them, the Driving1 sequence ( Figure 18 (a)) is a challenging dataset with a path length exceeding 500m and many dynamic targets. The ORB-SLAM2 method fails to track shortly after initialization, and SVO2.0 loses feature tracking when there is a large turn in the sequence and dynamic targets pass by. DSO performs better and shows a similar RMSE to the proposed system, however, it still has a relatively large error.

[0241] Figure 18 (b-f) It can be seen that the method of this embodiment has much higher positioning accuracy than other methods and the highest degree of coincidence with the ground truth.

[0242] In contrast, the system of this embodiment still runs smoothly in this sequence, provides more robust attitude estimation, and obtains the longest tracking distance and the highest positioning accuracy. Similarly, in the remaining sequence tests, the trajectory of the system of this embodiment basically coincides with the RTK-GNSS trajectory, showing the best performance.

[0243] In addition, another experiment was conducted to evaluate the accuracy of loop closures. We recorded the Driving5 loop closure sequence in a campus environment, and its path length exceeded 470m. Due to the lack of loop closures, both DSO and SVO2.0 exhibited significant scale drifts. The performance of the ORB-SLAM2 method was the closest to ours, but it was difficult for the ORB-SLAM2 method to detect the correct loop closures in the thermal scenario, and it was more difficult to initialize the map. We successfully detected the loop closures (Figure 19(a)) by using the retrained bag-of-words for thermal infrared images, and based on this, scale correction and global BA were performed. Our trajectory map was similar to the bird's-eye view of Google Maps (Figure 19(a)). Figure 19(b) depicts the position estimates of each SLAM system in the Driving5 sequence and the comparison with the ground truth. Compared with other systems, the trajectory estimated by MonoThermal-SLAM was more consistent with the true values of high-precision devices. Figure 19(c) shows the comparison of the tracking and position estimation performance of MonoThermal SLAM, DSO, SVO2.0, ORB-SLAM3, and RTK-GNSS on the Driving5 sequence.

[0244] In this embodiment, the influence of the line feature selection strategy:

[0245] The weak texture of thermal infrared images is not conducive to the extraction and matching of point features, but it becomes an advantage for line segments, which avoids detecting a large number of complex and invalid line segments like visual images. Our specific line feature selection strategy for thermal infrared image sequences contributed to the success of the system in this embodiment. First, real-time full noise removal of the images significantly improved the detection rate of line segments. Second, we utilized geometric constraints to enhance the robustness of line features during the line feature extraction and matching stages. And descriptor matching was performed by adjusting strategies such as the search range. When the number of matches was small, we increased the threshold of the parameters in the matching process to find more matching pairs, which improved both efficiency and reduced the false matching rate. Finally, for different scenarios, we dynamically adjusted the proportion of line features in localization and mapping, which was crucial for robust localization in certain scenarios.

[0246] Regarding the structural map:

[0247] Considering the instability of line features, we usually select line features more strictly to ensure the accuracy of tracking and positioning. Therefore, the map sometimes cannot well reflect the characteristics of the scene. However, buildings in the urban environment have rich straight-line structures, and using line features as the underlying features to express their geometric structures is more effective. Thanks to the accurate pose estimation of the system in this embodiment, we use Line3D (an existing method) with the pose and image of the camera as inputs, and output a denser scene geometric model as the 3D reconstruction result to better represent the characteristics of the scene, which is beneficial to the establishment of the semantic map and the simplified representation of the structural scene.

[0248] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A monocular thermal imaging simultaneous localization and mapping system, characterized in that: It includes a collector and a processor; the processor includes a denoising module, a static feature extraction module, an initialization module, a feature tracking module, a local mapping module, and a loop detection module; where: The collector collects thermal infrared video sequence images that are not interrupted by non-uniformity correction; The denoising module performs real-time denoising on the collected thermal infrared video sequence images; The static feature extraction module extracts the key points and descriptors of point features and the key line segments and descriptors of line features in the image respectively; uses an object detection neural network for object detection, adopts a moving object tracking method to improve accuracy, uses an instance segmentation method to reduce bounding boxes, realizes pixel-level segmentation of moving objects, and then uses epipolar geometry constraints to screen out static features for subsequent tracking; The initialization module performs the following processing: recover the poses of two adjacent frames based on the static point features and static line features extracted in the static feature extraction module; The feature tracking module performs the following processing: map point and map line tracking; key frame selection; In map point and map line tracking, data association is performed using inter-frame tracking and local map tracking in sequence to update the pose of the current frame; In key frame selection, for a thermal infrared image, the key frame with the most common observations is used as the reference key frame; The local mapping module updates the connection relationship between key frames through the co-visibility relationship of map points and map lines, and eliminates poor-quality map points and map lines according to the observation situation; The loop detection module uses image appearance-based similarity to determine whether a loop occurs, and uses the descriptors of point features and the descriptors of line features to offline train a bag-of-words model based on thermal infrared images through clustering.

2. The system according to claim 1, characterized in that: In real-time scene denoising: First, perform scene-based non-uniformity correction on the thermal infrared image; at the same time, when information is lost during the compression conversion of the thermal infrared image, use the adaptive histogram equalization method to retain the original information of the image to the greatest extent; Then remove the random noise of the thermal infrared image.

3. The system according to claim 2, characterized in that: For random noise, use FFDNet based on a convolutional neural network to denoise the thermal infrared image; Reshape the original image into multiple downsampled sub-images, and connect them with an adjustable noise level map and input them into the convolutional neural network to perform convolution, and finally generate a denoised image of the same size as the original input image.

4. The system according to claim 1, characterized in that: When performing line segment feature extraction, filter out extremely short invalid line features; connect long line segments that are separated; if the lengths and distances of two line segments are very close and the direction difference is small enough, merge such line segments; use descriptors to describe the new line segments, and add geometric constraints to remove possible abnormal matches. The geometric constraints include: (1) The length difference and angle difference of the matching line segment pair are less than a certain value; (2) The distance between the matching line segment pair is less than a certain value; (3) The descriptor distance of the matching line segment pair is less than a certain value; In the moving object tracking method, the target state detected by the target detection neural network is defined as: Among them, x and y represent the center coordinates of the detected target, and a and p represent the area and aspect ratio of the bounding box. It represents the center coordinates and area of the predicted target in the next frame; during the tracking process, the motion prediction and data association of the target are carried out through the Kalman filter and the Hungarian algorithm, and the intersection-over-union distance of the detection boxes is used to calculate the cost matrix.

5. The system according to claim 1, characterized in that: In the epipolar geometry constraint, judge the dynamic situation of the key points in the prior target area in the thermal infrared image; if there are more than a certain number of dynamic key points in the object area, the object is considered dynamic: First, obtain the detection and segmentation results of the target; second, obtain the key point matching results of consecutive frames by using coarse-to-fine optical flow; third, use the random sample consensus algorithm with the most points in the non-target area to solve the fundamental matrix, then calculate the epipolar line with the fundamental matrix, and finally determine that the key points with a distance greater than the threshold from the epipolar line are moving.

6. The system according to claim 1, characterized in that: The initialization module processes as follows: For a pair of parallel 3D lines, their 2D projections l 1 , l 2 intersect at the vanishing point l 1 × l 2 , whose normalized direction Similarly, its 2D projection m of the second frame image captured by the collector 1 , m 2 vanishing point normalization direction In addition, assume that the vanishing point normalization directions of the 2D projections of another pair of parallel 3D lines in the first frame image and the second frame image are v p3 and v p4 respectively. Then, for these two pairs of parallel 3D lines, the rotation matrix R between the two images ideally satisfies the following formula: where g 1 and g 2 are constants; Based on the above formula, solve for the rotation matrix R when ‖g 1 ‖ 2 +‖g 2 ‖ 2 is minimized; Consider two feature two-dimensional points p 1 and p 2 detected by a detector using point features in the first image. After feature matching, it is determined that they correspond to points q 1 and q 2 in the second image. The actual translation vector t is defined as: Simultaneously and in parallel use the initialization method in ORB-SLAM2, and use the random sample consensus method to sample point pairs and line pairs in the initialization method to select the best result.

7. The system according to claim 1, characterized in that: In map point and map line tracking, first, assuming the system moves at a constant speed, pose estimation is performed through the projection matching relationship between adjacent frames; the map points and map lines are projected onto the current frame, and point features and line features that meet the requirements are searched within a given range. For the matching of point features, in addition to requiring the shortest descriptor distance, rotational consistency also needs to be satisfied. If the number of effectively matched features after the above search is still less than the given threshold, the search area is expanded until the requirements are met; the above method obtains the initial value of the current pose. Next, the pose of the current frame is optimized using the 3D-2D projection relationship, that is, the reprojection error of 3D points and 3D lines is minimized through bundle adjustment; for the inter-frame pose estimation of the constant velocity model, the rotation matrix R * and the translation vector t * are used as state variables to be optimized to construct a graph optimization model, and the Levenberg-Marquardt method is used to iteratively solve by minimizing the following cost function: where ρ p and ρ l are Huber robust cost functions, and the Huber function is introduced to reduce the outliers of the cost function; are the minimum reprojection errors of points and lines, respectively; ∑ p and ∑ l represent the observation covariance matrices of points and lines; Represents a set of matching pairs between consecutive image frames in a video sequence. Considering the computational complexity, the Jacobian matrix is directly used for solution; after performing pose optimization, outliers and outlier lines in the map are removed; if the number of successfully matched map points and map lines is greater than a certain value, it is considered successful; Adopt the matching strategy between images; convert the feature descriptors into bag-of-words vectors to accelerate the matching of the current frame and the reference key frame, and the reference key frame is the key frame with the highest co-visibility degree with the current frame; if the number of matching features still does not meet the standard, then select to use the nearest neighbor matching algorithm, and use the random sample consensus method to calculate the homography matrix between images to obtain enough correct matching features; finally project the map points and map lines to the current frame, and use the pose of the previous frame as the initial value to optimize the pose according to the equation; If the above methods all fail to track and lead to positioning failure, repositioning is required; first convert the current frame into a bag-of-words, find the candidate key frame group similar to the current frame in the key frame database, and select the candidate key frames that meet the requirements from them; once the matching requirements are met, estimate the camera pose of the current frame by solving the PnP problem and optimize the pose according to the equation; if the number of inliers after optimization is too small, project the un-matched map points and map lines in the key frame to the current frame by projection to generate new matching relationships; according to the results of the projection matching, optimize the pose again; as long as one candidate key frame is successfully repositioned, the remaining candidate key frames are no longer considered, otherwise, repeat the next frame until repositioning is successful; In key frame selection, it is selected as a key frame in any of the following situations: (1) Some frames have passed since the last global repositioning; (2) Some frames have passed since the insertion of the last key frame or the local map construction thread is idle; (3) The number of map points and map lines tracked by the current frame is less than a certain proportion of the number of map points and map lines tracked by the reference key frame; (4) There is a certain degree of pose transformation from the last key frame; (5) The current frame tracks at least a certain number of feature points and spatial lines.

8. The system according to claim 7, characterized in that: The local construction module uses the current key frame and its adjacent co-visible key frames to generate new map points and map lines through triangulation to ensure more stable tracking; finally, it checks and fuses duplicate map points and map lines. When the number of key frames in the local map is greater than a certain number, local bundle adjustment is performed according to the following equation to adjust the camera pose in the local map Map point and map line Among them, and represent the set of matching pairs of points and lines in the local map; After optimization and adjustment, determine whether map points and map lines above a certain percentage of key-frame tracking can be tracked by other key frames, and accordingly remove redundant key frames, and then add the current frame to the loop closure detection queue.

9. The system according to claim 1, wherein: In the loop closure detection module, first calculate the similarity score between the word bag vectors of the current key frame and each co-visible key frame, and define the similarity as: where p and l respectively represent weight coefficients, s p (η a , η b ) represents the point feature similarity between images, s l (η a , η b ) represents the line feature similarity between images; by finding the set of loop candidate frames among all key frames, it is judged whether the loop is successfully closed; if successful, the pose, map points and lines are adjusted through the solved similarity transformation, and finally the global bundle adjustment method is carried out to achieve the optimum.

10. A monocular thermal imaging simultaneous localization and mapping method, wherein: Applicable to the system described in any one of claims 1-9, including: collecting thermal infrared images, image denoising, feature extraction, initialization, feature tracking, local mapping, and loop closure detection.