All-time infrared inertial SLAM method and device for dynamic environment
Through thermal infrared image sequence feature extraction and visual inertial fusion, combined with BA optimization processing of SLAM system in dynamic environments, the problem of inaccurate positioning in dynamic environments is solved, and accurate and robust positioning is achieved throughout the day.
Patent Information
- Application Number
- CN202510419313.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing SLAM systems have problems with light changes, visual occlusions and dynamic objects in dynamic environments, and the compatibility and information utilization of thermal infrared cameras are insufficient, and deep learning methods are very limited.
Thermal infrared image sequence is used for feature extraction and matching, combined with visual inertial fusion and BA optimization, and robust optimization is performed through visual reprojection residual, IMU pre-integrated residual and marginalized residual, dynamic features are eliminated, and global closed-loop optimization is performed.
It realizes accurate and robust positioning of drones throughout the day in dynamic environments, improves the accuracy of SLAM map positioning and the robustness of the system, and eliminates cumulative drift errors.
Smart Images

Figure CN120467322A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and in particular relates to an all-day infrared inertial SLAM method and device for dynamic environments. Background Art
[0002] Existing visible light vision-based simultaneous localization and mapping (SLAM) frameworks are mature in stable environments. However, certain environments can experience extreme light distribution variations, dynamic lighting changes, or visual obstructions such as dust, fog, and smoke. This visual degradation invariably reduces the reliability of SLAM state estimation solutions. Compared to visible light cameras, thermal infrared cameras offer significant advantages in visually degraded scenarios due to their all-day perception capabilities.
[0003] Current thermal-inertial odometry (TIO) solutions are essentially improvements upon conventional visual-inertial odometry (VIO) and are used to process image feature information for SLAM state estimation solutions. However, using thermal infrared cameras with existing VIO frameworks presents compatibility issues. For example, image data captured by thermal infrared cameras typically has low resolution and contrast. Because thermal radiation from the surrounding area cannot be distinguished, much visually observable texture information, such as color and stripes, is lost in thermal images. Furthermore, thermal infrared cameras require non-uniformity correction or flat-field correction during operation to eliminate imaging artifacts between consecutive frames caused by data interruptions resulting from accumulated non-zero mean noise. These issues have led most studies to integrate thermal infrared camera information as a supplement to other sensors.
[0004] Furthermore, most visual SLAM systems operate based on the assumption of static scenes. However, in the real world, SLAM systems operate in the presence of a wide variety of dynamic objects. Many visual SLAM methods still face potential risks when interacting with real-world environments containing these dynamic objects. Feature points located on dynamic objects can lead to mismatches, and the presence of dynamic objects can cause the algorithm to generate erroneous data correlations, reducing the accuracy of the SLAM system's pose estimation. Furthermore, in practical applications, temporary static objects, which appear static when observed but move when unobserved, can cause false positives during loop closure detection, leading to critical failures. Related research has addressed this issue by using deep learning-assisted methods to detect dynamic object regions. However, these methods are limited in that they only work for predefined objects. Furthermore, some researchers have incorporated the dynamic characteristics of objects into their model optimization frameworks. However, geometry-based methods require precise camera poses and can therefore only handle a limited range of dynamic objects. Summary of the Invention
[0005] In response to the above-mentioned problems existing in the prior art, the embodiments of the present invention provide an all-day infrared inertial SLAM method and device for dynamic environments; this method can achieve accurate and robust positioning of drones all day based on thermal infrared images, thereby improving the accuracy of SLAM map positioning.
[0006] According to a first aspect of an embodiment of the present invention, a full-time infrared inertial SLAM method for dynamic environments is provided, the method comprising: determining a motion trajectory of a target UAV based on a thermal infrared image sequence of a target scene; wherein the motion trajectory comprises visual pose information corresponding to the target UAV at different moments; based on any current key frame in a key frame set corresponding to the thermal infrared image sequence: performing pre-integration based on IMU data of the current key frame and a previous key frame adjacent to the current key frame to generate a pose estimate of the target UAV; selecting visual pose information corresponding to the current moment from the motion trajectory; and performing visual-inertial fusion processing on the visual pose information and the pose estimate to obtain visual reprojection residuals and IMU pre-integration residuals; performing BA optimization processing on the key frame set according to a preset sliding window based on the visual reprojection residuals, the IMU pre-integration residuals, and the marginalization residuals, and outputting an optimized key frame set; performing global closed-loop optimization processing on the optimized key frame set to output a global SLAM map.
[0007] Optionally, according to the visual reprojection residuals, IMU pre-integration residuals, and marginalization residuals, the key frame set is subjected to BA optimization processing according to a preset sliding window, and an optimized key frame set is output; including: for any current key frame in the preset sliding window: using a regularization factor and a weight momentum factor to correct the visual reprojection residuals corresponding to the current key frame to obtain a corrected visual reprojection residual; based on the marginalization residuals corresponding to all current key frames in the preset sliding window, and the corrected visual reprojection residuals and IMU pre-integration residuals corresponding to each current key frame, a BA optimization model is constructed; when the BA optimization model tends to be minimum, the static features corresponding to each current key frame in the preset sliding window are obtained to generate an optimized key frame; based on the optimized key frame corresponding to each preset sliding window in the key frame set, an optimized key frame set is generated.
[0008] Optionally, the motion trajectory of the target UAV is determined based on the thermal infrared image sequence of the target scene; including: for any thermal infrared image in the thermal infrared image sequence of the target scene: performing feature extraction processing on the thermal infrared image based on a lightweight SuperPoint network to generate a feature point cloud and a descriptor corresponding to each feature point; based on the descriptors corresponding to the feature points, performing feature matching and tracking processing on all thermal infrared images in the thermal infrared image sequence to generate a tracking feature point cloud; performing visual SFM processing on the tracking feature point cloud to generate a three-dimensional feature point cloud; and determining the motion trajectory of the target UAV based on the three-dimensional feature point cloud.
[0009] Optionally, the method further includes: selecting a thermal infrared image having a number of tracking feature points greater than a preset threshold from the thermal infrared image sequence as a key frame to obtain a key frame set.
[0010] Optionally, the optimized key frame set is subjected to global closed-loop optimization processing to output a global SLAM map; including: based on tracking feature points, the optimized key frame set is divided into several key frame groups; for any key frame group in the several key frame groups: a key frame having a loop relationship with the current key frame is selected from the key frame group, and the loop relationship is clustered to obtain a cluster group; based on the cluster group corresponding to each of the key frame groups in the several key frame groups, several cluster groups are obtained; based on the loop relationship similarity corresponding to each of the loop relationships in the cluster group, the average similarity corresponding to the cluster group is determined; the cluster groups with the top two average similarities are selected from the several cluster groups as hypothetical clusters, and the current key frame is subjected to BA optimization processing using the two hypothetical clusters; the hypothetical cluster with the highest weight is selected from the BA optimization processing results to perform global pose optimization processing to generate a global SLAM map.
[0011] Optionally, the method comprises: selecting a key frame having a loop relationship with the current key frame from the key frame group, and clustering the loop relationship to obtain a cluster group; comprising: using the DBoW2 bag-of-words model to identify key frames similar to the current key frame from the key frame group; and taking the key frames whose similarity in the identification results is greater than a preset threshold as associated key frames to obtain at least one associated key frame; determining the Euclidean distance between the associated key frames and the current key frame based on the relative posture information between the two; establishing a loop relationship between the associated key frame having a Euclidean distance less than a preset threshold in the at least one associated key frame and the current key frame to obtain at least one loop relationship; and clustering the at least one loop relationship to obtain a cluster group.
[0012] Optionally, based on tracking feature points, the optimized key frame set is divided into several key frame groups; including: for any target key frame in the optimized key frame set: if the number of tracking feature points shared between a first key frame adjacent to the target key frame and the target key frame is not less than a preset threshold, the first key frame is added to the group corresponding to the target key frame; and the first key frame is used as the next target key frame, and if the number of tracking feature points shared between a second key frame adjacent to the first key frame and the first key frame is not less than a preset threshold, the second key frame is added to the group of the target key frame, until the number of tracking feature points shared between two adjacent key frames is less than the preset threshold, then the grouping of the target key frame is ended to generate a target key frame group; based on several target key frame groups in the optimized key frame set, several key frame groups are generated.
[0013] Optionally, feature extraction processing is performed on the thermal infrared image based on a lightweight SuperPoint network to generate a feature point cloud and a descriptor corresponding to each feature point; including: feature extraction processing is performed on the thermal infrared image based on a GhostNet network to generate extracted features; weighted processing is performed on the extracted features based on a long-distance attention mechanism to output a feature point cloud; description processing is performed on each feature point in the feature point cloud to generate a descriptor corresponding to the feature point.
[0014] According to the second aspect of an embodiment of the present invention, an all-day infrared inertial SLAM device for dynamic environments is also provided, including: a determination module for determining the motion trajectory of a target UAV based on a thermal infrared image sequence of a target scene; wherein the motion trajectory includes visual pose information corresponding to the target UAV at different times; a fusion processing module for: based on any current key frame in a key frame set corresponding to the thermal infrared image sequence: performing pre-integration based on the IMU data of the current key frame and the previous key frame adjacent to the current key frame to generate a pose estimate of the target UAV; selecting the visual pose information corresponding to the current moment from the motion trajectory; and performing visual-inertial fusion processing on the visual pose information and the pose estimate to obtain visual reprojection residuals and IMU pre-integration residuals; a BA optimization processing module for performing BA optimization processing on the key frame set according to a preset sliding window based on the visual reprojection residuals, IMU pre-integration residuals, and marginalization residuals, and outputting an optimized key frame set; a closed-loop optimization processing module for performing global closed-loop optimization processing on the optimized key frame set to output a global SLAM map.
[0015] According to a third aspect of an embodiment of the present invention, a computer-readable medium is further provided, on which a computer program is stored. When the program is executed by a processor, the method described in the first aspect is implemented.
[0016] An embodiment of the present invention provides an all-day infrared inertial SLAM method and device for dynamic environments. The method first determines the motion trajectory of a target UAV based on a thermal infrared image sequence of a target scene; wherein the motion trajectory includes visual pose information corresponding to the target UAV at different times; secondly, based on any current key frame in a key frame set corresponding to the thermal infrared image sequence: pre-integration is performed based on the IMU data of the current key frame and the previous key frame adjacent to the current key frame to generate a pose estimate of the target UAV; the visual pose information corresponding to the current moment is selected from the motion trajectory; and the visual pose information and the pose estimate are subjected to visual-inertial fusion processing to obtain visual reprojection residuals and IMU pre-integration residuals; then, according to the visual reprojection residuals, IMU pre-integration residuals, and marginalization residuals, the key frame set is subjected to BA optimization processing according to a preset sliding window, and an optimized key frame set is output; finally, the optimized key frame set is subjected to global closed-loop optimization processing to output a global SLAM map. This embodiment uses a thermal infrared camera as a sensor in the SLAM system, and optimizes the key frames in the collected thermal infrared image sequence based on visual inertial fusion technology and BA optimization technology; thereby, the features of dynamic objects that significantly deviate from the motion prior can be discarded when estimating the UAV's posture, thereby improving the accuracy of SLAM map positioning; then, by performing global closed-loop optimization processing on the optimized key frame set, closed-loop detection interference from temporary static objects can be rejected, thereby eliminating the global accumulated drift error, and providing more reliable support for the autonomous navigation and positioning of the UAV, thereby improving the robustness and accuracy of the SLAM system state estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Hereinafter, some specific embodiments of the present invention will be described in detail in an exemplary and non-limiting manner with reference to the accompanying drawings. The same reference numerals in the accompanying drawings indicate the same or similar components or parts. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. In the accompanying drawings:
[0018] Figure 1 A flowchart of an all-day infrared inertial SLAM method for dynamic environments provided by one embodiment of the present invention;
[0019] Figure 2 A schematic diagram of the GhostNet V2 network structure provided by one embodiment of the present invention;
[0020] Figure 3 A schematic diagram of a process for performing global closed-loop optimization on an optimized key frame set according to an embodiment of the present invention;
[0021] Figure 4 A schematic diagram of a preset sliding window robust BA optimization process provided by an embodiment of the present invention;
[0022] Figure 5 A dynamic environment thermal infrared inertial SLAM system framework with a neural network front end provided by one embodiment of the present invention;
[0023] Figure 6 A schematic diagram of the structure of an all-day infrared inertial SLAM device for dynamic environments provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the purposes, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0025] Visible light SLAM systems face the problem of poor robustness in dark or low-light environments, while thermal infrared cameras can distinguish objects by temperature differences, which is particularly useful in complex environments, such as in thermal fog, rainy days, or low-field conditions. Using thermal infrared cameras as sensors in SLAM systems is not only unaffected by changes in lighting, but also allows the system to obtain key information for positioning and map construction in low-light or even completely dark environments, and its performance is more stable. However, at the same time, infrared imaging has disadvantages such as poor planar texture, low contrast, and low signal-to-noise ratio. In order to address the difficulties and high time complexity of feature extraction from infrared images, it is necessary to reduce the amount of computation and the number of parameters while ensuring the accuracy of feature extraction.
[0026] Furthermore, visual SLAM typically assumes a static environment, and dynamic objects may be mistaken for part of the environment, leading to map drift or positioning errors. To address these issues, robust SLAM methods are needed to handle undefined dynamic objects that cannot be addressed by learning-based or vision-based methods alone. Real-world scenarios also contain temporary static objects that appear static when observed but move when unobserved. Therefore, ensuring that changes in dynamic objects do not affect the map quality and positioning accuracy of the static parts becomes a major challenge in maintaining map consistency.
[0027] like Figure 1 As shown in FIG, a flow chart of an all-day infrared inertial SLAM method for dynamic environments provided by an embodiment of the present invention. Figure 2 , which is a schematic diagram of the GhostNet V2 network structure provided by an embodiment of the present invention.
[0028] The all-day infrared inertial SLAM method for dynamic environments includes at least the following steps:
[0029] S101, determining a motion trajectory of a target UAV based on a thermal infrared image sequence of a target scene; wherein the motion trajectory includes visual pose information corresponding to the target UAV at different times;
[0030] S102: Based on any current keyframe in the keyframe set corresponding to the thermal infrared image sequence, pre-integrate the IMU data of the current keyframe and the previous keyframe adjacent to the current keyframe to generate a pose estimate of the target UAV; select the visual pose information corresponding to the current moment from the motion trajectory; and perform visual-inertial fusion processing on the visual pose information and the pose estimate to obtain a visual reprojection residual and an IMU pre-integration residual.
[0031] S103, performing BA optimization on the key frame set according to the preset sliding window based on the visual reprojection residual, the IMU pre-integration residual, and the marginalization residual, and outputting the optimized key frame set;
[0032] S104: Perform global closed-loop optimization on the optimized key frame set and output a global SLAM map.
[0033] In S101, based on preset rules or algorithm models, the visual pose information corresponding to the target UAV at different times is determined according to the thermal infrared image sequence of the target scene, and then the motion trajectory of the target UAV is generated based on the visual pose information corresponding to the different times.
[0034] Exemplary: For any thermal infrared image in the thermal infrared image sequence of the target scene: perform feature extraction processing on the thermal infrared image based on the lightweight SuperPoint network to generate a feature point cloud and a descriptor corresponding to each feature point; based on the descriptors corresponding to the feature points, perform feature matching and tracking processing on all thermal infrared images in the thermal infrared image sequence to generate a tracking feature point cloud; perform visual SFM processing on the tracking feature point cloud to generate a three-dimensional feature point cloud; based on the three-dimensional feature point cloud, determine the motion trajectory of the target UAV.
[0035] Furthermore, feature extraction processing is performed on the thermal infrared image based on the lightweight SuperPoint network to generate a feature point cloud and a descriptor corresponding to each feature point; including: feature extraction processing is performed on the thermal infrared image based on the GhostNet network to generate extracted features; weighted processing is performed on the extracted features based on the long-range attention mechanism to output a feature point cloud; description processing is performed on each feature point in the feature point cloud to generate a descriptor corresponding to the feature point.
[0036] For example, SuperPoint is an end-to-end feature point and descriptor extraction network that uses an encoder-decoder structure similar to a semantic segmentation network. Its workflow takes a complete image as input, extracts deep features through a shared encoder, and then outputs feature points and their descriptors through two decoders. The VGG-like encoder used in the original SuperPoint network, while simple in structure, results in high computational cost and parameter count due to its large number of layers and channels. To overcome the shortcomings of the VGG architecture in the original SuperPoint network, particularly the high number of parameters and computational complexity, GhostNetV2 was adopted to replace the original VGG encoding layers. GhostNet's overall structure is composed of GhostBottlenecks and other structures, which significantly reduce computational requirements while maintaining accuracy. GhostNetV2 is formed by introducing a long-range attention mechanism based on GhostNet.
[0037] like Figure 2 As shown in the figure, the GhostNet V2 network structure provided by an embodiment of the present invention. When GhostNetV2 processes an input image (such as 224×224×3), it first passes through a 3×3 convolution block containing convolution, normalization and activation functions, and then stacks multiple Ghost Bottlenecks to finally obtain a 7×7×160 feature layer. After that, the number of channels is adjusted through a 1×1 convolution block to obtain a 7×7×960 feature layer; then global average pooling is performed, and then a 1×1 convolution block is used to obtain a 1×1×1280 feature layer, and finally classification is performed through a fully connected layer. The advantage of this structural design is that GhostNetV2 introduces a long-distance attention mechanism based on GhostNet, which further improves the representation ability, thereby achieving better performance while reducing the number of parameters.
[0038] In S102 and S103, a key frame set corresponding to the thermal infrared image sequence is determined based on a preset rule or algorithm model. Exemplarily, a thermal infrared image with a number of tracking feature points greater than a preset threshold is selected from the thermal infrared image sequence as a key frame to obtain a key frame set.
[0039] For example, the pose estimate and visual pose information of the target drone at each moment are aligned visually and inertially, ensuring that the scale information from the pure vision matches the IMU measurements. This combined approach improves the overall accuracy and robustness of the SFM system and is of great significance in practical applications.
[0040] Exemplarily, according to the visual reprojection residual, IMU pre-integration residual, and marginalization residual, the key frame set is subjected to BA optimization processing according to a preset sliding window, and an optimized key frame set is output; including: for any current key frame in the preset sliding window: using a regularization factor and a weight momentum factor to correct the visual reprojection residual corresponding to the current key frame to obtain a corrected visual reprojection residual; based on the marginalization residual corresponding to all current key frames in the preset sliding window, and the corrected visual reprojection residual and IMU pre-integration residual corresponding to each current key frame, a BA optimization model is constructed; when the BA optimization model tends to be minimum, the static features corresponding to each current key frame in the preset sliding window are obtained to generate an optimized key frame; based on the optimized key frame corresponding to each preset sliding window in the key frame set, an optimized key frame set is generated.
[0041] The application of a sliding window robust BA optimization in SLAM systems aims to improve the accuracy of pose estimation while effectively handling outliers in dynamic scenes. The following describes the robust BA optimization model in detail.
[0042] ① Visual-inertial BA optimization model
[0043] In the visual-inertial state estimation based on the Visual-Inertial Navigation System (VINS), the Maximum A Posteriori (MAP) is achieved by minimizing the sum of the prior and Mahalanobis distance norms of all measurement residuals. The BA optimization model of visual-inertial is defined as follows:
[0044]
[0045] Among them, ρ H (.) is the Huber loss function; r p 、 and They represent marginalization residual, IMU pre-integration residual, and visual reprojection residual respectively; are IMU observation values and feature point observation values; H p represents the marginalized measurement estimate matrix, P represents the covariance of each item, and X represents the pose information of the current keyframe. As the proportion of outliers in dynamic scenes increases, the Huber loss function cannot completely reject the outlier residuals, and the system cannot function successfully. In other words, the Huber loss function is used to process these residuals to enhance the model's robustness to outliers.
[0046] ②Regularization factor
[0047] For simplicity, item, Items are omitted and represented as In order to robustly estimate the pose while rejecting abnormal features, a new residual term inspired by the BR (Black-Rangarajan) duality is constructed. :
[0048]
[0049] Among them, w j ∈[0, 1] represents each feature f j The corresponding weight, determine w j Feature f close to 1 j It is a static feature; is a constant parameter; Φ(w j ) is the weight w j The regularization factor is defined as follows
[0050] φ(w j )=1-w j Formula (3);
[0051] ③Weighted Momentum Factor
[0052] When motion becomes intense, IMU pre-integration becomes imprecise, leading to inaccurate pose estimates. In this case, the feature reprojection residuals for static objects become large; these features are then ignored in the bundle adjustment process through the regularization factor, resulting in inaccurate bundle adjustment results even if the previous weights are close to 1. An additional factor, the weight momentum factor, is constructed to insulate previously estimated feature weights from the effects of intense motion.
[0053] Because features are tracked continuously, each feature f j Use its previous weight Conducted n j Second optimization. In order to make the current weight tend to remain And with n j The increase in the weight momentum factor ψ(w j ) is designed as follows:
[0054]
[0055] The corrected visual reprojection residual can be obtained as follows:
[0056]
[0057] in, Represents a constant parameter used to adjust the impact of the momentum factor on BA.
[0058] Using (5) Instead of the Huber norm in the visual reprojection residual term in (1), the robust BA optimization model can be expressed as
[0059]
[0060] This embodiment introduces feature weights to handle visual reprojection residuals and uses pre-integrated IMU data to calculate IMU pre-integration residuals. Each feature is assigned a weight, and these weights are updated and optimized by introducing a weight momentum factor and a regularization factor. During the optimization process, the weight momentum factor uses the weights of previously tracked features, while the regularization factor is adjusted based on the weights of all features in the current preset sliding window. This strategy solves the problem through an alternating optimization approach. Since the pose estimate X of the current keyframe can be estimated using IMU pre-integration and the previously optimized state, feature weights are first optimized based on the estimated state. Therefore, features with large visual reprojection residuals start with smaller weights to reduce their impact on the overall optimization. The optimization steps are repeated until both the state and weights converge. During this process, the weights of outlier features are reduced, thereby flattening their losses. The regularization factor effectively filters out outliers by adaptively adjusting the weights, but these outliers are not completely ignored during the optimization process. Using weights and regularization factors inspired by the BR duality approach can reduce the impact of features with high reprojection error on pose estimation while maintaining state estimation performance. This approach enhances the robustness of pose estimation, making the optimization process more stable and accurate, thereby improving the system's performance in complex environments. Overall, this optimization strategy demonstrates good performance when handling dynamic scenes and interference, helping to improve the accuracy and reliability of visual-inertial odometry.
[0061] In S104 , a global closed-loop optimization process is performed on the optimized key frame set based on a preset rule or algorithm model, and a global SLAM map is output.
[0062] Performing global closed-loop optimization on the optimized keyframe set effectively groups closed-loop constraints, reducing the interference of temporary static objects on the global closed-loop optimization, thereby improving the robustness and accuracy of the SLAM system. This approach not only improves the optimization efficiency of the optimized keyframe set, but also better handles the interference of dynamic and static objects in complex environments, providing more reliable support for autonomous navigation and robot positioning.
[0063] This embodiment uses a thermal infrared camera as a sensor for the SLAM system, and extracts and matches feature points of the thermal infrared image by adopting a lightweight SuperPoint neural network with a GhostNetV2 encoding structure; thereby, the computational complexity and parameter quantity of the SLAM system can be reduced while ensuring the accuracy of image feature extraction; by adding a regularization factor and a weight momentum factor to the BA optimization model within a preset sliding window, and using the weights of previously tracked features in the weight momentum factor, and using the weights of all features in the current window in the regularization factor; thereby, the features of dynamic objects that significantly deviate from the motion prior can be discarded when estimating the UAV's pose, thereby improving the accuracy of UAV positioning; in the global closed-loop optimization, the loop constraints are grouped into multiple hypotheses; thereby, closed-loop detection interference from temporary static objects can be rejected, the global cumulative drift error can be eliminated, and a global SLAM map can be generated; thereby, more reliable support is provided for the autonomous navigation and positioning of the UAV, and the robustness and accuracy of the SLAM system state estimation are improved.
[0064] Loop closure detection in the SLAM framework can eliminate cumulative errors. However, false positive loop closure features caused by temporary static objects can cause global loop closure optimization to fail. In fact, features from temporary static objects and real static objects may exist in the same keyframe. To this end, a robust global loop closure optimization method is designed to group loop closure constraints into multiple hypotheses to reject loop closure detection interference from temporary static objects. Loops from the same feature are grouped, even if they come from different keyframes, and only one weight is used for each group, resulting in faster optimization.
[0065] like Figure 3 FIG. 1 is a flow chart of performing global closed-loop optimization on an optimized key frame set according to an embodiment of the present invention.
[0066] Performing global closed-loop optimization on the optimized keyframe set includes at least the following steps:
[0067] S301, dividing the optimized key frame set into several key frame groups based on the tracking feature points;
[0068] S302, for any key frame group among the plurality of key frame groups: selecting a key frame having a loop relationship with the current key frame from the key frame group, and clustering the loop relationship to obtain a cluster group;
[0069] S303: obtaining a plurality of cluster groups based on the cluster groups corresponding to each key frame group in the plurality of key frame groups; and determining an average similarity corresponding to the cluster groups based on the similarity corresponding to each loop relationship in the cluster groups;
[0070] S304: Select the top two clusters in terms of average similarity from the plurality of cluster groups as hypothetical clusters, and perform BA optimization on the current key frame using the two hypothetical clusters;
[0071] S205: Select the hypothesis cluster with the highest weight from the BA optimization processing results to perform global pose optimization processing to generate a global SLAM map.
[0072] In S301, exemplarily, based on tracking feature points, the optimized key frame set is divided into several key frame groups; including: for any target key frame in the optimized key frame set: if the number of tracking feature points shared between a first key frame adjacent to the target key frame and the target key frame is not less than a preset threshold, the first key frame is added to the group corresponding to the target key frame; and the first key frame is used as the next target key frame, and if the number of tracking feature points shared between a second key frame adjacent to the first key frame and the first key frame is not less than a preset threshold, the second key frame is added to the group of the target key frame, until the number of tracking feature points shared between two adjacent key frames is less than the preset threshold, then the grouping of the target key frame is ended to generate a target key frame group; based on several target key frame groups in the optimized key frame set, several key frame groups are generated.
[0073] In S302, exemplarily, a key frame having a loop relationship with the current key frame is selected from the key frame group, and the loop relationship is clustered to obtain a cluster group; including: using the DBoW2 bag-of-words model to identify key frames similar to the current key frame from the key frame group; and taking the key frames in the identification results whose similarity is greater than a preset threshold as associated key frames to obtain at least one associated key frame; based on the relative posture information between the associated key frames and the current key frame, determining the Euclidean distance between the two; establishing a loop relationship between the associated key frame having a Euclidean distance less than a preset threshold in the at least one associated key frame and the current key frame to obtain at least one loop relationship; clustering the at least one loop relationship to obtain a cluster group.
[0074] In S303, for any loop relationship in the cluster group, the similarity between the current key frame and the key frame in the loop relationship is obtained, and the similarity is used as the loop relationship similarity. The loop relationship similarities corresponding to each loop relationship in the cluster group are summed and averaged to obtain the average similarity of the cluster group.
[0075] In S304, since there are several key frame groups before the current key frame, when detecting the loop relationship of the current key frame, it is necessary to perform loop relationship detection on each of the several key frame groups, thereby obtaining several cluster groups. The cluster groups are sorted from largest to smallest according to their average similarity, and the cluster groups ranked first and second in the sorting are used as hypothetical clusters.
[0076] Before grouping loop closures, we must group adjacent keyframes that share the least number of tracking features. i The starting group is defined as
[0077]
[0078] Among them, α represents the minimum number of tracking features, Indicates that from C i to C k For the sake of simplicity, the key frames are grouped into Group (C i ) is denoted as G i Then perform multi-hypothesis clustering. For the current key frame, use DBoW2 bag of words recognition and key frame grouping G i Each key frame C in k Similar keyframe C m If there is no similar keyframe, skip C k After identifying k at most 3 different m, in C k By matching features between these key frames, the relative pose T can be obtained. If the features used for matching come from the same object, even if the matched C k and C m Different, the matching estimated poses will also be located close to each other. Therefore, by calculating the Euclidean distance between the loop poses, similar closed loops with smaller Euclidean distances can be clustered. Each cluster group can be called a hypothesis. In order to reduce the computational cost, the first two hypotheses are adopted to group the key frames into G i The two hypothetical clusters are denoted as and
[0079] The BR dual method is used to construct a BR optimization model for two hypothetical clusters. The BR optimization model is shown as follows:
[0080]
[0081] in, Represents two adjacent key frames C i and C i+1 The local posture between It is in a closed loop and Ck The relative positions between and P L denote the covariance of local pose and closed loop respectively; is a constant parameter, the regularization factor Φ of the loop closure i The definition is as follows
[0082]
[0083] in To ensure that the weight is not affected by the number of loop closures in the hypothesis cluster, the weight is divided by the cardinality of each hypothesis cluster. Then, the same optimization as (7) is performed on Equation (9). Finally, only the hypothesis clusters with higher weights are used for global pose optimization to eliminate the global cumulative drift error and form a global trajectory map.
[0084] This embodiment uses a thermal infrared camera as a sensor for the SLAM system, which is not affected by changes in illumination. In low-light or even completely dark environments, the system can still obtain key information for positioning and map construction, and its performance stability is higher. A SuperPoint neural network with a lightweight GhostNetV2 encoding structure is used for feature point extraction, which reduces the amount of calculation and parameters while ensuring the accuracy of image feature extraction. A robust VI-SLAM method is used to handle undefined dynamic objects that cannot be solved by learning-based or vision-based methods alone. A new bundle adjustment (BA) method is adopted, which uses the regularization factor of IMU pre-integration and considers the momentum factor of the previous state of each weight to cover the temporary inaccuracy of pre-integration. The weights of previously tracked features are used in the weight momentum factor, and the weights of all features in the current window are used in the regularization factor to simultaneously estimate the camera pose and discard the features of dynamic objects that deviate significantly from the motion prior. A robust global optimization method is used to group loop constraints into multiple hypotheses to reject closed-loop detection interference from temporary static objects, eliminate global accumulated drift errors, form a global trajectory map, ensure map updates and maintain global consistency. This approach not only improves the optimization efficiency of the optimized keyframe set, but also better handles the interference of dynamic and static objects in complex environments, providing more reliable support for autonomous navigation and robot positioning; thereby improving the robustness and accuracy of the SLAM system.
[0085] like Figure 4 , which is a schematic diagram of a preset sliding window robust BA optimization process provided by an embodiment of the present invention.
[0086] Each feature f is assigned a weight, which is used for the visual reprojection residual (i.e., visual residual); the IMU pre-integration residual (i.e., IMU residual) uses IMU pre-integration. The weight momentum factor (i.e., momentum factor) tracks the feature weight using a previously preset sliding window (i.e., previous window), while the regularization factor uses the weights of all features in the current preset sliding window (i.e., current window). Each weight is optimized using the regularization factor and momentum factor. An alternating optimization approach is used for the solution. Since the current state X can be estimated from the IMU pre-integration and previously optimized states, the weights are first optimized based on the estimated state. Therefore, features with large visual reprojection errors start with small weights. As the optimization steps are repeated until the states and weights converge, the weights of outlier features are reduced, flattening their losses. The regularization factor effectively filters out outliers by adaptively adjusting the weights, but outliers are not completely ignored during the optimization process. Based on the robust BA optimization model, the visual reprojection residual, IMU pre-integration residual, marginalization residual, momentum factor and regularization factor are optimized, thereby reducing the impact of features with high visual reprojection error relative to the estimated state while maintaining the state estimation performance.
[0087] The following describes in detail the all-day infrared inertial SLAM method for dynamic environments provided by this embodiment in conjunction with specific application scenarios.
[0088] The all-day infrared inertial SLAM method for dynamic environments includes at least the following steps:
[0089] S1, for any thermal infrared image in the thermal infrared image sequence of the target scene: perform feature extraction processing on the thermal infrared image based on the GhostNet network to generate extracted features; perform weighted processing on the extracted features based on the long-range attention mechanism to output a feature point cloud; perform description processing on each feature point in the feature point cloud to generate a descriptor corresponding to the feature point. Based on the descriptors corresponding to the feature points, perform feature matching and tracking processing on all thermal infrared images in the thermal infrared image sequence to generate a tracking feature point cloud; perform visual SFM processing on the tracking feature point cloud to generate a three-dimensional feature point cloud; based on the three-dimensional feature point cloud, determine the motion trajectory of the target UAV; wherein, the motion trajectory includes the visual pose information corresponding to the target UAV at different times;
[0090] S2, selecting thermal infrared images whose number of tracking feature points is greater than a preset threshold from the thermal infrared image sequence as key frames to obtain a key frame set;
[0091] S3, based on any current key frame in the key frame set corresponding to the thermal infrared image sequence: pre-integrate IMU data of the current key frame and the previous key frame adjacent to the current key frame to generate a pose estimate of the target UAV; select visual pose information corresponding to the current moment from the motion trajectory; and perform visual-inertial fusion processing on the visual pose information and the pose estimate to obtain a visual reprojection residual and an IMU pre-integration residual;
[0092] S4, for any current key frame in the preset sliding window: use the regularization factor and the weight momentum factor to correct the visual reprojection residual corresponding to the current key frame to obtain the corrected visual reprojection residual; based on the marginalization residuals corresponding to all current key frames in the preset sliding window, and the corrected visual reprojection residual and IMU pre-integration residual corresponding to each current key frame, build a BA optimization model; when the BA optimization model tends to the minimum, obtain the static features corresponding to each current key frame in the preset sliding window and generate an optimized key frame; based on the optimized key frame corresponding to each preset sliding window in the key frame set, generate an optimized key frame set;
[0093] S5, for any target key frame in the optimized key frame set: if the number of shared tracking feature points between a first key frame adjacent to the target key frame and the target key frame is not less than a preset threshold, then the first key frame is added to the group corresponding to the target key frame; and the first key frame is used as the next target key frame; if the number of shared tracking feature points between a second key frame adjacent to the first key frame and the first key frame is not less than a preset threshold, then the second key frame is added to the group of the target key frame, until the number of shared tracking feature points between two adjacent key frames is less than the preset threshold, then the grouping of the target key frames is terminated, and a target key frame group is generated;
[0094] S6, based on the target key frame groups in the optimized key frame set, generate several key frame groups. For any key frame group in the several key frame groups: use the DBoW2 bag-of-words model to identify key frames similar to the current key frame from the key frame group; and use the key frames with similarity greater than a preset threshold in the identification results as associated key frames to obtain at least one associated key frame; based on the relative posture information between the associated key frames and the current key frame, determine the Euclidean distance between the two; establish a loop relationship between the associated key frames with a Euclidean distance less than a preset threshold in the at least one associated key frame and the current key frame, to obtain at least one loop relationship; cluster the at least one loop relationship to obtain a cluster group;
[0095] S7, based on the cluster group corresponding to each of the key frame groups, obtain several cluster groups; based on the loop relationship similarity corresponding to each of the loop relationships in the cluster groups, determine the average similarity corresponding to the cluster groups; select the cluster groups with the top two average similarities from the several cluster groups as hypothetical clusters, and use the two hypothetical clusters to perform BA optimization processing on the current key frame; select the hypothetical cluster with the highest weight from the BA optimization processing results to perform global pose optimization processing to generate a global SLAM map.
[0096] This embodiment uses a thermal infrared camera as a sensor for the SLAM system. By designing a lightweight SuperPoint neural network with a GhostNetV2 encoding structure, feature points of the thermal infrared image are extracted and matched. This reduces the amount of computation and parameters while ensuring the accuracy of image feature extraction. This embodiment adds a regularization factor and a weight momentum factor to the sliding window optimization model, uses the weights of previously tracked features in the weight momentum factor, and uses the weights of all features in the current window in the regularization factor. This allows the camera pose to be estimated simultaneously and features of dynamic objects that deviate significantly from the motion prior are discarded. In the global loop closure optimization, this embodiment groups loop constraints into multiple hypothesis clusters to reject closed-loop detection interference from temporary static objects, eliminate global cumulative drift errors, and form a global SLAM map.
[0097] like Figure 5 As shown, a dynamic environment thermal infrared inertial SLAM system framework with a neural network front end is provided in one embodiment of the present invention.
[0098] The system includes a front-end measurement processing module, a sliding window optimization module and a global optimization module.
[0099] The front-end measurement and processing module first uses the Superpoint network to perform feature extraction and tracking on the thermal infrared image sequence captured by the thermal infrared camera; then performs visual SFM processing on the feature extraction and tracking results; secondly, obtains the key frame set corresponding to the thermal infrared image sequence based on the visual SFM processing results; then, for any key frame in the key frame set: obtains the IMU data corresponding to the key frame from the inertial measurement unit; and performs IMU pre-integration on the IMU data; finally, performs visual-inertial alignment processing based on the IMU pre-integration results and the visual SFM processing results, and outputs the visual reprojection residual and the IMU pre-integration residual.
[0100] The sliding window optimization module first uses the regularization factor and the weight momentum factor to correct the visual reprojection residual to generate the corrected visual reprojection residual; secondly, it constructs the BA optimization model based on the corrected visual reprojection residual, the IMU pre-integration residual, and the marginalization residual; then, it performs dynamic robust BA optimization on the key frames in the sliding window to obtain key frames with static features; finally, based on all key features with static features in each sliding window, it obtains the optimized key frame set.
[0101] The global optimization module first divides the optimized keyframe set into several keyframe groups; then, for the current keyframe received, it uses the DBoW2 bag-of-words to match the current keyframe with the keyframe group; and performs multi-hypothesis clustering based on the matching results; secondly, it selectively optimizes the loop hypothesis in the multi-hypothesis clustering results to obtain the correct loop; finally, it generates a global SLAM map based on the correct loop.
[0102] like Figure 6 , which is a schematic structural diagram of an all-day infrared inertial SLAM device for dynamic environments provided by one embodiment of the present invention.
[0103] An all-day infrared inertial SLAM device for dynamic environments, the device 600 includes: a determination module 601, which is used to determine the motion trajectory of a target UAV based on a thermal infrared image sequence of a target scene; wherein the motion trajectory includes visual pose information corresponding to the target UAV at different moments; a fusion processing module 602, which is used to: based on any current key frame in a key frame set corresponding to the thermal infrared image sequence: pre-integrate IMU data of the current key frame and the previous key frame adjacent to the current key frame to generate a pose estimate of the target UAV; select the visual pose information corresponding to the current moment from the motion trajectory; and perform visual-inertial fusion processing on the visual pose information and the pose estimate to obtain visual reprojection residuals and IMU pre-integration residuals; a BA optimization processing module 603, which is used to perform BA optimization processing on the key frame set according to a preset sliding window based on the visual reprojection residuals, IMU pre-integration residuals, and marginalization residuals, and output an optimized key frame set; a closed-loop optimization processing module 604, which is used to perform global closed-loop optimization processing on the optimized key frame set to output a global SLAM map.
[0104] In a preferred implementation manner of this embodiment, the BA optimization processing module includes: a correction processing unit, which is used to correct the visual reprojection residual corresponding to any current key frame in the preset sliding window using a regularization factor and a weight momentum factor to obtain a corrected visual reprojection residual; a first generation unit, which is used to construct a BA optimization model based on the marginalization residuals corresponding to all current key frames in the preset sliding window, and the corrected visual reprojection residuals and IMU pre-integration residuals corresponding to each of the current key frames; when the BA optimization model tends to be minimum, the static features corresponding to each of the current key frames in the preset sliding window are obtained to generate an optimized key frame; a second generation unit, which is used to generate an optimized key frame set based on the optimized key frame corresponding to each of the preset sliding windows in the key frame set.
[0105] In a preferred implementation manner of this embodiment, the determination module includes: a first generation unit, which is used to perform feature extraction processing on any thermal infrared image in the thermal infrared image sequence of the target scene based on a lightweight SuperPoint network to generate a feature point cloud and a descriptor corresponding to each feature point; a second generation unit, which is used to perform feature matching and tracking processing on all thermal infrared images in the thermal infrared image sequence based on the descriptors corresponding to the feature points to generate a tracking feature point cloud; a third generation unit, which is used to perform visual SFM processing on the tracking feature point cloud to generate a three-dimensional feature point cloud; and a determination unit, which is used to determine the motion trajectory of the target drone based on the three-dimensional feature point cloud.
[0106] In a preferred implementation of this embodiment, the device further includes: a selection module for selecting thermal infrared images with a number of tracking feature points greater than a preset threshold from the thermal infrared image sequence as key frames to obtain a key frame set.
[0107] In a preferred implementation of this embodiment, the closed-loop optimization processing module includes: a division unit for dividing the optimized key frame set into several key frame groups based on tracking feature points; a clustering unit for: selecting a key frame having a loop relationship with the current key frame from the key frame group, and clustering the loop relationship to obtain a cluster group for any key frame group in the several key frame groups; a determination unit for obtaining several cluster groups based on the cluster group corresponding to each of the key frame groups in the several key frame groups; determining the average similarity corresponding to the cluster group based on the similarity of the loop relationship corresponding to each of the loop relationships in the cluster group; a BA optimization processing unit for selecting the cluster groups with the top two average similarities from the several cluster groups as hypothetical clusters, and performing BA optimization processing on the current key frame using the two hypothetical clusters; a global optimization processing unit for selecting the hypothetical cluster with the highest weight from the BA optimization processing results for global pose optimization processing to generate a global SLAM map.
[0108] In a preferred implementation of this embodiment, the clustering unit includes: an identification subunit, which is used to use the DBoW2 bag-of-words model to identify key frames similar to the current key frame from the key frame group; and use the key frames in the identification results whose similarity is greater than a preset threshold as associated key frames to obtain at least one associated key frame; a determination subunit, which is used to determine the Euclidean distance between the associated key frames and the current key frame based on the relative posture information between the two; an establishment subunit, which is used to establish a loop relationship between the associated key frame in the at least one associated key frame whose Euclidean distance is less than the preset threshold and the current key frame, to obtain at least one loop relationship; a clustering subunit, which is used to cluster the at least one loop relationship to obtain a cluster group.
[0109] In a preferred implementation manner of this embodiment, the division unit includes: a first generation sub-unit, which is used for any target key frame in the optimized key frame set: if the number of shared tracking feature points between a first key frame adjacent to the target key frame and the target key frame is not less than a preset threshold, then the first key frame is added to the group corresponding to the target key frame; and the first key frame is used as the next target key frame, and if the number of shared tracking feature points between a second key frame adjacent to the first key frame and the first key frame is not less than a preset threshold, then the second key frame is added to the group of the target key frame, until the number of shared tracking feature points between two adjacent key frames is less than the preset threshold, then the grouping of the target key frame is terminated and a target key frame group is generated; a second generation sub-unit, which is used to generate several key frame groups based on several target key frame groups in the optimized key frame set.
[0110] In a preferred implementation of this embodiment, the first generation unit includes: a feature extraction subunit, which is used to perform feature extraction processing on the thermal infrared image based on the GhostNet network to generate extracted features; a weighted processing subunit, which is used to perform weighted processing on the extracted features based on the long-distance attention mechanism to output a feature point cloud; and a description processing subunit, which is used to perform description processing on each feature point in the feature point cloud to generate a descriptor corresponding to the feature point.
[0111] The above-mentioned device can execute the all-day infrared inertial SLAM method for dynamic environments provided by one embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the all-day infrared inertial SLAM method for dynamic environments. For technical details not fully described in this embodiment, please refer to the all-day infrared inertial SLAM method for dynamic environments provided by one embodiment of the present invention.
[0112] The present invention also provides an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the all-day infrared inertial SLAM method for dynamic environments described in the present invention.
[0113] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present application described in the above-mentioned "Exemplary Method" section of this specification.
[0114] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0115] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enable the processor to execute the steps of the method according to the following embodiments of the present application described in the above "Exemplary Method" section of this specification.
[0116] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0117] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.
[0118] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0119] It should also be noted that in the apparatus, device, and method of the present application, each component or each step can be decomposed and / or recombined, and such decomposition and / or recombination should be regarded as equivalent solutions of the present application.
[0120] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0121] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
[0122] In the description of this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and integrate different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless otherwise inconsistent.
[0123] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0124] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. An all-day infrared inertial SLAM method for dynamic environments, characterized by: include: Determine the motion trajectory of the target UAV based on a sequence of thermal infrared images of the target scene; wherein the motion trajectory includes visual pose information corresponding to the target UAV at different times; Based on any current key frame in the key frame set corresponding to the thermal infrared image sequence: pre-integrate the IMU data of the current key frame and the previous key frame adjacent to the current key frame to generate a pose estimate of the target UAV; select the visual pose information corresponding to the current moment from the motion trajectory; and perform visual-inertial fusion processing on the visual pose information and the pose estimate to obtain a visual reprojection residual and an IMU pre-integration residual; According to the visual reprojection residual, the IMU pre-integration residual, and the marginalization residual, the key frame set is subjected to BA optimization processing according to a preset sliding window, and an optimized key frame set is output; Performing global closed-loop optimization processing on the optimized key frame set and outputting a global SLAM map.
2. The method according to claim 1, characterized in that The method includes performing BA optimization processing on the key frame set according to the visual reprojection residual, the IMU pre-integration residual, and the marginalization residual according to a preset sliding window, and outputting an optimized key frame set; comprising: For any current keyframe in the preset sliding window: using a regularization factor and a weight momentum factor to correct the visual reprojection residual corresponding to the current keyframe to obtain a corrected visual reprojection residual; Based on the marginalized residuals corresponding to all current key frames in the preset sliding window, and the corrected visual reprojection residuals and IMU pre-integration residuals corresponding to each current key frame, a BA optimization model is constructed; when the BA optimization model tends to be minimum, the static features corresponding to each current key frame in the preset sliding window are obtained to generate an optimized key frame; An optimized key frame set is generated based on the optimized key frame corresponding to each preset sliding window in the key frame set.
3. The method according to claim 1, characterized in that The method of determining the motion trajectory of the target UAV based on the thermal infrared image sequence of the target scene comprises: For any thermal infrared image in the thermal infrared image sequence of the target scene: performing feature extraction processing on the thermal infrared image based on a lightweight SuperPoint network to generate a feature point cloud and a descriptor corresponding to each feature point; Based on the descriptors corresponding to the feature points, feature matching and tracking processing are performed on all thermal infrared images in the thermal infrared image sequence to generate a tracking feature point cloud; Performing visual SFM processing on the tracking feature point cloud to generate a three-dimensional feature point cloud; Based on the three-dimensional feature point cloud, the motion trajectory of the target UAV is determined.
4. The method according to claim 3, characterized in that Also includes: A thermal infrared image having a number of tracking feature points greater than a preset threshold is selected from the thermal infrared image sequence as a key frame to obtain a key frame set.
5. The method according to claim 3, characterized in that The method of performing global closed-loop optimization processing on the optimized key frame set and outputting a global SLAM map comprises: Based on the tracking feature points, the optimized key frame set is divided into a plurality of key frame groups; For any key frame group among the plurality of key frame groups: selecting a key frame having a loop relationship with the current key frame from the key frame group, and clustering the loop relationship to obtain a cluster group; Based on the cluster group corresponding to each of the key frame groups in the plurality of key frame groups, a plurality of cluster groups are obtained; based on the loop relationship similarity corresponding to each of the loop relationships in the cluster groups, an average similarity corresponding to the cluster groups is determined; Selecting the top two cluster groups in terms of average similarity from the plurality of cluster groups as hypothetical clusters, and performing BA optimization processing on the current key frame using the two hypothetical clusters; The hypothesis cluster with the highest weight is selected from the BA optimization results for global pose optimization to generate a global SLAM map.
6. The method according to claim 5, characterized in that selecting a key frame having a loop relationship with the current key frame from the key frame group, and clustering the loop relationship to obtain a cluster group; include: Using the DBoW2 bag-of-words model, identify key frames similar to the current key frame from the key frame group; and use key frames with similarity greater than a preset threshold in the identification results as associated key frames to obtain at least one associated key frame; Determining a Euclidean distance between the associated key frame and the current key frame based on relative pose information between the two; Establishing a loop relationship between an associated key frame having a Euclidean distance less than a preset threshold in the at least one associated key frame and the current key frame to obtain at least one loop relationship; Cluster the at least one loop relationship to obtain a cluster group.
7. The method according to claim 5, characterized in that Based on the tracking feature points, the optimized key frame set is divided into a number of key frame groups; including: For any target key frame in the optimized key frame set: if the number of shared tracking feature points between a first key frame adjacent to the target key frame and the target key frame is not less than a preset threshold, the first key frame is added to the group corresponding to the target key frame; and the first key frame is used as the next target key frame. If the number of shared tracking feature points between a second key frame adjacent to the first key frame and the first key frame is not less than a preset threshold, the second key frame is added to the group of the target key frame, until the number of shared tracking feature points between two adjacent key frames is less than the preset threshold, then the grouping of the target key frames is terminated to generate a target key frame group. Based on the plurality of target key frame groups in the optimized key frame set, a plurality of key frame groups are generated.
8. The method according to claim 3, characterized in that The thermal infrared image is subjected to feature extraction processing based on a lightweight SuperPoint network to generate a feature point cloud and a descriptor corresponding to each feature point; including: Performing feature extraction processing on the thermal infrared image based on the GhostNet network to generate extracted features; Performing weighted processing on the extracted features based on a long-range attention mechanism and outputting a feature point cloud; A description process is performed on each of the feature points in the feature point cloud to generate a descriptor corresponding to the feature point.
9. All-day infrared inertial SLAM device for dynamic environment, characterized by: include: A determination module is used to determine the motion trajectory of the target UAV based on the thermal infrared image sequence of the target scene; wherein the motion trajectory includes the visual pose information corresponding to the target UAV at different times; A fusion processing module is configured to: based on any current key frame in the key frame set corresponding to the thermal infrared image sequence, perform pre-integration based on IMU data of the current key frame and the previous key frame adjacent to the current key frame to generate a pose estimate of the target UAV; select visual pose information corresponding to the current moment from the motion trajectory; and perform visual-inertial fusion processing on the visual pose information and the pose estimate to obtain a visual reprojection residual and an IMU pre-integration residual; A BA optimization processing module is used to perform BA optimization processing on the key frame set according to the visual reprojection residual, the IMU pre-integration residual, and the marginalization residual according to a preset sliding window, and output an optimized key frame set; The closed-loop optimization processing module is used to perform global closed-loop optimization processing on the optimized key frame set and output a global SLAM map.
10. A computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Visual inertia SLAM method based on image edge features
CN112749665A
Monocular vision through-guided SLAM method and device based on point-line feature fusion
CN113720323A
Methods of attitude and misalignment estimation for constraint free portable navigation
US20230041831A1