End-to-end thermal image repositioning method, device and equipment based on deep learning
By constructing the ThermalLoc model, combining the efficiency network and transformer module to extract the local and global features of the thermal image, the accuracy problems of traditional thermal camera relocation methods under light changes and low resolution are solved, and high-precision and robust thermal image relocation are achieved.
Patent Information
- Application Number
- CN202510442177.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional thermal camera repositioning methods are difficult to effectively extract feature under light changes, textureless surfaces and dynamic lighting conditions. The thermal image resolution is low and the texture and contrast are reduced, resulting in poor positioning accuracy, high global feature matching calculation cost, and environmental interference leads to spatial and temporal inconsistency.
The end-to-end thermal image relocation method based on deep learning is adopted, and local thermal features are extracted through the efficiency network module, the bridge module performs dimensional transformation and embedding, the transformer module captures global correlation features, the regressor performs pose regression prediction, and a ThermalLoc model is constructed for thermal image relocation.
It improves the accuracy and robustness of thermal image relocation, enhances texture and detail characteristics, adapts to various scenarios and tasks, has flexibility and scalability, and realizes real-time prediction.
Smart Images

Figure CN120374927A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of thermal image relocalization, and particularly to an end-to-end thermal image relocalization method, device, and equipment based on deep learning. Background Art
[0002] Camera relocalization or absolute pose estimation is crucial in fields such as autonomous driving, robot navigation, and virtual and augmented reality. Traditional camera relocalization methods are achieved by constructing a feature map and relying on feature extraction between the image and the map. However, these methods often struggle to cope with challenging situations such as lighting changes, textureless surfaces, and dynamic lighting conditions. Radar-based methods also degrade in performance in environments with air particulates.
[0003] Thermal imaging addresses many of these problems by detecting the distribution of thermal radiation intensity emitted from the surface of an object. It has two key advantages: being unaffected by environmental lighting conditions as it directly measures thermal features, and being able to better penetrate atmospheric obscurants (such as fog, smoke, and dust). Therefore, thermal camera relocalization has advantages such as all-weather adaptability, stable operation in low-light or smoky environments, and robustness to dynamic lighting and certain weather conditions (especially during day-night changes or vehicle light interference).
[0004] However, developing a thermal camera-based positioning system is challenging because thermal images have a lower resolution (usually converted from 14 bits to 8 bits), reduced texture and contrast, making traditional feature extraction ineffective. In large-scale scenes, thermal radiation attenuation reduces the signal-to-noise ratio (SNR), impairing feature extraction of distant objects. Moreover, the periodic non-uniformity correction (NUC) that the thermal camera needs to perform every 500 milliseconds introduces perspective changes and increases odometer trajectory loss. Additionally, as the scene scale and feature density increase, global feature matching becomes computationally expensive. Thermal imaging also suffers from spatio-temporal inconsistency due to environmental interference, and materials exhibit different thermal features depending on the time of day and the viewing angle. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide an end-to-end thermal image relocalization method, device, and equipment based on deep learning, which can, after enhancing the details of the thermal image, first use an efficiency network module to process low-texture detail features, and then use a transformer module to perform global feature correlation extraction to generate a feature representation that has both local detail features and global context awareness for subsequent regression tasks, thereby improving the accuracy and robustness of thermal camera relocalization.
[0006] An end-to-end thermal image relocalization method based on deep learning, the method comprising:
[0007] Preprocess the original thermal image output by the thermal imager;
[0008] Construct an end-to-end ThermalLoc model, and input the preprocessed thermal image dataset into the model for training to obtain a trained ThermalLoc model; wherein, the model consists of an EfficientNet module, a bridging module, a Transformer module, and a regressor connected in sequence; the EfficientNet module is used to extract local thermal features of the preprocessed thermal image; the bridging module is used to perform dimensional transformation and embedding operations on the output of the EfficientNet module according to the shape priority mechanism, and then convert it into the input of the Transformer module; the Transformer module is used to capture the global correlation thermal features corresponding to the input by using multiple layers of Transformer modules; the regressor is used to obtain the output of the Transformer module and perform thermal imager pose regression prediction using two MLPs;
[0009] Input the preprocessed thermal image into the trained ThermalLoc model for thermal image relocalization, and predict and output the 6-degree-of-freedom pose vector of the thermal imager.
[0010] In one embodiment, preprocessing the original thermal image output by the thermal imager includes:
[0011] Use histogram equalization based on linear transformation and a Gaussian low-pass filter to perform brightness adjustment, contrast adjustment, and detail enhancement on the original thermal image in sequence.
[0012] In one embodiment, using histogram equalization based on linear transformation and a Gaussian low-pass filter to perform brightness adjustment, contrast adjustment, and detail enhancement on the original thermal image in sequence includes:
[0013] Perform a linear transformation on the gray value of the original thermal image using histogram equalization based on linear transformation to obtain a thermal image after brightness adjustment and contrast adjustment, denoted as
[0014] P′ = a·P + b;
[0015] where P represents the gray value of the original thermal image output by the thermal imager based on temperature imaging, P′ represents the gray value of the thermal image after brightness adjustment and contrast adjustment, the parameter a is used to adjust the contrast, and the parameter b is used to control the brightness;
[0016] Use a Gaussian low-pass filter to perform sharpening processing on the thermal image after brightness adjustment and contrast adjustment to obtain the preprocessed thermal image, denoted as
[0017] T′ = P′ + h×(P′ - P′*G);
[0018] where T′ represents the preprocessed thermal image, the parameter h is the sharpening intensity coefficient set manually, and P′*G represents the convolution of P′; is the Gaussian kernel function, where x and y represent the position coordinates where the Gaussian kernel function acts, and σ is the standard deviation, which is used to control the intensity of the Gaussian low-pass filter.
[0019] In one embodiment, the efficiency network module consists of a 3*3 convolutional layer, 7 MBConv blocks, and a 1*1 convolutional layer; among them, the 7 MBConv blocks include one MBConv1 block and six MBConv6 blocks. The architectures of these two types of MBConv blocks are the same, only the expansion ratio s1 and the value of the convolutional kernel size k in the architecture are different. In the MBConv1 block, s1 is 1 and k is 3, while in the MBConv6 block, s1 is 6 and k is 3 or 5. The compression ratio s2 in both types of MBConv blocks is 4;
[0020] Among them, the process of the MBConv block for local thermal feature extraction is expressed as
[0021] y = x + Dropout(Conv(SE(DWConv(Conv(x)))));
[0022] Among them, y represents the feature map output by the MBConv block, x represents the input of the MBConv block obtained from the preprocessed thermal image, Dropout is the dropout layer, Conv is the convolutional layer, SE is the squeeze-and-excitation module, and DWConv is the depthwise convolutional layer.
[0023] In one embodiment, the bridging module consists of a dimension transformation unit and an embedding operation unit; among them, the dimension transformation unit is used to sequentially perform data rearrangement, linear transformation, and tensor expansion and splicing on the feature map output by the efficiency network module according to the shape priority mechanism, and output two feature vectors that match the required sizes of the Transformer modules; the embedding operation unit is used to embed the two feature vectors into the Transformer module.
[0024] In one embodiment, the transformer module includes six Transformer modules. The Transformer module consists of a multi-head attention mechanism module and a feed-forward module, and there is a normalization layer in front of both the multi-head attention mechanism module and the feed-forward module to normalize the input data;
[0025] The processing process of the multi-head attention mechanism module is expressed as
[0026] attn(y′) = softmax((Q·K T )×scale)·V;
[0027] Among them, y′ represents the input of the Transformer module, Q, K, and V respectively represent the query, key, and value obtained through linear transformation processing by the normalization layer before the multi-head attention mechanism module, and the superscript T represents transpose; softmax is the activation function, and scale is the scaling factor;
[0028] The feed-forward module consists of two linear layers Linear and a GELU activation function, and the processing process is expressed as
[0029] f forward (o) = Linear(Dropout(GELU(Linear(o))));
[0030] Among them, o represents the input of the feed-forward module, and Dropout is the dropout layer;
[0031] The output of the entire Transformer module is expressed as
[0032] z = f forward (attn(y′) + y′) + attn(y′) + y′;
[0033] After global correlation hot feature extraction through 6 Transformer modules, the output of the transformer module is expressed as
[0034] F = Transformer(y′) = Norm(z 6 );
[0035] Among them, Norm represents normalization processing, and z 6 represents the output after passing through 6 Transformer modules.
[0036] In one embodiment, the regressor uses two MLPs for thermal imager pose regression prediction, expressed as
[0037] [l, q] = MLPs(Transformer(EfficientNet(T′)));
[0038] Among them, l ∈ R 3 and q ∈ R 4 respectively represent the predicted thermal imager position and rotation quaternion, R is the vector space, EfficientNet represents the efficient network module, and T′ is the preprocessed thermal image.
[0039] In one embodiment, the ThermalLoc model is optimized and trained by minimizing the loss between the predicted values of the thermal imager position and rotation quaternion and the ground truth, and the loss function is defined as
[0040]
[0041] Among them, and respectively represent the ground truth of the thermal imager position and the rotation quaternion. β and γ are learning parameters used to balance the contributions of the thermal imager position and the rotation quaternion during the training process. During the training process, the logarithm logq of the unit rotation quaternion is directly used instead of q, and all rotation quaternions are constrained within a hemisphere. The logarithmic transformation of the unit rotation quaternion is defined as
[0042]
[0043] where u and v respectively represent the real part and the imaginary part of q, u is a scalar, v is a three-dimensional vector, and 0 is a zero vector.
[0044] An end-to-end thermal image relocalization device based on deep learning, the device includes:
[0045] A preprocessing module for preprocessing the original thermal image output by the thermal imager;
[0046] A model construction and training module for constructing an end-to-end ThermalLoc model, and inputting the preprocessed thermal image dataset into the model for training to obtain a trained ThermalLoc model. Among them, the model is composed of an EfficientNet module, a bridging module, a Transformer module, and a regressor connected in sequence. The EfficientNet module is used to extract local thermal features of the preprocessed thermal image. The bridging module is used to perform dimensional transformation and embedding operations on the output of the EfficientNet module according to the shape priority mechanism, and then convert it into the input of the Transformer module. The Transformer module is used to capture the global correlation thermal features corresponding to the input by using multiple layers of Transformer modules. The regressor is used to obtain the output of the Transformer module and perform thermal imager pose regression prediction using two MLPs;
[0047] A thermal image relocalization module for inputting the preprocessed thermal image into the trained ThermalLoc model for thermal image relocalization, and predicting and outputting the 6-degree-of-freedom pose vector of the thermal imager.
[0048] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0049] Preprocess the original thermal image output by the thermal imager;
[0050] Build an end-to-end ThermalLoc model, and input the preprocessed thermal image dataset into the model for training to obtain a trained ThermalLoc model; wherein, the model consists of an EfficientNet module, a bridging module, a Transformer module, and a regressor connected in sequence; the EfficientNet module is used to extract local thermal features of the preprocessed thermal images; the bridging module is used to perform dimensional transformation and embedding operations on the output of the EfficientNet module according to the shape priority mechanism, and then convert it into the input of the Transformer module; the Transformer module is used to capture the global associated thermal features of the corresponding input by using multiple Transformer layers; the regressor is used to obtain the output of the Transformer module and perform thermal imager pose regression prediction by using two MLPs.
[0051] Input the preprocessed thermal images into the trained ThermalLoc model for thermal image relocalization, and predict and output the 6-degree-of-freedom pose vector of the thermal imager.
[0052] The above end-to-end thermal image relocalization method, device, and equipment based on deep learning have the following beneficial effects:
[0053] 1. By preprocessing the original thermal images with brightness adjustment, contrast adjustment, and detail enhancement, the edge information in the thermal images becomes clearer, enhancing the texture and detail features of the thermal images and improving the accuracy of subsequent thermal image relocalization.
[0054] 2. By using the end-to-end combined EfficientNet module and Transformer module in the ThermalLoc model, the local details and global associations of thermal image features can be captured simultaneously, improving the accuracy of thermal image relocalization and the robustness of the model in different scenarios, and having flexibility and scalability for various tasks and datasets; moreover, the bridging module built between the two modules is based on the shape priority mechanism, converting the output of the EfficientNet module into the input of the required size of the Transformer module, which can ensure that the feature data is transmitted to the Transformer module in the optimal form, improving the compatibility and maintainability of the model. Brief Description of the Drawings
[0055] Figure 1 It is a schematic flowchart of an end-to-end thermal image relocalization method based on deep learning in an embodiment;
[0056] Figure 2 It is a schematic overall architecture diagram of an end-to-end thermal image relocalization method based on deep learning in an embodiment; wherein, Figure 2 (a) is a schematic diagram of preprocessing the original thermal images; Figure 2 (b) is a schematic overall architecture diagram of the ThermalLoc model; Figure 2 (c) is a schematic structural diagram of the EfficientNet module; Figure 2 (d) is a schematic structural diagram of the Transformer module;
[0057] Figure 3 It is the internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0058] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0059] In one embodiment, as Figure 1 and Figure 2 shown, an end-to-end thermal image relocalization method based on deep learning is provided, including the following steps:
[0060] Step S1, preprocess the original thermal image output by the thermal imager.
[0061] Compared with a trichromatic image, the thermal imager mainly uses color coding to display the temperature distribution map, where the thermal image corresponds to the temperature of the measured environment or object. Although this characteristic is beneficial to temperature analysis, it often faces challenges such as low resolution, poor contrast, limited details, sparse texture information, and edge noise. To solve these problems, in step S1, histogram equalization based on linear transformation and a Gaussian low-pass filter are specifically used to sequentially perform brightness adjustment, contrast adjustment, and detail enhancement on the original thermal image output by the thermal imager. As can be seen from Figure 2 (a), the preprocessing technology applied to the thermal image significantly enhances the representation of detailed texture features. Inputting the preprocessed thermal image into the model enables the model to more accurately capture the local details and global associations of the features, improving the richness of the model feature extraction and the accuracy of relocalization.
[0062] Among them, the brightness adjustment ensures that the brightness range is fixed according to the gray level to prevent flickering in the 10Hz thermal dynamic data collected. The contrast adjustment mainly stretches the gray range to enhance the temperature difference between different objects in the same image. This process makes the edge information clearer and enhances the texture and detail features of the thermal image. The adjustments of brightness and contrast are both achieved by linearly transforming the gray values of the thermal image. Specifically, histogram equalization based on linear transformation is used to linearly transform the gray values of the original thermal image to obtain the thermal image after brightness adjustment and contrast adjustment, expressed as
[0063] P′ = a·P + b;
[0064] where P represents the gray value of the original thermal image output by the thermal imager based on temperature imaging, P′ represents the gray value of the thermal image after brightness adjustment and contrast adjustment, the parameter a is used to adjust the contrast, and the parameter b is used to control the brightness.
[0065] To further enhance the texture detail information of the thermal image, a Gaussian low-pass filter is used to sharpen the thermal image after brightness adjustment and contrast adjustment, and the preprocessed thermal image is obtained, which is expressed as
[0066] T′ = P′ + h × (P′ - P′ * G);
[0067] where T′ represents the preprocessed thermal image, the parameter h is the sharpening intensity coefficient set manually, and P′ * G represents the convolution of P′; is the Gaussian kernel function, where x and y represent the position coordinates where the Gaussian kernel function acts, and σ is the standard deviation, which is used to control the intensity of the Gaussian low-pass filter.
[0068] Step S2, construct an end-to-end ThermalLoc model, and input the preprocessed thermal image dataset into the model for training to obtain a trained ThermalLoc model.
[0069] Among them, Figure 2 (b) shows the overall architecture of the ThermalLoc model. The model consists of an EfficientNet module, a bridging module, a transformer module, and a regressor connected in sequence; the EfficientNet module is used to extract local thermal features of the preprocessed thermal image; the bridging module is used to perform dimensional transformation and embedding operations on the output of the EfficientNet module according to the shape priority mechanism, and then convert it into the input of the transformer module; the transformer module is used to capture the global correlation thermal features corresponding to the input by using multiple layers of Transformer modules; the regressor is used to obtain the output of the transformer module and perform thermal imager pose regression prediction by using two MLPs.
[0070] In this embodiment, EffcientNet-B0 is selected as the EfficientNet module, and the specific structure is as Figure 2 (c) shown, which consists of a 3*3 convolutional layer, 7 MBConv blocks (Inverter Residual Blocks), and a 1*1 convolutional layer. The feature extraction ability of the EfficientNet module is attributed to its use of MBConv blocks. The MBConv block integrates a depth convolutional layer (DWConv) and a squeeze-and-excitation module (SE) to minimize parameters while improving channel representation. Among the 7 MBConv blocks, there is one MBConv1 block and 6 MBConv6 blocks. The architectures of these two types of MBConv blocks are the same, only the values of the expansion ratio s1 and the convolutional kernel size k in the architecture are different. In the MBConv1 block, s1 is 1 and k is 3, while in the MBConv6 block, s1 is 6 and k is 3 or 5. The compression ratio s2 in both types of MBConv blocks is 4.
[0071] Among them, the process of the MBConv block for local thermal feature extraction is expressed as
[0072] y = x + Dropout(Conv(SE(DWConv(Conv(x)))));
[0073] Among them, y represents the feature map output by the MBConv block, x represents the input of the MBConv block obtained from the preprocessed thermal image, Dropout is the dropout layer, Conv is the convolutional layer, SE is the squeeze-and-excitation module, and DWConv is the depthwise convolutional layer. And a batch normalization layer (BN) and a Swish activation function are added after the convolutional layer and the depthwise convolutional layer
[0074] In this embodiment, the bridging module consists of a dimension transformation unit and an embedding operation unit. The dimension transformation unit is used to sequentially perform data rearrangement, linear transformation, and tensor expansion and splicing on the feature map output by the efficient network module according to the shape priority mechanism, and output two feature vectors with the sizes required by the matching Transformer modules. Among them, data rearrangement is used to rearrange the dimensions or order of the feature map output by the efficient network module, aiming to make the format of the feature map match the input requirements of the subsequent Transformer module. Linear transformation is used to perform a linear transformation on the dimensions of the feature map output by the efficient network module to make it consistent with the input dimensions of the subsequent Transformer module. Tensor expansion and splicing is used to assume that the value output after linear transformation is Y (this is a multi-dimensional tensor), and then randomly generate a learnable parameter cls-token with the same dimension as Y, and then splice these two tensors together. y and cls-token are two independent tokens (corresponding Figure 2 to the 2 tokens in
[0075] ), and they are spliced into a sequence as the input of the embedding operation unit.
[0076] In this embodiment, Vision Transformer is selected as the transformer module, which includes 6 Transformer modules based on the sub-attention mechanism. The structure of the Transformer module is as Figure 2As shown in (d), it is composed of a multi-head attention mechanism module (Multi-Head Attention) and a feedforward module (Feedforward), and there is a normalization layer (Norm) before both the multi-head attention mechanism module and the feedforward module to normalize the input data. Moreover, different from the traditional Transformer model, the Transformer module constructed in this application is different in the following two key aspects: (1) In this application, the hidden attention mechanism in the traditional Transformer model is deleted, simplifying the model structure, reducing the computational overhead of the attention mechanism, and improving the running efficiency of the model. (2) In this application, the linear transformation process of the normalization layer is used to calculate Q (query), K (key), and V (value), rather than using multiple MLPs (multi-layer perceptrons) to calculate. The linear transformation process further reduces the computational complexity of the model, and through the normalization layer, the training process can be stabilized, and the problem of gradient disappearance or explosion can be alleviated.
[0077] Specifically, the processing process of the multi-head attention mechanism module is expressed as
[0078] attn(y′) = softmax((Q · K T ) × scale) · V;
[0079] where y′ represents the input of the Transformer module, Q, K, and V respectively represent the query, key, and value obtained through the linear transformation process of the normalization layer before the multi-head attention mechanism module, and the superscript T represents transpose; softmax is the activation function, and scale is the scaling factor.
[0080] The feedforward module consists of two linear layers Linear and a GELU activation function, and the processing process is expressed as
[0081] f forward (o) = Linear(Dropout(GELU(Linear(o))));
[0082] where o represents the input of the feedforward module, and Dropout is the dropout layer.
[0083] The output of the entire Transformer module is expressed as
[0084] z = f forward (attn(y′) + y′) + attn(y′) + y′.
[0085] After global correlation hot feature extraction through 6 Transformer modules, the output of the transformer module is expressed as
[0086]
[0087] Among them, Norm represents normalization processing, and z 6 represents the output after passing through 6 Transformer modules.
[0088] Through the end-to-end combined EfficientNet module and Transformer module, the model can simultaneously capture the local details and global correlations of thermal image features, improve the thermal image relocalization accuracy and the robustness of the model in different scenarios, and has flexibility and scalability for various tasks and datasets.
[0089] In this embodiment, the regressor (Pose Regressor) uses two MLPs for thermal imager pose regression prediction, denoted as
[0090] [l, q] = MLPs(Transformer(EfficientNet(T')));
[0091] where l ∈ R 3 and q ∈ R 4 respectively represent the predicted thermal imager position and rotation quaternion, R is the vector space, EfficientNet represents the EfficientNet module, and T' is the preprocessed thermal image.
[0092] Furthermore, the ThermalLoc model is optimized and trained by minimizing the loss between the predicted values of the thermal imager position and rotation quaternion and the ground truth, and the loss function is defined as
[0093]
[0094] where and respectively represent the ground truth of the thermal imager position and rotation quaternion, β and γ are learning parameters, which are used to balance the contributions of the thermal imager position and rotation quaternion during the training process, and during the training process, the logarithm logq of the unit rotation quaternion is directly used instead of so as to avoid the normalization problem in the loss function. This method reduces the influence of outliers, enhances the robustness to atypical observations, and encourages parameter and feature sparsity. The logarithmic transformation of the unit rotation quaternion is defined as
[0095]
[0096] Among them, u and v respectively represent the real and imaginary parts of q, where u is a scalar and v is a three-dimensional vector; 0 is the zero vector. Further, the rotation vector is mapped to a three-dimensional space and converted into a normalized unit rotation quaternion, which represents the rotation of the thermal imager pose estimation. However, due to the non-uniqueness of the rotation quaternion (i.e., -q and q can represent the same rotation), this application constrains all selected quaternions to a hemisphere to ensure a unique representation.
[0097] Step S3: Input the preprocessed thermal image into the trained ThermalLoc model for thermal image relocalization, and predict and output the 6-degree-of-freedom pose vector of the thermal imager.
[0098] In summary, an end-to-end thermal image relocalization method based on deep learning provided by this application combines an efficiency network module and a transformer module by constructing a ThermalLoc model, can effectively extract local and global correlation thermal features from the preprocessed thermal image, and realizes 6-degree-of-freedom pose regression prediction of the thermal imager based on two MLPs, improving the accuracy and robustness of thermal image relocalization.
[0099] The beneficial effects of an end-to-end thermal image relocalization method based on deep learning provided by this application have been verified by experiments. The experimental settings are as follows: To ensure consistent training across different datasets, we resize the thermal image to 224×224 pixels and normalize it within the intensity range of -1 to 1. The EffcientNet-B0 module of the ThermalLoc model is initialized using a pre-trained model on the ImageNet dataset, while the remaining modules are randomly initialized. Then, the features extracted from EffcientNet-B0 are fed into the Transformer module for further processing. The model is trained 300 times from scratch using an NVIDIA RTX 3090 GPU.
[0100] During the test, on an NVIDIA RTX 4060 GPU, the processing time for each thermal image is 6 milliseconds. Using an infrared camera with a frequency of 10 Hz, the ThermalLoc model can achieve real-time prediction. When training the model, the ADAM optimizer is used, the learning rate is 5×10 -5 , the batch size is 8, the dropout rate is 0.5, and the weights are initialized as β0 = -3.0 and γ0 = 0.0.
[0101] Table 1 Thermal Image Dataset
[0102] Scene sequence Season Label Distance Number of images IF-1 Spring Daytime 3.5km 22329 IF-2 Summer Night 1.0km 7735 IF-3 Summer Night 3.3km 20828 IF-4 Summer Night 1.7km 11521
[0103] Due to the scarcity of relocalization data for large-scale thermal cameras, the experiment collected a thermal image dataset across an urban-scale environment from four different scenarios. This dataset covers a variety of environmental conditions, including different seasons (spring and summer), different lighting conditions (daytime and nighttime), and different traffic dynamics. The dataset consists of thermal images, each with a resolution of 480×270 pixels. Ground truth data for position and orientation was obtained using high-precision RTK (Real-Time Kinematic) differential positioning, providing a position accuracy of up to 0.1 meters. Table 1 provides an overview of the thermal image dataset, which has dynamic elements such as moving and stationary vehicles, cyclists, and pedestrians, posing additional challenges for the visual relocalization task.
[0104] To evaluate the advantages of the method proposed in this application, the ThermalLoc model was compared with existing end-to-end learning-based visual localization methods, including PoseNet (a CNN-based pose network), AtLoc (a CNN+attention-based localization network), and Robust-Loc (a CNN+Transformer-based robust localization network). RobustLoc uses a robust neural differential equation diffusion block module for information interaction and diffusion, enhancing the robustness to environmental interference. These end-to-end methods can be directly applied to the processed thermal images without considering the conversion of image patterns. They address the challenges posed by the lack of texture and details in thermal images. The same thermal image dataset was used for training and testing in all network models, and position error and angle error were used to evaluate the relocalization performance.
[0105] Table 2 calculates the median and average errors of position and orientation for each trajectory in four scenarios in the thermal image dataset for four models
[0106]
[0107] Table 2 presents the comparison of the ThermalLoc model constructed in this application with PoseNet, AtLoc, and RobustLoc on the thermal image dataset. As can be seen from Table 2, compared with PoseNet, the ThermalLoc model significantly improves the localization accuracy. For example, in the IF-2 scenario, the average localization accuracy increases from 5.93 meters to 3.48 meters. In the IF-4 scenario, the average localization accuracy increases from 8.54 meters to 5.27 meters, representing improvements of 41.3% and 38.3% respectively. For a more intuitive comparison, the IF-1 scenario is selected as a typical example. In this case, the ThermalLoc model shows a substantial improvement in localization accuracy compared to other models. Compared with PoseNet, the ThermalLoc model reduces the average position error by 68.7%. Compared with AtLoc, it reduces by 42.4%, and compared with RobustLoc, it reduces by 11.4%.
[0108] Furthermore, an ablation study was conducted using the IF-1 scenario in the thermal image dataset shown in Table 1 to evaluate the impact of different building components on the ThermalLoc model.
[0109] First, an ablation study was carried out to evaluate the contribution of the EfficientNet module in the ThermalLoc model. As shown in Table 3, "ResNet-based" in Table 3 means replacing the EfficientNet module in the ThermalLoc model with ResNet (Residual Network) while keeping all other components of ThermalLoc unchanged. "Only EfficientNet" in Table 3 indicates using only the EfficientNet module as the backbone for thermal image feature extraction. The results show that replacing the EfficientNet module with ResNet in ThermalLoc significantly reduces the performance. Specifically, in the IF-1 scenario, the localization accuracy of ThermalLoc is 44.3% higher than that of the ResNet-based model. In addition, the configuration of the EfficientNet module is superior to AtLoc in terms of re-localization accuracy, demonstrating the applicability of the EfficientNet module in thermal image feature extraction.
[0110] Table 3 calculates the average errors of ResNet-based, only EfficientNet, and AtLoc in the IF-1 scenario
[0111] Model Based on residual network Only efficiency network AtLoc IF-1 7.51m,2.26° 5.01m,2.53° 7.26m,3.24°
[0112] Then, to improve the integration of the efficiency network module and the transformer module, an ablation study was conducted on the bridging module, as shown in Table 4. Table 4 shows the impact of the bridging module based on different mechanisms on the relocalization results. Among them, Patch-First (patch-first mechanism) is used to process the feature maps output by the efficiency network module into patches, convert them into embeddings, and then adjust the embedding size to fit the Transformer model. The VINT-like mechanism is used to follow the connection strategy outlined in the VINT network, where the feature maps are masked before being input into the Transformer model. As can be seen from Table 4, the shape-first mechanism in the ThermalLoc model constructed in this application is much better than the other two mechanisms, indicating that the bridging module designed in this application can improve the thermal image relocalization accuracy.
[0113] Table 4 Comparison of the average errors of position and rotation for relocalization using bridging modules based on different mechanisms
[0114] Mechanism Patch priority mechanism Shape priority mechanism VINT-like mechanism IF-1 6.53m,2.06° 4.18m,1.93° 17.30m,3.77° IF-2 3.92m,16.48° 3.48m,14.81° 10.76m,18.07° IF-3 7.95m,9.83° 6.10m,9.09° 19.41m,12.51° IF-4 6.05m,12.11° 5.27m,11.84° 11.45m,16.29° Mean 6.11m,10.12° 4.76m,9.42° 14.73m,12.66°
[0115] Finally, to determine the most effective Transformer configuration, a series of ablation studies were conducted on different types and models of Transformer modules with different depths. As shown in Table 5, "mask-based" in Table 5 means using Masked Goal Attention to mask the feature maps, and "depth = n" means using an n-layer Transformer module to extract global feature correlations from thermal images. The results shown in Table 5 indicate that the Transformer module without a mask constructed in this application performs better in thermal image relocalization, and the 6-layer Transformer module achieves the best balance between efficiency and effectiveness.
[0116] Table 5 Comparison of the average errors of position and rotation for relocalization in the IF-1 scenario using Transformer modules with different depths and masked Transformer modules
[0117]
[0118] In one embodiment, an end-to-end thermal image relocalization device based on deep learning is provided, including:
[0119] A preprocessing module for preprocessing the original thermal image output by the thermal imager;
[0120] The model construction and training module is used to construct an end-to-end ThermalLoc model, and input the preprocessed thermal image dataset into the model for training to obtain a trained ThermalLoc model. Among them, the model consists of an efficientnet module, a bridging module, a transformer module, and a regressor connected in sequence. The efficientnet module is used to extract the local thermal features of the preprocessed thermal images. The bridging module is used to perform dimensional transformation and embedding operations on the output of the efficientnet module according to the shape priority mechanism, and then convert it into the input of the transformer module. The transformer module is used to capture the global associated thermal features of the corresponding input by using multiple layers of Transformer modules. The regressor is used to obtain the output of the transformer module and perform thermal imager pose regression prediction by using two MLPs.
[0121] The thermal image relocalization module is used to input the preprocessed thermal image into the trained ThermalLoc model for thermal image relocalization, and predict and output the 6-degree-of-freedom pose vector of the thermal imager.
[0122] For the specific limitations of the end-to-end thermal image relocalization device based on deep learning, reference can be made to the limitations of the end-to-end thermal image relocalization method based on deep learning in the above text, which will not be elaborated here. Each module in the above end-to-end thermal image relocalization device based on deep learning can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0123] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an end-to-end thermal image relocalization method based on deep learning. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0124] Those skilled in the art can understand, Figure 3The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. Specifically, the computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0125] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0126] Preprocess the original thermal image output by the thermal imager;
[0127] Construct an end-to-end ThermalLoc model, and input the preprocessed thermal image dataset into the model for training to obtain a trained ThermalLoc model; wherein, the model is composed of an EfficientNet module, a bridging module, a Transformer module, and a regressor connected in sequence; the EfficientNet module is used to extract local thermal features of the preprocessed thermal image; the bridging module is used to perform dimensional transformation and embedding operations on the output of the EfficientNet module according to the shape priority mechanism, and then convert it into the input of the Transformer module; the Transformer module is used to capture the global associated thermal features of the corresponding input by using multiple layers of Transformer modules; the regressor is used to obtain the output of the Transformer module and perform thermal imager pose regression prediction by using two MLPs;
[0128] Input the preprocessed thermal image into the trained ThermalLoc model for thermal image relocalization, and predict and output the 6-degree-of-freedom pose vector of the thermal imager.
[0129] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0130] The above-described embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. An end-to-end thermal image relocalization method based on deep learning, characterized in that, The method includes: Preprocessing the original thermal image output by the thermal imager; Constructing an end-to-end ThermalLoc model, and inputting the preprocessed thermal image dataset into the model for training to obtain a trained ThermalLoc model; wherein, the model consists of an efficientnet module, a bridging module, a transformer module, and a regressor connected in sequence; the efficientnet module is used to extract local thermal features of the preprocessed thermal image; the bridging module is used to perform dimensional transformation and embedding operations on the output of the efficientnet module according to the shape priority mechanism, and then convert it into the input of the transformer module; the transformer module is used to capture global correlation thermal features corresponding to the input by using multiple Transformer modules; the regressor is used to obtain the output of the transformer module and perform thermal imager pose regression prediction by using two MLPs; Inputting the preprocessed thermal image into the trained ThermalLoc model for thermal image relocalization, and predicting and outputting a 6-degree-of-freedom pose vector of the thermal imager.
2. The method according to claim 1, wherein Preprocessing the original thermal image output by the thermal imager, including: Performing brightness adjustment, contrast adjustment, and detail enhancement on the original thermal image in sequence by using histogram equalization based on linear transformation and a Gaussian low-pass filter.
3. The method according to claim 2, wherein Performing brightness adjustment, contrast adjustment, and detail enhancement on the original thermal image in sequence by using histogram equalization based on linear transformation and a Gaussian low-pass filter, including: Performing linear transformation on the gray value of the original thermal image by using histogram equalization based on linear transformation to obtain a thermal image after brightness adjustment and contrast adjustment, denoted as P′ = a·P + b; wherein, P represents the gray value of the original thermal image output by the thermal imager based on temperature imaging, P′ represents the gray value of the thermal image after brightness adjustment and contrast adjustment, the parameter a is used to adjust the contrast, and the parameter b is used to control the brightness; Performing sharpening processing on the thermal image after brightness adjustment and contrast adjustment by using a Gaussian low-pass filter to obtain a preprocessed thermal image, denoted as T′ = P′ + h×(P′ - P′*G); Among them, T′ represents the preprocessed thermal image, the parameter h is the sharpening intensity coefficient set manually, and P′*G represents the convolution of P′; is the Gaussian kernel function, where x and y represent the position coordinates where the Gaussian kernel function acts, and σ is the standard deviation, which is used to control the intensity of the Gaussian low-pass filter.
4. The method according to claim 1, characterized in that The efficientnet module consists of a 3*3 convolutional layer, 7 MBConv blocks, and a 1*1 convolutional layer; wherein, the 7 MBConv blocks include one MBConv1 block and 6 MBConv6 blocks, and the architectures of these two types of MBConv blocks are the same, only the values of the expansion ratio s1 and the convolutional kernel size k in the architecture are different. In the MBConv1 block, s1 is 1 and k is 3, while in the MBConv6 block, s1 is 6 and k is 3 or 5, and the compression ratio s2 in both types of MBConv blocks is 4; wherein, the process of the MBConv block for extracting local thermal features is denoted as y = x + Dropout(Conv(SE(DWConv(Conv(x))))); Among them, y represents the feature map output by the MBConv block, x represents the input of the MBConv block obtained from the preprocessed thermal image, Dropout is the dropout layer, Conv is the convolutional layer, SE is the squeeze-and-excitation module, and DWConv is the depthwise convolutional layer.
5. The method according to claim 1, wherein The bridging module consists of a dimension transformation unit and an embedding operation unit; among them, the dimension transformation unit is used to sequentially perform data rearrangement, linear transformation, and tensor expansion and splicing on the feature map output by the efficiency network module according to the shape priority mechanism, and output two feature vectors that match the dimensions required by the Transformer module; the embedding operation unit is used to embed the two feature vectors into the Transformer module.
6. The method according to claim 1, wherein The transformer module includes 6 Transformer modules, and each Transformer module consists of a multi-head attention mechanism module and a feed-forward module. There is a normalization layer before both the multi-head attention mechanism module and the feed-forward module to normalize the input data. The processing process of the multi-head attention mechanism module is expressed as attn(y′) = softmax((Q·K T ) × scale)·V; Among them, y′ represents the input of the Transformer module, Q, K, and V respectively represent the query, key, and value obtained by linear transformation through the normalization layer before the multi-head attention mechanism module, and the superscript T represents transpose; softmax is the activation function, and scale is the scaling factor. The feed-forward module consists of two linear layers Linear and a GELU activation function, and the processing process is expressed as f forward (o) = Linear(Dropout(GELU(Linear(o)))) Among them, o represents the input of the feed-forward module, and Dropout is the dropout layer. The output of the entire Transformer module is expressed as z = f forward (attn(y′) + y′) + attn(y′) + y′; After global correlation thermal feature extraction through 6 Transformer modules, the output of the transformer module is expressed as F = Transformer(y′) = Norm(z 6 ); Among them, Norm represents normalization processing, and z 6 represents the output after passing through 6 Transformer modules.
7. The method according to claim 1, characterized in that, The regressor uses two MLPs for thermal imager pose regression prediction, expressed as [l,q]=MLPs(Transformer(EfficientNet(T′))); where \(l\in\mathbb{R}\) 3 and \(q\in\mathbb{R}\) 4 represent the predicted thermal imager position and the rotation quaternion respectively, \(\mathbb{R}\) is the vector space, EfficientNet represents the EfficientNet module, and \(T'\) is the preprocessed thermal image.
8. The method according to claim 7, wherein The ThermalLoc model is optimized and trained by minimizing the loss between the predicted values of the thermal imager position and rotation quaternion and the ground truth, and the loss function is defined as Among them, and respectively represent the ground truth of the thermal imager position and the rotation quaternion. β and γ are learning parameters, which are used to balance the contributions of the thermal imager position and the rotation quaternion during the training process. During the training process, the logarithm logq of the unit rotation quaternion is directly used instead of q, and all rotation quaternions are constrained within a hemisphere; the logarithmic transformation of the unit rotation quaternion is defined as Among them, u and v respectively represent the real part and the imaginary part of q, and u is a scalar and v is a three-dimensional vector; 0 is the zero vector.
9. An end-to-end thermal image relocalization device based on deep learning, characterized in that, The device includes: A preprocessing module for preprocessing the original thermal image output by the thermal imager. The model construction and training module is used to construct an end-to-end ThermalLoc model, and input the preprocessed thermal image dataset into the model for training to obtain a trained ThermalLoc model; wherein, the model is composed of an EfficientNet module, a bridging module, a Transformer module, and a regressor connected in sequence; the EfficientNet module is used to extract local thermal features of the preprocessed thermal images; the bridging module is used to perform dimensionality transformation and embedding operations on the output of the EfficientNet module according to the shape priority mechanism, and then convert it into the input of the Transformer module; the Transformer module is used to capture global correlation thermal features corresponding to the input by using multiple layers of Transformer modules; the regressor is used to obtain the output of the Transformer module and perform thermal imager pose regression prediction by using two MLPs. The thermal image relocalization module is used to input the preprocessed thermal image into the trained ThermalLoc model for thermal image relocalization, and predict and output the 6-degree-of-freedom pose vector of the thermal imager.
10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.