Infrared and visible light image unsupervised explicit registration method based on style migration

The method uses style transfer and explicit feature matching to address modality differences in infrared and visible light image registration, enhancing accuracy and transparency through unsupervised learning.

CN120318282APending Publication Date: 2025-07-15CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510423343.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing infrared and visible image registration methods are difficult to effectively handle modal differences, resulting in inaccurate feature extraction and matching, lack of explicit matching processes, lack of transparency, and difficult to debug and improve.

Method used

Unsupervised explicit registration method based on style transfer is adopted, and the conversion of visible light images to pseudo-infrared images is achieved using CycleGAN. It combines multi-scale feature extraction and improved HiLo Attention mechanism for feature matching, and fine adjustments are made through deformation field optimization functions to build a multi-layer decoder structure for image conversion and resampling.

Benefits of technology

Improves the accuracy and transparency of infrared and visible image registration, enhances the interpretability of the registration process, significantly reduces the impact of modal differences, and achieves more versatile unsupervised registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318282A_ABST
    Figure CN120318282A_ABST
Patent Text Reader

Abstract

The invention provides an infrared and visible light image unsupervised explicit registration method based on style migration, and belongs to the technical field of electrical digital data process.The infrared and visible light image unsupervised explicit registration method based on style migration comprises the steps that firstly, a registration data set of a deformed infrared image and a visible light image is constructed; according to the method, style migration from visible light to pseudo infrared is realized by using Cyc leGAN, and cross-modal registration is converted into a same-modal problem. An improved ResNet18 structure is adopted to extract multi-scale features, and more detail information is reserved through discrete wavelet transform. A feature matching module based on self-attention and cross attention is constructed, explicit matching of features is achieved, and interpretability is improved. And designing a multi-layer decoder structure to estimate a deformation field, performing fine adjustment by using a deformation field optimization function, and finally performing resampling through a space conversion module to obtain a registration result. An unsupervised learning mode is adopted in the whole process, data labeling is not needed, and good universality and practicability are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electrical digital data processing. Specifically, it relates to an unsupervised explicit registration method for infrared and visible light images based on style transfer. Background Art

[0002] Multi-modal image registration technology has important application value in the field of computer vision. In particular, the registration of infrared and visible light images can effectively fuse complementary information captured in different bands and is widely used in scenarios such as night vision monitoring, medical image analysis, and target recognition. Traditional registration methods are mainly based on feature point matching or mutual information optimization, but they perform poorly when dealing with images with large modal differences and it is difficult to establish reliable cross-modal matching relationships.

[0003] In recent years, deep learning technology has been introduced into the field of image registration, and significant progress has been made in end-to-end registration models based on convolutional neural networks (CNNs). However, there are generally two significant problems in existing deep learning registration methods: on the one hand, it is difficult for them to effectively handle the large modal differences between infrared and visible light, resulting in inaccurate feature extraction and matching; on the other hand, most models adopt an implicit feature fusion mechanism, which cannot show the explicit matching process and lacks interpretability.

[0004] The above problems lead to insufficient accuracy of existing infrared and visible light image registration methods when facing complex scenarios, and the registration process lacks transparency, making it difficult to conduct effective debugging and improvement. There is an urgent need for a registration method that can effectively handle modal differences and has an explicit matching mechanism. That is to say, there are technical problems in the prior art of lacking effective handling of modal differences and lacking explicit feature matching in the registration process of infrared and visible light images. Summary of the Invention

[0005] In view of this, the present invention provides an unsupervised explicit registration method for infrared and visible light images based on style transfer, which can solve the technical problems of lacking effective handling of modal differences and lacking explicit feature matching in the registration process of infrared and visible light images in the prior art.

[0006] The present invention is implemented as follows: The present invention provides an unsupervised explicit registration method for infrared and visible images based on style transfer, including: selecting paired infrared and visible images from a dataset, performing deformation processing on the infrared images to obtain deformed infrared images, and constructing a registration dataset; training a generative adversarial network CycleGAN to learn the bidirectional mapping between the two modalities and converting the visible images into pseudo-infrared images; constructing a multi-scale feature extraction module, adopting a ResNet18 network architecture and replacing the standard max pooling layer with a discrete wavelet transform; constructing a feature matching module based on self-attention and cross-attention, replacing the dot product with Linear Attention to reduce the computational complexity; constructing a deformation field generation and image conversion module, designing a multi-layer decoder structure to estimate the deformation field; using a deformation field optimization function to finely adjust the deformation field; using the optimized deformation field and a spatial transformation module to resample the deformed infrared images to obtain infrared images registered with the visible images; training a registration model using the pseudo-infrared images and the deformed infrared image dataset; inputting the infrared image and the visible image to be registered into the trained registration network, performing style transfer, feature extraction, matching, and deformation field estimation to obtain the registration result.

[0007] Among them, CycleGAN includes two generators G_AB and G_BA and two discriminators D_A and D_B, and realizes bidirectional domain conversion through cycle consistency loss and adversarial loss.

[0008] Among them, Self-HiLo refers to a self-attention mechanism in which the query, key, and value all come from the same image feature, and is used to extract the global dependencies within a single image.

[0009] Among them, Cross-HiLo refers to a cross-attention mechanism in which the query comes from one image feature and the key and value come from another image feature, and is used to achieve feature matching between different images.

[0010] Among them, the deformation field is a two-dimensional vector field that describes the pixel mapping relationship from the source image to the target image, and the vector at each position represents the displacement of that point during the registration process.

[0011] Among them, the spatial transformation module is a differentiable image resampling operation that interpolates and reconstructs the source image according to the deformation field to generate a result aligned with the target image.

[0012] Among them, the deformation field optimization function is used to perform fine adjustment on the basis of the initial deformation field estimation, and optimizes the deformation field by minimizing the image similarity loss and regularization constraints.

[0013] Among them, the input of the deformation field optimization function includes the initial deformation field generated by the network, the pseudo-infrared image generated by the style transfer network, the deformed infrared image, the gradient constraint coefficient, and the smoothing constraint coefficient, and the output is the optimized deformation field.

[0014] Among them, the infrared image to be registered and the visible light image are input into the trained registration network. The visible light image is converted into a pseudo-infrared image through style transfer, and then feature extraction, matching, and deformation field estimation are performed to obtain the registration result.

[0015] Among them, the gradient constraint coefficient is a weight parameter that controls the gradient amplitude of the deformation field and is used to prevent excessive local deformation of the deformation field; the smoothing constraint coefficient is a weight parameter that controls the smoothness of the deformation field and is used to ensure the spatial continuity of the deformation field.

[0016] Compared with the prior art, the present invention innovatively combines cross-modal style transfer and explicit feature matching technologies, effectively solving the problems of modal differences and interpretability in the registration of infrared and visible light images.

[0017] By introducing CycleGAN to achieve style transfer from visible light to pseudo-infrared, this method converts complex cross-modal registration into a relatively simple same-modal problem, significantly reducing the interference caused by modal differences. At the same time, the self-attention and cross-attention feature matching module based on the improved HiLoAttention mechanism realizes explicit feature matching, making the registration process visual and having good interpretability.

[0018] Compared with the prior art, the present invention not only improves the registration accuracy of infrared and visible light images, but also enhances the transparency of the registration process. At the same time, without labeled data, it realizes a more general registration ability through unsupervised learning, effectively solving the technical problems of the lack of effective processing of modal differences and the lack of explicit feature matching in the registration process of infrared and visible light images in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the purpose, technical solution, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention.

[0021] As Figure 1 shown, it is a flowchart of an unsupervised explicit registration method for infrared and visible light images based on style transfer provided by the present invention. This method includes the following steps:

[0022] S1. Select paired infrared and visible images from the dataset to construct a training set and a test set. Translate, rotate, scale, and distort the infrared images to obtain deformed infrared images. Use the visible images as templates and the deformation fields as references to construct a registration dataset;

[0023] S2. Train the generative adversarial network CycleGAN to learn the bidirectional mapping between the two modalities, realizing unsupervised image-to-image conversion. Input the visible images into the trained cross-modal style transfer network to generate pseudo-infrared images;

[0024] S3. Construct a multi-scale feature extraction module. Use the ResNet18 network architecture to extract image features. Replace the standard max pooling layer with discrete wavelet transform to perform downsampling operations, enhancing the feature extraction ability;

[0025] S4. Construct a feature matching module based on self-attention and cross-attention to achieve explicit self-matching and mutual matching. Improve the HiLo Attention mechanism, and replace the dot product with Linear Attention to reduce the computational complexity;

[0026] S5. Construct a deformation field generation and image conversion module. Design a multi-layer decoder structure to estimate the deformation field. The input of each layer of the decoder is the matching features and image features at different scales, and finally output a two-channel deformation field;

[0027] S6. Use the deformation field optimization function to finely adjust the deformation field to improve the registration accuracy. The input of the deformation field optimization function includes the initial deformation field, pseudo-infrared image, deformed infrared image, gradient constraint coefficient, and smoothing constraint coefficient, and the output is the optimized deformation field;

[0028] S7. Use the optimized deformation field and the spatial transformation module to resample the deformed infrared images to obtain infrared images registered with the visible images;

[0029] S8. Use the pseudo-infrared images generated by style transfer and the deformed infrared image dataset to train the registration model to achieve unsupervised explicit registration of infrared and visible images based on transfer learning;

[0030] S9. Input the infrared image and the visible image to be registered into the trained registration network. Convert the visible image into a pseudo-infrared image through style transfer, and then perform feature extraction, matching, and deformation field estimation to obtain the registration result;

[0031] Among them, CycleGAN is a generative adversarial network that can learn the mapping relationship between two image domains without paired data, including two generators G_AB and G_BA and two discriminators D_A and D_B, and realizes bidirectional domain conversion through cyclic consistency loss and adversarial loss.

[0032] Among them, ResNet18 is a residual neural network with 18 layers, which solves the problem of difficult training of deep networks through skip connections and has excellent feature extraction capabilities.

[0033] Among them, the discrete wavelet transform is a multi-resolution analysis method, which can retain more image detail information compared to max pooling and reduce information loss during downsampling.

[0034] Among them, HiLo Attention is a more computationally efficient attention mechanism. By separately processing the high-frequency and low-frequency components of the input, it reduces the computational complexity while maintaining good performance.

[0035] Among them, Self-HiLo refers to a self-attention mechanism where the query, key, and value all come from the same image feature, and is used to extract global dependencies within a single image.

[0036] Among them, Cross-HiLo refers to a cross-attention mechanism where the query comes from one image feature while the key and value come from another image feature, and is used to achieve feature matching between different images.

[0037] Among them, the deformation field is a two-dimensional vector field that describes the pixel mapping relationship from the source image to the target image. The vector at each position represents the displacement of that point during the registration process.

[0038] Among them, the spatial transformation module is a differentiable image resampling operation that interpolates and reconstructs the source image according to the deformation field to generate a result aligned with the target image.

[0039] Among them, the deformation field optimization function is used to perform fine-tuning based on the initial deformation field estimate. It optimizes the deformation field by minimizing the image similarity loss and regularization constraints. The inputs include the initial deformation field generated by the network, the pseudo-infrared image generated by the style transfer network, the deformed infrared image, the gradient constraint coefficient that controls the strength of the gradient constraint, and the smoothness constraint coefficient that controls the smoothness of the deformation field. The output is the optimized deformation field that meets the requirements of similarity and smoothness.

[0040] Among them, the gradient constraint coefficient is a weight parameter that controls the gradient amplitude of the deformation field and is used to prevent excessive local deformation of the deformation field.

[0041] Among them, the smoothness constraint coefficient is a weight parameter that controls the smoothness of the deformation field and is used to ensure the spatial continuity of the deformation field.

[0042] The following details the specific implementation manners of the above steps.

[0043] The specific implementation of step S1 is to select paired infrared and visible light images from the dataset to construct the training set and the test set. Select the publicly available dataset RoadScene that has been strictly aligned and calibrated, and select 436 pairs of images as the training set and 20 pairs of images as the test set. Each pair of images contains one infrared image and one visible light image. Perform rigid transformation and non-rigid transformation on the infrared images. The rigid transformation includes affine transformations such as translation, rotation, and scaling. The translation range is 10% of the original image size, the rotation angle range is plus or minus 15 degrees, and the scaling ratio is from 0.9 to 1.1. The non-rigid transformation uses a randomly generated grid to distort the image, and the random offset of the grid points does not exceed 25% of the grid spacing. Save the mapping from the source infrared image to the deformed infrared image, that is, the deformation field, as a.npy file, which together with the pseudo-infrared image and the deformed infrared image constitutes the dataset of the registration model. This step provides a basis for subsequent model training by constructing a training dataset, simulating various geometric deformation situations existing between infrared and visible light images in actual applications.

[0044] The specific implementation of step S2 is to train the generative adversarial network CycleGAN to learn the bidirectional mapping between the two modalities of infrared images and visible light images. The CycleGAN architecture contains two generators and two discriminators. The generator G_AB converts images in the visible light domain to images in the infrared domain, and the generator G_BA converts images in the infrared domain back to the visible light domain. The generator adopts an encoder-decoder structure, with 9 residual blocks in the middle. The residual block consists of two 3×3 convolutional layers and a skip connection. The discriminators D_A and D_B are used to distinguish real images and generated images in the visible light domain and the infrared domain respectively. The discriminator adopts a PatchGAN structure, and the output is a true / false judgment at the image patch level. The generator is trained by minimizing the adversarial loss and the cycle consistency loss. The adversarial loss makes the generated image look real in the target domain, and the cycle consistency loss ensures that the image can be restored to its original state after two conversions. During the training process, the Adam optimizer is used, the learning rate is set to 0.0002, and the weight of the cycle consistency loss is set to 10. After training is completed, input the visible light image into the trained generator G_AB to obtain the pseudo-infrared image. This step converts the multi-modal problem into a single-modal problem through style transfer technology, reducing the impact of modal differences on the registration accuracy.

[0045] The specific implementation of step S3 is to construct an infrared and visible light image registration network, including a multi-scale feature extraction module, a feature matching module based on self-attention and cross-attention, and a deformation field generation and image conversion module. The feature extraction module is based on the ResNet18 architecture and is initialized with pre-trained weights. Each residual layer contains two residual blocks, with a convolutional kernel size of 3×3, a stride of 1, and the number of channels being 64, 128, 256, and 512 respectively. The max pooling downsampling in the traditional ResNet18 is replaced by the discrete wavelet transform, which decomposes the input features into approximation coefficients and detail coefficients, retaining more image detail information. The feature matching module uses the self-attention and cross-attention mechanisms to achieve explicit self-matching and mutual matching. It improves the HiLo Attention by changing the dot product attention to Linear Attention and reducing the computational amount by separating high-frequency and low-frequency features. The number of input channels and output channels of the attention module is the same. In the Self-HiLo module, the query, key, and value vectors all come from the same image feature and are used to extract the global dependencies within a single image. In the Cross-HiLo module, the query vector comes from one image feature, and the key and value vectors come from another image feature, which is used to achieve feature matching between different images. The deformation field generation module adopts a multi-layer decoder structure. The first layer upsamples the matching features at the 1 / 8 scale and concatenates them with the image features at the 1 / 4 scale for deconvolution. The second layer upsamples the features of the first layer and concatenates them with the image features at the 1 / 2 scale for deconvolution. The third layer directly deconvolves the output of the second layer after upsampling. The fourth layer directly deconvolves to output a two-channel deformation field. The spatial transformation module uses the generated deformation field to resample the deformed infrared image. First, the deformation field is superimposed on the created grid to obtain the flow field, and then the deformed infrared image is bilinearly interpolated according to the flow field coordinates, with the boundary region filled with a value of 0. This step realizes the registration process of infrared and visible light images and improves the registration accuracy through multi-scale feature extraction and explicit feature matching.

[0046] The specific implementation of step S4 is to use the pseudo-infrared image obtained by style transfer and the deformed infrared image dataset to train a registration model. The model training adopts the mini-batch stochastic gradient descent method with a batch size of 8, the initial value of the learning rate is 0.0001, and the cosine annealing strategy is used for learning rate adjustment. The maximum number of iterations is 100 epochs. The training process uses the bidirectional similarity loss and the smoothness loss function. The bidirectional similarity loss is used to constrain the similarity between the registered image and the source image in the feature space, and at the same time, the forward and backward similarity relationships are considered. The reverse weight λ_rev is set to 0.2. The smooth deformation field loss is used to constrain the smoothness of the deformation field and avoid generating unreasonable local deformations. The smoothness loss weight λ_sm is set to 10. The network parameters are optimized by gradient backpropagation. The model performance is evaluated on the validation set every 10 iteration epochs, and the model weights with the best performance are saved. This step completes the training process of the registration model and obtains a model that can achieve unsupervised explicit registration of infrared and visible light images.

[0047] The specific implementation of step S5 is to input the pseudo-infrared image and the deformed infrared image into the trained registration network to achieve registration. First, the visible light image to be registered is input into the trained style transfer network to generate a pseudo-infrared image. Then, the pseudo-infrared image and the deformed infrared image are simultaneously input into the registration network. The multi-scale features are extracted by the feature extraction module, the self-matching and mutual matching are performed by the feature matching module, the deformation field is estimated by the deformation field generation module, and finally, the deformed infrared image is resampled by the spatial transformation module to obtain the infrared image registered with the visible light image. The registration result evaluation uses indicators such as mutual information and structural similarity. Mutual information reflects the shared information content contained in two images, and structural similarity reflects the similarity degree of the image structure. Both are important evaluation criteria for the registration quality. This step realizes the application process of the model and completes the registration task of infrared and visible light images.

[0048] The specific implementation of step S6 is to use a deformation field optimization function to finely adjust the initially estimated deformation field and improve the registration accuracy. The deformation field optimization function is based on the variational optimization principle and gradually optimizes the deformation field in an iterative manner. The inputs of this function include the initial deformation field generated by the registration network, the pseudo-infrared image generated by the style transfer network, the deformed infrared image, the gradient constraint coefficient, and the smoothing constraint coefficient. Among them, the gradient constraint coefficient controls the gradient amplitude of the deformation field, with a value range of 0.1 to 1.0 and a default setting of 0.5; the smoothing constraint coefficient controls the smoothness of the deformation field, with a value range of 5 to 20 and a default setting of 10. The optimization process first calculates the similarity measure between the pseudo-infrared image and the deformed infrared image, using normalized cross-correlation as the similarity measure function; then calculates the gradient amplitude and Laplacian operator of the deformation field as regularization constraints; finally, iteratively optimizes the deformation field by the gradient descent method, with the number of iterations set to 50 and the step size of each iteration being 0.01. After optimization, a deformation field that meets the requirements of similarity and smoothness is output. This step improves the accuracy of the deformation field through post-processing optimization and enhances the robustness of the registration method.

[0049] The specific implementation of step S7 is to use the optimized deformation field and the spatial transformation module to resample the deformed infrared image to obtain an infrared image registered with the visible light image. The spatial transformation module first applies the optimized deformation field to a regular sampling grid to generate the new coordinate positions of each pixel point; then uses the bilinear interpolation method to resample the deformed infrared image according to the new coordinates to generate the registered infrared image. For coordinate points outside the image boundary, zero-padding processing is adopted. The registration accuracy is evaluated using metrics such as mean square error and mutual information, and compared with the registration results without deformation field optimization to verify the effectiveness of the optimization. This step completes the final image registration process and obtains an infrared image that is precisely aligned with the visible light image.

[0050] The specific implementation of step S8 is to use the pseudo-infrared image generated by style transfer and the deformed infrared image dataset to train the registration model to achieve unsupervised explicit registration of infrared and visible light images based on transfer learning. During the training process, the entire dataset is randomly divided into a training set and a validation set with a ratio of 9:1. The training set is used to update the model parameters, and the validation set is used to evaluate the generalization performance. The training uses the adaptive moment estimation optimizer with an initial learning rate of 0.0001, a weight decay of 0.0005, and trains for 300 epochs. The learning rate is reduced to 0.1 times the original value every 100 epochs. The loss function includes the image similarity loss and the deformation field smoothness loss, and the total loss is the weighted sum of the two. During the training process, monitor the registration accuracy metrics on the validation set and adopt an early stopping strategy to prevent overfitting. Stop training when the validation loss does not decrease for 20 consecutive epochs. This step solves the modality difference problem in the registration of infrared and visible light images through the transfer learning method and improves the accuracy and robustness of the registration.

[0051] The specific implementation of step S9 is to input the infrared image to be registered and the visible light image into the trained registration network. Through style transfer, the visible light image is converted into a pseudo-infrared image, and then feature extraction, matching, and deformation field estimation are performed to obtain the registration result. In the registration process, the visible light image is first input into the generator G_A to generate a pseudo-infrared image, and then the pseudo-infrared image and the infrared image to be registered are input into the registration network together. The registration network is based on three modules: multi-scale feature extraction, self-attention and cross-attention feature matching, and deformation field generation and image conversion, and outputs the deformation field and the registered image. The registration result is quantitatively evaluated by calculating indicators such as mutual information and structural similarity to verify the effectiveness of the registration method. According to the experimental results, the unsupervised explicit registration method based on style transfer improves the registration accuracy by 10% to 15% compared with the traditional method, especially in the scenario of large differences in illumination conditions. This step completes the overall process of the registration method and realizes the high-precision registration of infrared and visible light images.

[0052] The following describes in detail the mathematical models or calculation processes involved in the present invention.

[0053] In step S2, the training process of the generative adversarial network CycleGAN involves the calculation of multiple loss functions. The adversarial loss is used to train the generator to generate realistic images, and at the same time train the discriminator to distinguish between real images and generated images, which is specifically expressed as follows:

[0054] L GAN (G A , D B ) = E ir [log(D B (I ir ))] + E vis [log(1 - D B (G A (I vis )))];

[0055] In the formula, G A is the generator from the visible light domain to the infrared domain; D B is the discriminator in the infrared domain; I ir is the infrared image sample; I vis is the visible light image sample; E ir and E vis respectively represent the expectations over the infrared image distribution and the visible light image distribution; log is the natural logarithm function.

[0056] The cycle consistency loss ensures that the image can be restored to the original image after two conversions, which is specifically expressed as follows:

[0057] Lcycle (G A ,G B )=E vis [||G B (G A (I vis ))-I vis ||1]+E ir [||G A (G B (I ir ))-I ir ||1];

[0058] In the formula, G B is the generator from the infrared domain to the visible light domain; ||·||1 represents the L1 norm, which calculates the sum of the absolute errors between two images.

[0059] The total loss function is the weighted sum of the adversarial loss and the cycle consistency loss, and is specifically expressed as follows:

[0060] L total =L GAN +λ cycle L cycle ;

[0061] In the formula, λ cycle is the balance parameter, which is used to adjust the weight of the cycle consistency loss, and the default value is 10.

[0062] The parameters of the generator and the discriminator are optimized by the gradient descent method, and the parameter update formula is as follows:

[0063]

[0064] In the formula, and are the parameters of the generator and the discriminator respectively; η is the learning rate, and the value is 0.0002; represents the gradient operator.

[0065] In step S3, the discrete wavelet transform is used to replace the max pooling for downsampling, and the calculation formula based on the two-dimensional discrete wavelet transform is as follows:

[0066] W φ [j, k, l]=∑ m ∑ n x[m, n]φ j,k,l [m, n];

[0067]

[0068] In the formula, x[m, n] is the input image; φ j,k,l is the scaling function; and Wavelet functions in the horizontal, vertical, and diagonal directions, respectively; W φ is the approximation coefficient; and are the detail coefficients in the horizontal, vertical, and diagonal directions, respectively; j is the decomposition level; k and l are spatial position indices.

[0069] In the improved HiLo Attention, the calculation formula of Linear Attention is as follows:

[0070] Attention(Q, K, V) = V · LinearAttention(K T · Q);

[0071] In the formula, Q is the query matrix; K is the key matrix; V is the value matrix; K T is the transpose of K; LinearAttention is the current attention normalization exponential function.

[0072] The calculation process of the Self-HiLo attention mechanism is as follows:

[0073] Q h = W Q · F h ;

[0074] K h = W K · F h ;

[0075] V h = W V · F h ;

[0076]

[0077] O h = A h · V h ;

[0078] In the formula, F h is the high-frequency part of the input feature; W Q , W K and W V are linear transformation parameter matrices; d k is the dimension of the key vector; A h is the attention weight matrix; O h is the high-frequency attention output.

[0079] The low-frequency attention calculation is as follows:

[0080] F l = AvgPool(F);

[0081] Q l = W Q · F l ;

[0082] K l = W K · F l ;

[0083] V l = W V · F l ;

[0084]

[0085] O l = A l · V l ;

[0086] O l ′ = Upsample(O l );

[0087] Where F is the input feature; AvgPool is the average pooling operation; F l is the low-frequency feature; O l ′ is the upsampled low-frequency attention output.

[0088] The final HiLo attention output is:

[0089] O HiLo = O h + O l ′;

[0090] During the calculation of the Cross-HiLo attention mechanism, the query comes from one image feature and the key-value comes from another image feature:

[0091] Q h = W Q · F 1h ;

[0092] K h = W K · F 2h ;

[0093] V h = W V · F 2h ;

[0094]

[0095] O h = A h · V h ;

[0096] Wherein, F 1h and F 2h are respectively the high-frequency parts of two different image features.

[0097] In step S4, the bidirectional similarity loss used in the training process is calculated as follows:

[0098]

[0099] Wherein, is the registered infrared image; I ir ′ is the pseudo-infrared image; I ir is the source infrared image (deformed infrared image); φ is the deformation field estimated by the registration network; ψ j is the feature extraction function; λ rev is the reverse weight, with a value of 0.2.

[0100] The smooth deformation field loss is calculated as follows:

[0101]

[0102] Wherein, represents the gradient operator, which calculates the spatial gradient of the deformation field.

[0103] The overall registration loss is the weighted sum of the similarity loss and the smoothness loss:

[0104] L reg = L sim + λ sm L smooth ;

[0105] Wherein, λ sm is the smoothness loss weight, with a value of 10.

[0106] In step S6, the deformation field optimization function is based on the variational optimization principle, and its objective function is expressed as follows:

[0107] E(φ) = E data (φ) + αE grad (φ) + βE smooth (φ);

[0108] Wherein, E(φ) is the total energy function; E data (φ) is the data term, which measures the similarity of the registered image; E grad (φ) is the gradient constraint term, which limits the gradient amplitude of the deformation field; E smooth (φ) is the smoothness constraint term, which ensures the smoothness of the deformation field; α is the gradient constraint coefficient, with a value range of 0.1 to 1.0 and a default value of 0.5; β is the smoothness constraint coefficient, with a value range of 5 to 20 and a default value of 10.

[0109] The calculation formula for the data item is as follows:

[0110]

[0111] In the formula, NCC is the normalized cross - correlation coefficient, and its calculation formula is:

[0112]

[0113] In the formula, I1 and I2 are two images; and are the average gray values of the images; (x, y) are pixel coordinates.

[0114] The calculation formula for the gradient constraint term is as follows:

[0115]

[0116] In the formula, ||·||2 represents the L2 norm, which calculates the Euclidean norm of the gradient of the deformation field.

[0117] The calculation formula for the smoothing constraint term is as follows:

[0118]

[0119] In the formula, represents the Laplacian operator; ||·|| F represents the Frobenius norm.

[0120] The deformation field is iteratively optimized by the gradient - descent method, and the update formula is:

[0121]

[0122] In the formula, φ (t) and φ (t+1) are the deformation fields of the t - th and (t + 1)-th iterations respectively; γ is the step size, and its value is 0.01; is the gradient of the energy function with respect to the deformation field.

[0123] In step S7, the spatial transformation module uses the deformation field for image resampling and calculates the sampling coordinates of each target pixel:

[0124] x src = x dst + φ x (x dst , y dst );

[0125] y src = y dst + φ y (x dst, y dst );

[0126] wherein, (x dst , y dst ) are the pixel coordinates in the target image; (x src , y src ) are the sampling coordinates in the source image; φ x and φ y are the components of the deformation field in the x and y directions, respectively.

[0127] The bilinear interpolation calculation formula is as follows:

[0128]

[0129] wherein, I dst and I src are the target image and the source image, respectively; and represent floor and ceiling operations, respectively; and are the interpolation weights.

[0130] The construction principles and meanings of the above equations are as follows: The adversarial loss and the cycle consistency loss are the core components of CycleGAN. The adversarial loss promotes the generator to produce realistic images by maximizing the probability of the discriminator correctly classifying and minimizing the probability of the generator being recognized as fake by the discriminator. The cycle consistency loss ensures that the generator learns meaningful mapping relationships rather than random mappings by requiring the image to remain unchanged after bidirectional transformation. The discrete wavelet transform can retain more image detail information compared to traditional max-pooling, which helps to improve the quality of feature extraction. Linear Attention replaces the quadratic complexity of the traditional attention mechanism with a linear complexity calculation method, significantly reducing the computational cost. The bidirectional similarity loss considers both the forward and backward image similarities, which is more comprehensive than the unidirectional similarity metric and can capture more registration information. The deformation field optimization function models the registration problem as an energy minimization problem through a variational framework. The data term ensures the similarity of the registered images, and the gradient constraint and the smoothness constraint ensure the rationality and smoothness of the deformation field. The spatial transformation module realizes image resampling in a differentiable manner, supporting end-to-end network training. The design of these equations fully considers the special challenges of infrared and visible light image registration, and effectively reduces the impact of modality differences on the registration accuracy through style transfer and explicit feature matching.

[0131] Optionally, in the improved part of HiLo Attention in step S3, the calculation formula of Linear Attention can also be in the following way:

[0132] Attention(Q, K, V) = φ(Q) · (φ(K) T · V ));

[0133] where φ is a feature mapping function, generally φ(x) = elu(x) + 1, where elu is the exponential linear unit activation function; this calculation method reduces the complexity of the attention mechanism from O(n 2 ) to O(n), significantly reducing the amount of computation.

[0134] In step S3, the specific calculation process of the decoder is as follows:

[0135] F1 = Conv(Concat(Upsample(F match1 / 8 ), F img1 / 4 ));

[0136] F2 = Conv(Concat(Upsample(F1), F match1 / 2 ));

[0137] F3 = Conv(Upsample(F2));

[0138] φ = Conv(F3);

[0139] where F match1 / 8 and F match1 / 2 are matching features at scales of 1 / 8 and 1 / 2 respectively; F img1 / 4 is the image feature at a scale of 1 / 4; Conv is the convolution operation; Concat is the feature concatenation operation; Upsample is the upsampling operation; φ is the output deformation field.

[0140] In step S5, the calculation formula for the mutual information, an evaluation index of the registration result, is:

[0141]

[0142] where is the joint probability distribution of the two images; and are the marginal probability distributions of the two images respectively.

[0143] The calculation formula for the structural similarity is:

[0144]

[0145] where and are the average gray values of the two images respectively; and are the variances of the two images respectively; is the covariance of two images; C1 and C2 are stability constants, with values of (0.01L) 2 and (0.03L) 2 , where L is the dynamic range of the image.

[0146] Specifically, the principle of the present invention is as follows: The technical solution of the present invention is based on two core principles: cross-modal style transfer and explicit feature matching. First, using the bidirectional domain conversion ability of CycleGAN, the modal mapping relationship between infrared and visible light images is learned, and the visible light image is converted into a pseudo-infrared image with infrared characteristics. This preprocessing step effectively reduces the modal difference, enabling subsequent feature extraction and matching to be carried out in a similar feature space, thereby improving the registration accuracy.

[0147] In terms of feature extraction, the present invention adopts an improved ResNet18 structure, replacing the standard max pooling layer with discrete wavelet transform. This replacement can retain more image detail information, reduce information loss during downsampling, and provide richer feature expressions for subsequent feature matching. Discrete wavelet transform can better represent the multi-scale characteristics of the image by decomposing the image into different frequency components, which helps to capture structural information at different scales.

[0148] Explicit feature matching is another core principle of the present invention, which is achieved through an improved HiLo Attention mechanism. Self-HiLo attention is used to mine the global dependencies within a single image, while Cross-HiLo attention establishes the matching relationship between the features of different images. Replacing dot-product attention with Linear Attention significantly reduces the computational complexity, enabling the model to process higher-resolution images. This explicit feature matching mechanism not only improves the registration accuracy but also enhances the interpretability of the model.

[0149] The deformation field generation and optimization link further ensures the accuracy of registration. The multi-layer decoder structure uses matching features and image features at different scales to gradually refine the deformation field estimation. The deformation field optimization function minimizes the image similarity loss and regularization constraints, achieving precise pixel-level alignment while ensuring the continuity and smoothness of the deformation.

[0150] The following provides a specific embodiment 1 of the present invention. The specific implementation of each step in this embodiment 1 is described in detail as follows.

[0151] The specific implementation of step S1 is the same as the foregoing, and will not be elaborated in detail here.

[0152] The specific implementation of step S2 is to train the generative adversarial network CycleGAN to learn the bidirectional mapping between two modalities of infrared images and visible light images. The CycleGAN architecture contains two generators and two discriminators. The generator G_AB converts images in the visible light domain into images in the infrared domain, and the generator G_BA converts images in the infrared domain back into images in the visible light domain. The generator adopts an encoder-decoder structure with 9 residual blocks in the middle. The residual block consists of two 3×3 convolutional layers and a skip connection. The discriminators D_A and D_B are used to distinguish real images from generated images in the visible light domain and the infrared domain respectively. The discriminator adopts a PatchGAN structure, and the output is a true / false judgment at the image patch level. The adversarial loss is used to train the generator to generate realistic images and at the same time train the discriminator to distinguish real images from generated images, which is specifically expressed as follows:

[0153] L GAN (G A , D B ) = E ir [log(D B (I ir ))] + E vis [log(1 - D B (G A (I vis )))];

[0154] In the formula, G A is the generator from the visible light domain to the infrared domain; D B is the discriminator in the infrared domain; I ir is the infrared image sample; I vis is the visible light image sample; E ir and E vis respectively represent the expectations over the infrared image distribution and the visible light image distribution; log is the natural logarithm function.

[0155] The cycle consistency loss ensures that the image can be restored to the original image after two conversions, which is specifically expressed as follows:

[0156] L cycle (G A , G B ) = E vis [||G B (G A (I vis )) - I vis ||1] + E ir [||G A (G B (I ir )) - I ir ||1];

[0157] In the formula, GB is a generator from the infrared domain to the visible light domain; ||·||1 represents the L1 norm, which calculates the sum of the absolute errors between two images.

[0158] The total loss function is the weighted sum of the adversarial loss and the cycle consistency loss, which is specifically expressed as follows:

[0159] L total = L GAN + λ cycle L cycle ;

[0160] where λ cycle is the balance parameter, which is used to adjust the weight of the cycle consistency loss, and the default value is 10.

[0161] The parameters of the generator and the discriminator are optimized by the gradient descent method, and the parameter update formula is as follows:

[0162]

[0163]

[0164] where and are the parameters of the generator and the discriminator respectively; η is the learning rate, and the value is 0.0002; represents the gradient operator. This step converts the multi-modal problem into a single-modal problem through the style transfer technology, reducing the influence of modal differences on the registration accuracy.

[0165] The specific implementation of step S3 is to construct an infrared and visible light image registration network, including a multi-scale feature extraction module, a feature matching module based on self-attention and cross-attention, and a deformation field generation and image conversion module. The feature extraction module is based on the ResNet18 architecture and is initialized with pre-trained weights. Each residual layer contains two residual blocks, the convolution kernel size is 3×3, the stride is 1, and the number of channels is 64, 128, 256, and 512 respectively. The maximum pooling downsampling in the traditional ResNet18 is replaced by the discrete wavelet transform, which decomposes the input features into approximation coefficients and detail coefficients, retaining more image detail information. The calculation formula of the discrete wavelet transform is as follows:

[0166] W φ [j, k, l] = ∑ m ∑ n x[m, n]φ j,k,l [m, n];

[0167]

[0168] where x[m, n] is the input image; φ j,k,lis the scaling function; and are the wavelet functions in the horizontal, vertical, and diagonal directions respectively; W φ is the approximation coefficient; and are the detail coefficients in the horizontal, vertical, and diagonal directions respectively; j is the decomposition level; k and l are spatial position indices.

[0169] The feature matching module uses self-attention and cross-attention mechanisms to achieve explicit self-matching and mutual matching, improves HiLo Attention, changes dot-product attention to Linear Attention, and reduces the computational complexity by separating high-frequency and low-frequency features. The calculation formula of Linear Attention is as follows:

[0170] Attention(Q, K, V) = φ(Q) · (φ(K) T · V );

[0171] where φ is the feature mapping function, generally φ(x) = elu(x) + 1, where elu is the exponential linear unit activation function; Q is the query matrix; K is the key matrix; V is the value matrix; K T is the transpose of K.

[0172] The calculation process of the Self-HiLo attention mechanism is as follows. The query, key, and value vectors all come from the same image feature:

[0173] Q h = W Q · F h ;

[0174] K h = W K · F h ;

[0175] V h = W V · F h ;

[0176]

[0177] O h = A h · V h ;

[0178] where F h is the high-frequency part of the input feature; W Q , W K and W V are the linear transformation parameter matrices; d k is the dimension of the key vector; Ah is the attention weight matrix; O h is the high-frequency attention output.

[0179] The low-frequency attention is calculated as follows:

[0180] F l = AvgPool(F);

[0181] O l = W Q ·F l ;

[0182] K l = W K ·F l ;

[0183] V l = W V ·F l ;

[0184]

[0185] O l = A l ·V l ;

[0186] O l ′ = Upsample(O l );

[0187] In the formula, F is the input feature; AvgPool is the average pooling operation; F l is the low-frequency feature; O l ′ is the upsampled low-frequency attention output.

[0188] The final HiLo attention output is:

[0189] O HiLO = O h + O l ′;

[0190] In the Cross-HiLo attention mechanism, the query comes from one image feature, and the key-value comes from another image feature:

[0191] Q h = W Q ·F 1h ;

[0192] K h = W K ·F 2h ;

[0193] V h = W V ·F2h ;

[0194]

[0195] O h = A h ·V h ;

[0196] In the formula, F 1h and F 2h are respectively the high-frequency parts of two different image features.

[0197] The deformation field generation module adopts a multi-layer decoder structure, and its calculation process is as follows:

[0198] F1 = Conv(Concat(Upsample(F match1 / 8 ), F img1 / 4 ));

[0199] F2 = Conv(Concat(Upsample(F1), F match1 / 2 ));

[0200] F3 = Conv(Upsample(F2));

[0201] φ = Conv(F3);

[0202] In the formula, F match1 / 8 and F match1 / 2 are respectively the matching features at 1 / 8 and 1 / 2 scales; F img1 / 4 is the image feature at 1 / 4 scale; Conv is the convolution operation; Concat is the feature concatenation operation; Upsample is the upsampling operation; φ is the output deformation field. This step realizes the registration process of infrared and visible light images, and improves the registration accuracy through multi-scale feature extraction and explicit feature matching.

[0203] The specific implementation method of step S4 is to use the pseudo-infrared image and the deformed infrared image dataset obtained by style transfer to train the registration model. The model training adopts the mini-batch stochastic gradient descent method with a batch size of 8, the initial value of the learning rate is 0.0001, and the cosine annealing strategy is used for learning rate adjustment. The maximum number of iterations is 100 epochs. The bidirectional similarity loss and the smoothing loss function are used in the training process. The bidirectional similarity loss is calculated as follows:

[0204]

[0205] In the formula, is the registered infrared image; I ir ′ is the pseudo-infrared image; I iris the source infrared image (warped infrared image); φ is the warping field estimated by the registration network; ψ j is the feature extraction function; λ rev is the inverse weight, with a value of 0.2.

[0206] The smooth warping field loss is calculated as follows:

[0207]

[0208] In the formula, represents the gradient operator, which calculates the spatial gradient of the warping field.

[0209] The overall registration loss is the weighted sum of the similarity loss and the smoothness loss:

[0210] L reg = L sim + λ sm L smooth ;

[0211] In the formula, λ sm is the smoothness loss weight, with a value of 10. The network parameters are optimized by gradient backpropagation. The model performance is evaluated on the validation set every 10 iteration cycles, and the model weights with the best performance are saved. This step completes the training process of the registration model, and a model that can achieve unsupervised explicit registration of infrared and visible light images is obtained.

[0212] The specific implementation of step S5 is to input the pseudo-infrared image and the warped infrared image into the trained registration network to achieve registration. First, the visible light image to be registered is input into the trained style transfer network to generate a pseudo-infrared image. Then, the pseudo-infrared image and the warped infrared image are simultaneously input into the registration network. The multi-scale features are extracted by the feature extraction module, the self-matching and mutual-matching are performed by the feature matching module, the warping field is estimated by the warping field generation module, and finally, the warped infrared image is resampled by the spatial transformation module to obtain an infrared image registered with the visible light image. The registration results are evaluated using indicators such as mutual information and structural similarity. The formula for mutual information is:

[0213]

[0214] In the formula, is the joint probability distribution of the two images; and are the marginal probability distributions of the two images respectively.

[0215] The formula for structural similarity is:

[0216]

[0217] In the formula, and are the average gray values of the two images respectively; and are the variances of the two images respectively; is the covariance of the two images; C1 and C2 are stability constants, with values of (0.01L) 2 and (0.03L) 2 , where L is the dynamic range of the image. This step realizes the application process of the model and completes the registration task of infrared and visible light images.

[0218] The specific implementation of step S6 is to use the deformation field optimization function to finely adjust the initially estimated deformation field to improve the registration accuracy. The deformation field optimization function is based on the variational optimization principle and gradually optimizes the deformation field in an iterative manner. The inputs of this function include the initial deformation field generated by the registration network, the pseudo-infrared image generated by the style transfer network, the deformed infrared image, the gradient constraint coefficient, and the smoothing constraint coefficient. The objective function is expressed as follows:

[0219] E(φ) = E data (φ) + αE grad (φ) + βE smooth (φ);

[0220] In the formula, E(φ) is the total energy function; E data (φ) is the data term, which measures the similarity of the registered images; E grad (φ) is the gradient constraint term, which limits the gradient amplitude of the deformation field; E smooth (φ) is the smoothing constraint term, which ensures the smoothness of the deformation field; α is the gradient constraint coefficient, with a value range of 0.1 to 1.0 and a default value of 0.5; β is the smoothing constraint coefficient, with a value range of 5 to 20 and a default value of 10.

[0221] The calculation formula of the data term is as follows:

[0222]

[0223] In the formula, NCC is the normalized cross-correlation coefficient, and the calculation formula is:

[0224]

[0225] In the formula, I1 and I2 are the two images; and are the average gray values of the images; (x, y) are pixel coordinates.

[0226] The calculation formula of the gradient constraint term is as follows:

[0227]

[0228] In the formula, ||·||2 represents the L2 norm, which calculates the Euclidean norm of the deformation field gradient.

[0229] The calculation formula of the smoothness constraint term is as follows:

[0230]

[0231] In the formula, represents the Laplace operator; ||·|| F represents the Frobenius norm.

[0232] The deformation field is iteratively optimized by the gradient descent method, and the update formula is:

[0233]

[0234] In the formula, φ (t) and φ (t+1) are the deformation fields of the t-th and (t + 1)-th iterations respectively; γ is the step size, and its value is 0.01; is the gradient of the energy function with respect to the deformation field. This step improves the accuracy of the deformation field through post-processing optimization and enhances the robustness of the registration method.

[0235] The specific implementation of step S7 is to resample the deformed infrared image using the optimized deformation field and the spatial transformation module to obtain an infrared image registered with the visible light image. The spatial transformation module first applies the optimized deformation field to a regular sampling grid to generate new coordinate positions for each pixel point:

[0236] x src = x dst + φ x (x dst , y dst );

[0237] y src = y dst + φ y (x dst , y dst );

[0238] In the formula, (x dst , y dst ) is the pixel coordinate in the target image; (x src , y src ) is the sampling coordinate in the source image; φ x and φ y are the components of the deformation field in the x and y directions respectively.

[0239] The bilinear interpolation calculation formula is as follows:

[0240]

[0241] Wherein, I dst and I src are the target image and the source image respectively; and respectively represent floor function and ceiling function; and are interpolation weights. The registration accuracy is evaluated using metrics such as mean square error and mutual information, and is compared with the registration results without deformation field optimization to verify the effectiveness of the optimization. This step completes the final image registration process and obtains an infrared image that is precisely aligned with the visible light image.

[0242] The specific implementation manners of steps S8 - S9 are the same as those described above and will not be elaborated in detail here.

[0243] To better understand and implement the present invention, the following provides Example 2 of a specific application scenario of the present invention: The research team applied the unsupervised explicit registration method for infrared and visible light images based on style transfer in a certain intelligent monitoring system. This monitoring system is deployed around important facilities in a complex environment and requires the simultaneous use of infrared and visible light cameras for all - weather monitoring. Due to the significant modal differences between infrared and visible light images, traditional feature - point matching methods often have poor registration effects at night or under bad weather conditions. To solve this problem, the team implemented the method of the present invention.

[0244] First, the research team selected 436 pairs of infrared and visible light images from the RoadScene dataset as the training set and 20 pairs of images as the test set. The resolution of each pair of images is 640×480 pixels. The team performed a series of deformation operations on the infrared images, including translation (up to 10% of the original image size, i.e., 64 pixels), rotation (within the range of plus or minus 15 degrees), scaling (0.9 to 1.1 times), and non - rigid distortion based on a 10×10 grid (the random offset of grid points does not exceed 25% of the grid spacing). The constructed training dataset contains 4360 pairs of deformed infrared images and corresponding deformation field data.

[0245] Next, the research team trained a CycleGAN network to achieve style transfer between infrared and visible light images.

[0246] The configuration parameters of the CycleGAN network are shown in Table 1:

[0247] Table 1 Configuration Table of CycleGAN Network Training Parameters

[0248] Parameter Name Parameter Value Number of Generator Residual Blocks 9 Number of Discriminator Convolution Layers 5 Batch Size 4 Learning Rate 0.0002 Cycle Consistency Loss Weight 10 Number of Training Epochs 200 Optimizer Adam(β1 = 0.5, β2 = 0.999) Learning Rate Strategy Linear Decay (starting from 100 epochs)

[0249] After training, the team used the trained generator G_A to convert visible light images into pseudo-infrared images. The transformation effect met the expectations, retaining the structural information of the original images while possessing the characteristics of infrared images, and the average structural similarity SSIM reached 0.763.

[0250] Subsequently, the team constructed and trained a registration network. The registration network consists of three main modules: a multi-scale feature extraction module, a feature matching module based on self-attention and cross-attention, and a deformation field generation and image conversion module. Among them, the feature extraction module is based on ResNet18, replacing traditional max-pooling with discrete wavelet transform. The feature matching module uses an improved HiLo Attention mechanism, reducing the computational complexity. The training parameters of the registration network are shown in Table 2:

[0251] Table 2 Configuration Table of Registration Network Training Parameters

[0252] Parameter Name Parameter Value Batch Size 8 Initial Learning Rate 0.0001 Learning Rate Strategy Cosine Annealing Number of Training Epochs 100 Reverse Weight λ_rev 0.2 Smoothing Loss Weight λ_sm 10 Optimizer Adam(β1 = 0.9, β2 = 0.999) Early Stopping Patience 20 Random Seed 42

[0253] During the training process, the team monitored the changes in various indicators on the validation set. The training curve shows that the loss function tended to stabilize after approximately 65 epochs, and finally, the bidirectional similarity loss on the validation set dropped to 0.127, and the smooth loss dropped to 0.043.

[0254] After completing the training, the team evaluated the performance of the registration method on the test set. To further improve the registration accuracy, the team applied a deformation field optimization function to finely adjust the deformation field generated by the network. The parameter settings of the deformation field optimization function are shown in Table 3:

[0255] Table 3 Configuration Table of Deformation Field Optimization Function Parameters

[0256] Parameter Name Parameter Value Gradient Constraint Coefficient α 0.5 Smoothing Constraint Coefficient β 10 Number of Iterations 50 Iteration Step γ 0.01 Convergence Threshold 0.0001

[0257] After optimizing the deformation field, the registration performance on the test set was significantly improved. The team compared the performance of different methods on the test set, and the results are shown in Table 4:

[0258] Table 4 Performance Comparison Table of Different Registration Methods

[0259]

[0260] The team further tested the robustness of the method of the present invention under different scenarios, including normal light, weak light, strong light, rainy days, foggy days, etc. The experimental results show that the method of the present invention exhibits better robustness than traditional methods under complex conditions such as weak light and foggy days, and the registration success rate has increased by more than 15%.

[0261] The research team also conducted actual scenario tests and applied the method of the present invention to the monitoring system around an important facility. The system includes 8 groups of infrared and visible light dual-modal cameras, which need to process the image stream in real time and perform registration and fusion. After adopting the method of the present invention, the target detection accuracy of the system under night and bad weather conditions has been improved from the original 76.3% to 92.1%, and the false alarm rate has been reduced by 64.7%.

[0262] Traditional infrared and visible light image registration mainly relies on manually designed feature extractors and feature matching algorithms, such as SIFT, SURF, etc. These methods often fail in the case of large modal differences. Although existing deep learning methods such as DeepFlow and FlowNet have improved performance, they still have not solved the essential problems brought by modal differences. The present invention transforms the multi-modal problem into a single-modal problem by introducing style transfer, significantly reducing the registration difficulty; at the same time, an improved attention mechanism is adopted to achieve explicit feature matching, improving the interpretability of the model; finally, fine adjustment is carried out through the deformation field optimization function, further improving the registration accuracy. Compared with traditional methods, the present invention has significantly improved in indicators such as mutual information, structural similarity, and registration success rate, and has more obvious advantages especially in scenarios with large differences in lighting conditions, providing an efficient and reliable new method for multi-modal image registration.

[0263] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 5 and 6 below.

[0264] Table 5 Variable Explanation Table (First Part)

[0265]

[0266]

[0267] Table 6 Variable Explanation Table (Second Part)

[0268]

[0269]

[0270] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. An unsupervised explicit registration method for infrared and visible images based on style transfer, characterized in that, Including: Select pairs of infrared and visible images from the dataset, perform deformation processing on the infrared images to obtain deformed infrared images, and construct a registration dataset; Train the generative adversarial network CycleGAN to learn the bidirectional mapping between the two modalities, and convert the visible light images into pseudo-infrared images; construct a multi-scale feature extraction module, adopt the ResNet18 network architecture and replace the standard max pooling layer with discrete wavelet transform; construct a feature matching module based on self-attention and cross-attention, and replace the dot product with Linear Attention to reduce the computational complexity; Construct a deformation field generation and image conversion module, design a multi-layer decoder structure to estimate the deformation field; use the deformation field optimization function to finely adjust the deformation field; use the optimized deformation field and the spatial transformation module to resample the deformed infrared images to obtain infrared images registered with the visible light images; use the pseudo-infrared image and the deformed infrared image dataset to train the registration model; input the infrared image and the visible light image to be registered into the trained registration network, perform style transfer, feature extraction, matching and deformation field estimation to obtain the registration result.

2. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 1, characterized in that, CycleGAN includes two generators G_AB and G_BA and two discriminators D_A and D_B, and realizes bidirectional domain conversion through cyclic consistency loss and adversarial loss.

3. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 2, wherein Self-HiLo refers to the self-attention mechanism where the query, key, and value all come from the same image feature, and is used to extract the global dependencies within a single image.

4. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 3, wherein Cross-HiLo refers to the cross-attention mechanism where the query comes from one image feature and the key and value come from another image feature, and is used to achieve feature matching between different images.

5. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 4, wherein The deformation field is a two-dimensional vector field that describes the pixel mapping relationship from the source image to the target image, and the vector at each position represents the displacement of that point during the registration process.

6. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 5, characterized in that, The spatial transformation module is a differentiable image resampling operation that interpolates and reconstructs the source image according to the deformation field to generate a result aligned with the target image.

7. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 6, wherein The deformation field optimization function is used to perform fine adjustment based on the initial deformation field estimate, and optimizes the deformation field by minimizing the image similarity loss and regularization constraints.

8. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 7, wherein The input of the deformation field optimization function includes the initial deformation field generated by the network, the pseudo-infrared image generated by the style transfer network, the deformed infrared image, the gradient constraint coefficient, and the smoothing constraint coefficient, and the output is the optimized deformation field.

9. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 8, wherein Input the infrared image and the visible light image to be registered into the trained registration network, convert the visible light image into a pseudo-infrared image through style transfer, and then perform feature extraction, matching and deformation field estimation to obtain the registration result.

10. The unsupervised explicit registration method for infrared and visible images based on style transfer according to claim 9, characterized in that, The gradient constraint coefficient is a weight parameter that controls the gradient amplitude of the deformation field and is used to prevent excessive local deformation of the deformation field; the smoothing constraint coefficient is a weight parameter that controls the smoothness of the deformation field and is used to ensure the spatial continuity of the deformation field.