Registration and fusion dual-drive misaligned infrared and visible light image fusion method
By employing a dual-drive approach of registration and fusion, and utilizing contrastive learning and collaborative attention fusion modules, the problems of unstable registration accuracy and artifacts in infrared and visible light images under non-rigid deformation were solved, thus achieving high-quality image fusion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-10
AI Technical Summary
Existing infrared and visible light image fusion methods suffer from unstable registration accuracy and the fusion result is affected by local registration errors when dealing with non-rigid deformations, leading to severe artifacts.
A dual-drive approach of registration and fusion is adopted, which improves registration accuracy and reduces artifacts by using a multi-scale feature extraction module based on contrastive learning and a collaborative attention fusion module. This includes a multi-scale feature extraction module CLMFE based on contrastive learning and a collaborative attention fusion module CAFM, which combines window attention and gradient channel attention to achieve fine alignment of features and suppression of misaligned and redundant features.
It significantly improves the registration accuracy and stability of unaligned infrared and visible light image fusion, reduces artifacts in the fusion results, and enhances the image fusion effect.
Smart Images

Figure CN121837041A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to image registration methods and image fusion methods, specifically to a dual-driven method for fusion of misaligned infrared and visible light images, involving both registration and fusion. Background Technology
[0002] Visible light sensors can provide rich texture details by capturing reflected light, but their image quality is easily affected by environmental factors such as low light and inclement weather. In contrast, infrared sensors can detect light-insensitive thermal radiation information and highlight important targets with high contrast, but they are insufficient in describing texture details. The natural complementarity of infrared and visible light images at the information level makes their fusion technology widely applicable in fields such as target detection, target tracking, semantic segmentation, security surveillance, and military reconnaissance.
[0003] Existing deep learning-based image fusion methods mainly include those based on autoencoders (AEs), convolutional neural networks (CNNs), generative adversarial networks (GANs), and Transformers. These methods typically require strict alignment of the input source images. However, due to differences in imaging principles and shooting locations, even infrared and visible light images acquired at the same time in the same scene often exhibit varying degrees of spatial misalignment. Therefore, in practical applications, it is usually necessary to preprocess the source images using image registration methods before fusion. However, when registration and fusion are considered two independent processes, fusion, as a downstream task of registration, can only "tolerate" rather than "counteract" the errors generated during pre-registration, inevitably leading to artifacts in the fusion result.
[0004] In recent years, some researchers have proposed a joint training framework for image registration and fusion to improve the fusion effect of misaligned images. However, such methods still have the following shortcomings: 1) Registration methods based on dense deformation field estimation are highly susceptible to noise interference, leading to unstable registration accuracy. 2) Existing registration methods are prone to local registration errors when dealing with non-rigid deformations, and existing fusion networks usually do not fully consider this problem in the design process, resulting in artifacts in the final fusion result.
[0005] Therefore, there is an urgent need in this field for a new method for fusing unaligned infrared and visible light images. This method can improve the stability of registration accuracy when dealing with non-rigid deformation, and can effectively reduce the impact of local registration errors on the fusion results, thereby significantly alleviating artifacts in the fused image. Summary of the Invention
[0006] This invention addresses the problems of unstable registration accuracy and fusion results affected by local registration errors in existing methods when dealing with non-rigid deformations. It proposes a dual-driven method for fusion of unaligned infrared and visible light images, involving both registration and fusion.
[0007] This invention is achieved using the following technical solution: a dual-driven method for fusion of misaligned infrared and visible light images, comprising the following steps:
[0008] Design and construct an unaligned infrared and visible light image fusion model: The unaligned infrared and visible light image fusion model includes a registration network and a fusion network. The registration network includes a multi-scale feature extraction module CLMFE based on contrastive learning and a multi-scale iterative flow estimation module MSIFE. The fusion network includes a feature extraction module, a collaborative attention fusion module CAFM, and a feature reconstruction module.
[0009] The CLMFE (Contrastive Learning-Based Multi-Scale Feature Extraction) module utilizes a contrastive learning mechanism to enhance the similarity between different modalities of images within the same scene, while simultaneously increasing the differences between images across scenes, thereby extracting more discriminative infrared and visible light multi-scale shared features. These extracted multi-scale shared features are input into the MSIFE (Multi-Scale Iterative Flow Estimation) module, which calculates correlations based on these features to generate a dense deformation field. This field is then used to resample the floating image, resulting in a registered image. Subsequently, the registered image and its corresponding fixed image are input into the feature extraction module to obtain their respective image features. Then, the CAFM (Collaborative Attention Fusion) module performs cross-modal fusion of the infrared and visible light image features. Finally, the fused features are processed by the feature reconstruction module to generate a fused image with clear structure and rich details.
[0010] The aforementioned registration and fusion dual-driven method for fusing unaligned infrared and visible light images includes a contrastive learning-based multi-scale feature extraction module (CLMFE) comprising a contrastive learning encoder (CLEN) and a parameter-shared feature extractor. CLEN consists of two convolutional blocks and a downsampling residual block (DRB), which, combined with a contrastive loss function, maps the two modal images into a modality-independent latent space, thereby extracting common features from the infrared and visible light images. The parameter-shared feature extractor, consisting of two DRBs and a convolutional block, is used to further extract multi-scale common features.
[0011] The aforementioned registration and fusion dual-driven method for fusion of misaligned infrared and visible light images comprises a Collaborative Attention Fusion (CAFM) module consisting of two sub-modules: window attention and gradient channel attention. In the window attention sub-module, intra-domain fusion of infrared and visible light image features is first performed using window self-attention. Then, the infrared image features are mapped to query vectors, and the visible light image features to key and value vectors, or vice versa. Multi-head cross-attention is then used to fuse features at the window level. In the gradient channel attention sub-module, the gradients of the dual-modal image features are processed through a global max-pooling layer and a global average-pooling layer. The results are then mapped through a linear layer, summed, and a sigmoid activation function is applied to generate channel attention weights. These weights are then applied to the fused feature map obtained through window attention to obtain the final fused features.
[0012] A contrastive learning mechanism is introduced into the registration network to enhance the discriminative power of shared features across modalities, thereby improving the stability of registration accuracy. In the design of the fusion network, to address the impact of local registration errors on the fusion result, the feedback characteristics of window attention and fusion consistency loss on registration are combined to achieve fine alignment of cross-modal features; the feedback characteristics of gradient channel attention and fusion consistency loss on registration are combined to effectively suppress misaligned redundant features. The synergistic effect of these two mechanisms significantly reduces artifacts generated in the fusion results of misaligned infrared and visible light images.
[0013] The aforementioned registration and fusion dual-driven method for fusion of misaligned infrared and visible light images uses floating infrared images as input during training of the misaligned infrared and visible light image fusion model. With fixed visible light images And floating visible light images With fixed infrared images After training, the input during actual multimodal image fusion is a floating infrared image. With fixed visible light images Or floating visible light image With fixed infrared images .
[0014] The above-described registration and fusion dual-driven method for fusing misaligned infrared and visible light images will work if the input to the misaligned infrared and visible light image fusion model is a floating infrared image. With fixed visible light images The input to the fusion network is a fixed visible light image. Image registered with infrared The fusion network first extracts features from infrared and visible light images using a feature extraction module consisting of three convolutional layers, obtaining infrared features. and visible light characteristics ; then, and Input the Collaborative Attention Fusion Module (CAFM) to generate fused features. .
[0015] The above-mentioned dual-driven registration and fusion method for fusion of misaligned infrared and visible light images, along with the CAFM attention fusion module, performs cross-modal fusion as follows:
[0016] (1) Use window attention to respectively and Intra-domain fusion is performed to effectively integrate image features within the same domain. Specifically, for and The region is divided into non-overlapping local windows. Standard multi-head self-attention (MSA) is performed on each window, and residual structures and layer normalization (LN) are applied to the MSA and feedforward network (FFN). Conventional window partitioning and moving window partitioning are used alternately to achieve cross-window connectivity and form infrared local window features. and visible light local window features ;
[0017] (2) By combining the feedback effect of window attention and fusion consistency loss on registration, the inter-domain fusion results between infrared and visible light image features are adaptively calculated, achieving fine alignment of features and effectively integrating complementary information from different modalities; specifically, given two local window features from different image domains and To fuse infrared features to enhance the representation of visible light features, Mapped to query vector , Mapped to key vector Sum value vector For each window, a multi-head cross-attention (MCA) mechanism is executed, namely: , , ,in, , , It is a learnable weight matrix. represent The output is after cross-domain fusion; similarly, this is to fuse visible light features to enhance the representation of infrared image features. Mapped to query vector , Mapped to key vector Sum value vector ,Right now: , , ,in, , and It is a learnable weight matrix. represent The output after cross-domain fusion;
[0018] The window attention mechanism has L layers, and the final output feature is defined as follows: , ;
[0019] (3) By combining the feedback effect of gradient channel attention and fusion consistency loss on registration, and through adaptive optimization of channel feature selection strategy, misaligned redundant features are effectively suppressed; specifically, and The gradients are processed by a global max pooling layer (GMP) and a global average pooling layer (GAP), and the results are mapped through a linear layer. Then, the results are summed and a sigmoid activation function is applied to generate channel attention weights. ,Will Multiplying this by the corresponding input features yields the final fused features. , , ,
[0020] in, Represents Sobel gradient computation. , This represents the learnable weight matrix of the linear layer. Represents the Sigmoid activation function. This represents element-wise multiplication.
[0021] The above-mentioned dual-driven registration and fusion method for fusion of misaligned infrared and visible light images, registering the image... The calculation process is as follows: The contrastive learning-based multi-scale feature extraction module CLMFE first uses the contrastive learning encoder CLEN to extract the floating infrared image. With fixed visible light images Mapping to a mode-independent latent space , Then, CLMFE utilizes a parameter-sharing feature extractor to further extract multi-scale common features, denoted as... and superscript Representing different scales, the common features across multiple scales are input into the multi-scale iterative flow estimation module MSIFE to estimate the dense deformation field. Based on this, the floating image is resampled to obtain the registered image. .
[0022] The above-mentioned registration and fusion dual-driven method for fusing unaligned infrared and visible light images has a loss function consisting of multiple parts during the training of the fusion model;
[0023] First, luminosity loss and endpoint loss As a registration loss function, it is used to ensure the accuracy of image registration: , ,
[0024] In the formula, The deformation field from a floating visible light image to a fixed infrared image. The deformation field from a floating infrared image to a fixed visible light image. For reference deformation field;
[0025] Secondly, strength loss gradient loss and structural similarity loss As a fusion loss function, it is used to improve the detail preservation and structural consistency of the fused image: , , ;
[0026] In the formula, For floating infrared images With fixed visible light images The fused image obtained after inputting the fusion model. For floating visible light images With fixed infrared images The fused image obtained after inputting the fusion model. For infrared registration images, For visible light registration images, Represents Sobel gradient computation;
[0027] In addition, contrast loss is introduced. To enhance the discriminative power of shared features between images of different modalities, given the significant modal differences between infrared and visible light images, it is necessary to optimize the contrast loss separately in the two image domains. Considering that using complete samples directly for comparative learning would impose significant storage and optimization burdens, this invention employs a sampling set. Training is performed. Image pairs of the same scene but different modalities are considered positive samples, while image pairs of different scenes with the same or different modalities are considered negative samples. By strengthening the similarity of positive samples and enhancing the difference of negative samples, the discriminative power of the features extracted during the registration stage can be effectively improved. Therefore, the contrast loss between infrared and visible light is considered. , The definition is as follows:
[0028] ,
[0029] ,in, It is a temperature coefficient used to adjust the dynamic range. Represents the similarity between image features; and The common features of the infrared images extracted from the m-th and n-th samples in the sampling set; and The common features of the visible light images extracted from the m-th and n-th samples in the sampling set;
[0030] Furthermore, a fusion consistency loss is introduced. , ,
[0031] In summary, the total loss function is: ,in , , , , , This represents the loss coefficient.
[0032] The aforementioned dual-driven registration and fusion method for fusion of misaligned infrared and visible light images obtains its training and test sets through the following process: Based on the MSRS dataset, images are subjected to synthetic random affine and elastic transformations to generate misaligned image pairs. The affine transformation includes rotations of [-10, 10] degrees and translations within the range of [-10, 10]. The elastic transformation uses six Gaussian filters to blur a dual-channel [-1, 1] noise map, with a sigma of 15 and a kernel size of 45×45. The training set contains 1083 image pairs, and the test set contains 361 image pairs. To enhance the diversity of the training set samples, data augmentation operations are performed on the training set images, including random cropping, rotation, and horizontal / vertical flipping.
[0033] To address the problems of unstable registration accuracy and fusion results affected by local registration errors in existing methods when handling non-rigid deformations, this invention proposes a dual-driven registration and fusion method for fusing misaligned infrared and visible light images. In this method, we utilize contrastive learning to enhance the similarity between different modalities of the same scene and increase the differences between images from different scenes, enabling the model to learn more discriminative features and improving the stability of registration accuracy. Simultaneously, a Collaborative Attention Fusion Module (CAFM) is designed, combining window attention, gradient channel attention, and the feedback characteristics of image fusion on registration (i.e., fusion consistency loss) to achieve fine feature alignment and suppression of misaligned redundant features, thus mitigating artifacts in the fusion results. Attached Figure Description
[0034] Figure 1 This is a diagram of the overall network structure.
[0035] Figure 2 This is a structural diagram of the contrastive learning-based multi-scale feature extraction module (CLMFE).
[0036] Figure 3 This is a structural diagram of the Collaborative Attention Fusion Module (CAFM).
[0037] Figure 4 The image shows the feature extraction results for infrared and visible light with and without contrast learning.
[0038] Figure 5 The fusion results are shown for different levels of window attention (WA).
[0039] Figure 6 The image shows the fusion results with and without gradient channel attention (GCA).
[0040] Figure 7 Images showing the fusion results of unaligned infrared and visible light images in different scenarios. Detailed Implementation
[0041] This invention provides a dual-driven registration and fusion method for fusing misaligned infrared and visible light images, referring to... Figure 1 This method is implemented through the following steps:
[0042] Step 1: Floating infrared image With fixed visible light images and floating visible light images With fixed infrared images The inputs are fed into two symmetrical network branches respectively.
[0043] Step 2: Utilize the contrastive learning-based multi-scale feature extraction module (CLMFE) to obtain modality-independent multi-scale common features. (CLMFE reference...) Figure 2 In its network structure, CLMFE first utilizes a contrastive learning encoder (CLEN) and its contrastive loss function to map infrared and visible light images into a mode-independent latent space, thereby obtaining common features. , (by and (For example, CLEN contains two convolutional blocks and one downsampling residual block (DRB): the convolutional blocks sequentially contain...) The system consists of convolutional layers, instance normalization layers, and the Leaky ReLU activation function; the DRB is a residual structure composed of three convolutional blocks and a max-pooling downsampling layer. Then, the CLMFE further extracts multi-scale common features using a parameter-shared feature extractor, denoted as... and superscript Representing different scales, the values range from 1, 2, and 3. This feature extractor consists of two DRBs and a convolutional block, and outputs the common features of that scale before the downsampling operation of the DRBs.
[0044] Step 3: Input the multi-scale common features into the multi-scale iterative flow estimation module (MSIFE) to estimate the dense deformation field. , Based on this, the floating image is resampled to obtain the registered image. and MSIFE utilizes shared features across multiple scales to refine the deformation field between the floating and fixed images scale by scale. Specifically, assuming the infrared image is the floating image and the visible light image is the fixed image, at the 1st... At each scale, the infrared features are first deformed based on the deformation field estimated at the previous scale to obtain coarsely registered infrared features. Then, the correlation between the coarsely registered infrared features and visible light features is calculated to obtain a correlation representation of the cross-modal features. Next, the correlation representation is concatenated with the two modal image features, and after convolution calculation, the deformation field at this scale is obtained. Finally, the deformation field at this scale is superimposed with the deformation field at the previous scale, thereby progressively refining the registration accuracy.
[0045] Step 4: With fixed visible light images ,as well as With fixed infrared images The input is fed into the fusion network to obtain the final fused image. and .by and Taking this as an example, the fusion network first extracts features from infrared and visible light images through a feature extraction module consisting of three convolutional layers, obtaining infrared features. and visible light characteristics Subsequently, and Input the Collaborative Attention Fusion Module (CAFM) to generate fused features. Specifically, CAFM aims to adaptively integrate and filter complementary information based on cross-modal feature alignment, thereby enhancing the detail representation of the fusion result. Its design includes the following key components:
[0046] (1) Use window attention to respectively and Intra-domain fusion is performed to effectively integrate image features within the same domain. Specifically, for a given size of... Features Divide it into non-overlapping groups. A partial window, resulting in a size of Features ,in This represents the total number of windows. Standard multi-head self-attention (MSA) is performed on each window, and residual structures and layer normalization (LN) are applied to both the MSA and the feedforward network (FFN). Regular window partitioning and moving window partitioning are used alternately to achieve cross-window connections. Therefore, window features... The intra-domain fusion process can be represented as:
[0047]
[0048]
[0049]
[0050] in, , , There are three learnable weight matrices. , , These represent the query vector, key vector, and value vector obtained through linear projection, respectively. Represents a local window feature The results of self-attention fusion.
[0051] (2) By combining the feedback effect of window attention and fusion consistency loss on registration, fine alignment of features is achieved through adaptive calculation of inter-domain fusion results between infrared and visible light image features, while effectively integrating complementary information from different modalities. Specifically, given two local window features from different image domains... and To fuse infrared features to enhance the representation of visible light features, Mapped to query vector , Mapped to key vector Sum value vector For each window, perform multi-head cross-attention (MCA) mechanism separately, that is:
[0052]
[0053]
[0054]
[0055] in, , , It is a learnable weight matrix. represent The output result after cross-domain fusion.
[0056] Similarly, to fuse visible light features to enhance the feature representation of infrared images, Mapped to query vector , Mapped to key vector Sum value vector ,Right now:
[0057]
[0058]
[0059]
[0060] in, , and It is a learnable weight matrix. represent The output result after cross-domain fusion.
[0061] It is worth noting that window attention has L layers, and its output features are defined as follows: , .
[0062] (3) By combining the feedback effect of gradient channel attention and fusion consistency loss on registration, and through an adaptive optimization strategy for channel feature selection, misaligned redundant features are effectively suppressed. Specifically, and The gradients are processed by a global max pooling (GMP) layer and a global average pooling (GAP) layer, and the corresponding results are mapped through a linear layer. Then, they are summed and a sigmoid activation function is applied to generate channel attention weights. .Will Multiplying this by the corresponding input features yields the final fused features. :
[0063]
[0064]
[0065] in, Represents Sobel gradient computation. , This represents the learnable weight matrix of the linear layer. Represents the Sigmoid activation function. This represents element-wise multiplication.
[0066] at last, The fused image is obtained through a feature reconstruction module consisting of three convolutional layers. .
[0067] Step 5: In the model training process of this invention, the loss function consists of multiple parts. First, the photometric loss... and endpoint loss As a registration loss function, it is used to ensure the accuracy of image registration:
[0068]
[0069]
[0070] Secondly, strength loss gradient loss and structural similarity loss As a fusion loss function, it is used to improve the detail preservation and structural consistency of the fused image:
[0071]
[0072]
[0073]
[0074] Furthermore, this invention introduces contrast loss. To enhance the discriminative power of shared features between images of different modalities, and given the significant modal differences between infrared and visible light images, it is necessary to optimize the contrast loss separately in the two image domains. Considering that using complete samples directly for comparative learning would impose significant storage and optimization burdens, this invention employs a sampling set. Training is performed. Image pairs of the same scene but different modalities are considered positive samples, while image pairs of different scenes with the same or different modalities are considered negative samples. By strengthening the similarity of positive samples and enhancing the difference of negative samples, the discriminative power of the features extracted during the registration stage can be effectively improved. Therefore, the contrast loss between infrared and visible light is considered. , The definition is as follows:
[0075]
[0076]
[0077] in, It is a temperature coefficient used to adjust the dynamic range. This represents the similarity between image features.
[0078] Furthermore, this invention introduces a fusion consistency loss. This is to enable the fusion to provide feedback on the registration task during the training process.
[0079]
[0080] In summary, the total loss function is:
[0081]
[0082] in , , , , , .
[0083] The overall model optimization process involves updating network parameters via backpropagation. To comprehensively evaluate the fusion effect, this invention employs multiple objective indicators for quantitative analysis, including: entropy (EN), spatial frequency (SF), mutual information (MI), sum of differential correlations (SCD), and visual information fidelity (VIF). And the Structural Similarity Index (SSIM).
[0084] Step 6: Model Training Parameter Settings: The CPU used in this invention is an Intel Xeon 24-core processor, the GPU is an NVIDIA RTX 3090 24G graphics card, the operating system is Windows 10, the testing software is PyCharm 2021.2.2, and the deep learning framework is PyTorch 1.8.1. The total number of training iterations is 380, the network uses the Adam optimizer, and the initial learning rate of the optimizer is set to 0.001. Before inputting into the network, all images are normalized so that their values are within the range of [0, 1].
Claims
1. A method for fusing misaligned infrared and visible light images driven by both registration and fusion, characterized in that: Includes the following steps: Design and construct an unaligned infrared and visible light image fusion model: The unaligned infrared and visible light image fusion model includes a registration network and a fusion network. The registration network includes a multi-scale feature extraction module CLMFE based on contrastive learning and a multi-scale iterative flow estimation module MSIFE. The fusion network includes a feature extraction module, a collaborative attention fusion module CAFM, and a feature reconstruction module. Among them, the contrastive learning-based multi-scale feature extraction module CLMFE uses a contrastive learning mechanism to enhance the similarity between images of different modalities in the same scene, while also enhancing the differences between images across scenes, thereby extracting more discriminative infrared and visible light multi-scale common features. The extracted multi-scale common features are input into the multi-scale iterative flow estimation module MSIFE, which calculates the correlation based on the multi-scale common features, generates a dense deformation field, and resamples the floating image accordingly to obtain the registered image. Subsequently, the registered image and the corresponding fixed image are input into the feature extraction module to obtain their respective image features. Then, the infrared and visible light image features are fused across modes through the Collaborative Attention Fusion (CAFM) module. Finally, the fused features are processed by the feature reconstruction module to generate a fused image with clear structure and rich details.
2. The registration and fusion dual-driven method for fusing misaligned infrared and visible light images according to claim 1, characterized in that: The contrastive learning-based multi-scale feature extraction module CLMFE includes a contrastive learning encoder CLEN and a parameter-shared feature extractor. CLEN consists of two convolutional blocks and a downsampling residual block DRB, which, combined with a contrastive loss function, maps the two modal images into a modality-independent latent space, thereby extracting common features between infrared and visible light images. The parameter-shared feature extractor consists of two DRBs and a convolutional block, used to further extract multi-scale common features.
3. The registration and fusion dual-driven method for fusing misaligned infrared and visible light images according to claim 1, characterized in that: The Collaborative Attention Fusion (CAFM) module consists of two sub-modules: window attention and gradient channel attention. In the window attention sub-module, the infrared and visible light image features are first fused within the domain using window self-attention. Then, the infrared image features are mapped to query vectors, and the visible light image features are mapped to key vectors and value vectors. Finally, the infrared image features are mapped to key vectors and value vectors, and the visible light image features are mapped to query vectors. Multi-head cross attention is used to fuse the features at the window level. In the gradient channel attention submodule, the gradients of the bimodal image features are processed by a global max pooling layer and a global average pooling layer. The corresponding results are then mapped through a linear layer, summed, and a sigmoid activation function is applied to generate channel attention weights. These weights are then applied to the fused feature map obtained by window attention to obtain the final fused features.
4. The registration and fusion dual-driven method for fusing misaligned infrared and visible light images according to claim 1, characterized in that: The unaligned infrared and visible light image fusion model is trained with floating infrared images as input. With fixed visible light images And floating visible light images With fixed infrared images After training, the input during actual multimodal image fusion is a floating infrared image. With fixed visible light images Or floating visible light image With fixed infrared images .
5. The registration and fusion dual-driven method for fusing misaligned infrared and visible light images according to claim 4, characterized in that: If the input is an unaligned infrared and visible light image fusion model, the infrared image will be floating. With fixed visible light images The input to the fusion network is a fixed visible light image. Image registered with infrared The fusion network first extracts features from infrared and visible light images using a feature extraction module consisting of three convolutional layers, obtaining infrared features. and visible light characteristics ; then, and Input the Collaborative Attention Fusion Module (CAFM) to generate fused features. .
6. The registration and fusion dual-driven method for fusing misaligned infrared and visible light images according to claim 5, characterized in that: The specific process of cross-modal fusion performed by the Collaborative Attention Fusion Module (CAFM) is as follows: (1) Use window attention to respectively and Intra-domain fusion is performed to effectively integrate image features within the same domain. Specifically, for and The network is divided into non-overlapping local windows. Standard multi-head self-attention (MSA) is performed on each window, and residual structures and layer normalization (LN) are applied to the MSA and feedforward network (FFN). Conventional window partitioning and moving window partitioning are used alternately to achieve cross-window connectivity and form local window features. and ; (2) By combining the feedback effect of window attention and fusion consistency loss on registration, the inter-domain fusion results between infrared and visible light image features are adaptively calculated, achieving fine alignment of features and effectively integrating complementary information from different modalities; specifically, given two local window features from different image domains and To fuse infrared features to enhance the representation of visible light features, Mapped to query vector , Mapped to key vector Sum value vector For each window, a multi-head cross-attention (MCA) mechanism is executed, namely: , , ,in, , , It is a learnable weight matrix. represent The output after cross-domain fusion; Similarly, to fuse visible light features to enhance the feature representation of infrared images, Mapped to query vector , Mapped to key vector Sum value vector ,Right now: , , ,in, , and It is a learnable weight matrix. represent The output after cross-domain fusion; The window attention mechanism has L layers, and the final output feature is defined as follows: , ; (3) By combining the feedback effect of gradient channel attention and fusion consistency loss on registration, and through adaptive optimization of channel feature selection strategy, misaligned redundant features are effectively suppressed; specifically, and The gradients are processed by a global max pooling layer (GMP) and a global average pooling layer (GAP), and the results are mapped through a linear layer. Then, the results are summed and a sigmoid activation function is applied to generate channel attention weights. ,Will Multiplying this by the corresponding input features yields the final fused features. , , , in, Represents Sobel gradient computation. , This represents the learnable weight matrix of the linear layer. Represents the Sigmoid activation function. This represents element-wise multiplication.
7. The registration and fusion dual-driven method for fusing misaligned infrared and visible light images according to claim 5 or 6, characterized in that: Registered images The calculation process is as follows: The contrastive learning-based multi-scale feature extraction module CLMFE first uses the contrastive learning encoder CLEN and its contrastive loss function to extract the floating infrared image. With fixed visible light images Mapping to a mode-independent latent space , Then, CLMFE utilizes a parameter-sharing feature extractor to further extract multi-scale common features, denoted as... and superscript Representing different scales, the common features across multiple scales are input into the multi-scale iterative flow estimation module MSIFE to estimate the dense deformation field. Based on this, the floating image is resampled to obtain the registered image. .
8. The registration and fusion dual-driven method for fusion of misaligned infrared and visible light images according to claim 5 or 6, characterized in that: The loss function during the training of the fusion model consists of multiple parts; First, luminosity loss and endpoint loss As a registration loss function, it is used to ensure the accuracy of image registration: , , In the formula, The deformation field from a floating visible light image to a fixed infrared image. The deformation field from a floating infrared image to a fixed visible light image. For reference deformation field; Secondly, strength loss gradient loss and structural similarity loss As a fusion loss function, it is used to improve the detail preservation and structural consistency of the fused image: , , ; In the formula, For floating infrared images With fixed visible light images The fused image obtained after inputting the fusion model. For floating visible light images With fixed infrared images The fused image obtained after inputting the fusion model. For infrared registration images, For visible light registration images, Represents Sobel gradient computation; In addition, contrast loss is introduced. To enhance the discriminative power of shared features between images of different modalities, given the significant modal differences between infrared and visible light images, it is necessary to optimize the contrast loss separately in the two image domains. ; Therefore, the contrast loss between infrared and visible light , The definition is as follows: , ,in, It is a temperature coefficient used to adjust the dynamic range. Represents the similarity between image features; and The common features of the infrared images extracted from the m-th and n-th samples in the sampling set; and The common features of the visible light images extracted from the m-th and n-th samples in the sampling set; Furthermore, a fusion consistency loss is introduced. , , In summary, the total loss function is: ,in , , , , , This represents the loss coefficient.
9. The registration and fusion dual-driven method for fusion of misaligned infrared and visible light images according to claim 5 or 6, characterized in that: The training and test sets of the fusion model are obtained through the following process: Based on the MSRS dataset, images are subjected to synthetic random affine and elastic transformations to generate unaligned image pairs. The affine transformation includes a rotation of [-10, 10] degrees and a translation range of [-10, 10]. The elastic transformation uses 6 Gaussian filters to blur a two-channel [-1, 1] noise map, where sigma is 15 and the kernel size is 45×45. The training set contains 1083 image pairs, and the test set contains 361 image pairs. To enhance the diversity of the training set samples, data augmentation operations are performed on the training set images, including random cropping, rotation, and horizontal / vertical flipping.